diff --git a/docs/agent-lifecycle/agent-hierarchy.md b/docs/agent-lifecycle/agent-hierarchy.md index 7da5b2bb..1e95f436 100644 --- a/docs/agent-lifecycle/agent-hierarchy.md +++ b/docs/agent-lifecycle/agent-hierarchy.md @@ -100,7 +100,8 @@ other agents don't: - **Naming/bootstrap** — the manager's broker recipient name, state-dir key, and nixos-container name are all `ruth` (container `h-ruth`). `hive-c0re` spawns it directly at boot if missing, with no operator - approval step — every other agent goes through a `Spawn` approval. + approval step — every other agent is created at swarm level + (`swarmctl agent create`). Roster-wise, `ruth` is just another entry. - **Wire-protocol** — the privileged `Request` variants (`Kill` / `Start` / `Restart` / `Update`; `GetLogs`) — marked diff --git a/docs/agent-lifecycle/approvals.md b/docs/agent-lifecycle/approvals.md index 563f3358..c211b587 100644 --- a/docs/agent-lifecycle/approvals.md +++ b/docs/agent-lifecycle/approvals.md @@ -24,14 +24,6 @@ CLI) before it takes effect. What you'll see, and what to do with it: anything in that chain fails, the change rolls back automatically — the agent stays on its last-good config, no recovery action needed from you. -- **New agent** (`Spawn`) — the one approval in creating a brand-new - agent. The swarm controller's `InitAgentConfigRepo` job - (`POST /api/agents`) scaffolds its config repo first, outside the - approval queue; `Spawn` then creates the container from that - config. Tailoring what the template seeded isn't a separate - mechanism — it's the config-change flow above, a PR you review like - any other. Every later change goes through that flow — there's no - repeat "spawn" for an existing agent. - **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump one or more Nix flake inputs (or all of them). Approving runs the update and commits the lock change; it doesn't rebuild anything by @@ -133,11 +125,11 @@ agent that lacks the `approvals` tool group: only an agent with that group submits approvals (for its direct children), so an agent without it has nothing of its own to withdraw. -The swarm controller's `InitAgentConfigRepo` job creates a brand-new -agent's config repo outside this queue, seeding it with a default -`agent.nix` template. The operator then **spawns** the agent -(the `Spawn` approval / `◆ R3QU3ST SP4WN` button), which creates the -container from that config. +Creating a brand-new agent isn't an approval: it's swarm-level +(`swarmctl agent create`, `POST /api/agents` on the swarm controller), +and no hive can originate an agent. The swarm controller's +`InitAgentConfigRepo` job seeds the agent's config repo with a default +`agent.nix` template, then asks the target hive to deploy it. Changing what the template seeded isn't a special case: like every later change, it's a PR on that config repo (`MergeConfigPr`), made @@ -147,7 +139,7 @@ through the web UI or the forge. ### Approval kinds (wire shapes) -`ApprovalKind` carries four variants; each maps to a different +`ApprovalKind` carries three variants; each maps to a different `commit_ref` encoding because `ApprovalKind` overloads that field as the kind-specific payload carrier. @@ -165,17 +157,6 @@ the kind-specific payload carrier. eval-verifies it; `DeployApply` fast-forward-merges the forge config repo's `main` to it (the merge) and runs `deploy_applied_target`; `DeployTail` compensates on failure. Never a first spawn. -- `Spawn` — direct container creation from the agent's config repo. - `commit_ref` is empty. Submitted via `HostRequest::RequestSpawn` - (operator-gated, the `◆ R3QU3ST SP4WN` dashboard button + - `hivectl agent request-create` CLI). The host-level `HostRequest::Spawn` - variant bypasses the approval queue entirely — privileged-context use - only (operator on the host shell, test scripts, one-off recoveries; - `hivectl agent create`). This is the **canonical first-spawn**: the - swarm controller's `InitAgentConfigRepo` job seeds the agent's config - repo, it gets customised through a PR, then the operator spawns to - create the container. Subsequent config changes go through a - `MergeConfigPr` PR. - `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs, etc.). hive-c0re sets the `agent` field to the requesting root agent. @@ -412,8 +393,8 @@ submitter pushes again (or closes it) to retry. ### Dispatch via the job queue -Long-running approval work — `MergeConfigPr`, `UpdateMetaInputs`, -`Spawn` — runs as a DAG on the global job queue +Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs` +— runs as a DAG on the global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler rather than run inline: @@ -421,7 +402,6 @@ rather than run inline: |---|---|---| | `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` | | `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` | -| `Spawn` | `spawn` (`Create → WriteDropin → Reconcile`) | `approval` | | `SchedulePrompt` | — runs inline (single sqlite insert) | — | The DAG carries the originating `approval_id`, surfaced on the node that @@ -433,7 +413,7 @@ state is the authoritative outcome. That hook fires the matching `HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's terminal tag (verifying the tag actually resolves in the applied repo — a pre-merge rejection plants none), posts the failing build log back to -the config PR, and for a spawn runs the post-spawn forge bookkeeping. +the config PR. Two visible consequences: @@ -612,12 +592,12 @@ event to the agent that originally submitted the approval (looked up from the `submitter` column on the `approvals` table). The harness delivers it as a regular `system` inbox message so it drives a normal claude turn. `finish_approval` fires an `ApprovalResolved` HelperEvent this way for -**every** approval kind's terminal state, `Spawn` included. A +**every** approval kind's terminal state. A "FYI, check when convenient" event doesn't need a message — those go through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct live dial of the target agent's in-container todo socket (same `UpsertTodo` request in-container producers use); `finish_approval` fires -one of these too for `Spawn`/`MergeConfigPr`, *in addition to* +one of these too for `MergeConfigPr`, *in addition to* the `ApprovalResolved` HelperEvent above, not instead of it. Legacy approval rows that predate the submitter column fall back to the root agent. Variants (`hive_sh4re::manager::HelperEvent`): @@ -651,7 +631,7 @@ hive-c0re-vouched commit sha. Optional `tag` carries the deploy bookkeeping tag — `deployed/` on a successful build or `failed/` on a failed one, planted by the `MergeConfigPr` deploy. Both fields are `Option`: `None` on the paths that don't deploy a new -commit (spawn / meta-update / deny, and the autoupdate +commit (meta-update / deny, and the autoupdate sweep's `job_queue::templates::rebuild` reapplying the existing main, or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set, `git show ` against `/applied//.git` inside the diff --git a/docs/agent-lifecycle/persistence.md b/docs/agent-lifecycle/persistence.md index af74b3ec..08eb802f 100644 --- a/docs/agent-lifecycle/persistence.md +++ b/docs/agent-lifecycle/persistence.md @@ -10,8 +10,8 @@ keeps its state, purging it doesn't.** - **`DESTR0Y`** (the default action) stops and removes the container but keeps everything on disk — config history, claude login, `/state/` notes, harness data. The agent shows up as a tombstone (K3PT ST4T3 on - the C0R3 page) with a `⊕ R3V1V3` button that recreates it from the - kept state, **no re-login needed**. + the C0R3 page); `swarmctl agent create` with the same name and hive + recreates it from the kept state, **no re-login needed**. - **`PURG3`** (opt-in, from the dashboard or `hivectl agent destroy --purge`) is `DESTR0Y` plus wiping all of it — config history, claude credentials, `/state/` notes, everything. **No @@ -432,8 +432,8 @@ See [For operators](#for-operators) above for what each action does to an agent's state. The mechanics, for completeness: - `DESTR0Y` also drops the systemd drop-in and fails any pending - approvals; the tombstone's `⊕ R3V1V3` button queues a Spawn approval - that reuses the kept state on approve. + approvals; reviving the agent is `swarmctl agent create` with the same + name and hive, which reuses the kept state. - `PURG3` wipes `/var/lib/hyperhive/{agents,applied}//` — the union of everything `DESTR0Y` left behind. diff --git a/docs/getting-started/setup.md b/docs/getting-started/setup.md index a3f9367e..197bb26a 100644 --- a/docs/getting-started/setup.md +++ b/docs/getting-started/setup.md @@ -303,17 +303,14 @@ a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`). ### 7 · Spawn sub-agents -Sub-agent creation is an operator action — agents have no tool for it. -Two steps: +Sub-agent creation is a swarm-level operator action — agents have no +tool for it, and no hive can create one on its own: ``` -# Step 1: scaffold the new agent's config repo. The swarm controller's -# InitAgentConfigRepo job does this (POST /api/agents), seeding -# /agents/iris/config/agent.nix from the default template. - -# Step 2: edit /agents/iris/config/agent.nix and commit it. Then spawn -# iris from the dashboard (◆ R3QU3ST SP4WN / Spawn approval), which -# builds + starts the container from that config. +# Create iris on hive pr1ma. The swarm controller seeds its config repo +# (agent-configs/iris) from the default template, then asks pr1ma to +# build + start the container from that config. +swarmctl agent create iris --hive pr1ma # Later config changes: open a PR on agent-configs/iris (hive-forge); # the operator reviews + approves it — no MCP tool call. diff --git a/docs/scheduler/coordinator.md b/docs/scheduler/coordinator.md index 853d41d7..25599287 100644 --- a/docs/scheduler/coordinator.md +++ b/docs/scheduler/coordinator.md @@ -322,7 +322,7 @@ the build slot while its parent held the meta window could block waiting for a resource its own parent already committed to, a lock-ordering hazard that one multi-resource root avoids by construction. -`Spawn` and `UpdateMetaInputs` approvals map onto the ordinary `spawn` / +`UpdateMetaInputs` approvals map onto the ordinary `meta-update` shapes. The scheduler fires `actions::resolve_approval_dag` exactly once when **any** approval-carrying DAG settles terminal — deploys included, since their outcome is now the DAG's own state (including diff --git a/docs/tools/hivectl-cli.md b/docs/tools/hivectl-cli.md index eb8702c0..50bbd89e 100644 --- a/docs/tools/hivectl-cli.md +++ b/docs/tools/hivectl-cli.md @@ -21,8 +21,6 @@ This document contains the help content for the `hivectl` command-line program. * [`hivectl agent pause`↴](#hivectl-agent-pause) * [`hivectl agent resume`↴](#hivectl-agent-resume) * [`hivectl agent start`↴](#hivectl-agent-start) -* [`hivectl agent create`↴](#hivectl-agent-create) -* [`hivectl agent request-create`↴](#hivectl-agent-request-create) * [`hivectl agent stop`↴](#hivectl-agent-stop) * [`hivectl agent kill`↴](#hivectl-agent-kill) * [`hivectl agent destroy`↴](#hivectl-agent-destroy) @@ -271,9 +269,7 @@ Everything here targets a single named agent (`hivectl agent foo restart`, `hive * `restart` — Stop and start this agent container without rebuilding config * `pause` — Park this agent's turn loop, leaving the container running * `resume` — Resume this paused agent — it drains whatever queued up while parked -* `start` — Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation. Use `create` for that -* `create` — Create this agent container from scratch (full first-time provisioning), bypassing the approval queue -* `request-create` — Queue a first-creation request for operator approval +* `start` — Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation, which is swarm-level (`swarmctl agent create`) * `stop` — Gracefully stop this agent container: signal → drain → reconcile. Never escalates to a hard kill — use `kill` for that * `kill` — Hard-stop this managed container * `destroy` — Tear down this sub-agent container, keeping its state by default. No undo @@ -326,7 +322,7 @@ Resume this paused agent — it drains whatever queued up while parked ## `hivectl agent start` -Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation. Use `create` for that +Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation, which is swarm-level (`swarmctl agent create`) **Usage:** `hivectl agent start [OPTIONS]` @@ -336,24 +332,6 @@ Start this EXISTING agent container. Fails immediately if `name` has no config/t -## `hivectl agent create` - -Create this agent container from scratch (full first-time provisioning), bypassing the approval queue. - -Operator-on-the-host only; use `request-create` for an approval-gated creation. - -**Usage:** `hivectl agent create` - - - -## `hivectl agent request-create` - -Queue a first-creation request for operator approval - -**Usage:** `hivectl agent request-create` - - - ## `hivectl agent stop` Gracefully stop this agent container: signal → drain → reconcile. Never escalates to a hard kill — use `kill` for that diff --git a/docs/tools/swarmctl-cli.md b/docs/tools/swarmctl-cli.md index 810ab195..caf28d89 100644 --- a/docs/tools/swarmctl-cli.md +++ b/docs/tools/swarmctl-cli.md @@ -70,7 +70,7 @@ No approval gate guards this: running this binary already means being root on th * `` — Name for the new agent: 1–63 characters of `[a-z0-9-]`. - Becomes an SSO subject, a forge user and a repository name, so it's validated here before queuing. + Becomes an SSO subject, a forge user and a repository name, so it's validated here before queuing. The controller refuses a name the swarm has already placed on a different hive; the same name on the same hive re-creates that agent. ###### **Options:** diff --git a/docs/web-ui/dashboard.md b/docs/web-ui/dashboard.md index a3f66cf0..6a6f2c7e 100644 --- a/docs/web-ui/dashboard.md +++ b/docs/web-ui/dashboard.md @@ -971,7 +971,6 @@ renderApprovals`) with three stacked sections: | `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) | | `update_meta_inputs` | `↻` | `meta-update` | — | | `schedule_prompt` | `⏱` | `schedule` | — | - | `spawn` | `⊕` | `spawn` | — | The chip ticks live every second via a `data-requested-at` @@ -985,7 +984,6 @@ renderApprovals`) with three stacked sections: config PR into `agent-configs//pulls/` (shown only when `forge_present` is true and `pr_number` has a value). The config diff lives on the forge PR itself — no inline diff side-panel. - - `spawn`: a one-line "container will be created" note instead. - **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a `prompt()` for an optional reason carried to the submitting agent as `HelperEvent::ApprovalResolved.note`. @@ -1043,7 +1041,6 @@ below — some endpoints aren't in it yet. - `POST /api/{rebuild,kill,restart,start,destroy}/{name}` — lifecycle. `destroy` accepts `purge=on` to also wipe state dirs. - `POST /api/purge-tombstone/{name}` — wipe a tombstone's state dirs. -- `POST /api/request-spawn` — queue a Spawn approval. - `POST /api/update-all` — rebuild every stale container. - `POST /api/rebuild-queue/{id}/cancel` — drop a `Queued` entry. Refuses `Running` / terminal-state entries (in-flight @@ -1243,10 +1240,9 @@ payload): - `container_state_changed` (container: ContainerView) / `container_removed` (name) — per-row container mutations, emitted by `Coordinator::rescan_containers_and_emit` from - many mutation sites — post-spawn approval bookkeeping - (`actions::approve`), the job queue's own node execution - (`job_queue::exec`, for example after a rebuild's stop/swap/start - steps or a destroy's teardown step) — and from the 10s + the job queue's own node execution (`job_queue::exec`, for example + after a rebuild's stop/swap/start steps or a destroy's teardown + step) — and from the 10s `crash_watch` poll. Client upserts/removes by name and reads the pending overlay from `transientsState` since the payload doesn't carry it. diff --git a/frontend/packages/dashboard/src/call.js b/frontend/packages/dashboard/src/call.js index 943962e9..f1c2c417 100644 --- a/frontend/packages/dashboard/src/call.js +++ b/frontend/packages/dashboard/src/call.js @@ -114,7 +114,7 @@ export function operatorInboxAppendFromEvent(ev) { onCountsChanged(); } -// ─── approvals — the operator config-change / spawn approval queue ──────── +// ─── approvals — the operator config-change approval queue ──────── const APPROVAL_TAB_KEY = "hyperhive:approvals:tab"; // Derived approval state — cold-loaded from /api/state, then mutated // live by `approval_added` / `approval_resolved` dashboard events. @@ -244,24 +244,14 @@ export function renderApprovals() { el( "span", { class: "glyph" }, - isMergePr ? "⇒" : isUpdateMeta ? "↻" : isSchedule ? "⏱" : "⊕", + isUpdateMeta ? "↻" : isSchedule ? "⏱" : "⇒", ), el("span", { class: "id" }, "#" + a.id), el("span", { class: "agent" }, a.agent), el( "span", - { - class: - "kind" + - (isMergePr || isUpdateMeta || isSchedule ? "" : " kind-spawn"), - }, - isMergePr - ? "merge-pr" - : isUpdateMeta - ? "meta-update" - : isSchedule - ? "schedule" - : "spawn", + { class: "kind" }, + isUpdateMeta ? "meta-update" : isSchedule ? "schedule" : "merge-pr", ), ); if (isMergePr && a.sha_short) head.append(el("code", {}, a.sha_short)); @@ -431,13 +421,11 @@ function renderApprovalHistory(root, history) { el( "span", { class: "kind" }, - a.kind === "merge_config_pr" - ? "merge-pr" - : a.kind === "update_meta_inputs" - ? "meta-update" - : a.kind === "schedule_prompt" - ? "schedule" - : "spawn", + a.kind === "update_meta_inputs" + ? "meta-update" + : a.kind === "schedule_prompt" + ? "schedule" + : "merge-pr", ), " ", ); diff --git a/frontend/packages/dashboard/src/dashboard.css b/frontend/packages/dashboard/src/dashboard.css index e87d1af7..587dd21c 100644 --- a/frontend/packages/dashboard/src/dashboard.css +++ b/frontend/packages/dashboard/src/dashboard.css @@ -671,10 +671,6 @@ ul form.inline { letter-spacing: 0.1em; text-transform: uppercase; } -.kind-spawn { - color: var(--amber); - border-color: var(--amber); -} details { margin-top: 0.5em; } diff --git a/frontend/packages/dashboard/src/tabs.js b/frontend/packages/dashboard/src/tabs.js index 6c53a445..d8b51162 100644 --- a/frontend/packages/dashboard/src/tabs.js +++ b/frontend/packages/dashboard/src/tabs.js @@ -81,10 +81,9 @@ window.marked = marked; for (const a of approvals) { if (seenApprovals.has(a.id)) continue; seenApprovals.add(a.id); - const verb = a.kind === "spawn" ? "spawn approval" : "config commit"; NOTIF.show( "◆ approval #" + a.id, - `${verb} for ${a.agent}`, + `config commit for ${a.agent}`, "hyperhive:approval:" + a.id, ); } diff --git a/hive-c0re/src/actions.rs b/hive-c0re/src/actions.rs index e2140da0..c57e8243 100644 --- a/hive-c0re/src/actions.rs +++ b/hive-c0re/src/actions.rs @@ -22,7 +22,6 @@ use crate::lifecycle; /// FinalizeDeploy`, plus an `AfterAny` `DeployTail`, under a /// resource-holding root; ~30-90s) /// - `UpdateMetaInputs` → a `MetaUpdate` DAG (fan-out on completion) -/// - `Spawn` → a `Spawn` DAG (`Create → WriteDropin → Reconcile`) /// /// Every queued kind — deploys included — resolves its approval row via /// [`resolve_approval_dag`] when the DAG settles terminal. @@ -55,25 +54,6 @@ pub async fn approve(coord: Arc, id: i64) -> Result<()> { coord.emit_rebuild_queue_snapshot(); Ok(()) } - ApprovalKind::Spawn => { - // The spawn's tail `Reconcile` starts the container, so the - // new agent's power intent is `Up` from the outset. - if let Err(e) = coord - .power - .set(approval.agent.as_str(), crate::power::Wanted::Up) - { - tracing::warn!(agent = %approval.agent, error = ?e, "agent_power: seed on spawn failed"); - } - let inserted = coord.job_queue.insert_job(|b| { - crate::job_queue::templates::spawn(b, approval.agent.as_str(), id); - Vec::new() - }); - if let Err(e) = inserted { - return Err(e.context("insert spawn dag")); - } - coord.emit_rebuild_queue_snapshot(); - Ok(()) - } ApprovalKind::SchedulePrompt => { // No queue card for SchedulePrompt — the work is a single // sqlite insert, the actual "running" lifetime lives on @@ -525,30 +505,16 @@ pub(crate) async fn resolve_approval_dag( TerminalState::Failed => Err(anyhow::anyhow!("{}", error.unwrap_or("job dag failed"))), }; let mut terminal_tag = None; - match approval.kind { - ApprovalKind::Spawn => { - // Post-spawn forge bookkeeping (config repo mirror, meta - // access) — warn-only, then the resolution events + a rescan so - // the dashboard reflects the post-spawn state either way. - if result.is_ok() { - forge_after_first_spawn(coord, approval.agent.as_str()).await; - } else { - coord.rescan_containers_and_emit().await; - crate::dashboard::emit_tombstones_snapshot(coord).await; - } + if approval.kind == ApprovalKind::MergeConfigPr { + terminal_tag = deploy_terminal_tag(approval.agent.as_str(), approval_id, outcome).await; + // On a failed deploy, surface the failing build log back onto the + // PR so the manager sees why it was rejected without leaving the + // forge. Posted here rather than inside a node because this is the + // one place that holds the DAG's definitive error — a `MergeVerify` + // rejection and a `DeployApply` build failure both land here. + if let Err(e) = &result { + post_merge_failure_to_pr(coord, &approval, e).await; } - ApprovalKind::MergeConfigPr => { - terminal_tag = deploy_terminal_tag(approval.agent.as_str(), approval_id, outcome).await; - // On a failed deploy, surface the failing build log back onto the - // PR so the manager sees why it was rejected without leaving the - // forge. Posted here rather than inside a node because this is the - // one place that holds the DAG's definitive error — a `MergeVerify` - // rejection and a `DeployApply` build failure both land here. - if let Err(e) = &result { - post_merge_failure_to_pr(coord, &approval, e).await; - } - } - _ => {} } if let Err(e) = finish_approval(coord, &approval, result, terminal_tag).await { tracing::warn!(approval_id, error = ?e, "approval dag resolved with failure"); @@ -604,26 +570,6 @@ fn fetch_approval_for_worker( Ok(approval) } -/// Forge bookkeeping run once after the very first container spawn: -/// mirror the applied repo and grant read access to core/meta. The -/// agent's forge user and token are swarm-controller's, not this -/// hive's. Also rescans containers so the dashboard reflects the post-spawn state. -async fn forge_after_first_spawn(coord: &Arc, agent: &str) { - if let Err(e) = crate::forge::ensure_config_repo(agent).await { - tracing::warn!(%agent, error = ?e, "forge: ensure_config_repo after first spawn failed"); - } - if let Some(core_token) = crate::forge::core_token() - && let Err(e) = crate::forge::meta_read_access(agent, &core_token).await - { - tracing::warn!(%agent, error = ?e, "forge: meta_read_access after first spawn failed"); - } - if let Err(e) = crate::forge::ensure_meta_remote(agent).await { - tracing::warn!(%agent, error = ?e, "forge: ensure_meta_remote after first spawn failed"); - } - coord.rescan_containers_and_emit().await; - crate::dashboard::emit_tombstones_snapshot(coord).await; -} - async fn finish_approval( coord: &Coordinator, approval: &hive_sh4re::approvals::Approval, @@ -670,35 +616,14 @@ async fn finish_approval( note: note.clone(), description: approval.description.clone(), }); - // For spawn/rebuild approvals, also surface the underlying action so the + // For rebuild approvals, also surface the underlying action so the // manager knows whether the lifecycle step succeeded. The // ApprovalResolved event already carries the same `ok` signal but // separating it lets the manager react to the lifecycle change // without having to special-case approvals. match approval.kind { - ApprovalKind::Spawn => { - let summary = if ok { - format!("agent '{}' spawned", approval.agent) - } else { - format!( - "agent '{}' spawn FAILED: {}", - approval.agent, - note.as_deref().unwrap_or("unknown error") - ) - }; - let _ = coord - .push_todo_submitter( - approval.id, - "core", - Some(format!("spawned:{}", approval.agent)), - summary, - None, - ) - .await; - } // MergeConfigPr ends in a container rebuild — surface a Rebuilt - // lifecycle event. (It is never a first spawn — the agent already - // exists — so it never needs the Spawned arm above.) + // lifecycle event. ApprovalKind::MergeConfigPr => { let summary = crate::coordinator::rebuilt_todo_summary( approval.agent.as_str(), @@ -738,8 +663,8 @@ async fn finish_approval( /// /// Caller-specific bits stay OUT of here: fetching the PR head, the /// `verify_commit` gate, and the ff-merge. The agent always already exists here -/// (a merge is never a first spawn), so there's no `sync_agents` step — the -/// operator `Spawn` flow owns first-time meta registration. +/// (a merge is never a first deploy), so there's no `sync_agents` step — the +/// first-deploy DAG's `Provision` node owns first-time meta registration. async fn prepare_applied_target( agent: &str, applied_dir: &std::path::Path, diff --git a/hive-c0re/src/coordinator.rs b/hive-c0re/src/coordinator.rs index 3a443d2c..e1417de3 100644 --- a/hive-c0re/src/coordinator.rs +++ b/hive-c0re/src/coordinator.rs @@ -841,7 +841,7 @@ impl Coordinator { /// Call after any mutation that could affect what /// `nixos-container list` returns or what a row's /// `running` / `needs_update` / `needs_login` / `deployed_sha` - /// resolves to — lifecycle ops, destroy, approve (post-spawn), + /// resolves to — lifecycle ops, destroy, /// rebuild, meta-update, and the crash-watcher's periodic poll. /// Cheap when nothing changed (one `nixos-container list` + a /// `HashMap` diff + zero emits). diff --git a/hive-c0re/src/dashboard/approvals.rs b/hive-c0re/src/dashboard/approvals.rs index df0e5803..2d6f5d85 100644 --- a/hive-c0re/src/dashboard/approvals.rs +++ b/hive-c0re/src/dashboard/approvals.rs @@ -83,11 +83,6 @@ pub(super) fn gc_orphans(coord: &Coordinator, approvals: Vec) -> Vec, - Form(form): Form, -) -> Response { - let name = form.name.trim().to_owned(); - if name.is_empty() { - return error_response("spawn: `name` required"); - } - match state.coord.approvals.submit_kind( - &name, - hive_sh4re::approvals::ApprovalKind::Spawn, - "", - None, - "operator", - None, - ) { - Ok(id) => { - tracing::info!(%id, %name, "operator: spawn approval queued via dashboard"); - // Phase 5b: notify the dashboard event channel so live - // subscribers can append the row without a snapshot - // refetch. Spawn approvals carry no sha. - state - .coord - .emit_approval_added(crate::coordinator::ApprovalAdded { - id, - agent: &name, - approval_kind: "spawn", - sha_short: None, - description: None, - pr_number: None, - }); - (StatusCode::OK, "ok").into_response() - } - Err(e) => error_response(&format!("request-spawn {name} failed: {e:#}")), - } -} diff --git a/hive-c0re/src/dashboard/mod.rs b/hive-c0re/src/dashboard/mod.rs index cbf7933a..6364f405 100644 --- a/hive-c0re/src/dashboard/mod.rs +++ b/hive-c0re/src/dashboard/mod.rs @@ -149,7 +149,6 @@ pub async fn serve( .routes(routes!(misc_api::api_stats_hive)) .routes(routes!(misc_api::api_container_resources)) .routes(routes!(misc_api::post_mark_all_read)) - .routes(routes!(misc_api::post_request_spawn)) .routes(routes!(misc_api::post_op_send)) .routes(routes!(build_logs::get_build_logs_all)) .routes(routes!(build_logs::get_build_log_for_node)) @@ -386,7 +385,6 @@ mod router_build_probe { .routes(routes!(misc_api::api_stats_hive)) .routes(routes!(misc_api::api_container_resources)) .routes(routes!(misc_api::post_mark_all_read)) - .routes(routes!(misc_api::post_request_spawn)) .routes(routes!(misc_api::post_op_send)) .routes(routes!(build_logs::get_build_logs_all)) .routes(routes!(build_logs::get_build_log_for_node)) diff --git a/hive-c0re/src/dashboard/state_snapshot.rs b/hive-c0re/src/dashboard/state_snapshot.rs index bd14e9d1..95239cca 100644 --- a/hive-c0re/src/dashboard/state_snapshot.rs +++ b/hive-c0re/src/dashboard/state_snapshot.rs @@ -140,7 +140,7 @@ struct ApprovalHistoryView { agent: String, kind: &'static str, /// First 12 chars of the canonical sha (preferred) or - /// manager-supplied ref. None for resolved spawn approvals. + /// manager-supplied ref. None when the approval carries neither. sha_short: Option, /// `approved` / `denied` / `failed`. status: &'static str, @@ -404,7 +404,7 @@ fn build_transient_views( } /// Render each pending approval into its dashboard view (short sha for -/// `MergeConfigPr`, just the name for `Spawn`). +/// `MergeConfigPr`). /// Project a resolved sqlite row into the lean shape the dashboard /// history tab consumes — no `diff_html` (rendering 30 of them /// per /api/state poll would mean 30 git diffs per refresh). @@ -439,16 +439,6 @@ fn build_approval_views(approvals: Vec) -> Vec { let mut out = Vec::with_capacity(approvals.len()); for a in approvals { out.push(match a.kind { - hive_sh4re::approvals::ApprovalKind::Spawn => ApprovalView { - id: a.id, - agent: a.agent.to_string(), - kind: "spawn", - sha_short: None, - description: a.description, - pr_number: None, - commit_ref: None, - requested_at: a.requested_at, - }, hive_sh4re::approvals::ApprovalKind::UpdateMetaInputs => ApprovalView { id: a.id, agent: a.agent.to_string(), diff --git a/hive-c0re/src/dashboard/tombstones.rs b/hive-c0re/src/dashboard/tombstones.rs index 365496ba..22a105c6 100644 --- a/hive-c0re/src/dashboard/tombstones.rs +++ b/hive-c0re/src/dashboard/tombstones.rs @@ -31,7 +31,7 @@ pub struct TombstoneView { } /// State-dir names that don't appear in the live container list. Each -/// one surfaces in the dashboard as a row with R3V1V3 + PURG3 actions. +/// one surfaces in the dashboard as a row with a PURG3 action. /// /// ⚠️ **This lists every agent whose container is absent, not only destroyed /// ones** — a mid-spawn agent (state dir seeded by `Provision`, container not diff --git a/hive-c0re/src/dashboard_events.rs b/hive-c0re/src/dashboard_events.rs index 12639d19..e856e890 100644 --- a/hive-c0re/src/dashboard_events.rs +++ b/hive-c0re/src/dashboard_events.rs @@ -52,7 +52,7 @@ pub enum DashboardEvent { /// enough to render the dashboard row without a `/api/state` /// refetch. /// - /// The approval's own kind (`"merge_config_pr"` / `"spawn"`) lives + /// The approval's own kind (`"merge_config_pr"` / `"schedule_prompt"`) lives /// on `approval_kind` rather than `kind` because the latter is taken /// by the serde tag identifying which `DashboardEvent` variant /// this is. @@ -119,8 +119,8 @@ pub enum DashboardEvent { name: String, transient_kind: String, }, - /// One container row changed — new container appeared (post-spawn - /// finalise), an existing one flipped `running` / `needs_update` / + /// One container row changed — a new container appeared, an existing + /// one flipped `running` / `needs_update` / /// `sha`, etc. Clients upsert by `container.name`. Payload carries /// the full row so cold-loaded clients and event-driven clients /// converge on the same render. @@ -136,10 +136,8 @@ pub enum DashboardEvent { /// `nixos-container destroy` (operator-driven or otherwise) on the /// next rescan. ContainerRemoved { seq: u64, name: String }, - /// Full snapshot of the tombstones list. Emitted on every - /// mutation that could add / remove a tombstone: destroy - /// (with or without purge), purge-tombstone, spawn approval - /// (which can consume a tombstone of the same name). Snapshot + /// Full snapshot of the tombstones list. Emitted by destroy (with + /// or without purge) and purge-tombstone. Snapshot /// shape (not diff) because the list is tiny (single-digit /// typical) and recomputing avoids the add/remove races a /// per-row event would have. diff --git a/hive-c0re/src/job_queue/model.rs b/hive-c0re/src/job_queue/model.rs index e0dd4c12..3056bb4d 100644 --- a/hive-c0re/src/job_queue/model.rs +++ b/hive-c0re/src/job_queue/model.rs @@ -122,7 +122,7 @@ pub enum NodeKind { /// node-inventory row for its three responsibilities and why it isn't /// named `AbortDeploy`. DeployTail { agent: String, approval_id: i64 }, - /// Tail node of an approval-carrying DAG (spawn / opaque deploy / + /// Tail node of an approval-carrying DAG (opaque deploy / /// config-PR merge): resolve the approval row from how the work ended. ResolveApproval { approval_id: i64, diff --git a/hive-c0re/src/job_queue/templates.rs b/hive-c0re/src/job_queue/templates.rs index 3f1ef5f3..1c3104e3 100644 --- a/hive-c0re/src/job_queue/templates.rs +++ b/hive-c0re/src/job_queue/templates.rs @@ -440,27 +440,21 @@ pub fn approval_deploy(builder: &JobBuilder, agent: &str, approval_id: i64) { resolve_approval_tails(builder, approval_id, window); } -/// First-deploy spawn (approval-driven): `Provision` (proposed/applied -/// repos, state subvolume, meta registration) then `Create` -/// (`nixos-container create`), drop-in write, then `Reconcile` starts -/// the container (`wanted = Up` written at approve time). All-or-nothing: -/// `Provision` (lease-exempt, precedes the container) is the group root; -/// `Create` (child) owns the agent lease; `WriteDropin` + `Reconcile` -/// (children of `Create`) borrow it. A failure cancel-cascades the rest — -/// unlike rebuild there's no recovery-reconcile (nothing to converge if the -/// container was never created). Closed by a `ResolveApproval` tail root edged -/// `AfterAny` onto `Provision` — the DAG's only other group-root, so its roll-up -/// already carries the whole cascade. -pub fn spawn(builder: &JobBuilder, agent: &str, approval_id: i64) { - let provision = spawn_nodes(builder, agent); - resolve_approval_tails(builder, approval_id, provision); -} - -/// The spawn subgraph with no tail, returning its group root. +/// First deploy of an agent this hive has never seen, asked for by the swarm: +/// `Provision` (proposed/applied repos, state subvolume, meta registration) +/// then `Create` (`nixos-container create`), drop-in write, then `Reconcile` +/// starts the container (`wanted = Up`, seeded by +/// `swarm_status::queue_first_deploy`). All-or-nothing: `Provision` +/// (lease-exempt, precedes the container) is the group root; `Create` (child) +/// owns the agent lease; `WriteDropin` + `Reconcile` (children of `Create`) +/// borrow it. A failure cancel-cascades the rest — unlike rebuild there's no +/// recovery-reconcile (nothing to converge if the container was never created). /// -/// Split out for the same reason [`rebuild_nodes`] is: two callers want the -/// same four nodes and disagree only about what closes them. -pub(crate) fn spawn_nodes<'a>(builder: &'a JobBuilder, agent: &str) -> Handle<'a> { +/// No approval tail: the operator authorised the creation at swarm level, and +/// the deploy request carries that authorisation. +/// +/// Returns the group root so a caller can wait on the whole subtree. +pub fn first_deploy(builder: &JobBuilder, agent: &str) -> Vec { let a = || agent.to_owned(); let provision = builder .node(NodeKind::Provision { agent: a() }) @@ -479,20 +473,7 @@ pub(crate) fn spawn_nodes<'a>(builder: &'a JobBuilder, agent: &str) -> Handle<'a .needs(Resource::Agent(a())) .part_of(create) .after_ok(dropin); - provision -} - -/// First deploy of an agent this hive has never seen, asked for by the swarm. -/// -/// [`spawn`] without the approval tail, and the absence is the point rather -/// than an omission: that flow exists because an operator used to approve the -/// spawn *at the hive*. When the swarm asks, the operator has already clicked -/// create at swarm level — the deploy request carries that authorisation, and a -/// second gate here would be asking the same person the same question twice. -/// -/// Returns the group root so a caller can wait on the whole subtree. -pub fn first_deploy(builder: &JobBuilder, agent: &str) -> Vec { - vec![spawn_nodes(builder, agent).guid()] + vec![provision.guid()] } /// Teardown: `Stop` → `DestroyContainer` → (`PurgeState`) → `DestroyBookkeeping`. diff --git a/hive-c0re/src/job_queue/tests.rs b/hive-c0re/src/job_queue/tests.rs index 8b3be7d2..55631416 100644 --- a/hive-c0re/src/job_queue/tests.rs +++ b/hive-c0re/src/job_queue/tests.rs @@ -1624,10 +1624,10 @@ fn pause_shape_signal_drain() { } #[test] -fn spawn_shape_provision_create_dropin_reconcile() { +fn first_deploy_shape_provision_create_dropin_reconcile() { let q = JobQueue::new(1); insert(&q, |builder| { - templates::spawn(builder, "newbie", 7); + templates::first_deploy(builder, "newbie"); }); assert_eq!( declared_shape(&q), @@ -1636,14 +1636,6 @@ fn spawn_shape_provision_create_dropin_reconcile() { row("create", Some("provision"), &[]), row("write_dropin", Some("create"), &[]), row("reconcile", Some("create"), &[("write_dropin", "done")]), - // One tail per outcome, each edged to accept only that one — so - // *which* tail the graph lets run already is the answer, and - // nothing branches at runtime. The three differ **only** in their - // accepted outcome, which is why `declared_shape` spells the - // outcome set out instead of bucketing it. - row("resolve_approval", None, &[("provision", "done")]), - row("resolve_approval", None, &[("provision", "failed")]), - row("resolve_approval", None, &[("provision", "cancelled")]), ] ); } diff --git a/hive-c0re/src/server.rs b/hive-c0re/src/server.rs index 9ecedb6e..06ccde65 100644 --- a/hive-c0re/src/server.rs +++ b/hive-c0re/src/server.rs @@ -109,20 +109,6 @@ async fn write_response( async fn dispatch(req: &HostRequest, coord: Arc) -> HostResponse { let result: anyhow::Result = async { Ok(match req { - HostRequest::Spawn { name } => handle_spawn(&coord, name.as_str()).await?, - HostRequest::RequestSpawn { name } => { - tracing::info!(%name, "request_spawn"); - let id = coord.approvals.submit_kind( - name.as_str(), - hive_sh4re::approvals::ApprovalKind::Spawn, - "", - None, - "operator", - None, - )?; - tracing::info!(%id, %name, "spawn approval queued"); - HostResponse::success() - } HostRequest::Kill { name } => submit_single(&coord, name.as_str(), Verb::Kill).await, HostRequest::Restart { name } => { submit_single(&coord, name.as_str(), Verb::Restart).await @@ -278,49 +264,6 @@ async fn dispatch(req: &HostRequest, coord: Arc) -> HostResponse { } } -/// Create + start the container for `name` and bind its MCP listener. On a -/// failed spawn nothing was registered: post a swarm notice and return the error. -async fn handle_spawn(coord: &Arc, name: &str) -> Result { - tracing::info!(%name, "spawn"); - let agent_dir = crate::paths::agent_runtime_dir(name); - let hive = coord.hive_env(); - let paths = Coordinator::agent_paths(name, agent_dir)?; - // lifecycle::spawn creates the runtime dir internally before start. - // MCP listener registration is event-driven: bind immediately on - // success so the harness can connect on its first turn without - // waiting for any poll interval. - match lifecycle::spawn(name, &hive, &paths).await { - Ok(()) => { - if let Err(e) = coord.power.set(name, crate::power::Wanted::Up) { - tracing::warn!(%name, error = ?e, "agent_power: set wanted=up failed"); - } - // Bind the MCP listener now that the container is starting up. - // The harness connects to this socket on its first turn. - coord.register_agent(name)?; - crate::swarm_notices::notify( - "core", - Some(format!("spawned:{name}")), - format!("agent '{name}' spawned"), - None, - ) - .await; - } - Err(e) => { - // Spawn failed: register_agent was never called, so there is - // nothing to unregister. Notify the swarm and propagate. - crate::swarm_notices::notify( - "core", - Some(format!("spawned:{name}")), - format!("agent '{name}' spawn FAILED: {e:#}"), - None, - ) - .await; - return Err(e); - } - } - Ok(HostResponse::success()) -} - /// `hivectl pause|resume` / the dashboard toggle: write or remove the /// agent's pause marker. /// diff --git a/hive-c0re/src/socket_server/config_approvals.rs b/hive-c0re/src/socket_server/config_approvals.rs index c1aeeb14..1c8a9a0c 100644 --- a/hive-c0re/src/socket_server/config_approvals.rs +++ b/hive-c0re/src/socket_server/config_approvals.rs @@ -31,7 +31,7 @@ pub(crate) async fn submit_merge_config_pr( anyhow::bail!( "applied repo missing for agent '{agent}' (expected at {}) — \ merge_config_pr requires the agent to be fully provisioned; \ - spawn the agent first (operator spawn) before opening config PRs", + create it first (`swarmctl agent create`) before opening config PRs", applied_dir.display() ); } diff --git a/hive-c0re/src/stores/approvals.rs b/hive-c0re/src/stores/approvals.rs index cc201f12..1e7be46f 100644 --- a/hive-c0re/src/stores/approvals.rs +++ b/hive-c0re/src/stores/approvals.rs @@ -1,7 +1,7 @@ //! Approval queue. Requests are submitted by an agent -//! (`RequestSchedulePrompt`), the config-PR webhook (`MergeConfigPr`), or -//! the operator (`Spawn`); the user approves/denies via the host admin CLI; -//! on approval the host runs the corresponding action. +//! (`RequestSchedulePrompt`) or the config-PR webhook (`MergeConfigPr`); the +//! user approves/denies via the host admin CLI; on approval the host runs the +//! corresponding action. //! //! `UpdateMetaInputs` rows are legacy: the MCP tool that queued them was //! removed and nothing produces the kind any more. The variant and @@ -78,7 +78,7 @@ impl Approvals { /// Insert a new pending approval row. `fetched_sha` may be supplied /// when the sha is already known at submission time (e.g. `MergeConfigPr` /// fetches the PR head before inserting), making the insert + sha-set - /// atomic. Pass `None` when the kind carries no sha (e.g. `Spawn`). + /// atomic. Pass `None` when the kind carries no sha (e.g. `SchedulePrompt`). pub fn submit_kind( &self, agent: &str, @@ -382,7 +382,6 @@ fn row_to_approval(row: &rusqlite::Row<'_>) -> rusqlite::Result { // Column order: id, agent, kind, commit_ref, requested_at, status, resolved_at, note, fetched_sha, description. let kind: String = row.get(2)?; let kind = match kind.as_str() { - "spawn" => ApprovalKind::Spawn, "update_meta_inputs" => ApprovalKind::UpdateMetaInputs, "schedule_prompt" => ApprovalKind::SchedulePrompt, "merge_config_pr" => ApprovalKind::MergeConfigPr, @@ -435,7 +434,6 @@ fn row_to_approval(row: &rusqlite::Row<'_>) -> rusqlite::Result { fn kind_from_str(s: &str) -> Result { Ok(match s { - "spawn" => ApprovalKind::Spawn, "update_meta_inputs" => ApprovalKind::UpdateMetaInputs, "schedule_prompt" => ApprovalKind::SchedulePrompt, "merge_config_pr" => ApprovalKind::MergeConfigPr, @@ -467,7 +465,7 @@ mod tests { None, ) .unwrap(); - db.submit_kind("b", ApprovalKind::Spawn, "", None, "b", None) + db.submit_kind("b", ApprovalKind::SchedulePrompt, "", None, "b", None) .unwrap(); db.submit_kind("c", ApprovalKind::UpdateMetaInputs, "[]", None, "c", None) .unwrap(); @@ -508,7 +506,14 @@ mod tests { // final — re-cancelling errors instead of silently overwriting. let (_dir, _path, db) = open_temp(); let id = db - .submit_kind("a", ApprovalKind::Spawn, "deadbeef", None, "a", None) + .submit_kind( + "a", + ApprovalKind::MergeConfigPr, + "deadbeef", + None, + "a", + None, + ) .unwrap(); db.mark_cancelled(id, "manager").expect("first cancel"); let err = db @@ -567,7 +572,7 @@ mod tests { let raw = Connection::open(&path).unwrap(); raw.execute( "INSERT INTO approvals (agent, kind, commit_ref, requested_at, status) - VALUES ('old', 'spawn', '', 0, 'pending')", + VALUES ('old', 'merge_config_pr', '', 0, 'pending')", [], ) .unwrap(); diff --git a/hive-c0re/src/swarm_status.rs b/hive-c0re/src/swarm_status.rs index 5c4b8516..49825a5a 100644 --- a/hive-c0re/src/swarm_status.rs +++ b/hive-c0re/src/swarm_status.rs @@ -370,10 +370,9 @@ async fn handle_deploy_request( /// (`power::Store::get_or_seed`), and at that point the container is freshly /// created but not started — so an unseeded row locks the agent to `Offline` /// on its very first reconcile and `Reconcile` never emits the `Start` node. -/// Setting the row up front closes that window, the same way -/// `actions::approve`'s `ApprovalKind::Spawn` arm does for the -/// operator-approved path. A second caller open-coding the insert would lose -/// exactly that, and the agent would come up stopped for no visible reason. +/// Setting the row up front closes that window. A second caller open-coding +/// the insert would lose exactly that, and the agent would come up stopped for +/// no visible reason. pub(crate) fn queue_first_deploy( coord: &std::sync::Arc, agent: &str, diff --git a/hive-c0re/src/workers/mcp_sockets.rs b/hive-c0re/src/workers/mcp_sockets.rs index 5e1c040f..9b4152f8 100644 --- a/hive-c0re/src/workers/mcp_sockets.rs +++ b/hive-c0re/src/workers/mcp_sockets.rs @@ -6,8 +6,8 @@ //! re-register all running agents. //! //! After startup, listeners are managed event-driven: -//! - `run_start` (the job queue's `Start` node) and `server::handle_spawn` -//! call `register_agent` once the container is started. +//! - `run_start` (the job queue's `Start` node) calls `register_agent` +//! once the container is started. //! - `kill`/`destroy` paths call `unregister_agent`. //! //! No recurring poll is needed because c0re owns the listener lifecycle. An diff --git a/hive-host-sock/src/lib.rs b/hive-host-sock/src/lib.rs index b1d133b2..4cb68b0c 100644 --- a/hive-host-sock/src/lib.rs +++ b/hive-host-sock/src/lib.rs @@ -117,20 +117,6 @@ pub enum ReconcileDirection { #[derive(Debug, Clone, Serialize, Deserialize)] #[serde(tag = "cmd", rename_all = "snake_case")] pub enum HostRequest { - /// Create and start a brand-new sub-agent container directly (full - /// first-time provisioning: proposed/applied repos, state subvolume, - /// meta-flake sync, `nixos-container create`), bypassing the approval - /// queue. Privileged-context only. Exposed on the CLI as `hivectl - /// agent create`. See `docs/agent-lifecycle/approvals.md::Approval kinds - /// (wire shapes)`. Wire name kept as `Spawn` (unrenamed underneath - /// the CLI-verb rename — `hivectl agent start` reuses the - /// existing scope-based [`HostRequest::Start`] below instead of a - /// new per-agent variant, see its doc comment). - Spawn { name: Ident }, - /// Submit a first-creation request for the operator to approve. See - /// `docs/agent-lifecycle/approvals.md::Approval kinds (wire shapes)` (`Spawn`). - /// Exposed on the CLI as `hivectl agent request-create`. - RequestSpawn { name: Ident }, /// Hard stop a managed container. Exposed on the CLI as `hivectl /// agent kill` — kept distinct from the new graceful-only /// `hivectl agent stop` (which reuses the scope-based diff --git a/hive-sh4re/src/approvals.rs b/hive-sh4re/src/approvals.rs index 8e0867ea..70398cc7 100644 --- a/hive-sh4re/src/approvals.rs +++ b/hive-sh4re/src/approvals.rs @@ -46,9 +46,6 @@ pub struct Approval { #[serde(rename_all = "snake_case")] #[strum(serialize_all = "snake_case")] pub enum ApprovalKind { - /// Create + start a new sub-agent container with the given name - /// (under the default `agent.nix` template). - Spawn, /// Run `nix flake update [inputs...]` on the meta flake and commit /// the resulting lock changes. UpdateMetaInputs, diff --git a/hivectl/src/agents.rs b/hivectl/src/agents.rs index 66c77215..f8400580 100644 --- a/hivectl/src/agents.rs +++ b/hivectl/src/agents.rs @@ -54,7 +54,7 @@ async fn agents_restart(socket: &Path, name: &str, no_wait: bool) -> Result<()> async fn agents_start(socket: &Path, name: &str, paused: bool) -> Result<()> { if !crate::util::agent_exists(socket, name).await? { bail!( - "no such agent: '{name}' (no state dir under {}/) — use 'hivectl agent {name} create' to provision a brand-new agent", + "no such agent: '{name}' (no state dir under {}/) — a brand-new agent is created at swarm level ('swarmctl agent create')", hive_host_sock::AGENTS_ROOT ); } @@ -269,14 +269,6 @@ pub(crate) async fn run_agent(socket: &Path, name: &str, cmd: AgentCmd) -> Resul AgentCmd::Pause { no_wait } => set_paused(socket, name, true, no_wait).await, AgentCmd::Resume => set_paused(socket, name, false, false).await, AgentCmd::Start { paused } => agents_start(socket, name, paused).await, - AgentCmd::Create => { - let name = crate::util::parse_ident(name)?; - render(crate::client::request(socket, HostRequest::Spawn { name }).await?) - } - AgentCmd::RequestCreate => { - let name = crate::util::parse_ident(name)?; - render(crate::client::request(socket, HostRequest::RequestSpawn { name }).await?) - } AgentCmd::Stop => agents_stop(socket, name).await, AgentCmd::Kill => { let name = crate::util::parse_ident(name)?; diff --git a/hivectl/src/cli.rs b/hivectl/src/cli.rs index fe05d67c..72452985 100644 --- a/hivectl/src/cli.rs +++ b/hivectl/src/cli.rs @@ -414,7 +414,7 @@ pub enum AgentCmd { Resume, /// Start this EXISTING agent container. Fails immediately if `name` /// has no config/topology entry at all — it never attempts - /// first-time creation. Use `create` for that. + /// first-time creation, which is swarm-level (`swarmctl agent create`). Start { /// Start (or leave) the agent paused: if it's currently down, hivectl /// writes the pause marker before the container boots, so it comes @@ -424,14 +424,6 @@ pub enum AgentCmd { #[arg(long)] paused: bool, }, - /// Create this agent container from scratch (full first-time - /// provisioning), bypassing the approval queue. - /// - /// Operator-on-the-host only; use `request-create` for an - /// approval-gated creation. - Create, - /// Queue a first-creation request for operator approval. - RequestCreate, /// Gracefully stop this agent container: signal → drain → reconcile. /// Never escalates to a hard kill — use `kill` for that. Stop, diff --git a/swarm-controller/src/main.rs b/swarm-controller/src/main.rs index 4adcde3d..053ca140 100644 --- a/swarm-controller/src/main.rs +++ b/swarm-controller/src/main.rs @@ -686,6 +686,10 @@ struct AppState { /// "no forge configured on this host" shape every other /// forge-backed field here uses. forge: Option>, + /// Held by `create_agent` from reading where a name is placed until its + /// graph is queued, so two creations of one name cannot both find it + /// unplaced. `tokio`'s mutex, unlike `jobq`'s: the read awaits the queue. + create_gate: Arc>, } /// Env var the controller's NixOS module sets from @@ -1352,7 +1356,9 @@ struct CreateAgentResponse { responses( (status = 200, description = "job chain queued", body = CreateAgentResponse), (status = 400, description = "`name` or `hive` is not a valid identifier, `name` is new and reserved or refused by the forge, or `hive` is not in this swarm (problem+json)", body = String), - (status = 500, description = "the job chain could not be queued (problem+json)", body = String), + (status = 409, description = "the swarm has already placed `name` on a different hive (problem+json)", body = String), + (status = 503, description = "the swarm queue is not connected, so whether `name` is placed on another hive is unknown (problem+json)", body = String), + (status = 500, description = "the job chain could not be queued, or another hive's wanted state could not be read (problem+json)", body = String), ), tag = "agents" )] @@ -1424,10 +1430,25 @@ async fn create_agent( return Err(error_problem(axum::http::StatusCode::BAD_REQUEST, &detail)); } + // Agent names are one swarm-wide namespace: identity, forge user, store + // path and matrix id all carry the bare name. The same name on the same + // hive is that agent being re-created, and goes through. + let _gate = state.create_gate.lock().await; + let declared = declarations_elsewhere(&state, &agent, &hive).await?; let mut sched = state .jobq .lock() .unwrap_or_else(std::sync::PoisonError::into_inner); + let queued = queued_placements(sched.graph()); + let elsewhere = placed_elsewhere(&agent, &hive, &declared, &queued); + if !elsewhere.is_empty() { + let detail = format!( + "agent name {agent:?} is already placed on hive {} — agent names are unique across \ + the swarm; destroy it there first to move it, or choose another name", + elsewhere.join(", ") + ); + return Err(error_problem(axum::http::StatusCode::CONFLICT, &detail)); + } let ids = sched .insert_job(None, |b| declare_agent_job(b, &agent, &hive)) .map_err(|e| { @@ -1445,6 +1466,81 @@ async fn create_agent( })) } +/// Every hive but `hive`'s published wanted state, for [`placed_elsewhere`]. +/// +/// No queue wired up means no hive has been sent a declaration, so there is +/// nothing to collide with. A queue that cannot be read refuses: an unread +/// declaration and an absent one look the same, and reading "absent" is how +/// a duplicate would get through. +async fn declarations_elsewhere( + state: &AppState, + agent: &str, + hive: &str, +) -> Result, problem_details::ProblemDetails> +{ + let Some(writer) = state.wanted.as_deref() else { + return Ok(Vec::new()); + }; + let mut declared = Vec::new(); + for other in state.hives.iter().filter(|h| h.name != hive) { + match writer.view(&other.name).await { + Ok(Some(declaration)) => declared.push((other.name.clone(), declaration)), + Ok(None) => {} + Err(e) => { + let detail = format!( + "cannot tell whether {agent:?} already exists on hive {:?}, so it was not \ + created: {e:#}", + other.name + ); + return Err(error_problem(wanted_error_status(&e), &detail)); + } + } + } + Ok(declared) +} + +/// `(hive, agent)` of every `SetAgentWanted` node not yet settled: a creation +/// accepted but not yet visible in the wanted state it is about to write. +fn queued_placements( + graph: &hive_jobq::Graph, +) -> Vec<(String, String)> { + graph + .nodes() + .filter(|node| !node.state.is_terminal()) + .filter_map(|node| match &node.payload { + SwarmNodeKind::SetAgentWanted { hive, agent } => Some((hive.clone(), agent.clone())), + _ => None, + }) + .collect() +} + +/// The hives other than `hive` the swarm has placed `agent` on, sorted. +/// +/// A placement is a declaration in any state but `Destroyed`, or a queued +/// one. A `Destroyed` agent is gone from that hive, so creating it on another +/// moves it. +fn placed_elsewhere( + agent: &str, + hive: &str, + declared: &[(String, swarm_queue_client::wanted::HiveWanted)], + queued: &[(String, String)], +) -> Vec { + use swarm_queue_client::wanted::AgentState; + let declared = declared.iter().filter(|(_, declaration)| { + declaration + .agents + .get(agent) + .is_some_and(|wanted| wanted.state != AgentState::Destroyed) + }); + let queued = queued.iter().filter(|(_, queued)| queued == agent); + let hives: std::collections::BTreeSet<&str> = declared + .map(|(h, _)| h.as_str()) + .chain(queued.map(|(h, _)| h.as_str())) + .filter(|h| *h != hive) + .collect(); + hives.into_iter().map(str::to_owned).collect() +} + /// Every naming rule `agent` breaks, each as a sentence naming the rule. /// /// `reserved` is [`hive_types::reserved_names_raw`]: the list nix owns in @@ -2554,6 +2650,7 @@ async fn main() -> Result<()> { swarm_name: load_swarm_name().map(Arc::from), auth, forge: state_forge, + create_gate: Arc::default(), }; let app = build_app(state); @@ -2710,6 +2807,7 @@ mod tests { // read is the verb that needs it, and it has its own test below. auth: None, forge: None, + create_gate: std::sync::Arc::default(), }; (state, sched) } @@ -2846,6 +2944,170 @@ mod tests { assert!(queued > 0, "an accepted creation must queue work"); } + /// `state_with_roster` with a second hive, `sec0nd`. + fn state_with_two_hives() -> (super::AppState, SharedSched) { + let (state, sched) = state_with_roster(); + let mut hives = (*state.hives).clone(); + hives.push(HiveEntry { + name: "sec0nd".to_owned(), + domain: "sec0nd.example".to_owned(), + }); + let state = super::AppState { + hives: std::sync::Arc::new(hives), + ..state + }; + (state, sched) + } + + async fn create( + state: &super::AppState, + name: &str, + hive: &str, + ) -> Result<(), problem_details::ProblemDetails> { + super::create_agent( + axum::extract::State(state.clone()), + axum::Json(super::CreateAgentRequest { + name: name.to_owned(), + hive: hive.to_owned(), + }), + ) + .await + .map(|_| ()) + } + + fn node_count(sched: &SharedSched) -> usize { + sched + .lock() + .unwrap_or_else(std::sync::PoisonError::into_inner) + .graph() + .nodes() + .count() + } + + /// A name queued for one hive is refused on another, by effect: 409 and + /// nothing more queued. The queued-but-unwritten placement is the one a + /// wanted-state read alone would miss. + #[tokio::test] + async fn a_name_placed_on_another_hive_is_refused_before_anything_is_queued() { + let (state, sched) = state_with_two_hives(); + create(&state, "atlas", "pr1ma") + .await + .expect("the first creation of a name must be accepted"); + let before = node_count(&sched); + + let err = create(&state, "atlas", "sec0nd") + .await + .expect_err("the same name on another hive must be refused"); + assert_eq!( + err.status, + Some(axum::http::StatusCode::CONFLICT), + "{err:?}" + ); + assert!( + format!("{err:?}").contains("pr1ma"), + "the refusal should name the hive holding the name, got: {err:?}" + ); + assert_eq!( + node_count(&sched), + before, + "a refused creation must queue no work" + ); + } + + /// The control for the refusal above: the same name on the SAME hive is + /// that agent being re-created, and queues again. + #[tokio::test] + async fn re_creating_an_agent_on_its_own_hive_is_accepted() { + let (state, sched) = state_with_two_hives(); + create(&state, "atlas", "pr1ma") + .await + .expect("the first creation of a name must be accepted"); + let before = node_count(&sched); + + create(&state, "atlas", "pr1ma") + .await + .expect("re-creating an agent on its own hive must be accepted"); + assert!(node_count(&sched) > before, "a re-creation must queue work"); + } + + /// A wanted state that cannot be read refuses rather than reading as + /// "placed nowhere", and queues nothing. + #[tokio::test] + async fn an_unreadable_wanted_state_refuses_creation() { + let (state, sched) = state_with_two_hives(); + let state = super::AppState { + wanted: Some(std::sync::Arc::new(wanted::WantedWriter::new( + disconnected_client().await, + ))), + ..state + }; + + let err = create(&state, "atlas", "pr1ma") + .await + .expect_err("an unreadable placement must refuse"); + assert_eq!( + err.status, + Some(axum::http::StatusCode::SERVICE_UNAVAILABLE), + "{err:?}" + ); + assert_eq!( + node_count(&sched), + 0, + "a refused creation must queue no work" + ); + } + + fn declaring( + hive: &str, + agent: &str, + state: swarm_queue_client::wanted::AgentState, + ) -> (String, swarm_queue_client::wanted::HiveWanted) { + let mut declaration = swarm_queue_client::wanted::HiveWanted::default(); + declaration.agents.insert( + agent.to_owned(), + swarm_queue_client::wanted::AgentWanted { state }, + ); + (hive.to_owned(), declaration) + } + + #[test] + fn a_live_declaration_on_another_hive_is_a_placement() { + use swarm_queue_client::wanted::AgentState; + for state in [AgentState::Up, AgentState::Offline, AgentState::Paused] { + let declared = [declaring("sec0nd", "atlas", state)]; + assert_eq!( + super::placed_elsewhere("atlas", "pr1ma", &declared, &[]), + ["sec0nd"], + "{state:?}" + ); + } + } + + /// A destroyed agent has left its hive, so creating it elsewhere moves it. + #[test] + fn a_destroyed_declaration_is_not_a_placement() { + let declared = [declaring( + "sec0nd", + "atlas", + swarm_queue_client::wanted::AgentState::Destroyed, + )]; + assert!(super::placed_elsewhere("atlas", "pr1ma", &declared, &[]).is_empty()); + } + + #[test] + fn placements_on_the_requested_hive_or_of_other_names_do_not_count() { + use swarm_queue_client::wanted::AgentState; + let declared = [ + declaring("pr1ma", "atlas", AgentState::Up), + declaring("sec0nd", "argus", AgentState::Up), + ]; + let queued = [ + ("pr1ma".to_owned(), "atlas".to_owned()), + ("sec0nd".to_owned(), "argus".to_owned()), + ]; + assert!(super::placed_elsewhere("atlas", "pr1ma", &declared, &queued).is_empty()); + } + /// Forgejo v16.0.5's reserved usernames our charset can spell, plus the /// bare `-` — the forge section of `nix/reserved-names.nix`. Read from /// the list nix exports, so dropping one from the file reds this. diff --git a/swarm-controller/src/matrix_account.rs b/swarm-controller/src/matrix_account.rs index 0cad702f..ad71fa48 100644 --- a/swarm-controller/src/matrix_account.rs +++ b/swarm-controller/src/matrix_account.rs @@ -580,6 +580,7 @@ mod tests { swarm_name: None, auth: None, forge: None, + create_gate: std::sync::Arc::default(), } } diff --git a/swarm-secret-client/src/policy.rs b/swarm-secret-client/src/policy.rs index f8b4460b..febbd9d6 100644 --- a/swarm-secret-client/src/policy.rs +++ b/swarm-secret-client/src/policy.rs @@ -83,10 +83,12 @@ pub const AGENT_PREFIX: &str = "hive-agent-"; /// attaches the policy and matches the subject by spelling all three the same. /// /// Injective in `agent` — the prefix is fixed and [`checked_segment`] has -/// already refused anything that could re-punctuate the suffix — and agent -/// names are one swarm-wide namespace (an agent's path is -/// `swarm/agents/`, with no hive segment to disambiguate two that -/// matched), so two agents cannot land on one name. +/// already refused anything that could re-punctuate the suffix. Agent names +/// are one swarm-wide namespace (an agent's path is `swarm/agents/`, +/// with no hive segment), so two agents sharing a name would share every +/// object this names. swarm-controller's `create_agent`, the one place an +/// agent is created, refuses a name the swarm has already placed on another +/// hive. An agent no hive's wanted state declares is invisible to that check. /// /// # Errors /// [`Error::PathSegment`] when `agent` holds anything but `[A-Za-z0-9_-]`. diff --git a/swarmctl/src/main.rs b/swarmctl/src/main.rs index 72276028..f086929b 100644 --- a/swarmctl/src/main.rs +++ b/swarmctl/src/main.rs @@ -206,7 +206,9 @@ struct AgentCreateArgs { /// Name for the new agent: 1–63 characters of `[a-z0-9-]`. /// /// Becomes an SSO subject, a forge user and a repository name, so - /// it's validated here before queuing. + /// it's validated here before queuing. The controller refuses a name + /// the swarm has already placed on a different hive; the same name on + /// the same hive re-creates that agent. name: String, /// Hive in this swarm to deploy the agent to. ///