Watch
0
0
Fork
You've already forked hyperhive
0

Make agent creation swarm-only and refuse a name placed on another hive

swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.

Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.

policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.

Refs #4396
This commit is contained in:
atlas 2026-09-29 15:11:57 +02:00
commit 5785c0024c
35 changed files with 376 additions and 434 deletions

View file

@ -100,7 +100,8 @@ other agents don't:
- **Naming/bootstrap** — the manager's broker recipient name, state-dir
key, and nixos-container name are all `ruth` (container `h-ruth`).
`hive-c0re` spawns it directly at boot if missing, with no operator
approval step — every other agent goes through a `Spawn` approval.
approval step — every other agent is created at swarm level
(`swarmctl agent create`).
Roster-wise, `ruth` is just another entry.
- **Wire-protocol** — the privileged `Request` variants
(`Kill` / `Start` / `Restart` / `Update`; `GetLogs`) — marked

View file

@ -24,14 +24,6 @@ CLI) before it takes effect. What you'll see, and what to do with it:
anything in that chain fails, the change rolls back automatically —
the agent stays on its last-good config, no recovery action needed
from you.
- **New agent** (`Spawn`) — the one approval in creating a brand-new
agent. The swarm controller's `InitAgentConfigRepo` job
(`POST /api/agents`) scaffolds its config repo first, outside the
approval queue; `Spawn` then creates the container from that
config. Tailoring what the template seeded isn't a separate
mechanism — it's the config-change flow above, a PR you review like
any other. Every later change goes through that flow — there's no
repeat "spawn" for an existing agent.
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
one or more Nix flake inputs (or all of them). Approving runs the
update and commits the lock change; it doesn't rebuild anything by
@ -133,11 +125,11 @@ agent that lacks the `approvals` tool group: only an agent with that
group submits approvals (for its direct children), so an agent
without it has nothing of its own to withdraw.
The swarm controller's `InitAgentConfigRepo` job creates a brand-new
agent's config repo outside this queue, seeding it with a default
`agent.nix` template. The operator then **spawns** the agent
(the `Spawn` approval / `◆ R3QU3ST SP4WN` button), which creates the
container from that config.
Creating a brand-new agent isn't an approval: it's swarm-level
(`swarmctl agent create`, `POST /api/agents` on the swarm controller),
and no hive can originate an agent. The swarm controller's
`InitAgentConfigRepo` job seeds the agent's config repo with a default
`agent.nix` template, then asks the target hive to deploy it.
Changing what the template seeded isn't a special case: like every
later change, it's a PR on that config repo (`MergeConfigPr`), made
@ -147,7 +139,7 @@ through the web UI or the forge.
### Approval kinds (wire shapes)
`ApprovalKind` carries four variants; each maps to a different
`ApprovalKind` carries three variants; each maps to a different
`commit_ref` encoding because `ApprovalKind` overloads that field as
the kind-specific payload carrier.
@ -165,17 +157,6 @@ the kind-specific payload carrier.
eval-verifies it; `DeployApply` fast-forward-merges the forge config
repo's `main` to it (the merge) and runs `deploy_applied_target`;
`DeployTail` compensates on failure. Never a first spawn.
- `Spawn` — direct container creation from the agent's config repo.
`commit_ref` is empty. Submitted via `HostRequest::RequestSpawn`
(operator-gated, the `◆ R3QU3ST SP4WN` dashboard button +
`hivectl agent <name> request-create` CLI). The host-level `HostRequest::Spawn`
variant bypasses the approval queue entirely — privileged-context use
only (operator on the host shell, test scripts, one-off recoveries;
`hivectl agent <name> create`). This is the **canonical first-spawn**: the
swarm controller's `InitAgentConfigRepo` job seeds the agent's config
repo, it gets customised through a PR, then the operator spawns to
create the container. Subsequent config changes go through a
`MergeConfigPr` PR.
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
etc.). hive-c0re sets the `agent` field to the requesting root agent.
@ -412,8 +393,8 @@ submitter pushes again (or closes it) to retry.
### Dispatch via the job queue
Long-running approval work — `MergeConfigPr`, `UpdateMetaInputs`,
`Spawn` — runs as a DAG on the global job queue
Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs`
— runs as a DAG on the global job queue
(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler
rather than run inline:
@ -421,7 +402,6 @@ rather than run inline:
|---|---|---|
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
| `Spawn` | `spawn` (`Create → WriteDropin → Reconcile`) | `approval` |
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
The DAG carries the originating `approval_id`, surfaced on the node that
@ -433,7 +413,7 @@ state is the authoritative outcome. That hook fires the matching
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
terminal tag (verifying the tag actually resolves in the applied repo —
a pre-merge rejection plants none), posts the failing build log back to
the config PR, and for a spawn runs the post-spawn forge bookkeeping.
the config PR.
Two visible consequences:
@ -612,12 +592,12 @@ event to the agent that originally submitted the approval (looked up from
the `submitter` column on the `approvals` table). The harness delivers it
as a regular `system` inbox message so it drives a normal claude turn.
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
**every** approval kind's terminal state, `Spawn` included. A
**every** approval kind's terminal state. A
"FYI, check when convenient" event doesn't need a message — those go
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct
live dial of the target agent's in-container todo socket (same
`UpsertTodo` request in-container producers use); `finish_approval` fires
one of these too for `Spawn`/`MergeConfigPr`, *in addition to*
one of these too for `MergeConfigPr`, *in addition to*
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
approval rows that predate the submitter column fall back to the
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
@ -651,7 +631,7 @@ hive-c0re-vouched commit sha. Optional `tag` carries the deploy
bookkeeping tag — `deployed/<id>` on a successful build or
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
Both fields are `Option`: `None` on the paths that don't deploy a new
commit (spawn / meta-update / deny, and the autoupdate
commit (meta-update / deny, and the autoupdate
sweep's `job_queue::templates::rebuild` reapplying the existing main,
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
`git show <sha>` against `/applied/<n>/.git` inside the

View file

@ -10,8 +10,8 @@ keeps its state, purging it doesn't.**
- **`DESTR0Y`** (the default action) stops and removes the container
but keeps everything on disk — config history, claude login, `/state/`
notes, harness data. The agent shows up as a tombstone (K3PT ST4T3 on
the C0R3 page) with a `⊕ R3V1V3` button that recreates it from the
kept state, **no re-login needed**.
the C0R3 page); `swarmctl agent create` with the same name and hive
recreates it from the kept state, **no re-login needed**.
- **`PURG3`** (opt-in, from the dashboard or `hivectl agent <name>
destroy --purge`) is `DESTR0Y` plus wiping all of it — config
history, claude credentials, `/state/` notes, everything. **No
@ -432,8 +432,8 @@ See [For operators](#for-operators) above for what each action does to
an agent's state. The mechanics, for completeness:
- `DESTR0Y` also drops the systemd drop-in and fails any pending
approvals; the tombstone's `⊕ R3V1V3` button queues a Spawn approval
that reuses the kept state on approve.
approvals; reviving the agent is `swarmctl agent create` with the same
name and hive, which reuses the kept state.
- `PURG3` wipes `/var/lib/hyperhive/{agents,applied}/<name>/` — the
union of everything `DESTR0Y` left behind.

View file

@ -303,17 +303,14 @@ a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`).
### 7 · Spawn sub-agents
Sub-agent creation is an operator action — agents have no tool for it.
Two steps:
Sub-agent creation is a swarm-level operator action — agents have no
tool for it, and no hive can create one on its own:
```
# Step 1: scaffold the new agent's config repo. The swarm controller's
# InitAgentConfigRepo job does this (POST /api/agents), seeding
# /agents/iris/config/agent.nix from the default template.
# Step 2: edit /agents/iris/config/agent.nix and commit it. Then spawn
# iris from the dashboard (◆ R3QU3ST SP4WN / Spawn approval), which
# builds + starts the container from that config.
# Create iris on hive pr1ma. The swarm controller seeds its config repo
# (agent-configs/iris) from the default template, then asks pr1ma to
# build + start the container from that config.
swarmctl agent create iris --hive pr1ma
# Later config changes: open a PR on agent-configs/iris (hive-forge);
# the operator reviews + approves it — no MCP tool call.

View file

@ -322,7 +322,7 @@ the build slot while its parent held the meta window could block waiting for
a resource its own parent already committed to, a lock-ordering hazard that
one multi-resource root avoids by construction.
`Spawn` and `UpdateMetaInputs` approvals map onto the ordinary `spawn` /
`UpdateMetaInputs` approvals map onto the ordinary
`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag`
exactly once when **any** approval-carrying DAG settles terminal — deploys
included, since their outcome is now the DAG's own state (including

View file

@ -21,8 +21,6 @@ This document contains the help content for the `hivectl` command-line program.
* [`hivectl agent pause`↴](#hivectl-agent-pause)
* [`hivectl agent resume`↴](#hivectl-agent-resume)
* [`hivectl agent start`↴](#hivectl-agent-start)
* [`hivectl agent create`↴](#hivectl-agent-create)
* [`hivectl agent request-create`↴](#hivectl-agent-request-create)
* [`hivectl agent stop`↴](#hivectl-agent-stop)
* [`hivectl agent kill`↴](#hivectl-agent-kill)
* [`hivectl agent destroy`↴](#hivectl-agent-destroy)
@ -271,9 +269,7 @@ Everything here targets a single named agent (`hivectl agent foo restart`, `hive
* `restart` — Stop and start this agent container without rebuilding config
* `pause` — Park this agent's turn loop, leaving the container running
* `resume` — Resume this paused agent — it drains whatever queued up while parked
* `start` — Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation. Use `create` for that
* `create` — Create this agent container from scratch (full first-time provisioning), bypassing the approval queue
* `request-create` — Queue a first-creation request for operator approval
* `start` — Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation, which is swarm-level (`swarmctl agent create`)
* `stop` — Gracefully stop this agent container: signal → drain → reconcile. Never escalates to a hard kill — use `kill` for that
* `kill` — Hard-stop this managed container
* `destroy` — Tear down this sub-agent container, keeping its state by default. No undo
@ -326,7 +322,7 @@ Resume this paused agent — it drains whatever queued up while parked
## `hivectl agent start`
Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation. Use `create` for that
Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation, which is swarm-level (`swarmctl agent create`)
**Usage:** `hivectl agent start [OPTIONS]`
@ -336,24 +332,6 @@ Start this EXISTING agent container. Fails immediately if `name` has no config/t
## `hivectl agent create`
Create this agent container from scratch (full first-time provisioning), bypassing the approval queue.
Operator-on-the-host only; use `request-create` for an approval-gated creation.
**Usage:** `hivectl agent create`
## `hivectl agent request-create`
Queue a first-creation request for operator approval
**Usage:** `hivectl agent request-create`
## `hivectl agent stop`
Gracefully stop this agent container: signal → drain → reconcile. Never escalates to a hard kill — use `kill` for that

View file

@ -70,7 +70,7 @@ No approval gate guards this: running this binary already means being root on th
* `<NAME>` — Name for the new agent: 1–63 characters of `[a-z0-9-]`.
Becomes an SSO subject, a forge user and a repository name, so it's validated here before queuing.
Becomes an SSO subject, a forge user and a repository name, so it's validated here before queuing. The controller refuses a name the swarm has already placed on a different hive; the same name on the same hive re-creates that agent.
###### **Options:**

View file

@ -971,7 +971,6 @@ renderApprovals`) with three stacked sections:
| `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) |
| `update_meta_inputs` | `↻` | `meta-update` | — |
| `schedule_prompt` | `⏱` | `schedule` | — |
| `spawn` | `⊕` | `spawn` | — |
<!-- vale write-good.Passive = NO -->
The chip ticks live every second via a `data-requested-at`
@ -985,7 +984,6 @@ renderApprovals`) with three stacked sections:
config PR into `agent-configs/<agent>/pulls/<pr_number>` (shown
only when `forge_present` is true and `pr_number` has a value). The config diff
lives on the forge PR itself — no inline diff side-panel.
- `spawn`: a one-line "container will be created" note instead.
- **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a
`prompt()` for an optional reason carried to the submitting agent as
`HelperEvent::ApprovalResolved.note`.
@ -1043,7 +1041,6 @@ below — some endpoints aren't in it yet.
- `POST /api/{rebuild,kill,restart,start,destroy}/{name}` — lifecycle.
`destroy` accepts `purge=on` to also wipe state dirs.
- `POST /api/purge-tombstone/{name}` — wipe a tombstone's state dirs.
- `POST /api/request-spawn` — queue a Spawn approval.
- `POST /api/update-all` — rebuild every stale container.
- `POST /api/rebuild-queue/{id}/cancel` — drop a `Queued` entry.
Refuses `Running` / terminal-state entries (in-flight
@ -1243,10 +1240,9 @@ payload):
- `container_state_changed` (container: ContainerView) /
`container_removed` (name) — per-row container mutations,
emitted by `Coordinator::rescan_containers_and_emit` from
many mutation sites — post-spawn approval bookkeeping
(`actions::approve`), the job queue's own node execution
(`job_queue::exec`, for example after a rebuild's stop/swap/start
steps or a destroy's teardown step) — and from the 10s
the job queue's own node execution (`job_queue::exec`, for example
after a rebuild's stop/swap/start steps or a destroy's teardown
step) — and from the 10s
`crash_watch` poll. Client upserts/removes by name and
reads the pending overlay from `transientsState` since the
payload doesn't carry it.