Watch
0
0
Fork
You've already forked hyperhive
0

Make agent creation swarm-only and refuse a name placed on another hive

swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.

Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.

policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.

Refs #4396
This commit is contained in:
atlas 2026-09-29 15:11:57 +02:00
commit 5785c0024c
35 changed files with 376 additions and 434 deletions

View file

@ -100,7 +100,8 @@ other agents don't:
- **Naming/bootstrap** — the manager's broker recipient name, state-dir
key, and nixos-container name are all `ruth` (container `h-ruth`).
`hive-c0re` spawns it directly at boot if missing, with no operator
approval step — every other agent goes through a `Spawn` approval.
approval step — every other agent is created at swarm level
(`swarmctl agent create`).
Roster-wise, `ruth` is just another entry.
- **Wire-protocol** — the privileged `Request` variants
(`Kill` / `Start` / `Restart` / `Update`; `GetLogs`) — marked

View file

@ -24,14 +24,6 @@ CLI) before it takes effect. What you'll see, and what to do with it:
anything in that chain fails, the change rolls back automatically —
the agent stays on its last-good config, no recovery action needed
from you.
- **New agent** (`Spawn`) — the one approval in creating a brand-new
agent. The swarm controller's `InitAgentConfigRepo` job
(`POST /api/agents`) scaffolds its config repo first, outside the
approval queue; `Spawn` then creates the container from that
config. Tailoring what the template seeded isn't a separate
mechanism — it's the config-change flow above, a PR you review like
any other. Every later change goes through that flow — there's no
repeat "spawn" for an existing agent.
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
one or more Nix flake inputs (or all of them). Approving runs the
update and commits the lock change; it doesn't rebuild anything by
@ -133,11 +125,11 @@ agent that lacks the `approvals` tool group: only an agent with that
group submits approvals (for its direct children), so an agent
without it has nothing of its own to withdraw.
The swarm controller's `InitAgentConfigRepo` job creates a brand-new
agent's config repo outside this queue, seeding it with a default
`agent.nix` template. The operator then **spawns** the agent
(the `Spawn` approval / `◆ R3QU3ST SP4WN` button), which creates the
container from that config.
Creating a brand-new agent isn't an approval: it's swarm-level
(`swarmctl agent create`, `POST /api/agents` on the swarm controller),
and no hive can originate an agent. The swarm controller's
`InitAgentConfigRepo` job seeds the agent's config repo with a default
`agent.nix` template, then asks the target hive to deploy it.
Changing what the template seeded isn't a special case: like every
later change, it's a PR on that config repo (`MergeConfigPr`), made
@ -147,7 +139,7 @@ through the web UI or the forge.
### Approval kinds (wire shapes)
`ApprovalKind` carries four variants; each maps to a different
`ApprovalKind` carries three variants; each maps to a different
`commit_ref` encoding because `ApprovalKind` overloads that field as
the kind-specific payload carrier.
@ -165,17 +157,6 @@ the kind-specific payload carrier.
eval-verifies it; `DeployApply` fast-forward-merges the forge config
repo's `main` to it (the merge) and runs `deploy_applied_target`;
`DeployTail` compensates on failure. Never a first spawn.
- `Spawn` — direct container creation from the agent's config repo.
`commit_ref` is empty. Submitted via `HostRequest::RequestSpawn`
(operator-gated, the `◆ R3QU3ST SP4WN` dashboard button +
`hivectl agent <name> request-create` CLI). The host-level `HostRequest::Spawn`
variant bypasses the approval queue entirely — privileged-context use
only (operator on the host shell, test scripts, one-off recoveries;
`hivectl agent <name> create`). This is the **canonical first-spawn**: the
swarm controller's `InitAgentConfigRepo` job seeds the agent's config
repo, it gets customised through a PR, then the operator spawns to
create the container. Subsequent config changes go through a
`MergeConfigPr` PR.
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
etc.). hive-c0re sets the `agent` field to the requesting root agent.
@ -412,8 +393,8 @@ submitter pushes again (or closes it) to retry.
### Dispatch via the job queue
Long-running approval work — `MergeConfigPr`, `UpdateMetaInputs`,
`Spawn` — runs as a DAG on the global job queue
Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs`
— runs as a DAG on the global job queue
(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler
rather than run inline:
@ -421,7 +402,6 @@ rather than run inline:
|---|---|---|
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
| `Spawn` | `spawn` (`Create → WriteDropin → Reconcile`) | `approval` |
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
The DAG carries the originating `approval_id`, surfaced on the node that
@ -433,7 +413,7 @@ state is the authoritative outcome. That hook fires the matching
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
terminal tag (verifying the tag actually resolves in the applied repo —
a pre-merge rejection plants none), posts the failing build log back to
the config PR, and for a spawn runs the post-spawn forge bookkeeping.
the config PR.
Two visible consequences:
@ -612,12 +592,12 @@ event to the agent that originally submitted the approval (looked up from
the `submitter` column on the `approvals` table). The harness delivers it
as a regular `system` inbox message so it drives a normal claude turn.
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
**every** approval kind's terminal state, `Spawn` included. A
**every** approval kind's terminal state. A
"FYI, check when convenient" event doesn't need a message — those go
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct
live dial of the target agent's in-container todo socket (same
`UpsertTodo` request in-container producers use); `finish_approval` fires
one of these too for `Spawn`/`MergeConfigPr`, *in addition to*
one of these too for `MergeConfigPr`, *in addition to*
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
approval rows that predate the submitter column fall back to the
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
@ -651,7 +631,7 @@ hive-c0re-vouched commit sha. Optional `tag` carries the deploy
bookkeeping tag — `deployed/<id>` on a successful build or
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
Both fields are `Option`: `None` on the paths that don't deploy a new
commit (spawn / meta-update / deny, and the autoupdate
commit (meta-update / deny, and the autoupdate
sweep's `job_queue::templates::rebuild` reapplying the existing main,
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
`git show <sha>` against `/applied/<n>/.git` inside the

View file

@ -10,8 +10,8 @@ keeps its state, purging it doesn't.**
- **`DESTR0Y`** (the default action) stops and removes the container
but keeps everything on disk — config history, claude login, `/state/`
notes, harness data. The agent shows up as a tombstone (K3PT ST4T3 on
the C0R3 page) with a `⊕ R3V1V3` button that recreates it from the
kept state, **no re-login needed**.
the C0R3 page); `swarmctl agent create` with the same name and hive
recreates it from the kept state, **no re-login needed**.
- **`PURG3`** (opt-in, from the dashboard or `hivectl agent <name>
destroy --purge`) is `DESTR0Y` plus wiping all of it — config
history, claude credentials, `/state/` notes, everything. **No
@ -432,8 +432,8 @@ See [For operators](#for-operators) above for what each action does to
an agent's state. The mechanics, for completeness:
- `DESTR0Y` also drops the systemd drop-in and fails any pending
approvals; the tombstone's `⊕ R3V1V3` button queues a Spawn approval
that reuses the kept state on approve.
approvals; reviving the agent is `swarmctl agent create` with the same
name and hive, which reuses the kept state.
- `PURG3` wipes `/var/lib/hyperhive/{agents,applied}/<name>/` — the
union of everything `DESTR0Y` left behind.

View file

@ -303,17 +303,14 @@ a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`).
### 7 · Spawn sub-agents
Sub-agent creation is an operator action — agents have no tool for it.
Two steps:
Sub-agent creation is a swarm-level operator action — agents have no
tool for it, and no hive can create one on its own:
```
# Step 1: scaffold the new agent's config repo. The swarm controller's
# InitAgentConfigRepo job does this (POST /api/agents), seeding
# /agents/iris/config/agent.nix from the default template.
# Step 2: edit /agents/iris/config/agent.nix and commit it. Then spawn
# iris from the dashboard (◆ R3QU3ST SP4WN / Spawn approval), which
# builds + starts the container from that config.
# Create iris on hive pr1ma. The swarm controller seeds its config repo
# (agent-configs/iris) from the default template, then asks pr1ma to
# build + start the container from that config.
swarmctl agent create iris --hive pr1ma
# Later config changes: open a PR on agent-configs/iris (hive-forge);
# the operator reviews + approves it — no MCP tool call.

View file

@ -322,7 +322,7 @@ the build slot while its parent held the meta window could block waiting for
a resource its own parent already committed to, a lock-ordering hazard that
one multi-resource root avoids by construction.
`Spawn` and `UpdateMetaInputs` approvals map onto the ordinary `spawn` /
`UpdateMetaInputs` approvals map onto the ordinary
`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag`
exactly once when **any** approval-carrying DAG settles terminal — deploys
included, since their outcome is now the DAG's own state (including

View file

@ -21,8 +21,6 @@ This document contains the help content for the `hivectl` command-line program.
* [`hivectl agent pause`↴](#hivectl-agent-pause)
* [`hivectl agent resume`↴](#hivectl-agent-resume)
* [`hivectl agent start`↴](#hivectl-agent-start)
* [`hivectl agent create`↴](#hivectl-agent-create)
* [`hivectl agent request-create`↴](#hivectl-agent-request-create)
* [`hivectl agent stop`↴](#hivectl-agent-stop)
* [`hivectl agent kill`↴](#hivectl-agent-kill)
* [`hivectl agent destroy`↴](#hivectl-agent-destroy)
@ -271,9 +269,7 @@ Everything here targets a single named agent (`hivectl agent foo restart`, `hive
* `restart` — Stop and start this agent container without rebuilding config
* `pause` — Park this agent's turn loop, leaving the container running
* `resume` — Resume this paused agent — it drains whatever queued up while parked
* `start` — Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation. Use `create` for that
* `create` — Create this agent container from scratch (full first-time provisioning), bypassing the approval queue
* `request-create` — Queue a first-creation request for operator approval
* `start` — Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation, which is swarm-level (`swarmctl agent create`)
* `stop` — Gracefully stop this agent container: signal → drain → reconcile. Never escalates to a hard kill — use `kill` for that
* `kill` — Hard-stop this managed container
* `destroy` — Tear down this sub-agent container, keeping its state by default. No undo
@ -326,7 +322,7 @@ Resume this paused agent — it drains whatever queued up while parked
## `hivectl agent start`
Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation. Use `create` for that
Start this EXISTING agent container. Fails immediately if `name` has no config/topology entry at all — it never attempts first-time creation, which is swarm-level (`swarmctl agent create`)
**Usage:** `hivectl agent start [OPTIONS]`
@ -336,24 +332,6 @@ Start this EXISTING agent container. Fails immediately if `name` has no config/t
## `hivectl agent create`
Create this agent container from scratch (full first-time provisioning), bypassing the approval queue.
Operator-on-the-host only; use `request-create` for an approval-gated creation.
**Usage:** `hivectl agent create`
## `hivectl agent request-create`
Queue a first-creation request for operator approval
**Usage:** `hivectl agent request-create`
## `hivectl agent stop`
Gracefully stop this agent container: signal → drain → reconcile. Never escalates to a hard kill — use `kill` for that

View file

@ -70,7 +70,7 @@ No approval gate guards this: running this binary already means being root on th
* `<NAME>` — Name for the new agent: 1–63 characters of `[a-z0-9-]`.
Becomes an SSO subject, a forge user and a repository name, so it's validated here before queuing.
Becomes an SSO subject, a forge user and a repository name, so it's validated here before queuing. The controller refuses a name the swarm has already placed on a different hive; the same name on the same hive re-creates that agent.
###### **Options:**

View file

@ -971,7 +971,6 @@ renderApprovals`) with three stacked sections:
| `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) |
| `update_meta_inputs` | `↻` | `meta-update` | — |
| `schedule_prompt` | `⏱` | `schedule` | — |
| `spawn` | `⊕` | `spawn` | — |
<!-- vale write-good.Passive = NO -->
The chip ticks live every second via a `data-requested-at`
@ -985,7 +984,6 @@ renderApprovals`) with three stacked sections:
config PR into `agent-configs/<agent>/pulls/<pr_number>` (shown
only when `forge_present` is true and `pr_number` has a value). The config diff
lives on the forge PR itself — no inline diff side-panel.
- `spawn`: a one-line "container will be created" note instead.
- **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a
`prompt()` for an optional reason carried to the submitting agent as
`HelperEvent::ApprovalResolved.note`.
@ -1043,7 +1041,6 @@ below — some endpoints aren't in it yet.
- `POST /api/{rebuild,kill,restart,start,destroy}/{name}` — lifecycle.
`destroy` accepts `purge=on` to also wipe state dirs.
- `POST /api/purge-tombstone/{name}` — wipe a tombstone's state dirs.
- `POST /api/request-spawn` — queue a Spawn approval.
- `POST /api/update-all` — rebuild every stale container.
- `POST /api/rebuild-queue/{id}/cancel` — drop a `Queued` entry.
Refuses `Running` / terminal-state entries (in-flight
@ -1243,10 +1240,9 @@ payload):
- `container_state_changed` (container: ContainerView) /
`container_removed` (name) — per-row container mutations,
emitted by `Coordinator::rescan_containers_and_emit` from
many mutation sites — post-spawn approval bookkeeping
(`actions::approve`), the job queue's own node execution
(`job_queue::exec`, for example after a rebuild's stop/swap/start
steps or a destroy's teardown step) — and from the 10s
the job queue's own node execution (`job_queue::exec`, for example
after a rebuild's stop/swap/start steps or a destroy's teardown
step) — and from the 10s
`crash_watch` poll. Client upserts/removes by name and
reads the pending overlay from `transientsState` since the
payload doesn't carry it.

View file

@ -114,7 +114,7 @@ export function operatorInboxAppendFromEvent(ev) {
onCountsChanged();
}
// ─── approvals — the operator config-change / spawn approval queue ────────
// ─── approvals — the operator config-change approval queue ────────
const APPROVAL_TAB_KEY = "hyperhive:approvals:tab";
// Derived approval state — cold-loaded from /api/state, then mutated
// live by `approval_added` / `approval_resolved` dashboard events.
@ -244,24 +244,14 @@ export function renderApprovals() {
el(
"span",
{ class: "glyph" },
isMergePr ? "⇒" : isUpdateMeta ? "↻" : isSchedule ? "⏱" : "⊕",
isUpdateMeta ? "↻" : isSchedule ? "⏱" : "⇒",
),
el("span", { class: "id" }, "#" + a.id),
el("span", { class: "agent" }, a.agent),
el(
"span",
{
class:
"kind" +
(isMergePr || isUpdateMeta || isSchedule ? "" : " kind-spawn"),
},
isMergePr
? "merge-pr"
: isUpdateMeta
? "meta-update"
: isSchedule
? "schedule"
: "spawn",
{ class: "kind" },
isUpdateMeta ? "meta-update" : isSchedule ? "schedule" : "merge-pr",
),
);
if (isMergePr && a.sha_short) head.append(el("code", {}, a.sha_short));
@ -431,13 +421,11 @@ function renderApprovalHistory(root, history) {
el(
"span",
{ class: "kind" },
a.kind === "merge_config_pr"
? "merge-pr"
: a.kind === "update_meta_inputs"
a.kind === "update_meta_inputs"
? "meta-update"
: a.kind === "schedule_prompt"
? "schedule"
: "spawn",
: "merge-pr",
),
" ",
);

View file

@ -671,10 +671,6 @@ ul form.inline {
letter-spacing: 0.1em;
text-transform: uppercase;
}
.kind-spawn {
color: var(--amber);
border-color: var(--amber);
}
details {
margin-top: 0.5em;
}

View file

@ -81,10 +81,9 @@ window.marked = marked;
for (const a of approvals) {
if (seenApprovals.has(a.id)) continue;
seenApprovals.add(a.id);
const verb = a.kind === "spawn" ? "spawn approval" : "config commit";
NOTIF.show(
"◆ approval #" + a.id,
`${verb} for ${a.agent}`,
`config commit for ${a.agent}`,
"hyperhive:approval:" + a.id,
);
}

View file

@ -22,7 +22,6 @@ use crate::lifecycle;
/// FinalizeDeploy`, plus an `AfterAny` `DeployTail`, under a
/// resource-holding root; ~30-90s)
/// - `UpdateMetaInputs` → a `MetaUpdate` DAG (fan-out on completion)
/// - `Spawn` → a `Spawn` DAG (`Create → WriteDropin → Reconcile`)
///
/// Every queued kind — deploys included — resolves its approval row via
/// [`resolve_approval_dag`] when the DAG settles terminal.
@ -55,25 +54,6 @@ pub async fn approve(coord: Arc<Coordinator>, id: i64) -> Result<()> {
coord.emit_rebuild_queue_snapshot();
Ok(())
}
ApprovalKind::Spawn => {
// The spawn's tail `Reconcile` starts the container, so the
// new agent's power intent is `Up` from the outset.
if let Err(e) = coord
.power
.set(approval.agent.as_str(), crate::power::Wanted::Up)
{
tracing::warn!(agent = %approval.agent, error = ?e, "agent_power: seed on spawn failed");
}
let inserted = coord.job_queue.insert_job(|b| {
crate::job_queue::templates::spawn(b, approval.agent.as_str(), id);
Vec::new()
});
if let Err(e) = inserted {
return Err(e.context("insert spawn dag"));
}
coord.emit_rebuild_queue_snapshot();
Ok(())
}
ApprovalKind::SchedulePrompt => {
// No queue card for SchedulePrompt — the work is a single
// sqlite insert, the actual "running" lifetime lives on
@ -525,19 +505,7 @@ pub(crate) async fn resolve_approval_dag(
TerminalState::Failed => Err(anyhow::anyhow!("{}", error.unwrap_or("job dag failed"))),
};
let mut terminal_tag = None;
match approval.kind {
ApprovalKind::Spawn => {
// Post-spawn forge bookkeeping (config repo mirror, meta
// access) — warn-only, then the resolution events + a rescan so
// the dashboard reflects the post-spawn state either way.
if result.is_ok() {
forge_after_first_spawn(coord, approval.agent.as_str()).await;
} else {
coord.rescan_containers_and_emit().await;
crate::dashboard::emit_tombstones_snapshot(coord).await;
}
}
ApprovalKind::MergeConfigPr => {
if approval.kind == ApprovalKind::MergeConfigPr {
terminal_tag = deploy_terminal_tag(approval.agent.as_str(), approval_id, outcome).await;
// On a failed deploy, surface the failing build log back onto the
// PR so the manager sees why it was rejected without leaving the
@ -548,8 +516,6 @@ pub(crate) async fn resolve_approval_dag(
post_merge_failure_to_pr(coord, &approval, e).await;
}
}
_ => {}
}
if let Err(e) = finish_approval(coord, &approval, result, terminal_tag).await {
tracing::warn!(approval_id, error = ?e, "approval dag resolved with failure");
}
@ -604,26 +570,6 @@ fn fetch_approval_for_worker(
Ok(approval)
}
/// Forge bookkeeping run once after the very first container spawn:
/// mirror the applied repo and grant read access to core/meta. The
/// agent's forge user and token are swarm-controller's, not this
/// hive's. Also rescans containers so the dashboard reflects the post-spawn state.
async fn forge_after_first_spawn(coord: &Arc<Coordinator>, agent: &str) {
if let Err(e) = crate::forge::ensure_config_repo(agent).await {
tracing::warn!(%agent, error = ?e, "forge: ensure_config_repo after first spawn failed");
}
if let Some(core_token) = crate::forge::core_token()
&& let Err(e) = crate::forge::meta_read_access(agent, &core_token).await
{
tracing::warn!(%agent, error = ?e, "forge: meta_read_access after first spawn failed");
}
if let Err(e) = crate::forge::ensure_meta_remote(agent).await {
tracing::warn!(%agent, error = ?e, "forge: ensure_meta_remote after first spawn failed");
}
coord.rescan_containers_and_emit().await;
crate::dashboard::emit_tombstones_snapshot(coord).await;
}
async fn finish_approval(
coord: &Coordinator,
approval: &hive_sh4re::approvals::Approval,
@ -670,35 +616,14 @@ async fn finish_approval(
note: note.clone(),
description: approval.description.clone(),
});
// For spawn/rebuild approvals, also surface the underlying action so the
// For rebuild approvals, also surface the underlying action so the
// manager knows whether the lifecycle step succeeded. The
// ApprovalResolved event already carries the same `ok` signal but
// separating it lets the manager react to the lifecycle change
// without having to special-case approvals.
match approval.kind {
ApprovalKind::Spawn => {
let summary = if ok {
format!("agent '{}' spawned", approval.agent)
} else {
format!(
"agent '{}' spawn FAILED: {}",
approval.agent,
note.as_deref().unwrap_or("unknown error")
)
};
let _ = coord
.push_todo_submitter(
approval.id,
"core",
Some(format!("spawned:{}", approval.agent)),
summary,
None,
)
.await;
}
// MergeConfigPr ends in a container rebuild — surface a Rebuilt
// lifecycle event. (It is never a first spawn — the agent already
// exists — so it never needs the Spawned arm above.)
// lifecycle event.
ApprovalKind::MergeConfigPr => {
let summary = crate::coordinator::rebuilt_todo_summary(
approval.agent.as_str(),
@ -738,8 +663,8 @@ async fn finish_approval(
///
/// Caller-specific bits stay OUT of here: fetching the PR head, the
/// `verify_commit` gate, and the ff-merge. The agent always already exists here
/// (a merge is never a first spawn), so there's no `sync_agents` step — the
/// operator `Spawn` flow owns first-time meta registration.
/// (a merge is never a first deploy), so there's no `sync_agents` step — the
/// first-deploy DAG's `Provision` node owns first-time meta registration.
async fn prepare_applied_target(
agent: &str,
applied_dir: &std::path::Path,

View file

@ -841,7 +841,7 @@ impl Coordinator {
/// Call after any mutation that could affect what
/// `nixos-container list` returns or what a row's
/// `running` / `needs_update` / `needs_login` / `deployed_sha`
/// resolves to — lifecycle ops, destroy, approve (post-spawn),
/// resolves to — lifecycle ops, destroy,
/// rebuild, meta-update, and the crash-watcher's periodic poll.
/// Cheap when nothing changed (one `nixos-container list` + a
/// `HashMap` diff + zero emits).

View file

@ -83,11 +83,6 @@ pub(super) fn gc_orphans(coord: &Coordinator, approvals: Vec<Approval>) -> Vec<A
approvals
.into_iter()
.filter(|a| {
// A Spawn approval is for a not-yet-existent agent; the proposed
// dir is supposed to be missing, so absence is not orphanhood.
if a.kind == hive_sh4re::approvals::ApprovalKind::Spawn {
return true;
}
if Coordinator::agent_proposed_dir(&a.agent).exists() {
true
} else {

View file

@ -232,56 +232,3 @@ pub(super) async fn post_op_send(
// no full-state refresh in between.
(axum::http::StatusCode::OK, "ok").into_response()
}
#[derive(Deserialize, ToSchema)]
pub(super) struct RequestSpawnForm {
name: String,
}
/// Queue a spawn approval for `name`.
#[utoipa::path(
post,
path = "/api/request-spawn",
request_body(content = RequestSpawnForm, content_type = "application/x-www-form-urlencoded"),
responses(
(status = 200, description = "spawn approval queued", body = String),
(status = 500, description = "missing name, or the approval submit failed"),
),
tag = "misc_api"
)]
pub(super) async fn post_request_spawn(
State(state): State<AppState>,
Form(form): Form<RequestSpawnForm>,
) -> Response {
let name = form.name.trim().to_owned();
if name.is_empty() {
return error_response("spawn: `name` required");
}
match state.coord.approvals.submit_kind(
&name,
hive_sh4re::approvals::ApprovalKind::Spawn,
"",
None,
"operator",
None,
) {
Ok(id) => {
tracing::info!(%id, %name, "operator: spawn approval queued via dashboard");
// Phase 5b: notify the dashboard event channel so live
// subscribers can append the row without a snapshot
// refetch. Spawn approvals carry no sha.
state
.coord
.emit_approval_added(crate::coordinator::ApprovalAdded {
id,
agent: &name,
approval_kind: "spawn",
sha_short: None,
description: None,
pr_number: None,
});
(StatusCode::OK, "ok").into_response()
}
Err(e) => error_response(&format!("request-spawn {name} failed: {e:#}")),
}
}

View file

@ -149,7 +149,6 @@ pub async fn serve(
.routes(routes!(misc_api::api_stats_hive))
.routes(routes!(misc_api::api_container_resources))
.routes(routes!(misc_api::post_mark_all_read))
.routes(routes!(misc_api::post_request_spawn))
.routes(routes!(misc_api::post_op_send))
.routes(routes!(build_logs::get_build_logs_all))
.routes(routes!(build_logs::get_build_log_for_node))
@ -386,7 +385,6 @@ mod router_build_probe {
.routes(routes!(misc_api::api_stats_hive))
.routes(routes!(misc_api::api_container_resources))
.routes(routes!(misc_api::post_mark_all_read))
.routes(routes!(misc_api::post_request_spawn))
.routes(routes!(misc_api::post_op_send))
.routes(routes!(build_logs::get_build_logs_all))
.routes(routes!(build_logs::get_build_log_for_node))

View file

@ -140,7 +140,7 @@ struct ApprovalHistoryView {
agent: String,
kind: &'static str,
/// First 12 chars of the canonical sha (preferred) or
/// manager-supplied ref. None for resolved spawn approvals.
/// manager-supplied ref. None when the approval carries neither.
sha_short: Option<String>,
/// `approved` / `denied` / `failed`.
status: &'static str,
@ -404,7 +404,7 @@ fn build_transient_views(
}
/// Render each pending approval into its dashboard view (short sha for
/// `MergeConfigPr`, just the name for `Spawn`).
/// `MergeConfigPr`).
/// Project a resolved sqlite row into the lean shape the dashboard
/// history tab consumes — no `diff_html` (rendering 30 of them
/// per /api/state poll would mean 30 git diffs per refresh).
@ -439,16 +439,6 @@ fn build_approval_views(approvals: Vec<Approval>) -> Vec<ApprovalView> {
let mut out = Vec::with_capacity(approvals.len());
for a in approvals {
out.push(match a.kind {
hive_sh4re::approvals::ApprovalKind::Spawn => ApprovalView {
id: a.id,
agent: a.agent.to_string(),
kind: "spawn",
sha_short: None,
description: a.description,
pr_number: None,
commit_ref: None,
requested_at: a.requested_at,
},
hive_sh4re::approvals::ApprovalKind::UpdateMetaInputs => ApprovalView {
id: a.id,
agent: a.agent.to_string(),

View file

@ -31,7 +31,7 @@ pub struct TombstoneView {
}
/// State-dir names that don't appear in the live container list. Each
/// one surfaces in the dashboard as a row with R3V1V3 + PURG3 actions.
/// one surfaces in the dashboard as a row with a PURG3 action.
///
/// ⚠️ **This lists every agent whose container is absent, not only destroyed
/// ones** — a mid-spawn agent (state dir seeded by `Provision`, container not

View file

@ -52,7 +52,7 @@ pub enum DashboardEvent {
/// enough to render the dashboard row without a `/api/state`
/// refetch.
///
/// The approval's own kind (`"merge_config_pr"` / `"spawn"`) lives
/// The approval's own kind (`"merge_config_pr"` / `"schedule_prompt"`) lives
/// on `approval_kind` rather than `kind` because the latter is taken
/// by the serde tag identifying which `DashboardEvent` variant
/// this is.
@ -119,8 +119,8 @@ pub enum DashboardEvent {
name: String,
transient_kind: String,
},
/// One container row changed — new container appeared (post-spawn
/// finalise), an existing one flipped `running` / `needs_update` /
/// One container row changed — a new container appeared, an existing
/// one flipped `running` / `needs_update` /
/// `sha`, etc. Clients upsert by `container.name`. Payload carries
/// the full row so cold-loaded clients and event-driven clients
/// converge on the same render.
@ -136,10 +136,8 @@ pub enum DashboardEvent {
/// `nixos-container destroy` (operator-driven or otherwise) on the
/// next rescan.
ContainerRemoved { seq: u64, name: String },
/// Full snapshot of the tombstones list. Emitted on every
/// mutation that could add / remove a tombstone: destroy
/// (with or without purge), purge-tombstone, spawn approval
/// (which can consume a tombstone of the same name). Snapshot
/// Full snapshot of the tombstones list. Emitted by destroy (with
/// or without purge) and purge-tombstone. Snapshot
/// shape (not diff) because the list is tiny (single-digit
/// typical) and recomputing avoids the add/remove races a
/// per-row event would have.

View file

@ -122,7 +122,7 @@ pub enum NodeKind {
/// node-inventory row for its three responsibilities and why it isn't
/// named `AbortDeploy`.
DeployTail { agent: String, approval_id: i64 },
/// Tail node of an approval-carrying DAG (spawn / opaque deploy /
/// Tail node of an approval-carrying DAG (opaque deploy /
/// config-PR merge): resolve the approval row from how the work ended.
ResolveApproval {
approval_id: i64,

View file

@ -440,27 +440,21 @@ pub fn approval_deploy(builder: &JobBuilder, agent: &str, approval_id: i64) {
resolve_approval_tails(builder, approval_id, window);
}
/// First-deploy spawn (approval-driven): `Provision` (proposed/applied
/// repos, state subvolume, meta registration) then `Create`
/// (`nixos-container create`), drop-in write, then `Reconcile` starts
/// the container (`wanted = Up` written at approve time). All-or-nothing:
/// `Provision` (lease-exempt, precedes the container) is the group root;
/// `Create` (child) owns the agent lease; `WriteDropin` + `Reconcile`
/// (children of `Create`) borrow it. A failure cancel-cascades the rest —
/// unlike rebuild there's no recovery-reconcile (nothing to converge if the
/// container was never created). Closed by a `ResolveApproval` tail root edged
/// `AfterAny` onto `Provision` — the DAG's only other group-root, so its roll-up
/// already carries the whole cascade.
pub fn spawn(builder: &JobBuilder, agent: &str, approval_id: i64) {
let provision = spawn_nodes(builder, agent);
resolve_approval_tails(builder, approval_id, provision);
}
/// The spawn subgraph with no tail, returning its group root.
/// First deploy of an agent this hive has never seen, asked for by the swarm:
/// `Provision` (proposed/applied repos, state subvolume, meta registration)
/// then `Create` (`nixos-container create`), drop-in write, then `Reconcile`
/// starts the container (`wanted = Up`, seeded by
/// `swarm_status::queue_first_deploy`). All-or-nothing: `Provision`
/// (lease-exempt, precedes the container) is the group root; `Create` (child)
/// owns the agent lease; `WriteDropin` + `Reconcile` (children of `Create`)
/// borrow it. A failure cancel-cascades the rest — unlike rebuild there's no
/// recovery-reconcile (nothing to converge if the container was never created).
///
/// Split out for the same reason [`rebuild_nodes`] is: two callers want the
/// same four nodes and disagree only about what closes them.
pub(crate) fn spawn_nodes<'a>(builder: &'a JobBuilder, agent: &str) -> Handle<'a> {
/// No approval tail: the operator authorised the creation at swarm level, and
/// the deploy request carries that authorisation.
///
/// Returns the group root so a caller can wait on the whole subtree.
pub fn first_deploy(builder: &JobBuilder, agent: &str) -> Vec<hive_jobq::NodeGuid> {
let a = || agent.to_owned();
let provision = builder
.node(NodeKind::Provision { agent: a() })
@ -479,20 +473,7 @@ pub(crate) fn spawn_nodes<'a>(builder: &'a JobBuilder, agent: &str) -> Handle<'a
.needs(Resource::Agent(a()))
.part_of(create)
.after_ok(dropin);
provision
}
/// First deploy of an agent this hive has never seen, asked for by the swarm.
///
/// [`spawn`] without the approval tail, and the absence is the point rather
/// than an omission: that flow exists because an operator used to approve the
/// spawn *at the hive*. When the swarm asks, the operator has already clicked
/// create at swarm level — the deploy request carries that authorisation, and a
/// second gate here would be asking the same person the same question twice.
///
/// Returns the group root so a caller can wait on the whole subtree.
pub fn first_deploy(builder: &JobBuilder, agent: &str) -> Vec<hive_jobq::NodeGuid> {
vec![spawn_nodes(builder, agent).guid()]
vec![provision.guid()]
}
/// Teardown: `Stop` → `DestroyContainer` → (`PurgeState`) → `DestroyBookkeeping`.

View file

@ -1624,10 +1624,10 @@ fn pause_shape_signal_drain() {
}
#[test]
fn spawn_shape_provision_create_dropin_reconcile() {
fn first_deploy_shape_provision_create_dropin_reconcile() {
let q = JobQueue::new(1);
insert(&q, |builder| {
templates::spawn(builder, "newbie", 7);
templates::first_deploy(builder, "newbie");
});
assert_eq!(
declared_shape(&q),
@ -1636,14 +1636,6 @@ fn spawn_shape_provision_create_dropin_reconcile() {
row("create", Some("provision"), &[]),
row("write_dropin", Some("create"), &[]),
row("reconcile", Some("create"), &[("write_dropin", "done")]),
// One tail per outcome, each edged to accept only that one — so
// *which* tail the graph lets run already is the answer, and
// nothing branches at runtime. The three differ **only** in their
// accepted outcome, which is why `declared_shape` spells the
// outcome set out instead of bucketing it.
row("resolve_approval", None, &[("provision", "done")]),
row("resolve_approval", None, &[("provision", "failed")]),
row("resolve_approval", None, &[("provision", "cancelled")]),
]
);
}

View file

@ -109,20 +109,6 @@ async fn write_response(
async fn dispatch(req: &HostRequest, coord: Arc<Coordinator>) -> HostResponse {
let result: anyhow::Result<HostResponse> = async {
Ok(match req {
HostRequest::Spawn { name } => handle_spawn(&coord, name.as_str()).await?,
HostRequest::RequestSpawn { name } => {
tracing::info!(%name, "request_spawn");
let id = coord.approvals.submit_kind(
name.as_str(),
hive_sh4re::approvals::ApprovalKind::Spawn,
"",
None,
"operator",
None,
)?;
tracing::info!(%id, %name, "spawn approval queued");
HostResponse::success()
}
HostRequest::Kill { name } => submit_single(&coord, name.as_str(), Verb::Kill).await,
HostRequest::Restart { name } => {
submit_single(&coord, name.as_str(), Verb::Restart).await
@ -278,49 +264,6 @@ async fn dispatch(req: &HostRequest, coord: Arc<Coordinator>) -> HostResponse {
}
}
/// Create + start the container for `name` and bind its MCP listener. On a
/// failed spawn nothing was registered: post a swarm notice and return the error.
async fn handle_spawn(coord: &Arc<Coordinator>, name: &str) -> Result<HostResponse> {
tracing::info!(%name, "spawn");
let agent_dir = crate::paths::agent_runtime_dir(name);
let hive = coord.hive_env();
let paths = Coordinator::agent_paths(name, agent_dir)?;
// lifecycle::spawn creates the runtime dir internally before start.
// MCP listener registration is event-driven: bind immediately on
// success so the harness can connect on its first turn without
// waiting for any poll interval.
match lifecycle::spawn(name, &hive, &paths).await {
Ok(()) => {
if let Err(e) = coord.power.set(name, crate::power::Wanted::Up) {
tracing::warn!(%name, error = ?e, "agent_power: set wanted=up failed");
}
// Bind the MCP listener now that the container is starting up.
// The harness connects to this socket on its first turn.
coord.register_agent(name)?;
crate::swarm_notices::notify(
"core",
Some(format!("spawned:{name}")),
format!("agent '{name}' spawned"),
None,
)
.await;
}
Err(e) => {
// Spawn failed: register_agent was never called, so there is
// nothing to unregister. Notify the swarm and propagate.
crate::swarm_notices::notify(
"core",
Some(format!("spawned:{name}")),
format!("agent '{name}' spawn FAILED: {e:#}"),
None,
)
.await;
return Err(e);
}
}
Ok(HostResponse::success())
}
/// `hivectl pause|resume` / the dashboard toggle: write or remove the
/// agent's pause marker.
///

View file

@ -31,7 +31,7 @@ pub(crate) async fn submit_merge_config_pr(
anyhow::bail!(
"applied repo missing for agent '{agent}' (expected at {}) — \
merge_config_pr requires the agent to be fully provisioned; \
spawn the agent first (operator spawn) before opening config PRs",
create it first (`swarmctl agent create`) before opening config PRs",
applied_dir.display()
);
}

View file

@ -1,7 +1,7 @@
//! Approval queue. Requests are submitted by an agent
//! (`RequestSchedulePrompt`), the config-PR webhook (`MergeConfigPr`), or
//! the operator (`Spawn`); the user approves/denies via the host admin CLI;
//! on approval the host runs the corresponding action.
//! (`RequestSchedulePrompt`) or the config-PR webhook (`MergeConfigPr`); the
//! user approves/denies via the host admin CLI; on approval the host runs the
//! corresponding action.
//!
//! `UpdateMetaInputs` rows are legacy: the MCP tool that queued them was
//! removed and nothing produces the kind any more. The variant and
@ -78,7 +78,7 @@ impl Approvals {
/// Insert a new pending approval row. `fetched_sha` may be supplied
/// when the sha is already known at submission time (e.g. `MergeConfigPr`
/// fetches the PR head before inserting), making the insert + sha-set
/// atomic. Pass `None` when the kind carries no sha (e.g. `Spawn`).
/// atomic. Pass `None` when the kind carries no sha (e.g. `SchedulePrompt`).
pub fn submit_kind(
&self,
agent: &str,
@ -382,7 +382,6 @@ fn row_to_approval(row: &rusqlite::Row<'_>) -> rusqlite::Result<Approval> {
// Column order: id, agent, kind, commit_ref, requested_at, status, resolved_at, note, fetched_sha, description.
let kind: String = row.get(2)?;
let kind = match kind.as_str() {
"spawn" => ApprovalKind::Spawn,
"update_meta_inputs" => ApprovalKind::UpdateMetaInputs,
"schedule_prompt" => ApprovalKind::SchedulePrompt,
"merge_config_pr" => ApprovalKind::MergeConfigPr,
@ -435,7 +434,6 @@ fn row_to_approval(row: &rusqlite::Row<'_>) -> rusqlite::Result<Approval> {
fn kind_from_str(s: &str) -> Result<ApprovalKind> {
Ok(match s {
"spawn" => ApprovalKind::Spawn,
"update_meta_inputs" => ApprovalKind::UpdateMetaInputs,
"schedule_prompt" => ApprovalKind::SchedulePrompt,
"merge_config_pr" => ApprovalKind::MergeConfigPr,
@ -467,7 +465,7 @@ mod tests {
None,
)
.unwrap();
db.submit_kind("b", ApprovalKind::Spawn, "", None, "b", None)
db.submit_kind("b", ApprovalKind::SchedulePrompt, "", None, "b", None)
.unwrap();
db.submit_kind("c", ApprovalKind::UpdateMetaInputs, "[]", None, "c", None)
.unwrap();
@ -508,7 +506,14 @@ mod tests {
// final — re-cancelling errors instead of silently overwriting.
let (_dir, _path, db) = open_temp();
let id = db
.submit_kind("a", ApprovalKind::Spawn, "deadbeef", None, "a", None)
.submit_kind(
"a",
ApprovalKind::MergeConfigPr,
"deadbeef",
None,
"a",
None,
)
.unwrap();
db.mark_cancelled(id, "manager").expect("first cancel");
let err = db
@ -567,7 +572,7 @@ mod tests {
let raw = Connection::open(&path).unwrap();
raw.execute(
"INSERT INTO approvals (agent, kind, commit_ref, requested_at, status)
VALUES ('old', 'spawn', '', 0, 'pending')",
VALUES ('old', 'merge_config_pr', '', 0, 'pending')",
[],
)
.unwrap();

View file

@ -370,10 +370,9 @@ async fn handle_deploy_request(
/// (`power::Store::get_or_seed`), and at that point the container is freshly
/// created but not started — so an unseeded row locks the agent to `Offline`
/// on its very first reconcile and `Reconcile` never emits the `Start` node.
/// Setting the row up front closes that window, the same way
/// `actions::approve`'s `ApprovalKind::Spawn` arm does for the
/// operator-approved path. A second caller open-coding the insert would lose
/// exactly that, and the agent would come up stopped for no visible reason.
/// Setting the row up front closes that window. A second caller open-coding
/// the insert would lose exactly that, and the agent would come up stopped for
/// no visible reason.
pub(crate) fn queue_first_deploy(
coord: &std::sync::Arc<crate::coordinator::Coordinator>,
agent: &str,

View file

@ -6,8 +6,8 @@
//! re-register all running agents.
//!
//! After startup, listeners are managed event-driven:
//! - `run_start` (the job queue's `Start` node) and `server::handle_spawn`
//! call `register_agent` once the container is started.
//! - `run_start` (the job queue's `Start` node) calls `register_agent`
//! once the container is started.
//! - `kill`/`destroy` paths call `unregister_agent`.
//!
//! No recurring poll is needed because c0re owns the listener lifecycle. An

View file

@ -117,20 +117,6 @@ pub enum ReconcileDirection {
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(tag = "cmd", rename_all = "snake_case")]
pub enum HostRequest {
/// Create and start a brand-new sub-agent container directly (full
/// first-time provisioning: proposed/applied repos, state subvolume,
/// meta-flake sync, `nixos-container create`), bypassing the approval
/// queue. Privileged-context only. Exposed on the CLI as `hivectl
/// agent <name> create`. See `docs/agent-lifecycle/approvals.md::Approval kinds
/// (wire shapes)`. Wire name kept as `Spawn` (unrenamed underneath
/// the CLI-verb rename — `hivectl agent <name> start` reuses the
/// existing scope-based [`HostRequest::Start`] below instead of a
/// new per-agent variant, see its doc comment).
Spawn { name: Ident },
/// Submit a first-creation request for the operator to approve. See
/// `docs/agent-lifecycle/approvals.md::Approval kinds (wire shapes)` (`Spawn`).
/// Exposed on the CLI as `hivectl agent <name> request-create`.
RequestSpawn { name: Ident },
/// Hard stop a managed container. Exposed on the CLI as `hivectl
/// agent <name> kill` — kept distinct from the new graceful-only
/// `hivectl agent <name> stop` (which reuses the scope-based

View file

@ -46,9 +46,6 @@ pub struct Approval {
#[serde(rename_all = "snake_case")]
#[strum(serialize_all = "snake_case")]
pub enum ApprovalKind {
/// Create + start a new sub-agent container with the given name
/// (under the default `agent.nix` template).
Spawn,
/// Run `nix flake update [inputs...]` on the meta flake and commit
/// the resulting lock changes.
UpdateMetaInputs,

View file

@ -54,7 +54,7 @@ async fn agents_restart(socket: &Path, name: &str, no_wait: bool) -> Result<()>
async fn agents_start(socket: &Path, name: &str, paused: bool) -> Result<()> {
if !crate::util::agent_exists(socket, name).await? {
bail!(
"no such agent: '{name}' (no state dir under {}/) — use 'hivectl agent {name} create' to provision a brand-new agent",
"no such agent: '{name}' (no state dir under {}/) — a brand-new agent is created at swarm level ('swarmctl agent create')",
hive_host_sock::AGENTS_ROOT
);
}
@ -269,14 +269,6 @@ pub(crate) async fn run_agent(socket: &Path, name: &str, cmd: AgentCmd) -> Resul
AgentCmd::Pause { no_wait } => set_paused(socket, name, true, no_wait).await,
AgentCmd::Resume => set_paused(socket, name, false, false).await,
AgentCmd::Start { paused } => agents_start(socket, name, paused).await,
AgentCmd::Create => {
let name = crate::util::parse_ident(name)?;
render(crate::client::request(socket, HostRequest::Spawn { name }).await?)
}
AgentCmd::RequestCreate => {
let name = crate::util::parse_ident(name)?;
render(crate::client::request(socket, HostRequest::RequestSpawn { name }).await?)
}
AgentCmd::Stop => agents_stop(socket, name).await,
AgentCmd::Kill => {
let name = crate::util::parse_ident(name)?;

View file

@ -414,7 +414,7 @@ pub enum AgentCmd {
Resume,
/// Start this EXISTING agent container. Fails immediately if `name`
/// has no config/topology entry at all — it never attempts
/// first-time creation. Use `create` for that.
/// first-time creation, which is swarm-level (`swarmctl agent create`).
Start {
/// Start (or leave) the agent paused: if it's currently down, hivectl
/// writes the pause marker before the container boots, so it comes
@ -424,14 +424,6 @@ pub enum AgentCmd {
#[arg(long)]
paused: bool,
},
/// Create this agent container from scratch (full first-time
/// provisioning), bypassing the approval queue.
///
/// Operator-on-the-host only; use `request-create` for an
/// approval-gated creation.
Create,
/// Queue a first-creation request for operator approval.
RequestCreate,
/// Gracefully stop this agent container: signal → drain → reconcile.
/// Never escalates to a hard kill — use `kill` for that.
Stop,

View file

@ -686,6 +686,10 @@ struct AppState {
/// "no forge configured on this host" shape every other
/// forge-backed field here uses.
forge: Option<Arc<forge::Client>>,
/// Held by `create_agent` from reading where a name is placed until its
/// graph is queued, so two creations of one name cannot both find it
/// unplaced. `tokio`'s mutex, unlike `jobq`'s: the read awaits the queue.
create_gate: Arc<tokio::sync::Mutex<()>>,
}
/// Env var the controller's NixOS module sets from
@ -1352,7 +1356,9 @@ struct CreateAgentResponse {
responses(
(status = 200, description = "job chain queued", body = CreateAgentResponse),
(status = 400, description = "`name` or `hive` is not a valid identifier, `name` is new and reserved or refused by the forge, or `hive` is not in this swarm (problem+json)", body = String),
(status = 500, description = "the job chain could not be queued (problem+json)", body = String),
(status = 409, description = "the swarm has already placed `name` on a different hive (problem+json)", body = String),
(status = 503, description = "the swarm queue is not connected, so whether `name` is placed on another hive is unknown (problem+json)", body = String),
(status = 500, description = "the job chain could not be queued, or another hive's wanted state could not be read (problem+json)", body = String),
),
tag = "agents"
)]
@ -1424,10 +1430,25 @@ async fn create_agent(
return Err(error_problem(axum::http::StatusCode::BAD_REQUEST, &detail));
}
// Agent names are one swarm-wide namespace: identity, forge user, store
// path and matrix id all carry the bare name. The same name on the same
// hive is that agent being re-created, and goes through.
let _gate = state.create_gate.lock().await;
let declared = declarations_elsewhere(&state, &agent, &hive).await?;
let mut sched = state
.jobq
.lock()
.unwrap_or_else(std::sync::PoisonError::into_inner);
let queued = queued_placements(sched.graph());
let elsewhere = placed_elsewhere(&agent, &hive, &declared, &queued);
if !elsewhere.is_empty() {
let detail = format!(
"agent name {agent:?} is already placed on hive {} — agent names are unique across \
the swarm; destroy it there first to move it, or choose another name",
elsewhere.join(", ")
);
return Err(error_problem(axum::http::StatusCode::CONFLICT, &detail));
}
let ids = sched
.insert_job(None, |b| declare_agent_job(b, &agent, &hive))
.map_err(|e| {
@ -1445,6 +1466,81 @@ async fn create_agent(
}))
}
/// Every hive but `hive`'s published wanted state, for [`placed_elsewhere`].
///
/// No queue wired up means no hive has been sent a declaration, so there is
/// nothing to collide with. A queue that cannot be read refuses: an unread
/// declaration and an absent one look the same, and reading "absent" is how
/// a duplicate would get through.
async fn declarations_elsewhere(
state: &AppState,
agent: &str,
hive: &str,
) -> Result<Vec<(String, swarm_queue_client::wanted::HiveWanted)>, problem_details::ProblemDetails>
{
let Some(writer) = state.wanted.as_deref() else {
return Ok(Vec::new());
};
let mut declared = Vec::new();
for other in state.hives.iter().filter(|h| h.name != hive) {
match writer.view(&other.name).await {
Ok(Some(declaration)) => declared.push((other.name.clone(), declaration)),
Ok(None) => {}
Err(e) => {
let detail = format!(
"cannot tell whether {agent:?} already exists on hive {:?}, so it was not \
created: {e:#}",
other.name
);
return Err(error_problem(wanted_error_status(&e), &detail));
}
}
}
Ok(declared)
}
/// `(hive, agent)` of every `SetAgentWanted` node not yet settled: a creation
/// accepted but not yet visible in the wanted state it is about to write.
fn queued_placements(
graph: &hive_jobq::Graph<SwarmNodeKind, SwarmResourceKind>,
) -> Vec<(String, String)> {
graph
.nodes()
.filter(|node| !node.state.is_terminal())
.filter_map(|node| match &node.payload {
SwarmNodeKind::SetAgentWanted { hive, agent } => Some((hive.clone(), agent.clone())),
_ => None,
})
.collect()
}
/// The hives other than `hive` the swarm has placed `agent` on, sorted.
///
/// A placement is a declaration in any state but `Destroyed`, or a queued
/// one. A `Destroyed` agent is gone from that hive, so creating it on another
/// moves it.
fn placed_elsewhere(
agent: &str,
hive: &str,
declared: &[(String, swarm_queue_client::wanted::HiveWanted)],
queued: &[(String, String)],
) -> Vec<String> {
use swarm_queue_client::wanted::AgentState;
let declared = declared.iter().filter(|(_, declaration)| {
declaration
.agents
.get(agent)
.is_some_and(|wanted| wanted.state != AgentState::Destroyed)
});
let queued = queued.iter().filter(|(_, queued)| queued == agent);
let hives: std::collections::BTreeSet<&str> = declared
.map(|(h, _)| h.as_str())
.chain(queued.map(|(h, _)| h.as_str()))
.filter(|h| *h != hive)
.collect();
hives.into_iter().map(str::to_owned).collect()
}
/// Every naming rule `agent` breaks, each as a sentence naming the rule.
///
/// `reserved` is [`hive_types::reserved_names_raw`]: the list nix owns in
@ -2554,6 +2650,7 @@ async fn main() -> Result<()> {
swarm_name: load_swarm_name().map(Arc::from),
auth,
forge: state_forge,
create_gate: Arc::default(),
};
let app = build_app(state);
@ -2710,6 +2807,7 @@ mod tests {
// read is the verb that needs it, and it has its own test below.
auth: None,
forge: None,
create_gate: std::sync::Arc::default(),
};
(state, sched)
}
@ -2846,6 +2944,170 @@ mod tests {
assert!(queued > 0, "an accepted creation must queue work");
}
/// `state_with_roster` with a second hive, `sec0nd`.
fn state_with_two_hives() -> (super::AppState, SharedSched) {
let (state, sched) = state_with_roster();
let mut hives = (*state.hives).clone();
hives.push(HiveEntry {
name: "sec0nd".to_owned(),
domain: "sec0nd.example".to_owned(),
});
let state = super::AppState {
hives: std::sync::Arc::new(hives),
..state
};
(state, sched)
}
async fn create(
state: &super::AppState,
name: &str,
hive: &str,
) -> Result<(), problem_details::ProblemDetails> {
super::create_agent(
axum::extract::State(state.clone()),
axum::Json(super::CreateAgentRequest {
name: name.to_owned(),
hive: hive.to_owned(),
}),
)
.await
.map(|_| ())
}
fn node_count(sched: &SharedSched) -> usize {
sched
.lock()
.unwrap_or_else(std::sync::PoisonError::into_inner)
.graph()
.nodes()
.count()
}
/// A name queued for one hive is refused on another, by effect: 409 and
/// nothing more queued. The queued-but-unwritten placement is the one a
/// wanted-state read alone would miss.
#[tokio::test]
async fn a_name_placed_on_another_hive_is_refused_before_anything_is_queued() {
let (state, sched) = state_with_two_hives();
create(&state, "atlas", "pr1ma")
.await
.expect("the first creation of a name must be accepted");
let before = node_count(&sched);
let err = create(&state, "atlas", "sec0nd")
.await
.expect_err("the same name on another hive must be refused");
assert_eq!(
err.status,
Some(axum::http::StatusCode::CONFLICT),
"{err:?}"
);
assert!(
format!("{err:?}").contains("pr1ma"),
"the refusal should name the hive holding the name, got: {err:?}"
);
assert_eq!(
node_count(&sched),
before,
"a refused creation must queue no work"
);
}
/// The control for the refusal above: the same name on the SAME hive is
/// that agent being re-created, and queues again.
#[tokio::test]
async fn re_creating_an_agent_on_its_own_hive_is_accepted() {
let (state, sched) = state_with_two_hives();
create(&state, "atlas", "pr1ma")
.await
.expect("the first creation of a name must be accepted");
let before = node_count(&sched);
create(&state, "atlas", "pr1ma")
.await
.expect("re-creating an agent on its own hive must be accepted");
assert!(node_count(&sched) > before, "a re-creation must queue work");
}
/// A wanted state that cannot be read refuses rather than reading as
/// "placed nowhere", and queues nothing.
#[tokio::test]
async fn an_unreadable_wanted_state_refuses_creation() {
let (state, sched) = state_with_two_hives();
let state = super::AppState {
wanted: Some(std::sync::Arc::new(wanted::WantedWriter::new(
disconnected_client().await,
))),
..state
};
let err = create(&state, "atlas", "pr1ma")
.await
.expect_err("an unreadable placement must refuse");
assert_eq!(
err.status,
Some(axum::http::StatusCode::SERVICE_UNAVAILABLE),
"{err:?}"
);
assert_eq!(
node_count(&sched),
0,
"a refused creation must queue no work"
);
}
fn declaring(
hive: &str,
agent: &str,
state: swarm_queue_client::wanted::AgentState,
) -> (String, swarm_queue_client::wanted::HiveWanted) {
let mut declaration = swarm_queue_client::wanted::HiveWanted::default();
declaration.agents.insert(
agent.to_owned(),
swarm_queue_client::wanted::AgentWanted { state },
);
(hive.to_owned(), declaration)
}
#[test]
fn a_live_declaration_on_another_hive_is_a_placement() {
use swarm_queue_client::wanted::AgentState;
for state in [AgentState::Up, AgentState::Offline, AgentState::Paused] {
let declared = [declaring("sec0nd", "atlas", state)];
assert_eq!(
super::placed_elsewhere("atlas", "pr1ma", &declared, &[]),
["sec0nd"],
"{state:?}"
);
}
}
/// A destroyed agent has left its hive, so creating it elsewhere moves it.
#[test]
fn a_destroyed_declaration_is_not_a_placement() {
let declared = [declaring(
"sec0nd",
"atlas",
swarm_queue_client::wanted::AgentState::Destroyed,
)];
assert!(super::placed_elsewhere("atlas", "pr1ma", &declared, &[]).is_empty());
}
#[test]
fn placements_on_the_requested_hive_or_of_other_names_do_not_count() {
use swarm_queue_client::wanted::AgentState;
let declared = [
declaring("pr1ma", "atlas", AgentState::Up),
declaring("sec0nd", "argus", AgentState::Up),
];
let queued = [
("pr1ma".to_owned(), "atlas".to_owned()),
("sec0nd".to_owned(), "argus".to_owned()),
];
assert!(super::placed_elsewhere("atlas", "pr1ma", &declared, &queued).is_empty());
}
/// Forgejo v16.0.5's reserved usernames our charset can spell, plus the
/// bare `-` — the forge section of `nix/reserved-names.nix`. Read from
/// the list nix exports, so dropping one from the file reds this.

View file

@ -580,6 +580,7 @@ mod tests {
swarm_name: None,
auth: None,
forge: None,
create_gate: std::sync::Arc::default(),
}
}

View file

@ -83,10 +83,12 @@ pub const AGENT_PREFIX: &str = "hive-agent-";
/// attaches the policy and matches the subject by spelling all three the same.
///
/// Injective in `agent` — the prefix is fixed and [`checked_segment`] has
/// already refused anything that could re-punctuate the suffix — and agent
/// names are one swarm-wide namespace (an agent's path is
/// `swarm/agents/<agent>`, with no hive segment to disambiguate two that
/// matched), so two agents cannot land on one name.
/// already refused anything that could re-punctuate the suffix. Agent names
/// are one swarm-wide namespace (an agent's path is `swarm/agents/<agent>`,
/// with no hive segment), so two agents sharing a name would share every
/// object this names. swarm-controller's `create_agent`, the one place an
/// agent is created, refuses a name the swarm has already placed on another
/// hive. An agent no hive's wanted state declares is invisible to that check.
///
/// # Errors
/// [`Error::PathSegment`] when `agent` holds anything but `[A-Za-z0-9_-]`.

View file

@ -206,7 +206,9 @@ struct AgentCreateArgs {
/// Name for the new agent: 1–63 characters of `[a-z0-9-]`.
///
/// Becomes an SSO subject, a forge user and a repository name, so
/// it's validated here before queuing.
/// it's validated here before queuing. The controller refuses a name
/// the swarm has already placed on a different hive; the same name on
/// the same hive re-creates that agent.
name: String,
/// Hive in this swarm to deploy the agent to.
///