Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/agent-lifecycle/approvals.md
atlas ed9c0f53ed docs: fix argus R4 review nits on config-PR merge docs
- hive-c0re/README.md: drop deleted webhook_secret.rs from the module
  list
- docs/swarm/README.md, docs/agent-lifecycle/approvals.md: rewrite
  temporal wording (legacy/older-release phrasing) as current
  behaviour
- docs/swarm/README.md: note that a lost closed delivery deploys
  nothing and how the operator recovers
2026-10-02 23:36:57 +02:00

29 KiB
Raw Blame History

Approvals + helper events

The approval queue is where an agent asks the operator for something it can't do itself: a scheduled prompt, or a meta-input bump. Config changes don't go through this queue: they're pull requests on the agent's config repo, which an operator merges on the forge, and that merge deploys them. The submitting agent — any agent with the approvals tool group, which manages the config of its direct children (the root agent for top-level agents; a sub-manager for its own subtree) — is the policy gate in front of both; helper events are how it stays informed about what happens after a decision lands.

For operators

  • Config change — an agent proposed a change to another agent's config (or its own, via a sub-manager) as a pull request on agent-configs/<agent>. Review it on the forge like any other PR, and merge it there to deploy it. See Config changes.
  • Scheduled prompt (SchedulePrompt) — an agent asked to schedule a message to one or more inboxes at a future time. It lands on your dashboard's Y3R C4LL tab (or hivectl approvals pending / approve <id> from the CLI). You can also add schedules yourself directly from the SCH3DUL3S tab, which skips this approval step entirely — the gate here is specifically for an agent asking to schedule something, not for you doing it.
  • Meta/flake update (UpdateMetaInputs) — an existing row reads back and you can approve it. Approving runs the update and commits the lock change; it doesn't rebuild anything by itself.

Don't want to approve something? Deny it (DENY on the dashboard card, or hivectl approvals deny <id>) — nothing runs. Either way the submitting agent is always notified that the operator denied their request; what's optional is only the reason text, which you can add on the dashboard's prompt (cancelling that prompt aborts the whole deny, not just the reason) but not from the CLI. Denying is final: a denied approval can't be re-approved later, the agent has to submit a fresh request.

Everything below this point is the implementation detail behind that flow.

Config changes

Config changes flow through a forge pull request on the agent's agent-configs/<name> repo — the same surface agents use for code PRs. No MCP tool exists for config changes: opening the PR IS the request.

  1. The submitting agent (the child's parent, holding the approvals tool group) clones agent-configs/<name>, edits it there (any tracked path, but agent.nix is the contract entry point), commits with its own git identity, and pushes a branch + opens a PR with hive-forge — the same way it would change any other repo. The bind-mounted /agents/<name>/config/ is a copy for reading a config, not the tree to edit: authoring in place there produces no PR. (it's currently mounted read-write, which is a defect tracked separately, not an authoring path.) Branch protection (the agent isn't on the main push allowlist; merge and approvals allowlist = operators team; see "Forge mirror" below) makes the agent a write collaborator that can't merge its own config PR.
  2. An operator reviews the PR on the forge (native diff, threaded comments, CI status) and merges it there. Merging needs membership of the operators team in agent-configs. The team starts empty: add each operator by hand in the forge UI.
  3. swarm-controller gets the agent-configs org's pull_request delivery. A closed event with merged: true and base branch main names the commit in merge_commit_sha.
  4. The controller looks up which hive's wanted state places the agent and publishes a deploy request carrying that commit on that hive's deploy subject. With no such hive, or more than one, it deploys nothing and logs a warning naming them.
  5. If the hive's applied/main already is that commit, it does nothing. Otherwise it fetches the config repo's main, requires the commit to descend from applied/main, fast-forwards applied/main to it, and queues a rebuild. A failure before the rebuild (the fetch, or a commit that doesn't descend) deploys nothing and posts the error as a comment on the merged PR. An agent with no container on that hive yet ignores the commit; its first deploy builds what it seeds.

This path runs no eval-verify. A merged config that doesn't evaluate or build fails the rebuild, and applied/main stays at that commit, so the agent's rebuilds keep failing until a fix merges. A failed rebuild shows on the hive like any other and isn't commented on the PR. A merge whose webhook delivery never reaches the controller deploys nothing.

Approval queue

Withdrawing a pending approval

The submitting agent can call cancel_loose_end(kind: "approval", id) to withdraw an approval the operator hasn't acted on yet. The row transitions to ApprovalStatus::Cancelled (distinct from Denied/Failed), the dashboard pulls the card out of the pending pane, and ApprovalResolved { status: "cancelled" } fires on the root agent + dashboard channels. Approvals that the operator (or a lifecycle failure) has already approved, denied, or failed return an error — the resolution is final once acted on.

The socket refuses the approval kind with a clear error for any agent that lacks the approvals tool group: only an agent with that group submits approvals (for its direct children), so an agent without it has nothing of its own to withdraw.

Creating a brand-new agent isn't an approval: it's swarm-level (swarmctl agent create, POST /api/agents on the swarm controller), and no hive can originate an agent. The swarm controller's InitAgentConfigRepo job seeds the agent's config repo with a default agent.nix template, then asks the target hive to deploy it.

Changing what the template seeded isn't a special case: like every later change, it's a PR on that config repo, made from a clone and merged by an operator on the forge (see Config changes).

Approval kinds (wire shapes)

ApprovalKind carries two variants; each maps to a different commit_ref encoding because ApprovalKind overloads that field as the kind-specific payload carrier.

  • UpdateMetaInputs — commit_ref stores the JSON-encoded inputs array ("[]" = all inputs, "[\"nixpkgs\"]" = just nixpkgs, etc.). hive-c0re sets the agent field to the requesting root agent. On approve hive-c0re runs nix flake update [inputs...] on the meta flake and commits the resulting lock changes.
  • SchedulePrompt — commit_ref stores the JSON-encoded SchedulePromptPayload (target list, body, schedule) so the approval row carries the full submission verbatim. On approve hive-c0re inserts a row into scheduled_prompts with source = approval:<id>; the worker fans the body out as inbox messages to each target at the scheduled time, recurring when interval_seconds is set.

Scheduled prompts (submit paths)

Two ways a row lands in scheduled_prompts:

  • Operator-direct (source = "operator"): the operator adds a schedule through the dashboard form. Lands in the table immediately, no approval gate — operator action is already the trust boundary.
  • Agent-requested (source = "approval:<id>"): an agent submits a RequestSchedulePrompt through its MCP socket (the request_schedule_prompt tool, scheduling group). hive-c0re queues an ApprovalKind::SchedulePrompt row; on approve, it inserts the schedule row with source = approval:<id> so the audit trail points back at the operator decision (above).

No self-target shortcut: even agent-self schedules need approval. The existing remind MCP tool stays the quick self-wake path (no approval, lands directly in the agent's own inbox); this module is the bigger, multi-recipient, operator-visible thing.

Scheduled prompt worker (catch-up clamp)

When hive-c0re comes back from being down, the worker sees rows whose next_fire_at_unix is well in the past. For recurring rows that would mean firing N delayed pulses in a row — spammy and useless. Instead the worker fires once per row and bumps next_fire_at_unix to the next interval slot ≥ now, recording how many cycles it skipped in last_result (per-target). Operators see "fired late, caught up from 17 skipped" instead of 17 wake-up storms.

The worker fires one-shot rows once (if past due, on the next worker pass) and deletes them; recurring rows survive until cancelled.

targets is its own table (scheduled_prompt_targets) so partial cancellation flips a single row and the dashboard can show last-fired / last-result per recipient. Cancelling every target reaps the parent row on the next worker pass.

Scheduled prompt delivery: todo, not a broker message

An agent target's delivery is push_todo (Coordinator::push_todo, docs/scheduler/coordinator.md covers the mechanism generally), not a broker Message — a scheduled prompt wakes its target with a todo instead of driving an immediate turn, by design. key = "schedule:<id>" per target drives push_todo's own upsert-by-key dedup: a re-fire of the same schedule against a target that hasn't reviewed the last one collapses into that one todo instead of stacking up.

Missing-target failure

When a target name doesn't resolve to a known agent (container destroyed, typo, etc.) the worker:

  1. Records last_result = "no such agent: <name>" on the per-target row.
  2. Sends a single advisory Message from system to operator naming the schedule, target, and reason. This one stays a Message regardless of target type — it's a to-operator advisory about a broken schedule, not the schedule's own delivery.
  3. Continues fanning out to the other live targets.

Transient broker errors (sqlite lock contention, etc.) get the same last_result annotation plus a tracing::warn, and then:

  • Recurring rows re-arm to the next interval slot — the retry self-heals on the next worker pass.
  • One-shot rows: the worker deletes them unconditionally after their single fan-out pass; a broker error on a one-shot isn't retried (the operator advisory and last_result are the only audit trail).

Reminder delivery: file-path semantics

A reminder may carry a file_path (the agent-visible path inside its container, for example /agents/<name>/state/foo.md). On delivery hive-c0re:

  1. Translates the container path to the host path (/var/lib/hyperhive/agents/<name>/state/foo.md) so c0re can write from outside the container.
  2. Validates the path: rejects anything outside the agent's own state subtree, containing .. (path traversal), or with an empty relative tail. On rejection hive-c0re skips the write and delivers the original message inline with a warning — the reminder still fires.
  3. Defends against symlink escape: after create_dir_all, hive-c0re canonicalizes the parent dir and re-verifies it lives under the agent's host state root. hive-c0re opens the final file with O_NOFOLLOW | O_CREAT | O_TRUNC so an existing symlink at the basename can't redirect the write to an arbitrary host path.
  4. Writes the body to disk and delivers a short pointer message in its place, keeping the agent's inbox / wake-prompt small while the agent reads the bulky payload out of band.

Broker::deliver_reminders_batch handles atomicity of the inbox INSERT + reminders.sent_at UPDATE; the scheduler only computes the body strings before calling it.

Destroy semantics

HostRequest::Destroy { name, purge } is the lifecycle tear-down, not an approval. Stops + removes the nspawn container, drops the systemd drop-in, fails any pending approvals. Persistent state (proposed/applied repos, claude credentials, /state/ notes) is kept by default — recreating the agent with the same name reuses prior config + login. With purge = true the agent's /var/lib/hyperhive/{agents,applied}/<name>/ trees are also wiped (config history + creds + notes gone forever). The root/bootstrap container is destroyable like any other — hive-c0re recreates it on the next startup if it's absent, so destroying it's transient.

Meta flake

The hive-c0re-owned repo at /var/lib/hyperhive/meta/ declares one flake input per agent (agent-<n>.url = "git+http://<forge>/agent-configs/<n>.git") and one nixosConfigurations.<n> output per agent. Each output wraps inputs.agent-<n>.nixosModules.default with the identity + HIVE_PORT / HIVE_LABEL / HIVE_DASHBOARD_PORT injection module. Containers run against --flake /var/lib/hyperhive/meta#<n>.

The declared input url is the agent's forge config repo (the same agent-configs/<n> config PRs merge into), so the meta flake references a reviewable, reproducible source rather than a local checkout. hive-c0re authenticates that git+http fetch via a git credential helper that reads the live forge-core token — no token in the url or the lock. The rebuild paths, however, do not re-lock from the forge: they --override-input agent-<n> git+file:///var/lib/hyperhive/applied/<n>, locking the exact config applied/<n>/main was fast-forwarded to. That keeps a rebuild reproducible and independent of forge reachability — rebuilds fire on crash-restart and meta bumps, not just config merges — while the declared url stays the forge. sync_agents re-renders + re-locks the persistent input; a plain nix flake lock leaves an existing applied override in place (it only re-locks when the declared url itself changes), so the forge-declared / applied-deployed split is stable.

Lock updates are single-phase and commit when the lock changed: meta::lock_update_for_rebuild(name) relocks one agent's input for a relocking rebuild (the manual ↻ R3BU1LD button, and the rebuild a merged config commit queues), and meta::lock_update_hyperhive() is the autoupdate flake-rev bump (one shot before per-agent rebuilds).

meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec]) — hive carries hyperhive_flake, dashboard_port, and the rest of the per-hive config — is the idempotent reconciler called by spawn, destroy, rebuild, and the startup migration. Renders flake.nix from the agent list; if it differs from disk, runs nix flake lock + commits as regenerate meta flake (or seed meta from N agent(s) on the first call).

The root agent has /meta RO-bound inside its container: git -C /meta log --oneline is the swarm-wide deploy log, cat /meta/flake.lock | jq '.nodes["agent-<n>"].locked' resolves which sha the flake pins each agent at right now. Dashboard surfaces the same info as a deployed:<sha12> chip per container row.

Two repos per agent

/var/lib/hyperhive/agents/<name>/config/    proposed — parent mount is RO
└── <anything>                              # any files the submitting
                                            # agent wants in the commit.
                                            # agent.nix is the
                                            # convention entry
                                            # point; flake.nix is
                                            # tracked boilerplate
                                            # (submitting agent doesn't
                                            # edit it).

/var/lib/hyperhive/applied/<name>/          applied — core-only
├── .git/                                   # tag-rich history
├── flake.nix                               # tracked, fixed
│                                           # boilerplate exporting
│                                           # nixosModules.default
├── agent.nix                               # working tree of main
└── <other committed files>                 # also tracked

/var/lib/hyperhive/meta/                    swarm-wide flake — core
├── .git/                                   # one commit per lock
│                                           # change
├── flake.nix                               # generated from agent set
└── flake.lock                              # pins each agent's sha

Why two physical repos: the submitting agent's /agents/<n>/config/ is RW — a buggy or hostile agent can git clean -fdx its own proposed tree. The applied repo is never bind-mounted (except the read-only .git exposure described below) so a destructive move inside the container can't reach it.

The container's --flake ref is /var/lib/hyperhive/meta#<name> (see "Meta flake" above). The agent's own applied/<n>/flake.nix is a fixed boilerplate that exports nixosModules.default = import ./agent.nix; the meta flake imports that module and wraps it with identity + HIVE_PORT / HIVE_LABEL / HIVE_DASHBOARD_PORT.

Tag state machine

hive-c0re plants deployed/0 on the seed commit at first spawn. A merged config commit that deploys plants no tag: the merged PR on the forge and meta's lock commits record it. A config PR nobody merges carries no extra state on the forge side: the PR stays open, and the submitter pushes again (or closes it) to retry.

Dispatch via the job queue

Long-running approval work — UpdateMetaInputs — runs as a DAG on the global job queue (docs/scheduler/coordinator.md::Job queue), submitted by the approval handler rather than run inline:

ApprovalKind DAG submitted source
UpdateMetaInputs meta_update (MetaLock + rebuild fan-out) approval
SchedulePrompt — runs inline (single sqlite insert) —

The DAG carries the originating approval_id. Every queued kind resolves through actions::resolve_approval_dag when its DAG settles terminal: the DAG's own terminal state is the authoritative outcome. That hook fires the matching HelperEvent::* via finish_approval.

Two visible consequences:

  • Operator dashboard: after selecting APPR0VE the work-in-progress shows up on the rebuild queue card (GET /api/jobq/graph, refetched on every rebuild_queue_changed tick), not on the approvals panel (which already moved the row to "approved"). A long meta-update cascade renders as a parent DAG with one child rebuild per affected agent — see docs/web-ui/dashboard.md for the layout.
  • Cancellation: the dashboard's × cancel button on a still-queued DAG calls POST /api/rebuild-queue/{id}/cancel, which flips it to Cancelled before any node runs (and fails the approval row instead of leaving it dangling). Returns {"cancelled": true} on success, {"cancelled": false} once any node started — terminal states can't be retroactively rewritten.

The approval source + approval_id mean a tail-end build failure surfaces back as a failed approval row, not just a silent queue entry. manual (dashboard ↻ R3BU1LD) and auto_update (boot reconcile) DAGs use the same queue but skip the approval plumbing.

Forge mirror

The bundled hive-forge container runs on the swarm's forge host (deploy.forgejo.enable, see ../swarm/services.md), and hive-c0re mirrors every agent's applied repo into a private agent-configs Forgejo org. forge::push_config(<name>) pushes every tag, then applied/main, to agent-configs/<name> on every startup sweep and every rebuild. Forge main is branch-protected, so the forge routinely refuses a push of an established main, which hive-c0re expects. Pushes are best-effort — a missing or stopped forge never blocks a deploy.

Each agent is a write collaborator on its own agent-configs/<name> repo — so it can push a branch and open a config PR — but not a member of any other agent's, so it can't reach another agent's config through the forge. Branch protection keeps the agent off the main push allowlist and allowlists merging to the operators team only, so an agent can't fast-forward its own config or self-merge its PR (see Config changes above). hive-c0re passes the tokenised push URL inline to git push, never writing it into applied/<n>/.git/config; that repo is RO-bind-mounted into the root agent, and a stored token would leak core's admin credential to an agent.

The dashboard deep-links into this org — a config repo link per container row. See docs/web-ui/dashboard.md.

Submitting agent's view of config repos

An agent holding ManageRootAgent has every other agent's config repo bind-mounted read-only (hive-c0re/src/lifecycle/host_config.rs calls bind_child_agent_dirs for each entry in topology::all_agents()). it's a copy to read another agent's current config — not an editing surface. An agent without the capability sees no other agent's config at all.

An agent with the approvals tool group submits a change the same way it makes any other change: clone the target agent's config repo from the forge into its own state dir, commit on a branch, open a PR, and let the operator review and approve it. By design, no second, mount-shaped path reaches the same file without the review.

Agents holding the manage_root_agent capability (granted per agent in capabilities.json; see hive-sh4re/src/permissions.rs) get additional host-side bind mounts via set_nspawn_flags:

  • /var/lib/hyperhive/agents/ → /agents/ (RW) — every agent's proposed repo, not just direct children. The capability means "may manage any agent", so the mounted set is every agent.
  • /var/lib/hyperhive/applied/ → /applied/ (RO) — every agent's authoritative applied repo, including .git.
  • /var/lib/hyperhive/meta/ → /meta/ (RO) — the swarm-wide deploy flake.

An agent without the capability only has its direct children's config dirs.

⚠️ nspawn bakes bind flags at container start, so granting or revoking this capability doesn't change any mount until that agent's container is rebuilt/restarted.

The root agent gets the capability by default, seeded on its autodeploy path (workers::auto_update::ensure_root_agent) so the recovery mounts are there from its first container. That seed only fires while capabilities.json doesn't exist yet: any grant or revoke through the dashboard creates the file, so a revoked root-agent grant stays revoked and isn't re-applied on the next hive-c0re restart.

Each proposed repo (/agents/<n>/config/) is pre-configured with applied as a git remote pointing at /applied/<n>/.git. Useful incantations from inside an agent with the full /applied mount:

git -C /agents/<n>/config fetch applied
git -C /agents/<n>/config log applied/main --oneline
git -C /agents/<n>/config show applied/refs/tags/deployed/0     # the seed commit
git -C /agents/<n>/config show applied/refs/tags/denied/<id>   # body = operator note
git -C /agents/<n>/config rebase applied/main                 # base in-flight work on what's deployed

git -C /meta log --oneline                                    # swarm-wide deploy history
cat /meta/flake.lock | jq '.nodes | with_entries(select(.key | startswith("agent-")))'

The RO binds block push at the kernel level — git plumbing inside the container can't corrupt either authoritative repo.

Startup migrations (older hosts)

hive-c0re runs a couple of idempotent migrations on every startup so a host set up before the tag-driven-deploy + meta-flake scheme (both described above) converges to it automatically. Each phase is a no-op once already applied:

  • Tags: hive-c0re tags agents from before the tag-driven scheme deployed/0 on main once. Non-destructive — it doesn't touch live containers, state dirs, or claude creds.
  • Meta flake: rewrites each applied/<n>/flake.nix to the module-only boilerplate, wires the applied remote in each proposed repo, and bootstraps the meta repo from the current agent list. Set HIVE_SKIP_META_MIGRATION=1 on the service to defer this phase.

No state loss in either migration: claude creds, /state/ notes, the events DB, and both proposed + applied history all survive. The root agent keeps its session; sub-agents stay logged in.

The root/bootstrap container is hive-c0re-managed

The root agent container runs through the same lifecycle as sub-agents. On hive-c0re serve startup, if ruth is missing, hive-c0re creates it. The root agent's flake lives at /var/lib/hyperhive/applied/ruth/; its proposed config at /var/lib/hyperhive/agents/ruth/config/. The root agent can edit its own agent.nix (visible inside the container at /agents/ruth/config/) and open a config PR on agent-configs/ruth for operator approval, same as any other agent.

Differences from sub-agents:

  • flake.nix extends hyperhive.nixosConfigurations.ruth (vs agent-base).
  • Web UI port via lifecycle::agent_web_port("ruth") — same FNV-1a hash as every other agent (8100..8999 range).
  • set_nspawn_flags adds two extra binds: /var/lib/hyperhive/agents → /agents (RW) so the root agent can edit per-agent proposed repos, and /var/lib/hyperhive/applied → /applied (RO) so the root agent can git fetch deployed/failed/denied tags from any agent's authoritative applied repo (see "Root-agent view of applied" below).
  • First-deploy spawn bypasses the approval queue (the root agent is required infrastructure).
  • socket_server::start_manager binds the root agent's socket, pure transport with no dedicated helpers — it uses the same per-agent runtime dir as any other agent (/run/hyperhive/agents/ruth/), not a special manager-only path.

Migration note (for older hosts): drop any containers.root = { ... } block from your host NixOS config. hyperhive creates and updates the root agent itself.

Root-agent policy

The system prompt (hive-agent/prompts/system.md, rendered by hive-agent/src/prompt.rs) is the same for every agent; what varies is which MCP tools it surfaces (gated by tool groups and capabilities in agent.nix). No role:manager block renders only for the root agent. The root agent's approval-gating behaviour comes from its CLAUDE.md / agent-specific instructions, not the system prompt template.

Helper events to the submitting agent

Coordinator::notify_submitter(approval_id, &HelperEvent) routes the event to the agent that originally submitted the approval (looked up from the submitter column on the approvals table). The harness delivers it as a regular system inbox message so it drives a normal claude turn. finish_approval fires an ApprovalResolved HelperEvent this way for every approval kind's terminal state. A "FYI, check when convenient" event doesn't need a message — those go through Coordinator::push_todo instead, a direct live dial of the target agent's in-container todo socket (same UpsertTodo request in-container producers use). Legacy approval rows that predate the submitter column fall back to the root agent. Variants (hive_sh4re::manager::HelperEvent):

  • ApprovalResolved { id, agent, commit_ref, status, note } — fired by actions::approve + actions::deny whenever an approval transitions to its terminal state.
  • ContainerCrash { agent, note } — crash_watch: a previously- running container went away with no operator-initiated transient state (Stopping / Restarting / Destroying / Rebuilding) AND nothing cleared that transient in the last 30s (RECENT_TRANSIENT_GRACE tombstone, three POLL_INTERVALs — closes the race where a lifecycle op finishes between two crash-watch polls and the container shows briefly as "stopped without transient" before the next start). The root agent escalates to the operator, who starts it again from the dashboard.
  • NeedsUpdate { agent } — sub-agent's recorded flake rev is stale. The operator rebuilds it from the dashboard.

The remaining lower-urgency lifecycle notices — Rebuilt, Killed, Destroyed, NeedsLogin, LoggedIn — are "FYI, check when convenient" events with no reason to drive an immediate turn, so they deliver via push_todo (see above) instead of HelperEvent: an agent_todo_socket push instead of a broker message, subsystem = "core", key = "<event>:<agent>" for dedup, and a single free-text summary (rebuilt_todo_summary renders Rebuilt's ok/note/sha/tag fields into that string).

To add a new lifecycle notice: if it needs to drive an immediate turn (something genuinely urgent, like ContainerCrash), add a HelperEvent variant + call sites + update prompts/system.md's message-event list. If it's "FYI, check when convenient," call push_todo directly instead — no new wire type needed.

Autoupdate on startup

hive-c0re serve runs auto_update::run in a background task right after opening the coordinator. It enumerates managed containers and rebuilds any whose recorded hyperhive rev differs from the current one — sub-agents and the root agent go through the same job_queue::templates::rebuild DAG.

"Rev" = canonical filesystem path of services.hyperhive.c0re.hyperhiveFlake. Marker file: /var/lib/hyperhive/applied/.<name>.hyperhive-rev. If the flake input has no canonical path (for example a github: URL), autoupdate is a no-op — rebuild manually.

The dashboard surfaces pending updates per agent: a clickable "needs update ↻" badge appears whenever the marker differs from current rev. The badge POSTs /api/rebuild/<name>, which inserts the same job_queue::templates::rebuild DAG so manual triggers and the startup scan can't drift. When at least one container is stale, a top-level ↻ UPD4TE 4LL button appears that loops over every stale container.