Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.
Refs #4850
570 lines
29 KiB
Markdown
570 lines
29 KiB
Markdown
# Approvals + helper events
|
||
|
||
The approval queue is where an agent asks the operator for something it
|
||
can't do itself: a scheduled prompt, or (legacy rows only) a meta-input
|
||
bump. Config changes don't go through this queue: they're pull requests
|
||
on the agent's config repo, which an operator merges on the forge, and
|
||
that merge deploys them. The submitting agent — any agent with the
|
||
`approvals` tool group, which manages the config of its **direct
|
||
children** (the root agent for top-level agents; a sub-manager for its
|
||
own subtree) — is the policy gate in front of both; helper events are how
|
||
it stays informed about what happens after a decision lands.
|
||
|
||
## For operators
|
||
|
||
- **Config change** — an agent proposed a change to another agent's
|
||
config (or its own, via a sub-manager) as a pull request on
|
||
`agent-configs/<agent>`. Review it on the forge like any other PR, and
|
||
merge it there to deploy it. See [Config changes](#config-changes).
|
||
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
|
||
a message to one or more inboxes at a future time. It lands on your
|
||
dashboard's Y3R C4LL tab (or `hivectl approvals pending` /
|
||
`approve <id>` from the CLI). You can also add schedules yourself
|
||
directly from the SCH3DUL3S tab, which skips this approval step
|
||
entirely — the gate here is specifically for an *agent* asking to
|
||
schedule something, not for you doing it.
|
||
- **Meta/flake update** (`UpdateMetaInputs`) — legacy: nothing queues
|
||
this kind, but an existing row still reads back and you can approve it.
|
||
Approving runs the update and commits the lock change; it doesn't
|
||
rebuild anything by itself.
|
||
|
||
Don't want to approve something? **Deny it** (`DENY` on the dashboard
|
||
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
|
||
submitting agent is always notified that the operator denied their request; what's
|
||
optional is only the reason text, which you can add on the dashboard's
|
||
prompt (cancelling that prompt aborts the whole deny, not just the
|
||
reason) but not from the CLI. Denying is final: a denied approval
|
||
can't be re-approved later, the agent has to submit a fresh request.
|
||
|
||
Everything below this point is the implementation detail behind that
|
||
flow.
|
||
|
||
## Config changes
|
||
|
||
Config changes flow through a **forge pull request** on the agent's
|
||
`agent-configs/<name>` repo — the same surface agents use for code PRs.
|
||
No MCP tool exists for config changes: opening the PR IS the request.
|
||
|
||
1. The submitting agent (the child's parent, holding the `approvals`
|
||
tool group) **clones** `agent-configs/<name>`, edits it there (any
|
||
tracked path, but `agent.nix` is the contract entry point), commits
|
||
with its own git identity, and pushes a branch + opens a PR with
|
||
`hive-forge` — the same way it would change any other repo.
|
||
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
|
||
config, not the tree to edit: authoring in place there produces no PR.
|
||
(it's currently mounted read-write, which is a defect tracked
|
||
separately, not an authoring path.)
|
||
Branch protection (the agent isn't on the `main` push allowlist;
|
||
merge and approvals allowlist = `operators` team; see "Forge mirror"
|
||
below) makes the agent a write collaborator that **can't merge its own
|
||
config PR**.
|
||
2. An operator reviews the PR **on the forge** (native diff, threaded
|
||
comments, CI status) and merges it there. Merging needs membership of
|
||
the `operators` team in `agent-configs`. The team starts empty: add
|
||
each operator by hand in the forge UI.
|
||
3. swarm-controller gets the `agent-configs` org's `pull_request`
|
||
delivery. A `closed` event with `merged: true` and base branch `main`
|
||
names the commit in `merge_commit_sha`.
|
||
4. The controller looks up which hive's wanted state places the agent
|
||
and publishes a deploy request carrying that commit on that hive's
|
||
deploy subject. With no such hive, or more than one, it deploys
|
||
nothing and logs a warning naming them.
|
||
5. If the hive's `applied/main` already is that commit, it does nothing.
|
||
Otherwise it fetches the config repo's `main`, requires the commit to
|
||
descend from `applied/main`, fast-forwards `applied/main` to it, and
|
||
queues a rebuild. A failure before the rebuild (the fetch, or a commit
|
||
that doesn't descend) deploys nothing and posts the error as a comment
|
||
on the merged PR. An agent with no container on that hive yet ignores
|
||
the commit; its first deploy builds what it seeds.
|
||
|
||
This path runs **no eval-verify**. A merged config that doesn't
|
||
evaluate or build fails the rebuild, and `applied/main` stays at that
|
||
commit, so the agent's rebuilds keep failing until a fix merges. A
|
||
failed rebuild shows on the hive like any other and isn't commented on
|
||
the PR. A merge whose webhook delivery never reaches the controller
|
||
deploys nothing.
|
||
|
||
## Approval queue
|
||
|
||
### Withdrawing a pending approval
|
||
|
||
The submitting agent can call `cancel_loose_end(kind: "approval", id)` to
|
||
withdraw an approval the operator hasn't acted on yet.
|
||
The row transitions to `ApprovalStatus::Cancelled` (distinct from
|
||
`Denied`/`Failed`), the dashboard pulls the card out of the
|
||
pending pane, and `ApprovalResolved { status: "cancelled" }` fires
|
||
on the root agent + dashboard channels. Approvals that the operator
|
||
(or a lifecycle failure) has already approved, denied, or failed return an error — the resolution is
|
||
final once acted on.
|
||
|
||
The socket refuses the `approval` kind with a clear error for any
|
||
agent that lacks the `approvals` tool group: only an agent with that
|
||
group submits approvals (for its direct children), so an agent
|
||
without it has nothing of its own to withdraw.
|
||
|
||
Creating a brand-new agent isn't an approval: it's swarm-level
|
||
(`swarmctl agent create`, `POST /api/agents` on the swarm controller),
|
||
and no hive can originate an agent. The swarm controller's
|
||
`InitAgentConfigRepo` job seeds the agent's config repo with a default
|
||
`agent.nix` template, then asks the target hive to deploy it.
|
||
|
||
Changing what the template seeded isn't a special case: like every
|
||
later change, it's a PR on that config repo, made from a clone and
|
||
merged by an operator on the forge (see [Config changes](#config-changes)).
|
||
|
||
### Approval kinds (wire shapes)
|
||
|
||
`ApprovalKind` carries two variants; each maps to a different
|
||
`commit_ref` encoding because `ApprovalKind` overloads that field as
|
||
the kind-specific payload carrier.
|
||
|
||
<!-- vale write-good.Passive = NO -->
|
||
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
|
||
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
|
||
etc.). hive-c0re sets the `agent` field to the requesting root agent.
|
||
On approve hive-c0re runs `nix flake update [inputs...]` on the
|
||
meta flake and commits the resulting lock changes.
|
||
- `SchedulePrompt` — `commit_ref` stores the JSON-encoded
|
||
`SchedulePromptPayload` (target list, body, schedule) so the
|
||
approval row carries the full submission verbatim. On approve
|
||
hive-c0re inserts a row into `scheduled_prompts` with
|
||
`source = approval:<id>`; the worker fans the body out as
|
||
inbox messages to each target at the scheduled time, recurring
|
||
when `interval_seconds` is set.
|
||
<!-- vale write-good.Passive = YES -->
|
||
|
||
### Scheduled prompts (submit paths)
|
||
|
||
Two ways a row lands in `scheduled_prompts`:
|
||
|
||
- **Operator-direct** (`source = "operator"`): the operator adds a schedule through the dashboard form. Lands in the table immediately, no approval gate — operator action is already the trust boundary.
|
||
- **Agent-requested** (`source = "approval:<id>"`): an agent submits a `RequestSchedulePrompt` through its MCP socket (the `request_schedule_prompt` tool, `scheduling` group). hive-c0re queues an `ApprovalKind::SchedulePrompt` row; on approve, it inserts the schedule row with `source = approval:<id>` so the audit trail points back at the operator decision (above).
|
||
|
||
No self-target shortcut: even agent-self schedules need approval. The existing `remind` MCP tool stays the quick self-wake path (no approval, lands directly in the agent's own inbox); this module is the bigger, multi-recipient, operator-visible thing.
|
||
|
||
### Scheduled prompt worker (catch-up clamp)
|
||
|
||
When hive-c0re comes back from being down, the worker sees rows whose `next_fire_at_unix` is well in the past. For recurring rows that would mean firing N delayed pulses in a row — spammy and useless. Instead the worker fires **once** per row and bumps `next_fire_at_unix` to the next interval slot ≥ `now`, recording how many cycles it skipped in `last_result` (per-target). Operators see "fired late, caught up from 17 skipped" instead of 17 wake-up storms.
|
||
|
||
The worker fires one-shot rows once (if past due, on the next worker pass) and deletes them; recurring rows survive until cancelled.
|
||
|
||
`targets` is its own table (`scheduled_prompt_targets`) so partial cancellation flips a single row and the dashboard can show last-fired / last-result per recipient. Cancelling every target reaps the parent row on the next worker pass.
|
||
|
||
### Scheduled prompt delivery: todo, not a broker message
|
||
|
||
An agent target's delivery is `push_todo` (`Coordinator::push_todo`,
|
||
`docs/scheduler/coordinator.md` covers the mechanism generally), not a broker
|
||
`Message` — a scheduled prompt wakes its target with a todo instead of
|
||
driving an immediate turn, by design. `key = "schedule:<id>"` per
|
||
target drives `push_todo`'s own upsert-by-key dedup: a re-fire of the
|
||
*same schedule* against a target that hasn't reviewed the last one
|
||
collapses into that one todo instead of stacking up.
|
||
|
||
### Missing-target failure
|
||
|
||
When a target name doesn't resolve to a known agent (container
|
||
destroyed, typo, etc.) the worker:
|
||
|
||
1. Records `last_result = "no such agent: <name>"` on the
|
||
per-target row.
|
||
2. Sends a single advisory `Message` from `system` to `operator`
|
||
naming the schedule, target, and reason. This one stays a `Message`
|
||
regardless of target type — it's a to-operator advisory about a
|
||
broken schedule, not the schedule's own delivery.
|
||
3. Continues fanning out to the other live targets.
|
||
|
||
Transient broker errors (sqlite lock contention, etc.) get the same
|
||
`last_result` annotation plus a `tracing::warn`, and then:
|
||
|
||
- **Recurring rows** re-arm to the next interval slot — the retry
|
||
self-heals on the next worker pass.
|
||
- **One-shot rows**: the worker deletes them unconditionally after their single
|
||
fan-out pass; a broker error on a one-shot isn't retried (the
|
||
operator advisory and `last_result` are the only audit trail).
|
||
|
||
### Reminder delivery: file-path semantics
|
||
|
||
A reminder may carry a `file_path` (the agent-visible path inside its
|
||
container, for example `/agents/<name>/state/foo.md`). On delivery hive-c0re:
|
||
|
||
1. **Translates** the container path to the host path
|
||
(`/var/lib/hyperhive/agents/<name>/state/foo.md`) so c0re can write
|
||
from outside the container.
|
||
2. **Validates** the path: rejects anything outside the agent's own state
|
||
subtree, containing `..` (path traversal), or with an empty relative
|
||
tail. On rejection hive-c0re skips the write and delivers the
|
||
original message inline with a warning — the reminder still fires.
|
||
3. **Defends against symlink escape**: after `create_dir_all`, hive-c0re
|
||
canonicalizes the parent dir and re-verifies it lives under the agent's host
|
||
state root. hive-c0re opens the final file with
|
||
`O_NOFOLLOW | O_CREAT | O_TRUNC` so an existing symlink at the
|
||
basename can't redirect the write to an arbitrary host path.
|
||
4. **Writes the body to disk** and delivers a short pointer message in its
|
||
place, keeping the agent's inbox / wake-prompt small while the agent
|
||
reads the bulky payload out of band.
|
||
|
||
`Broker::deliver_reminders_batch` handles atomicity of the inbox INSERT + `reminders.sent_at` UPDATE; the scheduler only computes the
|
||
body strings before calling it.
|
||
|
||
### Destroy semantics
|
||
|
||
`HostRequest::Destroy { name, purge }` is the lifecycle tear-down,
|
||
not an approval. Stops + removes the nspawn container, drops the
|
||
systemd drop-in, fails any pending approvals. Persistent state
|
||
(proposed/applied repos, claude credentials, `/state/` notes) is
|
||
**kept by default** — recreating the agent with the same name
|
||
reuses prior config + login. With `purge = true` the agent's
|
||
`/var/lib/hyperhive/{agents,applied}/<name>/` trees are also
|
||
wiped (config history + creds + notes gone forever). The
|
||
root/bootstrap container is destroyable like any other — hive-c0re
|
||
recreates it on the next startup if it's absent, so destroying it's
|
||
transient.
|
||
|
||
## Meta flake
|
||
|
||
The hive-c0re-owned repo at `/var/lib/hyperhive/meta/`
|
||
declares one flake input per agent (`agent-<n>.url =
|
||
"git+http://<forge>/agent-configs/<n>.git"`) and one
|
||
`nixosConfigurations.<n>` output per agent. Each output wraps
|
||
`inputs.agent-<n>.nixosModules.default` with the identity +
|
||
`HIVE_PORT` / `HIVE_LABEL` / `HIVE_DASHBOARD_PORT` injection module.
|
||
Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
|
||
|
||
The declared input url is the agent's **forge config repo** (the
|
||
same `agent-configs/<n>` config PRs merge into), so the meta flake
|
||
references a reviewable, reproducible source rather than a local
|
||
checkout. hive-c0re authenticates that `git+http` fetch via a git
|
||
credential helper that reads the live forge-core token — no token in
|
||
the url or the lock. The rebuild paths, however, do **not** re-lock
|
||
from the forge: they `--override-input agent-<n>
|
||
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
|
||
`applied/<n>/main` was fast-forwarded to. That keeps a rebuild
|
||
reproducible and independent of forge reachability — rebuilds fire on
|
||
crash-restart and meta bumps, not just config merges — while the
|
||
declared url stays the forge. `sync_agents` re-renders + re-locks the
|
||
persistent input; a plain `nix flake lock` leaves an existing applied
|
||
override in place (it only re-locks when the declared url itself
|
||
changes), so the forge-declared / applied-deployed split is stable.
|
||
|
||
Lock updates are single-phase and commit when the lock changed:
|
||
`meta::lock_update_for_rebuild(name)` relocks one agent's input for a
|
||
relocking rebuild (the manual `↻ R3BU1LD` button, and the rebuild a
|
||
merged config commit queues), and `meta::lock_update_hyperhive()` is the
|
||
autoupdate flake-rev bump (one shot before per-agent rebuilds).
|
||
|
||
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
|
||
carries `hyperhive_flake`, `dashboard_port`, and the rest of the
|
||
per-hive config — is the idempotent reconciler called by `spawn`,
|
||
`destroy`, `rebuild`, and the startup migration. Renders `flake.nix`
|
||
from the agent list; if it differs from disk, runs
|
||
`nix flake lock` + commits as `regenerate meta flake` (or
|
||
`seed meta from N agent(s)` on the first call).
|
||
|
||
The root agent has `/meta` RO-bound inside its container:
|
||
`git -C /meta log --oneline` is the swarm-wide deploy log,
|
||
`cat /meta/flake.lock | jq '.nodes["agent-<n>"].locked'`
|
||
resolves which sha the flake pins each agent at right now.
|
||
Dashboard surfaces the same info as a `deployed:<sha12>` chip
|
||
per container row.
|
||
|
||
## Two repos per agent
|
||
|
||
```
|
||
/var/lib/hyperhive/agents/<name>/config/ proposed — parent mount is RO
|
||
└── <anything> # any files the submitting
|
||
# agent wants in the commit.
|
||
# agent.nix is the
|
||
# convention entry
|
||
# point; flake.nix is
|
||
# tracked boilerplate
|
||
# (submitting agent doesn't
|
||
# edit it).
|
||
|
||
/var/lib/hyperhive/applied/<name>/ applied — core-only
|
||
├── .git/ # tag-rich history
|
||
├── flake.nix # tracked, fixed
|
||
│ # boilerplate exporting
|
||
│ # nixosModules.default
|
||
├── agent.nix # working tree of main
|
||
└── <other committed files> # also tracked
|
||
|
||
/var/lib/hyperhive/meta/ swarm-wide flake — core
|
||
├── .git/ # one commit per lock
|
||
│ # change
|
||
├── flake.nix # generated from agent set
|
||
└── flake.lock # pins each agent's sha
|
||
```
|
||
|
||
Why two physical repos: the submitting agent's `/agents/<n>/config/` is
|
||
RW — a buggy or hostile agent can `git clean -fdx` its own
|
||
proposed tree. The applied repo is never bind-mounted (except
|
||
the read-only `.git` exposure described below) so a destructive
|
||
move inside the container can't reach it.
|
||
|
||
The container's `--flake` ref is `/var/lib/hyperhive/meta#<name>`
|
||
(see "Meta flake" above). The agent's own `applied/<n>/flake.nix`
|
||
is a fixed boilerplate that exports `nixosModules.default =
|
||
import ./agent.nix`; the meta flake imports that module and
|
||
wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` /
|
||
`HIVE_DASHBOARD_PORT`.
|
||
|
||
### Tag state machine
|
||
|
||
hive-c0re plants `deployed/0` on the seed commit at first spawn. A
|
||
merged config commit that deploys plants no tag: the merged PR on the
|
||
forge and meta's lock commits record it. A config PR nobody merges
|
||
carries no extra state on the forge side: the PR stays open, and
|
||
the submitter pushes again (or closes it) to retry.
|
||
|
||
### Dispatch via the job queue
|
||
|
||
Long-running approval work — `UpdateMetaInputs` — runs as a DAG on the
|
||
global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted
|
||
by the approval handler rather than run inline:
|
||
|
||
| `ApprovalKind` | DAG submitted | source |
|
||
|---|---|---|
|
||
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
|
||
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
|
||
|
||
The DAG carries the originating `approval_id`. **Every** queued kind
|
||
resolves through `actions::resolve_approval_dag` when its DAG settles
|
||
terminal: the DAG's own terminal state is the authoritative outcome.
|
||
That hook fires the matching `HelperEvent::*` via `finish_approval`.
|
||
|
||
Two visible consequences:
|
||
|
||
- **Operator dashboard**: after selecting APPR0VE the work-in-progress
|
||
shows up on the *rebuild queue* card (`GET /api/jobq/graph`, refetched
|
||
on every `rebuild_queue_changed` tick), not on the approvals panel
|
||
(which already moved the row to "approved"). A long meta-update
|
||
cascade renders as a parent DAG with one child rebuild per affected
|
||
agent — see `docs/web-ui/dashboard.md` for the layout.
|
||
- **Cancellation**: the dashboard's *× cancel* button on a still-queued
|
||
DAG calls `POST /api/rebuild-queue/{id}/cancel`, which flips it to
|
||
`Cancelled` before any node runs (and fails the approval row instead
|
||
of leaving it dangling). Returns `{"cancelled": true}` on success,
|
||
`{"cancelled": false}` once any node started — terminal states can't
|
||
be retroactively rewritten.
|
||
|
||
The `approval` source + `approval_id` mean a tail-end build failure
|
||
surfaces back as a failed approval row, not just a silent queue
|
||
entry. `manual` (dashboard ↻ R3BU1LD) and `auto_update` (boot
|
||
reconcile) DAGs use the same queue but skip the approval plumbing.
|
||
|
||
### Forge mirror
|
||
|
||
The bundled `hive-forge` container runs on the swarm's forge host
|
||
(`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)),
|
||
and hive-c0re mirrors every agent's applied repo into a
|
||
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes
|
||
every tag, then `applied/main`, to `agent-configs/<name>` on every
|
||
startup sweep and every rebuild. Forge `main` is branch-protected, so
|
||
the forge routinely refuses a push of an established `main`, which
|
||
hive-c0re expects. Pushes are best-effort — a missing or stopped forge never
|
||
blocks a deploy.
|
||
|
||
Each agent is a **write collaborator on its own** `agent-configs/<name>`
|
||
repo — so it can push a branch and open a config PR — but not a member
|
||
of any other agent's, so it can't reach another agent's config through
|
||
the forge. Branch protection keeps the agent off the `main` push
|
||
allowlist and allowlists merging to the `operators` team only, so an
|
||
agent can't fast-forward its own config or self-merge its PR (see
|
||
[Config changes](#config-changes) above). hive-c0re passes the tokenised
|
||
push URL inline to `git push`, never writing it into
|
||
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
|
||
agent, and a stored token would leak core's admin credential to an
|
||
agent.
|
||
|
||
The dashboard deep-links into this org — a `config repo` link
|
||
per container row. See `docs/web-ui/dashboard.md`.
|
||
|
||
### Submitting agent's view of config repos
|
||
|
||
An agent holding `ManageRootAgent` has every other agent's config repo
|
||
bind-mounted **read-only** (`hive-c0re/src/lifecycle/host_config.rs`
|
||
calls `bind_child_agent_dirs` for each entry in
|
||
`topology::all_agents()`). it's a copy to *read* another agent's
|
||
current config — not an editing surface. An agent without the
|
||
capability sees no other agent's config at all.
|
||
|
||
An agent with the `approvals` tool group submits a change the same way
|
||
it makes any other change: **clone the target agent's config repo from the
|
||
forge into its own state dir, commit on a branch, open a PR**, and let
|
||
the operator review and approve it. By design, no second, mount-shaped
|
||
path reaches the same file without the review.
|
||
|
||
Agents holding the `manage_root_agent` capability (granted per agent in
|
||
`capabilities.json`; see `hive-sh4re/src/permissions.rs`) get additional
|
||
host-side bind mounts via `set_nspawn_flags`:
|
||
|
||
- `/var/lib/hyperhive/agents/` → `/agents/` (RW) — **every** agent's
|
||
proposed repo, not just direct children. The capability means "may
|
||
manage any agent", so the mounted set is every agent.
|
||
- `/var/lib/hyperhive/applied/` → `/applied/` (RO) — every agent's
|
||
authoritative applied repo, including `.git`.
|
||
- `/var/lib/hyperhive/meta/` → `/meta/` (RO) — the swarm-wide
|
||
deploy flake.
|
||
|
||
An agent without the capability only has its direct children's config
|
||
dirs.
|
||
|
||
⚠️ nspawn bakes bind flags at container start, so granting or
|
||
revoking this capability doesn't change any mount until that agent's
|
||
container is rebuilt/restarted.
|
||
|
||
The root agent gets the capability by default, seeded on its autodeploy
|
||
path (`workers::auto_update::ensure_root_agent`) so the recovery mounts
|
||
are there from its first container. That seed only fires while
|
||
`capabilities.json` doesn't exist yet: any grant or revoke through
|
||
the dashboard creates the file, so a revoked root-agent grant stays
|
||
revoked and isn't re-applied on the next hive-c0re restart.
|
||
|
||
Each proposed repo (`/agents/<n>/config/`) is pre-configured
|
||
with `applied` as a git remote pointing at
|
||
`/applied/<n>/.git`. Useful incantations from inside an agent with
|
||
the full `/applied` mount:
|
||
|
||
```sh
|
||
git -C /agents/<n>/config fetch applied
|
||
git -C /agents/<n>/config log applied/main --oneline
|
||
git -C /agents/<n>/config show applied/refs/tags/deployed/0 # the seed commit
|
||
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
|
||
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
|
||
|
||
git -C /meta log --oneline # swarm-wide deploy history
|
||
cat /meta/flake.lock | jq '.nodes | with_entries(select(.key | startswith("agent-")))'
|
||
```
|
||
|
||
The RO binds block push at the kernel level — git plumbing inside the
|
||
container can't corrupt either authoritative repo.
|
||
|
||
## Startup migrations (older hosts)
|
||
|
||
hive-c0re runs a couple of idempotent migrations on every startup so a
|
||
host set up before the tag-driven-deploy + meta-flake scheme (both
|
||
described above) converges to it automatically. Each phase is a no-op
|
||
once already applied:
|
||
|
||
- **Tags**: hive-c0re tags agents from before the tag-driven scheme
|
||
`deployed/0` on `main` once. Non-destructive — it doesn't touch live
|
||
containers, state dirs, or claude creds.
|
||
- **Meta flake**: rewrites each `applied/<n>/flake.nix` to the
|
||
module-only boilerplate, wires the `applied` remote in each proposed
|
||
repo, and bootstraps the meta repo from the current agent list. Set
|
||
`HIVE_SKIP_META_MIGRATION=1` on the service to defer this phase.
|
||
|
||
No state loss in either migration: claude creds, `/state/` notes, the
|
||
events DB, and both proposed + applied history all survive. The root
|
||
agent keeps its session; sub-agents stay logged in.
|
||
|
||
## The root/bootstrap container is hive-c0re-managed
|
||
|
||
The root agent container runs through the **same lifecycle as
|
||
sub-agents**. On `hive-c0re serve` startup, if `ruth` is missing,
|
||
hive-c0re creates it. The root agent's flake lives at
|
||
`/var/lib/hyperhive/applied/ruth/`; its proposed config at
|
||
`/var/lib/hyperhive/agents/ruth/config/`. The root agent can edit its own
|
||
`agent.nix` (visible inside the container at `/agents/ruth/config/`)
|
||
and open a config PR on `agent-configs/ruth` for operator approval,
|
||
same as any other agent.
|
||
|
||
Differences from sub-agents:
|
||
|
||
- `flake.nix` extends `hyperhive.nixosConfigurations.ruth`
|
||
(vs `agent-base`).
|
||
- Web UI port via `lifecycle::agent_web_port("ruth")` — same
|
||
FNV-1a hash as every other agent (8100..8999 range).
|
||
- `set_nspawn_flags` adds two extra binds: `/var/lib/hyperhive/agents`
|
||
→ `/agents` (RW) so the root agent can edit per-agent proposed repos,
|
||
and `/var/lib/hyperhive/applied` → `/applied` (RO) so the root agent
|
||
can `git fetch` deployed/failed/denied tags from any agent's
|
||
authoritative applied repo (see "Root-agent view of applied" below).
|
||
- First-deploy spawn bypasses the approval queue (the root agent is
|
||
required infrastructure).
|
||
- `socket_server::start_manager` binds the root agent's socket,
|
||
pure transport with no dedicated helpers — it uses the same
|
||
per-agent runtime dir as any other agent (`/run/hyperhive/agents/ruth/`),
|
||
not a special manager-only path.
|
||
|
||
**Migration note** (for older hosts): drop any `containers.root =
|
||
{ ... }` block from your host NixOS config. hyperhive creates and
|
||
updates the root agent itself.
|
||
|
||
## Root-agent policy
|
||
|
||
The system prompt (`hive-agent/prompts/system.md`, rendered by
|
||
`hive-agent/src/prompt.rs`) is the **same for every agent**; what
|
||
varies is which MCP tools it surfaces (gated by tool groups and
|
||
capabilities in `agent.nix`). No `role:manager` block renders only
|
||
for the root agent. The root agent's approval-gating
|
||
behaviour comes from its CLAUDE.md / agent-specific instructions, not
|
||
the system prompt template.
|
||
|
||
## Helper events to the submitting agent
|
||
|
||
`Coordinator::notify_submitter(approval_id, &HelperEvent)` routes the
|
||
event to the agent that originally submitted the approval (looked up from
|
||
the `submitter` column on the `approvals` table). The harness delivers it
|
||
as a regular `system` inbox message so it drives a normal claude turn.
|
||
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
|
||
**every** approval kind's terminal state. A
|
||
"FYI, check when convenient" event doesn't need a message — those go
|
||
through `Coordinator::push_todo` instead, a direct live dial of the target
|
||
agent's in-container todo socket (same `UpsertTodo` request in-container
|
||
producers use). Legacy
|
||
approval rows that predate the submitter column fall back to the
|
||
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
||
|
||
- `ApprovalResolved { id, agent, commit_ref, status, note }` —
|
||
fired by `actions::approve` + `actions::deny` whenever an
|
||
approval transitions to its terminal state.
|
||
- `ContainerCrash { agent, note }` — `crash_watch`: a previously-
|
||
running container went away with no operator-initiated transient
|
||
state (Stopping / Restarting / Destroying / Rebuilding) AND nothing
|
||
cleared that transient in the last 30s (`RECENT_TRANSIENT_GRACE`
|
||
tombstone, three `POLL_INTERVAL`s — closes the race where a
|
||
lifecycle op finishes between two crash-watch polls and the
|
||
container shows briefly as "stopped without transient" before
|
||
the next start). The root agent escalates to the operator, who
|
||
starts it again from the dashboard.
|
||
- `NeedsUpdate { agent }` — sub-agent's recorded flake rev is
|
||
stale. The operator rebuilds it from the dashboard.
|
||
|
||
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
|
||
`Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check
|
||
when convenient" events with no reason to drive an immediate turn, so
|
||
they deliver via `push_todo` (see above) instead
|
||
of `HelperEvent`: an `agent_todo_socket` push instead of a broker
|
||
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
|
||
and a single free-text `summary` (`rebuilt_todo_summary` renders
|
||
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
|
||
|
||
To add a new lifecycle notice: if it needs to drive an immediate turn
|
||
(something genuinely urgent, like `ContainerCrash`), add a
|
||
`HelperEvent` variant + call sites + update `prompts/system.md`'s
|
||
message-event list. If it's "FYI, check when convenient," call
|
||
`push_todo` directly instead — no new wire type
|
||
needed.
|
||
|
||
## Autoupdate on startup
|
||
|
||
`hive-c0re serve` runs `auto_update::run` in a background task right
|
||
after opening the coordinator. It enumerates managed containers and
|
||
rebuilds any whose recorded hyperhive rev differs from the current
|
||
one — sub-agents and the root agent go through the same
|
||
`job_queue::templates::rebuild` DAG.
|
||
|
||
"Rev" = canonical filesystem path of
|
||
`services.hyperhive.c0re.hyperhiveFlake`. Marker
|
||
file: `/var/lib/hyperhive/applied/.<name>.hyperhive-rev`. If the
|
||
flake input has no canonical path (for example a `github:` URL),
|
||
autoupdate is a no-op — rebuild manually.
|
||
|
||
The dashboard surfaces pending updates per agent: a clickable
|
||
"needs update ↻" badge appears whenever the marker differs from
|
||
current rev. The badge POSTs `/api/rebuild/<name>`, which inserts the
|
||
same `job_queue::templates::rebuild` DAG so manual triggers and the
|
||
startup scan can't drift. When at least one container is stale, a
|
||
top-level `↻ UPD4TE 4LL` button appears that loops over every
|
||
stale container.
|