docs: restructure into topic subdirectories, collapse duplicated index
Per mara's go-ahead on hyperhive#3902 ("getting started is good, but
terminal rendering does not go in there i think"):
Moved 21 top-level docs/*.md files into 7 new topic subdirectories
(existing web-ui/, turn-loop/, swarm/, tools/, crates/ untouched):
getting-started/ setup.md
agent-lifecycle/ agent-hierarchy.md, approvals.md, persistence.md
trust-boundary/ boundary.md, security.md
integrations/ forge.md, matrix.md, github.md, knowledge.md
networking/ gateway.md, network.md, snapshot-store.md
scheduler/ jobq.md, coordinator.md, ci.md, observability.md
process/ conventions.md, gotchas.md, pr-review-gate.md
web-ui/ terminal-rendering.md (moved into the EXISTING dir,
per mara's correction to the original getting-started
guess -- it's UI implementation detail, not onboarding)
The physical layout now matches docs/README.md's own topical headers,
which already amounted to this taxonomy -- see the scoping comment on
the issue for the two findings that motivated this (a genuine
duplication between CLAUDE.md's old "Reading paths" list and
docs/README.md's grouped one, since drifted out of sync with each
other; and the flat layout not matching the grouping we already had).
Fixed every cross-reference this moved across the whole repo (~120
files: docs/ internal links at every depth, Rust doc comments, nix
module option docs, crate READMEs) -- verified two ways: a grep sweep
confirming zero remaining references to any old path, and a script
that resolves every markdown link in docs/**/*.md + CLAUDE.md +
README.md against the filesystem and reports anything that doesn't
exist (zero broken links).
Collapsed CLAUDE.md's "Reading paths" section (the duplicate) down to
a pointer at docs/README.md, now the single index. Rewrote
docs/README.md itself to use the new subdirectory paths and added the
one doc it was missing that CLAUDE.md's old copy had (pr-review-gate.md).
Classified all 22 docs/*.md files first via a haiku subagent (mara's
suggestion) on two axes -- proposed grouping and operator-vs-
implementation focus -- before finalizing the taxonomy; spot-checked
the report and found internal inconsistencies (its classification
table disagreed with its own summary section for a few files), so this
taxonomy is my original proposal + the one correction mara gave
directly, not a blind application of the subagent's table. The
operator-focus data it gathered is still useful for a follow-up
content pass (docs skewing 'mixed' rather than pure operator-facing),
not addressed in this PR -- structure only.
nix fmt clean, both pre-push lints clean.
This commit is contained in:
parent
e4a22b4190
commit
07b62612b0
124 changed files with 301 additions and 377 deletions
679
docs/agent-lifecycle/approvals.md
Normal file
679
docs/agent-lifecycle/approvals.md
Normal file
|
|
@ -0,0 +1,679 @@
|
|||
# Approvals + helper events
|
||||
|
||||
The approval queue is hyperhive's pivot: nothing that changes the
|
||||
shape of an agent (its config, whether it exists) happens without an
|
||||
operator click. The submitting agent — any agent with the `approvals`
|
||||
tool group, which manages the config of its **direct children** (the
|
||||
root agent for top-level agents; a sub-manager for its own subtree) — is
|
||||
the policy gate in front of that queue; helper events are how it stays
|
||||
informed about what happens after a decision lands.
|
||||
|
||||
## For operators
|
||||
|
||||
Every add/remove/change to an agent lands on your dashboard's Y3R
|
||||
C4LL tab (or `hivectl approvals pending` / `approve <id>` from the
|
||||
CLI) before it takes effect. What you'll see, and what to do with it:
|
||||
|
||||
- **Config change** (`MergeConfigPr`) — an agent proposed a change to
|
||||
another agent's config (or its own, via a sub-manager) as a forge
|
||||
pull request. Review the diff on the forge — the dashboard card
|
||||
links straight to it, same as reviewing any other PR. Approving
|
||||
triggers the deploy automatically: hive-c0re re-verifies the PR
|
||||
hasn't moved since you looked at it, evaluates it (a dry run,
|
||||
nothing applied yet), merges it, and rebuilds the container. If
|
||||
anything in that chain fails, the change rolls back automatically —
|
||||
the agent stays on its last-good config, no recovery action needed
|
||||
from you.
|
||||
- **New agent** (`InitConfig` then `Spawn`) — creating a brand-new
|
||||
agent is two approvals. `InitConfig` creates the config repo and
|
||||
seeds it from a template; `Spawn` creates the container from that
|
||||
config. Tailoring the template first is not a separate mechanism —
|
||||
it's the config-change flow above, a PR you review like any other. Every later change goes through the config-change flow
|
||||
above — there's no repeat "spawn" for an existing agent.
|
||||
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
|
||||
one or more Nix flake inputs (or all of them). Approving runs the
|
||||
update and commits the lock change; it doesn't rebuild anything by
|
||||
itself.
|
||||
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
|
||||
a message to one or more inboxes at a future time. You can also add
|
||||
schedules yourself directly from the SCH3DUL3S tab, which skips this
|
||||
approval step entirely — the gate here is specifically for an
|
||||
*agent* asking to schedule something, not for you doing it.
|
||||
|
||||
Don't want to approve something? **Deny it** (`DENY` on the dashboard
|
||||
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
|
||||
submitting agent is always notified their request was denied; what's
|
||||
optional is only the reason text, which you can add on the dashboard's
|
||||
prompt (cancelling that prompt aborts the whole deny, not just the
|
||||
reason) but not from the CLI. Denying is final: a denied approval
|
||||
can't be re-approved later, the agent has to submit a fresh one (a new
|
||||
PR, a new request).
|
||||
|
||||
Everything below this point is the implementation detail behind that
|
||||
flow.
|
||||
|
||||
## End-to-end approval flow
|
||||
|
||||
Config changes flow through a **forge pull request** on the agent's
|
||||
`agent-configs/<name>` repo — the same surface agents use for code PRs.
|
||||
There is no bespoke MCP tool for config changes: opening the PR IS the
|
||||
request.
|
||||
|
||||
1. The submitting agent (the child's parent, holding the `approvals`
|
||||
tool group) **clones** `agent-configs/<name>`, edits it there (any
|
||||
tracked path, but `agent.nix` is the contract entry point), commits
|
||||
with its own git identity, and pushes a branch + opens a PR with
|
||||
`hive-forge` — the same way it would change any other repo.
|
||||
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
|
||||
config, not the tree to edit: authoring in place there produces no PR
|
||||
and no approval. (It is currently mounted read-write, which is a
|
||||
defect tracked separately, not an authoring path.)
|
||||
Branch protection (push/merge whitelist = `core`, approvals whitelist
|
||||
= operator team; see "Forge mirror" and #1787) makes the agent a
|
||||
write collaborator that **cannot merge its own config PR**.
|
||||
2. hive-c0re's `/webhook/config-pr` endpoint receives the Forgejo
|
||||
`pull_request` event (opened / synchronized / reopened) and queues a
|
||||
`MergeConfigPr` approval; a poll fallback catches any missed webhook.
|
||||
The approval row stores the PR **number** (`commit_ref`) and the PR
|
||||
**head sha at queue time** (`fetched_sha` — the "reviewed" sha). If
|
||||
the PR head later moves, the stale approval is superseded by a fresh
|
||||
one pinned to the new head, so the operator always reviews what will
|
||||
actually deploy.
|
||||
3. The operator reviews the PR **on the forge** (native diff, threaded
|
||||
comments, CI status) and sees a matching card on the dashboard with a
|
||||
"review PR on forge" deep link. They click ◆ APPR0VE (or
|
||||
`hivectl approvals approve <id>` on the CLI) once satisfied.
|
||||
4. On approve, a deploy DAG runs three phases under a resource-holding
|
||||
`DeployWindow` root (see *Queue templates* below):
|
||||
- `MergeVerify` re-reads the live PR head and **aborts if it drifted**
|
||||
from the reviewed `fetched_sha` (the submitter must push again,
|
||||
which queues a fresh approval); then fetches that head into the
|
||||
applied repo and **eval-verifies** it — a flake eval on a throwaway
|
||||
checkout. This is the trust gate: it relies on c0re's own eval, not
|
||||
on any in-repo (agent-forgeable) signal like a CI status. Nothing is
|
||||
mutated in this phase, so a rejection here leaves the forge and the
|
||||
applied repo exactly as they were.
|
||||
- `DeployApply` parks the pre-merge `applied/main` in
|
||||
`refs/hyperhive/rollback/<approval-id>`, then fast-forward-merges
|
||||
the reviewed head to the forge config repo's `main` (this IS the
|
||||
merge — a `core`-authenticated ff-merge pinned to the reviewed sha,
|
||||
so a moved PR head can't substitute bytes), and runs the deploy
|
||||
proper (`deploy_applied_target`): ff `applied/main`, two-phase meta
|
||||
deploy, container rebuild. On success it drops the rollback ref and
|
||||
plants `deployed/<id>`.
|
||||
- `DeployTail` runs on **every** outcome, including a cancel-cascade.
|
||||
If the rollback ref survived, the deploy never confirmed good: it
|
||||
rolls `applied/main` back, resyncs the working tree, and aborts the
|
||||
staged meta lock, so the agent stays on its last-good tree. Then it
|
||||
mirrors the config repo (and its new deploy tag) to the forge.
|
||||
|
||||
The rollback state lives in a **git ref, not a local variable**, on
|
||||
purpose: hive-c0re can restart between the apply and the tail, and the
|
||||
tail still has to know what to undo when it does.
|
||||
5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the
|
||||
**submitting agent's** inbox via `notify_submitter`, carrying both the
|
||||
canonical sha and the terminal tag (the approval row carries a
|
||||
`submitter` column recording the agent the change is for).
|
||||
|
||||
### Withdrawing a pending approval
|
||||
|
||||
The submitting agent can call `cancel_loose_end(kind: "approval", id)` to
|
||||
withdraw an approval that hasn't been acted on yet.
|
||||
The row transitions to `ApprovalStatus::Cancelled` (distinct from
|
||||
`Denied`/`Failed`), the dashboard pulls the card out of the
|
||||
pending pane, and `ApprovalResolved { status: "cancelled" }` fires
|
||||
on the root agent + dashboard channels. Approvals that have already
|
||||
been approved/denied/failed return an error — the resolution is
|
||||
final once the operator (or a lifecycle failure) acted on the row.
|
||||
|
||||
The socket refuses the `approval` kind with a clear error for any
|
||||
agent that lacks the `approvals` tool group: only an agent with that
|
||||
group submits approvals (for its direct children), so an agent
|
||||
without it has nothing of its own to withdraw.
|
||||
|
||||
`InitConfig` approvals create a brand-new agent's config repo. On
|
||||
approve, hive-c0re seeds it with a default `agent.nix` template and
|
||||
pushes a todo (`push_todo_submitter`) into the submitting agent's
|
||||
in-container store. The operator then **spawns** the agent (the
|
||||
`Spawn` approval / `◆ R3QU3ST SP4WN` button), which creates the
|
||||
container from that config.
|
||||
|
||||
Changing what the template seeded is not a special case: like every
|
||||
later change, it's a PR on that config repo (`MergeConfigPr`), made
|
||||
from a clone, reviewed and approved by the operator. The PR flow is
|
||||
the one path — an operator can equally drive both steps herself
|
||||
through the web UI or the forge.
|
||||
|
||||
### Approval kinds (wire shapes)
|
||||
|
||||
`ApprovalKind` carries five variants; each maps to a different
|
||||
`commit_ref` encoding because that field is overloaded as the
|
||||
kind-specific payload carrier.
|
||||
|
||||
- `MergeConfigPr` — the config-change flow. Triggered automatically:
|
||||
when an agent opens (or force-pushes) a PR on its
|
||||
`agent-configs/<agent>` forge repo, hive-c0re's `/webhook/config-pr`
|
||||
endpoint receives the Forgejo pull_request event and queues this
|
||||
approval row. No MCP tool call needed — the forge PR IS the request.
|
||||
`commit_ref` stores the **PR number** (decimal), and `fetched_sha` is
|
||||
the PR **head sha at queue time** (the "reviewed" sha). On approve,
|
||||
the deploy DAG's `MergeVerify` phase re-reads the live PR head and
|
||||
aborts if it drifted from `fetched_sha` (submitter must push again to
|
||||
re-trigger), then fetches that head into the applied repo and
|
||||
eval-verifies it; `DeployApply` fast-forward-merges the forge config
|
||||
repo's `main` to it (the merge) and runs `deploy_applied_target`;
|
||||
`DeployTail` compensates on failure. Never a first spawn.
|
||||
- `Spawn` — direct container creation from the agent's config repo.
|
||||
`commit_ref` is empty. Submitted via `HostRequest::RequestSpawn`
|
||||
(operator-gated, the `◆ R3QU3ST SP4WN` dashboard button +
|
||||
`hivectl agent <name> request-create` CLI). The host-level `HostRequest::Spawn`
|
||||
variant bypasses the approval queue entirely — privileged-context use
|
||||
only (operator on the host shell, test scripts, one-off recoveries;
|
||||
`hivectl agent <name> create`). This is the **canonical first-spawn**: a new agent's `InitConfig`
|
||||
seeds its config repo, the submitting agent customises it, then the
|
||||
operator spawns to create the container. Subsequent config changes go
|
||||
through a `MergeConfigPr` PR.
|
||||
- `InitConfig` — `commit_ref` is empty; the variant just gates
|
||||
"seed the proposed repo with the default template" against
|
||||
operator approval. Step 1 of the two-step spawn flow above.
|
||||
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
|
||||
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
|
||||
etc.). `agent` field is set to the requesting root agent.
|
||||
On approve hive-c0re runs `nix flake update [inputs...]` on the
|
||||
meta flake and commits the resulting lock changes.
|
||||
- `SchedulePrompt` — `commit_ref` stores the JSON-encoded
|
||||
`SchedulePromptPayload` (target list, body, schedule) so the
|
||||
approval row carries the full submission verbatim. On approve
|
||||
hive-c0re inserts a row into `scheduled_prompts` with
|
||||
`source = approval:<id>`; the worker fans the body out as
|
||||
inbox messages to each target at the scheduled time, recurring
|
||||
when `interval_seconds` is set.
|
||||
|
||||
### Scheduled prompts (submit paths)
|
||||
|
||||
Two ways a row lands in `scheduled_prompts`:
|
||||
|
||||
- **Operator-direct** (`source = "operator"`): the operator adds a schedule through the dashboard form. Lands in the table immediately, no approval gate — operator action is already the trust boundary.
|
||||
- **Agent-requested** (`source = "approval:<id>"`): an agent submits a `RequestSchedulePrompt` through its MCP socket (the `request_schedule_prompt` tool, `scheduling` group). An `ApprovalKind::SchedulePrompt` row is queued; on approve, hive-c0re inserts the schedule row with `source = approval:<id>` so the audit trail points back at the operator decision (above).
|
||||
|
||||
No self-target shortcut: even agent-self schedules need approval. The existing `remind` MCP tool stays the quick self-wake path (no approval, lands directly in the agent's own inbox); this module is the bigger, multi-recipient, operator-visible thing.
|
||||
|
||||
### Scheduled prompt worker (catch-up clamp)
|
||||
|
||||
When hive-c0re comes back from being down, the worker sees rows whose `next_fire_at_unix` is well in the past. For recurring rows that would mean firing N delayed pulses in a row — spammy and useless. Instead the worker fires **once** per row and bumps `next_fire_at_unix` to the next interval slot ≥ `now`, recording how many cycles were skipped in `last_result` (per-target). Operators see "fired late, caught up from 17 skipped" instead of 17 wake-up storms.
|
||||
|
||||
One-shot rows fire once (if past due, on the next worker pass) and are deleted by the worker; recurring rows survive until cancelled.
|
||||
|
||||
`targets` is its own table (`scheduled_prompt_targets`) so partial cancellation flips a single row and the dashboard can show last-fired / last-result per recipient. Cancelling every target reaps the parent row on the next worker pass.
|
||||
|
||||
### Scheduled prompt delivery: todo, not a broker message
|
||||
|
||||
An agent target's delivery is `push_todo` (`Coordinator::push_todo`,
|
||||
`docs/coordinator.md` covers the mechanism generally), not a broker
|
||||
`Message` — a scheduled prompt wakes its target with a todo instead of
|
||||
driving an immediate turn, by design. `key = "schedule:<id>"` per
|
||||
target drives `push_todo`'s own upsert-by-key dedup: a re-fire of the
|
||||
*same schedule* against a target that hasn't reviewed the last one
|
||||
collapses into that one todo instead of stacking up.
|
||||
|
||||
**`operator` is the one exception** — it's a valid schedule target but
|
||||
has no in-container todo inbox, so it keeps the original broker
|
||||
`Message` path (the dashboard mirrors `to == operator` into its own
|
||||
pane, same as before).
|
||||
|
||||
### Missing-target failure
|
||||
|
||||
When a target name doesn't resolve to a known agent (container
|
||||
destroyed, operator typo, etc.) the worker:
|
||||
|
||||
1. Records `last_result = "no such agent: <name>"` on the
|
||||
per-target row.
|
||||
2. Sends a single advisory `Message` from `system` to `operator`
|
||||
naming the schedule, target, and reason. This one stays a `Message`
|
||||
regardless of target type — it's a to-operator advisory about a
|
||||
broken schedule, not the schedule's own delivery.
|
||||
3. Continues fanning out to the other live targets.
|
||||
|
||||
Transient broker errors (sqlite lock contention, etc.) get the same
|
||||
`last_result` annotation plus a `tracing::warn`, and then:
|
||||
|
||||
- **Recurring rows** re-arm to the next interval slot — the retry
|
||||
self-heals on the next worker pass.
|
||||
- **One-shot rows** are deleted unconditionally after their single
|
||||
fan-out pass; a broker error on a one-shot is not retried (the
|
||||
operator advisory and `last_result` are the only audit trail).
|
||||
|
||||
### Reminder delivery: file-path semantics
|
||||
|
||||
A reminder may carry a `file_path` (the agent-visible path inside its
|
||||
container, e.g. `/agents/<name>/state/foo.md`). On delivery hive-c0re:
|
||||
|
||||
1. **Translates** the container path to the host path
|
||||
(`/var/lib/hyperhive/agents/<name>/state/foo.md`) so c0re can write
|
||||
from outside the container.
|
||||
2. **Validates** the path: rejects anything outside the agent's own state
|
||||
subtree, containing `..` (path traversal), or with an empty relative
|
||||
tail. On rejection the write is skipped and the original message is
|
||||
delivered inline with a warning — the reminder still fires.
|
||||
3. **Defends against symlink escape**: after `create_dir_all`, the parent
|
||||
dir is canonicalized and re-verified to live under the agent's host
|
||||
state root. The final file is opened with
|
||||
`O_NOFOLLOW | O_CREAT | O_TRUNC` so an existing symlink at the
|
||||
basename cannot redirect the write to an arbitrary host path.
|
||||
4. **Writes the body to disk** and delivers a short pointer message in its
|
||||
place, keeping the agent's inbox / wake-prompt small while the bulky
|
||||
payload is read out of band.
|
||||
|
||||
Atomicity of the inbox INSERT + `reminders.sent_at` UPDATE is handled
|
||||
inside `Broker::deliver_reminders_batch`; the scheduler only computes the
|
||||
body strings before calling it.
|
||||
|
||||
### Destroy semantics
|
||||
|
||||
`HostRequest::Destroy { name, purge }` is the lifecycle tear-down,
|
||||
not an approval. Stops + removes the nspawn container, drops the
|
||||
systemd drop-in, fails any pending approvals. Persistent state
|
||||
(proposed/applied repos, claude credentials, `/state/` notes) is
|
||||
**kept by default** — recreating the agent with the same name
|
||||
reuses prior config + login. With `purge = true` the agent's
|
||||
`/var/lib/hyperhive/{agents,applied}/<name>/` trees are also
|
||||
wiped (config history + creds + notes gone forever). The
|
||||
root/bootstrap container is destroyable like any other — hive-c0re
|
||||
recreates it on the next startup if it's absent, so destroying it is
|
||||
transient.
|
||||
|
||||
## Meta flake
|
||||
|
||||
The hive-c0re-owned repo at `/var/lib/hyperhive/meta/`
|
||||
declares one flake input per agent (`agent-<n>.url =
|
||||
"git+http://<forge>/agent-configs/<n>.git"`) and one
|
||||
`nixosConfigurations.<n>` output per agent. Each output wraps
|
||||
`inputs.agent-<n>.nixosModules.default` with the identity +
|
||||
`HIVE_PORT` / `HIVE_LABEL` / `HIVE_DASHBOARD_PORT` injection module.
|
||||
Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
|
||||
|
||||
The declared input url is the agent's **forge config repo** (the
|
||||
same `agent-configs/<n>` the config-PR flow lands approved changes
|
||||
on), so the meta flake references a reviewable, reproducible source
|
||||
rather than a local checkout. hive-c0re authenticates that
|
||||
`git+http` fetch via a git credential helper that reads the live
|
||||
forge-core token — no token in the url or the lock. The deploy and
|
||||
manual-rebuild paths, however, do **not** re-lock from the forge:
|
||||
they `--override-input agent-<n>
|
||||
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
|
||||
that `verify_commit` gated and `applied/<n>/main` was
|
||||
fast-forwarded to. That keeps a deploy/rebuild reproducible and
|
||||
independent of forge reachability — rebuilds fire on crash-restart
|
||||
and meta bumps, not just config PRs — while the declared url stays
|
||||
the forge. `sync_agents` re-renders + re-locks the persistent input;
|
||||
a plain `nix flake lock` leaves an existing applied override in
|
||||
place (it only re-locks when the declared url itself changes), so
|
||||
the forge-declared / applied-deployed split is stable.
|
||||
|
||||
Per-deploy lock flow (two-phase), spread across the deploy subtree's
|
||||
nodes — each phase is its own node, so the queue can show which one is
|
||||
running and a restart resumes at node granularity:
|
||||
|
||||
1. `DeployApply` → `meta::prepare_deploy(name)` runs
|
||||
`nix flake lock --update-input agent-<n>` without
|
||||
committing. Working tree of meta now points the input at
|
||||
`applied/<n>/main` (which the deploy already fast-forwarded to
|
||||
the reviewed PR head).
|
||||
2. The rebuild subgraph `DeployApply` grows into the DAG builds and
|
||||
swaps the container (`AgentWindow` bracing `Prebuild → StopForUpdate
|
||||
→ Swap → RebuildBookkeeping`, plus `Reconcile`). Nix evaluates
|
||||
against the staged lock.
|
||||
3. On success — `FinalizeDeploy` drops the rollback ref, plants
|
||||
`deployed/<id>`, then `meta::finalize_deploy(name, sha, "deployed/
|
||||
<id>")` stages `flake.lock` and commits with
|
||||
`deploy <n> deployed/<id> <sha12>`. Meta's git log gains
|
||||
one entry per successful deploy.
|
||||
4. On failure — the `DeployTail` node runs `meta::abort_deploy()`
|
||||
(`git restore flake.lock`) so the meta history shows only
|
||||
successes; the failure stays as an annotated `failed/<id>`
|
||||
tag in `applied/<n>`. The tail runs on every outcome, so this
|
||||
also covers a hive-c0re restart mid-build: the staged lock is
|
||||
dropped and `applied/main` rolled back from the parked
|
||||
`refs/hyperhive/rollback/<id>`.
|
||||
|
||||
Single-phase variants exist for paths without
|
||||
rollback semantics: `meta::lock_update_for_rebuild(name)` for
|
||||
the manual `↻ R3BU1LD` button (commits if the lock changed)
|
||||
and `meta::lock_update_hyperhive()` for the
|
||||
auto-update flake-rev bump (one shot before per-agent
|
||||
rebuilds, commits if the lock changed).
|
||||
|
||||
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
|
||||
carries `hyperhive_flake`, `dashboard_port`, and the rest of the
|
||||
per-hive config — is the idempotent reconciler called by `spawn`,
|
||||
`destroy`, `rebuild`, and the startup migration. Renders `flake.nix`
|
||||
from the agent list; if it differs from disk, runs
|
||||
`nix flake lock` + commits as `regenerate meta flake` (or
|
||||
`seed meta from N agent(s)` on the very first call).
|
||||
|
||||
The root agent has `/meta` RO-bound inside its container:
|
||||
`git -C /meta log --oneline` is the swarm-wide deploy log,
|
||||
`cat /meta/flake.lock | jq '.nodes["agent-<n>"].locked'`
|
||||
resolves which sha each agent is pinned at right now.
|
||||
Dashboard surfaces the same info as a `deployed:<sha12>` chip
|
||||
per container row.
|
||||
|
||||
## Two repos per agent
|
||||
|
||||
```
|
||||
/var/lib/hyperhive/agents/<name>/config/ proposed — parent mount is RO
|
||||
└── <anything> # any files the submitting
|
||||
# agent wants in the commit.
|
||||
# agent.nix is the
|
||||
# convention entry
|
||||
# point; flake.nix is
|
||||
# tracked boilerplate
|
||||
# (submitting agent doesn't
|
||||
# edit it).
|
||||
|
||||
/var/lib/hyperhive/applied/<name>/ applied — core-only
|
||||
├── .git/ # tag-rich history
|
||||
├── flake.nix # tracked, fixed
|
||||
│ # boilerplate exporting
|
||||
│ # nixosModules.default
|
||||
├── agent.nix # working tree of main
|
||||
└── <other committed files> # also tracked
|
||||
|
||||
/var/lib/hyperhive/meta/ swarm-wide flake — core
|
||||
├── .git/ # one commit per successful
|
||||
│ # deploy
|
||||
├── flake.nix # generated from agent set
|
||||
└── flake.lock # pins each agent's sha
|
||||
```
|
||||
|
||||
Why two physical repos: the submitting agent's `/agents/<n>/config/` is
|
||||
RW — a buggy or hostile agent can `git clean -fdx` its own
|
||||
proposed tree. The applied repo is never bind-mounted (except
|
||||
the read-only `.git` exposure described below) so a destructive
|
||||
move inside the container cannot reach it.
|
||||
|
||||
The container's `--flake` ref is `/var/lib/hyperhive/meta#<name>`
|
||||
(see "Meta flake" above). The agent's own `applied/<n>/flake.nix`
|
||||
is a fixed boilerplate that exports `nixosModules.default =
|
||||
import ./agent.nix`; the meta flake imports that module and
|
||||
wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` /
|
||||
`HIVE_DASHBOARD_PORT`.
|
||||
|
||||
### Tag state machine
|
||||
|
||||
Each deploy leaves a tag on the underlying commit inside the applied
|
||||
repo:
|
||||
|
||||
| Tag | When | Annotated? |
|
||||
|---|---|---|
|
||||
| `deployed/<id>` | rebuild succeeded — `main` ff's here | no |
|
||||
| `failed/<id>` | rebuild failed | yes (body = error) |
|
||||
|
||||
`deployed/0` is planted at first spawn. `applied/main` is always the
|
||||
latest `deployed/*`. A `failed/` tree stays browsable forever — `git log
|
||||
--tags` in the applied repo is the audit trail. A denied or failed config
|
||||
PR carries no extra state on the forge side: the PR stays open, and the
|
||||
submitter pushes again (or closes it) to retry.
|
||||
|
||||
### Dispatch via the job queue
|
||||
|
||||
Long-running approval work — `MergeConfigPr`, `UpdateMetaInputs`,
|
||||
`Spawn` — runs as a DAG on the global job queue
|
||||
(`docs/coordinator.md::Job queue`), submitted by the approval handler
|
||||
rather than run inline:
|
||||
|
||||
| `ApprovalKind` | DAG submitted | source |
|
||||
|---|---|---|
|
||||
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
|
||||
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
|
||||
| `Spawn` | `spawn` (`Create → WriteDropin → Reconcile`) | `approval` |
|
||||
| `InitConfig` | — runs inline (sub-second git seed) | — |
|
||||
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
|
||||
|
||||
The DAG carries the originating `approval_id`, surfaced on the node that
|
||||
owns it — for a deploy that's the `DeployWindow` root, so the dashboard
|
||||
renders one approval card, not four. **Every** queued kind resolves
|
||||
through `actions::resolve_approval_dag` when its DAG settles terminal:
|
||||
the deploy's phases are ordinary queue nodes, so the DAG's own terminal
|
||||
state is the authoritative outcome. That hook fires the matching
|
||||
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
|
||||
terminal tag (verifying the tag actually resolves in the applied repo —
|
||||
a pre-merge rejection plants none), posts the failing build log back to
|
||||
the config PR, and for a spawn runs the post-spawn forge bookkeeping.
|
||||
|
||||
Two visible consequences:
|
||||
|
||||
- **Operator dashboard**: after clicking APPR0VE the work-in-progress
|
||||
shows up on the *rebuild queue* card (`GET /api/jobq/graph`, refetched
|
||||
on every `rebuild_queue_changed` tick), not on the approvals panel
|
||||
(which already moved the row to "approved"). A long meta-update
|
||||
cascade renders as a parent DAG with one child rebuild per affected
|
||||
agent — see `docs/web-ui.md` for the layout.
|
||||
- **Cancellation**: the dashboard's *× cancel* button on a still-queued
|
||||
DAG calls `POST /api/rebuild-queue/{id}/cancel`, which flips it to
|
||||
`Cancelled` before any node runs (and fails the approval row instead
|
||||
of leaving it dangling). Returns `{"cancelled": true}` on success,
|
||||
`{"cancelled": false}` once any node started — terminal states can't
|
||||
be retroactively rewritten.
|
||||
|
||||
The `approval` source + `approval_id` mean a tail-end build failure
|
||||
surfaces back as a failed approval row, not just a silent queue
|
||||
entry. `manual` (dashboard ↻ R3BU1LD) and `auto_update` (boot
|
||||
reconcile) DAGs use the same queue but skip the approval plumbing.
|
||||
|
||||
### Forge mirror
|
||||
|
||||
The bundled `hive-forge` container is mandatory (it deploys with
|
||||
hyperhive), and hive-c0re mirrors every agent's applied repo into a
|
||||
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes `applied/main` plus
|
||||
every tag to `agent-configs/<name>` after each ref mutation:
|
||||
the spawn that seeds `deployed/0`, every successful deploy (which
|
||||
plants `deployed/<id>`) or failed build (`failed/<id>`), and a
|
||||
sweep at startup. Pushes are best-effort — a missing or stopped
|
||||
forge never blocks a deploy.
|
||||
|
||||
Each agent is a **write collaborator on its own** `agent-configs/<name>`
|
||||
repo — so it can push a branch and open a config PR — but not a member
|
||||
of any other agent's, so it can't reach another agent's config through
|
||||
the forge. Branch protection keeps `main` push/merge `core`-only with
|
||||
operator-team approval, so an agent can't fast-forward its own config or
|
||||
self-merge its PR (see the End-to-end flow + #1787). The tokenised push
|
||||
URL is passed inline to `git push`, never written into
|
||||
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
|
||||
agent, and a stored token would leak core's admin credential to an
|
||||
agent.
|
||||
|
||||
The dashboard deep-links into this org — a `config repo` link
|
||||
per container row and a `review PR on forge` link per config-PR
|
||||
approval card. See `docs/web-ui.md`.
|
||||
|
||||
### Submitting agent's view of config repos
|
||||
|
||||
Every parent agent's container has its **direct children's** config
|
||||
repos bind-mounted **read-only** (topology-driven:
|
||||
`hive-c0re/src/lifecycle/host_config.rs` calls `bind_child_agent_dirs`
|
||||
for each entry in
|
||||
`topology::children_of(agent_name)`). It is a copy to *read* a child's
|
||||
current config — not an editing surface.
|
||||
|
||||
An agent with the `approvals` tool group submits a change the same way
|
||||
any other change is made: **clone the child's config repo from the
|
||||
forge into its own state dir, commit on a branch, open a PR**, and let
|
||||
the operator review and approve it. There is deliberately no second,
|
||||
mount-shaped path that reaches the same file without the review.
|
||||
|
||||
Agents holding the `can_manage_top_level_agents` topology role (see
|
||||
`hive-c0re/src/agent_config/topology.rs`) get additional host-side
|
||||
bind mounts via `set_nspawn_flags`:
|
||||
|
||||
- `/var/lib/hyperhive/agents/` → `/agents/` (RW) — all top-level
|
||||
agents' proposed repos (not just direct children).
|
||||
- `/var/lib/hyperhive/applied/` → `/applied/` (RO) — every agent's
|
||||
authoritative applied repo, including `.git`.
|
||||
- `/var/lib/hyperhive/meta/` → `/meta/` (RO) — the swarm-wide
|
||||
deploy flake.
|
||||
|
||||
The root agent holds this role; a sub-manager that only manages a
|
||||
subtree does not, and only has its direct children's config dirs.
|
||||
|
||||
Each proposed repo (`/agents/<n>/config/`) is pre-configured
|
||||
with `applied` as a git remote pointing at
|
||||
`/applied/<n>/.git`. Useful incantations from inside an agent with
|
||||
the full `/applied` mount:
|
||||
|
||||
```sh
|
||||
git -C /agents/<n>/config fetch applied
|
||||
git -C /agents/<n>/config log applied/main --oneline
|
||||
git -C /agents/<n>/config show applied/refs/tags/deployed/<id>
|
||||
git -C /agents/<n>/config show applied/refs/tags/failed/<id> # body = build error
|
||||
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
|
||||
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
|
||||
|
||||
git -C /meta log --oneline # swarm-wide deploy history
|
||||
cat /meta/flake.lock | jq '.nodes | with_entries(select(.key | startswith("agent-")))'
|
||||
```
|
||||
|
||||
The RO binds block push at the kernel level — git plumbing inside the
|
||||
container cannot corrupt either authoritative repo.
|
||||
|
||||
## Startup migrations (older hosts)
|
||||
|
||||
hive-c0re runs a couple of idempotent migrations on every startup so a
|
||||
host set up before the tag-driven-deploy + meta-flake scheme (both
|
||||
described above) converges to it automatically. Each phase is a no-op
|
||||
once already applied:
|
||||
|
||||
- **Tags**: agents from before the tag-driven scheme are tagged
|
||||
`deployed/0` on `main` once. Non-destructive — it doesn't touch live
|
||||
containers, state dirs, or claude creds.
|
||||
- **Meta flake**: rewrites each `applied/<n>/flake.nix` to the
|
||||
module-only boilerplate, wires the `applied` remote in each proposed
|
||||
repo, and bootstraps the meta repo from the current agent list. Set
|
||||
`HIVE_SKIP_META_MIGRATION=1` on the service to defer this phase.
|
||||
|
||||
No state loss in either migration: claude creds, `/state/` notes, the
|
||||
events DB, and both proposed + applied history all survive. The root
|
||||
agent keeps its session; sub-agents stay logged in.
|
||||
|
||||
## The root/bootstrap container is hive-c0re-managed
|
||||
|
||||
The root agent container runs through the **same lifecycle as
|
||||
sub-agents**. On `hive-c0re serve` startup, if `ruth` is missing,
|
||||
hive-c0re creates it. The root agent's flake lives at
|
||||
`/var/lib/hyperhive/applied/ruth/`; its proposed config at
|
||||
`/var/lib/hyperhive/agents/ruth/config/`. The root agent can edit its own
|
||||
`agent.nix` (visible inside the container at `/agents/ruth/config/`)
|
||||
and open a config PR on `agent-configs/ruth` for operator approval,
|
||||
same as any other agent.
|
||||
|
||||
Differences from sub-agents:
|
||||
|
||||
- `flake.nix` extends `hyperhive.nixosConfigurations.ruth`
|
||||
(vs `agent-base`).
|
||||
- Web UI port via `lifecycle::agent_web_port("ruth")` — same
|
||||
FNV-1a hash as every other agent (8100..8999 range).
|
||||
- `set_nspawn_flags` adds two extra binds: `/var/lib/hyperhive/agents`
|
||||
→ `/agents` (RW) so the root agent can edit per-agent proposed repos,
|
||||
and `/var/lib/hyperhive/applied` → `/applied` (RO) so the root agent
|
||||
can `git fetch` deployed/failed/denied tags from any agent's
|
||||
authoritative applied repo (see "Root-agent view of applied" below).
|
||||
- First-deploy spawn bypasses the approval queue (the root agent is
|
||||
required infrastructure).
|
||||
- The root agent's socket is bound by `socket_server::start_manager`,
|
||||
pure transport with no dedicated helpers — it uses the same
|
||||
per-agent runtime dir as any other agent (`/run/hyperhive/agents/ruth/`),
|
||||
not a special manager-only path.
|
||||
|
||||
**Migration note** (for older hosts): drop any `containers.root =
|
||||
{ ... }` block from your host NixOS config. hyperhive creates and
|
||||
updates the root agent itself.
|
||||
|
||||
## Root-agent policy
|
||||
|
||||
The system prompt (`hive-agent/prompts/system.md`, rendered by
|
||||
`hive-agent/src/prompt.rs`) is the **same for every agent**; what
|
||||
varies is which MCP tools are surfaced (gated by tool groups and
|
||||
capabilities in `agent.nix`). There is no `role:manager` block that
|
||||
renders only for the root agent. The root agent's approval-gating
|
||||
behaviour comes from its CLAUDE.md / agent-specific instructions, not
|
||||
the system prompt template.
|
||||
|
||||
## Helper events to the submitting agent
|
||||
|
||||
`Coordinator::notify_submitter(approval_id, &HelperEvent)` routes the
|
||||
event to the agent that originally submitted the approval (looked up from
|
||||
the `submitter` column on the `approvals` table). The harness delivers it
|
||||
as a regular `system` inbox message so it drives a normal claude turn.
|
||||
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
|
||||
**every** approval kind's terminal state, `Spawn` included. A
|
||||
"FYI, check when convenient" event doesn't need a message — those go
|
||||
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct
|
||||
live dial of the target agent's in-container todo socket (same
|
||||
`UpsertTodo` request in-container producers use); `finish_approval` fires
|
||||
one of these too for `InitConfig`/`Spawn`/`MergeConfigPr`, *in addition to*
|
||||
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
|
||||
approval rows that predate the submitter column fall back to the
|
||||
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
||||
|
||||
- `ApprovalResolved { id, agent, commit_ref, status, note }` —
|
||||
fired by `actions::approve` + `actions::deny` whenever an
|
||||
approval transitions to its terminal state.
|
||||
- `ContainerCrash { agent, note }` — `crash_watch`: a previously-
|
||||
running container went away with no operator-initiated transient
|
||||
state (Stopping / Restarting / Destroying / Rebuilding) AND no
|
||||
such transient was cleared in the last 30s (`RECENT_TRANSIENT_GRACE`
|
||||
tombstone, three `POLL_INTERVAL`s — closes the race where a
|
||||
lifecycle op finishes between two crash-watch polls and the
|
||||
container shows briefly as "stopped without transient" before
|
||||
the next start). The root agent can `start` it again or escalate.
|
||||
- `NeedsUpdate { agent }` — sub-agent's recorded flake rev is
|
||||
stale. The root agent calls `update(name)` to rebuild — idempotent,
|
||||
no approval required.
|
||||
|
||||
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
|
||||
`Destroyed`, `NeedsLogin`, `LoggedIn`, `ConfigReady` — are "FYI, check
|
||||
when convenient" events with no reason to drive an immediate turn, so
|
||||
they deliver via `push_todo`/`push_todo_submitter` (see above) instead
|
||||
of `HelperEvent`: an `agent_todo_socket` push instead of a broker
|
||||
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
|
||||
and a single free-text `summary` (`rebuilt_todo_summary` renders
|
||||
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
|
||||
|
||||
Optional `sha` field on `ApprovalResolved` carries the canonical
|
||||
hive-c0re-vouched commit sha. Optional `tag` carries the deploy
|
||||
bookkeeping tag — `deployed/<id>` on a successful build or
|
||||
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
|
||||
Both fields are `Option`: `None` on the paths that don't deploy a new
|
||||
commit (spawn / init_config / meta-update / deny, and the auto-update
|
||||
sweep's `job_queue::templates::rebuild` reapplying the existing main,
|
||||
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
|
||||
`git show <sha>` against `/applied/<n>/.git` inside the
|
||||
bootstrap container yields the exact tree that was referenced.
|
||||
|
||||
To add a new lifecycle notice: if it needs to drive an immediate turn
|
||||
(something genuinely urgent, like `ContainerCrash`), add a
|
||||
`HelperEvent` variant + call sites + update `prompts/system.md`'s
|
||||
message-event list. If it's "FYI, check when convenient," call
|
||||
`push_todo`/`push_todo_submitter` directly instead — no new wire type
|
||||
needed.
|
||||
|
||||
## Auto-update on startup
|
||||
|
||||
`hive-c0re serve` runs `auto_update::run` in a background task right
|
||||
after opening the coordinator. It enumerates managed containers and
|
||||
rebuilds any whose recorded hyperhive rev differs from the current
|
||||
one — sub-agents and the root agent go through the same
|
||||
`job_queue::templates::rebuild` DAG.
|
||||
|
||||
"Rev" = canonical filesystem path of `cfg.hyperhiveFlake`. Marker
|
||||
file: `/var/lib/hyperhive/applied/.<name>.hyperhive-rev`. If the
|
||||
flake input has no canonical path (e.g. a `github:` URL),
|
||||
auto-update is a no-op — rebuild manually.
|
||||
|
||||
The dashboard surfaces pending updates per agent: a clickable
|
||||
"needs update ↻" badge appears whenever the marker differs from
|
||||
current rev. The badge POSTs `/api/rebuild/<name>`, which inserts the
|
||||
same `job_queue::templates::rebuild` DAG so manual triggers and the
|
||||
startup scan can't drift. When at least one container is stale, a
|
||||
top-level `↻ UPD4TE 4LL` button appears that loops over every
|
||||
stale container.
|
||||
Loading…
Reference in a new issue