Watch
0
0
Fork
You've already forked hyperhive
0

docs: config changes are operator merges on the forge

Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.

Refs #4850
This commit is contained in:
atlas 2026-10-02 22:32:17 +02:00
commit a88ed9f24e
14 changed files with 195 additions and 396 deletions

View file

@ -31,7 +31,7 @@ architecture change.
| **identity** | every agent is a swarm-wide principal — SSO subject, forge user, matrix account, secret-store cert identity — addressable as `name@hive.domain` | | **identity** | every agent is a swarm-wide principal — SSO subject, forge user, matrix account, secret-store cert identity — addressable as `name@hive.domain` |
| **secrets** | one OpenBao store; the operator places one mTLS identity per host, and everything else — every agent's credentials included — is fetched from the store under an identity rather than copied by hand | | **secrets** | one OpenBao store; the operator places one mTLS identity per host, and everything else — every agent's credentials included — is fetched from the store under an identity rather than copied by hand |
| **shared services** | one forge, homeserver, SSO, message queue and metrics/logs stack per swarm, each on whichever host you put it | | **shared services** | one forge, homeserver, SSO, message queue and metrics/logs stack per swarm, each on whichever host you put it |
| **config** | git: an agent proposes, the operator approves, the deploy lands as a `deployed/<id>` tag | | **config** | git: an agent opens a PR on its config repo, an operator merges it on the forge, and the merge deploys |
| **runtime** | `claude --print` by default; any [ACP](https://agentclientprotocol.com) agent (e.g. opencode) per agent with `services.hyperhive.agent.runtime = "acp"` | | **runtime** | `claude --print` by default; any [ACP](https://agentclientprotocol.com) agent (e.g. opencode) per agent with `services.hyperhive.agent.runtime = "acp"` |
| **substrate** | each hive runs one `nixos-container` per agent, under an unprivileged daemon with a tiny socket-activated helper for the few root ops | | **substrate** | each hive runs one `nixos-container` per agent, under an unprivileged daemon with a tiny socket-activated helper for the few root ops |
| **watching** | the swarm UI (hives, agents with live terminals, jobs, a cross-repo issue report), Grafana over OTEL, and a per-hive dashboard for host-level detail | | **watching** | the swarm UI (hives, agents with live terminals, jobs, a cross-repo issue report), Grafana over OTEL, and a per-hive dashboard for host-level detail |

View file

@ -1,38 +1,32 @@
# Approvals + helper events # Approvals + helper events
The approval queue is hyperhive's pivot: nothing that changes the The approval queue is where an agent asks the operator for something it
shape of an agent (its config, whether it exists) happens without an can't do itself: a scheduled prompt, or (legacy rows only) a meta-input
operator selection. The submitting agent — any agent with the `approvals` bump. Config changes don't go through this queue: they're pull requests
tool group, which manages the config of its **direct children** (the on the agent's config repo, which an operator merges on the forge, and
root agent for top-level agents; a sub-manager for its own subtree) — is that merge deploys them. The submitting agent — any agent with the
the policy gate in front of that queue; helper events are how it stays `approvals` tool group, which manages the config of its **direct
informed about what happens after a decision lands. children** (the root agent for top-level agents; a sub-manager for its
own subtree) — is the policy gate in front of both; helper events are how
it stays informed about what happens after a decision lands.
## For operators ## For operators
Every add/remove/change to an agent lands on your dashboard's Y3R - **Config change** — an agent proposed a change to another agent's
C4LL tab (or `hivectl approvals pending` / `approve <id>` from the config (or its own, via a sub-manager) as a pull request on
CLI) before it takes effect. What you'll see, and what to do with it: `agent-configs/<agent>`. Review it on the forge like any other PR, and
merge it there to deploy it. See [Config changes](#config-changes).
- **Config change** (`MergeConfigPr`) — an agent proposed a change to
another agent's config (or its own, via a sub-manager) as a forge
pull request. Review the diff on the forge — the dashboard card
links straight to it, same as reviewing any other PR. Approving
triggers the deploy automatically: hive-c0re re-verifies the PR
hasn't moved since you looked at it, evaluates it (a dry run,
nothing applied yet), merges it, and rebuilds the container. If
anything in that chain fails, the change rolls back automatically —
the agent stays on its last-good config, no recovery action needed
from you.
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
one or more Nix flake inputs (or all of them). Approving runs the
update and commits the lock change; it doesn't rebuild anything by
itself.
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule - **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
a message to one or more inboxes at a future time. You can also add a message to one or more inboxes at a future time. It lands on your
schedules yourself directly from the SCH3DUL3S tab, which skips this dashboard's Y3R C4LL tab (or `hivectl approvals pending` /
approval step entirely — the gate here is specifically for an `approve <id>` from the CLI). You can also add schedules yourself
*agent* asking to schedule something, not for you doing it. directly from the SCH3DUL3S tab, which skips this approval step
entirely — the gate here is specifically for an *agent* asking to
schedule something, not for you doing it.
- **Meta/flake update** (`UpdateMetaInputs`) — legacy: nothing queues
this kind, but an existing row still reads back and you can approve it.
Approving runs the update and commits the lock change; it doesn't
rebuild anything by itself.
Don't want to approve something? **Deny it** (`DENY` on the dashboard Don't want to approve something? **Deny it** (`DENY` on the dashboard
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
@ -40,18 +34,16 @@ submitting agent is always notified that the operator denied their request; what
optional is only the reason text, which you can add on the dashboard's optional is only the reason text, which you can add on the dashboard's
prompt (cancelling that prompt aborts the whole deny, not just the prompt (cancelling that prompt aborts the whole deny, not just the
reason) but not from the CLI. Denying is final: a denied approval reason) but not from the CLI. Denying is final: a denied approval
can't be re-approved later, the agent has to submit a fresh one (a new can't be re-approved later, the agent has to submit a fresh request.
PR, a new request).
Everything below this point is the implementation detail behind that Everything below this point is the implementation detail behind that
flow. flow.
## End-to-end approval flow ## Config changes
Config changes flow through a **forge pull request** on the agent's Config changes flow through a **forge pull request** on the agent's
`agent-configs/<name>` repo — the same surface agents use for code PRs. `agent-configs/<name>` repo — the same surface agents use for code PRs.
No bespoke MCP tool exists for config changes: opening the PR IS the No MCP tool exists for config changes: opening the PR IS the request.
request.
1. The submitting agent (the child's parent, holding the `approvals` 1. The submitting agent (the child's parent, holding the `approvals`
tool group) **clones** `agent-configs/<name>`, edits it there (any tool group) **clones** `agent-configs/<name>`, edits it there (any
@ -59,89 +51,40 @@ request.
with its own git identity, and pushes a branch + opens a PR with with its own git identity, and pushes a branch + opens a PR with
`hive-forge` — the same way it would change any other repo. `hive-forge` — the same way it would change any other repo.
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
config, not the tree to edit: authoring in place there produces no PR config, not the tree to edit: authoring in place there produces no PR.
and no approval. (it's currently mounted read-write, which is a (it's currently mounted read-write, which is a defect tracked
defect tracked separately, not an authoring path.) separately, not an authoring path.)
Branch protection (the agent isn't on the `main` push allowlist; Branch protection (the agent isn't on the `main` push allowlist;
merge allowlist = `core` user + `operators` team; approvals allowlist merge and approvals allowlist = `operators` team; see "Forge mirror"
= `operators` team; see "Forge mirror" below) makes the agent a write below) makes the agent a write collaborator that **can't merge its own
collaborator that **can't merge its own config PR**. config PR**.
2. hive-c0re's `/webhook/config-pr` endpoint receives the Forgejo 2. An operator reviews the PR **on the forge** (native diff, threaded
`pull_request` event (opened / synchronized / reopened) and queues a comments, CI status) and merges it there. Merging needs membership of
`MergeConfigPr` approval; a poll fallback catches any missed webhook. the `operators` team in `agent-configs`. The team starts empty: add
The approval row stores the PR **number** (`commit_ref`) and the PR each operator by hand in the forge UI.
**head sha at queue time** (`fetched_sha` — the "reviewed" sha). If 3. swarm-controller gets the `agent-configs` org's `pull_request`
the PR head later moves, a fresh approval pinned to the new head supersedes
the stale one, so the operator always reviews what will
actually deploy.
3. The operator reviews the PR **on the forge** (native diff, threaded
comments, CI status) and sees a matching card on the dashboard with a
"review PR on forge" deep link. They select ◆ APPR0VE (or
`hivectl approvals approve <id>` on the CLI) once satisfied.
4. On approve, a deploy DAG runs three phases under a resource-holding
`DeployWindow` root (see *Queue templates* below):
- `MergeVerify` re-reads the live PR head and **aborts if it drifted**
from the reviewed `fetched_sha` (the submitter must push again,
which queues a fresh approval); then fetches that head into the
applied repo and **eval-verifies** it — a flake eval on a throwaway
checkout. This is the trust gate: it relies on c0re's own eval, not
on any in-repo (agent-forgeable) signal like a CI status. Nothing is
mutated in this phase, so a rejection here leaves the forge and the
applied repo exactly as they were.
- `DeployApply` parks the pre-merge `applied/main` in
`refs/hyperhive/rollback/<approval-id>`, then fast-forward-merges
the reviewed head to the forge config repo's `main` (this IS the
merge — a `core`-authenticated ff-merge pinned to the reviewed sha,
so a moved PR head can't substitute bytes), and runs the deploy
proper (`deploy_applied_target`): ff `applied/main`, two-phase meta
deploy, container rebuild. On success it drops the rollback ref and
plants `deployed/<id>`.
- `DeployTail` runs on **every** outcome, including a cancel-cascade.
If the rollback ref survived, the deploy never confirmed good: it
rolls `applied/main` back, resyncs the working tree, and aborts the
staged meta lock, so the agent stays on its last-good tree. Then it
mirrors the config repo (and its new deploy tag) to the forge.
The rollback state lives in a **git ref, not a local variable**, on
purpose: hive-c0re can restart between the apply and the tail, and the
tail still has to know what to undo when it does.
5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the
**submitting agent's** inbox via `notify_submitter`, carrying both the
canonical sha and the terminal tag (the approval row carries a
`submitter` column recording the agent the change is for).
### Operator merge in the forge UI
An operator can also merge a config PR straight in the Forgejo UI. That
deploys the merged commit on the hive that runs the agent:
1. swarm-controller gets the `agent-configs` org's `pull_request`
delivery. A `closed` event with `merged: true` and base branch `main` delivery. A `closed` event with `merged: true` and base branch `main`
names the commit in `merge_commit_sha`. names the commit in `merge_commit_sha`.
2. The controller looks up which hive's wanted state places the agent 4. The controller looks up which hive's wanted state places the agent
and publishes a deploy request carrying that commit on that hive's and publishes a deploy request carrying that commit on that hive's
deploy subject. With no such hive, or more than one, it deploys deploy subject. With no such hive, or more than one, it deploys
nothing and logs a warning naming them. nothing and logs a warning naming them.
3. If the hive's `applied/main` already is that commit — a dashboard 5. If the hive's `applied/main` already is that commit, it does nothing.
approval merges and deploys its own PR — it does nothing. Otherwise Otherwise it fetches the config repo's `main`, requires the commit to
it fetches the config repo's `main`, requires the commit to descend descend from `applied/main`, fast-forwards `applied/main` to it, and
from `applied/main`, fast-forwards `applied/main` to it, and queues a queues a rebuild. A failure before the rebuild (the fetch, or a commit
rebuild. A failure before the rebuild (the fetch, or a commit that that doesn't descend) deploys nothing and posts the error as a comment
doesn't descend) deploys nothing and posts the error as a comment on on the merged PR. An agent with no container on that hive yet ignores
the merged PR. An agent with no container on that hive yet ignores
the commit; its first deploy builds what it seeds. the commit; its first deploy builds what it seeds.
This path runs **no eval-verify**. A merged config that doesn't This path runs **no eval-verify**. A merged config that doesn't
evaluate or build fails the rebuild, and `applied/main` stays at that evaluate or build fails the rebuild, and `applied/main` stays at that
commit, so the agent's rebuilds keep failing until a fix merges. A commit, so the agent's rebuilds keep failing until a fix merges. A
failed rebuild shows on the hive like any other and isn't commented on failed rebuild shows on the hive like any other and isn't commented on
the PR. the PR. A merge whose webhook delivery never reaches the controller
deploys nothing.
Merging needs membership of the `operators` team in `agent-configs`. ## Approval queue
The team starts empty: add each operator by hand in the forge UI. A
merge whose webhook delivery never reaches the controller deploys
nothing. The hive's own poll cancels the dashboard card for that PR
with the note `PR merged/closed outside the approval`.
### Withdrawing a pending approval ### Withdrawing a pending approval
@ -166,31 +109,16 @@ and no hive can originate an agent. The swarm controller's
`agent.nix` template, then asks the target hive to deploy it. `agent.nix` template, then asks the target hive to deploy it.
Changing what the template seeded isn't a special case: like every Changing what the template seeded isn't a special case: like every
later change, it's a PR on that config repo (`MergeConfigPr`), made later change, it's a PR on that config repo, made from a clone and
from a clone, reviewed and approved by the operator. The PR flow is merged by an operator on the forge (see [Config changes](#config-changes)).
the one path — an operator can equally drive both steps herself
through the web UI or the forge.
### Approval kinds (wire shapes) ### Approval kinds (wire shapes)
`ApprovalKind` carries three variants; each maps to a different `ApprovalKind` carries two variants; each maps to a different
`commit_ref` encoding because `ApprovalKind` overloads that field as `commit_ref` encoding because `ApprovalKind` overloads that field as
the kind-specific payload carrier. the kind-specific payload carrier.
<!-- vale write-good.Passive = NO --> <!-- vale write-good.Passive = NO -->
- `MergeConfigPr` — the config-change flow. Triggered automatically:
when an agent opens (or force-pushes) a PR on its
`agent-configs/<agent>` forge repo, hive-c0re's `/webhook/config-pr`
endpoint receives the Forgejo pull_request event and queues this
approval row. No MCP tool call needed — the forge PR IS the request.
`commit_ref` stores the **PR number** (decimal), and `fetched_sha` is
the PR **head sha at queue time** (the "reviewed" sha). On approve,
the deploy DAG's `MergeVerify` phase re-reads the live PR head and
aborts if it drifted from `fetched_sha` (submitter must push again to
re-trigger), then fetches that head into the applied repo and
eval-verifies it; `DeployApply` fast-forward-merges the forge config
repo's `main` to it (the merge) and runs `deploy_applied_target`;
`DeployTail` compensates on failure. Never a first spawn.
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs - `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs, array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
etc.). hive-c0re sets the `agent` field to the requesting root agent. etc.). hive-c0re sets the `agent` field to the requesting root agent.
@ -303,55 +231,26 @@ declares one flake input per agent (`agent-<n>.url =
Containers run against `--flake /var/lib/hyperhive/meta#<n>`. Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
The declared input url is the agent's **forge config repo** (the The declared input url is the agent's **forge config repo** (the
same `agent-configs/<n>` the config-PR flow lands approved changes same `agent-configs/<n>` config PRs merge into), so the meta flake
on), so the meta flake references a reviewable, reproducible source references a reviewable, reproducible source rather than a local
rather than a local checkout. hive-c0re authenticates that checkout. hive-c0re authenticates that `git+http` fetch via a git
`git+http` fetch via a git credential helper that reads the live credential helper that reads the live forge-core token — no token in
forge-core token — no token in the url or the lock. The deploy and the url or the lock. The rebuild paths, however, do **not** re-lock
manual-rebuild paths, however, do **not** re-lock from the forge: from the forge: they `--override-input agent-<n>
they `--override-input agent-<n>
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
that `verify_commit` gated and `applied/<n>/main` was `applied/<n>/main` was fast-forwarded to. That keeps a rebuild
fast-forwarded to. That keeps a deploy/rebuild reproducible and reproducible and independent of forge reachability — rebuilds fire on
independent of forge reachability — rebuilds fire on crash-restart crash-restart and meta bumps, not just config merges — while the
and meta bumps, not just config PRs — while the declared url stays declared url stays the forge. `sync_agents` re-renders + re-locks the
the forge. `sync_agents` re-renders + re-locks the persistent input; persistent input; a plain `nix flake lock` leaves an existing applied
a plain `nix flake lock` leaves an existing applied override in override in place (it only re-locks when the declared url itself
place (it only re-locks when the declared url itself changes), so changes), so the forge-declared / applied-deployed split is stable.
the forge-declared / applied-deployed split is stable.
Per-deploy lock flow (two-phase), spread across the deploy subtree's Lock updates are single-phase and commit when the lock changed:
nodes — each phase is its own node, so the queue can show which one is `meta::lock_update_for_rebuild(name)` relocks one agent's input for a
running and a restart resumes at node granularity: relocking rebuild (the manual `↻ R3BU1LD` button, and the rebuild a
merged config commit queues), and `meta::lock_update_hyperhive()` is the
1. `DeployApply` → `meta::prepare_deploy(name)` runs autoupdate flake-rev bump (one shot before per-agent rebuilds).
`nix flake lock --update-input agent-<n>` without
committing. Working tree of meta now points the input at
`applied/<n>/main` (which the deploy already fast-forwarded to
the reviewed PR head).
2. The rebuild subgraph `DeployApply` grows into the DAG builds and
swaps the container (`AgentWindow` bracing `Prebuild → StopForUpdate
→ Swap → RebuildBookkeeping`, plus `Reconcile`). Nix evaluates
against the staged lock.
3. On success — `FinalizeDeploy` drops the rollback ref, plants
`deployed/<id>`, then `meta::finalize_deploy(name, sha, "deployed/
<id>")` stages `flake.lock` and commits with
`deploy <n> deployed/<id> <sha12>`. Meta's git log gains
one entry per successful deploy.
4. On failure — the `DeployTail` node runs `meta::abort_deploy()`
(`git restore flake.lock`) so the meta history shows only
successes; the failure stays as an annotated `failed/<id>`
tag in `applied/<n>`. The tail runs on every outcome, so this
also covers a hive-c0re restart mid-build: the staged lock is
dropped and `applied/main` rolled back from the parked
`refs/hyperhive/rollback/<id>`.
Single-phase variants exist for paths without
rollback semantics: `meta::lock_update_for_rebuild(name)` for
the manual `↻ R3BU1LD` button (commits if the lock changed)
and `meta::lock_update_hyperhive()` for the
autoupdate flake-rev bump (one shot before per-agent
rebuilds, commits if the lock changed).
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive` `meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
carries `hyperhive_flake`, `dashboard_port`, and the rest of the carries `hyperhive_flake`, `dashboard_port`, and the rest of the
@ -390,8 +289,8 @@ per container row.
└── <other committed files> # also tracked └── <other committed files> # also tracked
/var/lib/hyperhive/meta/ swarm-wide flake — core /var/lib/hyperhive/meta/ swarm-wide flake — core
├── .git/ # one commit per successful ├── .git/ # one commit per lock
│ # deploy │ # change
├── flake.nix # generated from agent set ├── flake.nix # generated from agent set
└── flake.lock # pins each agent's sha └── flake.lock # pins each agent's sha
``` ```
@ -411,43 +310,27 @@ wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` /
### Tag state machine ### Tag state machine
Each deploy leaves a tag on the underlying commit inside the applied hive-c0re plants `deployed/0` on the seed commit at first spawn. A
repo: merged config commit that deploys plants no tag: the merged PR on the
forge and meta's lock commits record it. A config PR nobody merges
| Tag | When | Annotated? | carries no extra state on the forge side: the PR stays open, and
|---|---|---| the submitter pushes again (or closes it) to retry.
| `deployed/<id>` | rebuild succeeded — `main` ff's here | no |
| `failed/<id>` | rebuild failed | yes (body = error) |
hive-c0re plants `deployed/0` at first spawn. `applied/main` is always the
latest `deployed/*`. A `failed/` tree stays browsable forever — `git log
--tags` in the applied repo is the audit trail. A denied or failed config
PR carries no extra state on the forge side: the PR stays open, and the
submitter pushes again (or closes it) to retry.
### Dispatch via the job queue ### Dispatch via the job queue
Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs` Long-running approval work — `UpdateMetaInputs` — runs as a DAG on the
— runs as a DAG on the global job queue global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted
(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler by the approval handler rather than run inline:
rather than run inline:
| `ApprovalKind` | DAG submitted | source | | `ApprovalKind` | DAG submitted | source |
|---|---|---| |---|---|---|
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` | | `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
| `SchedulePrompt` | — runs inline (single sqlite insert) | — | | `SchedulePrompt` | — runs inline (single sqlite insert) | — |
The DAG carries the originating `approval_id`, surfaced on the node that The DAG carries the originating `approval_id`. **Every** queued kind
owns it — for a deploy that's the `DeployWindow` root, so the dashboard resolves through `actions::resolve_approval_dag` when its DAG settles
renders one approval card, not four. **Every** queued kind resolves terminal: the DAG's own terminal state is the authoritative outcome.
through `actions::resolve_approval_dag` when its DAG settles terminal: That hook fires the matching `HelperEvent::*` via `finish_approval`.
the deploy's phases are ordinary queue nodes, so the DAG's own terminal
state is the authoritative outcome. That hook fires the matching
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
terminal tag (verifying the tag actually resolves in the applied repo —
a pre-merge rejection plants none), posts the failing build log back to
the config PR.
Two visible consequences: Two visible consequences:
@ -474,30 +357,27 @@ reconcile) DAGs use the same queue but skip the approval plumbing.
The bundled `hive-forge` container runs on the swarm's forge host The bundled `hive-forge` container runs on the swarm's forge host
(`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)), (`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)),
and hive-c0re mirrors every agent's applied repo into a and hive-c0re mirrors every agent's applied repo into a
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes `applied/main` plus private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes
every tag to `agent-configs/<name>` after each ref mutation: every tag, then `applied/main`, to `agent-configs/<name>` on every
the spawn that seeds `deployed/0`, every successful deploy (which startup sweep and every rebuild. Forge `main` is branch-protected, so
plants `deployed/<id>`) or failed build (`failed/<id>`), and a the forge routinely refuses a push of an established `main`, which
sweep at startup. Pushes are best-effort — a missing or stopped hive-c0re expects. Pushes are best-effort — a missing or stopped forge never
forge never blocks a deploy. blocks a deploy.
Each agent is a **write collaborator on its own** `agent-configs/<name>` Each agent is a **write collaborator on its own** `agent-configs/<name>`
repo — so it can push a branch and open a config PR — but not a member repo — so it can push a branch and open a config PR — but not a member
of any other agent's, so it can't reach another agent's config through of any other agent's, so it can't reach another agent's config through
the forge. Branch protection keeps the agent off the `main` push the forge. Branch protection keeps the agent off the `main` push
allowlist and whitelists merging to the `core` user and the `operators` allowlist and allowlists merging to the `operators` team only, so an
team, so an agent can't fast-forward its own config or self-merge its agent can't fast-forward its own config or self-merge its PR (see
PR (see the End-to-end flow and [Config changes](#config-changes) above). hive-c0re passes the tokenised
[Operator merge in the forge UI](#operator-merge-in-the-forge-ui) above). push URL inline to `git push`, never writing it into
hive-c0re passes the tokenised push
URL inline to `git push`, never writing it into
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root `applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
agent, and a stored token would leak core's admin credential to an agent, and a stored token would leak core's admin credential to an
agent. agent.
The dashboard deep-links into this org — a `config repo` link The dashboard deep-links into this org — a `config repo` link
per container row and a `review PR on forge` link per config-PR per container row. See `docs/web-ui/dashboard.md`.
approval card. See `docs/web-ui/dashboard.md`.
### Submitting agent's view of config repos ### Submitting agent's view of config repos
@ -548,8 +428,7 @@ the full `/applied` mount:
```sh ```sh
git -C /agents/<n>/config fetch applied git -C /agents/<n>/config fetch applied
git -C /agents/<n>/config log applied/main --oneline git -C /agents/<n>/config log applied/main --oneline
git -C /agents/<n>/config show applied/refs/tags/deployed/<id> git -C /agents/<n>/config show applied/refs/tags/deployed/0 # the seed commit
git -C /agents/<n>/config show applied/refs/tags/failed/<id> # body = build error
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
@ -631,11 +510,9 @@ as a regular `system` inbox message so it drives a normal claude turn.
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for `finish_approval` fires an `ApprovalResolved` HelperEvent this way for
**every** approval kind's terminal state. A **every** approval kind's terminal state. A
"FYI, check when convenient" event doesn't need a message — those go "FYI, check when convenient" event doesn't need a message — those go
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct through `Coordinator::push_todo` instead, a direct live dial of the target
live dial of the target agent's in-container todo socket (same agent's in-container todo socket (same `UpsertTodo` request in-container
`UpsertTodo` request in-container producers use); `finish_approval` fires producers use). Legacy
one of these too for `MergeConfigPr`, *in addition to*
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
approval rows that predate the submitter column fall back to the approval rows that predate the submitter column fall back to the
root agent. Variants (`hive_sh4re::manager::HelperEvent`): root agent. Variants (`hive_sh4re::manager::HelperEvent`):
@ -657,28 +534,17 @@ root agent. Variants (`hive_sh4re::manager::HelperEvent`):
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`, The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
`Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check `Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check
when convenient" events with no reason to drive an immediate turn, so when convenient" events with no reason to drive an immediate turn, so
they deliver via `push_todo`/`push_todo_submitter` (see above) instead they deliver via `push_todo` (see above) instead
of `HelperEvent`: an `agent_todo_socket` push instead of a broker of `HelperEvent`: an `agent_todo_socket` push instead of a broker
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup, message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
and a single free-text `summary` (`rebuilt_todo_summary` renders and a single free-text `summary` (`rebuilt_todo_summary` renders
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string). `Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
Optional `sha` field on `ApprovalResolved` carries the canonical
hive-c0re-vouched commit sha. Optional `tag` carries the deploy
bookkeeping tag — `deployed/<id>` on a successful build or
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
Both fields are `Option`: `None` on the paths that don't deploy a new
commit (meta-update / deny, and the autoupdate
sweep's `job_queue::templates::rebuild` reapplying the existing main,
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
`git show <sha>` against `/applied/<n>/.git` inside the
bootstrap container yields the exact tree the sha referenced.
To add a new lifecycle notice: if it needs to drive an immediate turn To add a new lifecycle notice: if it needs to drive an immediate turn
(something genuinely urgent, like `ContainerCrash`), add a (something genuinely urgent, like `ContainerCrash`), add a
`HelperEvent` variant + call sites + update `prompts/system.md`'s `HelperEvent` variant + call sites + update `prompts/system.md`'s
message-event list. If it's "FYI, check when convenient," call message-event list. If it's "FYI, check when convenient," call
`push_todo`/`push_todo_submitter` directly instead — no new wire type `push_todo` directly instead — no new wire type
needed. needed.
## Autoupdate on startup ## Autoupdate on startup

View file

@ -56,8 +56,8 @@ power-intent registry:
per-agent store — see [`/harness/` contents per-agent store — see [`/harness/` contents
below](#state-dirs-per-agent) for where reminders (and todos) below](#state-dirs-per-agent) for where reminders (and todos)
live. live.
- `approvals` — the queue. `agent / kind (merge_config_pr | spawn | - `approvals` — the queue. `agent / kind (update_meta_inputs |
update_meta_inputs | schedule_prompt) / schedule_prompt) /
commit_ref / requested_at / status / resolved_at / note`. commit_ref / requested_at / status / resolved_at / note`.
- `scheduled_prompts` — recurring + one-shot prompt queue. - `scheduled_prompts` — recurring + one-shot prompt queue.
`owner / body / interval_seconds (NULL = one-shot) / `owner / body / interval_seconds (NULL = one-shot) /

View file

@ -127,8 +127,9 @@ The controller provisions iris's identity, forge user and config repo
and start the container. It returns once the scheduler queues the job — watch the and start the container. It returns once the scheduler queues the job — watch the
swarm UI's job view for progress. swarm UI's job view for progress.
Later config changes are PRs on `agent-configs/iris`, approved by you. → Later config changes are PRs on `agent-configs/iris`; you merge them on the
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md) forge, and the merge deploys them. →
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md#config-changes)
## Optional · Lock the hive dashboard ## Optional · Lock the hive dashboard

View file

@ -67,22 +67,15 @@ Two things live in the `agent-configs` Forgejo organization:
- A config repo per agent (`agent-configs/<name>`). The - A config repo per agent (`agent-configs/<name>`). The
agent is a **write collaborator on its own** repo — it can push agent is a **write collaborator on its own** repo — it can push
config-change branches and open config PRs (Forgejo `pull_request` config-change branches and open config PRs — but `main` is
webhook at `/webhook/config-pr` queues a `MergeConfigPr` approval; branch-protected by swarm-controller: the merge and approval
`hive-c0re/src/forge/config_pr_poll.rs` re-scans every 5 minutes as a allowlists are the `operators` team, and the agent can neither push
fault-tolerance backstop) — but `main` directly nor self-merge. An operator's merge in the Forgejo UI
`main` is branch-protected: the merge whitelist is the `core` user deploys the merged commit (see
(hive-c0re's merge of an approved `MergeConfigPr`) and the `operators` [approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes)).
team (an operator merging in the Forgejo UI, which deploys the merged hive-c0re never force-pushes: the `push_config` mirror pushes the
commit — see add-only status tags and `main` without force, and treats a refused
[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui)), `main` push as expected.
the approval whitelist is the `operators` team, and the agent can neither
push `main` directly nor self-merge. hive-c0re's own merge is
fast-forward-only, and hive-c0re never force-pushes (the
`push_config` mirror pushes `main` + the add-only
status tags without force, and treats a non-fast-forward rejection of
`main` after a rolled-back deploy as expected — the forge keeps the
approved history, the `failed/<id>` tag records the divergence).
Repos stay private, so an agent can't read another Repos stay private, so an agent can't read another
agent's config. (Agents remain read-only collaborators on `core/meta`.) agent's config. (Agents remain read-only collaborators on `core/meta`.)
hive-c0re also references this repo as the agent's **persistent meta hive-c0re also references this repo as the agent's **persistent meta

View file

@ -25,7 +25,7 @@ You rarely switch it on yourself. `gateway.enable` defaults to off, and every mo
| URL | upstream | when | | URL | upstream | when |
| --- | --- | --- | | --- | --- | --- |
| `<hive>/` | dashboard dist (static, from `servedFrontend`) | always | | `<hive>/` | dashboard dist (static, from `servedFrontend`) | always |
| `<hive>/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always | | `<hive>/api/`, `/health/` | hive-c0re (`7000`) | always |
| `<hive>/api/docs/` | themed Swagger UI dist (static) | always | | `<hive>/api/docs/` | themed Swagger UI dist (static) | always |
| `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) | | `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) |
| `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` | | `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` |
@ -130,7 +130,7 @@ services.hyperhive.gateway.auth = {
`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history. `hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history.
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`. **What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first. A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first.
@ -261,10 +261,9 @@ Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }`
#### Dashboard: path-based routing (not Accept-header) #### Dashboard: path-based routing (not Accept-header)
hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header: hive-c0re serves exactly two prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live. - `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live.
- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded.
- `location /health/` → hive-c0re: liveness and readiness. - `location /health/` → hive-c0re: liveness and readiness.
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`. - `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`.

View file

@ -43,22 +43,18 @@ because there is no malformed spec to reject.
Nix-heavy — hold one of the `buildSlots` permits for the node's duration: Nix-heavy — hold one of the `buildSlots` permits for the node's duration:
| Node | Wraps | | Node | Wraps |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead | | `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead |
| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node | | `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node |
| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta | | `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta |
| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion | | `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion |
| `DeployWindow` | resource-holding root of the merge-config-PR deploy subtree — declares the build slot, the lease and the meta window, then completes immediately so its children run under them (see _Approvals_ below) |
| `DeployApply` | the deploy's irreversible half: ff-merge the reviewed PR head, two-phase meta deploy, container rebuild |
Cheap — no build slot: Cheap — no build slot:
<!-- vale write-good.Passive = NO --> <!-- vale write-good.Passive = NO -->
| Node | Behavior | | Node | Behavior |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MergeVerify` | the deploy's pre-merge gate — PR-head drift check, fetch, `verify_commit` eval. Mutates nothing, so a rejection here needs no compensation |
| `DeployTail` | the deploy's `AfterAny` compensation + bookkeeping tail: (1) rolls `applied/main` back from the parked `refs/hyperhive/rollback/<id>` and aborts the staged meta lock when the deploy never confirmed good; (2) mirrors whichever deploy tag landed to the forge config repo, always, best-effort; (3) posts the failing build log back onto the config PR when the deploy failed. Named for (2)/(3), which run on the success path too — not `AbortDeploy`. Infallible by construction |
| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build | | `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build |
| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#<name>` ref resolves. Store/meta-only, no container yet | | `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#<name>` ref resolves. Store/meta-only, no container yet |
| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop | | `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop |
@ -75,14 +71,12 @@ Cheap — no build slot:
| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step | | `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step |
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s | | `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles | | `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles |
| `FinalizeDeploy` | deploy phase 3 — drop the rollback ref, plant `deployed/<id>`, commit the staged `flake.lock`. The first two git steps are fatal on purpose, so a confirmed-good deploy's outcome and the repo's state can't disagree |
| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent | | `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent |
| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome | | `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome |
| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload | | `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload |
| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots | | `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots |
| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless | | `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless |
| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless | | `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless |
| `WebhookRegister` | one-shot boot-time Forgejo webhook registration (`internal/knowledge` push→pull, `agent-configs` PR→approval). No-op until the core token, hive domain, and HMAC secret are all available. Agentless |
| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless | | `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless |
| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless | | `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless |
@ -93,26 +87,21 @@ with its commit under its internal `META_LOCK` mutex, so a standalone commit
node would open a dirty-working-tree window between nodes. node would open a dirty-working-tree window between nodes.
Two further layers protect the meta repo across _windows_ that span multiple Two further layers protect the meta repo across _windows_ that span multiple
`META_LOCK` acquisitions — above all the approval deploy's prepare→finalize `META_LOCK` acquisitions:
span, which keeps a bumped `flake.lock` **staged uncommitted** for the whole
container build:
- **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue - **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue
resource declared by every node kind that mutates the meta repo — `MetaSync`, resource declared by every node kind that mutates the meta repo — `MetaSync`,
`MetaLock`, `WritePermFile`, `Provision`'s agent registration, and `MetaLock`, `WritePermFile` and `Provision`'s agent registration. Two meta
`DeployWindow` — the deploy subtree's root, which holds it across every
phase below it (it declares `Resource::MetaWindow`). Two meta
mutations can therefore never interleave, so no commit lands inside another mutations can therefore never interleave, so no commit lands inside another
node's staged window. It's a queue resource rather than a runtime mutex node's window. It's a queue resource rather than a runtime mutex
because a subtree root holds a resource across its whole subtree, which because a subtree root holds a resource across its whole subtree, which
a `MutexGuard` (bounded by one executor fn) can't — that's what lets a a `MutexGuard` (bounded by one executor fn) can't. For the same reason the window must stay
multi-node deploy own one window. For the same reason the window must stay
_off_ long store-only work: the rebuild's meta preamble is its own _off_ long store-only work: the rebuild's meta preamble is its own
`MetaSync` node, a sibling of (never a parent of) `Prebuild`, so the `MetaSync` node, a sibling of (never a parent of) `Prebuild`, so the
toplevel build runs outside the window and `buildSlots > 1` still gives toplevel build runs outside the window and `buildSlots > 1` still gives
concurrent rebuilds across agents. concurrent rebuilds across agents.
- **Path-limited commits**: the targeted meta committers (perm files, - **Path-limited commits**: the targeted meta committers (perm files,
topology, lock bumps, finalize) commit `-- <their paths>` with path-scoped topology, lock bumps) commit `-- <their paths>` with path-scoped
dirty checks, so even a non-queue caller (boot migration, destroy's dirty checks, so even a non-queue caller (boot migration, destroy's
`sync_agents`) can never sweep someone else's staged content into its `sync_agents`) can never sweep someone else's staged content into its
commit. commit.
@ -220,8 +209,8 @@ resources are free. Resources:
2. **Per-agent lifecycle lease** — keyed on the **node's** agent (agent is 2. **Per-agent lifecycle lease** — keyed on the **node's** agent (agent is
per-node; a DAG can span agents) and globally exclusive per agent across per-node; a DAG can span agents) and globally exclusive per agent across
all DAGs: acquired either at a container-affecting node (`SetWanted`, all DAGs: acquired either at a container-affecting node (`SetWanted`,
`Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`, `Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`) on
`DeployWindow`) on behalf of a whole coordinated subtree; held by the owning behalf of a whole coordinated subtree; held by the owning
DAG until it's terminal, so two DAGs never interleave container ops on the DAG until it's terminal, so two DAGs never interleave container ops on the
same agent. A DAG touching multiple agents holds one lease per agent. same agent. A DAG touching multiple agents holds one lease per agent.
(`SetWanted` is a store write, not a container op, but takes the lease anyway (`SetWanted` is a store write, not a container op, but takes the lease anyway
@ -293,36 +282,9 @@ the dashboard renders one recent-builds list and one number bounds it.
### Approvals ### Approvals
`MergeConfigPr` approvals ride as a four-node deploy subtree: `UpdateMetaInputs` approvals map onto the ordinary `meta-update` shapes.
The scheduler fires `actions::resolve_approval_dag` exactly once when
``` **any** approval-carrying DAG settles terminal (including
DeployWindow (root — build slot + lease + meta window, no work of its own)
├── MergeVerify drift gate, fetch, verify_commit
├── DeployApply AfterOk(verify) park rollback ref, ff-merge, deploy
└── DeployTail AfterAny(apply) compensate, mirror to forge
```
The root holds its resources across the whole subtree, so the two-phase
`prepare_deploy` / `finalize_deploy` span keeps its staged `flake.lock`
protected even though the phases are separate nodes. Splitting them buys
three things a single opaque node couldn't have: per-phase visibility on the
dashboard, a `MergeVerify` failure that provably mutated nothing, and a
compensation step that survives a hive-c0re restart — `DeployApply` parks the pre-merge
`applied/main` in `refs/hyperhive/rollback/<approval-id>`, not in a
local variable, so `DeployTail` can still undo a half-finished deploy after a
crash.
`DeployWindow` declares all three resources (build slot, lease, meta window)
on itself rather than letting each phase declare its own, because the queue
acquires a node's resources atomically (all-or-nothing): a child that took
the build slot while its parent held the meta window could block waiting for
a resource its own parent already committed to, a lock-ordering hazard that
one multi-resource root avoids by construction.
`UpdateMetaInputs` approvals map onto the ordinary
`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag`
exactly once when **any** approval-carrying DAG settles terminal — deploys
included, since their outcome is the DAG's own state (including
cancelled-while-queued, which fails the approval instead of dangling it). cancelled-while-queued, which fails the approval instead of dangling it).
### Wire shape ### Wire shape
@ -398,16 +360,12 @@ Key operations:
- **`sync_agents`** (idempotent) — render `flake.nix` for the current agent set, - **`sync_agents`** (idempotent) — render `flake.nix` for the current agent set,
init the repo on first call, relock if the rendered contents changed, commit. init the repo on first call, relock if the rendered contents changed, commit.
Called by spawn / destroy / startup migration. Called by spawn / destroy / startup migration.
- **`prepare_deploy` + `finalize_deploy` / `abort_deploy`** — two-phase for the
`MergeConfigPr` deploy path so a failed `nixos-container update` leaves no orphan
commit in meta. Prepare writes the new lock without committing; finalize commits
with the deploy message; abort restores the lock.
- **`lock_update_hyperhive`** — one-shot for the boot-reconcile path (the - **`lock_update_hyperhive`** — one-shot for the boot-reconcile path (the
sweep DAG's `MetaLock` node): bumps the `hyperhive` input lock and commits; sweep DAG's `MetaLock` node): bumps the `hyperhive` input lock and commits;
the scheduler fans out the agent rebuilds on completion. the scheduler fans out the agent rebuilds on completion.
Every public `meta.rs` operation takes the module's internal `META_LOCK` Every public `meta.rs` operation takes the module's internal `META_LOCK`
mutex, so concurrent job-queue nodes (and the approval deploy pipeline) never mutex, so concurrent job-queue nodes never
race on the repo's `.git/index.lock`. race on the repo's `.git/index.lock`.
--- ---
@ -449,18 +407,6 @@ Sequence for a rebuild DAG (each step is its own queue node):
in-container activation script transitions old → new. Holds no build in-container activation script transitions old → new. Holds no build
slot, so the next DAG's `Prebuild` overlaps the container boot. slot, so the next DAG's `Prebuild` overlaps the container boot.
The approval deploy uses this same chain rather than a rebuild path of its
own. Its `DeployApply` node doesn't build: it merges, opens the two-phase
meta deploy, and returns the chain above as a subgraph the scheduler grafts
into the live DAG under that node. A `FinalizeDeploy` node gated on the
graft's completion then plants the deploy tag — so `Reconcile`'s success
answers "did the agent come back up?" the same way it does for every
other rebuild, instead of a fused inline start.
The grafted nodes land _inside_ `DeployWindow`'s subtree, so they re-enter
the meta window and build slot it already holds rather than deadlocking
against it.
### Cold-start fallback ### Cold-start fallback
`start` after `update` can exit non-zero when packages are **removed** between `start` after `update` can exit non-zero when packages are **removed** between

View file

@ -3,7 +3,7 @@
Long-running work runs through a job graph. The swarm controller keeps one Long-running work runs through a job graph. The swarm controller keeps one
for swarm-level work — creating an agent's identity, forge user and config for swarm-level work — creating an agent's identity, forge user and config
repo. Each hive's hive-c0re keeps its own for container operations — repo. Each hive's hive-c0re keeps its own for container operations —
rebuild, first-spawn, a config-PR deploy, power changes. This page explains rebuild, first-spawn, power changes. This page explains
what the job queue _is_, as a general idea, independent of what either uses what the job queue _is_, as a general idea, independent of what either uses
it for. For the hive-c0re step catalogue and the engineering internals it for. For the hive-c0re step catalogue and the engineering internals
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md) (scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)

View file

@ -227,9 +227,9 @@ one per hive. It ensures them at start and every five minutes after
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's - the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
owner org; owner org;
- the empty `operators` merge-gate team in `agents` and `agent-configs`; - the empty `operators` merge-gate team in `agents` and `agent-configs`;
- the `main` merge gate on every `agent-configs` repo: merge whitelist = - the `main` merge gate on every `agent-configs` repo: merge and approval
the `operators` team and the `core` user, approval whitelist = the whitelists = the `operators` team, and no user. The controller leaves a
`operators` team. The controller leaves a repo with no `main` rule alone; repo with no `main` rule alone;
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's - the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds); host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
- `internal/docs` (private) and `internal/knowledge` (public, with a - `internal/docs` (private) and `internal/knowledge` (public, with a
@ -258,7 +258,7 @@ decision, not an event to adjudicate.
A `config-pr` delivery reporting a PR merged into `main` queues a deploy A `config-pr` delivery reporting a PR merged into `main` queues a deploy
of its `merge_commit_sha` on the one hive whose wanted state places the of its `merge_commit_sha` on the one hive whose wanted state places the
agent; with no such hive, or several, the controller deploys nothing. See agent; with no such hive, or several, the controller deploys nothing. See
[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui). [approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes).
**`internal/knowledge` is on that path.** The controller's is the only **`internal/knowledge` is on that path.** The controller's is the only
hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a
@ -267,10 +267,10 @@ would take delivery away from the first rather than add a recipient.
<!-- vale write-good.Passive = NO --> <!-- vale write-good.Passive = NO -->
**The `agent-configs` org isn't.** Each hive registers its own **The `agent-configs` org is on it too.** The controller's hook is the
`pull_request` hook there, so that repo has two — the hive's and the only one that acts on config PRs. A hive's `/webhook/config-pr` hook left
controller's — and **both are expected; don't delete either.** Removing on the org by an older release delivers to a route no hive serves; delete
a hive's stops it acting on config PRs; removing the controller's stops it in the org's webhook settings. Removing the controller's hook stops
forge-UI merges from deploying until its next start recreates it. forge-UI merges from deploying until its next start recreates it.
<!-- vale write-good.Passive = YES --> <!-- vale write-good.Passive = YES -->

View file

@ -40,8 +40,8 @@ hivectl forge reconcile-config iris --verbose # include the full diff, not
applied config checkout and its forge `agent-configs/<agent>` `main`, then applied config checkout and its forge `agent-configs/<agent>` `main`, then
reconciles. `--from forge` resets the local checkout to forge `main` (takes reconciles. `--from forge` resets the local checkout to forge `main` (takes
effect on the next deploy — it doesn't autorebuild). `--from local` isn't effect on the next deploy — it doesn't autorebuild). `--from local` isn't
supported yet (forge `main` is core-only branch-protected; resolve via a supported yet (forge `main` is branch-protected; resolve via a config
config PR). With no `--from` it prompts for the direction after the diff. PR). With no `--from` it prompts for the direction after the diff.
## Matrix ## Matrix

View file

@ -123,14 +123,17 @@ checkpoints**, not about sandboxing the agent from its own tools:
highest-value action. On the **internal forge this is technically enforced, highest-value action. On the **internal forge this is technically enforced,
not just convention**: agents can't create repos (`max_repo_creation = 0`), not just convention**: agents can't create repos (`max_repo_creation = 0`),
and `main` on an `agent-configs/<name>` repo carries swarm-controller's and `main` on an `agent-configs/<name>` repo carries swarm-controller's
branch protection: merge allowlisted to the `operators` team, with one branch protection: merge and approval allowlisted to the `operators` team,
approval from it. Existing `agents/<repo>` repos carry the same merge gate. which swarm-controller converges on every config repo. An operator's merge
there is also what deploys the config. Existing `agents/<repo>` repos carry
the same merge gate.
An agent (a write collaborator, not a repo admin) can neither change those An agent (a write collaborator, not a repo admin) can neither change those
settings nor merge its own PR. It's **not** set up for external VCS (GitHub settings nor merge its own PR. It's **not** set up for external VCS (GitHub
etc.), though — there, operator-merge is process + accepted risk, not a etc.), though — there, operator-merge is process + accepted risk, not a
technical control. technical control.
- **Approvals** — config changes, schedule additions, and other - **Approvals** — schedule additions and other blast-radius-y operations
blast-radius-y operations route through the operator approval queue route through the operator approval queue; config changes are config PRs
an operator merges on the forge
(see [`approvals.md`](../agent-lifecycle/approvals.md)). (see [`approvals.md`](../agent-lifecycle/approvals.md)).
### Capability = accepted risk ### Capability = accepted risk

View file

@ -70,12 +70,7 @@ approval (`Coordinator::notify_submitter`, looked up from the
authenticated socket caller at submit time; a row with no authenticated socket caller at submit time; a row with no
recorded submitter falls back to the manager, `ruth`). recorded submitter falls back to the manager, `ruth`).
`ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`, `ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`,
hardcoded — `hive-c0re/src/workers/crash_watch.rs`). A `MergeConfigPr` hardcoded — `hive-c0re/src/workers/crash_watch.rs`). Lifecycle
approval's rebuild additionally pushes a `rebuilt:<agent>` todo to that
same submitter (`Coordinator::push_todo_submitter`, `subsystem =
"core"`), which wakes a turn (the todo-wake path — see [Turn
outcomes](README.md#turn-outcomes)) via a generic "call
`get_loose_ends`" prompt rather than the event body itself. Lifecycle
transitions the job-queue scheduler or crash watcher drive directly — transitions the job-queue scheduler or crash watcher drive directly —
stop/kill, destroy, a flake-rev login or logout state change — reach stop/kill, destroy, a flake-rev login or logout state change — reach
no individual agent: they publish onto a swarm-wide NATS no individual agent: they publish onto a swarm-wide NATS

View file

@ -48,24 +48,24 @@ and quick links (stats, screen, forge profile). Select the name to open
its terminal and watch it work in real time. its terminal and watch it work in real time.
**Approve something an agent is waiting on.** Y3R C4LL is the one tab **Approve something an agent is waiting on.** Y3R C4LL is the one tab
worth checking regularly — it's everything that needs _you_: approvals worth checking regularly — it's everything that needs _you_: an agent's
for config changes. The tab's count pill tells you at a glance if request to schedule a prompt. The tab's count pill tells you at a glance if
anything's pending. anything's pending.
**Approve or reject a config change.** Agent config changes (new **Review a config change.** Agent config changes (new packages, env
packages, env vars, MCP servers) go through an approval queue rather vars, MCP servers) are pull requests on the agent's `agent-configs`
than landing automatically — you'll see them on Y3R C4LL, with a diff repo. Review and merge them on the forge; the merge deploys the change.
of what's changing. See [`approvals.md`](../agent-lifecycle/approvals.md#config-changes).
**Start, stop, restart, or rebuild an agent.** Select one or more **Start, stop, restart, or rebuild an agent.** Select one or more
agents on SW4RM (select the icon) and use the selection bar, or use the agents on SW4RM (select the icon) and use the selection bar, or use the
per-agent `⋮` menu on a single row. Rebuilding re-applies that agent's per-agent `⋮` menu on a single row. Rebuilding re-applies that agent's
current config; use it after approving a change, or whenever an agent current config; use it whenever an agent
shows as "needs update." shows as "needs update."
**Watch a build.** BU1LDS shows the rebuild queue live, plus a **Watch a build.** BU1LDS shows the rebuild queue live, plus a
streaming log of whatever's currently building. Useful right after streaming log of whatever's currently building. Useful right after
approving a change or bumping a flake input. merging a config change or bumping a flake input.
**Grant or revoke a tool/capability.** P3RM1SS10NS is a checkbox matrix **Grant or revoke a tool/capability.** P3RM1SS10NS is a checkbox matrix
— rows are agents, columns are tool groups or capabilities. Nothing — rows are agents, columns are tool groups or capabilities. Nothing

View file

@ -877,11 +877,10 @@ renderApprovals`) with three stacked sections:
right-aligned `requested <N> ago` relative time from right-aligned `requested <N> ago` relative time from
`ApprovalView.requested_at`. Glyph and chip vary by kind: `ApprovalView.requested_at`. Glyph and chip vary by kind:
| kind | glyph | chip | sha shown | | kind | glyph | chip |
|---|---|---|---| |---|---|---|
| `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) | | `update_meta_inputs` | `↻` | `meta-update` |
| `update_meta_inputs` | `↻` | `meta-update` | — | | `schedule_prompt` | `⏱` | `schedule` |
| `schedule_prompt` | `⏱` | `schedule` | — |
<!-- vale write-good.Passive = NO --> <!-- vale write-good.Passive = NO -->
The chip ticks live every second via a `data-requested-at` The chip ticks live every second via a `data-requested-at`
@ -889,12 +888,9 @@ renderApprovals`) with three stacked sections:
the request has been pending ≥ 1h so a stale approval stands out; the request has been pending ≥ 1h so a stale approval stands out;
the `.stale` class flips precisely at the 3600s boundary rather the `.stale` class flips precisely at the 3600s boundary rather
than at the next `renderApprovals` call. than at the next `renderApprovals` call.
- **what-changed body** — the submitting agent's description, then - **what-changed body** — the submitting agent's description, then the
kind-specific drill-in triggers: kind's payload: the inputs to bump (`update_meta_inputs`) or the prompt to
- `merge_config_pr`: `↳ review PR on forge ↗` deep-links the schedule (`schedule_prompt`).
config PR into `agent-configs/<agent>/pulls/<pr_number>` (shown
only when `forge_present` is true and `pr_number` has a value). The config diff
lives on the forge PR itself — no inline diff side-panel.
- **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a - **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a
`prompt()` for an optional reason carried to the submitting agent as `prompt()` for an optional reason carried to the submitting agent as
`HelperEvent::ApprovalResolved.note`. `HelperEvent::ApprovalResolved.note`.