Watch
0
0
Fork
You've already forked hyperhive
0

docs: config changes are operator merges on the forge

Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.

Refs #4850
This commit is contained in:
atlas 2026-10-02 22:32:17 +02:00
commit a88ed9f24e
14 changed files with 195 additions and 396 deletions

View file

@ -1,38 +1,32 @@
# Approvals + helper events
The approval queue is hyperhive's pivot: nothing that changes the
shape of an agent (its config, whether it exists) happens without an
operator selection. The submitting agent — any agent with the `approvals`
tool group, which manages the config of its **direct children** (the
root agent for top-level agents; a sub-manager for its own subtree) — is
the policy gate in front of that queue; helper events are how it stays
informed about what happens after a decision lands.
The approval queue is where an agent asks the operator for something it
can't do itself: a scheduled prompt, or (legacy rows only) a meta-input
bump. Config changes don't go through this queue: they're pull requests
on the agent's config repo, which an operator merges on the forge, and
that merge deploys them. The submitting agent — any agent with the
`approvals` tool group, which manages the config of its **direct
children** (the root agent for top-level agents; a sub-manager for its
own subtree) — is the policy gate in front of both; helper events are how
it stays informed about what happens after a decision lands.
## For operators
Every add/remove/change to an agent lands on your dashboard's Y3R
C4LL tab (or `hivectl approvals pending` / `approve <id>` from the
CLI) before it takes effect. What you'll see, and what to do with it:
- **Config change** (`MergeConfigPr`) — an agent proposed a change to
another agent's config (or its own, via a sub-manager) as a forge
pull request. Review the diff on the forge — the dashboard card
links straight to it, same as reviewing any other PR. Approving
triggers the deploy automatically: hive-c0re re-verifies the PR
hasn't moved since you looked at it, evaluates it (a dry run,
nothing applied yet), merges it, and rebuilds the container. If
anything in that chain fails, the change rolls back automatically —
the agent stays on its last-good config, no recovery action needed
from you.
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
one or more Nix flake inputs (or all of them). Approving runs the
update and commits the lock change; it doesn't rebuild anything by
itself.
- **Config change** — an agent proposed a change to another agent's
config (or its own, via a sub-manager) as a pull request on
`agent-configs/<agent>`. Review it on the forge like any other PR, and
merge it there to deploy it. See [Config changes](#config-changes).
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
a message to one or more inboxes at a future time. You can also add
schedules yourself directly from the SCH3DUL3S tab, which skips this
approval step entirely — the gate here is specifically for an
*agent* asking to schedule something, not for you doing it.
a message to one or more inboxes at a future time. It lands on your
dashboard's Y3R C4LL tab (or `hivectl approvals pending` /
`approve <id>` from the CLI). You can also add schedules yourself
directly from the SCH3DUL3S tab, which skips this approval step
entirely — the gate here is specifically for an *agent* asking to
schedule something, not for you doing it.
- **Meta/flake update** (`UpdateMetaInputs`) — legacy: nothing queues
this kind, but an existing row still reads back and you can approve it.
Approving runs the update and commits the lock change; it doesn't
rebuild anything by itself.
Don't want to approve something? **Deny it** (`DENY` on the dashboard
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
@ -40,18 +34,16 @@ submitting agent is always notified that the operator denied their request; what
optional is only the reason text, which you can add on the dashboard's
prompt (cancelling that prompt aborts the whole deny, not just the
reason) but not from the CLI. Denying is final: a denied approval
can't be re-approved later, the agent has to submit a fresh one (a new
PR, a new request).
can't be re-approved later, the agent has to submit a fresh request.
Everything below this point is the implementation detail behind that
flow.
## End-to-end approval flow
## Config changes
Config changes flow through a **forge pull request** on the agent's
`agent-configs/<name>` repo — the same surface agents use for code PRs.
No bespoke MCP tool exists for config changes: opening the PR IS the
request.
No MCP tool exists for config changes: opening the PR IS the request.
1. The submitting agent (the child's parent, holding the `approvals`
tool group) **clones** `agent-configs/<name>`, edits it there (any
@ -59,89 +51,40 @@ request.
with its own git identity, and pushes a branch + opens a PR with
`hive-forge` — the same way it would change any other repo.
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
config, not the tree to edit: authoring in place there produces no PR
and no approval. (it's currently mounted read-write, which is a
defect tracked separately, not an authoring path.)
config, not the tree to edit: authoring in place there produces no PR.
(it's currently mounted read-write, which is a defect tracked
separately, not an authoring path.)
Branch protection (the agent isn't on the `main` push allowlist;
merge allowlist = `core` user + `operators` team; approvals allowlist
= `operators` team; see "Forge mirror" below) makes the agent a write
collaborator that **can't merge its own config PR**.
2. hive-c0re's `/webhook/config-pr` endpoint receives the Forgejo
`pull_request` event (opened / synchronized / reopened) and queues a
`MergeConfigPr` approval; a poll fallback catches any missed webhook.
The approval row stores the PR **number** (`commit_ref`) and the PR
**head sha at queue time** (`fetched_sha` — the "reviewed" sha). If
the PR head later moves, a fresh approval pinned to the new head supersedes
the stale one, so the operator always reviews what will
actually deploy.
3. The operator reviews the PR **on the forge** (native diff, threaded
comments, CI status) and sees a matching card on the dashboard with a
"review PR on forge" deep link. They select ◆ APPR0VE (or
`hivectl approvals approve <id>` on the CLI) once satisfied.
4. On approve, a deploy DAG runs three phases under a resource-holding
`DeployWindow` root (see *Queue templates* below):
- `MergeVerify` re-reads the live PR head and **aborts if it drifted**
from the reviewed `fetched_sha` (the submitter must push again,
which queues a fresh approval); then fetches that head into the
applied repo and **eval-verifies** it — a flake eval on a throwaway
checkout. This is the trust gate: it relies on c0re's own eval, not
on any in-repo (agent-forgeable) signal like a CI status. Nothing is
mutated in this phase, so a rejection here leaves the forge and the
applied repo exactly as they were.
- `DeployApply` parks the pre-merge `applied/main` in
`refs/hyperhive/rollback/<approval-id>`, then fast-forward-merges
the reviewed head to the forge config repo's `main` (this IS the
merge — a `core`-authenticated ff-merge pinned to the reviewed sha,
so a moved PR head can't substitute bytes), and runs the deploy
proper (`deploy_applied_target`): ff `applied/main`, two-phase meta
deploy, container rebuild. On success it drops the rollback ref and
plants `deployed/<id>`.
- `DeployTail` runs on **every** outcome, including a cancel-cascade.
If the rollback ref survived, the deploy never confirmed good: it
rolls `applied/main` back, resyncs the working tree, and aborts the
staged meta lock, so the agent stays on its last-good tree. Then it
mirrors the config repo (and its new deploy tag) to the forge.
The rollback state lives in a **git ref, not a local variable**, on
purpose: hive-c0re can restart between the apply and the tail, and the
tail still has to know what to undo when it does.
5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the
**submitting agent's** inbox via `notify_submitter`, carrying both the
canonical sha and the terminal tag (the approval row carries a
`submitter` column recording the agent the change is for).
### Operator merge in the forge UI
An operator can also merge a config PR straight in the Forgejo UI. That
deploys the merged commit on the hive that runs the agent:
1. swarm-controller gets the `agent-configs` org's `pull_request`
merge and approvals allowlist = `operators` team; see "Forge mirror"
below) makes the agent a write collaborator that **can't merge its own
config PR**.
2. An operator reviews the PR **on the forge** (native diff, threaded
comments, CI status) and merges it there. Merging needs membership of
the `operators` team in `agent-configs`. The team starts empty: add
each operator by hand in the forge UI.
3. swarm-controller gets the `agent-configs` org's `pull_request`
delivery. A `closed` event with `merged: true` and base branch `main`
names the commit in `merge_commit_sha`.
2. The controller looks up which hive's wanted state places the agent
4. The controller looks up which hive's wanted state places the agent
and publishes a deploy request carrying that commit on that hive's
deploy subject. With no such hive, or more than one, it deploys
nothing and logs a warning naming them.
3. If the hive's `applied/main` already is that commit — a dashboard
approval merges and deploys its own PR — it does nothing. Otherwise
it fetches the config repo's `main`, requires the commit to descend
from `applied/main`, fast-forwards `applied/main` to it, and queues a
rebuild. A failure before the rebuild (the fetch, or a commit that
doesn't descend) deploys nothing and posts the error as a comment on
the merged PR. An agent with no container on that hive yet ignores
5. If the hive's `applied/main` already is that commit, it does nothing.
Otherwise it fetches the config repo's `main`, requires the commit to
descend from `applied/main`, fast-forwards `applied/main` to it, and
queues a rebuild. A failure before the rebuild (the fetch, or a commit
that doesn't descend) deploys nothing and posts the error as a comment
on the merged PR. An agent with no container on that hive yet ignores
the commit; its first deploy builds what it seeds.
This path runs **no eval-verify**. A merged config that doesn't
evaluate or build fails the rebuild, and `applied/main` stays at that
commit, so the agent's rebuilds keep failing until a fix merges. A
failed rebuild shows on the hive like any other and isn't commented on
the PR.
the PR. A merge whose webhook delivery never reaches the controller
deploys nothing.
Merging needs membership of the `operators` team in `agent-configs`.
The team starts empty: add each operator by hand in the forge UI. A
merge whose webhook delivery never reaches the controller deploys
nothing. The hive's own poll cancels the dashboard card for that PR
with the note `PR merged/closed outside the approval`.
## Approval queue
### Withdrawing a pending approval
@ -166,31 +109,16 @@ and no hive can originate an agent. The swarm controller's
`agent.nix` template, then asks the target hive to deploy it.
Changing what the template seeded isn't a special case: like every
later change, it's a PR on that config repo (`MergeConfigPr`), made
from a clone, reviewed and approved by the operator. The PR flow is
the one path — an operator can equally drive both steps herself
through the web UI or the forge.
later change, it's a PR on that config repo, made from a clone and
merged by an operator on the forge (see [Config changes](#config-changes)).
### Approval kinds (wire shapes)
`ApprovalKind` carries three variants; each maps to a different
`ApprovalKind` carries two variants; each maps to a different
`commit_ref` encoding because `ApprovalKind` overloads that field as
the kind-specific payload carrier.
<!-- vale write-good.Passive = NO -->
- `MergeConfigPr` — the config-change flow. Triggered automatically:
when an agent opens (or force-pushes) a PR on its
`agent-configs/<agent>` forge repo, hive-c0re's `/webhook/config-pr`
endpoint receives the Forgejo pull_request event and queues this
approval row. No MCP tool call needed — the forge PR IS the request.
`commit_ref` stores the **PR number** (decimal), and `fetched_sha` is
the PR **head sha at queue time** (the "reviewed" sha). On approve,
the deploy DAG's `MergeVerify` phase re-reads the live PR head and
aborts if it drifted from `fetched_sha` (submitter must push again to
re-trigger), then fetches that head into the applied repo and
eval-verifies it; `DeployApply` fast-forward-merges the forge config
repo's `main` to it (the merge) and runs `deploy_applied_target`;
`DeployTail` compensates on failure. Never a first spawn.
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
etc.). hive-c0re sets the `agent` field to the requesting root agent.
@ -303,55 +231,26 @@ declares one flake input per agent (`agent-<n>.url =
Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
The declared input url is the agent's **forge config repo** (the
same `agent-configs/<n>` the config-PR flow lands approved changes
on), so the meta flake references a reviewable, reproducible source
rather than a local checkout. hive-c0re authenticates that
`git+http` fetch via a git credential helper that reads the live
forge-core token — no token in the url or the lock. The deploy and
manual-rebuild paths, however, do **not** re-lock from the forge:
they `--override-input agent-<n>
same `agent-configs/<n>` config PRs merge into), so the meta flake
references a reviewable, reproducible source rather than a local
checkout. hive-c0re authenticates that `git+http` fetch via a git
credential helper that reads the live forge-core token — no token in
the url or the lock. The rebuild paths, however, do **not** re-lock
from the forge: they `--override-input agent-<n>
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
that `verify_commit` gated and `applied/<n>/main` was
fast-forwarded to. That keeps a deploy/rebuild reproducible and
independent of forge reachability — rebuilds fire on crash-restart
and meta bumps, not just config PRs — while the declared url stays
the forge. `sync_agents` re-renders + re-locks the persistent input;
a plain `nix flake lock` leaves an existing applied override in
place (it only re-locks when the declared url itself changes), so
the forge-declared / applied-deployed split is stable.
`applied/<n>/main` was fast-forwarded to. That keeps a rebuild
reproducible and independent of forge reachability — rebuilds fire on
crash-restart and meta bumps, not just config merges — while the
declared url stays the forge. `sync_agents` re-renders + re-locks the
persistent input; a plain `nix flake lock` leaves an existing applied
override in place (it only re-locks when the declared url itself
changes), so the forge-declared / applied-deployed split is stable.
Per-deploy lock flow (two-phase), spread across the deploy subtree's
nodes — each phase is its own node, so the queue can show which one is
running and a restart resumes at node granularity:
1. `DeployApply` → `meta::prepare_deploy(name)` runs
`nix flake lock --update-input agent-<n>` without
committing. Working tree of meta now points the input at
`applied/<n>/main` (which the deploy already fast-forwarded to
the reviewed PR head).
2. The rebuild subgraph `DeployApply` grows into the DAG builds and
swaps the container (`AgentWindow` bracing `Prebuild → StopForUpdate
→ Swap → RebuildBookkeeping`, plus `Reconcile`). Nix evaluates
against the staged lock.
3. On success — `FinalizeDeploy` drops the rollback ref, plants
`deployed/<id>`, then `meta::finalize_deploy(name, sha, "deployed/
<id>")` stages `flake.lock` and commits with
`deploy <n> deployed/<id> <sha12>`. Meta's git log gains
one entry per successful deploy.
4. On failure — the `DeployTail` node runs `meta::abort_deploy()`
(`git restore flake.lock`) so the meta history shows only
successes; the failure stays as an annotated `failed/<id>`
tag in `applied/<n>`. The tail runs on every outcome, so this
also covers a hive-c0re restart mid-build: the staged lock is
dropped and `applied/main` rolled back from the parked
`refs/hyperhive/rollback/<id>`.
Single-phase variants exist for paths without
rollback semantics: `meta::lock_update_for_rebuild(name)` for
the manual `↻ R3BU1LD` button (commits if the lock changed)
and `meta::lock_update_hyperhive()` for the
autoupdate flake-rev bump (one shot before per-agent
rebuilds, commits if the lock changed).
Lock updates are single-phase and commit when the lock changed:
`meta::lock_update_for_rebuild(name)` relocks one agent's input for a
relocking rebuild (the manual `↻ R3BU1LD` button, and the rebuild a
merged config commit queues), and `meta::lock_update_hyperhive()` is the
autoupdate flake-rev bump (one shot before per-agent rebuilds).
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
carries `hyperhive_flake`, `dashboard_port`, and the rest of the
@ -390,8 +289,8 @@ per container row.
└── <other committed files> # also tracked
/var/lib/hyperhive/meta/ swarm-wide flake — core
├── .git/ # one commit per successful
│ # deploy
├── .git/ # one commit per lock
│ # change
├── flake.nix # generated from agent set
└── flake.lock # pins each agent's sha
```
@ -411,43 +310,27 @@ wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` /
### Tag state machine
Each deploy leaves a tag on the underlying commit inside the applied
repo:
| Tag | When | Annotated? |
|---|---|---|
| `deployed/<id>` | rebuild succeeded — `main` ff's here | no |
| `failed/<id>` | rebuild failed | yes (body = error) |
hive-c0re plants `deployed/0` at first spawn. `applied/main` is always the
latest `deployed/*`. A `failed/` tree stays browsable forever — `git log
--tags` in the applied repo is the audit trail. A denied or failed config
PR carries no extra state on the forge side: the PR stays open, and the
submitter pushes again (or closes it) to retry.
hive-c0re plants `deployed/0` on the seed commit at first spawn. A
merged config commit that deploys plants no tag: the merged PR on the
forge and meta's lock commits record it. A config PR nobody merges
carries no extra state on the forge side: the PR stays open, and
the submitter pushes again (or closes it) to retry.
### Dispatch via the job queue
Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs`
— runs as a DAG on the global job queue
(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler
rather than run inline:
Long-running approval work — `UpdateMetaInputs` — runs as a DAG on the
global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted
by the approval handler rather than run inline:
| `ApprovalKind` | DAG submitted | source |
|---|---|---|
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
The DAG carries the originating `approval_id`, surfaced on the node that
owns it — for a deploy that's the `DeployWindow` root, so the dashboard
renders one approval card, not four. **Every** queued kind resolves
through `actions::resolve_approval_dag` when its DAG settles terminal:
the deploy's phases are ordinary queue nodes, so the DAG's own terminal
state is the authoritative outcome. That hook fires the matching
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
terminal tag (verifying the tag actually resolves in the applied repo —
a pre-merge rejection plants none), posts the failing build log back to
the config PR.
The DAG carries the originating `approval_id`. **Every** queued kind
resolves through `actions::resolve_approval_dag` when its DAG settles
terminal: the DAG's own terminal state is the authoritative outcome.
That hook fires the matching `HelperEvent::*` via `finish_approval`.
Two visible consequences:
@ -474,30 +357,27 @@ reconcile) DAGs use the same queue but skip the approval plumbing.
The bundled `hive-forge` container runs on the swarm's forge host
(`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)),
and hive-c0re mirrors every agent's applied repo into a
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes `applied/main` plus
every tag to `agent-configs/<name>` after each ref mutation:
the spawn that seeds `deployed/0`, every successful deploy (which
plants `deployed/<id>`) or failed build (`failed/<id>`), and a
sweep at startup. Pushes are best-effort — a missing or stopped
forge never blocks a deploy.
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes
every tag, then `applied/main`, to `agent-configs/<name>` on every
startup sweep and every rebuild. Forge `main` is branch-protected, so
the forge routinely refuses a push of an established `main`, which
hive-c0re expects. Pushes are best-effort — a missing or stopped forge never
blocks a deploy.
Each agent is a **write collaborator on its own** `agent-configs/<name>`
repo — so it can push a branch and open a config PR — but not a member
of any other agent's, so it can't reach another agent's config through
the forge. Branch protection keeps the agent off the `main` push
allowlist and whitelists merging to the `core` user and the `operators`
team, so an agent can't fast-forward its own config or self-merge its
PR (see the End-to-end flow and
[Operator merge in the forge UI](#operator-merge-in-the-forge-ui) above).
hive-c0re passes the tokenised push
URL inline to `git push`, never writing it into
allowlist and allowlists merging to the `operators` team only, so an
agent can't fast-forward its own config or self-merge its PR (see
[Config changes](#config-changes) above). hive-c0re passes the tokenised
push URL inline to `git push`, never writing it into
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
agent, and a stored token would leak core's admin credential to an
agent.
The dashboard deep-links into this org — a `config repo` link
per container row and a `review PR on forge` link per config-PR
approval card. See `docs/web-ui/dashboard.md`.
per container row. See `docs/web-ui/dashboard.md`.
### Submitting agent's view of config repos
@ -548,8 +428,7 @@ the full `/applied` mount:
```sh
git -C /agents/<n>/config fetch applied
git -C /agents/<n>/config log applied/main --oneline
git -C /agents/<n>/config show applied/refs/tags/deployed/<id>
git -C /agents/<n>/config show applied/refs/tags/failed/<id> # body = build error
git -C /agents/<n>/config show applied/refs/tags/deployed/0 # the seed commit
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
@ -631,11 +510,9 @@ as a regular `system` inbox message so it drives a normal claude turn.
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
**every** approval kind's terminal state. A
"FYI, check when convenient" event doesn't need a message — those go
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct
live dial of the target agent's in-container todo socket (same
`UpsertTodo` request in-container producers use); `finish_approval` fires
one of these too for `MergeConfigPr`, *in addition to*
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
through `Coordinator::push_todo` instead, a direct live dial of the target
agent's in-container todo socket (same `UpsertTodo` request in-container
producers use). Legacy
approval rows that predate the submitter column fall back to the
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
@ -657,28 +534,17 @@ root agent. Variants (`hive_sh4re::manager::HelperEvent`):
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
`Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check
when convenient" events with no reason to drive an immediate turn, so
they deliver via `push_todo`/`push_todo_submitter` (see above) instead
they deliver via `push_todo` (see above) instead
of `HelperEvent`: an `agent_todo_socket` push instead of a broker
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
and a single free-text `summary` (`rebuilt_todo_summary` renders
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
Optional `sha` field on `ApprovalResolved` carries the canonical
hive-c0re-vouched commit sha. Optional `tag` carries the deploy
bookkeeping tag — `deployed/<id>` on a successful build or
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
Both fields are `Option`: `None` on the paths that don't deploy a new
commit (meta-update / deny, and the autoupdate
sweep's `job_queue::templates::rebuild` reapplying the existing main,
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
`git show <sha>` against `/applied/<n>/.git` inside the
bootstrap container yields the exact tree the sha referenced.
To add a new lifecycle notice: if it needs to drive an immediate turn
(something genuinely urgent, like `ContainerCrash`), add a
`HelperEvent` variant + call sites + update `prompts/system.md`'s
message-event list. If it's "FYI, check when convenient," call
`push_todo`/`push_todo_submitter` directly instead — no new wire type
`push_todo` directly instead — no new wire type
needed.
## Autoupdate on startup