docs: config changes are operator merges on the forge
Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.
Refs #4850
This commit is contained in:
parent
0cee0382e9
commit
a88ed9f24e
14 changed files with 195 additions and 396 deletions
|
|
@ -31,7 +31,7 @@ architecture change.
|
|||
| **identity** | every agent is a swarm-wide principal — SSO subject, forge user, matrix account, secret-store cert identity — addressable as `name@hive.domain` |
|
||||
| **secrets** | one OpenBao store; the operator places one mTLS identity per host, and everything else — every agent's credentials included — is fetched from the store under an identity rather than copied by hand |
|
||||
| **shared services** | one forge, homeserver, SSO, message queue and metrics/logs stack per swarm, each on whichever host you put it |
|
||||
| **config** | git: an agent proposes, the operator approves, the deploy lands as a `deployed/<id>` tag |
|
||||
| **config** | git: an agent opens a PR on its config repo, an operator merges it on the forge, and the merge deploys |
|
||||
| **runtime** | `claude --print` by default; any [ACP](https://agentclientprotocol.com) agent (e.g. opencode) per agent with `services.hyperhive.agent.runtime = "acp"` |
|
||||
| **substrate** | each hive runs one `nixos-container` per agent, under an unprivileged daemon with a tiny socket-activated helper for the few root ops |
|
||||
| **watching** | the swarm UI (hives, agents with live terminals, jobs, a cross-repo issue report), Grafana over OTEL, and a per-hive dashboard for host-level detail |
|
||||
|
|
|
|||
|
|
@ -1,38 +1,32 @@
|
|||
# Approvals + helper events
|
||||
|
||||
The approval queue is hyperhive's pivot: nothing that changes the
|
||||
shape of an agent (its config, whether it exists) happens without an
|
||||
operator selection. The submitting agent — any agent with the `approvals`
|
||||
tool group, which manages the config of its **direct children** (the
|
||||
root agent for top-level agents; a sub-manager for its own subtree) — is
|
||||
the policy gate in front of that queue; helper events are how it stays
|
||||
informed about what happens after a decision lands.
|
||||
The approval queue is where an agent asks the operator for something it
|
||||
can't do itself: a scheduled prompt, or (legacy rows only) a meta-input
|
||||
bump. Config changes don't go through this queue: they're pull requests
|
||||
on the agent's config repo, which an operator merges on the forge, and
|
||||
that merge deploys them. The submitting agent — any agent with the
|
||||
`approvals` tool group, which manages the config of its **direct
|
||||
children** (the root agent for top-level agents; a sub-manager for its
|
||||
own subtree) — is the policy gate in front of both; helper events are how
|
||||
it stays informed about what happens after a decision lands.
|
||||
|
||||
## For operators
|
||||
|
||||
Every add/remove/change to an agent lands on your dashboard's Y3R
|
||||
C4LL tab (or `hivectl approvals pending` / `approve <id>` from the
|
||||
CLI) before it takes effect. What you'll see, and what to do with it:
|
||||
|
||||
- **Config change** (`MergeConfigPr`) — an agent proposed a change to
|
||||
another agent's config (or its own, via a sub-manager) as a forge
|
||||
pull request. Review the diff on the forge — the dashboard card
|
||||
links straight to it, same as reviewing any other PR. Approving
|
||||
triggers the deploy automatically: hive-c0re re-verifies the PR
|
||||
hasn't moved since you looked at it, evaluates it (a dry run,
|
||||
nothing applied yet), merges it, and rebuilds the container. If
|
||||
anything in that chain fails, the change rolls back automatically —
|
||||
the agent stays on its last-good config, no recovery action needed
|
||||
from you.
|
||||
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
|
||||
one or more Nix flake inputs (or all of them). Approving runs the
|
||||
update and commits the lock change; it doesn't rebuild anything by
|
||||
itself.
|
||||
- **Config change** — an agent proposed a change to another agent's
|
||||
config (or its own, via a sub-manager) as a pull request on
|
||||
`agent-configs/<agent>`. Review it on the forge like any other PR, and
|
||||
merge it there to deploy it. See [Config changes](#config-changes).
|
||||
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
|
||||
a message to one or more inboxes at a future time. You can also add
|
||||
schedules yourself directly from the SCH3DUL3S tab, which skips this
|
||||
approval step entirely — the gate here is specifically for an
|
||||
*agent* asking to schedule something, not for you doing it.
|
||||
a message to one or more inboxes at a future time. It lands on your
|
||||
dashboard's Y3R C4LL tab (or `hivectl approvals pending` /
|
||||
`approve <id>` from the CLI). You can also add schedules yourself
|
||||
directly from the SCH3DUL3S tab, which skips this approval step
|
||||
entirely — the gate here is specifically for an *agent* asking to
|
||||
schedule something, not for you doing it.
|
||||
- **Meta/flake update** (`UpdateMetaInputs`) — legacy: nothing queues
|
||||
this kind, but an existing row still reads back and you can approve it.
|
||||
Approving runs the update and commits the lock change; it doesn't
|
||||
rebuild anything by itself.
|
||||
|
||||
Don't want to approve something? **Deny it** (`DENY` on the dashboard
|
||||
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
|
||||
|
|
@ -40,18 +34,16 @@ submitting agent is always notified that the operator denied their request; what
|
|||
optional is only the reason text, which you can add on the dashboard's
|
||||
prompt (cancelling that prompt aborts the whole deny, not just the
|
||||
reason) but not from the CLI. Denying is final: a denied approval
|
||||
can't be re-approved later, the agent has to submit a fresh one (a new
|
||||
PR, a new request).
|
||||
can't be re-approved later, the agent has to submit a fresh request.
|
||||
|
||||
Everything below this point is the implementation detail behind that
|
||||
flow.
|
||||
|
||||
## End-to-end approval flow
|
||||
## Config changes
|
||||
|
||||
Config changes flow through a **forge pull request** on the agent's
|
||||
`agent-configs/<name>` repo — the same surface agents use for code PRs.
|
||||
No bespoke MCP tool exists for config changes: opening the PR IS the
|
||||
request.
|
||||
No MCP tool exists for config changes: opening the PR IS the request.
|
||||
|
||||
1. The submitting agent (the child's parent, holding the `approvals`
|
||||
tool group) **clones** `agent-configs/<name>`, edits it there (any
|
||||
|
|
@ -59,89 +51,40 @@ request.
|
|||
with its own git identity, and pushes a branch + opens a PR with
|
||||
`hive-forge` — the same way it would change any other repo.
|
||||
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
|
||||
config, not the tree to edit: authoring in place there produces no PR
|
||||
and no approval. (it's currently mounted read-write, which is a
|
||||
defect tracked separately, not an authoring path.)
|
||||
config, not the tree to edit: authoring in place there produces no PR.
|
||||
(it's currently mounted read-write, which is a defect tracked
|
||||
separately, not an authoring path.)
|
||||
Branch protection (the agent isn't on the `main` push allowlist;
|
||||
merge allowlist = `core` user + `operators` team; approvals allowlist
|
||||
= `operators` team; see "Forge mirror" below) makes the agent a write
|
||||
collaborator that **can't merge its own config PR**.
|
||||
2. hive-c0re's `/webhook/config-pr` endpoint receives the Forgejo
|
||||
`pull_request` event (opened / synchronized / reopened) and queues a
|
||||
`MergeConfigPr` approval; a poll fallback catches any missed webhook.
|
||||
The approval row stores the PR **number** (`commit_ref`) and the PR
|
||||
**head sha at queue time** (`fetched_sha` — the "reviewed" sha). If
|
||||
the PR head later moves, a fresh approval pinned to the new head supersedes
|
||||
the stale one, so the operator always reviews what will
|
||||
actually deploy.
|
||||
3. The operator reviews the PR **on the forge** (native diff, threaded
|
||||
comments, CI status) and sees a matching card on the dashboard with a
|
||||
"review PR on forge" deep link. They select ◆ APPR0VE (or
|
||||
`hivectl approvals approve <id>` on the CLI) once satisfied.
|
||||
4. On approve, a deploy DAG runs three phases under a resource-holding
|
||||
`DeployWindow` root (see *Queue templates* below):
|
||||
- `MergeVerify` re-reads the live PR head and **aborts if it drifted**
|
||||
from the reviewed `fetched_sha` (the submitter must push again,
|
||||
which queues a fresh approval); then fetches that head into the
|
||||
applied repo and **eval-verifies** it — a flake eval on a throwaway
|
||||
checkout. This is the trust gate: it relies on c0re's own eval, not
|
||||
on any in-repo (agent-forgeable) signal like a CI status. Nothing is
|
||||
mutated in this phase, so a rejection here leaves the forge and the
|
||||
applied repo exactly as they were.
|
||||
- `DeployApply` parks the pre-merge `applied/main` in
|
||||
`refs/hyperhive/rollback/<approval-id>`, then fast-forward-merges
|
||||
the reviewed head to the forge config repo's `main` (this IS the
|
||||
merge — a `core`-authenticated ff-merge pinned to the reviewed sha,
|
||||
so a moved PR head can't substitute bytes), and runs the deploy
|
||||
proper (`deploy_applied_target`): ff `applied/main`, two-phase meta
|
||||
deploy, container rebuild. On success it drops the rollback ref and
|
||||
plants `deployed/<id>`.
|
||||
- `DeployTail` runs on **every** outcome, including a cancel-cascade.
|
||||
If the rollback ref survived, the deploy never confirmed good: it
|
||||
rolls `applied/main` back, resyncs the working tree, and aborts the
|
||||
staged meta lock, so the agent stays on its last-good tree. Then it
|
||||
mirrors the config repo (and its new deploy tag) to the forge.
|
||||
|
||||
The rollback state lives in a **git ref, not a local variable**, on
|
||||
purpose: hive-c0re can restart between the apply and the tail, and the
|
||||
tail still has to know what to undo when it does.
|
||||
5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the
|
||||
**submitting agent's** inbox via `notify_submitter`, carrying both the
|
||||
canonical sha and the terminal tag (the approval row carries a
|
||||
`submitter` column recording the agent the change is for).
|
||||
|
||||
### Operator merge in the forge UI
|
||||
|
||||
An operator can also merge a config PR straight in the Forgejo UI. That
|
||||
deploys the merged commit on the hive that runs the agent:
|
||||
|
||||
1. swarm-controller gets the `agent-configs` org's `pull_request`
|
||||
merge and approvals allowlist = `operators` team; see "Forge mirror"
|
||||
below) makes the agent a write collaborator that **can't merge its own
|
||||
config PR**.
|
||||
2. An operator reviews the PR **on the forge** (native diff, threaded
|
||||
comments, CI status) and merges it there. Merging needs membership of
|
||||
the `operators` team in `agent-configs`. The team starts empty: add
|
||||
each operator by hand in the forge UI.
|
||||
3. swarm-controller gets the `agent-configs` org's `pull_request`
|
||||
delivery. A `closed` event with `merged: true` and base branch `main`
|
||||
names the commit in `merge_commit_sha`.
|
||||
2. The controller looks up which hive's wanted state places the agent
|
||||
4. The controller looks up which hive's wanted state places the agent
|
||||
and publishes a deploy request carrying that commit on that hive's
|
||||
deploy subject. With no such hive, or more than one, it deploys
|
||||
nothing and logs a warning naming them.
|
||||
3. If the hive's `applied/main` already is that commit — a dashboard
|
||||
approval merges and deploys its own PR — it does nothing. Otherwise
|
||||
it fetches the config repo's `main`, requires the commit to descend
|
||||
from `applied/main`, fast-forwards `applied/main` to it, and queues a
|
||||
rebuild. A failure before the rebuild (the fetch, or a commit that
|
||||
doesn't descend) deploys nothing and posts the error as a comment on
|
||||
the merged PR. An agent with no container on that hive yet ignores
|
||||
5. If the hive's `applied/main` already is that commit, it does nothing.
|
||||
Otherwise it fetches the config repo's `main`, requires the commit to
|
||||
descend from `applied/main`, fast-forwards `applied/main` to it, and
|
||||
queues a rebuild. A failure before the rebuild (the fetch, or a commit
|
||||
that doesn't descend) deploys nothing and posts the error as a comment
|
||||
on the merged PR. An agent with no container on that hive yet ignores
|
||||
the commit; its first deploy builds what it seeds.
|
||||
|
||||
This path runs **no eval-verify**. A merged config that doesn't
|
||||
evaluate or build fails the rebuild, and `applied/main` stays at that
|
||||
commit, so the agent's rebuilds keep failing until a fix merges. A
|
||||
failed rebuild shows on the hive like any other and isn't commented on
|
||||
the PR.
|
||||
the PR. A merge whose webhook delivery never reaches the controller
|
||||
deploys nothing.
|
||||
|
||||
Merging needs membership of the `operators` team in `agent-configs`.
|
||||
The team starts empty: add each operator by hand in the forge UI. A
|
||||
merge whose webhook delivery never reaches the controller deploys
|
||||
nothing. The hive's own poll cancels the dashboard card for that PR
|
||||
with the note `PR merged/closed outside the approval`.
|
||||
## Approval queue
|
||||
|
||||
### Withdrawing a pending approval
|
||||
|
||||
|
|
@ -166,31 +109,16 @@ and no hive can originate an agent. The swarm controller's
|
|||
`agent.nix` template, then asks the target hive to deploy it.
|
||||
|
||||
Changing what the template seeded isn't a special case: like every
|
||||
later change, it's a PR on that config repo (`MergeConfigPr`), made
|
||||
from a clone, reviewed and approved by the operator. The PR flow is
|
||||
the one path — an operator can equally drive both steps herself
|
||||
through the web UI or the forge.
|
||||
later change, it's a PR on that config repo, made from a clone and
|
||||
merged by an operator on the forge (see [Config changes](#config-changes)).
|
||||
|
||||
### Approval kinds (wire shapes)
|
||||
|
||||
`ApprovalKind` carries three variants; each maps to a different
|
||||
`ApprovalKind` carries two variants; each maps to a different
|
||||
`commit_ref` encoding because `ApprovalKind` overloads that field as
|
||||
the kind-specific payload carrier.
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
- `MergeConfigPr` — the config-change flow. Triggered automatically:
|
||||
when an agent opens (or force-pushes) a PR on its
|
||||
`agent-configs/<agent>` forge repo, hive-c0re's `/webhook/config-pr`
|
||||
endpoint receives the Forgejo pull_request event and queues this
|
||||
approval row. No MCP tool call needed — the forge PR IS the request.
|
||||
`commit_ref` stores the **PR number** (decimal), and `fetched_sha` is
|
||||
the PR **head sha at queue time** (the "reviewed" sha). On approve,
|
||||
the deploy DAG's `MergeVerify` phase re-reads the live PR head and
|
||||
aborts if it drifted from `fetched_sha` (submitter must push again to
|
||||
re-trigger), then fetches that head into the applied repo and
|
||||
eval-verifies it; `DeployApply` fast-forward-merges the forge config
|
||||
repo's `main` to it (the merge) and runs `deploy_applied_target`;
|
||||
`DeployTail` compensates on failure. Never a first spawn.
|
||||
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
|
||||
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
|
||||
etc.). hive-c0re sets the `agent` field to the requesting root agent.
|
||||
|
|
@ -303,55 +231,26 @@ declares one flake input per agent (`agent-<n>.url =
|
|||
Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
|
||||
|
||||
The declared input url is the agent's **forge config repo** (the
|
||||
same `agent-configs/<n>` the config-PR flow lands approved changes
|
||||
on), so the meta flake references a reviewable, reproducible source
|
||||
rather than a local checkout. hive-c0re authenticates that
|
||||
`git+http` fetch via a git credential helper that reads the live
|
||||
forge-core token — no token in the url or the lock. The deploy and
|
||||
manual-rebuild paths, however, do **not** re-lock from the forge:
|
||||
they `--override-input agent-<n>
|
||||
same `agent-configs/<n>` config PRs merge into), so the meta flake
|
||||
references a reviewable, reproducible source rather than a local
|
||||
checkout. hive-c0re authenticates that `git+http` fetch via a git
|
||||
credential helper that reads the live forge-core token — no token in
|
||||
the url or the lock. The rebuild paths, however, do **not** re-lock
|
||||
from the forge: they `--override-input agent-<n>
|
||||
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
|
||||
that `verify_commit` gated and `applied/<n>/main` was
|
||||
fast-forwarded to. That keeps a deploy/rebuild reproducible and
|
||||
independent of forge reachability — rebuilds fire on crash-restart
|
||||
and meta bumps, not just config PRs — while the declared url stays
|
||||
the forge. `sync_agents` re-renders + re-locks the persistent input;
|
||||
a plain `nix flake lock` leaves an existing applied override in
|
||||
place (it only re-locks when the declared url itself changes), so
|
||||
the forge-declared / applied-deployed split is stable.
|
||||
`applied/<n>/main` was fast-forwarded to. That keeps a rebuild
|
||||
reproducible and independent of forge reachability — rebuilds fire on
|
||||
crash-restart and meta bumps, not just config merges — while the
|
||||
declared url stays the forge. `sync_agents` re-renders + re-locks the
|
||||
persistent input; a plain `nix flake lock` leaves an existing applied
|
||||
override in place (it only re-locks when the declared url itself
|
||||
changes), so the forge-declared / applied-deployed split is stable.
|
||||
|
||||
Per-deploy lock flow (two-phase), spread across the deploy subtree's
|
||||
nodes — each phase is its own node, so the queue can show which one is
|
||||
running and a restart resumes at node granularity:
|
||||
|
||||
1. `DeployApply` → `meta::prepare_deploy(name)` runs
|
||||
`nix flake lock --update-input agent-<n>` without
|
||||
committing. Working tree of meta now points the input at
|
||||
`applied/<n>/main` (which the deploy already fast-forwarded to
|
||||
the reviewed PR head).
|
||||
2. The rebuild subgraph `DeployApply` grows into the DAG builds and
|
||||
swaps the container (`AgentWindow` bracing `Prebuild → StopForUpdate
|
||||
→ Swap → RebuildBookkeeping`, plus `Reconcile`). Nix evaluates
|
||||
against the staged lock.
|
||||
3. On success — `FinalizeDeploy` drops the rollback ref, plants
|
||||
`deployed/<id>`, then `meta::finalize_deploy(name, sha, "deployed/
|
||||
<id>")` stages `flake.lock` and commits with
|
||||
`deploy <n> deployed/<id> <sha12>`. Meta's git log gains
|
||||
one entry per successful deploy.
|
||||
4. On failure — the `DeployTail` node runs `meta::abort_deploy()`
|
||||
(`git restore flake.lock`) so the meta history shows only
|
||||
successes; the failure stays as an annotated `failed/<id>`
|
||||
tag in `applied/<n>`. The tail runs on every outcome, so this
|
||||
also covers a hive-c0re restart mid-build: the staged lock is
|
||||
dropped and `applied/main` rolled back from the parked
|
||||
`refs/hyperhive/rollback/<id>`.
|
||||
|
||||
Single-phase variants exist for paths without
|
||||
rollback semantics: `meta::lock_update_for_rebuild(name)` for
|
||||
the manual `↻ R3BU1LD` button (commits if the lock changed)
|
||||
and `meta::lock_update_hyperhive()` for the
|
||||
autoupdate flake-rev bump (one shot before per-agent
|
||||
rebuilds, commits if the lock changed).
|
||||
Lock updates are single-phase and commit when the lock changed:
|
||||
`meta::lock_update_for_rebuild(name)` relocks one agent's input for a
|
||||
relocking rebuild (the manual `↻ R3BU1LD` button, and the rebuild a
|
||||
merged config commit queues), and `meta::lock_update_hyperhive()` is the
|
||||
autoupdate flake-rev bump (one shot before per-agent rebuilds).
|
||||
|
||||
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
|
||||
carries `hyperhive_flake`, `dashboard_port`, and the rest of the
|
||||
|
|
@ -390,8 +289,8 @@ per container row.
|
|||
└── <other committed files> # also tracked
|
||||
|
||||
/var/lib/hyperhive/meta/ swarm-wide flake — core
|
||||
├── .git/ # one commit per successful
|
||||
│ # deploy
|
||||
├── .git/ # one commit per lock
|
||||
│ # change
|
||||
├── flake.nix # generated from agent set
|
||||
└── flake.lock # pins each agent's sha
|
||||
```
|
||||
|
|
@ -411,43 +310,27 @@ wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` /
|
|||
|
||||
### Tag state machine
|
||||
|
||||
Each deploy leaves a tag on the underlying commit inside the applied
|
||||
repo:
|
||||
|
||||
| Tag | When | Annotated? |
|
||||
|---|---|---|
|
||||
| `deployed/<id>` | rebuild succeeded — `main` ff's here | no |
|
||||
| `failed/<id>` | rebuild failed | yes (body = error) |
|
||||
|
||||
hive-c0re plants `deployed/0` at first spawn. `applied/main` is always the
|
||||
latest `deployed/*`. A `failed/` tree stays browsable forever — `git log
|
||||
--tags` in the applied repo is the audit trail. A denied or failed config
|
||||
PR carries no extra state on the forge side: the PR stays open, and the
|
||||
submitter pushes again (or closes it) to retry.
|
||||
hive-c0re plants `deployed/0` on the seed commit at first spawn. A
|
||||
merged config commit that deploys plants no tag: the merged PR on the
|
||||
forge and meta's lock commits record it. A config PR nobody merges
|
||||
carries no extra state on the forge side: the PR stays open, and
|
||||
the submitter pushes again (or closes it) to retry.
|
||||
|
||||
### Dispatch via the job queue
|
||||
|
||||
Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs`
|
||||
— runs as a DAG on the global job queue
|
||||
(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler
|
||||
rather than run inline:
|
||||
Long-running approval work — `UpdateMetaInputs` — runs as a DAG on the
|
||||
global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted
|
||||
by the approval handler rather than run inline:
|
||||
|
||||
| `ApprovalKind` | DAG submitted | source |
|
||||
|---|---|---|
|
||||
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
|
||||
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
|
||||
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
|
||||
|
||||
The DAG carries the originating `approval_id`, surfaced on the node that
|
||||
owns it — for a deploy that's the `DeployWindow` root, so the dashboard
|
||||
renders one approval card, not four. **Every** queued kind resolves
|
||||
through `actions::resolve_approval_dag` when its DAG settles terminal:
|
||||
the deploy's phases are ordinary queue nodes, so the DAG's own terminal
|
||||
state is the authoritative outcome. That hook fires the matching
|
||||
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
|
||||
terminal tag (verifying the tag actually resolves in the applied repo —
|
||||
a pre-merge rejection plants none), posts the failing build log back to
|
||||
the config PR.
|
||||
The DAG carries the originating `approval_id`. **Every** queued kind
|
||||
resolves through `actions::resolve_approval_dag` when its DAG settles
|
||||
terminal: the DAG's own terminal state is the authoritative outcome.
|
||||
That hook fires the matching `HelperEvent::*` via `finish_approval`.
|
||||
|
||||
Two visible consequences:
|
||||
|
||||
|
|
@ -474,30 +357,27 @@ reconcile) DAGs use the same queue but skip the approval plumbing.
|
|||
The bundled `hive-forge` container runs on the swarm's forge host
|
||||
(`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)),
|
||||
and hive-c0re mirrors every agent's applied repo into a
|
||||
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes `applied/main` plus
|
||||
every tag to `agent-configs/<name>` after each ref mutation:
|
||||
the spawn that seeds `deployed/0`, every successful deploy (which
|
||||
plants `deployed/<id>`) or failed build (`failed/<id>`), and a
|
||||
sweep at startup. Pushes are best-effort — a missing or stopped
|
||||
forge never blocks a deploy.
|
||||
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes
|
||||
every tag, then `applied/main`, to `agent-configs/<name>` on every
|
||||
startup sweep and every rebuild. Forge `main` is branch-protected, so
|
||||
the forge routinely refuses a push of an established `main`, which
|
||||
hive-c0re expects. Pushes are best-effort — a missing or stopped forge never
|
||||
blocks a deploy.
|
||||
|
||||
Each agent is a **write collaborator on its own** `agent-configs/<name>`
|
||||
repo — so it can push a branch and open a config PR — but not a member
|
||||
of any other agent's, so it can't reach another agent's config through
|
||||
the forge. Branch protection keeps the agent off the `main` push
|
||||
allowlist and whitelists merging to the `core` user and the `operators`
|
||||
team, so an agent can't fast-forward its own config or self-merge its
|
||||
PR (see the End-to-end flow and
|
||||
[Operator merge in the forge UI](#operator-merge-in-the-forge-ui) above).
|
||||
hive-c0re passes the tokenised push
|
||||
URL inline to `git push`, never writing it into
|
||||
allowlist and allowlists merging to the `operators` team only, so an
|
||||
agent can't fast-forward its own config or self-merge its PR (see
|
||||
[Config changes](#config-changes) above). hive-c0re passes the tokenised
|
||||
push URL inline to `git push`, never writing it into
|
||||
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
|
||||
agent, and a stored token would leak core's admin credential to an
|
||||
agent.
|
||||
|
||||
The dashboard deep-links into this org — a `config repo` link
|
||||
per container row and a `review PR on forge` link per config-PR
|
||||
approval card. See `docs/web-ui/dashboard.md`.
|
||||
per container row. See `docs/web-ui/dashboard.md`.
|
||||
|
||||
### Submitting agent's view of config repos
|
||||
|
||||
|
|
@ -548,8 +428,7 @@ the full `/applied` mount:
|
|||
```sh
|
||||
git -C /agents/<n>/config fetch applied
|
||||
git -C /agents/<n>/config log applied/main --oneline
|
||||
git -C /agents/<n>/config show applied/refs/tags/deployed/<id>
|
||||
git -C /agents/<n>/config show applied/refs/tags/failed/<id> # body = build error
|
||||
git -C /agents/<n>/config show applied/refs/tags/deployed/0 # the seed commit
|
||||
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
|
||||
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
|
||||
|
||||
|
|
@ -631,11 +510,9 @@ as a regular `system` inbox message so it drives a normal claude turn.
|
|||
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
|
||||
**every** approval kind's terminal state. A
|
||||
"FYI, check when convenient" event doesn't need a message — those go
|
||||
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct
|
||||
live dial of the target agent's in-container todo socket (same
|
||||
`UpsertTodo` request in-container producers use); `finish_approval` fires
|
||||
one of these too for `MergeConfigPr`, *in addition to*
|
||||
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
|
||||
through `Coordinator::push_todo` instead, a direct live dial of the target
|
||||
agent's in-container todo socket (same `UpsertTodo` request in-container
|
||||
producers use). Legacy
|
||||
approval rows that predate the submitter column fall back to the
|
||||
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
||||
|
||||
|
|
@ -657,28 +534,17 @@ root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
|||
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
|
||||
`Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check
|
||||
when convenient" events with no reason to drive an immediate turn, so
|
||||
they deliver via `push_todo`/`push_todo_submitter` (see above) instead
|
||||
they deliver via `push_todo` (see above) instead
|
||||
of `HelperEvent`: an `agent_todo_socket` push instead of a broker
|
||||
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
|
||||
and a single free-text `summary` (`rebuilt_todo_summary` renders
|
||||
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
|
||||
|
||||
Optional `sha` field on `ApprovalResolved` carries the canonical
|
||||
hive-c0re-vouched commit sha. Optional `tag` carries the deploy
|
||||
bookkeeping tag — `deployed/<id>` on a successful build or
|
||||
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
|
||||
Both fields are `Option`: `None` on the paths that don't deploy a new
|
||||
commit (meta-update / deny, and the autoupdate
|
||||
sweep's `job_queue::templates::rebuild` reapplying the existing main,
|
||||
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
|
||||
`git show <sha>` against `/applied/<n>/.git` inside the
|
||||
bootstrap container yields the exact tree the sha referenced.
|
||||
|
||||
To add a new lifecycle notice: if it needs to drive an immediate turn
|
||||
(something genuinely urgent, like `ContainerCrash`), add a
|
||||
`HelperEvent` variant + call sites + update `prompts/system.md`'s
|
||||
message-event list. If it's "FYI, check when convenient," call
|
||||
`push_todo`/`push_todo_submitter` directly instead — no new wire type
|
||||
`push_todo` directly instead — no new wire type
|
||||
needed.
|
||||
|
||||
## Autoupdate on startup
|
||||
|
|
|
|||
|
|
@ -56,8 +56,8 @@ power-intent registry:
|
|||
per-agent store — see [`/harness/` contents
|
||||
below](#state-dirs-per-agent) for where reminders (and todos)
|
||||
live.
|
||||
- `approvals` — the queue. `agent / kind (merge_config_pr | spawn |
|
||||
update_meta_inputs | schedule_prompt) /
|
||||
- `approvals` — the queue. `agent / kind (update_meta_inputs |
|
||||
schedule_prompt) /
|
||||
commit_ref / requested_at / status / resolved_at / note`.
|
||||
- `scheduled_prompts` — recurring + one-shot prompt queue.
|
||||
`owner / body / interval_seconds (NULL = one-shot) /
|
||||
|
|
|
|||
|
|
@ -127,8 +127,9 @@ The controller provisions iris's identity, forge user and config repo
|
|||
and start the container. It returns once the scheduler queues the job — watch the
|
||||
swarm UI's job view for progress.
|
||||
|
||||
Later config changes are PRs on `agent-configs/iris`, approved by you. →
|
||||
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md)
|
||||
Later config changes are PRs on `agent-configs/iris`; you merge them on the
|
||||
forge, and the merge deploys them. →
|
||||
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md#config-changes)
|
||||
|
||||
## Optional · Lock the hive dashboard
|
||||
|
||||
|
|
|
|||
|
|
@ -67,22 +67,15 @@ Two things live in the `agent-configs` Forgejo organization:
|
|||
|
||||
- A config repo per agent (`agent-configs/<name>`). The
|
||||
agent is a **write collaborator on its own** repo — it can push
|
||||
config-change branches and open config PRs (Forgejo `pull_request`
|
||||
webhook at `/webhook/config-pr` queues a `MergeConfigPr` approval;
|
||||
`hive-c0re/src/forge/config_pr_poll.rs` re-scans every 5 minutes as a
|
||||
fault-tolerance backstop) — but
|
||||
`main` is branch-protected: the merge whitelist is the `core` user
|
||||
(hive-c0re's merge of an approved `MergeConfigPr`) and the `operators`
|
||||
team (an operator merging in the Forgejo UI, which deploys the merged
|
||||
commit — see
|
||||
[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui)),
|
||||
the approval whitelist is the `operators` team, and the agent can neither
|
||||
push `main` directly nor self-merge. hive-c0re's own merge is
|
||||
fast-forward-only, and hive-c0re never force-pushes (the
|
||||
`push_config` mirror pushes `main` + the add-only
|
||||
status tags without force, and treats a non-fast-forward rejection of
|
||||
`main` after a rolled-back deploy as expected — the forge keeps the
|
||||
approved history, the `failed/<id>` tag records the divergence).
|
||||
config-change branches and open config PRs — but `main` is
|
||||
branch-protected by swarm-controller: the merge and approval
|
||||
allowlists are the `operators` team, and the agent can neither push
|
||||
`main` directly nor self-merge. An operator's merge in the Forgejo UI
|
||||
deploys the merged commit (see
|
||||
[approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes)).
|
||||
hive-c0re never force-pushes: the `push_config` mirror pushes the
|
||||
add-only status tags and `main` without force, and treats a refused
|
||||
`main` push as expected.
|
||||
Repos stay private, so an agent can't read another
|
||||
agent's config. (Agents remain read-only collaborators on `core/meta`.)
|
||||
hive-c0re also references this repo as the agent's **persistent meta
|
||||
|
|
|
|||
|
|
@ -25,7 +25,7 @@ You rarely switch it on yourself. `gateway.enable` defaults to off, and every mo
|
|||
| URL | upstream | when |
|
||||
| --- | --- | --- |
|
||||
| `<hive>/` | dashboard dist (static, from `servedFrontend`) | always |
|
||||
| `<hive>/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always |
|
||||
| `<hive>/api/`, `/health/` | hive-c0re (`7000`) | always |
|
||||
| `<hive>/api/docs/` | themed Swagger UI dist (static) | always |
|
||||
| `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) |
|
||||
| `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` |
|
||||
|
|
@ -130,7 +130,7 @@ services.hyperhive.gateway.auth = {
|
|||
|
||||
`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history.
|
||||
|
||||
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
|
||||
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
|
||||
|
||||
A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first.
|
||||
|
||||
|
|
@ -261,10 +261,9 @@ Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }`
|
|||
|
||||
#### Dashboard: path-based routing (not Accept-header)
|
||||
|
||||
hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
|
||||
hive-c0re serves exactly two prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
|
||||
|
||||
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live.
|
||||
- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded.
|
||||
- `location /health/` → hive-c0re: liveness and readiness.
|
||||
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`.
|
||||
|
||||
|
|
|
|||
|
|
@ -42,49 +42,43 @@ because there is no malformed spec to reject.
|
|||
|
||||
Nix-heavy — hold one of the `buildSlots` permits for the node's duration:
|
||||
|
||||
| Node | Wraps |
|
||||
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead |
|
||||
| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node |
|
||||
| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta |
|
||||
| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion |
|
||||
| `DeployWindow` | resource-holding root of the merge-config-PR deploy subtree — declares the build slot, the lease and the meta window, then completes immediately so its children run under them (see _Approvals_ below) |
|
||||
| `DeployApply` | the deploy's irreversible half: ff-merge the reviewed PR head, two-phase meta deploy, container rebuild |
|
||||
| Node | Wraps |
|
||||
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead |
|
||||
| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node |
|
||||
| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta |
|
||||
| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion |
|
||||
|
||||
Cheap — no build slot:
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
| Node | Behavior |
|
||||
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MergeVerify` | the deploy's pre-merge gate — PR-head drift check, fetch, `verify_commit` eval. Mutates nothing, so a rejection here needs no compensation |
|
||||
| `DeployTail` | the deploy's `AfterAny` compensation + bookkeeping tail: (1) rolls `applied/main` back from the parked `refs/hyperhive/rollback/<id>` and aborts the staged meta lock when the deploy never confirmed good; (2) mirrors whichever deploy tag landed to the forge config repo, always, best-effort; (3) posts the failing build log back onto the config PR when the deploy failed. Named for (2)/(3), which run on the success path too — not `AbortDeploy`. Infallible by construction |
|
||||
| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build |
|
||||
| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#<name>` ref resolves. Store/meta-only, no container yet |
|
||||
| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop |
|
||||
| `Start` | mechanical container start — runtime dir + drop-ins, `start_with_fallback`, MCP listener registration, the manager kick. Fanned out by a `Reconcile` that observed `wanted = Up` and the container down |
|
||||
| `Stop` | mechanical container stop — `nixos-container` kill, MCP listener unregister, the `Killed` manager notify. Fanned out by a `Reconcile` that observed `wanted = Offline` and up |
|
||||
| `StopForUpdate` | mechanical `nixos-container stop` for the profile swap; never touches `wanted`; noop if already stopped |
|
||||
| `RebuildBookkeeping` | the swap's Ok-only bookkeeping tail — rev marker, forge/matrix sync, manager kick, rescan, meta-inputs snapshot; `AfterOk(Swap)` so it runs only on a successful swap (the DAG's `EmitRebuilt` tail node emits the `Rebuilt` manager event, not here). Split out of `Swap` for dashboard visibility + retry granularity, declares no resources of its own — a coordinated child of the `AgentWindow` brace |
|
||||
| `AgentWindow` | pure resource holder — the brace for one agent's rebuild. Declares the build slot + agent lease atomically and holds both for its whole subtree, so `Prebuild` and the `Signal`→`Drain` quiesce window run concurrently instead of one nested under the other. Performs no work; see _Braces_ |
|
||||
| `Signal` | set the graceful fence + kick, so the harness runs one stop-checkpoint turn |
|
||||
| `Drain` | await the harness clearing the fence, bounded by the 3-min graceful-stop timeout; resolves ok either way |
|
||||
| `PauseSignal` | write the pause marker + mark `pause_pending`. No kick, unlike `Signal` — the harness's between-turns poll is already responsive enough, and `Signal`'s kick-message body ("you were just (re)started") would be actively misleading here |
|
||||
| `PauseDrain` | await the harness reporting `PauseAcknowledged`, bounded timeout; best-effort like `Drain` |
|
||||
| `DestroyContainer` | `nixos-container destroy` + un-registration (drop from the roster, clear the ephemeral runtime dir). Runs downstream of a `Stop`, so deliberately excluded from `takes_container_down` — the container is already down by the time it claims |
|
||||
| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step |
|
||||
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
|
||||
| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles |
|
||||
| `FinalizeDeploy` | deploy phase 3 — drop the rollback ref, plant `deployed/<id>`, commit the staged `flake.lock`. The first two git steps are fatal on purpose, so a confirmed-good deploy's outcome and the repo's state can't disagree |
|
||||
| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent |
|
||||
| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome |
|
||||
| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload |
|
||||
| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots |
|
||||
| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless |
|
||||
| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless |
|
||||
| `WebhookRegister` | one-shot boot-time Forgejo webhook registration (`internal/knowledge` push→pull, `agent-configs` PR→approval). No-op until the core token, hive domain, and HMAC secret are all available. Agentless |
|
||||
| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless |
|
||||
| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless |
|
||||
| Node | Behavior |
|
||||
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build |
|
||||
| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#<name>` ref resolves. Store/meta-only, no container yet |
|
||||
| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop |
|
||||
| `Start` | mechanical container start — runtime dir + drop-ins, `start_with_fallback`, MCP listener registration, the manager kick. Fanned out by a `Reconcile` that observed `wanted = Up` and the container down |
|
||||
| `Stop` | mechanical container stop — `nixos-container` kill, MCP listener unregister, the `Killed` manager notify. Fanned out by a `Reconcile` that observed `wanted = Offline` and up |
|
||||
| `StopForUpdate` | mechanical `nixos-container stop` for the profile swap; never touches `wanted`; noop if already stopped |
|
||||
| `RebuildBookkeeping` | the swap's Ok-only bookkeeping tail — rev marker, forge/matrix sync, manager kick, rescan, meta-inputs snapshot; `AfterOk(Swap)` so it runs only on a successful swap (the DAG's `EmitRebuilt` tail node emits the `Rebuilt` manager event, not here). Split out of `Swap` for dashboard visibility + retry granularity, declares no resources of its own — a coordinated child of the `AgentWindow` brace |
|
||||
| `AgentWindow` | pure resource holder — the brace for one agent's rebuild. Declares the build slot + agent lease atomically and holds both for its whole subtree, so `Prebuild` and the `Signal`→`Drain` quiesce window run concurrently instead of one nested under the other. Performs no work; see _Braces_ |
|
||||
| `Signal` | set the graceful fence + kick, so the harness runs one stop-checkpoint turn |
|
||||
| `Drain` | await the harness clearing the fence, bounded by the 3-min graceful-stop timeout; resolves ok either way |
|
||||
| `PauseSignal` | write the pause marker + mark `pause_pending`. No kick, unlike `Signal` — the harness's between-turns poll is already responsive enough, and `Signal`'s kick-message body ("you were just (re)started") would be actively misleading here |
|
||||
| `PauseDrain` | await the harness reporting `PauseAcknowledged`, bounded timeout; best-effort like `Drain` |
|
||||
| `DestroyContainer` | `nixos-container destroy` + un-registration (drop from the roster, clear the ephemeral runtime dir). Runs downstream of a `Stop`, so deliberately excluded from `takes_container_down` — the container is already down by the time it claims |
|
||||
| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step |
|
||||
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
|
||||
| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles |
|
||||
| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent |
|
||||
| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome |
|
||||
| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload |
|
||||
| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots |
|
||||
| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless |
|
||||
| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless |
|
||||
| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless |
|
||||
| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless |
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
|
|
@ -93,26 +87,21 @@ with its commit under its internal `META_LOCK` mutex, so a standalone commit
|
|||
node would open a dirty-working-tree window between nodes.
|
||||
|
||||
Two further layers protect the meta repo across _windows_ that span multiple
|
||||
`META_LOCK` acquisitions — above all the approval deploy's prepare→finalize
|
||||
span, which keeps a bumped `flake.lock` **staged uncommitted** for the whole
|
||||
container build:
|
||||
`META_LOCK` acquisitions:
|
||||
|
||||
- **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue
|
||||
resource declared by every node kind that mutates the meta repo — `MetaSync`,
|
||||
`MetaLock`, `WritePermFile`, `Provision`'s agent registration, and
|
||||
`DeployWindow` — the deploy subtree's root, which holds it across every
|
||||
phase below it (it declares `Resource::MetaWindow`). Two meta
|
||||
`MetaLock`, `WritePermFile` and `Provision`'s agent registration. Two meta
|
||||
mutations can therefore never interleave, so no commit lands inside another
|
||||
node's staged window. It's a queue resource rather than a runtime mutex
|
||||
node's window. It's a queue resource rather than a runtime mutex
|
||||
because a subtree root holds a resource across its whole subtree, which
|
||||
a `MutexGuard` (bounded by one executor fn) can't — that's what lets a
|
||||
multi-node deploy own one window. For the same reason the window must stay
|
||||
a `MutexGuard` (bounded by one executor fn) can't. For the same reason the window must stay
|
||||
_off_ long store-only work: the rebuild's meta preamble is its own
|
||||
`MetaSync` node, a sibling of (never a parent of) `Prebuild`, so the
|
||||
toplevel build runs outside the window and `buildSlots > 1` still gives
|
||||
concurrent rebuilds across agents.
|
||||
- **Path-limited commits**: the targeted meta committers (perm files,
|
||||
topology, lock bumps, finalize) commit `-- <their paths>` with path-scoped
|
||||
topology, lock bumps) commit `-- <their paths>` with path-scoped
|
||||
dirty checks, so even a non-queue caller (boot migration, destroy's
|
||||
`sync_agents`) can never sweep someone else's staged content into its
|
||||
commit.
|
||||
|
|
@ -220,8 +209,8 @@ resources are free. Resources:
|
|||
2. **Per-agent lifecycle lease** — keyed on the **node's** agent (agent is
|
||||
per-node; a DAG can span agents) and globally exclusive per agent across
|
||||
all DAGs: acquired either at a container-affecting node (`SetWanted`,
|
||||
`Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`,
|
||||
`DeployWindow`) on behalf of a whole coordinated subtree; held by the owning
|
||||
`Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`) on
|
||||
behalf of a whole coordinated subtree; held by the owning
|
||||
DAG until it's terminal, so two DAGs never interleave container ops on the
|
||||
same agent. A DAG touching multiple agents holds one lease per agent.
|
||||
(`SetWanted` is a store write, not a container op, but takes the lease anyway
|
||||
|
|
@ -293,36 +282,9 @@ the dashboard renders one recent-builds list and one number bounds it.
|
|||
|
||||
### Approvals
|
||||
|
||||
`MergeConfigPr` approvals ride as a four-node deploy subtree:
|
||||
|
||||
```
|
||||
DeployWindow (root — build slot + lease + meta window, no work of its own)
|
||||
├── MergeVerify drift gate, fetch, verify_commit
|
||||
├── DeployApply AfterOk(verify) park rollback ref, ff-merge, deploy
|
||||
└── DeployTail AfterAny(apply) compensate, mirror to forge
|
||||
```
|
||||
|
||||
The root holds its resources across the whole subtree, so the two-phase
|
||||
`prepare_deploy` / `finalize_deploy` span keeps its staged `flake.lock`
|
||||
protected even though the phases are separate nodes. Splitting them buys
|
||||
three things a single opaque node couldn't have: per-phase visibility on the
|
||||
dashboard, a `MergeVerify` failure that provably mutated nothing, and a
|
||||
compensation step that survives a hive-c0re restart — `DeployApply` parks the pre-merge
|
||||
`applied/main` in `refs/hyperhive/rollback/<approval-id>`, not in a
|
||||
local variable, so `DeployTail` can still undo a half-finished deploy after a
|
||||
crash.
|
||||
|
||||
`DeployWindow` declares all three resources (build slot, lease, meta window)
|
||||
on itself rather than letting each phase declare its own, because the queue
|
||||
acquires a node's resources atomically (all-or-nothing): a child that took
|
||||
the build slot while its parent held the meta window could block waiting for
|
||||
a resource its own parent already committed to, a lock-ordering hazard that
|
||||
one multi-resource root avoids by construction.
|
||||
|
||||
`UpdateMetaInputs` approvals map onto the ordinary
|
||||
`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag`
|
||||
exactly once when **any** approval-carrying DAG settles terminal — deploys
|
||||
included, since their outcome is the DAG's own state (including
|
||||
`UpdateMetaInputs` approvals map onto the ordinary `meta-update` shapes.
|
||||
The scheduler fires `actions::resolve_approval_dag` exactly once when
|
||||
**any** approval-carrying DAG settles terminal (including
|
||||
cancelled-while-queued, which fails the approval instead of dangling it).
|
||||
|
||||
### Wire shape
|
||||
|
|
@ -398,16 +360,12 @@ Key operations:
|
|||
- **`sync_agents`** (idempotent) — render `flake.nix` for the current agent set,
|
||||
init the repo on first call, relock if the rendered contents changed, commit.
|
||||
Called by spawn / destroy / startup migration.
|
||||
- **`prepare_deploy` + `finalize_deploy` / `abort_deploy`** — two-phase for the
|
||||
`MergeConfigPr` deploy path so a failed `nixos-container update` leaves no orphan
|
||||
commit in meta. Prepare writes the new lock without committing; finalize commits
|
||||
with the deploy message; abort restores the lock.
|
||||
- **`lock_update_hyperhive`** — one-shot for the boot-reconcile path (the
|
||||
sweep DAG's `MetaLock` node): bumps the `hyperhive` input lock and commits;
|
||||
the scheduler fans out the agent rebuilds on completion.
|
||||
|
||||
Every public `meta.rs` operation takes the module's internal `META_LOCK`
|
||||
mutex, so concurrent job-queue nodes (and the approval deploy pipeline) never
|
||||
mutex, so concurrent job-queue nodes never
|
||||
race on the repo's `.git/index.lock`.
|
||||
|
||||
---
|
||||
|
|
@ -449,18 +407,6 @@ Sequence for a rebuild DAG (each step is its own queue node):
|
|||
in-container activation script transitions old → new. Holds no build
|
||||
slot, so the next DAG's `Prebuild` overlaps the container boot.
|
||||
|
||||
The approval deploy uses this same chain rather than a rebuild path of its
|
||||
own. Its `DeployApply` node doesn't build: it merges, opens the two-phase
|
||||
meta deploy, and returns the chain above as a subgraph the scheduler grafts
|
||||
into the live DAG under that node. A `FinalizeDeploy` node gated on the
|
||||
graft's completion then plants the deploy tag — so `Reconcile`'s success
|
||||
answers "did the agent come back up?" the same way it does for every
|
||||
other rebuild, instead of a fused inline start.
|
||||
|
||||
The grafted nodes land _inside_ `DeployWindow`'s subtree, so they re-enter
|
||||
the meta window and build slot it already holds rather than deadlocking
|
||||
against it.
|
||||
|
||||
### Cold-start fallback
|
||||
|
||||
`start` after `update` can exit non-zero when packages are **removed** between
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
Long-running work runs through a job graph. The swarm controller keeps one
|
||||
for swarm-level work — creating an agent's identity, forge user and config
|
||||
repo. Each hive's hive-c0re keeps its own for container operations —
|
||||
rebuild, first-spawn, a config-PR deploy, power changes. This page explains
|
||||
rebuild, first-spawn, power changes. This page explains
|
||||
what the job queue _is_, as a general idea, independent of what either uses
|
||||
it for. For the hive-c0re step catalogue and the engineering internals
|
||||
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)
|
||||
|
|
|
|||
|
|
@ -227,9 +227,9 @@ one per hive. It ensures them at start and every five minutes after
|
|||
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
|
||||
owner org;
|
||||
- the empty `operators` merge-gate team in `agents` and `agent-configs`;
|
||||
- the `main` merge gate on every `agent-configs` repo: merge whitelist =
|
||||
the `operators` team and the `core` user, approval whitelist = the
|
||||
`operators` team. The controller leaves a repo with no `main` rule alone;
|
||||
- the `main` merge gate on every `agent-configs` repo: merge and approval
|
||||
whitelists = the `operators` team, and no user. The controller leaves a
|
||||
repo with no `main` rule alone;
|
||||
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
|
||||
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
|
||||
- `internal/docs` (private) and `internal/knowledge` (public, with a
|
||||
|
|
@ -258,7 +258,7 @@ decision, not an event to adjudicate.
|
|||
A `config-pr` delivery reporting a PR merged into `main` queues a deploy
|
||||
of its `merge_commit_sha` on the one hive whose wanted state places the
|
||||
agent; with no such hive, or several, the controller deploys nothing. See
|
||||
[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui).
|
||||
[approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes).
|
||||
|
||||
**`internal/knowledge` is on that path.** The controller's is the only
|
||||
hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a
|
||||
|
|
@ -267,10 +267,10 @@ would take delivery away from the first rather than add a recipient.
|
|||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
**The `agent-configs` org isn't.** Each hive registers its own
|
||||
`pull_request` hook there, so that repo has two — the hive's and the
|
||||
controller's — and **both are expected; don't delete either.** Removing
|
||||
a hive's stops it acting on config PRs; removing the controller's stops
|
||||
**The `agent-configs` org is on it too.** The controller's hook is the
|
||||
only one that acts on config PRs. A hive's `/webhook/config-pr` hook left
|
||||
on the org by an older release delivers to a route no hive serves; delete
|
||||
it in the org's webhook settings. Removing the controller's hook stops
|
||||
forge-UI merges from deploying until its next start recreates it.
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
|
|
|||
|
|
@ -40,8 +40,8 @@ hivectl forge reconcile-config iris --verbose # include the full diff, not
|
|||
applied config checkout and its forge `agent-configs/<agent>` `main`, then
|
||||
reconciles. `--from forge` resets the local checkout to forge `main` (takes
|
||||
effect on the next deploy — it doesn't autorebuild). `--from local` isn't
|
||||
supported yet (forge `main` is core-only branch-protected; resolve via a
|
||||
config PR). With no `--from` it prompts for the direction after the diff.
|
||||
supported yet (forge `main` is branch-protected; resolve via a config
|
||||
PR). With no `--from` it prompts for the direction after the diff.
|
||||
|
||||
## Matrix
|
||||
|
||||
|
|
|
|||
|
|
@ -123,14 +123,17 @@ checkpoints**, not about sandboxing the agent from its own tools:
|
|||
highest-value action. On the **internal forge this is technically enforced,
|
||||
not just convention**: agents can't create repos (`max_repo_creation = 0`),
|
||||
and `main` on an `agent-configs/<name>` repo carries swarm-controller's
|
||||
branch protection: merge allowlisted to the `operators` team, with one
|
||||
approval from it. Existing `agents/<repo>` repos carry the same merge gate.
|
||||
branch protection: merge and approval allowlisted to the `operators` team,
|
||||
which swarm-controller converges on every config repo. An operator's merge
|
||||
there is also what deploys the config. Existing `agents/<repo>` repos carry
|
||||
the same merge gate.
|
||||
An agent (a write collaborator, not a repo admin) can neither change those
|
||||
settings nor merge its own PR. It's **not** set up for external VCS (GitHub
|
||||
etc.), though — there, operator-merge is process + accepted risk, not a
|
||||
technical control.
|
||||
- **Approvals** — config changes, schedule additions, and other
|
||||
blast-radius-y operations route through the operator approval queue
|
||||
- **Approvals** — schedule additions and other blast-radius-y operations
|
||||
route through the operator approval queue; config changes are config PRs
|
||||
an operator merges on the forge
|
||||
(see [`approvals.md`](../agent-lifecycle/approvals.md)).
|
||||
|
||||
### Capability = accepted risk
|
||||
|
|
|
|||
|
|
@ -70,12 +70,7 @@ approval (`Coordinator::notify_submitter`, looked up from the
|
|||
authenticated socket caller at submit time; a row with no
|
||||
recorded submitter falls back to the manager, `ruth`).
|
||||
`ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`,
|
||||
hardcoded — `hive-c0re/src/workers/crash_watch.rs`). A `MergeConfigPr`
|
||||
approval's rebuild additionally pushes a `rebuilt:<agent>` todo to that
|
||||
same submitter (`Coordinator::push_todo_submitter`, `subsystem =
|
||||
"core"`), which wakes a turn (the todo-wake path — see [Turn
|
||||
outcomes](README.md#turn-outcomes)) via a generic "call
|
||||
`get_loose_ends`" prompt rather than the event body itself. Lifecycle
|
||||
hardcoded — `hive-c0re/src/workers/crash_watch.rs`). Lifecycle
|
||||
transitions the job-queue scheduler or crash watcher drive directly —
|
||||
stop/kill, destroy, a flake-rev login or logout state change — reach
|
||||
no individual agent: they publish onto a swarm-wide NATS
|
||||
|
|
|
|||
|
|
@ -48,24 +48,24 @@ and quick links (stats, screen, forge profile). Select the name to open
|
|||
its terminal and watch it work in real time.
|
||||
|
||||
**Approve something an agent is waiting on.** Y3R C4LL is the one tab
|
||||
worth checking regularly — it's everything that needs _you_: approvals
|
||||
for config changes. The tab's count pill tells you at a glance if
|
||||
worth checking regularly — it's everything that needs _you_: an agent's
|
||||
request to schedule a prompt. The tab's count pill tells you at a glance if
|
||||
anything's pending.
|
||||
|
||||
**Approve or reject a config change.** Agent config changes (new
|
||||
packages, env vars, MCP servers) go through an approval queue rather
|
||||
than landing automatically — you'll see them on Y3R C4LL, with a diff
|
||||
of what's changing.
|
||||
**Review a config change.** Agent config changes (new packages, env
|
||||
vars, MCP servers) are pull requests on the agent's `agent-configs`
|
||||
repo. Review and merge them on the forge; the merge deploys the change.
|
||||
See [`approvals.md`](../agent-lifecycle/approvals.md#config-changes).
|
||||
|
||||
**Start, stop, restart, or rebuild an agent.** Select one or more
|
||||
agents on SW4RM (select the icon) and use the selection bar, or use the
|
||||
per-agent `⋮` menu on a single row. Rebuilding re-applies that agent's
|
||||
current config; use it after approving a change, or whenever an agent
|
||||
current config; use it whenever an agent
|
||||
shows as "needs update."
|
||||
|
||||
**Watch a build.** BU1LDS shows the rebuild queue live, plus a
|
||||
streaming log of whatever's currently building. Useful right after
|
||||
approving a change or bumping a flake input.
|
||||
merging a config change or bumping a flake input.
|
||||
|
||||
**Grant or revoke a tool/capability.** P3RM1SS10NS is a checkbox matrix
|
||||
— rows are agents, columns are tool groups or capabilities. Nothing
|
||||
|
|
|
|||
|
|
@ -877,11 +877,10 @@ renderApprovals`) with three stacked sections:
|
|||
right-aligned `requested <N> ago` relative time from
|
||||
`ApprovalView.requested_at`. Glyph and chip vary by kind:
|
||||
|
||||
| kind | glyph | chip | sha shown |
|
||||
|---|---|---|---|
|
||||
| `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) |
|
||||
| `update_meta_inputs` | `↻` | `meta-update` | — |
|
||||
| `schedule_prompt` | `⏱` | `schedule` | — |
|
||||
| kind | glyph | chip |
|
||||
|---|---|---|
|
||||
| `update_meta_inputs` | `↻` | `meta-update` |
|
||||
| `schedule_prompt` | `⏱` | `schedule` |
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
The chip ticks live every second via a `data-requested-at`
|
||||
|
|
@ -889,12 +888,9 @@ renderApprovals`) with three stacked sections:
|
|||
the request has been pending ≥ 1h so a stale approval stands out;
|
||||
the `.stale` class flips precisely at the 3600s boundary rather
|
||||
than at the next `renderApprovals` call.
|
||||
- **what-changed body** — the submitting agent's description, then
|
||||
kind-specific drill-in triggers:
|
||||
- `merge_config_pr`: `↳ review PR on forge ↗` deep-links the
|
||||
config PR into `agent-configs/<agent>/pulls/<pr_number>` (shown
|
||||
only when `forge_present` is true and `pr_number` has a value). The config diff
|
||||
lives on the forge PR itself — no inline diff side-panel.
|
||||
- **what-changed body** — the submitting agent's description, then the
|
||||
kind's payload: the inputs to bump (`update_meta_inputs`) or the prompt to
|
||||
schedule (`schedule_prompt`).
|
||||
- **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a
|
||||
`prompt()` for an optional reason carried to the submitting agent as
|
||||
`HelperEvent::ApprovalResolved.note`.
|
||||
|
|
|
|||
Loading…
Reference in a new issue