docs: config changes are operator merges on the forge
Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.
Refs #4850
This commit is contained in:
parent
0cee0382e9
commit
a88ed9f24e
14 changed files with 195 additions and 396 deletions
|
|
@ -31,7 +31,7 @@ architecture change.
|
||||||
| **identity** | every agent is a swarm-wide principal — SSO subject, forge user, matrix account, secret-store cert identity — addressable as `name@hive.domain` |
|
| **identity** | every agent is a swarm-wide principal — SSO subject, forge user, matrix account, secret-store cert identity — addressable as `name@hive.domain` |
|
||||||
| **secrets** | one OpenBao store; the operator places one mTLS identity per host, and everything else — every agent's credentials included — is fetched from the store under an identity rather than copied by hand |
|
| **secrets** | one OpenBao store; the operator places one mTLS identity per host, and everything else — every agent's credentials included — is fetched from the store under an identity rather than copied by hand |
|
||||||
| **shared services** | one forge, homeserver, SSO, message queue and metrics/logs stack per swarm, each on whichever host you put it |
|
| **shared services** | one forge, homeserver, SSO, message queue and metrics/logs stack per swarm, each on whichever host you put it |
|
||||||
| **config** | git: an agent proposes, the operator approves, the deploy lands as a `deployed/<id>` tag |
|
| **config** | git: an agent opens a PR on its config repo, an operator merges it on the forge, and the merge deploys |
|
||||||
| **runtime** | `claude --print` by default; any [ACP](https://agentclientprotocol.com) agent (e.g. opencode) per agent with `services.hyperhive.agent.runtime = "acp"` |
|
| **runtime** | `claude --print` by default; any [ACP](https://agentclientprotocol.com) agent (e.g. opencode) per agent with `services.hyperhive.agent.runtime = "acp"` |
|
||||||
| **substrate** | each hive runs one `nixos-container` per agent, under an unprivileged daemon with a tiny socket-activated helper for the few root ops |
|
| **substrate** | each hive runs one `nixos-container` per agent, under an unprivileged daemon with a tiny socket-activated helper for the few root ops |
|
||||||
| **watching** | the swarm UI (hives, agents with live terminals, jobs, a cross-repo issue report), Grafana over OTEL, and a per-hive dashboard for host-level detail |
|
| **watching** | the swarm UI (hives, agents with live terminals, jobs, a cross-repo issue report), Grafana over OTEL, and a per-hive dashboard for host-level detail |
|
||||||
|
|
|
||||||
|
|
@ -1,38 +1,32 @@
|
||||||
# Approvals + helper events
|
# Approvals + helper events
|
||||||
|
|
||||||
The approval queue is hyperhive's pivot: nothing that changes the
|
The approval queue is where an agent asks the operator for something it
|
||||||
shape of an agent (its config, whether it exists) happens without an
|
can't do itself: a scheduled prompt, or (legacy rows only) a meta-input
|
||||||
operator selection. The submitting agent — any agent with the `approvals`
|
bump. Config changes don't go through this queue: they're pull requests
|
||||||
tool group, which manages the config of its **direct children** (the
|
on the agent's config repo, which an operator merges on the forge, and
|
||||||
root agent for top-level agents; a sub-manager for its own subtree) — is
|
that merge deploys them. The submitting agent — any agent with the
|
||||||
the policy gate in front of that queue; helper events are how it stays
|
`approvals` tool group, which manages the config of its **direct
|
||||||
informed about what happens after a decision lands.
|
children** (the root agent for top-level agents; a sub-manager for its
|
||||||
|
own subtree) — is the policy gate in front of both; helper events are how
|
||||||
|
it stays informed about what happens after a decision lands.
|
||||||
|
|
||||||
## For operators
|
## For operators
|
||||||
|
|
||||||
Every add/remove/change to an agent lands on your dashboard's Y3R
|
- **Config change** — an agent proposed a change to another agent's
|
||||||
C4LL tab (or `hivectl approvals pending` / `approve <id>` from the
|
config (or its own, via a sub-manager) as a pull request on
|
||||||
CLI) before it takes effect. What you'll see, and what to do with it:
|
`agent-configs/<agent>`. Review it on the forge like any other PR, and
|
||||||
|
merge it there to deploy it. See [Config changes](#config-changes).
|
||||||
- **Config change** (`MergeConfigPr`) — an agent proposed a change to
|
|
||||||
another agent's config (or its own, via a sub-manager) as a forge
|
|
||||||
pull request. Review the diff on the forge — the dashboard card
|
|
||||||
links straight to it, same as reviewing any other PR. Approving
|
|
||||||
triggers the deploy automatically: hive-c0re re-verifies the PR
|
|
||||||
hasn't moved since you looked at it, evaluates it (a dry run,
|
|
||||||
nothing applied yet), merges it, and rebuilds the container. If
|
|
||||||
anything in that chain fails, the change rolls back automatically —
|
|
||||||
the agent stays on its last-good config, no recovery action needed
|
|
||||||
from you.
|
|
||||||
- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump
|
|
||||||
one or more Nix flake inputs (or all of them). Approving runs the
|
|
||||||
update and commits the lock change; it doesn't rebuild anything by
|
|
||||||
itself.
|
|
||||||
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
|
- **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule
|
||||||
a message to one or more inboxes at a future time. You can also add
|
a message to one or more inboxes at a future time. It lands on your
|
||||||
schedules yourself directly from the SCH3DUL3S tab, which skips this
|
dashboard's Y3R C4LL tab (or `hivectl approvals pending` /
|
||||||
approval step entirely — the gate here is specifically for an
|
`approve <id>` from the CLI). You can also add schedules yourself
|
||||||
*agent* asking to schedule something, not for you doing it.
|
directly from the SCH3DUL3S tab, which skips this approval step
|
||||||
|
entirely — the gate here is specifically for an *agent* asking to
|
||||||
|
schedule something, not for you doing it.
|
||||||
|
- **Meta/flake update** (`UpdateMetaInputs`) — legacy: nothing queues
|
||||||
|
this kind, but an existing row still reads back and you can approve it.
|
||||||
|
Approving runs the update and commits the lock change; it doesn't
|
||||||
|
rebuild anything by itself.
|
||||||
|
|
||||||
Don't want to approve something? **Deny it** (`DENY` on the dashboard
|
Don't want to approve something? **Deny it** (`DENY` on the dashboard
|
||||||
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
|
card, or `hivectl approvals deny <id>`) — nothing runs. Either way the
|
||||||
|
|
@ -40,18 +34,16 @@ submitting agent is always notified that the operator denied their request; what
|
||||||
optional is only the reason text, which you can add on the dashboard's
|
optional is only the reason text, which you can add on the dashboard's
|
||||||
prompt (cancelling that prompt aborts the whole deny, not just the
|
prompt (cancelling that prompt aborts the whole deny, not just the
|
||||||
reason) but not from the CLI. Denying is final: a denied approval
|
reason) but not from the CLI. Denying is final: a denied approval
|
||||||
can't be re-approved later, the agent has to submit a fresh one (a new
|
can't be re-approved later, the agent has to submit a fresh request.
|
||||||
PR, a new request).
|
|
||||||
|
|
||||||
Everything below this point is the implementation detail behind that
|
Everything below this point is the implementation detail behind that
|
||||||
flow.
|
flow.
|
||||||
|
|
||||||
## End-to-end approval flow
|
## Config changes
|
||||||
|
|
||||||
Config changes flow through a **forge pull request** on the agent's
|
Config changes flow through a **forge pull request** on the agent's
|
||||||
`agent-configs/<name>` repo — the same surface agents use for code PRs.
|
`agent-configs/<name>` repo — the same surface agents use for code PRs.
|
||||||
No bespoke MCP tool exists for config changes: opening the PR IS the
|
No MCP tool exists for config changes: opening the PR IS the request.
|
||||||
request.
|
|
||||||
|
|
||||||
1. The submitting agent (the child's parent, holding the `approvals`
|
1. The submitting agent (the child's parent, holding the `approvals`
|
||||||
tool group) **clones** `agent-configs/<name>`, edits it there (any
|
tool group) **clones** `agent-configs/<name>`, edits it there (any
|
||||||
|
|
@ -59,89 +51,40 @@ request.
|
||||||
with its own git identity, and pushes a branch + opens a PR with
|
with its own git identity, and pushes a branch + opens a PR with
|
||||||
`hive-forge` — the same way it would change any other repo.
|
`hive-forge` — the same way it would change any other repo.
|
||||||
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
|
The bind-mounted `/agents/<name>/config/` is a **copy for reading** a
|
||||||
config, not the tree to edit: authoring in place there produces no PR
|
config, not the tree to edit: authoring in place there produces no PR.
|
||||||
and no approval. (it's currently mounted read-write, which is a
|
(it's currently mounted read-write, which is a defect tracked
|
||||||
defect tracked separately, not an authoring path.)
|
separately, not an authoring path.)
|
||||||
Branch protection (the agent isn't on the `main` push allowlist;
|
Branch protection (the agent isn't on the `main` push allowlist;
|
||||||
merge allowlist = `core` user + `operators` team; approvals allowlist
|
merge and approvals allowlist = `operators` team; see "Forge mirror"
|
||||||
= `operators` team; see "Forge mirror" below) makes the agent a write
|
below) makes the agent a write collaborator that **can't merge its own
|
||||||
collaborator that **can't merge its own config PR**.
|
config PR**.
|
||||||
2. hive-c0re's `/webhook/config-pr` endpoint receives the Forgejo
|
2. An operator reviews the PR **on the forge** (native diff, threaded
|
||||||
`pull_request` event (opened / synchronized / reopened) and queues a
|
comments, CI status) and merges it there. Merging needs membership of
|
||||||
`MergeConfigPr` approval; a poll fallback catches any missed webhook.
|
the `operators` team in `agent-configs`. The team starts empty: add
|
||||||
The approval row stores the PR **number** (`commit_ref`) and the PR
|
each operator by hand in the forge UI.
|
||||||
**head sha at queue time** (`fetched_sha` — the "reviewed" sha). If
|
3. swarm-controller gets the `agent-configs` org's `pull_request`
|
||||||
the PR head later moves, a fresh approval pinned to the new head supersedes
|
|
||||||
the stale one, so the operator always reviews what will
|
|
||||||
actually deploy.
|
|
||||||
3. The operator reviews the PR **on the forge** (native diff, threaded
|
|
||||||
comments, CI status) and sees a matching card on the dashboard with a
|
|
||||||
"review PR on forge" deep link. They select ◆ APPR0VE (or
|
|
||||||
`hivectl approvals approve <id>` on the CLI) once satisfied.
|
|
||||||
4. On approve, a deploy DAG runs three phases under a resource-holding
|
|
||||||
`DeployWindow` root (see *Queue templates* below):
|
|
||||||
- `MergeVerify` re-reads the live PR head and **aborts if it drifted**
|
|
||||||
from the reviewed `fetched_sha` (the submitter must push again,
|
|
||||||
which queues a fresh approval); then fetches that head into the
|
|
||||||
applied repo and **eval-verifies** it — a flake eval on a throwaway
|
|
||||||
checkout. This is the trust gate: it relies on c0re's own eval, not
|
|
||||||
on any in-repo (agent-forgeable) signal like a CI status. Nothing is
|
|
||||||
mutated in this phase, so a rejection here leaves the forge and the
|
|
||||||
applied repo exactly as they were.
|
|
||||||
- `DeployApply` parks the pre-merge `applied/main` in
|
|
||||||
`refs/hyperhive/rollback/<approval-id>`, then fast-forward-merges
|
|
||||||
the reviewed head to the forge config repo's `main` (this IS the
|
|
||||||
merge — a `core`-authenticated ff-merge pinned to the reviewed sha,
|
|
||||||
so a moved PR head can't substitute bytes), and runs the deploy
|
|
||||||
proper (`deploy_applied_target`): ff `applied/main`, two-phase meta
|
|
||||||
deploy, container rebuild. On success it drops the rollback ref and
|
|
||||||
plants `deployed/<id>`.
|
|
||||||
- `DeployTail` runs on **every** outcome, including a cancel-cascade.
|
|
||||||
If the rollback ref survived, the deploy never confirmed good: it
|
|
||||||
rolls `applied/main` back, resyncs the working tree, and aborts the
|
|
||||||
staged meta lock, so the agent stays on its last-good tree. Then it
|
|
||||||
mirrors the config repo (and its new deploy tag) to the forge.
|
|
||||||
|
|
||||||
The rollback state lives in a **git ref, not a local variable**, on
|
|
||||||
purpose: hive-c0re can restart between the apply and the tail, and the
|
|
||||||
tail still has to know what to undo when it does.
|
|
||||||
5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the
|
|
||||||
**submitting agent's** inbox via `notify_submitter`, carrying both the
|
|
||||||
canonical sha and the terminal tag (the approval row carries a
|
|
||||||
`submitter` column recording the agent the change is for).
|
|
||||||
|
|
||||||
### Operator merge in the forge UI
|
|
||||||
|
|
||||||
An operator can also merge a config PR straight in the Forgejo UI. That
|
|
||||||
deploys the merged commit on the hive that runs the agent:
|
|
||||||
|
|
||||||
1. swarm-controller gets the `agent-configs` org's `pull_request`
|
|
||||||
delivery. A `closed` event with `merged: true` and base branch `main`
|
delivery. A `closed` event with `merged: true` and base branch `main`
|
||||||
names the commit in `merge_commit_sha`.
|
names the commit in `merge_commit_sha`.
|
||||||
2. The controller looks up which hive's wanted state places the agent
|
4. The controller looks up which hive's wanted state places the agent
|
||||||
and publishes a deploy request carrying that commit on that hive's
|
and publishes a deploy request carrying that commit on that hive's
|
||||||
deploy subject. With no such hive, or more than one, it deploys
|
deploy subject. With no such hive, or more than one, it deploys
|
||||||
nothing and logs a warning naming them.
|
nothing and logs a warning naming them.
|
||||||
3. If the hive's `applied/main` already is that commit — a dashboard
|
5. If the hive's `applied/main` already is that commit, it does nothing.
|
||||||
approval merges and deploys its own PR — it does nothing. Otherwise
|
Otherwise it fetches the config repo's `main`, requires the commit to
|
||||||
it fetches the config repo's `main`, requires the commit to descend
|
descend from `applied/main`, fast-forwards `applied/main` to it, and
|
||||||
from `applied/main`, fast-forwards `applied/main` to it, and queues a
|
queues a rebuild. A failure before the rebuild (the fetch, or a commit
|
||||||
rebuild. A failure before the rebuild (the fetch, or a commit that
|
that doesn't descend) deploys nothing and posts the error as a comment
|
||||||
doesn't descend) deploys nothing and posts the error as a comment on
|
on the merged PR. An agent with no container on that hive yet ignores
|
||||||
the merged PR. An agent with no container on that hive yet ignores
|
|
||||||
the commit; its first deploy builds what it seeds.
|
the commit; its first deploy builds what it seeds.
|
||||||
|
|
||||||
This path runs **no eval-verify**. A merged config that doesn't
|
This path runs **no eval-verify**. A merged config that doesn't
|
||||||
evaluate or build fails the rebuild, and `applied/main` stays at that
|
evaluate or build fails the rebuild, and `applied/main` stays at that
|
||||||
commit, so the agent's rebuilds keep failing until a fix merges. A
|
commit, so the agent's rebuilds keep failing until a fix merges. A
|
||||||
failed rebuild shows on the hive like any other and isn't commented on
|
failed rebuild shows on the hive like any other and isn't commented on
|
||||||
the PR.
|
the PR. A merge whose webhook delivery never reaches the controller
|
||||||
|
deploys nothing.
|
||||||
|
|
||||||
Merging needs membership of the `operators` team in `agent-configs`.
|
## Approval queue
|
||||||
The team starts empty: add each operator by hand in the forge UI. A
|
|
||||||
merge whose webhook delivery never reaches the controller deploys
|
|
||||||
nothing. The hive's own poll cancels the dashboard card for that PR
|
|
||||||
with the note `PR merged/closed outside the approval`.
|
|
||||||
|
|
||||||
### Withdrawing a pending approval
|
### Withdrawing a pending approval
|
||||||
|
|
||||||
|
|
@ -166,31 +109,16 @@ and no hive can originate an agent. The swarm controller's
|
||||||
`agent.nix` template, then asks the target hive to deploy it.
|
`agent.nix` template, then asks the target hive to deploy it.
|
||||||
|
|
||||||
Changing what the template seeded isn't a special case: like every
|
Changing what the template seeded isn't a special case: like every
|
||||||
later change, it's a PR on that config repo (`MergeConfigPr`), made
|
later change, it's a PR on that config repo, made from a clone and
|
||||||
from a clone, reviewed and approved by the operator. The PR flow is
|
merged by an operator on the forge (see [Config changes](#config-changes)).
|
||||||
the one path — an operator can equally drive both steps herself
|
|
||||||
through the web UI or the forge.
|
|
||||||
|
|
||||||
### Approval kinds (wire shapes)
|
### Approval kinds (wire shapes)
|
||||||
|
|
||||||
`ApprovalKind` carries three variants; each maps to a different
|
`ApprovalKind` carries two variants; each maps to a different
|
||||||
`commit_ref` encoding because `ApprovalKind` overloads that field as
|
`commit_ref` encoding because `ApprovalKind` overloads that field as
|
||||||
the kind-specific payload carrier.
|
the kind-specific payload carrier.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
- `MergeConfigPr` — the config-change flow. Triggered automatically:
|
|
||||||
when an agent opens (or force-pushes) a PR on its
|
|
||||||
`agent-configs/<agent>` forge repo, hive-c0re's `/webhook/config-pr`
|
|
||||||
endpoint receives the Forgejo pull_request event and queues this
|
|
||||||
approval row. No MCP tool call needed — the forge PR IS the request.
|
|
||||||
`commit_ref` stores the **PR number** (decimal), and `fetched_sha` is
|
|
||||||
the PR **head sha at queue time** (the "reviewed" sha). On approve,
|
|
||||||
the deploy DAG's `MergeVerify` phase re-reads the live PR head and
|
|
||||||
aborts if it drifted from `fetched_sha` (submitter must push again to
|
|
||||||
re-trigger), then fetches that head into the applied repo and
|
|
||||||
eval-verifies it; `DeployApply` fast-forward-merges the forge config
|
|
||||||
repo's `main` to it (the merge) and runs `deploy_applied_target`;
|
|
||||||
`DeployTail` compensates on failure. Never a first spawn.
|
|
||||||
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
|
- `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs
|
||||||
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
|
array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs,
|
||||||
etc.). hive-c0re sets the `agent` field to the requesting root agent.
|
etc.). hive-c0re sets the `agent` field to the requesting root agent.
|
||||||
|
|
@ -303,55 +231,26 @@ declares one flake input per agent (`agent-<n>.url =
|
||||||
Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
|
Containers run against `--flake /var/lib/hyperhive/meta#<n>`.
|
||||||
|
|
||||||
The declared input url is the agent's **forge config repo** (the
|
The declared input url is the agent's **forge config repo** (the
|
||||||
same `agent-configs/<n>` the config-PR flow lands approved changes
|
same `agent-configs/<n>` config PRs merge into), so the meta flake
|
||||||
on), so the meta flake references a reviewable, reproducible source
|
references a reviewable, reproducible source rather than a local
|
||||||
rather than a local checkout. hive-c0re authenticates that
|
checkout. hive-c0re authenticates that `git+http` fetch via a git
|
||||||
`git+http` fetch via a git credential helper that reads the live
|
credential helper that reads the live forge-core token — no token in
|
||||||
forge-core token — no token in the url or the lock. The deploy and
|
the url or the lock. The rebuild paths, however, do **not** re-lock
|
||||||
manual-rebuild paths, however, do **not** re-lock from the forge:
|
from the forge: they `--override-input agent-<n>
|
||||||
they `--override-input agent-<n>
|
|
||||||
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
|
git+file:///var/lib/hyperhive/applied/<n>`, locking the exact config
|
||||||
that `verify_commit` gated and `applied/<n>/main` was
|
`applied/<n>/main` was fast-forwarded to. That keeps a rebuild
|
||||||
fast-forwarded to. That keeps a deploy/rebuild reproducible and
|
reproducible and independent of forge reachability — rebuilds fire on
|
||||||
independent of forge reachability — rebuilds fire on crash-restart
|
crash-restart and meta bumps, not just config merges — while the
|
||||||
and meta bumps, not just config PRs — while the declared url stays
|
declared url stays the forge. `sync_agents` re-renders + re-locks the
|
||||||
the forge. `sync_agents` re-renders + re-locks the persistent input;
|
persistent input; a plain `nix flake lock` leaves an existing applied
|
||||||
a plain `nix flake lock` leaves an existing applied override in
|
override in place (it only re-locks when the declared url itself
|
||||||
place (it only re-locks when the declared url itself changes), so
|
changes), so the forge-declared / applied-deployed split is stable.
|
||||||
the forge-declared / applied-deployed split is stable.
|
|
||||||
|
|
||||||
Per-deploy lock flow (two-phase), spread across the deploy subtree's
|
Lock updates are single-phase and commit when the lock changed:
|
||||||
nodes — each phase is its own node, so the queue can show which one is
|
`meta::lock_update_for_rebuild(name)` relocks one agent's input for a
|
||||||
running and a restart resumes at node granularity:
|
relocking rebuild (the manual `↻ R3BU1LD` button, and the rebuild a
|
||||||
|
merged config commit queues), and `meta::lock_update_hyperhive()` is the
|
||||||
1. `DeployApply` → `meta::prepare_deploy(name)` runs
|
autoupdate flake-rev bump (one shot before per-agent rebuilds).
|
||||||
`nix flake lock --update-input agent-<n>` without
|
|
||||||
committing. Working tree of meta now points the input at
|
|
||||||
`applied/<n>/main` (which the deploy already fast-forwarded to
|
|
||||||
the reviewed PR head).
|
|
||||||
2. The rebuild subgraph `DeployApply` grows into the DAG builds and
|
|
||||||
swaps the container (`AgentWindow` bracing `Prebuild → StopForUpdate
|
|
||||||
→ Swap → RebuildBookkeeping`, plus `Reconcile`). Nix evaluates
|
|
||||||
against the staged lock.
|
|
||||||
3. On success — `FinalizeDeploy` drops the rollback ref, plants
|
|
||||||
`deployed/<id>`, then `meta::finalize_deploy(name, sha, "deployed/
|
|
||||||
<id>")` stages `flake.lock` and commits with
|
|
||||||
`deploy <n> deployed/<id> <sha12>`. Meta's git log gains
|
|
||||||
one entry per successful deploy.
|
|
||||||
4. On failure — the `DeployTail` node runs `meta::abort_deploy()`
|
|
||||||
(`git restore flake.lock`) so the meta history shows only
|
|
||||||
successes; the failure stays as an annotated `failed/<id>`
|
|
||||||
tag in `applied/<n>`. The tail runs on every outcome, so this
|
|
||||||
also covers a hive-c0re restart mid-build: the staged lock is
|
|
||||||
dropped and `applied/main` rolled back from the parked
|
|
||||||
`refs/hyperhive/rollback/<id>`.
|
|
||||||
|
|
||||||
Single-phase variants exist for paths without
|
|
||||||
rollback semantics: `meta::lock_update_for_rebuild(name)` for
|
|
||||||
the manual `↻ R3BU1LD` button (commits if the lock changed)
|
|
||||||
and `meta::lock_update_hyperhive()` for the
|
|
||||||
autoupdate flake-rev bump (one shot before per-agent
|
|
||||||
rebuilds, commits if the lock changed).
|
|
||||||
|
|
||||||
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
|
`meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive`
|
||||||
carries `hyperhive_flake`, `dashboard_port`, and the rest of the
|
carries `hyperhive_flake`, `dashboard_port`, and the rest of the
|
||||||
|
|
@ -390,8 +289,8 @@ per container row.
|
||||||
└── <other committed files> # also tracked
|
└── <other committed files> # also tracked
|
||||||
|
|
||||||
/var/lib/hyperhive/meta/ swarm-wide flake — core
|
/var/lib/hyperhive/meta/ swarm-wide flake — core
|
||||||
├── .git/ # one commit per successful
|
├── .git/ # one commit per lock
|
||||||
│ # deploy
|
│ # change
|
||||||
├── flake.nix # generated from agent set
|
├── flake.nix # generated from agent set
|
||||||
└── flake.lock # pins each agent's sha
|
└── flake.lock # pins each agent's sha
|
||||||
```
|
```
|
||||||
|
|
@ -411,43 +310,27 @@ wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` /
|
||||||
|
|
||||||
### Tag state machine
|
### Tag state machine
|
||||||
|
|
||||||
Each deploy leaves a tag on the underlying commit inside the applied
|
hive-c0re plants `deployed/0` on the seed commit at first spawn. A
|
||||||
repo:
|
merged config commit that deploys plants no tag: the merged PR on the
|
||||||
|
forge and meta's lock commits record it. A config PR nobody merges
|
||||||
| Tag | When | Annotated? |
|
carries no extra state on the forge side: the PR stays open, and
|
||||||
|---|---|---|
|
the submitter pushes again (or closes it) to retry.
|
||||||
| `deployed/<id>` | rebuild succeeded — `main` ff's here | no |
|
|
||||||
| `failed/<id>` | rebuild failed | yes (body = error) |
|
|
||||||
|
|
||||||
hive-c0re plants `deployed/0` at first spawn. `applied/main` is always the
|
|
||||||
latest `deployed/*`. A `failed/` tree stays browsable forever — `git log
|
|
||||||
--tags` in the applied repo is the audit trail. A denied or failed config
|
|
||||||
PR carries no extra state on the forge side: the PR stays open, and the
|
|
||||||
submitter pushes again (or closes it) to retry.
|
|
||||||
|
|
||||||
### Dispatch via the job queue
|
### Dispatch via the job queue
|
||||||
|
|
||||||
Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs`
|
Long-running approval work — `UpdateMetaInputs` — runs as a DAG on the
|
||||||
— runs as a DAG on the global job queue
|
global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted
|
||||||
(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler
|
by the approval handler rather than run inline:
|
||||||
rather than run inline:
|
|
||||||
|
|
||||||
| `ApprovalKind` | DAG submitted | source |
|
| `ApprovalKind` | DAG submitted | source |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` |
|
|
||||||
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
|
| `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` |
|
||||||
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
|
| `SchedulePrompt` | — runs inline (single sqlite insert) | — |
|
||||||
|
|
||||||
The DAG carries the originating `approval_id`, surfaced on the node that
|
The DAG carries the originating `approval_id`. **Every** queued kind
|
||||||
owns it — for a deploy that's the `DeployWindow` root, so the dashboard
|
resolves through `actions::resolve_approval_dag` when its DAG settles
|
||||||
renders one approval card, not four. **Every** queued kind resolves
|
terminal: the DAG's own terminal state is the authoritative outcome.
|
||||||
through `actions::resolve_approval_dag` when its DAG settles terminal:
|
That hook fires the matching `HelperEvent::*` via `finish_approval`.
|
||||||
the deploy's phases are ordinary queue nodes, so the DAG's own terminal
|
|
||||||
state is the authoritative outcome. That hook fires the matching
|
|
||||||
`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's
|
|
||||||
terminal tag (verifying the tag actually resolves in the applied repo —
|
|
||||||
a pre-merge rejection plants none), posts the failing build log back to
|
|
||||||
the config PR.
|
|
||||||
|
|
||||||
Two visible consequences:
|
Two visible consequences:
|
||||||
|
|
||||||
|
|
@ -474,30 +357,27 @@ reconcile) DAGs use the same queue but skip the approval plumbing.
|
||||||
The bundled `hive-forge` container runs on the swarm's forge host
|
The bundled `hive-forge` container runs on the swarm's forge host
|
||||||
(`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)),
|
(`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)),
|
||||||
and hive-c0re mirrors every agent's applied repo into a
|
and hive-c0re mirrors every agent's applied repo into a
|
||||||
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes `applied/main` plus
|
private `agent-configs` Forgejo org. `forge::push_config(<name>)` pushes
|
||||||
every tag to `agent-configs/<name>` after each ref mutation:
|
every tag, then `applied/main`, to `agent-configs/<name>` on every
|
||||||
the spawn that seeds `deployed/0`, every successful deploy (which
|
startup sweep and every rebuild. Forge `main` is branch-protected, so
|
||||||
plants `deployed/<id>`) or failed build (`failed/<id>`), and a
|
the forge routinely refuses a push of an established `main`, which
|
||||||
sweep at startup. Pushes are best-effort — a missing or stopped
|
hive-c0re expects. Pushes are best-effort — a missing or stopped forge never
|
||||||
forge never blocks a deploy.
|
blocks a deploy.
|
||||||
|
|
||||||
Each agent is a **write collaborator on its own** `agent-configs/<name>`
|
Each agent is a **write collaborator on its own** `agent-configs/<name>`
|
||||||
repo — so it can push a branch and open a config PR — but not a member
|
repo — so it can push a branch and open a config PR — but not a member
|
||||||
of any other agent's, so it can't reach another agent's config through
|
of any other agent's, so it can't reach another agent's config through
|
||||||
the forge. Branch protection keeps the agent off the `main` push
|
the forge. Branch protection keeps the agent off the `main` push
|
||||||
allowlist and whitelists merging to the `core` user and the `operators`
|
allowlist and allowlists merging to the `operators` team only, so an
|
||||||
team, so an agent can't fast-forward its own config or self-merge its
|
agent can't fast-forward its own config or self-merge its PR (see
|
||||||
PR (see the End-to-end flow and
|
[Config changes](#config-changes) above). hive-c0re passes the tokenised
|
||||||
[Operator merge in the forge UI](#operator-merge-in-the-forge-ui) above).
|
push URL inline to `git push`, never writing it into
|
||||||
hive-c0re passes the tokenised push
|
|
||||||
URL inline to `git push`, never writing it into
|
|
||||||
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
|
`applied/<n>/.git/config`; that repo is RO-bind-mounted into the root
|
||||||
agent, and a stored token would leak core's admin credential to an
|
agent, and a stored token would leak core's admin credential to an
|
||||||
agent.
|
agent.
|
||||||
|
|
||||||
The dashboard deep-links into this org — a `config repo` link
|
The dashboard deep-links into this org — a `config repo` link
|
||||||
per container row and a `review PR on forge` link per config-PR
|
per container row. See `docs/web-ui/dashboard.md`.
|
||||||
approval card. See `docs/web-ui/dashboard.md`.
|
|
||||||
|
|
||||||
### Submitting agent's view of config repos
|
### Submitting agent's view of config repos
|
||||||
|
|
||||||
|
|
@ -548,8 +428,7 @@ the full `/applied` mount:
|
||||||
```sh
|
```sh
|
||||||
git -C /agents/<n>/config fetch applied
|
git -C /agents/<n>/config fetch applied
|
||||||
git -C /agents/<n>/config log applied/main --oneline
|
git -C /agents/<n>/config log applied/main --oneline
|
||||||
git -C /agents/<n>/config show applied/refs/tags/deployed/<id>
|
git -C /agents/<n>/config show applied/refs/tags/deployed/0 # the seed commit
|
||||||
git -C /agents/<n>/config show applied/refs/tags/failed/<id> # body = build error
|
|
||||||
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
|
git -C /agents/<n>/config show applied/refs/tags/denied/<id> # body = operator note
|
||||||
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
|
git -C /agents/<n>/config rebase applied/main # base in-flight work on what's deployed
|
||||||
|
|
||||||
|
|
@ -631,11 +510,9 @@ as a regular `system` inbox message so it drives a normal claude turn.
|
||||||
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
|
`finish_approval` fires an `ApprovalResolved` HelperEvent this way for
|
||||||
**every** approval kind's terminal state. A
|
**every** approval kind's terminal state. A
|
||||||
"FYI, check when convenient" event doesn't need a message — those go
|
"FYI, check when convenient" event doesn't need a message — those go
|
||||||
through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct
|
through `Coordinator::push_todo` instead, a direct live dial of the target
|
||||||
live dial of the target agent's in-container todo socket (same
|
agent's in-container todo socket (same `UpsertTodo` request in-container
|
||||||
`UpsertTodo` request in-container producers use); `finish_approval` fires
|
producers use). Legacy
|
||||||
one of these too for `MergeConfigPr`, *in addition to*
|
|
||||||
the `ApprovalResolved` HelperEvent above, not instead of it. Legacy
|
|
||||||
approval rows that predate the submitter column fall back to the
|
approval rows that predate the submitter column fall back to the
|
||||||
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
||||||
|
|
||||||
|
|
@ -657,28 +534,17 @@ root agent. Variants (`hive_sh4re::manager::HelperEvent`):
|
||||||
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
|
The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`,
|
||||||
`Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check
|
`Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check
|
||||||
when convenient" events with no reason to drive an immediate turn, so
|
when convenient" events with no reason to drive an immediate turn, so
|
||||||
they deliver via `push_todo`/`push_todo_submitter` (see above) instead
|
they deliver via `push_todo` (see above) instead
|
||||||
of `HelperEvent`: an `agent_todo_socket` push instead of a broker
|
of `HelperEvent`: an `agent_todo_socket` push instead of a broker
|
||||||
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
|
message, `subsystem = "core"`, `key = "<event>:<agent>"` for dedup,
|
||||||
and a single free-text `summary` (`rebuilt_todo_summary` renders
|
and a single free-text `summary` (`rebuilt_todo_summary` renders
|
||||||
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
|
`Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string).
|
||||||
|
|
||||||
Optional `sha` field on `ApprovalResolved` carries the canonical
|
|
||||||
hive-c0re-vouched commit sha. Optional `tag` carries the deploy
|
|
||||||
bookkeeping tag — `deployed/<id>` on a successful build or
|
|
||||||
`failed/<id>` on a failed one, planted by the `MergeConfigPr` deploy.
|
|
||||||
Both fields are `Option`: `None` on the paths that don't deploy a new
|
|
||||||
commit (meta-update / deny, and the autoupdate
|
|
||||||
sweep's `job_queue::templates::rebuild` reapplying the existing main,
|
|
||||||
or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set,
|
|
||||||
`git show <sha>` against `/applied/<n>/.git` inside the
|
|
||||||
bootstrap container yields the exact tree the sha referenced.
|
|
||||||
|
|
||||||
To add a new lifecycle notice: if it needs to drive an immediate turn
|
To add a new lifecycle notice: if it needs to drive an immediate turn
|
||||||
(something genuinely urgent, like `ContainerCrash`), add a
|
(something genuinely urgent, like `ContainerCrash`), add a
|
||||||
`HelperEvent` variant + call sites + update `prompts/system.md`'s
|
`HelperEvent` variant + call sites + update `prompts/system.md`'s
|
||||||
message-event list. If it's "FYI, check when convenient," call
|
message-event list. If it's "FYI, check when convenient," call
|
||||||
`push_todo`/`push_todo_submitter` directly instead — no new wire type
|
`push_todo` directly instead — no new wire type
|
||||||
needed.
|
needed.
|
||||||
|
|
||||||
## Autoupdate on startup
|
## Autoupdate on startup
|
||||||
|
|
|
||||||
|
|
@ -56,8 +56,8 @@ power-intent registry:
|
||||||
per-agent store — see [`/harness/` contents
|
per-agent store — see [`/harness/` contents
|
||||||
below](#state-dirs-per-agent) for where reminders (and todos)
|
below](#state-dirs-per-agent) for where reminders (and todos)
|
||||||
live.
|
live.
|
||||||
- `approvals` — the queue. `agent / kind (merge_config_pr | spawn |
|
- `approvals` — the queue. `agent / kind (update_meta_inputs |
|
||||||
update_meta_inputs | schedule_prompt) /
|
schedule_prompt) /
|
||||||
commit_ref / requested_at / status / resolved_at / note`.
|
commit_ref / requested_at / status / resolved_at / note`.
|
||||||
- `scheduled_prompts` — recurring + one-shot prompt queue.
|
- `scheduled_prompts` — recurring + one-shot prompt queue.
|
||||||
`owner / body / interval_seconds (NULL = one-shot) /
|
`owner / body / interval_seconds (NULL = one-shot) /
|
||||||
|
|
|
||||||
|
|
@ -127,8 +127,9 @@ The controller provisions iris's identity, forge user and config repo
|
||||||
and start the container. It returns once the scheduler queues the job — watch the
|
and start the container. It returns once the scheduler queues the job — watch the
|
||||||
swarm UI's job view for progress.
|
swarm UI's job view for progress.
|
||||||
|
|
||||||
Later config changes are PRs on `agent-configs/iris`, approved by you. →
|
Later config changes are PRs on `agent-configs/iris`; you merge them on the
|
||||||
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md)
|
forge, and the merge deploys them. →
|
||||||
|
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md#config-changes)
|
||||||
|
|
||||||
## Optional · Lock the hive dashboard
|
## Optional · Lock the hive dashboard
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -67,22 +67,15 @@ Two things live in the `agent-configs` Forgejo organization:
|
||||||
|
|
||||||
- A config repo per agent (`agent-configs/<name>`). The
|
- A config repo per agent (`agent-configs/<name>`). The
|
||||||
agent is a **write collaborator on its own** repo — it can push
|
agent is a **write collaborator on its own** repo — it can push
|
||||||
config-change branches and open config PRs (Forgejo `pull_request`
|
config-change branches and open config PRs — but `main` is
|
||||||
webhook at `/webhook/config-pr` queues a `MergeConfigPr` approval;
|
branch-protected by swarm-controller: the merge and approval
|
||||||
`hive-c0re/src/forge/config_pr_poll.rs` re-scans every 5 minutes as a
|
allowlists are the `operators` team, and the agent can neither push
|
||||||
fault-tolerance backstop) — but
|
`main` directly nor self-merge. An operator's merge in the Forgejo UI
|
||||||
`main` is branch-protected: the merge whitelist is the `core` user
|
deploys the merged commit (see
|
||||||
(hive-c0re's merge of an approved `MergeConfigPr`) and the `operators`
|
[approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes)).
|
||||||
team (an operator merging in the Forgejo UI, which deploys the merged
|
hive-c0re never force-pushes: the `push_config` mirror pushes the
|
||||||
commit — see
|
add-only status tags and `main` without force, and treats a refused
|
||||||
[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui)),
|
`main` push as expected.
|
||||||
the approval whitelist is the `operators` team, and the agent can neither
|
|
||||||
push `main` directly nor self-merge. hive-c0re's own merge is
|
|
||||||
fast-forward-only, and hive-c0re never force-pushes (the
|
|
||||||
`push_config` mirror pushes `main` + the add-only
|
|
||||||
status tags without force, and treats a non-fast-forward rejection of
|
|
||||||
`main` after a rolled-back deploy as expected — the forge keeps the
|
|
||||||
approved history, the `failed/<id>` tag records the divergence).
|
|
||||||
Repos stay private, so an agent can't read another
|
Repos stay private, so an agent can't read another
|
||||||
agent's config. (Agents remain read-only collaborators on `core/meta`.)
|
agent's config. (Agents remain read-only collaborators on `core/meta`.)
|
||||||
hive-c0re also references this repo as the agent's **persistent meta
|
hive-c0re also references this repo as the agent's **persistent meta
|
||||||
|
|
|
||||||
|
|
@ -25,7 +25,7 @@ You rarely switch it on yourself. `gateway.enable` defaults to off, and every mo
|
||||||
| URL | upstream | when |
|
| URL | upstream | when |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `<hive>/` | dashboard dist (static, from `servedFrontend`) | always |
|
| `<hive>/` | dashboard dist (static, from `servedFrontend`) | always |
|
||||||
| `<hive>/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always |
|
| `<hive>/api/`, `/health/` | hive-c0re (`7000`) | always |
|
||||||
| `<hive>/api/docs/` | themed Swagger UI dist (static) | always |
|
| `<hive>/api/docs/` | themed Swagger UI dist (static) | always |
|
||||||
| `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) |
|
| `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) |
|
||||||
| `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` |
|
| `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` |
|
||||||
|
|
@ -130,7 +130,7 @@ services.hyperhive.gateway.auth = {
|
||||||
|
|
||||||
`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history.
|
`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history.
|
||||||
|
|
||||||
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
|
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
|
||||||
|
|
||||||
A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first.
|
A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first.
|
||||||
|
|
||||||
|
|
@ -261,10 +261,9 @@ Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }`
|
||||||
|
|
||||||
#### Dashboard: path-based routing (not Accept-header)
|
#### Dashboard: path-based routing (not Accept-header)
|
||||||
|
|
||||||
hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
|
hive-c0re serves exactly two prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
|
||||||
|
|
||||||
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live.
|
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live.
|
||||||
- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded.
|
|
||||||
- `location /health/` → hive-c0re: liveness and readiness.
|
- `location /health/` → hive-c0re: liveness and readiness.
|
||||||
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`.
|
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -43,22 +43,18 @@ because there is no malformed spec to reject.
|
||||||
Nix-heavy — hold one of the `buildSlots` permits for the node's duration:
|
Nix-heavy — hold one of the `buildSlots` permits for the node's duration:
|
||||||
|
|
||||||
| Node | Wraps |
|
| Node | Wraps |
|
||||||
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead |
|
| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead |
|
||||||
| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node |
|
| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node |
|
||||||
| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta |
|
| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta |
|
||||||
| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion |
|
| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion |
|
||||||
| `DeployWindow` | resource-holding root of the merge-config-PR deploy subtree — declares the build slot, the lease and the meta window, then completes immediately so its children run under them (see _Approvals_ below) |
|
|
||||||
| `DeployApply` | the deploy's irreversible half: ff-merge the reviewed PR head, two-phase meta deploy, container rebuild |
|
|
||||||
|
|
||||||
Cheap — no build slot:
|
Cheap — no build slot:
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
|
|
||||||
| Node | Behavior |
|
| Node | Behavior |
|
||||||
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
| `MergeVerify` | the deploy's pre-merge gate — PR-head drift check, fetch, `verify_commit` eval. Mutates nothing, so a rejection here needs no compensation |
|
|
||||||
| `DeployTail` | the deploy's `AfterAny` compensation + bookkeeping tail: (1) rolls `applied/main` back from the parked `refs/hyperhive/rollback/<id>` and aborts the staged meta lock when the deploy never confirmed good; (2) mirrors whichever deploy tag landed to the forge config repo, always, best-effort; (3) posts the failing build log back onto the config PR when the deploy failed. Named for (2)/(3), which run on the success path too — not `AbortDeploy`. Infallible by construction |
|
|
||||||
| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build |
|
| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build |
|
||||||
| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#<name>` ref resolves. Store/meta-only, no container yet |
|
| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#<name>` ref resolves. Store/meta-only, no container yet |
|
||||||
| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop |
|
| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop |
|
||||||
|
|
@ -75,14 +71,12 @@ Cheap — no build slot:
|
||||||
| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step |
|
| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step |
|
||||||
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
|
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
|
||||||
| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles |
|
| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles |
|
||||||
| `FinalizeDeploy` | deploy phase 3 — drop the rollback ref, plant `deployed/<id>`, commit the staged `flake.lock`. The first two git steps are fatal on purpose, so a confirmed-good deploy's outcome and the repo's state can't disagree |
|
|
||||||
| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent |
|
| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent |
|
||||||
| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome |
|
| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome |
|
||||||
| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload |
|
| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload |
|
||||||
| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots |
|
| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots |
|
||||||
| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless |
|
| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless |
|
||||||
| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless |
|
| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless |
|
||||||
| `WebhookRegister` | one-shot boot-time Forgejo webhook registration (`internal/knowledge` push→pull, `agent-configs` PR→approval). No-op until the core token, hive domain, and HMAC secret are all available. Agentless |
|
|
||||||
| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless |
|
| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless |
|
||||||
| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless |
|
| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless |
|
||||||
|
|
||||||
|
|
@ -93,26 +87,21 @@ with its commit under its internal `META_LOCK` mutex, so a standalone commit
|
||||||
node would open a dirty-working-tree window between nodes.
|
node would open a dirty-working-tree window between nodes.
|
||||||
|
|
||||||
Two further layers protect the meta repo across _windows_ that span multiple
|
Two further layers protect the meta repo across _windows_ that span multiple
|
||||||
`META_LOCK` acquisitions — above all the approval deploy's prepare→finalize
|
`META_LOCK` acquisitions:
|
||||||
span, which keeps a bumped `flake.lock` **staged uncommitted** for the whole
|
|
||||||
container build:
|
|
||||||
|
|
||||||
- **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue
|
- **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue
|
||||||
resource declared by every node kind that mutates the meta repo — `MetaSync`,
|
resource declared by every node kind that mutates the meta repo — `MetaSync`,
|
||||||
`MetaLock`, `WritePermFile`, `Provision`'s agent registration, and
|
`MetaLock`, `WritePermFile` and `Provision`'s agent registration. Two meta
|
||||||
`DeployWindow` — the deploy subtree's root, which holds it across every
|
|
||||||
phase below it (it declares `Resource::MetaWindow`). Two meta
|
|
||||||
mutations can therefore never interleave, so no commit lands inside another
|
mutations can therefore never interleave, so no commit lands inside another
|
||||||
node's staged window. It's a queue resource rather than a runtime mutex
|
node's window. It's a queue resource rather than a runtime mutex
|
||||||
because a subtree root holds a resource across its whole subtree, which
|
because a subtree root holds a resource across its whole subtree, which
|
||||||
a `MutexGuard` (bounded by one executor fn) can't — that's what lets a
|
a `MutexGuard` (bounded by one executor fn) can't. For the same reason the window must stay
|
||||||
multi-node deploy own one window. For the same reason the window must stay
|
|
||||||
_off_ long store-only work: the rebuild's meta preamble is its own
|
_off_ long store-only work: the rebuild's meta preamble is its own
|
||||||
`MetaSync` node, a sibling of (never a parent of) `Prebuild`, so the
|
`MetaSync` node, a sibling of (never a parent of) `Prebuild`, so the
|
||||||
toplevel build runs outside the window and `buildSlots > 1` still gives
|
toplevel build runs outside the window and `buildSlots > 1` still gives
|
||||||
concurrent rebuilds across agents.
|
concurrent rebuilds across agents.
|
||||||
- **Path-limited commits**: the targeted meta committers (perm files,
|
- **Path-limited commits**: the targeted meta committers (perm files,
|
||||||
topology, lock bumps, finalize) commit `-- <their paths>` with path-scoped
|
topology, lock bumps) commit `-- <their paths>` with path-scoped
|
||||||
dirty checks, so even a non-queue caller (boot migration, destroy's
|
dirty checks, so even a non-queue caller (boot migration, destroy's
|
||||||
`sync_agents`) can never sweep someone else's staged content into its
|
`sync_agents`) can never sweep someone else's staged content into its
|
||||||
commit.
|
commit.
|
||||||
|
|
@ -220,8 +209,8 @@ resources are free. Resources:
|
||||||
2. **Per-agent lifecycle lease** — keyed on the **node's** agent (agent is
|
2. **Per-agent lifecycle lease** — keyed on the **node's** agent (agent is
|
||||||
per-node; a DAG can span agents) and globally exclusive per agent across
|
per-node; a DAG can span agents) and globally exclusive per agent across
|
||||||
all DAGs: acquired either at a container-affecting node (`SetWanted`,
|
all DAGs: acquired either at a container-affecting node (`SetWanted`,
|
||||||
`Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`,
|
`Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`) on
|
||||||
`DeployWindow`) on behalf of a whole coordinated subtree; held by the owning
|
behalf of a whole coordinated subtree; held by the owning
|
||||||
DAG until it's terminal, so two DAGs never interleave container ops on the
|
DAG until it's terminal, so two DAGs never interleave container ops on the
|
||||||
same agent. A DAG touching multiple agents holds one lease per agent.
|
same agent. A DAG touching multiple agents holds one lease per agent.
|
||||||
(`SetWanted` is a store write, not a container op, but takes the lease anyway
|
(`SetWanted` is a store write, not a container op, but takes the lease anyway
|
||||||
|
|
@ -293,36 +282,9 @@ the dashboard renders one recent-builds list and one number bounds it.
|
||||||
|
|
||||||
### Approvals
|
### Approvals
|
||||||
|
|
||||||
`MergeConfigPr` approvals ride as a four-node deploy subtree:
|
`UpdateMetaInputs` approvals map onto the ordinary `meta-update` shapes.
|
||||||
|
The scheduler fires `actions::resolve_approval_dag` exactly once when
|
||||||
```
|
**any** approval-carrying DAG settles terminal (including
|
||||||
DeployWindow (root — build slot + lease + meta window, no work of its own)
|
|
||||||
├── MergeVerify drift gate, fetch, verify_commit
|
|
||||||
├── DeployApply AfterOk(verify) park rollback ref, ff-merge, deploy
|
|
||||||
└── DeployTail AfterAny(apply) compensate, mirror to forge
|
|
||||||
```
|
|
||||||
|
|
||||||
The root holds its resources across the whole subtree, so the two-phase
|
|
||||||
`prepare_deploy` / `finalize_deploy` span keeps its staged `flake.lock`
|
|
||||||
protected even though the phases are separate nodes. Splitting them buys
|
|
||||||
three things a single opaque node couldn't have: per-phase visibility on the
|
|
||||||
dashboard, a `MergeVerify` failure that provably mutated nothing, and a
|
|
||||||
compensation step that survives a hive-c0re restart — `DeployApply` parks the pre-merge
|
|
||||||
`applied/main` in `refs/hyperhive/rollback/<approval-id>`, not in a
|
|
||||||
local variable, so `DeployTail` can still undo a half-finished deploy after a
|
|
||||||
crash.
|
|
||||||
|
|
||||||
`DeployWindow` declares all three resources (build slot, lease, meta window)
|
|
||||||
on itself rather than letting each phase declare its own, because the queue
|
|
||||||
acquires a node's resources atomically (all-or-nothing): a child that took
|
|
||||||
the build slot while its parent held the meta window could block waiting for
|
|
||||||
a resource its own parent already committed to, a lock-ordering hazard that
|
|
||||||
one multi-resource root avoids by construction.
|
|
||||||
|
|
||||||
`UpdateMetaInputs` approvals map onto the ordinary
|
|
||||||
`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag`
|
|
||||||
exactly once when **any** approval-carrying DAG settles terminal — deploys
|
|
||||||
included, since their outcome is the DAG's own state (including
|
|
||||||
cancelled-while-queued, which fails the approval instead of dangling it).
|
cancelled-while-queued, which fails the approval instead of dangling it).
|
||||||
|
|
||||||
### Wire shape
|
### Wire shape
|
||||||
|
|
@ -398,16 +360,12 @@ Key operations:
|
||||||
- **`sync_agents`** (idempotent) — render `flake.nix` for the current agent set,
|
- **`sync_agents`** (idempotent) — render `flake.nix` for the current agent set,
|
||||||
init the repo on first call, relock if the rendered contents changed, commit.
|
init the repo on first call, relock if the rendered contents changed, commit.
|
||||||
Called by spawn / destroy / startup migration.
|
Called by spawn / destroy / startup migration.
|
||||||
- **`prepare_deploy` + `finalize_deploy` / `abort_deploy`** — two-phase for the
|
|
||||||
`MergeConfigPr` deploy path so a failed `nixos-container update` leaves no orphan
|
|
||||||
commit in meta. Prepare writes the new lock without committing; finalize commits
|
|
||||||
with the deploy message; abort restores the lock.
|
|
||||||
- **`lock_update_hyperhive`** — one-shot for the boot-reconcile path (the
|
- **`lock_update_hyperhive`** — one-shot for the boot-reconcile path (the
|
||||||
sweep DAG's `MetaLock` node): bumps the `hyperhive` input lock and commits;
|
sweep DAG's `MetaLock` node): bumps the `hyperhive` input lock and commits;
|
||||||
the scheduler fans out the agent rebuilds on completion.
|
the scheduler fans out the agent rebuilds on completion.
|
||||||
|
|
||||||
Every public `meta.rs` operation takes the module's internal `META_LOCK`
|
Every public `meta.rs` operation takes the module's internal `META_LOCK`
|
||||||
mutex, so concurrent job-queue nodes (and the approval deploy pipeline) never
|
mutex, so concurrent job-queue nodes never
|
||||||
race on the repo's `.git/index.lock`.
|
race on the repo's `.git/index.lock`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
@ -449,18 +407,6 @@ Sequence for a rebuild DAG (each step is its own queue node):
|
||||||
in-container activation script transitions old → new. Holds no build
|
in-container activation script transitions old → new. Holds no build
|
||||||
slot, so the next DAG's `Prebuild` overlaps the container boot.
|
slot, so the next DAG's `Prebuild` overlaps the container boot.
|
||||||
|
|
||||||
The approval deploy uses this same chain rather than a rebuild path of its
|
|
||||||
own. Its `DeployApply` node doesn't build: it merges, opens the two-phase
|
|
||||||
meta deploy, and returns the chain above as a subgraph the scheduler grafts
|
|
||||||
into the live DAG under that node. A `FinalizeDeploy` node gated on the
|
|
||||||
graft's completion then plants the deploy tag — so `Reconcile`'s success
|
|
||||||
answers "did the agent come back up?" the same way it does for every
|
|
||||||
other rebuild, instead of a fused inline start.
|
|
||||||
|
|
||||||
The grafted nodes land _inside_ `DeployWindow`'s subtree, so they re-enter
|
|
||||||
the meta window and build slot it already holds rather than deadlocking
|
|
||||||
against it.
|
|
||||||
|
|
||||||
### Cold-start fallback
|
### Cold-start fallback
|
||||||
|
|
||||||
`start` after `update` can exit non-zero when packages are **removed** between
|
`start` after `update` can exit non-zero when packages are **removed** between
|
||||||
|
|
|
||||||
|
|
@ -3,7 +3,7 @@
|
||||||
Long-running work runs through a job graph. The swarm controller keeps one
|
Long-running work runs through a job graph. The swarm controller keeps one
|
||||||
for swarm-level work — creating an agent's identity, forge user and config
|
for swarm-level work — creating an agent's identity, forge user and config
|
||||||
repo. Each hive's hive-c0re keeps its own for container operations —
|
repo. Each hive's hive-c0re keeps its own for container operations —
|
||||||
rebuild, first-spawn, a config-PR deploy, power changes. This page explains
|
rebuild, first-spawn, power changes. This page explains
|
||||||
what the job queue _is_, as a general idea, independent of what either uses
|
what the job queue _is_, as a general idea, independent of what either uses
|
||||||
it for. For the hive-c0re step catalogue and the engineering internals
|
it for. For the hive-c0re step catalogue and the engineering internals
|
||||||
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)
|
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)
|
||||||
|
|
|
||||||
|
|
@ -227,9 +227,9 @@ one per hive. It ensures them at start and every five minutes after
|
||||||
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
|
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
|
||||||
owner org;
|
owner org;
|
||||||
- the empty `operators` merge-gate team in `agents` and `agent-configs`;
|
- the empty `operators` merge-gate team in `agents` and `agent-configs`;
|
||||||
- the `main` merge gate on every `agent-configs` repo: merge whitelist =
|
- the `main` merge gate on every `agent-configs` repo: merge and approval
|
||||||
the `operators` team and the `core` user, approval whitelist = the
|
whitelists = the `operators` team, and no user. The controller leaves a
|
||||||
`operators` team. The controller leaves a repo with no `main` rule alone;
|
repo with no `main` rule alone;
|
||||||
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
|
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
|
||||||
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
|
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
|
||||||
- `internal/docs` (private) and `internal/knowledge` (public, with a
|
- `internal/docs` (private) and `internal/knowledge` (public, with a
|
||||||
|
|
@ -258,7 +258,7 @@ decision, not an event to adjudicate.
|
||||||
A `config-pr` delivery reporting a PR merged into `main` queues a deploy
|
A `config-pr` delivery reporting a PR merged into `main` queues a deploy
|
||||||
of its `merge_commit_sha` on the one hive whose wanted state places the
|
of its `merge_commit_sha` on the one hive whose wanted state places the
|
||||||
agent; with no such hive, or several, the controller deploys nothing. See
|
agent; with no such hive, or several, the controller deploys nothing. See
|
||||||
[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui).
|
[approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes).
|
||||||
|
|
||||||
**`internal/knowledge` is on that path.** The controller's is the only
|
**`internal/knowledge` is on that path.** The controller's is the only
|
||||||
hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a
|
hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a
|
||||||
|
|
@ -267,10 +267,10 @@ would take delivery away from the first rather than add a recipient.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
|
|
||||||
**The `agent-configs` org isn't.** Each hive registers its own
|
**The `agent-configs` org is on it too.** The controller's hook is the
|
||||||
`pull_request` hook there, so that repo has two — the hive's and the
|
only one that acts on config PRs. A hive's `/webhook/config-pr` hook left
|
||||||
controller's — and **both are expected; don't delete either.** Removing
|
on the org by an older release delivers to a route no hive serves; delete
|
||||||
a hive's stops it acting on config PRs; removing the controller's stops
|
it in the org's webhook settings. Removing the controller's hook stops
|
||||||
forge-UI merges from deploying until its next start recreates it.
|
forge-UI merges from deploying until its next start recreates it.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
<!-- vale write-good.Passive = YES -->
|
||||||
|
|
|
||||||
|
|
@ -40,8 +40,8 @@ hivectl forge reconcile-config iris --verbose # include the full diff, not
|
||||||
applied config checkout and its forge `agent-configs/<agent>` `main`, then
|
applied config checkout and its forge `agent-configs/<agent>` `main`, then
|
||||||
reconciles. `--from forge` resets the local checkout to forge `main` (takes
|
reconciles. `--from forge` resets the local checkout to forge `main` (takes
|
||||||
effect on the next deploy — it doesn't autorebuild). `--from local` isn't
|
effect on the next deploy — it doesn't autorebuild). `--from local` isn't
|
||||||
supported yet (forge `main` is core-only branch-protected; resolve via a
|
supported yet (forge `main` is branch-protected; resolve via a config
|
||||||
config PR). With no `--from` it prompts for the direction after the diff.
|
PR). With no `--from` it prompts for the direction after the diff.
|
||||||
|
|
||||||
## Matrix
|
## Matrix
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -123,14 +123,17 @@ checkpoints**, not about sandboxing the agent from its own tools:
|
||||||
highest-value action. On the **internal forge this is technically enforced,
|
highest-value action. On the **internal forge this is technically enforced,
|
||||||
not just convention**: agents can't create repos (`max_repo_creation = 0`),
|
not just convention**: agents can't create repos (`max_repo_creation = 0`),
|
||||||
and `main` on an `agent-configs/<name>` repo carries swarm-controller's
|
and `main` on an `agent-configs/<name>` repo carries swarm-controller's
|
||||||
branch protection: merge allowlisted to the `operators` team, with one
|
branch protection: merge and approval allowlisted to the `operators` team,
|
||||||
approval from it. Existing `agents/<repo>` repos carry the same merge gate.
|
which swarm-controller converges on every config repo. An operator's merge
|
||||||
|
there is also what deploys the config. Existing `agents/<repo>` repos carry
|
||||||
|
the same merge gate.
|
||||||
An agent (a write collaborator, not a repo admin) can neither change those
|
An agent (a write collaborator, not a repo admin) can neither change those
|
||||||
settings nor merge its own PR. It's **not** set up for external VCS (GitHub
|
settings nor merge its own PR. It's **not** set up for external VCS (GitHub
|
||||||
etc.), though — there, operator-merge is process + accepted risk, not a
|
etc.), though — there, operator-merge is process + accepted risk, not a
|
||||||
technical control.
|
technical control.
|
||||||
- **Approvals** — config changes, schedule additions, and other
|
- **Approvals** — schedule additions and other blast-radius-y operations
|
||||||
blast-radius-y operations route through the operator approval queue
|
route through the operator approval queue; config changes are config PRs
|
||||||
|
an operator merges on the forge
|
||||||
(see [`approvals.md`](../agent-lifecycle/approvals.md)).
|
(see [`approvals.md`](../agent-lifecycle/approvals.md)).
|
||||||
|
|
||||||
### Capability = accepted risk
|
### Capability = accepted risk
|
||||||
|
|
|
||||||
|
|
@ -70,12 +70,7 @@ approval (`Coordinator::notify_submitter`, looked up from the
|
||||||
authenticated socket caller at submit time; a row with no
|
authenticated socket caller at submit time; a row with no
|
||||||
recorded submitter falls back to the manager, `ruth`).
|
recorded submitter falls back to the manager, `ruth`).
|
||||||
`ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`,
|
`ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`,
|
||||||
hardcoded — `hive-c0re/src/workers/crash_watch.rs`). A `MergeConfigPr`
|
hardcoded — `hive-c0re/src/workers/crash_watch.rs`). Lifecycle
|
||||||
approval's rebuild additionally pushes a `rebuilt:<agent>` todo to that
|
|
||||||
same submitter (`Coordinator::push_todo_submitter`, `subsystem =
|
|
||||||
"core"`), which wakes a turn (the todo-wake path — see [Turn
|
|
||||||
outcomes](README.md#turn-outcomes)) via a generic "call
|
|
||||||
`get_loose_ends`" prompt rather than the event body itself. Lifecycle
|
|
||||||
transitions the job-queue scheduler or crash watcher drive directly —
|
transitions the job-queue scheduler or crash watcher drive directly —
|
||||||
stop/kill, destroy, a flake-rev login or logout state change — reach
|
stop/kill, destroy, a flake-rev login or logout state change — reach
|
||||||
no individual agent: they publish onto a swarm-wide NATS
|
no individual agent: they publish onto a swarm-wide NATS
|
||||||
|
|
|
||||||
|
|
@ -48,24 +48,24 @@ and quick links (stats, screen, forge profile). Select the name to open
|
||||||
its terminal and watch it work in real time.
|
its terminal and watch it work in real time.
|
||||||
|
|
||||||
**Approve something an agent is waiting on.** Y3R C4LL is the one tab
|
**Approve something an agent is waiting on.** Y3R C4LL is the one tab
|
||||||
worth checking regularly — it's everything that needs _you_: approvals
|
worth checking regularly — it's everything that needs _you_: an agent's
|
||||||
for config changes. The tab's count pill tells you at a glance if
|
request to schedule a prompt. The tab's count pill tells you at a glance if
|
||||||
anything's pending.
|
anything's pending.
|
||||||
|
|
||||||
**Approve or reject a config change.** Agent config changes (new
|
**Review a config change.** Agent config changes (new packages, env
|
||||||
packages, env vars, MCP servers) go through an approval queue rather
|
vars, MCP servers) are pull requests on the agent's `agent-configs`
|
||||||
than landing automatically — you'll see them on Y3R C4LL, with a diff
|
repo. Review and merge them on the forge; the merge deploys the change.
|
||||||
of what's changing.
|
See [`approvals.md`](../agent-lifecycle/approvals.md#config-changes).
|
||||||
|
|
||||||
**Start, stop, restart, or rebuild an agent.** Select one or more
|
**Start, stop, restart, or rebuild an agent.** Select one or more
|
||||||
agents on SW4RM (select the icon) and use the selection bar, or use the
|
agents on SW4RM (select the icon) and use the selection bar, or use the
|
||||||
per-agent `⋮` menu on a single row. Rebuilding re-applies that agent's
|
per-agent `⋮` menu on a single row. Rebuilding re-applies that agent's
|
||||||
current config; use it after approving a change, or whenever an agent
|
current config; use it whenever an agent
|
||||||
shows as "needs update."
|
shows as "needs update."
|
||||||
|
|
||||||
**Watch a build.** BU1LDS shows the rebuild queue live, plus a
|
**Watch a build.** BU1LDS shows the rebuild queue live, plus a
|
||||||
streaming log of whatever's currently building. Useful right after
|
streaming log of whatever's currently building. Useful right after
|
||||||
approving a change or bumping a flake input.
|
merging a config change or bumping a flake input.
|
||||||
|
|
||||||
**Grant or revoke a tool/capability.** P3RM1SS10NS is a checkbox matrix
|
**Grant or revoke a tool/capability.** P3RM1SS10NS is a checkbox matrix
|
||||||
— rows are agents, columns are tool groups or capabilities. Nothing
|
— rows are agents, columns are tool groups or capabilities. Nothing
|
||||||
|
|
|
||||||
|
|
@ -877,11 +877,10 @@ renderApprovals`) with three stacked sections:
|
||||||
right-aligned `requested <N> ago` relative time from
|
right-aligned `requested <N> ago` relative time from
|
||||||
`ApprovalView.requested_at`. Glyph and chip vary by kind:
|
`ApprovalView.requested_at`. Glyph and chip vary by kind:
|
||||||
|
|
||||||
| kind | glyph | chip | sha shown |
|
| kind | glyph | chip |
|
||||||
|---|---|---|---|
|
|---|---|---|
|
||||||
| `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) |
|
| `update_meta_inputs` | `↻` | `meta-update` |
|
||||||
| `update_meta_inputs` | `↻` | `meta-update` | — |
|
| `schedule_prompt` | `⏱` | `schedule` |
|
||||||
| `schedule_prompt` | `⏱` | `schedule` | — |
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
The chip ticks live every second via a `data-requested-at`
|
The chip ticks live every second via a `data-requested-at`
|
||||||
|
|
@ -889,12 +888,9 @@ renderApprovals`) with three stacked sections:
|
||||||
the request has been pending ≥ 1h so a stale approval stands out;
|
the request has been pending ≥ 1h so a stale approval stands out;
|
||||||
the `.stale` class flips precisely at the 3600s boundary rather
|
the `.stale` class flips precisely at the 3600s boundary rather
|
||||||
than at the next `renderApprovals` call.
|
than at the next `renderApprovals` call.
|
||||||
- **what-changed body** — the submitting agent's description, then
|
- **what-changed body** — the submitting agent's description, then the
|
||||||
kind-specific drill-in triggers:
|
kind's payload: the inputs to bump (`update_meta_inputs`) or the prompt to
|
||||||
- `merge_config_pr`: `↳ review PR on forge ↗` deep-links the
|
schedule (`schedule_prompt`).
|
||||||
config PR into `agent-configs/<agent>/pulls/<pr_number>` (shown
|
|
||||||
only when `forge_present` is true and `pr_number` has a value). The config diff
|
|
||||||
lives on the forge PR itself — no inline diff side-panel.
|
|
||||||
- **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a
|
- **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a
|
||||||
`prompt()` for an optional reason carried to the submitting agent as
|
`prompt()` for an optional reason carried to the submitting agent as
|
||||||
`HelperEvent::ApprovalResolved.note`.
|
`HelperEvent::ApprovalResolved.note`.
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue