From a88ed9f24e71d792db446505da9f1e4e0ba81c97 Mon Sep 17 00:00:00 2001 From: atlas Date: Fri, 2 Oct 2026 22:32:17 +0200 Subject: [PATCH] docs: config changes are operator merges on the forge Rewrites the config-change flow around the forge merge and the DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy DAG, the hive's `/webhook/` route and the `core` merge allowlist from the docs, and states that operators join the `operators` team by hand. Refs #4850 --- README.md | 2 +- docs/agent-lifecycle/approvals.md | 332 +++++++++------------------- docs/agent-lifecycle/persistence.md | 4 +- docs/getting-started/setup.md | 5 +- docs/integrations/forge.md | 25 +-- docs/networking/gateway.md | 7 +- docs/scheduler/coordinator.md | 140 ++++-------- docs/scheduler/jobq.md | 2 +- docs/swarm/README.md | 16 +- docs/tools/hivectl.md | 4 +- docs/trust-boundary/security.md | 11 +- docs/turn-loop/mcp.md | 7 +- docs/web-ui/README.md | 16 +- docs/web-ui/dashboard.md | 18 +- 14 files changed, 194 insertions(+), 395 deletions(-) diff --git a/README.md b/README.md index 1d25a23b..35aea61c 100644 --- a/README.md +++ b/README.md @@ -31,7 +31,7 @@ architecture change. | **identity** | every agent is a swarm-wide principal — SSO subject, forge user, matrix account, secret-store cert identity — addressable as `name@hive.domain` | | **secrets** | one OpenBao store; the operator places one mTLS identity per host, and everything else — every agent's credentials included — is fetched from the store under an identity rather than copied by hand | | **shared services** | one forge, homeserver, SSO, message queue and metrics/logs stack per swarm, each on whichever host you put it | -| **config** | git: an agent proposes, the operator approves, the deploy lands as a `deployed/` tag | +| **config** | git: an agent opens a PR on its config repo, an operator merges it on the forge, and the merge deploys | | **runtime** | `claude --print` by default; any [ACP](https://agentclientprotocol.com) agent (e.g. opencode) per agent with `services.hyperhive.agent.runtime = "acp"` | | **substrate** | each hive runs one `nixos-container` per agent, under an unprivileged daemon with a tiny socket-activated helper for the few root ops | | **watching** | the swarm UI (hives, agents with live terminals, jobs, a cross-repo issue report), Grafana over OTEL, and a per-hive dashboard for host-level detail | diff --git a/docs/agent-lifecycle/approvals.md b/docs/agent-lifecycle/approvals.md index bec3e346..25ddea72 100644 --- a/docs/agent-lifecycle/approvals.md +++ b/docs/agent-lifecycle/approvals.md @@ -1,38 +1,32 @@ # Approvals + helper events -The approval queue is hyperhive's pivot: nothing that changes the -shape of an agent (its config, whether it exists) happens without an -operator selection. The submitting agent — any agent with the `approvals` -tool group, which manages the config of its **direct children** (the -root agent for top-level agents; a sub-manager for its own subtree) — is -the policy gate in front of that queue; helper events are how it stays -informed about what happens after a decision lands. +The approval queue is where an agent asks the operator for something it +can't do itself: a scheduled prompt, or (legacy rows only) a meta-input +bump. Config changes don't go through this queue: they're pull requests +on the agent's config repo, which an operator merges on the forge, and +that merge deploys them. The submitting agent — any agent with the +`approvals` tool group, which manages the config of its **direct +children** (the root agent for top-level agents; a sub-manager for its +own subtree) — is the policy gate in front of both; helper events are how +it stays informed about what happens after a decision lands. ## For operators -Every add/remove/change to an agent lands on your dashboard's Y3R -C4LL tab (or `hivectl approvals pending` / `approve ` from the -CLI) before it takes effect. What you'll see, and what to do with it: - -- **Config change** (`MergeConfigPr`) — an agent proposed a change to - another agent's config (or its own, via a sub-manager) as a forge - pull request. Review the diff on the forge — the dashboard card - links straight to it, same as reviewing any other PR. Approving - triggers the deploy automatically: hive-c0re re-verifies the PR - hasn't moved since you looked at it, evaluates it (a dry run, - nothing applied yet), merges it, and rebuilds the container. If - anything in that chain fails, the change rolls back automatically — - the agent stays on its last-good config, no recovery action needed - from you. -- **Meta/flake update** (`UpdateMetaInputs`) — an agent asked to bump - one or more Nix flake inputs (or all of them). Approving runs the - update and commits the lock change; it doesn't rebuild anything by - itself. +- **Config change** — an agent proposed a change to another agent's + config (or its own, via a sub-manager) as a pull request on + `agent-configs/`. Review it on the forge like any other PR, and + merge it there to deploy it. See [Config changes](#config-changes). - **Scheduled prompt** (`SchedulePrompt`) — an agent asked to schedule - a message to one or more inboxes at a future time. You can also add - schedules yourself directly from the SCH3DUL3S tab, which skips this - approval step entirely — the gate here is specifically for an - *agent* asking to schedule something, not for you doing it. + a message to one or more inboxes at a future time. It lands on your + dashboard's Y3R C4LL tab (or `hivectl approvals pending` / + `approve ` from the CLI). You can also add schedules yourself + directly from the SCH3DUL3S tab, which skips this approval step + entirely — the gate here is specifically for an *agent* asking to + schedule something, not for you doing it. +- **Meta/flake update** (`UpdateMetaInputs`) — legacy: nothing queues + this kind, but an existing row still reads back and you can approve it. + Approving runs the update and commits the lock change; it doesn't + rebuild anything by itself. Don't want to approve something? **Deny it** (`DENY` on the dashboard card, or `hivectl approvals deny `) — nothing runs. Either way the @@ -40,18 +34,16 @@ submitting agent is always notified that the operator denied their request; what optional is only the reason text, which you can add on the dashboard's prompt (cancelling that prompt aborts the whole deny, not just the reason) but not from the CLI. Denying is final: a denied approval -can't be re-approved later, the agent has to submit a fresh one (a new -PR, a new request). +can't be re-approved later, the agent has to submit a fresh request. Everything below this point is the implementation detail behind that flow. -## End-to-end approval flow +## Config changes Config changes flow through a **forge pull request** on the agent's `agent-configs/` repo — the same surface agents use for code PRs. -No bespoke MCP tool exists for config changes: opening the PR IS the -request. +No MCP tool exists for config changes: opening the PR IS the request. 1. The submitting agent (the child's parent, holding the `approvals` tool group) **clones** `agent-configs/`, edits it there (any @@ -59,89 +51,40 @@ request. with its own git identity, and pushes a branch + opens a PR with `hive-forge` — the same way it would change any other repo. The bind-mounted `/agents//config/` is a **copy for reading** a - config, not the tree to edit: authoring in place there produces no PR - and no approval. (it's currently mounted read-write, which is a - defect tracked separately, not an authoring path.) + config, not the tree to edit: authoring in place there produces no PR. + (it's currently mounted read-write, which is a defect tracked + separately, not an authoring path.) Branch protection (the agent isn't on the `main` push allowlist; - merge allowlist = `core` user + `operators` team; approvals allowlist - = `operators` team; see "Forge mirror" below) makes the agent a write - collaborator that **can't merge its own config PR**. -2. hive-c0re's `/webhook/config-pr` endpoint receives the Forgejo - `pull_request` event (opened / synchronized / reopened) and queues a - `MergeConfigPr` approval; a poll fallback catches any missed webhook. - The approval row stores the PR **number** (`commit_ref`) and the PR - **head sha at queue time** (`fetched_sha` — the "reviewed" sha). If - the PR head later moves, a fresh approval pinned to the new head supersedes - the stale one, so the operator always reviews what will - actually deploy. -3. The operator reviews the PR **on the forge** (native diff, threaded - comments, CI status) and sees a matching card on the dashboard with a - "review PR on forge" deep link. They select ◆ APPR0VE (or - `hivectl approvals approve ` on the CLI) once satisfied. -4. On approve, a deploy DAG runs three phases under a resource-holding - `DeployWindow` root (see *Queue templates* below): - - `MergeVerify` re-reads the live PR head and **aborts if it drifted** - from the reviewed `fetched_sha` (the submitter must push again, - which queues a fresh approval); then fetches that head into the - applied repo and **eval-verifies** it — a flake eval on a throwaway - checkout. This is the trust gate: it relies on c0re's own eval, not - on any in-repo (agent-forgeable) signal like a CI status. Nothing is - mutated in this phase, so a rejection here leaves the forge and the - applied repo exactly as they were. - - `DeployApply` parks the pre-merge `applied/main` in - `refs/hyperhive/rollback/`, then fast-forward-merges - the reviewed head to the forge config repo's `main` (this IS the - merge — a `core`-authenticated ff-merge pinned to the reviewed sha, - so a moved PR head can't substitute bytes), and runs the deploy - proper (`deploy_applied_target`): ff `applied/main`, two-phase meta - deploy, container rebuild. On success it drops the rollback ref and - plants `deployed/`. - - `DeployTail` runs on **every** outcome, including a cancel-cascade. - If the rollback ref survived, the deploy never confirmed good: it - rolls `applied/main` back, resyncs the working tree, and aborts the - staged meta lock, so the agent stays on its last-good tree. Then it - mirrors the config repo (and its new deploy tag) to the forge. - - The rollback state lives in a **git ref, not a local variable**, on - purpose: hive-c0re can restart between the apply and the tail, and the - tail still has to know what to undo when it does. -5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the - **submitting agent's** inbox via `notify_submitter`, carrying both the - canonical sha and the terminal tag (the approval row carries a - `submitter` column recording the agent the change is for). - -### Operator merge in the forge UI - -An operator can also merge a config PR straight in the Forgejo UI. That -deploys the merged commit on the hive that runs the agent: - -1. swarm-controller gets the `agent-configs` org's `pull_request` + merge and approvals allowlist = `operators` team; see "Forge mirror" + below) makes the agent a write collaborator that **can't merge its own + config PR**. +2. An operator reviews the PR **on the forge** (native diff, threaded + comments, CI status) and merges it there. Merging needs membership of + the `operators` team in `agent-configs`. The team starts empty: add + each operator by hand in the forge UI. +3. swarm-controller gets the `agent-configs` org's `pull_request` delivery. A `closed` event with `merged: true` and base branch `main` names the commit in `merge_commit_sha`. -2. The controller looks up which hive's wanted state places the agent +4. The controller looks up which hive's wanted state places the agent and publishes a deploy request carrying that commit on that hive's deploy subject. With no such hive, or more than one, it deploys nothing and logs a warning naming them. -3. If the hive's `applied/main` already is that commit — a dashboard - approval merges and deploys its own PR — it does nothing. Otherwise - it fetches the config repo's `main`, requires the commit to descend - from `applied/main`, fast-forwards `applied/main` to it, and queues a - rebuild. A failure before the rebuild (the fetch, or a commit that - doesn't descend) deploys nothing and posts the error as a comment on - the merged PR. An agent with no container on that hive yet ignores +5. If the hive's `applied/main` already is that commit, it does nothing. + Otherwise it fetches the config repo's `main`, requires the commit to + descend from `applied/main`, fast-forwards `applied/main` to it, and + queues a rebuild. A failure before the rebuild (the fetch, or a commit + that doesn't descend) deploys nothing and posts the error as a comment + on the merged PR. An agent with no container on that hive yet ignores the commit; its first deploy builds what it seeds. This path runs **no eval-verify**. A merged config that doesn't evaluate or build fails the rebuild, and `applied/main` stays at that commit, so the agent's rebuilds keep failing until a fix merges. A failed rebuild shows on the hive like any other and isn't commented on -the PR. +the PR. A merge whose webhook delivery never reaches the controller +deploys nothing. -Merging needs membership of the `operators` team in `agent-configs`. -The team starts empty: add each operator by hand in the forge UI. A -merge whose webhook delivery never reaches the controller deploys -nothing. The hive's own poll cancels the dashboard card for that PR -with the note `PR merged/closed outside the approval`. +## Approval queue ### Withdrawing a pending approval @@ -166,31 +109,16 @@ and no hive can originate an agent. The swarm controller's `agent.nix` template, then asks the target hive to deploy it. Changing what the template seeded isn't a special case: like every -later change, it's a PR on that config repo (`MergeConfigPr`), made -from a clone, reviewed and approved by the operator. The PR flow is -the one path — an operator can equally drive both steps herself -through the web UI or the forge. +later change, it's a PR on that config repo, made from a clone and +merged by an operator on the forge (see [Config changes](#config-changes)). ### Approval kinds (wire shapes) -`ApprovalKind` carries three variants; each maps to a different +`ApprovalKind` carries two variants; each maps to a different `commit_ref` encoding because `ApprovalKind` overloads that field as the kind-specific payload carrier. -- `MergeConfigPr` — the config-change flow. Triggered automatically: - when an agent opens (or force-pushes) a PR on its - `agent-configs/` forge repo, hive-c0re's `/webhook/config-pr` - endpoint receives the Forgejo pull_request event and queues this - approval row. No MCP tool call needed — the forge PR IS the request. - `commit_ref` stores the **PR number** (decimal), and `fetched_sha` is - the PR **head sha at queue time** (the "reviewed" sha). On approve, - the deploy DAG's `MergeVerify` phase re-reads the live PR head and - aborts if it drifted from `fetched_sha` (submitter must push again to - re-trigger), then fetches that head into the applied repo and - eval-verifies it; `DeployApply` fast-forward-merges the forge config - repo's `main` to it (the merge) and runs `deploy_applied_target`; - `DeployTail` compensates on failure. Never a first spawn. - `UpdateMetaInputs` — `commit_ref` stores the JSON-encoded inputs array (`"[]"` = all inputs, `"[\"nixpkgs\"]"` = just nixpkgs, etc.). hive-c0re sets the `agent` field to the requesting root agent. @@ -303,55 +231,26 @@ declares one flake input per agent (`agent-.url = Containers run against `--flake /var/lib/hyperhive/meta#`. The declared input url is the agent's **forge config repo** (the -same `agent-configs/` the config-PR flow lands approved changes -on), so the meta flake references a reviewable, reproducible source -rather than a local checkout. hive-c0re authenticates that -`git+http` fetch via a git credential helper that reads the live -forge-core token — no token in the url or the lock. The deploy and -manual-rebuild paths, however, do **not** re-lock from the forge: -they `--override-input agent- +same `agent-configs/` config PRs merge into), so the meta flake +references a reviewable, reproducible source rather than a local +checkout. hive-c0re authenticates that `git+http` fetch via a git +credential helper that reads the live forge-core token — no token in +the url or the lock. The rebuild paths, however, do **not** re-lock +from the forge: they `--override-input agent- git+file:///var/lib/hyperhive/applied/`, locking the exact config -that `verify_commit` gated and `applied//main` was -fast-forwarded to. That keeps a deploy/rebuild reproducible and -independent of forge reachability — rebuilds fire on crash-restart -and meta bumps, not just config PRs — while the declared url stays -the forge. `sync_agents` re-renders + re-locks the persistent input; -a plain `nix flake lock` leaves an existing applied override in -place (it only re-locks when the declared url itself changes), so -the forge-declared / applied-deployed split is stable. +`applied//main` was fast-forwarded to. That keeps a rebuild +reproducible and independent of forge reachability — rebuilds fire on +crash-restart and meta bumps, not just config merges — while the +declared url stays the forge. `sync_agents` re-renders + re-locks the +persistent input; a plain `nix flake lock` leaves an existing applied +override in place (it only re-locks when the declared url itself +changes), so the forge-declared / applied-deployed split is stable. -Per-deploy lock flow (two-phase), spread across the deploy subtree's -nodes — each phase is its own node, so the queue can show which one is -running and a restart resumes at node granularity: - -1. `DeployApply` → `meta::prepare_deploy(name)` runs - `nix flake lock --update-input agent-` without - committing. Working tree of meta now points the input at - `applied//main` (which the deploy already fast-forwarded to - the reviewed PR head). -2. The rebuild subgraph `DeployApply` grows into the DAG builds and - swaps the container (`AgentWindow` bracing `Prebuild → StopForUpdate - → Swap → RebuildBookkeeping`, plus `Reconcile`). Nix evaluates - against the staged lock. -3. On success — `FinalizeDeploy` drops the rollback ref, plants - `deployed/`, then `meta::finalize_deploy(name, sha, "deployed/ - ")` stages `flake.lock` and commits with - `deploy deployed/ `. Meta's git log gains - one entry per successful deploy. -4. On failure — the `DeployTail` node runs `meta::abort_deploy()` - (`git restore flake.lock`) so the meta history shows only - successes; the failure stays as an annotated `failed/` - tag in `applied/`. The tail runs on every outcome, so this - also covers a hive-c0re restart mid-build: the staged lock is - dropped and `applied/main` rolled back from the parked - `refs/hyperhive/rollback/`. - -Single-phase variants exist for paths without -rollback semantics: `meta::lock_update_for_rebuild(name)` for -the manual `↻ R3BU1LD` button (commits if the lock changed) -and `meta::lock_update_hyperhive()` for the -autoupdate flake-rev bump (one shot before per-agent -rebuilds, commits if the lock changed). +Lock updates are single-phase and commit when the lock changed: +`meta::lock_update_for_rebuild(name)` relocks one agent's input for a +relocking rebuild (the manual `↻ R3BU1LD` button, and the rebuild a +merged config commit queues), and `meta::lock_update_hyperhive()` is the +autoupdate flake-rev bump (one shot before per-agent rebuilds). `meta::sync_agents(hive: &HiveEnv, agents: &[AgentSpec])` — `hive` carries `hyperhive_flake`, `dashboard_port`, and the rest of the @@ -390,8 +289,8 @@ per container row. └── # also tracked /var/lib/hyperhive/meta/ swarm-wide flake — core -├── .git/ # one commit per successful -│ # deploy +├── .git/ # one commit per lock +│ # change ├── flake.nix # generated from agent set └── flake.lock # pins each agent's sha ``` @@ -411,43 +310,27 @@ wraps it with identity + `HIVE_PORT` / `HIVE_LABEL` / ### Tag state machine -Each deploy leaves a tag on the underlying commit inside the applied -repo: - -| Tag | When | Annotated? | -|---|---|---| -| `deployed/` | rebuild succeeded — `main` ff's here | no | -| `failed/` | rebuild failed | yes (body = error) | - -hive-c0re plants `deployed/0` at first spawn. `applied/main` is always the -latest `deployed/*`. A `failed/` tree stays browsable forever — `git log ---tags` in the applied repo is the audit trail. A denied or failed config -PR carries no extra state on the forge side: the PR stays open, and the -submitter pushes again (or closes it) to retry. +hive-c0re plants `deployed/0` on the seed commit at first spawn. A +merged config commit that deploys plants no tag: the merged PR on the +forge and meta's lock commits record it. A config PR nobody merges +carries no extra state on the forge side: the PR stays open, and +the submitter pushes again (or closes it) to retry. ### Dispatch via the job queue -Long-running approval work — `MergeConfigPr` and `UpdateMetaInputs` -— runs as a DAG on the global job queue -(`docs/scheduler/coordinator.md::Job queue`), submitted by the approval handler -rather than run inline: +Long-running approval work — `UpdateMetaInputs` — runs as a DAG on the +global job queue (`docs/scheduler/coordinator.md::Job queue`), submitted +by the approval handler rather than run inline: | `ApprovalKind` | DAG submitted | source | |---|---|---| -| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` | | `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` | | `SchedulePrompt` | — runs inline (single sqlite insert) | — | -The DAG carries the originating `approval_id`, surfaced on the node that -owns it — for a deploy that's the `DeployWindow` root, so the dashboard -renders one approval card, not four. **Every** queued kind resolves -through `actions::resolve_approval_dag` when its DAG settles terminal: -the deploy's phases are ordinary queue nodes, so the DAG's own terminal -state is the authoritative outcome. That hook fires the matching -`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's -terminal tag (verifying the tag actually resolves in the applied repo — -a pre-merge rejection plants none), posts the failing build log back to -the config PR. +The DAG carries the originating `approval_id`. **Every** queued kind +resolves through `actions::resolve_approval_dag` when its DAG settles +terminal: the DAG's own terminal state is the authoritative outcome. +That hook fires the matching `HelperEvent::*` via `finish_approval`. Two visible consequences: @@ -474,30 +357,27 @@ reconcile) DAGs use the same queue but skip the approval plumbing. The bundled `hive-forge` container runs on the swarm's forge host (`deploy.forgejo.enable`, see [`../swarm/services.md`](../swarm/services.md)), and hive-c0re mirrors every agent's applied repo into a -private `agent-configs` Forgejo org. `forge::push_config()` pushes `applied/main` plus -every tag to `agent-configs/` after each ref mutation: -the spawn that seeds `deployed/0`, every successful deploy (which -plants `deployed/`) or failed build (`failed/`), and a -sweep at startup. Pushes are best-effort — a missing or stopped -forge never blocks a deploy. +private `agent-configs` Forgejo org. `forge::push_config()` pushes +every tag, then `applied/main`, to `agent-configs/` on every +startup sweep and every rebuild. Forge `main` is branch-protected, so +the forge routinely refuses a push of an established `main`, which +hive-c0re expects. Pushes are best-effort — a missing or stopped forge never +blocks a deploy. Each agent is a **write collaborator on its own** `agent-configs/` repo — so it can push a branch and open a config PR — but not a member of any other agent's, so it can't reach another agent's config through the forge. Branch protection keeps the agent off the `main` push -allowlist and whitelists merging to the `core` user and the `operators` -team, so an agent can't fast-forward its own config or self-merge its -PR (see the End-to-end flow and -[Operator merge in the forge UI](#operator-merge-in-the-forge-ui) above). -hive-c0re passes the tokenised push -URL inline to `git push`, never writing it into +allowlist and allowlists merging to the `operators` team only, so an +agent can't fast-forward its own config or self-merge its PR (see +[Config changes](#config-changes) above). hive-c0re passes the tokenised +push URL inline to `git push`, never writing it into `applied//.git/config`; that repo is RO-bind-mounted into the root agent, and a stored token would leak core's admin credential to an agent. The dashboard deep-links into this org — a `config repo` link -per container row and a `review PR on forge` link per config-PR -approval card. See `docs/web-ui/dashboard.md`. +per container row. See `docs/web-ui/dashboard.md`. ### Submitting agent's view of config repos @@ -548,8 +428,7 @@ the full `/applied` mount: ```sh git -C /agents//config fetch applied git -C /agents//config log applied/main --oneline -git -C /agents//config show applied/refs/tags/deployed/ -git -C /agents//config show applied/refs/tags/failed/ # body = build error +git -C /agents//config show applied/refs/tags/deployed/0 # the seed commit git -C /agents//config show applied/refs/tags/denied/ # body = operator note git -C /agents//config rebase applied/main # base in-flight work on what's deployed @@ -631,11 +510,9 @@ as a regular `system` inbox message so it drives a normal claude turn. `finish_approval` fires an `ApprovalResolved` HelperEvent this way for **every** approval kind's terminal state. A "FYI, check when convenient" event doesn't need a message — those go -through `Coordinator::push_todo`/`push_todo_submitter` instead, a direct -live dial of the target agent's in-container todo socket (same -`UpsertTodo` request in-container producers use); `finish_approval` fires -one of these too for `MergeConfigPr`, *in addition to* -the `ApprovalResolved` HelperEvent above, not instead of it. Legacy +through `Coordinator::push_todo` instead, a direct live dial of the target +agent's in-container todo socket (same `UpsertTodo` request in-container +producers use). Legacy approval rows that predate the submitter column fall back to the root agent. Variants (`hive_sh4re::manager::HelperEvent`): @@ -657,28 +534,17 @@ root agent. Variants (`hive_sh4re::manager::HelperEvent`): The remaining lower-urgency lifecycle notices — `Rebuilt`, `Killed`, `Destroyed`, `NeedsLogin`, `LoggedIn` — are "FYI, check when convenient" events with no reason to drive an immediate turn, so -they deliver via `push_todo`/`push_todo_submitter` (see above) instead +they deliver via `push_todo` (see above) instead of `HelperEvent`: an `agent_todo_socket` push instead of a broker message, `subsystem = "core"`, `key = ":"` for dedup, and a single free-text `summary` (`rebuilt_todo_summary` renders `Rebuilt`'s `ok`/`note`/`sha`/`tag` fields into that string). -Optional `sha` field on `ApprovalResolved` carries the canonical -hive-c0re-vouched commit sha. Optional `tag` carries the deploy -bookkeeping tag — `deployed/` on a successful build or -`failed/` on a failed one, planted by the `MergeConfigPr` deploy. -Both fields are `Option`: `None` on the paths that don't deploy a new -commit (meta-update / deny, and the autoupdate -sweep's `job_queue::templates::rebuild` reapplying the existing main, -or the dashboard `↻ R3BU1LD` button when the lock didn't move). When set, -`git show ` against `/applied//.git` inside the -bootstrap container yields the exact tree the sha referenced. - To add a new lifecycle notice: if it needs to drive an immediate turn (something genuinely urgent, like `ContainerCrash`), add a `HelperEvent` variant + call sites + update `prompts/system.md`'s message-event list. If it's "FYI, check when convenient," call -`push_todo`/`push_todo_submitter` directly instead — no new wire type +`push_todo` directly instead — no new wire type needed. ## Autoupdate on startup diff --git a/docs/agent-lifecycle/persistence.md b/docs/agent-lifecycle/persistence.md index 7d5e76d1..48181a57 100644 --- a/docs/agent-lifecycle/persistence.md +++ b/docs/agent-lifecycle/persistence.md @@ -56,8 +56,8 @@ power-intent registry: per-agent store — see [`/harness/` contents below](#state-dirs-per-agent) for where reminders (and todos) live. -- `approvals` — the queue. `agent / kind (merge_config_pr | spawn | - update_meta_inputs | schedule_prompt) / +- `approvals` — the queue. `agent / kind (update_meta_inputs | + schedule_prompt) / commit_ref / requested_at / status / resolved_at / note`. - `scheduled_prompts` — recurring + one-shot prompt queue. `owner / body / interval_seconds (NULL = one-shot) / diff --git a/docs/getting-started/setup.md b/docs/getting-started/setup.md index a2504ac4..c62f9e2e 100644 --- a/docs/getting-started/setup.md +++ b/docs/getting-started/setup.md @@ -127,8 +127,9 @@ The controller provisions iris's identity, forge user and config repo and start the container. It returns once the scheduler queues the job — watch the swarm UI's job view for progress. -Later config changes are PRs on `agent-configs/iris`, approved by you. → -[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md) +Later config changes are PRs on `agent-configs/iris`; you merge them on the +forge, and the merge deploys them. → +[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md#config-changes) ## Optional · Lock the hive dashboard diff --git a/docs/integrations/forge.md b/docs/integrations/forge.md index 21e38fd8..b395cf7a 100644 --- a/docs/integrations/forge.md +++ b/docs/integrations/forge.md @@ -67,22 +67,15 @@ Two things live in the `agent-configs` Forgejo organization: - A config repo per agent (`agent-configs/`). The agent is a **write collaborator on its own** repo — it can push - config-change branches and open config PRs (Forgejo `pull_request` - webhook at `/webhook/config-pr` queues a `MergeConfigPr` approval; - `hive-c0re/src/forge/config_pr_poll.rs` re-scans every 5 minutes as a - fault-tolerance backstop) — but - `main` is branch-protected: the merge whitelist is the `core` user - (hive-c0re's merge of an approved `MergeConfigPr`) and the `operators` - team (an operator merging in the Forgejo UI, which deploys the merged - commit — see - [approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui)), - the approval whitelist is the `operators` team, and the agent can neither - push `main` directly nor self-merge. hive-c0re's own merge is - fast-forward-only, and hive-c0re never force-pushes (the - `push_config` mirror pushes `main` + the add-only - status tags without force, and treats a non-fast-forward rejection of - `main` after a rolled-back deploy as expected — the forge keeps the - approved history, the `failed/` tag records the divergence). + config-change branches and open config PRs — but `main` is + branch-protected by swarm-controller: the merge and approval + allowlists are the `operators` team, and the agent can neither push + `main` directly nor self-merge. An operator's merge in the Forgejo UI + deploys the merged commit (see + [approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes)). + hive-c0re never force-pushes: the `push_config` mirror pushes the + add-only status tags and `main` without force, and treats a refused + `main` push as expected. Repos stay private, so an agent can't read another agent's config. (Agents remain read-only collaborators on `core/meta`.) hive-c0re also references this repo as the agent's **persistent meta diff --git a/docs/networking/gateway.md b/docs/networking/gateway.md index 0397f172..e6dc0cf0 100644 --- a/docs/networking/gateway.md +++ b/docs/networking/gateway.md @@ -25,7 +25,7 @@ You rarely switch it on yourself. `gateway.enable` defaults to off, and every mo | URL | upstream | when | | --- | --- | --- | | `/` | dashboard dist (static, from `servedFrontend`) | always | -| `/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always | +| `/api/`, `/health/` | hive-c0re (`7000`) | always | | `/api/docs/` | themed Swagger UI dist (static) | always | | `/agent//` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) | | `/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` | @@ -130,7 +130,7 @@ services.hyperhive.gateway.auth = { `hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password ` also works but lands in shell history. -**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent//` routes, which come from `agents.conf` and inherit no auth from `/`. +**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent//` routes, which come from `agents.conf` and inherit no auth from `/`. A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first. @@ -261,10 +261,9 @@ Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }` #### Dashboard: path-based routing (not Accept-header) -hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header: +hive-c0re serves exactly two prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header: - `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live. -- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded. - `location /health/` → hive-c0re: liveness and readiness. - `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`. diff --git a/docs/scheduler/coordinator.md b/docs/scheduler/coordinator.md index 42f187ae..e77713a2 100644 --- a/docs/scheduler/coordinator.md +++ b/docs/scheduler/coordinator.md @@ -42,49 +42,43 @@ because there is no malformed spec to reject. Nix-heavy — hold one of the `buildSlots` permits for the node's duration: -| Node | Wraps | -| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead | -| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node | -| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta | -| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion | -| `DeployWindow` | resource-holding root of the merge-config-PR deploy subtree — declares the build slot, the lease and the meta window, then completes immediately so its children run under them (see _Approvals_ below) | -| `DeployApply` | the deploy's irreversible half: ff-merge the reviewed PR head, two-phase meta deploy, container rebuild | +| Node | Wraps | +| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `Prebuild` | `lifecycle::prebuild_toplevel` — build the toplevel out-of-band while the container keeps serving (its meta preamble is the upstream `MetaSync` node). Skipped when the container is already down; `Swap` builds inline instead | +| `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `RebuildBookkeeping` node | +| `Create` | first-spawn `nixos-container create` proper; assumes the upstream `Provision` node already registered the agent in meta | +| `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion | Cheap — no build slot: -| Node | Behavior | -| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `MergeVerify` | the deploy's pre-merge gate — PR-head drift check, fetch, `verify_commit` eval. Mutates nothing, so a rejection here needs no compensation | -| `DeployTail` | the deploy's `AfterAny` compensation + bookkeeping tail: (1) rolls `applied/main` back from the parked `refs/hyperhive/rollback/` and aborts the staged meta lock when the deploy never confirmed good; (2) mirrors whichever deploy tag landed to the forge config repo, always, best-effort; (3) posts the failing build log back onto the config PR when the deploy failed. Named for (2)/(3), which run on the success path too — not `AbortDeploy`. Infallible by construction | -| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build | -| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#` ref resolves. Store/meta-only, no container yet | -| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop | -| `Start` | mechanical container start — runtime dir + drop-ins, `start_with_fallback`, MCP listener registration, the manager kick. Fanned out by a `Reconcile` that observed `wanted = Up` and the container down | -| `Stop` | mechanical container stop — `nixos-container` kill, MCP listener unregister, the `Killed` manager notify. Fanned out by a `Reconcile` that observed `wanted = Offline` and up | -| `StopForUpdate` | mechanical `nixos-container stop` for the profile swap; never touches `wanted`; noop if already stopped | -| `RebuildBookkeeping` | the swap's Ok-only bookkeeping tail — rev marker, forge/matrix sync, manager kick, rescan, meta-inputs snapshot; `AfterOk(Swap)` so it runs only on a successful swap (the DAG's `EmitRebuilt` tail node emits the `Rebuilt` manager event, not here). Split out of `Swap` for dashboard visibility + retry granularity, declares no resources of its own — a coordinated child of the `AgentWindow` brace | -| `AgentWindow` | pure resource holder — the brace for one agent's rebuild. Declares the build slot + agent lease atomically and holds both for its whole subtree, so `Prebuild` and the `Signal`→`Drain` quiesce window run concurrently instead of one nested under the other. Performs no work; see _Braces_ | -| `Signal` | set the graceful fence + kick, so the harness runs one stop-checkpoint turn | -| `Drain` | await the harness clearing the fence, bounded by the 3-min graceful-stop timeout; resolves ok either way | -| `PauseSignal` | write the pause marker + mark `pause_pending`. No kick, unlike `Signal` — the harness's between-turns poll is already responsive enough, and `Signal`'s kick-message body ("you were just (re)started") would be actively misleading here | -| `PauseDrain` | await the harness reporting `PauseAcknowledged`, bounded timeout; best-effort like `Drain` | -| `DestroyContainer` | `nixos-container destroy` + un-registration (drop from the roster, clear the ephemeral runtime dir). Runs downstream of a `Stop`, so deliberately excluded from `takes_container_down` — the container is already down by the time it claims | -| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step | -| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s | -| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles | -| `FinalizeDeploy` | deploy phase 3 — drop the rollback ref, plant `deployed/`, commit the staged `flake.lock`. The first two git steps are fatal on purpose, so a confirmed-good deploy's outcome and the repo's state can't disagree | -| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent | -| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome | -| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload | -| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots | -| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless | -| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless | -| `WebhookRegister` | one-shot boot-time Forgejo webhook registration (`internal/knowledge` push→pull, `agent-configs` PR→approval). No-op until the core token, hive domain, and HMAC secret are all available. Agentless | -| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless | -| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless | +| Node | Behavior | +| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build | +| `Provision` | first-spawn pre-create provisioning — proposed/applied repos, state subvolume, meta registration (`sync_agents`); runs ahead of `Create` so the `nixos-container create --flake meta#` ref resolves. Store/meta-only, no container yet | +| `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop | +| `Start` | mechanical container start — runtime dir + drop-ins, `start_with_fallback`, MCP listener registration, the manager kick. Fanned out by a `Reconcile` that observed `wanted = Up` and the container down | +| `Stop` | mechanical container stop — `nixos-container` kill, MCP listener unregister, the `Killed` manager notify. Fanned out by a `Reconcile` that observed `wanted = Offline` and up | +| `StopForUpdate` | mechanical `nixos-container stop` for the profile swap; never touches `wanted`; noop if already stopped | +| `RebuildBookkeeping` | the swap's Ok-only bookkeeping tail — rev marker, forge/matrix sync, manager kick, rescan, meta-inputs snapshot; `AfterOk(Swap)` so it runs only on a successful swap (the DAG's `EmitRebuilt` tail node emits the `Rebuilt` manager event, not here). Split out of `Swap` for dashboard visibility + retry granularity, declares no resources of its own — a coordinated child of the `AgentWindow` brace | +| `AgentWindow` | pure resource holder — the brace for one agent's rebuild. Declares the build slot + agent lease atomically and holds both for its whole subtree, so `Prebuild` and the `Signal`→`Drain` quiesce window run concurrently instead of one nested under the other. Performs no work; see _Braces_ | +| `Signal` | set the graceful fence + kick, so the harness runs one stop-checkpoint turn | +| `Drain` | await the harness clearing the fence, bounded by the 3-min graceful-stop timeout; resolves ok either way | +| `PauseSignal` | write the pause marker + mark `pause_pending`. No kick, unlike `Signal` — the harness's between-turns poll is already responsive enough, and `Signal`'s kick-message body ("you were just (re)started") would be actively misleading here | +| `PauseDrain` | await the harness reporting `PauseAcknowledged`, bounded timeout; best-effort like `Drain` | +| `DestroyContainer` | `nixos-container destroy` + un-registration (drop from the roster, clear the ephemeral runtime dir). Runs downstream of a `Stop`, so deliberately excluded from `takes_container_down` — the container is already down by the time it claims | +| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step | +| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s | +| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles | +| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent | +| `EmitRebuilt` | tail of a rebuild/perm-change — emit the agent's `Rebuilt` manager event (ok/fail per outcome, nothing on cancel). One node per agent _and_ per outcome | +| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload | +| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots | +| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless | +| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; each tick queues its own pass, which waits for the resource if one is already live. Agentless | +| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree; each trigger queues its own pass, which waits for the resource if one is already live. Agentless | +| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless | @@ -93,26 +87,21 @@ with its commit under its internal `META_LOCK` mutex, so a standalone commit node would open a dirty-working-tree window between nodes. Two further layers protect the meta repo across _windows_ that span multiple -`META_LOCK` acquisitions — above all the approval deploy's prepare→finalize -span, which keeps a bumped `flake.lock` **staged uncommitted** for the whole -container build: +`META_LOCK` acquisitions: - **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue resource declared by every node kind that mutates the meta repo — `MetaSync`, - `MetaLock`, `WritePermFile`, `Provision`'s agent registration, and - `DeployWindow` — the deploy subtree's root, which holds it across every - phase below it (it declares `Resource::MetaWindow`). Two meta + `MetaLock`, `WritePermFile` and `Provision`'s agent registration. Two meta mutations can therefore never interleave, so no commit lands inside another - node's staged window. It's a queue resource rather than a runtime mutex + node's window. It's a queue resource rather than a runtime mutex because a subtree root holds a resource across its whole subtree, which - a `MutexGuard` (bounded by one executor fn) can't — that's what lets a - multi-node deploy own one window. For the same reason the window must stay + a `MutexGuard` (bounded by one executor fn) can't. For the same reason the window must stay _off_ long store-only work: the rebuild's meta preamble is its own `MetaSync` node, a sibling of (never a parent of) `Prebuild`, so the toplevel build runs outside the window and `buildSlots > 1` still gives concurrent rebuilds across agents. - **Path-limited commits**: the targeted meta committers (perm files, - topology, lock bumps, finalize) commit `-- ` with path-scoped + topology, lock bumps) commit `-- ` with path-scoped dirty checks, so even a non-queue caller (boot migration, destroy's `sync_agents`) can never sweep someone else's staged content into its commit. @@ -220,8 +209,8 @@ resources are free. Resources: 2. **Per-agent lifecycle lease** — keyed on the **node's** agent (agent is per-node; a DAG can span agents) and globally exclusive per agent across all DAGs: acquired either at a container-affecting node (`SetWanted`, - `Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`, - `DeployWindow`) on behalf of a whole coordinated subtree; held by the owning + `Reconcile`, `WriteDropin`, `Create`) or at a **brace** (`AgentWindow`) on + behalf of a whole coordinated subtree; held by the owning DAG until it's terminal, so two DAGs never interleave container ops on the same agent. A DAG touching multiple agents holds one lease per agent. (`SetWanted` is a store write, not a container op, but takes the lease anyway @@ -293,36 +282,9 @@ the dashboard renders one recent-builds list and one number bounds it. ### Approvals -`MergeConfigPr` approvals ride as a four-node deploy subtree: - -``` -DeployWindow (root — build slot + lease + meta window, no work of its own) -├── MergeVerify drift gate, fetch, verify_commit -├── DeployApply AfterOk(verify) park rollback ref, ff-merge, deploy -└── DeployTail AfterAny(apply) compensate, mirror to forge -``` - -The root holds its resources across the whole subtree, so the two-phase -`prepare_deploy` / `finalize_deploy` span keeps its staged `flake.lock` -protected even though the phases are separate nodes. Splitting them buys -three things a single opaque node couldn't have: per-phase visibility on the -dashboard, a `MergeVerify` failure that provably mutated nothing, and a -compensation step that survives a hive-c0re restart — `DeployApply` parks the pre-merge -`applied/main` in `refs/hyperhive/rollback/`, not in a -local variable, so `DeployTail` can still undo a half-finished deploy after a -crash. - -`DeployWindow` declares all three resources (build slot, lease, meta window) -on itself rather than letting each phase declare its own, because the queue -acquires a node's resources atomically (all-or-nothing): a child that took -the build slot while its parent held the meta window could block waiting for -a resource its own parent already committed to, a lock-ordering hazard that -one multi-resource root avoids by construction. - -`UpdateMetaInputs` approvals map onto the ordinary -`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag` -exactly once when **any** approval-carrying DAG settles terminal — deploys -included, since their outcome is the DAG's own state (including +`UpdateMetaInputs` approvals map onto the ordinary `meta-update` shapes. +The scheduler fires `actions::resolve_approval_dag` exactly once when +**any** approval-carrying DAG settles terminal (including cancelled-while-queued, which fails the approval instead of dangling it). ### Wire shape @@ -398,16 +360,12 @@ Key operations: - **`sync_agents`** (idempotent) — render `flake.nix` for the current agent set, init the repo on first call, relock if the rendered contents changed, commit. Called by spawn / destroy / startup migration. -- **`prepare_deploy` + `finalize_deploy` / `abort_deploy`** — two-phase for the - `MergeConfigPr` deploy path so a failed `nixos-container update` leaves no orphan - commit in meta. Prepare writes the new lock without committing; finalize commits - with the deploy message; abort restores the lock. - **`lock_update_hyperhive`** — one-shot for the boot-reconcile path (the sweep DAG's `MetaLock` node): bumps the `hyperhive` input lock and commits; the scheduler fans out the agent rebuilds on completion. Every public `meta.rs` operation takes the module's internal `META_LOCK` -mutex, so concurrent job-queue nodes (and the approval deploy pipeline) never +mutex, so concurrent job-queue nodes never race on the repo's `.git/index.lock`. --- @@ -449,18 +407,6 @@ Sequence for a rebuild DAG (each step is its own queue node): in-container activation script transitions old → new. Holds no build slot, so the next DAG's `Prebuild` overlaps the container boot. -The approval deploy uses this same chain rather than a rebuild path of its -own. Its `DeployApply` node doesn't build: it merges, opens the two-phase -meta deploy, and returns the chain above as a subgraph the scheduler grafts -into the live DAG under that node. A `FinalizeDeploy` node gated on the -graft's completion then plants the deploy tag — so `Reconcile`'s success -answers "did the agent come back up?" the same way it does for every -other rebuild, instead of a fused inline start. - -The grafted nodes land _inside_ `DeployWindow`'s subtree, so they re-enter -the meta window and build slot it already holds rather than deadlocking -against it. - ### Cold-start fallback `start` after `update` can exit non-zero when packages are **removed** between diff --git a/docs/scheduler/jobq.md b/docs/scheduler/jobq.md index 0bf926fc..990a3362 100644 --- a/docs/scheduler/jobq.md +++ b/docs/scheduler/jobq.md @@ -3,7 +3,7 @@ Long-running work runs through a job graph. The swarm controller keeps one for swarm-level work — creating an agent's identity, forge user and config repo. Each hive's hive-c0re keeps its own for container operations — -rebuild, first-spawn, a config-PR deploy, power changes. This page explains +rebuild, first-spawn, power changes. This page explains what the job queue _is_, as a general idea, independent of what either uses it for. For the hive-c0re step catalogue and the engineering internals (scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md) diff --git a/docs/swarm/README.md b/docs/swarm/README.md index ad22eec1..34e7e3c1 100644 --- a/docs/swarm/README.md +++ b/docs/swarm/README.md @@ -227,9 +227,9 @@ one per hive. It ensures them at start and every five minutes after - the orgs `agent-configs`, `internal` and `agents`, plus each mirror's owner org; - the empty `operators` merge-gate team in `agents` and `agent-configs`; -- the `main` merge gate on every `agent-configs` repo: merge whitelist = - the `operators` team and the `core` user, approval whitelist = the - `operators` team. The controller leaves a repo with no `main` rule alone; +- the `main` merge gate on every `agent-configs` repo: merge and approval + whitelists = the `operators` team, and no user. The controller leaves a + repo with no `main` rule alone; - the pull-mirrors from `deploy.forgejo.mirrors` on the controller's host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds); - `internal/docs` (private) and `internal/knowledge` (public, with a @@ -258,7 +258,7 @@ decision, not an event to adjudicate. A `config-pr` delivery reporting a PR merged into `main` queues a deploy of its `merge_commit_sha` on the one hive whose wanted state places the agent; with no such hive, or several, the controller deploys nothing. See -[approvals.md § Operator merge in the forge UI](../agent-lifecycle/approvals.md#operator-merge-in-the-forge-ui). +[approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes). **`internal/knowledge` is on that path.** The controller's is the only hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a @@ -267,10 +267,10 @@ would take delivery away from the first rather than add a recipient. -**The `agent-configs` org isn't.** Each hive registers its own -`pull_request` hook there, so that repo has two — the hive's and the -controller's — and **both are expected; don't delete either.** Removing -a hive's stops it acting on config PRs; removing the controller's stops +**The `agent-configs` org is on it too.** The controller's hook is the +only one that acts on config PRs. A hive's `/webhook/config-pr` hook left +on the org by an older release delivers to a route no hive serves; delete +it in the org's webhook settings. Removing the controller's hook stops forge-UI merges from deploying until its next start recreates it. diff --git a/docs/tools/hivectl.md b/docs/tools/hivectl.md index 5bae57dd..293ed79e 100644 --- a/docs/tools/hivectl.md +++ b/docs/tools/hivectl.md @@ -40,8 +40,8 @@ hivectl forge reconcile-config iris --verbose # include the full diff, not applied config checkout and its forge `agent-configs/` `main`, then reconciles. `--from forge` resets the local checkout to forge `main` (takes effect on the next deploy — it doesn't autorebuild). `--from local` isn't - supported yet (forge `main` is core-only branch-protected; resolve via a - config PR). With no `--from` it prompts for the direction after the diff. + supported yet (forge `main` is branch-protected; resolve via a config + PR). With no `--from` it prompts for the direction after the diff. ## Matrix diff --git a/docs/trust-boundary/security.md b/docs/trust-boundary/security.md index f9cd3dca..3b8dd5e7 100644 --- a/docs/trust-boundary/security.md +++ b/docs/trust-boundary/security.md @@ -123,14 +123,17 @@ checkpoints**, not about sandboxing the agent from its own tools: highest-value action. On the **internal forge this is technically enforced, not just convention**: agents can't create repos (`max_repo_creation = 0`), and `main` on an `agent-configs/` repo carries swarm-controller's - branch protection: merge allowlisted to the `operators` team, with one - approval from it. Existing `agents/` repos carry the same merge gate. + branch protection: merge and approval allowlisted to the `operators` team, + which swarm-controller converges on every config repo. An operator's merge + there is also what deploys the config. Existing `agents/` repos carry + the same merge gate. An agent (a write collaborator, not a repo admin) can neither change those settings nor merge its own PR. It's **not** set up for external VCS (GitHub etc.), though — there, operator-merge is process + accepted risk, not a technical control. -- **Approvals** — config changes, schedule additions, and other - blast-radius-y operations route through the operator approval queue +- **Approvals** — schedule additions and other blast-radius-y operations + route through the operator approval queue; config changes are config PRs + an operator merges on the forge (see [`approvals.md`](../agent-lifecycle/approvals.md)). ### Capability = accepted risk diff --git a/docs/turn-loop/mcp.md b/docs/turn-loop/mcp.md index cf3b3139..3bf22965 100644 --- a/docs/turn-loop/mcp.md +++ b/docs/turn-loop/mcp.md @@ -70,12 +70,7 @@ approval (`Coordinator::notify_submitter`, looked up from the authenticated socket caller at submit time; a row with no recorded submitter falls back to the manager, `ruth`). `ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`, -hardcoded — `hive-c0re/src/workers/crash_watch.rs`). A `MergeConfigPr` -approval's rebuild additionally pushes a `rebuilt:` todo to that -same submitter (`Coordinator::push_todo_submitter`, `subsystem = -"core"`), which wakes a turn (the todo-wake path — see [Turn -outcomes](README.md#turn-outcomes)) via a generic "call -`get_loose_ends`" prompt rather than the event body itself. Lifecycle +hardcoded — `hive-c0re/src/workers/crash_watch.rs`). Lifecycle transitions the job-queue scheduler or crash watcher drive directly — stop/kill, destroy, a flake-rev login or logout state change — reach no individual agent: they publish onto a swarm-wide NATS diff --git a/docs/web-ui/README.md b/docs/web-ui/README.md index e4adb0d0..07d02307 100644 --- a/docs/web-ui/README.md +++ b/docs/web-ui/README.md @@ -48,24 +48,24 @@ and quick links (stats, screen, forge profile). Select the name to open its terminal and watch it work in real time. **Approve something an agent is waiting on.** Y3R C4LL is the one tab -worth checking regularly — it's everything that needs _you_: approvals -for config changes. The tab's count pill tells you at a glance if +worth checking regularly — it's everything that needs _you_: an agent's +request to schedule a prompt. The tab's count pill tells you at a glance if anything's pending. -**Approve or reject a config change.** Agent config changes (new -packages, env vars, MCP servers) go through an approval queue rather -than landing automatically — you'll see them on Y3R C4LL, with a diff -of what's changing. +**Review a config change.** Agent config changes (new packages, env +vars, MCP servers) are pull requests on the agent's `agent-configs` +repo. Review and merge them on the forge; the merge deploys the change. +See [`approvals.md`](../agent-lifecycle/approvals.md#config-changes). **Start, stop, restart, or rebuild an agent.** Select one or more agents on SW4RM (select the icon) and use the selection bar, or use the per-agent `⋮` menu on a single row. Rebuilding re-applies that agent's -current config; use it after approving a change, or whenever an agent +current config; use it whenever an agent shows as "needs update." **Watch a build.** BU1LDS shows the rebuild queue live, plus a streaming log of whatever's currently building. Useful right after -approving a change or bumping a flake input. +merging a config change or bumping a flake input. **Grant or revoke a tool/capability.** P3RM1SS10NS is a checkbox matrix — rows are agents, columns are tool groups or capabilities. Nothing diff --git a/docs/web-ui/dashboard.md b/docs/web-ui/dashboard.md index 16400404..87dfb675 100644 --- a/docs/web-ui/dashboard.md +++ b/docs/web-ui/dashboard.md @@ -877,11 +877,10 @@ renderApprovals`) with three stacked sections: right-aligned `requested ago` relative time from `ApprovalView.requested_at`. Glyph and chip vary by kind: - | kind | glyph | chip | sha shown | - |---|---|---|---| - | `merge_config_pr` | `⇒` | `merge-pr` | PR-head sha (`sha_short`) | - | `update_meta_inputs` | `↻` | `meta-update` | — | - | `schedule_prompt` | `⏱` | `schedule` | — | + | kind | glyph | chip | + |---|---|---| + | `update_meta_inputs` | `↻` | `meta-update` | + | `schedule_prompt` | `⏱` | `schedule` | The chip ticks live every second via a `data-requested-at` @@ -889,12 +888,9 @@ renderApprovals`) with three stacked sections: the request has been pending ≥ 1h so a stale approval stands out; the `.stale` class flips precisely at the 3600s boundary rather than at the next `renderApprovals` call. -- **what-changed body** — the submitting agent's description, then - kind-specific drill-in triggers: - - `merge_config_pr`: `↳ review PR on forge ↗` deep-links the - config PR into `agent-configs//pulls/` (shown - only when `forge_present` is true and `pr_number` has a value). The config diff - lives on the forge PR itself — no inline diff side-panel. +- **what-changed body** — the submitting agent's description, then the + kind's payload: the inputs to bump (`update_meta_inputs`) or the prompt to + schedule (`schedule_prompt`). - **decision actions** — `◆ APPR0VE` and `DENY`. Deny pops a `prompt()` for an optional reason carried to the submitting agent as `HelperEvent::ApprovalResolved.note`.