mara's review on #2436: no submit-await-submit composition, even server-side. Adds Template::GracefulRestart (Signal -> Drain -> StopForUpdate -> Reconcile, wanted=Up) mirroring how Restart already does StopForUpdate -> Reconcile, plus submit::graceful_restart and templates::graceful_restart. handle_restart_scoped now submits exactly one DAG per agent up front for both the graceful and non-graceful case -- no await_dags in the loop anymore.
21 KiB
hive-c0re coordinator internals
Architecture notes for the hive-c0re coordinator daemon's internal subsystems.
For the public API surface (dashboard, socket protocol) see docs/conventions.md
and docs/persistence.md.
Job queue
Every container/meta operation (rebuild, meta-update, first-spawn, power
changes) is submitted to the global job-DAG queue (hive-c0re/src/job_queue/)
as a DAG of primitive nodes. One scheduler task drives all DAGs;
concurrency comes from the resource classes below, not from multiple workers.
The old special cases — the graceful-stop watcher thread, the deferred-start
fast-lane follow-up, the meta-update cascade pre-enqueue — are all just DAG
shapes now.
Two levels: DAG and node
The DAG is the unit of dedup / cancel / approval-resolution and the
dashboard group; the node is the unit of scheduling / execution /
build-log / step label. Deps are intra-DAG edges only (AfterOk by default:
the dep must succeed, a failed/cancelled dep cancels the dependent —
cancel-downstream). Cross-DAG ordering comes from the per-agent lease + dedup,
never from edges between DAGs. Submit-time validation (petgraph toposort)
rejects cyclic specs outright, fixing the old queue's "circular dep silently
deadlocks" caveat.
Node inventory (primitives)
Nix-heavy — hold one of the buildSlots permits for the node's duration:
| Node | Wraps |
|---|---|
Prebuild |
meta sync_agents + optional per-agent relock + lifecycle::prebuild_toplevel — build the toplevel out-of-band while the container keeps serving |
Swap |
drop-in rewrite + nixos-container update profile-swap (requires the container stopped) + the post-rebuild bookkeeping tail (rev marker, Rebuilt event, forge/matrix sync, kick, rescan) |
Create |
first-spawn provisioning + nixos-container create (atomic build+create) |
MetaLock |
meta flake lock bump (lock_update / boot-sweep lock_update_hyperhive, commit fused — see below); fans out child Rebuild DAGs on completion |
ApprovalDeploy |
the opaque apply-commit / merge-config-PR pipeline (see Approvals below) |
Cheap — no build slot:
| Node | Behavior |
|---|---|
Reconcile |
idempotent power converge: read wanted (below) + observed state; start if Up & down (cold-start fallback included), stop if Offline & up, else noop |
StopForUpdate |
mechanical nixos-container stop for the profile swap; never touches wanted; noop if already stopped |
Signal |
set the graceful fence + kick, so the harness runs one stop-checkpoint turn |
Drain |
await the harness clearing the fence, bounded by the 3-min graceful-stop timeout; resolves ok either way |
WriteDropin |
set_nspawn_flags + set_resource_limits + daemon-reload |
WritePermFile |
commit tool-groups.json / capabilities.json (single git commit under META_LOCK) + emit the P3RM1SS10NS snapshots |
There is deliberately no GitCommit node: meta.rs fuses each mutation
with its commit under its internal META_LOCK mutex, so a standalone commit
node would open a dirty-working-tree window between nodes.
Two further layers protect the meta repo across windows that span multiple
META_LOCK acquisitions — above all the approval deploy's prepare→finalize
span, which keeps a bumped flake.lock staged uncommitted for the whole
container build:
- The deploy-window gate (
meta::exclusive()): every executor that mutates the meta repo (Prebuild's sync+relock,MetaLock,WritePermFile,Create's agent registration, andApprovalDeployfor its whole span) holds this async mutex for its mutation span, so no commit can land inside another node's staged window.Prebuilddrops it before the long toplevel build (store reads only), preservingbuildSlots > 1concurrency. - Path-limited commits: the targeted meta committers (perm files,
topology, lock bumps, finalize) commit
-- <their paths>with path-scoped dirty checks, so even a non-queue caller (boot migration, destroy'ssync_agents) can never sweep someone else's staged content into its commit.
Every operation as a DAG
rebuild(a): Prebuild(a) → StopForUpdate(a) → Swap(a) →(after-any) Reconcile(a)
graceful-stop(a): [wanted=Offline] Signal(a) → Drain(a) → Reconcile(a)
restart(a): [wanted=Up] StopForUpdate(a) → Reconcile(a)
graceful-restart(a): [wanted=Up] Signal(a) → Drain(a) → StopForUpdate(a) → Reconcile(a)
start(a): [wanted=Up] Reconcile(a) (stale rev ⇒ upgraded to rebuild)
stop(a): [wanted=Offline] Reconcile(a)
spawn(a): [wanted=Up] Create(a) → WriteDropin(a) → Reconcile(a)
perm-change(a): WritePermFile(a) → «rebuild subgraph»
meta-update(inp): MetaLock(inp) → «fan-out rebuild(a) per affected agent»
boot: (if any rev marker stale) MetaLock(hyperhive) → «fan-out rebuild»;
plus Reconcile(a) for every drifted agent
Notable collapses:
rebuildis one uniform shape — nowas_runningbranch.StopForUpdatenoops when already down; the tailReconcileauto-noops the start whenwanted = Offline(a rebuild of a deliberately-stopped agent leaves it stopped).- The swap-failure recovery-start is structural:
Reconciledeps onSwapwith the oneAfterAnyedge in the system — it runs afterSwapterminal ok or fail, bringing a wanted-up agent back on its old config. - Deferred start is automatic:
Reconcileholds no build slot, so the next DAG'sPrebuildstarts as soon asSwapfrees the slot. - Graceful stop needs no watcher thread:
Signal/Drainare cheap, so a whole-hive graceful stop fires every agent's signal immediately and all drains overlap; each DAG's tailReconciledoes the actual stop. - The meta-update cascade fans out on completion:
MetaLock's executor computes the affected agent set after the bump lands and appends childrebuildDAGs (parent_idset,relock = falseso the children don't revert the bump). A failed bump fans out nothing — no cancel-children dance.
Desired-state (spec vs status)
Per-agent power intent — wanted: Up | Offline — is durable as the
agent_power table in the coordinator DB (hive-c0re/src/stores/power.rs).
container_view remains the observed status; Reconcile nodes converge the
two. Setting wanted is never a queued node: the submit layer
(job_queue/submit.rs) writes the row synchronously, then submits the DAG
whose Reconcile reads the fresh value — rapid toggles are last-writer-wins.
Power toggles never commit to the meta repo. Every operator power surface —
dashboard buttons, the MCP tools, and hivectl stop/start/restart/kill —
rides the queue through that submit layer, so intent, lease serialization,
and crash-watch suppression can't drift per surface; the only direct starts
left are the root-agent bootstrap and infra containers (no lease, no
harness). Cancelling a still-queued power DAG reverts wanted to the
observed state — a cancel means "don't do it", not "do it later". Agents
without a row are seeded from observed state on first touch (running ⇒
Up); destroy removes the row.
The admin-socket responses carry the submitted DAG ids; hivectl polls
HostRequest::QueueDag (~1s) and prints a progress line per DAG — roll-up
glyph, template, agent, node chain with the running node's step label — so
CLI verbs block until their jobs finish (--no-wait opts out; failures exit
non-zero). Fan-out children joining a polled parent show up in the same
loop.
Scheduler semantics
A node is ready when it's Queued, every dep is satisfied, and its
resources are free. Resources:
- Build slots —
services.hyperhive.c0re.buildSlotspermits (default 1), held by nix-heavy nodes for the node's duration. - Per-agent lifecycle lease — DAG-scoped: acquired at the DAG's first
container-affecting node (
StopForUpdate,Swap,Signal,Drain,Reconcile,WriteDropin,Create,ApprovalDeploy), held until the DAG is terminal, so two lifecycle DAGs for one agent never interleave their container ops. Lease-exempt:Prebuild,MetaLock,WritePermFile— they touch the store / meta, not the running container, which is exactly why a stop can land while another DAG's prebuild is still building.
Among simultaneously-ready nodes competing for a resource, DAG-submit order wins (FIFO) so bulk operations drain predictably. The scheduler also owns the DAG-lifetime transient guard (dashboard pill + crash-watch suppression), created on lease acquisition and dropped when the DAG settles terminal.
The queue is in-memory only and lost on hive-c0re restart — deliberate: desired state is re-derived at boot from the DB + rev markers (see Boot reconcile), so there is no durable-recovery machinery to go wrong.
Dedup, cancel, history
Dedup at DAG granularity: a repeat submit against a DAG whose roll-up is
still Queued with the same (template, agent, parent_id, approval_id) —
plus inputs for meta-updates and the perm-type discriminant for perm
changes — returns the existing id and appends an "also requested by …" line.
parent_id in the key keeps a cascade child from collapsing into a
standalone or sweep rebuild. Running/terminal DAGs never dedup.
Cancel only applies to still-fully-queued DAGs (an in-flight nix build isn't
interruptible); cancel_children cancels a parent's still-queued child DAGs.
Roll-up state: Failed if any node failed, else Running / Queued /
Cancelled / Done. The snapshot retains the 5 most recent terminal DAGs
per template.
Approvals
ApplyCommit / MergeConfigPr approvals ride as single-node
ApprovalDeploy DAGs: the two-phase prepare_deploy / finalize_deploy /
abort_deploy meta orchestration stays inside actions.rs in v1
(deliberately not modeled as scheduler nodes) and resolves the approval
itself. Spawn and UpdateMetaInputs approvals map onto the ordinary
spawn / meta-update shapes; the scheduler fires
actions::resolve_approval_dag exactly once when such a DAG settles
terminal (including cancelled-while-queued, which fails the approval instead
of dangling it).
Wire shape
RebuildQueueChanged { seq, queue: [DagView…] } (event name kept). Each
DagView carries the old entry-level fields (id, kind = template string,
roll-up state, agent, source, parent_id, reason, timestamps,
inputs, approval_id) plus nodes: [NodeView…] — per-node kind, deps,
state, step, build_log_id, timestamps, error. Step labels and build
logs are per-node; the dashboard renders the node chain on each queue
card and keys the live-log panel off the running node.
Container view
container_view.rs maintains an in-memory snapshot of every nixos-container's
systemd service state. It is polled on coordinator startup and re-scanned after
every lifecycle operation (spawn, rebuild, kill) so the dashboard always reflects
the actual container status without a live nixos-container list call on each
render.
Boot reconcile
On startup, auto_update::run classifies every agent by rev freshness (the
per-agent .{name}.hyperhive-rev marker under /var/lib/hyperhive/applied/
vs the current flake path) and persisted wanted intent, then:
-
Config path — when any marker is stale, submit one
StartupSweepDAG: aMetaLock(hyperhive input bump, non-fatal) that fans outRebuildchildren for the stale agents whosewanted = Up(topology-sorted, parents first). Stale but wanted-offline agents get no boot-time nix work — their rebuild happens on their next start (the start submit path upgrades a stale start to rebuild+start), which is also why the lock bump runs even when every stale agent is offline: those later start-upgrades must build against the bumped lock. Each child rebuild's tailReconcilebrings the agent (back) up, covering both the running-stale and stopped-but-wanted-up cases. -
Power path — every agent whose observed state drifted from
wantedgets a plainReconcileDAG (kind = reconcile, sourceauto_update).
Booting with no config change performs no meta commit — only reconciles.
The sweep reason records the rebuild / deferred / up-to-date counts so the
operator sees at a glance how much work the boot triggered. Agents without an
agent_power row are seeded from observed state during classification (the
one-time migration; thereafter the DB is authoritative).
Meta flake
meta.rs owns the single coordinator-managed flake at /var/lib/hyperhive/meta/.
This flake consumes every agent's applied config repo as a flake input and exports
one nixosConfiguration per agent. Container lifecycle ops drive the lock file so
meta's git log is the system-wide deploy audit trail.
Key operations:
sync_agents(idempotent) — renderflake.nixfor the current agent set, init the repo on first call, relock if the rendered contents changed, commit. Called by spawn / destroy / startup migration.prepare_deploy+finalize_deploy/abort_deploy— two-phase for theRequestApplyCommitpath so a failednixos-container updateleaves no orphan commit in meta. Prepare writes the new lock without committing; finalize commits with the deploy message; abort restores the lock.lock_update_hyperhive— one-shot for the boot-reconcile path (the sweep DAG'sMetaLocknode): bumps thehyperhiveinput lock and commits; the scheduler fans out the agent rebuilds on completion.
Every public meta.rs operation takes the module's internal META_LOCK
mutex, so concurrent job-queue nodes (and the approval deploy pipeline) never
race on the repo's .git/index.lock.
Container lifecycle (lifecycle.rs)
Every container operation ultimately calls into lifecycle.rs. Two paths exist:
rebuild (existing container) and spawn (first-time creation).
Rebuild path (existing container)
Goal: apply the new system profile and any EXTRA_NSPAWN_FLAGS / drop-in changes
in a single start, with minimum downtime.
nixos-container update only runs systemctl reload container@<c> when the
container is already up (per isContainerRunning in nixos-container.pl). Stopping
first turns update into a boot-style operation: it builds + nix-env --sets the
new profile and skips the in-container switch-to-configuration. The subsequent
start then applies both the new profile and any EXTRA_NSPAWN_FLAGS changes in
one go, rather than the double-bounce a live update would trigger.
Sequence for a rebuild DAG (each step is its own queue node):
Prebuild— build the newsystem.build.toplevelbefore stopping. The container keeps serving the previous generation while eval + fetch + build happen out-of-band.nixos-container updatethen finds the result cached and skips straight to the profile-swap. Build failures surface here, before the running container is touched. (Runs even for a stopped container — same total nix work, one uniform DAG shape.)StopForUpdate— bring the container down (noop when already stopped).Swap—nixos-container update --flake meta#<name>profile-swap (near-instant after the prebuild).Reconcile— boot into the new generation whenwanted = Up; the in-container activation script transitions old → new. Holds no build slot, so the next DAG'sPrebuildoverlaps the container boot — the old "deferred start" split, now structural.
The approval apply-commit pipeline still drives lifecycle::rebuild_no_meta
(the fused stop/update/start path with an inline start) inside its
ApprovalDeploy node, because it verifies the agent comes back up before
finalizing the deploy tag.
Cold-start fallback
start after update can exit non-zero when packages are removed between
generations: the old-generation activation script references units that no longer
exist in the new closure, causing systemd to exit non-zero. The container may be
half-started at that point.
Fallback: stop (graceful SIGTERM drain) → kill (SIGKILL any lingering processes)
→ start (clean cold-start, no generation transition, new activation runs cleanly).
Both errors are preserved and surfaced if the cold-start also fails. The fallback
lives in lifecycle::start_with_fallback, shared by the apply-commit deploy's
inline start and every Reconcile node's start action.
Spawn path (new container)
For a first-time create, nixos-container create is atomic: if the build fails,
no container record is left to clean up. A separate prebuild would just duplicate
the eval, so it's skipped. Sequence: create --flake meta#<name> → write nspawn
flags → systemctl daemon-reload → start.
Prebuild attr path
nix build does not auto-resolve meta#<name> against nixosConfigurations the
way nixos-container does internally. The explicit attr path
<flake-root>#nixosConfigurations.<name>.config.system.build.toplevel is required;
using the bare meta#<name> ref would make nix look in packages, legacyPackages,
or the flake root directly — none of which exist in the rendered meta flake.
Host-level resource + performance options
A handful of services.hyperhive.c0re.* options tune container resource
limits, build parallelism, and first-spawn latency.
Build slots
buildSlots (default 1) sets how many nix-heavy job-queue nodes
(prebuilds, profile swaps, first-spawn creates, meta lock bumps) run
concurrently. The default serializes all heavy nix work like the pre-DAG
rebuild queue did; raise it on hosts with the cores/RAM to build several
agent toplevels at once. Per-agent correctness is independent of the count —
each agent's container-affecting ops serialize on its lifecycle lease
regardless.
Container resource limits
agentCpuQuota and agentMemoryMax map directly to systemd
CPUQuota= and MemoryMax=. hive-c0re writes a
container@h-<name>.service.d/ drop-in file on each spawn and
rebuild, so changes take effect on the next lifecycle op without
requiring a host rebuild.
| Option | Default | Description |
|---|---|---|
services.hyperhive.c0re.agentCpuQuota |
"200%" |
CPU cap per agent, as a percentage of one core ("200%" = 2 cores). Raise if agents hit CPU limits during builds or heavy tool use. |
services.hyperhive.c0re.agentMemoryMax |
"4G" |
Memory cap per agent. Raise for agents that run large nix builds or hold big in-memory data. |
For a hive-wide cap across all containers together, set
systemd.slices.machine.serviceConfig.CPUQuota in your NixOS
config — all nspawn containers live in machine.slice.
Pre-building agent templates
preBuildAgentTemplates (default false) causes the host NixOS
build to pre-fetch the per-container system closures
(agent-base + manager toplevels) into /nix/store, instead of
leaving that work to the first nixos-container start. The
trade-off:
- On (recommended for x86_64 hosts that care about first-spawn
latency): the first
nixos-container startfor any new agent completes in seconds because nothing is left to fetch. Cost: the full nixpkgs runtime closure + claude-code + the harness binary are added to the host system closure (low single-digit GB additional). - Off (default): the host closure stays lean; the first spawn does all the eval + fetch work at runtime (can take several minutes on a fresh store).
Note: toplevels are pinned to x86_64-linux. Enabling on an
aarch64 host forces a cross-compilation or remote-builder build,
which is almost never desired. Leave off on non-x86 hosts.
See also
docs/approvals.md— approval flow + scheduled promptsdocs/persistence.md— SQLite schema, state-dir layoutdocs/conventions.md— wire protocol, recipient sentinelsdocs/agent-hierarchy.md— topology and parent/child relations