hyperhive/docs/coordinator.md
damocles eae0e875cf feat(#1086): serialize perm changes through rebuild queue
add QueueKind::PermChange — dashboard tool-group and capability
handlers no longer write the shared JSON files inline. instead they
enqueue a PermChange entry; the FIFO worker applies the file write
then calls rebuild_agent so the updated env var takes effect.

concurrent batch-apply actions for different agents previously raced
on tool-groups.json / capabilities.json (last write wins, earlier
change silently dropped). serialising through the queue prevents this.

dedup check extended with perm-type discriminant so tool-groups and
capabilities changes for the same agent are kept as distinct entries
and never collapse into one slot.
2026-06-02 16:42:22 +02:00

9.3 KiB

hive-c0re coordinator internals

Architecture notes for the hive-c0re coordinator daemon's internal subsystems. For the public API surface (dashboard, socket protocol) see docs/conventions.md and docs/persistence.md.


Rebuild queue

Every long-running container/meta operation (rebuild, meta-update, first-spawn) goes through the global rebuild queue (hive-c0re/src/rebuild_queue.rs). A single background worker drains it in FIFO order so two nixos-container update runs on the same agent never overlap, and a fresh agent rebuild never races a meta-update's lock bump.

Why one queue

Before the rebuild queue landed, four independent call paths could fire auto_update::rebuild_agent concurrently:

  • Dashboard manual rebuild button
  • update-all / meta-update cascade
  • Approval handler (apply-commit / spawn)
  • Startup auto-update sweep

Nothing serialised them. nix-daemon serialises the actual store ops, but the rest of rebuild_agent (token sync, kick, rescan, lock-bump emit) interleaved unpredictably. The single-worker queue gives operators a visible, ordered runway and lets the UI render "what's about to happen" instead of "something might be happening somewhere."

Queue kinds

Kind Description
Rebuild Single-agent rebuild. Covers manual, approval-driven, auto-update, and meta-update cascade variants — all funnel through the same path.
MetaUpdate nix flake update on the meta flake. The worker runs the lock bump itself, then enqueues a cascade of Rebuild entries with parent_id set to the meta-update's id.
Spawn First-deploy of a new agent (approval-driven). Same serialisation as Rebuild from the operator's POV.
Destroy For future use (destroy --purge does real I/O). Variant exists so the wire shape doesn't change later; not currently routed through the queue.
Restart Stop + start a container without touching config (~5-10s). Routed through the queue so it serialises against in-flight rebuilds for the same agent — prevents a restart racing a rebuild mid-flight. Sources: dashboard ↺ button, manager restart MCP tool.
PermChange Write a tool-group or capability change to the shared JSON file (tool-groups.json / capabilities.json), then rebuild the agent so the updated HIVE_TOOL_GROUPS / HIVE_CAPABILITIES env var takes effect. Serialising the file write through the queue prevents concurrent dashboard batch-apply actions from racing on the shared file.

Intentionally not queued (sub-second ops): start, stop, kill.

Dedup

Enqueueing (kind, agent) that already has a Queued entry returns the existing entry's id and appends the new request as an "also requested by …" line. Running entries do not dedup — a re-queue during a run is legitimate (something changed since the current run started).

Sources

Source Meaning
Manual Operator clicked rebuild / update-all / meta-update on the dashboard, or any other direct human action (CLI, manager tool).
AutoUpdate Legacy startup-sweep source (flat, no parent). Replaced by StartupSweep for new boots.
StartupSweep Child of a StartupSweep parent entry; boot-time per-agent rebuild with the sweep as the visual group header.
Approval Triggered by an operator-approved ApprovalKind::{Spawn, ApplyCommit}.

Cascade parent tracking

MetaUpdate and StartupSweep entries fan out Rebuild children, each carrying parent_id = <parent_id>. The dashboard groups children under their parent in the queue panel so the operator sees the whole cascade as a tree, not a flat list.

Step labels

Each queue entry has a mutable step: Option<String> field that the worker updates as it progresses through lifecycle phases ("nix build", "nixos-container stop", "nixos-container update", "nixos-container start"). The dashboard polls /api/state and renders the current step beneath the running entry so the operator can see which phase is taking time.


Container view

container_view.rs maintains an in-memory snapshot of every nixos-container's systemd service state. It is polled on coordinator startup and re-scanned after every lifecycle operation (spawn, rebuild, kill) so the dashboard always reflects the actual container status without a live nixos-container list call on each render.


Auto-update sweep

On startup, auto_update.rs rebuilds every known container unconditionally. nixos-container update is a no-op at the nix level when nothing changed (same store path), so the cost is low and avoids rev-marker staleness — all agents always need an update pass when any meta commit lands.

auto_update::run enqueues a single StartupSweep parent entry (kind = startup_sweep, agent = "hyperhive") followed by per-agent Rebuild children (source = startup_sweep, parent_id = sweep_id). The worker processes the parent by bumping the meta hyperhive input lock, then transitions it to Done. The child rebuilds drain sequentially through the queue; the dashboard renders them nested under the parent so the operator can see the whole boot-time sweep in one group.

Before this change, each boot enqueued flat Rebuild entries with source = AutoUpdate and no parent — visible but ungrouped.

Meta flake

meta.rs owns the single coordinator-managed flake at /var/lib/hyperhive/meta/. This flake consumes every agent's applied config repo as a flake input and exports one nixosConfiguration per agent. Container lifecycle ops drive the lock file so meta's git log is the system-wide deploy audit trail.

Key operations:

  • sync_agents (idempotent) — render flake.nix for the current agent set, init the repo on first call, relock if the rendered contents changed, commit. Called by spawn / destroy / startup migration.
  • prepare_deploy + finalize_deploy / abort_deploy — two-phase for the RequestApplyCommit path so a failed nixos-container update leaves no orphan commit in meta. Prepare writes the new lock without committing; finalize commits with the deploy message; abort restores the lock.
  • lock_update_hyperhive — one-shot for the auto-update path: bumps the hyperhive input lock, commits, cascades agent rebuilds.

Container lifecycle (lifecycle.rs)

Every container operation ultimately calls into lifecycle.rs. Two paths exist: rebuild (existing container) and spawn (first-time creation).

Rebuild path (existing container)

Goal: apply the new system profile and any EXTRA_NSPAWN_FLAGS / drop-in changes in a single start, with minimum downtime.

nixos-container update only runs systemctl reload container@<c> when the container is already up (per isContainerRunning in nixos-container.pl). Stopping first turns update into a boot-style operation: it builds + nix-env --sets the new profile and skips the in-container switch-to-configuration. The subsequent start then applies both the new profile and any EXTRA_NSPAWN_FLAGS changes in one go, rather than the double-bounce a live update would trigger.

Sequence for a running container:

  1. prebuild_toplevel — build the new system.build.toplevel before stopping. The container keeps serving the previous generation while eval + fetch + build happen out-of-band. nixos-container update then finds the result cached and skips straight to the profile-swap. Build failures surface here, before the running container is touched.
  2. nixos-container stop — bring the container down.
  3. nixos-container update --flake meta#<name> — profile-swap (near-instant after the prebuild).
  4. nixos-container start — boot into the new generation; the in-container activation script transitions old → new.

If the container is already stopped, step 1 is skipped (no downtime to shave — no point evaluating the flake twice).

Cold-start fallback

start after update can exit non-zero when packages are removed between generations: the old-generation activation script references units that no longer exist in the new closure, causing systemd to exit non-zero. The container may be half-started at that point.

Fallback: stop (graceful SIGTERM drain) → kill (SIGKILL any lingering processes) → start (clean cold-start, no generation transition, new activation runs cleanly). Both errors are preserved and surfaced if the cold-start also fails.

Spawn path (new container)

For a first-time create, nixos-container create is atomic: if the build fails, no container record is left to clean up. A separate prebuild would just duplicate the eval, so it's skipped. Sequence: create --flake meta#<name> → write nspawn flags → systemctl daemon-reloadstart.

Prebuild attr path

nix build does not auto-resolve meta#<name> against nixosConfigurations the way nixos-container does internally. The explicit attr path <flake-root>#nixosConfigurations.<name>.config.system.build.toplevel is required; using the bare meta#<name> ref would make nix look in packages, legacyPackages, or the flake root directly — none of which exist in the rendered meta flake.


See also

  • docs/approvals.md — approval flow + scheduled prompts
  • docs/persistence.md — SQLite schema, state-dir layout
  • docs/conventions.md — wire protocol, recipient sentinels
  • docs/agent-hierarchy.md — topology and parent/child relations