Four host glue units fetched a secret from swarm-bao and rendered it with `> path; chmod`: a reader racing the write could see a truncated file, and briefly one at the wrong mode before the chmod landed. glue-matrix-bao-token.nix, glue-queue-agent-credential.nix (both files), swarm-grafana.nix and swarm-otel.nix now write to a same- directory temp file, set its final mode/owner, then `mv -f` it over the target — a shared `atomic_write_secret` helper (nix/host-modules/lib/atomic-write-secret.nix) so the five call sites share one implementation. The first-boot `/root/.claude` migration in nix/agent-modules/user.nix wrote its done-marker unconditionally, so a failed `cp` (disk full, permission error) left the marker behind and no boot ever retried the copy. The marker is now written only when there was nothing to migrate or the copy succeeded; `cp -an`'s no-clobber semantics already make a retry after a partial copy safe. Refs #4723
620 lines
31 KiB
Markdown
620 lines
31 KiB
Markdown
# Persistence + retention
|
|
|
|
Where state lives, what survives what, and how it's bounded.
|
|
|
|
## For operators
|
|
|
|
The short answer to "will I lose anything": **destroying an agent
|
|
keeps its state, purging it doesn't.**
|
|
|
|
- **`DESTR0Y`** (the default action) stops and removes the container
|
|
but keeps everything on disk — config history, claude login, `/state/`
|
|
notes, harness data. The agent shows up as a tombstone (K3PT ST4T3 on
|
|
the C0R3 page) with a `⊕ R3V1V3` button that recreates it from the
|
|
kept state, **no re-login needed**.
|
|
- **`PURG3`** (opt-in, from the dashboard or `hivectl agent <name>
|
|
destroy --purge`) is `DESTR0Y` plus wiping all of it — config
|
|
history, claude credentials, `/state/` notes, everything. **No
|
|
undo.** Only reach for this when you actually want the agent gone
|
|
for good.
|
|
|
|
Beyond that:
|
|
|
|
- **hive-c0re keeps approvals forever** — they're an audit trail, not a
|
|
cache. Nothing about them ever ages out.
|
|
- **Broker messages**: acked ones vacuum after 30 days; anything
|
|
undelivered or delivered-but-not-yet-acked is always kept, however
|
|
old.
|
|
- **An agent's own `/state/` notes and claude login survive every
|
|
restart and rebuild** — only an explicit purge (or a hive
|
|
`--purge`-style host operation, or the agent's own choices) touches
|
|
them.
|
|
- The **root/bootstrap agent is special**: it isn't really destroyable
|
|
in practice — hive-c0re recreates it automatically on its next
|
|
startup if it's ever gone.
|
|
|
|
Everything below this point is implementation detail: exact table
|
|
schemas, file layouts, and internal migration mechanics.
|
|
|
|
## Sqlite databases
|
|
|
|
### `/var/lib/hyperhive/db/broker.sqlite` (host)
|
|
|
|
Seven tables, all in one file — three queues, a small key/value
|
|
table, the schedule header/targets split, and the per-agent
|
|
power-intent registry:
|
|
|
|
- `messages` — every inter-agent / operator-bound message.
|
|
`sender / recipient / body / sent_at / delivered_at / acked_at /
|
|
in_reply_to / priority`. `in_reply_to` links a reply to its parent
|
|
row id; the dashboard and per-agent inbox render these as threaded
|
|
rows.
|
|
- `kv` — small persistent key/value store (`key PK / value`) for
|
|
host-side bookkeeping that doesn't warrant its own table.
|
|
|
|
⚠️ The `mcp__hyperhive__remind` queue is **not** here any more: it
|
|
moved to a harness-local, per-agent store as part of the
|
|
loose-ends-v2 migration — see [`/harness/` contents
|
|
below](#state-dirs-per-agent) for where reminders (and todos)
|
|
actually live now.
|
|
- `approvals` — the queue. `agent / kind (merge_config_pr | spawn |
|
|
update_meta_inputs | schedule_prompt) /
|
|
commit_ref / requested_at / status / resolved_at / note`.
|
|
- `scheduled_prompts` — recurring + one-shot prompt queue.
|
|
`owner / body / interval_seconds (NULL = one-shot) /
|
|
next_fire_at_unix / created_at_unix / source ("operator" or
|
|
"approval:<id>") / cancelled_at_unix / description`. `owner`
|
|
drives cancel-permission checks (operator vs the submitting
|
|
agent). The worker tombstones and reaps cancelled rows
|
|
on its next pass.
|
|
- `scheduled_prompt_targets` — per-target state for each schedule.
|
|
`schedule_id / target / cancelled_at_unix /
|
|
last_fired_at_unix / last_result`. `ON DELETE CASCADE` from
|
|
`scheduled_prompts(id)` — requires `PRAGMA foreign_keys = ON`
|
|
per connection (set at open).
|
|
- `agent_power` — one row per agent: `agent PK / wanted (up |
|
|
offline) / updated_at`, owned by `hive-c0re/src/stores/power.rs`.
|
|
This is the durable power *intent* the job queue reconciles the
|
|
observed container state against; intent survives hive-c0re
|
|
restarts even though in-flight queue work doesn't. See
|
|
[`docs/scheduler/coordinator.md`'s Desired-state
|
|
section](../scheduler/coordinator.md#desired-state-spec-vs-status) for who
|
|
writes and reads it and how reconciliation works.
|
|
|
|
Retention:
|
|
|
|
- `Broker::vacuum_delivered` runs hourly via a tokio task in
|
|
`hive-c0re::main`. Drops acked message rows older than 30 days
|
|
(`acked_at IS NOT NULL`). Undelivered + delivered-but-not-acked
|
|
rows are always kept — the harness `ack_turn`s only after a
|
|
successful turn, so `requeue_inflight` can still requeue
|
|
an unacked row on a crash.
|
|
- hive-c0re keeps approvals indefinitely — an audit trail. `actions::destroy`
|
|
rows stay visible to anything that queries by id.
|
|
- Scheduled prompts: the worker deletes one-shot rows on fire;
|
|
recurring rows live until the operator cancels them
|
|
(`cancel_schedule` MCP / dashboard ✗) which tombstones via
|
|
`cancelled_at_unix`, then `reap_cancelled` drops the row on
|
|
the next worker pass.
|
|
- `agent_power` rows live until the operator destroys the agent (one row per
|
|
agent — nothing to vacuum).
|
|
|
|
### `/harness/hyperhive-events.sqlite` (per agent)
|
|
|
|
Lives inside each container's bind-mounted `/harness/` dir (host
|
|
path: `/var/lib/hyperhive/agents/<name>/harness/hyperhive-events.sqlite`).
|
|
One table:
|
|
|
|
- `events(id, ts, kind, payload_json)` — every `LiveEvent` the
|
|
harness emits during turn loop execution.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
The harness both writes and vacuums it — this used to be a host-side
|
|
sweep, but hive-c0re runs as the unprivileged `hive-core` user under
|
|
privsep and can't delete agent-owned files (host-side deletes hit
|
|
`PermissionDenied` on the bash-task trio and a readonly database error
|
|
here), so cleanup moved in-container. `hive-agent`'s `vacuum::run`
|
|
(`hive-agent/src/vacuum.rs`) sweeps hourly. Retention is
|
|
**type-scoped**: it deletes only the verbose `stream` rows (the raw
|
|
claude `stream-json` deltas — one per text chunk / tool use, the bulk
|
|
of the file's size) older than 14 days, and keeps every other kind
|
|
(`turn_start`, `turn_end`, `note`, `status_changed`, `model_changed`,
|
|
`token_usage_changed`, `turn_state_changed`) indefinitely — those are
|
|
small and carry the semantic per-turn history the operator scrolls
|
|
back through when debugging a regression. Age-only within the
|
|
`stream` kind — no row cap — so a chatty turn doesn't lose its stream
|
|
history sooner than a quiet one. The trade-off (accepted): a
|
|
misbehaving harness could now skip its own cleanup, which the old
|
|
host-side sweep was meant to prevent — but a compromised harness is
|
|
already inside the container trust boundary
|
|
([`docs/trust-boundary/security.md`](../trust-boundary/security.md)), and these are ephemeral local
|
|
artifacts, so cleaning them up where they live is the honest fix.
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
Path overridable via `HYPERHIVE_EVENTS_DB` (for dev / no-`/harness`
|
|
setups). On open failure the `Bus` falls back to no-store mode
|
|
rather than crashing the harness — events still broadcast over SSE,
|
|
just nothing persisted.
|
|
|
|
### `/harness/hyperhive-turn-stats.sqlite` (per agent)
|
|
|
|
Per-turn analytics sink. One row per claude turn captures
|
|
identity (`model`, `wake_from`, `result_kind`), timing
|
|
(`started_at`, `ended_at`, `duration_ms`), cost (input / output /
|
|
cache_read / cache_creation token counts), behaviour
|
|
(`tool_call_count` + `tool_call_breakdown_json`), and post-turn
|
|
snapshot metrics (`open_threads_count`,
|
|
`open_reminders_count` — fetched via the same socket the harness
|
|
already uses for `GetOpenThreads` + `CountPendingReminders`).
|
|
Bin-loop helpers `build_row` + `record` land each row at
|
|
`turn_end`; writes are best-effort, a sqlite hiccup logs + lets
|
|
the turn loop continue.
|
|
|
|
The `hive-bash-daemon` (not the harness) writes a sibling
|
|
`bash_commands(ts INTEGER, head TEXT)` table in the same
|
|
file: one row per executed bash task recording the normalised command head -
|
|
the basename of the first real command, looking past `cd repo &&`
|
|
prefixes, env-assignments, and prefix-runners like `sudo`/`env`. It
|
|
backs the "favorite tools" view on the /stats page (aggregated
|
|
host-side). Best-effort and created on first write
|
|
(`CREATE TABLE IF NOT EXISTS`), so it's absent until a bash
|
|
task runs.
|
|
|
|
turn-stats.sqlite has **no vacuum** — it's one small row per turn
|
|
(~hundreds of KB even over months), read directly by the `/stats` page
|
|
and the hive-wide stats view, so pruning it would only lose trend
|
|
history for no space gain.
|
|
|
|
### `/state/hyperhive-harness.json` (per agent)
|
|
|
|
Consolidated harness state file written atomically (`.tmp` + rename) by
|
|
`Bus::emit_status` whenever rate-limited or login-failed flags change.
|
|
Shape:
|
|
|
|
```json
|
|
{ "rate_limited": false, "needs_login": false, "active_model": "…" }
|
|
```
|
|
|
|
- `rate_limited` — set when the harness detects a 429 from the Claude
|
|
API; cleared by any subsequent status emit. Drives
|
|
`ContainerView.rate_limited` on the dashboard.
|
|
- `needs_login` — set when a turn hits 401 (expired OAuth credentials);
|
|
cleared by `"online"` status (re-auth completed). Drives the
|
|
`needs_login` flag alongside the `claude_has_session` check.
|
|
- `active_model` — the resolved Claude model for the dashboard badge.
|
|
|
|
The turn loop is the only writer today, but it still goes
|
|
read-modify-write under a shared in-process lock and merges into the
|
|
existing object rather than reconstructing it — so a second writer
|
|
would preserve fields it doesn't own, and the lock closes the
|
|
lost-update window between a writer's read and its rename. The lock
|
|
is in-process only, so it wouldn't serialise a writer running as a
|
|
separate process; none of today's writers are.
|
|
|
|
hive-c0re reads this file on each `build_all` sweep (~10s) via
|
|
`container_view::read_harness_flags`. Falls back to the legacy individual
|
|
sentinel files (`hyperhive-rate-limited`, `hyperhive-needs-login`) if the
|
|
JSON is absent, so existing containers keep working through the transition
|
|
window before their next rebuild.
|
|
|
|
### `/var/lib/hyperhive/db/build_logs.sqlite` (host)
|
|
|
|
Full stdout + stderr capture for every `nixos-container` / `nix
|
|
build` invocation the lifecycle layer fires. One row per invocation;
|
|
the row accumulates lines as the child runs.
|
|
|
|
Capturing the full stream (rather than a short tail buffer) matters
|
|
because real eval errors routinely run long — "tried alternatives"
|
|
blocks alone are often 30+ lines — so a truncated tail would cut off
|
|
the actual failure and leave only the host journal holding the
|
|
complete output. With this table the dashboard can surface the entire
|
|
log.
|
|
|
|
Three indices:
|
|
- `(agent, started_at)` — backs the per-agent latest-N lookup used
|
|
by the agent card chip.
|
|
- `(status, finished_at)` — backs the retention sweep that runs
|
|
as part of the existing hourly vacuum.
|
|
- `(node_id)` — added by a later migration so a build log row can be
|
|
looked up by the job-queue node it belongs to (a `hive_jobq` node is
|
|
immutable after insert, so hive-c0re records the link on the log row
|
|
instead); legacy rows predating the column keep `node_id IS NULL`.
|
|
|
|
Writes are best-effort: `append_stdout` / `append_stderr` / `finish`
|
|
log a warning on sqlite error and let the build continue. A failed
|
|
log row never blocks a rebuild.
|
|
|
|
### `/harness/hyperhive-model` (per agent)
|
|
|
|
Single-line text file holding the claude model name currently
|
|
selected for this agent (default `haiku` when absent). Written by
|
|
`Bus::set_model` whenever the operator flips it via `/model
|
|
<name>` in the web terminal. Read once at harness boot in
|
|
`Bus::new`. Path overridable via `HYPERHIVE_MODEL_FILE`.
|
|
Survives destroy/recreate, gone on `--purge`.
|
|
|
|
### `/harness/paused` (per agent)
|
|
|
|
Empty marker file. Its presence parks the agent's turn loop: the
|
|
harness keeps serving its web UI and MCP daemons but drives no turns,
|
|
and inbox messages queue unacked until it's removed (see
|
|
[turn loop](../turn-loop/README.md#the-loop)).
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
Unusually, it's read and written from **both** sides of the harness
|
|
bind-mount, and that's the whole design: the harness stats it
|
|
in-container via `hive-agent`'s `paths::paused_marker`, while hive-c0re
|
|
stats it on the host (`Coordinator::is_paused`) to populate the
|
|
`paused` field on the agent card, and creates/removes it
|
|
(`Coordinator::set_paused`) for `hivectl agent <name> pause|resume` and the
|
|
dashboard toggle. Because the file itself is the only shared state
|
|
there's no protocol between them, no round-trip into the container, and
|
|
pause keeps working when the harness is wedged or the container is
|
|
stopped.
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
It lives in `/harness/` rather than `/state/` deliberately: `/state/`
|
|
is the agent's own space to fill, and this is harness control state.
|
|
Survives destroy/recreate, gone on `--purge` — so a paused agent comes
|
|
back paused after a restart, which is the intended behaviour rather
|
|
than an accident of storage.
|
|
|
|
## State dirs (per agent)
|
|
|
|
Under `/var/lib/hyperhive/agents/<name>/`:
|
|
|
|
- `config/` — the proposed nix repo (root-agent-editable). Bind-mounted
|
|
**read-only** to `/agents/<name>/config` inside the sub-agent's own
|
|
container so the agent can inspect what defines it and request
|
|
precise changes from the root agent; RW into the root agent via the
|
|
`/agents` tree bind.
|
|
- `claude/` — claude OAuth credentials, bind-mounted RW to
|
|
`/home/<name>/.claude` inside the container.
|
|
- `state/` — durable notes and `hyperhive-harness.json`. Bind-mounted
|
|
to `/agents/<name>/state` inside the container (uniform for
|
|
all agents). The `$HYPERHIVE_STATE_DIR` env var exposes
|
|
the same path to in-container scripts. Notable files written here
|
|
by the harness:
|
|
- `hyperhive-status` — single-line free-text status string written
|
|
by `set_status`; cleared on explicit `set_status("")`. Read by
|
|
hive-c0re and the per-agent `/api/dashboard-state` endpoint to
|
|
surface the status chip on the dashboard. Absent when no status
|
|
has a value.
|
|
- `hyperhive-harness.json` — rate-limited / needs-login flags read
|
|
by the dashboard's async container-state fetch. See
|
|
`docs/web-ui/dashboard.md::Container row`.
|
|
- `harness/` — harness-internal ephemeral state; not intended for
|
|
agent consumption. Bind-mounted to `/agents/<name>/harness`
|
|
inside the container (`$HYPERHIVE_HARNESS_DIR`). Contents:
|
|
- `bash-tasks/` — task JSON + stdout/stderr files for
|
|
background `mcp__bash__run` jobs. JSON files are
|
|
`<id>.json` (status + tails), `<id>.out` / `<id>.err`
|
|
(full captured output). The harness's own hourly sweep
|
|
(`hive-agent`'s `vacuum::run`, same one that ages out `stream`
|
|
event rows above) deletes terminal task trios older than 48
|
|
hours; non-terminal (still-running) tasks are never deleted. This
|
|
used to be a host-side `hive-c0re` vacuum, moved in-container for
|
|
the same privsep-ownership reason as the events vacuum above.
|
|
- `hyperhive-state.sqlite` — consolidated loose-ends-v2 store: todos
|
|
and reminders, one small table each in a single file (in-container
|
|
daemons — `hive-bash-daemon`, `hive-matrix-daemon`,
|
|
`hive-forge-notify` — upsert keyed todos here over the harness's
|
|
in-agent socket, `HIVE_AGENT_SOCKET`; the harness merges them into
|
|
`get_loose_ends` output and clears a row on `mark_todo_done`).
|
|
Replaces three formerly separate files
|
|
(`hyperhive-todos.sqlite`, `hyperhive-reminders.sqlite`, and the
|
|
old file-based `mcp-loose-ends/` scanner before that) — a one-time
|
|
boot migration (`db_migrate::run`) folds the legacy files into this
|
|
path the first time a harness boots after the upgrade. Also backs
|
|
the `mcp__hyperhive__remind` queue, which moved from a host-side
|
|
`broker.sqlite` table to this per-agent store as part of the same
|
|
migration.
|
|
|
|
The harness itself is also a producer, not just the socket server:
|
|
boot wiring's `spawn_todo_socket` starts `todo_server::run` (the
|
|
socket the out-of-process daemons above dial) alongside
|
|
`disk_watch::run` — an *in-process* todo producer that shares the
|
|
store + wake `Notify` directly rather than dialling its own socket.
|
|
`disk_watch` raises a keyed `disk` todo when the filesystem backing
|
|
this agent's state gets tight, naming the agent's own biggest
|
|
directories; it buckets the summary and carries no raw byte
|
|
counts, so an unchanged situation re-upserts as `changed == false`
|
|
and never re-wakes.
|
|
|
|
Retention, same hourly `hive-agent::vacuum::run` sweep as
|
|
`hyperhive-events.sqlite` below: it reaps delivered (soft-deleted) reminder
|
|
rows 14 days after delivery, kept that long only to serve
|
|
the trailing-window `ReminderRollup` stats, and reaps acked todo rows
|
|
30 days after acking (long enough that only a genuinely quiet
|
|
month triggers the "one spurious re-announcement" fallback a
|
|
reconciled producer like `disk_watch` relies on — see `todos.rs`'s
|
|
module doc). Un-acked todos and undelivered reminders are never
|
|
swept — same "audit trail, not cache" treatment as the c0re-side
|
|
tables above.
|
|
- `hyperhive-events.sqlite` — turn-loop event log.
|
|
- `hyperhive-turn-stats.sqlite` — per-turn timing stats.
|
|
- `hyperhive-model` — single-line model name override file.
|
|
|
|
### Cross-agent access to state
|
|
|
|
An agent holding the `ManageRootAgent` capability gets every other
|
|
agent's `state` dir bind-mounted **read-write** and its `config` dir
|
|
**read-only** (`bind_child_agent_dirs` in `lifecycle/host_config.rs`).
|
|
The RW on `state` is deliberate, not an oversight: the holder recovers
|
|
other agents, which includes writing into their state (for example
|
|
seeding notes, clearing a stuck sentinel) as well as reading it.
|
|
|
|
This is the **only** cross-agent mount. Dropping the topology parent
|
|
field took with it the unconditional grant every agent used to
|
|
get over its own direct children — an agent holding no capability now
|
|
sees its own dirs and nothing else.
|
|
|
|
**`harness` isn't mounted at all.** It holds that agent's own runtime
|
|
material — `bash-tasks/`, the turn-stats and event sqlite dbs — and
|
|
nothing argues for anyone else reading it, let alone writing it.
|
|
hive-c0re reads a harness dir **directly on the host** when it wants
|
|
those stats, which needs no mount into another container.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
**`config` is read-only, including for the holder.** A config change is
|
|
a PR on that agent's config repo, made from a clone and merged after
|
|
review — so the bind-mounted `config` dir is a *copy to read*, never a
|
|
tree anyone edits in place. Mounting it writable would leave a second
|
|
path to the same file that skips the review entirely, which makes the
|
|
boundary a convention rather than a permission.
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
⚠️ Don't confuse it with the config-repo seeding hive-c0re does at
|
|
spawn (`lifecycle::setup_proposed`): that writes the agent's initial
|
|
config repo as **hive-c0re, against the host path**, and `read_only` on a bind
|
|
constrains writers *inside* a container only. The two are unrelated —
|
|
conflating them can lead you to reason your way into thinking this
|
|
mount should be writable when it shouldn't.
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
Isolation still holds by default: a container has its *own* dirs and,
|
|
unless it holds the capability, nothing else.
|
|
|
|
Under `/var/lib/hyperhive/applied/<name>/` — the hive-c0re-only
|
|
applied repo. Tracks `flake.nix` (module-only boilerplate; never
|
|
edited after first spawn) + `agent.nix` (the actual config; the
|
|
root agent's edits land here via the approval flow) + any other
|
|
files committed via the approval flow. `.git/` carries the proposal /
|
|
approved / building / deployed / failed / denied tag history.
|
|
|
|
Under `/var/lib/hyperhive/meta/` — the swarm-wide deploy flake plus
|
|
system-level config files. Single git repo for the whole host; hive-c0re
|
|
commits every mutation that should survive a restart here.
|
|
Contents:
|
|
|
|
- `flake.nix` — declares one `nixpkgs` input per agent + one
|
|
`nixosConfigurations.<n>` output per agent. `flake.lock` is the
|
|
canonical "what's deployed where." The git log is the deploy
|
|
audit trail (one commit per successful deploy or hyperhive bump).
|
|
- `topology.json` — the agent roster (`["alice", "bob", "ruth"]`).
|
|
Written by `topology::reconcile` on every meta sync; read by
|
|
`topology::all_agents`, which is the set the `ManageRootAgent`
|
|
capability grants mounts over. Carried a `parent` per agent in the
|
|
legacy format; the reader still accepts that shape and keeps its keys.
|
|
- `tool-groups.json` — per-agent MCP tool group grants
|
|
(`{ "alice": ["messaging", "inbox", "execution"] }`). Written by
|
|
`tool_groups::set_groups`; injected as `HIVE_TOOL_GROUPS` env
|
|
var into each agent's container.
|
|
- `capabilities.json` — per-agent capability grants
|
|
(`{ "ruth": ["manage_root_agent"] }`). Written by
|
|
`capabilities::set_caps`; injected as `HIVE_CAPABILITIES` env
|
|
var. Absent agents have no extra capabilities.
|
|
- `resource-limits.json` — per-agent container resource overrides
|
|
(`{ "sock": { "cpu_quota": "400%", "memory_max": "8G" } }`).
|
|
Written by `resource_limits::set_limits`; read where
|
|
`lifecycle::write_dropins` generates the systemd drop-in, **not** injected
|
|
into the container — these are host-side caps on the container, so
|
|
the capped party never sees or sets them. Fallback is per *field*:
|
|
an absent file, absent agent, or absent field falls back to the
|
|
hive-wide `services.hyperhive.c0re.agentCpuQuota` / `agentMemoryMax`,
|
|
so an agent can override only its memory and still track the hive
|
|
default for CPU. The `CPUWeight=` / `IOWeight=` shares in the same
|
|
drop-in have **no** per-agent override — they're hive-wide only and
|
|
come straight off `HiveEnv`, so this file has no field for them.
|
|
|
|
The root agent has the meta dir RO-mounted at `/meta/`.
|
|
|
|
The `.meta-migration-done` marker no longer exists: the
|
|
one-shot container repoint it guarded no longer exists either, since
|
|
hive-c0re renders containers onto `meta#<n>` at creation. A stale
|
|
marker file left over from an older hive is inert and the operator can
|
|
delete it.
|
|
|
|
## Destroy vs purge
|
|
|
|
See [For operators](#for-operators) above for what each action does to
|
|
an agent's state. The mechanics, for completeness:
|
|
|
|
- `DESTR0Y` also drops the systemd drop-in and fails any pending
|
|
approvals; the tombstone's `⊕ R3V1V3` button queues a Spawn approval
|
|
that reuses the kept state on approve.
|
|
- `PURG3` wipes `/var/lib/hyperhive/{agents,applied}/<name>/` — the
|
|
union of everything `DESTR0Y` left behind.
|
|
|
|
`actions::destroy` implements the root/bootstrap agent's specialness as a soft policy
|
|
guard that refuses to destroy it, backstopped by
|
|
`auto_update::ensure_root_agent`, which recreates it on the next
|
|
hive-c0re startup if it's ever absent (bypassing the approval queue,
|
|
as required infrastructure) — so even without the guard, destroying it
|
|
would only be transient.
|
|
|
|
### btrfs subvolumes for `/var/lib/hyperhive/agents/<name>`
|
|
|
|
On a btrfs host, `lifecycle::ensure_agent_state_subvolume` creates a brand-new agent's state root as a
|
|
**btrfs subvolume** instead of a plain directory (progressive
|
|
enhancement). This is a no-op fallback on
|
|
non-btrfs hosts and for any agent whose root already exists, so
|
|
hive-c0re automigrates nothing: existing agents keep their plain dirs
|
|
until an explicit opt-in upgrade.
|
|
|
|
- **Creation:** `lifecycle::ensure_agent_state_subvolume` runs before
|
|
hive-c0re creates the per-agent subdirs (spawn / rebuild).
|
|
It skips the work when the root already exists; otherwise it asks
|
|
hive-priv (`EnsureAgentSubvolume`) to `btrfs subvolume create` the
|
|
root when the FS is btrfs (`statfs` magic gate) and chown it to the
|
|
`hive-core` user so the normal `state/` `claude/` `harness/` mkdirs
|
|
succeed inside it.
|
|
- **DESTR0Y keeps the subvolume** exactly like a plain dir — revival
|
|
reuses it untouched.
|
|
- **PURG3 deletes it correctly:** `rmdir`/`remove_dir_all` can't remove
|
|
a subvolume root, so purge first calls hive-priv
|
|
(`DeleteAgentSubvolume`) which `btrfs subvolume delete`s it iff it's
|
|
actually a subvolume, then the normal `remove_dir_all` sweep covers
|
|
plain-dir agents + the applied dir.
|
|
|
|
Per-subvolume disk-usage accounting and optional quotas have since
|
|
landed as the qgroup work: `hivectl quota-enable` turns on btrfs
|
|
qgroup accounting hive-wide (opt-in, no-op on non-btrfs hosts), and
|
|
`hivectl agent <name> quota show|set` reads/limits one agent's
|
|
subvolume usage through the same hive-priv-mediated path as
|
|
subvolume creation/deletion above.
|
|
|
|
This is the same subvolume `hivectl agent <name> subvol snapshot push`
|
|
sends to the swarm's snapshot store — see
|
|
[`docs/networking/snapshot-store.md`](../networking/snapshot-store.md) for what a pushed
|
|
snapshot contains and how the store authenticates a sender.
|
|
|
|
## `/var/lib/swarm-controller/` (swarm-controller host only)
|
|
|
|
Only present on the one host running
|
|
`services.hyperhive.deploy.swarm-controller.enable`. systemd `StateDirectory=`,
|
|
so it survives restarts and redeploys.
|
|
|
|
- `webhook-secret` — the HMAC key Forgejo signs the swarm's forge webhooks
|
|
with. **Keep it.** It's handed to Forgejo when swarm-controller registers a hook,
|
|
so replacing the file means every subsequent delivery fails
|
|
verification until the hook is re-registered with the new value. It's
|
|
generated automatically on first start; there's nothing to configure.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
If the file is unreadable at startup the daemon still starts and logs
|
|
`webhook secret unavailable`; the webhook endpoint then answers 503
|
|
rather than accepting deliveries it can't verify. Everything else the
|
|
controller serves is unaffected.
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
## Run-time dirs
|
|
|
|
`/run/hyperhive/` is tmpfs-backed (systemd `RuntimeDirectory=`) but
|
|
preserved across hive-c0re restarts via `RuntimeDirectoryPreserve=yes`.
|
|
Without that, every restart wipes bind sources and existing
|
|
containers can't start.
|
|
|
|
- `/run/hyperhive/host.sock` — admin socket (host-side CLI).
|
|
- `/run/hyperhive/agents/<name>/mcp.sock` — per-agent socket
|
|
(bind-mounted into the container as `/run/hive/mcp.sock`).
|
|
|
|
On startup, `Coordinator::register_agent` drops any prior socket
|
|
task before rebinding — idempotent so a hive-c0re restart followed
|
|
by `rebuild alice` recreates the agent's socket without a clean
|
|
reinstall.
|
|
|
|
## First-boot agent-user migration
|
|
|
|
The harness runs as a per-agent unix user inside the container
|
|
(`services.hyperhive.agent.user.name`, defaults to the agent's logical label so each
|
|
container has a uniquely named user). Operators with legacy root-owned
|
|
state dirs need a one-time data shuffle so they don't lose their claude
|
|
session.
|
|
|
|
`system.activationScripts.hive-agent-user-migrate` (in
|
|
`nix/agent-modules/user.nix`) runs on every activation,
|
|
marker-guarded so the substantive moves only happen once per
|
|
container lifetime:
|
|
|
|
1. **`${homeDir}` exists with the right ownership** — covers the
|
|
the first boot before `useradd`'s `createHome` has had a
|
|
chance to chown. Also re-applies on every rebuild in case the
|
|
meta-flake's per-agent name evolves (rare).
|
|
2. **Migrate any leftover `/root/.claude` content into
|
|
`${homeDir}/.claude`** — legacy `claude` wrote to root's
|
|
empty home; the bind mount didn't exist yet. Marker
|
|
(`/var/lib/hive-agent-user-migrated`) guards single-shot; it is
|
|
only written once there is nothing left to migrate, so a `cp`
|
|
failure leaves it absent and the next boot retries.
|
|
`cp -an` (no-clobber) so any pre-existing files at the new
|
|
location win — never blow over data already there, and so a
|
|
retry after a partial copy is as safe as the first attempt.
|
|
3. **Chown the bind-mounted state dir** (`/agents/*/state`)
|
|
recursively so the agent user can read/write it. Wildcard
|
|
matches the single agent that container sees; `-h` skips
|
|
symlinks the agent might have planted.
|
|
4. **Chown the `~/.claude/` bind-mount** recursively. Legacy
|
|
`claude` wrote `.credentials.json` 0600 root:root; the
|
|
current harness reads `~/.claude/` as the agent user to decide
|
|
Online vs NeedsLogin in `login::has_session`. Without the
|
|
chown the existing credentials get silently treated as "no
|
|
session" and the operator re-prompts every boot.
|
|
|
|
The activation script will eventually become unnecessary once no
|
|
operators have legacy root-owned state dirs left to migrate; drop
|
|
the body + marker check at that point.
|
|
|
|
## Matrix per-agent daemon + token-arrival trigger
|
|
|
|
`hive-matrix-daemon` is a long-running matrix-sdk Client + sync
|
|
process per agent. Serves its MCP tools directly over
|
|
streamable-http (`services.hyperhive.agent.mcp.matrixHttpPort`, no stdio bridge —
|
|
same shape as `hive-bash-daemon`), emits hyperhive wake signals
|
|
on incoming room events via `/run/hive/mcp.sock`. Conditional on the
|
|
agent having at least one `services.hyperhive.agent.matrixAccounts`
|
|
entry — there is no separate enable switch, and the same condition gates
|
|
the daemon, its path watcher AND the autoinjected
|
|
`extraMcpServers.matrix` entry. The module declares the hive-internal
|
|
`main` entry whenever `services.hyperhive.agent.matrix.url` is non-null,
|
|
so on a real hive (where hive-c0re renders that URL per agent) every
|
|
agent has one; a `null` URL with no operator-declared account is the
|
|
"this agent has no matrix" state.
|
|
|
|
**First-boot ordering**: a token can arrive after the container comes
|
|
up — a file the hive delivers, or a token the swarm mints into the store.
|
|
Without the path-trigger sibling
|
|
(`systemd.paths.hive-matrix-daemon`, `PathExistsGlob =
|
|
<this agent's state dir>/matrix-token*` — the trailing `*` also catches
|
|
a secondary multi-account token like `matrix-token-ccc`), the daemon
|
|
would exit 0 quietly the first time it ran and the MCP would have no
|
|
daemon until the next restart. The `.path` unit makes the appearance
|
|
of the token re-fire the service so the daemon comes alive in the
|
|
same boot cycle as
|
|
provisioning. A token in the store changes no file, so a store-backed agent
|
|
also gets a timer that restarts the daemon five minutes after it last
|
|
exited. The same token watcher also drives avatar setting: on a
|
|
restart the daemon re-runs each account's bring-up, which sets the
|
|
avatar (see below).
|
|
|
|
### matrix avatar (set by the daemon over the live Client)
|
|
|
|
`hive-matrix-daemon` itself publishes the agent icon (`services.hyperhive.agent.icon`, an SVG) as each matrix
|
|
account's profile avatar
|
|
(`hive-matrix-mcp::client::sync_avatar`), not a separate oneshot. After
|
|
the daemon builds + restores an account's `Client` (authenticated,
|
|
pointed at that account's resolved homeserver), it calls matrix-sdk's
|
|
`account().upload_avatar()` — one call that uploads the media and sets
|
|
`avatar_url`. Because it reuses the live Client, there is no hardcoded
|
|
homeserver URL, no token re-read, and no token-file globbing: the daemon
|
|
already iterates every configured + dashboard-discovered account in its
|
|
bring-up loop, so it sets the avatar for **every** account.
|
|
|
|
Nix rasterizes the SVG to a 512x512 PNG at build time (`iconPng`, via
|
|
librsvg) and forwards its store path as `HIVE_ICON_PNG` on the daemon
|
|
unit, gated on `services.hyperhive.agent.icon != null`. No icon configured → the env is
|
|
unset → `sync_avatar` returns early and sets no avatar.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
Idempotency is **per-account**: an `avatar-icon-hash` file in each
|
|
account's matrix-sdk `state_dir`. The daemon hashes the PNG bytes and
|
|
skips the upload when unchanged, because every upload mints a fresh
|
|
`mxc://` URI that emits a profile state event in every joined room —
|
|
re-uploading identical bytes is timeline spam. A dashboard-provisioned
|
|
account gets its avatar when the `systemd.paths.hive-matrix-daemon` token
|
|
watcher restarts the daemon (which re-runs the per-account bring-up), so
|
|
no separate avatar trigger is needed. The daemon swallows avatar failures
|
|
(logged, non-fatal) so they never break account bring-up or sync.
|
|
<!-- vale write-good.Passive = YES -->
|
|
|