From 66442d41ffaa9ae569ea2010c6b3e8e5cf328a73 Mon Sep 17 00:00:00 2001 From: iris Date: Sat, 15 Aug 2026 11:47:26 +0200 Subject: [PATCH] docs(agent-hierarchy): restructure audit doc into current/planned, trim internals --- docs/agent-hierarchy.md | 318 +++++++++++++++++----------------------- 1 file changed, 137 insertions(+), 181 deletions(-) diff --git a/docs/agent-hierarchy.md b/docs/agent-hierarchy.md index d3c11ff4..7c24a665 100644 --- a/docs/agent-hierarchy.md +++ b/docs/agent-hierarchy.md @@ -1,11 +1,13 @@ # Agent hierarchy & privileges -Design + audit doc for the agent-privileges + tree-shape milestone -(the [issue tree](http://localhost:3000/hyperhive/hyperhive/issues/361)). -The implementation lands in pieces; this doc tracks what's done, what's -planned, and what currently special-cases the manager. +Every agent has a place in an operator-editable parent/child tree, used +to scope which agents can manage which others. This doc covers how the +tree is stored and edited today, the rules that are meant to run on top +of it once enforcement is finished, and where the manager still gets +special-cased in the meantime. Tracking issue: +[hyperhive#361](http://localhost:3000/hyperhive/hyperhive/issues/361). -## Current state (as of this PR) +## Where the tree lives Topology lives in the hive-c0re-owned **meta repo**, alongside `flake.nix`, at `/var/lib/hyperhive/meta/topology.json`: @@ -18,45 +20,35 @@ Topology lives in the hive-c0re-owned **meta repo**, alongside } ``` -`null` = root-level agent. New agents **default to root** (`null` parent) — -there is no structural manager that everything hangs under. Hierarchy is -built explicitly: an agent that requests a sub-agent gets a -requester-as-parent edge written at its `init_config` approval (so `bob` -above was spawned by `alice`), and the operator can reparent any agent. The -bootstrap container (`ruth`) is just another root. Re-parenting is -operator-driven: +`null` = root-level agent. New agents **default to root** — there is no +structural manager that everything hangs under. Hierarchy is built +explicitly: an agent that requests a sub-agent gets a +requester-as-parent edge written at its `init_config` approval (so +`bob` above was spawned by `alice`), and the operator can reparent any +agent, including the bootstrap container (`ruth`) — it's just another +root. The manager is reparentable like any other agent; there's no +"structurally root" carve-out. Its privileges live on its MCP socket, +not its tree position (see *Manager special-casing today* below). -- CLI: `hivectl agent set-parent --parent ` (or `--root` to - promote). Exactly one of `--parent` / `--root` is required. +### Reparenting + +- CLI: `hivectl agent set-parent --parent ` (or `--root` + to promote). Exactly one of `--parent` / `--root` is required. - Dashboard: `POST /api/topology/set-parent` (form fields `child`, optional `new_parent` — absent / empty ⇒ promote to root). - Wire: `HostRequest::SetParent { child, new_parent: Option }`. -All three converge on `topology::set_parent`, which delegates the -validation rules to a pure `apply_set_parent` helper. Refuses: +All three go through the same validation, which refuses: - unknown `child` / `new_parent` (typo guard), - self-parenting, -- cycles (32-hop ancestor walk, mirroring `is_descendant_of`). +- cycles (a bounded ancestor walk — moving the manager under one of + its own descendants is the only real safety concern here, and it's + caught the same way as any other agent). -The manager is reparentable like any other agent — there's no -"structurally root" carve-out; the manager's privileges live on its -MCP socket, not its tree position, and the cycle walk above catches -the only real safety concern (moving the manager under one of its -own descendants). - -Idempotent no-op fast path skips the disk write when the parent is -already what's requested. After a successful write the surfaces call -`Coordinator::rescan_containers_and_emit` so connected dashboard -viewers see the tree repaint without polling -(`ContainerView.parent` is sourced from `topology.json`). - -**Today's caveat:** the move is purely a JSON edit. Only the -top-level manager (`root`) gets `/var/lib/hyperhive/agents` -bind-mounted at `/agents` in its container, so sub-agents don't yet -see their would-be children's state. Once sub-manager bind mounts -land alongside cap enforcement, `set_parent` grows a companion -umount-old / mount-new / restart-cascade step. +Setting a parent to its current value is a no-op (no disk write). A +successful change triggers an immediate rescan, so connected dashboard +viewers see the tree repaint without polling. ### Why meta, not per-agent `agent.nix` @@ -65,37 +57,39 @@ consent, and operator-driven re-parenting shouldn't require touching the moved agent's config. Topology IS a system-level concern; meta is where system-level facts live. -### Flow +### How `topology.json` gets updated -1. **Read**: `topology::read()` parses `topology.json` into a - `BTreeMap>`. Missing / unparsable file → - empty map → every agent treated as root (safe degradation for - fresh installs that haven't run `meta::sync_agents` yet). -2. **Reconcile**: `meta::sync_agents` calls `topology::reconcile` - alongside its `flake.nix` regeneration. New agents default to root - (null parent) — an agent-requested sub-agent already carries an - explicit requester-as-parent edge from its `init_config` approval, so - only user/operator-initiated spawns hit this default, and those are - roots; removed agents drop. Existing entries are preserved as-is so - operator overrides stick across regenerations. Pending-init agents - (an operator-approved proposed config repo but no container yet — - `Coordinator::pending_init_names`) are kept too, so the - `child -> parent` edge written when `request_init_config` is approved - survives the gap until the first apply-commit spawns the container. -3. **Inject**: `meta::render_flake` looks up each agent's parent and - passes it to `mkAgent`. When non-null, the mkAgent body sets - `HIVE_PARENT = parent` in the agent's systemd service environment - so the harness / claude prompts can see it. -4. **Surface**: `container_view::build_all` reads `topology.json` and - populates `ContainerView.parent: Option` on every rescan. - The dashboard renders the field as a tree. +- **Read** — parsed into an agent→parent map; a missing or unparsable + file degrades safely to "every agent is root" (covers a fresh + install that hasn't synced yet). +- **Reconcile** — runs alongside the periodic meta/flake regeneration. + New agents default to root unless they already carry an explicit + parent edge from an `init_config` approval; existing entries + (including operator overrides) are preserved; removed agents drop. + Agents that are approved but not yet spawned keep their edge too, so + it survives the gap until the container actually appears. +- **Inject** — each container's parent (if any) is exposed to its own + environment as `HIVE_PARENT`, so the harness / system-prompt + renderer can see it. +- **Surface** — every rescan re-reads `topology.json` and populates + `ContainerView.parent`, which the dashboard renders as a tree. -## Target topology semantics +See `hive-c0re/src/topology.rs` and `hive-c0re/src/meta.rs`'s module +docs for the exact call chain. -Once enforcement lands the rules collapse into: +### Current limitation: state-dir visibility lags topology + +Reparenting today is purely a JSON edit. Only the top-level manager +(`root`) gets `/var/lib/hyperhive/agents` bind-mounted at `/agents` in +its container, so sub-agents don't yet see their would-be children's +state dirs. Once sub-manager bind mounts land alongside capability +enforcement, reparenting will grow a companion +umount-old / mount-new / restart-cascade step. + +## Planned topology semantics (once ancestor-based enforcement lands) | operation | who can do it | -| ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | +| ----------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | | `kill` / `start` / `restart` / `update` (any descendant) | any ancestor | | `request_init_config` (spawn a new child) | any agent, child added under self | | config change via forge PR (any descendant's config) | any ancestor | @@ -105,106 +99,71 @@ Once enforcement lands the rules collapse into: | `request_update_meta_inputs` (bump meta lock) | root agents only (today: just `manager`) | "Ancestor" walks `ContainerView.parent` chains; cycles are guarded by a -visited-set at dispatch time (a malformed topology.json can't lock the -dispatcher into a loop). +visited-set at dispatch time (a malformed `topology.json` can't lock +the dispatcher into a loop). -## Current manager special-casings — the audit +## Manager special-casing today -What currently makes the manager different from every other agent, and -which axis the post-milestone version reads each special-case along: +Enforcement of the ancestor rules above isn't fully wired yet, so the +**manager (`ruth`) still gets some hard-coded special treatment** +other agents don't: -### A — naming + bootstrap +- **Naming/bootstrap** — the manager's broker recipient name, state-dir + key, and nixos-container name are all `ruth` (container `h-ruth`). + `hive-c0re` spawns it directly at boot if missing, with no operator + approval step — every other agent goes through `request_init_config` + → approval. Topology-wise, `ruth` is still just another root agent. +- **Wire-protocol** — the `ManagerRequest::*` operations + (`RequestInitConfig`; `Kill` / `Start` / `Restart` / `Update`; + `GetLogs`; `RequestUpdateMetaInputs`) are reachable only from the + manager's socket flavour today. Planned rule for each is in the + table above ("any agent, child added under self" for init-config, + "any ancestor" for lifecycle/logs); `RequestUpdateMetaInputs` stays + a root-only capability even post-milestone, not a topology rule. + One exception: `Wake` (inject a `from: ` message into the + caller's own inbox) isn't really privileged — every per-agent daemon + (e.g. `hive-forge-notify`) needs it, and sub-agents already have the + equivalent on their own socket. +- **Storage/mounts** — only the manager container gets + `/var/lib/hyperhive/agents` bind-mounted RW at `/agents` (so it can + manage any agent's state dir — config is not authored there, since a + real config change is a PR from a clone), plus RO mounts for + `/applied` (diff against what's deployed) and `/meta` (system-wide + deploy log). Planned: each agent gets RW to `/agents//` + for just its own subtree — the manager's full-forest RW becomes the + "root's subtree is everything" case of that same rule. RO `/meta` + access will be gated on a "meta read" capability; only + `request_update_meta_inputs` writes `flake.lock`, gated by its own + capability. +- **Prompt/tools** — the system prompt uses `` / + `` marker blocks, and a `Flavor::{Agent, + Manager}` switch picks the MCP tool allow-list claude sees. Both are + already parametrised on a single flavour value, so the planned + per-capability-group version (`cap:` prompt blocks + a + matching tool allow-list) is additive rather than a rewrite. +- **State dirs** — *not* special-cased: `HYPERHIVE_STATE_DIR` is + injected uniformly via `systemd.globalEnvironment` for every + container including the manager, so all token/state paths resolve + through it the same way everywhere. +- **Scattered ownership checks** — a handful of independent + manager-only overrides exist across `hive-c0re` today: loose-ends + visibility (manager sees hive-wide, sub-agents only their own), + "manager can cancel any question/reminder" overrides on the owner + check, `destroy` refusing to act on the manager, and crash-watch + skipping the manager (it auto-restarts via systemd instead of going + through the crash-watch loop). Each is planned to become an + ancestor/descendant check instead of a manager-name check — see the + module docs for `loose_ends.rs`, `operator_questions.rs`, + `broker.rs`, `reminder_scheduler.rs`, `actions.rs`, and + `crash_watch.rs` for the current owner-check logic in each. -- `MANAGER_AGENT = "ruth"` (broker recipient name), - `MANAGER_NAME = "ruth"` (logical name, state-dir key), and - `MANAGER_CONTAINER = "h-ruth"` (nixos-container name); the `h-` - prefix lets `lifecycle::list()` use a single `starts_with("h-")` - filter. -- `auto_update::ensure_manager` runs at hive-c0re boot and spawns - `h-ruth` if missing. **Topology**: ruth defaults to root-level (no - parent); hive-c0re handles the bootstrap lifecycle directly. +None of the above is a stable interface — treat the module doc +comments as the source of truth for exactly which checks exist today. -### B — wire-protocol privileges +## Future work: sub-agents inside the same container -The `ManagerRequest::*` variants in `hive-sh4re/src/lib.rs` are -operations the manager flavour socket can make that sub-agent sockets -can't: - -| variant | semantic | post-milestone | -| --------------------------------------- | ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `RequestInitConfig` | seed an agent's proposed config repo | **topology** — existing direct child (re-init) or a brand-new name (child added under self on approval); a name owned by a different parent is refused | -| `Kill` / `Start` / `Restart` / `Update` | container lifecycle on an existing agent | **topology** — descendants only | -| `RequestUpdateMetaInputs` | bump meta `flake.lock` | **per-agent cap** (root-only today; a future "let coder bump its own input" might grant it) | -| `GetLogs` | journalctl scrape of a sub-agent | **topology** — descendants only | -| `Wake` | inject a `from: ` message into self's inbox | **not really privileged** — the wire surface exists because the per-agent daemons (e.g. `hive-forge-notify`) need it. Sub-agents have the same via their own socket. | - -### C — storage / mounts (`hive-c0re::lifecycle`) - -The manager container's nspawn bind set: - -- `HOST_AGENTS_ROOT (/var/lib/hyperhive/agents) → /agents` RW — so the - manager can manage any agent's state dir. Config is **not** authored - here: a config change is a PR from a clone, and `/config/` is - a copy for reading (its write access is a defect tracked separately) -- `HOST_APPLIED_ROOT (/var/lib/hyperhive/applied) → /applied` RO — so - the manager can diff against what's deployed -- `HOST_META_ROOT (/var/lib/hyperhive/meta) → /meta` RO — so the - manager can read the system-wide deploy log - -Tree-shape version: - -- Each agent gets RW to `/agents//` for every descendant in - its subtree. The root agent (today: manager) gets RW to the full - forest as a special case of "the root has every other agent as a - descendant". -- RO `/meta` access if the agent holds a "meta read" cap. -- `request_update_meta_inputs` is the only path that actually writes - `flake.lock`, gated by the cap; everyone else stays RO. - -### D — drop legacy `/state` for manager ✓ done - -`lifecycle.rs` no longer binds `/state` for the manager. -`HYPERHIVE_STATE_DIR` is now injected uniformly via -`systemd.globalEnvironment` in `meta.rs` for every container -(manager included), so all token/state paths resolve through -`$HYPERHIVE_STATE_DIR`. The agent-module shell scripts -(tea-login, forge-avatar-sync) simplified from glob+for loops to a -direct `$HYPERHIVE_STATE_DIR/` read. - -### E — prompt + tools - -- `prompts/system.md` with `` / `` - marker blocks, assembled by `hive_ag3nt::prompt::render` based on - flavor. **Per-agent cap list** of what the agent can do — already - a single parametrised prompt; once per-agent cap groups land the - marker grammar grows `cap:` blocks the renderer reads from - the per-agent ToolGroup set. -- `mcp.rs::Flavor::{Agent, Manager}` controls which MCP tools claude - sees. Already structured this way internally — the per-flavour - allow-list becomes a per-cap-set lookup. - -### F — drive-by checks across c0re - -- `loose_ends.rs`: manager sees hive-wide loose-ends, sub-agents only - their own. **Topology** — every agent sees its own + its - descendants'. -- `operator_questions.rs` + `broker.rs`: "manager can cancel any - question" override on the owner check. **Topology** — agents can - moderate threads of their descendants. -- `reminder_scheduler.rs`: same override pattern for reminder cancel. - **Topology** — descendants only. -- `actions.rs`: `destroy` refuses to act on `MANAGER_NAME` (no - foot-shooting). **Topology** — agents can destroy descendants but - never themselves or ancestors. -- `crash_watch.rs`: skips `ContainerCrash` for the manager (it - auto-restarts via systemd). **Topology** — the root container has - different recovery semantics, every other agent falls into the same - watch loop. - -### G — sub-agents inside the same container - -Future work: when enabled for an agent, it can spawn temporary -"sub-agents" that run inside its own container. Lighter than a full +When enabled for an agent, it will be able to spawn temporary +"sub-agents" that run inside its own container — lighter than a full nspawn agent. Open questions, not yet wired: - Inherit caps from parent, or take an explicit narrower set? @@ -216,17 +175,16 @@ nspawn agent. Open questions, not yet wired: ## Harness systemd unit shape One harness serve binary (`hive-agent`, with its `hive-agent-mcp` -sibling), one shared `nix/agent-modules/` tree, one -service unit (`systemd.services.hive-agent`) for all agents. There -is no longer a separate manager service name or role distinction in -the harness — privilege differences live server-side in the broker -socket (which tool groups and manager-surface calls each agent -receives). +sibling), one shared `nix/agent-modules/` tree, one service unit +(`systemd.services.hive-agent`) for all agents. There is no separate +manager service name or role distinction in the harness — privilege +differences live server-side in the broker socket (which tool groups +and manager-surface calls each agent receives). `agent.nix` and `ruth.nix` both import the shared `nix/agent-modules/`. `ruth.nix` additionally sets forge defaults to suppress the -subscription/participation firehose so ruth's inbox stays focused -on direct mentions, reviews, and assignments. +subscription/participation firehose so ruth's inbox stays focused on +direct mentions, reviews, and assignments. ### Environment variables set on the unit @@ -248,32 +206,30 @@ on direct mentions, reviews, and assignments. path = [ "/run/wrappers" "/run/current-system/sw" ]; ``` -`/run/wrappers` comes first so setuid wrappers (notably `sudo`) -resolve before bare nix-store binaries. NixOS's -`systemd.services..path` appends `/bin` to every entry via -`lib.makeBinPath`; passing `/run/wrappers/bin` directly produces -`/run/wrappers/bin/bin` which doesn't exist (`docs/gotchas.md:: -systemd.services.*.path appends /bin to every entry`). With the -harness running as the per-agent user this matters: without the -wrapper dir on PATH, `sudo` resolves to the un-setuid nix-store -binary and rejects with `must be owned by uid 0 and have the setuid -bit set` regardless of `hyperhive.user.passwordlessSudo`. +`/run/wrappers` (not `/run/wrappers/bin`) comes first so setuid +wrappers — notably `sudo` — resolve before bare nix-store binaries; see +[`docs/gotchas.md`](gotchas.md) ("`systemd.services.*.path` appends +`/bin` to every entry") for why the trailing `/bin` matters in +general. It's load-bearing here because the harness runs as the +per-agent user: without the wrapper dir on `PATH`, `sudo` resolves to +the non-setuid nix-store binary and every +`hyperhive.user.passwordlessSudo` grant fails with "must be owned by +uid 0 and have the setuid bit set." ### `serviceConfig` highlights -- `ExecStart = pkgs.hyperhive/bin/hive-agent` — same binary for - every agent. +- `ExecStart = pkgs.hyperhive/bin/hive-agent` — same binary for every + agent. - `Restart = on-failure`, `RestartSec = 2` — keeps the harness resilient across transient crashes without thundering retries. - `RuntimeDirectory = "hive-config"` → `/run/hive-config/` owned by `User=`, auto-cleared on stop. The harness writes regenerated `claude-{mcp-config,settings,system-prompt}` files there - (`paths::config_dir`). Deliberately separate from `/run/hive`, - which the host bind-mounts in root-owned and which holds - hive-c0re's `mcp.sock`. -- `User = Group = userName` — drops root inside the container; sudo - is the explicit escalation surface - (`hyperhive.user.passwordlessSudo`). + (`paths::config_dir`). Deliberately separate from `/run/hive`, which + the host bind-mounts in root-owned and which holds hive-c0re's + `mcp.sock`. +- `User = Group = userName` — drops root inside the container; sudo is + the explicit escalation surface (`hyperhive.user.passwordlessSudo`). ## Cross-references