Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/agent-lifecycle/agent-hierarchy.md

206 lines
9.7 KiB
Markdown

# Agent roster & privileges
<!-- vale write-good.Passive = NO -->
Agents are a **flat set**: the capability store scopes which agents can
manage which others.
<!-- vale write-good.Passive = YES -->
This doc covers what the roster file holds, the map-shaped alternative
`topology::all_agents` accepts, and what's hard-coded for the manager.
## Where the roster lives
"The roster" means two different, non-overlapping things depending on
scope. The [README](../../README.md)'s control-plane row ("`swarm-controller`
holds the hive directory, the agent roster and the job graph") is the
**swarm-wide** one: `GET /api/agents` on swarm-controller reads
straight from the swarm's authelia identity store over
`swarm-authelia-bridge` (`swarm-controller/src/auth.rs::list_agent_identities`)
— "every agent the swarm holds an identity for" (`swarm-controller/src/main.rs`'s
`get_agents` doc comment). It's authoritative: an agent exists in the
swarm if and only if it has an identity there, and `swarmctl agent
create` (or the swarm UI, via `POST /api/agents`) writes that identity
once, by calling `ensure_agent_identity`.
This page is about a different, **hive-local** file: the hive-c0re-owned
**meta repo**, alongside `flake.nix`, at
`/var/lib/hyperhive/meta/topology.json`:
```json
["alice", "bob", "ruth"]
```
One entry per agent _this hive_ has state/config for, in name order.
Unlike the swarm roster, it's a derived cache: the reconcile pass below
rebuilds it from what the hive observes locally (config repos cloned, containers spawned), and it
answers a narrower question than "does this agent exist": _which of
this hive's local agents the `ManageRootAgent` capability's bind-mounts
should cover._ `topology::all_agents` is the only reader that matters.
That reader is a permission boundary: the set it returns is what an agent
holding the `ManageRootAgent` capability gets bind-mounted (each other
agent's `state` read-write and `config` read-only; never `harness`). An
agent holding no capability sees its own dirs and nothing else:
`ManageRootAgent` is the only grant that reaches another agent's state
dir. See
[`persistence.md`](persistence.md)'s _Cross-agent access to state._
### Reading the map-shaped format
`topology.json` may also be a map of `name → parent | null`; the reader
accepts that shape and takes its keys as the roster. An empty
roster costs more than a cosmetic gap: every capability holder loses its
mounts until the next reconcile pass writes the array form.
### Why meta, not per-agent `agent.nix`
An agent shouldn't be able to add itself to a set that governs who may
reach its state dir. The roster IS a system-level fact; meta is where
system-level facts live.
### How `topology.json` gets updated
- **Read** — parsed into a set of names; a missing or unparsable file
degrades safely to "no agents" (covers a fresh install that hasn't
synced yet).
- **Reconcile** — runs alongside the periodic meta/flake regeneration.
Adds newly-spawned agents, drops removed ones. Reconcile keeps agents
whose config repo exists but that haven't spawned yet, so the gap until
the container appears doesn't churn the file.
No write API and no operator verb reach this file. Reconcile derives it
from which agents exist, so the next pass overwrites a hand edit.
See `hive-c0re/src/agent_config/topology.rs` and `hive-c0re/src/meta.rs`'s
module docs for the exact call chain.
## Hard-coded manager behaviour
The **manager (`ruth`)** differs from other agents in naming/bootstrap,
its socket flavour and its default capability grant; its prompt, tool
allow-list and state dirs work as for every other agent. A hive with
`services.hyperhive.ruthless = true` runs no manager: `hive-c0re` skips
its root-agent create/start sweep
(`hive-c0re/src/workers/auto_update.rs`), and this section doesn't
apply.
<!-- vale write-good.Passive = NO -->
- **Naming/bootstrap** — the manager's broker recipient name, state-dir
key, and nixos-container name are all `ruth` (container `h-ruth`).
`hive-c0re` spawns it directly at boot if missing, with no operator
approval step — every other agent is created at swarm level
(`swarmctl agent create`).
Roster-wise, `ruth` is just another entry.
- **Wire-protocol** — the `*(privileged)*` `Request` variants in
`hive-core-agent-sock`'s unified enum are the scheduling ops
(`RequestSchedulePrompt`, `CancelSchedule`, `FireScheduleNow`,
`EditSchedule`), reachable only from the manager's socket flavour,
matching the `scheduling` tool group
([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Container lifecycle
ops (kill/start/restart/rebuild) go through the separate host-admin
socket `hivectl` speaks. `Wake` (inject a `from: <X>` message straight
into the caller's own inbox) is on this socket too, unprivileged on
both flavours; matrix, bash and forge don't call it — their
notifications push a todo on the harness's in-agent socket (see
[`docs/turn-loop/mcp.md`](../turn-loop/mcp.md#waking-the-agent-from-inside-the-container)).
- **Storage/mounts** — the manager container gets
`/var/lib/hyperhive/agents` bind-mounted RW at `/agents` (so it can
manage any agent's state dir — config isn't authored there, since a
real config change is a PR from a clone), plus RO mounts for
`/applied` (diff against what's deployed) and `/meta` (system-wide
deploy log). That grant is the `ManageRootAgent` capability: ruth
holds it by default, and any agent holding it gets the same mounts.
The operator dashboard's `POST /api/meta-update` is the only path
that writes `flake.lock`.
- **Prompt** — identical for every container: `prompts/system.md`
carries `<!-- role:agent -->` / `<!-- role:manager -->` marker
blocks, and `prompt::render` filters for `"agent"` unconditionally,
manager included ("always `agent` role — there is only one role",
`hive-agent/src/prompt.rs`'s own module doc). The `role:manager`
blocks are dead in production, exercised only by a unit test
(`filter_role_blocks(SAMPLE, "manager")`).
- **Tool allow-list** — the MCP tools claude sees come from
`HIVE_TOOL_GROUPS` alone, the same mechanism for every agent
([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Ruth's wider default
surface is a wider default grant (`ToolGroup::MANAGER_DEFAULT`, seeded
by `auto_update.rs` whenever its groups aren't already set).
- **State dirs** — _not_ special-cased: `HYPERHIVE_STATE_DIR` is
injected uniformly via `systemd.globalEnvironment` for every
container including the manager, so all token/state paths resolve
through it the same way everywhere.
<!-- vale write-good.Passive = YES -->
None of the above is a stable interface — treat the module doc
comments as the source of truth for exactly which checks exist today.
## Harness systemd unit shape
One harness serve binary (`hive-agent`, with its `hive-agent-mcp`
sibling), one shared `nix/agent-modules/` tree, one service unit
(`systemd.services.hive-agent`) for all agents. No separate manager
service name or role distinction exists in the harness — privilege
differences live server-side in the broker socket (which tool groups
and manager-surface calls each agent receives).
`agent.nix` and `ruth.nix` both import the shared `nix/agent-modules/`.
`ruth.nix` additionally sets forge defaults to suppress the
subscription/participation firehose so ruth's inbox stays focused on
direct mentions, reviews, and assignments.
### Environment variables set on the unit
- `HOME = /home/<userName>` — systemd defaults `HOME` to `/` for
services without `User=` set; with the per-agent user the harness
needs the right home so claude finds its bind-mounted `~/.claude/`
session dir.
- `HIVE_STATIC_DIR = <mergedDist>` — `tower_http::ServeDir` root for
the per-agent web UI; merged dist = agent default + every
`services.hyperhive.agent.frontend.extraFiles` overlay.
- `HIVE_ASSETS_DIR = pkgs.hyperhive-assets/share/hyperhive` — set
directly on the unit, **not** via `environment.variables`, because
the latter only populates `/etc/profile` which systemd services
don't inherit.
### `PATH` setup (the wrapper-dir trick)
```nix
path = [ "/run/wrappers" "/run/current-system/sw" ];
```
<!-- vale write-good.Passive = NO -->
`/run/wrappers` (not `/run/wrappers/bin`) comes first so setuid
wrappers — notably `sudo` — resolve before bare nix-store binaries; see
[`docs/process/gotchas.md`](../process/gotchas.md) ("`systemd.services.*.path` appends
`/bin` to every entry") for why the trailing `/bin` matters in
general. It's load-bearing here because the harness runs as the
per-agent user: without the wrapper dir on `PATH`, `sudo` resolves to
the non-setuid nix-store binary and every
`services.hyperhive.agent.user.passwordlessSudo` grant fails with "must be owned by
uid 0 and have the setuid bit set."
<!-- vale write-good.Passive = YES -->
### `serviceConfig` highlights
- `ExecStart = pkgs.hyperhive/bin/hive-agent` — same binary for every
agent.
- `Restart = on-failure`, `RestartSec = 2` — keeps the harness
resilient across transient crashes without thundering retries.
- `RuntimeDirectory = "hive-config"` → `/run/hive-config/` owned by
`User=`, autocleared on stop. The harness writes regenerated
`claude-{mcp-config,settings,system-prompt}` files there
(`paths::config_dir`). Deliberately separate from `/run/hive`, which
the host bind-mounts in root-owned and which holds hive-c0re's
`mcp.sock`.
- `User = Group = userName` — drops root inside the container; sudo is
the explicit escalation surface (`services.hyperhive.agent.user.passwordlessSudo`).
## Cross-references
- Milestone: "Agent privileges and sub-agents" (tracked internally)
- Audit table source: milestone comment (tracked internally)
- Operator/agent trust boundary (orthogonal axis): [`boundary.md`](../trust-boundary/boundary.md)