Watch
0
0
Fork
You've already forked hyperhive
0

docs(agents): facts pass on agent-hierarchy.md and mcp.md

mcp.md:
- matrix and subagent extra MCP servers are http (hive-matrix-daemon,
  hive-subagent-daemon), not stdio; screen is the one entry that still
  uses the stdio default (nix/agent-modules/matrix.nix:290-297,
  screen.nix:22-25)
- set_status is always-on, not meta-group-gated; mark_todos_done (also
  always-on) was undocumented (hive-sh4re/src/permissions.rs:151,
  hive-agent-mcp/src/mcp/mod.rs:425)
- System messages: HelperEvent has 3 variants, not the 8 previously
  listed; ApprovalResolved/ContainerCrash routing and the swarm-wide
  NATS notices stream (swarm_notices.rs) replace the old per-agent
  todo-wake description for rebuilt/killed/destroyed/logged_in/needs_login
- get_loose_ends's approval rows are manager-only; PendingMessages and
  UnreadMatrix were missing from the description
  (hive-sh4re/src/inbox.rs:127-184)
- subagent spawning runs on hive-runtime (claude or ACP), not
  claude-only (hive-subagent-mcp/src/session.rs:77)
- Waking section: matrix/bash/forge all moved to the in-agent todo
  socket; the host Wake request has no built-in caller left today

agent-hierarchy.md:
- distinguished the swarm-wide agent roster (swarm-controller's
  identity store, authoritative) from the hive-local topology.json
  (a derived, reconciled cache scoping ManageRootAgent's bind-mounts),
  linking README's framing
- noted services.hyperhive.ruthless (a hive can run with no manager at
  all)
- Wire-protocol bullet: the only privileged Request variants left are
  the scheduling ops; Kill/Start/Restart/Update/GetLogs don't exist on
  this socket
- Prompt/tools: prompt::render hardcodes the agent role for every
  container today (role:manager blocks are dead code); the tool
  allow-list has no Flavor switch, it's HIVE_TOOL_GROUPS same as any
  agent

Not touched: agent-hierarchy.md:140-200 (Harness systemd unit shape,
kept in place — see PR follow-ups) and docs/agent-lifecycle/approvals.md
(blocked on #4853).
This commit is contained in:
atlas 2026-10-02 00:14:17 +02:00 • committed by mara
commit 4e31550dad
2 changed files with 129 additions and 71 deletions

View file

@ -15,17 +15,33 @@ cleanup.
## Where the roster lives
The roster lives in the hive-c0re-owned **meta repo**, alongside
`flake.nix`, at `/var/lib/hyperhive/meta/topology.json`:
"The roster" means two different, non-overlapping things depending on
scope. The [README](../../README.md)'s control-plane row ("`swarm-controller`
holds the hive directory, the agent roster and the job graph") is the
**swarm-wide** one: `GET /api/agents` on swarm-controller reads
straight from the swarm's authelia identity store over
`swarm-authelia-bridge` (`swarm-controller/src/auth.rs::list_agent_identities`)
— "every agent the swarm holds an identity for" (`swarm-controller/src/main.rs`'s
`get_agents` doc comment). It's authoritative: an agent exists in the
swarm if and only if it has an identity there, and `swarmctl agent
create` (or the swarm UI, via `POST /api/agents`) writes that identity
once, by calling `ensure_agent_identity`.
This page is about a different, **hive-local** file: the hive-c0re-owned
**meta repo**, alongside `flake.nix`, at
`/var/lib/hyperhive/meta/topology.json`:
```json
["alice", "bob", "ruth"]
```
One entry per agent the hive knows about, in name order. The file
carries no per-agent value any more, and encodes no ordering or
grouping — it answers exactly one question, _which agents exist,_ and
`topology::all_agents` is the only reader that matters.
One entry per agent _this hive_ currently has state/config for, in name
order. Unlike the swarm roster, nothing writes this file directly —
it's a derived cache, rebuilt by the reconcile pass below from what the
hive observes locally (config repos cloned, containers spawned), and it
answers a narrower question than "does this agent exist": _which of
this hive's local agents the `ManageRootAgent` capability's bind-mounts
should cover._ `topology::all_agents` is the only reader that matters.
That reader is a permission boundary: the set it returns is what an agent
holding the `ManageRootAgent` capability gets bind-mounted (each other
@ -93,7 +109,10 @@ consequence of removing the field, not a side effect of it.
Capability enforcement isn't fully wired yet, so the
**manager (`ruth`) still gets some hard-coded special treatment**
other agents don't:
other agents don't. A hive can opt out of having one at all
(`services.hyperhive.ruthless = true` skips `hive-c0re`'s root-agent
create/start sweep entirely, `hive-c0re/src/workers/auto_update.rs`);
everything below applies only when it doesn't:
<!-- vale write-good.Passive = NO -->
@ -103,14 +122,20 @@ other agents don't:
approval step — every other agent is created at swarm level
(`swarmctl agent create`).
Roster-wise, `ruth` is just another entry.
- **Wire-protocol** — the privileged `Request` variants
(`Kill` / `Start` / `Restart` / `Update`; `GetLogs`) — marked
`*(privileged)*` in `hive-core-agent-sock`'s unified `Request` enum —
are reachable only from the manager's socket flavour today; each is
planned to become a capability check. One exception: `Wake` (inject a `from: <X>` message into the
caller's own inbox) isn't really privileged — every per-agent daemon
(for example `hive-forge-notify`) needs it, and sub-agents already have the
equivalent on their own socket.
- **Wire-protocol** — the only `*(privileged)*` `Request` variants left
in `hive-core-agent-sock`'s unified enum are the scheduling ops
(`RequestSchedulePrompt`, `CancelSchedule`, `FireScheduleNow`,
`EditSchedule`), reachable only from the manager's socket flavour
today, matching the `scheduling` tool group
([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Container lifecycle
ops (kill/start/restart/rebuild) never lived on this socket — they go
through the separate host-admin socket `hivectl` speaks. `Wake`
(inject a `from: <X>` message straight into the caller's own inbox)
is on this socket too but isn't privileged to either flavour; no
built-in in-container producer calls it today — matrix, bash and
forge notifications all moved to pushing a todo on the harness's
in-agent socket instead (see
[`docs/turn-loop/mcp.md`](../turn-loop/mcp.md#waking-the-agent-from-inside-the-container)).
- **Storage/mounts** — only the manager container gets
`/var/lib/hyperhive/agents` bind-mounted RW at `/agents` (so it can
manage any agent's state dir — config isn't authored there, since a
@ -122,12 +147,20 @@ other agents don't:
path writes `flake.lock` any more — `request_update_meta_inputs` was
removed, leaving the operator dashboard's `POST
/api/meta-update` as the only entry point.
- **Prompt/tools** — the system prompt uses `<!-- role:agent -->` /
`<!-- role:manager -->` marker blocks, and a `Flavor::{Agent,
Manager}` switch picks the MCP tool allow-list claude sees. Both are
already parametrised on a single flavour value, so the planned
per-capability-group version (`cap:<group>` prompt blocks + a
matching tool allow-list) is additive rather than a rewrite.
- **Prompt** — _not_ treated differently in practice any more: `prompts/system.md`
still carries `<!-- role:agent -->` / `<!-- role:manager -->` marker
blocks, but `prompt::render` filters for `"agent"` unconditionally for
every container, manager included ("always `agent` role — there is
only one role", `hive-agent/src/prompt.rs`'s own module doc). The
`role:manager` blocks are dead in production, exercised only by a
unit test (`filter_role_blocks(SAMPLE, "manager")`).
- **Tool allow-list** — also not a flavour switch: the MCP tools claude
sees come from `HIVE_TOOL_GROUPS` alone, the same mechanism for every
agent ([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Ruth's wider
default surface is just a wider default grant
(`ToolGroup::MANAGER_DEFAULT`, seeded by `auto_update.rs` whenever its
groups aren't already set), not anything keyed off its name or
container.
- **State dirs** — _not_ special-cased: `HYPERHIVE_STATE_DIR` is
injected uniformly via `systemd.globalEnvironment` for every
container including the manager, so all token/state paths resolve