diff --git a/docs/agent-lifecycle/agent-hierarchy.md b/docs/agent-lifecycle/agent-hierarchy.md index 8b555de9..953b04b4 100644 --- a/docs/agent-lifecycle/agent-hierarchy.md +++ b/docs/agent-lifecycle/agent-hierarchy.md @@ -2,15 +2,13 @@ -Agents are a **flat set**, with no parent/child tree: `topology.json` -carries no `parent` field, and the capability store — not tree -position — scopes which agents can manage which others. +Agents are a **flat set**: the capability store scopes which agents can +manage which others. -This doc covers what the roster file holds today, the map-shaped -alternative `topology::all_agents` still accepts, and where the manager -still gets special-cased, as a tracked cleanup. +This doc covers what the roster file holds, the map-shaped alternative +`topology::all_agents` accepts, and what's hard-coded for the manager. ## Where the roster lives @@ -34,10 +32,9 @@ This page is about a different, **hive-local** file: the hive-c0re-owned ["alice", "bob", "ruth"] ``` -One entry per agent _this hive_ currently has state/config for, in name -order. Unlike the swarm roster, nothing writes this file directly — -it's a derived cache, rebuilt by the reconcile pass below from what the -hive observes locally (config repos cloned, containers spawned), and it +One entry per agent _this hive_ has state/config for, in name order. +Unlike the swarm roster, it's a derived cache: the reconcile pass below +rebuilds it from what the hive observes locally (config repos cloned, containers spawned), and it answers a narrower question than "does this agent exist": _which of this hive's local agents the `ManageRootAgent` capability's bind-mounts should cover._ `topology::all_agents` is the only reader that matters. @@ -45,14 +42,15 @@ should cover._ `topology::all_agents` is the only reader that matters. That reader is a permission boundary: the set it returns is what an agent holding the `ManageRootAgent` capability gets bind-mounted (each other agent's `state` read-write and `config` read-only; never `harness`). An -agent holding no capability sees its own dirs and nothing else. See +agent holding no capability sees its own dirs and nothing else: +`ManageRootAgent` is the only grant that reaches another agent's state +dir. See [`persistence.md`](persistence.md)'s _Cross-agent access to state._ ### Reading the map-shaped format `topology.json` may also be a map of `name → parent | null`; the reader -accepts that shape too and keeps its keys, so a hive with a file in that -shape still reads the same roster rather than an empty one. An empty +accepts that shape and takes its keys as the roster. An empty roster costs more than a cosmetic gap: every capability holder loses its mounts until the next reconcile pass writes the array form. @@ -78,40 +76,15 @@ from which agents exist, so the next pass overwrites a hand edit. See `hive-c0re/src/agent_config/topology.rs` and `hive-c0re/src/meta.rs`'s module docs for the exact call chain. -## What replaced the parent field +## Hard-coded manager behaviour -Recorded so a reader who finds one of these in an old branch, an issue -thread or a stale comment knows each one went away rather than moved: - -| gone | what replaced it | -| -------------------------------------------- | --------------------------------------------------- | -| `` recipient sentinel | address `operator` directly | -| `` fan-out recipient | nothing — name the recipients, or broadcast to `*` | -| `hivectl agent set-parent` | nothing | -| `POST /api/topology/set-parent{,-bulk}` | nothing | -| `HostRequest::SetParent` | nothing | -| `NodeKind::Reparent` and its DAG template | nothing | -| `HIVE_PARENT` on the container | nothing — no consumer ever read it | -| every agent's grant over its direct children | the `ManageRootAgent` capability, for every agent | -| rebuild ordering by topology depth | alphabetical, which the depth sort already produced | - - - -The last row is the one with teeth: reaching a child's state dir -requires holding `ManageRootAgent`; parentage grants nothing. That -narrowing is the intended consequence of removing the field, not a -side effect of it. - - - -## Manager special-casing today - -Capability enforcement isn't fully wired yet, so the -**manager (`ruth`) still gets some hard-coded special treatment** -other agents don't. A hive can opt out of having one at all -(`services.hyperhive.ruthless = true` skips `hive-c0re`'s root-agent -create/start sweep entirely, `hive-c0re/src/workers/auto_update.rs`); -everything below applies only when it doesn't: +The **manager (`ruth`)** differs from other agents in naming/bootstrap, +its socket flavour and its default capability grant; its prompt, tool +allow-list and state dirs work as for every other agent. A hive with +`services.hyperhive.ruthless = true` runs no manager: `hive-c0re` skips +its root-agent create/start sweep +(`hive-c0re/src/workers/auto_update.rs`), and this section doesn't +apply. @@ -121,44 +94,39 @@ everything below applies only when it doesn't: approval step — every other agent is created at swarm level (`swarmctl agent create`). Roster-wise, `ruth` is just another entry. -- **Wire-protocol** — the only `*(privileged)*` `Request` variants left - in `hive-core-agent-sock`'s unified enum are the scheduling ops +- **Wire-protocol** — the `*(privileged)*` `Request` variants in + `hive-core-agent-sock`'s unified enum are the scheduling ops (`RequestSchedulePrompt`, `CancelSchedule`, `FireScheduleNow`, - `EditSchedule`), reachable only from the manager's socket flavour - today, matching the `scheduling` tool group + `EditSchedule`), reachable only from the manager's socket flavour, + matching the `scheduling` tool group ([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Container lifecycle - ops (kill/start/restart/rebuild) never lived on this socket — they go - through the separate host-admin socket `hivectl` speaks. `Wake` - (inject a `from: ` message straight into the caller's own inbox) - is on this socket too but isn't privileged to either flavour; no - built-in in-container producer calls it today — matrix, bash and - forge notifications push a todo on the harness's - in-agent socket instead (see + ops (kill/start/restart/rebuild) go through the separate host-admin + socket `hivectl` speaks. `Wake` (inject a `from: ` message straight + into the caller's own inbox) is on this socket too, unprivileged on + both flavours; matrix, bash and forge don't call it — their + notifications push a todo on the harness's in-agent socket (see [`docs/turn-loop/mcp.md`](../turn-loop/mcp.md#waking-the-agent-from-inside-the-container)). -- **Storage/mounts** — only the manager container gets +- **Storage/mounts** — the manager container gets `/var/lib/hyperhive/agents` bind-mounted RW at `/agents` (so it can manage any agent's state dir — config isn't authored there, since a real config change is a PR from a clone), plus RO mounts for `/applied` (diff against what's deployed) and `/meta` (system-wide - deploy log). That grant is the `ManageRootAgent` capability, which - ruth holds; nothing checks the agent's name. hive-c0re will - gate RO `/meta` access on a "meta read" capability; no agent-facing - path writes `flake.lock` — the operator dashboard's `POST -/api/meta-update` is the only entry point. -- **Prompt** — treated identically: `prompt::render` always filters for `agent`. `prompts/system.md` - still carries `` / `` marker - blocks, but `prompt::render` filters for `"agent"` unconditionally for - every container, manager included ("always `agent` role — there is - only one role", `hive-agent/src/prompt.rs`'s own module doc). The - `role:manager` blocks are dead in production, exercised only by a - unit test (`filter_role_blocks(SAMPLE, "manager")`). -- **Tool allow-list** — also not a flavour switch: the MCP tools claude - sees come from `HIVE_TOOL_GROUPS` alone, the same mechanism for every - agent ([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Ruth's wider - default surface is just a wider default grant - (`ToolGroup::MANAGER_DEFAULT`, seeded by `auto_update.rs` whenever its - groups aren't already set), not anything keyed off its name or - container. + deploy log). That grant is the `ManageRootAgent` capability: ruth + holds it by default, and any agent holding it gets the same mounts. + The operator dashboard's `POST /api/meta-update` is the only path + that writes `flake.lock`. +- **Prompt** — identical for every container: `prompts/system.md` + carries `` / `` marker + blocks, and `prompt::render` filters for `"agent"` unconditionally, + manager included ("always `agent` role — there is only one role", + `hive-agent/src/prompt.rs`'s own module doc). The `role:manager` + blocks are dead in production, exercised only by a unit test + (`filter_role_blocks(SAMPLE, "manager")`). +- **Tool allow-list** — the MCP tools claude sees come from + `HIVE_TOOL_GROUPS` alone, the same mechanism for every agent + ([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Ruth's wider default + surface is a wider default grant (`ToolGroup::MANAGER_DEFAULT`, seeded + by `auto_update.rs` whenever its groups aren't already set). - **State dirs** — _not_ special-cased: `HYPERHIVE_STATE_DIR` is injected uniformly via `systemd.globalEnvironment` for every container including the manager, so all token/state paths resolve diff --git a/docs/process/conventions.md b/docs/process/conventions.md index 6fcf1e1a..872d12a6 100644 --- a/docs/process/conventions.md +++ b/docs/process/conventions.md @@ -44,7 +44,7 @@ its leftmost label by convention, but the convention isn't machine-readable, and federated hives at different DNS domains can share a swarm name. Humans want both: the address (`@darkest.space`) AND the prose name (`pr1ma`). Matrix MXIDs -still use the domain-based convention untouched. +use the domain-based convention. `qualify(label)` is the same shape as `qualified_label()` but applies to an arbitrary label the caller already has (for example a peer @@ -65,11 +65,10 @@ wake-event-injection surface on the host-served per-agent socket. Recipient is implicit — the agent the socket belongs to — and `from` is caller-chosen so the wake prompt can label the source verbatim (`"matrix: new message in #general"`, `"forge: PR #42 opened"`, etc.). -`hive-c0re` still handles it, but no built-in -in-container producer calls it today — matrix, bash and forge -notifications all upsert a todo on the harness's in-agent socket -(`HIVE_AGENT_SOCKET`) instead, which signals the same turn loop without -a hive-c0re round-trip. See +`hive-c0re` handles it; matrix, bash and forge don't call it — their +notifications upsert a todo on the harness's in-agent socket +(`HIVE_AGENT_SOCKET`), which signals the same turn loop without a +hive-c0re round-trip. See [`docs/turn-loop/mcp.md` § Waking the agent from inside the container](../turn-loop/mcp.md#waking-the-agent-from-inside-the-container). @@ -91,11 +90,7 @@ angle-bracket and asterisk shapes below are structurally safe. - `operator` — the human at the dashboard. Messages accumulate in the inbox view; no agent ever `recv`'s them. -`` and `` aren't valid recipients: address -`operator` directly instead of ``, and name the recipients (or -broadcast to `*`) instead of ``. Nothing rewrites a -recipient at send time — what an agent passes is what the broker -stores. +The broker stores the recipient exactly as the agent passed it. ## Wire protocol @@ -237,10 +232,8 @@ status_text, status_set_at, hive_name, swarm_name, matrix_accounts }`: `false`, the host clears `status_text` / `status_set_at` — on-disk values from before the stop are stale snapshots, not live status. Defaults to `true` on the - wire: the deserializer treats a payload lacking the field — from a - harness that never serialises it, or a host that only knows how to - ask about live containers — as running, keeping compatibility with - pre-running-field payloads. + wire: the deserializer treats a payload without the field as + running. - `status_text` / `status_set_at`: last value written via `SetStatus`, plus its unix timestamp. Both `None` when the target has never set a status, when the agent name is unknown, @@ -302,10 +295,10 @@ binary flavor. | `meta` | `get_agent_meta` (`set_status` is always-on, see below) | | `inbox` | `get_loose_ends`, `cancel_loose_end`, `remind` | | `execution` | vestigial — `mcp__bash__run` / `mcp__bash__status` are always available unconditionally via `extraMcpServers`; this group's entries expand to non-existent `mcp__hyperhive__run` / `mcp__hyperhive__status` and have no effect. See `docs/tools/bash.md`. | -| `lifecycle` | none — `list_containers` isn't a tool; the variant survives only so existing grants parse. | -| `approvals` | none — `request_update_meta_inputs` isn't a tool. Still a live server-side gate: `cancel_loose_end`'s approval-cancel arm requires it. | +| `lifecycle` | none | +| `approvals` | none — gates `cancel_loose_end`'s approval-cancel arm server-side. | | `scheduling` | `request_schedule_prompt`, `fire_schedule_now`, `cancel_schedule`, `edit_schedule`, `list_schedules` *(privileged)* | -| `forge` | none — `create_repo` isn't a tool; the variant survives only so existing grants parse. | +| `forge` | none | | `web_tools` | none (gates the Claude built-ins `WebFetch`/`WebSearch`, not an MCP tool) | **Always-on tools** — `ToolGroup::ALWAYS_ON_TOOLS` exposes `set_status`, @@ -411,8 +404,7 @@ rewrite — `PRIVATE_NETWORK=1`, `HOST_ADDRESS` = the bridge gateway IP, sets `EXTRA_NSPAWN_FLAGS` — plus the systemd resource-limits drop-in) into the `Swap` node, then runs `nixos-container update` + stop + start across the `StopForUpdate → Swap → RebuildBookkeeping` -brace and the tail `Reconcile` node. `flake.nix` itself isn't -regenerated host-side on rebuild — it's tracked in the agent's +brace and the tail `Reconcile` node. `flake.nix` itself lives in the agent's proposed/applied repos and rides along on every fetch (see `docs/agent-lifecycle/approvals.md::Two repos per agent`). diff --git a/docs/turn-loop/mcp.md b/docs/turn-loop/mcp.md index 1f42d807..cf3b3139 100644 --- a/docs/turn-loop/mcp.md +++ b/docs/turn-loop/mcp.md @@ -14,7 +14,7 @@ entry (`type = "stdio" | "http"`, default `"stdio"`): `matrix` and (`hive-matrix-daemon`, `hive-subagent-daemon`), same reasoning as the built-in surface and `bash`'s `hive-bash-daemon`. The GUI bridge (`screen`, gated on `services.hyperhive.agent.gui.enable`) is the one -entry that still takes the schema's stdio default — it spawns +entry on the schema's stdio default — it spawns `hive-screen-mcp` fresh each turn. The server name is `hyperhive`, so the tools land in claude as `mcp__hyperhive__`. Each entry also has an `availableToSubagents` @@ -62,8 +62,8 @@ are opt-in via the P3RM1SS10NS tab. **System messages** (from sender `system`): the broker delivers `hive_sh4re::manager::HelperEvent` (three variants: `ApprovalResolved`, -`ContainerCrash`, `NeedsUpdate` — the last declared but constructed by -no call site today) as regular inbox messages (same `recv` +`ContainerCrash`, `NeedsUpdate` — declared, but no call site +constructs it) as regular inbox messages (same `recv` path; body is a JSON object with an `event` discriminant field). `ApprovalResolved` goes to whichever agent actually submitted the approval (`Coordinator::notify_submitter`, looked up from the @@ -73,7 +73,7 @@ recorded submitter falls back to the manager, `ruth`). hardcoded — `hive-c0re/src/workers/crash_watch.rs`). A `MergeConfigPr` approval's rebuild additionally pushes a `rebuilt:` todo to that same submitter (`Coordinator::push_todo_submitter`, `subsystem = -"core"`), which still wakes a turn (the todo-wake path — see [Turn +"core"`), which wakes a turn (the todo-wake path — see [Turn outcomes](README.md#turn-outcomes)) via a generic "call `get_loose_ends`" prompt rather than the event body itself. Lifecycle transitions the job-queue scheduler or crash watcher drive directly — @@ -91,8 +91,7 @@ at_unix_timestamp?)`. bash, forge, or any user-configured MCP server), an undelivered-inbox count, and unread matrix notifications; `approval` rows only appear when the caller is the manager (`ruth`) — sub-agents don't submit - approvals. Always the caller's own rows — there is no way to target - another agent. Full per-variant wire shape in + approvals. Always the caller's own rows. Full per-variant wire shape in [`docs/process/conventions.md` § Loose-ends wire shape](../process/conventions.md#loose-ends-wire-shape). - `cancel_loose_end` — hard-delete a `reminder`, cancel a pending @@ -138,8 +137,7 @@ hive_name?, swarm_name?, matrix_accounts? }`. `matrix_accounts` is a broker — see `hive-agent/src/todo_server.rs`. - `mark_todos_done(ids)` — bulk-clear specific loose-ends-v2 todo rows by id in one call, instead of `cancel_loose_end`ing each one. - List-based, not range-based (no `ack_until`-style "clear below id - N"): pass the ids `get_loose_ends` actually showed you reviewed; + List-based: pass the ids `get_loose_ends` actually showed you reviewed; unknown or already-done ids are silently skipped. Dials the same in-agent todo socket as `cancel_loose_end`'s `todo` kind. @@ -149,16 +147,16 @@ hive_name?, swarm_name?, matrix_accounts? }`. `matrix_accounts` is a [`docs/tools/bash.md`](../tools/bash.md). - **Subagent spawning** — nested sessions on the agent's own runtime (claude or ACP, via `hive-runtime`) as background tasks, shipped - default-on like bash execution (no tool group gates it yet). See + default-on like bash execution; no tool group gates it. See [`docs/tools/subagent.md`](../tools/subagent.md). -- **Lifecycle + config** (`lifecycle`, `approvals`) — neither group - carries an MCP tool. `approvals` survives as a server-side gate on - `cancel_loose_end`'s approval-cancel arm; `lifecycle` gates nothing. +- **Lifecycle + config** (`lifecycle`, `approvals`) — `approvals` + gates `cancel_loose_end`'s approval-cancel arm server-side; + `lifecycle` gates nothing. Config changes go through a forge PR on `agent-configs/` — see [`docs/agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md). - **Scheduling** (`scheduling`) — scheduled prompts. See [`docs/tools/scheduling.md`](../tools/scheduling.md). -- **Forge repos** (`forge`) — carries no MCP tool; `forge` gates nothing. +- **Forge repos** (`forge`) — gates nothing. - **Web egress** (`web_tools`) — enables Claude's built-in `WebFetch` and `WebSearch` tools (not MCP tools; added directly to the `--allowedTools` list). Off by default; add the group in the @@ -177,20 +175,17 @@ The built-in producers (matrix, bash, forge) upsert a todo on the harness's in-agent socket (`HIVE_AGENT_SOCKET`, `UpsertTodo` — see [`docs/tools/bash.md` § Completion as a todo (loose-ends v2)](../tools/bash.md#completion-as-a-todo-loose-ends-v2) -for the full upsert/signal/clear mechanism); none dials a direct -wake: a new or changed summary +for the full upsert/signal/clear mechanism): a new or changed summary makes the harness signal its own turn loop, with no hive-c0re round-trip. An external MCP server (or any other in-container process) -can push its own todo the same way — `subsystem` is a plain string, -not a closed set. +can push its own todo the same way — `subsystem` is any string. -The host-served per-agent socket at `/run/hive/mcp.sock` still carries -a lower-level `Wake` request (`hive_core_agent_sock::Request::Wake { +The host-served per-agent socket at `/run/hive/mcp.sock` carries a +lower-level `Wake` request (`hive_core_agent_sock::Request::Wake { from, body }` — JSON-line, `{"cmd":"wake","from":"matrix","body":"new dm from @alice"}\n`) that drops the body straight into the broker inbox as `{from: