From 4e31550dad209c73fb35ff110caee5e076e8fa2c Mon Sep 17 00:00:00 2001 From: atlas Date: Fri, 2 Oct 2026 00:14:17 +0200 Subject: [PATCH] docs(agents): facts pass on agent-hierarchy.md and mcp.md MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit mcp.md: - matrix and subagent extra MCP servers are http (hive-matrix-daemon, hive-subagent-daemon), not stdio; screen is the one entry that still uses the stdio default (nix/agent-modules/matrix.nix:290-297, screen.nix:22-25) - set_status is always-on, not meta-group-gated; mark_todos_done (also always-on) was undocumented (hive-sh4re/src/permissions.rs:151, hive-agent-mcp/src/mcp/mod.rs:425) - System messages: HelperEvent has 3 variants, not the 8 previously listed; ApprovalResolved/ContainerCrash routing and the swarm-wide NATS notices stream (swarm_notices.rs) replace the old per-agent todo-wake description for rebuilt/killed/destroyed/logged_in/needs_login - get_loose_ends's approval rows are manager-only; PendingMessages and UnreadMatrix were missing from the description (hive-sh4re/src/inbox.rs:127-184) - subagent spawning runs on hive-runtime (claude or ACP), not claude-only (hive-subagent-mcp/src/session.rs:77) - Waking section: matrix/bash/forge all moved to the in-agent todo socket; the host Wake request has no built-in caller left today agent-hierarchy.md: - distinguished the swarm-wide agent roster (swarm-controller's identity store, authoritative) from the hive-local topology.json (a derived, reconciled cache scoping ManageRootAgent's bind-mounts), linking README's framing - noted services.hyperhive.ruthless (a hive can run with no manager at all) - Wire-protocol bullet: the only privileged Request variants left are the scheduling ops; Kill/Start/Restart/Update/GetLogs don't exist on this socket - Prompt/tools: prompt::render hardcodes the agent role for every container today (role:manager blocks are dead code); the tool allow-list has no Flavor switch, it's HIVE_TOOL_GROUPS same as any agent Not touched: agent-hierarchy.md:140-200 (Harness systemd unit shape, kept in place — see PR follow-ups) and docs/agent-lifecycle/approvals.md (blocked on #4853). --- docs/agent-lifecycle/agent-hierarchy.md | 75 +++++++++++---- docs/turn-loop/mcp.md | 123 ++++++++++++++---------- 2 files changed, 128 insertions(+), 70 deletions(-) diff --git a/docs/agent-lifecycle/agent-hierarchy.md b/docs/agent-lifecycle/agent-hierarchy.md index 1e95f436..573cb1f9 100644 --- a/docs/agent-lifecycle/agent-hierarchy.md +++ b/docs/agent-lifecycle/agent-hierarchy.md @@ -15,17 +15,33 @@ cleanup. ## Where the roster lives -The roster lives in the hive-c0re-owned **meta repo**, alongside -`flake.nix`, at `/var/lib/hyperhive/meta/topology.json`: +"The roster" means two different, non-overlapping things depending on +scope. The [README](../../README.md)'s control-plane row ("`swarm-controller` +holds the hive directory, the agent roster and the job graph") is the +**swarm-wide** one: `GET /api/agents` on swarm-controller reads +straight from the swarm's authelia identity store over +`swarm-authelia-bridge` (`swarm-controller/src/auth.rs::list_agent_identities`) +— "every agent the swarm holds an identity for" (`swarm-controller/src/main.rs`'s +`get_agents` doc comment). It's authoritative: an agent exists in the +swarm if and only if it has an identity there, and `swarmctl agent +create` (or the swarm UI, via `POST /api/agents`) writes that identity +once, by calling `ensure_agent_identity`. + +This page is about a different, **hive-local** file: the hive-c0re-owned +**meta repo**, alongside `flake.nix`, at +`/var/lib/hyperhive/meta/topology.json`: ```json ["alice", "bob", "ruth"] ``` -One entry per agent the hive knows about, in name order. The file -carries no per-agent value any more, and encodes no ordering or -grouping — it answers exactly one question, _which agents exist,_ and -`topology::all_agents` is the only reader that matters. +One entry per agent _this hive_ currently has state/config for, in name +order. Unlike the swarm roster, nothing writes this file directly — +it's a derived cache, rebuilt by the reconcile pass below from what the +hive observes locally (config repos cloned, containers spawned), and it +answers a narrower question than "does this agent exist": _which of +this hive's local agents the `ManageRootAgent` capability's bind-mounts +should cover._ `topology::all_agents` is the only reader that matters. That reader is a permission boundary: the set it returns is what an agent holding the `ManageRootAgent` capability gets bind-mounted (each other @@ -93,7 +109,10 @@ consequence of removing the field, not a side effect of it. Capability enforcement isn't fully wired yet, so the **manager (`ruth`) still gets some hard-coded special treatment** -other agents don't: +other agents don't. A hive can opt out of having one at all +(`services.hyperhive.ruthless = true` skips `hive-c0re`'s root-agent +create/start sweep entirely, `hive-c0re/src/workers/auto_update.rs`); +everything below applies only when it doesn't: @@ -103,14 +122,20 @@ other agents don't: approval step — every other agent is created at swarm level (`swarmctl agent create`). Roster-wise, `ruth` is just another entry. -- **Wire-protocol** — the privileged `Request` variants - (`Kill` / `Start` / `Restart` / `Update`; `GetLogs`) — marked - `*(privileged)*` in `hive-core-agent-sock`'s unified `Request` enum — - are reachable only from the manager's socket flavour today; each is - planned to become a capability check. One exception: `Wake` (inject a `from: ` message into the - caller's own inbox) isn't really privileged — every per-agent daemon - (for example `hive-forge-notify`) needs it, and sub-agents already have the - equivalent on their own socket. +- **Wire-protocol** — the only `*(privileged)*` `Request` variants left + in `hive-core-agent-sock`'s unified enum are the scheduling ops + (`RequestSchedulePrompt`, `CancelSchedule`, `FireScheduleNow`, + `EditSchedule`), reachable only from the manager's socket flavour + today, matching the `scheduling` tool group + ([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Container lifecycle + ops (kill/start/restart/rebuild) never lived on this socket — they go + through the separate host-admin socket `hivectl` speaks. `Wake` + (inject a `from: ` message straight into the caller's own inbox) + is on this socket too but isn't privileged to either flavour; no + built-in in-container producer calls it today — matrix, bash and + forge notifications all moved to pushing a todo on the harness's + in-agent socket instead (see + [`docs/turn-loop/mcp.md`](../turn-loop/mcp.md#waking-the-agent-from-inside-the-container)). - **Storage/mounts** — only the manager container gets `/var/lib/hyperhive/agents` bind-mounted RW at `/agents` (so it can manage any agent's state dir — config isn't authored there, since a @@ -122,12 +147,20 @@ other agents don't: path writes `flake.lock` any more — `request_update_meta_inputs` was removed, leaving the operator dashboard's `POST /api/meta-update` as the only entry point. -- **Prompt/tools** — the system prompt uses `` / - `` marker blocks, and a `Flavor::{Agent, -Manager}` switch picks the MCP tool allow-list claude sees. Both are - already parametrised on a single flavour value, so the planned - per-capability-group version (`cap:` prompt blocks + a - matching tool allow-list) is additive rather than a rewrite. +- **Prompt** — _not_ treated differently in practice any more: `prompts/system.md` + still carries `` / `` marker + blocks, but `prompt::render` filters for `"agent"` unconditionally for + every container, manager included ("always `agent` role — there is + only one role", `hive-agent/src/prompt.rs`'s own module doc). The + `role:manager` blocks are dead in production, exercised only by a + unit test (`filter_role_blocks(SAMPLE, "manager")`). +- **Tool allow-list** — also not a flavour switch: the MCP tools claude + sees come from `HIVE_TOOL_GROUPS` alone, the same mechanism for every + agent ([`docs/turn-loop/mcp.md`](../turn-loop/mcp.md)). Ruth's wider + default surface is just a wider default grant + (`ToolGroup::MANAGER_DEFAULT`, seeded by `auto_update.rs` whenever its + groups aren't already set), not anything keyed off its name or + container. - **State dirs** — _not_ special-cased: `HYPERHIVE_STATE_DIR` is injected uniformly via `systemd.globalEnvironment` for every container including the manager, so all token/state paths resolve diff --git a/docs/turn-loop/mcp.md b/docs/turn-loop/mcp.md index 2db07c13..c4bae992 100644 --- a/docs/turn-loop/mcp.md +++ b/docs/turn-loop/mcp.md @@ -9,10 +9,14 @@ URL survives the per-turn claude re-spawn (and a host-side hive-c0re restart) — there is no per-turn MCP re-registration race for the built-in surface. HTTP is the sole transport for it (no stdio fallback). Extra servers (`services.hyperhive.agent.extraMcpServers`) pick their own transport per -entry (`type = "stdio" | "http"`, default `"stdio"`): `matrix` stays a -stdio bridge, `bash` runs its own persistent streamable-http listener -(`hive-bash-daemon`) — same reasoning as the built-in surface. The -server name is `hyperhive`, so the tools land in claude as +entry (`type = "stdio" | "http"`, default `"stdio"`): `matrix` and +`subagent` each run their own persistent streamable-http listener +(`hive-matrix-daemon`, `hive-subagent-daemon`), same reasoning as the +built-in surface and `bash`'s `hive-bash-daemon`. The GUI bridge +(`screen`, gated on `services.hyperhive.agent.gui.enable`) is the one +entry that still takes the schema's stdio default — it spawns +`hive-screen-mcp` fresh each turn. The server name is `hyperhive`, so +the tools land in claude as `mcp__hyperhive__`. Each entry also has an `availableToSubagents` toggle (default `false`) — see [`docs/tools/subagent.md`](../tools/subagent.md#mcp-servers-available-to-a-subagent) @@ -56,32 +60,43 @@ are opt-in via the P3RM1SS10NS tab. `ack_until(N)` to prevent re-pop. Acked rows never redeliver. Transient pings (sentinel id 0) have nothing to ack and show no marker. -**System messages** (from sender `system`): the broker delivers the -higher-urgency lifecycle events (`hive_sh4re::manager::HelperEvent`) -as regular inbox messages (same `recv` path; body is a JSON -object with an `event` discriminant field). The **submitting agent** -(the root agent for top-level containers; an agent with the `approvals` -tool group for its own subtree) receives `container_crash`, -`needs_update`, and `approval_resolved` this way. The remaining, lower-urgency lifecycle -notices — `spawned`, `rebuilt`, `killed`, `destroyed`, `needs_login`, -`logged_in` — skip the inbox entirely: they land as -todos on the submitting agent's in-container todo socket instead -(`Coordinator::push_todo`/`push_todo_submitter`, `subsystem = "core"`), -which still wakes a turn (the todo-wake path — see [Turn -outcomes](README.md#turn-outcomes)) but via a generic "call -`get_loose_ends`" prompt rather than the event body itself. Full -payload shapes and routing logic in -[`docs/agent-lifecycle/approvals.md` § Helper events](../agent-lifecycle/approvals.md#helper-events-to-the-submitting-agent). +**System messages** (from sender `system`): the broker delivers +`hive_sh4re::manager::HelperEvent` (three variants: `ApprovalResolved`, +`ContainerCrash`, `NeedsUpdate` — the last declared but constructed by +no call site today) as regular inbox messages (same `recv` +path; body is a JSON object with an `event` discriminant field). +`ApprovalResolved` goes to whichever agent actually submitted the +approval (`Coordinator::notify_submitter`, looked up from the +authenticated socket caller at submit time; a legacy row with no +recorded submitter falls back to the manager, `ruth`). +`ContainerCrash` always goes to `ruth` (`Coordinator::notify_manager`, +hardcoded — `hive-c0re/src/workers/crash_watch.rs`). A `MergeConfigPr` +approval's rebuild additionally pushes a `rebuilt:` todo to that +same submitter (`Coordinator::push_todo_submitter`, `subsystem = +"core"`), which still wakes a turn (the todo-wake path — see [Turn +outcomes](README.md#turn-outcomes)) via a generic "call +`get_loose_ends`" prompt rather than the event body itself. Lifecycle +transitions the job-queue scheduler or crash watcher drive directly — +stop/kill, destroy, a flake-rev login or logout state change — reach +no individual agent any more: they publish onto a swarm-wide NATS +JetStream stream instead (`swarm_notices::notify`, +`hive-c0re/src/swarm_notices.rs`), since every hive is swarm-controlled +now and there is no manager-agent fallback left to push an in-container +todo to. **Inbox** (`inbox` group): `get_loose_ends()`, `cancel_loose_end(kind, id)`, `remind(message, delay_seconds? | at_unix_timestamp?)`. -- `get_loose_ends()` — list scheduled reminders, pending approvals you - submitted, and active local tasks published by external MCP daemons - (for example running bash tasks from `hive-bash-daemon`). Each row - carries an id + kind for `cancel_loose_end`. Always your own threads — - there is no way to target another agent. +- `get_loose_ends()` — list this agent's own open threads: pending + reminders, open todos pushed by an in-container subsystem (matrix, + bash, forge, or any user-configured MCP server), an undelivered-inbox + count, and unread matrix notifications; `approval` rows only appear + when the caller is the manager (`ruth`) — sub-agents don't submit + approvals. Always the caller's own rows — there is no way to target + another agent. Full per-variant wire shape in + [`docs/process/conventions.md` § Loose-ends wire + shape](../process/conventions.md#loose-ends-wire-shape). - `cancel_loose_end` — hard-delete a `reminder`, cancel a pending `approval` row, or clear a `todo` row (loose-ends-v2). Agents may only cancel rows they own; the `approval` kind is further restricted @@ -96,11 +111,8 @@ No same-turn self-continue tool exists — see (bash-task completion, forge notification, matrix activity) for work already in flight. -**Meta** (`meta` group): `set_status(text)`, `get_agent_meta(name?)`. +**Meta** (`meta` group): `get_agent_meta(name?)`. -- `set_status` — set a free-text status string visible on the - dashboard. Single line, ≤ 200 chars. Persisted to - `{state_dir}/hyperhive-status`. Pass `""` to clear. - `get_agent_meta` — fetch identity + status metadata for an agent: `{ name, hyperhive_rev, running, status_text, status_set_at, hive_name?, swarm_name?, matrix_accounts? }`. `matrix_accounts` is a @@ -108,8 +120,14 @@ hive_name?, swarm_name?, matrix_accounts? }`. `matrix_accounts` is a `homeserver`); omitted for agents with no matrix provisioning. Omit `name` to query self. -**Always-on, no tool group** (like `set_status`): `compact()`. +**Always-on, no tool group**: `set_status(text)`, `compact()`, +`mark_todos_done(ids)`. +- `set_status` — set a free-text status string visible on the + dashboard. Single line, ≤ 200 chars. Persisted to + `{state_dir}/hyperhive-status`. Pass `""` to clear. Not gated by the + `meta` group — every agent needs to keep its dashboard status chip + current regardless of which optional groups it holds. - `compact` — agent self-service equivalent of the operator dashboard's `/compact` button. No args. Gated server-side on the last completed turn's context usage: refused (with an explanation, no side effect) @@ -120,14 +138,21 @@ hive_name?, swarm_name?, matrix_accounts? }`. `matrix_accounts` is a notes-checkpoint turn still fires first. Dispatched through the in-agent socket (`hive-agent-sock::Request::Compact`), not the broker — see `hive-agent/src/todo_server.rs`. +- `mark_todos_done(ids)` — bulk-clear specific loose-ends-v2 todo rows + by id in one call, instead of `cancel_loose_end`ing each one. + List-based, not range-based (no `ack_until`-style "clear below id + N"): pass the ids `get_loose_ends` actually showed you reviewed; + unknown or already-done ids are silently skipped. Dials the same + in-agent todo socket as `cancel_loose_end`'s `todo` kind. ## Privileged tools (by tool group) - **Bash execution** (`execution`) — background shell tasks. See [`docs/tools/bash.md`](../tools/bash.md). -- **Subagent spawning** — headless claude sub-instances as background - tasks, shipped default-on like bash execution (no tool group gates it - yet). See [`docs/tools/subagent.md`](../tools/subagent.md). +- **Subagent spawning** — nested sessions on the agent's own runtime + (claude or ACP, via `hive-runtime`) as background tasks, shipped + default-on like bash execution (no tool group gates it yet). See + [`docs/tools/subagent.md`](../tools/subagent.md). - **Lifecycle + config** (`lifecycle`, `approvals`) — neither group carries an MCP tool any more: `list_containers` and `request_update_meta_inputs` no longer exist, with no @@ -153,23 +178,23 @@ hive_name?, swarm_name?, matrix_accounts? }`. `matrix_accounts` is a ## Waking the agent from inside the container -External MCP servers (and any other in-container process) can -inject a wake-up event into the agent's inbox via the per-agent -socket at `/run/hive/mcp.sock`. Speak the wire protocol directly — -JSON-line over the unix socket: `{"cmd":"wake","from":"matrix","body": -"new dm from @alice"}\n`. Same shape as any other request on this -socket; see `hive_core_agent_sock::Request::Wake`. Every built-in producer that wakes -the harness (matrix, bash) dials the socket directly — there is no -CLI wrapper, just the raw protocol. +The built-in producers (matrix, bash, forge) no longer dial a direct +wake — all three now upsert a todo on the harness's in-agent socket +(`HIVE_AGENT_SOCKET`, `UpsertTodo`, see [Inbox](#core-tools-always-available) +above): a new or changed summary makes the harness signal its own turn +loop, with no hive-c0re round-trip. An external MCP server (or any +other in-container process) can push its own todo the same way — +`subsystem` is a plain string, not a closed set. -The wake event lands in the broker as `{from: