hyperhive/docs/turn-loop.md
lexis 13fe5ef9d3 docs(turn-loop.md): document get_loose_ends agent parameter scoping
Add clarification for the agent parameter: omit to list own threads,
pass agent name for direct children (always accessible), or query_agent_state
capability for non-children. Note that hive-wide '*' query unavailable
on agent socket.
2026-06-04 18:38:19 +02:00

20 KiB

Turn loop + MCP

How the harness wakes up, what it asks claude to do, and what tools claude has access to in return.

The loop

Each agent harness (hive serve — one binary for all agents) runs:

  1. Long-poll Recv on its socket. The host-side broker (broker.rs::recv_blocking_batch) returns immediately if there's a pending message, otherwise waits up to 30 s for a broker Sent event for this recipient.
  2. Pop one message. Peek the remaining inbox depth with Status.
  3. Emit LiveEvent::TurnStart { from, body, unread } onto the SSE bus.
  4. Spawn claude (one process per turn) and pipe the wake prompt over stdin.
  5. Stream stdout (JSON lines) into the bus as LiveEvent::Stream(value). Pump stderr as Note.
  6. Wait for claude to exit. Compaction is two-pronged — reactive on Prompt is too long and proactive on a context watermark (see Compaction below). Rate-limit detection: on stderr the harness does a raw-line match for 429 / rate_limit markers; on stdout it only fires on parsed {"type":"error"} JSON events (avoiding false positives when agents discuss rate_limit_error in conversation text). On detection the harness sets the rate_limited sentinel (Bus::emit_status("rate_limited")), sleeps HIVE_RATE_LIMIT_SLEEP_SECS (default 300), then retries. The dashboard and per-agent page show a ⊘ rate limited badge while the harness is parked. Auth-failed detection: both stdout and stderr pumps also match AUTH_FAIL_MARKERS ("authentication_failed", 401, etc.). On the first 401, drive_turn retries the same prompt once immediately (transient token-refresh races and brief API hiccups can cause a 401 that clears on retry). Only if the retry also returns AuthFailed does drive_turn bubble it up to the serve loop, which then writes {state_dir}/hyperhive-needs-login, emits needs_login_idle status, requeues the inflight message (so it replays after re-auth), and parks in wait_for_login — the same path used at boot. The operator re-authenticates via the per-agent web UI login flow; on success the sentinel is cleared and the queued message drives the next turn normally. Mtime-snapshot resumption: wait_for_login snapshots the ~/.claude/ dir (newest file mtime + file count) at entry and only resumes when that snapshot advances — not just when credentials exist on disk. This prevents a silent infinite-401 loop: stale credentials already on disk at the time of the 401 no longer cause an immediate false-resume. The DirSnapshot struct tracks both axes; either a mtime advance OR a file-count change triggers resume (the count axis handles filesystems where modified() errors on every file).
  7. Emit LiveEvent::TurnEnd { ok, note }. Sleep poll_ms to avoid tight loops on transient failures.

Harness binary shape

One hive binary for all agents. The earlier split into hive-ag3nt + hive-m1nd was collapsed because the privilege boundary lives server-side at the broker socket (/run/hive/mcp.sock): ManagerRequest calls are refused by the standard agent socket regardless of who sends them.

Three subcommands:

  • serve — long-running harness loop (the inbox poll + claude-pump + ack/requeue cycle described above).
  • mcp — stdio MCP server claude spawns via --mcp-config per turn. Same binary, different mode.
  • wake --from <name> --body <body> — push a message into our own inbox so the next turn fires with the given body. Used by co-process daemons (matrix bridge, scraper, webhook listeners) to nudge claude on external events. --body - reads from stdin.

Surface trait + zero-sized type tags

AgentRequest / AgentResponse (= ManagerRequest / ManagerResponse — type aliases) are the wire types. There is one role: agent. bin/hive.rs factors the turn loop through a Surface trait with one zero-sized impl (AgentSurface) wrapping:

  • One async method per wire op: ack_turn, requeue_inflight, inbox_unread, post_turn_counts, send_to_parent, self_wake, recv_next, wake_external.

main() calls serve_main::<AgentSurface> for all roles. The turn loop (serve_loop / handle_turn / wake) has no per-role branches.

Boot wiring

serve_main reads HIVE_PORT (default DEFAULT_WEB_PORT) + HIVE_LABEL (default "hive" for standalone runs; the meta flake sets it unconditionally for any container-deployed agent; see docs/conventions.md::Hive identity for the env stack), opens turn-stats sqlite, prepares the on-boot files (see below), installs claude plugins, spawns forge_notify::run + web_ui::serve, and either drops into serve_loop directly (Online) or parks on the login flow first (NeedsLogin).

Plugin install failures are not fatal: each entry comes back as a human-readable failure string that gets routed via Surface::send_to_parent to the agent's topology parent (the broker resolves <parent> per topology::parent_of; root agents and the manager fall through to operator).

Turn outcomes

turn::TurnOutcome drives the post-claude branch:

Outcome Action
Ok / Compacted ack_turn
RateLimited sleep HIVE_RATE_LIMIT_SLEEP_SECS (default 300), requeue inflight, status back to online
AuthFailed emit needs_login_idle sentinel, requeue inflight, park in wait_for_login
Failed(err) route [system] \` claude turn failed:\ntoviasend_to_parent`

After the outcome handler, the stats sink records a row and the hyperhive-continue sentinel (dropped by the request_next_turn MCP tool) is consumed if present, firing self_wake so the next turn starts with { from: "self", body: "continue" } even if the inbox is empty.

The claude invocation

claude --print --verbose --output-format stream-json --model <name> \
  --continue --settings /run/hive/claude-settings.json \
  --system-prompt-file /run/hive/claude-system-prompt.md \
  --mcp-config /run/hive/claude-mcp-config.json --strict-mcp-config \
  --tools <builtins> --allowedTools <builtins+mcp>
# wake prompt piped over stdin

<name> is read from Bus::model() on each turn. The initial default is set by hyperhive.model in the agent's agent.nix (NixOS option; propagates via HIVE_DEFAULT_MODEL env var; falls back to "haiku" if unset). The operator can flip it at runtime with /model <name> in the web terminal — the next turn picks it up. The choice is persisted to /harness/hyperhive-model so it survives restart; override path: HYPERHIVE_MODEL_FILE env var for tests.

Context-window size is looked up per-model via events::context_window_tokens(model). Resolution order (first match wins):

  1. HIVE_CONTEXT_WINDOW_TOKENS_<KEY> env var, where KEY (lowercased) is a substring of the active model name. Injected by the meta flake from services.hyperhive.c0re.contextWindowTokens (host-level NixOS option, defaults: haiku=200k, sonnet=1M, opus=1M). Override these for all agents at once without a per-agent config change.
  2. HIVE_CONTEXT_WINDOW_TOKENS — single global override for any model (useful in dev / test).
  3. Hard fallback: 200_000 (conservative; only reached outside NixOS where the env vars aren't set).

The effective window drives watermarks and is exposed at runtime via /api/state.context_window_tokens so the UI can show a percentage-of-window ctx badge.

--continue keeps a persistent session per agent (claude stores sessions in ~/.claude/projects/, which is bind-mounted persistently). Auto-compact and auto-memory are disabled via --settings because hyperhive owns compaction — see Compaction below. A one-shot --continue suppression is available via POST /api/new-session (or /new-session slash command in the per-agent terminal) — Bus::take_skip_continue() flips an AtomicBool once per turn, the next claude invocation drops --continue, every subsequent turn resumes normal behaviour.

Compaction

claude's own in-session auto-compact is off (--settings); hyperhive owns it explicitly in turn::drive_turn. There are two triggers:

  • Reactive — claude-code prints Prompt is too long (the PROMPT_TOO_LONG_MARKER). The session is already past the context window, so no turn can run on it — drive_turn runs /compact straight away and retries the same wake-up prompt once. No notes-checkpoint turn is possible here: the detail is gone.
  • Proactive — a turn finishes cleanly but the last inference's context size (Bus::last_ctx_usage().context_tokens()) is at or above a watermark. While the session is still healthy, drive_turn injects one synthetic notes-checkpoint turn (CHECKPOINT_PROMPT — "context is filling up, flush durable state into /state now") and then runs /compact. This gives the agent a chance to persist in-flight task state, decisions, and file paths before the conversation detail collapses into a summary.

The compact watermark defaults to 75% of context_window_tokens(model) (dynamically derived — 150k for haiku, 750k for sonnet/opus). Override with HIVE_COMPACT_WATERMARK_TOKENS (absolute token count); set to 0 to disable proactive compaction entirely (the reactive path always applies). The proactive path is best-effort — a failed checkpoint turn or /compact is surfaced as a Note but never fails the turn that already succeeded. The operator can also force a compaction any time via /api/compact.

  • Auto session-reset — a third path that fires when both conditions hold: context is ≥ a watermark (HIVE_AUTO_RESET_WATERMARK_TOKENS, default 50% of context_window_tokens(model)) AND the time since the last turn exceeds the assumed prompt-cache TTL (HIVE_CACHE_TTL_SECS, default 3600). Claude's prompt cache lives ~5 minutes; if the cache is already cold, resuming with --continue pays the full re-upload cost of the current context with no benefit over starting fresh. So: drive_turn injects one AUTO_RESET_CHECKPOINT_PROMPT notes turn ("flush state to files, cache is cold") then arms Bus::take_skip_continue() for the real turn — the next turn runs without --continue, starting a fresh session. Unlike proactive compaction the session is dropped entirely, not compacted. Set HIVE_AUTO_RESET_WATERMARK_TOKENS=0 to disable.

The child runs with cwd = /state (when the bind exists; falls back to the parent's cwd in dev), so any relative path in a tool call (Read foo.md, Bash ls, Write notes.md) lands in the agent's durable bind-mounted dir. CLAUDE.md auto-load walks upward from /state — drop a per-agent CLAUDE.md there if you want long-term hints that survive destroy/recreate.

The wake prompt is intentionally minimal: just the popped message's from/body, plus an inline ({unread} more pending — drain via …) hint when unread > 0. Claude drives any further recv/send itself via the embedded MCP server.

Whenever hive-c0re starts / restarts / rebuilds a container, it also drops a system message into the agent's inbox via Coordinator::kick_agent — a one-line "you were just (re)started, check /state/ for your notes, --continue session is intact". The next turn picks it up like any other inbox message.

On-boot files

hive_ag3nt::turn::write_* writes three files next to the per-agent socket at /run/hive/ once at startup:

  • claude-mcp-config.json — re-invokes the running binary as mcp child (so the same binary serves as harness + as claude's MCP child process).

  • claude-settings.json — the --settings blob (auto-compact and auto-memory off, effortLevel medium).

  • claude-system-prompt.md — rendered from hive-ag3nt/prompts/system.md by hive_ag3nt::prompt::render: HTML-comment markers (<!-- role:agent -->...<!-- /role:agent -->, same for role:manager) gate the role-specific blocks; everything else is shared. Five placeholders are then substituted: {label} (short agent name), {qualified_label} (hive-qualified name@domain form), {operator_pronouns}, {hive_identity} (e.g. on hive `pr1ma`; empty when hyperhive.hiveName is unset), and {swarm_identity} (same shape for the swarm). Pronouns come from HIVE_OPERATOR_PRONOUNS env (set by the meta flake from services.hyperhive.c0re.operatorPronouns, default she/her). Passed via --system-prompt-file.

    Marker grammar. <!-- role:X --> opens a block; matching <!-- /role:X --> closes it. The renderer always uses role agent. Blocks with other role tags are elided. Nesting is NOT supported — a stray opener overrides until its closing tag (or end of file). A mismatched closer is elided from the output but does NOT pop the active role. Whitespace inside markers is tolerated (<!--role:foo--> parses the same as <!-- role:foo -->). Content outside any marker is always included.

    hive_identity / swarm_identity shape. Each carries a leading space + backticked name ( on hive \pr1ma`, in swarm `constellat1on`) when the corresponding env var is set, otherwise empty string. The independence lets the template drop one or both into the opener prose without breaking single-hive deployments that never set the option; the renderer also treats Some("")from a caller asNone` so empty-string env vars and missing env vars round-trip the same way.

The shared per-turn plumbing lives in hive_ag3nt::turn::{write_mcp_config, write_settings, write_system_prompt, run_turn, drive_turn, emit_turn_end, wait_for_login, compact_session} so the two binaries can't drift.

MCP surface

The harness ships an embedded MCP server (rmcp 1.7). Claude launches it as a stdio child via --mcp-config. The hyperhive socket name is hyperhive, so the tools land in claude as mcp__hyperhive__<tool>.

Tool access is gated by tool groups (HIVE_TOOL_GROUPS). The default preset (AGENT_DEFAULT) includes messaging, meta, inbox, and execution. Privileged groups (lifecycle, approvals, scheduling, diagnostics) are opt-in via the P3RM1SS10NS tab.

Core tools (always available)

Messaging (messaging group): send(to, body, in_reply_to?), recv(wait_seconds?, max?), ask(question, options?, multi?, ttl_seconds?, to?), answer(id, answer).

  • send — message a peer (logical name) or the operator (to: "operator"). Use to: "<parent>" to address the topology parent without hardcoding the label; the broker resolves the sentinel at delivery time. Optional in_reply_to: i64 links the message to a prior id for thread rendering.
  • recv — drain inbox. Without wait_seconds (or 0) returns immediately. Positive value parks the turn up to that many seconds (cap 180) — incoming messages wake instantly. max (default 1, cap 32) drains up to N rows; wait_seconds applies to the first, then drains up to max total.
  • ask — surface a structured question to the operator (default) or a peer agent (to: "<agent>"). Non-blocking — returns a question id; the answer arrives as a question_answered system event in the asker's inbox. options is advisory; multi=true renders as checkboxes; ttl_seconds auto-cancels with answer [expired].
  • answer — respond to a question_asked event routed to this agent. Strict authorisation: only the declared target can answer.

Inbox (inbox group): get_loose_ends(agent?), cancel_loose_end(kind, id), remind(message, delay_seconds? | at_unix_timestamp?), request_next_turn().

  • get_loose_ends(agent?) — list pending questions (asked/owed) and scheduled reminders. Each row carries an id + kind for cancel_loose_end. Omit agent to list your own threads. Pass agent: "<name>" to inspect a direct child agent (always accessible per topology enforcement); non-children require the query_agent_state capability. The "*" hive-wide query is not available on the agent socket.
  • cancel_loose_end — withdraw a question (posts [cancelled by <self>]), hard-delete a reminder, or cancel a pending approval row. Agents may only cancel rows they own; the approval kind is further restricted to the root agent (ruth) server-side.
  • remind — schedule a reminder in this agent's own inbox. Large payloads spill to /agents/<self>/state/reminders/. Pending count capped at 50 per agent (HIVE_REMIND_MAX_PENDING_PER_AGENT).
  • request_next_turn — ask the harness to start another turn immediately after this one ends, even if the inbox is empty. Next turn fires with from: "self" and body: "continue".

Meta (meta group): set_status(text), get_agent_meta(name?).

  • set_status — set a free-text status string visible on the dashboard. Single line, ≤ 200 chars. Persisted to {state_dir}/hyperhive-status. Pass "" to clear.
  • get_agent_meta — fetch identity + status metadata for an agent: { name, hyperhive_rev, running, status_text, status_set_at, hive_name?, swarm_name? }. Omit name to query self.

Privileged tools (by tool group)

  • Bash execution (execution) — background shell tasks. See docs/tools/bash.md.
  • Lifecycle + config (lifecycle, approvals) — manage child agents, spawn new ones, apply config commits. See docs/tools/lifecycle.md.
  • Scheduling + diagnostics (scheduling, diagnostics) — scheduled prompts, get_logs. See docs/tools/scheduling.md.
  • Matrix MCP + extra serversmcp__matrix__* tools and per-agent extra MCP config. See docs/tools/matrix.md.

Waking the agent from inside the container

External MCP servers (and any other in-container process) can inject a wake-up event into the agent's inbox via the per-agent socket at /run/hive/mcp.sock. Two equivalent paths:

  • Shell out to hive wake --from <label> --body <text> (use --body - to read body from stdin). Already on the container's PATH since the harness binary is in systemPackages. Convenient for shell-script integrations and co-process daemons (matrix bridge, webhook listeners, scrapers).

  • Speak the wire protocol directly — JSON-line over the unix socket: {"cmd":"wake","from":"matrix","body":"new dm from @alice"}\n. Same shape as any other AgentRequest; see hive-sh4re::AgentRequest::Wake.

The wake event lands in the broker as {from:<label>, to:<agent>, body}, waking whatever recv call the harness is currently blocked on. The next turn fires with the wake prompt formed from that message.

Identity = socket: anything that can connect to /run/hive/mcp.sock is implicitly trusted to inject these — the bind-mount is the agent's own container only.

Authoritative state

hive_ag3nt::events::Bus carries the current turn-loop state in addition to the broadcast channel and the events history. Variants:

  • Idle — sitting on Recv waiting for mail.
  • Thinkingclaude --print is running for a turn.
  • Compacting — operator-triggered /compact is in flight.

The harness flips state at the relevant transitions (set_state(Thinking) before drive_turn, set_state(Idle) after; set_state(Compacting) around compact_session). Exposed via /api/state.turn_state + turn_state_since (unix seconds); the agent page renders this rather than deriving from SSE events.

Tool envelope

mcp::run_tool_envelope: every MCP tool handler logs the request, runs the body, logs the result. Pre-/post-log only — the inbox status hint moved to the wake prompt + UI header.

Tool whitelist (mcp::ALLOWED_BUILTIN_TOOLS)

  • Allowed built-ins: Edit, Glob, Grep, Read, Write.
  • Tool-group-gated built-ins: WebFetch, WebSearch (added when the web_tools tool group is enabled — see P3RM1SS10NS tab).
  • Denied by omission or claude-settings.json deny list: Bash, Task, NotebookEdit, TodoWrite.
  • Allowed MCP tools: as listed above (by tool group).

Bash is disallowed — shell execution goes through mcp__bash__run (background tasks with structured output + task-id tracking) instead of an interactive shell. The run / status MCP tools (mcp__bash__run / mcp__bash__status) are always in the --allowedTools list.

WebFetch / WebSearch are off by default; enable the web_tools tool group in the P3RM1SS10NS tab and rebuild the agent to enable them.