20 KiB
Turn loop + MCP
How the harness wakes up, what it asks claude to do, and what tools claude has access to in return.
The loop
Each agent harness (hive serve, role set via $HIVE_ROLE — always
"agent", one binary) runs:
- Long-poll
Recvon its socket. The host-side broker (broker.rs::recv_blocking_batch) returns immediately if there's a pending message, otherwise waits up to 30 s for a brokerSentevent for this recipient. - Pop one message. Peek the remaining inbox depth with
Status. - Emit
LiveEvent::TurnStart { from, body, unread }onto the SSE bus. - Spawn claude (one process per turn) and pipe the wake prompt over stdin.
- Stream stdout (JSON lines) into the bus as
LiveEvent::Stream(value). Pump stderr asNote. - Wait for claude to exit. Compaction is two-pronged — reactive
on
Prompt is too longand proactive on a context watermark (see Compaction below). Rate-limit detection: on stderr the harness does a raw-line match for429/rate_limitmarkers; on stdout it only fires on parsed{"type":"error"}JSON events (avoiding false positives when agents discussrate_limit_errorin conversation text). On detection the harness sets therate_limitedsentinel (Bus::emit_status("rate_limited")), sleepsHIVE_RATE_LIMIT_SLEEP_SECS(default 300), then retries. The dashboard and per-agent page show a⊘ rate limitedbadge while the harness is parked. Auth-failed detection: both stdout and stderr pumps also matchAUTH_FAIL_MARKERS("authentication_failed",401, etc.). On the first 401,drive_turnretries the same prompt once immediately (transient token-refresh races and brief API hiccups can cause a 401 that clears on retry). Only if the retry also returnsAuthFaileddoesdrive_turnbubble it up to the serve loop, which then writes{state_dir}/hyperhive-needs-login, emitsneeds_login_idlestatus, requeues the inflight message (so it replays after re-auth), and parks inwait_for_login— the same path used at boot. The operator re-authenticates via the per-agent web UI login flow; on success the sentinel is cleared and the queued message drives the next turn normally. Mtime-snapshot resumption:wait_for_loginsnapshots the~/.claude/dir (newest file mtime + file count) at entry and only resumes when that snapshot advances — not just when credentials exist on disk. This prevents a silent infinite-401 loop: stale credentials already on disk at the time of the 401 no longer cause an immediate false-resume. TheDirSnapshotstruct tracks both axes; either a mtime advance OR a file-count change triggers resume (the count axis handles filesystems wheremodified()errors on every file). - Emit
LiveEvent::TurnEnd { ok, note }. Sleeppoll_msto avoid tight loops on transient failures.
Harness binary shape
One hive binary serves both roles. The split into
hive-ag3nt + hive-m1nd was collapsed because the privilege
boundary lives server-side at the broker socket
(/run/hive/mcp.sock): an agent-flavor socket refuses
ManagerRequest calls regardless of who sends them, so there's no
escalation risk in shipping the same code to both. main() reads
$HIVE_ROLE (set by harness-base.nix from hyperhive.role;
defaults to "agent" for standalone nix run invocations) and
dispatches.
Three subcommands:
serve— long-running harness loop (the inbox poll + claude-pump + ack/requeue cycle described above).mcp— stdio MCP server claude spawns via--mcp-configper turn. Same binary, different mode.wake --from <name> --body <body>— push a message into our own inbox so the next turn fires with the given body. Used by co-process daemons (matrix bridge, scraper, webhook listeners) to nudge claude on external events.--body -reads from stdin.
Surface trait + zero-sized type tags
AgentRequest / AgentResponse (= ManagerRequest / ManagerResponse —
type aliases) are the wire types. There is one role: agent.
bin/hive.rs factors the turn loop through a Surface trait with one
zero-sized impl (AgentSurface) wrapping:
- One async method per wire op:
ack_turn,requeue_inflight,inbox_unread,post_turn_counts,send_to_parent,self_wake,recv_next,wake_external.
main() calls serve_main::<AgentSurface> for all roles. The turn
loop (serve_loop / handle_turn / wake) has no per-role branches.
Boot wiring
serve_main reads HIVE_PORT (default DEFAULT_WEB_PORT) +
HIVE_LABEL (default "hive" for standalone runs; the meta
flake sets it unconditionally for any container-deployed agent;
see docs/conventions.md::Hive identity for the env stack),
opens turn-stats sqlite, prepares the on-boot files (see below),
installs claude plugins, spawns forge_notify::run + web_ui::serve,
and either drops into serve_loop directly (Online) or parks on
the login flow first (NeedsLogin).
Plugin install failures are not fatal: each entry comes back as a
human-readable failure string that gets routed via
Surface::send_to_parent to the agent's topology parent (the
broker resolves <parent> per topology::parent_of; root agents
and the manager fall through to operator).
Turn outcomes
turn::TurnOutcome drives the post-claude branch:
| Outcome | Action |
|---|---|
Ok / Compacted |
ack_turn |
RateLimited |
sleep HIVE_RATE_LIMIT_SLEEP_SECS (default 300), requeue inflight, status back to online |
AuthFailed |
emit needs_login_idle sentinel, requeue inflight, park in wait_for_login |
Failed(err) |
route [system] \` claude turn failed:\ntoviasend_to_parent` |
After the outcome handler, the stats sink records a row and the
hyperhive-continue sentinel (dropped by the request_next_turn
MCP tool) is consumed if present, firing self_wake so the next
turn starts with { from: "self", body: "continue" } even if the
inbox is empty.
The claude invocation
claude --print --verbose --output-format stream-json --model <name> \
--continue --settings /run/hive/claude-settings.json \
--system-prompt-file /run/hive/claude-system-prompt.md \
--mcp-config /run/hive/claude-mcp-config.json --strict-mcp-config \
--tools <builtins> --allowedTools <builtins+mcp>
# wake prompt piped over stdin
<name> is read from Bus::model() on each turn. The initial
default is set by hyperhive.model in the agent's agent.nix
(NixOS option; propagates via HIVE_DEFAULT_MODEL env var; falls
back to "haiku" if unset). The operator can flip it at runtime
with /model <name> in the web terminal — the next turn picks it
up. The choice is persisted to /harness/hyperhive-model so it
survives restart; override path: HYPERHIVE_MODEL_FILE env var
for tests.
Context-window size is looked up per-model via
events::context_window_tokens(model). Resolution order (first
match wins):
HIVE_CONTEXT_WINDOW_TOKENS_<KEY>env var, whereKEY(lowercased) is a substring of the active model name. Injected by the meta flake fromservices.hyperhive.c0re.contextWindowTokens(host-level NixOS option, defaults: haiku=200k, sonnet=1M, opus=1M). Override these for all agents at once without a per-agent config change.HIVE_CONTEXT_WINDOW_TOKENS— single global override for any model (useful in dev / test).- Hard fallback:
200_000(conservative; only reached outside NixOS where the env vars aren't set).
The effective window drives watermarks and is exposed at runtime
via /api/state.context_window_tokens so the UI can show a
percentage-of-window ctx badge.
--continue keeps a persistent session per agent (claude stores
sessions in ~/.claude/projects/, which is bind-mounted
persistently). Auto-compact and auto-memory are disabled via
--settings because hyperhive owns compaction — see
Compaction below.
A one-shot --continue suppression is available via
POST /api/new-session (or /new-session slash command in the
per-agent terminal) — Bus::take_skip_continue() flips an
AtomicBool once per turn, the next claude invocation drops
--continue, every subsequent turn resumes normal behaviour.
Compaction
claude's own in-session auto-compact is off (--settings); hyperhive
owns it explicitly in turn::drive_turn. There are two triggers:
- Reactive — claude-code prints
Prompt is too long(thePROMPT_TOO_LONG_MARKER). The session is already past the context window, so no turn can run on it —drive_turnruns/compactstraight away and retries the same wake-up prompt once. No notes-checkpoint turn is possible here: the detail is gone. - Proactive — a turn finishes cleanly but the last inference's
context size (
Bus::last_ctx_usage().context_tokens()) is at or above a watermark. While the session is still healthy,drive_turninjects one synthetic notes-checkpoint turn (CHECKPOINT_PROMPT— "context is filling up, flush durable state into/statenow") and then runs/compact. This gives the agent a chance to persist in-flight task state, decisions, and file paths before the conversation detail collapses into a summary.
The compact watermark defaults to 75% of context_window_tokens(model)
(dynamically derived — 150k for haiku, 750k for sonnet/opus). Override
with HIVE_COMPACT_WATERMARK_TOKENS (absolute token count); set to 0
to disable proactive compaction entirely (the reactive path always
applies). The proactive path is best-effort — a failed checkpoint turn
or /compact is surfaced as a Note but never fails the turn that
already succeeded. The operator can also force a compaction any time
via /api/compact.
- Auto session-reset — a third path that fires when both
conditions hold: context is ≥ a watermark (
HIVE_AUTO_RESET_WATERMARK_TOKENS, default 50% ofcontext_window_tokens(model)) AND the time since the last turn exceeds the assumed prompt-cache TTL (HIVE_CACHE_TTL_SECS, default3600). Claude's prompt cache lives ~5 minutes; if the cache is already cold, resuming with--continuepays the full re-upload cost of the current context with no benefit over starting fresh. So:drive_turninjects oneAUTO_RESET_CHECKPOINT_PROMPTnotes turn ("flush state to files, cache is cold") then armsBus::take_skip_continue()for the real turn — the next turn runs without--continue, starting a fresh session. Unlike proactive compaction the session is dropped entirely, not compacted. SetHIVE_AUTO_RESET_WATERMARK_TOKENS=0to disable.
The child runs with cwd = /state (when the bind exists; falls
back to the parent's cwd in dev), so any relative path in a tool
call (Read foo.md, Bash ls, Write notes.md) lands in the
agent's durable bind-mounted dir. CLAUDE.md auto-load walks
upward from /state — drop a per-agent CLAUDE.md there if you
want long-term hints that survive destroy/recreate.
The wake prompt is intentionally minimal: just the popped message's
from/body, plus an inline ({unread} more pending — drain via …) hint when unread > 0. Claude drives any further recv/send
itself via the embedded MCP server.
Whenever hive-c0re starts / restarts / rebuilds a container, it
also drops a system message into the agent's inbox via
Coordinator::kick_agent — a one-line "you were just (re)started,
check /state/ for your notes, --continue session is intact". The
next turn picks it up like any other inbox message.
On-boot files
hive_ag3nt::turn::write_* writes three files next to the per-agent
socket at /run/hive/ once at startup:
-
claude-mcp-config.json— re-invokes the running binary asmcpchild (so the same binary serves as harness + as claude's MCP child process). -
claude-settings.json— the--settingsblob (auto-compact and auto-memory off, effortLevel medium). -
claude-system-prompt.md— rendered fromhive-ag3nt/prompts/system.mdbyhive_ag3nt::prompt::render: HTML-comment markers (<!-- role:agent -->...<!-- /role:agent -->, same forrole:manager) gate the role-specific blocks; everything else is shared. Five placeholders are then substituted:{label}(short agent name),{qualified_label}(hive-qualifiedname@domainform),{operator_pronouns},{hive_identity}(e.g.on hive `pr1ma`; empty whenhyperhive.hiveNameis unset), and{swarm_identity}(same shape for the swarm). Pronouns come fromHIVE_OPERATOR_PRONOUNSenv (set by the meta flake fromservices.hyperhive.c0re.operatorPronouns, defaultshe/her). Passed via--system-prompt-file.Marker grammar.
<!-- role:X -->opens a block; matching<!-- /role:X -->closes it. The renderer always uses roleagent. Blocks with other role tags are elided. Nesting is NOT supported — a stray opener overrides until its closing tag (or end of file). A mismatched closer is elided from the output but does NOT pop the active role. Whitespace inside markers is tolerated (<!--role:foo-->parses the same as<!-- role:foo -->). Content outside any marker is always included.hive_identity/swarm_identityshape. Each carries a leading space + backticked name (on hive \pr1ma`,in swarm `constellat1on`) when the corresponding env var is set, otherwise empty string. The independence lets the template drop one or both into the opener prose without breaking single-hive deployments that never set the option; the renderer also treatsSome("")from a caller asNone` so empty-string env vars and missing env vars round-trip the same way.
The shared per-turn plumbing lives in hive_ag3nt::turn::{write_mcp_config, write_settings, write_system_prompt, run_turn, drive_turn, emit_turn_end, wait_for_login, compact_session} so the two binaries
can't drift.
MCP surface
The harness ships an embedded MCP server (rmcp 1.7). Claude launches
it as a stdio child via --mcp-config. The hyperhive socket name is
hyperhive, so the tools land in claude as mcp__hyperhive__<tool>.
Tool access is gated by tool groups (HIVE_TOOL_GROUPS). The default
preset (AGENT_DEFAULT) includes messaging, meta, inbox, and
execution. Privileged groups (lifecycle, approvals, scheduling,
diagnostics) are opt-in via the P3RM1SS10NS tab.
Core tools (always available)
Messaging (messaging group): send(to, body, in_reply_to?),
recv(wait_seconds?, max?), ask(question, options?, multi?, ttl_seconds?, to?), answer(id, answer).
send— message a peer (logical name) or the operator (to: "operator"). Useto: "<parent>"to address the topology parent without hardcoding the label; the broker resolves the sentinel at delivery time. Optionalin_reply_to: i64links the message to a prior id for thread rendering.recv— drain inbox. Withoutwait_seconds(or0) returns immediately. Positive value parks the turn up to that many seconds (cap 180) — incoming messages wake instantly.max(default 1, cap 32) drains up to N rows;wait_secondsapplies to the first, then drains up tomaxtotal.ask— surface a structured question to the operator (default) or a peer agent (to: "<agent>"). Non-blocking — returns a question id; the answer arrives as aquestion_answeredsystem event in the asker's inbox.optionsis advisory;multi=truerenders as checkboxes;ttl_secondsauto-cancels with answer[expired].answer— respond to aquestion_askedevent routed to this agent. Strict authorisation: only the declared target can answer.
Inbox (inbox group): get_loose_ends(),
cancel_loose_end(kind, id), remind(message, delay_seconds? | at_unix_timestamp?), request_next_turn().
get_loose_ends— list pending questions (asked/owed) and scheduled reminders. Each row carries an id + kind forcancel_loose_end.cancel_loose_end— withdraw aquestion(posts[cancelled by <self>]), hard-delete areminder, or cancel a pendingapprovalrow. Agents may only cancel rows they own; theapprovalkind is further restricted to the root agent (ruth) server-side.remind— schedule a reminder in this agent's own inbox. Large payloads spill to/agents/<self>/state/reminders/. Pending count capped at 50 per agent (HIVE_REMIND_MAX_PENDING_PER_AGENT).request_next_turn— ask the harness to start another turn immediately after this one ends, even if the inbox is empty. Next turn fires withfrom: "self"andbody: "continue".
Meta (meta group): set_status(text), get_agent_meta(name?).
set_status— set a free-text status string visible on the dashboard. Single line, ≤ 200 chars. Persisted to{state_dir}/hyperhive-status. Pass""to clear.get_agent_meta— fetch identity + status metadata for an agent:{ name, hyperhive_rev, running, status_text, status_set_at, hive_name?, swarm_name? }. Omitnameto query self.
Privileged tools (by tool group)
- Bash execution (
execution) — background shell tasks. Seedocs/tools/bash.md. - Lifecycle + config (
lifecycle,approvals) — manage child agents, spawn new ones, apply config commits. Seedocs/tools/lifecycle.md. - Scheduling + diagnostics (
scheduling,diagnostics) — scheduled prompts,get_logs. Seedocs/tools/scheduling.md. - Matrix MCP + extra servers —
mcp__matrix__*tools and per-agent extra MCP config. Seedocs/tools/matrix.md.
Waking the agent from inside the container
External MCP servers (and any other in-container process) can
inject a wake-up event into the agent's inbox via the per-agent
socket at /run/hive/mcp.sock. Two equivalent paths:
-
Shell out to
hive wake --from <label> --body <text>(use--body -to read body from stdin). Already on the container'sPATHsince the harness binary is insystemPackages. Convenient for shell-script integrations and co-process daemons (matrix bridge, webhook listeners, scrapers). -
Speak the wire protocol directly — JSON-line over the unix socket:
{"cmd":"wake","from":"matrix","body":"new dm from @alice"}\n. Same shape as any otherAgentRequest; seehive-sh4re::AgentRequest::Wake.
The wake event lands in the broker as {from:<label>, to:<agent>, body}, waking whatever recv call the harness
is currently blocked on. The next turn fires with the wake
prompt formed from that message.
Identity = socket: anything that can connect to
/run/hive/mcp.sock is implicitly trusted to inject these —
the bind-mount is the agent's own container only.
Authoritative state
hive_ag3nt::events::Bus carries the current turn-loop state in
addition to the broadcast channel and the events history. Variants:
Idle— sitting onRecvwaiting for mail.Thinking—claude --printis running for a turn.Compacting— operator-triggered/compactis in flight.
The harness flips state at the relevant transitions
(set_state(Thinking) before drive_turn, set_state(Idle)
after; set_state(Compacting) around compact_session). Exposed
via /api/state.turn_state + turn_state_since (unix seconds);
the agent page renders this rather than deriving from SSE events.
Tool envelope
mcp::run_tool_envelope: every MCP tool handler logs the request,
runs the body, logs the result. Pre-/post-log only — the inbox
status hint moved to the wake prompt + UI header.
Tool whitelist (mcp::ALLOWED_BUILTIN_TOOLS)
- Allowed built-ins:
Edit,Glob,Grep,Read,Write. - Tool-group-gated built-ins:
WebFetch,WebSearch(added when theweb_toolstool group is enabled — see P3RM1SS10NS tab). - Denied by omission or
claude-settings.jsondeny list:Bash,Task,NotebookEdit,TodoWrite. - Allowed MCP tools: as listed above (by tool group).
Bash is disallowed — shell execution goes through
mcp__bash__run (background tasks with structured output +
task-id tracking) instead of an interactive shell. The run /
status MCP tools (mcp__bash__run / mcp__bash__status) are always
in the --allowedTools list.
WebFetch / WebSearch are off by default; enable the web_tools
tool group in the P3RM1SS10NS tab and rebuild the agent to enable them.