Apply contraction fixes across ~40 doc files (setup, integrations, lifecycle, networking, scheduler, swarm, tools, trust-boundary, UI, etc.). Skipped 14 hits: - 10 where words appear in ALL CAPS for deliberate emphasis (is NOT, do NOT, etc.) - 4 where text could not be safely located due to markdown formatting or column position Applied via systematic scan with checks for fenced code blocks, inline code spans, and intentional caps. Preserves sentence-initial capitalization throughout.
9 KiB
Turn loop + MCP
How the harness wakes up, what it asks claude to do, and what tools claude has access to in return.
The loop
Each agent harness (hive-agent — one serve-loop binary for all
agents) runs:
- Check the pause marker (
<harness>/paused). While it exists the loop does nothing but re-stat it every 5 s — no broker poll, no claude process. Because step 1 is never reached, messages stay queued and unacked, so a resume drains the backlog instead of losing it; reminders and todo wakes buffer in their channels. Set it withhivectl agent <name> pauseor the dashboard toggle; see persistence. - Long-poll
Recvon its socket. The host-side broker (broker.rs::recv_blocking_batch) returns immediately if there's a pending message, otherwise waits up to 30 s for a brokerSentevent for this recipient. - Pop one message. Peek the remaining inbox depth with
Status. - Emit
LiveEvent::TurnStart { from, body, unread }onto the SSE bus. - Spawn claude (one process per turn) and pipe the wake prompt over stdin.
- Stream stdout (JSON lines) into the bus as
LiveEvent::Stream(value). Pump stderr asNote. - Wait for claude to exit and classify the turn's outcome from the stream + exit — success, compaction, rate-limit, auth-failure, or hard failure. The outcome drives the post-turn action (see Turn outcomes); compaction is handled inside the session (see Compaction). Rate-limit and auth-failure detection is described below.
- Emit
LiveEvent::TurnEnd { ok, note }. Sleeppoll_msto avoid tight loops on transient failures.
Failure detection and login
- Rate limit — a
429/rate_limitmarker on stderr, or a parsed{"type":"error"}rate-limit event on stdout (conversation-text mentions don't count), sets therate_limitedsentinel, parks forHIVE_RATE_LIMIT_SLEEP_SECS(default 300), then retries. The UI shows a⊘ rate limitedbadge while parked. - Auth failure (401) —
drive_turnretries once (transient token-refresh races clear on retry); a secondAuthFailedwrites{state_dir}/hyperhive-needs-login, requeues the message, and parks inwait_for_login— the same path as a cold boot with no session. The operator re-auths via the per-agent web UI; the queued message then drives the next turn. - Login detection — both boot (
login::has_session, Online vs NeedsLogin) andwait_for_login's resume check key off the credential files inlogin::CRED_FILE_NAMES(the set/logoutdeletes).wait_for_logintakes asince: SystemTimebaseline (the instant of the 401 that parked it, orlogin::NO_PRIOR_FAILUREat cold boot) and resumes only once a credential file's mtime postdates it — so stale credentials already on disk at the 401 don't trigger an instant false-resume, and a login that lands beforewait_for_logineven starts polling still resumes correctly (baselining on a fixed instant rather than an entry-time directory snapshot is what closes that race). Leftover session-history files still don't read as a live session after a logout + container recreate (has_session/is_cred_filescope toCRED_FILE_NAMESeither way).
Harness binary shape
Two sibling crates, both role-agnostic (there is one role: agent —
the privilege boundary lives server-side at the broker socket
(/run/hive/mcp.sock), which refuses privileged Request variants
regardless of who sends them):
hive-agent— long-running harness loop (the inbox poll + claude-pump + ack/requeue cycle described above).hive-agent-mcp— MCP server for the built-inhyperhivesurface. Run with--http <addr>as a persistent streamable-HTTP daemon (thehive-mcp-httpsystemd unit, onhyperhive.mcp.httpPort, default 8790); claude connects to its URL via--mcp-config. HTTP is the sole transport — no per-turn stdio child (eliminates the re-registration race).
hive-agent's wire types (hive_core_agent_sock::{Request, Response} —
one unified enum shared by the agent and manager sockets) and its turn
loop are factored through a small Surface trait with one zero-sized
impl, so the loop itself has no per-role branches. See
hive-agent/src/main.rs's module
doc for the trait shape.
Boot wiring
serve_main reads HIVE_PORT (default DEFAULT_WEB_PORT) +
HIVE_LABEL (default "hive" for standalone runs; the meta
flake sets it unconditionally for any container-deployed agent;
see docs/process/conventions.md::Hive identity for the env stack),
opens turn-stats sqlite, prepares the on-boot files (see
claude-invocation),
installs claude plugins, spawns web_ui::serve + vacuum::run,
and either drops into serve_loop directly (Online) or parks on
the login flow first (NeedsLogin). Forge notifications are polled by
their own process, not this loop — see hive-forge-notify in
forge.md.
Boot also opens the todos store and the socket in-container producers
dial. Matrix / bash / forge-notify daemons and the in-process
disk_watch todo producer (low state-disk space) are the built-in
producers, but the socket accepts any subsystem marker — a
user-configured MCP server can push its own todos the same way. See
docs/agent-lifecycle/persistence.md for
what each built-in todo producer watches and how the store + get_loose_ends
merge work.
Plugin install failures aren't fatal: each entry comes back as a
human-readable failure string that gets routed via
Surface::send_to_parent to the agent's topology parent (the
broker resolves <parent> per topology::resolve_recipient; root
agents and the manager fall through to operator).
Turn outcomes
turn::TurnOutcome (Result<bool, TurnError> — Ok(compacted) on success,
else a TurnError) drives the post-claude branch:
| Outcome | Action |
|---|---|
Ok(_) (false normal / true compacted) |
ack_turn |
Err(PromptTooLong) |
drive_turn archived the session (the lib already compacted + retried and it still overflowed); requeue inflight so the message redelivers into a fresh session that fits — no status park |
Err(RateLimited) |
sleep HIVE_RATE_LIMIT_SLEEP_SECS (default 300), requeue inflight, status back to online |
Err(AuthFailed) |
emit needs_login_idle sentinel, requeue inflight, park in wait_for_login |
Err(SessionNotFound) |
resume + create self-heal both missed ("shouldn't happen"); requeue inflight so the next turn creates fresh — no status park, message not dropped |
Err(ApiStall) |
idle watchdog killed claude after HIVE_TURN_IDLE_SECS (default 600) of output silence; sleep HIVE_STALL_SLEEP_SECS (default 60), requeue inflight, status back to online |
Err(Failed(err)) |
route [system] \` claude turn failed:\ntoviasend_to_parent` |
ApiStall catches an Anthropic API stall — a multi-retry connection storm where
the stream goes silent for minutes. The idle watchdog lives in hive-claude's
driver (Config::idle_timeout, enforced around child.wait()): the timer
resets on every stdout line, so a large but still-streaming turn is never cut —
only complete output silence for the window trips it. The harness sets the
window from HIVE_TURN_IDLE_SECS (0 disables) and maps the driver's
Error::IdleTimeout onto TurnError::ApiStall.
After the outcome handler, the stats sink records a row. handle_turn
reports the result to serve_loop via TurnControl { auth_failed } —
on auth failure the loop parks in wait_for_login; otherwise it loops
straight back to the idle wait (step 1). There is no same-turn
self-continue mechanism: every multi-step continuation rides an
external wake instead — a new inbox message, a remind, or an
in-container todo wake (bash-task completion, forge notification,
matrix activity). Ending the turn and letting one of
those drive the next one is strictly better than parking in-process:
it checkpoints the session and observes wakes that only reach the
harness between turns.
Sub-pages
The rest lives alongside this page, in three topic files:
- claude-invocation.md — how the harness
spawns
claude --printeach turn, the two-pronged compaction (reactive + proactive), and the on-boot files it materialises (--mcp-config,--system-prompt-file). - config.md — the optional per-agent knobs the meta flake wires in (reference docs, icon, passwordless sudo, dashboard links, custom static files, connectivity overrides, claude plugins, cargo message filtering).
- mcp.md — the MCP tool surface claude sees: core tools, privileged tool groups, self-wake, authoritative state, the tool envelope, and the built-in tool allowlist.
Per-subsystem impl detail lives in each module's //! doc-comment; these pages
describe present-state behaviour + wiring, not line-level mechanics.