Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/turn-loop
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 2b2608a491 docs(turn-loop): move harness systemd unit shape out of agent-roster.md
Moves the "Harness systemd unit shape" section from
docs/agent-lifecycle/agent-roster.md into docs/turn-loop/README.md: it
describes the per-agent harness systemd unit (env vars, PATH wiring,
serviceConfig), which is turn-loop material, not roster material.

Fixes two facts while moving: the ExecStart package is `hive-agent`,
not `hyperhive` (no package by that name exists); and `ruth.nix`
doesn't set any forge subscription default — it only defaults
`services.hyperhive.agent.docs.enable`.

Updates the inbound pointers in docs/turn-loop/config.md and the
module comment at nix/agent-modules/agent-service.nix.

Refs #3902
2026-10-02 14:47:31 +02:00
..
claude-invocation.md docs(turn-loop): facts + structure pass 2026-10-02 07:52:13 +02:00
config.md docs(turn-loop): move harness systemd unit shape out of agent-roster.md 2026-10-02 14:47:31 +02:00
mcp.md docs(agents): drop change-log section and temporal wording 2026-10-02 12:52:38 +02:00
README.md docs(turn-loop): move harness systemd unit shape out of agent-roster.md 2026-10-02 14:47:31 +02:00

Turn loop + MCP

How the harness wakes up, what it asks the agent's runtime to do, and what tools the agent has access to in return.

Runtimes

Every turn runs through hive-runtime, on the runtime services.hyperhive.agent.runtime selects:

Both report the turn in claude's stream-json shape, so the loop below is the same for either.

The loop

Each agent harness (hive-agent — one serve-loop binary for all agents) runs:

  1. Check the pause marker (<harness>/paused). While it exists the loop does nothing but re-stat it every 5 s — no broker poll, no claude process. Because step 1 is never reached, messages stay queued and unacked, so a resume drains the backlog instead of losing it; reminders and todo wakes buffer in their channels. Set it with hivectl agent <name> pause or the dashboard toggle; see persistence.
  2. Long-poll Recv on its socket. The host-side broker (broker.rs::recv_blocking_batch) returns immediately if there's a pending message, otherwise waits up to 30 s for a broker Sent event for this recipient.
  3. Pop one message. Peek the remaining inbox depth with Status.
  4. Emit LiveEvent::TurnStart { from, body, unread } onto the SSE bus.
  5. Hand the wake prompt to the runtime: on claude, spawn claude --print and pipe it over stdin; on acp, send it as a prompt on the agent's session.
  6. Stream the turn's stream-json lines into the bus as LiveEvent::Stream(value). Pump stderr as Note.
  7. Wait for the turn to end and classify its outcome from the stream + result — success, compaction, rate-limit, auth-failure, stall, or hard failure. The outcome drives the post-turn action (see Turn outcomes); the session handles compaction internally (see Compaction). This page describes rate-limit and auth-failure detection below.
  8. Emit LiveEvent::TurnEnd { ok, note }. Sleep poll_ms to avoid tight loops on transient failures.

Failure detection and login (claude)

  • Rate limit — a 429 / rate_limit marker on stderr, or a parsed {"type":"error"} rate-limit event on stdout (conversation-text mentions don't count), sets the rate_limited sentinel, parks for HIVE_RATE_LIMIT_SLEEP_SECS (default 300), then retries. The UI shows a ⊘ rate limited badge while parked.
  • Auth failure (401) — drive_turn retries once (transient token-refresh races clear on retry); a second AuthFailed writes {state_dir}/hyperhive-needs-login, requeues the message, and parks in wait_for_login — the same path as a cold boot with no session. The operator re-auths via the per-agent web UI; the queued message then drives the next turn.
  • Login detection — both boot (login::has_session, Online vs NeedsLogin) and wait_for_login's resume check key off the credential files in login::CRED_FILE_NAMES (the set /logout deletes). wait_for_login takes a since: SystemTime baseline (the instant of the 401 that parked it, or login::NO_PRIOR_FAILURE at cold boot) and resumes only once a credential file's mtime postdates it — so stale credentials already on disk at the 401 don't trigger an instant false-resume, and a login that lands before wait_for_login even starts polling still resumes correctly (baselining on a fixed instant rather than an entry-time directory snapshot is what closes that race). Leftover session-history files still don't read as a live session after a logout + container recreate (has_session/is_cred_file scope to CRED_FILE_NAMES either way).

Harness binary shape

Two sibling binaries over one runtime library, all role-agnostic (there is one role: agent — the privilege boundary lives server-side at the broker socket (/run/hive/mcp.sock), which refuses privileged Request variants regardless of who sends them):

  • hive-agent — long-running harness loop (the inbox poll + runtime turn + ack/requeue cycle described above).
  • hive-runtime — the library every turn runs through, one Runtime impl per runtime (ClaudeRuntime, AcpRuntime); hive-subagent-mcp drives its nested sessions through it too.
  • hive-agent-mcp — MCP server for the built-in hyperhive surface. Run with --http <addr> as a persistent streamable-HTTP daemon (the hive-mcp-http systemd unit, on services.hyperhive.agent.mcp.httpPort, default 8790); claude connects to its URL via --mcp-config. HTTP is the sole transport — no per-turn stdio child (eliminates the re-registration race).

A small Surface trait, with one zero-sized impl, factors hive-agent's wire types (hive_core_agent_sock::{Request, Response} — one unified enum shared by the agent and manager sockets) and its turn loop, so the loop itself has no per-role branches. See hive-agent/src/main.rs's module doc for the trait shape.

Boot wiring

serve_main reads HIVE_PORT (default DEFAULT_WEB_PORT) + HIVE_LABEL (default "hive" for standalone runs; the meta flake sets it unconditionally for any container-deployed agent; see docs/process/conventions.md::Hive identity for the env stack), opens turn-stats sqlite, prepares the on-boot files (see claude-invocation), installs claude plugins, spawns web_ui::serve + vacuum::run, and either drops into serve_loop directly (Online) or parks on the login flow first (NeedsLogin). hive-forge-notify polls forge notifications on its own, not this loop — see forge.md.

Boot also opens the todos store and the socket in-container producers dial. Matrix / bash / forge-notify daemons and the in-process disk_watch todo producer (low state-disk space) are the built-in producers, but the socket accepts any subsystem marker — a user-configured MCP server can push its own todos the same way. See docs/agent-lifecycle/persistence.md for what each built-in todo producer watches and how the store + get_loose_ends merge work.

Plugin install failures aren't fatal: each entry comes back as a human-readable failure string that gets routed via Surface::send_to_operator to the operator.

Turn outcomes

turn::TurnOutcome (Result<bool, TurnError> — Ok(compacted) on success, else a TurnError) drives the post-claude branch:

Outcome Action
Ok(_) (false normal / true compacted) ack_turn
Err(PromptTooLong) drive_turn archived the session (the lib already compacted + retried and it still overflowed); requeue inflight so the message redelivers into a fresh session that fits — no status park
Err(RateLimited) sleep HIVE_RATE_LIMIT_SLEEP_SECS (default 300), requeue inflight, status back to online
Err(AuthFailed) emit needs_login_idle sentinel, requeue inflight, park in wait_for_login
Err(SessionNotFound) resume + create self-heal both missed ("shouldn't happen"); requeue inflight so the next turn creates fresh — no status park, message not dropped
Err(ApiStall) idle watchdog killed claude after HIVE_TURN_IDLE_SECS (default 600) of output silence; sleep HIVE_STALL_SLEEP_SECS (default 60), requeue inflight, status back to online
Err(AgentStall(note)) the same, for an ACP agent: after HIVE_TURN_IDLE_SECS with no session/update the harness sent the agent session/cancel, and killed it if it ignored that; note says which
Err(Failed(err)) route [system] \` claude turn failed:\ntooperatorviasend_to_operator`

ApiStall catches an Anthropic API stall — a multi-retry connection storm where the stream goes silent for minutes. On claude the idle watchdog lives in hive-claude's driver (Config::idle_timeout, enforced around child.wait()): the timer resets on every stdout line, so a large but still-streaming turn is never cut — only complete output silence for the window trips it. The harness sets the window from HIVE_TURN_IDLE_SECS (0 disables) and maps the driver's Error::IdleTimeout onto TurnError::ApiStall.

hive-runtime applies the same window to an ACP agent's turns. That's also how a provider HTTP 429 ends an ACP turn when the agent retries it without reporting it: the turn goes silent until the watchdog stops it.

After the outcome handler, the stats sink records a row. handle_turn reports the result to serve_loop via TurnControl { auth_failed } — on auth failure the loop parks in wait_for_login; otherwise it loops straight back to the idle wait (step 1). No same-turn self-continue mechanism exists: every multi-step continuation rides an external wake instead — a new inbox message, a remind, or an in-container todo wake (bash-task completion, forge notification, matrix activity). Ending the turn and letting one of those drive the next one is strictly better than parking in-process: it checkpoints the session and observes wakes that only reach the harness between turns.

Harness systemd unit shape

One harness serve binary (hive-agent, with its hive-agent-mcp sibling), one shared nix/agent-modules/ tree, one service unit (systemd.services.hive-agent) for all agents. Privilege differences live server-side in the broker socket (which tool groups and manager-surface calls each agent receives).

agent.nix and ruth.nix both import the shared nix/agent-modules/. ruth.nix additionally defaults services.hyperhive.agent.docs.enable to true, so the root/manager gets the hyperhive reference docs readable at $HIVE_DOCS_DIR by default; other agents opt in per agent.nix (see config.md, "Reference docs").

Environment variables set on the unit

  • HOME = /home/<userName> — systemd defaults HOME to / for services without User= set; with the per-agent user the harness needs the right home so claude finds its bind-mounted ~/.claude/ session dir.
  • HIVE_STATIC_DIR = <mergedDist> — tower_http::ServeDir root for the per-agent web UI; merged dist = agent default + every services.hyperhive.agent.frontend.extraFiles overlay.
  • HIVE_ASSETS_DIR = "${config.services.hyperhive.agent.packages.assets}/share/hyperhive" (nix/agent-modules/agent-service.nix:480; packages.assets resolves to the hyperhive-assets derivation) — set directly on the unit, not via environment.variables, because the latter only populates /etc/profile which systemd services don't inherit.

PATH setup (the wrapper-dir trick)

path = [ "/run/wrappers" "/run/current-system/sw" ];

/run/wrappers (not /run/wrappers/bin) comes first so setuid wrappers — notably sudo — resolve before bare nix-store binaries; see docs/process/gotchas.md ("systemd.services.*.path appends /bin to every entry") for why the trailing /bin matters in general. It's load-bearing here because the harness runs as the per-agent user: without the wrapper dir on PATH, sudo resolves to the non-setuid nix-store binary and every services.hyperhive.agent.user.passwordlessSudo grant fails with "must be owned by uid 0 and have the setuid bit set."

serviceConfig highlights

  • ExecStart = "${config.services.hyperhive.agent.packages.hive-agent}/bin/${binary}" (nix/agent-modules/agent-service.nix:535); packages.hive-agent is wired to the flake's hyperhive.packages.<system>.hive-agent output (nix/agent-modules/packages.nix:8-22) — same binary for every agent.
  • Restart = on-failure, RestartSec = 2 — keeps the harness resilient across transient crashes without thundering retries.
  • RuntimeDirectory = "hive-config" → /run/hive-config/ owned by User=, autocleared on stop. The harness writes regenerated claude-{mcp-config,settings,system-prompt} files there (paths::config_dir). Deliberately separate from /run/hive, which the host bind-mounts in root-owned and which holds hive-c0re's mcp.sock.
  • User = Group = userName — drops root inside the container; sudo is the explicit escalation surface (services.hyperhive.agent.user.passwordlessSudo).

Sub-pages

The rest lives alongside this page, in three topic files:

  • claude-invocation.md — how the harness spawns claude --print each turn, the two-pronged compaction (reactive + proactive), and the on-boot files it materialises (--mcp-config, --system-prompt-file).
  • config.md — the optional per-agent knobs the meta flake wires in (reference docs, icon, passwordless sudo, dashboard links, custom static files, connectivity overrides, claude plugins, cargo message filtering).
  • mcp.md — the MCP tool surface claude sees: core tools, privileged tool groups, self-wake, authoritative state, the tool envelope, and the built-in tool allowlist.

Per-subsystem impl detail lives in each module's //! doc-comment; these pages describe present-state behaviour + wiring, not line-level mechanics.