hyperhive/docs/tools/subagent.md
atlas f817e27d4c subagent daemon: report a killed session as killed, not idle
A subagent whose claude process died on a signal — the kernel's OOM
killer, a stopped unit, an `interrupt` — was indistinguishable from one
that finished its turn: its entry left the `running` map, `status` fell
through to "a session exists on disk" and answered `idle`, and the
end-of-turn todo said the subagent had "finished". The usual next move
on that reading is `continue`, which resumes work that was cut mid-turn
with nothing having recorded that it was cut.

The driver already preserves how the child ended — `RunningClaude::wait`
returns `Error::Exit` carrying the `ExitStatus`, whose `signal()` is the
whole answer — so this reads it rather than having to recover it:
`classify_end` turns the outcome into `Complete` / `Killed { signal }` /
`Failed`, and `State::finish_turn` remembers a kill against the name
(cleared by the next confirmed spawn under it).

What an agent sees as a result:

- `status` reports the session killed, naming the signal, instead of idle;
- the todo the daemon pushes without being asked says the subagent was
  KILLED mid-turn rather than that it finished;
- `continue` still resumes such a session, but its reply says the
  previous turn was killed, so no caller carries on from cut-off work
  believing it was complete.

Refs #4326
2026-09-13 14:58:45 +02:00

4.5 KiB

Subagent daemon

hive-subagent-daemon (crate hive-subagent-mcp) spawns nested headless claude sessions on request. Own process, own systemd unit, own MCP server (subagent, not hyperhive) — independent of hive-bash-daemon: a subagent is a full nested claude process, a materially heavier capability than a background shell command, so it gets its own deployable/restartable unit rather than living inside the bash daemon.

Shipped default-on for every agent — nix/agent-modules/mcp.nix injects subagent into hyperhive.extraMcpServers via lib.mkDefault (allowedTools = ["*"]), same as bash. Default-on rather than unconditional: an agent.nix can override or drop the entry, which is what mkDefault is there for. The operator's own framing: default-on for now, a real opt-in capability later.

For what the tools do and when an agent should reach for them, see the subagent MCP server's own tool descriptions and the base:claude-subagents skill — this page covers the daemon as deployed infrastructure, not the agent-facing API.

Tools

Served under the subagent MCP server (mcp__subagent__<tool>): start, continue, status, interrupt.

State

In-memory only: a map of currently running processes, live only as long as the daemon process is. A daemon restart stops whatever was running rather than adopting it. The durable record of a subagent's existence is claude's own on-disk session (hive_claude::SessionStore), which continue reattaches to independent of the daemon's own lifetime — a restart loses the in-flight turn, not the subagent's history.

A killed turn

A subagent whose claude process dies on a signal — the kernel's OOM killer under memory pressure, a stopped unit, an interrupt from the owning agent — did not finish its turn, and the daemon says so rather than letting it settle back into idle:

  • status reports the session killed, naming the signal, instead of the idle it reports for a turn that ended on its own;
  • the end-of-turn todo the daemon pushes without being asked says the subagent was killed mid-turn, not that it finished;
  • continue still resumes such a session — often what you want — but its reply says the previous turn was killed, so nobody carries on from cut-off work believing it was complete.

The record is per-name, in memory with the rest of this daemon's state, and the next confirmed spawn under that name clears it. A daemon restart loses it along with everything else — the daemon has to outlive the kill to report it, which is what OOMPolicy=continue on the unit is for.

Compaction trade-off

Built on hive_claude::Claude::spawn + RunningClaude::wait directly rather than InfiniteSession::run, since only the low-level driver exposes a cancel handle to stop a turn mid-flight — that's what makes interrupt genuinely stop a running turn rather than only cancelling a still-pending one. The cost: a turn that overflows the context window surfaces as an error rather than self-healing via reactive compaction. Subagents are meant to be bounded, single-batch work, not sessions long-lived enough to need in-place compaction — a real follow-up if that assumption stops holding.

Configuration

hyperhive.mcp.subagentHttpPort — the daemon's streamable-http listen port. Same pattern as bashHttpPort/matrixHttpPort: a per-agent default assigned by nix/agent-modules/mcp.nix, only worth overriding for an agent that needs a stable or non-default port.

Own systemd unit, defined alongside the other per-agent MCP daemons in nix/agent-modules/mcp.nix.

MCP servers available to a subagent

A subagent runs with --strict-mcp-config and no --mcp-config by default — zero MCP servers, full stop; it falls back to claude's own native tools (Bash, WebFetch, etc.), not the parent's mcp__bash__* / mcp__hyperhive__* surface. Nothing implicit reaches it: the built-in hyperhive surface (todos/messaging) isn't an extraMcpServers entry at all, and the automatically injected bash/subagent entries default to excluded too (a subagent can't spawn hive-bash tasks or its own nested subagents unless an operator opts them in explicitly, same as anything else).

Set hyperhive.extraMcpServers.<name>.availableToSubagents = true on a specific entry to hand that one server to subagents as well — useful for, say, a read-only lookup or scraper MCP a subagent's bounded, single-batch task might need. hive-subagent-mcp's mcp_config module renders the opted-in subset into its own --mcp-config file per turn; an entry left at the default false never appears there.