hyperhive/docs/tools/subagent.md
atlas b18348bc9a subagent: give a run a goal, turns toward it, and a reason it stopped
`start` takes an optional `goal`. With one set a session stops being a
single turn: when a turn ends and nothing has said to stop, the daemon
spawns another turn re-prompting the subagent toward that goal, up to
`max_turns` (default 5, per-session). Without a goal nothing changes —
one turn, one todo, same as before.

Four things end a run, each recorded distinctly and reported by `status`:
the turn ending with no goal, `goal_reached`, `need_help`, and the turn
cap. The last says so out loud rather than stopping quietly — the todo
states the harness limit was reached and the goal was never reported
reached. Every stop extends the done message rather than replacing it,
and lands in the session's report file when it has one. The path is
never inferred: it comes from `start`'s `report_file` or from the
subagent naming where it wrote.

`goal_reached` and `need_help` are the subagent's own, served on a second
route (`/signal/mcp`) that carries those two tools and nothing else, so
reporting on a run can't become starting one. `goal_reached` is built as
a label, never a gate: it is self-reported by a subagent that has just
been re-prompted with "you haven't reached the goal", which is exactly
the incentive to claim it — the same failure class as a build report
asserting the tests pass. Every surface that renders it says so.
`need_help` is the blocking signal, and shows in `status` as its own
state so a parent polling it sees the block without reading a file.

`status` also carries `turn N of M`: with 4330's last-event age, that
separates working from wedged from out of turns off one answer.

Two bugs the new tests caught: a `tokio::fs::File` was dropped without
flushing, so the report line was written to nothing, and the plain idle
answer dropped the turn counter.

Also documents `await_resume`'s third case — a closed channel with no
send, which fails open the same as `Underway` — per argus on #4411.

Refs #4403
2026-09-14 21:46:59 +02:00

217 lines
11 KiB
Markdown

# Subagent daemon
`hive-subagent-daemon` (crate `hive-subagent-mcp`) spawns nested headless
`claude` sessions on request. Own process, own systemd unit, own MCP
server (`subagent`, not `hyperhive`) — independent of `hive-bash-daemon`:
a subagent is a full nested claude process, a materially heavier
capability than a background shell command, so it gets its own
deployable/restartable unit rather than living inside the bash daemon.
Shipped default-on for every agent — `nix/agent-modules/mcp.nix` injects
`subagent` into `hyperhive.extraMcpServers` via `lib.mkDefault`
(`allowedTools = ["*"]`), same as `bash`. Default-on rather than
unconditional: an `agent.nix` can override or drop the entry, which is
what `mkDefault` is there for. The operator's own framing: default-on for
now, a real opt-in capability later.
For what the tools do and when an agent should reach for them, see the
`subagent` MCP server's own tool descriptions and the
`base:claude-subagents` skill — this page covers the daemon as deployed
infrastructure, not the agent-facing API.
## Tools
Served under the `subagent` MCP server (`mcp__subagent__<tool>`): `start`,
`continue`, `status`, `interrupt`.
A second route on the same port serves the two tools a **subagent** calls
about its own run — `goal_reached` and `need_help`. It isn't part of the
`subagent` server an agent's own config points at; the daemon writes it
into each subagent's `--mcp-config` itself, under `subagent_control`. Two
routes rather than six tools on one, so that being able to report on a run
never carries the ability to start one: there's no route a subagent holds
that `start` is reachable from.
## State
In-memory only: what's running now, where each name's session lives, how
each name's last turn ended, when each running turn last produced output,
what each session is working toward, how far through its turn budget it
is, why its run stopped, and where it writes its report. All of it lives
only as long as the daemon process does. A daemon
restart stops whatever was running rather than adopting it. The durable
record of a subagent's existence is claude's own on-disk session
(`hive_claude::SessionStore`), which `continue` reattaches to independent
of the daemon's own lifetime — a restart loses the _in-flight turn_, not
the subagent's history.
Because the remembered directory goes with the rest of it, a `continue`
after a restart has to re-supply `dir` when the session lives anywhere
other than the daemon's own working directory — and a restart is the
situation you reach for `continue` in most often.
## Goals, and turns toward them
`start` takes an optional `goal`. Without one a session is a single turn,
exactly as it always was. With one, the daemon keeps the session going:
when a turn ends and nothing has said to stop, it starts another turn
re-prompting the subagent toward that goal, quoting it verbatim and saying
which turn of the budget this is. `max_turns` caps that, defaulting to
**5**.
`status` reports `Turn N of M` for such a session in every state it
reaches. Read alongside the last-event age below, it's what separates a
subagent that's working from one that's wedged from one that's out of
turns — without `ps` and without opening a file.
Four things stop a run, and each is recorded distinctly, reported by
`status`, and appended to the one todo the daemon pushes when the run ends:
- **the turn ended and there was no goal** — the single-turn case;
- **`goal_reached`**, which the subagent calls itself;
- **`need_help`**, likewise;
- **the turn cap**, which says so rather than stopping quietly: the todo
states that the harness limit was reached and the goal was never
reported reached, so the work stopped where it had got to.
A killed or failed turn ends the run too, and keeps the records it already
had — see [A killed turn](#a-killed-turn). `interrupt` therefore stops a
whole goal run, not just the turn in flight.
When the session was told where its report goes — `start`'s `report_file`,
or the path the subagent names when it signals — the stop reason is
appended to that file as well, so the artifact you were going to read
anyway also says how the run ended. Nothing is inferred: with no path
given, no file is touched.
## `goal_reached` is a label, not a gate
Both signals stop the continuation and **extend** the done message. Extend,
not replace: the turn's observed end and the reason the run stopped are
different facts, and the second never stands in for the first.
`goal_reached` is **self-reported**, by a subagent that has just been
re-prompted with "you haven't reached the goal" — which is precisely the
incentive to claim it. It's the same failure class as a build report
asserting "done, tests pass": a claim about an artifact, not the artifact.
Nothing in this daemon treats it as verification, and every surface that
renders it says so. Read the diff and the gate output regardless.
`need_help` is the blocking signal. It stops the run and shows up in
`status` as its own state — blocked, with the subagent's reason — so a
parent that polls `status` sees the block without reading anything else.
`continue` is how you answer it.
## Is it working, or is it wedged?
`status` reporting **running** says a process is tracked, which a wedged
subagent satisfies as fully as a busy one. A running answer therefore
carries the age of that turn's last event too: seconds means it's working,
an age climbing into the minutes means it's stuck. That one number
replaces inferring the same thing from `ps` output and CPU-time deltas.
It resets at each turn's spawn, so on a goal run it describes the turn in
flight rather than the run — which is what you want, since a run that's
making progress spends several perfectly healthy minutes.
Every line the subagent's `claude` process writes bumps the timestamp —
stream-json events, plain stdout chatter and stderr alike — and what the
subagent actually said is never read. The record says the child is alive,
not what it's doing. The clock starts at the spawn, so a subagent that wedged
before it ever emitted anything still reports a climbing age rather than
no age at all. It's dropped when the turn ends, since a finished turn has
no progress left to describe.
## A `continue` that finds no session
`continue` doesn't check for the session before spawning. claude's own
`--resume` is the authority, and it exits non-zero rather than quietly
starting a fresh session, so the check could only duplicate the lookup
the driver was about to do — while answering as though the session were
gone. The usual truth is that the session exists somewhere else.
`continue` waits for that answer instead. Where `start` returns the
instant the process exists — it creates its session, so the spawn
succeeding is the whole story — a resumed turn can fail a moment _after_
a pid exists, and a pid is no proof that turn began. `continue` holds the
tool call open until the turn is underway or the resume has come back
missed, and reports a miss as the call's own error, carrying claude's
message plus the location this daemon searched:
```
continue error: claude error: no session matched the requested id or title (searched /home/agent/.claude for cwd /home/agent/work; if it was started elsewhere, pass the `dir` it was started in)
```
The directory is the part claude's own message never names, and the part
that resolves the confusion: pass `dir` to point `continue` at the
directory the session was started in.
Nothing here is a fixed delay on the way to a successful turn. The wait
ends on whichever comes first — the turn's first stream event or its
early exit — and on this box both land inside a second, so a `continue`
that works answers about as fast as it did before. A five-second cap
bounds the one case neither covers: a child that neither speaks nor
exits, reported as started, with the end-of-turn todo left to say how it
goes. That todo still carries every failure that happens later in the
turn, exactly as before; the only one it no longer repeats is the miss
the caller has just been handed to its face.
## A killed turn
A subagent whose `claude` process dies on a signal — the kernel's OOM
killer under memory pressure, a stopped unit, an `interrupt` from the
owning agent — didn't finish its turn, and the daemon says so rather
than letting it settle back into `idle`:
- `status` reports the session **killed**, naming the signal, instead of
the `idle` it reports for a turn that ended on its own;
- the end-of-turn todo the daemon pushes without being asked says the
subagent was killed mid-turn, not that it finished;
- `continue` still resumes such a session — often what you want — but
its reply says the previous turn was killed, so nobody carries on from
cut-off work believing it was complete.
The record is per-name, in memory with the rest of this daemon's state,
and the next confirmed spawn under that name clears it. A daemon restart
loses it along with everything else — the daemon has to outlive the kill
to report it, which is what `OOMPolicy=continue` on the unit is for.
## Compaction trade-off
Built on `hive_claude::Claude::spawn` + `RunningClaude::wait` directly
rather than `InfiniteSession::run`, since only the low-level driver
exposes a cancel handle to stop a turn mid-flight — that's what makes
`interrupt` genuinely stop a running turn rather than only cancelling a
still-pending one. The cost: a turn that overflows the context window
surfaces as an error rather than self-healing via reactive compaction.
Subagents are meant to be bounded, single-batch work, not sessions
long-lived enough to need in-place compaction — a real follow-up if that
assumption stops holding.
## Configuration
`hyperhive.mcp.subagentHttpPort` — the daemon's streamable-http listen
port. Same pattern as `bashHttpPort`/`matrixHttpPort`: a per-agent default
assigned by `nix/agent-modules/mcp.nix`, only worth overriding for an
agent that needs a stable or non-default port.
Own systemd unit, defined alongside the other per-agent MCP daemons in
`nix/agent-modules/mcp.nix`.
## MCP servers available to a subagent
A subagent runs with `--strict-mcp-config` and, by default, exactly one
MCP server: the two-tool `subagent_control` route above. It otherwise
falls back to claude's own native tools (`Bash`, `WebFetch`, etc.), not
the parent's `mcp__bash__*` / `mcp__hyperhive__*` surface. Nothing
implicit reaches it: the built-in
hyperhive surface (todos/messaging) isn't an `extraMcpServers` entry at
all, and the automatically injected `bash`/`subagent` entries default to excluded
too (a subagent can't spawn hive-bash tasks or its own nested subagents
unless an operator opts them in explicitly, same as anything else).
Set `hyperhive.extraMcpServers.<name>.availableToSubagents = true` on a
specific entry to hand that one server to subagents as well — useful for,
say, a read-only lookup or scraper MCP a subagent's bounded, single-batch
task might need. `hive-subagent-mcp`'s `mcp_config` module renders the
opted-in subset into its own `--mcp-config` file per turn; an entry left
at the default `false` never appears there.