`goal_reached`/`need_help` took the session name as a tool argument, so identity was an assertion by the caller and the only guard on it was `occupancy()` — "does that name have a turn in flight", which two concurrently running siblings both satisfy for each other. A subagent could stop its sibling's run by naming it. Identity moves into the URL. Each spawned run is minted an unguessable token (`Uuid::new_v4`, the OS CSPRNG), the URL carrying it goes into that one subagent's own `--mcp-config`, and the route resolves it back to a session before dispatching to a handler bound to that session. Neither tool takes a `name` any more: a subagent has no field in which to name a sibling, and a sibling's name — which a brief may well mention — is not a token. One route with a path parameter, not a route per session: the `Router` is built once at startup and subagents come and go for the daemon's whole life. An unminted or revoked token gets a bare 404, the same answer either way, so nothing enumerates. A run's token is revoked when the run ends (`finish_turn`) or when a call never reached a spawn. Two things fall out of that: - the config file becomes one per session. A single shared path was already a race between two `start`s; with a per-session URL in it, the loser would read the winner's identity. - `occupancy()` stops being the identity guard and is gone from the signal path entirely rather than kept "just in case" — a revoked token can't reach it, and it never answered the question it was standing in for. It still backs `status`, which is what it was always actually for. Refs #4403 Refs #4413
241 lines
12 KiB
Markdown
241 lines
12 KiB
Markdown
# Subagent daemon
|
|
|
|
`hive-subagent-daemon` (crate `hive-subagent-mcp`) spawns nested headless
|
|
`claude` sessions on request. Own process, own systemd unit, own MCP
|
|
server (`subagent`, not `hyperhive`) — independent of `hive-bash-daemon`:
|
|
a subagent is a full nested claude process, a materially heavier
|
|
capability than a background shell command, so it gets its own
|
|
deployable/restartable unit rather than living inside the bash daemon.
|
|
|
|
Shipped default-on for every agent — `nix/agent-modules/mcp.nix` injects
|
|
`subagent` into `hyperhive.extraMcpServers` via `lib.mkDefault`
|
|
(`allowedTools = ["*"]`), same as `bash`. Default-on rather than
|
|
unconditional: an `agent.nix` can override or drop the entry, which is
|
|
what `mkDefault` is there for. The operator's own framing: default-on for
|
|
now, a real opt-in capability later.
|
|
|
|
For what the tools do and when an agent should reach for them, see the
|
|
`subagent` MCP server's own tool descriptions and the
|
|
`base:claude-subagents` skill — this page covers the daemon as deployed
|
|
infrastructure, not the agent-facing API.
|
|
|
|
## Tools
|
|
|
|
Served under the `subagent` MCP server (`mcp__subagent__<tool>`): `start`,
|
|
`continue`, `status`, `interrupt`.
|
|
|
|
A second route on the same port serves the two tools a **subagent** calls
|
|
about its own run — `goal_reached` and `need_help`. It isn't part of the
|
|
`subagent` server an agent's own config points at; the daemon writes it
|
|
into each subagent's `--mcp-config` itself, under `subagent_control`. Two
|
|
routes rather than six tools on one, so that being able to report on a run
|
|
never carries the ability to start one: there's no route a subagent holds
|
|
that `start` is reachable from.
|
|
|
|
## A subagent can't name a session, not even its own
|
|
|
|
Neither signal tool takes a session name. **The URL is the identity.** At
|
|
each spawn the daemon mints that run an unguessable token, serves it at
|
|
`/signal/mcp/<token>`, and writes that one URL into that one subagent's own
|
|
`--mcp-config` — a file per session, not a shared one. A request is
|
|
resolved to a session before it's dispatched, and the tools read the
|
|
session off the resolution.
|
|
|
|
A subagent therefore has no field in which to name a sibling, and knowing
|
|
a sibling's name buys nothing: a name isn't a token. Two subagents running
|
|
concurrently can't signal each other — which the earlier shape, a shared
|
|
route plus a `name` argument guarded by a liveness check, allowed.
|
|
|
|
A token that resolves to nothing — never minted, or revoked when its run
|
|
ended — gets a bare **404**, the same answer for every token, so nothing
|
|
about the refusal says whether some other session exists.
|
|
|
|
## State
|
|
|
|
In-memory only: what's running now, where each name's session lives, how
|
|
each name's last turn ended, when each running turn last produced output,
|
|
what each session is working toward, how far through its turn budget it
|
|
is, why its run stopped, where it writes its report, and which signal
|
|
token belongs to it. All of it lives only as long as the daemon process
|
|
does — so a daemon restart invalidates every signal URL it had issued,
|
|
which is the same thing as it having stopped the runs those URLs belonged
|
|
to. A daemon
|
|
restart stops whatever was running rather than adopting it. The durable
|
|
record of a subagent's existence is claude's own on-disk session
|
|
(`hive_claude::SessionStore`), which `continue` reattaches to independent
|
|
of the daemon's own lifetime — a restart loses the _in-flight turn_, not
|
|
the subagent's history.
|
|
|
|
Because the remembered directory goes with the rest of it, a `continue`
|
|
after a restart has to re-supply `dir` when the session lives anywhere
|
|
other than the daemon's own working directory — and a restart is the
|
|
situation you reach for `continue` in most often.
|
|
|
|
## Goals, and turns toward them
|
|
|
|
`start` takes an optional `goal`. Without one a session is a single turn,
|
|
exactly as it always was. With one, the daemon keeps the session going:
|
|
when a turn ends and nothing has said to stop, it starts another turn
|
|
re-prompting the subagent toward that goal, quoting it verbatim and saying
|
|
which turn of the budget this is. `max_turns` caps that, defaulting to
|
|
**5**.
|
|
|
|
`status` reports `Turn N of M` for such a session in every state it
|
|
reaches. Read alongside the last-event age below, it's what separates a
|
|
subagent that's working from one that's wedged from one that's out of
|
|
turns — without `ps` and without opening a file.
|
|
|
|
Four things stop a run, and each is recorded distinctly, reported by
|
|
`status`, and appended to the one todo the daemon pushes when the run ends:
|
|
|
|
- **the turn ended and there was no goal** — the single-turn case;
|
|
- **`goal_reached`**, which the subagent calls itself;
|
|
- **`need_help`**, likewise;
|
|
- **the turn cap**, which says so rather than stopping quietly: the todo
|
|
states that the harness limit was reached and the goal was never
|
|
reported reached, so the work stopped where it had got to.
|
|
|
|
A killed or failed turn ends the run too, and keeps the records it already
|
|
had — see [A killed turn](#a-killed-turn). `interrupt` therefore stops a
|
|
whole goal run, not just the turn in flight.
|
|
|
|
When the session was told where its report goes — `start`'s `report_file`,
|
|
or the path the subagent names when it signals — the stop reason is
|
|
appended to that file as well, so the artifact you were going to read
|
|
anyway also says how the run ended. Nothing is inferred: with no path
|
|
given, no file is touched.
|
|
|
|
## `goal_reached` is a label, not a gate
|
|
|
|
Both signals stop the continuation and **extend** the done message. Extend,
|
|
not replace: the turn's observed end and the reason the run stopped are
|
|
different facts, and the second never stands in for the first.
|
|
|
|
`goal_reached` is **self-reported**, by a subagent that has just been
|
|
re-prompted with "you haven't reached the goal" — which is precisely the
|
|
incentive to claim it. It's the same failure class as a build report
|
|
asserting "done, tests pass": a claim about an artifact, not the artifact.
|
|
Nothing in this daemon treats it as verification, and every surface that
|
|
renders it says so. Read the diff and the gate output regardless.
|
|
|
|
`need_help` is the blocking signal. It stops the run and shows up in
|
|
`status` as its own state — blocked, with the subagent's reason — so a
|
|
parent that polls `status` sees the block without reading anything else.
|
|
`continue` is how you answer it.
|
|
|
|
## Is it working, or is it wedged?
|
|
|
|
`status` reporting **running** says a process is tracked, which a wedged
|
|
subagent satisfies as fully as a busy one. A running answer therefore
|
|
carries the age of that turn's last event too: seconds means it's working,
|
|
an age climbing into the minutes means it's stuck. That one number
|
|
replaces inferring the same thing from `ps` output and CPU-time deltas.
|
|
It resets at each turn's spawn, so on a goal run it describes the turn in
|
|
flight rather than the run — which is what you want, since a run that's
|
|
making progress spends several perfectly healthy minutes.
|
|
|
|
Every line the subagent's `claude` process writes bumps the timestamp —
|
|
stream-json events, plain stdout chatter and stderr alike — and what the
|
|
subagent actually said is never read. The record says the child is alive,
|
|
not what it's doing. The clock starts at the spawn, so a subagent that wedged
|
|
before it ever emitted anything still reports a climbing age rather than
|
|
no age at all. It's dropped when the turn ends, since a finished turn has
|
|
no progress left to describe.
|
|
|
|
## A `continue` that finds no session
|
|
|
|
`continue` doesn't check for the session before spawning. claude's own
|
|
`--resume` is the authority, and it exits non-zero rather than quietly
|
|
starting a fresh session, so the check could only duplicate the lookup
|
|
the driver was about to do — while answering as though the session were
|
|
gone. The usual truth is that the session exists somewhere else.
|
|
|
|
`continue` waits for that answer instead. Where `start` returns the
|
|
instant the process exists — it creates its session, so the spawn
|
|
succeeding is the whole story — a resumed turn can fail a moment _after_
|
|
a pid exists, and a pid is no proof that turn began. `continue` holds the
|
|
tool call open until the turn is underway or the resume has come back
|
|
missed, and reports a miss as the call's own error, carrying claude's
|
|
message plus the location this daemon searched:
|
|
|
|
```
|
|
continue error: claude error: no session matched the requested id or title (searched /home/agent/.claude for cwd /home/agent/work; if it was started elsewhere, pass the `dir` it was started in)
|
|
```
|
|
|
|
The directory is the part claude's own message never names, and the part
|
|
that resolves the confusion: pass `dir` to point `continue` at the
|
|
directory the session was started in.
|
|
|
|
Nothing here is a fixed delay on the way to a successful turn. The wait
|
|
ends on whichever comes first — the turn's first stream event or its
|
|
early exit — and on this box both land inside a second, so a `continue`
|
|
that works answers about as fast as it did before. A five-second cap
|
|
bounds the one case neither covers: a child that neither speaks nor
|
|
exits, reported as started, with the end-of-turn todo left to say how it
|
|
goes. That todo still carries every failure that happens later in the
|
|
turn, exactly as before; the only one it no longer repeats is the miss
|
|
the caller has just been handed to its face.
|
|
|
|
## A killed turn
|
|
|
|
A subagent whose `claude` process dies on a signal — the kernel's OOM
|
|
killer under memory pressure, a stopped unit, an `interrupt` from the
|
|
owning agent — didn't finish its turn, and the daemon says so rather
|
|
than letting it settle back into `idle`:
|
|
|
|
- `status` reports the session **killed**, naming the signal, instead of
|
|
the `idle` it reports for a turn that ended on its own;
|
|
- the end-of-turn todo the daemon pushes without being asked says the
|
|
subagent was killed mid-turn, not that it finished;
|
|
- `continue` still resumes such a session — often what you want — but
|
|
its reply says the previous turn was killed, so nobody carries on from
|
|
cut-off work believing it was complete.
|
|
|
|
The record is per-name, in memory with the rest of this daemon's state,
|
|
and the next confirmed spawn under that name clears it. A daemon restart
|
|
loses it along with everything else — the daemon has to outlive the kill
|
|
to report it, which is what `OOMPolicy=continue` on the unit is for.
|
|
|
|
## Compaction trade-off
|
|
|
|
Built on `hive_claude::Claude::spawn` + `RunningClaude::wait` directly
|
|
rather than `InfiniteSession::run`, since only the low-level driver
|
|
exposes a cancel handle to stop a turn mid-flight — that's what makes
|
|
`interrupt` genuinely stop a running turn rather than only cancelling a
|
|
still-pending one. The cost: a turn that overflows the context window
|
|
surfaces as an error rather than self-healing via reactive compaction.
|
|
Subagents are meant to be bounded, single-batch work, not sessions
|
|
long-lived enough to need in-place compaction — a real follow-up if that
|
|
assumption stops holding.
|
|
|
|
## Configuration
|
|
|
|
`hyperhive.mcp.subagentHttpPort` — the daemon's streamable-http listen
|
|
port. Same pattern as `bashHttpPort`/`matrixHttpPort`: a per-agent default
|
|
assigned by `nix/agent-modules/mcp.nix`, only worth overriding for an
|
|
agent that needs a stable or non-default port.
|
|
|
|
Own systemd unit, defined alongside the other per-agent MCP daemons in
|
|
`nix/agent-modules/mcp.nix`.
|
|
|
|
## MCP servers available to a subagent
|
|
|
|
A subagent runs with `--strict-mcp-config` and, by default, exactly one
|
|
MCP server: the two-tool `subagent_control` route above. It otherwise
|
|
falls back to claude's own native tools (`Bash`, `WebFetch`, etc.), not
|
|
the parent's `mcp__bash__*` / `mcp__hyperhive__*` surface. Nothing
|
|
implicit reaches it: the built-in
|
|
hyperhive surface (todos/messaging) isn't an `extraMcpServers` entry at
|
|
all, and the automatically injected `bash`/`subagent` entries default to excluded
|
|
too (a subagent can't spawn hive-bash tasks or its own nested subagents
|
|
unless an operator opts them in explicitly, same as anything else).
|
|
|
|
Set `hyperhive.extraMcpServers.<name>.availableToSubagents = true` on a
|
|
specific entry to hand that one server to subagents as well — useful for,
|
|
say, a read-only lookup or scraper MCP a subagent's bounded, single-batch
|
|
task might need. `hive-subagent-mcp`'s `mcp_config` module renders the
|
|
opted-in subset into a `--mcp-config` file per session, rewritten each
|
|
turn; an entry left at the default `false` never appears there. Per
|
|
session rather than one shared file, because the `subagent_control` entry
|
|
in it carries that session's own signal URL — one file for everyone would
|
|
be a race over whose identity each subagent reads at startup.
|