`status` could only answer running / starting / idle / killed / none, because every turn ran against `&NoopSink` and the whole stream-json stream was discarded. "Running" describes a wedged subagent exactly as well as a busy one, leaving a caller to tell them apart from `ps` output and CPU-time deltas. So the daemon now keeps a `name -> last_event_at` clock, bumped by `LivenessSink` on every line of every stream — stream-json events, plain stdout chatter and stderr alike — and `status` reports its age on a running answer: a few seconds means working, an age climbing into the minutes with no end-of-turn todo means wedged. Nothing is read out of the content; classifying *what* a subagent is doing is a separate question and waits on its own driver work. In memory with the rest of this daemon's state, dropped when the turn ends, no persistence. The clock is seeded at the spawn rather than at the first line, so a subagent that wedged before emitting anything still reports a climbing age rather than no age at all — the case an age is worth most in. Separately, `continue`'s existence pre-check is gone. It could only repeat the lookup `Claude::spawn` was about to do, and its message — "no session named `x` exists" — was false in the common failure: the session existed, just not under the claude home + cwd `build_store` resolved from. claude's own `--resume` is the authority and exits non-zero (`does not match any session title`) rather than quietly starting a fresh session, so the turn fails on its own. `classify_end` appends the one fact the CLI's message lacks — the directory searched: claude error: no session matched the requested id or title (searched <claude_home> for cwd <cwd>; if the session was started elsewhere, pass `dir`) The `dirs` map's durability is untouched; whether to persist it stays an open operator decision. Module doc, `docs/tools/subagent.md`, the `continue`/`status` tool descriptions and the `base:claude-subagents` skill all updated — including `continue`'s `dir` doc, which said "the daemon remembers it" without saying that a restart is both when it forgets and when you most want it. Refs #4330 Refs #4405
162 lines
8.2 KiB
Markdown
162 lines
8.2 KiB
Markdown
---
|
|
name: claude-subagents
|
|
description: Spin up a short-lived headless `claude` sub-instance to grind through a well-scoped, mechanical batch (bulk relabeling, a repetitive find/replace, a mechanical migration) instead of burning your own context on it inline. Use this when a task has a clear, describable recipe and is either big enough to eat your context or long enough that you'd rather not babysit it. Not for judgement-heavy work, anything needing operator back-and-forth, or a task whose blast radius you can't bound up front.
|
|
---
|
|
|
|
# Ephemeral Sub-Agents
|
|
|
|
A sub-instance you spawn inherits the same filesystem and credentials you
|
|
have. This is a first-class tool for offloading a bounded, mechanical
|
|
batch.
|
|
|
|
## When to use it
|
|
|
|
- **Yes:** mechanical batches with a clear recipe (relabel N issues,
|
|
rewrite a call-site pattern, migrate a config field), especially when
|
|
the batch would eat your context or take long enough that you'd
|
|
rather not babysit it turn-by-turn.
|
|
- **No:** judgement-heavy work, anything needing back-and-forth with a
|
|
human, or a task whose blast radius you can't bound up front.
|
|
|
|
## The `subagent` MCP tools
|
|
|
|
Your container's `subagent` MCP server (`hive-subagent-daemon`) runs this
|
|
skill's recipe for you:
|
|
|
|
```
|
|
start(name, prompt_file, model?, effort?, trigger?)
|
|
continue(name, prompt, model?, effort?)
|
|
status(name)
|
|
interrupt(name, force?)
|
|
```
|
|
|
|
`start` and `continue` return as soon as the process is confirmed
|
|
running, not once it finishes — a completion lands as a todo
|
|
(`get_loose_ends`), same as any other producer. Use `status` for a
|
|
zero-cost "is it still going" check — for a running subagent it also
|
|
reports how long since that turn last produced output, so a few seconds
|
|
means it's working and an age climbing into the minutes means it's
|
|
wedged and worth an `interrupt`. Reach for `continue` only once you
|
|
actually have a new instruction for it, since that spends a turn.
|
|
`interrupt` genuinely stops a running turn (`force: true` for SIGKILL).
|
|
|
|
Model choice, prompt hygiene, splitting big batches, verify-then-report —
|
|
everything else in this skill — applies exactly the same whether you're
|
|
calling the tool or thinking through the recipe by hand.
|
|
|
|
## Effort
|
|
|
|
An omitted `effort` defaults to `medium` here — cheaper than claude's own
|
|
model default (`high` on most models), a deliberate cost-conscious choice
|
|
for subagent work specifically, same spirit as "cheaper-than-you" model
|
|
choice above. Raise it explicitly when the task is complex enough to
|
|
actually need deeper reasoning, not as a reflex.
|
|
|
|
Anthropic's own levels (per current Claude Code docs — names/availability
|
|
are model-dependent, check before relying on an exact list): `low`,
|
|
`medium`, `high`, `xhigh`, `max`. Each trades token spend for capability:
|
|
|
|
- **`low`** — short, scoped, latency-sensitive tasks that aren't
|
|
intelligence-sensitive.
|
|
- **`medium`** — cost-sensitive work that can trade off some intelligence.
|
|
This skill's default for subagent work.
|
|
- **`high`** — balances token usage and intelligence; Anthropic's own
|
|
recommended default for most _interactive_ coding tasks (not what this
|
|
skill defaults subagents to — see above).
|
|
- **`xhigh`** — deeper reasoning at higher token spend.
|
|
- **`max`** — demanding tasks only; diminishing returns and overthinking
|
|
are a real risk, per Anthropic's own guidance — don't reach for it as a
|
|
default.
|
|
|
|
Anthropic's guidance: treat effort as a general preference, not a
|
|
task-by-task dial — raise it if a subagent keeps skipping files, not
|
|
running tests, or not double-checking its own work; lower it for routine
|
|
work where quality hasn't suffered. Changing effort between turns on the
|
|
_same_ session invalidates prompt caching, so pick a level for the whole
|
|
session rather than flipping it turn-to-turn.
|
|
|
|
## Prompt hygiene - this is where batches succeed or fail
|
|
|
|
- **Concrete constants, not "figure it out":** exact ids, field names,
|
|
option values, endpoints.
|
|
- **Critical ordering rules, spelled out** - the sub-agent won't infer
|
|
invariants you don't state.
|
|
- **Explicit scope + exemptions**, plus a fallback rule for ambiguous
|
|
cases: "leave it as-is where genuinely unclear; do not guess."
|
|
- **Ask for a report file** - per-item results + anything skipped and
|
|
why, so you can verify without re-deriving.
|
|
- **Tune on ONE item first**, eyeball the result, fix the prompt, _then_
|
|
turn it loose on the full batch. A prompt bug replicated across 100
|
|
items is 100 cleanups.
|
|
|
|
## Split a big independent batch across parallel subagents
|
|
|
|
One subagent is a sequential worker — a 50-item batch takes roughly 50
|
|
items' worth of wall-clock even though nothing about the recipe forces
|
|
serialization. If the items are independent (each one touches its own
|
|
file/issue/row, nothing depends on another item's result) and the batch
|
|
is big enough that wall-clock matters, chunk it across N subagents
|
|
running at once instead of handing the whole thing to one.
|
|
|
|
- **Split by natural boundaries**, not an arbitrary item count — one
|
|
subagent per file, per directory, per module, per label, whatever
|
|
grouping keeps each worker's slice self-contained. A worker that has
|
|
to coordinate with another worker mid-task isn't actually independent
|
|
work; re-scope the split until it is.
|
|
- **Isolate each worker's writes.** For a git-based batch, give each
|
|
worker its own `git worktree` (own working directory, own branch, same
|
|
underlying repo — cheap, no full reclone) so N workers editing
|
|
different files never race on the same working tree or index. For a
|
|
non-git batch (issue relabeling, API calls), independent items don't
|
|
need filesystem isolation at all — just launch N in parallel.
|
|
- **Launch all N and don't babysit any single one** — same as the
|
|
one-subagent case, just N `start` calls instead of one. Check on them
|
|
as a batch, not by polling each individually in a loop.
|
|
- **Mind the container's memory cap before picking N.** Your whole
|
|
container shares one `MemoryMax` (a few GB by default) with every
|
|
subagent you spawn _and_ your own process. A `claude` process plus its
|
|
MCP servers can hold several hundred MB to ~1GB depending on the task;
|
|
spawning a dozen at once on a small container doesn't just slow
|
|
things down, it can OOM the whole container — taking your own
|
|
in-flight turn down with it, not just the subagents. Rule of thumb:
|
|
**2-4 concurrent workers** on a default-sized container; check
|
|
`free -h` (or ask whoever owns the container's config for its
|
|
`MemoryMax`) before going higher, and chunk a bigger batch into
|
|
successive waves of that size rather than firing everything at once.
|
|
- **You own the merge.** Once all N report done, review + verify each
|
|
worker's slice (same "don't trust the self-report blind" rule as
|
|
below), then combine — for a git-based split, that's you merging N
|
|
branches (or cherry-picking) into one, not each worker pushing/PRing
|
|
its own slice.
|
|
|
|
Skip this for a batch small enough to finish in a couple minutes single-
|
|
threaded — the coordination overhead isn't worth it below that size.
|
|
|
|
## Resume for follow-ups
|
|
|
|
`continue` reuses the same session, so the sub-agent keeps every constant
|
|
and gotcha it already discovered instead of re-learning the surface from
|
|
a cold prompt. Reach for `start` under a new name only for a genuinely
|
|
unrelated task.
|
|
|
|
## Verify, then report
|
|
|
|
On completion: read the report file, spot-check a handful of results
|
|
yourself (don't trust the self-report blind), then summarize. Surface
|
|
anything the sub-agent left ambiguous or exempted for a human call.
|
|
|
|
## Pitfalls (all observed in practice)
|
|
|
|
- Using a model bigger than yourself → paying premium cost for
|
|
mechanical work that didn't need it.
|
|
- Skipping the one-item tuning pass → a systematic mistake smeared
|
|
across the whole batch.
|
|
- `start`-ing fresh instead of `continue`-ing for a follow-up → throws
|
|
away all the context the first pass earned.
|
|
- Baking an unverified assumption into the recipe - if a quick check
|
|
suggests something "isn't possible" or "doesn't exist," confirm it
|
|
before writing that conclusion into the prompt; a wrong assumption
|
|
gets replicated across the whole batch.
|
|
- Handing a big independent batch to one subagent instead of splitting
|
|
it across several in parallel - a 50-item sequential run burns wall-
|
|
clock the split-by-worktree approach above would've avoided for free.
|