subagent: give a run a goal, turns toward it, and a reason it stopped
`start` takes an optional `goal`. With one set a session stops being a single turn: when a turn ends and nothing has said to stop, the daemon spawns another turn re-prompting the subagent toward that goal, up to `max_turns` (default 5, per-session). Without a goal nothing changes — one turn, one todo, same as before. Four things end a run, each recorded distinctly and reported by `status`: the turn ending with no goal, `goal_reached`, `need_help`, and the turn cap. The last says so out loud rather than stopping quietly — the todo states the harness limit was reached and the goal was never reported reached. Every stop extends the done message rather than replacing it, and lands in the session's report file when it has one. The path is never inferred: it comes from `start`'s `report_file` or from the subagent naming where it wrote. `goal_reached` and `need_help` are the subagent's own, served on a second route (`/signal/mcp`) that carries those two tools and nothing else, so reporting on a run can't become starting one. `goal_reached` is built as a label, never a gate: it is self-reported by a subagent that has just been re-prompted with "you haven't reached the goal", which is exactly the incentive to claim it — the same failure class as a build report asserting the tests pass. Every surface that renders it says so. `need_help` is the blocking signal, and shows in `status` as its own state so a parent polling it sees the block without reading a file. `status` also carries `turn N of M`: with 4330's last-event age, that separates working from wedged from out of turns off one answer. Two bugs the new tests caught: a `tokio::fs::File` was dropped without flushing, so the report line was written to nothing, and the plain idle answer dropped the turn counter. Also documents `await_resume`'s third case — a closed channel with no send, which fails open the same as `Underway` — per argus on #4411. Refs #4403
This commit is contained in:
parent
6e2de33f26
commit
b18348bc9a
10 changed files with 1758 additions and 253 deletions
14
CLAUDE.md
14
CLAUDE.md
|
|
@ -101,12 +101,14 @@ hand-maintained per-file tree drifts out of sync with the code.
|
|||
- **`hive-subagent-mcp/`** — per-agent claude-subagent runner daemon
|
||||
(`hive-subagent-daemon`); spawns nested claude sessions on request and
|
||||
serves `start`/`continue`/`status`/`interrupt` directly over
|
||||
streamable-http (no stdio bridge). Independent of `hive-bash-mcp` (a
|
||||
subagent is a much heavier capability than a bash command). No task
|
||||
files — the daemon's only state is an in-memory map of currently-running
|
||||
processes, live only as long as the process is; the actual claude
|
||||
session survives a daemon restart independently (see `session.rs`'s
|
||||
module doc).
|
||||
streamable-http (no stdio bridge), plus a second subagent-facing route
|
||||
carrying `goal_reached`/`need_help`. Independent of `hive-bash-mcp` (a
|
||||
subagent is a much heavier capability than a bash command). A `start`
|
||||
with a `goal` is a multi-turn run: the daemon re-prompts the subagent
|
||||
toward the goal each time a turn ends, up to a per-session turn cap. No
|
||||
task files — the daemon's state is in-memory only, live only as long as
|
||||
the process is; the actual claude session survives a daemon restart
|
||||
independently (see `session.rs`'s module doc).
|
||||
- **`hive-sh4re/`** — shared wire types (Agent / Manager request +
|
||||
response, `Message`, `Approval`, `HelperEvent`) used across the unix
|
||||
sockets. Host-admin-socket and hive-priv-socket wire types have been
|
||||
|
|
|
|||
|
|
@ -24,11 +24,21 @@ infrastructure, not the agent-facing API.
|
|||
Served under the `subagent` MCP server (`mcp__subagent__<tool>`): `start`,
|
||||
`continue`, `status`, `interrupt`.
|
||||
|
||||
A second route on the same port serves the two tools a **subagent** calls
|
||||
about its own run — `goal_reached` and `need_help`. It isn't part of the
|
||||
`subagent` server an agent's own config points at; the daemon writes it
|
||||
into each subagent's `--mcp-config` itself, under `subagent_control`. Two
|
||||
routes rather than six tools on one, so that being able to report on a run
|
||||
never carries the ability to start one: there's no route a subagent holds
|
||||
that `start` is reachable from.
|
||||
|
||||
## State
|
||||
|
||||
In-memory only: what's running now, where each name's session lives, how
|
||||
each name's last turn ended, and when each running turn last produced
|
||||
output. All of it lives only as long as the daemon process does. A daemon
|
||||
each name's last turn ended, when each running turn last produced output,
|
||||
what each session is working toward, how far through its turn budget it
|
||||
is, why its run stopped, and where it writes its report. All of it lives
|
||||
only as long as the daemon process does. A daemon
|
||||
restart stops whatever was running rather than adopting it. The durable
|
||||
record of a subagent's existence is claude's own on-disk session
|
||||
(`hive_claude::SessionStore`), which `continue` reattaches to independent
|
||||
|
|
@ -40,14 +50,68 @@ after a restart has to re-supply `dir` when the session lives anywhere
|
|||
other than the daemon's own working directory — and a restart is the
|
||||
situation you reach for `continue` in most often.
|
||||
|
||||
## Goals, and turns toward them
|
||||
|
||||
`start` takes an optional `goal`. Without one a session is a single turn,
|
||||
exactly as it always was. With one, the daemon keeps the session going:
|
||||
when a turn ends and nothing has said to stop, it starts another turn
|
||||
re-prompting the subagent toward that goal, quoting it verbatim and saying
|
||||
which turn of the budget this is. `max_turns` caps that, defaulting to
|
||||
**5**.
|
||||
|
||||
`status` reports `Turn N of M` for such a session in every state it
|
||||
reaches. Read alongside the last-event age below, it's what separates a
|
||||
subagent that's working from one that's wedged from one that's out of
|
||||
turns — without `ps` and without opening a file.
|
||||
|
||||
Four things stop a run, and each is recorded distinctly, reported by
|
||||
`status`, and appended to the one todo the daemon pushes when the run ends:
|
||||
|
||||
- **the turn ended and there was no goal** — the single-turn case;
|
||||
- **`goal_reached`**, which the subagent calls itself;
|
||||
- **`need_help`**, likewise;
|
||||
- **the turn cap**, which says so rather than stopping quietly: the todo
|
||||
states that the harness limit was reached and the goal was never
|
||||
reported reached, so the work stopped where it had got to.
|
||||
|
||||
A killed or failed turn ends the run too, and keeps the records it already
|
||||
had — see [A killed turn](#a-killed-turn). `interrupt` therefore stops a
|
||||
whole goal run, not just the turn in flight.
|
||||
|
||||
When the session was told where its report goes — `start`'s `report_file`,
|
||||
or the path the subagent names when it signals — the stop reason is
|
||||
appended to that file as well, so the artifact you were going to read
|
||||
anyway also says how the run ended. Nothing is inferred: with no path
|
||||
given, no file is touched.
|
||||
|
||||
## `goal_reached` is a label, not a gate
|
||||
|
||||
Both signals stop the continuation and **extend** the done message. Extend,
|
||||
not replace: the turn's observed end and the reason the run stopped are
|
||||
different facts, and the second never stands in for the first.
|
||||
|
||||
`goal_reached` is **self-reported**, by a subagent that has just been
|
||||
re-prompted with "you haven't reached the goal" — which is precisely the
|
||||
incentive to claim it. It's the same failure class as a build report
|
||||
asserting "done, tests pass": a claim about an artifact, not the artifact.
|
||||
Nothing in this daemon treats it as verification, and every surface that
|
||||
renders it says so. Read the diff and the gate output regardless.
|
||||
|
||||
`need_help` is the blocking signal. It stops the run and shows up in
|
||||
`status` as its own state — blocked, with the subagent's reason — so a
|
||||
parent that polls `status` sees the block without reading anything else.
|
||||
`continue` is how you answer it.
|
||||
|
||||
## Is it working, or is it wedged?
|
||||
|
||||
`status` reporting **running** says a process is tracked, which a wedged
|
||||
subagent satisfies as fully as a busy one. A running answer therefore
|
||||
carries the age of that turn's last event too: seconds means it's working,
|
||||
an age climbing into the minutes with no end-of-turn todo means it's
|
||||
stuck. That one number replaces inferring the same thing from `ps` output
|
||||
and CPU-time deltas.
|
||||
an age climbing into the minutes means it's stuck. That one number
|
||||
replaces inferring the same thing from `ps` output and CPU-time deltas.
|
||||
It resets at each turn's spawn, so on a goal run it describes the turn in
|
||||
flight rather than the run — which is what you want, since a run that's
|
||||
making progress spends several perfectly healthy minutes.
|
||||
|
||||
Every line the subagent's `claude` process writes bumps the timestamp —
|
||||
stream-json events, plain stdout chatter and stderr alike — and what the
|
||||
|
|
@ -135,10 +199,11 @@ Own systemd unit, defined alongside the other per-agent MCP daemons in
|
|||
|
||||
## MCP servers available to a subagent
|
||||
|
||||
A subagent runs with `--strict-mcp-config` and no `--mcp-config` by
|
||||
default — zero MCP servers, full stop; it falls back to claude's own
|
||||
native tools (`Bash`, `WebFetch`, etc.), not the parent's `mcp__bash__*` /
|
||||
`mcp__hyperhive__*` surface. Nothing implicit reaches it: the built-in
|
||||
A subagent runs with `--strict-mcp-config` and, by default, exactly one
|
||||
MCP server: the two-tool `subagent_control` route above. It otherwise
|
||||
falls back to claude's own native tools (`Bash`, `WebFetch`, etc.), not
|
||||
the parent's `mcp__bash__*` / `mcp__hyperhive__*` surface. Nothing
|
||||
implicit reaches it: the built-in
|
||||
hyperhive surface (todos/messaging) isn't an `extraMcpServers` entry at
|
||||
all, and the automatically injected `bash`/`subagent` entries default to excluded
|
||||
too (a subagent can't spawn hive-bash tasks or its own nested subagents
|
||||
|
|
|
|||
|
|
@ -33,7 +33,9 @@ tokio = { workspace = true, features = ["test-util"] }
|
|||
# `hive-subagent-daemon` — long-running per-agent claude-subagent runner.
|
||||
# Independent of `hive-bash-mcp` (own crate, own binary, own MCP server) —
|
||||
# see lib.rs's module doc for why. Serves its MCP tools (`start`/
|
||||
# `continue`/`interrupt`) directly over streamable-http — no stdio bridge.
|
||||
# `continue`/`status`/`interrupt`, plus the subagent-facing
|
||||
# `goal_reached`/`need_help` route) directly over streamable-http — no
|
||||
# stdio bridge.
|
||||
[[bin]]
|
||||
name = "hive-subagent-daemon"
|
||||
path = "src/main.rs"
|
||||
|
|
|
|||
|
|
@ -2,9 +2,10 @@
|
|||
|
||||
Per-agent daemon (`hive-subagent-daemon`) that spawns nested headless
|
||||
`claude` sessions on request and serves the tool surface
|
||||
(`start`/`continue`/`interrupt`) directly over streamable-http. No
|
||||
stdio bridge, no per-turn respawn — claude reconnects to the same
|
||||
stable URL every turn.
|
||||
(`start`/`continue`/`status`/`interrupt`, plus a separate
|
||||
subagent-facing `goal_reached`/`need_help` route) directly over
|
||||
streamable-http. No stdio bridge, no per-turn respawn — claude
|
||||
reconnects to the same stable URL every turn.
|
||||
|
||||
Independent of `hive-bash-mcp` — a subagent spawns a full nested
|
||||
`claude` session, a much heavier capability than a bash command, worth
|
||||
|
|
@ -18,10 +19,12 @@ own lib (`src/lib.rs`):
|
|||
- **`session.rs`** — the actual claude-facing logic: `Claude::spawn` +
|
||||
`RunningClaude::wait`/`cancel_handle` (not `InfiniteSession::run`,
|
||||
which has no cancel handle to reach in — see the module doc for the
|
||||
v1 scope this trades away), the in-memory `running` map that's the
|
||||
_only_ state this daemon keeps (no task files — a restart stops
|
||||
whatever's running; the actual claude session is the durable store,
|
||||
found again by name via `hive_claude::SessionStore`).
|
||||
- **`mcp.rs`** — the `rmcp` tool router (`start`/`continue`/`interrupt`)
|
||||
- `serve_http`.
|
||||
v1 scope this trades away), the turn-continuation loop a `goal`
|
||||
switches on, and the in-memory maps that are the _only_ state this
|
||||
daemon keeps (no task files — a restart stops whatever's running;
|
||||
the actual claude session is the durable store, found again by name
|
||||
via `hive_claude::SessionStore`).
|
||||
- **`mcp.rs`** — the `rmcp` tool routers (the parent's
|
||||
`start`/`continue`/`status`/`interrupt` on `/mcp`, the subagent's
|
||||
`goal_reached`/`need_help` on `/signal/mcp`) + `serve_http`.
|
||||
- **`paths.rs`** — the in-agent todo-socket path.
|
||||
|
|
|
|||
|
|
@ -1,13 +1,18 @@
|
|||
//! Library for `hive-subagent-daemon`: spawns nested claude sessions on
|
||||
//! request and serves the `start`/`continue`/`interrupt` MCP tool surface
|
||||
//! directly over streamable-http — no stdio bridge, no round-trip socket.
|
||||
//! Independent of `hive-bash-mcp` — a subagent is a much heavier capability
|
||||
//! than a bash command (a full nested `claude` process), worth its own
|
||||
//! deployable/restartable unit rather than sharing one.
|
||||
//! request and serves the `start`/`continue`/`status`/`interrupt` MCP tool
|
||||
//! surface directly over streamable-http — no stdio bridge, no round-trip
|
||||
//! socket. Independent of `hive-bash-mcp` — a subagent is a much heavier
|
||||
//! capability than a bash command (a full nested `claude` process), worth
|
||||
//! its own deployable/restartable unit rather than sharing one.
|
||||
//!
|
||||
//! A second route on the same listener serves `goal_reached`/`need_help` to
|
||||
//! the *subagents*, which is how a run says it's done or stuck; see
|
||||
//! [`mcp`]'s module doc for why that is a separate surface rather than two
|
||||
//! more tools on the parent's.
|
||||
//!
|
||||
//! See [`session`]'s module doc for the actual design: no task files, no
|
||||
//! restart recovery, no mid-turn compaction — the daemon's only state is an
|
||||
//! in-memory `name -> Cancel` map, live only as long as the process is.
|
||||
//! restart recovery, no mid-turn compaction — the daemon's only state is a
|
||||
//! handful of in-memory maps, live only as long as the process is.
|
||||
|
||||
pub mod mcp;
|
||||
pub mod mcp_config;
|
||||
|
|
|
|||
|
|
@ -1,7 +1,9 @@
|
|||
//! `hive-subagent-daemon` binary — spawns nested claude sessions on
|
||||
//! request and serves the `start`/`continue`/`interrupt` MCP tool surface
|
||||
//! directly over streamable-http on `--http <addr>` — no stdio bridge, no
|
||||
//! separate bin claude has to respawn every turn.
|
||||
//! request and serves the `start`/`continue`/`status`/`interrupt` MCP tool
|
||||
//! surface directly over streamable-http on `--http <addr>` — no stdio
|
||||
//! bridge, no separate bin claude has to respawn every turn. The same
|
||||
//! address also carries the subagent-facing `goal_reached`/`need_help`
|
||||
//! route, which is the only place the daemon can learn its own URL from.
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
|
|
@ -42,7 +44,14 @@ async fn main() -> Result<()> {
|
|||
"hive-subagent-daemon starting"
|
||||
);
|
||||
|
||||
let state = Arc::new(hive_subagent_mcp::session::State::new(todo_socket));
|
||||
// The one place the subagent-facing signal URL can come from: the
|
||||
// address this process was told to listen on. Anything else would be a
|
||||
// guess at the deployment's own port assignment.
|
||||
let signal_url = format!("http://{}{}", cli.http, hive_subagent_mcp::mcp::SIGNAL_PATH);
|
||||
let state = Arc::new(hive_subagent_mcp::session::State::new(
|
||||
todo_socket,
|
||||
signal_url,
|
||||
));
|
||||
|
||||
// Serve the MCP tools over streamable-http forever. No background poll
|
||||
// loop to start — unlike the bash daemon's task-file queue, `start`/
|
||||
|
|
|
|||
|
|
@ -1,6 +1,12 @@
|
|||
//! The MCP tool surface: `start` / `continue` / `status` / `interrupt`,
|
||||
//! served directly over streamable-http — no stdio bridge, no round-trip
|
||||
//! socket.
|
||||
//!
|
||||
//! Two surfaces, two routes. `/mcp` is the parent's: the four tools above.
|
||||
//! `/signal/mcp` is the *subagent's*, and carries `goal_reached` and
|
||||
//! `need_help` only — it is what a subagent is handed in its own
|
||||
//! `--mcp-config` (see [`crate::mcp_config`]), so the ability to report on
|
||||
//! its own run can't be the ability to spawn a nested one.
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
|
|
@ -57,12 +63,64 @@ struct StartArgs {
|
|||
/// this daemon tracks in memory.
|
||||
#[serde(default)]
|
||||
dir: Option<String>,
|
||||
/// What this subagent is working *toward*, in its own words — set it and
|
||||
/// the daemon keeps giving it turns until it says it's done, says it's
|
||||
/// stuck, or runs out. Omit it and the session is a single turn, exactly
|
||||
/// as before. Written for the subagent to read: it's quoted back at it
|
||||
/// verbatim at the start of every continued turn, so "get `cargo clippy
|
||||
/// --workspace` to pass with no warnings" continues far better than "fix
|
||||
/// the lints".
|
||||
#[serde(default)]
|
||||
goal: Option<String>,
|
||||
/// How many turns the continuation may spend before the harness stops it
|
||||
/// itself. Default 5. Only meaningful alongside `goal` — without one
|
||||
/// there's nothing to re-prompt toward, so nothing to cap. The cap
|
||||
/// bounds *unattended* re-prompting: a `continue` you issue yourself
|
||||
/// starts the allowance over.
|
||||
#[serde(default)]
|
||||
max_turns: Option<u32>,
|
||||
/// Where the instructions in `prompt_file` told this subagent to write
|
||||
/// its report. The daemon appends the run's stop reason to that file when
|
||||
/// the run ends, so the artifact you were going to read anyway also says
|
||||
/// how it stopped. Nothing is inferred: if you don't pass it (and the
|
||||
/// subagent doesn't name it when it signals), no file is touched.
|
||||
#[serde(default)]
|
||||
report_file: Option<String>,
|
||||
}
|
||||
|
||||
fn default_trigger() -> String {
|
||||
"Carry out the task described in your instructions.".to_owned()
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize, JsonSchema)]
|
||||
struct GoalReachedArgs {
|
||||
/// Your own session name — the one your brief and your continuation
|
||||
/// prompts address you by.
|
||||
name: String,
|
||||
/// Optionally, what you did. It's shown to whoever spawned you.
|
||||
#[serde(default)]
|
||||
msg: Option<String>,
|
||||
/// Optionally, the path you wrote your report to, so the stop reason
|
||||
/// gets appended to it.
|
||||
#[serde(default)]
|
||||
report_file: Option<String>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize, JsonSchema)]
|
||||
struct NeedHelpArgs {
|
||||
/// Your own session name — the one your brief and your continuation
|
||||
/// prompts address you by.
|
||||
name: String,
|
||||
/// What is blocking you, specifically enough for someone else to act on
|
||||
/// it. This is the whole content of the signal, which is why it's
|
||||
/// required.
|
||||
msg: String,
|
||||
/// Optionally, the path you wrote your report to, so the stop reason
|
||||
/// gets appended to it.
|
||||
#[serde(default)]
|
||||
report_file: Option<String>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize, JsonSchema)]
|
||||
struct ContinueArgs {
|
||||
/// The existing session's name (from a prior `start`).
|
||||
|
|
@ -127,17 +185,26 @@ impl SubagentMcp {
|
|||
first (real fresh start, not a silent resume); a *currently running* one is \
|
||||
refused. Runs unattended — every tool-call permission prompt is pre-approved rather \
|
||||
than interactively confirmed — with its MCP server set fixed to what this daemon \
|
||||
configures for it. See the `base:claude-subagents` skill for when to reach for this."
|
||||
configures for it. Pass `goal` to make this a multi-turn run: the daemon re-prompts \
|
||||
the subagent toward that goal each time a turn ends, up to `max_turns` (default 5), \
|
||||
stopping early when the subagent reports the goal reached or asks for help. Whichever \
|
||||
way it stops, one todo is pushed at the end and `status` says which. See the \
|
||||
`base:claude-subagents` skill for when to reach for this."
|
||||
)]
|
||||
fn start(&self, Parameters(args): Parameters<StartArgs>) -> String {
|
||||
match session::start(
|
||||
&self.state,
|
||||
&args.name,
|
||||
args.model,
|
||||
args.effort,
|
||||
&args.prompt_file,
|
||||
args.trigger,
|
||||
args.dir.as_deref(),
|
||||
session::StartRequest {
|
||||
name: args.name,
|
||||
model: args.model,
|
||||
effort: args.effort,
|
||||
prompt_file: args.prompt_file,
|
||||
trigger: args.trigger,
|
||||
dir: args.dir,
|
||||
goal: args.goal,
|
||||
max_turns: args.max_turns,
|
||||
report_file: args.report_file,
|
||||
},
|
||||
) {
|
||||
Ok(msg) => msg,
|
||||
Err(e) => format!("start error: {e:#}"),
|
||||
|
|
@ -193,7 +260,9 @@ impl SubagentMcp {
|
|||
never launches a process, unlike `continue`. The answer says what state it found \
|
||||
and what to do about it. For a running one it also reports how long since that turn \
|
||||
last produced any output, which is how you tell a subagent that's working from one \
|
||||
that has wedged without resorting to `ps`."
|
||||
that has wedged without resorting to `ps`. For a session started with a `goal` it \
|
||||
reports which turn of the budget it is on, and — once the run has stopped — why it \
|
||||
stopped: goal reported reached, blocked and needing help, or out of turns."
|
||||
)]
|
||||
fn status(&self, Parameters(args): Parameters<StatusArgs>) -> String {
|
||||
match session::status(&self.state, &args.name, args.dir.as_deref()) {
|
||||
|
|
@ -206,10 +275,76 @@ impl SubagentMcp {
|
|||
#[tool_handler]
|
||||
impl ServerHandler for SubagentMcp {}
|
||||
|
||||
/// The surface a *subagent* gets, served on its own route: two tools, both
|
||||
/// of them ways for a running subagent to say how its own run should end.
|
||||
/// Separate handler rather than two more tools on [`SubagentMcp`] so the
|
||||
/// split is structural — there is no route a subagent holds that `start` is
|
||||
/// reachable from.
|
||||
#[derive(Clone)]
|
||||
struct SubagentSignalMcp {
|
||||
state: Arc<State>,
|
||||
}
|
||||
|
||||
#[tool_router]
|
||||
impl SubagentSignalMcp {
|
||||
#[tool(
|
||||
description = "Report that you have reached the goal you were given. Stops the harness \
|
||||
from starting another turn to re-prompt you toward it, and extends the \"subagent \
|
||||
done\" message your parent gets with what you say here. This records a claim, not a \
|
||||
result: whoever spawned you reads the diff and the gate output regardless, so \
|
||||
calling it does not make unfinished work finished. Call it when the goal is actually \
|
||||
met — otherwise keep working, or call `need_help` if you can't proceed."
|
||||
)]
|
||||
fn goal_reached(&self, Parameters(args): Parameters<GoalReachedArgs>) -> String {
|
||||
match session::goal_reached(
|
||||
&self.state,
|
||||
&args.name,
|
||||
args.msg,
|
||||
args.report_file.as_deref(),
|
||||
) {
|
||||
Ok(msg) => msg,
|
||||
Err(e) => format!("goal_reached error: {e:#}"),
|
||||
}
|
||||
}
|
||||
|
||||
#[tool(
|
||||
description = "Report that you cannot proceed, and why. Stops the harness from starting \
|
||||
another turn to re-prompt you toward your goal, marks your session as blocked so \
|
||||
whoever spawned you sees it in `status` without reading any file, and extends the \
|
||||
\"subagent done\" message with your reason. Use it for a genuine block — a missing \
|
||||
credential, a decision that isn't yours, an instruction that contradicts what you \
|
||||
found — not for work that is merely hard. Say enough that someone else can act on it."
|
||||
)]
|
||||
fn need_help(&self, Parameters(args): Parameters<NeedHelpArgs>) -> String {
|
||||
match session::need_help(
|
||||
&self.state,
|
||||
&args.name,
|
||||
args.msg,
|
||||
args.report_file.as_deref(),
|
||||
) {
|
||||
Ok(msg) => msg,
|
||||
Err(e) => format!("need_help error: {e:#}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[tool_handler]
|
||||
impl ServerHandler for SubagentSignalMcp {}
|
||||
|
||||
/// Path the subagent-facing signal surface is served at, and the tail of the
|
||||
/// URL [`crate::session::State`] hands to every subagent it spawns. Kept
|
||||
/// here, next to the route that answers it, so the two can't drift.
|
||||
pub const SIGNAL_PATH: &str = "/signal/mcp";
|
||||
|
||||
/// Run the MCP server over HTTP (rmcp streamable-http transport) on `addr`.
|
||||
/// Loopback-only bind, one long-lived session — same shape as the bash and
|
||||
/// matrix daemons' own `serve_http`.
|
||||
///
|
||||
/// Two routes off one listener: `/mcp` for the parent's four tools, and
|
||||
/// [`SIGNAL_PATH`] for the subagent's two. Separate session managers because
|
||||
/// they're separate MCP servers to separate clients — the parent's harness
|
||||
/// on one, each subagent's own claude on the other.
|
||||
///
|
||||
/// # Errors
|
||||
///
|
||||
/// Returns an error if the listener cannot bind `addr` or the HTTP server
|
||||
|
|
@ -218,23 +353,41 @@ pub async fn serve_http(addr: std::net::SocketAddr, state: Arc<State>) -> anyhow
|
|||
use rmcp::transport::streamable_http_server::{
|
||||
StreamableHttpServerConfig, StreamableHttpService, session::local::LocalSessionManager,
|
||||
};
|
||||
let mut session_manager = LocalSessionManager::default();
|
||||
// A subagent turn can run considerably longer than a bash command —
|
||||
// same 24h keep-alive rationale as the bash/matrix daemons.
|
||||
let manager = || {
|
||||
let mut session_manager = LocalSessionManager::default();
|
||||
session_manager.session_config.keep_alive = Some(std::time::Duration::from_hours(24));
|
||||
let session_manager = std::sync::Arc::new(session_manager);
|
||||
std::sync::Arc::new(session_manager)
|
||||
};
|
||||
let parent_state = Arc::clone(&state);
|
||||
let service = StreamableHttpService::new(
|
||||
move || {
|
||||
Ok(SubagentMcp {
|
||||
state: Arc::clone(&parent_state),
|
||||
})
|
||||
},
|
||||
manager(),
|
||||
StreamableHttpServerConfig::default(),
|
||||
);
|
||||
let signal_service = StreamableHttpService::new(
|
||||
move || {
|
||||
Ok(SubagentSignalMcp {
|
||||
state: Arc::clone(&state),
|
||||
})
|
||||
},
|
||||
session_manager,
|
||||
manager(),
|
||||
StreamableHttpServerConfig::default(),
|
||||
);
|
||||
let app = axum::Router::new().nest_service("/mcp", service);
|
||||
let app = axum::Router::new()
|
||||
.nest_service("/mcp", service)
|
||||
.nest_service(SIGNAL_PATH, signal_service);
|
||||
let listener = tokio::net::TcpListener::bind(addr).await?;
|
||||
tracing::info!(%addr, "serving hive-subagent MCP over streamable-http at /mcp");
|
||||
tracing::info!(
|
||||
%addr,
|
||||
signal = SIGNAL_PATH,
|
||||
"serving hive-subagent MCP over streamable-http at /mcp"
|
||||
);
|
||||
axum::serve(listener, app).await?;
|
||||
Ok(())
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,35 +1,61 @@
|
|||
//! Builds a subagent's own filtered `--mcp-config`: only
|
||||
//! `hyperhive.extraMcpServers` entries with `availableToSubagents = true`
|
||||
//! (see `hive_agent_sock::extra_mcp::ExtraMcpServer::available_to_subagents`) ever reach
|
||||
//! a subagent's claude invocation. Everything else — the built-in hyperhive
|
||||
//! a subagent's claude invocation, plus this daemon's own two-tool signal
|
||||
//! surface. Everything else — the built-in hyperhive
|
||||
//! surface (todos/messaging), the automatically injected `bash` and `subagent`
|
||||
//! entries — stays unreachable by construction: none of those default to
|
||||
//! opted in, and the built-in surface isn't an `extraMcpServers` entry at
|
||||
//! all, so there's no name for an operator to opt it in under even if they
|
||||
//! wanted to.
|
||||
//!
|
||||
//! The signal surface is the one server a subagent always gets. It is a
|
||||
//! *different route* on this daemon's listener from the one the parent
|
||||
//! uses, serving `goal_reached` and `need_help` and nothing else — so
|
||||
//! "a subagent can say it is done or stuck" never widens into "a subagent
|
||||
//! can spawn subagents", which is what handing it the parent's route would
|
||||
//! have meant.
|
||||
|
||||
use std::path::PathBuf;
|
||||
|
||||
/// Filename the rendered config lives at, under [`crate::paths::harness_dir`].
|
||||
const CONFIG_FILE: &str = "subagent-mcp-config.json";
|
||||
|
||||
/// Render the subagent-eligible extra-MCP servers to a `--mcp-config` file,
|
||||
/// returning its path — or `None` when no entry opts in (or the render/write
|
||||
/// fails), so [`hive_claude::Config::mcp_config`] stays unset and the
|
||||
/// subagent gets literally zero MCP servers, the same default as before this
|
||||
/// toggle existed. Re-rendered on every call (cheap: a filter plus a small
|
||||
/// Name the signal surface appears under in a subagent's own MCP config,
|
||||
/// and therefore the prefix its tools are called by
|
||||
/// (`mcp__subagent_control__goal_reached`).
|
||||
const SIGNAL_SERVER: &str = "subagent_control";
|
||||
|
||||
/// Render a subagent's `--mcp-config` file, returning its path — or `None`
|
||||
/// when there is nothing to put in it (or the render/write fails), so
|
||||
/// [`hive_claude::Config::mcp_config`] stays unset and the subagent gets
|
||||
/// literally zero MCP servers.
|
||||
///
|
||||
/// `signal_url` is this daemon's own `goal_reached`/`need_help` route; it is
|
||||
/// always included when given, since a subagent that can't say it's done or
|
||||
/// stuck is exactly the one the turn cap has to stop on its behalf. `None`
|
||||
/// reproduces the pre-continuation shape — only the opted-in extras, and no
|
||||
/// file at all when none opt in.
|
||||
///
|
||||
/// Re-rendered on every call (cheap: a filter plus a small
|
||||
/// file write) rather than cached once at daemon startup, so a config change
|
||||
/// takes effect on this subagent's next `start`/`continue` without needing
|
||||
/// the daemon itself restarted.
|
||||
#[must_use]
|
||||
pub fn build() -> Option<PathBuf> {
|
||||
pub fn build(signal_url: Option<&str>) -> Option<PathBuf> {
|
||||
let state_dir = crate::paths::state_dir();
|
||||
let servers: serde_json::Map<String, serde_json::Value> =
|
||||
let mut servers: serde_json::Map<String, serde_json::Value> =
|
||||
hive_agent_sock::extra_mcp::load_extra_mcp()
|
||||
.into_iter()
|
||||
.filter(|(_, spec)| spec.available_to_subagents())
|
||||
.map(|(name, spec)| (name, spec.to_json_entry(&state_dir)))
|
||||
.collect();
|
||||
if let Some(url) = signal_url {
|
||||
servers.insert(
|
||||
SIGNAL_SERVER.to_owned(),
|
||||
serde_json::json!({ "type": "http", "url": url }),
|
||||
);
|
||||
}
|
||||
if servers.is_empty() {
|
||||
return None;
|
||||
}
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -349,11 +349,14 @@ in
|
|||
# Subagent task runner daemon — independent of `hive-bash-daemon` (own
|
||||
# crate, own process): spawns nested claude sessions on request, serves
|
||||
# the `start`/`continue`/`status`/`interrupt` MCP tools directly over
|
||||
# streamable-http on `hyperhive.mcp.subagentHttpPort`. No task files —
|
||||
# this daemon's only state is an in-memory map of currently-running
|
||||
# processes, live only as long as the process is (see
|
||||
# `hive-subagent-mcp/src/session.rs`'s module doc); a restart stops
|
||||
# whatever's running, the actual claude session survives independently.
|
||||
# streamable-http on `hyperhive.mcp.subagentHttpPort`. The same port also
|
||||
# serves a second, subagent-facing route (`/signal/mcp`:
|
||||
# `goal_reached`/`need_help`) that the daemon hands each subagent it
|
||||
# spawns — not something an agent's own config points at. No task files —
|
||||
# this daemon's only state is in-memory, live only as long as the process
|
||||
# is (see `hive-subagent-mcp/src/session.rs`'s module doc); a restart
|
||||
# stops whatever's running, the actual claude session survives
|
||||
# independently.
|
||||
systemd.services.hive-subagent-daemon = {
|
||||
description = "subagent task runner + MCP daemon for hive-subagent";
|
||||
wantedBy = [ "multi-user.target" ];
|
||||
|
|
|
|||
Loading…
Reference in a new issue