//! The claude-facing half of this daemon: spawn a subagent turn, keep giving //! it turns until it says it's done or runs out of them, track it only while //! it's alive, and push exactly one todo when the whole run stops — saying //! whether it finished, was killed, or stopped for one of the four reasons //! the continuation loop records. //! //! **No task files, no restart recovery.** The daemon's only state is an //! in-memory `name -> Option` map (see `State`'s own doc for the //! rest of the maps and for what the `None`/`Some` split is for) — all of //! it living for exactly as long as the process is. A daemon restart means //! whatever was running gets killed //! with it (`tokio`'s own child-process drop semantics), not adopted, and //! the remembered `dir` is gone too. The durable record of a subagent's //! existence is `hive_claude::SessionStore` — claude's own on-disk //! session, found again by name. `continue` is how a caller reattaches to //! it, whether that's "give it a new turn" or "the daemon restarted and I //! want to pick this back up" — after a restart, re-supply `dir` once if //! the session isn't in the daemon's own default directory. //! //! **A running turn is not necessarily a working turn.** A wedged child //! satisfies "running" as fully as a busy one, so every line of the child's //! streams bumps `last_event_at` (`LivenessSink`) and `status` reports its //! age — the timestamp says the child is alive, never what it said. //! //! **`continue` doesn't pre-check the session's existence — it waits for //! the answer instead.** claude's own `--resume` is the authority, so //! `continue` holds its tool call open for up to `RESUME_GRACE` and reports //! a missed resume as its own error, naming the directory searched //! (`classify_end`). See `docs/tools/subagent.md`. //! **One `start` can be more than one turn.** With a `goal` set, a turn that //! ends without a stop signal is followed by another re-prompting the //! subagent toward it, up to `max_turns` (default five). The loop lives in //! `spawn_and_track`'s background task, so both tool-call contracts are //! unchanged and only one todo is pushed, when the *run* stops. No goal //! means what it always did: one turn, one todo. //! //! **A subagent's own stop signals are labels, not gates.** `goal_reached` //! and `need_help` stop the loop and extend the done message; neither //! verifies anything. `goal_reached` is self-reported by a subagent that has //! just been re-prompted with "you have not reached the goal", which is //! exactly the incentive to claim it — the same failure class as a build //! report asserting "done, tests pass". Everything here treats it as a claim //! about the work, never as the work, and says so. See //! `docs/tools/subagent.md`. //! **A subagent cannot say who it is.** Neither signal tool takes a session //! name. Each run is minted an unguessable token at spawn //! (`State::mint_signal_url`), the URL carrying it is written into that one //! subagent's own `--mcp-config`, and the route resolves it back to a session //! before dispatching — so the identity of a signal is a property of the //! endpoint it arrived on, not a field its sender filled in. An unminted or //! revoked token is a 404. The `occupancy()` liveness check that used to //! stand here instead was a guard on an assertion: two siblings running //! concurrently could each satisfy it for the other's name. //! **A killed turn is not a finished turn.** A child that died on a signal //! arrives as a `hive_claude::Error::Exit` carrying its `ExitStatus`, so the //! "how" is there to be read: `classify_end` takes the signal out of it and //! `State::finish_turn` remembers it against the name, which is what lets //! `status`, the end-of-turn todo and `continue`'s own reply all say the //! session was killed rather than let it read as finished. See //! `docs/tools/subagent.md`. //! //! **No mid-turn compaction.** Building on `hive_claude::Claude::spawn` + //! `RunningClaude::wait` directly (not `InfiniteSession::run`) is what makes //! `interrupt` possible at all — `InfiniteSession` has no cancel handle to //! reach in from the outside, only `RunningClaude::cancel_handle` does. The //! trade: this daemon doesn't get `InfiniteSession`'s reactive-compact-on- //! overflow or proactive-checkpoint-compact for free: a turn that overflows //! the context window surfaces as a plain `Error::PromptTooLong` to the //! caller instead of self-healing. Subagents are meant to be bounded, //! single-batch work (see the `base:claude-subagents` skill), not sessions //! long-lived enough to need in-place compaction — a real follow-up if that //! assumption stops holding, not shipped here. use std::collections::HashMap; use std::os::unix::process::ExitStatusExt as _; use std::path::PathBuf; use std::sync::{Arc, Mutex, PoisonError}; use std::time::{Duration, Instant}; use hive_claude::{Attach, Cancel, Claude, Config, SessionStore}; use tokio::sync::oneshot; /// How long a `continue` holds its tool call open waiting to find out /// whether the resume landed. A cap, not a delay: both real outcomes settle /// it well inside this, and it is only ever reached by a child that neither /// speaks nor exits. /// /// Measured rather than guessed, on this box, running the driver's own /// invocation (`--print --verbose --output-format stream-json --resume /// `) against a session name that matches nothing: 14 runs, /// 550–1087 ms wall from spawn to exit. A healthy turn's first stream event /// lands at roughly 500 ms, so the two paths finish within ~100 ms of each /// other and neither waits on this bound. Five seconds is ~4.6× the slowest /// miss observed — headroom for a loaded box starting node far slower than /// any of those runs, while still bounding the one case that reaches it. const RESUME_GRACE: Duration = Duration::from_secs(5); /// How many turns a goal-continued session gets before the harness stops it /// itself. Five is the number the feature was specified with, not one tuned /// here: enough for a bounded batch to converge, short enough that a /// subagent which has misunderstood its goal can't re-attempt it forever on /// someone else's budget. `start`'s `max_turns` overrides it per session, /// which is where a caller that genuinely needs a longer leash says so. const DEFAULT_MAX_TURNS: u32 = 5; /// Why a session's turn continuation stopped — recorded per name, reported /// by `status`, appended to the end-of-turn todo, and written into the /// session's report file when it has one. /// /// Only the four ends of the *continuation loop* live here. A turn that was /// killed or that failed outright never reaches the loop's decision at all: /// those keep the records they already had (`State::killed`, the todo's own /// killed wording), and giving them a second home here would have handed /// `status` two rival answers for one fact. #[derive(Debug, Clone, PartialEq, Eq)] enum StopReason { /// The turn ended and there was no goal to continue toward — the /// single-turn shape every session had before continuation existed. Done, /// The subagent called `goal_reached`. **Self-reported.** It records /// that the subagent claimed the goal, never that the goal was met. GoalReached(Option), /// The subagent called `need_help`: it can't proceed, and says why. NeedHelp(String), /// `turns` turns ran and the goal was never reported reached. TurnCap { turns: u32 }, } impl StopReason { /// The one sentence that says why, shared by the todo extension and the /// report-file line so a parent reading either sees the same words. /// /// `GoalReached` carries its "self-reported" caveat in the sentence /// itself rather than leaving it to whichever surface renders it: the /// caveat is the load-bearing half of that claim, and a surface that /// forgot to add it would read as verification. fn sentence(&self) -> String { match self { Self::Done => "its turn ended and there was no goal to continue toward".to_owned(), Self::GoalReached(msg) => { let said = msg .as_deref() .map_or_else(String::new, |m| format!(": {m}")); format!( "the subagent reported its goal reached{said} — self-reported, not verified, \ so read what it actually changed before acting on it" ) } Self::NeedHelp(msg) => { format!("the subagent called `need_help` and can't proceed: {msg}") } Self::TurnCap { turns } => format!( "the harness turn limit was reached ({turns} turns) without the goal ever being \ reported reached, so the work stopped where it had got to" ), } } } /// How a subagent's turn ended, as far as the daemon can tell from what /// [`hive_claude::RunningClaude::wait`] returned. /// /// The distinction that matters is `Killed` vs everything else: a child that /// died on a signal was cut off mid-turn by something outside this daemon /// (the kernel's OOM killer, a `systemctl stop`, an operator's `kill`) or by /// an `interrupt` call here — either way its work stopped wherever it got /// to, which is not what `Complete` means. #[derive(Debug, Clone, PartialEq, Eq)] enum TurnEnd { /// The turn ran to completion. Complete, /// The child was terminated by `signal` instead of exiting on its own. Killed { signal: i32 }, /// claude exited on its own but did not complete the turn — a /// recognized sentinel (rate limit, prompt too long, …) or a plain /// non-zero exit. Carries the message shown to the caller. Failed(String), } /// What settled a resumed turn's bounded wait — the one thing `continue` /// blocks on before it answers. enum ResumeVerdict { /// The child emitted a stream event that isn't the turn's own terminal /// `result` — which a resume that matched nothing never gets as far as /// emitting. The turn is genuinely running; there is nothing left to /// wait for. Underway, /// The turn ended before that happened. A `Failed` end here is the /// resume miss `continue` exists to report. Ended(TurnEnd), } /// Write half of [`ResumeVerdict`]'s channel, shared between the turn's sink /// (which reports `Underway`) and its background task (which reports /// `Ended`). `Option` because a `oneshot::Sender` is consumed by its single /// send: whichever side gets there first takes it, and the loser finds /// nothing left to send on — exactly the "first answer wins" the wait wants. /// The whole thing is `None` on a `start`, which has no resume to miss. type VerdictTx = Arc>>>; /// Report `verdict`, if nothing has been reported yet. /// /// Returns `true` only when this call both won that race *and* found a /// receiver still listening — i.e. the verdict genuinely reached a /// `continue` that is about to act on it. The background task reads that /// answer to decide whether its end-of-turn todo would be a second /// notification for something the caller has already been told to its face. fn settle(tx: Option<&VerdictTx>, verdict: ResumeVerdict) -> bool { let Some(tx) = tx else { return false }; let taken = tx.lock().unwrap_or_else(PoisonError::into_inner).take(); taken.is_some_and(|tx| tx.send(verdict).is_ok()) } /// Name a signal number for a human: `9` reads as `SIGKILL (signal 9)`. Only /// the three that end a subagent in practice are named — anything else keeps /// the number, which is still enough to look up. fn describe_signal(signal: i32) -> String { match signal { libc::SIGKILL => "SIGKILL (signal 9)".to_owned(), libc::SIGTERM => "SIGTERM (signal 15)".to_owned(), libc::SIGINT => "SIGINT (signal 2)".to_owned(), other => format!("signal {other}"), } } /// Read how the turn ended out of what the driver returned. A signalled /// child surfaces as [`hive_claude::Error::Exit`] whose `ExitStatus` has a /// `signal()` — the driver's own docs point at exactly this check — so the /// "how" is preserved by the time it reaches here rather than having to be /// recovered. /// /// `searched` is the resume lookup's location, present only for a turn that /// attached with [`Attach::Resume`] (see `searched_location`). A missed /// resume arrives as [`hive_claude::Error::SessionNotFound`], whose message /// names the *value* that matched nothing but not the *directory* it was /// looked for in — and "the session lives in another directory" is the way /// this actually fails in practice, so the caller gets told where the daemon /// looked rather than left to conclude the session is gone. fn classify_end(outcome: hive_claude::Result<()>, searched: Option<&str>) -> TurnEnd { let Err(error) = outcome else { return TurnEnd::Complete; }; if let hive_claude::Error::Exit { status, .. } = &error && let Some(signal) = status.signal() { return TurnEnd::Killed { signal }; } if matches!(error, hive_claude::Error::SessionNotFound) && let Some(searched) = searched { return TurnEnd::Failed(format!("claude error: {error} {searched}")); } TurnEnd::Failed(format!("claude error: {error}")) } /// Where a `--resume` against `config` would have looked for the session, /// phrased to append to claude's own not-found message. Both halves come /// from the same resolution `build_store` and the driver itself use, so this /// names the directory that was genuinely searched rather than a guess at /// it. `None` when either half fails to resolve (no `HOME`, a `cwd` that /// doesn't exist) — an unresolvable location is worse than none, and the /// underlying error still reaches the caller unadorned. fn searched_location(config: &Config) -> Option { let home = config.resolved_claude_home().ok()?; let cwd = config.resolved_cwd().ok()?; Some(format!( "(searched {} for cwd {}; if it was started elsewhere, pass the `dir` it was started in)", home.display(), cwd.display() )) } /// This daemon's whole state: which names have a live process or a /// reservation in flight, what each is working toward and how it stopped, /// and where to push the completion todo. `Arc`-wrapped so the background /// task driving a run outlives the tool call that started it. /// /// The map value is `Option`: `None` means `name` is reserved for /// an in-flight `start`/`continue` that hasn't reached a confirmed /// `Claude::spawn` yet; `Some(cancel)` means a real process is tracked and /// interruptible. The `None` state exists to close a real TOCTOU window a /// reviewer caught in the original check-then-insert version: checking "is /// `name` free" and committing to it are two different lock acquisitions /// unless the check *is* the reservation — see `reserve`. /// /// `killed` is the other half of that map's story: `running` says what is in /// flight *now*, `killed` remembers the names whose last turn ended on a /// signal rather than on its own, keyed to the signal number. Without it a /// killed session is indistinguishable from a finished one the moment its /// entry leaves `running` — the whole point of this record. /// /// `last_event` splits `running` a second way — *in flight* against /// *making progress*, since a wedged child is tracked exactly like a busy /// one (`note_event` / `last_event_age`). /// /// ⚠️ This concurrency guard is keyed by `name` alone — a `start`/`continue` /// for `name` with a *different* `dir` than one already in flight under /// that name is refused as "already running," even though the two would /// resolve to entirely separate on-disk sessions. Deliberate: `name` is the /// caller's one chosen identity for a subagent, not `(name, dir)` — reuse a /// name across directories at your own risk, the tool doesn't disambiguate /// it (flagged in review when `dir` was added). pub struct State { running: Mutex>>, dirs: Mutex>, killed: Mutex>, last_event: Mutex>, /// Turn continuation's three records. Unlike the four maps above they /// deliberately **outlive the turn** — a stop reason that vanished with /// the process it described would be unreadable by the time anyone /// asked — so they're cleared by the next `start` under the same name, /// not by a turn ending. goals: Mutex>, stops: Mutex>, reports: Mutex>, socket: PathBuf, /// The prefix of every subagent's signal URL — this daemon's own `--http` /// address with [`crate::mcp::SIGNAL_PATH`] on the end, and *not* a /// reachable route by itself. It lives here because the daemon can only /// learn it from its own `--http` argument — deriving it from a /// convention would be the same inference the report path is careful not /// to make. signal_base: String, signal_tokens: Mutex, } /// Which opaque URL segment belongs to which session — the whole of a /// subagent's identity, as far as `goal_reached`/`need_help` are concerned. /// /// A subagent is told one URL, in its own `--mcp-config`, and that URL is /// what says who it is: it has no field to name a session in and no second /// session's URL to reach for. Both directions live under one lock because a /// half-updated pair is exactly the state in which a token could resolve to /// a session that has since minted another one. #[derive(Default)] struct SignalTokens { /// The token currently minted for a session, so a re-mint can retire it. by_name: HashMap, /// The resolution the route does: token -> the session it speaks for. by_token: HashMap, } /// What a session is being continued toward, and how far through its turn /// budget it is. Present only for a session `start`ed with a `goal`; its /// absence is what makes a session single-turn. struct GoalState { /// Verbatim from `start` — re-prompted at the subagent each turn rather /// than paraphrased, since the caller wrote it for the subagent to read. goal: String, /// The cap this session runs under (`start`'s `max_turns`). max_turns: u32, /// Turns started so far, counting the first — so the very first turn is /// `1 of max_turns`, not `0`. turn: u32, } impl State { /// `signal_base` is where the streamable-http endpoint a subagent's own /// claude reaches `goal_reached`/`need_help` on *starts* — this daemon's /// `--http` address with the signal route appended (see /// `crate::mcp::serve_http`). Each session's actual URL is that plus its /// own token; see `State::mint_signal_url`. #[must_use] pub fn new(socket: PathBuf, signal_base: String) -> Self { Self { running: Mutex::new(HashMap::new()), dirs: Mutex::new(HashMap::new()), killed: Mutex::new(HashMap::new()), last_event: Mutex::new(HashMap::new()), goals: Mutex::new(HashMap::new()), stops: Mutex::new(HashMap::new()), reports: Mutex::new(HashMap::new()), socket, signal_base, signal_tokens: Mutex::new(SignalTokens::default()), } } /// Mint `name` a fresh signal URL: an unguessable token appended to /// [`State::signal_base`], resolvable back to this one session and to no /// other. Called once per spawned run, and the result goes into exactly /// one place — that subagent's own `--mcp-config` (`crate::mcp_config`). /// /// A v4 UUID's 122 bits come from the OS CSPRNG, so the segment is not /// derived from the name, the port or anything else a sibling subagent /// could compute; a subagent that wants to signal as somebody else has /// nothing to guess *from*. Minting replaces any token the name held /// before, which is what stops a name's old URL surviving the run it was /// issued for. pub(crate) fn mint_signal_url(&self, name: &str) -> String { let token = uuid::Uuid::new_v4().simple().to_string(); let mut tokens = self .signal_tokens .lock() .unwrap_or_else(PoisonError::into_inner); if let Some(previous) = tokens.by_name.insert(name.to_owned(), token.clone()) { tokens.by_token.remove(&previous); } tokens.by_token.insert(token.clone(), name.to_owned()); drop(tokens); format!("{}/{token}", self.signal_base) } /// Which session an incoming signal request speaks for, or `None` for a /// token this daemon never minted or has since revoked — which the route /// answers with a bare 404 (see `crate::mcp::serve_http`). `None` is the /// only failure shape there is: nothing about the answer distinguishes /// "never existed" from "that run is over", so a caller holding a wrong /// token learns nothing from being refused. pub(crate) fn session_for_signal_token(&self, token: &str) -> Option { self.signal_tokens .lock() .unwrap_or_else(PoisonError::into_inner) .by_token .get(token) .cloned() } /// Retire `name`'s signal URL — its run is over (or never started), so /// the route it was handed stops resolving and answers 404 from here on. /// This is the expiry half of "unknown or expired token ⇒ 404": without /// it a finished subagent's config file would still name a live route. fn revoke_signal_token(&self, name: &str) { let mut tokens = self .signal_tokens .lock() .unwrap_or_else(PoisonError::into_inner); if let Some(token) = tokens.by_name.remove(name) { tokens.by_token.remove(&token); } } /// `Some(true)` — a live process is tracked, interruptible. `Some(false)` /// — the name is claimed but no process is confirmed under it: an /// in-flight start/continue that hasn't spawned yet, or the gap between /// one continued turn's child exiting and the next one spawning (see /// `between_turns`). `None` — nothing tracked under `name` at all. fn occupancy(&self, name: &str) -> Option { self.running .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .map(Option::is_some) } /// Atomically claim `name` for an in-flight start/continue: check /// "is anything tracked under this name" and "commit to this call /// owning it" in the *same* lock acquisition, so two calls racing the /// same name can't both pass a check before either commits (the exact /// same-name concurrent-run hazard this module's doc warns about). /// Returns `false` (reserving nothing) if `name` is already reserved /// or running. fn reserve(&self, name: &str) -> bool { let mut running = self.running.lock().unwrap_or_else(PoisonError::into_inner); if running.contains_key(name) { return false; } running.insert(name.to_owned(), None); true } /// Release a reservation that never made it to a real spawn (an error /// on the slow path between `reserve` and `Claude::spawn` succeeding). /// A no-op if the entry was already upgraded to `Some` — this only ever /// clears a still-`None` placeholder, never a live process. /// /// The signal token the failed call minted goes with it: no process ever /// read that URL, and a token outliving the call that minted it is the /// one way a route could resolve to a session that isn't there. fn release_reservation(&self, name: &str) { let mut running = self.running.lock().unwrap_or_else(PoisonError::into_inner); let released = matches!(running.get(name), Some(None)); if released { running.remove(name); } drop(running); // Only alongside a reservation this actually released: a call that // found a live process left it running, and revoking that run's URL // would cut off a subagent which is still using it. if released { self.revoke_signal_token(name); } } /// Retire the finished turn's tracking for `name` and remember how it /// ended: a `Killed` end is recorded so `status`, the end-of-turn todo /// and a later `continue` can all say so; any other end clears a stale /// record from an earlier turn. /// /// The kill is recorded *before* the `running` entry goes away, so there /// is no instant in which `name` is neither running nor known-killed — /// a `status` landing between the two would otherwise read the session /// as plainly idle, which is the exact confusion this record exists to /// remove. /// /// The liveness clock goes with the `running` entry, for the same reason /// it's kept at all: it answers "is this turn still making progress", /// and a turn that has ended has no progress left to make. /// /// So does the session's signal token. Every caller of this is a point /// where the *run* stops — a continued run's own turn boundary goes /// through `between_turns` instead, and keeps its URL because the next /// turn is the same subagent against the same rendered config. Revoking /// here is what makes a signal for a session that already ended a 404 /// rather than a late stop reason recorded against it. fn finish_turn(&self, name: &str, end: &TurnEnd) { self.revoke_signal_token(name); let mut killed = self.killed.lock().unwrap_or_else(PoisonError::into_inner); match end { TurnEnd::Killed { signal } => killed.insert(name.to_owned(), *signal), TurnEnd::Complete | TurnEnd::Failed(_) => killed.remove(name), }; drop(killed); self.running .lock() .unwrap_or_else(PoisonError::into_inner) .remove(name); self.last_event .lock() .unwrap_or_else(PoisonError::into_inner) .remove(name); } /// The signal `name`'s last turn died on, if it died on one. fn killed_by(&self, name: &str) -> Option { self.killed .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .copied() } /// Forget that `name`'s last turn was killed — called once a new turn is /// confirmed spawned, since the record describes the turn before it and /// would otherwise keep flagging a session that has since run again. fn clear_kill(&self, name: &str) { self.killed .lock() .unwrap_or_else(PoisonError::into_inner) .remove(name); } /// `name`'s turn just produced output — record *when*, and nothing else. /// Called from `LivenessSink` for every line of every stream, so it runs /// far more often than anything else in this type and stays a single /// map write for that reason. /// /// A monotonic [`Instant`], not a wall clock: what's ever read back out /// is an elapsed age, which a clock adjustment must not be able to /// distort into a stall that never happened. /// /// Also called by `spawn_and_track` the moment each turn's child exists, /// so the clock starts at the spawn rather than at the first line. /// Without that seed a subagent that wedged before emitting anything at /// all would report no age forever — the one case where an age is most /// worth having. The value therefore reads as "how long since the daemon /// last heard anything from this child", counting the spawn itself as /// the first thing it heard. A continued run re-seeds it per turn, for /// the same reason: the age describes the turn in flight, not the run. fn note_event(&self, name: &str) { self.last_event .lock() .unwrap_or_else(PoisonError::into_inner) .insert(name.to_owned(), Instant::now()); } /// How long since `name`'s turn last produced output. `None` once the /// turn is over (`finish_turn` drops the entry) or for a name that never /// reached a spawn — in both cases there's no *running* turn whose /// progress the age would describe, and a leftover age from a turn that /// has already ended would read as a stall that isn't one. fn last_event_age(&self, name: &str) -> Option { self.last_event .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .map(Instant::elapsed) } /// An explicit `dir` is remembered against `name` and returned as-is; an /// omitted one (`None`) falls back to whatever was last remembered for /// `name`, so a `start` that gave a `dir` doesn't force every later /// `continue`/`status` against the same name to repeat it. Explicit /// always wins and updates the memory — there's no way to say "forget /// it, use the default" once something's been remembered short of /// giving a genuinely different `dir`, a deliberate simplicity trade, /// not an oversight. In-memory only, same durability envelope as /// `running`: a daemon restart forgets it same as it forgets everything /// else here — the caller re-supplies `dir` once, same as it always /// could. /// /// ⚠️ **Callers must only call this once the call is known to proceed** /// (i.e. after `reserve` has already succeeded in `start`/`continue_`) — /// see `peek_dir` for the non-committing variant a call that might still /// be refused (or that shouldn't persist its `dir` at all, like /// `status`) needs instead. This one argus caught in review: calling it /// unconditionally before `reserve`'s check meant a *rejected* concurrent /// call still overwrote the remembered `dir` for the name it was refused /// against. fn resolve_dir(&self, name: &str, dir: Option<&str>) -> Option { let mut dirs = self.dirs.lock().unwrap_or_else(PoisonError::into_inner); match dir { Some(d) => { dirs.insert(name.to_owned(), d.to_owned()); Some(d.to_owned()) } None => dirs.get(name).cloned(), } } /// Same resolution as `resolve_dir` (explicit wins, omitted falls back /// to the remembered value) but never writes — for a call that must not /// change what a later bare `continue`/`status` resolves to. `status`'s /// own doc offers `dir` as a one-off "check a different directory" /// knob; if that call committed the same way `start`/`continue_` do, the /// one-off peek would silently become the new remembered default (the /// second bug argus flagged in the same review). fn peek_dir(&self, name: &str, dir: Option<&str>) -> Option { match dir { Some(d) => Some(d.to_owned()), None => self .dirs .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .cloned(), } } /// Upgrade `name`'s `None` reservation to a real, interruptible process. /// Same key as the reservation, so there is no window in which `name` /// reads as unoccupied between the two. fn track(&self, name: &str, cancel: Cancel) { self.running .lock() .unwrap_or_else(PoisonError::into_inner) .insert(name.to_owned(), Some(cancel)); } /// Drop `name`'s cancel handle but keep the name claimed — the moment /// between a continued turn's child exiting and its successor being /// spawned. Without it the name would read as free mid-loop and a /// concurrent `start` could take it out from under the continuation; /// with it, `status` says "starting" and `interrupt` says "still /// starting, retry shortly", both of which are true of the sub-second /// gap it covers. fn between_turns(&self, name: &str) { self.running .lock() .unwrap_or_else(PoisonError::into_inner) .insert(name.to_owned(), None); } /// Set (or clear) `name`'s goal and start its turn budget over at turn /// one. Called by `start` only: the goal is a property of the session /// being created, and a later `continue` re-prompts toward whatever /// `start` set rather than redefining it. fn set_goal(&self, name: &str, goal: Option, max_turns: u32) { let mut goals = self.goals.lock().unwrap_or_else(PoisonError::into_inner); match goal { Some(goal) => { goals.insert( name.to_owned(), GoalState { goal, max_turns, turn: 1, }, ); } None => { goals.remove(name); } } } /// Put `name` back at turn one, keeping whatever goal `start` set. A /// `continue` is the parent's own deliberate turn, and the cap exists to /// bound *unattended* continuation — so its allowance starts over rather /// than a capped session being permanently un-continuable. The parent /// was always the authority on whether more turns are worth spending. fn restart_turns(&self, name: &str) { if let Some(goal) = self .goals .lock() .unwrap_or_else(PoisonError::into_inner) .get_mut(name) { goal.turn = 1; } } /// `(turn, max_turns)` for `name`. `None` for a session with no goal, /// which has no turn budget to be partway through — reporting `1 of 1` /// there would invent a cap that isn't enforced. fn turns(&self, name: &str) -> Option<(u32, u32)> { self.goals .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .map(|g| (g.turn, g.max_turns)) } /// Record why `name` stopped. Idempotent by design: the subagent's own /// signal lands here mid-turn and the loop re-records the same reason at /// the turn's end, so both paths can write without checking. fn record_stop(&self, name: &str, stop: StopReason) { self.stops .lock() .unwrap_or_else(PoisonError::into_inner) .insert(name.to_owned(), stop); } fn stop_reason(&self, name: &str) -> Option { self.stops .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .cloned() } /// Forget why `name` last stopped — called when a new turn is starting, /// since the record describes the run before it and would otherwise make /// a session that has since been given another turn still read as /// blocked or out of turns. fn clear_stop(&self, name: &str) { self.stops .lock() .unwrap_or_else(PoisonError::into_inner) .remove(name); } /// Remember where `name` writes its report, so a stop reason can be /// appended to the artifact a parent already reads rather than living /// only in a todo. `None` leaves whatever was remembered alone, so a /// signal tool that doesn't name a path doesn't erase the one `start` /// gave. /// /// Never inferred: `start` carries what the brief named, and the signal /// tools carry where the subagent says it actually wrote. A daemon that /// derived this from a path convention would be guessing about someone /// else's layout. fn set_report_file(&self, name: &str, path: Option<&str>) { if let Some(path) = path { self.reports .lock() .unwrap_or_else(PoisonError::into_inner) .insert(name.to_owned(), PathBuf::from(path)); } } /// Forget `name`'s report path — a fresh `start` under the same name is /// a different piece of work, and inheriting the last one's artifact /// path would append its stop reason to a file this run never wrote. fn clear_report_file(&self, name: &str) { self.reports .lock() .unwrap_or_else(PoisonError::into_inner) .remove(name); } fn report_file(&self, name: &str) -> Option { self.reports .lock() .unwrap_or_else(PoisonError::into_inner) .get(name) .cloned() } /// What the continuation loop does after a turn that ran to completion: /// stop with a reason, or spend another turn re-prompting toward the /// goal. Advances the turn counter itself, since deciding to continue /// and consuming a turn are the same act. /// /// A signal the subagent raised mid-turn is already in `stops` and wins /// outright — that is what "both stop goal continues" means, and it's /// checked before the goal so a `goal_reached` on the last allowed turn /// reads as reached rather than as capped. fn plan_after_turn(&self, name: &str) -> Continuation { if let Some(stop) = self.stop_reason(name) { return Continuation::Stop(stop); } let mut goals = self.goals.lock().unwrap_or_else(PoisonError::into_inner); let Some(state) = goals.get_mut(name) else { return Continuation::Stop(StopReason::Done); }; if state.turn >= state.max_turns { return Continuation::Stop(StopReason::TurnCap { turns: state.turn }); } state.turn += 1; Continuation::Continue { prompt: continuation_prompt(name, &state.goal, state.turn, state.max_turns), } } } /// `plan_after_turn`'s answer: the loop either stops with a reason to /// report, or has the next turn's prompt ready to spawn against. #[derive(Debug, PartialEq, Eq)] enum Continuation { Stop(StopReason), Continue { prompt: String }, } /// A caller-chosen name, validated the same way `hive-bash-mcp`'s task ids /// are: a single safe [`hive_types::Ident`] segment, which doubles as /// claude's own `--name`/`--resume` session title. fn validate_name(name: &str) -> anyhow::Result<()> { hive_types::Ident::parse(name) .map(|_| ()) .map_err(|e| anyhow::anyhow!("invalid subagent name {name:?}: {e}")) } /// Extend the ambient `OTEL_RESOURCE_ATTRIBUTES` with a `subagent=` /// attribute, so every token/cost/tool-call data point this subagent's own /// claude process emits carries it alongside the parent's `agent=` /// label. `Config.env` applies after the inherited environment, so this one /// entry overriding the ambient value is the intended shape, not a /// wholesale replacement. fn subagent_otel_attrs(name: &str) -> String { match std::env::var("OTEL_RESOURCE_ATTRIBUTES") { Ok(existing) if !existing.is_empty() => format!("{existing},subagent={name}"), _ => format!("subagent={name}"), } } /// Build the `Config` one subagent turn runs against. `model` maps straight /// onto `Config::model` — `--model`, omitted when `None` so claude falls /// back to its own default. `effort` does not: an omitted `effort` defaults /// to `"medium"` here rather than falling through to claude's own default /// (`high` on most models) — a deliberate hive policy for subagent work /// specifically, not a reflection of Anthropic's own recommendation, on the /// same "cheaper than you" cost-consciousness the `base:claude-subagents` /// skill already asks of `model`. `prompt_file`, when /// given, becomes `--append-system-prompt-file` — the subagent's task /// instructions. `dir`, when given, becomes `Config::cwd` (e.g. a worktree /// the caller already prepared); `None` inherits this daemon's own working /// directory, same as before this field existed. Always /// `--dangerously-skip-permissions --strict-mcp-config` — the safety /// property is `strict_mcp_config: true` with no ambient MCP discovery, not /// an unconditional absence of `--mcp-config`: a subagent gets exactly the /// `hyperhive.extraMcpServers` entries an operator has explicitly opted in /// via `availableToSubagents = true` (`crate::mcp_config::build`), plus this /// daemon's own two-tool signal surface when `signal_url` is given — nothing /// implicit and nothing more. /// /// `signal_url` is what makes `goal_reached`/`need_help` callable at all: a /// subagent reaches them over the same streamable-http listener its parent /// uses, on a route that serves those two tools and nothing else, so being /// able to say "I'm done" never carries the ability to spawn a subagent of /// its own. It is also *this* session's own URL — minted per run by /// [`State::mint_signal_url`] and written only into this session's config /// file — which is where the signal tools get the identity they no longer /// ask the caller for. `None` — which only `status` passes, building a /// config purely to resolve the session store — leaves the surface out /// entirely. fn build_config( name: &str, model: Option, effort: Option, prompt_file: Option<&str>, dir: Option<&str>, signal_url: Option<&str>, ) -> Config { let mut extra_args = vec!["--dangerously-skip-permissions".to_owned()]; if let Some(path) = prompt_file { extra_args.push("--append-system-prompt-file".to_owned()); extra_args.push(path.to_owned()); } Config { model, effort: Some(effort.unwrap_or_else(|| "medium".to_owned())), cwd: dir.map(PathBuf::from), mcp_config: crate::mcp_config::build(name, signal_url), strict_mcp_config: true, extra_args, env: vec![( "OTEL_RESOURCE_ATTRIBUTES".to_owned(), subagent_otel_attrs(name), )], ..Default::default() } } /// The [`SessionStore`] a subagent's turn actually runs against — same /// resolution `hive_claude::Claude` itself uses, so a lookup here can't /// disagree with what the driver does a moment later. fn build_store(config: &Config) -> std::io::Result { Ok(SessionStore::new( config.resolved_claude_home()?, config.resolved_cwd()?, )) } /// Everything a `start` needs, as one struct rather than a parameter list: /// the call already carried six mostly-optional values before goals were /// added, and nine positional arguments is both unreadable at the call site /// and a lint. pub struct StartRequest { /// Session name — the tracking key and claude's own session title. pub name: String, pub model: Option, pub effort: Option, /// File holding the subagent's task instructions. pub prompt_file: String, /// The first turn's prompt. A goal, when given, is appended to it. pub trigger: String, /// Working directory for the session; `None` inherits the daemon's. pub dir: Option, /// What this session is being continued *toward*. `None` keeps the /// pre-continuation shape: one turn, one todo, no re-prompting. pub goal: Option, /// Turn cap for the continuation, defaulting to `DEFAULT_MAX_TURNS`. /// Ignored without a `goal`, which is what continuation continues /// toward — there is nothing to re-prompt against otherwise. pub max_turns: Option, /// Where this session's brief told it to write its report, so the stop /// reason can be appended to that artifact. Never inferred — see /// `State::set_report_file`. pub report_file: Option, } /// Start a fresh subagent under `name`. A prior *finished* session under /// the same name is archived first (so this is a real fresh start, not a /// silent resume of old history) — a *currently running* one is refused /// outright, since `hive_claude::InfiniteSession`'s own docs warn that two /// concurrent runs against the same name corrupt both. /// /// With a `goal` set this starts a whole *run*, not a single turn: see /// `spawn_and_track` for the continuation loop. The return is unchanged /// either way — it reports the first turn's spawn, not the run's outcome. /// /// # Errors /// /// A name already running, an invalid name, an archive failure, or the /// underlying `Claude::spawn` failing (binary missing, etc.) — the last /// case is the only one that can happen *after* commit-to-run, and it's /// exactly why nothing is registered in `running` until spawn actually /// succeeds. pub fn start(state: &Arc, req: StartRequest) -> anyhow::Result { let name = req.name.as_str(); validate_name(name)?; if !state.reserve(name) { anyhow::bail!("subagent `{name}` is already running — use `continue` or `interrupt`"); } // Only commit the remembered `dir` now that `reserve` has actually // claimed `name` — see `resolve_dir`'s doc for why the order matters, // and the same reasoning governs the three records below it. let dir = state.resolve_dir(name, req.dir.as_deref()); // A fresh start owns none of the previous run's records: its goal, its // turn budget, why it stopped and where it wrote are all about work this // call is deliberately replacing. state.clear_stop(name); state.clear_report_file(name); state.set_report_file(name, req.report_file.as_deref()); let max_turns = req.max_turns.unwrap_or(DEFAULT_MAX_TURNS).max(1); state.set_goal(name, req.goal.clone(), max_turns); let trigger = match req.goal.as_deref() { None => req.trigger, Some(goal) => format!("{}{}", req.trigger, goal_briefing(name, goal, max_turns)), }; let result = start_reserved( state, name, req.model, req.effort, &req.prompt_file, trigger, dir.as_deref(), ); if result.is_err() { state.release_reservation(name); } result } /// The slow, fallible part of `start`, run only after `reserve` has /// already closed the TOCTOU window — split out so `start` can release the /// reservation on any error path here without duplicating that logic per /// failure site. fn start_reserved( state: &Arc, name: &str, model: Option, effort: Option, prompt_file: &str, trigger: String, dir: Option<&str>, ) -> anyhow::Result { let signal_url = state.mint_signal_url(name); let config = build_config( name, model, effort, Some(prompt_file), dir, Some(&signal_url), ); let store = build_store(&config)?; if store.find_by_title(name).is_some() { tracing::info!( name, "start: archiving a finished prior session for a fresh start" ); store .archive_by_title(name) .map_err(|e| anyhow::anyhow!("archiving the prior `{name}` session failed: {e}"))?; } // No verdict channel: a `start` creates its session, so there is no // resume to miss and nothing for the caller to wait on past the spawn — // see `spawn_and_track`'s doc. spawn_and_track( state, name, &config, &Attach::Create(name.to_owned()), trigger, None, ) } /// Give an existing named session a new turn — resuming it whether that /// means "the previous turn finished, here's the next instruction" or "the /// daemon restarted, reattaching." Refuses a name already running (same /// concurrent-run hazard as `start`). /// /// A name with **no** session to resume is still not *pre-checked* here: /// claude's own `--resume` answers that, and the pre-check that used to live /// here could only repeat the lookup the driver was about to do anyway, /// while telling the caller a session didn't exist when the true answer was /// almost always that it exists somewhere else. What changed is that /// `continue` now waits for the driver's answer instead of returning ahead /// of it: a missed resume comes back as this call's own `Err`, carrying /// claude's message and the directory that was searched (`classify_end`), /// rather than only as an end-of-turn todo the caller had already stopped /// looking for. A `continue` that reports "started" therefore means the turn /// started. See `await_resume` for the bound, and `spawn_and_track`'s doc for /// why `start` keeps the older, unconditional contract. /// /// Resuming a session whose last turn was *killed* is allowed — that is /// often exactly what the caller wants — but never silent: the reply says /// so, since a caller that never ran `status` and missed the todo would /// otherwise carry on from cut-off work believing it was finished work. /// /// A goal session re-enters its continuation loop here with a fresh turn /// budget — the cap bounds *unattended* re-prompting, not the parent's own. /// /// # Errors /// /// An invalid name, one already running, `Claude::spawn` failing, or the /// resumed turn failing within `RESUME_GRACE` — in practice a missed resume. pub async fn continue_( state: &Arc, name: &str, prompt: String, model: Option, effort: Option, dir: Option<&str>, ) -> anyhow::Result { validate_name(name)?; if !state.reserve(name) { anyhow::bail!( "subagent `{name}` is already running — use `interrupt` first if you meant to redirect it" ); } // Read before the spawn, which clears the record as soon as the new turn // is confirmed — this reply is the last chance to mention it. let killed = state.killed_by(name); // Only commit the remembered `dir` now that `reserve` has actually // claimed `name` — see `resolve_dir`'s doc for why the order matters. let dir = state.resolve_dir(name, dir); // A new turn supersedes why the last run stopped, and gives a // goal-continued session its allowance back — see `restart_turns`. state.clear_stop(name); state.restart_turns(name); let (tx, rx) = oneshot::channel(); let verdict: VerdictTx = Arc::new(Mutex::new(Some(tx))); let started = continue_reserved(state, name, prompt, model, effort, dir.as_deref(), &verdict); if started.is_err() { state.release_reservation(name); return started; } // Nothing to release on this path either way: a turn that ended inside // the grace has already been through `finish_turn`, which clears the // tracking the spawn put there, and one that a goal carried into another // turn still owns the name via `between_turns`. await_resume(rx).await?; started.map(|msg| note_resumed_after_kill(&msg, killed)) } /// Hold a `continue` open until its resume is known to have landed — or /// until [`RESUME_GRACE`] says that waiting any longer costs more than the /// answer is worth. /// /// Returns as soon as *either* side of the race reports, so a successful /// `continue` pays no fixed delay: its first stream event settles the wait /// at roughly the same moment a miss's exit would have. The timeout is the /// floor under a child that does neither, and a turn that reaches it is /// reported as started — which it is, with the end-of-turn todo left to say /// how it goes. /// /// The third case is the channel closing with nothing ever sent — both the /// sink and the background task dropping their `Arc` without calling /// `settle`, which only a panic in the task can produce. That reads here as /// "started", the same as `Underway`, and the fail-open is deliberate: the /// one thing already known is that `Claude::spawn` returned a live pid, so /// answering "the resume missed" would be a claim about the session that /// nothing observed. The turn's real end still reaches the caller as a todo. /// /// # Errors /// /// The turn's own failure message, verbatim: the caller asked claude to /// resume a session and claude said why it couldn't, which is a better /// answer than anything this daemon could paraphrase it into. async fn await_resume(rx: oneshot::Receiver) -> anyhow::Result<()> { match tokio::time::timeout(RESUME_GRACE, rx).await { Ok(Ok(ResumeVerdict::Ended(TurnEnd::Failed(e)))) => anyhow::bail!("{e}"), // Everything else is a turn that started: it spoke (`Underway`), or // it ended on its own terms within the grace — `Complete`, or a // `Killed` that some concurrent `interrupt` asked for and whose todo // says so — or it is still going when the grace runs out, or nobody // ever sent at all, which only a panicked task produces and which // this deliberately fails open on. See the doc above. _ => Ok(()), } } /// Append the "you are resuming a killed session" note to `continue`'s reply /// when its previous turn was signalled; pass the reply through unchanged /// otherwise. fn note_resumed_after_kill(msg: &str, killed: Option) -> String { match killed { None => msg.to_owned(), Some(signal) => format!( "{msg} — note: its previous turn was killed ({}) rather than finishing, so this turn \ resumes from work that was cut off mid-way", describe_signal(signal) ), } } /// The slow, fallible part of `continue_`, run only after `reserve` has /// already closed the TOCTOU window — same split rationale as /// `start_reserved`. fn continue_reserved( state: &Arc, name: &str, prompt: String, model: Option, effort: Option, dir: Option<&str>, verdict: &VerdictTx, ) -> anyhow::Result { // A fresh URL for the resumed run, not the one the last run was handed: // a token is per run, and this is a new one. let signal_url = state.mint_signal_url(name); let config = build_config(name, model, effort, None, dir, Some(&signal_url)); // No existence pre-check: claude's own `--resume` is the authority on // whether the session is there, and it errors rather than quietly // starting a fresh one. `verdict` is how that answer gets back to the // caller in time to be its error. See this module's doc. spawn_and_track( state, name, &config, &Attach::Resume(name.to_owned()), prompt, Some(Arc::clone(verdict)), ) } /// Bumps `name`'s liveness clock on every line of the turn's output, and /// does nothing else with it. Replaces the `NoopSink` this daemon used to /// run turns against, which discarded the stream wholesale and left `status` /// unable to tell a working child from a wedged one. /// /// **All three callbacks, deliberately.** A stderr line or a stdout line /// that didn't parse as JSON is proof the child is alive every bit as much /// as a stream-json event is, and the failure that matters here is reporting /// a live subagent as wedged — so anything the child says counts, and what /// it said is never read: classifying *what* the subagent is doing is a /// separate question from whether it's doing anything. /// /// Sink methods are called synchronously from the driver's stream readers as /// lines arrive, so the body has to stay cheap — one uncontended map write /// is, and forwarding to a channel to do the same write elsewhere would cost /// more than it saved. `settle` adds a second uncontended lock on a /// `resume`d turn only, and finds an already-emptied slot after the first /// event — strictly less work than the map write next to it. /// /// **Liveness counts every callback; "the turn is underway" does not.** A /// resume that matched nothing is not silent: claude writes the reason to /// stderr *and* emits a terminal stream-json `result` event before exiting, /// so treating any callback at all as proof the turn began would report /// every missed resume as a successful start. The narrowest fact that /// separates the two is the event's own kind — a `result` is stream-json's /// end-of-turn marker, so an event that isn't one is a turn still in /// progress. That's the envelope, not the content: nothing here reads what /// the subagent said. struct LivenessSink { state: Arc, name: String, /// Where to report the first non-terminal event, on a `resume` whose /// caller is still waiting to hear whether it landed. `None` on a /// `start`, which has nobody waiting. verdict: Option, } impl hive_claude::Sink for LivenessSink { fn on_event(&self, event: &serde_json::Value) { self.state.note_event(&self.name); if event.get("type").and_then(serde_json::Value::as_str) != Some("result") { settle(self.verdict.as_ref(), ResumeVerdict::Underway); } } fn on_stdout_line(&self, _line: &str) { self.state.note_event(&self.name); } fn on_stderr_line(&self, _line: &str) { self.state.note_event(&self.name); } } /// Spawn the child (synchronous — `Claude::spawn` returns with a real pid /// the instant the process exists), track it in `running`, and hand the /// actual turn off to a background task so nobody blocks on the whole turn. /// /// **A pid is "confirmed running" for a `start`, and only for a `start`.** /// A `start` creates its session, so the only thing that can go wrong at /// attach time is the spawn itself, which has already either succeeded or /// returned here as an error — there is no stronger signal to wait for, and /// waiting would slow every call down for no reason. A `resume` breaks that /// reasoning: the session it names may not be there, and claude only says so /// a fraction of a second *after* the process exists, so for a `continue` a /// pid is not proof the turn began. That's what `verdict` is for — `Some` /// only on the resume path, carrying the first real answer back to a caller /// that is waiting for it (`await_resume`); `None` on a `start`, which keeps /// the immediate-return contract unchanged. /// /// **The turn *continuation* loop lives in that background task, and only /// there.** A session `start`ed with a goal runs turn after turn until /// something stops it, and every one of those turns is a fresh /// `Claude::spawn` against `Attach::Resume` — the driver's `wait` consumes /// its child, so there is no other shape it could take. Keeping the loop /// behind the same `tokio::spawn` is what leaves both tool-call contracts /// untouched: `start` still returns at the first spawn, `continue` still /// returns when its own turn is underway, and neither waits on turns two /// through five. /// /// Not `async` itself — `tokio::spawn` needs an active runtime to spawn /// *onto*, not an `async` caller to spawn *from*. fn spawn_and_track( state: &Arc, name: &str, config: &Config, attach: &Attach, prompt: String, verdict: Option, ) -> anyhow::Result { let running = Claude::spawn(config, attach) .map_err(|e| anyhow::anyhow!("starting the subagent process failed: {e}"))?; state.track(name, running.cancel_handle()); // This turn supersedes whatever the previous one did, including having // been killed — the record is about the turn before this one. state.clear_kill(name); // Start the liveness clock at the spawn, so the age is already an answer // before the child's first line — see `note_event`. state.note_event(name); // Resolved here, on the calling thread, while the config is still to // hand: only a resume can miss, and only `classify_end` finding a // `SessionNotFound` ever uses it. Every continuation turn is a resume, // so this is worth having even when the *first* attach is a `Create`. let resume_searched = searched_location(config); let mut searched = if matches!(attach, Attach::Resume(_)) { resume_searched.clone() } else { None }; let config = config.clone(); let state = Arc::clone(state); let task_name = name.to_owned(); tokio::spawn(async move { let mut running = running; let mut prompt = prompt; loop { let sink = LivenessSink { state: Arc::clone(&state), name: task_name.clone(), verdict: verdict.clone(), }; let end = classify_end(running.wait(&prompt, &sink).await, searched.as_deref()); log_turn_end(&task_name, &end); // The todo is how a turn's end reaches an agent that is no // longer looking — so it is pushed for every end *except* the // one the caller is being handed as a tool-call error right now. // `settle` saying the verdict was delivered is what makes that // certain: a `continue` whose grace had already run out gets // `false` here and its todo, same as before. Only the first turn // can ever win this — the sender is consumed — which is right, // since only the first turn is one a caller is still waiting on. let reported = settle(verdict.as_ref(), ResumeVerdict::Ended(end.clone())); if !matches!(end, TurnEnd::Complete) { // A killed or failed turn ends the run, goal or not: there is // nothing to re-prompt a child that isn't there any more, and // these two ends already have records of their own. state.finish_turn(&task_name, &end); if !(matches!(end, TurnEnd::Failed(_)) && reported) { push_turn_end_todo(&state.socket, &task_name, &end, None).await; } return; } match state.plan_after_turn(&task_name) { Continuation::Stop(stop) => { state.record_stop(&task_name, stop.clone()); state.finish_turn(&task_name, &end); write_stop_to_report(state.report_file(&task_name), &task_name, &stop).await; push_turn_end_todo(&state.socket, &task_name, &end, Some(&stop)).await; return; } Continuation::Continue { prompt: next } => { // The name stays claimed across the gap — see // `between_turns` for what a concurrent `start` would // otherwise be able to do with it. state.between_turns(&task_name); match Claude::spawn(&config, &Attach::Resume(task_name.clone())) { Ok(next_running) => { state.track(&task_name, next_running.cancel_handle()); state.note_event(&task_name); running = next_running; prompt = next; searched = resume_searched.clone(); } Err(e) => { let end = TurnEnd::Failed(format!( "claude error: starting the next goal turn failed: {e}" )); log_turn_end(&task_name, &end); state.finish_turn(&task_name, &end); push_turn_end_todo(&state.socket, &task_name, &end, None).await; return; } } } } } }); Ok(format!("subagent `{name}` started")) } /// Log a turn's end at the level its severity deserves: a completion is /// unremarkable, the other two are not. fn log_turn_end(name: &str, end: &TurnEnd) { match end { TurnEnd::Complete => {} TurnEnd::Killed { signal } => { tracing::warn!( name = %name, signal, "subagent: turn killed — the child died on a signal, it did not finish" ); } TurnEnd::Failed(e) => { tracing::warn!(name = %name, error = %e, "subagent: turn failed"); } } } /// Append the stop reason to the session's own report file, so the artifact /// a parent already reads is where the run's ending is recorded too — rather /// than the parent having to correlate a todo against a file. /// /// Best-effort and appended, never rewritten: the subagent wrote that file, /// and this adds a line under what it wrote instead of taking a position on /// the rest of it. A path this daemon can't write to is logged and dropped — /// the todo still carries the same sentence, so nothing is only here. async fn write_stop_to_report(path: Option, name: &str, stop: &StopReason) { let Some(path) = path else { return }; let line = format!("\n**Subagent `{name}` stopped:** {}\n", stop.sentence()); let appended = tokio::fs::OpenOptions::new() .create(true) .append(true) .open(&path) .await; let result = match appended { Ok(mut file) => { use tokio::io::AsyncWriteExt as _; // `flush`, not just `write_all`: a `tokio::fs::File` buffers, and // dropping one discards whatever hasn't been handed to the // blocking pool — so without this the line is written to nothing // and the failure is silent. A unit test caught exactly that. match file.write_all(line.as_bytes()).await { Ok(()) => file.flush().await, Err(e) => Err(e), } } Err(e) => Err(e), }; if let Err(e) = result { tracing::warn!( name, path = %path.display(), error = ?e, "subagent: could not record the stop reason in the session's report file", ); } } /// The goal contract appended to a `start`'s first prompt when a goal was /// given. Spelled out to the subagent rather than left implicit: it is about /// to be re-prompted by something it can't see, and the two tools that stop /// that are the only way it has to say "done" or "stuck". fn goal_briefing(name: &str, goal: &str, max_turns: u32) -> String { format!( "\n\nYou are the subagent session `{name}`.\n\nYour goal for this session: {goal}\n\nYou \ have up to {max_turns} turns to reach it. When a turn of yours ends and you haven't \ reported the goal reached, the harness starts another turn re-prompting you toward it. \ Call the `goal_reached` tool once you've genuinely reached it, or `need_help` with what \ is blocking you if you can't proceed — either one stops the re-prompting. Neither takes \ a session name: the endpoint you call them on is yours alone, so they always apply to \ this session and can't be aimed at another one. Running out of turns stops it too, with \ the work left wherever it had got to." ) } /// The prompt a continuation turn opens with. It states the one fact the /// subagent can't observe for itself — that its last turn ended without the /// goal being reported reached — and says plainly that claiming the goal /// isn't the same as reaching it, since being re-prompted is precisely the /// pressure to claim it. fn continuation_prompt(name: &str, goal: &str, turn: u32, max_turns: u32) -> String { format!( "Your previous turn ended and you have not reported the goal reached.\n\nYou are the \ subagent session `{name}`.\n\nGoal: {goal}\n\nThis is turn {turn} of {max_turns}. Carry \ on toward the goal. If you have in fact reached it, call `goal_reached`; if you are \ blocked, call `need_help` with what is blocking you. Neither is a substitute for \ the work: whoever spawned you reads what you actually changed, not what you claim about \ it." ) } /// Record the subagent's own "I have reached the goal" signal and stop its /// turn continuation. Called by the subagent, from inside its own turn, over /// the signal route this daemon hands it (see `build_config`). /// /// **`name` is not a parameter of the tool.** It is whatever the route's /// token resolved to (`State::session_for_signal_token`), so a subagent /// records a stop against its own session because that is the only session /// its URL can reach — not because it addressed the right one. /// /// **This verifies nothing**, and the answer it returns says so to the /// subagent's face. It stops the loop and extends the done message; whether /// the goal was actually reached is a question about the diff and the gate /// output, which the parent reads for itself. /// /// `report_file` is the subagent saying where it wrote its report, which is /// the only reason this daemon ever knows that path — see /// `State::set_report_file`. pub fn goal_reached( state: &State, name: &str, msg: Option, report_file: Option<&str>, ) -> String { signal_stop(state, name, StopReason::GoalReached(msg), report_file); format!( "noted — `{name}`'s goal is recorded as reported reached, so this turn finishes and no \ further goal turn is started. It is recorded as your claim, not as verification: whoever \ spawned you still reads what you changed." ) } /// Record that the subagent can't proceed, and stop its turn continuation. /// The blocking signal the run has otherwise no way to raise: without it a /// stuck subagent would be re-prompted toward a goal it has already told /// nobody it can't reach, until the turn cap. /// /// `msg` is required, unlike `goal_reached`'s — "I'm stuck" with no reason /// gives the parent nothing to act on, and acting on it is the entire point. /// /// `name` comes from the route's token, exactly as in [`goal_reached`]. pub fn need_help(state: &State, name: &str, msg: String, report_file: Option<&str>) -> String { signal_stop(state, name, StopReason::NeedHelp(msg), report_file); format!( "noted — `{name}` is recorded as blocked, so this turn finishes and no further goal turn \ is started. Write down what you have done so far where your brief told you to; whoever \ spawned you sees the block in `status` and in this run's todo." ) } /// The half [`goal_reached`] and [`need_help`] share: remember where the /// subagent says it wrote, and record the stop. /// /// It checks nothing, and has nothing left to check. `name` reached it by /// being resolved from the route's own token, so "is this signal about the /// session it claims" is answered before the request is dispatched at all — /// a token that names no live session never reaches this function, it gets a /// 404. What used to stand here was an `occupancy()` liveness check standing /// in for identity, which two concurrently-running siblings could each /// satisfy for the other's name. fn signal_stop(state: &State, name: &str, stop: StopReason, report_file: Option<&str>) { state.set_report_file(name, report_file); state.record_stop(name, stop); } /// Report whether `name` is currently running — a zero-cost check that /// never launches a process, unlike `continue`. Distinguishes running, /// starting (reserved, not yet a confirmed spawn — see `State`'s doc), /// killed (its last turn died on a signal), each of the four ways a goal /// run stops, idle (a session exists, its last turn finished, nothing is in /// flight), and no such session at all. /// /// A *running* answer also carries how long since that turn last produced /// output, which is the part of this answer a caller can act on: "running" /// describes a wedged child and a busy one identically, and the age /// separates them (see `State`'s `last_event` doc). /// /// A goal-continued session carries `turn N of M` alongside that, in every /// state. With the age, it is what lets a caller tell *working* from /// *wedged* from *out of turns* off one answer, without reaching for `ps` /// or reading any file. /// /// # Errors /// /// An invalid name, or no session — running, killed, stopped or on disk — /// under `name`. pub fn status(state: &State, name: &str, dir: Option<&str>) -> anyhow::Result { validate_name(name)?; // Read-only: an explicit `dir` here is a one-off "check this other // directory's session" per this fn's own doc, not a new remembered // default — `peek_dir` resolves the same way but never writes. let dir = state.peek_dir(name, dir); let facts = StatusFacts { occupancy: state.occupancy(name), killed: state.killed_by(name), last_event_age: state.last_event_age(name), turns: state.turns(name), stop: state.stop_reason(name), ..StatusFacts::new(name) }; // Nothing on disk to look for while something is tracked in memory — // those states answer on their own, and the store read is the only // expensive part of this call. A recorded stop reason counts: it is // proof this daemon ran the session, which is what the lookup asks. let session_exists = if facts.occupancy.is_none() && facts.killed.is_none() && facts.stop.is_none() { // No signal surface in this config: it exists only to resolve the // session store, and rendering a subagent's MCP config for a // read-only status check would be writing a file for nobody. let config = build_config(name, None, None, None, dir.as_deref(), None); build_store(&config)?.find_by_title(name).is_some() } else { false }; describe_status(&StatusFacts { session_exists, ..facts }) } /// The liveness sentence appended to a *running* answer, and the whole point /// of recording the timestamp: an age in seconds a caller can read a verdict /// off directly, rather than one it has to go and measure itself with `ps` /// and CPU-time deltas. /// /// Empty when there's no age to report — a turn that ended between `status` /// reading `running` and reading the clock. Saying nothing is right there: /// the alternative is an age that describes a turn which has already /// finished, which reads as a stall that never happened. fn describe_liveness(age: Option) -> String { match age { None => String::new(), Some(age) => format!( " Last event {}s ago — how long since this turn's claude process produced any output \ at all, whatever it was: a few seconds means it's working, an age that keeps \ climbing into the minutes means it's wedged. It resets at each turn's spawn, so on a \ goal run it describes the turn in flight, not the run.", age.as_secs() ), } } /// The progress sentence a goal-continued session carries in every state: /// which turn of its budget it is on. Empty for a session with no goal, /// which has no budget to be partway through. /// /// On a *running* answer this is the turn in flight; on a stopped one it is /// the turn it stopped on — the same number either way, since the counter /// only advances when the loop decides to spend another turn. fn describe_turns(turns: Option<(u32, u32)>) -> String { match turns { None => String::new(), Some((turn, max_turns)) => format!(" Turn {turn} of {max_turns}."), } } /// Everything `status` gathers, as one value — seven separate parameters /// read as noise at both the call site and the test sites, and every one of /// them is a fact about the same session. struct StatusFacts<'a> { name: &'a str, occupancy: Option, killed: Option, session_exists: bool, last_event_age: Option, turns: Option<(u32, u32)>, stop: Option, } impl<'a> StatusFacts<'a> { /// The nothing-known baseline for `name`: no process, no records, no /// session. Both the real gathering in `status` and the tests build on /// it, so neither has to spell out the fields it isn't exercising. fn new(name: &'a str) -> Self { Self { name, occupancy: None, killed: None, session_exists: false, last_event_age: None, turns: None, stop: None, } } } /// Render `status`'s answer from the facts it gathers. Split out from the /// gathering so the killed-versus-idle distinction — and the liveness age, /// the turn counter and each stop reason — are exercisable without a real /// spawn, a real signal, a real stream of events and a real on-disk claude /// session. /// /// A recorded kill outranks both the stop reason and the on-disk session: /// the session file exists either way and a stop reason may be left over /// from the signal a subagent raised just before something killed it, so /// neither can tell a killed turn from a finished one. The kill can. /// /// Every answer is self-contained — it names the one state the caller got /// and what to do next — because the tool description deliberately doesn't /// enumerate the state space (the operator's ruling on this surface: /// describe the tool, explain the state when returning it). Keep the split: /// a terser answer here has nowhere left to be explained from. fn describe_status(facts: &StatusFacts<'_>) -> anyhow::Result { let name = facts.name; let turns = describe_turns(facts.turns); match facts.occupancy { Some(true) => { return Ok(format!( "subagent `{name}` is running — its turn is still in flight, so there's nothing \ to do but let it work: the daemon pushes a todo when the run ends, or \ `interrupt` it if you want it stopped early.{turns}{}{}", describe_liveness(facts.last_event_age), describe_pending_signal(facts.stop.as_ref()), )); } Some(false) => { return Ok(format!( "subagent `{name}` is starting — the name is claimed but no process is confirmed \ under it yet, either because a `start`/`continue` hasn't spawned one or because \ a goal run is between turns. Normally over in well under a second: check again \ shortly rather than starting anything else under this name.{turns}" )); } None => {} } if let Some(signal) = facts.killed { return Ok(format!( "subagent `{name}` was killed — its last turn died on {}, so its work stopped \ wherever it had got to rather than finishing. `continue` still resumes it, but \ whatever it was told to do is unfinished: check what it actually left behind before \ trusting it.{turns}", describe_signal(signal) )); } if let Some(stop) = &facts.stop { return Ok(describe_stopped(name, stop, &turns)); } if facts.session_exists { Ok(format!( "subagent `{name}` is idle — its session exists, its last turn ended on its own \ rather than being cut off, and nothing is in flight: that turn's own todo says how \ it went, and `continue` gives it another.{turns}" )) } else { anyhow::bail!( "no subagent named `{name}` exists — nothing is running under that name and there's \ no session on disk to resume, so `start` is what creates one (check the name if you \ expected something here)." ) } } /// The note a *running* answer carries when the subagent has already raised /// a stop signal for the turn still in flight. Without it a parent polling /// `status` would read plain "running" for the whole stretch between the /// subagent saying it is blocked and its turn actually ending — the one /// stretch where "it's working, leave it alone" is the wrong conclusion. fn describe_pending_signal(stop: Option<&StopReason>) -> String { match stop { None | Some(StopReason::Done | StopReason::TurnCap { .. }) => String::new(), Some(stop) => format!( " It has already signalled how this run ends — {} — so this is its last turn.", stop.sentence() ), } } /// The answer for a session whose run has stopped, one per [`StopReason`]. /// Each names the state, why the continuation stopped, and what `continue` /// would do about it — and the `GoalReached` one is deliberately the least /// reassuring of the four, because it is the one a caller is most likely to /// read as "finished successfully" when it means "said so". fn describe_stopped(name: &str, stop: &StopReason, turns: &str) -> String { match stop { StopReason::Done => format!( "subagent `{name}` is idle — its session exists, its last turn ended on its own \ rather than being cut off, and nothing is in flight: that turn's own todo says how \ it went, and `continue` gives it another.{turns}" ), StopReason::GoalReached(msg) => format!( "subagent `{name}` stopped: it reported its goal reached{}. Nothing is in flight and \ no further goal turn will start.{turns} That report is the subagent's own claim, not \ a verification of anything — read the diff and whatever gate the work was supposed \ to pass before you treat the goal as met, exactly as you would a build report saying \ the tests passed.", msg.as_deref() .map_or_else(String::new, |m| format!(": {m}")) ), StopReason::NeedHelp(msg) => format!( "subagent `{name}` is BLOCKED and needs help: {msg}. It called `need_help`, which \ stopped its goal continuation, and nothing is in flight — it stays blocked until you \ answer it.{turns} `continue` is how you answer: give it what it asked for as the \ next turn's prompt." ), StopReason::TurnCap { turns: spent } => format!( "subagent `{name}` ran out of turns — the harness limit of {spent} was reached and it \ never reported its goal reached, so the work stopped wherever it had got to rather \ than finishing.{turns} Check what it actually left behind; `continue` gives it a \ fresh allowance if carrying on is worth it." ), } } /// Signal `name`'s running process — `force` picks SIGKILL over SIGINT (see /// `hive_claude::Cancel::cancel`). Refuses a name with nothing running: no /// entry at all, or one still in the brief not-yet-spawned window (nothing /// to signal yet — the reservation is put back so a concurrent /// `start`/`continue` for the same name still gets refused). /// /// This stops a goal run, not just the turn in it: the signalled child ends /// as `TurnEnd::Killed`, which the continuation loop treats as the end of /// the whole run rather than something to re-prompt past. That falls out of /// there being no child left to continue, and it is the answer you want — /// `interrupt` would be useless if the harness immediately started turn /// three of five. /// /// # Errors /// /// An invalid name, nothing tracked under `name`, or `name` has no confirmed /// process right now (a spawn in flight, or a goal run between turns). pub fn interrupt(state: &State, name: &str, force: bool) -> anyhow::Result { validate_name(name)?; let mut running = state.running.lock().unwrap_or_else(PoisonError::into_inner); match running.remove(name) { None => anyhow::bail!("no subagent named `{name}` is currently running"), Some(None) => { running.insert(name.to_owned(), None); anyhow::bail!( "subagent `{name}` is still starting — not yet confirmed running, try again \ shortly" ); } Some(Some(cancel)) => { drop(running); cancel.cancel(force); Ok(format!("interrupt sent to subagent `{name}`")) } } } /// Push `name`'s one-shot end-of-turn todo. Best-effort: a connect/write /// failure is logged and swallowed, matching every other in-agent-socket /// producer in this codebase — there's no retry queue to fall back to, and /// the caller has already moved on by the time this fires. async fn push_turn_end_todo( socket: &std::path::Path, name: &str, end: &TurnEnd, stop: Option<&StopReason>, ) { let req = hive_agent_sock::Request::UpsertTodo { subsystem: "subagent".to_owned(), key: Some(name.to_owned()), summary: turn_end_summary(name, end, stop), source: None, reopen_if_acked: false, }; if let Err(e) = hive_sock_client::notify(socket, &req, hive_sock_client::Retry::None).await { tracing::warn!(name, error = ?e, "subagent: end-of-turn todo push failed"); } } /// The todo text for a run that has ended. A killed turn deliberately does /// not use the "finished" wording the other two share: this todo is the only /// thing the owning agent is shown without asking, so it has to read as the /// interruption it is rather than as one more completed subagent. /// /// A `stop` **extends** that message, it never replaces it. Both halves are /// load-bearing and neither substitutes for the other: the turn's own end is /// what the daemon observed, the stop reason is why the run went no further /// — and for `GoalReached` that second half is a claim, which would read as /// a verdict if it were allowed to stand where the observed end belongs. /// `StopReason::Done` adds nothing, since "there was no goal" is exactly /// what the unextended message already describes. fn turn_end_summary(name: &str, end: &TurnEnd, stop: Option<&StopReason>) -> String { let base = match end { TurnEnd::Complete => format!("subagent `{name}` finished: turn complete"), TurnEnd::Failed(e) => format!("subagent `{name}` finished: {e}"), TurnEnd::Killed { signal } => format!( "subagent `{name}` was KILLED mid-turn ({}) — it did not finish, and its work stopped \ wherever it had got to. Check what it left behind before you act on it; `continue` \ resumes the session if you want it carried on.", describe_signal(*signal) ), }; match stop { None | Some(StopReason::Done) => base, Some(stop) => format!("{base} — and the run stopped there: {}", stop.sentence()), } } #[cfg(test)] mod tests { use super::*; // `Cancel` can only be constructed from a real spawned process (no test // fixture in `hive_claude` for it), so these exercise the reservation // half of `State` directly — the actual TOCTOU-closure logic — rather // than the full start/spawn path. /// A stand-in for the signal route a real daemon would hand its /// subagents — the *prefix*, as `State::new` takes it. Nothing in these /// tests dials it: what `State` does with it is append a minted token and /// carry the result into `build_config`, both asserted on directly. fn signal_url() -> String { "http://127.0.0.1:1/signal/mcp".to_owned() } /// The token out of a minted URL — what the route would have parsed out /// of the path before resolving it. fn token_of(url: &str) -> String { url.rsplit('/') .next() .expect("a minted URL always has a last segment") .to_owned() } /// A `StartRequest` with only the fields a test cares about set — the /// other six are the same "nothing asked for" every time. fn start_request(name: &str) -> StartRequest { StartRequest { name: name.to_owned(), model: None, effort: None, prompt_file: "/tmp/prompt.md".to_owned(), trigger: "trigger".to_owned(), dir: None, goal: None, max_turns: None, report_file: None, } } #[test] fn reserve_is_exclusive_for_the_same_name() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert!(state.reserve("dup"), "first reservation should succeed"); assert!( !state.reserve("dup"), "a second reservation for the same name must be refused — this is the exact race \ argus found: two calls both passing a check before either commits" ); } #[test] fn reserve_does_not_cross_block_different_names() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert!(state.reserve("a")); assert!( state.reserve("b"), "unrelated names must not block each other" ); } #[test] fn release_reservation_frees_the_name_for_reuse() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert!(state.reserve("n")); state.release_reservation("n"); assert!( state.reserve("n"), "releasing a still-`None` reservation must free the name again" ); } #[test] fn occupancy_reflects_the_reserved_but_not_running_state() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert_eq!(state.occupancy("never-reserved"), None); state.reserve("n"); assert_eq!( state.occupancy("n"), Some(false), "reserved-but-not-yet-spawned must read as occupied-but-not-running" ); } // One test, not two: `subagent_otel_attrs` reads the real process-wide // `OTEL_RESOURCE_ATTRIBUTES` env var, and cargo runs tests in parallel // threads by default — two separate tests each mutating that global // raced each other (and, in this container, lost to the agent's own // real ambient value). Sequencing both assertions in one test removes // the race instead of papering over it with a mutex. #[test] fn otel_attrs_append_ambient_value_or_stand_alone() { // SAFETY: test-only env mutation; sequenced within this one test so // no other test's concurrent read/write of the same var can race it. unsafe { std::env::remove_var("OTEL_RESOURCE_ATTRIBUTES"); } assert_eq!(subagent_otel_attrs("batch-1"), "subagent=batch-1"); unsafe { std::env::set_var("OTEL_RESOURCE_ATTRIBUTES", "agent=damocles"); } assert_eq!( subagent_otel_attrs("batch-1"), "agent=damocles,subagent=batch-1" ); unsafe { std::env::remove_var("OTEL_RESOURCE_ATTRIBUTES"); } } #[test] fn build_config_only_appends_system_prompt_when_given() { let with = build_config("n", None, None, Some("/tmp/p.md"), None, None); assert!( with.extra_args .contains(&"--append-system-prompt-file".to_owned()) ); assert!(with.extra_args.contains(&"/tmp/p.md".to_owned())); let without = build_config("n", None, None, None, None, None); assert!( !without .extra_args .contains(&"--append-system-prompt-file".to_owned()) ); } #[test] fn resolve_dir_remembers_an_explicit_dir_and_falls_back_to_it_when_omitted() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert_eq!( state.resolve_dir("n", None), None, "nothing remembered yet — omitted dir has nothing to fall back to" ); assert_eq!( state.resolve_dir("n", Some("/tmp/worktree")), Some("/tmp/worktree".to_owned()), "an explicit dir is returned as-is" ); assert_eq!( state.resolve_dir("n", None), Some("/tmp/worktree".to_owned()), "a later omitted dir falls back to what was just remembered" ); assert_eq!( state.resolve_dir("n", Some("/tmp/other")), Some("/tmp/other".to_owned()), "a later explicit dir overrides the memory, not just reads it" ); assert_eq!( state.resolve_dir("n", None), Some("/tmp/other".to_owned()), "the fallback now reflects the override" ); } #[test] fn resolve_dir_does_not_cross_names() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.resolve_dir("a", Some("/tmp/a")); assert_eq!( state.resolve_dir("b", None), None, "an unrelated name's memory must stay empty" ); } #[test] fn peek_dir_resolves_like_resolve_dir_but_never_writes() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.resolve_dir("n", Some("/tmp/remembered")); assert_eq!( state.peek_dir("n", Some("/tmp/one-off")), Some("/tmp/one-off".to_owned()), "an explicit dir still resolves as given" ); assert_eq!( state.resolve_dir("n", None), Some("/tmp/remembered".to_owned()), "but the one-off peek must not have overwritten the remembered value" ); } #[test] fn start_does_not_corrupt_the_remembered_dir_when_the_name_is_already_running() { // The exact scenario argus caught in review: a rejected concurrent // `start` used to commit its `dir` before the reservation check // refused it, corrupting what a later bare `continue` would resolve // to. Reordering `reserve` before `resolve_dir` in `start` closes it // — a name that's already reserved must never reach `resolve_dir` at // all, so `dirs` stays exactly as a caller left it. let state = Arc::new(State::new(PathBuf::from("/dev/null"), signal_url())); state.resolve_dir("dup", Some("/tmp/original")); assert!( state.reserve("dup"), "simulate an in-flight start/continue already owning the name" ); let result = start( &state, StartRequest { dir: Some("/tmp/rejected".to_owned()), ..start_request("dup") }, ); assert!( result.is_err(), "a name already reserved must be refused, not spawned" ); assert_eq!( state.resolve_dir("dup", None), Some("/tmp/original".to_owned()), "the rejected call's dir must not have overwritten the remembered one" ); } /// The error the driver hands back for a child that died on `signal` — /// `ExitStatus::from_raw` takes a raw `wait(2)` status, whose low seven /// bits are the terminating signal, so this is the same value /// `RunningClaude::wait` would have produced from a real kill. fn signalled_exit(signal: i32) -> hive_claude::Error { hive_claude::Error::Exit { status: std::process::ExitStatus::from_raw(signal), stderr_tail: String::new(), } } #[test] fn classify_end_separates_a_signalled_child_from_a_clean_one() { assert_eq!(classify_end(Ok(()), None), TurnEnd::Complete); assert_eq!( classify_end(Err(signalled_exit(libc::SIGKILL)), None), TurnEnd::Killed { signal: 9 }, "a SIGKILLed child must not read as a turn that ended on its own" ); assert_eq!( classify_end(Err(signalled_exit(libc::SIGTERM)), None), TurnEnd::Killed { signal: 15 } ); assert!( matches!( classify_end(Err(hive_claude::Error::PromptTooLong), None), TurnEnd::Failed(_) ), "claude exiting on its own is a failure, not a kill" ); } #[test] fn status_reports_killed_not_idle_for_a_signalled_session() { // Both sessions exist on disk and neither is running: the only thing // that can tell them apart is the recorded signal. let killed = describe_status(&StatusFacts { killed: Some(libc::SIGKILL), session_exists: true, ..StatusFacts::new("n") }) .expect("killed status"); let idle = describe_status(&StatusFacts { session_exists: true, ..StatusFacts::new("n") }) .expect("idle status"); assert!( killed.contains("killed") && killed.contains("SIGKILL (signal 9)"), "a killed session must say so, and name the signal: {killed}" ); assert!( !killed.contains("idle"), "a killed session must not also read as idle: {killed}" ); assert!(idle.contains("idle"), "a finished turn still reads idle"); assert_ne!( killed, idle, "collapsing the two back together is the bug this reports" ); } #[test] fn every_status_answer_names_its_state_and_the_next_move() { // The tool description no longer lists the states, so each answer // has to carry its own explanation — checked one state at a time, // which is all a caller ever gets back. let running = describe_status(&StatusFacts { occupancy: Some(true), ..StatusFacts::new("n") }) .expect("running status"); assert!( running.contains("running") && running.contains("`interrupt`"), "a running answer must say the turn is in flight and how to stop it: {running}" ); let starting = describe_status(&StatusFacts { occupancy: Some(false), ..StatusFacts::new("n") }) .expect("starting status"); assert!( starting.contains("starting") && starting.contains("check again"), "a starting answer must say the spawn isn't confirmed yet and to retry: {starting}" ); let idle = describe_status(&StatusFacts { session_exists: true, ..StatusFacts::new("n") }) .expect("idle status"); assert!( idle.contains("idle") && idle.contains("`continue`"), "an idle answer must say the last turn ended on its own and how to give it another: \ {idle}" ); let killed = describe_status(&StatusFacts { killed: Some(libc::SIGKILL), session_exists: true, ..StatusFacts::new("n") }) .expect("killed status"); assert!( killed.contains("killed") && killed.contains("`continue`"), "a killed answer must name the kill and say resuming is still possible: {killed}" ); let missing = describe_status(&StatusFacts::new("n")) .expect_err("nothing tracked and nothing on disk is an error, not a state"); assert!( missing.to_string().contains("`start`"), "the no-such-session answer must point at what creates one: {missing}" ); } #[test] fn a_killed_turn_and_a_clean_one_do_not_land_in_the_same_state() { // The whole path a real turn takes, minus the process: what the // driver returned -> what the daemon records -> what `status` says. let state = State::new(PathBuf::from("/dev/null"), signal_url()); for (name, outcome) in [ ("gone", Err(signalled_exit(libc::SIGKILL))), ("done", Ok(())), ] { assert!(state.reserve(name)); state.finish_turn(name, &classify_end(outcome, None)); assert_eq!( state.occupancy(name), None, "the turn is over either way — nothing stays tracked as running" ); } assert_eq!(state.killed_by("gone"), Some(9)); assert_eq!(state.killed_by("done"), None); let gone = describe_status(&StatusFacts { killed: state.killed_by("gone"), session_exists: true, ..StatusFacts::new("gone") }) .expect("status"); let done = describe_status(&StatusFacts { killed: state.killed_by("done"), session_exists: true, ..StatusFacts::new("done") }) .expect("status"); assert!(gone.contains("killed"), "{gone}"); assert!(done.contains("idle"), "{done}"); } #[test] fn a_new_turn_clears_the_previous_turn_s_kill() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.finish_turn("n", &TurnEnd::Killed { signal: 9 }); assert_eq!(state.killed_by("n"), Some(9)); // What `spawn_and_track` does once the next turn is confirmed // spawned, and then what its own clean end does. state.clear_kill("n"); assert_eq!(state.killed_by("n"), None); state.finish_turn("n", &TurnEnd::Complete); assert_eq!( state.killed_by("n"), None, "a turn that finished must not leave the old kill standing" ); } #[test] fn the_killed_todo_does_not_read_like_a_completion() { let complete = turn_end_summary("n", &TurnEnd::Complete, None); let killed = turn_end_summary("n", &TurnEnd::Killed { signal: 9 }, None); assert_eq!( complete, "subagent `n` finished: turn complete", "the ordinary completion todo is unchanged" ); assert!( killed.contains("KILLED mid-turn") && killed.contains("SIGKILL (signal 9)"), "the killed todo must state the kill and the signal: {killed}" ); assert!( !killed.contains("finished"), "the one notification pushed without being asked must not read as a finished turn: \ {killed}" ); assert_ne!(complete, killed); } #[test] fn continue_tells_the_caller_it_is_resuming_a_killed_session() { let plain = note_resumed_after_kill("subagent `n` started", None); assert_eq!( plain, "subagent `n` started", "an ordinary resume is unchanged" ); let after_kill = note_resumed_after_kill("subagent `n` started", Some(libc::SIGKILL)); assert!( after_kill.contains("killed") && after_kill.contains("SIGKILL (signal 9)"), "resuming a killed session is allowed, but the caller has to be told: {after_kill}" ); } #[test] fn describe_signal_names_the_ones_that_end_a_subagent() { assert_eq!(describe_signal(9), "SIGKILL (signal 9)"); assert_eq!(describe_signal(15), "SIGTERM (signal 15)"); assert_eq!(describe_signal(2), "SIGINT (signal 2)"); assert_eq!(describe_signal(7), "signal 7"); } #[test] fn build_config_sets_cwd_only_when_a_dir_is_given() { let with = build_config("n", None, None, None, Some("/tmp/some-worktree"), None); assert_eq!(with.cwd, Some(PathBuf::from("/tmp/some-worktree"))); let without = build_config("n", None, None, None, None, None); assert_eq!(without.cwd, None); } #[test] fn build_config_sets_effort_alongside_model() { let config = build_config( "n", Some("opus".to_owned()), Some("high".to_owned()), None, None, None, ); assert_eq!(config.model, Some("opus".to_owned())); assert_eq!(config.effort, Some("high".to_owned())); } #[test] fn an_event_starts_the_liveness_clock_and_the_turn_ending_stops_it() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert_eq!( state.last_event_age("n"), None, "a name that never spawned has no clock to read" ); state.note_event("n"); assert!( state.last_event_age("n").is_some(), "the first event must give `status` an age to report" ); state.finish_turn("n", &TurnEnd::Complete); assert_eq!( state.last_event_age("n"), None, "the age describes a turn in flight — a finished turn's leftover age would read as a \ stall that never happened" ); } #[test] fn a_later_event_resets_the_age_rather_than_letting_it_climb() { // The distinction the whole record exists for: a child still // producing output must not accumulate the age of a wedged one. let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.note_event("n"); std::thread::sleep(Duration::from_millis(20)); let before = state.last_event_age("n").expect("clock started"); state.note_event("n"); let after = state.last_event_age("n").expect("clock still running"); assert!( after < before, "a fresh event must reset the age, not extend it: {after:?} vs {before:?}" ); } #[test] fn the_liveness_clock_does_not_cross_names() { let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.note_event("a"); assert_eq!( state.last_event_age("b"), None, "one subagent's output says nothing about another's" ); } #[test] fn every_sink_callback_counts_as_liveness() { // All three, deliberately: a stderr line or a non-JSON stdout line is // proof the child is alive exactly as much as a stream-json event, // and reporting a live subagent as wedged is the failure that costs. use hive_claude::Sink as _; fn fresh() -> (Arc, LivenessSink) { let state = Arc::new(State::new(PathBuf::from("/dev/null"), signal_url())); let sink = LivenessSink { state: Arc::clone(&state), name: "n".to_owned(), verdict: None, }; (state, sink) } let (state, sink) = fresh(); sink.on_event(&serde_json::json!({"type": "system"})); assert!( state.last_event_age("n").is_some(), "a stream-json event is liveness" ); let (state, sink) = fresh(); sink.on_stdout_line("not json"); assert!( state.last_event_age("n").is_some(), "so is a stdout line that didn't parse as JSON" ); let (state, sink) = fresh(); sink.on_stderr_line("some chatter"); assert!( state.last_event_age("n").is_some(), "so is a stderr line — anything the child says at all counts" ); } #[test] fn a_running_status_reports_the_age_and_an_idle_one_does_not() { let running = describe_status(&StatusFacts { occupancy: Some(true), last_event_age: Some(Duration::from_secs(4)), ..StatusFacts::new("n") }) .expect("running status"); assert!( running.contains("Last event 4s ago"), "a running answer must carry the age — the one part of it a caller can act on: \ {running}" ); let wedged = describe_status(&StatusFacts { occupancy: Some(true), last_event_age: Some(Duration::from_mins(15)), ..StatusFacts::new("n") }) .expect("running status"); assert!( wedged.contains("Last event 900s ago"), "the wedged case is the one this exists for: {wedged}" ); assert_ne!( running, wedged, "4s-ago and 900s-ago must not render identically — collapsing them back together is \ the bug this reports" ); let idle = describe_status(&StatusFacts { session_exists: true, ..StatusFacts::new("n") }) .expect("idle status"); assert!( !idle.contains("Last event"), "an idle session has no in-flight turn whose progress an age would describe: {idle}" ); } #[test] fn a_running_status_still_answers_when_the_turn_ended_mid_call() { // `status` reads `running` and the clock under separate locks, so a // turn can finish between the two. The answer drops the age rather // than inventing one. let raced = describe_status(&StatusFacts { occupancy: Some(true), ..StatusFacts::new("n") }) .expect("running status"); assert!( raced.contains("is running") && !raced.contains("Last event"), "a missing age must cost the sentence, not the answer: {raced}" ); } #[test] fn a_missed_resume_names_the_directory_that_was_searched() { // The failure that cost an afternoon: the session existed, just not // where this daemon looked. claude's own message names the value it // failed to match but never the directory, which is the fact that // was missing. let searched = "(searched /home/agent/.claude for cwd /home/agent/work; if it was started \ elsewhere, pass the `dir` it was started in)"; let end = classify_end(Err(hive_claude::Error::SessionNotFound), Some(searched)); let TurnEnd::Failed(msg) = end else { panic!("a missed resume is a failed turn"); }; assert!( msg.contains("no session matched"), "claude's own diagnosis must survive: {msg}" ); assert!( msg.contains("/home/agent/.claude") && msg.contains("/home/agent/work"), "and the search location must be appended to it: {msg}" ); assert!( msg.contains("pass the `dir` it was started in"), "naming the directory is only half of it — say what to do about it: {msg}" ); } #[test] fn only_a_missed_resume_gets_the_search_location() { // A `start` never passes one (nothing to resume), and an unrelated // failure must not be decorated with a location that had no part in // it. let no_hint = classify_end(Err(hive_claude::Error::SessionNotFound), None); assert_eq!( no_hint, TurnEnd::Failed( "claude error: no session matched the requested id or title".to_owned() ), "with no resolvable location the underlying error still reaches the caller, plain" ); let unrelated = classify_end(Err(hive_claude::Error::PromptTooLong), Some("(searched …)")); let TurnEnd::Failed(msg) = unrelated else { panic!("an overflowed context is a failed turn"); }; assert!( !msg.contains("searched"), "a context overflow has nothing to do with where the session lives: {msg}" ); } #[test] fn searched_location_names_the_same_place_the_store_would_look() { // The point of reading both halves off the `Config` rather than // re-deriving them: the location reported must be the one actually // searched, so it can't drift from `build_store`/the driver. let cwd = std::env::current_dir().expect("a cwd"); let config = build_config("n", None, None, None, Some(&cwd.to_string_lossy()), None); let Some(located) = searched_location(&config) else { // No `HOME` in this environment — nothing to compare against. return; }; let home = config.resolved_claude_home().expect("home resolved"); let resolved_cwd = config.resolved_cwd().expect("cwd resolved"); assert!( located.contains(&home.display().to_string()) && located.contains(&resolved_cwd.display().to_string()), "both halves must come from the config's own resolution: {located}" ); } /// The `Failed` end a real missed resume produces, message and all — /// `classify_end`'s own output for the error the driver raises when /// `--resume` matches nothing. fn missed_resume() -> TurnEnd { classify_end( Err(hive_claude::Error::SessionNotFound), Some( "(searched /home/agent/.claude for cwd /home/agent/work; if it was started \ elsewhere, pass the `dir` it was started in)", ), ) } #[tokio::test] async fn a_continue_whose_resume_missed_is_an_error_not_a_started_message() { // The whole point of the bounded wait: a `continue` naming a session // that isn't there must fail the tool call, not answer "started" and // leave the real answer to a todo nobody is waiting on any more. let (tx, rx) = oneshot::channel(); tx.send(ResumeVerdict::Ended(missed_resume())) .map_err(|_| ()) .expect("the waiter is still listening"); let err = await_resume(rx) .await .expect_err("a missed resume must reach the caller as an error"); let msg = err.to_string(); assert!( msg.contains("no session matched"), "claude's own diagnosis is the error: {msg}" ); assert!( msg.contains("pass the `dir` it was started in"), "and the searched location travels with it: {msg}" ); } #[tokio::test(start_paused = true)] async fn a_turn_that_got_underway_answers_without_waiting_out_the_grace() { // A successful `continue` must not pay a fixed delay — it returns on // the first sign of the turn, not on the bound. Paused time makes // that measurable: the clock only advances if something awaits it. let (tx, rx) = oneshot::channel(); tx.send(ResumeVerdict::Underway) .map_err(|_| ()) .expect("the waiter is still listening"); let started = tokio::time::Instant::now(); await_resume(rx) .await .expect("an underway turn is not an error"); assert!( started.elapsed() < RESUME_GRACE, "a turn that spoke must settle the wait immediately, not on the bound: {:?}", started.elapsed() ); } #[tokio::test(start_paused = true)] async fn a_turn_that_neither_speaks_nor_exits_is_reported_as_started() { // The bound is a floor under a child that does nothing at all. The // sender is held open for the whole wait, so only the timeout can // end it. let (tx, rx) = oneshot::channel::(); let started = tokio::time::Instant::now(); await_resume(rx) .await .expect("a silent child is not a failed resume"); assert!( started.elapsed() >= RESUME_GRACE, "the wait must actually run to the bound: {:?}", started.elapsed() ); drop(tx); } #[tokio::test] async fn an_early_end_that_is_not_a_failure_still_reads_as_started() { // A turn that completed, or that a concurrent `interrupt` killed, // inside the grace: both ended a turn that genuinely began, and both // have a todo of their own to explain themselves. for end in [TurnEnd::Complete, TurnEnd::Killed { signal: 9 }] { let (tx, rx) = oneshot::channel(); tx.send(ResumeVerdict::Ended(end.clone())) .map_err(|_| ()) .expect("the waiter is still listening"); assert!( await_resume(rx).await.is_ok(), "{end:?} ended a turn that started — only a failure is the caller's error" ); } } #[test] fn only_the_first_verdict_is_reported_and_only_to_a_live_caller() { // `settle`'s two jobs: the sink and the background task race for one // sender, and the winner's `true` is what tells the task the caller // has already been handed the failure. let (tx, rx) = oneshot::channel(); let verdict: VerdictTx = Arc::new(Mutex::new(Some(tx))); assert!( settle(Some(&verdict), ResumeVerdict::Underway), "the first report reaches a listening caller" ); assert!( !settle(Some(&verdict), ResumeVerdict::Ended(missed_resume())), "the loser of the race has nothing left to send on — and so must still push its todo" ); drop(rx); let (tx, rx) = oneshot::channel(); let verdict: VerdictTx = Arc::new(Mutex::new(Some(tx))); drop(rx); assert!( !settle(Some(&verdict), ResumeVerdict::Ended(missed_resume())), "a caller whose grace already expired is not listening, so its todo must still be \ pushed" ); assert!( !settle(None, ResumeVerdict::Underway), "a `start` has no verdict channel and nobody waiting on one" ); } #[test] fn a_terminal_result_event_is_liveness_but_not_proof_the_turn_began() { // Measured, and the reason the underway signal reads the event's // kind at all: a resume that matched nothing still emits a // stream-json `result` event (and stderr chatter) before exiting, so // treating any callback as "underway" would report every missed // resume as a successful start. use hive_claude::Sink as _; fn fresh() -> (oneshot::Receiver, LivenessSink) { let (tx, rx) = oneshot::channel(); let sink = LivenessSink { state: Arc::new(State::new(PathBuf::from("/dev/null"), signal_url())), name: "n".to_owned(), verdict: Some(Arc::new(Mutex::new(Some(tx)))), }; (rx, sink) } let (mut rx, sink) = fresh(); sink.on_event(&serde_json::json!({"type": "result", "is_error": true})); assert!( rx.try_recv().is_err(), "the turn's own end-of-turn marker must not settle the wait as underway" ); assert!( sink.state.last_event_age("n").is_some(), "it is still liveness — the child did say something" ); let (mut rx, sink) = fresh(); sink.on_stderr_line("Error: --resume requires a valid session ID or session title"); assert!( rx.try_recv().is_err(), "nor does stderr, which a missed resume writes to before it exits" ); let (mut rx, sink) = fresh(); sink.on_event(&serde_json::json!({"type": "system", "subtype": "init"})); assert!( matches!(rx.try_recv(), Ok(ResumeVerdict::Underway)), "a non-terminal event is a turn in progress, which is the signal continue waits for" ); } #[test] fn build_config_defaults_effort_to_medium_when_omitted() { // Unlike `model`, an omitted `effort` does not fall through to // claude's own default (`high` on most models) — this daemon picks // `medium` itself, a deliberate cost-conscious choice for subagent // work. let config = build_config("n", None, None, None, None, None); assert_eq!(config.effort, Some("medium".to_owned())); } // ---- turn continuation ------------------------------------------------- /// A state with `name` mid-run against `goal`, as `start` would have left /// it: reserved, goal registered, on turn one of `max_turns`. fn mid_run(name: &str, goal: &str, max_turns: u32) -> State { let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.reserve(name); state.set_goal(name, Some(goal.to_owned()), max_turns); state } #[test] fn a_session_with_no_goal_stops_after_one_turn() { // The pre-continuation shape, and the reason `goal` is what switches // the loop on rather than a separate flag: nothing to continue // toward is the same fact as nothing to continue. let state = mid_run("n", "unused", 5); state.set_goal("n", None, 5); assert!( matches!( state.plan_after_turn("n"), Continuation::Stop(StopReason::Done) ), "without a goal the first completed turn ends the run" ); assert_eq!( state.turns("n"), None, "and there is no turn budget to report being partway through" ); } #[test] fn a_goal_keeps_spending_turns_until_the_cap_and_then_stops() { let state = mid_run("n", "make the gate pass", 3); assert_eq!(state.turns("n"), Some((1, 3)), "the first turn is 1 of 3"); for expected in [2, 3] { let Continuation::Continue { prompt } = state.plan_after_turn("n") else { panic!("turn {expected} of 3 must still be spent"); }; assert!( prompt.contains("make the gate pass"), "the re-prompt continues toward the goal verbatim: {prompt}" ); assert!( prompt.contains("have not reported the goal reached"), "and says the one thing the subagent can't observe for itself: {prompt}" ); assert_eq!(state.turns("n"), Some((expected, 3))); } assert_eq!( state.plan_after_turn("n"), Continuation::Stop(StopReason::TurnCap { turns: 3 }), "the cap is the number of turns run, not one more" ); assert_eq!( state.turns("n"), Some((3, 3)), "a capped run must not advance past its own cap" ); } #[test] fn the_default_cap_is_five_turns() { // The number is the feature's own, not a value tuned here — pinned so // a later edit to `DEFAULT_MAX_TURNS` has to be deliberate. assert_eq!(DEFAULT_MAX_TURNS, 5); } #[test] fn both_signals_stop_the_continuation_before_the_cap_is_reached() { // The issue's own words: "both stop goal continues". Turn one of // five, so only the signal can be what stopped it. for stop in [ StopReason::GoalReached(Some("wrote the fix".to_owned())), StopReason::NeedHelp("no credential for the registry".to_owned()), ] { let state = mid_run("n", "a goal", 5); state.record_stop("n", stop.clone()); assert_eq!( state.plan_after_turn("n"), Continuation::Stop(stop.clone()), "{stop:?} must end the run with turns still on the clock" ); assert_eq!( state.turns("n"), Some((1, 5)), "and must not have spent one on the way out" ); } } #[test] fn a_signal_on_the_last_allowed_turn_outranks_the_cap() { // Both are true at once, and which one is reported is the difference // between "it says it finished" and "it ran out of road". let state = mid_run("n", "a goal", 1); state.record_stop("n", StopReason::GoalReached(None)); assert_eq!( state.plan_after_turn("n"), Continuation::Stop(StopReason::GoalReached(None)) ); } #[test] fn a_subagent_cannot_signal_a_different_session() { // The requirement itself, and the reason it's a test: two siblings // running concurrently, each holding exactly one signal URL. Half the // answer is in `mcp.rs` — neither tool has a `name` argument to put a // sibling's name in (`the_signal_tools_take_no_session_name` pins // that). The other half is here: whether the only identity a subagent // *does* hold, its token, can be made to resolve to anyone else. let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.reserve("alpha"); state.reserve("beta"); let alpha = token_of(&state.mint_signal_url("alpha")); let beta = token_of(&state.mint_signal_url("beta")); assert_ne!(alpha, beta, "two sessions must not share a token"); assert_eq!( state.session_for_signal_token(&alpha).as_deref(), Some("alpha") ); assert_eq!( state.session_for_signal_token(&beta).as_deref(), Some("beta") ); // `alpha` signals the only way it can: on its own endpoint, with the // session resolved from the token rather than supplied by the caller. let resolved = state .session_for_signal_token(&alpha) .expect("alpha's own route resolves"); need_help(&state, &resolved, "no credential".to_owned(), None); assert_eq!( state.stop_reason("alpha"), Some(StopReason::NeedHelp("no credential".to_owned())) ); assert_eq!( state.stop_reason("beta"), None, "a sibling's run must be untouched — there is no route `alpha` holds that reaches it" ); // A subagent does know its siblings' *names* (a brief can mention // them) — and a name is not a token, which is the whole point. assert_eq!(state.session_for_signal_token("beta"), None); assert_eq!(state.session_for_signal_token(&format!("{alpha}0")), None); } #[test] fn a_token_stops_resolving_once_its_run_is_over() { // The expiry half of "unknown or expired token ⇒ 404": a finished // run's config file still names its URL, and that URL must be dead. let state = State::new(PathBuf::from("/dev/null"), signal_url()); state.reserve("n"); let token = token_of(&state.mint_signal_url("n")); state.finish_turn("n", &TurnEnd::Complete); assert_eq!( state.session_for_signal_token(&token), None, "the run ended, so the route it was issued must resolve to nothing" ); // Same for a call that never reached a spawn at all. state.reserve("n"); let unspawned = token_of(&state.mint_signal_url("n")); state.release_reservation("n"); assert_eq!(state.session_for_signal_token(&unspawned), None); } #[test] fn a_re_minted_url_retires_the_previous_one() { // A `continue` mints the resumed run its own token; the run before it // is over, so the URL that run was handed must not still work. let state = State::new(PathBuf::from("/dev/null"), signal_url()); let first = token_of(&state.mint_signal_url("n")); let second = token_of(&state.mint_signal_url("n")); assert_eq!(state.session_for_signal_token(&first), None); assert_eq!( state.session_for_signal_token(&second).as_deref(), Some("n") ); } #[test] fn need_help_is_a_state_a_parent_can_see_without_reading_anything() { // Requirement in full: it stops the session *and* shows up in // `status` as its own state, distinct from idle and from killed. let blocked = describe_status(&StatusFacts { stop: Some(StopReason::NeedHelp( "the brief contradicts the code".to_owned(), )), turns: Some((2, 5)), ..StatusFacts::new("n") }) .expect("a blocked session is a state, not an error"); assert!( blocked.contains("BLOCKED") && blocked.contains("the brief contradicts the code"), "the block and its reason must both be in the answer: {blocked}" ); assert!( blocked.contains("Turn 2 of 5"), "with the progress that says how far it got: {blocked}" ); assert!( !blocked.contains("idle"), "a blocked subagent must not also read as idle: {blocked}" ); } #[test] fn a_reported_goal_never_reads_as_a_verified_one() { // The failure this is built against: `goal_reached` is self-reported // by a subagent that has just been told it hasn't reached the goal. // Every surface that renders it has to say so. let status = describe_status(&StatusFacts { stop: Some(StopReason::GoalReached(Some( "refactored the parser".to_owned(), ))), ..StatusFacts::new("n") }) .expect("status"); assert!( status.contains("refactored the parser") && status.contains("claim"), "status must carry both the report and the fact it's only a report: {status}" ); let todo = turn_end_summary( "n", &TurnEnd::Complete, Some(&StopReason::GoalReached(Some( "refactored the parser".to_owned(), ))), ); assert!( todo.contains("self-reported, not verified"), "and so must the todo, which is the half a parent reads unprompted: {todo}" ); } #[test] fn a_stop_reason_extends_the_done_message_rather_than_replacing_it() { // "Extend", not "replace": the observed end of the turn and the // reason the run stopped are different facts, and dropping the first // would let a self-reported claim stand where an observation was. let plain = turn_end_summary("n", &TurnEnd::Complete, None); for stop in [ StopReason::GoalReached(None), StopReason::NeedHelp("blocked".to_owned()), StopReason::TurnCap { turns: 5 }, ] { let extended = turn_end_summary("n", &TurnEnd::Complete, Some(&stop)); assert!( extended.starts_with(&plain), "{stop:?} must extend the done message, not rewrite it: {extended}" ); assert!( extended.len() > plain.len(), "{stop:?} must actually add something: {extended}" ); } assert_eq!( turn_end_summary("n", &TurnEnd::Complete, Some(&StopReason::Done)), plain, "`Done` has nothing to add — the unextended message already says exactly that" ); } #[test] fn the_turn_cap_todo_says_the_harness_limit_was_what_stopped_it() { // Not a silent stop: the one notification a parent gets unprompted // has to distinguish "it finished" from "we stopped asking". let todo = turn_end_summary( "n", &TurnEnd::Complete, Some(&StopReason::TurnCap { turns: 5 }), ); assert!( todo.contains("harness turn limit was reached (5 turns)"), "the todo must name the limit as the cause: {todo}" ); assert!( todo.contains("without the goal ever being reported reached"), "and say what that means for the work: {todo}" ); } #[test] fn status_reports_the_turn_counter_in_every_state_a_goal_run_reaches() { for occupancy in [Some(true), Some(false), None] { let answer = describe_status(&StatusFacts { occupancy, session_exists: true, turns: Some((3, 5)), ..StatusFacts::new("n") }) .expect("status"); assert!( answer.contains("Turn 3 of 5"), "progress is the point of the counter — it can't be absent from {occupancy:?}: \ {answer}" ); } let goalless = describe_status(&StatusFacts { occupancy: Some(true), ..StatusFacts::new("n") }) .expect("status"); assert!( !goalless.contains("Turn "), "a session with no goal has no budget to be partway through: {goalless}" ); } #[test] fn a_running_turn_that_has_already_signalled_says_so() { // The window a parent would otherwise misread: the subagent has said // it is blocked, its turn hasn't ended yet, and plain "running" would // tell the parent to leave it alone. let answer = describe_status(&StatusFacts { occupancy: Some(true), stop: Some(StopReason::NeedHelp("no credential".to_owned())), ..StatusFacts::new("n") }) .expect("status"); assert!( answer.contains("is running") && answer.contains("no credential"), "both facts are true at once and both have to be in the answer: {answer}" ); let quiet = describe_status(&StatusFacts { occupancy: Some(true), stop: Some(StopReason::Done), ..StatusFacts::new("n") }) .expect("status"); assert!( !quiet.contains("already signalled"), "`Done` is not a signal the subagent raised — nothing to announce: {quiet}" ); } #[test] fn a_kill_outranks_a_stop_reason_the_subagent_had_already_raised() { // Both records can be set at once — a subagent calls `goal_reached` // and something SIGKILLs it before the turn ends. The observed kill // is the one that can't be a claim. let answer = describe_status(&StatusFacts { killed: Some(libc::SIGKILL), stop: Some(StopReason::GoalReached(None)), session_exists: true, ..StatusFacts::new("n") }) .expect("status"); assert!( answer.contains("was killed"), "the kill is what happened: {answer}" ); assert!( !answer.contains("goal reached"), "a claim made just before being killed must not be the headline: {answer}" ); } #[test] fn a_continue_gives_a_capped_session_its_allowance_back() { // `TurnCap` stops the loop, it doesn't retire the session — and the // parent spending a turn on purpose is not what the cap bounds. let state = mid_run("n", "a goal", 2); state.record_stop("n", StopReason::TurnCap { turns: 2 }); state.clear_stop("n"); state.restart_turns("n"); assert_eq!(state.turns("n"), Some((1, 2))); assert_eq!(state.stop_reason("n"), None); assert!( matches!(state.plan_after_turn("n"), Continuation::Continue { .. }), "with the allowance back, the loop has a turn to spend again" ); } #[test] fn a_fresh_start_inherits_nothing_from_the_run_before_it() { // `start` archives the prior session precisely so this is a fresh // start; a leftover stop reason or report path would make the new run // report the old one's ending, into the old one's file. let state = Arc::new(State::new(PathBuf::from("/dev/null"), signal_url())); state.record_stop("n", StopReason::NeedHelp("old block".to_owned())); state.set_report_file("n", Some("/tmp/old-report.md")); state.reserve("n"); let refused = start(&state, start_request("n")); assert!(refused.is_err(), "the reserved name is refused as before"); // The same clearing `start` does once it owns the name. state.clear_stop("n"); state.clear_report_file("n"); assert_eq!(state.stop_reason("n"), None); assert_eq!(state.report_file("n"), None); } #[test] fn the_report_path_is_taken_from_the_session_and_never_guessed() { // Both halves of where it can come from: the brief `start` carried, // and the subagent saying where it actually wrote. A signal that // names no path leaves the remembered one alone rather than erasing // it. let state = State::new(PathBuf::from("/dev/null"), signal_url()); assert_eq!( state.report_file("n"), None, "a session nobody told about a report file has none — nothing is inferred" ); state.set_report_file("n", Some("/tmp/brief-said.md")); goal_reached(&state, "n", None, None); assert_eq!( state.report_file("n"), Some(PathBuf::from("/tmp/brief-said.md")), "a signal with no path must not erase what the brief named" ); goal_reached(&state, "n", None, Some("/tmp/actually-wrote.md")); assert_eq!( state.report_file("n"), Some(PathBuf::from("/tmp/actually-wrote.md")), "and the subagent saying where it wrote is what wins" ); } #[tokio::test] async fn the_stop_reason_is_appended_to_the_report_file_the_session_named() { let dir = std::env::temp_dir().join(format!("hive-subagent-report-{}", std::process::id())); std::fs::create_dir_all(&dir).expect("scratch dir"); let path = dir.join("report.md"); std::fs::write(&path, "# what the subagent wrote\n").expect("seed the report"); write_stop_to_report(Some(path.clone()), "n", &StopReason::TurnCap { turns: 5 }).await; let body = std::fs::read_to_string(&path).expect("report still readable"); assert!( body.starts_with("# what the subagent wrote\n"), "appended, never rewritten — the subagent wrote that file: {body}" ); assert!( body.contains("harness turn limit was reached (5 turns)"), "and the stop reason lands in the artifact a parent already reads: {body}" ); std::fs::remove_dir_all(&dir).ok(); } #[tokio::test] async fn an_unwritable_report_path_costs_the_line_and_nothing_else() { // Best-effort by design: the same sentence is in the todo, so a path // this daemon can't write is a missing convenience, not a lost fact. write_stop_to_report( Some(PathBuf::from("/proc/definitely/not/writable/report.md")), "n", &StopReason::Done, ) .await; write_stop_to_report(None, "n", &StopReason::Done).await; } #[test] fn a_goal_briefing_tells_the_subagent_what_it_cannot_otherwise_know() { let briefing = goal_briefing("batch-1", "get the gate to pass", 5); assert!( briefing.contains("get the gate to pass") && briefing.contains("up to 5 turns"), "the goal and the budget both have to reach the subagent: {briefing}" ); assert!( briefing.contains("goal_reached") && briefing.contains("need_help"), "as do the two ways it has to stop the re-prompting: {briefing}" ); assert!( briefing.contains("batch-1"), "and its own name, which is what those tools are called with: {briefing}" ); } #[test] fn a_signal_url_reaches_the_subagent_and_a_status_check_renders_no_config() { // `turn` actually writes an --mcp-config file (crate::mcp_config::build), // which resolves `hive_agent_sock::paths::harness_dir` — normally the // container's injected `HYPERHIVE_HARNESS_DIR`, unset in a plain `cargo // test` sandbox. Point it at a scratch dir for just this test rather // than depending on ambient environment; only this test's call path // reads the var, so there's no cross-test race to sequence around // (contrast `otel_attrs_append_ambient_value_or_stand_alone` above). let dir = std::env::temp_dir().join(format!( "hive-subagent-mcp-test-harness-dir-{}", std::process::id() )); // SAFETY: test-only env mutation; no other test reads this var. unsafe { std::env::set_var("HYPERHIVE_HARNESS_DIR", &dir); } // The signal surface is how `goal_reached` is callable at all, so a // spawned turn's config has to carry it; `status` builds a config // purely to resolve the store and has no subagent to hand it to. let turn = build_config("n", None, None, None, None, Some(&signal_url())); let checked = build_config("n", None, None, None, None, None); // SAFETY: see above. unsafe { std::env::remove_var("HYPERHIVE_HARNESS_DIR"); } let _ = std::fs::remove_dir_all(&dir); assert!( turn.mcp_config.is_some(), "a turn's config must carry an --mcp-config with the signal surface in it" ); assert!( turn.strict_mcp_config && checked.strict_mcp_config, "the safety property is unchanged: no ambient MCP discovery either way" ); } }