mara on #778: 'remove the manager special case argus nitted about'.
`plugins::install_configured` no longer takes a `notify_recipient`
hardcoding "manager". Now returns a Vec<String> of failure messages;
serve_main<S> iterates them and routes each through S::send_to_parent
— the same <parent> sentinel failure-notify uses everywhere else
(#703). Manager plugin failures now reach operator via root → operator
fallback (improvement on the pre-PR silent-drop).
Also rename FORGE_MENTIONS_ONLY → FORGE_IS_MANAGER to fix the misnomer:
the boolean picks which wire enum (AgentRequest::Wake vs
ManagerRequest::Wake) the forge_notify poller uses, not anything about
mentions-only filtering (that's a separate nix-side option). Real fix
is to lift Surface into the lib crate and make forge_notify::run
generic; deferred to its own issue.
Net: -20 LOC.
mara on #778: "I would have expected the manager and agents to share
the exact same turn function, making one obsolete. I don't see that
in the code, why not?" — fair. went further.
introduces a Surface trait + AgentSurface / ManagerSurface zero-sized
impls wrapping the disjoint Request/Response enums + boot-time
constants (FLAVOR / DEFAULT_LABEL / PLUGINS_PARENT / FORGE_MENTIONS_ONLY).
the turn loop itself collapses to one generic implementation:
- serve_main<S> replaces agent_serve_main + manager_serve_main
- serve_loop<S> replaces agent_serve_loop + manager_serve_loop
- handle_turn<S> replaces handle_agent_turn + handle_manager_turn
- wake<S> replaces agent_wake + manager_wake
RecvOutcome enum decouples the per-role Response shape from the loop's
match arms so serve_loop never sees either enum.
main's dispatch picks the type parameter from HIVE_ROLE; everything
downstream is identical by construction.
net: -62 LOC vs main even with the new manager notify-on-failure +
continue-sentinel features kept.
three shared helpers replace the duplicated pre-#598 patterns:
- `log_system_event` lifts the HelperEvent parse + bus emit out of
handle_manager_turn so agents log QuestionAnswered/ContainerCrash/
reparent notifications the same way (#692 part 1).
- `format_turn_failure` produces the failure-notification body using
identity::qualified_label() instead of a label param threaded through
three layers. drops `label` from handle_agent_turn, agent_serve_loop,
agent_check_and_inject_continue.
- `consume_continue_sentinel` lifts the file-probe so both surfaces
reuse it (#692 part 2 — sentinel now works for manager too).
agent_notify_manager_of_failure → agent_notify_parent_of_failure: routes
via the <parent> sentinel landed in #703 instead of the literal string
'manager'. mirrored on manager side; root-manager failures resolve to
operator via topology::resolve_recipient.
handle_*_turn signatures now identical modulo the wire-type prefix
(part 3 acceptance from the issue).
The mara: 'many agents stuck at needs login? seem to do turns fine, no
needs login on agent page. dash shows needs login tho'.
Diagnosed: hive-ag3nt::events::Bus::emit_status writes
{state_dir}/hyperhive-needs-login when status flips to needs_login_idle
and only removes it when status flips to online. But the boot flow only
emits 'online' if the harness was previously parked in wait_for_login —
the LoginState::Online branch went straight into serve() without
touching emit_status. So a sentinel written on a prior boot (e.g. a
401-triggered park) survived a healthy re-spawn, and the dashboard's
auth_failed_sentinel(name) read it as still-needing-login forever after.
Fix: add bus.emit_status('online') at the top of the LoginState::Online
boot branch in both hive-ag3nt.rs (sub-agent) and hive-m1nd.rs (manager).
Idempotent — emit_status is a one-line write/remove on a tiny empty
file; calling it for an already-clean state is a no-op.
This addresses one half of #682. The other half (claude_has_session()
may EACCES on the agent's 0700-perm ~/.claude dir post-#658 root drop)
is c0re-side and tracked separately in the issue thread for damocles.
refs #682
mara on #563: 'we fixed the agent to not start turns in that
state (it fell back to online before), but this does not show on
dashboard properly'.
Root cause: the per-agent harness flips LoginState::NeedsLogin in
memory on three entry paths (cold-boot without a session, 401
mid-turn, /api/logout) and parks in wait_for_login. But
wait_for_login itself never called bus.emit_status('needs_login_idle')
at entry — only the /api/logout handler does that today. So:
- Cold boot: agent has no session, harness shows 'needs login'
on its own web UI (via LoginState mutex), but the dashboard's
needs_login field stays false because the
{state_dir}/hyperhive-needs-login sentinel was never written.
- 401 mid-turn: same — the 'after a turn failed' path in
hive-ag3nt.rs / hive-m1nd.rs flips LoginState directly without
emitting status, then calls wait_for_login, which now waits
silently with no sentinel write.
Fix: hoist the emit_status('needs_login_idle') call into
wait_for_login itself. All three entry paths get the sentinel
write for free; the /api/logout handler's explicit call (web_ui.rs
line 966) becomes redundant but idempotent — no behaviour change
there. The 'online' clear at session refresh stays exactly where
it was at the loop's exit.
Both hive-ag3nt and hive-m1nd binaries share wait_for_login, so the
manager harness benefits without a separate change.
Cargo's default test runner parallelises tests within a binary,
so the original 'tests run serially' comment was wrong — two
`with_env` calls running concurrently would race the
process-wide HIVE_LABEL / HYPERHIVE_HIVE_DOMAIN state.
Added a module-scope `static ENV_LOCK: Mutex<()>` and acquire it
at the top of `with_env` so each set / run / restore window is
exclusive. Poison recovery via `unwrap_or_else(into_inner)` so a
single test panic doesn't cascade through the rest of the module.
Lighter than pulling in serial_test for one module. No new deps.
First chunk of #589 v0 phase A: plumbing the hive-qualified
'name@hive' form through the per-agent surfaces that the harness
itself owns. Broker from/to + dashboard rendering + container_view
follow in subsequent PRs once damocles ships the HYPERHIVE_HIVE_DOMAIN
env var in harness-base.nix.
- new hive_ag3nt::identity module: label() / hive_domain() /
qualified_label() / qualify(label). Reads HYPERHIVE_HIVE_DOMAIN
(set by hive-c0re.nix module from hyperhive.domain) — when unset
or empty, qualified_label degrades to just the short label so
existing single-hive deployments are unchanged. Six unit tests
cover the set / unset / empty / arbitrary-label paths.
- prompt::render gains {qualified_label} substitution alongside
the existing {label}. system.md template uses both: the agent
intro now reads 'You are hyperhive agent iris (qualified:
iris@darkest.space) in a multi-agent system. ... When you're
talking to or about a peer on a different hive, use the
qualified form (name@hive) so the operator + the manager can
disambiguate'. Manager flavor gets the same treatment.
- /api/state gains qualified_label: String. Always present, equals
label when no domain is configured.
- frontend setHeader takes the qualified_label, drives the browser
tab title (so two tabs from different hives are
distinguishable in the tab bar) while the glyphic #title stays
short for the cinematic header.
Gated on env var presence — no behaviour change for single-hive
deployments. Pairs with damocles's upcoming harness-base.nix
HYPERHIVE_HIVE_DOMAIN ship; safe to land in either order.
Per mara's review on PR #561: the previous commit kept
`./hive-ag3nt/prompts` in `cleanSrc` because
`hive-ag3nt::prompt::tests` had a compile-time
`include_str!("../prompts/system.md")`. That meant a prompt edit
still busted the cargo cache.
This change:
- Replaces the test-side `include_str!` with a runtime read from
`$HIVE_ASSETS_DIR/prompts/system.md` (with a CARGO_MANIFEST_DIR
fallback for plain `cargo test` from a checked-out repo).
- Drops `./hive-ag3nt/prompts` from `cleanSrc` — it's now
`craneLib.cleanCargoSource ./.` (Cargo.* + *.rs only).
- Sets `doCheck = false` on `packages.default` and lifts
`cargo test` into a separate `checks.cargo-test` derivation
that carries the `hyperhive-assets` build input. That scopes the
asset rebuild blast radius to the test check — `nix flake check`
still exercises the suite, but the binary derivation no longer
carries the assets dep.
Verified cache-invariance matrix (via `echo '' >> <f>; nix eval
.#default.outPath`):
| edit | default | cargo-test | clippy |
|-------------------------|---------|------------|--------|
| README.md | stable | stable | stable |
| branding/hyperhive.svg | stable | CHANGED | stable |
| nix/modules/* | stable | stable | stable |
| prompts/system.md | stable | CHANGED | stable |
| hive-c0re/src/main.rs | CHANGED | CHANGED | CHANGED |
(`cargo-test` CHANGED on prompts/branding is correct — tests
read the production template + need the assets output.)
Cuts every `include_bytes!`/`include_str!` of a non-rust path in
the workspace over to runtime file loads from `$HIVE_ASSETS_DIR`
(the `hyperhive-assets` derivation introduced in the previous
commit). After this commit the rust derivation has no compile-time
dependency on `branding/*` or `hive-ag3nt/prompts/*` anymore.
Call-site flips:
- `hive-c0re/src/forge.rs::CORE_AVATAR_PNG` /
`CONFIG_ORG_AVATAR_PNG`: were `include_bytes!` of
`branding/hyperhive.png` and `$OUT_DIR/agent-configs.png`. Now
`ensure_core_avatar` / `ensure_config_org_avatar` `tokio::fs::read`
via `hive_sh4re::assets::{core_avatar_png, config_org_avatar_png}`
at startup. The `agent-configs.png` is now rendered by the
`hyperhive-assets` derivation's rsvg-convert step (was
`hive-c0re/build.rs` + librsvg on the rust derivation's
nativeBuildInputs — both gone in the next commit).
- `hive-ag3nt/src/prompt.rs::TEMPLATE`: `render` now takes the
template as an argument; `write_system_prompt` reads it once from
`$HIVE_ASSETS_DIR/prompts/system.md` before calling render. The
test module still `include_str!`s the production template so
`cargo test --workspace` doesn't need `HIVE_ASSETS_DIR` set —
this is the only remaining compile-time reference to the file
from the rust workspace, gated to `#[cfg(test)]`.
- `hive-ag3nt/src/turn.rs::CLAUDE_SETTINGS`: was `include_str!`'d
and written via `tokio::fs::write`; now `tokio::fs::copy` from
`$HIVE_ASSETS_DIR/prompts/claude-settings.json` into the
per-agent socket dir.
- `hive-ag3nt/src/web_ui.rs::DEFAULT_ICON`: was `include_str!`'d;
now read on-demand from `$HIVE_ASSETS_DIR/branding/hyperhive.svg`
inside `serve_icon`. Falls back to an empty body if missing so
the endpoint never panics on a misconfigured container (matches
the existing "per-agent icon.svg override" fallthrough).
`HIVE_ASSETS_DIR` wiring:
- Inside containers: `nix/templates/harness-base.nix`
`environment.variables` sets it to
`${pkgs.hyperhive-assets}/share/hyperhive` (resolved through
the default overlay applied in `mkContainer`). Verified by
building `agent-base-toplevel` and grepping the resulting
`/etc/set-environment`.
- Host-side: `nix/modules/hive-c0re.nix` adds an `assets` option
defaulting to `hyperhive.packages.${system}.assets`, threaded
in from the flake's nixosModules wiring, and sets the same env
var on the `hive-c0re` systemd unit so the daemon's
`forge::ensure_*_avatar` startup hooks find the PNGs.
`hive-c0re/build.rs` deleted entirely; `[package].build` removed
from `hive-c0re/Cargo.toml`; rsvg-convert dependency lives in the
assets derivation only.
Validated: `nix build .#default .#checks.x86_64-linux.clippy
.#agent-base-toplevel .#manager-toplevel --fallback` all succeed.
`/etc/set-environment` in the toplevel shows
`HIVE_ASSETS_DIR="/nix/store/.../hyperhive-assets-0.1.0/share/hyperhive"`.
The naersk → crane swap in the parent commit flips clippy from
silently passing to actually failing on `-D warnings` (naersk's
`mode = "clippy"` mangled the `--` separator so the deny never took
effect). This commit clears the surfaced lints so the workspace
builds clean under the new enforcement — every fix is mechanical and
preserves behaviour. Tests still pass (160 across the workspace).
Auto-fixes via `cargo clippy --fix`:
- `doc_markdown` (19 sites): bare identifiers in doc comments
wrapped in backticks
- `format_in_format_args`, `explicit_into_iter_loop`,
`redundant_closure_for_method_calls`, `useless_conversion`, and
a few more — mechanical rewrites of the kind cargo can apply
safely.
Hand-fixed:
- `match_same_arms` (forge_notify::is_atx_heading): two arms returning
`true` collapsed into a single `matches!` pattern.
- `cast_sign_loss` + `format_push_string` (mcp.rs status formatter):
guarded `i64 → u64` through `u64::try_from(…).unwrap_or(0)` (status
timestamps are always positive in practice; clamp the skew edge to
0) and swapped `out.push_str(&format!(…))` for `write!` into the
buffer with an infallible-writer `let _ =`.
- `doc_lazy_continuation` in turn.rs + manager_server.rs + sh4re/lib.rs:
doc paragraphs that the markdown parser was treating as list-item
continuations got either a separating blank line or a `/`-for-`+`
word swap so the parser stops seeing a list.
- `unused_async` (manager_server::handle_request_schedule_prompt):
function has no `.await`; dropped the `async` and its `.await` call
site.
- `needless_pass_by_value` (scheduled_prompts::submit): take
`&NewSchedule` instead of moving the struct in; updated two prod
callers and eight test sites to pass references.
- `type_complexity` (approvals::mark_cancelled): hoisted the
7-tuple SELECT row shape into a `type CancelLookupRow = (…);` alias.
Allow-with-reason for intentional patterns:
- `option_option` (6 sites across dashboard / scheduled_prompts /
manager_server): `Option<Option<T>>` carries three-state PATCH
semantics (missing key = leave alone, `Some(None)` = clear,
`Some(Some(v))` = set). Collapsing to `Option<T>` loses the
"clear" state.
- `dead_code` (rebuild_queue::QueueKind::Destroy /
QueueSource::CrashRecover; topology::parent_of / default_seed):
wire-shape variants + API surfaces kept for the upcoming features
(#361 follow-ups, future `Destroy` queue routing, crash-recovery
path). Allowed at the variant / function level with the rationale
in `reason = "…"`.
- `too_many_lines` on three specific call-sites: a 117-line
exhaustive-variant test (dashboard_events::kind_tag_matches_…),
the meta-flake string template renderer
(meta::render_flake_with_lookup), and the notification poll loop
(forge_notify::poll_once) — splitting any of them would just hide
the contiguous shape they exist to keep visible.
`nix flake check` formatting target is still broken on main itself
(pre-existing nixfmt drift across ~28 files unrelated to this PR);
left alone here so the scope stays "crane port + lints the port
exposed" and the operator's review doesn't have to triage drive-by
nixfmt churn.
per mara's review on #433, move the gating from the dashboard into
the host so a stopped container's stale on-disk state (rate_limited
sentinel, hyperhive-needs-login, last-turn-stats row, status blob)
never reaches the wire in the first place. when build_all sees
is_running == false:
- needs_login → false
- ctx_tokens / context_window_tokens → None
- rate_limited → false
- status_text / status_set_at → None
static / declared fields (extra_links, deployed_sha,
pending_reminders, needs_update, parent) stay populated regardless
of run state.
extend AgentMeta (both AgentResponse + ManagerResponse) with a
`running: bool` field so get_agent_meta callers can tell whether
the target is up — answers the second half of #432 ("agent meta
should probably show the info that it is not running as well").
read_agent_status_live wraps the existing read_agent_status with
the same is_running gate so the manager/agent socket handlers don't
have to know about sentinel semantics.
format_agent_meta now prints a `running: yes|no` line so claude
sees the run state in plain text alongside hyperhive_rev.
frontend follow-up in the same commit: drop the redundant
`c.running &&` guards on ctx_tokens / status_text in
renderContainers — the backend now guarantees those fields are
absent when the container is stopped, so the existing
truthy-check is sufficient. the `■ not running` badge + icon /
links fetch short-circuits stay (those are pure presentation /
network-noise wins the backend can't address).
Phase 4 of #273 — the actual switch. Both axum routers now serve
their static surface via `tower_http::services::ServeDir` mounted
as a fallback service, reading the dist path from `HIVE_STATIC_DIR`
(set by Phase 3's NixOS module wiring).
Deletes:
- `hive-c0re/assets/{index.html, app.js, dashboard.css}`
- `hive-ag3nt/assets/{index.html, app.js, agent.css, stats.html,
stats.js, screen.html}`
- The whole `hive-fr0nt/` crate (workspace member dropped, both
hive-c0re and hive-ag3nt drop their `hive-fr0nt.workspace = true`
dep). Its contents now live as `@hive/shared` under
`frontend/packages/shared/`.
Rust changes:
- `hive-c0re/src/dashboard.rs`: remove `serve_index`, `serve_css`,
`serve_app_js`, `serve_shared_js`, `serve_marked_js`,
`serve_favicon` (all six `include_str!` handlers); replace their
routes with a single `.fallback_service(ServeDir::new(static_dir))`
on the router. Fail closed (anyhow::bail) if `HIVE_STATIC_DIR` is
unset or not a directory at startup.
- `hive-ag3nt/src/web_ui.rs`: remove `serve_index`, `serve_css`,
`serve_app_js`, `serve_shared_js`, `serve_marked_js`,
`serve_stats`, `serve_stats_js`, `serve_screen`; same
`fallback_service` pattern. `serve_icon` stays (consumes
`/etc/hyperhive/icon.svg` + `branding/hyperhive.svg` fallback,
neither of which lives under the frontend dist).
- `AgentLink` URLs for stats/screen switched from `/stats` / `/screen`
to `/stats.html` / `/screen.html` since ServeDir doesn't auto-
append the extension and the on-disk filename is the natural URL
post-cutover.
- `Cargo.toml` (workspace): drop `hive-fr0nt` member + workspace
dep, add `tower-http = { version = "0.6", features = ["fs"] }`.
- `hive-c0re/Cargo.toml` + `hive-ag3nt/Cargo.toml`: drop the
`hive-fr0nt.workspace = true` dep, add `tower-http.workspace =
true`.
Docs updated:
- `CLAUDE.md`: file map reflects `frontend/` (was `hive-fr0nt/` +
`assets/`) and the ServeDir/HIVE_STATIC_DIR shape.
- `docs/web-ui.md` 'Shape (shared by both)' section: describes the
ServeDir fallback + bundled-by-esbuild surface, no more
`include_str!` references.
- `docs/terminal-rendering.md`: src paths point at
`frontend/packages/{agent,shared}/src/`; marked is the npm dep,
not vendored UMD.
Validation:
- `cargo check --workspace` — clean (5 warnings, all pre-existing
in `rebuild_queue.rs`, none on changed files).
- `cargo clippy --workspace --all-targets` — clean (11 warnings,
same pre-existing source).
- `cd frontend && npm run build` from the prior commit's lockfile
produces the dist directories the new routers consume:
dashboard: `dist/{index.html, static/{app.js, dashboard.css}}`
agent: `dist/{index.html, stats.html, screen.html,
static/{app.js, stats.js, agent.css}}`
(favicon.svg lands in dashboard/ during the nix build —
`nix/frontend.nix` install phase copies `branding/hyperhive.svg`
there, since it's outside the npm tree.)
Refs #273.