refactor(hive-agent): split the forge notification poller into its own crate

The poller was a `tokio::spawn` inside the `hive-agent` serve loop. It
never needed anything from that loop except a socket path, so being
in-process bought nothing and cost two things: a harness restart took
forge notifications down with it, and the whole forge/HTTP dependency
tree was linked into the serve-loop binary.

It is now `hive-forge-notify`, a per-agent daemon with its own systemd
unit, a sibling of `hive-bash-daemon` and `hive-matrix-daemon`. Same
contract as those two: it reaches the harness only by upserting todos on
the in-agent socket, and nowhere else.

The module moves verbatim (`notify.rs`) — the formatters, the activation
gates, the dedupe map and all 33 tests are unchanged. Only the socket
call sites are rewritten, onto a small local `todo_client` rather than
the harness's. That mirrors what both sibling daemons already do, and
the etiquette differs on purpose: the harness's client carries a 60s
backoff schedule sized to ride out a hive-c0re restart, which its
callers need because they have no retry of their own. This poller's two
call sites both sit inside the 30s poll loop and both treat a failure as
"leave the thread unread, try next tick", so the poll interval already
is the retry; a second backoff would only stack sleeps and delay the
rest of the batch.

The unit is `Restart=on-failure`, not `always`. An agent with no forge
account is a supported configuration and the poller reports it by
logging why and exiting 0 — under `always` that clean exit would be a
restart loop on every forge-less agent.

`forgejo-api`, `url` and `time` drop out of `hive-agent`'s dependencies
with the module.

Also corrects docs that outlived the code they described: the persisted
`forge_cursor` field is long gone (forge's own read-state is the durable
record of what has been delivered), but `docs/persistence.md` and the
`harness_state` module docs still documented it as live.
This commit is contained in:
atlas 2026-07-26 20:44:54 +02:00 committed by mara
commit 246c9471b1
18 changed files with 301 additions and 48 deletions

File diff suppressed because it is too large Load diff

View file

@ -1,9 +1,8 @@
//! File-backed harness state, split out of [`crate::events`] (which owns
//! the live event bus + sqlite store). This module holds the runtime
//! model/effort selection, the consolidated `hyperhive-harness.json`
//! state file, and the `forge_notify` delivery-dedupe cursor — the
//! persisted knobs the harness reads/writes across turns, none of which
//! are about the live event stream.
//! model/effort selection and the consolidated `hyperhive-harness.json`
//! state file — the persisted knobs the harness reads/writes across
//! turns, none of which are about the live event stream.
use std::path::PathBuf;
@ -84,12 +83,14 @@ fn harness_json_path() -> PathBuf {
crate::paths::state_dir().join(HARNESS_JSON)
}
// Serialises the read-modify-write of `hyperhive-harness.json`. Two
// harness tasks touch it in the same process — the turn loop (rate-limit
// / needs-login / active-model) and the forge_notify poller (the
// delivery-dedupe cursor) — writing disjoint fields, so each writer must
// preserve the other's. The lock closes the lost-update window between a
// writer's read and its rename.
// Serialises the read-modify-write of `hyperhive-harness.json`. Every
// writer merges into the existing object rather than reconstructing it,
// so a future second writer preserves fields it doesn't own; the lock
// closes the lost-update window between a writer's read and its rename.
// (The forge poller used to be that second writer, for a delivery-dedupe
// cursor. It no longer persists one — forge's own read-state is the
// durable record — and it is a separate process now, which an
// in-process mutex could not have serialised anyway.)
static HARNESS_JSON_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(());
/// Read the consolidated state file as a JSON object, or an empty object

View file

@ -13,7 +13,6 @@ mod client;
mod db_migrate;
mod disk_watch;
mod events;
mod forge_notify;
mod harness_state;
mod identity;
mod login;
@ -550,15 +549,12 @@ async fn serve_main<S: Surface>(socket: &Path, poll_ms: u64) -> Result<()> {
for failure in plugins::install_configured().await {
S::send_to_parent(socket, failure).await;
}
// forge_notify pushes forge notifications as todos on the in-agent
// socket (loose-ends v2), not direct wakes — so it dials
// `HIVE_AGENT_SOCKET`, not the host mcp.sock.
tokio::spawn(crate::forge_notify::run(
std::env::var_os("HIVE_AGENT_SOCKET").map_or_else(
|| std::path::PathBuf::from(hive_agent_sock::DEFAULT_AGENT_SOCKET),
std::path::PathBuf::from,
),
));
// The forge notification poller used to be spawned here. It is its own
// process now (`hive-forge-notify`, its own systemd unit) so a harness
// restart doesn't take forge notifications down with it; it reaches the
// harness the same way the bash and matrix daemons do, by upserting
// todos on the in-agent socket.
//
// Agent-side cleanup of this agent's own harness artifacts (completed
// bash-task files + verbose event rows). Runs here, not host-side in
// hive-c0re, because the files are agent-owned — see `vacuum` module docs.