refactor(hive-agent): split the forge notification poller into its own crate
The poller was a `tokio::spawn` inside the `hive-agent` serve loop. It never needed anything from that loop except a socket path, so being in-process bought nothing and cost two things: a harness restart took forge notifications down with it, and the whole forge/HTTP dependency tree was linked into the serve-loop binary. It is now `hive-forge-notify`, a per-agent daemon with its own systemd unit, a sibling of `hive-bash-daemon` and `hive-matrix-daemon`. Same contract as those two: it reaches the harness only by upserting todos on the in-agent socket, and nowhere else. The module moves verbatim (`notify.rs`) — the formatters, the activation gates, the dedupe map and all 33 tests are unchanged. Only the socket call sites are rewritten, onto a small local `todo_client` rather than the harness's. That mirrors what both sibling daemons already do, and the etiquette differs on purpose: the harness's client carries a 60s backoff schedule sized to ride out a hive-c0re restart, which its callers need because they have no retry of their own. This poller's two call sites both sit inside the 30s poll loop and both treat a failure as "leave the thread unread, try next tick", so the poll interval already is the retry; a second backoff would only stack sleeps and delay the rest of the batch. The unit is `Restart=on-failure`, not `always`. An agent with no forge account is a supported configuration and the poller reports it by logging why and exiting 0 — under `always` that clean exit would be a restart loop on every forge-less agent. `forgejo-api`, `url` and `time` drop out of `hive-agent`'s dependencies with the module. Also corrects docs that outlived the code they described: the persisted `forge_cursor` field is long gone (forge's own read-state is the durable record of what has been delivered), but `docs/persistence.md` and the `harness_state` module docs still documented it as live.
This commit is contained in:
parent
d153d1d35f
commit
246c9471b1
18 changed files with 301 additions and 48 deletions
File diff suppressed because it is too large
Load diff
|
|
@ -1,9 +1,8 @@
|
|||
//! File-backed harness state, split out of [`crate::events`] (which owns
|
||||
//! the live event bus + sqlite store). This module holds the runtime
|
||||
//! model/effort selection, the consolidated `hyperhive-harness.json`
|
||||
//! state file, and the `forge_notify` delivery-dedupe cursor — the
|
||||
//! persisted knobs the harness reads/writes across turns, none of which
|
||||
//! are about the live event stream.
|
||||
//! model/effort selection and the consolidated `hyperhive-harness.json`
|
||||
//! state file — the persisted knobs the harness reads/writes across
|
||||
//! turns, none of which are about the live event stream.
|
||||
|
||||
use std::path::PathBuf;
|
||||
|
||||
|
|
@ -84,12 +83,14 @@ fn harness_json_path() -> PathBuf {
|
|||
crate::paths::state_dir().join(HARNESS_JSON)
|
||||
}
|
||||
|
||||
// Serialises the read-modify-write of `hyperhive-harness.json`. Two
|
||||
// harness tasks touch it in the same process — the turn loop (rate-limit
|
||||
// / needs-login / active-model) and the forge_notify poller (the
|
||||
// delivery-dedupe cursor) — writing disjoint fields, so each writer must
|
||||
// preserve the other's. The lock closes the lost-update window between a
|
||||
// writer's read and its rename.
|
||||
// Serialises the read-modify-write of `hyperhive-harness.json`. Every
|
||||
// writer merges into the existing object rather than reconstructing it,
|
||||
// so a future second writer preserves fields it doesn't own; the lock
|
||||
// closes the lost-update window between a writer's read and its rename.
|
||||
// (The forge poller used to be that second writer, for a delivery-dedupe
|
||||
// cursor. It no longer persists one — forge's own read-state is the
|
||||
// durable record — and it is a separate process now, which an
|
||||
// in-process mutex could not have serialised anyway.)
|
||||
static HARNESS_JSON_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(());
|
||||
|
||||
/// Read the consolidated state file as a JSON object, or an empty object
|
||||
|
|
|
|||
|
|
@ -13,7 +13,6 @@ mod client;
|
|||
mod db_migrate;
|
||||
mod disk_watch;
|
||||
mod events;
|
||||
mod forge_notify;
|
||||
mod harness_state;
|
||||
mod identity;
|
||||
mod login;
|
||||
|
|
@ -550,15 +549,12 @@ async fn serve_main<S: Surface>(socket: &Path, poll_ms: u64) -> Result<()> {
|
|||
for failure in plugins::install_configured().await {
|
||||
S::send_to_parent(socket, failure).await;
|
||||
}
|
||||
// forge_notify pushes forge notifications as todos on the in-agent
|
||||
// socket (loose-ends v2), not direct wakes — so it dials
|
||||
// `HIVE_AGENT_SOCKET`, not the host mcp.sock.
|
||||
tokio::spawn(crate::forge_notify::run(
|
||||
std::env::var_os("HIVE_AGENT_SOCKET").map_or_else(
|
||||
|| std::path::PathBuf::from(hive_agent_sock::DEFAULT_AGENT_SOCKET),
|
||||
std::path::PathBuf::from,
|
||||
),
|
||||
));
|
||||
// The forge notification poller used to be spawned here. It is its own
|
||||
// process now (`hive-forge-notify`, its own systemd unit) so a harness
|
||||
// restart doesn't take forge notifications down with it; it reaches the
|
||||
// harness the same way the bash and matrix daemons do, by upserting
|
||||
// todos on the in-agent socket.
|
||||
//
|
||||
// Agent-side cleanup of this agent's own harness artifacts (completed
|
||||
// bash-task files + verbose event rows). Runs here, not host-side in
|
||||
// hive-c0re, because the files are agent-owned — see `vacuum` module docs.
|
||||
|
|
|
|||
Loading…
Reference in a new issue