Watch
0
0
Fork
You've already forked hyperhive
0

hive-c0re: serialise the matrix and knowledge sweeps through the job queue

The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
This commit is contained in:
atlas 2026-09-27 02:55:59 +02:00
commit 1d4c77d2c8
14 changed files with 352 additions and 105 deletions

View file

@ -6,7 +6,7 @@
//! `Swap` tail); DAG-level failure handling is cancel-downstream in
//! the queue.
use std::sync::Arc;
use std::sync::{Arc, Mutex, OnceLock};
use anyhow::{Context as _, Result};
@ -15,6 +15,7 @@ use hive_jobq::{NodeId, TerminalState};
use super::model::NodeKind;
use crate::coordinator::Coordinator;
use crate::power::{ReconcileAction, reconcile_action};
use crate::stats::sweep_health::{self, SweepHealth};
/// Max time `Drain` waits for the harness to run its stop-checkpoint
/// turn before falling back to the hard stop. Generous — a checkpoint
@ -166,19 +167,57 @@ async fn run_forge_sweep() -> Result<()> {
Ok(())
}
/// Boot-time matrix user/space sweep as a DAG node — see
/// [`NodeKind::MatrixSweep`]. Reports failure as the node's own error so a
/// failed boot sweep is visible on the dashboard; the debounced
/// `sweep_health`-driven warning banner is a separate concern owned by the
/// periodic loop in `main.rs`, unaffected by this node's own outcome.
/// Matrix user/space sweep as a DAG node — see [`NodeKind::MatrixSweep`].
/// Every pass reports here, whoever submitted it: as the node's own error on
/// the dashboard, and into the debounced banner.
async fn run_matrix_sweep() -> Result<()> {
if crate::matrix::ensure_all().await {
let ok = tokio::time::timeout(MATRIX_SWEEP_DEADLINE, crate::matrix::ensure_all())
.await
.unwrap_or_else(|_| {
tracing::warn!(deadline = ?MATRIX_SWEEP_DEADLINE, "matrix: sweep timed out");
false
});
let mut health = matrix_sweep_health()
.lock()
.unwrap_or_else(std::sync::PoisonError::into_inner);
if ok {
health.record_ok();
Ok(())
} else {
health.record_err(matrix_sweep_banner);
anyhow::bail!("matrix ensure_all: one or more agents failed sync (see logs)")
}
}
/// A pass still running holds [`Resource::MatrixSweep`] against every later
/// one. Each HTTP call has its own 10 s timeout, but the call count grows with
/// the agent count and `lifecycle::list()` has none. A pass dropped here leaves
/// at worst a created but unpersisted room, which the next pass adopts by name.
///
/// [`Resource::MatrixSweep`]: super::resource::Resource::MatrixSweep
const MATRIX_SWEEP_DEADLINE: std::time::Duration = std::time::Duration::from_mins(10);
/// Banner for a matrix sweep that keeps failing. Two misses in a row, so one
/// bad pass (homeserver mid-restart, a transient HTTP blip) doesn't flap the
/// dashboard; cleared by the next clean pass.
fn matrix_sweep_health() -> &'static Mutex<SweepHealth> {
static HEALTH: OnceLock<Mutex<SweepHealth>> = OnceLock::new();
HEALTH.get_or_init(|| Mutex::new(SweepHealth::new("matrix_ensure_all", "warn", 2)))
}
fn matrix_sweep_banner(ctx: sweep_health::SweepFailure) -> String {
let age = ctx.since_last_ok.map_or_else(
|| "no success this session".to_owned(),
|d| format!("last ok {} ago", sweep_health::fmt_age(d)),
);
format!(
"matrix user/space sweep failing ({} consecutive, {age}) \
— some agents may be missing matrix accounts, space membership, \
or chat-room invites",
ctx.consecutive
)
}
/// Boot-time Forgejo webhook management as a DAG node — see
/// [`NodeKind::WebhookRegister`]. Mirrors the guard chain the
/// `tokio::spawn` block it replaced used: no-op (not an error) when the
@ -207,13 +246,55 @@ async fn run_webhook_register() -> Result<()> {
Ok(())
}
/// Boot-time `/knowledge` pull as a DAG node — see
/// [`NodeKind::KnowledgePull`]. Unlike the `main.rs` periodic loop's
/// startup call, a failure here is *not* swallowed to debug level: the node
/// exists so a failed boot pull is visible on the dashboard rather than
/// only in the journal.
/// `/knowledge` pull as a DAG node — see [`NodeKind::KnowledgePull`]. Every
/// pass reports here, whoever submitted it: as the node's own error on the
/// dashboard, and into the debounced banner.
async fn run_knowledge_pull(coord: &Arc<Coordinator>) -> Result<()> {
crate::workers::knowledge::pull(coord).await
let result = tokio::time::timeout(
KNOWLEDGE_PULL_DEADLINE,
crate::workers::knowledge::pull(coord),
)
.await
.unwrap_or_else(|_| {
Err(anyhow::anyhow!(
"knowledge pull timed out after {KNOWLEDGE_PULL_DEADLINE:?}"
))
});
let mut health = knowledge_pull_health()
.lock()
.unwrap_or_else(std::sync::PoisonError::into_inner);
match &result {
Ok(()) => health.record_ok(),
Err(e) => {
let err = format!("{e:#}");
health.record_err(|ctx| {
let age = ctx.since_last_ok.map_or_else(
|| "no success this session".to_owned(),
|d| format!("last ok {} ago", sweep_health::fmt_age(d)),
);
format!(
"knowledge repo pull failing ({} consecutive, {age}) \
— /knowledge is stale until it recovers: {err}",
ctx.consecutive
)
});
}
}
result
}
/// A pull still running holds [`Resource::KnowledgeTree`] against every later
/// one. The `reset` / `clean` / `pull` git children are `kill_on_drop`, so the
/// deadline kills them.
///
/// [`Resource::KnowledgeTree`]: super::resource::Resource::KnowledgeTree
const KNOWLEDGE_PULL_DEADLINE: std::time::Duration = std::time::Duration::from_mins(5);
/// Banner for a `/knowledge` pull that keeps failing: three misses in a row,
/// cleared by the next successful pull.
fn knowledge_pull_health() -> &'static Mutex<SweepHealth> {
static HEALTH: OnceLock<Mutex<SweepHealth>> = OnceLock::new();
HEALTH.get_or_init(|| Mutex::new(SweepHealth::new("knowledge_pull", "warn", 3)))
}
/// Boot-time swarm wanted-state pull as a DAG node — see

View file

@ -192,6 +192,39 @@ impl JobQueue {
Ok(named)
}
/// Insert the single node `declare` names, unless a node of `kind` is
/// already in one of the `fold_into` states. `None` means the caller's
/// request folded into that node and nothing was inserted.
///
/// For a standalone sweep, where a pass that has not started yet already
/// covers "one more pass": it reads whatever is current when it runs. The
/// read and the insert share one lock, so two racing callers cannot both
/// miss the other.
///
/// # Errors
/// Propagates a graph-insert error.
pub fn insert_unless_live(
&self,
kind: &NodeKind,
fold_into: &[State],
declare: impl FnOnce(&JobBuilder) -> Handle<'_>,
) -> anyhow::Result<Option<NodeId>> {
let mut inner = self.lock();
let live = inner
.graph()
.nodes()
.any(|n| &n.payload == kind && fold_into.contains(&n.state));
if live {
return Ok(None);
}
let named = inner
.insert_job(None, |b| vec![declare(b).guid()])
.map_err(|e| anyhow::anyhow!("job_queue: graph insert failed: {e}"))?;
drop(inner);
self.notify.notify_one();
Ok(named.first().copied())
}
/// The scheduler itself, for `hive_jobq`'s run-loop seam
/// (`Scheduler::claim_next`), which takes exactly this type.
///

View file

@ -144,12 +144,13 @@ pub enum NodeKind {
/// One-shot boot-time forge user/token sweep for every existing
/// container (`forge::ensure_all`). Agentless.
ForgeSweep,
/// One-shot boot-time matrix user/space sweep (`matrix::ensure_all`).
/// Agentless.
/// Matrix user/space sweep (`matrix::ensure_all`): at boot and every 30
/// minutes. Agentless.
MatrixSweep,
/// One-shot boot-time Forgejo webhook registration. Agentless.
WebhookRegister,
/// One-shot boot-time `/knowledge` pull (`knowledge::pull`). Agentless.
/// `/knowledge` pull (`knowledge::pull`): at boot, on the swarm
/// knowledge-changed event, and hourly. Agentless.
KnowledgePull,
/// One-shot boot-time pull of the agent set the swarm controller
/// declares for this hive (`wanted::pull`). Agentless.

View file

@ -31,6 +31,15 @@ pub enum Resource {
/// The meta-repo mutation window — a global singleton held by any node
/// that mutates the meta repo, so two meta mutations never interleave.
MetaWindow,
/// The `/knowledge` working tree, held by every `KnowledgePull`. Two
/// overlapping `reset` / `clean` / `pull` passes there fail on
/// `.git/index.lock` and the remote-tracking ref lock.
KnowledgeTree,
/// The hive's matrix provisioning, held by every `MatrixSweep`. Two
/// overlapping passes on a hive with no persisted Space / chat-room id both
/// miss the by-name lookup and both `createRoom`, leaving a duplicate room
/// with agents invited to both.
MatrixSweep,
}
/// This resource's name on the generic graph wire.
@ -45,6 +54,8 @@ impl hive_jobq_wire::WireResource for Resource {
Resource::BuildSlot => "build-slot".to_owned(),
Resource::Agent(agent) => format!("agent:{agent}"),
Resource::MetaWindow => "meta-window".to_owned(),
Resource::KnowledgeTree => "knowledge-tree".to_owned(),
Resource::MatrixSweep => "matrix-sweep".to_owned(),
}
}
}

View file

@ -205,3 +205,104 @@ fn reconcile_transients(coord: &Arc<Coordinator>, prev: &mut TransientSeen) {
prev.insert(key, t.takes_container_down);
}
}
#[cfg(test)]
mod tests {
use hive_jobq::scheduler::{Outcome, Scheduler};
use super::super::{JobQueue, NodeKind, State, templates};
/// Claim one runnable node through the same seam [`super::run_worker`]
/// uses. The returned future completes the node when awaited. `None` when
/// nothing is runnable.
fn claim(q: &JobQueue) -> Option<impl Future<Output = ()>> {
Scheduler::claim_next(q.sched(), |_, _, builder| async move {
(builder, Outcome::Done)
})
.map(|run| async move {
run.await.1.expect("a sweep grows nothing");
})
}
/// Kinds of the nodes currently `Running`, sorted.
fn running(q: &JobQueue) -> Vec<&'static str> {
let sched = q.sched().lock().expect("job_queue mutex poisoned");
let mut kinds: Vec<&'static str> = sched
.graph()
.nodes()
.filter(|n| n.state == State::Running)
.map(|n| <&'static str>::from(&n.payload))
.collect();
kinds.sort_unstable();
kinds
}
fn count(q: &JobQueue, kind: &str) -> usize {
let sched = q.sched().lock().expect("job_queue mutex poisoned");
sched
.graph()
.nodes()
.filter(|n| <&str>::from(&n.payload) == kind)
.count()
}
#[tokio::test]
async fn two_passes_of_one_sweep_never_run_together() {
let q = JobQueue::new(1);
for _ in 0..2 {
q.insert_job(|b| vec![templates::matrix_sweep(b).guid()])
.expect("insert");
}
q.insert_job(|b| vec![templates::knowledge_pull(b).guid()])
.expect("insert");
let first = claim(&q).expect("a matrix pass is runnable");
let other = claim(&q).expect("a different sweep runs alongside it");
assert!(
claim(&q).is_none(),
"the second matrix pass waits for the first"
);
assert_eq!(running(&q), ["knowledge_pull", "matrix_sweep"]);
first.await;
other.await;
let second = claim(&q).expect("the second matrix pass runs once the first is done");
assert_eq!(running(&q), ["matrix_sweep"]);
second.await;
}
#[tokio::test]
async fn a_tick_folds_into_a_live_pull_and_an_event_queues_one_behind_it() {
let q = JobQueue::new(1);
let kind = NodeKind::KnowledgePull;
let queued = [State::Pending];
let live = [State::Pending, State::Running];
let submit = |fold_into: &[State]| {
q.insert_unless_live(&kind, fold_into, templates::knowledge_pull)
.expect("insert")
.is_some()
};
assert!(submit(&live), "nothing live: a tick inserts");
let pull = claim(&q).expect("the pull is runnable");
assert!(!submit(&live), "a tick folds into the running pull");
assert!(
submit(&queued),
"an event queues a pull behind the running one"
);
assert!(
!submit(&queued),
"a second event folds into the queued pull"
);
assert!(!submit(&live), "a tick folds into the queued pull");
assert_eq!(count(&q, "knowledge_pull"), 2);
assert!(
q.insert_unless_live(&NodeKind::MatrixSweep, &live, templates::matrix_sweep)
.expect("insert")
.is_some(),
"a live pull does not fold a different sweep"
);
pull.await;
}
}

View file

@ -601,3 +601,18 @@ pub fn meta_update(builder: &JobBuilder, inputs: Vec<String>, approval_id: Optio
// as ONE `Boot` DAG (a sweep `MetaLock` root that grows rebuild subgraphs
// in-DAG, plus a `Reconcile` root per drifted agent) — no anchor node and no
// per-agent child DAGs.
/// One matrix user/space sweep. Holding [`Resource::MatrixSweep`] is what keeps
/// two passes from overlapping, whichever caller submitted each.
pub fn matrix_sweep(builder: &JobBuilder) -> Handle<'_> {
builder
.node(NodeKind::MatrixSweep)
.needs(Resource::MatrixSweep)
}
/// One `/knowledge` pull, holding [`Resource::KnowledgeTree`] for its duration.
pub fn knowledge_pull(builder: &JobBuilder) -> Handle<'_> {
builder
.node(NodeKind::KnowledgePull)
.needs(Resource::KnowledgeTree)
}

View file

@ -1773,3 +1773,26 @@ fn perm_change_shape_prefixes_rebuild_chain() {
"the perm write prefixes an otherwise ordinary rebuild chain"
);
}
// ---- standalone sweeps ----
#[test]
fn each_sweep_holds_its_own_resource() {
// Same kind, same capacity-1 resource: two passes serialise. Different
// kinds, disjoint resources and no edges: they run side by side. That the
// scheduler honours this is `hive_jobq`'s
// `unrelated_nodes_needing_the_same_resource_are_serialized`.
let q = JobQueue::new(1);
insert(&q, |builder| {
let _ = templates::matrix_sweep(builder);
let _ = templates::knowledge_pull(builder);
});
assert_eq!(
declared_resources_of_kind(&q, "matrix_sweep"),
vec![vec![Resource::MatrixSweep]]
);
assert_eq!(
declared_resources_of_kind(&q, "knowledge_pull"),
vec![vec![Resource::KnowledgeTree]]
);
}