hive-c0re: serialise the matrix and knowledge sweeps through the job queue
The matrix sweep and the /knowledge pull each had concurrent callers (#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock and the remote-tracking ref lock: 30 of 30 concurrent replays of the reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential ones did. Two overlapping matrix sweeps on a hive with no persisted Space / chat-room id both miss the by-name lookup and both createRoom (from reading the code, not reproduced against a homeserver). On every boot the MatrixSweep DAG node and the main.rs loop's immediate first call ran at once. Every sweep now runs as a job node, and each sweep's node holds its own capacity-1 queue resource (Resource::MatrixSweep, Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never starts a second pass of one sweep while the first holds the resource, and different sweeps still run side by side. - templates::matrix_sweep / templates::knowledge_pull build the node with its resource; boot, the periodic loops and the swarm event all use them. - JobQueue::insert_unless_live folds a submission into a live node of the same kind instead of queueing another. Periodic ticks fold into a queued or running pass. The swarm knowledge event folds into a queued pull only, and queues one behind a running pull, which may have fetched before the push. - The main.rs matrix loop no longer sweeps immediately at startup; the boot MatrixSweep node is the startup pass, as KnowledgePull already was for knowledge. - The executors bound each pass (10 min matrix, 5 min knowledge), since a hung pass would otherwise hold its resource against every later one, and own the sweep-health banners, so every pass reports to them. Replaces the SweepLock version of this branch, per review. Refs #4723
This commit is contained in:
parent
2252c55df8
commit
1d4c77d2c8
14 changed files with 352 additions and 105 deletions
|
|
@ -6,7 +6,7 @@
|
|||
//! `Swap` tail); DAG-level failure handling is cancel-downstream in
|
||||
//! the queue.
|
||||
|
||||
use std::sync::Arc;
|
||||
use std::sync::{Arc, Mutex, OnceLock};
|
||||
|
||||
use anyhow::{Context as _, Result};
|
||||
|
||||
|
|
@ -15,6 +15,7 @@ use hive_jobq::{NodeId, TerminalState};
|
|||
use super::model::NodeKind;
|
||||
use crate::coordinator::Coordinator;
|
||||
use crate::power::{ReconcileAction, reconcile_action};
|
||||
use crate::stats::sweep_health::{self, SweepHealth};
|
||||
|
||||
/// Max time `Drain` waits for the harness to run its stop-checkpoint
|
||||
/// turn before falling back to the hard stop. Generous — a checkpoint
|
||||
|
|
@ -166,19 +167,57 @@ async fn run_forge_sweep() -> Result<()> {
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Boot-time matrix user/space sweep as a DAG node — see
|
||||
/// [`NodeKind::MatrixSweep`]. Reports failure as the node's own error so a
|
||||
/// failed boot sweep is visible on the dashboard; the debounced
|
||||
/// `sweep_health`-driven warning banner is a separate concern owned by the
|
||||
/// periodic loop in `main.rs`, unaffected by this node's own outcome.
|
||||
/// Matrix user/space sweep as a DAG node — see [`NodeKind::MatrixSweep`].
|
||||
/// Every pass reports here, whoever submitted it: as the node's own error on
|
||||
/// the dashboard, and into the debounced banner.
|
||||
async fn run_matrix_sweep() -> Result<()> {
|
||||
if crate::matrix::ensure_all().await {
|
||||
let ok = tokio::time::timeout(MATRIX_SWEEP_DEADLINE, crate::matrix::ensure_all())
|
||||
.await
|
||||
.unwrap_or_else(|_| {
|
||||
tracing::warn!(deadline = ?MATRIX_SWEEP_DEADLINE, "matrix: sweep timed out");
|
||||
false
|
||||
});
|
||||
let mut health = matrix_sweep_health()
|
||||
.lock()
|
||||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||||
if ok {
|
||||
health.record_ok();
|
||||
Ok(())
|
||||
} else {
|
||||
health.record_err(matrix_sweep_banner);
|
||||
anyhow::bail!("matrix ensure_all: one or more agents failed sync (see logs)")
|
||||
}
|
||||
}
|
||||
|
||||
/// A pass still running holds [`Resource::MatrixSweep`] against every later
|
||||
/// one. Each HTTP call has its own 10 s timeout, but the call count grows with
|
||||
/// the agent count and `lifecycle::list()` has none. A pass dropped here leaves
|
||||
/// at worst a created but unpersisted room, which the next pass adopts by name.
|
||||
///
|
||||
/// [`Resource::MatrixSweep`]: super::resource::Resource::MatrixSweep
|
||||
const MATRIX_SWEEP_DEADLINE: std::time::Duration = std::time::Duration::from_mins(10);
|
||||
|
||||
/// Banner for a matrix sweep that keeps failing. Two misses in a row, so one
|
||||
/// bad pass (homeserver mid-restart, a transient HTTP blip) doesn't flap the
|
||||
/// dashboard; cleared by the next clean pass.
|
||||
fn matrix_sweep_health() -> &'static Mutex<SweepHealth> {
|
||||
static HEALTH: OnceLock<Mutex<SweepHealth>> = OnceLock::new();
|
||||
HEALTH.get_or_init(|| Mutex::new(SweepHealth::new("matrix_ensure_all", "warn", 2)))
|
||||
}
|
||||
|
||||
fn matrix_sweep_banner(ctx: sweep_health::SweepFailure) -> String {
|
||||
let age = ctx.since_last_ok.map_or_else(
|
||||
|| "no success this session".to_owned(),
|
||||
|d| format!("last ok {} ago", sweep_health::fmt_age(d)),
|
||||
);
|
||||
format!(
|
||||
"matrix user/space sweep failing ({} consecutive, {age}) \
|
||||
— some agents may be missing matrix accounts, space membership, \
|
||||
or chat-room invites",
|
||||
ctx.consecutive
|
||||
)
|
||||
}
|
||||
|
||||
/// Boot-time Forgejo webhook management as a DAG node — see
|
||||
/// [`NodeKind::WebhookRegister`]. Mirrors the guard chain the
|
||||
/// `tokio::spawn` block it replaced used: no-op (not an error) when the
|
||||
|
|
@ -207,13 +246,55 @@ async fn run_webhook_register() -> Result<()> {
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Boot-time `/knowledge` pull as a DAG node — see
|
||||
/// [`NodeKind::KnowledgePull`]. Unlike the `main.rs` periodic loop's
|
||||
/// startup call, a failure here is *not* swallowed to debug level: the node
|
||||
/// exists so a failed boot pull is visible on the dashboard rather than
|
||||
/// only in the journal.
|
||||
/// `/knowledge` pull as a DAG node — see [`NodeKind::KnowledgePull`]. Every
|
||||
/// pass reports here, whoever submitted it: as the node's own error on the
|
||||
/// dashboard, and into the debounced banner.
|
||||
async fn run_knowledge_pull(coord: &Arc<Coordinator>) -> Result<()> {
|
||||
crate::workers::knowledge::pull(coord).await
|
||||
let result = tokio::time::timeout(
|
||||
KNOWLEDGE_PULL_DEADLINE,
|
||||
crate::workers::knowledge::pull(coord),
|
||||
)
|
||||
.await
|
||||
.unwrap_or_else(|_| {
|
||||
Err(anyhow::anyhow!(
|
||||
"knowledge pull timed out after {KNOWLEDGE_PULL_DEADLINE:?}"
|
||||
))
|
||||
});
|
||||
let mut health = knowledge_pull_health()
|
||||
.lock()
|
||||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||||
match &result {
|
||||
Ok(()) => health.record_ok(),
|
||||
Err(e) => {
|
||||
let err = format!("{e:#}");
|
||||
health.record_err(|ctx| {
|
||||
let age = ctx.since_last_ok.map_or_else(
|
||||
|| "no success this session".to_owned(),
|
||||
|d| format!("last ok {} ago", sweep_health::fmt_age(d)),
|
||||
);
|
||||
format!(
|
||||
"knowledge repo pull failing ({} consecutive, {age}) \
|
||||
— /knowledge is stale until it recovers: {err}",
|
||||
ctx.consecutive
|
||||
)
|
||||
});
|
||||
}
|
||||
}
|
||||
result
|
||||
}
|
||||
|
||||
/// A pull still running holds [`Resource::KnowledgeTree`] against every later
|
||||
/// one. The `reset` / `clean` / `pull` git children are `kill_on_drop`, so the
|
||||
/// deadline kills them.
|
||||
///
|
||||
/// [`Resource::KnowledgeTree`]: super::resource::Resource::KnowledgeTree
|
||||
const KNOWLEDGE_PULL_DEADLINE: std::time::Duration = std::time::Duration::from_mins(5);
|
||||
|
||||
/// Banner for a `/knowledge` pull that keeps failing: three misses in a row,
|
||||
/// cleared by the next successful pull.
|
||||
fn knowledge_pull_health() -> &'static Mutex<SweepHealth> {
|
||||
static HEALTH: OnceLock<Mutex<SweepHealth>> = OnceLock::new();
|
||||
HEALTH.get_or_init(|| Mutex::new(SweepHealth::new("knowledge_pull", "warn", 3)))
|
||||
}
|
||||
|
||||
/// Boot-time swarm wanted-state pull as a DAG node — see
|
||||
|
|
|
|||
|
|
@ -192,6 +192,39 @@ impl JobQueue {
|
|||
Ok(named)
|
||||
}
|
||||
|
||||
/// Insert the single node `declare` names, unless a node of `kind` is
|
||||
/// already in one of the `fold_into` states. `None` means the caller's
|
||||
/// request folded into that node and nothing was inserted.
|
||||
///
|
||||
/// For a standalone sweep, where a pass that has not started yet already
|
||||
/// covers "one more pass": it reads whatever is current when it runs. The
|
||||
/// read and the insert share one lock, so two racing callers cannot both
|
||||
/// miss the other.
|
||||
///
|
||||
/// # Errors
|
||||
/// Propagates a graph-insert error.
|
||||
pub fn insert_unless_live(
|
||||
&self,
|
||||
kind: &NodeKind,
|
||||
fold_into: &[State],
|
||||
declare: impl FnOnce(&JobBuilder) -> Handle<'_>,
|
||||
) -> anyhow::Result<Option<NodeId>> {
|
||||
let mut inner = self.lock();
|
||||
let live = inner
|
||||
.graph()
|
||||
.nodes()
|
||||
.any(|n| &n.payload == kind && fold_into.contains(&n.state));
|
||||
if live {
|
||||
return Ok(None);
|
||||
}
|
||||
let named = inner
|
||||
.insert_job(None, |b| vec![declare(b).guid()])
|
||||
.map_err(|e| anyhow::anyhow!("job_queue: graph insert failed: {e}"))?;
|
||||
drop(inner);
|
||||
self.notify.notify_one();
|
||||
Ok(named.first().copied())
|
||||
}
|
||||
|
||||
/// The scheduler itself, for `hive_jobq`'s run-loop seam
|
||||
/// (`Scheduler::claim_next`), which takes exactly this type.
|
||||
///
|
||||
|
|
|
|||
|
|
@ -144,12 +144,13 @@ pub enum NodeKind {
|
|||
/// One-shot boot-time forge user/token sweep for every existing
|
||||
/// container (`forge::ensure_all`). Agentless.
|
||||
ForgeSweep,
|
||||
/// One-shot boot-time matrix user/space sweep (`matrix::ensure_all`).
|
||||
/// Agentless.
|
||||
/// Matrix user/space sweep (`matrix::ensure_all`): at boot and every 30
|
||||
/// minutes. Agentless.
|
||||
MatrixSweep,
|
||||
/// One-shot boot-time Forgejo webhook registration. Agentless.
|
||||
WebhookRegister,
|
||||
/// One-shot boot-time `/knowledge` pull (`knowledge::pull`). Agentless.
|
||||
/// `/knowledge` pull (`knowledge::pull`): at boot, on the swarm
|
||||
/// knowledge-changed event, and hourly. Agentless.
|
||||
KnowledgePull,
|
||||
/// One-shot boot-time pull of the agent set the swarm controller
|
||||
/// declares for this hive (`wanted::pull`). Agentless.
|
||||
|
|
|
|||
|
|
@ -31,6 +31,15 @@ pub enum Resource {
|
|||
/// The meta-repo mutation window — a global singleton held by any node
|
||||
/// that mutates the meta repo, so two meta mutations never interleave.
|
||||
MetaWindow,
|
||||
/// The `/knowledge` working tree, held by every `KnowledgePull`. Two
|
||||
/// overlapping `reset` / `clean` / `pull` passes there fail on
|
||||
/// `.git/index.lock` and the remote-tracking ref lock.
|
||||
KnowledgeTree,
|
||||
/// The hive's matrix provisioning, held by every `MatrixSweep`. Two
|
||||
/// overlapping passes on a hive with no persisted Space / chat-room id both
|
||||
/// miss the by-name lookup and both `createRoom`, leaving a duplicate room
|
||||
/// with agents invited to both.
|
||||
MatrixSweep,
|
||||
}
|
||||
|
||||
/// This resource's name on the generic graph wire.
|
||||
|
|
@ -45,6 +54,8 @@ impl hive_jobq_wire::WireResource for Resource {
|
|||
Resource::BuildSlot => "build-slot".to_owned(),
|
||||
Resource::Agent(agent) => format!("agent:{agent}"),
|
||||
Resource::MetaWindow => "meta-window".to_owned(),
|
||||
Resource::KnowledgeTree => "knowledge-tree".to_owned(),
|
||||
Resource::MatrixSweep => "matrix-sweep".to_owned(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -205,3 +205,104 @@ fn reconcile_transients(coord: &Arc<Coordinator>, prev: &mut TransientSeen) {
|
|||
prev.insert(key, t.takes_container_down);
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use hive_jobq::scheduler::{Outcome, Scheduler};
|
||||
|
||||
use super::super::{JobQueue, NodeKind, State, templates};
|
||||
|
||||
/// Claim one runnable node through the same seam [`super::run_worker`]
|
||||
/// uses. The returned future completes the node when awaited. `None` when
|
||||
/// nothing is runnable.
|
||||
fn claim(q: &JobQueue) -> Option<impl Future<Output = ()>> {
|
||||
Scheduler::claim_next(q.sched(), |_, _, builder| async move {
|
||||
(builder, Outcome::Done)
|
||||
})
|
||||
.map(|run| async move {
|
||||
run.await.1.expect("a sweep grows nothing");
|
||||
})
|
||||
}
|
||||
|
||||
/// Kinds of the nodes currently `Running`, sorted.
|
||||
fn running(q: &JobQueue) -> Vec<&'static str> {
|
||||
let sched = q.sched().lock().expect("job_queue mutex poisoned");
|
||||
let mut kinds: Vec<&'static str> = sched
|
||||
.graph()
|
||||
.nodes()
|
||||
.filter(|n| n.state == State::Running)
|
||||
.map(|n| <&'static str>::from(&n.payload))
|
||||
.collect();
|
||||
kinds.sort_unstable();
|
||||
kinds
|
||||
}
|
||||
|
||||
fn count(q: &JobQueue, kind: &str) -> usize {
|
||||
let sched = q.sched().lock().expect("job_queue mutex poisoned");
|
||||
sched
|
||||
.graph()
|
||||
.nodes()
|
||||
.filter(|n| <&str>::from(&n.payload) == kind)
|
||||
.count()
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn two_passes_of_one_sweep_never_run_together() {
|
||||
let q = JobQueue::new(1);
|
||||
for _ in 0..2 {
|
||||
q.insert_job(|b| vec![templates::matrix_sweep(b).guid()])
|
||||
.expect("insert");
|
||||
}
|
||||
q.insert_job(|b| vec![templates::knowledge_pull(b).guid()])
|
||||
.expect("insert");
|
||||
|
||||
let first = claim(&q).expect("a matrix pass is runnable");
|
||||
let other = claim(&q).expect("a different sweep runs alongside it");
|
||||
assert!(
|
||||
claim(&q).is_none(),
|
||||
"the second matrix pass waits for the first"
|
||||
);
|
||||
assert_eq!(running(&q), ["knowledge_pull", "matrix_sweep"]);
|
||||
|
||||
first.await;
|
||||
other.await;
|
||||
let second = claim(&q).expect("the second matrix pass runs once the first is done");
|
||||
assert_eq!(running(&q), ["matrix_sweep"]);
|
||||
second.await;
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_tick_folds_into_a_live_pull_and_an_event_queues_one_behind_it() {
|
||||
let q = JobQueue::new(1);
|
||||
let kind = NodeKind::KnowledgePull;
|
||||
let queued = [State::Pending];
|
||||
let live = [State::Pending, State::Running];
|
||||
let submit = |fold_into: &[State]| {
|
||||
q.insert_unless_live(&kind, fold_into, templates::knowledge_pull)
|
||||
.expect("insert")
|
||||
.is_some()
|
||||
};
|
||||
|
||||
assert!(submit(&live), "nothing live: a tick inserts");
|
||||
let pull = claim(&q).expect("the pull is runnable");
|
||||
assert!(!submit(&live), "a tick folds into the running pull");
|
||||
assert!(
|
||||
submit(&queued),
|
||||
"an event queues a pull behind the running one"
|
||||
);
|
||||
assert!(
|
||||
!submit(&queued),
|
||||
"a second event folds into the queued pull"
|
||||
);
|
||||
assert!(!submit(&live), "a tick folds into the queued pull");
|
||||
assert_eq!(count(&q, "knowledge_pull"), 2);
|
||||
|
||||
assert!(
|
||||
q.insert_unless_live(&NodeKind::MatrixSweep, &live, templates::matrix_sweep)
|
||||
.expect("insert")
|
||||
.is_some(),
|
||||
"a live pull does not fold a different sweep"
|
||||
);
|
||||
pull.await;
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -601,3 +601,18 @@ pub fn meta_update(builder: &JobBuilder, inputs: Vec<String>, approval_id: Optio
|
|||
// as ONE `Boot` DAG (a sweep `MetaLock` root that grows rebuild subgraphs
|
||||
// in-DAG, plus a `Reconcile` root per drifted agent) — no anchor node and no
|
||||
// per-agent child DAGs.
|
||||
|
||||
/// One matrix user/space sweep. Holding [`Resource::MatrixSweep`] is what keeps
|
||||
/// two passes from overlapping, whichever caller submitted each.
|
||||
pub fn matrix_sweep(builder: &JobBuilder) -> Handle<'_> {
|
||||
builder
|
||||
.node(NodeKind::MatrixSweep)
|
||||
.needs(Resource::MatrixSweep)
|
||||
}
|
||||
|
||||
/// One `/knowledge` pull, holding [`Resource::KnowledgeTree`] for its duration.
|
||||
pub fn knowledge_pull(builder: &JobBuilder) -> Handle<'_> {
|
||||
builder
|
||||
.node(NodeKind::KnowledgePull)
|
||||
.needs(Resource::KnowledgeTree)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1773,3 +1773,26 @@ fn perm_change_shape_prefixes_rebuild_chain() {
|
|||
"the perm write prefixes an otherwise ordinary rebuild chain"
|
||||
);
|
||||
}
|
||||
|
||||
// ---- standalone sweeps ----
|
||||
|
||||
#[test]
|
||||
fn each_sweep_holds_its_own_resource() {
|
||||
// Same kind, same capacity-1 resource: two passes serialise. Different
|
||||
// kinds, disjoint resources and no edges: they run side by side. That the
|
||||
// scheduler honours this is `hive_jobq`'s
|
||||
// `unrelated_nodes_needing_the_same_resource_are_serialized`.
|
||||
let q = JobQueue::new(1);
|
||||
insert(&q, |builder| {
|
||||
let _ = templates::matrix_sweep(builder);
|
||||
let _ = templates::knowledge_pull(builder);
|
||||
});
|
||||
assert_eq!(
|
||||
declared_resources_of_kind(&q, "matrix_sweep"),
|
||||
vec![vec![Resource::MatrixSweep]]
|
||||
);
|
||||
assert_eq!(
|
||||
declared_resources_of_kind(&q, "knowledge_pull"),
|
||||
vec![vec![Resource::KnowledgeTree]]
|
||||
);
|
||||
}
|
||||
|
|
|
|||
Loading…
Reference in a new issue