hive-c0re: serialise the matrix and knowledge sweeps through the job queue

The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
This commit is contained in:
atlas 2026-09-27 02:55:59 +02:00
commit 1d4c77d2c8
14 changed files with 352 additions and 105 deletions

View file

@ -238,8 +238,16 @@ async fn drain_swarm_events(
return;
}
tracing::info!(%subject, "swarm events: knowledge change announced, pulling");
if let Err(e) = crate::workers::knowledge::pull(&coord).await {
tracing::warn!(error = ?e, "swarm events: knowledge pull failed");
// Folds into a pull that hasn't started, but queues behind a
// running one: that pull may have fetched before this push.
match coord.job_queue.insert_unless_live(
&crate::job_queue::NodeKind::KnowledgePull,
&[crate::job_queue::State::Pending],
crate::job_queue::templates::knowledge_pull,
) {
Ok(Some(_)) => {}
Ok(None) => tracing::debug!("swarm events: a knowledge pull is already queued"),
Err(e) => tracing::warn!(error = ?e, "swarm events: knowledge pull submit failed"),
}
}
msg = deploy_sub.next() => {