hive-c0re: serialise the matrix and knowledge sweeps through the job queue

The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
This commit is contained in:
atlas 2026-09-27 02:55:59 +02:00
commit 1d4c77d2c8
14 changed files with 352 additions and 105 deletions

View file

@ -62,14 +62,18 @@ pull`, so agents see the new content on their next turn.
points at a hive's own `/webhook/knowledge`.
2. **Periodic pull** — a background task in `hive-c0re::main`
pulls on a fixed cadence as a fallback (webhook missed, c0re
queues a pull on a fixed cadence as a fallback (webhook missed, c0re
restarted between pushes). The pull is best-effort — a failure
logs a warning and doesn't affect the rest of the daemon.
shows as a failed `KnowledgePull` node on the job queue, counts toward
the knowledge-pull banner, and doesn't affect the rest of the daemon.
Both paths share the same `knowledge::pull()` function, which also
handles the change notice below — neither path can forget to wire it
in since the broadcast logic lives once, in `pull()` itself, not at
each call site.
Both paths, and the one pull at boot, queue the same `KnowledgePull` job
node. It holds the working tree for its duration, so two pulls never run
over each other. An event that arrives while a pull is waiting to start
folds into it; one that arrives while a pull is running queues one more
behind it, since the running pull may have fetched before the push. The
node runs `knowledge::pull()`, which also handles the change notice
below.
### Change notice

View file

@ -81,9 +81,9 @@ Cheap — no build slot:
| `WriteDropin` | `set_nspawn_flags` + `set_resource_limits` + daemon-reload |
| `WritePermFile` | commit `tool-groups.json` / `capabilities.json` (single git commit under `META_LOCK`) + emit the P3RM1SS10NS snapshots |
| `ForgeSweep` | one-shot boot-time forge user/token sweep for every container (`forge::ensure_all`) as a first-class node, so it shows as real work on the dashboard instead of running invisibly in a bare `tokio::spawn`. Agentless |
| `MatrixSweep` | same as `ForgeSweep`, for matrix (`matrix::ensure_all`). The periodic 30-min re-sweep stays a background loop in `main.rs`; only the boot-time instance is a node |
| `MatrixSweep` | matrix user/space sweep (`matrix::ensure_all`): the boot-time instance, plus one every 30 min from a loop in `main.rs`. Holds `Resource::MatrixSweep` (capacity 1), so two passes never overlap; a periodic tick finding one queued or running folds into it instead of stacking. Agentless |
| `WebhookRegister` | one-shot boot-time Forgejo webhook registration (`internal/knowledge` push→pull, `agent-configs` PR→approval). No-op until the core token, hive domain, and HMAC secret are all available. Agentless |
| `KnowledgePull` | one-shot boot-time `/knowledge` pull (`knowledge::pull`), reconciling commits that landed while `hive-c0re` was down. Same rationale as `MatrixSweep`: the periodic hourly re-pull stays a background loop |
| `KnowledgePull` | `/knowledge` pull (`knowledge::pull`): at boot (commits that landed while `hive-c0re` was down), on the swarm knowledge-changed event, and hourly as a fallback. Holds `Resource::KnowledgeTree` (capacity 1), so two pulls never overlap on the working tree. The event folds into a queued pull but queues behind a running one; the hourly tick folds into either. Agentless |
| `WantedPull` | one-shot boot-time pull of the agent set the swarm controller declares for this hive (`wanted::pull`), converging the agents it names. No background loop behind this one — boot is the whole cadence; the deploy event (`swarm_status`) is the fast path, this repairs a missed one. Agentless |
<!-- vale write-good.Passive = YES -->