hyperhive/docs/integrations/knowledge.md
atlas 1d4c77d2c8 hive-c0re: serialise the matrix and knowledge sweeps through the job queue
The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
2026-09-27 20:15:55 +02:00

6.2 KiB

Hive-wide knowledge repository

internal/knowledge on the forge is a shared reference doc repo readable by every agent. hive-c0re clones it to the host and bind-mounts the clone read-only into every agent container at /knowledge.

Agent access

Inside any agent container:

/knowledge/          # read-only bind-mount of the local clone
/knowledge/README.md # table of contents (seeded on first use)

Agents read documents directly from that path. The mount is read-only — agents never write through it. To contribute, use the hive-forge AGit flow (no fork needed — see Contributing); the operator reviews and merges, and the local clone updates automatically (see Sync mechanism below).

Repository layout

Canonical forge location: internal/knowledge (org internal, repo knowledge). The repo is public, so every agent's forge account has read access without an explicit per-agent collaborator grant; only the core account has push access.

The swarm controller creates the repo if it doesn't exist, at startup and every few minutes after, and seeds an empty one with a README.md containing a contribution guide and a blank table of contents. Add an entry to that ToC each time you create a new document.

Sync mechanism

hive-c0re maintains the local clone at /var/lib/hyperhive/knowledge via two paths:

  1. Swarm event — the swarm controller holds the single push hook on internal/knowledge (see docs/swarm/README.md § Swarm-wide forge webhooks). On any push to main, including merge commits, it sends an event to every hive over the swarm queue and each hive runs git pull, so agents see the new content on their next turn.

    A hive that's offline when the swarm controller sends the event doesn't get it on reconnect — the periodic pull below is what closes that gap. One hive briefly showing older /knowledge content than another is expected, and resolves by itself within the fallback interval.

    don't add a per-hive hook. A webhook has exactly one target URL, so a second registration against the same repo doesn't add a recipient — it takes delivery away from whoever registered first. Earlier versions had each hive register its own, and later ones removed that leftover at startup. A hive upgraded straight past those versions still holds its old hook: delete any hook on the repo that points at a hive's own /webhook/knowledge.

  2. Periodic pull — a background task in hive-c0re::main queues a pull on a fixed cadence as a fallback (webhook missed, c0re restarted between pushes). The pull is best-effort — a failure shows as a failed KnowledgePull node on the job queue, counts toward the knowledge-pull banner, and doesn't affect the rest of the daemon.

Both paths, and the one pull at boot, queue the same KnowledgePull job node. It holds the working tree for its duration, so two pulls never run over each other. An event that arrives while a pull is waiting to start folds into it; one that arrives while a pull is running queues one more behind it, since the running pull may have fetched before the push. The node runs knowledge::pull(), which also handles the change notice below.

Change notice

When a pull actually moves the local clone's HEAD (a real change, not a no-op — for example the periodic pull finding nothing new), hive-c0re broadcasts a short notice to every currently registered agent's inbox: sender system, body [system] /knowledge updated: followed by a git diff --stat <old>..<new> summary of what changed (or a generic "see the repo" fallback if computing the diff itself fails). This is the same broadcast mechanism used for other hive-wide notices — inbox message only, no forced wake, and it carries the standard "this was a broadcast" hint. hive-c0re logs a diff or per-agent send failure without blocking the pull itself.

knowledge::ensure_local_clone creates (or refreshes) the local clone once at startup. If the repo is brand new and empty, ensure_local_clone seeds it with the default README before returning.

State

  • Host clone: /var/lib/hyperhive/knowledge — persists across hive-c0re restarts and agent destroy/recreate. Deleted only by manual operator action.
  • In-container mount: /knowledge — bind-mounted read-only from the host clone on every container start. Gone when container is stopped; reappears on next start with the current clone state.

The mount deliberately excludes .git: the host clone embeds the core token in .git/config (it rides the clone URL), so hive-priv overlays an empty tmpfs at /knowledge/.git — agents see the documents, not the repo metadata or token.

Contributing

Agents have read access to internal/knowledge (it's public) but no write access, so they can't push a branch directly. The supported path is Forgejo's AGit flow through the hive-forge CLI — no fork required.

  1. Clone the repo (hive-forge injects credentials automatically; the -r flag selects the repo, the clone lands in ./knowledge):

    hive-forge -r internal/knowledge clone
    cd knowledge
    
  2. Create a branch and add or update a document, then commit normally.

  3. Open (or update) a PR with --agit. This pushes the current HEAD to refs/for/<base>/<topic>, which Forgejo turns into a PR even though you can't push a branch:

    hive-forge -r internal/knowledge pr create --agit \
      --title "docs: add the X runbook" \
      --topic add-x-runbook \
      --body-file - <<'EOF'
    What this document adds and why.
    EOF
    

    Re-running with the same --topic updates the open PR (it force-pushes the AGit scratch ref). The base defaults to main; in --agit mode the push goes to the origin remote that hive-forge clone set up.

  4. The operator reviews and merges. The webhook fires on merge; every running container sees the updated content within seconds (see Sync mechanism).

Do not try to git push a branch directly — lacking write access, it's rejected. The --agit flow above is the no-fork path that works from any agent.