hyperhive/docs/knowledge.md
atlas d2a550e685 feat(#3255): hives stop owning the knowledge webhook, and clean up their own
A webhook has exactly one target URL, so every hive registering one
against the shared internal/knowledge repository was last-writer-wins
rather than idempotent: all but the most recent silently stopped
receiving deliveries. The swarm controller holds the single registration
and now addresses an event to each hive over the queue instead.

This is a migration, not a deletion. Not registering any more fixes
nothing on a hive that has already run — the hook it created persists on
the forge, so the contention would survive on exactly the deployments
that have it while fresh installs looked fixed. The hive that created a
hook removes it.

It removes only its OWN, matched on the full URL rather than the
/webhook/knowledge suffix. A hook with that suffix and a different base
belongs to another hive, possibly one not yet upgraded, and deleting it
would break that hive's knowledge sync until it caught up. Reaping a
neighbour's registration is the behaviour being removed here; doing it
while fixing it would only invert the direction.

The predecessor did reap by suffix, to clear loopback hooks left by an
older single-hive layout. That was safe when a hive was alone on its
forge and is not safe now. The hive-side registrars also acted as reapers
of hooks under their own path, which is why the swarm hook lives under
/webhook/forge/; removing this registrar removes that reaper too.
Intended, and stated because no reviewer would infer it from the diff.

The receive endpoint goes with it. A live HMAC-verified
/webhook/knowledge that nothing can legitimately reach would tell the
next reader that this is how a hive learns about knowledge changes.

Docs move in the same commit: docs/swarm/README.md said two hooks exist
per swarm-wide repo and neither should be deleted, which is now true for
agent-configs and wrong for internal/knowledge — a half-correct
description being worse than an uncorrected one.
2026-08-19 21:05:52 +02:00

142 lines
5.6 KiB
Markdown

# Hive-wide knowledge repository
`internal/knowledge` on the forge is a shared reference doc repo
readable by every agent. hive-c0re clones it to the host and
bind-mounts the clone read-only into every agent container at
`/knowledge`.
## Agent access
Inside any agent container:
```
/knowledge/ # read-only bind-mount of the local clone
/knowledge/README.md # table of contents (seeded on first use)
```
Agents read documents directly from that path. The mount is
read-only — agents never write through it. To contribute, use the
`hive-forge` AGit flow (no fork needed — see
[Contributing](#contributing)); the operator reviews and merges, and
the local clone updates automatically (see
[Sync mechanism](#sync-mechanism) below).
## Repository layout
Canonical forge location: `internal/knowledge` (org `internal`,
repo `knowledge`). The repo is public, so every agent's forge account
has read access without an explicit per-agent collaborator grant;
only the `core` account has push access, for auto-seeding.
The repo is auto-created at hive-c0re startup if it doesn't exist,
seeded with a `README.md` containing a contribution guide and a
blank table of contents. Add an entry to that ToC each time you
create a new document.
## Sync mechanism
hive-c0re maintains the local clone at
`/var/lib/hyperhive/knowledge` via two paths:
1. **Swarm event** — the swarm controller holds the single push hook on
`internal/knowledge` (see `docs/swarm/README.md` § Swarm-wide forge
webhooks). On any push to main, including merge commits, it sends an
event to every hive over the swarm queue and each hive runs `git
pull`, so agents see the new content on their next turn.
A hive that is offline when the event is sent does not get it on
reconnect — the periodic pull below is what closes that gap. So one
hive briefly showing older `/knowledge` content than another is
expected, and resolves by itself within the fallback interval.
**Do not add a per-hive hook.** A webhook has exactly one target
URL, so a second registration against the same repo does not add a
recipient — it takes delivery away from whoever registered first.
Earlier versions had each hive register its own; hive-c0re now
removes its own leftover at startup, so no operator step is needed
to migrate.
2. **Periodic pull** — a background task in `hive-c0re::main`
pulls on a fixed cadence as a fallback (webhook missed, c0re
restarted between pushes). The pull is best-effort — a failure
logs a warning and does not affect the rest of the daemon.
Both paths share the same `knowledge::pull()` function, which also
handles the change notice below — neither path can forget to wire it
in since the broadcast logic lives once, in `pull()` itself, not at
each call site.
### Change notice
When a pull actually moves the local clone's `HEAD` (a real change,
not a no-op — e.g. the periodic pull finding nothing new), hive-c0re
broadcasts a short notice to every currently-registered agent's inbox:
sender `system`, body `[system] /knowledge updated:` followed by a
`git diff --stat <old>..<new>` summary of what changed (or a generic
"see the repo" fallback if computing the diff itself fails). This is
the same broadcast mechanism used for other hive-wide notices — inbox
message only, no forced wake, and it carries the standard "this was a
broadcast" hint. A diff or per-agent send failure is logged but never
blocks the pull itself.
The local clone is created (or refreshed) once at startup via
`knowledge::ensure_local_clone`. If the repo is brand new and
empty, `ensure_local_clone` seeds it with the default README
before returning.
## State
- **Host clone**: `/var/lib/hyperhive/knowledge` — persists across
hive-c0re restarts and agent destroy/recreate. Deleted only by
manual operator action.
- **In-container mount**: `/knowledge` — bind-mounted read-only
from the host clone on every container start. Gone when container
is stopped; reappears on next start with the current clone state.
The mount deliberately **excludes `.git`**: the host clone embeds the `core`
token in `.git/config` (it rides the clone URL), so hive-priv overlays an empty
tmpfs at `/knowledge/.git` — agents see the documents, not the repo metadata or
token.
## Contributing
Agents have read access to `internal/knowledge` (it's public) but no
write access, so they can't push a branch directly. The supported path
is Forgejo's **AGit flow** through the `hive-forge` CLI — no fork
required.
1. Clone the repo (credentials are injected automatically; the `-r`
flag selects the repo, the clone lands in `./knowledge`):
```sh
hive-forge -r internal/knowledge clone
cd knowledge
```
2. Create a branch and add or update a document, then commit normally.
3. Open (or update) a PR with `--agit`. This pushes the current `HEAD`
to `refs/for/<base>/<topic>`, which Forgejo turns into a PR even
though you can't push a branch:
```sh
hive-forge -r internal/knowledge pr-create --agit \
--title "docs: add the X runbook" \
--topic add-x-runbook \
--body-file - <<'EOF'
What this document adds and why.
EOF
```
Re-running with the same `--topic` updates the open PR (it
force-pushes the AGit scratch ref). The base defaults to `main`;
in `--agit` mode the push goes to the `origin` remote that
`hive-forge clone` set up.
4. The operator reviews and merges. The webhook fires on merge; every
running container sees the updated content within seconds (see
[Sync mechanism](#sync-mechanism)).
Do **not** try to `git push` a branch directly — lacking write access,
it's rejected. The `--agit` flow above is the no-fork path that works
from any agent.