hyperhive/docs/integrations/knowledge.md
atlas 20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00

150 lines
5.9 KiB
Markdown

# Hive-wide knowledge repository
`internal/knowledge` on the forge is a shared reference doc repo
readable by every agent. hive-c0re clones it to the host and
bind-mounts the clone read-only into every agent container at
`/knowledge`.
## Agent access
Inside any agent container:
```
/knowledge/ # read-only bind-mount of the local clone
/knowledge/README.md # table of contents (seeded on first use)
```
<!-- vale write-good.Passive = NO -->
Agents read documents directly from that path. The mount is
read-only — agents never write through it. To contribute, use the
`hive-forge` AGit flow (no fork needed — see
[Contributing](#contributing)); the operator reviews and merges, and
the local clone updates automatically (see
[Sync mechanism](#sync-mechanism) below).
<!-- vale write-good.Passive = YES -->
## Repository layout
Canonical forge location: `internal/knowledge` (org `internal`,
repo `knowledge`). The repo is public, so every agent's forge account
has read access without an explicit per-agent collaborator grant;
only the `core` account has push access.
The swarm controller creates the repo if it doesn't exist, at startup
and every few minutes after, and seeds an empty one with a `README.md`
containing a contribution guide and a blank table of contents. Add an entry to that ToC each time you
create a new document.
## Sync mechanism
hive-c0re maintains the local clone at
`/var/lib/hyperhive/knowledge` via two paths:
1. **Swarm event** — the swarm controller holds the single push hook on
`internal/knowledge` (see `docs/swarm/README.md` § Swarm-wide forge
webhooks). On any push to main, including merge commits, it sends an
event to every hive over the swarm queue and each hive runs `git
pull`, so agents see the new content on their next turn.
A hive that's offline when the swarm controller sends the event
doesn't get it on reconnect — the periodic pull below is what closes that gap. One
hive briefly showing older `/knowledge` content than another is
expected, and resolves by itself within the fallback interval.
**don't add a per-hive hook.** A webhook has exactly one target
URL, so a second registration against the same repo doesn't add a
recipient — it takes delivery away from whoever registered first.
Earlier versions had each hive register its own, and later ones
removed that leftover at startup. A hive upgraded straight past those
versions still holds its old hook: delete any hook on the repo that
points at a hive's own `/webhook/knowledge`.
2. **Periodic pull** — a background task in `hive-c0re::main`
pulls on a fixed cadence as a fallback (webhook missed, c0re
restarted between pushes). The pull is best-effort — a failure
logs a warning and doesn't affect the rest of the daemon.
Both paths share the same `knowledge::pull()` function, which also
handles the change notice below — neither path can forget to wire it
in since the broadcast logic lives once, in `pull()` itself, not at
each call site.
### Change notice
When a pull actually moves the local clone's `HEAD` (a real change,
not a no-op — for example the periodic pull finding nothing new), hive-c0re
broadcasts a short notice to every currently registered agent's inbox:
sender `system`, body `[system] /knowledge updated:` followed by a
`git diff --stat <old>..<new>` summary of what changed (or a generic
"see the repo" fallback if computing the diff itself fails). This is
the same broadcast mechanism used for other hive-wide notices — inbox
message only, no forced wake, and it carries the standard "this was a
broadcast" hint. hive-c0re logs a diff or per-agent send failure
without blocking the pull itself.
`knowledge::ensure_local_clone` creates (or refreshes) the local
clone once at startup. If the repo is brand new and
empty, `ensure_local_clone` seeds it with the default README
before returning.
## State
<!-- vale write-good.Passive = NO -->
- **Host clone**: `/var/lib/hyperhive/knowledge` — persists across
hive-c0re restarts and agent destroy/recreate. Deleted only by
manual operator action.
- **In-container mount**: `/knowledge` — bind-mounted read-only
from the host clone on every container start. Gone when container
is stopped; reappears on next start with the current clone state.
<!-- vale write-good.Passive = YES -->
The mount deliberately **excludes `.git`**: the host clone embeds the `core`
token in `.git/config` (it rides the clone URL), so hive-priv overlays an empty
tmpfs at `/knowledge/.git` — agents see the documents, not the repo metadata or
token.
## Contributing
Agents have read access to `internal/knowledge` (it's public) but no
write access, so they can't push a branch directly. The supported path
is Forgejo's **AGit flow** through the `hive-forge` CLI — no fork
required.
1. Clone the repo (`hive-forge` injects credentials automatically;
the `-r` flag selects the repo, the clone lands in `./knowledge`):
```sh
hive-forge -r internal/knowledge clone
cd knowledge
```
2. Create a branch and add or update a document, then commit normally.
3. Open (or update) a PR with `--agit`. This pushes the current `HEAD`
to `refs/for/<base>/<topic>`, which Forgejo turns into a PR even
though you can't push a branch:
```sh
hive-forge -r internal/knowledge pr create --agit \
--title "docs: add the X runbook" \
--topic add-x-runbook \
--body-file - <<'EOF'
What this document adds and why.
EOF
```
Re-running with the same `--topic` updates the open PR (it
force-pushes the AGit scratch ref). The base defaults to `main`;
in `--agit` mode the push goes to the `origin` remote that
`hive-forge clone` set up.
4. The operator reviews and merges. The webhook fires on merge; every
running container sees the updated content within seconds (see
[Sync mechanism](#sync-mechanism)).
Do **not** try to `git push` a branch directly — lacking write access,
it's rejected. The `--agit` flow above is the no-fork path that works
from any agent.