Each agent card now leads with the agent's icon, loaded as an `<img>` from `GET /api/agents/<name>/icon`: the same 5em square, background and fallback as the hive dashboard's container row. An agent with no icon (the route's 404), or any other failed load, shows the dimmed hyperhive mark (`/favicon.svg`) instead of a broken image. Only ever an `<img>`, never inline markup: the body is an agent-authored SVG, and an image load does not run its script.
543 lines
26 KiB
Markdown
543 lines
26 KiB
Markdown
# Multi-hive swarms
|
|
|
|
A **swarm** is a collection of agents that share an identity and
|
|
coordinate across one or more hives. A single hyperhive instance
|
|
running on one host is already a swarm (one hive). This doc covers
|
|
the additional config needed when the swarm spans multiple hosts.
|
|
|
|
For the full option reference rather than prose: `services.hyperhive.swarm.*`
|
|
(swarm-wide facts, identical on every host) and `services.hyperhive.deploy.*`
|
|
(this host's own deployment decisions — does _this_ machine run grafana,
|
|
the swarm controller, authelia, …) are separate generated pages, `nix
|
|
build .#docs-swarm` / `.#docs-deploy` or the website's `/options/swarm.html`
|
|
/ `/options/deploy.html`.
|
|
|
|
## Terminology
|
|
|
|
- **hive** — a single hyperhive installation on one host. Has its
|
|
own `services.hyperhive.domain` DNS name and its own set of agent
|
|
containers.
|
|
- **swarm** — one or more hives whose operators have declared them
|
|
as peers. You can qualify an agent as `agent@hive-domain`.
|
|
- **peer hive** — any hive in `services.hyperhive.swarm.hives` other
|
|
than this one. Peers are _derived_, not declared: the directory lists
|
|
every hive including yourself, and `hiveName` says which one you are.
|
|
|
|
## Hive identity config
|
|
|
|
```nix
|
|
services.hyperhive = {
|
|
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
|
hiveName = "pr1ma"; # required — this hive's label in it
|
|
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
|
|
|
# required — the directory, identical on every host in the swarm.
|
|
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
|
swarm.hives = {
|
|
pr1ma = { };
|
|
edge = { };
|
|
};
|
|
};
|
|
```
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
`swarm.domain` and `hiveName` are **required** whenever hyperhive is
|
|
enabled; eval fails with a hint naming each. Neither defaults,
|
|
because a guessed value here is a wrong hostname that evaluates cleanly
|
|
and deploys — an eval failure asking the operator to write the address
|
|
down is the cheaper outcome. **Upgrading past this release means setting
|
|
both once.**
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
You must still set `domain` too, but you no longer _write_ it: it's read from
|
|
this hive's own entry in the directory, whose `domain` defaults to
|
|
`<name>.<swarm.domain>`. A conventional swarm states no addresses at
|
|
all, and a hive addressed by something else states it in the one place
|
|
the other hives read — `swarm.hives.edge.domain = "edge.elsewhere.example";`.
|
|
|
|
Setting `services.hyperhive.domain` directly still works and still wins,
|
|
with a **deprecation warning**. The reason it's deprecated isn't tidiness:
|
|
that option is local to one host, and the operator copies the directory
|
|
to every host, so a value written only there leaves every peer pointing
|
|
somewhere else with nothing detecting the disagreement.
|
|
|
|
⚠️ **Upgrading:** a hive that has been running on `swarm.domain` +
|
|
`hiveName` alone now needs its own directory entry —
|
|
`services.hyperhive.swarm.hives.<hiveName> = { };`, one line, no value.
|
|
Eval fails naming it if you forget.
|
|
|
|
`domain` drives `HYPERHIVE_HIVE_DOMAIN` in every container so agents can
|
|
form qualified labels (`iris@pr1ma.example.com`).
|
|
|
|
`swarm.name` is purely display — it surfaces in the dashboard chrome
|
|
header and per-agent system prompts, and federated hives at different
|
|
domains can share one. `hiveName` surfaces in the same places but is
|
|
_not_ only display: it's the leftmost label of the hive's domain. That
|
|
`swarm.name` sits under `swarm` and `hiveName` doesn't is the whole
|
|
distinction — one names this hive, the other names the group it belongs
|
|
to.
|
|
|
|
See `docs/process/conventions.md` § Hive identity for the env-var chain
|
|
and `qualify()` / `qualified_label()` semantics.
|
|
|
|
## Swarm CA
|
|
|
|
A hive's internal TLS chains to a **swarm root CA**, so a peer that
|
|
trusts the root validates every hive in the swarm rather than pinning
|
|
to each one by hand. Provisioning modes, what to hand a peer
|
|
(`trust-bundle.pem`, never `ca.pem`), the name constraints on a hive
|
|
CA, and how an existing hive adopts the hierarchy: [`ca.md`](ca.md).
|
|
|
|
## Running the swarm's shared services
|
|
|
|
One authelia, one matrix, one forge per swarm — which host runs them,
|
|
and what a hive that runs none of them configures instead:
|
|
[`services.md`](services.md).
|
|
|
|
## Single sign-on
|
|
|
|
Which secrets the SSO provider generates, which one has a reader in
|
|
another container, and the three ways that one gets delivered:
|
|
[`sso.md`](sso.md).
|
|
|
|
## Secrets
|
|
|
|
Every credential the swarm holds, who mints it, where it must live, and
|
|
which of the three topologies makes it the operator's job to place:
|
|
[`secrets.md`](secrets.md).
|
|
|
|
Where that shape is **going** — the per-secret minter/reader/renewal
|
|
contract, the target of one mTLS identity per host and everything else
|
|
through the store, and the test a change has to pass to count as movement
|
|
toward it: [`credentials.md`](credentials.md). It supersedes `secrets.md`
|
|
when the migration completes.
|
|
|
|
## Swarm UI
|
|
|
|
The operator-only web surface on the swarm apex, why reaching it needs
|
|
the `admins` group rather than just a session, and the four sites you
|
|
wire a swarm service name into: [`ui.md`](ui.md).
|
|
|
|
## The swarm's hive directory
|
|
|
|
```nix
|
|
services.hyperhive.swarm.hives = {
|
|
pr1ma = { }; # this host, per hiveName
|
|
lab = { }; # a second hive in the swarm
|
|
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
|
};
|
|
```
|
|
|
|
One attrset describing **every** hive in the swarm, **including this
|
|
one**, keyed by that hive's `hiveName`. It's meant to be _identical on
|
|
every host_ — write it once, share it, and each host reads it correctly
|
|
because `services.hyperhive.hiveName` says which entry is itself.
|
|
|
|
Empty (the default) means this host isn't in a swarm. Once non-empty it
|
|
**must** contain an entry for `hiveName`; eval fails naming the missing
|
|
hive. That assertion is load-bearing rather than pedantic — "my peers"
|
|
comes from _everything that isn't me_, so a directory that doesn't
|
|
contain you derives every hive as a peer and you peer with yourself.
|
|
|
|
`domain` defaults to `<name>.<swarm.domain>`, the convention every hive
|
|
follows, so a conventional directory is names only. The default is a
|
|
derivation from two values the operator already had to state — the swarm's
|
|
domain and the entry's own name — rather than a guess, which is what makes
|
|
it safe here when a guessed hostname wouldn't be. Set it only for a hive
|
|
addressed by something else.
|
|
|
|
> **No per-hive CA field exists, and no per-hive cert pinning.** Trust
|
|
> inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every
|
|
> hive chains to it, so one anchor replaces per-hive pinning entirely.
|
|
> What that genuinely drops is trusting a hive whose root this swarm
|
|
> does _not_ own — another swarm's, or one keeping its own CA. That's
|
|
> a cross-swarm problem and wants a mechanism designed for it. (An
|
|
> earlier `certFingerprint` field existed for exactly that gap, pinning
|
|
> a peer's TLS leaf for hive-c0re's own peer HTTPS checks — removed
|
|
> along with the dashboard feature it existed to serve, since nothing
|
|
> else ever consumed it.)
|
|
|
|
## What the config does at runtime
|
|
|
|
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
|
directory and serves it at `GET /api/hives`; `swarm-ui`'s overview
|
|
page renders it (`docs/swarm/ui.md`). This is the operator-facing
|
|
"what hives exist" surface — a per-hive dashboard "peer hives"
|
|
display existed here once; it no longer exists, in favour of this.
|
|
|
|
2. **Matrix federation** — when `matrix.enable` is on, tuwunel
|
|
federates with the peer's matrix server (discovered via the peer's
|
|
`.well-known/matrix/server` delegation, which the gateway serves).
|
|
Federation validates the peer's TLS certificate against the matrix
|
|
**container's** trust bundle, independent of this directory.
|
|
|
|
⚠️ **That container currently trusts no swarm-internal CA**, so a
|
|
self-signed gateway certificate doesn't federate. You can't list the
|
|
swarm root there: `security.pki.certificateFiles` is
|
|
read when the system is _built_, and the root is a runtime file (its
|
|
key must never enter the store), so there is no build-time name for
|
|
it. Bridging that needs a runtime mechanism; a separate issue tracks
|
|
it. Until then, federation needs CA-issued certs (ACME). See
|
|
`docs/integrations/matrix.md` for federation firewall + TLS requirements.
|
|
|
|
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
|
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
|
to configure `wg-hive`. See "WireGuard inter-hive mesh" below.
|
|
|
|
## One directory, not a bilateral declaration
|
|
|
|
Both hives hold the **same** `hives` attrset; neither declares the
|
|
other. What differs between the two hosts is only `hiveName`:
|
|
|
|
```
|
|
# hive A # hive B
|
|
hiveName = "pr1ma"; hiveName = "edge";
|
|
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
|
```
|
|
|
|
That's the point of the shape, and it removes a class of bug rather
|
|
than saving typing: a per-host peer list let two hosts hold _different_
|
|
facts about the same third hive — a stale endpoint, a rotated
|
|
fingerprint — with nothing to detect the disagreement. One entry per
|
|
hive makes it unrepresentable.
|
|
|
|
## WireGuard inter-hive mesh (optional)
|
|
|
|
The peer config above uses public HTTPS for all inter-hive traffic.
|
|
For private deployments — or to reduce latency and TLS overhead on
|
|
intra-swarm traffic — hive-c0re can configure a host-to-host
|
|
WireGuard mesh.
|
|
|
|
### Generating keys
|
|
|
|
On each hive host:
|
|
|
|
```bash
|
|
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
|
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
|
```
|
|
|
|
### Config example (two hives)
|
|
|
|
```nix
|
|
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
|
services.hyperhive = {
|
|
deploy.wireguard = {
|
|
enable = true;
|
|
privateKeyFile = "/etc/wireguard/hive.key";
|
|
address = "10.100.0.1/24";
|
|
listenPort = 51820; # optional, default 51820
|
|
};
|
|
|
|
# The same `hives` attrset both hosts hold — mesh fields included,
|
|
# since "where this hive can be dialled" is a fact about that hive.
|
|
swarm.hives = {
|
|
pr1ma = {
|
|
domain = "pr1ma.example.com";
|
|
wireguardPublicKey = "base64keyA=";
|
|
wireguardEndpoint = "198.51.100.1:51820";
|
|
wireguardAddress = "10.100.0.1/32";
|
|
};
|
|
edge = {
|
|
domain = "edge.corp";
|
|
wireguardPublicKey = "base64keyB=";
|
|
wireguardEndpoint = "203.0.113.42:51820";
|
|
wireguardAddress = "10.100.0.2/32";
|
|
};
|
|
};
|
|
};
|
|
|
|
# hive B (edge.corp, mesh IP 10.100.0.2)
|
|
services.hyperhive = {
|
|
deploy.wireguard = {
|
|
enable = true;
|
|
privateKeyFile = "/etc/wireguard/hive.key";
|
|
address = "10.100.0.2/24";
|
|
};
|
|
|
|
swarm.hives = { /* … identical to hive A's … */ };
|
|
};
|
|
```
|
|
|
|
### What the mesh does
|
|
|
|
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
|
host (not inside agent containers; containers reach peers via the
|
|
host's routing table).
|
|
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
|
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
|
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
|
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
|
rather than the public domain.
|
|
- It sets `persistentKeepalive = 25` by default; override or null to
|
|
disable (not needed when both sides have public IPs and no NAT).
|
|
|
|
### NAT / one-sided endpoints
|
|
|
|
If one host is behind NAT and can't accept incoming connections, only
|
|
that host needs a null `wireguardEndpoint` on the peer config — the
|
|
other side initiates. With keepalive on, the NAT hole stays open.
|
|
|
|
If both hosts are behind NAT, you need a STUN relay or a third host
|
|
(exit node). Out of scope for v0.
|
|
|
|
## Snapshot store
|
|
|
|
One further option lives in this namespace but its docs live with the
|
|
service it points at: `services.hyperhive.swarm.snapshotStore.{address,
|
|
port}` tells this hive where the swarm's `btrfs receive` endpoint is, so
|
|
`hivectl agent <name> subvol snapshot push` has somewhere to stream to.
|
|
|
|
It's genuinely swarm-scoped rather than per-peer — a swarm has exactly
|
|
one store, because the receiver keys destinations by _agent_ so a
|
|
migrating agent keeps one unbroken incremental chain. See
|
|
[snapshot-store.md](../networking/snapshot-store.md).
|
|
|
|
## Swarm controller
|
|
|
|
`services.hyperhive.deploy.swarm-controller.enable` runs the `swarm-controller`
|
|
daemon on this host. **Off by default and deliberately not derived from
|
|
`services.hyperhive.deploy.hive-controller.enable`**: a swarm has one
|
|
controller, so enabling it states a fact about swarm topology, not about
|
|
whether this host runs a hive. Every hive runs `hive-c0re` (the agents on that host); one
|
|
hive additionally runs this (what's true across hives).
|
|
|
|
What it serves, why it's a unix socket rather than a port, and the
|
|
socket-directory constraint that governs where `socketPath` may point:
|
|
[`swarm-controller/README.md`](../../swarm-controller/README.md).
|
|
|
|
### Per-hive status (`GET /api/hives/status`)
|
|
|
|
One row per hive in `swarm.hives`, saying when it last reported and what
|
|
it said. Hives publish upward through the swarm queue; the controller
|
|
never reaches down to collect, so a hive that can't reach the swarm
|
|
still knows its own state — you just can't see it from here.
|
|
|
|
A hive publishes only once it holds all three status-publish coordinates
|
|
below. A hive without them reads
|
|
`never_reported` — it's not broken, it just has nothing to say upward.
|
|
|
|
| freshness | what to do about it |
|
|
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
| `fresh` | nothing — reported within `staleAfterSeconds` |
|
|
| `stale` | the hive stopped reporting. Its last payload is still shown, so check `age_seconds` and the payload for what it managed to say |
|
|
| `never_reported` | this hive has never reported at all — normally a deployment that hasn't happened, not an outage |
|
|
| `unknown` | something is publishing under a name that's not in `swarm.hives` — a typo in the roster, or a hive removed from the roster but still running |
|
|
|
|
Every row also carries `last_seen_unix` and `age_seconds` if you want to
|
|
apply your own threshold. The timestamp is the one the queue recorded on
|
|
arrival, not one the hive put in its own payload.
|
|
|
|
Set `services.hyperhive.swarm.controller.staleAfterSeconds` (default
|
|
`120`) **above the rate hives publish at**, or everything reads `stale`
|
|
between reports. Hives publish once a minute, so the default tolerates
|
|
one missed report and flags two. It takes effect on the next request;
|
|
nothing has to re-publish.
|
|
|
|
### Making a hive report
|
|
|
|
Three options on the **hive**. They sit in two namespaces, because two of them
|
|
are facts about _this machine_ and one is the swarm's single address:
|
|
|
|
| option | what to set it to |
|
|
| ------------------------------------------------------- | --------------------------------------------- |
|
|
| `deploy.hive-controller.statusPublish.natsUrl` | `tls://<swarm.nats.domain>:<swarm.nats.port>` |
|
|
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` |
|
|
| `deploy.hive-controller.statusPublish.clientSecretFile` | path to this hive's client secret, plaintext |
|
|
|
|
The queue URL and the token endpoint default to the swarm's own addresses on
|
|
every hive, so there is nothing to set for them. The secret is what turns
|
|
publishing on: a hive without it doesn't publish. A secret without the other
|
|
two is an eval error rather than a hive that quietly never reports.
|
|
|
|
On a host that runs the queue and the IdP, the secret defaults to the local one.
|
|
Any other hive needs the secret to physically be there, because the swarm
|
|
doesn't distribute it. Copy `hive-<hiveName>.secret` out of the swarm host's
|
|
`deploy.authelia.hostClientSecretDir` with whatever secret management the
|
|
deployment already uses.
|
|
|
|
The queue URL is the same string on every hive. The queue accepts TLS only,
|
|
with a certificate for `swarm.nats.domain` (default `nats.<swarm.domain>`) and
|
|
no other name or address, so a URL with an IP address or `nats://` fails.
|
|
The queue's host resolves the name itself. **A multi-host swarm needs one
|
|
upstream DNS record**, `nats.<swarm.domain>` pointing at the queue host's mesh
|
|
address, the same contract as `bao.<swarm.domain>`. The queue's port is open
|
|
on `wg-hive` when that host is on the mesh.
|
|
|
|
The identity isn't a choice — a hive authenticates as `hive-<hiveName>`
|
|
and publishes under `hiveName`, the same name that keys `swarm.hives`.
|
|
|
|
If a hive stops reporting, its own dashboard is the place to look: a
|
|
failure to publish raises a warning banner there after three consecutive
|
|
misses. It stays `warn` rather than `crit` on purpose — a hive that
|
|
can't reach the queue isn't itself unhealthy, so it doesn't start
|
|
calling itself degraded for being unable to say it's fine.
|
|
|
|
The endpoint answers **503** when this host has no swarm queue
|
|
configured, or has one and can't read it — deliberately not an empty
|
|
list, which would look like a silent swarm rather than a controller that
|
|
can't see. The body says which. Status survives a controller restart:
|
|
it's stored in the queue, not in the daemon.
|
|
|
|
### Giving the agents the queue too
|
|
|
|
An agent authenticates as its own client, not as its hive, so it needs its
|
|
own coordinates. Two of them are options on the hive; the other two arrive
|
|
with the credential itself and aren't configurable.
|
|
|
|
| option | what to set it to |
|
|
| ------------------------------------------- | -------------------------------------------------------------- |
|
|
| `deploy.hive-controller.queue.agentNatsUrl` | where the queue listens, as an agent **container** reaches it |
|
|
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` — the same one the hive uses |
|
|
|
|
On every hive, `agentNatsUrl` defaults to the same
|
|
`tls://<swarm.nats.domain>:<swarm.nats.port>` as the hive's own. On the queue's
|
|
host the name resolves inside a container to the bridge address, where the
|
|
firewall opens the port; on any other hive it resolves through the host's DNS,
|
|
like the store's name. ⚠️ **Never a loopback address here**: an agent has its
|
|
own network namespace, so `127.0.0.1` reaches the agent.
|
|
|
|
The harness sees four variables, and treats them as all-or-none:
|
|
`HIVE_AGENT_NATS_URL` and `HIVE_AGENT_OIDC_TOKEN_ENDPOINT` from the two options
|
|
above, plus `HIVE_AGENT_OIDC_CLIENT_SECRET_FILE` and
|
|
`HIVE_AGENT_OIDC_CLIENT_ID_FILE`, which point into the unit's own credentials
|
|
directory. The last two come from the delivered credential rather than from
|
|
config — see [`secrets.md`](secrets.md#hive-level--one-of-each-per-hive) for
|
|
how it gets there. A hive lacking the queue's address for its
|
|
agents sets none of the four and each agent logs that it has none; a half-set
|
|
environment logs an error and the harness keeps serving.
|
|
|
|
What an agent does with that connection is publish its terminal. Every row its
|
|
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
|
|
a swarm-level terminal can follow one agent without subscribing to the swarm's
|
|
whole traffic. An agent connected with its own queue credential gets that
|
|
subject. An agent without one, or whose own credential the queue
|
|
refused, connects with its hive's shared client and publishes to
|
|
`$SWARM.term.<hive>.<agent>` instead, the `<hive>` being the one that client id
|
|
names. The swarm controller relays both. Publishing only: an
|
|
agent talks about itself here and reads nothing. Rows aren't retained — a
|
|
subscriber that wasn't listening missed them, the same as on the agent's own
|
|
live stream.
|
|
|
|
The queue would refuse a row too large for its `max_payload` outright and
|
|
take the connection down with it, so the harness drops such a row's body before
|
|
sending and leaves a marker in its place; the summary, level and icon still
|
|
arrive. The harness logs and skips a row that's too large even without its body.
|
|
|
|
The second thing an agent publishes is its **turn-state header**, on
|
|
`$SWARM.agent-state.<agent>` (or `$SWARM.agent-state.<hive>.<agent>`) — same shape of subject, same grant
|
|
mechanics, same lack of retention. It carries what a header bar wants: what the
|
|
turn loop is doing (`turn_state`, plus `turn_state_since` as an ISO 8601 UTC
|
|
stamp), which model (`model` and the resolved id the last turn actually ran on),
|
|
the context budget and the last turn's context and cost token blocks, and
|
|
`agent_state`.
|
|
|
|
`agent_state` reuses the swarm's own wanted-state vocabulary
|
|
(`up`/`offline`/`paused`/`destroyed`) so a reader can compare what an agent _is_
|
|
against what the swarm declared it should be without translating between two
|
|
spellings. ⚠️ From inside the container only two of those four are sayable: the
|
|
harness reports `up`, or `paused` when the pause marker is present. `offline` and
|
|
`destroyed` are hive-c0re's observations — a stopped agent publishes nothing and
|
|
a destroyed one doesn't exist — so a view that needs the full four-state picture
|
|
takes them from the `agent-status` bucket and uses this subject to sharpen the
|
|
rest.
|
|
|
|
Headers go out **on transition, not on a timer**: the harness rebuilds the
|
|
header whenever its event bus moves and publishes only when the result differs
|
|
from what it last sent. That's the whole point of the subject — the
|
|
`agent-status` bucket already republishes once a minute, which is far too slow
|
|
for "is this agent thinking right now." The cost of a core subject is that a
|
|
subscriber attaching mid-idle sees nothing until the next change, so a renderer
|
|
opens with the bucket's snapshot and lets this stream refine it.
|
|
|
|
Swarm-side, `GET /api/agents/<name>/state/stream` relays that subject as SSE,
|
|
resolving the agent's hive at request time exactly as the terminal stream does.
|
|
The payload passes through opaquely — the controller never parses a header.
|
|
|
|
The third thing an agent publishes is its **icon**, the same SVG its own
|
|
`GET /icon` serves. It goes into the `agent-icons` KV bucket under the key
|
|
`<agent>`, with no hive in it, so the swarm can show the icon of an agent that's
|
|
stopped or has moved hives. Only an agent connected with its own queue
|
|
credential publishes it: the queue grants that credential
|
|
`$KV.agent-icons.<agent>` and no other key, and grants the hive's shared client
|
|
none of the bucket. The harness writes once per start, because a config change
|
|
reaches an agent by restarting its container. An agent with no icon deletes its
|
|
key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
|
|
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
|
|
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
|
|
|
|
### Swarm-wide forge objects
|
|
|
|
The controller also keeps the forge objects that are one per swarm, not
|
|
one per hive. It ensures them at start and every five minutes after
|
|
(`swarm-controller/src/forge/objects.rs`):
|
|
|
|
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
|
|
owner org;
|
|
- the empty `operators` merge-gate team in `agents` and `agent-configs`;
|
|
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
|
|
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
|
|
- `internal/docs` (private) and `internal/knowledge` (public, with a
|
|
seed `README.md` while empty);
|
|
- the `agent-configs` org avatar
|
|
(`deploy.swarm-controller.configOrgAvatarPng`).
|
|
|
|
hive-c0re no longer creates any of them. A pass that can't finish logs a
|
|
`warn` line per object plus `swarm forge objects: pass incomplete` in
|
|
`journalctl -u swarm-controller`, and retries on the next tick. While the
|
|
controller is down the objects stay as they are.
|
|
|
|
### Swarm-wide forge webhooks
|
|
|
|
At startup the controller registers two Forgejo hooks pointing at
|
|
itself — a `push` hook on `internal/knowledge` and a `pull_request` hook
|
|
on the `agent-configs` org, both under
|
|
`https://<swarm.domain>/webhook/forge/`.
|
|
|
|
The controller **interprets** a delivery and sends hives a specific
|
|
message — _the knowledge repo changed_, _deploy agent `foo` at rev
|
|
`abc123`_ — rather than forwarding forge payloads for each hive to
|
|
re-derive. Approval happens once, at the swarm level: a hive receives a
|
|
decision, not an event to adjudicate.
|
|
|
|
**`internal/knowledge` is on that path.** The controller's is the only
|
|
hook on it: hives no longer register their own (see
|
|
`docs/integrations/knowledge.md` for clearing a leftover). A webhook has exactly one target URL, so per-hive
|
|
registration never added a recipient — it took delivery away from
|
|
whichever hive registered before it.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
**The `agent-configs` org isn't yet.** Each hive still registers its own
|
|
`pull_request` hook there, so that repo has two — the hive's and the
|
|
controller's — and **both are expected; don't delete either.** Removing
|
|
a hive's stops it acting on config PRs; removing the controller's just
|
|
gets recreated on its next start.
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
Nothing to configure. The controller registers the hooks only when this
|
|
host also serves the swarm UI vhost — that's what publishes the
|
|
endpoint, and a hook the forge can't reach would collect failed
|
|
deliveries while looking healthy. The controller generates the HMAC
|
|
secret on first start and keeps it
|
|
(see [`docs/agent-lifecycle/persistence.md`](../agent-lifecycle/persistence.md)).
|
|
|
|
To check it's working, push to `internal/knowledge` and look for
|
|
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
|
|
refused delivery logs `webhook: refused delivery` with the reason.
|
|
|
|
## Cross-references
|
|
|
|
- `docs/networking/snapshot-store.md` — the swarm's `btrfs receive` endpoint, and
|
|
the `swarm.snapshotStore` option that points a hive at it
|
|
- `docs/process/conventions.md` § Hive identity — env vars, qualified labels
|
|
- `docs/integrations/matrix.md` — matrix federation, TLS cert autogeneration,
|
|
firewall posture
|
|
- `docs/swarm/ui.md` — the swarm-wide hive roster, now the operator
|
|
surface for "what hives exist" (superseded the per-hive dashboard's
|
|
old "peer hives" display)
|
|
- `docs/networking/gateway.md` — nginx vhosts and the `.well-known/matrix/`
|
|
autodiscovery scheme
|