Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/README.md
atlas fb2fff0668 fix(nix): require swarm domain only when a hyperhive service is enabled
The swarm.domain assertion in hive-network.nix fired on every host that
imported the module, so a host that enables nothing failed eval. It now
fires only when one of the hyperhive service switches is on (every
deploy.*.enable that runs something, gateway, gateway.dns, network,
otel, snapshotStore). The requirement itself is unchanged: any host that
runs a hyperhive service still needs swarm.domain.

The core-toggle module-eval suite gains a case: a missing swarm.domain is
refused on a hive and on a swarm-service-only host, and a host enabling
nothing passes every assertion.

Closes #4887
2026-10-02 13:09:45 +02:00

471 lines
24 KiB
Markdown

# The swarm
The **swarm** is where things live: agent identities and accounts, secrets,
the job graph that creates and places agents, telemetry, and the UI you
drive it all from. **Hives are the substrate** — NixOS hosts that run agent
containers on the swarm's behalf. Every hive belongs to a swarm; a single
host is a swarm of one.
This page is for the operator. It covers the control plane, how a hive
joins the directory, and how hives and agents report upward. The steps for a
fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local
config is the [README quick start](../../README.md#quick-start-an-all-local-swarm).
Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical
on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs
grafana, the controller, authelia, …) →
[options reference](https://hyperhive.darkest.space/options/), or
`nix build .#docs-swarm` / `.#docs-deploy`.
## Where each piece lives
- **control plane** — `swarm-controller`: hive directory, agent roster, job
graph, agent creation. → [below](#swarm-controller)
- **swarm UI** — the operator's day-to-day surface, on the swarm apex,
`admins` only. → [`ui.md`](ui.md)
- **shared services** — one forge, homeserver, SSO, queue and metrics/logs
stack, each on whichever host you put it. → [`services.md`](services.md)
- **SSO** — the secrets authelia generates and how each reaches its reader.
→ [`sso.md`](sso.md)
- **secrets** — every credential the swarm holds, who mints it and where it
lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal
contract → [`credentials.md`](credentials.md) · how the store comes up,
who writes its grants and how it unseals → [`bao.md`](bao.md)
- **swarm CA** — the root every hive's internal TLS chains to, and what to
hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md)
- **snapshot store** — the swarm's one `btrfs receive` endpoint,
`swarm.snapshotStore.{address,port}`. →
[`snapshot-store.md`](../networking/snapshot-store.md)
## Swarm controller
`swarm-controller` is the swarm's control plane: it holds the hive directory,
the agent roster and their wanted state, and the job graph. Creating an agent —
from the swarm UI or `swarmctl agent create --hive <h>` — queues its SSO
identity, forge user, config repo, store identity and matrix account, then
sends the hive a deploy message. A new agent starts `paused`.
`services.hyperhive.deploy.swarm-controller.enable` runs it on this host;
`singleHostSwarm` turns it on. **Otherwise off: set it on the one host that
runs the controller.** A swarm has one controller, so enabling it states a
fact about swarm topology, not about whether this host runs a hive.
What it serves, why it's a unix socket rather than a port, and the
socket-directory constraint that governs where `socketPath` may point:
[`swarm-controller/README.md`](../../swarm-controller/README.md).
### Per-hive status (`GET /api/hives/status`)
One row per hive in `swarm.hives`, saying when it last reported and what
it said. Hives publish upward through the swarm queue; the controller
never reaches down to collect, so a hive that can't reach the swarm
still knows its own state — you just can't see it from here.
A hive publishes only once it holds all three status-publish coordinates
below. A hive without them reads
`never_reported` — it's not broken, it just has nothing to say upward.
| freshness | what to do about it |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `fresh` | nothing — reported within `staleAfterSeconds` |
| `stale` | the hive stopped reporting. Its last payload is still shown, so check `age_seconds` and the payload for what it managed to say |
| `never_reported` | this hive has never reported at all — normally a deployment that hasn't happened, not an outage |
| `unknown` | something is publishing under a name that's not in `swarm.hives` — a typo in the roster, or a hive removed from the roster but still running |
Every row also carries `last_seen_unix` and `age_seconds` if you want to
apply your own threshold. The timestamp is the one the queue recorded on
arrival, not one the hive put in its own payload.
Set `services.hyperhive.swarm.controller.staleAfterSeconds` (default
`120`) **above the rate hives publish at**, or everything reads `stale`
between reports. Hives publish once a minute, so the default tolerates
one missed report and flags two. It takes effect on the next request;
nothing has to re-publish.
### Making a hive report
Three options on the **hive**. They sit in two namespaces, because two of them
are facts about _this machine_ and one is the swarm's single address:
| option | what to set it to |
| ------------------------------------------------------- | --------------------------------------------- |
| `deploy.hive-controller.statusPublish.natsUrl` | `tls://<swarm.nats.domain>:<swarm.nats.port>` |
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` |
| `deploy.hive-controller.statusPublish.clientSecretFile` | path to this hive's client secret, plaintext |
The queue URL and the token endpoint default to the swarm's own addresses on
every hive, so there is nothing to set for them. The secret is what turns
publishing on: a hive without it doesn't publish. A secret without the other
two is an eval error rather than a hive that quietly never reports.
On a host that runs the queue and the IdP, the secret defaults to the local one.
Any other hive needs the secret to physically be there, because the swarm
doesn't distribute it. Copy `hive-<hiveName>.secret` out of the swarm host's
`deploy.authelia.hostClientSecretDir` with whatever secret management the
deployment already uses.
The queue URL is the same string on every hive. The queue accepts TLS only,
with a certificate for `swarm.nats.domain` (default `nats.<swarm.domain>`) and
no other name or address, so a URL with an IP address or `nats://` fails.
The queue's host resolves the name itself. **A multi-host swarm needs one
upstream DNS record**, `nats.<swarm.domain>` pointing at the queue host's mesh
address, the same contract as `bao.<swarm.domain>`. The queue's port is open
on `wg-hive` when that host is on the mesh.
The identity isn't a choice — a hive authenticates as `hive-<hiveName>`
and publishes under `hiveName`, the same name that keys `swarm.hives`.
If a hive stops reporting, its own dashboard is the place to look: a
failure to publish raises a warning banner there after three consecutive
misses. It stays `warn` rather than `crit` on purpose — a hive that
can't reach the queue isn't itself unhealthy, so it doesn't start
calling itself degraded for being unable to say it's fine.
The endpoint answers **503** when this host has no swarm queue
configured, or has one and can't read it — deliberately not an empty
list, which would look like a silent swarm rather than a controller that
can't see. The body says which. Status survives a controller restart:
it's stored in the queue, not in the daemon.
### Giving the agents the queue too
An agent authenticates as its own client, not as its hive, so it needs its
own coordinates. Two of them are options on the hive; the other two arrive
with the credential itself and aren't configurable.
| option | what to set it to |
| ------------------------------------------- | -------------------------------------------------------------- |
| `deploy.hive-controller.queue.agentNatsUrl` | where the queue listens, as an agent **container** reaches it |
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` — the same one the hive uses |
On every hive, `agentNatsUrl` defaults to the same
`tls://<swarm.nats.domain>:<swarm.nats.port>` as the hive's own. On the queue's
host the name resolves inside a container to the bridge address, where the
firewall opens the port; on any other hive it resolves through the host's DNS,
like the store's name. ⚠️ **Never a loopback address here**: an agent has its
own network namespace, so `127.0.0.1` reaches the agent.
The harness sees four variables, and treats them as all-or-none:
`HIVE_AGENT_NATS_URL` and `HIVE_AGENT_OIDC_TOKEN_ENDPOINT` from the two options
above, plus `HIVE_AGENT_OIDC_CLIENT_SECRET_FILE` and
`HIVE_AGENT_OIDC_CLIENT_ID_FILE`, which point into the unit's own credentials
directory. The last two come from the delivered credential rather than from
config — see [`secrets.md`](secrets.md#hive-level--one-of-each-per-hive) for
how it gets there. A hive lacking the queue's address for its
agents sets none of the four and each agent logs that it has none; a half-set
environment logs an error and the harness keeps serving.
<details><summary>What an agent publishes over the queue</summary>
What an agent does with that connection is publish its terminal. Every row its
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
a swarm-level terminal can follow one agent without subscribing to the swarm's
whole traffic. An agent connected with its own queue credential gets that
subject. An agent without one, or whose own credential the queue
refused, connects with its hive's shared client and publishes to
`$SWARM.term.<hive>.<agent>` instead, the `<hive>` being the one that client id
names. The swarm controller relays both. Publishing only: an
agent talks about itself here and reads nothing. Rows aren't retained — a
subscriber that wasn't listening missed them, the same as on the agent's own
live stream.
The queue would refuse a row too large for its `max_payload` outright and
take the connection down with it, so the harness drops such a row's body before
sending and leaves a marker in its place; the summary, level and icon still
arrive. The harness logs and skips a row that's too large even without its body.
The second thing an agent publishes is its **turn-state header**, on
`$SWARM.agent-state.<agent>` (or `$SWARM.agent-state.<hive>.<agent>`) — same shape of subject, same grant
mechanics, same lack of retention. It carries what a header bar wants: what the
turn loop is doing (`turn_state`, plus `turn_state_since` as an ISO 8601 UTC
stamp), which model (`model` and the resolved id the last turn actually ran on),
the context budget and the last turn's context and cost token blocks, and
`agent_state`.
`agent_state` reuses the swarm's own wanted-state vocabulary
(`up`/`offline`/`paused`/`destroyed`) so a reader can compare what an agent _is_
against what the swarm declared it should be without translating between two
spellings. ⚠️ From inside the container only two of those four are sayable: the
harness reports `up`, or `paused` when the pause marker is present. `offline` and
`destroyed` are hive-c0re's observations — a stopped agent publishes nothing and
a destroyed one doesn't exist — so a view that needs the full four-state picture
takes them from the `agent-status` bucket and uses this subject to sharpen the
rest.
Headers go out **on transition, not on a timer**: the harness rebuilds the
header whenever its event bus moves and publishes only when the result differs
from what it last sent. That's the whole point of the subject — the
`agent-status` bucket already republishes once a minute, which is far too slow
for "is this agent thinking right now." The cost of a core subject is that a
subscriber attaching mid-idle sees nothing until the next change, so a renderer
opens with the bucket's snapshot and lets this stream refine it.
Swarm-side, `GET /api/agents/<name>/state/stream` relays that subject as SSE,
resolving the agent's hive at request time exactly as the terminal stream does.
The payload passes through opaquely — the controller never parses a header.
The third thing an agent publishes is its **icon**, the same SVG its own
`GET /icon` serves. It goes into the `agent-icons` KV bucket under the key
`<agent>`, with no hive in it, so the swarm can show the icon of an agent that's
stopped or has moved hives. Only an agent connected with its own queue
credential publishes it: the queue grants that credential
`$KV.agent-icons.<agent>` and no other key, and grants the hive's shared client
none of the bucket. The harness writes once per start, because a config change
reaches an agent by restarting its container. An agent with no icon deletes its
key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
</details>
### Swarm-wide forge objects
The controller also keeps the forge objects that are one per swarm, not
one per hive. It ensures them at start and every five minutes after
(`swarm-controller/src/forge/objects.rs`):
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
owner org;
- the empty `operators` merge-gate team in `agents` and `agent-configs`;
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
- `internal/docs` (private) and `internal/knowledge` (public, with a
seed `README.md` while empty);
- the `agent-configs` org avatar
(`deploy.swarm-controller.configOrgAvatarPng`).
A pass that can't finish logs a
`warn` line per object plus `swarm forge objects: pass incomplete` in
`journalctl -u swarm-controller`, and retries on the next tick. While the
controller is down the objects stay as they are.
### Swarm-wide forge webhooks
At startup the controller registers two Forgejo hooks pointing at
itself — a `push` hook on `internal/knowledge` and a `pull_request` hook
on the `agent-configs` org, both under
`https://<swarm.domain>/webhook/forge/`.
The controller **interprets** a delivery and sends hives a specific
message — _the knowledge repo changed_, _deploy agent `foo` at rev
`abc123`_ — rather than forwarding forge payloads for each hive to
re-derive. Approval happens once, at the swarm level: a hive receives a
decision, not an event to adjudicate.
**`internal/knowledge` is on that path.** The controller's is the only
hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a
leftover). A webhook has exactly one target URL, so a second registration
would take delivery away from the first rather than add a recipient.
<!-- vale write-good.Passive = NO -->
**The `agent-configs` org isn't.** Each hive registers its own
`pull_request` hook there, so that repo has two — the hive's and the
controller's — and **both are expected; don't delete either.** Removing
a hive's stops it acting on config PRs; removing the controller's just
gets recreated on its next start.
<!-- vale write-good.Passive = YES -->
Nothing to configure. The controller registers the hooks only when this
host also serves the swarm UI vhost — that's what publishes the
endpoint, and a hook the forge can't reach would collect failed
deliveries while looking healthy. The controller generates the HMAC
secret on first start and keeps it
(see [`docs/agent-lifecycle/persistence.md`](../agent-lifecycle/persistence.md)).
To check it's working, push to `internal/knowledge` and look for
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
refused delivery logs `webhook: refused delivery` with the reason.
## Hives: the substrate
- **hive** — one host running `hive-c0re` and its agent containers
(`deploy.hive-controller.enable`). Addressed as `<hiveName>.<swarm.domain>`.
- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the
shared services and controller. You can qualify an agent as
`agent@hive-domain`.
- **peer hive** — any hive in the directory other than this one. Peers are
_derived_, not declared: the directory lists every hive including
yourself, and `hiveName` says which one you are.
### Hive identity config
```nix
services.hyperhive = {
swarm.domain = "example.com"; # required — the swarm's DNS domain
hiveName = "pr1ma"; # required — this hive's label in it
swarm.name = "constellat1on"; # shared swarm display name (optional)
# required — the directory, identical on every host in the swarm.
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
swarm.hives = {
pr1ma = { }; # this host, per hiveName
lab = { }; # a second hive
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
};
};
```
<!-- vale write-good.Passive = NO -->
Swarm options are identical on every host in the swarm, so `swarm.domain` is
**required** on every host that runs any hyperhive service; a host that
imports the module and enables nothing evaluates without it. `hiveName` is
required on a host that runs a hive, the secret store or the homeserver. Eval fails with a hint naming
each. Neither defaults, because a guessed value here is a wrong hostname
that evaluates cleanly and deploys.
<!-- vale write-good.Passive = YES -->
**The directory is one attrset, identical on every host.** It describes
every hive in the swarm, **including this one**, keyed by `hiveName`; what
differs between hosts is only `hiveName`. It **must** contain an entry for
this host's `hiveName`, and an empty directory fails that check too — eval
names the missing hive. Peers are _every entry but this host's_, so a
directory without this host would make every hive a peer, itself included.
```
# hive A # hive B
hiveName = "pr1ma"; hiveName = "edge";
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
```
One entry per hive means two hosts can't hold _different_ facts about the
same third hive, such as a stale endpoint.
`domain` defaults to `<name>.<swarm.domain>`, so a conventional directory
is names only. Set it only for a hive addressed by something else. This
hive's own `services.hyperhive.domain` comes from its entry; it drives
`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified
labels (`iris@pr1ma.example.com`). Setting `services.hyperhive.domain`
directly overrides the directory entry, with a deprecation warning: the
option is local to this host while every host shares the directory, so a
value written only here is invisible to the rest of the swarm.
`swarm.name` is display only — the dashboard chrome header and per-agent
system prompts — and federated hives at different domains can share one.
`hiveName` surfaces in the same places but is also the leftmost label of the
hive's domain. `swarm.name` names the group; `hiveName` names this hive.
The env-var chain and `qualify()` / `qualified_label()` semantics:
[`conventions.md`](../process/conventions.md) § Hive identity.
Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive
chains to it. Trusting a hive whose root this swarm doesn't own — another
swarm's — has no mechanism.
### What the directory feeds
1. **Swarm-wide hive roster** — swarm-controller reads this same
directory and serves it at `GET /api/hives`; the swarm UI's front page
renders it ([`ui.md`](ui.md)).
2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel
federates with the peer's matrix server (discovered via the peer's
`.well-known/matrix/server` delegation, which the gateway serves).
Federation validates the peer's TLS certificate against the matrix
**container's** trust bundle, independent of this directory.
⚠️ On a hive with self-signed gateway certificates, the container
also trusts this hive's `trust-bundle.pem`, bound in at runtime and
added to the public CAs; it ends at the swarm root, so a peer whose
certificate chains to the same root validates. A peer outside this
swarm's root needs a CA-issued certificate (ACME). Federation firewall
and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md).
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional).
### WireGuard inter-hive mesh (optional)
The peer config above uses public HTTPS for all inter-hive traffic.
For private deployments — or to reduce latency and TLS overhead on
intra-swarm traffic — hyperhive can configure a host-to-host
WireGuard mesh.
#### Generating keys
On each hive host:
```bash
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
```
#### Config example (two hives)
```nix
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
services.hyperhive = {
deploy.wireguard = {
enable = true;
privateKeyFile = "/etc/wireguard/hive.key";
address = "10.100.0.1/24";
listenPort = 51820; # optional, default 51820
};
# The same `hives` attrset both hosts hold — mesh fields included,
# since "where this hive can be dialled" is a fact about that hive.
swarm.hives = {
pr1ma = {
domain = "pr1ma.example.com";
wireguardPublicKey = "base64keyA=";
wireguardEndpoint = "198.51.100.1:51820";
wireguardAddress = "10.100.0.1/32";
};
edge = {
domain = "edge.corp";
wireguardPublicKey = "base64keyB=";
wireguardEndpoint = "203.0.113.42:51820";
wireguardAddress = "10.100.0.2/32";
};
};
};
# hive B (edge.corp, mesh IP 10.100.0.2)
services.hyperhive = {
deploy.wireguard = {
enable = true;
privateKeyFile = "/etc/wireguard/hive.key";
address = "10.100.0.2/24";
};
swarm.hives = { /* … identical to hive A's … */ };
};
```
#### What the mesh does
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
host (not inside agent containers; containers reach peers via the
host's routing table).
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
`allowedIPs`, so intra-swarm traffic can route over the mesh address
rather than the public domain.
- It sets `persistentKeepalive = 25` by default; override or null to
disable (not needed when both sides have public IPs and no NAT).
#### NAT / one-sided endpoints
If one host is behind NAT and can't accept incoming connections, only
that host needs a null `wireguardEndpoint` on the peer config — the
other side initiates. With keepalive on, the NAT hole stays open.
If both hosts are behind NAT, you need a STUN relay or a third host
(exit node); hyperhive sets up neither.
## Cross-references
- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist"
- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the
swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option
- [`../process/conventions.md`](../process/conventions.md) § Hive identity —
env vars, qualified labels
- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix
federation, TLS cert autogeneration, firewall posture
- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and
the `.well-known/matrix/` autodiscovery scheme