Watch
0
0
Fork
You've already forked hyperhive
0

docs(swarm): facts + structure pass

swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
This commit is contained in:
atlas 2026-10-01 23:25:24 +02:00 • committed by mara
commit 270430a4b4
5 changed files with 408 additions and 402 deletions

View file

@ -1,308 +1,55 @@
# Multi-hive swarms
# The swarm
A **swarm** is a collection of agents that share an identity and
coordinate across one or more hives. A single hyperhive instance
running on one host is already a swarm (one hive). This doc covers
the additional config needed when the swarm spans multiple hosts.
The **swarm** is where things live: agent identities and accounts, secrets,
the job graph that creates and places agents, telemetry, and the UI you
drive it all from. **Hives are the substrate** — NixOS hosts that run agent
containers on the swarm's behalf. Every hive belongs to a swarm; a single
host is a swarm of one.
For the full option reference rather than prose: `services.hyperhive.swarm.*`
(swarm-wide facts, identical on every host) and `services.hyperhive.deploy.*`
(this host's own deployment decisions — does _this_ machine run grafana,
the swarm controller, authelia, …) are separate generated pages, `nix
build .#docs-swarm` / `.#docs-deploy` or the website's `/options/swarm.html`
/ `/options/deploy.html`.
This page is for the operator. It covers the control plane, how a hive
joins the directory, and how hives and agents report upward. The steps for a
fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local
config is the [README quick start](../../README.md#quick-start-an-all-local-swarm).
## Terminology
Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical
on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs
grafana, the controller, authelia, …) →
[options reference](https://hyperhive.darkest.space/options/), or
`nix build .#docs-swarm` / `.#docs-deploy`.
- **hive** — a single hyperhive installation on one host. Has its
own `services.hyperhive.domain` DNS name and its own set of agent
containers.
- **swarm** — one or more hives whose operators have declared them
as peers. You can qualify an agent as `agent@hive-domain`.
- **peer hive** — any hive in `services.hyperhive.swarm.hives` other
than this one. Peers are _derived_, not declared: the directory lists
every hive including yourself, and `hiveName` says which one you are.
## Where each piece lives
## Hive identity config
```nix
services.hyperhive = {
swarm.domain = "example.com"; # required — the swarm's DNS domain
hiveName = "pr1ma"; # required — this hive's label in it
swarm.name = "constellat1on"; # shared swarm display name (optional)
# required — the directory, identical on every host in the swarm.
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
swarm.hives = {
pr1ma = { };
edge = { };
};
};
```
<!-- vale write-good.Passive = NO -->
`swarm.domain` and `hiveName` are **required** whenever hyperhive is
enabled; eval fails with a hint naming each. Neither defaults,
because a guessed value here is a wrong hostname that evaluates cleanly
and deploys — an eval failure asking the operator to write the address
down is the cheaper outcome. **Upgrading past this release means setting
both once.**
<!-- vale write-good.Passive = YES -->
You must still set `domain` too, but you no longer _write_ it: it's read from
this hive's own entry in the directory, whose `domain` defaults to
`<name>.<swarm.domain>`. A conventional swarm states no addresses at
all, and a hive addressed by something else states it in the one place
the other hives read — `swarm.hives.edge.domain = "edge.elsewhere.example";`.
Setting `services.hyperhive.domain` directly still works and still wins,
with a **deprecation warning**. The reason it's deprecated isn't tidiness:
that option is local to one host, and the operator copies the directory
to every host, so a value written only there leaves every peer pointing
somewhere else with nothing detecting the disagreement.
⚠️ **Upgrading:** a hive that has been running on `swarm.domain` +
`hiveName` alone now needs its own directory entry —
`services.hyperhive.swarm.hives.<hiveName> = { };`, one line, no value.
Eval fails naming it if you forget.
`domain` drives `HYPERHIVE_HIVE_DOMAIN` in every container so agents can
form qualified labels (`iris@pr1ma.example.com`).
`swarm.name` is purely display — it surfaces in the dashboard chrome
header and per-agent system prompts, and federated hives at different
domains can share one. `hiveName` surfaces in the same places but is
_not_ only display: it's the leftmost label of the hive's domain. That
`swarm.name` sits under `swarm` and `hiveName` doesn't is the whole
distinction — one names this hive, the other names the group it belongs
to.
See `docs/process/conventions.md` § Hive identity for the env-var chain
and `qualify()` / `qualified_label()` semantics.
## Swarm CA
A hive's internal TLS chains to a **swarm root CA**, so a peer that
trusts the root validates every hive in the swarm rather than pinning
to each one by hand. Provisioning modes, what to hand a peer
(`trust-bundle.pem`, never `ca.pem`), the name constraints on a hive
CA, and how an existing hive adopts the hierarchy: [`ca.md`](ca.md).
## Running the swarm's shared services
One authelia, one matrix, one forge per swarm — which host runs them,
and what a hive that runs none of them configures instead:
[`services.md`](services.md).
## Single sign-on
Which secrets the SSO provider generates, which one has a reader in
another container, and the three ways that one gets delivered:
[`sso.md`](sso.md).
## Secrets
Every credential the swarm holds, who mints it, where it must live, and
which of the three topologies makes it the operator's job to place:
[`secrets.md`](secrets.md).
Where that shape is **going** — the per-secret minter/reader/renewal
contract, the target of one mTLS identity per host and everything else
through the store, and the test a change has to pass to count as movement
toward it: [`credentials.md`](credentials.md). It supersedes `secrets.md`
when the migration completes.
## Swarm UI
The operator-only web surface on the swarm apex, why reaching it needs
the `admins` group rather than just a session, and the four sites you
wire a swarm service name into: [`ui.md`](ui.md).
## The swarm's hive directory
```nix
services.hyperhive.swarm.hives = {
pr1ma = { }; # this host, per hiveName
lab = { }; # a second hive in the swarm
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
};
```
One attrset describing **every** hive in the swarm, **including this
one**, keyed by that hive's `hiveName`. It's meant to be _identical on
every host_ — write it once, share it, and each host reads it correctly
because `services.hyperhive.hiveName` says which entry is itself.
Empty (the default) means this host isn't in a swarm. Once non-empty it
**must** contain an entry for `hiveName`; eval fails naming the missing
hive. That assertion is load-bearing rather than pedantic — "my peers"
comes from _everything that isn't me_, so a directory that doesn't
contain you derives every hive as a peer and you peer with yourself.
`domain` defaults to `<name>.<swarm.domain>`, the convention every hive
follows, so a conventional directory is names only. The default is a
derivation from two values the operator already had to state — the swarm's
domain and the entry's own name — rather than a guess, which is what makes
it safe here when a guessed hostname wouldn't be. Set it only for a hive
addressed by something else.
> **No per-hive CA field exists, and no per-hive cert pinning.** Trust
> inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every
> hive chains to it, so one anchor replaces per-hive pinning entirely.
> What that genuinely drops is trusting a hive whose root this swarm
> does _not_ own — another swarm's, or one keeping its own CA. That's
> a cross-swarm problem and wants a mechanism designed for it. (An
> earlier `certFingerprint` field existed for exactly that gap, pinning
> a peer's TLS leaf for hive-c0re's own peer HTTPS checks — removed
> along with the dashboard feature it existed to serve, since nothing
> else ever consumed it.)
## What the config does at runtime
1. **Swarm-wide hive roster** — swarm-controller reads this same
directory and serves it at `GET /api/hives`; `swarm-ui`'s overview
page renders it (`docs/swarm/ui.md`). This is the operator-facing
"what hives exist" surface — a per-hive dashboard "peer hives"
display existed here once; it no longer exists, in favour of this.
2. **Matrix federation** — when `matrix.enable` is on, tuwunel
federates with the peer's matrix server (discovered via the peer's
`.well-known/matrix/server` delegation, which the gateway serves).
Federation validates the peer's TLS certificate against the matrix
**container's** trust bundle, independent of this directory.
⚠️ **That container currently trusts no swarm-internal CA**, so a
self-signed gateway certificate doesn't federate. You can't list the
swarm root there: `security.pki.certificateFiles` is
read when the system is _built_, and the root is a runtime file (its
key must never enter the store), so there is no build-time name for
it. Bridging that needs a runtime mechanism; a separate issue tracks
it. Until then, federation needs CA-issued certs (ACME). See
`docs/integrations/matrix.md` for federation firewall + TLS requirements.
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
to configure `wg-hive`. See "WireGuard inter-hive mesh" below.
## One directory, not a bilateral declaration
Both hives hold the **same** `hives` attrset; neither declares the
other. What differs between the two hosts is only `hiveName`:
```
# hive A # hive B
hiveName = "pr1ma"; hiveName = "edge";
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
```
That's the point of the shape, and it removes a class of bug rather
than saving typing: a per-host peer list let two hosts hold _different_
facts about the same third hive — a stale endpoint, a rotated
fingerprint — with nothing to detect the disagreement. One entry per
hive makes it unrepresentable.
## WireGuard inter-hive mesh (optional)
The peer config above uses public HTTPS for all inter-hive traffic.
For private deployments — or to reduce latency and TLS overhead on
intra-swarm traffic — hive-c0re can configure a host-to-host
WireGuard mesh.
### Generating keys
On each hive host:
```bash
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
```
### Config example (two hives)
```nix
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
services.hyperhive = {
deploy.wireguard = {
enable = true;
privateKeyFile = "/etc/wireguard/hive.key";
address = "10.100.0.1/24";
listenPort = 51820; # optional, default 51820
};
# The same `hives` attrset both hosts hold — mesh fields included,
# since "where this hive can be dialled" is a fact about that hive.
swarm.hives = {
pr1ma = {
domain = "pr1ma.example.com";
wireguardPublicKey = "base64keyA=";
wireguardEndpoint = "198.51.100.1:51820";
wireguardAddress = "10.100.0.1/32";
};
edge = {
domain = "edge.corp";
wireguardPublicKey = "base64keyB=";
wireguardEndpoint = "203.0.113.42:51820";
wireguardAddress = "10.100.0.2/32";
};
};
};
# hive B (edge.corp, mesh IP 10.100.0.2)
services.hyperhive = {
deploy.wireguard = {
enable = true;
privateKeyFile = "/etc/wireguard/hive.key";
address = "10.100.0.2/24";
};
swarm.hives = { /* … identical to hive A's … */ };
};
```
### What the mesh does
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
host (not inside agent containers; containers reach peers via the
host's routing table).
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
`allowedIPs`, so intra-swarm traffic can route over the mesh address
rather than the public domain.
- It sets `persistentKeepalive = 25` by default; override or null to
disable (not needed when both sides have public IPs and no NAT).
### NAT / one-sided endpoints
If one host is behind NAT and can't accept incoming connections, only
that host needs a null `wireguardEndpoint` on the peer config — the
other side initiates. With keepalive on, the NAT hole stays open.
If both hosts are behind NAT, you need a STUN relay or a third host
(exit node). Out of scope for v0.
## Snapshot store
One further option lives in this namespace but its docs live with the
service it points at: `services.hyperhive.swarm.snapshotStore.{address,
port}` tells this hive where the swarm's `btrfs receive` endpoint is, so
`hivectl agent <name> subvol snapshot push` has somewhere to stream to.
It's genuinely swarm-scoped rather than per-peer — a swarm has exactly
one store, because the receiver keys destinations by _agent_ so a
migrating agent keeps one unbroken incremental chain. See
[snapshot-store.md](../networking/snapshot-store.md).
- **control plane** — `swarm-controller`: hive directory, agent roster, job
graph, agent creation. → [below](#swarm-controller)
- **swarm UI** — the operator's day-to-day surface, on the swarm apex,
`admins` only. → [`ui.md`](ui.md)
- **shared services** — one forge, homeserver, SSO, queue and metrics/logs
stack, each on whichever host you put it. → [`services.md`](services.md)
- **SSO** — the secrets authelia generates and how each reaches its reader.
→ [`sso.md`](sso.md)
- **secrets** — every credential the swarm holds, who mints it and where it
lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal
contract → [`credentials.md`](credentials.md) · how the store comes up,
who writes its grants and how it unseals → [`bao.md`](bao.md)
- **swarm CA** — the root every hive's internal TLS chains to, and what to
hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md)
- **snapshot store** — the swarm's one `btrfs receive` endpoint,
`swarm.snapshotStore.{address,port}`. →
[`snapshot-store.md`](../networking/snapshot-store.md)
## Swarm controller
`services.hyperhive.deploy.swarm-controller.enable` runs the `swarm-controller`
daemon on this host. **Off by default and deliberately not derived from
`services.hyperhive.deploy.hive-controller.enable`**: a swarm has one
`swarm-controller` is the swarm's control plane: it holds the hive directory,
the agent roster and their wanted state, and the job graph. Creating an agent —
from the swarm UI or `swarmctl agent create --hive <h>` — queues its SSO
identity, forge user, config repo, store identity and matrix account, then
sends the hive a deploy message. A new agent starts `paused`.
`services.hyperhive.deploy.swarm-controller.enable` runs it on this host;
`singleHostSwarm` turns it on. **Otherwise off by default and deliberately
not derived from `deploy.hive-controller.enable`**: a swarm has one
controller, so enabling it states a fact about swarm topology, not about
whether this host runs a hive. Every hive runs `hive-c0re` (the agents on that host); one
hive additionally runs this (what's true across hives).
whether this host runs a hive.
What it serves, why it's a unix socket rather than a port, and the
socket-directory constraint that governs where `socketPath` may point:
@ -409,6 +156,8 @@ how it gets there. A hive lacking the queue's address for its
agents sets none of the four and each agent logs that it has none; a half-set
environment logs an error and the harness keeps serving.
<details><summary>What an agent publishes over the queue</summary>
What an agent does with that connection is publish its terminal. Every row its
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
a swarm-level terminal can follow one agent without subscribing to the swarm's
@ -468,6 +217,8 @@ key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
</details>
### Swarm-wide forge objects
The controller also keeps the forge objects that are one per swarm, not
@ -484,7 +235,7 @@ one per hive. It ensures them at start and every five minutes after
- the `agent-configs` org avatar
(`deploy.swarm-controller.configOrgAvatarPng`).
hive-c0re no longer creates any of them. A pass that can't finish logs a
A pass that can't finish logs a
`warn` line per object plus `swarm forge objects: pass incomplete` in
`journalctl -u swarm-controller`, and retries on the next tick. While the
controller is down the objects stay as they are.
@ -503,14 +254,14 @@ re-derive. Approval happens once, at the swarm level: a hive receives a
decision, not an event to adjudicate.
**`internal/knowledge` is on that path.** The controller's is the only
hook on it: hives no longer register their own (see
`docs/integrations/knowledge.md` for clearing a leftover). A webhook has exactly one target URL, so per-hive
registration never added a recipient — it took delivery away from
whichever hive registered before it.
hook on it; hives register none of their own
([`knowledge.md`](../integrations/knowledge.md) covers clearing a leftover).
A webhook has exactly one target URL, so a second registration would take
delivery away from the first rather than add a recipient.
<!-- vale write-good.Passive = NO -->
**The `agent-configs` org isn't yet.** Each hive still registers its own
**The `agent-configs` org isn't.** Each hive registers its own
`pull_request` hook there, so that repo has two — the hive's and the
controller's — and **both are expected; don't delete either.** Removing
a hive's stops it acting on config PRs; removing the controller's just
@ -529,15 +280,206 @@ To check it's working, push to `internal/knowledge` and look for
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
refused delivery logs `webhook: refused delivery` with the reason.
## Hives: the substrate
- **hive** — one host running `hive-c0re` and its agent containers
(`deploy.hive-controller.enable`). Addressed as `<hiveName>.<swarm.domain>`.
- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the
shared services and controller. You can qualify an agent as
`agent@hive-domain`.
- **peer hive** — any hive in the directory other than this one. Peers are
_derived_, not declared: the directory lists every hive including
yourself, and `hiveName` says which one you are.
### Hive identity config
```nix
services.hyperhive = {
swarm.domain = "example.com"; # required — the swarm's DNS domain
hiveName = "pr1ma"; # required — this hive's label in it
swarm.name = "constellat1on"; # shared swarm display name (optional)
# required — the directory, identical on every host in the swarm.
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
swarm.hives = {
pr1ma = { }; # this host, per hiveName
lab = { }; # a second hive
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
};
};
```
<!-- vale write-good.Passive = NO -->
`swarm.domain` is **required** on a host that runs a hive, and `hiveName` on
a host that runs a hive, the secret store or the homeserver; eval fails with
a hint naming each. Neither defaults, because a guessed value here is a
wrong hostname that evaluates cleanly and deploys.
<!-- vale write-good.Passive = YES -->
**The directory is one attrset, identical on every host.** It describes
every hive in the swarm, **including this one**, keyed by `hiveName`; what
differs between hosts is only `hiveName`. It **must** contain an entry for
this host's `hiveName`, and an empty directory fails that check too — eval
names the missing hive. Peers are _every entry but this host's_, so a
directory without this host would make every hive a peer, itself included.
```
# hive A # hive B
hiveName = "pr1ma"; hiveName = "edge";
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
```
One entry per hive means two hosts can't hold _different_ facts about the
same third hive, such as a stale endpoint.
`domain` defaults to `<name>.<swarm.domain>`, so a conventional directory
is names only. Set it only for a hive addressed by something else. This
hive's own `services.hyperhive.domain` comes from its entry; it drives
`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified
labels (`iris@pr1ma.example.com`).
`swarm.name` is display only — the dashboard chrome header and per-agent
system prompts — and federated hives at different domains can share one.
`hiveName` surfaces in the same places but is also the leftmost label of the
hive's domain. `swarm.name` names the group; `hiveName` names this hive.
The env-var chain and `qualify()` / `qualified_label()` semantics:
[`conventions.md`](../process/conventions.md) § Hive identity.
Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive
chains to it, so there is no per-hive CA field and no per-hive cert pinning.
Trusting a hive whose root this swarm doesn't own — another swarm's — has no
mechanism.
<details><summary>Upgrading an existing hive</summary>
- Set `swarm.domain` and `hiveName` once; eval fails naming each until you do.
- Add the hive's own directory entry,
`services.hyperhive.swarm.hives.<hiveName> = { };` — one line, no value.
Eval fails naming it if you forget.
- Setting `services.hyperhive.domain` directly still works and still wins,
with a **deprecation warning**: the option is local to one host while
every host holds the directory, so a value written only there leaves every
peer pointing somewhere else. Move it into the hive's directory entry, or
drop it if it's the conventional `<hiveName>.<swarm.domain>`.
- Setting `swarm.hives.<name>.certFingerprint` fails eval; the field no
longer exists. Trust comes from the swarm root instead.
</details>
### What the directory feeds
1. **Swarm-wide hive roster** — swarm-controller reads this same
directory and serves it at `GET /api/hives`; the swarm UI's front page
renders it ([`ui.md`](ui.md)).
2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel
federates with the peer's matrix server (discovered via the peer's
`.well-known/matrix/server` delegation, which the gateway serves).
Federation validates the peer's TLS certificate against the matrix
**container's** trust bundle, independent of this directory.
⚠️ On a hive with self-signed gateway certificates, the container
also trusts this hive's `trust-bundle.pem`, bound in at runtime and
added to the public CAs; it ends at the swarm root, so a peer whose
certificate chains to the same root validates. A peer outside this
swarm's root needs a CA-issued certificate (ACME). Federation firewall
and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md).
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional).
### WireGuard inter-hive mesh (optional)
The peer config above uses public HTTPS for all inter-hive traffic.
For private deployments — or to reduce latency and TLS overhead on
intra-swarm traffic — hyperhive can configure a host-to-host
WireGuard mesh.
#### Generating keys
On each hive host:
```bash
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
```
#### Config example (two hives)
```nix
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
services.hyperhive = {
deploy.wireguard = {
enable = true;
privateKeyFile = "/etc/wireguard/hive.key";
address = "10.100.0.1/24";
listenPort = 51820; # optional, default 51820
};
# The same `hives` attrset both hosts hold — mesh fields included,
# since "where this hive can be dialled" is a fact about that hive.
swarm.hives = {
pr1ma = {
domain = "pr1ma.example.com";
wireguardPublicKey = "base64keyA=";
wireguardEndpoint = "198.51.100.1:51820";
wireguardAddress = "10.100.0.1/32";
};
edge = {
domain = "edge.corp";
wireguardPublicKey = "base64keyB=";
wireguardEndpoint = "203.0.113.42:51820";
wireguardAddress = "10.100.0.2/32";
};
};
};
# hive B (edge.corp, mesh IP 10.100.0.2)
services.hyperhive = {
deploy.wireguard = {
enable = true;
privateKeyFile = "/etc/wireguard/hive.key";
address = "10.100.0.2/24";
};
swarm.hives = { /* … identical to hive A's … */ };
};
```
#### What the mesh does
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
host (not inside agent containers; containers reach peers via the
host's routing table).
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
`allowedIPs`, so intra-swarm traffic can route over the mesh address
rather than the public domain.
- It sets `persistentKeepalive = 25` by default; override or null to
disable (not needed when both sides have public IPs and no NAT).
#### NAT / one-sided endpoints
If one host is behind NAT and can't accept incoming connections, only
that host needs a null `wireguardEndpoint` on the peer config — the
other side initiates. With keepalive on, the NAT hole stays open.
If both hosts are behind NAT, you need a STUN relay or a third host
(exit node); hyperhive sets up neither.
## Cross-references
- `docs/networking/snapshot-store.md` — the swarm's `btrfs receive` endpoint, and
the `swarm.snapshotStore` option that points a hive at it
- `docs/process/conventions.md` § Hive identity — env vars, qualified labels
- `docs/integrations/matrix.md` — matrix federation, TLS cert autogeneration,
firewall posture
- `docs/swarm/ui.md` — the swarm-wide hive roster, now the operator
surface for "what hives exist" (superseded the per-hive dashboard's
old "peer hives" display)
- `docs/networking/gateway.md` — nginx vhosts and the `.well-known/matrix/`
autodiscovery scheme
- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist"
- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the
swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option
- [`../process/conventions.md`](../process/conventions.md) § Hive identity —
env vars, qualified labels
- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix
federation, TLS cert autogeneration, firewall posture
- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and
the `.well-known/matrix/` autodiscovery scheme

View file

@ -1,7 +1,12 @@
# Swarm-wide services
Some things exist once per **swarm** rather than once per hive. Two
options say where the optional ones live, and everything else derives:
Some things exist once per **swarm** rather than once per hive: the forge,
the matrix homeserver, SSO, the secret store, the queue, and the metrics and
log stack. This page says which host runs them and what a hive that runs none
of them configures instead. The all-local quick start sets everything with one
line → [README](../../README.md#quick-start-an-all-local-swarm).
Two options say where they live, and everything else derives:
```nix
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
@ -14,10 +19,15 @@ here" means: every once-per-swarm service takes its `enable` from it.** That's t
sections below don't repeat it, so a service that stops deriving is a
visible difference rather than one more paragraph saying the same thing.
`singleHostSwarm` is the all-on-one-box switch above it: it defaults
both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host
generates the swarm CA here). You can still set each derived toggle on its own,
which wins, so "all local except X" needs no further option.
`singleHostSwarm` is the all-on-one-box mode above it. It defaults
`deploy.allSwarmServices`, the swarm CA (`swarm.ca.autoConfigure`, generated
on this host), the swarm controller (`deploy.swarm-controller.enable`), the
host's `/etc/hosts` entries for the names it serves
(`gateway.localHostsEntry`), the queue's auth-callout keys
(`deploy.nats.autoGenerateCallout`) and where the secret store's bootstrap
token goes (`deploy.bao.bootstrapTokenFile`). You can still set each derived
toggle on its own, which wins, so "all local except X" needs no further
option.
**Both default to off**, and that's deliberate: a host can't tell
whether it's meant to be the swarm's service host, so this is an
@ -38,23 +48,29 @@ answers its name from its own resolver, so on a swarm spread over
more than one host, the operator's DNS has to resolve those names to that
host.
<details><summary>Moving an existing hive's forge to the swarm's</summary>
A hive that stops running the forge keeps the old container's state at
`/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's
forge: push anything worth keeping there by hand. Its
`/var/lib/hyperhive/forge-core-token` came from that old forge and
fails against the swarm's one.
</details>
## Deployment shapes
Those two options are what makes the difference between deployments, so
the shapes worth naming are the ones they produce:
- **All-local.** Everything on one machine:
`singleHostSwarm = true`. Setup is automatic apart from
choosing a domain and creating the first user.
`singleHostSwarm = true`, plus `deploy.hive-controller.enable = true` for
a hive to run agents on. After the first switch, the steps in
[`setup.md`](../getting-started/setup.md) remain.
- **Services on the swarm controller host.**
`deploy.allSwarmServices = true` there; the required services
deploy together on that host, with hives elsewhere.
Set `deploy.allSwarmServices` and `deploy.swarm-controller.enable` there,
with hives elsewhere. The controller doesn't derive from
`allSwarmServices`.
- **Fully spread out.** One container / VM / machine per service,
somewhere.
@ -88,12 +104,10 @@ there is one IdP and one auth path.
container. Set it explicitly when joining a swarm whose IdP is under
another name.
swarm-controller writes the users database, not by hand: hive-c0re
creates and destroys agents continuously, so the subject set is dynamic.
This module only guarantees the file exists and parses, so authelia
starts with nobody in it rather than failing to start — a provider with
no subjects yet is the correct state before anything has provisioned
them. Authelia generates session and storage keys in the container on
Agent subjects come from swarm-controller's agent-creation job, written
into the users database by `swarm-authelia-bridge`; human ones come from
`swarmctl user add` → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
On first boot this module seeds an empty users database. Authelia generates session and storage keys in the container on
first boot and never rotates them automatically; replacing one
invalidates data already written (sessions, the encrypted store), so
that's an operator action.
@ -196,9 +210,9 @@ gateway either way.
**Both store exporters are unconditional**, and `deploy.victoriametrics.enable`
doesn't gate them: that option says this host _runs_ the store, while the swarm
has one either way, reached by its swarm name through the gateway. Gating on it
once left a collector on any other host with no exporter at all — receiving from
every hive and dropping it, silently, because an absent exporter isn't an error.
has one either way, reached by its swarm name through the gateway. A collector
with no exporter would receive from every hive and drop it silently, because an
absent exporter isn't an error.
Agent-side configuration, and what a hive's own collector does, are in
[`../scheduler/observability.md`](../scheduler/observability.md).

View file

@ -1,10 +1,23 @@
# Swarm UI
The swarm's own web surface, served by the gateway on the **swarm apex**
(`services.hyperhive.swarm.domain`) and readable only by operators.
The swarm's own web surface and the operator's day-to-day view: served by
the gateway on the **swarm apex** (`services.hyperhive.swarm.domain`),
readable only by operators. The per-hive dashboard, on each hive's own
domain, covers host-level detail for one hive.
Distinct from the per-hive dashboard, which lives on the hive domain and
answers for one host. This one is the view _across_ hives.
## What it shows
| route | what |
| ------------------------- | ------------------------------------------------------------------------------------------- |
| `/` | the hive directory, each hive with its last reported status |
| `/agents` | every agent: status, config PR, wanted state; create agents, link forge and matrix accounts |
| `/agents/<name>/terminal` | one agent's live terminal |
| `/jobs` | the controller's job graph — where agent creation and credential mints show progress |
| `/issues` | a cross-repo issue report |
Everything it shows comes from [`swarm-controller`](../../swarm-controller/README.md).
An agent created here or with `swarmctl agent create` starts `paused`; set it
`up` from its card.
## Enabling
@ -35,28 +48,16 @@ requiring `group:admins`. An account without that group authenticates
fine and still gets bounced.
```sh
swarmctl user add <you> --group admins
swarmctl user add <you> --email <you>@example.com --group admins
swarmctl user update <you> --add-group admins # an account that already exists
```
<!-- vale write-good.Passive = NO -->
`--email` isn't needed for the UI, but the forge won't create your account
without one → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
`admins` deliberately, not a new word: [`../getting-started/setup.md`](../getting-started/setup.md) has
told every operator to create exactly that group since the bootstrap step
existed, so an account made by following the guide already passes. This
is the first rule that _consumes_ a group name — inventing a second one
would have meant those accounts silently failing a check they were
supposed to pass.
<!-- vale write-good.Passive = YES -->
An account created without any group needs re-adding with the flag —
`swarmctl` reads the existing entry out of `users.yml,` so the group is
what changes.
Why a group and not a list of usernames: agents are getting authelia
accounts of their own (matrix SSO), and _authenticated_ would then
include every agent in the hive. The group is the only thing standing
between "an operator's page" and "anyone with a session."
Why a group and not "any session": agents are authelia subjects too, so
_authenticated_ includes every agent in the swarm. The group is the only
thing standing between "an operator's page" and "anyone with a session."
## What it costs to be reachable
@ -66,7 +67,22 @@ not a hole: **reachability isn't the access control here.** An agent
that resolves the name and connects still has no operator session, and
the subrequest denies it.
## Two wiring sites
## Quick links
The swarm UI's header carries a single 🔗 button, visible on every route,
opening a popover of links to other swarm-wide services. Backed by
`GET /api/links` (swarm-controller), which serves
`services.hyperhive.swarm.controller.links` (a `listOf { label, icon, url }`,
same shape as the per-agent `services.hyperhive.agent.dashboardLinks`).
Each service's own module contributes its entry when it's enabled on the
controller's host — `swarm-authelia.nix`, `hive-matrix.nix`,
`hive-forge/default.nix`, `swarm-grafana.nix`, `swarm-victorialogs.nix`, and
`swarm-ui.nix` for this UI's own API docs. Adding a link for a new service is
a nix-only change to that service's module, or an operator adding an entry
directly. An empty list hides the button.
<details><summary>Adding a swarm service name: the two wiring sites</summary>
Adding a swarm service name means touching two things. Missing the
second ships as a different flavour of "works from the host, broken from
@ -99,23 +115,7 @@ and the apex is a **sibling** of `forge.<swarm>` / `chat.<swarm>` /
implicitly. Left out, the vhost falls back to the hive leaf and the
swarm's front page opens with a name mismatch.
## Quick links
The swarm UI's header carries a single 🔗 button, visible on every route,
opening a popover of links to other swarm-wide services — authelia,
matrix, forge, this UI's own swagger docs. Backed by `GET /api/links`
(swarm-controller), which serves `services.hyperhive.swarm.controller.links`
(a `listOf { label, icon, url }`, same shape as the per-agent
`services.hyperhive.agent.dashboardLinks`).
Rather than one central hardcoded list, each service's own module
contributes its own entry when it's actually enabled on the controller's
host — `swarm-authelia.nix`, `hive-matrix.nix` and `hive-forge/default.nix`
all do, the same list-merge idiom `gateway.localNames` uses above. Adding a
link for a new service is a nix-only change to that service's own module
(or an operator adding an entry directly); no swarm-controller or swarm-ui
change needed. Empty list hides the button rather than showing an empty
popover.
</details>
## Cross-references