docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity and the directory follow as the substrate. Upgrade notes move into a <details> block, the per-agent queue publishing detail into another, and the one-paragraph pointer sections collapse into a link list. Fact fixes, checked against origin/main: - an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean "not in a swarm" - swarm.domain is required with a hive (hive-network.nix:156,188), hiveName with a hive, store or homeserver (hyperhive.nix:161-166) - the matrix container trusts the hive's trust-bundle.pem at runtime under self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85) - singleHostSwarm also defaults the controller, localHostsEntry, the nats callout keys and the bao bootstrap token path (local-defaults.nix:72-129) - swarm-controller serves far more than /health: roster, wanted state, job graph, agent creation and credential mints (main.rs:2874-2899) - swarmctl user add needs --email for the forge account and refuses an existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document agent mint-identity and mint-forge-token - agent creation also mints store identity, forge token and matrix account, and declares the agent paused (main.rs:1822-1920, 247-248) Refs #3902
This commit is contained in:
parent
f688cfcdf0
commit
270430a4b4
5 changed files with 408 additions and 402 deletions
|
|
@ -1,308 +1,55 @@
|
|||
# Multi-hive swarms
|
||||
# The swarm
|
||||
|
||||
A **swarm** is a collection of agents that share an identity and
|
||||
coordinate across one or more hives. A single hyperhive instance
|
||||
running on one host is already a swarm (one hive). This doc covers
|
||||
the additional config needed when the swarm spans multiple hosts.
|
||||
The **swarm** is where things live: agent identities and accounts, secrets,
|
||||
the job graph that creates and places agents, telemetry, and the UI you
|
||||
drive it all from. **Hives are the substrate** — NixOS hosts that run agent
|
||||
containers on the swarm's behalf. Every hive belongs to a swarm; a single
|
||||
host is a swarm of one.
|
||||
|
||||
For the full option reference rather than prose: `services.hyperhive.swarm.*`
|
||||
(swarm-wide facts, identical on every host) and `services.hyperhive.deploy.*`
|
||||
(this host's own deployment decisions — does _this_ machine run grafana,
|
||||
the swarm controller, authelia, …) are separate generated pages, `nix
|
||||
build .#docs-swarm` / `.#docs-deploy` or the website's `/options/swarm.html`
|
||||
/ `/options/deploy.html`.
|
||||
This page is for the operator. It covers the control plane, how a hive
|
||||
joins the directory, and how hives and agents report upward. The steps for a
|
||||
fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local
|
||||
config is the [README quick start](../../README.md#quick-start-an-all-local-swarm).
|
||||
|
||||
## Terminology
|
||||
Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical
|
||||
on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs
|
||||
grafana, the controller, authelia, …) →
|
||||
[options reference](https://hyperhive.darkest.space/options/), or
|
||||
`nix build .#docs-swarm` / `.#docs-deploy`.
|
||||
|
||||
- **hive** — a single hyperhive installation on one host. Has its
|
||||
own `services.hyperhive.domain` DNS name and its own set of agent
|
||||
containers.
|
||||
- **swarm** — one or more hives whose operators have declared them
|
||||
as peers. You can qualify an agent as `agent@hive-domain`.
|
||||
- **peer hive** — any hive in `services.hyperhive.swarm.hives` other
|
||||
than this one. Peers are _derived_, not declared: the directory lists
|
||||
every hive including yourself, and `hiveName` says which one you are.
|
||||
## Where each piece lives
|
||||
|
||||
## Hive identity config
|
||||
|
||||
```nix
|
||||
services.hyperhive = {
|
||||
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
||||
hiveName = "pr1ma"; # required — this hive's label in it
|
||||
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
||||
|
||||
# required — the directory, identical on every host in the swarm.
|
||||
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
||||
swarm.hives = {
|
||||
pr1ma = { };
|
||||
edge = { };
|
||||
};
|
||||
};
|
||||
```
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
`swarm.domain` and `hiveName` are **required** whenever hyperhive is
|
||||
enabled; eval fails with a hint naming each. Neither defaults,
|
||||
because a guessed value here is a wrong hostname that evaluates cleanly
|
||||
and deploys — an eval failure asking the operator to write the address
|
||||
down is the cheaper outcome. **Upgrading past this release means setting
|
||||
both once.**
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
You must still set `domain` too, but you no longer _write_ it: it's read from
|
||||
this hive's own entry in the directory, whose `domain` defaults to
|
||||
`<name>.<swarm.domain>`. A conventional swarm states no addresses at
|
||||
all, and a hive addressed by something else states it in the one place
|
||||
the other hives read — `swarm.hives.edge.domain = "edge.elsewhere.example";`.
|
||||
|
||||
Setting `services.hyperhive.domain` directly still works and still wins,
|
||||
with a **deprecation warning**. The reason it's deprecated isn't tidiness:
|
||||
that option is local to one host, and the operator copies the directory
|
||||
to every host, so a value written only there leaves every peer pointing
|
||||
somewhere else with nothing detecting the disagreement.
|
||||
|
||||
⚠️ **Upgrading:** a hive that has been running on `swarm.domain` +
|
||||
`hiveName` alone now needs its own directory entry —
|
||||
`services.hyperhive.swarm.hives.<hiveName> = { };`, one line, no value.
|
||||
Eval fails naming it if you forget.
|
||||
|
||||
`domain` drives `HYPERHIVE_HIVE_DOMAIN` in every container so agents can
|
||||
form qualified labels (`iris@pr1ma.example.com`).
|
||||
|
||||
`swarm.name` is purely display — it surfaces in the dashboard chrome
|
||||
header and per-agent system prompts, and federated hives at different
|
||||
domains can share one. `hiveName` surfaces in the same places but is
|
||||
_not_ only display: it's the leftmost label of the hive's domain. That
|
||||
`swarm.name` sits under `swarm` and `hiveName` doesn't is the whole
|
||||
distinction — one names this hive, the other names the group it belongs
|
||||
to.
|
||||
|
||||
See `docs/process/conventions.md` § Hive identity for the env-var chain
|
||||
and `qualify()` / `qualified_label()` semantics.
|
||||
|
||||
## Swarm CA
|
||||
|
||||
A hive's internal TLS chains to a **swarm root CA**, so a peer that
|
||||
trusts the root validates every hive in the swarm rather than pinning
|
||||
to each one by hand. Provisioning modes, what to hand a peer
|
||||
(`trust-bundle.pem`, never `ca.pem`), the name constraints on a hive
|
||||
CA, and how an existing hive adopts the hierarchy: [`ca.md`](ca.md).
|
||||
|
||||
## Running the swarm's shared services
|
||||
|
||||
One authelia, one matrix, one forge per swarm — which host runs them,
|
||||
and what a hive that runs none of them configures instead:
|
||||
[`services.md`](services.md).
|
||||
|
||||
## Single sign-on
|
||||
|
||||
Which secrets the SSO provider generates, which one has a reader in
|
||||
another container, and the three ways that one gets delivered:
|
||||
[`sso.md`](sso.md).
|
||||
|
||||
## Secrets
|
||||
|
||||
Every credential the swarm holds, who mints it, where it must live, and
|
||||
which of the three topologies makes it the operator's job to place:
|
||||
[`secrets.md`](secrets.md).
|
||||
|
||||
Where that shape is **going** — the per-secret minter/reader/renewal
|
||||
contract, the target of one mTLS identity per host and everything else
|
||||
through the store, and the test a change has to pass to count as movement
|
||||
toward it: [`credentials.md`](credentials.md). It supersedes `secrets.md`
|
||||
when the migration completes.
|
||||
|
||||
## Swarm UI
|
||||
|
||||
The operator-only web surface on the swarm apex, why reaching it needs
|
||||
the `admins` group rather than just a session, and the four sites you
|
||||
wire a swarm service name into: [`ui.md`](ui.md).
|
||||
|
||||
## The swarm's hive directory
|
||||
|
||||
```nix
|
||||
services.hyperhive.swarm.hives = {
|
||||
pr1ma = { }; # this host, per hiveName
|
||||
lab = { }; # a second hive in the swarm
|
||||
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
||||
};
|
||||
```
|
||||
|
||||
One attrset describing **every** hive in the swarm, **including this
|
||||
one**, keyed by that hive's `hiveName`. It's meant to be _identical on
|
||||
every host_ — write it once, share it, and each host reads it correctly
|
||||
because `services.hyperhive.hiveName` says which entry is itself.
|
||||
|
||||
Empty (the default) means this host isn't in a swarm. Once non-empty it
|
||||
**must** contain an entry for `hiveName`; eval fails naming the missing
|
||||
hive. That assertion is load-bearing rather than pedantic — "my peers"
|
||||
comes from _everything that isn't me_, so a directory that doesn't
|
||||
contain you derives every hive as a peer and you peer with yourself.
|
||||
|
||||
`domain` defaults to `<name>.<swarm.domain>`, the convention every hive
|
||||
follows, so a conventional directory is names only. The default is a
|
||||
derivation from two values the operator already had to state — the swarm's
|
||||
domain and the entry's own name — rather than a guess, which is what makes
|
||||
it safe here when a guessed hostname wouldn't be. Set it only for a hive
|
||||
addressed by something else.
|
||||
|
||||
> **No per-hive CA field exists, and no per-hive cert pinning.** Trust
|
||||
> inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every
|
||||
> hive chains to it, so one anchor replaces per-hive pinning entirely.
|
||||
> What that genuinely drops is trusting a hive whose root this swarm
|
||||
> does _not_ own — another swarm's, or one keeping its own CA. That's
|
||||
> a cross-swarm problem and wants a mechanism designed for it. (An
|
||||
> earlier `certFingerprint` field existed for exactly that gap, pinning
|
||||
> a peer's TLS leaf for hive-c0re's own peer HTTPS checks — removed
|
||||
> along with the dashboard feature it existed to serve, since nothing
|
||||
> else ever consumed it.)
|
||||
|
||||
## What the config does at runtime
|
||||
|
||||
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
||||
directory and serves it at `GET /api/hives`; `swarm-ui`'s overview
|
||||
page renders it (`docs/swarm/ui.md`). This is the operator-facing
|
||||
"what hives exist" surface — a per-hive dashboard "peer hives"
|
||||
display existed here once; it no longer exists, in favour of this.
|
||||
|
||||
2. **Matrix federation** — when `matrix.enable` is on, tuwunel
|
||||
federates with the peer's matrix server (discovered via the peer's
|
||||
`.well-known/matrix/server` delegation, which the gateway serves).
|
||||
Federation validates the peer's TLS certificate against the matrix
|
||||
**container's** trust bundle, independent of this directory.
|
||||
|
||||
⚠️ **That container currently trusts no swarm-internal CA**, so a
|
||||
self-signed gateway certificate doesn't federate. You can't list the
|
||||
swarm root there: `security.pki.certificateFiles` is
|
||||
read when the system is _built_, and the root is a runtime file (its
|
||||
key must never enter the store), so there is no build-time name for
|
||||
it. Bridging that needs a runtime mechanism; a separate issue tracks
|
||||
it. Until then, federation needs CA-issued certs (ACME). See
|
||||
`docs/integrations/matrix.md` for federation firewall + TLS requirements.
|
||||
|
||||
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
||||
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
||||
to configure `wg-hive`. See "WireGuard inter-hive mesh" below.
|
||||
|
||||
## One directory, not a bilateral declaration
|
||||
|
||||
Both hives hold the **same** `hives` attrset; neither declares the
|
||||
other. What differs between the two hosts is only `hiveName`:
|
||||
|
||||
```
|
||||
# hive A # hive B
|
||||
hiveName = "pr1ma"; hiveName = "edge";
|
||||
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
||||
```
|
||||
|
||||
That's the point of the shape, and it removes a class of bug rather
|
||||
than saving typing: a per-host peer list let two hosts hold _different_
|
||||
facts about the same third hive — a stale endpoint, a rotated
|
||||
fingerprint — with nothing to detect the disagreement. One entry per
|
||||
hive makes it unrepresentable.
|
||||
|
||||
## WireGuard inter-hive mesh (optional)
|
||||
|
||||
The peer config above uses public HTTPS for all inter-hive traffic.
|
||||
For private deployments — or to reduce latency and TLS overhead on
|
||||
intra-swarm traffic — hive-c0re can configure a host-to-host
|
||||
WireGuard mesh.
|
||||
|
||||
### Generating keys
|
||||
|
||||
On each hive host:
|
||||
|
||||
```bash
|
||||
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
||||
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
||||
```
|
||||
|
||||
### Config example (two hives)
|
||||
|
||||
```nix
|
||||
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.1/24";
|
||||
listenPort = 51820; # optional, default 51820
|
||||
};
|
||||
|
||||
# The same `hives` attrset both hosts hold — mesh fields included,
|
||||
# since "where this hive can be dialled" is a fact about that hive.
|
||||
swarm.hives = {
|
||||
pr1ma = {
|
||||
domain = "pr1ma.example.com";
|
||||
wireguardPublicKey = "base64keyA=";
|
||||
wireguardEndpoint = "198.51.100.1:51820";
|
||||
wireguardAddress = "10.100.0.1/32";
|
||||
};
|
||||
edge = {
|
||||
domain = "edge.corp";
|
||||
wireguardPublicKey = "base64keyB=";
|
||||
wireguardEndpoint = "203.0.113.42:51820";
|
||||
wireguardAddress = "10.100.0.2/32";
|
||||
};
|
||||
};
|
||||
};
|
||||
|
||||
# hive B (edge.corp, mesh IP 10.100.0.2)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.2/24";
|
||||
};
|
||||
|
||||
swarm.hives = { /* … identical to hive A's … */ };
|
||||
};
|
||||
```
|
||||
|
||||
### What the mesh does
|
||||
|
||||
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
||||
host (not inside agent containers; containers reach peers via the
|
||||
host's routing table).
|
||||
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
||||
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
||||
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
||||
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
||||
rather than the public domain.
|
||||
- It sets `persistentKeepalive = 25` by default; override or null to
|
||||
disable (not needed when both sides have public IPs and no NAT).
|
||||
|
||||
### NAT / one-sided endpoints
|
||||
|
||||
If one host is behind NAT and can't accept incoming connections, only
|
||||
that host needs a null `wireguardEndpoint` on the peer config — the
|
||||
other side initiates. With keepalive on, the NAT hole stays open.
|
||||
|
||||
If both hosts are behind NAT, you need a STUN relay or a third host
|
||||
(exit node). Out of scope for v0.
|
||||
|
||||
## Snapshot store
|
||||
|
||||
One further option lives in this namespace but its docs live with the
|
||||
service it points at: `services.hyperhive.swarm.snapshotStore.{address,
|
||||
port}` tells this hive where the swarm's `btrfs receive` endpoint is, so
|
||||
`hivectl agent <name> subvol snapshot push` has somewhere to stream to.
|
||||
|
||||
It's genuinely swarm-scoped rather than per-peer — a swarm has exactly
|
||||
one store, because the receiver keys destinations by _agent_ so a
|
||||
migrating agent keeps one unbroken incremental chain. See
|
||||
[snapshot-store.md](../networking/snapshot-store.md).
|
||||
- **control plane** — `swarm-controller`: hive directory, agent roster, job
|
||||
graph, agent creation. → [below](#swarm-controller)
|
||||
- **swarm UI** — the operator's day-to-day surface, on the swarm apex,
|
||||
`admins` only. → [`ui.md`](ui.md)
|
||||
- **shared services** — one forge, homeserver, SSO, queue and metrics/logs
|
||||
stack, each on whichever host you put it. → [`services.md`](services.md)
|
||||
- **SSO** — the secrets authelia generates and how each reaches its reader.
|
||||
→ [`sso.md`](sso.md)
|
||||
- **secrets** — every credential the swarm holds, who mints it and where it
|
||||
lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal
|
||||
contract → [`credentials.md`](credentials.md) · how the store comes up,
|
||||
who writes its grants and how it unseals → [`bao.md`](bao.md)
|
||||
- **swarm CA** — the root every hive's internal TLS chains to, and what to
|
||||
hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md)
|
||||
- **snapshot store** — the swarm's one `btrfs receive` endpoint,
|
||||
`swarm.snapshotStore.{address,port}`. →
|
||||
[`snapshot-store.md`](../networking/snapshot-store.md)
|
||||
|
||||
## Swarm controller
|
||||
|
||||
`services.hyperhive.deploy.swarm-controller.enable` runs the `swarm-controller`
|
||||
daemon on this host. **Off by default and deliberately not derived from
|
||||
`services.hyperhive.deploy.hive-controller.enable`**: a swarm has one
|
||||
`swarm-controller` is the swarm's control plane: it holds the hive directory,
|
||||
the agent roster and their wanted state, and the job graph. Creating an agent —
|
||||
from the swarm UI or `swarmctl agent create --hive <h>` — queues its SSO
|
||||
identity, forge user, config repo, store identity and matrix account, then
|
||||
sends the hive a deploy message. A new agent starts `paused`.
|
||||
|
||||
`services.hyperhive.deploy.swarm-controller.enable` runs it on this host;
|
||||
`singleHostSwarm` turns it on. **Otherwise off by default and deliberately
|
||||
not derived from `deploy.hive-controller.enable`**: a swarm has one
|
||||
controller, so enabling it states a fact about swarm topology, not about
|
||||
whether this host runs a hive. Every hive runs `hive-c0re` (the agents on that host); one
|
||||
hive additionally runs this (what's true across hives).
|
||||
whether this host runs a hive.
|
||||
|
||||
What it serves, why it's a unix socket rather than a port, and the
|
||||
socket-directory constraint that governs where `socketPath` may point:
|
||||
|
|
@ -409,6 +156,8 @@ how it gets there. A hive lacking the queue's address for its
|
|||
agents sets none of the four and each agent logs that it has none; a half-set
|
||||
environment logs an error and the harness keeps serving.
|
||||
|
||||
<details><summary>What an agent publishes over the queue</summary>
|
||||
|
||||
What an agent does with that connection is publish its terminal. Every row its
|
||||
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
|
||||
a swarm-level terminal can follow one agent without subscribing to the swarm's
|
||||
|
|
@ -468,6 +217,8 @@ key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
|
|||
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
|
||||
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
|
||||
|
||||
</details>
|
||||
|
||||
### Swarm-wide forge objects
|
||||
|
||||
The controller also keeps the forge objects that are one per swarm, not
|
||||
|
|
@ -484,7 +235,7 @@ one per hive. It ensures them at start and every five minutes after
|
|||
- the `agent-configs` org avatar
|
||||
(`deploy.swarm-controller.configOrgAvatarPng`).
|
||||
|
||||
hive-c0re no longer creates any of them. A pass that can't finish logs a
|
||||
A pass that can't finish logs a
|
||||
`warn` line per object plus `swarm forge objects: pass incomplete` in
|
||||
`journalctl -u swarm-controller`, and retries on the next tick. While the
|
||||
controller is down the objects stay as they are.
|
||||
|
|
@ -503,14 +254,14 @@ re-derive. Approval happens once, at the swarm level: a hive receives a
|
|||
decision, not an event to adjudicate.
|
||||
|
||||
**`internal/knowledge` is on that path.** The controller's is the only
|
||||
hook on it: hives no longer register their own (see
|
||||
`docs/integrations/knowledge.md` for clearing a leftover). A webhook has exactly one target URL, so per-hive
|
||||
registration never added a recipient — it took delivery away from
|
||||
whichever hive registered before it.
|
||||
hook on it; hives register none of their own
|
||||
([`knowledge.md`](../integrations/knowledge.md) covers clearing a leftover).
|
||||
A webhook has exactly one target URL, so a second registration would take
|
||||
delivery away from the first rather than add a recipient.
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
**The `agent-configs` org isn't yet.** Each hive still registers its own
|
||||
**The `agent-configs` org isn't.** Each hive registers its own
|
||||
`pull_request` hook there, so that repo has two — the hive's and the
|
||||
controller's — and **both are expected; don't delete either.** Removing
|
||||
a hive's stops it acting on config PRs; removing the controller's just
|
||||
|
|
@ -529,15 +280,206 @@ To check it's working, push to `internal/knowledge` and look for
|
|||
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
|
||||
refused delivery logs `webhook: refused delivery` with the reason.
|
||||
|
||||
## Hives: the substrate
|
||||
|
||||
- **hive** — one host running `hive-c0re` and its agent containers
|
||||
(`deploy.hive-controller.enable`). Addressed as `<hiveName>.<swarm.domain>`.
|
||||
- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the
|
||||
shared services and controller. You can qualify an agent as
|
||||
`agent@hive-domain`.
|
||||
- **peer hive** — any hive in the directory other than this one. Peers are
|
||||
_derived_, not declared: the directory lists every hive including
|
||||
yourself, and `hiveName` says which one you are.
|
||||
|
||||
### Hive identity config
|
||||
|
||||
```nix
|
||||
services.hyperhive = {
|
||||
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
||||
hiveName = "pr1ma"; # required — this hive's label in it
|
||||
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
||||
|
||||
# required — the directory, identical on every host in the swarm.
|
||||
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
||||
swarm.hives = {
|
||||
pr1ma = { }; # this host, per hiveName
|
||||
lab = { }; # a second hive
|
||||
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
||||
};
|
||||
};
|
||||
```
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
`swarm.domain` is **required** on a host that runs a hive, and `hiveName` on
|
||||
a host that runs a hive, the secret store or the homeserver; eval fails with
|
||||
a hint naming each. Neither defaults, because a guessed value here is a
|
||||
wrong hostname that evaluates cleanly and deploys.
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
**The directory is one attrset, identical on every host.** It describes
|
||||
every hive in the swarm, **including this one**, keyed by `hiveName`; what
|
||||
differs between hosts is only `hiveName`. It **must** contain an entry for
|
||||
this host's `hiveName`, and an empty directory fails that check too — eval
|
||||
names the missing hive. Peers are _every entry but this host's_, so a
|
||||
directory without this host would make every hive a peer, itself included.
|
||||
|
||||
```
|
||||
# hive A # hive B
|
||||
hiveName = "pr1ma"; hiveName = "edge";
|
||||
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
||||
```
|
||||
|
||||
One entry per hive means two hosts can't hold _different_ facts about the
|
||||
same third hive, such as a stale endpoint.
|
||||
|
||||
`domain` defaults to `<name>.<swarm.domain>`, so a conventional directory
|
||||
is names only. Set it only for a hive addressed by something else. This
|
||||
hive's own `services.hyperhive.domain` comes from its entry; it drives
|
||||
`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified
|
||||
labels (`iris@pr1ma.example.com`).
|
||||
|
||||
`swarm.name` is display only — the dashboard chrome header and per-agent
|
||||
system prompts — and federated hives at different domains can share one.
|
||||
`hiveName` surfaces in the same places but is also the leftmost label of the
|
||||
hive's domain. `swarm.name` names the group; `hiveName` names this hive.
|
||||
|
||||
The env-var chain and `qualify()` / `qualified_label()` semantics:
|
||||
[`conventions.md`](../process/conventions.md) § Hive identity.
|
||||
|
||||
Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive
|
||||
chains to it, so there is no per-hive CA field and no per-hive cert pinning.
|
||||
Trusting a hive whose root this swarm doesn't own — another swarm's — has no
|
||||
mechanism.
|
||||
|
||||
<details><summary>Upgrading an existing hive</summary>
|
||||
|
||||
- Set `swarm.domain` and `hiveName` once; eval fails naming each until you do.
|
||||
- Add the hive's own directory entry,
|
||||
`services.hyperhive.swarm.hives.<hiveName> = { };` — one line, no value.
|
||||
Eval fails naming it if you forget.
|
||||
- Setting `services.hyperhive.domain` directly still works and still wins,
|
||||
with a **deprecation warning**: the option is local to one host while
|
||||
every host holds the directory, so a value written only there leaves every
|
||||
peer pointing somewhere else. Move it into the hive's directory entry, or
|
||||
drop it if it's the conventional `<hiveName>.<swarm.domain>`.
|
||||
- Setting `swarm.hives.<name>.certFingerprint` fails eval; the field no
|
||||
longer exists. Trust comes from the swarm root instead.
|
||||
|
||||
</details>
|
||||
|
||||
### What the directory feeds
|
||||
|
||||
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
||||
directory and serves it at `GET /api/hives`; the swarm UI's front page
|
||||
renders it ([`ui.md`](ui.md)).
|
||||
|
||||
2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel
|
||||
federates with the peer's matrix server (discovered via the peer's
|
||||
`.well-known/matrix/server` delegation, which the gateway serves).
|
||||
Federation validates the peer's TLS certificate against the matrix
|
||||
**container's** trust bundle, independent of this directory.
|
||||
|
||||
⚠️ On a hive with self-signed gateway certificates, the container
|
||||
also trusts this hive's `trust-bundle.pem`, bound in at runtime and
|
||||
added to the public CAs; it ends at the swarm root, so a peer whose
|
||||
certificate chains to the same root validates. A peer outside this
|
||||
swarm's root needs a CA-issued certificate (ACME). Federation firewall
|
||||
and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md).
|
||||
|
||||
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
||||
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
||||
to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional).
|
||||
|
||||
### WireGuard inter-hive mesh (optional)
|
||||
|
||||
The peer config above uses public HTTPS for all inter-hive traffic.
|
||||
For private deployments — or to reduce latency and TLS overhead on
|
||||
intra-swarm traffic — hyperhive can configure a host-to-host
|
||||
WireGuard mesh.
|
||||
|
||||
#### Generating keys
|
||||
|
||||
On each hive host:
|
||||
|
||||
```bash
|
||||
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
||||
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
||||
```
|
||||
|
||||
#### Config example (two hives)
|
||||
|
||||
```nix
|
||||
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.1/24";
|
||||
listenPort = 51820; # optional, default 51820
|
||||
};
|
||||
|
||||
# The same `hives` attrset both hosts hold — mesh fields included,
|
||||
# since "where this hive can be dialled" is a fact about that hive.
|
||||
swarm.hives = {
|
||||
pr1ma = {
|
||||
domain = "pr1ma.example.com";
|
||||
wireguardPublicKey = "base64keyA=";
|
||||
wireguardEndpoint = "198.51.100.1:51820";
|
||||
wireguardAddress = "10.100.0.1/32";
|
||||
};
|
||||
edge = {
|
||||
domain = "edge.corp";
|
||||
wireguardPublicKey = "base64keyB=";
|
||||
wireguardEndpoint = "203.0.113.42:51820";
|
||||
wireguardAddress = "10.100.0.2/32";
|
||||
};
|
||||
};
|
||||
};
|
||||
|
||||
# hive B (edge.corp, mesh IP 10.100.0.2)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.2/24";
|
||||
};
|
||||
|
||||
swarm.hives = { /* … identical to hive A's … */ };
|
||||
};
|
||||
```
|
||||
|
||||
#### What the mesh does
|
||||
|
||||
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
||||
host (not inside agent containers; containers reach peers via the
|
||||
host's routing table).
|
||||
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
||||
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
||||
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
||||
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
||||
rather than the public domain.
|
||||
- It sets `persistentKeepalive = 25` by default; override or null to
|
||||
disable (not needed when both sides have public IPs and no NAT).
|
||||
|
||||
#### NAT / one-sided endpoints
|
||||
|
||||
If one host is behind NAT and can't accept incoming connections, only
|
||||
that host needs a null `wireguardEndpoint` on the peer config — the
|
||||
other side initiates. With keepalive on, the NAT hole stays open.
|
||||
|
||||
If both hosts are behind NAT, you need a STUN relay or a third host
|
||||
(exit node); hyperhive sets up neither.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- `docs/networking/snapshot-store.md` — the swarm's `btrfs receive` endpoint, and
|
||||
the `swarm.snapshotStore` option that points a hive at it
|
||||
- `docs/process/conventions.md` § Hive identity — env vars, qualified labels
|
||||
- `docs/integrations/matrix.md` — matrix federation, TLS cert autogeneration,
|
||||
firewall posture
|
||||
- `docs/swarm/ui.md` — the swarm-wide hive roster, now the operator
|
||||
surface for "what hives exist" (superseded the per-hive dashboard's
|
||||
old "peer hives" display)
|
||||
- `docs/networking/gateway.md` — nginx vhosts and the `.well-known/matrix/`
|
||||
autodiscovery scheme
|
||||
- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist"
|
||||
- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the
|
||||
swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option
|
||||
- [`../process/conventions.md`](../process/conventions.md) § Hive identity —
|
||||
env vars, qualified labels
|
||||
- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix
|
||||
federation, TLS cert autogeneration, firewall posture
|
||||
- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and
|
||||
the `.well-known/matrix/` autodiscovery scheme
|
||||
|
|
|
|||
Loading…
Reference in a new issue