The controller-created term-sub-<agent> stream had max_age only, so a publishing agent could grow it without bound for 24h. Add a 64 MiB max_bytes cap with discard: Old (oldest rows drop first, publish never fails on the cap), and size max_message_size off the queue's live max_payload rather than a hardcoded guess.
498 lines
25 KiB
Markdown
498 lines
25 KiB
Markdown
# The swarm
|
|
|
|
The **swarm** is where things live: agent identities and accounts, secrets,
|
|
the job graph that creates and places agents, telemetry, and the UI you
|
|
drive it all from. **Hives are the substrate** — NixOS hosts that run agent
|
|
containers on the swarm's behalf. Every hive belongs to a swarm; a single
|
|
host is a swarm of one.
|
|
|
|
This page is for the operator. It covers the control plane, how a hive
|
|
joins the directory, and how hives and agents report upward. The steps for a
|
|
fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local
|
|
config is the [README quick start](../../README.md#quick-start-an-all-local-swarm).
|
|
|
|
Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical
|
|
on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs
|
|
grafana, the controller, authelia, …) →
|
|
[options reference](https://hyperhive.darkest.space/options/), or
|
|
`nix build .#docs-swarm` / `.#docs-deploy`.
|
|
|
|
## Where each piece lives
|
|
|
|
- **control plane** — `swarm-controller`: hive directory, agent roster, job
|
|
graph, agent creation. → [below](#swarm-controller)
|
|
- **swarm UI** — the operator's day-to-day surface, on the swarm apex,
|
|
`admins` only. → [`ui.md`](ui.md)
|
|
- **shared services** — one forge, homeserver, SSO, queue and metrics/logs
|
|
stack, each on whichever host you put it. → [`services.md`](services.md)
|
|
- **SSO** — the secrets authelia generates and how each reaches its reader.
|
|
→ [`sso.md`](sso.md)
|
|
- **secrets** — every credential the swarm holds, who mints it and where it
|
|
lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal
|
|
contract → [`credentials.md`](credentials.md) · how the store comes up,
|
|
who writes its grants and how it unseals → [`bao.md`](bao.md)
|
|
- **swarm CA** — the root every hive's internal TLS chains to, and what to
|
|
hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md)
|
|
- **snapshot store** — the swarm's one `btrfs receive` endpoint,
|
|
`swarm.snapshotStore.{address,port}`. →
|
|
[`snapshot-store.md`](../networking/snapshot-store.md)
|
|
|
|
## Swarm controller
|
|
|
|
`swarm-controller` is the swarm's control plane: it holds the hive directory,
|
|
the agent roster and their wanted state, and the job graph. Creating an agent —
|
|
from the swarm UI or `swarmctl agent create --hive <h>` — queues its SSO
|
|
identity, forge user, config repo, store identity and matrix account, then
|
|
sends the hive a deploy message. A new agent starts `paused`.
|
|
|
|
`services.hyperhive.deploy.swarm-controller.enable` runs it on this host;
|
|
`singleHostSwarm` turns it on. **Otherwise off: set it on the one host that
|
|
runs the controller.** A swarm has one controller, so enabling it states a
|
|
fact about swarm topology, not about whether this host runs a hive.
|
|
|
|
What it serves, why it's a unix socket rather than a port, and the
|
|
socket-directory constraint that governs where `socketPath` may point:
|
|
[`swarm-controller/README.md`](../../swarm-controller/README.md).
|
|
|
|
### Per-hive status (`GET /api/hives/status`)
|
|
|
|
One row per hive in `swarm.hives`, saying when it last reported and what
|
|
it said. Hives publish upward through the swarm queue; the controller
|
|
never reaches down to collect, so a hive that can't reach the swarm
|
|
still knows its own state — you just can't see it from here.
|
|
|
|
A hive publishes only once it holds all three status-publish coordinates
|
|
below. A hive without them reads
|
|
`never_reported` — it's not broken, it just has nothing to say upward.
|
|
|
|
| freshness | what to do about it |
|
|
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
| `fresh` | nothing — reported within `staleAfterSeconds` |
|
|
| `stale` | the hive stopped reporting. Its last payload is still shown, so check `age_seconds` and the payload for what it managed to say |
|
|
| `never_reported` | this hive has never reported at all — normally a deployment that hasn't happened, not an outage |
|
|
| `unknown` | something is publishing under a name that's not in `swarm.hives` — a typo in the roster, or a hive removed from the roster but still running |
|
|
|
|
Every row also carries `last_seen_unix` and `age_seconds` if you want to
|
|
apply your own threshold. The timestamp is the one the queue recorded on
|
|
arrival, not one the hive put in its own payload.
|
|
|
|
Set `services.hyperhive.swarm.controller.staleAfterSeconds` (default
|
|
`120`) **above the rate hives publish at**, or everything reads `stale`
|
|
between reports. Hives publish once a minute, so the default tolerates
|
|
one missed report and flags two. It takes effect on the next request;
|
|
nothing has to re-publish.
|
|
|
|
### Making a hive report
|
|
|
|
Three options on the **hive**. They sit in two namespaces, because two of them
|
|
are facts about _this machine_ and one is the swarm's single address:
|
|
|
|
| option | what to set it to |
|
|
| ------------------------------------------------------- | --------------------------------------------- |
|
|
| `deploy.hive-controller.statusPublish.natsUrl` | `tls://<swarm.nats.domain>:<swarm.nats.port>` |
|
|
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` |
|
|
| `deploy.hive-controller.statusPublish.clientSecretFile` | path to this hive's client secret, plaintext |
|
|
|
|
The queue URL and the token endpoint default to the swarm's own addresses on
|
|
every hive, so there is nothing to set for them. The secret is what turns
|
|
publishing on: a hive without it doesn't publish. A secret without the other
|
|
two is an eval error rather than a hive that quietly never reports.
|
|
|
|
On a host that runs the queue and the IdP, the secret defaults to the local one.
|
|
Any other hive needs the secret to physically be there, because the swarm
|
|
doesn't distribute it. Copy `hive-<hiveName>.secret` out of the swarm host's
|
|
`deploy.authelia.hostClientSecretDir` with whatever secret management the
|
|
deployment already uses.
|
|
|
|
The queue URL is the same string on every hive. The queue accepts TLS only,
|
|
with a certificate for `swarm.nats.domain` (default `nats.<swarm.domain>`) and
|
|
no other name or address, so a URL with an IP address or `nats://` fails.
|
|
The queue's host resolves the name itself. **A multi-host swarm needs one
|
|
upstream DNS record**, `nats.<swarm.domain>` pointing at the queue host's mesh
|
|
address, the same contract as `bao.<swarm.domain>`. The queue's port is open
|
|
on `wg-hive` when that host is on the mesh.
|
|
|
|
The identity isn't a choice — a hive authenticates as `hive-<hiveName>`
|
|
and publishes under `hiveName`, the same name that keys `swarm.hives`.
|
|
|
|
If a hive stops reporting, its own dashboard is the place to look: a
|
|
failure to publish raises a warning banner there after three consecutive
|
|
misses. It stays `warn` rather than `crit` on purpose — a hive that
|
|
can't reach the queue isn't itself unhealthy, so it doesn't start
|
|
calling itself degraded for being unable to say it's fine.
|
|
|
|
The endpoint answers **503** when this host has no swarm queue
|
|
configured, or has one and can't read it — deliberately not an empty
|
|
list, which would look like a silent swarm rather than a controller that
|
|
can't see. The body says which. Status survives a controller restart:
|
|
it's stored in the queue, not in the daemon.
|
|
|
|
### Giving the agents the queue too
|
|
|
|
An agent authenticates as its own client, not as its hive, so it needs its
|
|
own coordinates. Two of them are options on the hive; the other two arrive
|
|
with the credential itself and aren't configurable.
|
|
|
|
| option | what to set it to |
|
|
| ------------------------------------------- | -------------------------------------------------------------- |
|
|
| `deploy.hive-controller.queue.agentNatsUrl` | where the queue listens, as an agent **container** reaches it |
|
|
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` — the same one the hive uses |
|
|
|
|
On every hive, `agentNatsUrl` defaults to the same
|
|
`tls://<swarm.nats.domain>:<swarm.nats.port>` as the hive's own. On the queue's
|
|
host the name resolves inside a container to the bridge address, where the
|
|
firewall opens the port; on any other hive it resolves through the host's DNS,
|
|
like the store's name. ⚠️ **Never a loopback address here**: an agent has its
|
|
own network namespace, so `127.0.0.1` reaches the agent.
|
|
|
|
The harness sees four variables, and treats them as all-or-none:
|
|
`HIVE_AGENT_NATS_URL` and `HIVE_AGENT_OIDC_TOKEN_ENDPOINT` from the two options
|
|
above, plus `HIVE_AGENT_OIDC_CLIENT_SECRET_FILE` and
|
|
`HIVE_AGENT_OIDC_CLIENT_ID_FILE`, which point into the unit's own credentials
|
|
directory. The last two come from the delivered credential rather than from
|
|
config — see [`secrets.md`](secrets.md#hive-level--one-of-each-per-hive) for
|
|
how it gets there. A hive lacking the queue's address for its
|
|
agents sets none of the four and each agent logs that it has none; a half-set
|
|
environment logs an error and the harness keeps serving.
|
|
|
|
<details><summary>What an agent publishes over the queue</summary>
|
|
|
|
What an agent does with that connection is publish its terminal. Every row its
|
|
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
|
|
a swarm-level terminal can follow one agent without subscribing to the swarm's
|
|
whole traffic. An agent connected with its own queue credential gets that
|
|
subject. An agent without one, or whose own credential the queue
|
|
refused, connects with its hive's shared client and publishes to
|
|
`$SWARM.term.<hive>.<agent>` instead, the `<hive>` being the one that client id
|
|
names. The swarm controller relays both. Publishing only: an
|
|
agent talks about itself here and reads nothing. Rows aren't retained — a
|
|
subscriber that wasn't listening missed them, the same as on the agent's own
|
|
live stream.
|
|
|
|
An agent's **subagents** publish their terminals too, from the agent's
|
|
subagent daemon: each subagent's rows go to `$SWARM.term.<agent>.sub.<subagent>`,
|
|
classified the same way. The queue grants that family to the agent's own queue
|
|
credential alone, so the daemon reads it from the store under the agent's store
|
|
identity, exactly as the harness does, and publishes nothing without it. The
|
|
queue keeps these rows in the stream `term-sub-<agent>` for 24 hours or 64 MiB,
|
|
whichever comes first, dropping the oldest rows once either limit kicks in so
|
|
a publish never fails because of it, and the swarm lists an agent's subagents
|
|
from that stream's subjects. The swarm controller creates that stream for
|
|
every agent a hive's wanted state names, within a minute of the agent
|
|
appearing there. The agent's grant is publish on its own
|
|
`$SWARM.term.<agent>.sub.>` and no `JetStream` subject, so the stream's
|
|
subjects and limits are the controller's and never the agent's. Subagents
|
|
publish output only and read nothing.
|
|
|
|
The queue would refuse a row too large for its `max_payload` outright and
|
|
take the connection down with it, so the harness drops such a row's body before
|
|
sending and leaves a marker in its place; the summary, level and icon still
|
|
arrive. The harness logs and skips a row that's too large even without its body.
|
|
|
|
The second thing an agent publishes is its **turn-state header**, on
|
|
`$SWARM.agent-state.<agent>` (or `$SWARM.agent-state.<hive>.<agent>`) — same shape of subject, same grant
|
|
mechanics, same lack of retention. It carries what a header bar wants: what the
|
|
turn loop is doing (`turn_state`, plus `turn_state_since` as an ISO 8601 UTC
|
|
stamp), which model (`model` and the resolved id the last turn actually ran on),
|
|
the context budget and the last turn's context and cost token blocks, and
|
|
`agent_state`.
|
|
|
|
`agent_state` reuses the swarm's own wanted-state vocabulary
|
|
(`up`/`offline`/`paused`/`destroyed`) so a reader can compare what an agent _is_
|
|
against what the swarm declared it should be without translating between two
|
|
spellings. ⚠️ From inside the container only two of those four are sayable: the
|
|
harness reports `up`, or `paused` when the pause marker is present. `offline` and
|
|
`destroyed` are hive-c0re's observations — a stopped agent publishes nothing and
|
|
a destroyed one doesn't exist — so a view that needs the full four-state picture
|
|
takes them from the `agent-status` bucket and uses this subject to sharpen the
|
|
rest.
|
|
|
|
Headers go out **on transition, not on a timer**: the harness rebuilds the
|
|
header whenever its event bus moves and publishes only when the result differs
|
|
from what it last sent. That's the whole point of the subject — the
|
|
`agent-status` bucket already republishes once a minute, which is far too slow
|
|
for "is this agent thinking right now." The cost of a core subject is that a
|
|
subscriber attaching mid-idle sees nothing until the next change, so a renderer
|
|
opens with the bucket's snapshot and lets this stream refine it.
|
|
|
|
Swarm-side, `GET /api/agents/<name>/state/stream` relays that subject as SSE,
|
|
resolving the agent's hive at request time exactly as the terminal stream does.
|
|
The payload passes through opaquely — the controller never parses a header.
|
|
|
|
The third thing an agent publishes is its **icon**, the same SVG its own
|
|
`GET /icon` serves. It goes into the `agent-icons` KV bucket under the key
|
|
`<agent>`, with no hive in it, so the swarm can show the icon of an agent that's
|
|
stopped or has moved hives. Only an agent connected with its own queue
|
|
credential publishes it: the queue grants that credential
|
|
`$KV.agent-icons.<agent>` and no other key, and grants the hive's shared client
|
|
none of the bucket. The harness writes once per start, because a config change
|
|
reaches an agent by restarting its container. An agent with no icon deletes its
|
|
key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
|
|
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
|
|
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
|
|
|
|
</details>
|
|
|
|
### Swarm-wide forge objects
|
|
|
|
The controller also keeps the forge objects that are one per swarm, not
|
|
one per hive. It ensures them at start and every five minutes after
|
|
(`swarm-controller/src/forge/objects.rs`):
|
|
|
|
- the orgs `agent-configs`, `internal` and `agents`, plus each mirror's
|
|
owner org;
|
|
- the empty `operators` merge-gate team in `agents` and `agent-configs`;
|
|
- the `main` merge gate on every `agent-configs` repo: merge and approval
|
|
whitelists = the `operators` team, and no user. The controller leaves a
|
|
repo with no `main` rule alone;
|
|
- the pull-mirrors from `deploy.forgejo.mirrors` on the controller's
|
|
host (with the `actions/checkout` one `deploy.forgejo.ci.enable` adds);
|
|
- `internal/docs` (private) and `internal/knowledge` (public, with a
|
|
seed `README.md` while empty);
|
|
- the `agent-configs` org avatar
|
|
(`deploy.swarm-controller.configOrgAvatarPng`).
|
|
|
|
A pass that can't finish logs a
|
|
`warn` line per object plus `swarm forge objects: pass incomplete` in
|
|
`journalctl -u swarm-controller`, and retries on the next tick. While the
|
|
controller is down the objects stay as they are.
|
|
|
|
### Swarm-wide forge webhooks
|
|
|
|
At startup the controller registers two Forgejo hooks pointing at
|
|
itself — a `push` hook on `internal/knowledge` and a `pull_request` hook
|
|
on the `agent-configs` org, both under
|
|
`https://<swarm.domain>/webhook/forge/`.
|
|
|
|
The controller **interprets** a delivery and sends hives a specific
|
|
message — _the knowledge repo changed_, _deploy agent `foo` at rev
|
|
`abc123`_ — rather than forwarding forge payloads for each hive to
|
|
re-derive. Approval happens once, at the swarm level: a hive receives a
|
|
decision, not an event to adjudicate.
|
|
|
|
A `config-pr` delivery reporting a PR merged into `main` queues a deploy
|
|
of its `merge_commit_sha` on the one hive whose wanted state places the
|
|
agent; with no such hive, or several, the controller deploys nothing. See
|
|
[approvals.md § Config changes](../agent-lifecycle/approvals.md#config-changes).
|
|
A merge whose `closed` delivery never arrives deploys nothing either; the
|
|
operator recovers by redelivering that delivery from the forge hook page,
|
|
which re-enters the same handler and re-queues the deploy.
|
|
|
|
**`internal/knowledge` is on that path.** The controller's is the only
|
|
hook on it ([`knowledge.md`](../integrations/knowledge.md) covers clearing a
|
|
leftover). A webhook has exactly one target URL, so a second registration
|
|
would take delivery away from the first rather than add a recipient.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
**The `agent-configs` org is on it too.** The controller's hook is the
|
|
only one that acts on config PRs. Each hive deletes its own
|
|
`/webhook/config-pr` org hook on startup, since a route no hive serves
|
|
would otherwise sit on the org collecting failed deliveries. Removing
|
|
the controller's hook stops forge-UI merges from deploying until its
|
|
next start recreates it.
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
Nothing to configure. The controller registers the hooks only when this
|
|
host also serves the swarm UI vhost — that's what publishes the
|
|
endpoint, and a hook the forge can't reach would collect failed
|
|
deliveries while looking healthy. The controller generates the HMAC
|
|
secret on first start and keeps it
|
|
(see [`docs/agent-lifecycle/persistence.md`](../agent-lifecycle/persistence.md)).
|
|
|
|
To check it's working, push to `internal/knowledge` and look for
|
|
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
|
|
refused delivery logs `webhook: refused delivery` with the reason.
|
|
|
|
## Hives: the substrate
|
|
|
|
- **hive** — one host running `hive-c0re` and its agent containers
|
|
(`deploy.hive-controller.enable`). Addressed as `<hiveName>.<swarm.domain>`.
|
|
- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the
|
|
shared services and controller. You can qualify an agent as
|
|
`agent@hive-domain`.
|
|
- **peer hive** — any hive in the directory other than this one. Peers are
|
|
_derived_, not declared: the directory lists every hive including
|
|
yourself, and `hiveName` says which one you are.
|
|
|
|
### Hive identity config
|
|
|
|
```nix
|
|
services.hyperhive = {
|
|
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
|
hiveName = "pr1ma"; # required — this hive's label in it
|
|
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
|
|
|
# required — the directory, identical on every host in the swarm.
|
|
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
|
swarm.hives = {
|
|
pr1ma = { }; # this host, per hiveName
|
|
lab = { }; # a second hive
|
|
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
|
};
|
|
};
|
|
```
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
Swarm options are identical on every host in the swarm, so `swarm.domain` is
|
|
**required** on every host that runs any hyperhive service; a host that
|
|
imports the module and enables nothing evaluates without it. `hiveName` is
|
|
required on a host that runs a hive, the secret store or the homeserver. Eval fails with a hint naming
|
|
each. Neither defaults, because a guessed value here is a wrong hostname
|
|
that evaluates cleanly and deploys.
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
**The directory is one attrset, identical on every host.** It describes
|
|
every hive in the swarm, **including this one**, keyed by `hiveName`; what
|
|
differs between hosts is only `hiveName`. It **must** contain an entry for
|
|
this host's `hiveName`, and an empty directory fails that check too — eval
|
|
names the missing hive. Peers are _every entry but this host's_, so a
|
|
directory without this host would make every hive a peer, itself included.
|
|
|
|
```
|
|
# hive A # hive B
|
|
hiveName = "pr1ma"; hiveName = "edge";
|
|
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
|
```
|
|
|
|
One entry per hive means two hosts can't hold _different_ facts about the
|
|
same third hive, such as a stale endpoint.
|
|
|
|
`domain` defaults to `<name>.<swarm.domain>`, so a conventional directory
|
|
is names only. Set it only for a hive addressed by something else. This
|
|
hive's own `services.hyperhive.domain` comes from its entry; it drives
|
|
`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified
|
|
labels (`iris@pr1ma.example.com`). Setting `services.hyperhive.domain`
|
|
directly overrides the directory entry, with a deprecation warning: the
|
|
option is local to this host while every host shares the directory, so a
|
|
value written only here is invisible to the rest of the swarm.
|
|
|
|
`swarm.name` is display only — the dashboard chrome header and per-agent
|
|
system prompts — and federated hives at different domains can share one.
|
|
`hiveName` surfaces in the same places but is also the leftmost label of the
|
|
hive's domain. `swarm.name` names the group; `hiveName` names this hive.
|
|
|
|
The env-var chain and `qualify()` / `qualified_label()` semantics:
|
|
[`conventions.md`](../process/conventions.md) § Hive identity.
|
|
|
|
Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive
|
|
chains to it. Trusting a hive whose root this swarm doesn't own — another
|
|
swarm's — has no mechanism.
|
|
|
|
### What the directory feeds
|
|
|
|
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
|
directory and serves it at `GET /api/hives`; the swarm UI's front page
|
|
renders it ([`ui.md`](ui.md)).
|
|
|
|
2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel
|
|
federates with the peer's matrix server (discovered via the peer's
|
|
`.well-known/matrix/server` delegation, which the gateway serves).
|
|
Federation validates the peer's TLS certificate against the matrix
|
|
**container's** trust bundle, independent of this directory.
|
|
|
|
⚠️ On a hive with self-signed gateway certificates, the container
|
|
also trusts this hive's `trust-bundle.pem`, bound in at runtime and
|
|
added to the public CAs; it ends at the swarm root, so a peer whose
|
|
certificate chains to the same root validates. A peer outside this
|
|
swarm's root needs a CA-issued certificate (ACME). Federation firewall
|
|
and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md).
|
|
|
|
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
|
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
|
to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional).
|
|
|
|
### WireGuard inter-hive mesh (optional)
|
|
|
|
The peer config above uses public HTTPS for all inter-hive traffic.
|
|
For private deployments — or to reduce latency and TLS overhead on
|
|
intra-swarm traffic — hyperhive can configure a host-to-host
|
|
WireGuard mesh.
|
|
|
|
#### Generating keys
|
|
|
|
On each hive host:
|
|
|
|
```bash
|
|
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
|
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
|
```
|
|
|
|
#### Config example (two hives)
|
|
|
|
```nix
|
|
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
|
services.hyperhive = {
|
|
deploy.wireguard = {
|
|
enable = true;
|
|
privateKeyFile = "/etc/wireguard/hive.key";
|
|
address = "10.100.0.1/24";
|
|
listenPort = 51820; # optional, default 51820
|
|
};
|
|
|
|
# The same `hives` attrset both hosts hold — mesh fields included,
|
|
# since "where this hive can be dialled" is a fact about that hive.
|
|
swarm.hives = {
|
|
pr1ma = {
|
|
domain = "pr1ma.example.com";
|
|
wireguardPublicKey = "base64keyA=";
|
|
wireguardEndpoint = "198.51.100.1:51820";
|
|
wireguardAddress = "10.100.0.1/32";
|
|
};
|
|
edge = {
|
|
domain = "edge.corp";
|
|
wireguardPublicKey = "base64keyB=";
|
|
wireguardEndpoint = "203.0.113.42:51820";
|
|
wireguardAddress = "10.100.0.2/32";
|
|
};
|
|
};
|
|
};
|
|
|
|
# hive B (edge.corp, mesh IP 10.100.0.2)
|
|
services.hyperhive = {
|
|
deploy.wireguard = {
|
|
enable = true;
|
|
privateKeyFile = "/etc/wireguard/hive.key";
|
|
address = "10.100.0.2/24";
|
|
};
|
|
|
|
swarm.hives = { /* … identical to hive A's … */ };
|
|
};
|
|
```
|
|
|
|
#### What the mesh does
|
|
|
|
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
|
host (not inside agent containers; containers reach peers via the
|
|
host's routing table).
|
|
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
|
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
|
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
|
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
|
rather than the public domain.
|
|
- It sets `persistentKeepalive = 25` by default; override or null to
|
|
disable (not needed when both sides have public IPs and no NAT).
|
|
|
|
#### NAT / one-sided endpoints
|
|
|
|
If one host is behind NAT and can't accept incoming connections, only
|
|
that host needs a null `wireguardEndpoint` on the peer config — the
|
|
other side initiates. With keepalive on, the NAT hole stays open.
|
|
|
|
If both hosts are behind NAT, you need a STUN relay or a third host
|
|
(exit node); hyperhive sets up neither.
|
|
|
|
## Cross-references
|
|
|
|
- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist"
|
|
- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the
|
|
swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option
|
|
- [`../process/conventions.md`](../process/conventions.md) § Hive identity —
|
|
env vars, qualified labels
|
|
- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix
|
|
federation, TLS cert autogeneration, firewall posture
|
|
- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and
|
|
the `.well-known/matrix/` autodiscovery scheme
|