docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity and the directory follow as the substrate. Upgrade notes move into a <details> block, the per-agent queue publishing detail into another, and the one-paragraph pointer sections collapse into a link list. Fact fixes, checked against origin/main: - an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean "not in a swarm" - swarm.domain is required with a hive (hive-network.nix:156,188), hiveName with a hive, store or homeserver (hyperhive.nix:161-166) - the matrix container trusts the hive's trust-bundle.pem at runtime under self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85) - singleHostSwarm also defaults the controller, localHostsEntry, the nats callout keys and the bao bootstrap token path (local-defaults.nix:72-129) - swarm-controller serves far more than /health: roster, wanted state, job graph, agent creation and credential mints (main.rs:2874-2899) - swarmctl user add needs --email for the forge account and refuses an existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document agent mint-identity and mint-forge-token - agent creation also mints store identity, forge token and matrix account, and declares the agent paused (main.rs:1822-1920, 247-248) Refs #3902
This commit is contained in:
parent
f688cfcdf0
commit
270430a4b4
5 changed files with 408 additions and 402 deletions
|
|
@ -1,308 +1,55 @@
|
|||
# Multi-hive swarms
|
||||
# The swarm
|
||||
|
||||
A **swarm** is a collection of agents that share an identity and
|
||||
coordinate across one or more hives. A single hyperhive instance
|
||||
running on one host is already a swarm (one hive). This doc covers
|
||||
the additional config needed when the swarm spans multiple hosts.
|
||||
The **swarm** is where things live: agent identities and accounts, secrets,
|
||||
the job graph that creates and places agents, telemetry, and the UI you
|
||||
drive it all from. **Hives are the substrate** — NixOS hosts that run agent
|
||||
containers on the swarm's behalf. Every hive belongs to a swarm; a single
|
||||
host is a swarm of one.
|
||||
|
||||
For the full option reference rather than prose: `services.hyperhive.swarm.*`
|
||||
(swarm-wide facts, identical on every host) and `services.hyperhive.deploy.*`
|
||||
(this host's own deployment decisions — does _this_ machine run grafana,
|
||||
the swarm controller, authelia, …) are separate generated pages, `nix
|
||||
build .#docs-swarm` / `.#docs-deploy` or the website's `/options/swarm.html`
|
||||
/ `/options/deploy.html`.
|
||||
This page is for the operator. It covers the control plane, how a hive
|
||||
joins the directory, and how hives and agents report upward. The steps for a
|
||||
fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local
|
||||
config is the [README quick start](../../README.md#quick-start-an-all-local-swarm).
|
||||
|
||||
## Terminology
|
||||
Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical
|
||||
on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs
|
||||
grafana, the controller, authelia, …) →
|
||||
[options reference](https://hyperhive.darkest.space/options/), or
|
||||
`nix build .#docs-swarm` / `.#docs-deploy`.
|
||||
|
||||
- **hive** — a single hyperhive installation on one host. Has its
|
||||
own `services.hyperhive.domain` DNS name and its own set of agent
|
||||
containers.
|
||||
- **swarm** — one or more hives whose operators have declared them
|
||||
as peers. You can qualify an agent as `agent@hive-domain`.
|
||||
- **peer hive** — any hive in `services.hyperhive.swarm.hives` other
|
||||
than this one. Peers are _derived_, not declared: the directory lists
|
||||
every hive including yourself, and `hiveName` says which one you are.
|
||||
## Where each piece lives
|
||||
|
||||
## Hive identity config
|
||||
|
||||
```nix
|
||||
services.hyperhive = {
|
||||
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
||||
hiveName = "pr1ma"; # required — this hive's label in it
|
||||
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
||||
|
||||
# required — the directory, identical on every host in the swarm.
|
||||
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
||||
swarm.hives = {
|
||||
pr1ma = { };
|
||||
edge = { };
|
||||
};
|
||||
};
|
||||
```
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
`swarm.domain` and `hiveName` are **required** whenever hyperhive is
|
||||
enabled; eval fails with a hint naming each. Neither defaults,
|
||||
because a guessed value here is a wrong hostname that evaluates cleanly
|
||||
and deploys — an eval failure asking the operator to write the address
|
||||
down is the cheaper outcome. **Upgrading past this release means setting
|
||||
both once.**
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
You must still set `domain` too, but you no longer _write_ it: it's read from
|
||||
this hive's own entry in the directory, whose `domain` defaults to
|
||||
`<name>.<swarm.domain>`. A conventional swarm states no addresses at
|
||||
all, and a hive addressed by something else states it in the one place
|
||||
the other hives read — `swarm.hives.edge.domain = "edge.elsewhere.example";`.
|
||||
|
||||
Setting `services.hyperhive.domain` directly still works and still wins,
|
||||
with a **deprecation warning**. The reason it's deprecated isn't tidiness:
|
||||
that option is local to one host, and the operator copies the directory
|
||||
to every host, so a value written only there leaves every peer pointing
|
||||
somewhere else with nothing detecting the disagreement.
|
||||
|
||||
⚠️ **Upgrading:** a hive that has been running on `swarm.domain` +
|
||||
`hiveName` alone now needs its own directory entry —
|
||||
`services.hyperhive.swarm.hives.<hiveName> = { };`, one line, no value.
|
||||
Eval fails naming it if you forget.
|
||||
|
||||
`domain` drives `HYPERHIVE_HIVE_DOMAIN` in every container so agents can
|
||||
form qualified labels (`iris@pr1ma.example.com`).
|
||||
|
||||
`swarm.name` is purely display — it surfaces in the dashboard chrome
|
||||
header and per-agent system prompts, and federated hives at different
|
||||
domains can share one. `hiveName` surfaces in the same places but is
|
||||
_not_ only display: it's the leftmost label of the hive's domain. That
|
||||
`swarm.name` sits under `swarm` and `hiveName` doesn't is the whole
|
||||
distinction — one names this hive, the other names the group it belongs
|
||||
to.
|
||||
|
||||
See `docs/process/conventions.md` § Hive identity for the env-var chain
|
||||
and `qualify()` / `qualified_label()` semantics.
|
||||
|
||||
## Swarm CA
|
||||
|
||||
A hive's internal TLS chains to a **swarm root CA**, so a peer that
|
||||
trusts the root validates every hive in the swarm rather than pinning
|
||||
to each one by hand. Provisioning modes, what to hand a peer
|
||||
(`trust-bundle.pem`, never `ca.pem`), the name constraints on a hive
|
||||
CA, and how an existing hive adopts the hierarchy: [`ca.md`](ca.md).
|
||||
|
||||
## Running the swarm's shared services
|
||||
|
||||
One authelia, one matrix, one forge per swarm — which host runs them,
|
||||
and what a hive that runs none of them configures instead:
|
||||
[`services.md`](services.md).
|
||||
|
||||
## Single sign-on
|
||||
|
||||
Which secrets the SSO provider generates, which one has a reader in
|
||||
another container, and the three ways that one gets delivered:
|
||||
[`sso.md`](sso.md).
|
||||
|
||||
## Secrets
|
||||
|
||||
Every credential the swarm holds, who mints it, where it must live, and
|
||||
which of the three topologies makes it the operator's job to place:
|
||||
[`secrets.md`](secrets.md).
|
||||
|
||||
Where that shape is **going** — the per-secret minter/reader/renewal
|
||||
contract, the target of one mTLS identity per host and everything else
|
||||
through the store, and the test a change has to pass to count as movement
|
||||
toward it: [`credentials.md`](credentials.md). It supersedes `secrets.md`
|
||||
when the migration completes.
|
||||
|
||||
## Swarm UI
|
||||
|
||||
The operator-only web surface on the swarm apex, why reaching it needs
|
||||
the `admins` group rather than just a session, and the four sites you
|
||||
wire a swarm service name into: [`ui.md`](ui.md).
|
||||
|
||||
## The swarm's hive directory
|
||||
|
||||
```nix
|
||||
services.hyperhive.swarm.hives = {
|
||||
pr1ma = { }; # this host, per hiveName
|
||||
lab = { }; # a second hive in the swarm
|
||||
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
||||
};
|
||||
```
|
||||
|
||||
One attrset describing **every** hive in the swarm, **including this
|
||||
one**, keyed by that hive's `hiveName`. It's meant to be _identical on
|
||||
every host_ — write it once, share it, and each host reads it correctly
|
||||
because `services.hyperhive.hiveName` says which entry is itself.
|
||||
|
||||
Empty (the default) means this host isn't in a swarm. Once non-empty it
|
||||
**must** contain an entry for `hiveName`; eval fails naming the missing
|
||||
hive. That assertion is load-bearing rather than pedantic — "my peers"
|
||||
comes from _everything that isn't me_, so a directory that doesn't
|
||||
contain you derives every hive as a peer and you peer with yourself.
|
||||
|
||||
`domain` defaults to `<name>.<swarm.domain>`, the convention every hive
|
||||
follows, so a conventional directory is names only. The default is a
|
||||
derivation from two values the operator already had to state — the swarm's
|
||||
domain and the entry's own name — rather than a guess, which is what makes
|
||||
it safe here when a guessed hostname wouldn't be. Set it only for a hive
|
||||
addressed by something else.
|
||||
|
||||
> **No per-hive CA field exists, and no per-hive cert pinning.** Trust
|
||||
> inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every
|
||||
> hive chains to it, so one anchor replaces per-hive pinning entirely.
|
||||
> What that genuinely drops is trusting a hive whose root this swarm
|
||||
> does _not_ own — another swarm's, or one keeping its own CA. That's
|
||||
> a cross-swarm problem and wants a mechanism designed for it. (An
|
||||
> earlier `certFingerprint` field existed for exactly that gap, pinning
|
||||
> a peer's TLS leaf for hive-c0re's own peer HTTPS checks — removed
|
||||
> along with the dashboard feature it existed to serve, since nothing
|
||||
> else ever consumed it.)
|
||||
|
||||
## What the config does at runtime
|
||||
|
||||
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
||||
directory and serves it at `GET /api/hives`; `swarm-ui`'s overview
|
||||
page renders it (`docs/swarm/ui.md`). This is the operator-facing
|
||||
"what hives exist" surface — a per-hive dashboard "peer hives"
|
||||
display existed here once; it no longer exists, in favour of this.
|
||||
|
||||
2. **Matrix federation** — when `matrix.enable` is on, tuwunel
|
||||
federates with the peer's matrix server (discovered via the peer's
|
||||
`.well-known/matrix/server` delegation, which the gateway serves).
|
||||
Federation validates the peer's TLS certificate against the matrix
|
||||
**container's** trust bundle, independent of this directory.
|
||||
|
||||
⚠️ **That container currently trusts no swarm-internal CA**, so a
|
||||
self-signed gateway certificate doesn't federate. You can't list the
|
||||
swarm root there: `security.pki.certificateFiles` is
|
||||
read when the system is _built_, and the root is a runtime file (its
|
||||
key must never enter the store), so there is no build-time name for
|
||||
it. Bridging that needs a runtime mechanism; a separate issue tracks
|
||||
it. Until then, federation needs CA-issued certs (ACME). See
|
||||
`docs/integrations/matrix.md` for federation firewall + TLS requirements.
|
||||
|
||||
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
||||
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
||||
to configure `wg-hive`. See "WireGuard inter-hive mesh" below.
|
||||
|
||||
## One directory, not a bilateral declaration
|
||||
|
||||
Both hives hold the **same** `hives` attrset; neither declares the
|
||||
other. What differs between the two hosts is only `hiveName`:
|
||||
|
||||
```
|
||||
# hive A # hive B
|
||||
hiveName = "pr1ma"; hiveName = "edge";
|
||||
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
||||
```
|
||||
|
||||
That's the point of the shape, and it removes a class of bug rather
|
||||
than saving typing: a per-host peer list let two hosts hold _different_
|
||||
facts about the same third hive — a stale endpoint, a rotated
|
||||
fingerprint — with nothing to detect the disagreement. One entry per
|
||||
hive makes it unrepresentable.
|
||||
|
||||
## WireGuard inter-hive mesh (optional)
|
||||
|
||||
The peer config above uses public HTTPS for all inter-hive traffic.
|
||||
For private deployments — or to reduce latency and TLS overhead on
|
||||
intra-swarm traffic — hive-c0re can configure a host-to-host
|
||||
WireGuard mesh.
|
||||
|
||||
### Generating keys
|
||||
|
||||
On each hive host:
|
||||
|
||||
```bash
|
||||
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
||||
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
||||
```
|
||||
|
||||
### Config example (two hives)
|
||||
|
||||
```nix
|
||||
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.1/24";
|
||||
listenPort = 51820; # optional, default 51820
|
||||
};
|
||||
|
||||
# The same `hives` attrset both hosts hold — mesh fields included,
|
||||
# since "where this hive can be dialled" is a fact about that hive.
|
||||
swarm.hives = {
|
||||
pr1ma = {
|
||||
domain = "pr1ma.example.com";
|
||||
wireguardPublicKey = "base64keyA=";
|
||||
wireguardEndpoint = "198.51.100.1:51820";
|
||||
wireguardAddress = "10.100.0.1/32";
|
||||
};
|
||||
edge = {
|
||||
domain = "edge.corp";
|
||||
wireguardPublicKey = "base64keyB=";
|
||||
wireguardEndpoint = "203.0.113.42:51820";
|
||||
wireguardAddress = "10.100.0.2/32";
|
||||
};
|
||||
};
|
||||
};
|
||||
|
||||
# hive B (edge.corp, mesh IP 10.100.0.2)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.2/24";
|
||||
};
|
||||
|
||||
swarm.hives = { /* … identical to hive A's … */ };
|
||||
};
|
||||
```
|
||||
|
||||
### What the mesh does
|
||||
|
||||
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
||||
host (not inside agent containers; containers reach peers via the
|
||||
host's routing table).
|
||||
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
||||
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
||||
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
||||
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
||||
rather than the public domain.
|
||||
- It sets `persistentKeepalive = 25` by default; override or null to
|
||||
disable (not needed when both sides have public IPs and no NAT).
|
||||
|
||||
### NAT / one-sided endpoints
|
||||
|
||||
If one host is behind NAT and can't accept incoming connections, only
|
||||
that host needs a null `wireguardEndpoint` on the peer config — the
|
||||
other side initiates. With keepalive on, the NAT hole stays open.
|
||||
|
||||
If both hosts are behind NAT, you need a STUN relay or a third host
|
||||
(exit node). Out of scope for v0.
|
||||
|
||||
## Snapshot store
|
||||
|
||||
One further option lives in this namespace but its docs live with the
|
||||
service it points at: `services.hyperhive.swarm.snapshotStore.{address,
|
||||
port}` tells this hive where the swarm's `btrfs receive` endpoint is, so
|
||||
`hivectl agent <name> subvol snapshot push` has somewhere to stream to.
|
||||
|
||||
It's genuinely swarm-scoped rather than per-peer — a swarm has exactly
|
||||
one store, because the receiver keys destinations by _agent_ so a
|
||||
migrating agent keeps one unbroken incremental chain. See
|
||||
[snapshot-store.md](../networking/snapshot-store.md).
|
||||
- **control plane** — `swarm-controller`: hive directory, agent roster, job
|
||||
graph, agent creation. → [below](#swarm-controller)
|
||||
- **swarm UI** — the operator's day-to-day surface, on the swarm apex,
|
||||
`admins` only. → [`ui.md`](ui.md)
|
||||
- **shared services** — one forge, homeserver, SSO, queue and metrics/logs
|
||||
stack, each on whichever host you put it. → [`services.md`](services.md)
|
||||
- **SSO** — the secrets authelia generates and how each reaches its reader.
|
||||
→ [`sso.md`](sso.md)
|
||||
- **secrets** — every credential the swarm holds, who mints it and where it
|
||||
lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal
|
||||
contract → [`credentials.md`](credentials.md) · how the store comes up,
|
||||
who writes its grants and how it unseals → [`bao.md`](bao.md)
|
||||
- **swarm CA** — the root every hive's internal TLS chains to, and what to
|
||||
hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md)
|
||||
- **snapshot store** — the swarm's one `btrfs receive` endpoint,
|
||||
`swarm.snapshotStore.{address,port}`. →
|
||||
[`snapshot-store.md`](../networking/snapshot-store.md)
|
||||
|
||||
## Swarm controller
|
||||
|
||||
`services.hyperhive.deploy.swarm-controller.enable` runs the `swarm-controller`
|
||||
daemon on this host. **Off by default and deliberately not derived from
|
||||
`services.hyperhive.deploy.hive-controller.enable`**: a swarm has one
|
||||
`swarm-controller` is the swarm's control plane: it holds the hive directory,
|
||||
the agent roster and their wanted state, and the job graph. Creating an agent —
|
||||
from the swarm UI or `swarmctl agent create --hive <h>` — queues its SSO
|
||||
identity, forge user, config repo, store identity and matrix account, then
|
||||
sends the hive a deploy message. A new agent starts `paused`.
|
||||
|
||||
`services.hyperhive.deploy.swarm-controller.enable` runs it on this host;
|
||||
`singleHostSwarm` turns it on. **Otherwise off by default and deliberately
|
||||
not derived from `deploy.hive-controller.enable`**: a swarm has one
|
||||
controller, so enabling it states a fact about swarm topology, not about
|
||||
whether this host runs a hive. Every hive runs `hive-c0re` (the agents on that host); one
|
||||
hive additionally runs this (what's true across hives).
|
||||
whether this host runs a hive.
|
||||
|
||||
What it serves, why it's a unix socket rather than a port, and the
|
||||
socket-directory constraint that governs where `socketPath` may point:
|
||||
|
|
@ -409,6 +156,8 @@ how it gets there. A hive lacking the queue's address for its
|
|||
agents sets none of the four and each agent logs that it has none; a half-set
|
||||
environment logs an error and the harness keeps serving.
|
||||
|
||||
<details><summary>What an agent publishes over the queue</summary>
|
||||
|
||||
What an agent does with that connection is publish its terminal. Every row its
|
||||
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
|
||||
a swarm-level terminal can follow one agent without subscribing to the swarm's
|
||||
|
|
@ -468,6 +217,8 @@ key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
|
|||
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
|
||||
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
|
||||
|
||||
</details>
|
||||
|
||||
### Swarm-wide forge objects
|
||||
|
||||
The controller also keeps the forge objects that are one per swarm, not
|
||||
|
|
@ -484,7 +235,7 @@ one per hive. It ensures them at start and every five minutes after
|
|||
- the `agent-configs` org avatar
|
||||
(`deploy.swarm-controller.configOrgAvatarPng`).
|
||||
|
||||
hive-c0re no longer creates any of them. A pass that can't finish logs a
|
||||
A pass that can't finish logs a
|
||||
`warn` line per object plus `swarm forge objects: pass incomplete` in
|
||||
`journalctl -u swarm-controller`, and retries on the next tick. While the
|
||||
controller is down the objects stay as they are.
|
||||
|
|
@ -503,14 +254,14 @@ re-derive. Approval happens once, at the swarm level: a hive receives a
|
|||
decision, not an event to adjudicate.
|
||||
|
||||
**`internal/knowledge` is on that path.** The controller's is the only
|
||||
hook on it: hives no longer register their own (see
|
||||
`docs/integrations/knowledge.md` for clearing a leftover). A webhook has exactly one target URL, so per-hive
|
||||
registration never added a recipient — it took delivery away from
|
||||
whichever hive registered before it.
|
||||
hook on it; hives register none of their own
|
||||
([`knowledge.md`](../integrations/knowledge.md) covers clearing a leftover).
|
||||
A webhook has exactly one target URL, so a second registration would take
|
||||
delivery away from the first rather than add a recipient.
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
**The `agent-configs` org isn't yet.** Each hive still registers its own
|
||||
**The `agent-configs` org isn't.** Each hive registers its own
|
||||
`pull_request` hook there, so that repo has two — the hive's and the
|
||||
controller's — and **both are expected; don't delete either.** Removing
|
||||
a hive's stops it acting on config PRs; removing the controller's just
|
||||
|
|
@ -529,15 +280,206 @@ To check it's working, push to `internal/knowledge` and look for
|
|||
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
|
||||
refused delivery logs `webhook: refused delivery` with the reason.
|
||||
|
||||
## Hives: the substrate
|
||||
|
||||
- **hive** — one host running `hive-c0re` and its agent containers
|
||||
(`deploy.hive-controller.enable`). Addressed as `<hiveName>.<swarm.domain>`.
|
||||
- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the
|
||||
shared services and controller. You can qualify an agent as
|
||||
`agent@hive-domain`.
|
||||
- **peer hive** — any hive in the directory other than this one. Peers are
|
||||
_derived_, not declared: the directory lists every hive including
|
||||
yourself, and `hiveName` says which one you are.
|
||||
|
||||
### Hive identity config
|
||||
|
||||
```nix
|
||||
services.hyperhive = {
|
||||
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
||||
hiveName = "pr1ma"; # required — this hive's label in it
|
||||
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
||||
|
||||
# required — the directory, identical on every host in the swarm.
|
||||
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
||||
swarm.hives = {
|
||||
pr1ma = { }; # this host, per hiveName
|
||||
lab = { }; # a second hive
|
||||
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
||||
};
|
||||
};
|
||||
```
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
`swarm.domain` is **required** on a host that runs a hive, and `hiveName` on
|
||||
a host that runs a hive, the secret store or the homeserver; eval fails with
|
||||
a hint naming each. Neither defaults, because a guessed value here is a
|
||||
wrong hostname that evaluates cleanly and deploys.
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
**The directory is one attrset, identical on every host.** It describes
|
||||
every hive in the swarm, **including this one**, keyed by `hiveName`; what
|
||||
differs between hosts is only `hiveName`. It **must** contain an entry for
|
||||
this host's `hiveName`, and an empty directory fails that check too — eval
|
||||
names the missing hive. Peers are _every entry but this host's_, so a
|
||||
directory without this host would make every hive a peer, itself included.
|
||||
|
||||
```
|
||||
# hive A # hive B
|
||||
hiveName = "pr1ma"; hiveName = "edge";
|
||||
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
||||
```
|
||||
|
||||
One entry per hive means two hosts can't hold _different_ facts about the
|
||||
same third hive, such as a stale endpoint.
|
||||
|
||||
`domain` defaults to `<name>.<swarm.domain>`, so a conventional directory
|
||||
is names only. Set it only for a hive addressed by something else. This
|
||||
hive's own `services.hyperhive.domain` comes from its entry; it drives
|
||||
`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified
|
||||
labels (`iris@pr1ma.example.com`).
|
||||
|
||||
`swarm.name` is display only — the dashboard chrome header and per-agent
|
||||
system prompts — and federated hives at different domains can share one.
|
||||
`hiveName` surfaces in the same places but is also the leftmost label of the
|
||||
hive's domain. `swarm.name` names the group; `hiveName` names this hive.
|
||||
|
||||
The env-var chain and `qualify()` / `qualified_label()` semantics:
|
||||
[`conventions.md`](../process/conventions.md) § Hive identity.
|
||||
|
||||
Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive
|
||||
chains to it, so there is no per-hive CA field and no per-hive cert pinning.
|
||||
Trusting a hive whose root this swarm doesn't own — another swarm's — has no
|
||||
mechanism.
|
||||
|
||||
<details><summary>Upgrading an existing hive</summary>
|
||||
|
||||
- Set `swarm.domain` and `hiveName` once; eval fails naming each until you do.
|
||||
- Add the hive's own directory entry,
|
||||
`services.hyperhive.swarm.hives.<hiveName> = { };` — one line, no value.
|
||||
Eval fails naming it if you forget.
|
||||
- Setting `services.hyperhive.domain` directly still works and still wins,
|
||||
with a **deprecation warning**: the option is local to one host while
|
||||
every host holds the directory, so a value written only there leaves every
|
||||
peer pointing somewhere else. Move it into the hive's directory entry, or
|
||||
drop it if it's the conventional `<hiveName>.<swarm.domain>`.
|
||||
- Setting `swarm.hives.<name>.certFingerprint` fails eval; the field no
|
||||
longer exists. Trust comes from the swarm root instead.
|
||||
|
||||
</details>
|
||||
|
||||
### What the directory feeds
|
||||
|
||||
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
||||
directory and serves it at `GET /api/hives`; the swarm UI's front page
|
||||
renders it ([`ui.md`](ui.md)).
|
||||
|
||||
2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel
|
||||
federates with the peer's matrix server (discovered via the peer's
|
||||
`.well-known/matrix/server` delegation, which the gateway serves).
|
||||
Federation validates the peer's TLS certificate against the matrix
|
||||
**container's** trust bundle, independent of this directory.
|
||||
|
||||
⚠️ On a hive with self-signed gateway certificates, the container
|
||||
also trusts this hive's `trust-bundle.pem`, bound in at runtime and
|
||||
added to the public CAs; it ends at the swarm root, so a peer whose
|
||||
certificate chains to the same root validates. A peer outside this
|
||||
swarm's root needs a CA-issued certificate (ACME). Federation firewall
|
||||
and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md).
|
||||
|
||||
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
||||
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
||||
to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional).
|
||||
|
||||
### WireGuard inter-hive mesh (optional)
|
||||
|
||||
The peer config above uses public HTTPS for all inter-hive traffic.
|
||||
For private deployments — or to reduce latency and TLS overhead on
|
||||
intra-swarm traffic — hyperhive can configure a host-to-host
|
||||
WireGuard mesh.
|
||||
|
||||
#### Generating keys
|
||||
|
||||
On each hive host:
|
||||
|
||||
```bash
|
||||
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
||||
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
||||
```
|
||||
|
||||
#### Config example (two hives)
|
||||
|
||||
```nix
|
||||
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.1/24";
|
||||
listenPort = 51820; # optional, default 51820
|
||||
};
|
||||
|
||||
# The same `hives` attrset both hosts hold — mesh fields included,
|
||||
# since "where this hive can be dialled" is a fact about that hive.
|
||||
swarm.hives = {
|
||||
pr1ma = {
|
||||
domain = "pr1ma.example.com";
|
||||
wireguardPublicKey = "base64keyA=";
|
||||
wireguardEndpoint = "198.51.100.1:51820";
|
||||
wireguardAddress = "10.100.0.1/32";
|
||||
};
|
||||
edge = {
|
||||
domain = "edge.corp";
|
||||
wireguardPublicKey = "base64keyB=";
|
||||
wireguardEndpoint = "203.0.113.42:51820";
|
||||
wireguardAddress = "10.100.0.2/32";
|
||||
};
|
||||
};
|
||||
};
|
||||
|
||||
# hive B (edge.corp, mesh IP 10.100.0.2)
|
||||
services.hyperhive = {
|
||||
deploy.wireguard = {
|
||||
enable = true;
|
||||
privateKeyFile = "/etc/wireguard/hive.key";
|
||||
address = "10.100.0.2/24";
|
||||
};
|
||||
|
||||
swarm.hives = { /* … identical to hive A's … */ };
|
||||
};
|
||||
```
|
||||
|
||||
#### What the mesh does
|
||||
|
||||
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
||||
host (not inside agent containers; containers reach peers via the
|
||||
host's routing table).
|
||||
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
||||
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
||||
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
||||
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
||||
rather than the public domain.
|
||||
- It sets `persistentKeepalive = 25` by default; override or null to
|
||||
disable (not needed when both sides have public IPs and no NAT).
|
||||
|
||||
#### NAT / one-sided endpoints
|
||||
|
||||
If one host is behind NAT and can't accept incoming connections, only
|
||||
that host needs a null `wireguardEndpoint` on the peer config — the
|
||||
other side initiates. With keepalive on, the NAT hole stays open.
|
||||
|
||||
If both hosts are behind NAT, you need a STUN relay or a third host
|
||||
(exit node); hyperhive sets up neither.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- `docs/networking/snapshot-store.md` — the swarm's `btrfs receive` endpoint, and
|
||||
the `swarm.snapshotStore` option that points a hive at it
|
||||
- `docs/process/conventions.md` § Hive identity — env vars, qualified labels
|
||||
- `docs/integrations/matrix.md` — matrix federation, TLS cert autogeneration,
|
||||
firewall posture
|
||||
- `docs/swarm/ui.md` — the swarm-wide hive roster, now the operator
|
||||
surface for "what hives exist" (superseded the per-hive dashboard's
|
||||
old "peer hives" display)
|
||||
- `docs/networking/gateway.md` — nginx vhosts and the `.well-known/matrix/`
|
||||
autodiscovery scheme
|
||||
- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist"
|
||||
- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the
|
||||
swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option
|
||||
- [`../process/conventions.md`](../process/conventions.md) § Hive identity —
|
||||
env vars, qualified labels
|
||||
- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix
|
||||
federation, TLS cert autogeneration, firewall posture
|
||||
- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and
|
||||
the `.well-known/matrix/` autodiscovery scheme
|
||||
|
|
|
|||
|
|
@ -1,7 +1,12 @@
|
|||
# Swarm-wide services
|
||||
|
||||
Some things exist once per **swarm** rather than once per hive. Two
|
||||
options say where the optional ones live, and everything else derives:
|
||||
Some things exist once per **swarm** rather than once per hive: the forge,
|
||||
the matrix homeserver, SSO, the secret store, the queue, and the metrics and
|
||||
log stack. This page says which host runs them and what a hive that runs none
|
||||
of them configures instead. The all-local quick start sets everything with one
|
||||
line → [README](../../README.md#quick-start-an-all-local-swarm).
|
||||
|
||||
Two options say where they live, and everything else derives:
|
||||
|
||||
```nix
|
||||
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
|
||||
|
|
@ -14,10 +19,15 @@ here" means: every once-per-swarm service takes its `enable` from it.** That's t
|
|||
sections below don't repeat it, so a service that stops deriving is a
|
||||
visible difference rather than one more paragraph saying the same thing.
|
||||
|
||||
`singleHostSwarm` is the all-on-one-box switch above it: it defaults
|
||||
both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host
|
||||
generates the swarm CA here). You can still set each derived toggle on its own,
|
||||
which wins, so "all local except X" needs no further option.
|
||||
`singleHostSwarm` is the all-on-one-box mode above it. It defaults
|
||||
`deploy.allSwarmServices`, the swarm CA (`swarm.ca.autoConfigure`, generated
|
||||
on this host), the swarm controller (`deploy.swarm-controller.enable`), the
|
||||
host's `/etc/hosts` entries for the names it serves
|
||||
(`gateway.localHostsEntry`), the queue's auth-callout keys
|
||||
(`deploy.nats.autoGenerateCallout`) and where the secret store's bootstrap
|
||||
token goes (`deploy.bao.bootstrapTokenFile`). You can still set each derived
|
||||
toggle on its own, which wins, so "all local except X" needs no further
|
||||
option.
|
||||
|
||||
**Both default to off**, and that's deliberate: a host can't tell
|
||||
whether it's meant to be the swarm's service host, so this is an
|
||||
|
|
@ -38,23 +48,29 @@ answers its name from its own resolver, so on a swarm spread over
|
|||
more than one host, the operator's DNS has to resolve those names to that
|
||||
host.
|
||||
|
||||
<details><summary>Moving an existing hive's forge to the swarm's</summary>
|
||||
|
||||
A hive that stops running the forge keeps the old container's state at
|
||||
`/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's
|
||||
forge: push anything worth keeping there by hand. Its
|
||||
`/var/lib/hyperhive/forge-core-token` came from that old forge and
|
||||
fails against the swarm's one.
|
||||
|
||||
</details>
|
||||
|
||||
## Deployment shapes
|
||||
|
||||
Those two options are what makes the difference between deployments, so
|
||||
the shapes worth naming are the ones they produce:
|
||||
|
||||
- **All-local.** Everything on one machine:
|
||||
`singleHostSwarm = true`. Setup is automatic apart from
|
||||
choosing a domain and creating the first user.
|
||||
`singleHostSwarm = true`, plus `deploy.hive-controller.enable = true` for
|
||||
a hive to run agents on. After the first switch, the steps in
|
||||
[`setup.md`](../getting-started/setup.md) remain.
|
||||
- **Services on the swarm controller host.**
|
||||
`deploy.allSwarmServices = true` there; the required services
|
||||
deploy together on that host, with hives elsewhere.
|
||||
Set `deploy.allSwarmServices` and `deploy.swarm-controller.enable` there,
|
||||
with hives elsewhere. The controller doesn't derive from
|
||||
`allSwarmServices`.
|
||||
- **Fully spread out.** One container / VM / machine per service,
|
||||
somewhere.
|
||||
|
||||
|
|
@ -88,12 +104,10 @@ there is one IdP and one auth path.
|
|||
container. Set it explicitly when joining a swarm whose IdP is under
|
||||
another name.
|
||||
|
||||
swarm-controller writes the users database, not by hand: hive-c0re
|
||||
creates and destroys agents continuously, so the subject set is dynamic.
|
||||
This module only guarantees the file exists and parses, so authelia
|
||||
starts with nobody in it rather than failing to start — a provider with
|
||||
no subjects yet is the correct state before anything has provisioned
|
||||
them. Authelia generates session and storage keys in the container on
|
||||
Agent subjects come from swarm-controller's agent-creation job, written
|
||||
into the users database by `swarm-authelia-bridge`; human ones come from
|
||||
`swarmctl user add` → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
|
||||
On first boot this module seeds an empty users database. Authelia generates session and storage keys in the container on
|
||||
first boot and never rotates them automatically; replacing one
|
||||
invalidates data already written (sessions, the encrypted store), so
|
||||
that's an operator action.
|
||||
|
|
@ -196,9 +210,9 @@ gateway either way.
|
|||
|
||||
**Both store exporters are unconditional**, and `deploy.victoriametrics.enable`
|
||||
doesn't gate them: that option says this host _runs_ the store, while the swarm
|
||||
has one either way, reached by its swarm name through the gateway. Gating on it
|
||||
once left a collector on any other host with no exporter at all — receiving from
|
||||
every hive and dropping it, silently, because an absent exporter isn't an error.
|
||||
has one either way, reached by its swarm name through the gateway. A collector
|
||||
with no exporter would receive from every hive and drop it silently, because an
|
||||
absent exporter isn't an error.
|
||||
|
||||
Agent-side configuration, and what a hive's own collector does, are in
|
||||
[`../scheduler/observability.md`](../scheduler/observability.md).
|
||||
|
|
|
|||
|
|
@ -1,10 +1,23 @@
|
|||
# Swarm UI
|
||||
|
||||
The swarm's own web surface, served by the gateway on the **swarm apex**
|
||||
(`services.hyperhive.swarm.domain`) and readable only by operators.
|
||||
The swarm's own web surface and the operator's day-to-day view: served by
|
||||
the gateway on the **swarm apex** (`services.hyperhive.swarm.domain`),
|
||||
readable only by operators. The per-hive dashboard, on each hive's own
|
||||
domain, covers host-level detail for one hive.
|
||||
|
||||
Distinct from the per-hive dashboard, which lives on the hive domain and
|
||||
answers for one host. This one is the view _across_ hives.
|
||||
## What it shows
|
||||
|
||||
| route | what |
|
||||
| ------------------------- | ------------------------------------------------------------------------------------------- |
|
||||
| `/` | the hive directory, each hive with its last reported status |
|
||||
| `/agents` | every agent: status, config PR, wanted state; create agents, link forge and matrix accounts |
|
||||
| `/agents/<name>/terminal` | one agent's live terminal |
|
||||
| `/jobs` | the controller's job graph — where agent creation and credential mints show progress |
|
||||
| `/issues` | a cross-repo issue report |
|
||||
|
||||
Everything it shows comes from [`swarm-controller`](../../swarm-controller/README.md).
|
||||
An agent created here or with `swarmctl agent create` starts `paused`; set it
|
||||
`up` from its card.
|
||||
|
||||
## Enabling
|
||||
|
||||
|
|
@ -35,28 +48,16 @@ requiring `group:admins`. An account without that group authenticates
|
|||
fine and still gets bounced.
|
||||
|
||||
```sh
|
||||
swarmctl user add <you> --group admins
|
||||
swarmctl user add <you> --email <you>@example.com --group admins
|
||||
swarmctl user update <you> --add-group admins # an account that already exists
|
||||
```
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
`--email` isn't needed for the UI, but the forge won't create your account
|
||||
without one → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
|
||||
|
||||
`admins` deliberately, not a new word: [`../getting-started/setup.md`](../getting-started/setup.md) has
|
||||
told every operator to create exactly that group since the bootstrap step
|
||||
existed, so an account made by following the guide already passes. This
|
||||
is the first rule that _consumes_ a group name — inventing a second one
|
||||
would have meant those accounts silently failing a check they were
|
||||
supposed to pass.
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
An account created without any group needs re-adding with the flag —
|
||||
`swarmctl` reads the existing entry out of `users.yml,` so the group is
|
||||
what changes.
|
||||
|
||||
Why a group and not a list of usernames: agents are getting authelia
|
||||
accounts of their own (matrix SSO), and _authenticated_ would then
|
||||
include every agent in the hive. The group is the only thing standing
|
||||
between "an operator's page" and "anyone with a session."
|
||||
Why a group and not "any session": agents are authelia subjects too, so
|
||||
_authenticated_ includes every agent in the swarm. The group is the only
|
||||
thing standing between "an operator's page" and "anyone with a session."
|
||||
|
||||
## What it costs to be reachable
|
||||
|
||||
|
|
@ -66,7 +67,22 @@ not a hole: **reachability isn't the access control here.** An agent
|
|||
that resolves the name and connects still has no operator session, and
|
||||
the subrequest denies it.
|
||||
|
||||
## Two wiring sites
|
||||
## Quick links
|
||||
|
||||
The swarm UI's header carries a single 🔗 button, visible on every route,
|
||||
opening a popover of links to other swarm-wide services. Backed by
|
||||
`GET /api/links` (swarm-controller), which serves
|
||||
`services.hyperhive.swarm.controller.links` (a `listOf { label, icon, url }`,
|
||||
same shape as the per-agent `services.hyperhive.agent.dashboardLinks`).
|
||||
|
||||
Each service's own module contributes its entry when it's enabled on the
|
||||
controller's host — `swarm-authelia.nix`, `hive-matrix.nix`,
|
||||
`hive-forge/default.nix`, `swarm-grafana.nix`, `swarm-victorialogs.nix`, and
|
||||
`swarm-ui.nix` for this UI's own API docs. Adding a link for a new service is
|
||||
a nix-only change to that service's module, or an operator adding an entry
|
||||
directly. An empty list hides the button.
|
||||
|
||||
<details><summary>Adding a swarm service name: the two wiring sites</summary>
|
||||
|
||||
Adding a swarm service name means touching two things. Missing the
|
||||
second ships as a different flavour of "works from the host, broken from
|
||||
|
|
@ -99,23 +115,7 @@ and the apex is a **sibling** of `forge.<swarm>` / `chat.<swarm>` /
|
|||
implicitly. Left out, the vhost falls back to the hive leaf and the
|
||||
swarm's front page opens with a name mismatch.
|
||||
|
||||
## Quick links
|
||||
|
||||
The swarm UI's header carries a single 🔗 button, visible on every route,
|
||||
opening a popover of links to other swarm-wide services — authelia,
|
||||
matrix, forge, this UI's own swagger docs. Backed by `GET /api/links`
|
||||
(swarm-controller), which serves `services.hyperhive.swarm.controller.links`
|
||||
(a `listOf { label, icon, url }`, same shape as the per-agent
|
||||
`services.hyperhive.agent.dashboardLinks`).
|
||||
|
||||
Rather than one central hardcoded list, each service's own module
|
||||
contributes its own entry when it's actually enabled on the controller's
|
||||
host — `swarm-authelia.nix`, `hive-matrix.nix` and `hive-forge/default.nix`
|
||||
all do, the same list-merge idiom `gateway.localNames` uses above. Adding a
|
||||
link for a new service is a nix-only change to that service's own module
|
||||
(or an operator adding an entry directly); no swarm-controller or swarm-ui
|
||||
change needed. Empty list hides the button rather than showing an empty
|
||||
popover.
|
||||
</details>
|
||||
|
||||
## Cross-references
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue