hyperhive/docs/swarm/secrets.md
atlas aad5d3638f swarm-nats: manual callout needs all four keys, not two
The two callout assertions guarded the halves the server verifies with.
The responder needs the other halves, and nothing related them: a config
supplying only the public keys satisfies both, renders a syntactically
valid auth_callout block, and defines no responder unit.

Callout with no responder is the fail-closed state, so that queue refuses
every client — and a NATS denial arrives as a timeout, so the symptom is
every consumer hanging with nothing logged anywhere.

The build-time config check does run in this case and passes, because the
config is valid; what is missing is a unit, and the absence of a unit is
not an event.
2026-08-24 23:06:59 +02:00

153 lines
9.9 KiB
Markdown

# Swarm secrets: what exists, and where each one lives
A swarm's credentials are generated in three different places and read in a
fourth, so "where does this file go" has a different answer per deployment.
This page is that answer, one row per secret.
Two rules run through all of it.
**Private key material and access tokens are paths, never values.** Every option
carrying one takes a file path (`*File`), because a literal written into a nix
expression is rendered into the nix store — world-readable and permanent. No
option in this tree accepts one inline, and adding one would be a leak rather
than a convenience.
The rule is about what must stay secret, not about credentials generally.
**Public material is a value**: a certificate, or a public nkey like
`swarm.nats.calloutUserPublicKey`, is published to every client that connects,
so the store is a perfectly good place for it.
**The generator and the reader are usually in different containers.** They share
the host's network namespace, which makes them feel co-located, but their
filesystem roots are separate. That is why delivery is a **host-side copy rather
than a bind mount**: `nixos-container` refuses to start when a bind source is
missing, and a secret minted on another container's first boot does not exist
yet. Binding it would make one container wait on a file that waits on a
container that starts after it.
## The three topologies
Every row below is read against one of these.
| topology | what it means | who places secrets |
|---|---|---|
| **all-local** | one host runs the swarm's shared services and its own hive | nobody — each secret is generated where it is read, or copied by a host unit |
| **swarm-managed** | the swarm's services run on a host with `swarmctl` | `swarmctl` writes what it owns; the rest is still generated in place |
| **hive elsewhere** | a hive that federates with a swarm it does not host | the operator provides the file and names it in config |
## Swarm-level — one of each per swarm
| secret | generated by | lives at | hive elsewhere |
|---|---|---|---|
| swarm root CA cert | `swarm-ca.nix` first-boot unit, when `autoConfigure` is set | `/var/lib/swarm-ca/root.pem` | operator copies the **cert** in; it is public |
| swarm root CA key | same unit | `/var/lib/swarm-ca/root-key.pem`, `0600` | stays on whichever host holds it — see the constraint below |
| swarm-services sub-CA (cert + key) | `swarm-ca.nix`, signed by the root | `/var/lib/swarm-ca/services-ca{,-key}.pem` | issued where the root lives |
| authelia session, JWT and storage-encryption keys | authelia's first-boot unit, in-container | `/var/lib/authelia-swarm/{session,jwt,storage-encryption}.key` | generated in place; nothing outside that container reads them |
| authelia OIDC HMAC key | same unit | `/var/lib/authelia-swarm/oidc-hmac.key` | same |
| authelia OIDC issuer key (RSA) | same unit | `/var/lib/authelia-swarm/oidc-issuer.key` | same — relying parties verify against the **public** half at `/jwks.json` |
| OIDC client secret, plaintext half | `authelia crypto hash generate --random` | `/var/lib/authelia-swarm/oidc-clients/<id>.secret` | operator provides the file and names it in whichever option reads it — `sso.clientSecretFile` for a service, `otel.clientSecretFile` for the hive's telemetry collector |
| OIDC client secret, digest half | the same mint | `oidc-clients/<id>.digest` | authelia's own half; merged at runtime via `settingsFiles` |
| the swarm collector's copy of its OIDC secret | `swarm-otel-oidc-secret.service` copies it from authelia's tree | `/var/lib/swarm-otel-oidc/<id>.secret` inside the `swarm-otel` container | n/a — this collector runs on the swarm's service host, beside authelia |
| authelia subject store | `swarmctl` and `swarm-authelia-bridge` | `users.yml` — one file, read and written by both | `swarmctl`, on the host that runs authelia |
| wireguard private key | **the operator**`wg genkey` | whatever `swarm.wireguard.privateKeyFile` names | always operator-provided; nothing generates this for you |
| queue auth-callout nkeys (user seed + account seed) | `swarm-nats-callout-keys` first-boot unit, when `nats.autoGenerateCallout` is set | `/var/lib/swarm-nats-callout/{callout-user,issuer}.seed`, `0600` | operator mints both with `nk` and names them in `nats.calloutUserSeedFile` / `nats.calloutIssuerSeedFile` |
The three keys authelia mints for itself are generated in-container precisely
because nothing outside that container ever reads them. **That is the test worth
applying to any secret added here** — and the client secret's plaintext half is
the one row that fails it, which is the entire reason a delivery step exists.
There are two telemetry collectors and they land on opposite sides of that test.
The **hive's** collector needs no delivery step. It authenticates to the swarm's
collector as its own hive, and it is a host unit rather than a container, so on
an all-local swarm it reads authelia's file where it lies and no second copy is
made. On any other topology it is an ordinary "operator provides the file"
case — see `services.hyperhive.otel.clientSecretFile`.
The **swarm's** collector does need one. It runs in a container, so its copy is
placed by `swarm-otel-oidc-secret.service` and lands at
`/var/lib/swarm-otel-oidc/<client-id>.secret` — the same shape as the forge and
homeserver rows below, and for the same reason: the container that mints the
secret is not the container that reads it.
There is no operator-provided variant of that one, and that is a property of
where it runs rather than an omission: the swarm's collector lives on the host
that runs the swarm's services, which is the host that runs authelia. The copy
is only made when authelia is enabled here and something published is being
scraped; otherwise no secret is needed and none is placed.
### Minting the queue's callout nkeys
`nats.autoGenerateCallout` mints both keypairs on the host before the queue
starts. It is on by default only under `enableAllLocalDefaults` — the one
topology where the queue, its responder and the operator are the same person. On
every other topology, mint them yourself:
```
nk -gen user > callout-user.seed # the responder's own identity
nk -gen account > issuer.seed # signs the user JWTs it hands out
nk -inkey callout-user.seed -pubout # → calloutUserPublicKey
nk -inkey issuer.seed -pubout # → calloutIssuerPublicKey
```
Keep both seeds at `0600` and name them in `calloutUserSeedFile` /
`calloutIssuerSeedFile`. Possession of the **issuer** seed is the authority to
admit anyone to the queue, so it belongs wherever the responder runs and nowhere
else.
A hive that sets neither the public keys nor `autoGenerateCallout` fails at
eval, naming the option it wants. That is deliberate: a queue that started
without them would accept `CONNECT {"user":"auth"}` from anyone sharing the
host's network namespace, and nothing would look wrong until somebody connected.
**All four or none** — the seed paths are required too, not just the public
keys. They are two halves of the same pair: the server verifies with the public
half, the responder signs with the private one. Supplying only the public keys
used to pass eval and leave the queue with an auth-callout nobody answers, which
refuses every client rather than degrading — and a refusal reaches the client as
a timeout, so the symptom is every consumer hanging with nothing logged.
One consequence of the generated path worth knowing before you debug it: with
`autoGenerateCallout` set, the queue's config is assembled at boot rather than at
build time, so a malformed one surfaces when the container starts instead of
when the system builds. The server names the offending file and refuses to run.
## Hive-level — one of each per hive
| secret | generated by | lives at |
|---|---|---|
| hive CA cert + key | `hive-tls.nix` first-boot unit | `<tls.stateDir>/ca.pem`, `ca-key.pem` (`0600`) |
| hive leaf certs | `hive-tls.nix`, signed by the hive CA | `<tls.stateDir>/<name>.pem` |
| matrix registration token | a host activation script, on first boot | `/var/lib/hyperhive/matrix-register-token` (`0600`) |
| the forge's copy of its OIDC secret | `hive-forge-oidc-secret.service` copies it from authelia's tree | `/var/lib/forgejo-oidc/<id>.secret` inside the forge container |
| the homeserver's copy of its OIDC secret | `hive-matrix-oidc-secret.service`, same shape | `/var/lib/tuwunel-oidc/<id>.secret`, handed to tuwunel through `LoadCredential` |
Both delivery units wait for authelia's first boot to mint the secret — a
bounded wait, 120s — and then **fail loudly** rather than skipping. A silent skip
produces a service whose login button always fails, which is a symptom several
layers from its cause.
## The constraint that decides where the root lives
A hive CA carries `nameConstraints=permitted;DNS:<hive domain>`, and **a swarm
service name is a sibling of the hive domain rather than a child** — `forge.<swarm>`
next to `<hive>.<swarm>`. So a hive CA cannot issue a certificate for a swarm
service. Not by policy: by construction, and openssl enforces it.
Whatever holds the swarm root is therefore what makes swarm-service certificates
possible at all. Two things follow:
- **The root's private key is a runtime file and must never enter the nix
store**, so nothing build-time can name it — `security.pki.certificateFiles` is
read when the system is built, and is the wrong tool here. Trust reaches
containers through a bind-mounted bundle assembled at boot instead.
- **On any topology other than all-local, placing that key is an operations
decision**, not something this module tree makes for you. A hive that hosts no
swarm services needs only the root's *cert*, to trust what others issue.
## Adding a secret
State three things, in the row you add above: **who mints it**, **which
container reads it**, and **what happens when they differ**. If they differ, it
needs a delivery unit, and the unit copies — it does not bind.