`services.hyperhive.tls.{stateDir,caValidityDays,leafValidityDays}` sat at
the top of `services.hyperhive`, which is meant to be everything about
hyperhive rather than the settings of one hive. Where the hive CA lives,
how long it lasts and how long the leaves it signs last are decisions of
the host holding the key — `deploy.*`, by the same rule as the switches
that moved before them.
`hive-controller` is hive-c0re's new name (mara on the issue), so the
knobs hang off the daemon that owns the CA rather than off a bare `tls`
at the root. mkRenamedOptionModule entries carry existing configs.
⚠️ Unlike the two switch renames, these names are NOT unique, so this was
swept by ALIAS BINDING rather than by identifier: hive-tls.nix alone holds
two options spelled `stateDir` — its own `cfg.stateDir` and the swarm CA's
`swarmCaCfg.stateDir`, four sites that must not move. Nine files bind an
alias to this config; the rename followed those bindings.
Two sites were invisible to the obvious check, and an unanchored sweep for
`hyperhive\.tls\b` is what found them: the option declaration (`= {` after
the path, so no trailing `.` or `;`) and the alias convention documented in
a comment in lib/hive-ca-trust.nix.
Also renamed the `<tls.stateDir>` shorthand in four docs and two Rust doc
comments, anchored on its delimiters — the new path contains the old one
as a substring, so an unanchored replace would have doubled the prefix.
182 lines
12 KiB
Markdown
182 lines
12 KiB
Markdown
# Swarm secrets: what exists, and where each one lives
|
|
|
|
A swarm's credentials are generated in three different places and read in a
|
|
fourth, so "where does this file go" has a different answer per deployment.
|
|
This page is that answer, one row per secret.
|
|
|
|
Two rules run through all of it.
|
|
|
|
**Private key material and access tokens are paths, never values.** Every option
|
|
carrying one takes a file path (`*File`), because a literal written into a nix
|
|
expression is rendered into the nix store — world-readable and permanent. No
|
|
option in this tree accepts one inline, and adding one would be a leak rather
|
|
than a convenience.
|
|
|
|
The rule is about what must stay secret, not about credentials generally.
|
|
**Public material is a value**: a certificate, or a public nkey like
|
|
`swarm.nats.calloutUserPublicKey`, is published to every client that connects,
|
|
so the store is a perfectly good place for it.
|
|
|
|
**The generator and the reader are usually in different containers.** They share
|
|
the host's network namespace, which makes them feel co-located, but their
|
|
filesystem roots are separate. That is why delivery is a **host-side copy rather
|
|
than a bind mount**: `nixos-container` refuses to start when a bind source is
|
|
missing, and a secret minted on another container's first boot does not exist
|
|
yet. Binding it would make one container wait on a file that waits on a
|
|
container that starts after it.
|
|
|
|
## Topologies, by who places secrets
|
|
|
|
Every row below is read against one of these. This is a different cut
|
|
from the [deployment shapes](services.md#deployment-shapes) --- those
|
|
say *where services run*, these say *who is responsible for a secret
|
|
file being there* --- so the two lists do not line up one-to-one, and
|
|
neither is a renaming of the other.
|
|
|
|
| topology | what it means | who places secrets |
|
|
|---|---|---|
|
|
| **all-local** | one host runs the swarm's shared services and its own hive | nobody — each secret is generated where it is read, or copied by a host unit |
|
|
| **swarm-managed** | the swarm's services run on a host with `swarmctl` | `swarmctl` writes what it owns; the rest is still generated in place |
|
|
| **hive elsewhere** | a hive that federates with a swarm it does not host | the operator provides the file and names it in config |
|
|
|
|
## Swarm-level — one of each per swarm
|
|
|
|
| secret | generated by | lives at | hive elsewhere |
|
|
|---|---|---|---|
|
|
| swarm root CA cert | `swarm-ca.nix` first-boot unit, when `autoConfigure` is set | `/var/lib/swarm-ca/root.pem` | operator copies the **cert** in; it is public |
|
|
| swarm root CA key | same unit | `/var/lib/swarm-ca/root-key.pem`, `0600` | stays on whichever host holds it — see the constraint below |
|
|
| swarm-services sub-CA (cert + key) | `swarm-ca.nix`, signed by the root | `/var/lib/swarm-ca/services-ca{,-key}.pem` | issued where the root lives |
|
|
| authelia session, JWT and storage-encryption keys | authelia's first-boot unit, in-container | `/var/lib/authelia-swarm/{session,jwt,storage-encryption}.key` | generated in place; nothing outside that container reads them |
|
|
| authelia OIDC HMAC key | same unit | `/var/lib/authelia-swarm/oidc-hmac.key` | same |
|
|
| authelia OIDC issuer key (RSA) | same unit | `/var/lib/authelia-swarm/oidc-issuer.key` | same — relying parties verify against the **public** half at `/jwks.json` |
|
|
| OIDC client secret, plaintext half | `authelia crypto hash generate --random` | `/var/lib/authelia-swarm/oidc-clients/<id>.secret` | operator provides the file and names it in whichever option reads it — `sso.clientSecretFile` for a service, `otel.clientSecretFile` for the hive's telemetry collector |
|
|
| OIDC client secret, digest half | the same mint | `oidc-clients/<id>.digest` | authelia's own half; merged at runtime via `settingsFiles` |
|
|
| the swarm collector's copy of its OIDC secret | `swarm-otel-oidc-secret.service` copies it from authelia's tree | `/var/lib/swarm-otel-oidc/<id>.secret` inside the `swarm-otel` container | n/a — this collector runs on the swarm's service host, beside authelia |
|
|
| authelia subject store | `swarmctl` and `swarm-authelia-bridge` | `users.yml` — one file, read and written by both | `swarmctl`, on the host that runs authelia |
|
|
| wireguard private key | **the operator** — `wg genkey` | whatever `swarm.wireguard.privateKeyFile` names | always operator-provided; nothing generates this for you |
|
|
| queue auth-callout nkeys (user seed + account seed) | `swarm-nats-callout-keys` first-boot unit, when `nats.autoGenerateCallout` is set | `/var/lib/swarm-nats-callout/{callout-user,issuer}.seed`, `0600` | operator mints both with `nk` and names them in `nats.calloutUserSeedFile` / `nats.calloutIssuerSeedFile` |
|
|
| the secret store's own contents | openbao, on first `bao operator init` — **an operator action, not a unit** | `/var/lib/swarm-bao` on the host of whoever runs the store, bind-mounted into the `swarm-bao` container | n/a — there is one store; a hive elsewhere is a *client* of it and holds none of this |
|
|
| the secret store's unseal material | the HSM/TPM under `deploy.bao.seal = "pkcs11"`; openbao itself under `"shamir"` | in the token; or held by whoever ran `bao operator init`, which is what `"shamir"` means and why it is stated rather than inferred | n/a — only the host running the store seals anything |
|
|
|
|
The three keys authelia mints for itself are generated in-container precisely
|
|
because nothing outside that container ever reads them. **That is the test worth
|
|
applying to any secret added here** — and the client secret's plaintext half is
|
|
the one row that fails it, which is the entire reason a delivery step exists.
|
|
|
|
There are two telemetry collectors and they land on opposite sides of that test.
|
|
|
|
The **hive's** collector needs no delivery step. It authenticates to the swarm's
|
|
collector as its own hive, and it is a host unit rather than a container, so on
|
|
an all-local swarm it reads authelia's file where it lies and no second copy is
|
|
made. On any other topology it is an ordinary "operator provides the file"
|
|
case — see `services.hyperhive.otel.clientSecretFile`.
|
|
|
|
The **swarm's** collector does need one. It runs in a container, so its copy is
|
|
placed by `swarm-otel-oidc-secret.service` and lands at
|
|
`/var/lib/swarm-otel-oidc/<client-id>.secret` — the same shape as the forge and
|
|
homeserver rows below, and for the same reason: the container that mints the
|
|
secret is not the container that reads it.
|
|
|
|
There is no operator-provided variant of that one, and that is a property of
|
|
where it runs rather than an omission: the swarm's collector lives on the host
|
|
that runs the swarm's services, which is the host that runs authelia. The copy
|
|
is only made when authelia is enabled here and something published is being
|
|
scraped; otherwise no secret is needed and none is placed.
|
|
|
|
### Minting the queue's callout nkeys
|
|
|
|
`nats.autoGenerateCallout` mints both keypairs on the host before the queue
|
|
starts. It is on by default only under `singleHostSwarm` — the one
|
|
topology where the queue, its responder and the operator are the same person. On
|
|
every other topology, mint them yourself:
|
|
|
|
```
|
|
nk -gen user > callout-user.seed # the responder's own identity
|
|
nk -gen account > issuer.seed # signs the user JWTs it hands out
|
|
nk -inkey callout-user.seed -pubout # → calloutUserPublicKey
|
|
nk -inkey issuer.seed -pubout # → calloutIssuerPublicKey
|
|
```
|
|
|
|
Keep both seeds at `0600` and name them in `calloutUserSeedFile` /
|
|
`calloutIssuerSeedFile`. Possession of the **issuer** seed is the authority to
|
|
admit anyone to the queue, so it belongs wherever the responder runs and nowhere
|
|
else.
|
|
|
|
A hive that sets neither the public keys nor `autoGenerateCallout` fails at
|
|
eval, naming the option it wants. That is deliberate: a queue that started
|
|
without them would accept `CONNECT {"user":"auth"}` from anyone sharing the
|
|
host's network namespace, and nothing would look wrong until somebody connected.
|
|
|
|
**All four or none** — the seed paths are required too, not just the public
|
|
keys. They are two halves of the same pair: the server verifies with the public
|
|
half, the responder signs with the private one. Supplying only the public keys
|
|
used to pass eval and leave the queue with an auth-callout nobody answers, which
|
|
refuses every client rather than degrading — and a refusal reaches the client as
|
|
a timeout, so the symptom is every consumer hanging with nothing logged.
|
|
|
|
One consequence of the generated path worth knowing before you debug it: with
|
|
`autoGenerateCallout` set, the queue's config is assembled at boot rather than at
|
|
build time, so a malformed one surfaces when the container starts instead of
|
|
when the system builds. The server names the offending file and refuses to run.
|
|
|
|
## Hive-level — one of each per hive
|
|
|
|
| secret | generated by | lives at |
|
|
|---|---|---|
|
|
| hive CA cert + key | `hive-tls.nix` first-boot unit | `<deploy.hive-controller.tls.stateDir>/ca.pem`, `ca-key.pem` (`0600`) |
|
|
| hive leaf certs | `hive-tls.nix`, signed by the hive CA | `<deploy.hive-controller.tls.stateDir>/<name>.pem` |
|
|
| matrix registration token | a host activation script, on first boot | `/var/lib/hyperhive/matrix-register-token` (`0600`) |
|
|
| the forge's copy of its OIDC secret | `hive-forge-oidc-secret.service` copies it from authelia's tree | `/var/lib/forgejo-oidc/<id>.secret` inside the forge container |
|
|
| the homeserver's copy of its OIDC secret | `hive-matrix-oidc-secret.service`, same shape | `/var/lib/tuwunel-oidc/<id>.secret`, handed to tuwunel through `LoadCredential` |
|
|
|
|
Both delivery units wait for authelia's first boot to mint the secret — a
|
|
bounded wait, 120s — and then **fail loudly** rather than skipping. A silent skip
|
|
produces a service whose login button always fails, which is a symptom several
|
|
layers from its cause.
|
|
|
|
The store's **first reader** is the matrix registration token, and it is worth
|
|
saying why that one: it is an opaque 32-byte value with no second file and no
|
|
format. Authelia's OIDC secret needs a `.secret` *and* a matching `.digest`, so
|
|
starting there would have meant debugging "can a reader authenticate and get
|
|
bytes back" and "did we write authelia's file format right" at once, with an
|
|
SSO outage as the failure mode.
|
|
|
|
`glue-matrix-bao-token.nix` fetches it and writes the file `hive-matrix.nix`
|
|
already reads, so the homeserver never learns the store exists. Every failure
|
|
path — no such key, sealed store, unreachable store, empty value — leaves the
|
|
locally minted token in place, so a hive with no store behaves exactly as it
|
|
did before.
|
|
|
|
⚠️ **Service↔store mTLS is its own trust domain.** A credential you must
|
|
already hold to authenticate cannot be fetched from the thing it authenticates
|
|
you to, so the store's identity cannot come from an authority the store
|
|
distributes — which excludes the hive CA and the swarm CA both, and has nothing
|
|
to do with the gateway's HTTPS certificates either way. `glue-bao-tls.nix`
|
|
mints a CA that signs exactly two things, the store's server certificate and a
|
|
reader's client certificate, and distributes nothing. A deployment with a real
|
|
internal CA deletes that file and names its own paths in
|
|
`deploy.bao.serverCertFile` / `clientCaFile`; the store itself has no opinion.
|
|
|
|
## The constraint that decides where the root lives
|
|
|
|
A hive CA carries `nameConstraints=permitted;DNS:<hive domain>`, and **a swarm
|
|
service name is a sibling of the hive domain rather than a child** — `forge.<swarm>`
|
|
next to `<hive>.<swarm>`. So a hive CA cannot issue a certificate for a swarm
|
|
service. Not by policy: by construction, and openssl enforces it.
|
|
|
|
Whatever holds the swarm root is therefore what makes swarm-service certificates
|
|
possible at all. Two things follow:
|
|
|
|
- **The root's private key is a runtime file and must never enter the nix
|
|
store**, so nothing build-time can name it — `security.pki.certificateFiles` is
|
|
read when the system is built, and is the wrong tool here. Trust reaches
|
|
containers through a bind-mounted bundle assembled at boot instead.
|
|
- **On any topology other than all-local, placing that key is an operations
|
|
decision**, not something this module tree makes for you. A hive that hosts no
|
|
swarm services needs only the root's *cert*, to trust what others issue.
|
|
|
|
## Adding a secret
|
|
|
|
State three things, in the row you add above: **who mints it**, **which
|
|
container reads it**, and **what happens when they differ**. If they differ, it
|
|
needs a delivery unit, and the unit copies — it does not bind.
|