hyperhive/docs/swarm/secrets.md
iris fab2a0dedc docs: fix genuine passive-voice hits in docs/swarm
Read all 94 write-good.Passive hits across docs/swarm/ (ca.md,
README.md, secrets.md, services.md, sso.md, ui.md) in context. 44 are
genuine catches with a nameable, usually already-established actor
(swarm-controller, authelia, swarmctl, the controller, the gateway,
this module, hyperhive itself, or 'the operator' for manual actions) —
rewritten to active. 50 are legitimate passives or false catches, left
alone: predicate-adjective state descriptions (is expected/misconfigured/
broken), negative-capability idioms (no X is needed/placed, can't be
Yed/listed/fetched), config-state conditionals (whenever/when X is
enabled/configured/set), requirement-list labels (is required),
'is tracked as' idiom, backward-looking changelog facts with no actor
(was removed/verified/introduced), ambiguous-actor statements left
conservatively alone (agents are created and destroyed — could be
hive-c0re or swarm-controller, doc doesn't say), and a couple of
deliberately-parallel idiom pairs.

Several sibling-inconsistency fixes: a passive clause sitting next to
an already-active sibling describing the same fact/mechanism (ca.md's
two-bullet consumer list, README's 4-item WireGuard-mesh bullet list,
README's controller-registers-hooks paragraph, sso.md's followed-a-302
sentence).

Verified via vale on the whole directory, diffed against main's exact
baseline (not just the Passive count): write-good.Passive 94 -> 50
exactly, every other category unchanged (1 pre-existing
Microsoft.Contractions error at services... at secrets.md:182,
8 TooWordy, 1 Microsoft.We, 1 Microsoft.FirstPerson — same counts,
same locations).
2026-09-08 15:54:16 +02:00

22 KiB

Swarm secrets: what exists, and where each one lives

A swarm's credentials are generated in three different places and read in a fourth, so "where does this file go" has a different answer per deployment. This page is that answer, one row per secret.

Two rules run through all of it.

Private key material and access tokens are paths, never values. Every option carrying one takes a file path (*File), because a literal written into a nix expression ends up in the nix store — world-readable and permanent. No option in this tree accepts one inline, and adding one would be a leak rather than a convenience.

The rule is about what must stay secret, not about credentials generally. Public material is a value: a certificate, or a public nkey like deploy.nats.calloutUserPublicKey, is published to every client that connects, so the store is a perfectly good place for it.

The generator and the reader typically live in different containers. They share the host's network namespace, which makes them feel co-located, but their filesystem roots are separate. That's why delivery is a host-side copy rather than a bind mount: nixos-container refuses to start when a bind source is missing, and a secret minted on another container's first boot doesn't exist yet. Binding it would make one container wait on a file that waits on a container that starts after it.

Topologies, by who places secrets

Read every row below against one of these. This is a different cut from the deployment shapes --- those say where services run, these say who is responsible for a secret file being there --- the two lists don't line up one-to-one, and neither is a renaming of the other.

topology what it means who places secrets
all-local one host runs the swarm's shared services and its own hive nobody — each secret is generated where it's read, or copied by a host unit
swarm-managed the swarm's services run on a host with swarmctl swarmctl writes what it owns; the rest is still generated in place
hive elsewhere a hive that federates with a swarm it doesn't host the operator provides the file and names it in config

Swarm-level — one of each per swarm

secret generated by lives at hive elsewhere
swarm root CA cert swarm-ca.nix first-boot unit, when autoConfigure is set /var/lib/swarm-ca/root.pem operator copies the cert in; it's public
swarm root CA key same unit /var/lib/swarm-ca/root-key.pem, 0600 stays on whichever host holds it — see the constraint below
swarm-services sub-CA (cert + key) swarm-ca.nix, signed by the root /var/lib/swarm-ca/services-ca{,-key}.pem issued where the root lives
authelia session, JWT and storage-encryption keys authelia's first-boot unit, in-container /var/lib/authelia-swarm/{session,jwt,storage-encryption}.key generated in place; nothing outside that container reads them
authelia OIDC HMAC key same unit /var/lib/authelia-swarm/oidc-hmac.key same
authelia OIDC issuer key (RSA) same unit /var/lib/authelia-swarm/oidc-issuer.key same — relying parties verify against the public half at /jwks.json
OIDC client secret, plaintext half authelia crypto hash generate --random /var/lib/authelia-swarm/oidc-clients/<id>.secret operator provides the file and names it in whichever option reads it — sso.clientSecretFile for a service, otel.clientSecretFile for the hive's telemetry collector
OIDC client secret, digest half the same mint oidc-clients/<id>.digest authelia's own half; merged at runtime via settingsFiles
the swarm collector's copy of its OIDC secret swarm-otel-oidc-secret.service copies it from authelia's tree, when authelia runs on this host /var/lib/swarm-otel-oidc/<id>.secret inside the swarm-otel container operator provides the file and names it in deploy.swarm-otel.clientSecretFile — the collector need not share a host with authelia
authelia subject store swarmctl and swarm-authelia-bridge users.yml — one file, read and written by both swarmctl, on the host that runs authelia
wireguard private key the operatorwg genkey whatever deploy.wireguard.privateKeyFile names always operator-provided; nothing generates this for you
queue auth-callout nkeys (user seed + account seed) swarm-nats-callout-keys first-boot unit, when deploy.nats.autoGenerateCallout is set /var/lib/swarm-nats-callout/{callout-user,issuer}.seed, 0600 operator mints both with nk and names them in deploy.nats.calloutUserSeedFile / deploy.nats.calloutIssuerSeedFile
the secret store's own contents openbao, on first bao operator initan operator action, not a unit inside the swarm-bao container, at its own /var/lib/openbao, kept across rebuilds by ephemeral = false. ⚠️ Not a host path: nixos-container destroy swarm-bao takes the raft data with it, so back up the container's tree, not /var/lib/. Only the store's TLS material (/var/lib/swarm-bao-tls) and its PKCS11 token (/var/lib/swarm-bao-token) are host-level n/a — there is one store; a hive elsewhere is a client of it and holds none of this
the secret store's unseal material the HSM/TPM under deploy.bao.seal = "pkcs11"; openbao itself under "shamir" in the token; or held by whoever ran bao operator init, which is what "shamir" means and why it's stated rather than inferred n/a — only the host running the store seals anything

Authelia mints the three keys for itself, in-container, precisely because nothing outside that container ever reads them. That's the test worth applying to any secret added here — and the client secret's plaintext half is the one row that fails it, which is the entire reason a delivery step exists.

Two telemetry collectors exist, and they land on opposite sides of that test.

The hive's collector needs no delivery step. It authenticates to the swarm's collector as its own hive, and it's a host unit rather than a container, so on an all-local swarm it reads authelia's file where it lies and no second copy is made. On any other topology it's an ordinary "operator provides the file" case — see services.hyperhive.otel.clientSecretFile.

The swarm's collector does need one. It runs in a container, so swarm-otel-oidc-secret.service places its copy, landing at /var/lib/swarm-otel-oidc/<client-id>.secret — the same shape as the forge and homeserver rows below, and for the same reason: the container that mints the secret isn't the container that reads it.

The copy is only made when authelia is enabled on this host and something published is being scraped; otherwise no secret is needed and none is placed.

⚠️ Don't read that delivery unit as the only way this collector is fed. Whether it authenticates follows the credential, never another service's placement: a swarm collector may run on a host that holds neither store and no authelia, and then the secret is an ordinary operator-provided file named in services.hyperhive.deploy.swarm-otel.clientSecretFile — the same shape as the hive collector's row above. The copy unit is the convenience for the co-located case, not the definition of the case.

Minting the queue's callout nkeys

deploy.nats.autoGenerateCallout mints both keypairs on the host before the queue starts. It's on by default only under singleHostSwarm — the one topology where the queue, its responder and the operator are the same person. On every other topology, mint them yourself:

nk -gen user    > callout-user.seed     # the responder's own identity
nk -gen account > issuer.seed           # signs the user JWTs it hands out
nk -inkey callout-user.seed -pubout     # → calloutUserPublicKey
nk -inkey issuer.seed       -pubout     # → calloutIssuerPublicKey

Keep both seeds at 0600 and name them in deploy.nats.calloutUserSeedFile / deploy.nats.calloutIssuerSeedFile. Possession of the issuer seed is the authority to admit anyone to the queue, so it belongs wherever the responder runs and nowhere else.

A hive that sets neither the public keys nor autoGenerateCallout fails at eval, naming the option it wants. That's deliberate: a queue that started without them would accept CONNECT {"user":"auth"} from anyone sharing the host's network namespace, and nothing would look wrong until somebody connected.

All four or none — the seed paths are required too, not just the public keys. They're two halves of the same pair: the server verifies with the public half, the responder signs with the private one. Supplying only the public keys used to pass eval and leave the queue with an auth-callout nobody answers, which refuses every client rather than degrading — and a refusal reaches the client as a timeout, so the symptom is every consumer hanging with nothing logged.

One consequence of the generated path worth knowing before you debug it: with autoGenerateCallout set, the queue assembles its config at boot rather than at build time, so a malformed one surfaces when the container starts instead of when the system builds. The server names the offending file and refuses to run.

Hive-level — one of each per hive

secret generated by lives at
hive CA cert + key hive-tls.nix first-boot unit <deploy.hive-controller.tls.stateDir>/ca.pem, ca-key.pem (0600)
hive leaf certs hive-tls.nix, signed by the hive CA <deploy.hive-controller.tls.stateDir>/<name>.pem
matrix registration token a host activation script, on first boot /var/lib/hyperhive/matrix-register-token (0600)
the forge's copy of its OIDC secret hive-forge-oidc-secret.service copies it from authelia's tree /var/lib/forgejo-oidc/<id>.secret inside the forge container
the homeserver's copy of its OIDC secret hive-matrix-oidc-secret.service, same shape /var/lib/tuwunel-oidc/<id>.secret, handed to tuwunel through LoadCredential

Both delivery units wait for authelia's first boot to mint the secret — a bounded wait, 120s — and then fail loudly rather than skipping. A silent skip produces a service whose login button always fails, which is a symptom many layers from its cause.

The store's first reader is the matrix registration token, and it's worth saying why that one: it's an opaque 32-byte value with no second file and no format. Authelia's OIDC secret needs a .secret and a matching .digest, so starting there would have meant debugging "can a reader authenticate and get bytes back" and "did we write authelia's file format right" at once, with an SSO outage as the failure mode.

glue-matrix-bao-token.nix fetches it and writes the file hive-matrix.nix already reads, so the homeserver never learns the store exists. Every failure path — no such key, sealed store, unreachable store, empty value — leaves the locally minted token in place, so a hive with no store behaves exactly as it did before.

⚠️ Service↔store mTLS is its own trust domain. A credential you must already hold to authenticate can't be fetched from the thing it authenticates you to, so the store's identity can't come from an authority the store distributes — which excludes the hive CA and the swarm CA both, and has nothing to do with the gateway's HTTPS certificates either way. glue-bao-tls.nix mints a CA that signs exactly two things, the store's server certificate and a reader's client certificate, and distributes nothing. A deployment with a real internal CA deletes that file and names its own paths in deploy.bao.serverCertFile / clientCaFile; the store itself has no opinion. A hive that reads from a store on another machine names the reader's half — clientCertFile, clientKeyFile, serverCaFile — and places that leaf by hand. It's the one credential that can't come out of the store, being what opens it; everything else a hive needs does.

The constraint that decides where the root lives

A hive CA carries nameConstraints=permitted;DNS:<hive domain>, and a swarm service name is a sibling of the hive domain rather than a childforge.<swarm> next to <hive>.<swarm>. A hive CA can't issue a certificate for a swarm service. Not by policy: by construction, and openssl enforces it.

Whatever holds the swarm root is therefore what makes swarm-service certificates possible at all. Two things follow:

  • The root's private key is a runtime file and must never enter the nix store, so nothing build-time can name it — security.pki.certificateFiles is read when the system is built, and is the wrong tool here. Trust reaches containers through a bind-mounted bundle assembled at boot instead.
  • On any topology other than all-local, placing that key is an operations decision, not something this module tree makes for you. A hive that hosts no swarm services needs only the root's cert, to trust what others issue.

Adding a secret

State three things, in the row you add above: who mints it, which container reads it, and what happens when they differ. If they differ, it needs a delivery unit, and the unit copies — it doesn't bind.