Refs #4374 - otelNoIdentity: name journaldUnits so the fixture trips the store- identity path it's testing instead of swarm-otel's unrelated journald-safety assertion (an empty list with log collection on is refused as "collect everything", not "collect nothing"). - otelNoStores: give it a bao client identity. The secret gate moved from deployCfg.authelia.enable to a real client cert/key pair, so a fixture meaning "no telemetry stores" now needs its own secret identity to keep exercising the exporter/authenticator wiring it was written for. - docs/swarm/secrets.md: two vale fixes — a contraction, and drop a condescending "simply".
302 lines
30 KiB
Markdown
302 lines
30 KiB
Markdown
# Swarm secrets: what exists, and where each one lives
|
|
|
|
A swarm's credentials are generated in three different places and read in a
|
|
fourth, so "where does this file go" has a different answer per deployment.
|
|
This page is that answer, one row per secret.
|
|
|
|
Two rules run through all of it.
|
|
|
|
**Private key material and access tokens are paths, never values.** Every option
|
|
carrying one takes a file path (`*File`), because a literal written into a nix
|
|
expression ends up in the nix store — world-readable and permanent. No
|
|
option in this tree accepts one inline, and adding one would be a leak rather
|
|
than a convenience.
|
|
|
|
The rule is about what must stay secret, not about credentials generally.
|
|
**Public material is a value**: a certificate, or a public nkey like
|
|
`deploy.nats.calloutUserPublicKey`, is published to every client that connects,
|
|
so the store is a perfectly good place for it.
|
|
|
|
**The generator and the reader typically live in different containers.** They share
|
|
the host's network namespace, which makes them feel co-located, but their
|
|
filesystem roots are separate. That's why delivery is a **host-side copy rather
|
|
than a bind mount**: `nixos-container` refuses to start when a bind source is
|
|
missing, and a secret minted on another container's first boot doesn't exist
|
|
yet. Binding it would make one container wait on a file that waits on a
|
|
container that starts after it.
|
|
|
|
## Topologies, by who places secrets
|
|
|
|
Read every row below against one of these. This is a different cut
|
|
from the [deployment shapes](services.md#deployment-shapes) --- those
|
|
say _where services run_, these say _who is responsible for a secret
|
|
file being there_ --- the two lists don't line up one-to-one, and
|
|
neither is a renaming of the other.
|
|
|
|
| topology | what it means | who places secrets |
|
|
| ------------------ | ---------------------------------------------------------- | --------------------------------------------------------------------------- |
|
|
| **all-local** | one host runs the swarm's shared services and its own hive | nobody — each secret is generated where it's read, or copied by a host unit |
|
|
| **swarm-managed** | the swarm's services run on a host with `swarmctl` | `swarmctl` writes what it owns; the rest is still generated in place |
|
|
| **hive elsewhere** | a hive that federates with a swarm it doesn't host | the operator provides the file and names it in config |
|
|
|
|
## Swarm-level — one of each per swarm
|
|
|
|
| secret | generated by | lives at | hive elsewhere |
|
|
| --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
| swarm root CA cert | `swarm-ca.nix` first-boot unit, when `autoConfigure` is set | `/var/lib/swarm-ca/root.pem` | operator copies the **cert** in; it's public |
|
|
| swarm root CA key | same unit | `/var/lib/swarm-ca/root-key.pem`, `0600` | stays on whichever host holds it — see the constraint below |
|
|
| swarm-services sub-CA (cert + key) | `swarm-ca.nix`, signed by the root | `/var/lib/swarm-ca/services-ca{,-key}.pem` | issued where the root lives |
|
|
| authelia session, JWT and storage-encryption keys | authelia's first-boot unit, in-container | `/var/lib/authelia-swarm/{session,jwt,storage-encryption}.key` | generated in place; nothing outside that container reads them |
|
|
| authelia OIDC HMAC key | same unit | `/var/lib/authelia-swarm/oidc-hmac.key` | same |
|
|
| authelia OIDC issuer key (RSA) | same unit | `/var/lib/authelia-swarm/oidc-issuer.key` | same — relying parties verify against the **public** half at `/jwks.json` |
|
|
| OIDC client secret, plaintext half | `authelia crypto hash generate --random` | `/var/lib/authelia-swarm/oidc-clients/<id>.secret` | operator provides the file and names it in whichever option reads it — `sso.clientSecretFile` for a service, `otel.clientSecretFile` for the hive's telemetry collector |
|
|
| OIDC client secret, digest half | the same mint | `oidc-clients/<id>.digest` | authelia's own half; merged at runtime via `settingsFiles` |
|
|
| the swarm collector's copy of its OIDC secret | `swarm-bao-otel-oidc.service` reads it out of the swarm secret store, **on every host that runs the collector and holds a store identity** | `/var/lib/swarm-otel-oidc/<id>.secret` inside the `swarm-otel` container | same unit, same path — one route, co-located or not. A collector with no store identity is left with `clientSecretFile == null`, its already-supported unauthenticated-push degrade — see below |
|
|
| Grafana's copy of its OIDC secret | `swarm-bao-grafana-oidc.service` reads it out of the swarm secret store, **on every host that runs Grafana** | `/var/lib/grafana-oidc/<id>.secret` inside the `swarm-grafana` container | same unit, same path — one route, co-located or not. Nothing for an operator to place beyond this host's store leaf, see below |
|
|
| authelia subject store | `swarmctl` and `swarm-authelia-bridge` | `users.yml` — one file, read and written by both | `swarmctl`, on the host that runs authelia |
|
|
| wireguard private key | **the operator** — `wg genkey` | whatever `deploy.wireguard.privateKeyFile` names | always operator-provided; nothing generates this for you |
|
|
| queue auth-callout nkeys (user seed + account seed) | `swarm-nats-callout-keys` first-boot unit, when `deploy.nats.autoGenerateCallout` is set | `/var/lib/swarm-nats-callout/{callout-user,issuer}.seed`, `0600` | operator mints both with `nk` and names them in `deploy.nats.calloutUserSeedFile` / `deploy.nats.calloutIssuerSeedFile` |
|
|
| the secret store's own contents | openbao, on first `bao operator init` — **an operator action, not a unit** | inside the `swarm-bao` container, at its own `/var/lib/openbao`, kept across rebuilds by `ephemeral = false`. ⚠️ **Not a host path**: `nixos-container destroy swarm-bao` takes the raft data with it, so back up the container's tree, not `/var/lib/`. Only the store's TLS material (`/var/lib/swarm-bao-tls`) and its PKCS11 token (`/var/lib/swarm-bao-token`) are host-level | n/a — there is one store; a hive elsewhere is a _client_ of it and holds none of this |
|
|
| the secret store's unseal material | the HSM/TPM under `deploy.bao.seal = "pkcs11"`; openbao itself under `"shamir"` | in the token; or held by whoever ran `bao operator init`, which is what `"shamir"` means and why it's stated rather than inferred | n/a — only the host running the store seals anything |
|
|
|
|
Authelia mints the three keys for itself, in-container, precisely
|
|
because nothing outside that container ever reads them. **That's the test worth
|
|
applying to any secret added here** — and the client secret's plaintext half is
|
|
the one row that fails it, which is the entire reason a delivery step exists.
|
|
|
|
Two telemetry collectors exist, and they land on opposite sides of that test.
|
|
|
|
The **hive's** collector needs no delivery step. It authenticates to the swarm's
|
|
collector as its own hive, and it's a host unit rather than a container, so on
|
|
an all-local swarm it reads authelia's file where it lies and no second copy is
|
|
made. On any other topology it's an ordinary "operator provides the file"
|
|
case — see `services.hyperhive.otel.clientSecretFile`.
|
|
|
|
The **swarm's** collector does need one, and its delivery route is the fourth
|
|
store reader — the same shape as Grafana's, described in full below. In short:
|
|
`swarm-bao-otel-oidc.service` reads `swarm/services/<client-id>/oidc/client`
|
|
out of the store on every host that runs the collector and holds a store
|
|
identity, whether authelia is a network away or in the container next door.
|
|
The client it reads for is registered unconditionally, the same call
|
|
`glue-grafana-oidc-client.nix` makes for Grafana's: authelia refuses a
|
|
bearer-authz client with no audience, and the push audiences supply one even
|
|
for a collector with nothing published to scrape.
|
|
|
|
⚠️ **A collector with no store identity isn't refused.** Unlike Grafana,
|
|
where an absent OIDC block is a container with no way in at all, a collector
|
|
with nothing to authenticate with still receives every hive's telemetry —
|
|
only its own pushes to the stores go out unauthenticated and get refused
|
|
there. `swarm-bao-otel-oidc.service` therefore doesn't render at all without
|
|
`deploy.bao.clientCertFile` / `clientKeyFile`, the same choice
|
|
`glue-matrix-bao-token.nix` and `glue-queue-agent-credential.nix` make for
|
|
their own optional readers, rather than the hard assertion `swarm-grafana.nix`
|
|
uses for its non-optional one. The secret can still be delivered by hand, at
|
|
`services.hyperhive.deploy.swarm-otel.clientSecretFile`, on a collector this
|
|
unit never reaches.
|
|
|
|
### Minting the queue's callout nkeys
|
|
|
|
`deploy.nats.autoGenerateCallout` mints both keypairs on the host before the queue
|
|
starts. It's on by default only under `singleHostSwarm` — the one
|
|
topology where the queue, its responder and the operator are the same person. On
|
|
every other topology, mint them yourself:
|
|
|
|
```
|
|
nk -gen user > callout-user.seed # the responder's own identity
|
|
nk -gen account > issuer.seed # signs the user JWTs it hands out
|
|
nk -inkey callout-user.seed -pubout # → calloutUserPublicKey
|
|
nk -inkey issuer.seed -pubout # → calloutIssuerPublicKey
|
|
```
|
|
|
|
Keep both seeds at `0600` and name them in `deploy.nats.calloutUserSeedFile` /
|
|
`deploy.nats.calloutIssuerSeedFile`. Possession of the **issuer** seed is the authority to
|
|
admit anyone to the queue, so it belongs wherever the responder runs and nowhere
|
|
else.
|
|
|
|
A hive that sets neither the public keys nor `autoGenerateCallout` fails at
|
|
eval, naming the option it wants. That's deliberate: a queue that started
|
|
without them would accept `CONNECT {"user":"auth"}` from anyone sharing the
|
|
host's network namespace, and nothing would look wrong until somebody connected.
|
|
|
|
**All four or none** — the seed paths are required too, not just the public
|
|
keys. They're two halves of the same pair: the server verifies with the public
|
|
half, the responder signs with the private one. Supplying only the public keys
|
|
used to pass eval and leave the queue with an auth-callout nobody answers, which
|
|
refuses every client rather than degrading — and a refusal reaches the client as
|
|
a timeout, so the symptom is every consumer hanging with nothing logged.
|
|
|
|
One consequence of the generated path worth knowing before you debug it: with
|
|
`autoGenerateCallout` set, the queue assembles its config at boot rather than at
|
|
build time, so a malformed one surfaces when the container starts instead of
|
|
when the system builds. The server names the offending file and refuses to run.
|
|
|
|
## Hive-level — one of each per hive
|
|
|
|
| secret | generated by | lives at |
|
|
| ---------------------------------------- | --------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
|
|
| hive CA cert + key | `hive-tls.nix` first-boot unit | `<deploy.hive-controller.tls.stateDir>/ca.pem`, `ca-key.pem` (`0600`) |
|
|
| hive leaf certs | `hive-tls.nix`, signed by the hive CA | `<deploy.hive-controller.tls.stateDir>/<name>.pem` |
|
|
| matrix registration token | a host activation script, on first boot | `/var/lib/hyperhive/matrix-register-token` (`0600`) |
|
|
| the forge's copy of its OIDC secret | `hive-forge-oidc-secret.service` copies it from authelia's tree | `/var/lib/forgejo-oidc/<id>.secret` inside the forge container |
|
|
| the homeserver's copy of its OIDC secret | `hive-matrix-oidc-secret.service`, same shape | `/var/lib/tuwunel-oidc/<id>.secret`, handed to tuwunel through `LoadCredential` |
|
|
| the agent containers' queue credential | authelia, published to the store by `swarm-secret-publish` | `<deploy.hive-controller.queue.agentCredentialDir>/secret` (`0600`) and `/client_id` (`0644`) |
|
|
|
|
Both delivery units wait for authelia's first boot to mint the secret — a
|
|
bounded wait, 120s — and then **fail loudly** rather than skipping. A silent skip
|
|
produces a service whose login button always fails, which is a symptom many
|
|
layers from its cause.
|
|
|
|
The store's **first reader** is the matrix registration token, and it's worth
|
|
saying why that one: it's an opaque 32-byte value with no second file and no
|
|
format. Authelia's OIDC secret needs a `.secret` _and_ a matching `.digest`, so
|
|
starting there would have meant debugging "can a reader authenticate and get
|
|
bytes back" and "is authelia's file format right" at once, with an
|
|
SSO outage as the failure mode.
|
|
|
|
`glue-matrix-bao-token.nix` fetches it and writes the file `hive-matrix.nix`
|
|
already reads, so the homeserver never learns the store exists. Every failure
|
|
path — no such key, sealed store, unreachable store, empty value — leaves the
|
|
locally minted token in place, so a hive with no store behaves exactly as it
|
|
did before.
|
|
|
|
The **second reader** is the agent containers' queue credential:
|
|
`glue-queue-agent-credential.nix` lands it as two files, the client secret and
|
|
the client id it authenticates, because that's how a queue client reads them —
|
|
the secret by path, the id by value. No local fallback exists, and none is
|
|
possible, so absent files mean this hive's agents don't connect, which is what
|
|
a swarm looks like before the publisher on the authelia host has run. The
|
|
reader runs before `hive-c0re.service` and is wanted (not required) by it, so
|
|
an agent container never renders ahead of the credential; an unreachable store
|
|
delays the daemon's start rather than failing it.
|
|
|
|
That credential still has one hop left, because the reader of it lives inside
|
|
an agent container. **It crosses as a systemd credential, not as a bind
|
|
mount**, and the mode decides that: the secret is `root:0600` and a harness
|
|
runs as its own unprivileged agent user, so a bind would deliver a file that
|
|
user can't open. `hive-c0re` stats the two files — the directory is `0755`, so
|
|
it needs no read access to either — and hands them to
|
|
`systemd-nspawn --load-credential` through `hive-priv`, which runs as root.
|
|
Inside, `hive-agent.service` names the same two ids in `LoadCredential=` and
|
|
reads them out of its own `$CREDENTIALS_DIRECTORY`, owned by the agent user
|
|
and by nobody else. `hive-c0re` never reads the bytes at any point: it runs as
|
|
`hive-core`, which is the reason a copy wasn't an option either. When the
|
|
files aren't there the daemon forwards nothing and says so in its journal, and
|
|
the harness logs that it has no queue coordinates — the same absent-and-legal state, twice,
|
|
rather than a container that refuses to start.
|
|
|
|
The **third reader** is Grafana's OIDC client secret, and it's the first one
|
|
that's a _swarm service's_ own credential rather than a hive's. It has exactly
|
|
one delivery route: `swarm-bao-grafana-oidc.service` reads
|
|
`swarm/services/<client-id>/oidc/client` out of the store, on **every** host
|
|
that runs Grafana. Where authelia is doesn't enter into it — the publisher on
|
|
authelia's host writes that path whether the reader is a network away or in the
|
|
container next door.
|
|
|
|
That's a deliberate refusal of the shortcut. Copying the plaintext
|
|
straight out of authelia's tree when the two happen to share a host is one
|
|
round trip cheaper and costs a second delivery unit, a second way for the file
|
|
to be wrong, and a gate deciding between them — and the gate is the expensive
|
|
part, because whatever it's wrong about is an outage nobody can read. The store
|
|
exists so a host holds **one** out-of-band secret, its client certificate, and
|
|
reads everything else with it.
|
|
|
|
Two things follow, and `swarm-grafana.nix` asserts both rather than degrading:
|
|
running Grafana requires `swarm.authelia.url` (its local login form is disabled
|
|
unconditionally, so SSO isn't a feature of some topologies — it's the only way
|
|
in), and it requires this host's `deploy.bao.clientCertFile` /
|
|
`clientKeyFile`. Each refusal names the option to set. Both used to be silent:
|
|
a null URL dropped the OIDC block, a missing leaf produced a warning, and
|
|
either one left a Grafana with no SSO and no password box — nothing failed, so
|
|
nothing said anything. Registration is separate and stays where authelia is
|
|
(`glue-grafana-oidc-client.nix`): a client is a row in authelia's own config,
|
|
so it's declared on the host running authelia, which isn't necessarily the host
|
|
running Grafana.
|
|
|
|
The **fourth reader** is the swarm collector's own OIDC client secret, the
|
|
same shape one level down: `swarm-bao-otel-oidc.service` reads
|
|
`swarm/services/<client-id>/oidc/client` out of the store on every host that
|
|
runs the collector, and registration lives separately in
|
|
`glue-swarm-otel-oidc-client.nix` for the same reason Grafana's does. Where it
|
|
differs is what an absent credential means: the collector still receives
|
|
telemetry with none, so `swarm-otel.nix` doesn't assert
|
|
`deploy.bao.clientCertFile` / `clientKeyFile` the way Grafana does — it
|
|
doesn't render the reading unit without them, the shape
|
|
`glue-matrix-bao-token.nix` and `glue-queue-agent-credential.nix` use for
|
|
their own optional reads.
|
|
|
|
A service's secret is one value for the whole swarm rather than one per hive, so
|
|
it lives under the `services` prefix, and a hive's read policy grants that prefix
|
|
whole. That's because nothing in a swarm's configuration records which hive runs
|
|
a given service — placement is a `deploy.*` fact, per-host by definition — so
|
|
there is nothing to scope the grant to. Every hive can therefore read every swarm
|
|
service's client secret; that's the cost of a shape where the service's host has
|
|
no store identity of its own to present, and it's stated in
|
|
`swarm-secret-client`'s `policy` module beside the grant itself. Absence behaves
|
|
as it does for the other readers: a store that says "not here" leaves the
|
|
file alone and says so, since there is no such thing as a locally valid OIDC
|
|
client secret to stand in.
|
|
|
|
⚠️ **Service↔store mTLS is its own trust domain.** A credential you must
|
|
already hold to authenticate can't be fetched from the thing it authenticates
|
|
you to, so the store's identity can't come from an authority the store
|
|
distributes — which excludes the hive CA and the swarm CA both, and has nothing
|
|
to do with the gateway's HTTPS certificates either way. `glue-bao-tls.nix`
|
|
mints a CA that signs exactly two things, the store's server certificate and a
|
|
reader's client certificate, and distributes nothing. A deployment with a real
|
|
internal CA deletes that file and names its own paths in
|
|
`deploy.bao.serverCertFile` / `clientCaFile`; the store itself has no opinion.
|
|
A hive that reads from a store on **another** machine names the reader's half —
|
|
`clientCertFile`, `clientKeyFile`, `serverCaFile` — and places that leaf by hand.
|
|
It's the one credential that can't come out of the store, being what opens it;
|
|
everything else a hive needs does.
|
|
|
|
## How a reader reaches the store
|
|
|
|
Every reader dials the same URL — `https://bao.<swarm domain>:<port>` — and on
|
|
the host that runs the store that name resolves twice. `/etc/hosts` answers a
|
|
host-side unit with loopback, where openbao binds; the hive's dnsmasq answers a
|
|
container with the bridge IP, where nginx holds the same port. One `BAO_ADDR`,
|
|
and the network namespace that asked decides which half of it answers.
|
|
|
|
The container's half goes through nginx and **still doesn't terminate TLS**. It
|
|
is an nginx `stream` server with `ssl_preread on`: it reads the SNI off the
|
|
ClientHello and splices the rest of the connection through byte for byte, so
|
|
openbao completes the handshake itself. A vhost would decrypt here instead, and
|
|
the client certificate — the store's whole authentication — would stop at nginx,
|
|
leaving openbao seeing one client for every hive in the swarm. That's why the
|
|
store has no vhost and gets a passthrough: not an exception to the no-vhost
|
|
rule, the one shape that keeps it.
|
|
|
|
The stream listener binds the bridge IP rather than every address, because
|
|
openbao already holds loopback on that port in the same network namespace and a
|
|
wildcard bind would fail with `EADDRINUSE` — taking nginx, and with it every
|
|
hive domain behind the gateway, down. `network.exposeHostPorts` opens the port
|
|
on the bridge firewall and nowhere else.
|
|
|
|
Reaching the port grants nothing by itself: openbao answers nothing without a
|
|
client certificate signed by `deploy.bao.clientCaFile`. The passthrough carries
|
|
whichever certificate the reader presents, unchanged.
|
|
|
|
## The constraint that decides where the root lives
|
|
|
|
A hive CA carries `nameConstraints=permitted;DNS:<hive domain>`, and **a swarm
|
|
service name is a sibling of the hive domain rather than a child** — `forge.<swarm>`
|
|
next to `<hive>.<swarm>`. A hive CA can't issue a certificate for a swarm
|
|
service. Not by policy: by construction, and openssl enforces it.
|
|
|
|
Whatever holds the swarm root is therefore what makes swarm-service certificates
|
|
possible at all. Two things follow:
|
|
|
|
- **The root's private key is a runtime file and must never enter the nix
|
|
store**, so nothing build-time can name it — `security.pki.certificateFiles` is
|
|
read when the system is built, and is the wrong tool here. Trust reaches
|
|
containers through a bind-mounted bundle assembled at boot instead.
|
|
- **On any topology other than all-local, placing that key is an operations
|
|
decision**, not something this module tree makes for you. A hive that hosts no
|
|
swarm services needs only the root's _cert_, to trust what others issue.
|
|
|
|
## Adding a secret
|
|
|
|
State three things, in the row you add above: **who mints it**, **which
|
|
container reads it**, and **what happens when they differ**. If they differ, it
|
|
needs a delivery unit, and the unit copies — it doesn't bind.
|