`swarm/sso.md` described the empty user store as a resting state — a
provider that is reachable but has nobody in it yet. It isn't. Authelia
validates the store at startup and treats zero users as fatal:
error reading the authentication database: could not validate the
schema: users: non zero value required
so it exits 1, systemd restarts it, and an enabled-but-unbootstrapped
swarm presents as a crash-looping container behind a vhost that is
working correctly. The observed symptom is `502 Bad Gateway`, which
reads as a proxy fault and is not one.
Says so, gives the error text to grep for, and marks the `swarmctl user
add` step as required to finish the install rather than as a first-login
convenience. `gateway.md` gains the same warning next to the vhost,
because that is where someone lands when the 502 is what they can see.
The reason the store ships empty is unchanged and still right: seeding
an account means a credential in a config file. What was wrong was
calling the resulting state harmless.
6 KiB
Swarm SSO
The swarm runs one authelia, and it is two things at once: the session
provider every protected vhost checks (auth_request), and — once any
client is declared — an OIDC provider issuing tokens to relying
parties like the forge.
The second role is derived rather than switched:
services.hyperhive.swarm.authelia.oidc.clients being non-empty turns it
on. authelia refuses to start with a provider that has no clients, so a
separate enable would be a second fact free to disagree with the first.
Getting in the first time
authelia binds loopback only. The gateway on the host running it
publishes it as auth.<swarm.domain> — vhost, dnsmasq record and TLS
name all follow swarm.authelia.enable, so there is nothing to turn on
separately. (Details, including why a client hive must not declare that
vhost: ../gateway.md.)
Authelia does not start until at least one user exists. The user store is generated empty — deliberately, since seeding a default account would put a credential in a config file — but authelia validates it at startup and treats "no users" as fatal:
error reading the authentication database: could not validate the schema:
users: non zero value required
It then exits 1 and systemd restarts it, so a swarm that has been
enabled but not bootstrapped shows a crash-looping unit and 502 Bad Gateway from the vhost — not a login page with nobody able to use it.
The gateway is working in that state; the upstream is not up.
⚠️ So the step below is required to finish the install, not an optional first-login convenience. Run it before concluding anything is wrong with the proxy: a 502 here means "no users yet" far more often than it means a routing fault.
Add the first subject on the host running authelia:
# swarmctl user add mara --display-name Mara --email mara@example.com --group admins
added mara to /var/lib/authelia-swarm/users.yml
password: <generated>
this password is stored nowhere — record it now
The password is generated, hashed, and printed once; only the hash is
kept. swarmctl writes its canonical users.json, re-renders authelia's
users.yml from it, and restarts authelia. Full reference:
../tools/swarmctl-cli.md.
This step stays manual on purpose. Bootstrapping an identity provider non-interactively means a secret arriving from somewhere — a file, an env var, a nix expression — and every one of those is worse than an operator typing one command once.
What secrets exist, and where each one lives
| secret | generated by | rests in | read by |
|---|---|---|---|
jwt.key, session.key, storage-encryption.key |
authelia's first-boot unit | /var/lib/authelia-swarm/ |
authelia |
oidc-hmac.key |
same unit | same directory | authelia |
oidc-issuer.key (RSA) |
same unit | same directory | authelia signs with it; clients verify the public half at /jwks.json |
oidc-clients/<id>.digest |
same unit, via authelia crypto hash generate |
same directory, merged in through settingsFiles |
authelia |
oidc-clients/<id>.secret |
the same mint — this is its plaintext half | same directory | the relying party, in another container |
Everything above the last row is generated in-container because nothing outside that container ever reads it. That is the test worth applying to any secret added here. The last row fails it, and that is the entire reason a delivery step exists.
None of it is ever written into a nix expression. authelia's
settings are rendered into the nix store, which is world-readable and
permanent, so the client digest reaches authelia through settingsFiles
(merged at runtime) and every other secret through a *File option
carrying a path rather than a value.
Getting the plaintext to the relying party
Three cases, and they are genuinely different mechanisms rather than one mechanism with flags.
1. All-local — one host runs both
Nothing to configure beyond swarm.forge.sso.enable = true. A host-side
unit waits for authelia's first boot to mint the secret and copies it
into the forge container, and the forge module contributes its own client
entry — callback URL included — to authelia's client list.
The callback is built from the same source name the registration uses, so the redirect URI authelia is told to allow and the one forgejo actually sends cannot drift apart. A mismatch there is a rejected login with no error text worth reading.
⚠️ The delivery is a copy, not a bindMounts entry, and deliberately so:
nixos-container refuses to start a container whose bind source is
missing, and this secret does not exist until authelia's first boot has
run. Binding it would make the forge wait on a file that waits on a
container that starts after it — on a fresh hive, a permanent stall
presenting as "the forge is broken", several layers from its cause.
2. Swarm-managed services
The controller side owns provisioning: swarmctl writes both halves, the
same way it already owns authelia's user store (users.json canonical,
users.yml a rendered artifact).
3. A hive elsewhere
No shared host, so no automatic path. The operator provides the file and names it:
services.hyperhive.swarm = {
authelia.url = "https://auth.example.com";
forge.sso = {
enable = true;
clientSecretFile = "/var/lib/hyperhive/forge-oidc-secret";
};
};
Both are asserted at eval. A hive that boots with SSO half-configured shows a login button that always fails — a symptom several layers from its cause, and far worse to diagnose than an evaluation error.
What this does not do
- It does not disable local login. The forge keeps its password database and gains a second door. An identity provider that can take the forge offline when it hiccups is a worse forge than one with two ways in.
- It does not provision users. Agents are created and destroyed
continuously, so the subject set belongs to a program rather than to a
config file; today that program is
swarmctl.