hyperhive/docs/swarm/sso.md
atlas 9b14014077 docs(sso): document the machine surface, and stop restating it in nix
`docs/swarm/sso.md` described a person in a browser. The swarm's other
callers — the telemetry collector, the queue's auth-callout responder, each
hive's agents — hold no session and follow no redirect, and nothing operator-
facing said how they authenticate. Its relying-party table is forge and
matrix, both browser surfaces.

The new section carries what `swarm-authelia.nix` was holding in comments:
one client per hive because identity belongs to the directory, the audience
being that client id rather than a parallel naming scheme, and signed rather
than opaque tokens because the collector verifies offline against
`/jwks.json` while the queue introspects.

It also states the fail-closed rule once, in the place a reader looks before
touching a vhost: an error page answers 200, and `auth_request` reads any 2xx
as access granted. That shape has now appeared three times — this module's
`/api/` prefix and both of victorialogs' routes — which is what makes it
documentation rather than a comment.

The two comment blocks those replace shrink to the part that is genuinely
local: the submodule-typing reason these clients are a definition rather than
an append, and a loud warning against folding the machine prefix back into
`/`. The security warning stays at the site; only its consequence list moves.

Comments 495 -> 465 lines. Option `description` strings are untouched: they
are the source `pkgs.nixosOptionsDoc` renders into the operator's options
reference, so trimming one would delete published documentation rather than a
duplicate.
2026-09-02 08:57:23 +02:00

252 lines
12 KiB
Markdown

# Swarm SSO
The swarm runs one authelia, and it is two things at once: the **session
provider** every protected vhost checks (`auth_request`), and — once any
client is declared — an **OIDC provider** issuing tokens to relying
parties: the forge and the matrix homeserver.
The second role is derived rather than switched:
`services.hyperhive.swarm.authelia.oidc.clients` being non-empty turns it
on. authelia refuses to start with a provider that has no clients, so a
separate `enable` would be a second fact free to disagree with the first.
## Getting in the first time
authelia binds loopback only. The **gateway** on the host running it
publishes it as `auth.<swarm.domain>` — vhost, dnsmasq record and TLS
name all follow `deploy.authelia`, so there is nothing to turn on
separately. (Details, including why a client hive must not declare that
vhost: [`../networking/gateway.md`](../networking/gateway.md).)
**Authelia does not start until at least one user exists.** The user
store is generated empty — deliberately, since seeding a default account
would put a credential in a config file — but authelia validates it at
startup and treats "no users" as fatal:
```
error reading the authentication database: could not validate the schema:
users: non zero value required
```
It then exits 1 and systemd restarts it, so a swarm that has been
enabled but not bootstrapped shows a **crash-looping unit** and `502 Bad
Gateway` from the vhost — not a login page with nobody able to use it.
The gateway is working in that state; the upstream is not up.
⚠️ So the step below is **required to finish the install**, not an
optional first-login convenience. Run it before concluding anything is
wrong with the proxy: a 502 here means "no users yet" far more often
than it means a routing fault.
Add the first subject on the host running authelia:
```console
# swarmctl user add mara --display-name Mara --email mara@example.com --group admins
added mara to /var/lib/authelia-swarm/users.yml
password: <generated>
this password is stored nowhere — record it now
```
The password is generated, hashed, and printed once; only the hash is
kept. `swarmctl` reads and writes authelia's `users.yml` directly — it is
the one user store, shared with `swarm-authelia-bridge`, which creates
agent identities in the same file. No restart: authelia watches it. Full
reference: [`../tools/swarmctl-cli.md`](../tools/swarmctl-cli.md).
You can edit `users.yml` by hand, and `swarmctl` will read what you
wrote. ⚠️ It rewrites the whole file on every change, so **comments and
formatting do not survive**; values and unrecognised keys do.
This step stays manual on purpose. Bootstrapping an identity provider
non-interactively means a secret arriving from somewhere — a file, an
env var, a nix expression — and every one of those is worse than an
operator typing one command once.
### Changing a subject afterwards
`user add` only ever adds: on a name that already exists it refuses,
rather than resurfacing as a second account or a silent overwrite.
Editing an existing subject is `user update`, and the flags compose, so
one call can change several things:
```console
# swarmctl user update mara --add-group admins --email mara@example.com
added to group "admins"
email: unset -> "mara@example.com"
mara is now in groups: admins
```
Two behaviours worth knowing before you rely on them:
- **`--remove-group` fails if the user is not in that group.** Every
other flag is idempotent — setting what is already set is fine, so a
"make these four things true" call does not break when one of them
already was. Revocation is the exception on purpose: a typo'd group
name that reported success would leave an account holding access you
believe you took away, and that is the one outcome nobody re-checks.
- **The resulting group list is printed** because group names have no
registry anywhere. A misspelled `--add-group` creates a real group that
no access-control rule mentions, so the user gains nothing and no error
is possible — reading the line back is the only check there is.
Passwords are deliberately out of scope here: regenerating a credential
is a different intent from editing an attribute, and folding them means
an attribute edit can invalidate a login by accident.
## What secrets exist, and where each one lives
Every secret in the swarm, with its generator and its path, is tabulated
in one place: [`secrets.md`](secrets.md), including authelia's own keys
(session, JWT, storage-encryption, OIDC HMAC, OIDC issuer) and the two
halves of each client secret. That page's two rules — a secret is always
a path, never a value, and the generator and the reader are usually in
different containers — are why the client secret's plaintext half needs
the delivery step below and the rest of authelia's keys don't.
## Getting the plaintext to the relying party
Three cases, and they are genuinely different mechanisms rather than one
mechanism with flags.
### 1. All-local — one host runs both
Nothing to configure at all. Per service, a host-side unit waits for
authelia's first boot to mint that client's secret and copies it into the
service's container, and the service's own module contributes its client
entry — callback URL included — to authelia's client list.
The callback is built once and read twice, so the redirect URI authelia is
told to allow and the one the service actually sends cannot drift apart. A
mismatch there is a rejected login with no error text worth reading.
⚠️ The delivery is a copy, not a `bindMounts` entry, and deliberately so:
nixos-container refuses to start a container whose bind source is
missing, and this secret does not exist until authelia's first boot has
run. Binding it would make the service wait on a file that waits on a
container that starts after it — on a fresh hive, a permanent stall
presenting as "the forge is broken", several layers from its cause.
### 2. Swarm-managed services
The controller side owns provisioning: `swarmctl` writes both halves, the
same way it already owns authelia's user store (`users.yml`, read and
written in place).
### 3. A hive elsewhere
No shared host, so no automatic path. The operator provides the file and
names it:
```nix
services.hyperhive.swarm = {
authelia.url = "https://auth.example.com";
forge.sso = {
enable = true;
clientSecretFile = "/var/lib/hyperhive/forge-oidc-secret";
};
};
```
**Both are asserted at eval.** A hive that boots with SSO
half-configured shows a login button that always fails — a symptom
several layers from its cause, and far worse to diagnose than an
evaluation error.
## Where each relying party differs
The registration half is identical; what each service does with the
result is not.
| | forge | matrix |
|---|---|---|
| how it learns the config | a oneshot calls `forgejo admin auth`, writing a login-source row into its database | tuwunel reads a `[[global.identity_provider]]` entry from its config file |
| how it reads the secret | a path inside its container | the same path, handed on by `LoadCredential` |
| callback URL | `<root>/user/oauth2/<source>/callback` | `<homeserver>/_matrix/client/unstable/login/sso/callback/<client_id>`, a shape tuwunel fixes rather than accepts |
| cost of a malformed entry | the login source is missing | the homeserver can refuse to start |
Two consequences worth stating plainly:
- **tuwunel re-reads its secret file on every OAuth exchange**, not only
at startup, and its own sandboxing hides most paths from it. It gets the
file through `LoadCredential` for the same reason the registration token
does — that keeps `DynamicUser` and `PrivateUsers` intact, with no
host-side ownership arrangement to maintain.
- **Matrix SSO lives inside the homeserver.** The client-server API is
spoken by non-browser clients holding matrix access tokens — every
agent's own daemon — as well as by federation, so `/_matrix/` is served
directly and authenticates itself. The forward-auth vhosts protect
browser surfaces; this is not one of them.
## Machine clients
Everything above is a person in a browser. A swarm also has callers that
hold no session and follow no redirect: the telemetry collector, the
queue's auth-callout responder, and each hive's own agents.
**One client per hive, not one per service.** A hive's identity belongs to
the directory rather than to whichever service happens to consume it, so a
hive holds a single OIDC client — `<hiveClientPrefix><hive>` — and mints a
different token per service from it. The alternative, letting each
consuming subsystem declare its own list, collides on the same client id
the moment a second consumer appears.
**The audience is that client id.** A swarm service that has to tell hives
apart needs one name both sides already agree on, and the client id is
already that name. A parallel per-hive naming scheme would be a second
thing to keep in step, and it drifts silently — a mismatch presents as a
valid token refused at the target, which reads like a broken credential
rather than a broken name.
**Tokens are signed (`RS256`), not opaque**, because a resource server
that cannot call the provider back is a real case here: the telemetry
collector verifies offline against `/jwks.json`, and an opaque token gives
it nothing to verify. The queue's responder introspects instead — a
different question asked of the same token, and the reason both
`/api/oidc/introspection` and `/jwks.json` have to stay reachable.
### Machine callers must fail closed
⚠️ **An error page that answers `200` is a security bug, not a cosmetic
one.** The browser surface intercepts upstream errors and serves a
friendly "SSO is unavailable" page; that page is a file, so it returns
`200`. Any machine caller routed through it receives a success carrying
HTML instead of the failure that actually happened:
- `/api/authz/auth-request` — nginx `auth_request` treats **any 2xx as
success**, so a down provider means *access granted*
- `/api/oidc/introspection` — a token check that answers `200`
- `/api/oidc/token`, `/.well-known/openid-configuration` — a client
parsing an error page as its JSON document
So authelia's `/api/` and `/.well-known/` prefixes are routed **without**
error interception. The split is by *audience*, not by an enumerated path
list: a human gets the page, every machine caller gets the status.
Enumerating endpoints individually would leave the next one added
silently intercepted.
The same shape bites any machine route behind a browser-shaped gate: a
`302` to a login page is followed, the login page answers `200`, and the
caller reports success while nothing happened. Log ingest hit exactly this
and lost eleven hours of delivery in silence.
**Checking it, if you change this routing.** Point the vhost at a dead
upstream and compare three requests, not one:
1. through `/` — must still serve the friendly page
2. through `/api/` — must deny
3. a direct dial to authelia — must match what (2) did
All three matter. A change that silently deleted the browser page would
pass a deny-only check, and one that quietly stopped denying would pass a
page-only check. This was verified that way when the split was introduced.
## What this does not do
- **It does not disable local login.** Each service keeps its password
database and gains a second door. An identity provider that can take a
service offline when it hiccups is worse than one with two ways in.
Making authelia the only path is a separate, reversible switch per
service (tuwunel's `login_with_password`, forgejo's own setting).
- **It does not provision users.** Agents are created and destroyed
continuously, so the subject set belongs to a program rather than to a
config file; today that program is `swarmctl`.