hyperhive/docs/swarm/sso.md
atlas c39e94758e docs(swarm): one page saying where every secret goes
Per mara on the CA question: outside all-local this is an ops problem,
and what is missing is documentation rather than machinery.

One row per secret, read against three topologies, because the same
credential is generated in place on one and handed over by an operator on
another. sso.md's table is replaced by a pointer -- two tables listing the
same secrets would drift, and its prose about why a secret is generated
in-container is the half worth keeping there.

States the constraint the whole thing rests on: a hive CA is name-
constrained to the hive domain and a swarm service name is a sibling of
it, so a hive CA cannot issue a swarm-service certificate at all. That is
why placing the swarm root is an operations decision.
2026-08-14 13:22:43 +02:00

195 lines
8.7 KiB
Markdown

# Swarm SSO
The swarm runs one authelia, and it is two things at once: the **session
provider** every protected vhost checks (`auth_request`), and — once any
client is declared — an **OIDC provider** issuing tokens to relying
parties: the forge and the matrix homeserver.
The second role is derived rather than switched:
`services.hyperhive.swarm.authelia.oidc.clients` being non-empty turns it
on. authelia refuses to start with a provider that has no clients, so a
separate `enable` would be a second fact free to disagree with the first.
## Getting in the first time
authelia binds loopback only. The **gateway** on the host running it
publishes it as `auth.<swarm.domain>` — vhost, dnsmasq record and TLS
name all follow `swarm.authelia.enable`, so there is nothing to turn on
separately. (Details, including why a client hive must not declare that
vhost: [`../gateway.md`](../gateway.md).)
**Authelia does not start until at least one user exists.** The user
store is generated empty — deliberately, since seeding a default account
would put a credential in a config file — but authelia validates it at
startup and treats "no users" as fatal:
```
error reading the authentication database: could not validate the schema:
users: non zero value required
```
It then exits 1 and systemd restarts it, so a swarm that has been
enabled but not bootstrapped shows a **crash-looping unit** and `502 Bad
Gateway` from the vhost — not a login page with nobody able to use it.
The gateway is working in that state; the upstream is not up.
⚠️ So the step below is **required to finish the install**, not an
optional first-login convenience. Run it before concluding anything is
wrong with the proxy: a 502 here means "no users yet" far more often
than it means a routing fault.
Add the first subject on the host running authelia:
```console
# swarmctl user add mara --display-name Mara --email mara@example.com --group admins
added mara to /var/lib/authelia-swarm/users.yml
password: <generated>
this password is stored nowhere — record it now
```
The password is generated, hashed, and printed once; only the hash is
kept. `swarmctl` writes its canonical `users.json`, re-renders authelia's
`users.yml` from it, and restarts authelia. Full reference:
[`../tools/swarmctl-cli.md`](../tools/swarmctl-cli.md).
This step stays manual on purpose. Bootstrapping an identity provider
non-interactively means a secret arriving from somewhere — a file, an
env var, a nix expression — and every one of those is worse than an
operator typing one command once.
### Changing a subject afterwards
`user add` only ever adds: on a name that already exists it refuses,
rather than resurfacing as a second account or a silent overwrite.
Editing an existing subject is `user update`, and the flags compose, so
one call can change several things:
```console
# swarmctl user update mara --add-group admins --email mara@example.com
added to group "admins"
email: unset -> "mara@example.com"
mara is now in groups: admins
```
Two behaviours worth knowing before you rely on them:
- **`--remove-group` fails if the user is not in that group.** Every
other flag is idempotent — setting what is already set is fine, so a
"make these four things true" call does not break when one of them
already was. Revocation is the exception on purpose: a typo'd group
name that reported success would leave an account holding access you
believe you took away, and that is the one outcome nobody re-checks.
- **The resulting group list is printed** because group names have no
registry anywhere. A misspelled `--add-group` creates a real group that
no access-control rule mentions, so the user gains nothing and no error
is possible — reading the line back is the only check there is.
Passwords are deliberately out of scope here: regenerating a credential
is a different intent from editing an attribute, and folding them means
an attribute edit can invalidate a login by accident.
## What secrets exist, and where each one lives
Every secret in the swarm, with its generator and its path, is tabulated
in one place: [`secrets.md`](secrets.md). The rows relevant here are
authelia's own keys (session, JWT, storage-encryption, OIDC HMAC, OIDC
issuer) plus the two halves of each client secret.
What matters for this page is the shape rather than the paths. Authelia's
own keys are generated **in-container**, because nothing outside that
container ever reads them — that is the test worth applying to any secret
added here. The plaintext half of a client secret is the one that fails
it: its reader lives in a different container, and that is the entire
reason a delivery step exists.
**None of it is ever written into a nix expression.** authelia's
`settings` are rendered into the nix store, which is world-readable and
permanent, so the client digest reaches authelia through `settingsFiles`
(merged at runtime) and every other secret through a `*File` option
carrying a path rather than a value.
## Getting the plaintext to the relying party
Three cases, and they are genuinely different mechanisms rather than one
mechanism with flags.
### 1. All-local — one host runs both
Nothing to configure beyond `swarm.forge.sso.enable = true` or
`swarm.matrix.sso.enable = true`. Per service, a host-side unit waits for
authelia's first boot to mint that client's secret and copies it into the
service's container, and the service's own module contributes its client
entry — callback URL included — to authelia's client list.
The callback is built once and read twice, so the redirect URI authelia is
told to allow and the one the service actually sends cannot drift apart. A
mismatch there is a rejected login with no error text worth reading.
⚠️ The delivery is a copy, not a `bindMounts` entry, and deliberately so:
nixos-container refuses to start a container whose bind source is
missing, and this secret does not exist until authelia's first boot has
run. Binding it would make the service wait on a file that waits on a
container that starts after it — on a fresh hive, a permanent stall
presenting as "the forge is broken", several layers from its cause.
### 2. Swarm-managed services
The controller side owns provisioning: `swarmctl` writes both halves, the
same way it already owns authelia's user store (`users.json` canonical,
`users.yml` a rendered artifact).
### 3. A hive elsewhere
No shared host, so no automatic path. The operator provides the file and
names it:
```nix
services.hyperhive.swarm = {
authelia.url = "https://auth.example.com";
forge.sso = {
enable = true;
clientSecretFile = "/var/lib/hyperhive/forge-oidc-secret";
};
};
```
**Both are asserted at eval.** A hive that boots with SSO
half-configured shows a login button that always fails — a symptom
several layers from its cause, and far worse to diagnose than an
evaluation error.
## Where each relying party differs
The registration half is identical; what each service does with the
result is not.
| | forge | matrix |
|---|---|---|
| how it learns the config | a oneshot calls `forgejo admin auth`, writing a login-source row into its database | tuwunel reads a `[[global.identity_provider]]` entry from its config file |
| how it reads the secret | a path inside its container | the same path, handed on by `LoadCredential` |
| callback URL | `<root>/user/oauth2/<source>/callback` | `<homeserver>/_matrix/client/unstable/login/sso/callback/<client_id>`, a shape tuwunel fixes rather than accepts |
| cost of a malformed entry | the login source is missing | the homeserver can refuse to start |
Two consequences worth stating plainly:
- **tuwunel re-reads its secret file on every OAuth exchange**, not only
at startup, and its own sandboxing hides most paths from it. It gets the
file through `LoadCredential` for the same reason the registration token
does — that keeps `DynamicUser` and `PrivateUsers` intact, with no
host-side ownership arrangement to maintain.
- **Matrix SSO lives inside the homeserver.** The client-server API is
spoken by non-browser clients holding matrix access tokens — every
agent's own daemon — as well as by federation, so `/_matrix/` is served
directly and authenticates itself. The forward-auth vhosts protect
browser surfaces; this is not one of them.
## What this does not do
- **It does not disable local login.** Each service keeps its password
database and gains a second door. An identity provider that can take a
service offline when it hiccups is worse than one with two ways in.
Making authelia the only path is a separate, reversible switch per
service (tuwunel's `login_with_password`, forgejo's own setting).
- **It does not provision users.** Agents are created and destroyed
continuously, so the subject set belongs to a program rather than to a
config file; today that program is `swarmctl`.