hyperhive/docs/swarm/sso.md
atlas 9b14014077 docs(sso): document the machine surface, and stop restating it in nix
`docs/swarm/sso.md` described a person in a browser. The swarm's other
callers — the telemetry collector, the queue's auth-callout responder, each
hive's agents — hold no session and follow no redirect, and nothing operator-
facing said how they authenticate. Its relying-party table is forge and
matrix, both browser surfaces.

The new section carries what `swarm-authelia.nix` was holding in comments:
one client per hive because identity belongs to the directory, the audience
being that client id rather than a parallel naming scheme, and signed rather
than opaque tokens because the collector verifies offline against
`/jwks.json` while the queue introspects.

It also states the fail-closed rule once, in the place a reader looks before
touching a vhost: an error page answers 200, and `auth_request` reads any 2xx
as access granted. That shape has now appeared three times — this module's
`/api/` prefix and both of victorialogs' routes — which is what makes it
documentation rather than a comment.

The two comment blocks those replace shrink to the part that is genuinely
local: the submodule-typing reason these clients are a definition rather than
an append, and a loud warning against folding the machine prefix back into
`/`. The security warning stays at the site; only its consequence list moves.

Comments 495 -> 465 lines. Option `description` strings are untouched: they
are the source `pkgs.nixosOptionsDoc` renders into the operator's options
reference, so trimming one would delete published documentation rather than a
duplicate.
2026-09-02 08:57:23 +02:00

12 KiB

Swarm SSO

The swarm runs one authelia, and it is two things at once: the session provider every protected vhost checks (auth_request), and — once any client is declared — an OIDC provider issuing tokens to relying parties: the forge and the matrix homeserver.

The second role is derived rather than switched: services.hyperhive.swarm.authelia.oidc.clients being non-empty turns it on. authelia refuses to start with a provider that has no clients, so a separate enable would be a second fact free to disagree with the first.

Getting in the first time

authelia binds loopback only. The gateway on the host running it publishes it as auth.<swarm.domain> — vhost, dnsmasq record and TLS name all follow deploy.authelia, so there is nothing to turn on separately. (Details, including why a client hive must not declare that vhost: ../networking/gateway.md.)

Authelia does not start until at least one user exists. The user store is generated empty — deliberately, since seeding a default account would put a credential in a config file — but authelia validates it at startup and treats "no users" as fatal:

error reading the authentication database: could not validate the schema:
  users: non zero value required

It then exits 1 and systemd restarts it, so a swarm that has been enabled but not bootstrapped shows a crash-looping unit and 502 Bad Gateway from the vhost — not a login page with nobody able to use it. The gateway is working in that state; the upstream is not up.

⚠️ So the step below is required to finish the install, not an optional first-login convenience. Run it before concluding anything is wrong with the proxy: a 502 here means "no users yet" far more often than it means a routing fault.

Add the first subject on the host running authelia:

# swarmctl user add mara --display-name Mara --email mara@example.com --group admins
added mara to /var/lib/authelia-swarm/users.yml
password: <generated>
this password is stored nowhere — record it now

The password is generated, hashed, and printed once; only the hash is kept. swarmctl reads and writes authelia's users.yml directly — it is the one user store, shared with swarm-authelia-bridge, which creates agent identities in the same file. No restart: authelia watches it. Full reference: ../tools/swarmctl-cli.md.

You can edit users.yml by hand, and swarmctl will read what you wrote. ⚠️ It rewrites the whole file on every change, so comments and formatting do not survive; values and unrecognised keys do.

This step stays manual on purpose. Bootstrapping an identity provider non-interactively means a secret arriving from somewhere — a file, an env var, a nix expression — and every one of those is worse than an operator typing one command once.

Changing a subject afterwards

user add only ever adds: on a name that already exists it refuses, rather than resurfacing as a second account or a silent overwrite. Editing an existing subject is user update, and the flags compose, so one call can change several things:

# swarmctl user update mara --add-group admins --email mara@example.com
added to group "admins"
email: unset -> "mara@example.com"
mara is now in groups: admins

Two behaviours worth knowing before you rely on them:

  • --remove-group fails if the user is not in that group. Every other flag is idempotent — setting what is already set is fine, so a "make these four things true" call does not break when one of them already was. Revocation is the exception on purpose: a typo'd group name that reported success would leave an account holding access you believe you took away, and that is the one outcome nobody re-checks.
  • The resulting group list is printed because group names have no registry anywhere. A misspelled --add-group creates a real group that no access-control rule mentions, so the user gains nothing and no error is possible — reading the line back is the only check there is.

Passwords are deliberately out of scope here: regenerating a credential is a different intent from editing an attribute, and folding them means an attribute edit can invalidate a login by accident.

What secrets exist, and where each one lives

Every secret in the swarm, with its generator and its path, is tabulated in one place: secrets.md, including authelia's own keys (session, JWT, storage-encryption, OIDC HMAC, OIDC issuer) and the two halves of each client secret. That page's two rules — a secret is always a path, never a value, and the generator and the reader are usually in different containers — are why the client secret's plaintext half needs the delivery step below and the rest of authelia's keys don't.

Getting the plaintext to the relying party

Three cases, and they are genuinely different mechanisms rather than one mechanism with flags.

1. All-local — one host runs both

Nothing to configure at all. Per service, a host-side unit waits for authelia's first boot to mint that client's secret and copies it into the service's container, and the service's own module contributes its client entry — callback URL included — to authelia's client list.

The callback is built once and read twice, so the redirect URI authelia is told to allow and the one the service actually sends cannot drift apart. A mismatch there is a rejected login with no error text worth reading.

⚠️ The delivery is a copy, not a bindMounts entry, and deliberately so: nixos-container refuses to start a container whose bind source is missing, and this secret does not exist until authelia's first boot has run. Binding it would make the service wait on a file that waits on a container that starts after it — on a fresh hive, a permanent stall presenting as "the forge is broken", several layers from its cause.

2. Swarm-managed services

The controller side owns provisioning: swarmctl writes both halves, the same way it already owns authelia's user store (users.yml, read and written in place).

3. A hive elsewhere

No shared host, so no automatic path. The operator provides the file and names it:

services.hyperhive.swarm = {
  authelia.url = "https://auth.example.com";
  forge.sso = {
    enable = true;
    clientSecretFile = "/var/lib/hyperhive/forge-oidc-secret";
  };
};

Both are asserted at eval. A hive that boots with SSO half-configured shows a login button that always fails — a symptom several layers from its cause, and far worse to diagnose than an evaluation error.

Where each relying party differs

The registration half is identical; what each service does with the result is not.

forge matrix
how it learns the config a oneshot calls forgejo admin auth, writing a login-source row into its database tuwunel reads a [[global.identity_provider]] entry from its config file
how it reads the secret a path inside its container the same path, handed on by LoadCredential
callback URL <root>/user/oauth2/<source>/callback <homeserver>/_matrix/client/unstable/login/sso/callback/<client_id>, a shape tuwunel fixes rather than accepts
cost of a malformed entry the login source is missing the homeserver can refuse to start

Two consequences worth stating plainly:

  • tuwunel re-reads its secret file on every OAuth exchange, not only at startup, and its own sandboxing hides most paths from it. It gets the file through LoadCredential for the same reason the registration token does — that keeps DynamicUser and PrivateUsers intact, with no host-side ownership arrangement to maintain.
  • Matrix SSO lives inside the homeserver. The client-server API is spoken by non-browser clients holding matrix access tokens — every agent's own daemon — as well as by federation, so /_matrix/ is served directly and authenticates itself. The forward-auth vhosts protect browser surfaces; this is not one of them.

Machine clients

Everything above is a person in a browser. A swarm also has callers that hold no session and follow no redirect: the telemetry collector, the queue's auth-callout responder, and each hive's own agents.

One client per hive, not one per service. A hive's identity belongs to the directory rather than to whichever service happens to consume it, so a hive holds a single OIDC client — <hiveClientPrefix><hive> — and mints a different token per service from it. The alternative, letting each consuming subsystem declare its own list, collides on the same client id the moment a second consumer appears.

The audience is that client id. A swarm service that has to tell hives apart needs one name both sides already agree on, and the client id is already that name. A parallel per-hive naming scheme would be a second thing to keep in step, and it drifts silently — a mismatch presents as a valid token refused at the target, which reads like a broken credential rather than a broken name.

Tokens are signed (RS256), not opaque, because a resource server that cannot call the provider back is a real case here: the telemetry collector verifies offline against /jwks.json, and an opaque token gives it nothing to verify. The queue's responder introspects instead — a different question asked of the same token, and the reason both /api/oidc/introspection and /jwks.json have to stay reachable.

Machine callers must fail closed

⚠️ An error page that answers 200 is a security bug, not a cosmetic one. The browser surface intercepts upstream errors and serves a friendly "SSO is unavailable" page; that page is a file, so it returns 200. Any machine caller routed through it receives a success carrying HTML instead of the failure that actually happened:

  • /api/authz/auth-request — nginx auth_request treats any 2xx as success, so a down provider means access granted
  • /api/oidc/introspection — a token check that answers 200
  • /api/oidc/token, /.well-known/openid-configuration — a client parsing an error page as its JSON document

So authelia's /api/ and /.well-known/ prefixes are routed without error interception. The split is by audience, not by an enumerated path list: a human gets the page, every machine caller gets the status. Enumerating endpoints individually would leave the next one added silently intercepted.

The same shape bites any machine route behind a browser-shaped gate: a 302 to a login page is followed, the login page answers 200, and the caller reports success while nothing happened. Log ingest hit exactly this and lost eleven hours of delivery in silence.

Checking it, if you change this routing. Point the vhost at a dead upstream and compare three requests, not one:

  1. through / — must still serve the friendly page
  2. through /api/ — must deny
  3. a direct dial to authelia — must match what (2) did

All three matter. A change that silently deleted the browser page would pass a deny-only check, and one that quietly stopped denying would pass a page-only check. This was verified that way when the split was introduced.

What this does not do

  • It does not disable local login. Each service keeps its password database and gains a second door. An identity provider that can take a service offline when it hiccups is worse than one with two ways in. Making authelia the only path is a separate, reversible switch per service (tuwunel's login_with_password, forgejo's own setting).
  • It does not provision users. Agents are created and destroyed continuously, so the subject set belongs to a program rather than to a config file; today that program is swarmctl.