Apply contraction fixes across ~40 doc files (setup, integrations, lifecycle, networking, scheduler, swarm, tools, trust-boundary, UI, etc.). Skipped 14 hits: - 10 where words appear in ALL CAPS for deliberate emphasis (is NOT, do NOT, etc.) - 4 where text could not be safely located due to markdown formatting or column position Applied via systematic scan with checks for fenced code blocks, inline code spans, and intentional caps. Preserves sentence-initial capitalization throughout.
252 lines
12 KiB
Markdown
252 lines
12 KiB
Markdown
# Swarm SSO
|
|
|
|
The swarm runs one authelia, and it's two things at once: the **session
|
|
provider** every protected vhost checks (`auth_request`), and — once any
|
|
client is declared — an **OIDC provider** issuing tokens to relying
|
|
parties: the forge and the matrix homeserver.
|
|
|
|
The second role is derived rather than switched:
|
|
`services.hyperhive.swarm.authelia.oidc.clients` being non-empty turns it
|
|
on. authelia refuses to start with a provider that has no clients, so a
|
|
separate `enable` would be a second fact free to disagree with the first.
|
|
|
|
## Getting in the first time
|
|
|
|
authelia binds loopback only. The **gateway** on the host running it
|
|
publishes it as `auth.<swarm.domain>` — vhost, dnsmasq record and TLS
|
|
name all follow `deploy.authelia`, so there is nothing to turn on
|
|
separately. (Details, including why a client hive must not declare that
|
|
vhost: [`../networking/gateway.md`](../networking/gateway.md).)
|
|
|
|
**Authelia doesn't start until at least one user exists.** The user
|
|
store is generated empty — deliberately, since seeding a default account
|
|
would put a credential in a config file — but authelia validates it at
|
|
startup and treats "no users" as fatal:
|
|
|
|
```
|
|
error reading the authentication database: could not validate the schema:
|
|
users: non zero value required
|
|
```
|
|
|
|
It then exits 1 and systemd restarts it, so a swarm that has been
|
|
enabled but not bootstrapped shows a **crash-looping unit** and `502 Bad
|
|
Gateway` from the vhost — not a login page with nobody able to use it.
|
|
The gateway is working in that state; the upstream isn't up.
|
|
|
|
⚠️ So the step below is **required to finish the install**, not an
|
|
optional first-login convenience. Run it before concluding anything is
|
|
wrong with the proxy: a 502 here means "no users yet" far more often
|
|
than it means a routing fault.
|
|
|
|
Add the first subject on the host running authelia:
|
|
|
|
```console
|
|
# swarmctl user add mara --display-name Mara --email mara@example.com --group admins
|
|
added mara to /var/lib/authelia-swarm/users.yml
|
|
password: <generated>
|
|
this password is stored nowhere — record it now
|
|
```
|
|
|
|
The password is generated, hashed, and printed once; only the hash is
|
|
kept. `swarmctl` reads and writes authelia's `users.yml` directly — it's
|
|
the one user store, shared with `swarm-authelia-bridge`, which creates
|
|
agent identities in the same file. No restart: authelia watches it. Full
|
|
reference: [`../tools/swarmctl-cli.md`](../tools/swarmctl-cli.md).
|
|
|
|
You can edit `users.yml` by hand, and `swarmctl` will read what you
|
|
wrote. ⚠️ It rewrites the whole file on every change, so **comments and
|
|
formatting don't survive**; values and unrecognised keys do.
|
|
|
|
This step stays manual on purpose. Bootstrapping an identity provider
|
|
non-interactively means a secret arriving from somewhere — a file, an
|
|
env var, a nix expression — and every one of those is worse than an
|
|
operator typing one command once.
|
|
|
|
### Changing a subject afterwards
|
|
|
|
`user add` only ever adds: on a name that already exists it refuses,
|
|
rather than resurfacing as a second account or a silent overwrite.
|
|
Editing an existing subject is `user update`, and the flags compose, so
|
|
one call can change several things:
|
|
|
|
```console
|
|
# swarmctl user update mara --add-group admins --email mara@example.com
|
|
added to group "admins"
|
|
email: unset -> "mara@example.com"
|
|
mara is now in groups: admins
|
|
```
|
|
|
|
Two behaviours worth knowing before you rely on them:
|
|
|
|
- **`--remove-group` fails if the user isn't in that group.** Every
|
|
other flag is idempotent — setting what's already set is fine, so a
|
|
"make these four things true" call doesn't break when one of them
|
|
already was. Revocation is the exception on purpose: a typo'd group
|
|
name that reported success would leave an account holding access you
|
|
believe you took away, and that's the one outcome nobody re-checks.
|
|
- **The resulting group list is printed** because group names have no
|
|
registry anywhere. A misspelled `--add-group` creates a real group that
|
|
no access-control rule mentions, so the user gains nothing and no error
|
|
is possible — reading the line back is the only check there is.
|
|
|
|
Passwords are deliberately out of scope here: regenerating a credential
|
|
is a different intent from editing an attribute, and folding them means
|
|
an attribute edit can invalidate a login by accident.
|
|
|
|
## What secrets exist, and where each one lives
|
|
|
|
Every secret in the swarm, with its generator and its path, is tabulated
|
|
in one place: [`secrets.md`](secrets.md), including authelia's own keys
|
|
(session, JWT, storage-encryption, OIDC HMAC, OIDC issuer) and the two
|
|
halves of each client secret. That page's two rules — a secret is always
|
|
a path, never a value, and the generator and the reader are usually in
|
|
different containers — are why the client secret's plaintext half needs
|
|
the delivery step below and the rest of authelia's keys don't.
|
|
|
|
## Getting the plaintext to the relying party
|
|
|
|
Three cases, and they're genuinely different mechanisms rather than one
|
|
mechanism with flags.
|
|
|
|
### 1. All-local — one host runs both
|
|
|
|
Nothing to configure at all. Per service, a host-side unit waits for
|
|
authelia's first boot to mint that client's secret and copies it into the
|
|
service's container, and the service's own module contributes its client
|
|
entry — callback URL included — to authelia's client list.
|
|
|
|
The callback is built once and read twice, so the redirect URI authelia is
|
|
told to allow and the one the service actually sends can't drift apart. A
|
|
mismatch there is a rejected login with no error text worth reading.
|
|
|
|
⚠️ The delivery is a copy, not a `bindMounts` entry, and deliberately so:
|
|
nixos-container refuses to start a container whose bind source is
|
|
missing, and this secret doesn't exist until authelia's first boot has
|
|
run. Binding it would make the service wait on a file that waits on a
|
|
container that starts after it — on a fresh hive, a permanent stall
|
|
presenting as "the forge is broken", several layers from its cause.
|
|
|
|
### 2. Swarm-managed services
|
|
|
|
The controller side owns provisioning: `swarmctl` writes both halves, the
|
|
same way it already owns authelia's user store (`users.yml`, read and
|
|
written in place).
|
|
|
|
### 3. A hive elsewhere
|
|
|
|
No shared host, so no automatic path. The operator provides the file and
|
|
names it:
|
|
|
|
```nix
|
|
services.hyperhive.swarm = {
|
|
authelia.url = "https://auth.example.com";
|
|
forge.sso = {
|
|
enable = true;
|
|
clientSecretFile = "/var/lib/hyperhive/forge-oidc-secret";
|
|
};
|
|
};
|
|
```
|
|
|
|
**Both are asserted at eval.** A hive that boots with SSO
|
|
half-configured shows a login button that always fails — a symptom
|
|
several layers from its cause, and far worse to diagnose than an
|
|
evaluation error.
|
|
|
|
## Where each relying party differs
|
|
|
|
The registration half is identical; what each service does with the
|
|
result isn't.
|
|
|
|
| | forge | matrix |
|
|
| ------------------------- | ---------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
|
|
| how it learns the config | a oneshot calls `forgejo admin auth`, writing a login-source row into its database | tuwunel reads a `[[global.identity_provider]]` entry from its config file |
|
|
| how it reads the secret | a path inside its container | the same path, handed on by `LoadCredential` |
|
|
| callback URL | `<root>/user/oauth2/<source>/callback` | `<homeserver>/_matrix/client/unstable/login/sso/callback/<client_id>`, a shape tuwunel fixes rather than accepts |
|
|
| cost of a malformed entry | the login source is missing | the homeserver can refuse to start |
|
|
|
|
Two consequences worth stating plainly:
|
|
|
|
- **tuwunel re-reads its secret file on every OAuth exchange**, not only
|
|
at startup, and its own sandboxing hides most paths from it. It gets the
|
|
file through `LoadCredential` for the same reason the registration token
|
|
does — that keeps `DynamicUser` and `PrivateUsers` intact, with no
|
|
host-side ownership arrangement to maintain.
|
|
- **Matrix SSO lives inside the homeserver.** The client-server API is
|
|
spoken by non-browser clients holding matrix access tokens — every
|
|
agent's own daemon — as well as by federation, so `/_matrix/` is served
|
|
directly and authenticates itself. The forward-auth vhosts protect
|
|
browser surfaces; this isn't one of them.
|
|
|
|
## Machine clients
|
|
|
|
Everything above is a person in a browser. A swarm also has callers that
|
|
hold no session and follow no redirect: the telemetry collector, the
|
|
queue's auth-callout responder, and each hive's own agents.
|
|
|
|
**One client per hive, not one per service.** A hive's identity belongs to
|
|
the directory rather than to whichever service happens to consume it, so a
|
|
hive holds a single OIDC client — `<hiveClientPrefix><hive>` — and mints a
|
|
different token per service from it. The alternative, letting each
|
|
consuming subsystem declare its own list, collides on the same client id
|
|
the moment a second consumer appears.
|
|
|
|
**The audience is that client id.** A swarm service that has to tell hives
|
|
apart needs one name both sides already agree on, and the client id is
|
|
already that name. A parallel per-hive naming scheme would be a second
|
|
thing to keep in step, and it drifts silently — a mismatch presents as a
|
|
valid token refused at the target, which reads like a broken credential
|
|
rather than a broken name.
|
|
|
|
**Tokens are signed (`RS256`), not opaque**, because a resource server
|
|
that can't call the provider back is a real case here: the telemetry
|
|
collector verifies offline against `/jwks.json`, and an opaque token gives
|
|
it nothing to verify. The queue's responder introspects instead — a
|
|
different question asked of the same token, and the reason both
|
|
`/api/oidc/introspection` and `/jwks.json` have to stay reachable.
|
|
|
|
### Machine callers must fail closed
|
|
|
|
⚠️ **An error page that answers `200` is a security bug, not a cosmetic
|
|
one.** The browser surface intercepts upstream errors and serves a
|
|
friendly "SSO is unavailable" page; that page is a file, so it returns
|
|
`200`. Any machine caller routed through it receives a success carrying
|
|
HTML instead of the failure that actually happened:
|
|
|
|
- `/api/authz/auth-request` — nginx `auth_request` treats **any 2xx as
|
|
success**, so a down provider means _access granted_
|
|
- `/api/oidc/introspection` — a token check that answers `200`
|
|
- `/api/oidc/token`, `/.well-known/openid-configuration` — a client
|
|
parsing an error page as its JSON document
|
|
|
|
So authelia's `/api/` and `/.well-known/` prefixes are routed **without**
|
|
error interception. The split is by _audience_, not by an enumerated path
|
|
list: a human gets the page, every machine caller gets the status.
|
|
Enumerating endpoints individually would leave the next one added
|
|
silently intercepted.
|
|
|
|
The same shape bites any machine route behind a browser-shaped gate: a
|
|
`302` to a login page is followed, the login page answers `200`, and the
|
|
caller reports success while nothing happened. Log ingest hit exactly this
|
|
and lost eleven hours of delivery in silence.
|
|
|
|
**Checking it, if you change this routing.** Point the vhost at a dead
|
|
upstream and compare three requests, not one:
|
|
|
|
1. through `/` — must still serve the friendly page
|
|
2. through `/api/` — must deny
|
|
3. a direct dial to authelia — must match what (2) did
|
|
|
|
All three matter. A change that silently deleted the browser page would
|
|
pass a deny-only check, and one that quietly stopped denying would pass a
|
|
page-only check. This was verified that way when the split was introduced.
|
|
|
|
## What this doesn't do
|
|
|
|
- **It doesn't disable local login.** Each service keeps its password
|
|
database and gains a second door. An identity provider that can take a
|
|
service offline when it hiccups is worse than one with two ways in.
|
|
Making authelia the only path is a separate, reversible switch per
|
|
service (tuwunel's `login_with_password`, forgejo's own setting).
|
|
- **It doesn't provision users.** Agents are created and destroyed
|
|
continuously, so the subject set belongs to a program rather than to a
|
|
config file; today that program is `swarmctl`.
|