hyperhive/swarm-secret-client
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas f778122f5a matrix: mint the appservice sender token in the matrix container
A swarm runs one homeserver and a homeserver has one appservice sender
account, so "mint it once" is a property of the thing being minted
rather than something a lock has to enforce. That is what makes this
account the one to move first: no trigger route, no controller change
and no agent list — a boot-time oneshot beside tuwunel is the whole
mechanism.

`swarm-matrix-minter` runs inside `containers.hive-matrix`, which
already holds the appservice token: the rendered registration is bound
in read-only because that is how tuwunel is handed it. What the
container lacked was an identity of its own, so this adds one — a leaf
from the store's CA with a grant of exactly one path, not the hive's
leaf, which reads every secret in the store.

Both ends of the credential ship here. The minter reads the path it
publishes to before it touches the homeserver, and returning on a
non-empty read IS the "only once"; `hive-c0re`'s `ensure_hive_user`
reads the same path, authenticating with the hive name already in
`HYPERHIVE_HIVE_NAME`. The existing mint-then-`M_USER_IN_USE`-login
ladder stays as the fallback for a store that is empty, unconfigured or
unreachable, which is every swarm deployed before this — so nothing
needs backfilling and nothing breaks if the rest of the sequence never
lands.

The credential is not an admin credential, and is not named like one.
It is the access token of the appservice registration's own
`sender_localpart` — `@hive:<server_name>`, an account the homeserver
creates for itself when it loads the registration. The store path is
`swarm/services/matrix/sender-token`, the host path is
`matrix/access-token`, and the homeserver no longer runs an
`admin_execute` promotion for that account at boot. Everything the hive
provisions with it — the Space, the chat room, their hierarchy and join
rules, the invites — rides on being the creator of those rooms at power
level 100, not on homeserver admin; there is no Synapse admin API here
to need, tuwunel has none.

Two operations do need an admin *sender* and therefore stop working:
`hivectl matrix promote-user` and `hivectl matrix reset-password`, both
`!admin …` messages into `#admins:<server>`, plus the password-reset
recovery path that an agent with a lost password file falls back to.
They are swarm-level operations and are left failing loudly rather than
served by an over-privileged token every other call site would also
carry. The sweep's own admin-rights check and self-repair go with them:
an account that is deliberately not an admin has nothing to check.

`ephemeral = false` stays, and hive root can still read the container's
filesystem. Accepted: what this buys is identity separation — no hive
*process* holds or reads the appservice token — not physical isolation.

Refs #4345
2026-09-20 22:07:16 +02:00
..
src matrix: mint the appservice sender token in the matrix container 2026-09-20 22:07:16 +02:00
Cargo.toml swarm-secret-client: derive Kind's segment strings via strum instead of a hand-written match 2026-09-12 00:06:31 +02:00
README.md swarm-secret-client: one module per kind of secret, not one struct 2026-09-08 15:53:50 +02:00

swarm-secret-client

Reading and writing a swarm credential in the secret store, over vaultrs. The HTTP is that crate's job. What this one owns is the agreements both ends of the store have to state identically: where a credential lives, which field its bytes are in, and how this deployment's environment becomes a logged-in client.

The surrounding picture — which secret is minted where, and why delivery is a copy rather than a bind mount — is docs/swarm/secrets.md.

Why a crate and not a module per binary

The store has two Rust ends and they are peers: the swarm controller writes a credential, a hive reads it. Neither is senior to the other, so a path formatted at each call site is an agreement with no owner — it holds right up until one side is edited alone, and then it fails as a missing key rather than as a mismatch.

The third end is what settles it. nix/host-modules/glue-matrix-bao-token.nix reads the store with bao kv get -field=value. That reader is a shell line in a nix module: it cannot be renamed by the same refactor as a Rust struct, and no Rust test reaches it. So the field name is pinned by a test against the serialised literal rather than left to the struct definition.

The identity is a certificate, and the role is the hive's name

Authentication is the store's cert auth method. nix/host-modules/glue-bao-tls.nix mints the client certificate with its CN set to the hive's name, because a cert-auth role matches on the CN. So cert_role is not a free choice for the caller: a hive passes its own name, and the policy attached to that role is what scopes what it may read.

The listener's tls_require_and_verify_client_cert is a different thing and not a substitute. It decides who may open a connection; it says nothing about who the connection belongs to, and a store with the option set and no cert mount configured refuses every login made here. That refusal is what Error::Vault out of connect means.

Configuration

Settings::from_env reads BAO_ADDR, BAO_CLIENT_CERT, BAO_CLIENT_KEY, and the optional BAO_CACERT.

The BAO_ spellings are read explicitly rather than left to vaultrs. Its own defaults look for VAULT_ADDR / VAULT_CLIENT_CERT / VAULT_CLIENT_KEY, which no unit in this tree sets. Falling through to them builds a client with no identity at all, and that surfaces as a TLS handshake failure — a place that names neither the variable nor the reason.

Empty is as absent as unset. systemd renders an unset nix option as Environment=BAO_CACERT=, so empty is the shape a missing value arrives in.

The certificate and key are paths, not values, and are read at connect time — same rule as every other credential in this tree, for the same reason: a value in a nix expression is rendered into the world-readable store.

Reading the environment is separate from connecting (Settings::from_lookup) because every one of those failures is a misconfiguration an operator has to read an error about, and none of them needs a reachable store to happen.

Names that arrive from elsewhere

matrix::account_path is fallible, which for a string formatter needs saying: its segments are an agent name from the topology and an account name from that agent's own config. A / turns one agent's segment into another agent's directory and .. walks out of the prefix entirely, so the charset it accepts is deliberately narrower than what the store would.

One module per kind of secret

client moves whatever type a caller names; it decodes nothing itself. What a stored object holds is stated in the module that also builds its path — matrix today, and a second kind of swarm secret gets a module beside it.

The split is deliberate. A single shared struct that grows one field per consumer ends up carrying, on every path, a field only one path's reader has ever heard of; and the two things a kind of secret must pin — where it lives and what is in it — are one agreement that reads worse split across modules.

What this crate does not do

It has no opinion on what a caller may read. That is the policy attached to the cert role, and it lives in the store.

It holds the token minted at login and renews nothing. A handle is built per credential, so the login is the cheap part of a rare operation — a caller that wanted to keep one alive across a token's lifetime would need more than this.