Compare commits

..
9 changed files with 10 additions and 941 deletions

View file

@ -44,41 +44,7 @@ echo "hunter2" | hivectl gateway create-user mara --password-stdin
hivectl gateway list-users
```
### 3 · Secret store (only when `deploy.bao`)
⚠️ **A sealed store still answers.** OpenBao starts uninitialised and
sealed, so the container is up and the port responds while every read
times out — the failure looks like a hang, not like a store that was
never initialised. Do this before anything is pointed at it.
```bash
# On the host that RUNS the store, once.
bao operator init # keep the keys it prints and the root token OFF this host
```
Whether anything more is needed depends on
`services.hyperhive.deploy.bao.seal`:
- **`pkcs11`** (the default) — the key is bound to the host's TPM and the
store unseals itself on every restart. `init` is the only manual step.
- **`shamir`** — no TPM, so `bao operator unseal` is needed again after
every restart, with the keys `init` printed.
The store serves TLS, and on a hive that deploys it you need do nothing: a
first-boot unit mints a CA of the store's own plus the two leaves it signs —
the store's server certificate and this host's client certificate — and points
`deploy.bao.serverCertFile`, `.serverKeyFile` and `.clientCaFile` at them.
Those are `mkDefault`s, so naming your own paths wins. Do that when your
certificates come from a real internal CA; the store has no opinion about
which. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is
**mTLS between services and the store**, a separate trust domain, because a
store that took its identity from an authority it will itself distribute could
never come up before that authority.
Making even the `init` unnecessary is tracked in issue #3768.
### 4 · Swarm SSO (only when `deploy.authelia`)
### 3 · Swarm SSO (only when `deploy.authelia`)
⚠️ **Required to finish the install, not optional.** Authelia treats an
empty user store as a fatal startup error, so until this runs the
@ -104,7 +70,7 @@ than amend — adding the group afterwards is `swarmctl user update mara
Detail, including what the password is and why this stays manual:
[`swarm/sso.md`](swarm/sso.md).
### 5 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
### 4 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
Nothing to run — it is served on the swarm apex
(`https://<swarm.domain>/`) as soon as the host rebuilds. Two things
@ -120,7 +86,7 @@ decide whether you can actually open it:
Detail, including why reachability is deliberately not the access
control: [`swarm/ui.md`](swarm/ui.md).
### 6 · Matrix
### 5 · Matrix
```bash
# 5a. Ensure the hive-internal admin account exists first
@ -141,7 +107,7 @@ hivectl matrix promote-user mara
The human operator's own matrix account is created via swarm SSO instead of
a manual `hivectl` step — see step 3 (`swarmctl user add`).
### 7 · Spawn sub-agents
### 6 · Spawn sub-agents
Sub-agent creation goes through the approval queue — ruth proposes, the
operator approves, the container builds. From ruth's own turn (inside
@ -162,7 +128,7 @@ request_init_config(name: "iris")
See [`approvals.md`](approvals.md) for the full flow.
### 8 · Useful host commands
### 7 · Useful host commands
```bash
# Roster: all agents, status, rev, parent, pending reminders

View file

@ -55,8 +55,6 @@ neither is a renaming of the other.
| authelia subject store | `swarmctl` and `swarm-authelia-bridge` | `users.yml` — one file, read and written by both | `swarmctl`, on the host that runs authelia |
| wireguard private key | **the operator**`wg genkey` | whatever `swarm.wireguard.privateKeyFile` names | always operator-provided; nothing generates this for you |
| queue auth-callout nkeys (user seed + account seed) | `swarm-nats-callout-keys` first-boot unit, when `nats.autoGenerateCallout` is set | `/var/lib/swarm-nats-callout/{callout-user,issuer}.seed`, `0600` | operator mints both with `nk` and names them in `nats.calloutUserSeedFile` / `nats.calloutIssuerSeedFile` |
| the secret store's own contents | openbao, on first `bao operator init`**an operator action, not a unit** | `/var/lib/swarm-bao` on the host of whoever runs the store, bind-mounted into the `swarm-bao` container | n/a — there is one store; a hive elsewhere is a *client* of it and holds none of this |
| the secret store's unseal material | the HSM/TPM under `deploy.bao.seal = "pkcs11"`; openbao itself under `"shamir"` | in the token; or held by whoever ran `bao operator init`, which is what `"shamir"` means and why it is stated rather than inferred | n/a — only the host running the store seals anything |
The three keys authelia mints for itself are generated in-container precisely
because nothing outside that container ever reads them. **That is the test worth
@ -134,29 +132,6 @@ bounded wait, 120s — and then **fail loudly** rather than skipping. A silent s
produces a service whose login button always fails, which is a symptom several
layers from its cause.
The store's **first reader** is the matrix registration token, and it is worth
saying why that one: it is an opaque 32-byte value with no second file and no
format. Authelia's OIDC secret needs a `.secret` *and* a matching `.digest`, so
starting there would have meant debugging "can a reader authenticate and get
bytes back" and "did we write authelia's file format right" at once, with an
SSO outage as the failure mode.
`glue-matrix-bao-token.nix` fetches it and writes the file `hive-matrix.nix`
already reads, so the homeserver never learns the store exists. Every failure
path — no such key, sealed store, unreachable store, empty value — leaves the
locally minted token in place, so a hive with no store behaves exactly as it
did before.
⚠️ **Service↔store mTLS is its own trust domain.** A credential you must
already hold to authenticate cannot be fetched from the thing it authenticates
you to, so the store's identity cannot come from an authority the store
distributes — which excludes the hive CA and the swarm CA both, and has nothing
to do with the gateway's HTTPS certificates either way. `glue-bao-tls.nix`
mints a CA that signs exactly two things, the store's server certificate and a
reader's client certificate, and distributes nothing. A deployment with a real
internal CA deletes that file and names its own paths in
`deploy.bao.serverCertFile` / `clientCaFile`; the store itself has no opinion.
## The constraint that decides where the root lives
A hive CA carries `nameConstraints=permitted;DNS:<hive domain>`, and **a swarm

View file

@ -23,10 +23,7 @@
./hive-priv.nix
./hive-tls.nix
./otel.nix
./glue-bao-tls.nix
./glue-matrix-bao-token.nix
./swarm-authelia.nix
./swarm-bao.nix
./swarm-ca.nix
./swarm-nats.nix
./swarm-controller.nix

View file

@ -10,9 +10,9 @@
# deploy.* — this machine's deployment decisions. Necessarily different
# on every host, because that is what a deployment is.
#
# The `enable` toggles live here rather than under `swarm.*` so the
# namespace that is identical everywhere does not carry the one thing
# that must differ per host.
# The `enable` toggles used to live under `swarm.*`, which made the
# namespace that is supposed to be identical everywhere carry the one
# thing that must differ.
#
# Flat and named for the thing deployed — `deploy.forgejo`, not
# `deploy.swarmServices.forgejo`: grouping by "swarm service" re-encodes
@ -161,32 +161,6 @@ in
'';
};
bao.enable = lib.mkOption {
type = lib.types.bool;
default = false;
example = true;
description = ''
Run the swarm's secret store in a `swarm-bao` container on this
host. A swarm has one store and it has to exist somewhere, so
this is asserted from
{option}`services.hyperhive.swarm.enableRequiredServices`
alongside the other once-per-swarm services.
That assertion is a `mkDefault`, which is what keeps *where* the
store runs a separate question from *that* it runs: set this
directly to put the store on a host of its own, and clients
still reach it by name at
{option}`services.hyperhive.swarm.bao.domain` rather than at a
local address.
With it off, this hive is a *client*: it still reads its own
secrets from whoever runs the store, authenticating with its own
client certificate. Every hive needs the client half; only one
runs the server half, which is why the two live in different
namespaces.
'';
};
authelia.enable = lib.mkOption {
type = lib.types.bool;
default = false;
@ -249,9 +223,8 @@ in
default = false;
description = ''
Run the swarm's message queue in a `swarm-nats` container on this
host. A swarm has one queue, so at most one host turns this on
but *which* host is its own decision, not necessarily the one
running the swarm's other shared services.
host. A swarm has one queue, so this belongs on the same host as
the rest of the shared services.
Off by default, and off means *absent*: no container is created
and nothing else in the evaluated config changes.

View file

@ -1,119 +0,0 @@
# Glue: give the secret store a PKI of its own, and point it at it.
#
# ONE PAIRING PER FILE — `glue-<consumer>-<what>.nix`. A single module holding
# every co-location default becomes the file nobody dares change, because a
# reader cannot tell which of its rules their deployment is subject to. Each
# of these should be deletable on its own, and deleting this one leaves a
# store that takes operator-provided certificates and nothing else.
#
# ⚠️ Why the PKI lives HERE and not in ./swarm-bao.nix: the store must have no
# opinion about where its identity comes from. Minting is an opinion — the
# most consequential one available — so it belongs to the glue that decides
# this deployment self-signs, not to the service that merely serves what it is
# handed. A deployment with a real internal CA drops this file and names its
# own paths; nothing in the store changes.
#
# ⚠️ Not the hive CA and not the swarm CA. The store will eventually
# distribute both, and an authority you must already hold a certificate from
# cannot be one the store hands out — reach the store to get the CA material,
# need a cert from that CA to reach the store. This CA signs exactly two
# things and distributes nothing, so it cannot enter that cycle.
#
# ⚠️ Files like this are the only place a `deploy.<foo>` value may derive from
# a `deploy.<bar>.enable`. Everywhere else that is forbidden. The exception
# earns itself: the derivation happens either way, and the alternative is
# having it spread through the service modules where it is invisible.
#
# Everything is `mkDefault`. An operator naming their own paths wins.
{
pkgs,
lib,
config,
...
}:
let
hyperhiveCfg = config.services.hyperhive;
deployCfg = hyperhiveCfg.deploy;
cfg = hyperhiveCfg.swarm.bao;
# Host-side, outside the container's tree, for the same reason the raft data
# is: `nixos-container destroy` must not take it. Losing the CA key means
# re-issuing every client certificate in the swarm.
pkiDir = "/var/lib/swarm-bao-pki";
# What a reader calls itself to the store. The hive's name, because a bao
# cert-auth role matches on the CN — this is an interface, not a label.
clientCn = if hyperhiveCfg.hiveName != null then hyperhiveCfg.hiveName else cfg.domain;
# $1 dir $2 basename $3 CN $4 SAN or "" $5 EKU
signLeaf = pkgs.writeShellScript "swarm-bao-sign-leaf" ''
set -euo pipefail
d="$1"; base="$2"; cn="$3"; sans="$4"; eku="$5"
csr="$(mktemp "$d/$base.csr.XXXXXX")"
ext="$(mktemp "$d/$base.ext.XXXXXX")"
trap 'rm -f "$csr" "$ext"' EXIT
openssl req -newkey rsa:4096 -nodes -sha256 \
-keyout "$d/$base-key.pem" -out "$csr" -subj "/CN=$cn"
{
[ -n "$sans" ] && printf 'subjectAltName=%s\n' "$sans"
printf 'basicConstraints=critical,CA:FALSE\n'
printf 'keyUsage=critical,digitalSignature,keyEncipherment\n'
printf 'extendedKeyUsage=%s\n' "$eku"
} > "$ext"
openssl x509 -req -in "$csr" -CA "$d/ca.pem" -CAkey "$d/ca-key.pem" \
-CAcreateserial -days 3650 -sha256 -extfile "$ext" -out "$d/$base.pem"
chmod 0600 "$d/$base-key.pem"
chmod 0644 "$d/$base.pem"
'';
in
{
config = lib.mkIf (hyperhiveCfg.enable && deployCfg.bao.enable) {
services.hyperhive.deploy.bao = {
serverCertFile = lib.mkDefault "${pkiDir}/server.pem";
serverKeyFile = lib.mkDefault "${pkiDir}/server-key.pem";
clientCaFile = lib.mkDefault "${pkiDir}/ca.pem";
};
# Idempotent on ABSENCE, never on content. Re-issuing the CA invalidates
# every client certificate already trusting it, so a rebuild that
# "refreshed" it would lock every reader in the swarm out at once — the
# same rule the store's TPM PIN unit follows, for a sharper reason.
systemd.services.swarm-bao-pki = {
description = "mint the swarm secret store's own CA and leaves";
before = [ "swarm-bao-certs.service" ];
requiredBy = [ "swarm-bao-certs.service" ];
path = [
pkgs.openssl
pkgs.coreutils
];
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
script = ''
set -euo pipefail
install -d -m 0700 ${pkiDir}
if [ ! -s ${pkiDir}/ca.pem ]; then
openssl req -x509 -newkey rsa:4096 -nodes -sha256 -days 3650 \
-keyout ${pkiDir}/ca-key.pem -out ${pkiDir}/ca.pem \
-subj "/CN=swarm-bao-ca ${cfg.domain}" \
-addext "basicConstraints=critical,CA:TRUE,pathlen:0" \
-addext "keyUsage=critical,keyCertSign,cRLSign"
chmod 0600 ${pkiDir}/ca-key.pem
chmod 0644 ${pkiDir}/ca.pem
fi
# The store's own identity, and the identity of a reader on this host.
# A reader elsewhere gets its leaf from this CA out of band — that is
# what makes the store reachable from another machine at all, and why
# the CA is a file rather than a service.
[ -s ${pkiDir}/server.pem ] || ${signLeaf} ${pkiDir} server \
${lib.escapeShellArg cfg.domain} ${lib.escapeShellArg "DNS:${cfg.domain}"} serverAuth
[ -s ${pkiDir}/client.pem ] || ${signLeaf} ${pkiDir} client \
${lib.escapeShellArg clientCn} "" clientAuth
'';
};
};
}

View file

@ -1,103 +0,0 @@
# Glue: the matrix registration token comes from the secret store.
#
# The store's first reader, and deliberately a small one. It fetches an opaque
# 32-byte value and writes it where ./hive-matrix.nix already looks — the
# homeserver never learns the store exists, and its config is unchanged.
#
# ⚠️ Why this credential first. It has no second file and no format: authelia's
# OIDC secret needs a `.secret` *and* a matching `.digest`, so shipping that
# one first would debug "can a reader authenticate and get bytes back" and
# "did we write authelia's file format right" at the same time, with an SSO
# outage as the failure mode. Here the failure is narrow — new agent accounts
# cannot be provisioned, existing ones are untouched, nothing crash-loops.
#
# ⚠️ The fallback is today's behaviour, not a new one. `hive-matrix.nix`'s
# activation script still mints a token when the file is absent; this unit
# overwrites it with the swarm's copy when the store has one. A store that is
# empty or unreachable leaves a working hive with a local token.
#
# 📌 LIMIT, stated rather than hidden: this gates on the store running HERE.
# A hive reading from a store on another machine needs the same unit with that
# machine's address and a client leaf issued out of band — the mechanism is
# identical, only `-address` and the cert's provenance differ. Deferred until
# there is a second hive to test it against, rather than shipped untested.
{
pkgs,
lib,
config,
...
}:
let
hyperhiveCfg = config.services.hyperhive;
deployCfg = hyperhiveCfg.deploy;
baoCfg = hyperhiveCfg.swarm.bao;
matrixCfg = hyperhiveCfg.swarm.matrix;
# Owned by ./glue-bao-tls.nix, which mints them. Named here rather than
# shared through a `let`: a cross-module binding would make these two files
# one file with a gap in the middle, and the whole point of a glue module is
# that it can be deleted on its own.
pkiDir = "/var/lib/swarm-bao-pki";
# Where the token lives in the store. A path, not a convention to guess at:
# whoever writes it and whoever reads it must agree, and the agreement
# belongs in one visible place.
tokenPath = "secret/swarm/matrix/registration-token";
# A literal, not an option — ./hive-matrix.nix names its container
# `containers.hive-matrix` directly and declares no `machine` to derive it
# from, which the trust-bundle call in that file already says out loud.
# ⚠️ `matrixCfg.machine` parses fine and fails at module-system resolution,
# so this is the kind of mistake only reading the target module catches.
matrixMachine = "hive-matrix";
in
{
config = lib.mkIf (hyperhiveCfg.enable && deployCfg.bao.enable && deployCfg.matrix.enable) {
systemd.services.swarm-bao-matrix-token = {
description = "fetch the matrix registration token from the swarm secret store";
after = [
"swarm-bao-pki.service"
"container@${baoCfg.machine}.service"
];
wants = [ "container@${baoCfg.machine}.service" ];
requires = [ "swarm-bao-pki.service" ];
before = [ "container@${matrixMachine}.service" ];
wantedBy = [ "container@${matrixMachine}.service" ];
path = [
deployCfg.bao.package
pkgs.coreutils
];
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
BAO_CACERT = "${pkiDir}/ca.pem";
BAO_CLIENT_CERT = "${pkiDir}/client.pem";
BAO_CLIENT_KEY = "${pkiDir}/client-key.pem";
};
script = ''
set -euo pipefail
# A sealed or uninitialised store answers on the port and times out on
# every read, so "the store is up" is not the same as "the store can
# answer" -- bound the wait rather than hanging the boot behind it.
if ! token="$(bao kv get -field=value ${lib.escapeShellArg tokenPath} 2>/dev/null)"; then
echo "swarm-bao holds no ${tokenPath}, or is sealed/unreachable." >&2
echo "Keeping the token hive-matrix already has." >&2
exit 0
fi
if [ -z "$token" ]; then
echo "swarm-bao returned an empty ${tokenPath}; keeping the local token." >&2
exit 0
fi
umask 077
printf '%s\n' "$token" > ${lib.escapeShellArg (toString matrixCfg.registrationTokenFile)}
chmod 0600 ${lib.escapeShellArg (toString matrixCfg.registrationTokenFile)}
'';
};
};
}

View file

@ -1,542 +0,0 @@
# The swarm's secret store: one OpenBao for the whole swarm, in a
# `swarm-bao` nixos-container.
#
# Today every credential in ./swarm-*.nix is minted where it is read or copied
# there by a delivery unit (docs/swarm/secrets.md), which ties each secret's
# lifetime to its container's.
#
# ⚠️ It authenticates hive clients with a CLIENT CERTIFICATE, not with the
# swarm's SSO — a boot-order fact rather than a preference. This store holds
# authelia's own OIDC client secret, so a client that had to obtain an authelia
# token first could never start from cold. `swarm-nats` can lean on authelia
# precisely because it does not store authelia's credentials.
#
# ⚠️ NO GATEWAY VHOST, and unlike `swarm-nats` that is not because this speaks
# a non-HTTP protocol. It speaks HTTPS, so nginx *could* front it: **the client
# certificate IS the authentication**, and a terminating proxy strips it,
# leaving bao seeing nginx as the client for every hive in the swarm — one
# identity where there must be many. Reach is loopback plus whatever
# `deploy.bao.extraListenAddresses` names.
#
# ⚠️ THIS MODULE HAS NO OPINION ABOUT WHERE THE STORE'S IDENTITY COMES FROM.
# A store must not take its certificates from an authority it will itself
# distribute: reach the store to get the CA material, need a cert from that CA
# to reach the store. Service↔store mTLS is therefore its own trust domain,
# separate from the gateway's HTTPS certificates and from both CAs in this
# tree. The cert paths are inputs with no defaults and nothing here fills them
# in; whatever comes to mint that identity is what they will point at.
{
pkgs,
lib,
config,
...
}:
let
cfg = config.services.hyperhive.swarm.bao;
hyperhiveCfg = config.services.hyperhive;
deployCfg = hyperhiveCfg.deploy;
baoDeploy = deployCfg.bao;
networkCfg = hyperhiveCfg.network;
swarmDomain = hyperhiveCfg.swarm.domain;
# Two names for one location. `stateDir` is where openbao writes inside the
# container — upstream's own default, kept so its documentation matches. The
# host path is a sibling of the other swarm services' state rather than a
# path inside the container's tree, so `nixos-container destroy` cannot take
# the swarm's secrets with it.
stateDir = "/var/lib/openbao";
hostStateDir = "/var/lib/swarm-bao";
# The PIN is deliberately absent here. It arrives as `BAO_HSM_PIN` from an
# EnvironmentFile the provisioning unit writes, because a value interpolated
# into a nix expression renders world-readable into the store.
#
# With no seal stanza openbao falls back to Shamir, so this attrset being
# empty is the difference between a store that unseals itself and one that
# needs a human after every restart.
# The PKCS11 token store and its PINs live on the HOST and are bind-mounted
# in. Losing them loses the sealed store, so they are a host-level fact an
# operator can back up — the same reasoning as the raft data below.
tokenStoreDir = "/var/lib/swarm-bao-token";
pinEnvFile = "${tokenStoreDir}/pin.env";
sealSettings = lib.optionalAttrs (baoDeploy.seal == "pkcs11") {
seal.pkcs11 = {
lib = "${pkgs.tpm2-pkcs11}/lib/libtpm2_pkcs11.so";
token_label = "swarm-bao";
key_label = "swarm-bao-seal";
};
};
# Total on a null swarm domain for the same reason every sibling module is:
# the required-domain assertion in hive-network.nix should be what an operator
# sees, not a coercion error from here.
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
# Where the leaf lands for openbao to read. Inside the container, because
# `hostStateDir` is already bind-mounted at `stateDir` — so the delivery
# below needs no second mount, and nothing has to bind `tls.stateDir`,
# which holds the hive CA's private key.
serverCertPath = "${stateDir}/server.pem";
serverKeyPath = "${stateDir}/server-key.pem";
# The host-side sources, verbatim from the options — no fallback, because a
# fallback is exactly the CA opinion this module must not hold. The units
# below only exist when both are set (see `haveServerTls`), so these are
# never forced while null.
serverCertSrc = baoDeploy.serverCertFile;
serverKeySrc = baoDeploy.serverKeyFile;
# Both or neither: a certificate without its key configures a listener that
# cannot start, and the failure would surface as openbao refusing to boot
# rather than as the missing setting it is.
haveServerTls = baoDeploy.serverCertFile != null && baoDeploy.serverKeyFile != null;
# Every listener serves the same identity: they differ in which address
# they answer on, not in who they are. Client verification is separate and
# optional — a store with no `clientCaFile` still serves TLS, it just does
# not authenticate the far end, which is the honest rendering of "nobody
# has said what to trust yet".
listenerTls = {
tls_cert_file = serverCertPath;
tls_key_file = serverKeyPath;
}
// lib.optionalAttrs (baoDeploy.clientCaFile != null) {
tls_client_ca_file = clientCaPath;
tls_require_and_verify_client_cert = true;
};
clientCaPath = "${stateDir}/client-ca.pem";
extraListeners = lib.listToAttrs (
lib.imap1 (
i: addr:
lib.nameValuePair "extra-${toString i}" (
{
type = "tcp";
address = "${addr}:${toString cfg.port}";
}
// listenerTls
)
) baoDeploy.extraListenAddresses
);
# Loopback is unconditional and everything else is declared, which is not
# symmetry for its own sake:
#
# A reader on this host reaches the store through loopback, and the host
# running the store is always one of its readers — so loopback is a property
# of what the store IS, not of where it sits. Every other address depends on
# which network the hives that read it share, and that is a deployment fact.
# Bind only loopback and no remote hive can reach the store; bind only a
# shared-network address and an all-local swarm cannot reach its own.
#
# Neither is a superset of the other, which is why this is not one address
# with a conditional value.
listeners = {
loopback = {
type = "tcp";
address = "127.0.0.1:${toString cfg.port}";
}
// listenerTls;
}
// extraListeners;
in
{
# One service, two namespaces, and the split decides who may set what.
#
# `deploy.bao.*` is what the host RUNNING the store decides: whether to run
# it (`enable`, declared in ./deploy.nix with its siblings), which build,
# how the root key is sealed, what it listens on. None of it means anything
# on a host that only reads secrets.
#
# `swarm.bao.*` below is what every host in the swarm has to agree on — the
# name the store answers to, its port, its container. A host that is purely
# a *client* needs all of that, because it is how the client finds the store.
options.services.hyperhive.deploy.bao = {
package = lib.mkOption {
type = lib.types.package;
default = pkgs.openbao;
defaultText = lib.literalExpression "pkgs.openbao";
description = ''
OpenBao package to run.
An assertion below refuses 2.7.0 or newer, which drops the
built-in PKCS11 seal.
'';
};
seal = lib.mkOption {
type = lib.types.enum [
"pkcs11"
"shamir"
];
default = "pkcs11";
example = "shamir";
description = ''
How the store's root key is sealed.
`pkcs11` is the default and binds the key to the host's TPM: the store
unseals itself at boot, and an attacker with the disk does not get the
secrets. `shamir` is openbao's own default unseal keys held by
whoever ran `bao operator init`, entered by hand after every restart
and is the honest choice for a host with no TPM.
This is a **declaration**, and nothing at evaluation time can check
it: nix runs on the build machine and cannot see the target's TPM.
Saying `pkcs11` on a host without one fails at activation, when the
provisioning unit cannot create the token. That is deliberate a
store that comes up sealed by software while the config says hardware
is weaker than it reads, and silently so.
'';
};
serverCertFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-bao/server.pem";
description = ''
Certificate the store serves, covering
{option}`services.hyperhive.swarm.bao.domain`.
No default, and this module deliberately does not know what could
provide one for the same reason
{option}`services.hyperhive.deploy.bao.clientCaFile` doesn't. The
deployment names the file; the store never reaches for an authority.
A path, never a value.
'';
};
serverKeyFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-bao/server-key.pem";
description = ''
Private key for {option}`services.hyperhive.deploy.bao.serverCertFile`.
Both or neither a certificate with no key is a listener that cannot
start.
'';
};
clientCaFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-ca/root.pem";
description = ''
Authority the store validates hive **client** certificates against.
Deliberately has no default, and does not reach for the hive CA:
the hive CA is a future *consumer* of the store, so a store that
authenticated against it could not come up before the thing it
issues. It is a value someone points at the swarm root for a
swarm that runs one, an operator's own CA otherwise.
`null` leaves client-certificate verification off, which is only
appropriate where something else authenticates the connection.
'';
};
extraListenAddresses = lib.mkOption {
type = lib.types.listOf lib.types.str;
default = [ ];
example = [ "10.100.0.1" ];
description = ''
Addresses the store listens on **in addition to loopback**, each on
{option}`services.hyperhive.swarm.bao.port`.
Loopback is unconditional and not listed here: the host running the
store is always one of its readers. Every other address depends on
which network the reading hives share with this one, and that is a
deployment fact no other module's config can be read to infer a
swarm meshed over wireguard names its mesh address, one on a trusted
LAN names that interface, and an all-local swarm names nothing at all.
Addresses only, no port: a store reachable on two ports is a
misconfiguration rather than a topology.
'';
};
};
options.services.hyperhive.swarm.bao = {
machine = lib.mkOption {
type = lib.types.str;
readOnly = true;
default = "swarm-bao";
description = ''
Container name. Read-only: the name appears in host paths and in
`machinectl`, so it is a fact other modules may read rather than a knob.
'';
};
domain = lib.mkOption {
type = lib.types.str;
default = "bao.${domainBase}";
defaultText = lib.literalExpression ''"bao.''${services.hyperhive.swarm.domain}"'';
description = ''
Name the store is reached on. A **sibling** of the swarm's other
service names, not a child of any hive domain: an authority whose
`nameConstraints` permit one hive's domain cannot issue for a sibling
of it, so the shape of this name decides which authorities could ever
sign for the store. That is a property of the name, not a choice of
issuer this module makes no such choice.
'';
};
port = lib.mkOption {
type = lib.types.port;
default = 8200;
description = ''
TCP port the store listens on. Upstream's own default, kept so an
operator reading OpenBao documentation finds what they expect.
Swarm-wide because a client has to know it to reach the store, and
the same port on every listener: which *addresses* the store answers
on is the running host's business
({option}`services.hyperhive.deploy.bao.extraListenAddresses`), but
which port it answers on is something the whole swarm agrees.
'';
};
};
# ⚠️ Gated on `deploy.bao.enable`, and that is load-bearing rather than
# tidiness: an unconditional `config` block would evaluate the seal
# assertion on EVERY hive, so a hive that runs no secret store at all
# would fail to build the day nixpkgs moves openbao past 2.7.0. A check
# about running this service has no business firing where it is not run.
config = lib.mkMerge [
# Assertions sit in their own arm, gated only on running the store, so
# they still fire when the cert paths are unset — the arm below is not
# evaluated in that case, and an assertion that disappears exactly when
# its subject is broken would be worse than none.
(lib.mkIf (hyperhiveCfg.enable && deployCfg.bao.enable) {
assertions = [
{
assertion = lib.versionOlder baoDeploy.package.version "2.7.0";
message = ''
The swarm secret store needs openbao older than 2.7.0 (this is
${baoDeploy.package.version}). 2.7.0 moves the PKCS11 seal out of the
distribution into a plugin nixpkgs does not package, so the store
would come up sealed by software without saying so.
See https://openbao.org/community/deprecation/
'';
}
{
assertion = haveServerTls;
message = ''
The swarm secret store has no server certificate: set both
services.hyperhive.deploy.bao.serverCertFile and .serverKeyFile.
Nothing defaults them, on purpose a store must not take its
identity from an authority it will itself distribute, and
service-to-store mTLS is a separate trust domain from the
gateway's certificates and from either CA in this tree.
'';
}
];
})
(lib.mkIf (hyperhiveCfg.enable && deployCfg.bao.enable && haveServerTls) {
# Provisions the TPM-backed token the seal above names. One-shot and
# idempotent on ABSENCE, never on content: regenerating a PIN would
# orphan an already-sealed store, so a rebuild must not rotate it.
#
# ⚠️ Untestable without a TPM, and the failure is deliberately at
# activation — nix evaluates on the build machine and cannot see the
# target's hardware, so `seal = "pkcs11"` is a declaration this unit
# either makes true or fails on.
# The store's server certificate, delivered rather than bind-mounted.
#
# ⚠️ A copy, for three separate reasons — the last one is the one that
# matters most and is the least obvious:
# 1. `nixos-container` refuses to start when a bind source is missing,
# and a certificate minted on this same boot does not exist yet when
# the container is ordered. Same trap `hostClientSecretDir`
# documents in ./swarm-authelia.nix. Ordering against whatever
# mints it belongs with whatever named the path, not here.
# 2. `hostStateDir` is already mounted at `stateDir`, so a copy needs
# no second mount.
# 3. A directory holding a leaf usually holds the CA's private key
# beside it. Binding that directory to reach one file inside it
# would hand the container authority to mint any name that CA can —
# which is why this takes a path to a FILE and copies it.
systemd.services.swarm-bao-certs = {
description = "deliver the swarm secret store's server certificate";
before = [ "container@${cfg.machine}.service" ];
requiredBy = [ "container@${cfg.machine}.service" ];
path = [ pkgs.coreutils ];
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
script = ''
set -euo pipefail
install -d -m 0700 ${hostStateDir}
# Fail loudly rather than start a store that cannot serve. The path
# is configured, so a missing file means whatever was supposed to
# produce it did not run or failed — either way this is where it is
# cheapest to notice. Otherwise it surfaces at the TLS handshake,
# several layers from the setting that caused it.
for f in ${lib.escapeShellArg serverCertSrc} ${lib.escapeShellArg serverKeySrc}; do
if [ ! -s "$f" ]; then
echo "swarm-bao has no server certificate: $f is missing or empty." >&2
echo "That path comes from deploy.bao.serverCertFile/serverKeyFile." >&2
exit 1
fi
done
install -m 0644 ${lib.escapeShellArg serverCertSrc} ${hostStateDir}/server.pem
install -m 0600 ${lib.escapeShellArg serverKeySrc} ${hostStateDir}/server-key.pem
''
+ lib.optionalString (baoDeploy.clientCaFile != null) ''
if [ ! -s ${lib.escapeShellArg baoDeploy.clientCaFile} ]; then
echo "deploy.bao.clientCaFile names ${baoDeploy.clientCaFile}, which is missing or empty." >&2
exit 1
fi
install -m 0644 ${lib.escapeShellArg baoDeploy.clientCaFile} ${hostStateDir}/client-ca.pem
'';
};
systemd.services.swarm-bao-token = lib.mkIf (baoDeploy.seal == "pkcs11") {
description = "provision the swarm secret store's TPM-backed PKCS11 token";
before = [ "container@${cfg.machine}.service" ];
requiredBy = [ "container@${cfg.machine}.service" ];
path = [
pkgs.openssl
pkgs.tpm2-pkcs11
pkgs.tpm2-tools
];
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
script = ''
set -euo pipefail
install -d -m 0700 ${tokenStoreDir}
# Required whenever the store is not at its default location, or the
# library cannot find the token the seal asks for.
export TPM2_PKCS11_STORE=${tokenStoreDir}
# Absence is the only trigger. `openssl rand` is the same generator
# the grafana admin key uses; the value never passes through a nix
# expression, which would render it world-readable into the store.
for p in so-pin user-pin; do
if [ ! -e ${tokenStoreDir}/$p ]; then
( umask 077; openssl rand -hex 16 > ${tokenStoreDir}/$p )
chmod 0400 ${tokenStoreDir}/$p
fi
done
if [ ! -e ${tokenStoreDir}/tpm2_pkcs11.sqlite3 ]; then
pid=$(tpm2_ptool init --path ${tokenStoreDir} | sed -n 's/.*id: //p')
tpm2_ptool addtoken --path ${tokenStoreDir} --pid="$pid" \
--label=swarm-bao \
--sopin="$(cat ${tokenStoreDir}/so-pin)" \
--userpin="$(cat ${tokenStoreDir}/user-pin)"
# AES rather than RSA on purpose: openbao discussion 1826
# reports an RSA keypair here yielding duplicate labels, so
# `bao operator init` fails with "got more than 1 key for the
# label" and then CKR_GENERAL_ERROR. The seal supports AES-GCM.
tpm2_ptool addkey --path ${tokenStoreDir} --label=swarm-bao \
--userpin="$(cat ${tokenStoreDir}/user-pin)" \
--algorithm=aes256 --key-label=swarm-bao-seal
fi
( umask 077; printf 'BAO_HSM_PIN=%s\n' "$(cat ${tokenStoreDir}/user-pin)" > ${pinEnvFile} )
chmod 0400 ${pinEnvFile}
'';
};
containers.${cfg.machine} = {
autoStart = true;
ephemeral = false;
# Shared host netns, like every sibling swarm container. Unlike them the
# gateway is NOT the client here (see the no-vhost note at the top), so
# sharing the netns is what lets the store bind the host's own addresses
# rather than a convenience for nginx.
privateNetwork = false;
# Raft state outlives the container. `ephemeral = false` keeps the
# container's own /var, but a bind makes the store's data a host-level
# fact an operator can back up and a `nixos-container destroy` cannot
# take with it — which for the swarm's secrets is the difference between
# a rebuild and an outage.
bindMounts = {
${stateDir} = {
hostPath = hostStateDir;
isReadOnly = false;
};
}
// lib.optionalAttrs (baoDeploy.seal == "pkcs11") {
# Writable: the library keeps its sqlite store here, and the seal
# reads the token through it on every unseal.
${tokenStoreDir} = {
hostPath = tokenStoreDir;
isReadOnly = false;
};
};
# The seal talks to the TPM through the kernel's resource manager, so
# the device has to cross the container boundary or the store cannot
# unseal itself — which is the whole point of pkcs11 over shamir.
allowedDevices = lib.optionals (baoDeploy.seal == "pkcs11") [
{
node = "/dev/tpmrm0";
modifier = "rw";
}
];
config =
{ ... }:
{
imports = [
(import ./swarm-container-resolver.nix {
inherit (networkCfg) bridgeIp;
dnsConsumers = [ "openbao.service" ];
})
];
system.stateVersion = "26.05";
# Shares the host netns, so its own firewall.service would rewrite
# the HOST ruleset at every boot. The host firewall owns filtering.
networking.firewall.enable = false;
# The resolver unit imported above owns /etc/resolv.conf; leaving
# resolvconf on would let host-tracking regenerate it empty.
networking.resolvconf.enable = lib.mkForce false;
services.openbao = {
enable = true;
package = baoDeploy.package;
settings = {
listener = listeners;
storage.raft.path = stateDir;
}
// sealSettings;
};
# The PIN reaches openbao as an environment variable read from a
# 0400 file the provisioning unit wrote — never as a value in this
# expression, which would render it world-readable into the store.
# `TPM2_PKCS11_STORE` is required because the store is not at the
# library's default location.
systemd.services.openbao.serviceConfig = lib.mkIf (baoDeploy.seal == "pkcs11") {
EnvironmentFile = pinEnvFile;
Environment = [ "TPM2_PKCS11_STORE=${tokenStoreDir}" ];
};
# ⚠️ Upstream sets `restartIfChanged = false` on this unit, on
# purpose: a restart SEALS the store and disconnects every client.
# So a change to the settings above does NOT take effect on
# `nixos-rebuild switch` — it lands in the config file and waits.
# Restarting is an operator action with an unseal on the far side of
# it, which is why nothing here tries to be clever about it.
};
};
})
];
}

View file

@ -90,11 +90,4 @@ in
# hive that is not the service host is a *client* of it, not a second one.
config.services.hyperhive.deploy.victorialogs.enable =
lib.mkDefault swarmCfg.enableRequiredServices;
# The secret store. Once per swarm and optional, so it belongs to the
# same switch: a hive that does not run it is a *client*, reading its
# own secrets from whoever does. `mkDefault` is what keeps the store
# placeable on a host of its own — it can be set directly here and
# turned off wherever this switch happens to be on.
config.services.hyperhive.deploy.bao.enable = lib.mkDefault swarmCfg.enableRequiredServices;
}

View file

@ -71,26 +71,6 @@ let
bare = hive { };
withCi = hive { deploy.forgejo.ci.enable = true; };
baoPkcs11 = hive {
deploy.bao.enable = true;
deploy.bao.seal = "pkcs11";
};
baoShamir = hive {
deploy.bao.enable = true;
deploy.bao.seal = "shamir";
};
baoExplicitCerts = hive {
deploy.bao.enable = true;
deploy.bao.serverCertFile = "/etc/pki/bao.pem";
deploy.bao.serverKeyFile = "/etc/pki/bao-key.pem";
};
# The store and a service that reads from it, versus the store alone. The
# pair is what makes the reader's absence arm mean anything.
baoWithMatrix = hive {
deploy.bao.enable = true;
deploy.matrix.enable = true;
};
# A priority collision is a property of the *option*, not
# of the merged value's interior — nix throws the moment the value is
# demanded at all, so `seq`-ing each `serviceConfig` value to WHNF is
@ -154,57 +134,6 @@ let
name = "the CI container's unit definitions merge without a priority collision";
ok = forceCiServiceConfigs;
}
{
# The store's seal is spread over five gates — the stanza, the
# provisioning unit, a bind mount, a device and an EnvironmentFile.
# Rendering only some of them is the dangerous state: a store that
# says hardware-backed and seals with a software key, which no
# assertion can catch because every value is individually valid.
name = "a shamir store renders no TPM provisioning unit";
ok = !(baoShamir.systemd.services ? swarm-bao-token);
}
{
# Presence control for the case above. Without it, a typo in the
# option name would satisfy the absence arm forever.
name = "a pkcs11 store renders the TPM provisioning unit";
ok = baoPkcs11.systemd.services ? swarm-bao-token;
}
{
# The store's mTLS identity is a separate trust domain from both CAs in
# this tree, because it must not come from an authority the store will
# itself distribute. What supplies it is the glue, which mints a CA of
# the store's own — so an enabled store has all three paths, and if this
# ever reads null again the store stops coming up on its own.
name = "a deployed store is given its own certificate, key and client CA";
ok =
let
b = baoPkcs11.services.hyperhive.deploy.bao;
in
b.serverCertFile != null && b.serverKeyFile != null && b.clientCaFile != null;
}
{
# Everything the glue sets is `mkDefault`, and this is the case that
# says so: a deployment whose certificates come from somewhere the glue
# has never heard of must win. Also the presence control for the case
# above — a renamed option would read `null` on both and satisfy
# neither, but only this one names a value.
name = "an operator's own certificate path beats the glue's default";
ok = baoExplicitCerts.services.hyperhive.deploy.bao.serverCertFile == "/etc/pki/bao.pem";
}
{
# The store's first reader. Its unit belongs to the pairing, not to
# either service: matrix must not learn the store exists, and the store
# must not know who reads it.
name = "a store deployed beside the homeserver fetches its registration token";
ok = baoWithMatrix.systemd.services ? swarm-bao-matrix-token;
}
{
# Absence arm. A store with nothing to serve renders no reader, so the
# unit is a function of the PAIRING rather than of the store — which is
# the property that makes it glue instead of a feature of either side.
name = "a store with no homeserver beside it renders no token reader";
ok = !(baoPkcs11.systemd.services ? swarm-bao-matrix-token);
}
];
bad = builtins.filter (c: !c.ok) cases;