Fix the 4 pre-existing vale error-level hits on main (docs/README.md:83, docs/getting-started/setup.md:11,127, docs/swarm/bao.md:4) that fail CI's prose-lint-errors job for every docs PR regardless of its own diff.
6.6 KiB
The swarm secret store (OpenBao)
The swarm runs one OpenBao store, swarm-bao. Every swarm-level
credential an agent or service needs comes from it under that
principal's own certificate identity. What lives in it and who reads what:
secrets.md. This page covers the store itself — how it
comes up, who may write its grants, and how it's sealed and reached.
The operator steps (init, the one-time granter bootstrap) are in
../getting-started/setup.md;
this page is the why behind them.
The granter
Cert auth answers a role, so nothing can authenticate until some role exists. A short-lived bootstrap token breaks that cycle once, and it's the only step that needs the root token.
The policy that token carries is nix/host-modules/bao-bootstrap-policy.hcl,
shipped on the store's host at /etc/hyperhive/bao-bootstrap-policy.hcl. It
covers the auth mounts and the granter's own policy and role, and nothing
else. CI fails when the unit using the token needs a path it lacks. The token
file is services.hyperhive.deploy.bao.bootstrapTokenFile, which all-local
names for you; on a store host that isn't all-local, set it and rebuild first.
swarm-bao-granter-role runs on the host. It enables the cert and oidc
auth methods (and disables approle if an older store still has it mounted),
writes the bao-granter policy, and creates the bao-granter role, which
accepts the leaf /var/lib/swarm-bao-pki/granter.pem. Every
swarm-bao-*-policy unit then logs in with that leaf. The controller's unit
mounts the KV and pki engines and writes the swarm-controller role, and each
sibling unit writes its own principal's policy and role. Every one runs on the
host, because every API listener but the loopback UI one demands a client
certificate and the host is the side that has one.
Until the granter exists, each swarm-bao-*-policy unit fails and logs
the bootstrap commands. It never skips. swarm-bao-granter-role skips while
the token file is absent, which is the steady state afterwards; the token's
TTL means a forgotten file expires rather than lingering.
After that, a new or changed swarm-* grant needs no operator step: the unit
that writes it changes, and the deploy restarts it. A root step comes back only
when the granter itself needs a path it lacks, such as a new mount.
⚠️ The granting units re-run on boot and whenever a deploy changes them, not on every deploy. When a grant drifts in the store and its unit stays the same, the next boot re-asserts it, not the next switch.
⚠️ Don't reach for bao read auth/cert/… to check a role. The host's bao
wrapper carries an address, a CA and a client certificate but deliberately
no token, so that read answers 403 whether or not the role exists. Read
the unit's journal instead.
Upgrading a swarm set up with the older swarm-bootstrap policy
A store set up before the granter existed has every grant, but no bao-granter
policy or role. After the deploy that introduces it, each swarm-bao-*-policy
unit fails and logs the one-time step. Run the setup step as it stands. The
old policy can go, with the root token again:
bao policy delete swarm-bootstrap
Residual risk
The granter is root-equivalent. It may write any swarm-* policy with any
content, and attach it to a role that accepts any certificate; no bao ACL can
constrain what a policy says. What bounds it:
nix/host-modules/swarm-bao.nixrenders every policy it writes, and module-eval pins each principal's grants. Merging a change to that policy text is granting it: it takes effect on the next deploy with no bao step, so code review is the only gate.- Its key sits permanently at
/var/lib/swarm-bao-pki/granter-key.pem,0600root in a0700directory, readable only by root units on the store's host. That host already holdsca-key.pem, which can mint a leaf with any subject, andcontroller-key.pem, whose policy is already root-equivalent. Root on that host gains nothing new. - Never copy
granter-key.pemoff the host the way operators copy the other leaves in that directory. That hands out root-equivalence. - Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
subject. Rotate the store's CA, or have root point the
bao-granterrole at a newdeploy.bao.granterCommonName. Deletinggranter{,-key}.pemand restartingswarm-bao-pkimints a new leaf, but doesn't invalidate the old one.
Sealing
services.hyperhive.deploy.bao.seal:
pkcs11(the default) — binds the key to the host's TPM, and the store unseals itself on every restart.initis the only manual step.shamir— no TPM, so runbao operator unsealagain after every restart, with the keysinitprinted.
⚠️ A sealed store still answers. The container is up and the port responds while every read times out — the failure looks like a hang, not like a store that was never initialised.
TLS
The store serves TLS, and on a host that deploys it you need do nothing: a
first-boot unit mints a CA of the store's own plus the leaves it signs — the
store's server certificate, this host's client certificate, and one per service
principal — and points deploy.bao.serverCertFile, .serverKeyFile and
.clientCaFile at the store's half, .clientCertFile, .clientKeyFile and
.serverCaFile at the reader's, and each principal's own pair at its own leaf.
Those are mkDefaults, so naming your own paths wins. Do that when your
certificates come from a real internal CA; the store has no opinion about
which. A host that does not deploy the store names the reader's three
itself, plus a pair for every principal it runs — see
per-principal identities for the list.
The operator issues those leaves out of band; they're the credentials the
store can't hand you, being what opens it.
⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is mTLS between services and the store, a separate trust domain, because a store that took its identity from an authority it itself distributes could never come up before that authority.
Inside the store's container (nixos-container root-login swarm-bao) the
bao CLI needs two extra pieces, because the certificate carries no IP SAN
and its DNS name resolves to the bridge from in there: export
BAO_ADDR=https://127.0.0.1:8200 alongside BAO_TLS_SERVER_NAME=bao.<swarm domain>. The host's bao wrapper already carries the address, CA and client
certificate, so the host is the shorter path.