Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/bao.md

6.6 KiB

The swarm secret store (OpenBao)

The swarm runs one OpenBao store, swarm-bao. Every swarm-level credential an agent or service needs is fetched from it under that principal's own certificate identity. What lives in it and who reads what: secrets.md. This page covers the store itself — how it comes up, who may write its grants, and how it's sealed and reached.

The operator steps (init, the one-time granter bootstrap) are in ../getting-started/setup.md; this page is the why behind them.

The granter

Cert auth answers a role, so nothing can authenticate until some role exists. A short-lived bootstrap token breaks that cycle once, and it's the only step that needs the root token.

The policy that token carries is nix/host-modules/bao-bootstrap-policy.hcl, shipped on the store's host at /etc/hyperhive/bao-bootstrap-policy.hcl. It covers the auth mounts and the granter's own policy and role, and nothing else. CI fails when the unit using the token needs a path it lacks. The token file is services.hyperhive.deploy.bao.bootstrapTokenFile, which all-local names for you; on a store host that isn't all-local, set it and rebuild first.

swarm-bao-granter-role runs on the host. It enables the cert and oidc auth methods (and disables approle if an older store still has it mounted), writes the bao-granter policy, and creates the bao-granter role, which accepts the leaf /var/lib/swarm-bao-pki/granter.pem. Every swarm-bao-*-policy unit then logs in with that leaf. The controller's unit mounts the KV and pki engines and writes the swarm-controller role, and each sibling unit writes its own principal's policy and role. Every one runs on the host, because every API listener but the loopback UI one demands a client certificate and the host is the side that has one.

Until the granter exists, each swarm-bao-*-policy unit fails and logs the bootstrap commands. It never skips. swarm-bao-granter-role skips while the token file is absent, which is the steady state afterwards; the token's TTL means a forgotten file expires rather than lingering.

After that, a new or changed swarm-* grant needs no operator step: the unit that writes it changes, and the deploy restarts it. A root step comes back only when the granter itself needs a path it lacks, such as a new mount.

⚠️ The granting units re-run on boot and whenever a deploy changes them, not on every deploy. When a grant drifts in the store and its unit stays the same, the next boot re-asserts it, not the next switch.

⚠️ Don't reach for bao read auth/cert/… to check a role. The host's bao wrapper carries an address, a CA and a client certificate but deliberately no token, so that read answers 403 whether or not the role exists. Read the unit's journal instead.

Upgrading a swarm set up with the older swarm-bootstrap policy

A store set up before the granter existed has every grant, but no bao-granter policy or role. After the deploy that introduces it, each swarm-bao-*-policy unit fails and logs the one-time step. Run the setup step as it stands. The old policy can go, with the root token again:

bao policy delete swarm-bootstrap

Residual risk

The granter is root-equivalent. It may write any swarm-* policy with any content, and attach it to a role that accepts any certificate; no bao ACL can constrain what a policy says. What bounds it:

  • nix/host-modules/swarm-bao.nix renders every policy it writes, and module-eval pins each principal's grants. Merging a change to that policy text is granting it: it takes effect on the next deploy with no bao step, so code review is the only gate.
  • Its key sits permanently at /var/lib/swarm-bao-pki/granter-key.pem, 0600 root in a 0700 directory, readable only by root units on the store's host. That host already holds ca-key.pem, which can mint a leaf with any subject, and controller-key.pem, whose policy is already root-equivalent. Root on that host gains nothing new.
  • Never copy granter-key.pem off the host the way operators copy the other leaves in that directory. That hands out root-equivalence.
  • Nothing revokes a stolen leaf on its own: the role trusts the CA plus the subject. Rotate the store's CA, or have root point the bao-granter role at a new deploy.bao.granterCommonName. Deleting granter{,-key}.pem and restarting swarm-bao-pki mints a new leaf, but doesn't invalidate the old one.

Sealing

services.hyperhive.deploy.bao.seal:

  • pkcs11 (the default) — binds the key to the host's TPM, and the store unseals itself on every restart. init is the only manual step.
  • shamir — no TPM, so run bao operator unseal again after every restart, with the keys init printed.

⚠️ A sealed store still answers. The container is up and the port responds while every read times out — the failure looks like a hang, not like a store that was never initialised.

TLS

The store serves TLS, and on a host that deploys it you need do nothing: a first-boot unit mints a CA of the store's own plus the leaves it signs — the store's server certificate, this host's client certificate, and one per service principal — and points deploy.bao.serverCertFile, .serverKeyFile and .clientCaFile at the store's half, .clientCertFile, .clientKeyFile and .serverCaFile at the reader's, and each principal's own pair at its own leaf.

Those are mkDefaults, so naming your own paths wins. Do that when your certificates come from a real internal CA; the store has no opinion about which. A host that does not deploy the store names the reader's three itself, plus a pair for every principal it runs — see per-principal identities for the list. The operator issues those leaves out of band; they're the credentials the store can't hand you, being what opens it.

⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is mTLS between services and the store, a separate trust domain, because a store that took its identity from an authority it itself distributes could never come up before that authority.

Inside the store's container (nixos-container root-login swarm-bao) the bao CLI needs two extra pieces, because the certificate carries no IP SAN and its DNS name resolves to the bridge from in there: export BAO_ADDR=https://127.0.0.1:8200 alongside BAO_TLS_SERVER_NAME=bao.<swarm domain>. The host's bao wrapper already carries the address, CA and client certificate, so the host is the shorter path.