Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/bao.md
atlas 9d804ae094 docs: fix vale errors on main
Fix the 4 pre-existing vale error-level hits on main (docs/README.md:83,
docs/getting-started/setup.md:11,127, docs/swarm/bao.md:4) that fail CI's
prose-lint-errors job for every docs PR regardless of its own diff.
2026-10-01 23:34:47 +02:00

134 lines
6.6 KiB
Markdown

# The swarm secret store (OpenBao)
The swarm runs one OpenBao store, `swarm-bao`. Every swarm-level
credential an agent or service needs comes from it under that
principal's own certificate identity. What lives in it and who reads what:
[`secrets.md`](secrets.md). This page covers the store itself — how it
comes up, who may write its grants, and how it's sealed and reached.
The operator steps (init, the one-time granter bootstrap) are in
[`../getting-started/setup.md`](../getting-started/setup.md#1--secret-store);
this page is the why behind them.
<!-- vale write-good.Passive = NO -->
## The granter
Cert auth answers a _role_, so nothing can authenticate until some role
exists. A short-lived bootstrap token breaks that cycle once, and it's the
only step that needs the root token.
The policy that token carries is `nix/host-modules/bao-bootstrap-policy.hcl`,
shipped on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It
covers the auth mounts and the granter's own policy and role, and nothing
else. CI fails when the unit using the token needs a path it lacks. The token
file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which all-local
names for you; on a store host that isn't all-local, set it and rebuild first.
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
auth methods (and disables `approle` if an older store still has it mounted),
writes the `bao-granter` policy, and creates the `bao-granter` role, which
accepts the leaf `/var/lib/swarm-bao-pki/granter.pem`. Every
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
mounts the KV and pki engines and writes the `swarm-controller` role, and each
sibling unit writes its own principal's policy and role. Every one runs on the
host, because every API listener but the loopback UI one demands a client
certificate and the host is the side that has one.
Until the granter exists, each `swarm-bao-*-policy` unit **fails** and logs
the bootstrap commands. It never skips. `swarm-bao-granter-role` skips while
the token file is absent, which is the steady state afterwards; the token's
TTL means a forgotten file expires rather than lingering.
After that, a new or changed `swarm-*` grant needs no operator step: the unit
that writes it changes, and the deploy restarts it. A root step comes back only
when the granter itself needs a path it lacks, such as a new mount.
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
them, not on every deploy. When a grant drifts in the store and its unit stays
the same, the next boot re-asserts it, not the next switch.
⚠️ Don't reach for `bao read auth/cert/…` to check a role. The host's `bao`
wrapper carries an address, a CA and a client certificate but deliberately
**no token**, so that read answers `403` whether or not the role exists. Read
the unit's journal instead.
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
A store set up before the granter existed has every grant, but no `bao-granter`
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
unit fails and logs the one-time step. Run the setup step as it stands. The
old policy can go, with the root token again:
```bash
bao policy delete swarm-bootstrap
```
</details>
### Residual risk
The granter is root-equivalent. It may write any `swarm-*` policy with any
content, and attach it to a role that accepts any certificate; no bao ACL can
constrain what a policy says. What bounds it:
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
module-eval pins each principal's grants. **Merging a change to that policy
text is granting it**: it takes effect on the next deploy with no bao step,
so code review is the only gate.
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
root in a `0700` directory, readable only by root units on the store's host.
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
and `controller-key.pem`, whose policy is already root-equivalent. Root on
that host gains nothing new.
- **Never copy `granter-key.pem` off the host** the way operators copy the
other leaves in that directory. That hands out root-equivalence.
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
subject. Rotate the store's CA, or have root point the `bao-granter` role at
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
one.
## Sealing
`services.hyperhive.deploy.bao.seal`:
- **`pkcs11`** (the default) — binds the key to the host's TPM, and the store
unseals itself on every restart. `init` is the only manual step.
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
restart, with the keys `init` printed.
⚠️ **A sealed store still answers.** The container is up and the port
responds while every read times out — the failure looks like a hang, not like
a store that was never initialised.
## TLS
The store serves TLS, and on a host that deploys it you need do nothing: a
first-boot unit mints a CA of the store's own plus the leaves it signs — the
store's server certificate, this host's client certificate, and one per service
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
Those are `mkDefault`s, so naming your own paths wins. Do that when your
certificates come from a real internal CA; the store has no opinion about
which. A host that does **not** deploy the store names the reader's three
itself, plus a pair for every principal it runs — see
[per-principal identities](secrets.md#per-principal-identities) for the list.
The operator issues those leaves out of band; they're the credentials the
store can't hand you, being what opens it.
⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is **mTLS
between services and the store**, a separate trust domain, because a store
that took its identity from an authority it itself distributes could never
come up before that authority.
Inside the store's container (`nixos-container root-login swarm-bao`) the
`bao` CLI needs two extra pieces, because the certificate carries no IP SAN
and its DNS name resolves to the bridge from in there: export
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
domain>`. The host's `bao` wrapper already carries the address, CA and client
certificate, so the host is the shorter path.
<!-- vale write-good.Passive = YES -->