Fix the 4 pre-existing vale error-level hits on main (docs/README.md:83, docs/getting-started/setup.md:11,127, docs/swarm/bao.md:4) that fail CI's prose-lint-errors job for every docs PR regardless of its own diff.
134 lines
6.6 KiB
Markdown
134 lines
6.6 KiB
Markdown
# The swarm secret store (OpenBao)
|
|
|
|
The swarm runs one OpenBao store, `swarm-bao`. Every swarm-level
|
|
credential an agent or service needs comes from it under that
|
|
principal's own certificate identity. What lives in it and who reads what:
|
|
[`secrets.md`](secrets.md). This page covers the store itself — how it
|
|
comes up, who may write its grants, and how it's sealed and reached.
|
|
|
|
The operator steps (init, the one-time granter bootstrap) are in
|
|
[`../getting-started/setup.md`](../getting-started/setup.md#1--secret-store);
|
|
this page is the why behind them.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
## The granter
|
|
|
|
Cert auth answers a _role_, so nothing can authenticate until some role
|
|
exists. A short-lived bootstrap token breaks that cycle once, and it's the
|
|
only step that needs the root token.
|
|
|
|
The policy that token carries is `nix/host-modules/bao-bootstrap-policy.hcl`,
|
|
shipped on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It
|
|
covers the auth mounts and the granter's own policy and role, and nothing
|
|
else. CI fails when the unit using the token needs a path it lacks. The token
|
|
file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which all-local
|
|
names for you; on a store host that isn't all-local, set it and rebuild first.
|
|
|
|
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
|
|
auth methods (and disables `approle` if an older store still has it mounted),
|
|
writes the `bao-granter` policy, and creates the `bao-granter` role, which
|
|
accepts the leaf `/var/lib/swarm-bao-pki/granter.pem`. Every
|
|
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
|
|
mounts the KV and pki engines and writes the `swarm-controller` role, and each
|
|
sibling unit writes its own principal's policy and role. Every one runs on the
|
|
host, because every API listener but the loopback UI one demands a client
|
|
certificate and the host is the side that has one.
|
|
|
|
Until the granter exists, each `swarm-bao-*-policy` unit **fails** and logs
|
|
the bootstrap commands. It never skips. `swarm-bao-granter-role` skips while
|
|
the token file is absent, which is the steady state afterwards; the token's
|
|
TTL means a forgotten file expires rather than lingering.
|
|
|
|
After that, a new or changed `swarm-*` grant needs no operator step: the unit
|
|
that writes it changes, and the deploy restarts it. A root step comes back only
|
|
when the granter itself needs a path it lacks, such as a new mount.
|
|
|
|
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
|
|
them, not on every deploy. When a grant drifts in the store and its unit stays
|
|
the same, the next boot re-asserts it, not the next switch.
|
|
|
|
⚠️ Don't reach for `bao read auth/cert/…` to check a role. The host's `bao`
|
|
wrapper carries an address, a CA and a client certificate but deliberately
|
|
**no token**, so that read answers `403` whether or not the role exists. Read
|
|
the unit's journal instead.
|
|
|
|
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
|
|
|
|
A store set up before the granter existed has every grant, but no `bao-granter`
|
|
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
|
|
unit fails and logs the one-time step. Run the setup step as it stands. The
|
|
old policy can go, with the root token again:
|
|
|
|
```bash
|
|
bao policy delete swarm-bootstrap
|
|
```
|
|
|
|
</details>
|
|
|
|
### Residual risk
|
|
|
|
The granter is root-equivalent. It may write any `swarm-*` policy with any
|
|
content, and attach it to a role that accepts any certificate; no bao ACL can
|
|
constrain what a policy says. What bounds it:
|
|
|
|
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
|
|
module-eval pins each principal's grants. **Merging a change to that policy
|
|
text is granting it**: it takes effect on the next deploy with no bao step,
|
|
so code review is the only gate.
|
|
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
|
|
root in a `0700` directory, readable only by root units on the store's host.
|
|
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
|
|
and `controller-key.pem`, whose policy is already root-equivalent. Root on
|
|
that host gains nothing new.
|
|
- **Never copy `granter-key.pem` off the host** the way operators copy the
|
|
other leaves in that directory. That hands out root-equivalence.
|
|
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
|
|
subject. Rotate the store's CA, or have root point the `bao-granter` role at
|
|
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
|
|
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
|
|
one.
|
|
|
|
## Sealing
|
|
|
|
`services.hyperhive.deploy.bao.seal`:
|
|
|
|
- **`pkcs11`** (the default) — binds the key to the host's TPM, and the store
|
|
unseals itself on every restart. `init` is the only manual step.
|
|
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
|
|
restart, with the keys `init` printed.
|
|
|
|
⚠️ **A sealed store still answers.** The container is up and the port
|
|
responds while every read times out — the failure looks like a hang, not like
|
|
a store that was never initialised.
|
|
|
|
## TLS
|
|
|
|
The store serves TLS, and on a host that deploys it you need do nothing: a
|
|
first-boot unit mints a CA of the store's own plus the leaves it signs — the
|
|
store's server certificate, this host's client certificate, and one per service
|
|
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
|
|
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
|
|
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
|
|
|
|
Those are `mkDefault`s, so naming your own paths wins. Do that when your
|
|
certificates come from a real internal CA; the store has no opinion about
|
|
which. A host that does **not** deploy the store names the reader's three
|
|
itself, plus a pair for every principal it runs — see
|
|
[per-principal identities](secrets.md#per-principal-identities) for the list.
|
|
The operator issues those leaves out of band; they're the credentials the
|
|
store can't hand you, being what opens it.
|
|
|
|
⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is **mTLS
|
|
between services and the store**, a separate trust domain, because a store
|
|
that took its identity from an authority it itself distributes could never
|
|
come up before that authority.
|
|
|
|
Inside the store's container (`nixos-container root-login swarm-bao`) the
|
|
`bao` CLI needs two extra pieces, because the certificate carries no IP SAN
|
|
and its DNS name resolves to the bridge from in there: export
|
|
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
|
|
domain>`. The host's `bao` wrapper already carries the address, CA and client
|
|
certificate, so the host is the shorter path.
|
|
|
|
<!-- vale write-good.Passive = YES -->
|