swarm-bao: write every swarm-* grant as a bao granter, not with a 24h token

Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).

A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.

swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.

The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.

module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.

Refs #4704
This commit is contained in:
atlas 2026-09-27 03:33:38 +02:00 • committed by mara
commit e9cec0da21
12 changed files with 898 additions and 443 deletions

View file

@ -88,66 +88,108 @@ its DNS name resolves to the bridge from in there: export
domain>` to verify the name while connecting on loopback. The host is the
shorter path.
While you still hold that root token, mint the one credential the swarm needs
to grant itself anything. Cert auth answers a _role_, so nothing can
authenticate until some role exists — this token is what breaks that cycle,
and it's the only step that needs the root token.
While you still hold that root token, set up the **granter**: the one principal
that writes every `swarm-*` policy and role from then on. Cert auth answers a
_role_, so nothing can authenticate until some role exists. A short-lived
bootstrap token breaks that cycle once, and it's the only step that needs the
root token.
The policy is `nix/host-modules/swarm-bao-bootstrap-policy.hcl` in this
repository, and CI fails when a unit using the token needs a path it lacks.
The policy it carries is `nix/host-modules/bao-bootstrap-policy.hcl`, shipped
on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It covers
the auth mounts and the granter's own policy and role, and nothing else. CI
fails when the unit using the token needs a path it lacks.
```bash
# The policy file, copied to wherever you run `bao`.
bao policy write swarm-bootstrap swarm-bao-bootstrap-policy.hcl
# A token holding it. `-orphan` so it outlives the session that made it.
bao token create -policy=swarm-bootstrap -ttl=24h -orphan -display-name=swarm-bootstrap
sudo -i
read -rs BAO_TOKEN && export BAO_TOKEN # paste the root token from `bao operator init`
bao policy write bao-bootstrap /etc/hyperhive/bao-bootstrap-policy.hcl
bao token create -policy=bao-bootstrap -ttl=24h -orphan -display-name=bao-bootstrap -field=token \
| install -D -m 0600 /dev/stdin /var/lib/swarm-bao-bootstrap/grant.token
unset BAO_TOKEN
systemctl restart swarm-bao-granter-role
```
Put the token's value at `services.hyperhive.deploy.bao.bootstrapTokenFile`
(all-local names that path for you), then rebuild. A one-shot unit **on the
host** reads it, writes the `swarm-controller` policy, enables the cert auth
method, mounts the KV engine the controller stores credentials in, and creates
the `swarm-controller` role that attaches policy to certificate. It runs there
because every API listener demands a client certificate, and the host is the
side that has one.
The token file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which
all-local names for you. On a store host that isn't all-local, set it and
rebuild first.
`swarm-bao-granter-role` runs **on the host**. It enables the cert auth
method, writes the `bao-granter` policy, and creates the `bao-granter` role,
which accepts the leaf `/var/lib/swarm-bao-pki/granter.pem`. Every
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
mounts the KV and pki engines and writes the `swarm-controller` role, and each
sibling unit writes its own principal's policy and role. Every one runs on the
host, because every API listener demands a client certificate and the host is
the side that has one.
**Confirm with `systemctl status swarm-bao-granter-role`**, which should log
`Uploaded policy: bao-granter` and `Data written to: auth/cert/certs/bao-granter`.
Then restart the granting units that failed while they waited:
```bash
systemctl reset-failed 'swarm-bao-*-policy.service'
systemctl restart 'swarm-bao-*-policy.service'
systemctl status swarm-bao-controller-policy # Uploaded policy, Data written to: auth/cert/certs/swarm-controller
rm /var/lib/swarm-bao-bootstrap/grant.token
```
⚠️ Don't reach for `bao read auth/cert/…` to check. The host's `bao`
wrapper carries an address, a CA and a client certificate but deliberately
**no token**, so that read answers `403` whether or not the role exists.
⏱️ **Expect the first attempt to fail if you rebuilt into this.** A rebuild
restarts the store, and the unit races it — the store answers `local node not
active` until it finishes coming up. It retries every 30s and the second
attempt is the one that usually lands. Nothing to do.
restarts the store, and the units race it: the store answers `local node not
active` until it finishes coming up. They retry every 30s for a day, so a
sealed or late store heals itself.
**Confirm with `systemctl status swarm-bao-controller-policy`**, which wants no
token — a successful run logs `Uploaded policy`, `Enabled cert auth method` and
`Data written to: auth/cert/certs/swarm-controller`. ⚠️ Do _not_ reach for `bao
read auth/cert/…` to check: the host's `bao` wrapper carries an address, a CA
and a client certificate but deliberately **no token**, so that read answers
`403` whether or not the role exists.
Until you set up the granter, each `swarm-bao-*-policy` unit **fails** and logs
the commands above. It never skips. Delete the token file only once
`swarm-bao-granter-role` has succeeded. That unit skips while the file is
absent, which is the steady state afterwards. The TTL above means a forgotten
token expires rather than lingering.
**Delete the token file only once that unit has succeeded.** It skips when the
token is absent, so a host that has finished bootstrapping stops carrying the
credential — but deleting it before the role
exists leaves the unit skipping forever with nothing to show for it, and looks
exactly like a store that was never bootstrapped. The TTL above means a
forgotten one expires rather than lingering.
After that, a new or changed `swarm-*` grant needs no operator step: the unit
that writes it changes, and the deploy restarts it. A root step comes back only
when the granter itself needs a path it lacks, such as a new mount.
<details><summary>Already bootstrapped before the KV mount existed?</summary>
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
them, not on every deploy. When a grant drifts in the store and its unit stays the
same, the next boot re-asserts it, not the next switch.
A store bootstrapped by an earlier version has the policy, the auth method and
the role, but no `secret/` engine — the controller's first credential write
answers `no handler for route "secret/data/…"`. The bootstrap token can't fix it
either: the policy that minted it names nothing under `sys/mounts`. Mount it
once with the root token from `init`:
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
A store set up before the granter existed has every grant, but no `bao-granter`
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
unit fails and logs the one-time step. Run the two blocks above as they stand.
The old policy can go, with the root token again:
```bash
sudo bash -c 'BAO_TOKEN="<root token>" bao secrets enable -path=secret kv-v2'
bao policy delete swarm-bootstrap
```
No rebuild needed — the unit's own check finds the mount on its next run and
leaves it alone.
</details>
**Residual risk, stated plainly.** The granter is root-equivalent. It may write
any `swarm-*` policy with any content, and attach it to a role that accepts any
certificate; no bao ACL can constrain what a policy says. What bounds it:
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
module-eval pins each principal's grants. **Merging a change to that policy
text is granting it**: it takes effect on the next deploy with no bao step,
so code review is the only gate.
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
root in a `0700` directory, readable only by root units on the store's host.
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
and `controller-key.pem`, whose policy is already root-equivalent. Root on
that host gains nothing new.
- **Never copy `granter-key.pem` off the host** the way operators copy the
other leaves in that directory. That hands out root-equivalence.
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
subject. Rotate the store's CA, or have root point the `bao-granter` role at
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
one.
What else you need depends on
`services.hyperhive.deploy.bao.seal`: