hyperhive/docs/getting-started/setup.md
atlas e9cec0da21 swarm-bao: write every swarm-* grant as a bao granter, not with a 24h token
Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).

A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.

swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.

The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.

module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.

Refs #4704
2026-09-27 22:57:46 +02:00

16 KiB

First-run setup (fresh-deploy bootstrap)

How to bring a fresh hyperhive hive online: provision accounts, open the gateway, bootstrap swarm SSO, make matrix reachable, and spawn the first sub-agents.

Aimed at ruth (the root/manager agent) on a fresh deploy, but it's a plain reference doc — read it whenever you need the bootstrap command sequence. All hivectl commands below run as root on the host (not inside an agent container); the request_* steps run from ruth's own turn via the MCP tools.

Bringing up a hive that doesn't host its own swarm services? Read swarm/secrets.md first. Everything below assumes each credential is generated where it's read, which is true on an all-local deploy and not otherwise — that page says which files an operator has to place, and where.

Step-by-step

1 · Forge

hive-c0re no longer creates agent forge users or mints agent tokens. swarm-controller does, for every agent that holds a store identity: a pass at start and every five minutes creates the forge user if it's missing, then mints the token into the swarm secret store, where the agent fetches it under that identity. An agent created with swarmctl agent create gets its identity then. Ruth doesn't: hive-c0re creates her on its own at startup, so she needs her identity minted by hand, once.

# On the swarm-controller host: give ruth her store identity. <hive> is the
# name of the hive she runs on.
swarmctl agent mint-identity ruth --hive <hive>

# On ruth's hive: re-apply her container config, which is when hive-c0re
# hands the new identity to the container.
hivectl agent ruth rebuild

You don't need to do anything else. The controller's next pass creates ruth's forge user and mints her token, and her container fetches it within about ten minutes. The same identity puts her in the matrix pass too (step 6). swarmctl agent mint-forge-token ruth skips the wait for the pass. A hive without a swarm secret store has no path to a forge token for ruth at all.

Swarm SSO creates the human operator's own forge account: the forge makes it on their first login through authelia, and swarmctl forge make-admin <you> then makes it a site admin — see Swarm SSO below.

2 · Gateway (HTTP Basic auth)

# Add an operator login to the gateway (reads password from stdin)
echo "hunter2" | hivectl gateway create-user mara --password-stdin

# List existing users
hivectl gateway list-users

3 · Secret store (only when deploy.bao)

⚠️ A sealed store still answers. OpenBao starts uninitialised and sealed, so the container is up and the port responds while every read times out — the failure looks like a hang, not like a store that was never initialised. Do this before you point anything at it.

Run this on the host. The bao there is a wrapper carrying this store's address, its CA, and the host's client certificate already, so nothing needs exporting:

sudo bao operator init   # keep the keys it prints and the root token OFF this host

sudo because the client certificate and key sit under /var/lib/swarm-bao-pki, which is mode 0700.

Inside the store's container (nixos-container root-login swarm-bao) the same command needs two extra pieces, because the certificate carries no IP SAN and its DNS name resolves to the bridge from in there: export BAO_ADDR=https://127.0.0.1:8200 alongside BAO_TLS_SERVER_NAME=bao.<swarm domain> to verify the name while connecting on loopback. The host is the shorter path.

While you still hold that root token, set up the granter: the one principal that writes every swarm-* policy and role from then on. Cert auth answers a role, so nothing can authenticate until some role exists. A short-lived bootstrap token breaks that cycle once, and it's the only step that needs the root token.

The policy it carries is nix/host-modules/bao-bootstrap-policy.hcl, shipped on the store's host at /etc/hyperhive/bao-bootstrap-policy.hcl. It covers the auth mounts and the granter's own policy and role, and nothing else. CI fails when the unit using the token needs a path it lacks.

sudo -i
read -rs BAO_TOKEN && export BAO_TOKEN          # paste the root token from `bao operator init`
bao policy write bao-bootstrap /etc/hyperhive/bao-bootstrap-policy.hcl
bao token create -policy=bao-bootstrap -ttl=24h -orphan -display-name=bao-bootstrap -field=token \
  | install -D -m 0600 /dev/stdin /var/lib/swarm-bao-bootstrap/grant.token
unset BAO_TOKEN
systemctl restart swarm-bao-granter-role

The token file is services.hyperhive.deploy.bao.bootstrapTokenFile, which all-local names for you. On a store host that isn't all-local, set it and rebuild first.

swarm-bao-granter-role runs on the host. It enables the cert auth method, writes the bao-granter policy, and creates the bao-granter role, which accepts the leaf /var/lib/swarm-bao-pki/granter.pem. Every swarm-bao-*-policy unit then logs in with that leaf. The controller's unit mounts the KV and pki engines and writes the swarm-controller role, and each sibling unit writes its own principal's policy and role. Every one runs on the host, because every API listener demands a client certificate and the host is the side that has one.

Confirm with systemctl status swarm-bao-granter-role, which should log Uploaded policy: bao-granter and Data written to: auth/cert/certs/bao-granter. Then restart the granting units that failed while they waited:

systemctl reset-failed 'swarm-bao-*-policy.service'
systemctl restart 'swarm-bao-*-policy.service'
systemctl status swarm-bao-controller-policy   # Uploaded policy, Data written to: auth/cert/certs/swarm-controller
rm /var/lib/swarm-bao-bootstrap/grant.token

⚠️ Don't reach for bao read auth/cert/… to check. The host's bao wrapper carries an address, a CA and a client certificate but deliberately no token, so that read answers 403 whether or not the role exists.

⏱️ Expect the first attempt to fail if you rebuilt into this. A rebuild restarts the store, and the units race it: the store answers local node not active until it finishes coming up. They retry every 30s for a day, so a sealed or late store heals itself.

Until you set up the granter, each swarm-bao-*-policy unit fails and logs the commands above. It never skips. Delete the token file only once swarm-bao-granter-role has succeeded. That unit skips while the file is absent, which is the steady state afterwards. The TTL above means a forgotten token expires rather than lingering.

After that, a new or changed swarm-* grant needs no operator step: the unit that writes it changes, and the deploy restarts it. A root step comes back only when the granter itself needs a path it lacks, such as a new mount.

⚠️ The granting units re-run on boot and whenever a deploy changes them, not on every deploy. When a grant drifts in the store and its unit stays the same, the next boot re-asserts it, not the next switch.

Upgrading a swarm set up with the older swarm-bootstrap policy

A store set up before the granter existed has every grant, but no bao-granter policy or role. After the deploy that introduces it, each swarm-bao-*-policy unit fails and logs the one-time step. Run the two blocks above as they stand. The old policy can go, with the root token again:

bao policy delete swarm-bootstrap

Residual risk, stated plainly. The granter is root-equivalent. It may write any swarm-* policy with any content, and attach it to a role that accepts any certificate; no bao ACL can constrain what a policy says. What bounds it:

  • nix/host-modules/swarm-bao.nix renders every policy it writes, and module-eval pins each principal's grants. Merging a change to that policy text is granting it: it takes effect on the next deploy with no bao step, so code review is the only gate.
  • Its key sits permanently at /var/lib/swarm-bao-pki/granter-key.pem, 0600 root in a 0700 directory, readable only by root units on the store's host. That host already holds ca-key.pem, which can mint a leaf with any subject, and controller-key.pem, whose policy is already root-equivalent. Root on that host gains nothing new.
  • Never copy granter-key.pem off the host the way operators copy the other leaves in that directory. That hands out root-equivalence.
  • Nothing revokes a stolen leaf on its own: the role trusts the CA plus the subject. Rotate the store's CA, or have root point the bao-granter role at a new deploy.bao.granterCommonName. Deleting granter{,-key}.pem and restarting swarm-bao-pki mints a new leaf, but doesn't invalidate the old one.

What else you need depends on services.hyperhive.deploy.bao.seal:

  • pkcs11 (the default) — pkcs11 binds the key to the host's TPM, and the store unseals itself on every restart. init is the only manual step.
  • shamir — no TPM, so run bao operator unseal again after every restart, with the keys init printed.

The store serves TLS, and on a hive that deploys it you need do nothing: a first-boot unit mints a CA of the store's own plus the leaves it signs — the store's server certificate, this host's client certificate, and one per service principal — and points deploy.bao.serverCertFile, .serverKeyFile and .clientCaFile at the store's half, .clientCertFile, .clientKeyFile and .serverCaFile at the reader's, and each principal's own pair at its own leaf.

Those are mkDefaults, so naming your own paths wins. Do that when your certificates come from a real internal CA; the store has no opinion about which. A hive that does not deploy the store names the reader's three itself, plus a pair for every principal it runs — see per-principal identities for the list. The operator issues those leaves out of band; they're the credentials the store can't hand you, being what opens it. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is mTLS between services and the store, a separate trust domain, because a store that took its identity from an authority it itself distributes could never come up before that authority.

4 · Swarm SSO (only when deploy.authelia)

⚠️ Required to finish the install, not optional. Authelia treats an empty user store as a fatal startup error, so until this runs the container crash-loops and auth.<swarm.domain> answers 502 Bad Gateway — a working vhost in front of an upstream that refuses to start. Skipping this step looks like a broken proxy.

# Runs as root on the host that RUNS authelia (not necessarily the
# controller host). Prints a generated password once — record it.
swarmctl user add mara --display-name Mara --email mara@example.com --group admins

⚠️ Keep --group admins. it's not decoration: operator-only surfaces (the swarm UI below) gate on that group, and an account without it authenticates successfully and is then refused — which reads like a broken login rather than a missing group.

If an account already exists without it, user add refuses rather than amends — adding the group afterwards is swarmctl user update mara --add-group admins.

Then sign in to the forge once through authelia, with that account. That first login creates your forge account, under the same username. Make it a site admin:

# On the swarm-controller's host. Fails until that first login has happened.
swarmctl forge make-admin mara

⚠️ Keep --email too. The forge won't create an account without an email: a subject that has none gets the forge's link-account page and no account. swarmctl user update mara --email … fixes it.

If the forge already has a local account with your username, the first SSO login asks for that account's forge password once, to link the two.

Detail, including what the password is and why this stays manual: swarm/sso.md.

5 · Swarm UI (only when deploy.swarm-ui, on by default with the controller)

Nothing to run — it's served on the swarm apex (https://<swarm.domain>/) as soon as the host rebuilds. Two things decide whether you can actually open it:

  • You are in admins (Swarm SSO above). The gateway asks authelia whether you have a session; the rule that makes it mean operator wants the group. Without it you log in and still get bounced.
  • The name resolves to this host. it's published to the hive's own resolver and to /etc/hosts when gateway.localHostsEntry is on; from anywhere else it needs a real DNS record like any other public name.

Detail, including why reachability is deliberately not the access control: swarm/ui.md.

6 · Matrix

# Ensure the appservice's sender account exists first
hivectl matrix sync-admin

# Invite the operator to the hive Space (and optionally to rooms)
hivectl matrix invite mara
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'

# Promote the operator to homeserver admin if needed
hivectl matrix promote-user mara

ruth's own matrix account comes from the swarm, like every agent's: swarm-controller creates it within five minutes of her holding a store identity (step 1), and her matrix daemon reads its token from the store. Without that identity she has no matrix account, and this hive no longer creates one.

Swarm SSO creates the human operator's own matrix account instead of a manual hivectl step — see Swarm SSO above (swarmctl user add).

7 · Spawn sub-agents

Sub-agent creation is an operator action — agents have no tool for it. Two steps:

# Step 1: scaffold the new agent's config repo. The swarm controller's
# InitAgentConfigRepo job does this (POST /api/agents), seeding
# /agents/iris/config/agent.nix from the default template.

# Step 2: edit /agents/iris/config/agent.nix and commit it. Then spawn
# iris from the dashboard (◆ R3QU3ST SP4WN / Spawn approval), which
# builds + starts the container from that config.

# Later config changes: open a PR on agent-configs/iris (hive-forge);
# the operator reviews + approves it — no MCP tool call.

See approvals.md for the full flow.

8 · Useful host commands

# Roster: all agents, status, rev, pending reminders
hivectl list-agents

# Restart a stuck container (no rebuild)
hivectl agent <agent> restart

# Open a Claude session inside an agent's container
hivectl agent <agent> choom

# Open hive web surfaces in a browser (or just print the URLs)
hivectl open           # operator dashboard
hivectl open forge     # Forgejo
hivectl open matrix    # Matrix GUI (fluffychat)

See tools/hivectl.md for every hivectl verb.

Security notes

  • No forge admin token is stored in any agent state dir. Agents hold a regular agent token, fetched from the swarm secret store into /run/hive-agent-forge-token/token (or, for an agent without a store identity, the forge-token file hive-c0re wrote before); sensitive creds (the core token) live on the host.
  • All config changes (forge PRs on agent-configs/<name>) go through operator approval — agents can't unilaterally rebuild containers, by design. See boundary.md and security.md.
  • Each hive authenticates its own telemetry ingest, and the hive label comes from which hive authenticated rather than from the payload — so no hive can report metrics as another. A first-run all-local hive gets this with nothing to configure; joining a swarm you don't host needs one secret copied across. See observability.md.

Once the hive is running, ruth records anything it needs to remember across restarts in /agents/ruth/state/notes.md.