Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/getting-started/setup.md
atlas ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00

358 lines
16 KiB
Markdown

# First-run setup (fresh-deploy bootstrap)
How to bring a fresh hyperhive hive online: provision accounts, open
the gateway, bootstrap swarm SSO, make matrix reachable, and spawn the
first sub-agents.
Aimed at `ruth` (the root/manager agent) on a fresh deploy, but it's a
plain reference doc — read it whenever you need the bootstrap command
sequence. All `hivectl` commands below run as **root on the host** (not
inside an agent container); the `request_*` steps run from ruth's own
turn via the MCP tools.
<!-- vale write-good.Passive = NO -->
**Bringing up a hive that doesn't host its own swarm services?** Read
[`swarm/secrets.md`](../swarm/secrets.md) first. Everything below assumes
each credential is generated where it's read, which is true on an
all-local deploy and not otherwise — that page says which files an
operator has to place, and where.
<!-- vale write-good.Passive = YES -->
## Step-by-step
### 1 · Forge
hive-c0re no longer creates agent forge users or mints agent tokens.
swarm-controller does, for every agent that holds a store identity: a pass
at start and every five minutes creates the forge user if it's missing,
then mints the token into the swarm secret store, where the agent fetches
it under that identity. An agent created with `swarmctl agent create` gets
its identity then. Ruth doesn't: hive-c0re creates her on its own at
startup, so she needs her identity minted by hand, once.
```bash
# On the swarm-controller host: give ruth her store identity.
swarmctl agent mint-identity ruth
# On ruth's hive: re-apply her container config, which is when hive-c0re
# hands the new identity to the container.
hivectl agent ruth rebuild
```
You don't need to do anything else. The controller's next pass creates
ruth's forge user and mints her token, and her container fetches it within
about ten minutes. The same identity puts her in the matrix pass too (step
6). `swarmctl agent mint-forge-token ruth` skips the wait for
the pass. A hive without a swarm secret store has no path to a forge token
for ruth at all.
Swarm SSO creates the human operator's own forge account: the forge
makes it on their first login through authelia, and `swarmctl forge
make-admin <you>` then makes it a site admin — see _Swarm SSO_ below.
### 2 · Gateway (HTTP Basic auth)
```bash
# Add an operator login to the gateway (reads password from stdin)
echo "hunter2" | hivectl gateway create-user mara --password-stdin
# List existing users
hivectl gateway list-users
```
### 3 · Secret store (only when `deploy.bao`)
⚠️ **A sealed store still answers.** OpenBao starts uninitialised and
sealed, so the container is up and the port responds while every read
times out — the failure looks like a hang, not like a store that was
never initialised. Do this before you point anything at it.
Run this **on the host**. The `bao` there is a wrapper carrying this store's
address, its CA, and the host's client certificate already, so nothing needs
exporting:
```bash
sudo bao operator init # keep the keys it prints and the root token OFF this host
```
`sudo` because the client certificate and key sit under
`/var/lib/swarm-bao-pki`, which is mode `0700`.
Inside the store's container (`nixos-container root-login swarm-bao`) the same
command needs two extra pieces, because the certificate carries no IP SAN and
its DNS name resolves to the bridge from in there: export
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
domain>` to verify the name while connecting on loopback. The host is the
shorter path.
While you still hold that root token, set up the **granter**: the one principal
that writes every `swarm-*` policy and role from then on. Cert auth answers a
_role_, so nothing can authenticate until some role exists. A short-lived
bootstrap token breaks that cycle once, and it's the only step that needs the
root token.
The policy it carries is `nix/host-modules/bao-bootstrap-policy.hcl`, shipped
on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It covers
the auth mounts and the granter's own policy and role, and nothing else. CI
fails when the unit using the token needs a path it lacks.
```bash
sudo -i
read -rs BAO_TOKEN && export BAO_TOKEN # paste the root token from `bao operator init`
bao policy write bao-bootstrap /etc/hyperhive/bao-bootstrap-policy.hcl
bao token create -policy=bao-bootstrap -ttl=24h -orphan -display-name=bao-bootstrap -field=token \
| install -D -m 0600 /dev/stdin /var/lib/swarm-bao-bootstrap/grant.token
unset BAO_TOKEN
systemctl restart swarm-bao-granter-role
```
The token file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which
all-local names for you. On a store host that isn't all-local, set it and
rebuild first.
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
auth methods (and disables `approle` if an older store still has it mounted),
writes the `bao-granter` policy, and creates the
`bao-granter` role, which accepts the leaf
`/var/lib/swarm-bao-pki/granter.pem`. Every
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
mounts the KV and pki engines and writes the `swarm-controller` role, and each
sibling unit writes its own principal's policy and role. Every one runs on the
host, because every API listener but the loopback UI one demands a client
certificate and the host is the side that has one.
**Confirm with `systemctl status swarm-bao-granter-role`**, which should log
`Uploaded policy: bao-granter` and `Data written to: auth/cert/certs/bao-granter`.
Then restart the granting units that failed while they waited. These two
names cover every unit that logs in as the granter, and CI fails when one
doesn't:
```bash
systemctl reset-failed 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
systemctl restart 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
systemctl status swarm-bao-controller-policy # Uploaded policy, Data written to: auth/cert/certs/swarm-controller
rm /var/lib/swarm-bao-bootstrap/grant.token
```
⚠️ Don't reach for `bao read auth/cert/…` to check. The host's `bao`
wrapper carries an address, a CA and a client certificate but deliberately
**no token**, so that read answers `403` whether or not the role exists.
⏱️ **Expect the first attempt to fail if you rebuilt into this.** A rebuild
restarts the store, and the units race it: the store answers `local node not
active` until it finishes coming up. They retry every 30s for a day, so a
sealed or late store heals itself.
Until you set up the granter, each `swarm-bao-*-policy` unit **fails** and logs
the commands above. It never skips. Delete the token file only once
`swarm-bao-granter-role` has succeeded. That unit skips while the file is
absent, which is the steady state afterwards. The TTL above means a forgotten
token expires rather than lingering.
After that, a new or changed `swarm-*` grant needs no operator step: the unit
that writes it changes, and the deploy restarts it. A root step comes back only
when the granter itself needs a path it lacks, such as a new mount.
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
them, not on every deploy. When a grant drifts in the store and its unit stays the
same, the next boot re-asserts it, not the next switch.
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
A store set up before the granter existed has every grant, but no `bao-granter`
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
unit fails and logs the one-time step. Run the two blocks above as they stand.
The old policy can go, with the root token again:
```bash
bao policy delete swarm-bootstrap
```
</details>
**Residual risk, stated plainly.** The granter is root-equivalent. It may write
any `swarm-*` policy with any content, and attach it to a role that accepts any
certificate; no bao ACL can constrain what a policy says. What bounds it:
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
module-eval pins each principal's grants. **Merging a change to that policy
text is granting it**: it takes effect on the next deploy with no bao step,
so code review is the only gate.
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
root in a `0700` directory, readable only by root units on the store's host.
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
and `controller-key.pem`, whose policy is already root-equivalent. Root on
that host gains nothing new.
- **Never copy `granter-key.pem` off the host** the way operators copy the
other leaves in that directory. That hands out root-equivalence.
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
subject. Rotate the store's CA, or have root point the `bao-granter` role at
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
one.
What else you need depends on
`services.hyperhive.deploy.bao.seal`:
- **`pkcs11`** (the default) — pkcs11 binds the key to the host's TPM, and
the store unseals itself on every restart. `init` is the only manual step.
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
restart, with the keys `init` printed.
The store serves TLS, and on a hive that deploys it you need do nothing: a
first-boot unit mints a CA of the store's own plus the leaves it signs — the
store's server certificate, this host's client certificate, and one per service
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
Those are `mkDefault`s, so naming your own paths wins. Do that when your
certificates come from a real internal CA; the store has no opinion about
which. A hive that does **not** deploy the store names the reader's three
itself, plus a pair for every principal it runs — see
[per-principal identities](../swarm/secrets.md#per-principal-identities) for the
list. The operator issues those leaves out of band; they're the credentials the
store can't hand you, being what opens it. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is
**mTLS between services and the store**, a separate trust domain, because a
store that took its identity from an authority it itself distributes could
never come up before that authority.
### 4 · Swarm SSO (only when `deploy.authelia`)
⚠️ **Required to finish the install, not optional.** Authelia treats an
empty user store as a fatal startup error, so until this runs the
container crash-loops and `auth.<swarm.domain>` answers `502 Bad
Gateway` — a working vhost in front of an upstream that refuses to
start. Skipping this step looks like a broken proxy.
```bash
# Runs as root on the host that RUNS authelia (not necessarily the
# controller host). Prints a generated password once — record it.
swarmctl user add mara --display-name Mara --email mara@example.com --group admins
```
⚠️ **Keep `--group admins`.** it's not decoration: operator-only
surfaces (the swarm UI below) gate on that group, and an account
without it authenticates successfully and is then refused — which reads
like a broken login rather than a missing group.
If an account already exists without it, `user add` refuses rather
than amends — adding the group afterwards is `swarmctl user update mara
--add-group admins`.
Then sign in to the forge once through authelia, with that account. That
first login creates your forge account, under the same username. Make it
a site admin:
```bash
# On the swarm-controller's host. Fails until that first login has happened.
swarmctl forge make-admin mara
```
⚠️ **Keep `--email` too.** The forge won't create an account without an
email: a subject that has none gets the forge's link-account page and no
account. `swarmctl user update mara --email …` fixes it.
If the forge already has a local account with your username, the first
SSO login asks for that account's forge password once, to link the two.
Detail, including what the password is and why this stays manual:
[`swarm/sso.md`](../swarm/sso.md).
### 5 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
Nothing to run — it's served on the swarm apex
(`https://<swarm.domain>/`) as soon as the host rebuilds. Two things
decide whether you can actually open it:
- **You are in `admins`** (_Swarm SSO_ above). The gateway asks authelia whether
you have a session; the rule that makes it mean _operator_ wants the
group. Without it you log in and still get bounced.
- **The name resolves to this host.** it's published to the hive's own
resolver and to `/etc/hosts` when `gateway.localHostsEntry` is on; from
anywhere else it needs a real DNS record like any other public name.
Detail, including why reachability is deliberately not the access
control: [`swarm/ui.md`](../swarm/ui.md).
### 6 · Matrix
```bash
# Invite the operator to the hive Space (and optionally to rooms)
hivectl matrix invite mara
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'
```
The operator's own matrix account comes from SSO, not `hivectl` — matrix
homeserver admin should eventually come from membership in authelia's
`admins` group; nobody has built that sync yet.
ruth's own matrix account comes from the swarm, like every agent's:
`swarm-controller` creates it within five minutes of her holding a store
identity (step 1), and her matrix daemon reads its token from the store.
Without that identity she has no matrix account, and this hive no longer
creates one.
Swarm SSO creates the human operator's own matrix account instead of
a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`).
### 7 · Spawn sub-agents
Sub-agent creation is a swarm-level operator action — agents have no
tool for it, and no hive can create one on its own:
```
# Create iris on hive pr1ma. The swarm controller seeds its config repo
# (agent-configs/iris) from the default template, then asks pr1ma to
# build + start the container from that config.
swarmctl agent create iris --hive pr1ma
# Later config changes: open a PR on agent-configs/iris (hive-forge);
# the operator reviews + approves it — no MCP tool call.
```
See [`approvals.md`](../agent-lifecycle/approvals.md) for the full flow.
### 8 · Useful host commands
```bash
# Roster: all agents, status, rev, pending reminders
hivectl list-agents
# Restart a stuck container (no rebuild)
hivectl agent <agent> restart
# Open a Claude session inside an agent's container
hivectl agent <agent> choom
# Open hive web surfaces in a browser (or just print the URLs)
hivectl open # operator dashboard
hivectl open forge # Forgejo
hivectl open matrix # Matrix GUI (fluffychat)
```
See [`tools/hivectl.md`](../tools/hivectl.md) for every `hivectl` verb.
## Security notes
<!-- vale write-good.Passive = NO -->
- **No forge admin token is stored in any agent state dir.** Agents
hold a regular agent token, fetched from the swarm secret store into
`/run/hive-agent-forge-token/token` (or, for an agent without a store
identity, the `forge-token` file hive-c0re wrote before); sensitive
creds (the core token) live on the host.
- All config changes (forge PRs on `agent-configs/<name>`) go through
operator approval — agents can't unilaterally rebuild containers, by design.
See [`boundary.md`](../trust-boundary/boundary.md) and [`security.md`](../trust-boundary/security.md).
- **Each hive authenticates its own telemetry ingest**, and the `hive` label comes
from which hive authenticated rather than from the payload — so no hive can
report metrics as another. A first-run all-local hive gets this with nothing
to configure; joining a swarm you don't host needs one secret copied across.
See [`observability.md`](../scheduler/observability.md#authenticated-ingest).
<!-- vale write-good.Passive = YES -->
Once the hive is running, ruth records anything it needs to remember
across restarts in `/agents/ruth/state/notes.md`.