Every hive is in a swarm and every swarm runs matrix, so every swarm has a swarm-controller, and since #4810 its hive_sender pass mints each hive's @hive-<hive>: sender token into the store every five minutes. The two other minters of that token go: - swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the hive-matrix container, Command::Mint and src/mint.rs. The binary, its appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert role stay. bao-matrix-reader's checks on the deleted unit are removed; the leaf-identity and no-token-in-env checks now look at swarm-matrix-appservice-publish, which runs under the same identity. - the hive-side mint ladder in hive-c0re's ensure_hive_user (register/appservice-login/password-login with the local as_token), with read_appservice_token, paths::matrix_appservice_token and the helpers only it used. ensure_hive_user now takes the store's token, keeps the file when the store has none or can't be reached, and fails otherwise. - hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is unchanged apart from no longer reading the local as_token. This removes the double-mint race #4810's review flagged: two minters logging in on one pinned device could leave a dead token in the store until the next pass. Closes #4813 Closes #4814
358 lines
16 KiB
Markdown
358 lines
16 KiB
Markdown
# First-run setup (fresh-deploy bootstrap)
|
|
|
|
How to bring a fresh hyperhive hive online: provision accounts, open
|
|
the gateway, bootstrap swarm SSO, make matrix reachable, and spawn the
|
|
first sub-agents.
|
|
|
|
Aimed at `ruth` (the root/manager agent) on a fresh deploy, but it's a
|
|
plain reference doc — read it whenever you need the bootstrap command
|
|
sequence. All `hivectl` commands below run as **root on the host** (not
|
|
inside an agent container); the `request_*` steps run from ruth's own
|
|
turn via the MCP tools.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
**Bringing up a hive that doesn't host its own swarm services?** Read
|
|
[`swarm/secrets.md`](../swarm/secrets.md) first. Everything below assumes
|
|
each credential is generated where it's read, which is true on an
|
|
all-local deploy and not otherwise — that page says which files an
|
|
operator has to place, and where.
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
## Step-by-step
|
|
|
|
### 1 · Forge
|
|
|
|
hive-c0re no longer creates agent forge users or mints agent tokens.
|
|
swarm-controller does, for every agent that holds a store identity: a pass
|
|
at start and every five minutes creates the forge user if it's missing,
|
|
then mints the token into the swarm secret store, where the agent fetches
|
|
it under that identity. An agent created with `swarmctl agent create` gets
|
|
its identity then. Ruth doesn't: hive-c0re creates her on its own at
|
|
startup, so she needs her identity minted by hand, once.
|
|
|
|
```bash
|
|
# On the swarm-controller host: give ruth her store identity.
|
|
swarmctl agent mint-identity ruth
|
|
|
|
# On ruth's hive: re-apply her container config, which is when hive-c0re
|
|
# hands the new identity to the container.
|
|
hivectl agent ruth rebuild
|
|
```
|
|
|
|
You don't need to do anything else. The controller's next pass creates
|
|
ruth's forge user and mints her token, and her container fetches it within
|
|
about ten minutes. The same identity puts her in the matrix pass too (step
|
|
6). `swarmctl agent mint-forge-token ruth` skips the wait for
|
|
the pass. A hive without a swarm secret store has no path to a forge token
|
|
for ruth at all.
|
|
|
|
Swarm SSO creates the human operator's own forge account: the forge
|
|
makes it on their first login through authelia, and `swarmctl forge
|
|
make-admin <you>` then makes it a site admin — see _Swarm SSO_ below.
|
|
|
|
### 2 · Gateway (HTTP Basic auth)
|
|
|
|
```bash
|
|
# Add an operator login to the gateway (reads password from stdin)
|
|
echo "hunter2" | hivectl gateway create-user mara --password-stdin
|
|
|
|
# List existing users
|
|
hivectl gateway list-users
|
|
```
|
|
|
|
### 3 · Secret store (only when `deploy.bao`)
|
|
|
|
⚠️ **A sealed store still answers.** OpenBao starts uninitialised and
|
|
sealed, so the container is up and the port responds while every read
|
|
times out — the failure looks like a hang, not like a store that was
|
|
never initialised. Do this before you point anything at it.
|
|
|
|
Run this **on the host**. The `bao` there is a wrapper carrying this store's
|
|
address, its CA, and the host's client certificate already, so nothing needs
|
|
exporting:
|
|
|
|
```bash
|
|
sudo bao operator init # keep the keys it prints and the root token OFF this host
|
|
```
|
|
|
|
`sudo` because the client certificate and key sit under
|
|
`/var/lib/swarm-bao-pki`, which is mode `0700`.
|
|
|
|
Inside the store's container (`nixos-container root-login swarm-bao`) the same
|
|
command needs two extra pieces, because the certificate carries no IP SAN and
|
|
its DNS name resolves to the bridge from in there: export
|
|
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
|
|
domain>` to verify the name while connecting on loopback. The host is the
|
|
shorter path.
|
|
|
|
While you still hold that root token, set up the **granter**: the one principal
|
|
that writes every `swarm-*` policy and role from then on. Cert auth answers a
|
|
_role_, so nothing can authenticate until some role exists. A short-lived
|
|
bootstrap token breaks that cycle once, and it's the only step that needs the
|
|
root token.
|
|
|
|
The policy it carries is `nix/host-modules/bao-bootstrap-policy.hcl`, shipped
|
|
on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It covers
|
|
the auth mounts and the granter's own policy and role, and nothing else. CI
|
|
fails when the unit using the token needs a path it lacks.
|
|
|
|
```bash
|
|
sudo -i
|
|
read -rs BAO_TOKEN && export BAO_TOKEN # paste the root token from `bao operator init`
|
|
bao policy write bao-bootstrap /etc/hyperhive/bao-bootstrap-policy.hcl
|
|
bao token create -policy=bao-bootstrap -ttl=24h -orphan -display-name=bao-bootstrap -field=token \
|
|
| install -D -m 0600 /dev/stdin /var/lib/swarm-bao-bootstrap/grant.token
|
|
unset BAO_TOKEN
|
|
systemctl restart swarm-bao-granter-role
|
|
```
|
|
|
|
The token file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which
|
|
all-local names for you. On a store host that isn't all-local, set it and
|
|
rebuild first.
|
|
|
|
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
|
|
auth methods (and disables `approle` if an older store still has it mounted),
|
|
writes the `bao-granter` policy, and creates the
|
|
`bao-granter` role, which accepts the leaf
|
|
`/var/lib/swarm-bao-pki/granter.pem`. Every
|
|
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
|
|
mounts the KV and pki engines and writes the `swarm-controller` role, and each
|
|
sibling unit writes its own principal's policy and role. Every one runs on the
|
|
host, because every API listener but the loopback UI one demands a client
|
|
certificate and the host is the side that has one.
|
|
|
|
**Confirm with `systemctl status swarm-bao-granter-role`**, which should log
|
|
`Uploaded policy: bao-granter` and `Data written to: auth/cert/certs/bao-granter`.
|
|
Then restart the granting units that failed while they waited. These two
|
|
names cover every unit that logs in as the granter, and CI fails when one
|
|
doesn't:
|
|
|
|
```bash
|
|
systemctl reset-failed 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
|
|
systemctl restart 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
|
|
systemctl status swarm-bao-controller-policy # Uploaded policy, Data written to: auth/cert/certs/swarm-controller
|
|
rm /var/lib/swarm-bao-bootstrap/grant.token
|
|
```
|
|
|
|
⚠️ Don't reach for `bao read auth/cert/…` to check. The host's `bao`
|
|
wrapper carries an address, a CA and a client certificate but deliberately
|
|
**no token**, so that read answers `403` whether or not the role exists.
|
|
|
|
⏱️ **Expect the first attempt to fail if you rebuilt into this.** A rebuild
|
|
restarts the store, and the units race it: the store answers `local node not
|
|
active` until it finishes coming up. They retry every 30s for a day, so a
|
|
sealed or late store heals itself.
|
|
|
|
Until you set up the granter, each `swarm-bao-*-policy` unit **fails** and logs
|
|
the commands above. It never skips. Delete the token file only once
|
|
`swarm-bao-granter-role` has succeeded. That unit skips while the file is
|
|
absent, which is the steady state afterwards. The TTL above means a forgotten
|
|
token expires rather than lingering.
|
|
|
|
After that, a new or changed `swarm-*` grant needs no operator step: the unit
|
|
that writes it changes, and the deploy restarts it. A root step comes back only
|
|
when the granter itself needs a path it lacks, such as a new mount.
|
|
|
|
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
|
|
them, not on every deploy. When a grant drifts in the store and its unit stays the
|
|
same, the next boot re-asserts it, not the next switch.
|
|
|
|
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
|
|
|
|
A store set up before the granter existed has every grant, but no `bao-granter`
|
|
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
|
|
unit fails and logs the one-time step. Run the two blocks above as they stand.
|
|
The old policy can go, with the root token again:
|
|
|
|
```bash
|
|
bao policy delete swarm-bootstrap
|
|
```
|
|
|
|
</details>
|
|
|
|
**Residual risk, stated plainly.** The granter is root-equivalent. It may write
|
|
any `swarm-*` policy with any content, and attach it to a role that accepts any
|
|
certificate; no bao ACL can constrain what a policy says. What bounds it:
|
|
|
|
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
|
|
module-eval pins each principal's grants. **Merging a change to that policy
|
|
text is granting it**: it takes effect on the next deploy with no bao step,
|
|
so code review is the only gate.
|
|
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
|
|
root in a `0700` directory, readable only by root units on the store's host.
|
|
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
|
|
and `controller-key.pem`, whose policy is already root-equivalent. Root on
|
|
that host gains nothing new.
|
|
- **Never copy `granter-key.pem` off the host** the way operators copy the
|
|
other leaves in that directory. That hands out root-equivalence.
|
|
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
|
|
subject. Rotate the store's CA, or have root point the `bao-granter` role at
|
|
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
|
|
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
|
|
one.
|
|
|
|
What else you need depends on
|
|
`services.hyperhive.deploy.bao.seal`:
|
|
|
|
- **`pkcs11`** (the default) — pkcs11 binds the key to the host's TPM, and
|
|
the store unseals itself on every restart. `init` is the only manual step.
|
|
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
|
|
restart, with the keys `init` printed.
|
|
|
|
The store serves TLS, and on a hive that deploys it you need do nothing: a
|
|
first-boot unit mints a CA of the store's own plus the leaves it signs — the
|
|
store's server certificate, this host's client certificate, and one per service
|
|
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
|
|
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
|
|
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
|
|
|
|
Those are `mkDefault`s, so naming your own paths wins. Do that when your
|
|
certificates come from a real internal CA; the store has no opinion about
|
|
which. A hive that does **not** deploy the store names the reader's three
|
|
itself, plus a pair for every principal it runs — see
|
|
[per-principal identities](../swarm/secrets.md#per-principal-identities) for the
|
|
list. The operator issues those leaves out of band; they're the credentials the
|
|
store can't hand you, being what opens it. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is
|
|
**mTLS between services and the store**, a separate trust domain, because a
|
|
store that took its identity from an authority it itself distributes could
|
|
never come up before that authority.
|
|
|
|
### 4 · Swarm SSO (only when `deploy.authelia`)
|
|
|
|
⚠️ **Required to finish the install, not optional.** Authelia treats an
|
|
empty user store as a fatal startup error, so until this runs the
|
|
container crash-loops and `auth.<swarm.domain>` answers `502 Bad
|
|
Gateway` — a working vhost in front of an upstream that refuses to
|
|
start. Skipping this step looks like a broken proxy.
|
|
|
|
```bash
|
|
# Runs as root on the host that RUNS authelia (not necessarily the
|
|
# controller host). Prints a generated password once — record it.
|
|
swarmctl user add mara --display-name Mara --email mara@example.com --group admins
|
|
```
|
|
|
|
⚠️ **Keep `--group admins`.** it's not decoration: operator-only
|
|
surfaces (the swarm UI below) gate on that group, and an account
|
|
without it authenticates successfully and is then refused — which reads
|
|
like a broken login rather than a missing group.
|
|
|
|
If an account already exists without it, `user add` refuses rather
|
|
than amends — adding the group afterwards is `swarmctl user update mara
|
|
--add-group admins`.
|
|
|
|
Then sign in to the forge once through authelia, with that account. That
|
|
first login creates your forge account, under the same username. Make it
|
|
a site admin:
|
|
|
|
```bash
|
|
# On the swarm-controller's host. Fails until that first login has happened.
|
|
swarmctl forge make-admin mara
|
|
```
|
|
|
|
⚠️ **Keep `--email` too.** The forge won't create an account without an
|
|
email: a subject that has none gets the forge's link-account page and no
|
|
account. `swarmctl user update mara --email …` fixes it.
|
|
|
|
If the forge already has a local account with your username, the first
|
|
SSO login asks for that account's forge password once, to link the two.
|
|
|
|
Detail, including what the password is and why this stays manual:
|
|
[`swarm/sso.md`](../swarm/sso.md).
|
|
|
|
### 5 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
|
|
|
|
Nothing to run — it's served on the swarm apex
|
|
(`https://<swarm.domain>/`) as soon as the host rebuilds. Two things
|
|
decide whether you can actually open it:
|
|
|
|
- **You are in `admins`** (_Swarm SSO_ above). The gateway asks authelia whether
|
|
you have a session; the rule that makes it mean _operator_ wants the
|
|
group. Without it you log in and still get bounced.
|
|
- **The name resolves to this host.** it's published to the hive's own
|
|
resolver and to `/etc/hosts` when `gateway.localHostsEntry` is on; from
|
|
anywhere else it needs a real DNS record like any other public name.
|
|
|
|
Detail, including why reachability is deliberately not the access
|
|
control: [`swarm/ui.md`](../swarm/ui.md).
|
|
|
|
### 6 · Matrix
|
|
|
|
```bash
|
|
# Invite the operator to the hive Space (and optionally to rooms)
|
|
hivectl matrix invite mara
|
|
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'
|
|
```
|
|
|
|
The operator's own matrix account comes from SSO, not `hivectl` — matrix
|
|
homeserver admin should eventually come from membership in authelia's
|
|
`admins` group; nobody has built that sync yet.
|
|
|
|
ruth's own matrix account comes from the swarm, like every agent's:
|
|
`swarm-controller` creates it within five minutes of her holding a store
|
|
identity (step 1), and her matrix daemon reads its token from the store.
|
|
Without that identity she has no matrix account, and this hive no longer
|
|
creates one.
|
|
|
|
Swarm SSO creates the human operator's own matrix account instead of
|
|
a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`).
|
|
|
|
### 7 · Spawn sub-agents
|
|
|
|
Sub-agent creation is a swarm-level operator action — agents have no
|
|
tool for it, and no hive can create one on its own:
|
|
|
|
```
|
|
# Create iris on hive pr1ma. The swarm controller seeds its config repo
|
|
# (agent-configs/iris) from the default template, then asks pr1ma to
|
|
# build + start the container from that config.
|
|
swarmctl agent create iris --hive pr1ma
|
|
|
|
# Later config changes: open a PR on agent-configs/iris (hive-forge);
|
|
# the operator reviews + approves it — no MCP tool call.
|
|
```
|
|
|
|
See [`approvals.md`](../agent-lifecycle/approvals.md) for the full flow.
|
|
|
|
### 8 · Useful host commands
|
|
|
|
```bash
|
|
# Roster: all agents, status, rev, pending reminders
|
|
hivectl list-agents
|
|
|
|
# Restart a stuck container (no rebuild)
|
|
hivectl agent <agent> restart
|
|
|
|
# Open a Claude session inside an agent's container
|
|
hivectl agent <agent> choom
|
|
|
|
# Open hive web surfaces in a browser (or just print the URLs)
|
|
hivectl open # operator dashboard
|
|
hivectl open forge # Forgejo
|
|
hivectl open matrix # Matrix GUI (fluffychat)
|
|
```
|
|
|
|
See [`tools/hivectl.md`](../tools/hivectl.md) for every `hivectl` verb.
|
|
|
|
## Security notes
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
- **No forge admin token is stored in any agent state dir.** Agents
|
|
hold a regular agent token, fetched from the swarm secret store into
|
|
`/run/hive-agent-forge-token/token` (or, for an agent without a store
|
|
identity, the `forge-token` file hive-c0re wrote before); sensitive
|
|
creds (the core token) live on the host.
|
|
- All config changes (forge PRs on `agent-configs/<name>`) go through
|
|
operator approval — agents can't unilaterally rebuild containers, by design.
|
|
See [`boundary.md`](../trust-boundary/boundary.md) and [`security.md`](../trust-boundary/security.md).
|
|
- **Each hive authenticates its own telemetry ingest**, and the `hive` label comes
|
|
from which hive authenticated rather than from the payload — so no hive can
|
|
report metrics as another. A first-run all-local hive gets this with nothing
|
|
to configure; joining a swarm you don't host needs one secret copied across.
|
|
See [`observability.md`](../scheduler/observability.md#authenticated-ingest).
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
Once the hive is running, ruth records anything it needs to remember
|
|
across restarts in `/agents/ruth/state/notes.md`.
|