docs: setup.md as a short all-local checklist; bao internals move to swarm/bao.md
This commit is contained in:
parent
3a0f74c1e7
commit
a7991c9242
3 changed files with 224 additions and 289 deletions
|
|
@ -1,102 +1,30 @@
|
|||
# First-run setup (fresh-deploy bootstrap)
|
||||
# First-run setup
|
||||
|
||||
How to bring a fresh hyperhive hive online: provision accounts, open
|
||||
the gateway, bootstrap swarm SSO, make matrix reachable, and spawn the
|
||||
first sub-agents.
|
||||
What's left to do by hand after the first `nixos-rebuild switch` of the
|
||||
[README's all-local quick start](../../README.md#quick-start-an-all-local-swarm)
|
||||
— from a freshly built host to your first agent. Do the steps in order;
|
||||
each one assumes the ones before it.
|
||||
|
||||
Aimed at `ruth` (the root/manager agent) on a fresh deploy, but it's a
|
||||
plain reference doc — read it whenever you need the bootstrap command
|
||||
sequence. All `hivectl` commands below run as **root on the host** (not
|
||||
inside an agent container); the `request_*` steps run from ruth's own
|
||||
turn via the MCP tools.
|
||||
Every command runs as **root on the host**.
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
> **Not all-local?** Each step says which host it runs on. A split swarm
|
||||
> also has credentials that can't be generated where they're read — each
|
||||
> host's mTLS identity at the secret store, and each hive's telemetry
|
||||
> ingest secret. Read [`swarm/secrets.md`](../swarm/secrets.md) first; it
|
||||
> says which files you place, and where.
|
||||
|
||||
**Bringing up a hive that doesn't host its own swarm services?** Read
|
||||
[`swarm/secrets.md`](../swarm/secrets.md) first. Everything below assumes
|
||||
each credential is generated where it's read, which is true on an
|
||||
all-local deploy and not otherwise — that page says which files an
|
||||
operator has to place, and where.
|
||||
## 1 · Secret store
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
## Step-by-step
|
||||
|
||||
### 1 · Forge
|
||||
|
||||
hive-c0re no longer creates agent forge users or mints agent tokens.
|
||||
swarm-controller does, for every agent that holds a store identity: a pass
|
||||
at start and every five minutes creates the forge user if it's missing,
|
||||
then mints the token into the swarm secret store, where the agent fetches
|
||||
it under that identity. An agent created with `swarmctl agent create` gets
|
||||
its identity then. Ruth doesn't: hive-c0re creates her on its own at
|
||||
startup, so she needs her identity minted by hand, once.
|
||||
|
||||
```bash
|
||||
# On the swarm-controller host: give ruth her store identity.
|
||||
swarmctl agent mint-identity ruth
|
||||
|
||||
# On ruth's hive: re-apply her container config, which is when hive-c0re
|
||||
# hands the new identity to the container.
|
||||
hivectl agent ruth rebuild
|
||||
```
|
||||
|
||||
You don't need to do anything else. The controller's next pass creates
|
||||
ruth's forge user and mints her token, and her container fetches it within
|
||||
about ten minutes. The same identity puts her in the matrix pass too (step
|
||||
6). `swarmctl agent mint-forge-token ruth` skips the wait for
|
||||
the pass. A hive without a swarm secret store has no path to a forge token
|
||||
for ruth at all.
|
||||
|
||||
Swarm SSO creates the human operator's own forge account: the forge
|
||||
makes it on their first login through authelia, and `swarmctl forge
|
||||
make-admin <you>` then makes it a site admin — see _Swarm SSO_ below.
|
||||
|
||||
### 2 · Gateway (HTTP Basic auth)
|
||||
|
||||
```bash
|
||||
# Add an operator login to the gateway (reads password from stdin)
|
||||
echo "hunter2" | hivectl gateway create-user mara --password-stdin
|
||||
|
||||
# List existing users
|
||||
hivectl gateway list-users
|
||||
```
|
||||
|
||||
### 3 · Secret store (only when `deploy.bao`)
|
||||
|
||||
⚠️ **A sealed store still answers.** OpenBao starts uninitialised and
|
||||
sealed, so the container is up and the port responds while every read
|
||||
times out — the failure looks like a hang, not like a store that was
|
||||
never initialised. Do this before you point anything at it.
|
||||
|
||||
Run this **on the host**. The `bao` there is a wrapper carrying this store's
|
||||
address, its CA, and the host's client certificate already, so nothing needs
|
||||
exporting:
|
||||
_On the host running `swarm-bao`._ Everything else fetches its credentials
|
||||
from the store, so it comes first.
|
||||
|
||||
```bash
|
||||
sudo bao operator init # keep the keys it prints and the root token OFF this host
|
||||
```
|
||||
|
||||
`sudo` because the client certificate and key sit under
|
||||
`/var/lib/swarm-bao-pki`, which is mode `0700`.
|
||||
|
||||
Inside the store's container (`nixos-container root-login swarm-bao`) the same
|
||||
command needs two extra pieces, because the certificate carries no IP SAN and
|
||||
its DNS name resolves to the bridge from in there: export
|
||||
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
|
||||
domain>` to verify the name while connecting on loopback. The host is the
|
||||
shorter path.
|
||||
|
||||
While you still hold that root token, set up the **granter**: the one principal
|
||||
that writes every `swarm-*` policy and role from then on. Cert auth answers a
|
||||
_role_, so nothing can authenticate until some role exists. A short-lived
|
||||
bootstrap token breaks that cycle once, and it's the only step that needs the
|
||||
root token.
|
||||
|
||||
The policy it carries is `nix/host-modules/bao-bootstrap-policy.hcl`, shipped
|
||||
on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It covers
|
||||
the auth mounts and the granter's own policy and role, and nothing else. CI
|
||||
fails when the unit using the token needs a path it lacks.
|
||||
Then bootstrap the **granter**, the one principal that writes every
|
||||
`swarm-*` policy and role from then on. This is the only step that needs
|
||||
the root token:
|
||||
|
||||
```bash
|
||||
sudo -i
|
||||
|
|
@ -108,26 +36,9 @@ unset BAO_TOKEN
|
|||
systemctl restart swarm-bao-granter-role
|
||||
```
|
||||
|
||||
The token file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which
|
||||
all-local names for you. On a store host that isn't all-local, set it and
|
||||
rebuild first.
|
||||
|
||||
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
|
||||
auth methods (and disables `approle` if an older store still has it mounted),
|
||||
writes the `bao-granter` policy, and creates the
|
||||
`bao-granter` role, which accepts the leaf
|
||||
`/var/lib/swarm-bao-pki/granter.pem`. Every
|
||||
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
|
||||
mounts the KV and pki engines and writes the `swarm-controller` role, and each
|
||||
sibling unit writes its own principal's policy and role. Every one runs on the
|
||||
host, because every API listener but the loopback UI one demands a client
|
||||
certificate and the host is the side that has one.
|
||||
|
||||
**Confirm with `systemctl status swarm-bao-granter-role`**, which should log
|
||||
`Uploaded policy: bao-granter` and `Data written to: auth/cert/certs/bao-granter`.
|
||||
Then restart the granting units that failed while they waited. These two
|
||||
names cover every unit that logs in as the granter, and CI fails when one
|
||||
doesn't:
|
||||
Check `systemctl status swarm-bao-granter-role` logs
|
||||
`Uploaded policy: bao-granter`, then restart the granting units that
|
||||
failed while they waited, and remove the token:
|
||||
|
||||
```bash
|
||||
systemctl reset-failed 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
|
||||
|
|
@ -136,223 +47,111 @@ systemctl status swarm-bao-controller-policy # Uploaded policy, Data written t
|
|||
rm /var/lib/swarm-bao-bootstrap/grant.token
|
||||
```
|
||||
|
||||
⚠️ Don't reach for `bao read auth/cert/…` to check. The host's `bao`
|
||||
wrapper carries an address, a CA and a client certificate but deliberately
|
||||
**no token**, so that read answers `403` whether or not the role exists.
|
||||
⏱️ A first attempt right after a rebuild may fail with `local node not
|
||||
active` while the store comes up; the units retry every 30s for a day.
|
||||
|
||||
⏱️ **Expect the first attempt to fail if you rebuilt into this.** A rebuild
|
||||
restarts the store, and the units race it: the store answers `local node not
|
||||
active` until it finishes coming up. They retry every 30s for a day, so a
|
||||
sealed or late store heals itself.
|
||||
With the default `pkcs11` seal the store unseals itself from here on. With
|
||||
`deploy.bao.seal = "shamir"`, run `bao operator unseal` after every restart.
|
||||
What the granter is, why it's root-equivalent, and the store's TLS:
|
||||
[`swarm/bao.md`](../swarm/bao.md).
|
||||
|
||||
Until you set up the granter, each `swarm-bao-*-policy` unit **fails** and logs
|
||||
the commands above. It never skips. Delete the token file only once
|
||||
`swarm-bao-granter-role` has succeeded. That unit skips while the file is
|
||||
absent, which is the steady state afterwards. The TTL above means a forgotten
|
||||
token expires rather than lingering.
|
||||
## 2 · Your SSO account
|
||||
|
||||
After that, a new or changed `swarm-*` grant needs no operator step: the unit
|
||||
that writes it changes, and the deploy restarts it. A root step comes back only
|
||||
when the granter itself needs a path it lacks, such as a new mount.
|
||||
|
||||
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
|
||||
them, not on every deploy. When a grant drifts in the store and its unit stays the
|
||||
same, the next boot re-asserts it, not the next switch.
|
||||
|
||||
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
|
||||
|
||||
A store set up before the granter existed has every grant, but no `bao-granter`
|
||||
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
|
||||
unit fails and logs the one-time step. Run the two blocks above as they stand.
|
||||
The old policy can go, with the root token again:
|
||||
_On the host running authelia._ Authelia refuses to start with no users,
|
||||
so until this runs `auth.<swarm.domain>` answers `502 Bad Gateway`.
|
||||
|
||||
```bash
|
||||
bao policy delete swarm-bootstrap
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
**Residual risk, stated plainly.** The granter is root-equivalent. It may write
|
||||
any `swarm-*` policy with any content, and attach it to a role that accepts any
|
||||
certificate; no bao ACL can constrain what a policy says. What bounds it:
|
||||
|
||||
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
|
||||
module-eval pins each principal's grants. **Merging a change to that policy
|
||||
text is granting it**: it takes effect on the next deploy with no bao step,
|
||||
so code review is the only gate.
|
||||
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
|
||||
root in a `0700` directory, readable only by root units on the store's host.
|
||||
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
|
||||
and `controller-key.pem`, whose policy is already root-equivalent. Root on
|
||||
that host gains nothing new.
|
||||
- **Never copy `granter-key.pem` off the host** the way operators copy the
|
||||
other leaves in that directory. That hands out root-equivalence.
|
||||
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
|
||||
subject. Rotate the store's CA, or have root point the `bao-granter` role at
|
||||
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
|
||||
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
|
||||
one.
|
||||
|
||||
What else you need depends on
|
||||
`services.hyperhive.deploy.bao.seal`:
|
||||
|
||||
- **`pkcs11`** (the default) — pkcs11 binds the key to the host's TPM, and
|
||||
the store unseals itself on every restart. `init` is the only manual step.
|
||||
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
|
||||
restart, with the keys `init` printed.
|
||||
|
||||
The store serves TLS, and on a hive that deploys it you need do nothing: a
|
||||
first-boot unit mints a CA of the store's own plus the leaves it signs — the
|
||||
store's server certificate, this host's client certificate, and one per service
|
||||
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
|
||||
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
|
||||
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
|
||||
|
||||
Those are `mkDefault`s, so naming your own paths wins. Do that when your
|
||||
certificates come from a real internal CA; the store has no opinion about
|
||||
which. A hive that does **not** deploy the store names the reader's three
|
||||
itself, plus a pair for every principal it runs — see
|
||||
[per-principal identities](../swarm/secrets.md#per-principal-identities) for the
|
||||
list. The operator issues those leaves out of band; they're the credentials the
|
||||
store can't hand you, being what opens it. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is
|
||||
**mTLS between services and the store**, a separate trust domain, because a
|
||||
store that took its identity from an authority it itself distributes could
|
||||
never come up before that authority.
|
||||
|
||||
### 4 · Swarm SSO (only when `deploy.authelia`)
|
||||
|
||||
⚠️ **Required to finish the install, not optional.** Authelia treats an
|
||||
empty user store as a fatal startup error, so until this runs the
|
||||
container crash-loops and `auth.<swarm.domain>` answers `502 Bad
|
||||
Gateway` — a working vhost in front of an upstream that refuses to
|
||||
start. Skipping this step looks like a broken proxy.
|
||||
|
||||
```bash
|
||||
# Runs as root on the host that RUNS authelia (not necessarily the
|
||||
# controller host). Prints a generated password once — record it.
|
||||
swarmctl user add mara --display-name Mara --email mara@example.com --group admins
|
||||
```
|
||||
|
||||
⚠️ **Keep `--group admins`.** it's not decoration: operator-only
|
||||
surfaces (the swarm UI below) gate on that group, and an account
|
||||
without it authenticates successfully and is then refused — which reads
|
||||
like a broken login rather than a missing group.
|
||||
It prints a generated password once — record it. Keep both flags:
|
||||
|
||||
If an account already exists without it, `user add` refuses rather
|
||||
than amends — adding the group afterwards is `swarmctl user update mara
|
||||
--add-group admins`.
|
||||
- **`--group admins`** — the swarm UI and other operator surfaces gate on
|
||||
it. Without it you log in fine and are then refused.
|
||||
- **`--email`** — the forge won't create an account without one.
|
||||
|
||||
Then sign in to the forge once through authelia, with that account. That
|
||||
first login creates your forge account, under the same username. Make it
|
||||
a site admin:
|
||||
Fix either afterwards with `swarmctl user update mara --add-group admins
|
||||
--email …`. Details: [`swarm/sso.md`](../swarm/sso.md).
|
||||
|
||||
The swarm UI is now at `https://<swarm.domain>/`. The name resolves on the
|
||||
box itself through `/etc/hosts`; from anywhere else it needs a real DNS
|
||||
record. → [`swarm/ui.md`](../swarm/ui.md)
|
||||
|
||||
## 3 · Forge admin
|
||||
|
||||
Sign in to the forge once through SSO — that first login creates your
|
||||
forge account under the same username. Then, _on the swarm-controller's
|
||||
host_:
|
||||
|
||||
```bash
|
||||
# On the swarm-controller's host. Fails until that first login has happened.
|
||||
swarmctl forge make-admin mara
|
||||
```
|
||||
|
||||
⚠️ **Keep `--email` too.** The forge won't create an account without an
|
||||
email: a subject that has none gets the forge's link-account page and no
|
||||
account. `swarmctl user update mara --email …` fixes it.
|
||||
|
||||
If the forge already has a local account with your username, the first
|
||||
SSO login asks for that account's forge password once, to link the two.
|
||||
SSO login asks for its forge password once, to link the two.
|
||||
|
||||
Detail, including what the password is and why this stays manual:
|
||||
[`swarm/sso.md`](../swarm/sso.md).
|
||||
## 4 · Ruth's store identity
|
||||
|
||||
### 5 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
|
||||
hive-c0re creates the manager agent, ruth, on its own at startup — so
|
||||
unlike agents made through the swarm, she has no store identity yet:
|
||||
|
||||
Nothing to run — it's served on the swarm apex
|
||||
(`https://<swarm.domain>/`) as soon as the host rebuilds. Two things
|
||||
decide whether you can actually open it:
|
||||
```bash
|
||||
swarmctl agent mint-identity ruth # on the swarm-controller's host
|
||||
hivectl agent ruth rebuild # on ruth's hive: hands her the identity
|
||||
```
|
||||
|
||||
- **You are in `admins`** (_Swarm SSO_ above). The gateway asks authelia whether
|
||||
you have a session; the rule that makes it mean _operator_ wants the
|
||||
group. Without it you log in and still get bounced.
|
||||
- **The name resolves to this host.** it's published to the hive's own
|
||||
resolver and to `/etc/hosts` when `gateway.localHostsEntry` is on; from
|
||||
anywhere else it needs a real DNS record like any other public name.
|
||||
Within about ten minutes the controller creates her forge user and matrix
|
||||
account and mints their tokens into the store, where she fetches them.
|
||||
`swarmctl agent mint-forge-token ruth` skips the wait for the forge token.
|
||||
|
||||
Detail, including why reachability is deliberately not the access
|
||||
control: [`swarm/ui.md`](../swarm/ui.md).
|
||||
## 5 · Matrix
|
||||
|
||||
### 6 · Matrix
|
||||
Your matrix account comes from SSO. Invite it to the hive Space, and
|
||||
optionally to rooms:
|
||||
|
||||
```bash
|
||||
# Invite the operator to the hive Space (and optionally to rooms)
|
||||
hivectl matrix invite mara
|
||||
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'
|
||||
```
|
||||
|
||||
The operator's own matrix account comes from SSO, not `hivectl` — matrix
|
||||
homeserver admin should eventually come from membership in authelia's
|
||||
`admins` group; nobody has built that sync yet.
|
||||
## 6 · Your first agent
|
||||
|
||||
ruth's own matrix account comes from the swarm, like every agent's:
|
||||
`swarm-controller` creates it within five minutes of her holding a store
|
||||
identity (step 1), and her matrix daemon reads its token from the store.
|
||||
Without that identity she has no matrix account, and this hive no longer
|
||||
creates one.
|
||||
|
||||
Swarm SSO creates the human operator's own matrix account instead of
|
||||
a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`).
|
||||
|
||||
### 7 · Spawn sub-agents
|
||||
|
||||
Sub-agent creation is a swarm-level operator action — agents have no
|
||||
tool for it, and no hive can create one on its own:
|
||||
|
||||
```
|
||||
# Create iris on hive pr1ma. The swarm controller seeds its config repo
|
||||
# (agent-configs/iris) from the default template, then asks pr1ma to
|
||||
# build + start the container from that config.
|
||||
swarmctl agent create iris --hive pr1ma
|
||||
|
||||
# Later config changes: open a PR on agent-configs/iris (hive-forge);
|
||||
# the operator reviews + approves it — no MCP tool call.
|
||||
```
|
||||
|
||||
See [`approvals.md`](../agent-lifecycle/approvals.md) for the full flow.
|
||||
|
||||
### 8 · Useful host commands
|
||||
Create it from the swarm UI, or:
|
||||
|
||||
```bash
|
||||
# Roster: all agents, status, rev, pending reminders
|
||||
hivectl list-agents
|
||||
|
||||
# Restart a stuck container (no rebuild)
|
||||
hivectl agent <agent> restart
|
||||
|
||||
# Open a Claude session inside an agent's container
|
||||
hivectl agent <agent> choom
|
||||
|
||||
# Open hive web surfaces in a browser (or just print the URLs)
|
||||
hivectl open # operator dashboard
|
||||
hivectl open forge # Forgejo
|
||||
hivectl open matrix # Matrix GUI (fluffychat)
|
||||
swarmctl agent create iris --hive pr1ma
|
||||
```
|
||||
|
||||
See [`tools/hivectl.md`](../tools/hivectl.md) for every `hivectl` verb.
|
||||
The controller provisions iris's identity, forge user and config repo
|
||||
(`agent-configs/iris`, from the default template), then has the hive build
|
||||
and start the container. It returns once the job is queued — watch the
|
||||
swarm UI's job view for progress.
|
||||
|
||||
## Security notes
|
||||
Later config changes are PRs on `agent-configs/iris`, approved by you. →
|
||||
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md)
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
## Optional · Lock the hive dashboard
|
||||
|
||||
- **No forge admin token is stored in any agent state dir.** Agents
|
||||
hold a regular agent token, fetched from the swarm secret store into
|
||||
`/run/hive-agent-forge-token/token` (or, for an agent without a store
|
||||
identity, the `forge-token` file hive-c0re wrote before); sensitive
|
||||
creds (the core token) live on the host.
|
||||
- All config changes (forge PRs on `agent-configs/<name>`) go through
|
||||
operator approval — agents can't unilaterally rebuild containers, by design.
|
||||
See [`boundary.md`](../trust-boundary/boundary.md) and [`security.md`](../trust-boundary/security.md).
|
||||
- **Each hive authenticates its own telemetry ingest**, and the `hive` label comes
|
||||
from which hive authenticated rather than from the payload — so no hive can
|
||||
report metrics as another. A first-run all-local hive gets this with nothing
|
||||
to configure; joining a swarm you don't host needs one secret copied across.
|
||||
See [`observability.md`](../scheduler/observability.md#authenticated-ingest).
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
The per-hive dashboard has no SSO in front of it. To put it behind HTTP
|
||||
Basic auth, set `services.hyperhive.gateway.auth.enable = true` and add
|
||||
logins:
|
||||
|
||||
Once the hive is running, ruth records anything it needs to remember
|
||||
across restarts in `/agents/ruth/state/notes.md`.
|
||||
```bash
|
||||
echo "hunter2" | hivectl gateway create-user mara --password-stdin
|
||||
```
|
||||
|
||||
→ [`networking/gateway.md`](../networking/gateway.md#http-basic-auth)
|
||||
|
||||
## Day to day
|
||||
|
||||
```bash
|
||||
hivectl list-agents # this hive's agents, status, rev
|
||||
hivectl agent <name> restart # stop + start, no rebuild
|
||||
hivectl agent <name> watch # follow its live event stream
|
||||
hivectl agent <name> choom # interactive claude session in its container
|
||||
hivectl approvals pending # what's waiting on you
|
||||
hivectl open # this hive's dashboard (or: forge, matrix)
|
||||
```
|
||||
|
||||
Every verb: [`tools/hivectl.md`](../tools/hivectl.md) ·
|
||||
[`tools/swarmctl-cli.md`](../tools/swarmctl-cli.md).
|
||||
|
|
|
|||
Loading…
Reference in a new issue