hyperhive/docs/getting-started/setup.md
atlas 6de51b7be4 setup: drop two vale errors my own callout introduced
`Microsoft.Contractions` fires on "is not" and "it is", and both were in
the sentence I added about the role having no consumer yet. Reworded so
neither appears, and the passive construction goes with them.

The vale job is red on every PR by design -- ci.yml says so, the docs
tree has a standing backlog against these rules -- which is exactly why
"vale is red" was not a safe thing to wave through. Checking per line
rather than per job is what found these.

Verified with the invocation ci.yml uses:
`XDG_DATA_HOME=$PWD/.vale-data vale sync && vale docs/getting-started/setup.md`
-> synced 3 packages, 0 errors, 14 warnings. The 2 errors CI reported at
103:39 and 104:7 are the presence control: same rule, same file, prior
commit.
2026-09-07 22:02:33 +02:00

254 lines
10 KiB
Markdown

# First-run setup (fresh-deploy bootstrap)
How to bring a fresh hyperhive hive online: provision accounts, open
the gateway, bootstrap swarm SSO, make matrix reachable, and spawn the
first sub-agents.
Aimed at `ruth` (the root/manager agent) on a fresh deploy, but it's a
plain reference doc — read it whenever you need the bootstrap command
sequence. All `hivectl` commands below run as **root on the host** (not
inside an agent container); the `request_*` steps run from ruth's own
turn via the MCP tools.
**Bringing up a hive that doesn't host its own swarm services?** Read
[`swarm/secrets.md`](../swarm/secrets.md) first. Everything below assumes
each credential is generated where it's read, which is true on an
all-local deploy and not otherwise — that page says which files an
operator has to place, and where.
## Step-by-step
### 1 · Forge
```bash
# Provision (or refresh) ruth's own forge account — do this first. Ruth's
# bootstrap bypasses the normal spawn-approval flow (see step 6), so unlike
# every other agent it does not get its forge account auto-provisioned —
# this manual step is still load-bearing.
hivectl forge create-user ruth
# Sub-agents spawned later (via the approval flow in step 6) get their
# forge accounts auto-provisioned — nothing to run here for them.
```
The human operator's own forge account is created via swarm SSO instead of
a manual `hivectl` step — see step 3 (`swarmctl user add`).
### 2 · Gateway (HTTP Basic auth)
```bash
# Add an operator login to the gateway (reads password from stdin)
echo "hunter2" | hivectl gateway create-user mara --password-stdin
# List existing users
hivectl gateway list-users
```
### 3 · Secret store (only when `deploy.bao`)
⚠️ **A sealed store still answers.** OpenBao starts uninitialised and
sealed, so the container is up and the port responds while every read
times out — the failure looks like a hang, not like a store that was
never initialised. Do this before anything is pointed at it.
```bash
# On the host that RUNS the store, once.
bao operator init # keep the keys it prints and the root token OFF this host
```
While you still hold that root token, mint the one credential the swarm needs
to grant itself anything. Cert auth answers a _role_, so nothing can
authenticate until some role exists — this token is what breaks that cycle,
and it's the only step that needs the root token.
```bash
# Exactly the four grants the bootstrap unit needs, and nothing else. Each was
# derived with `bao <cmd> -output-policy`, which prints what a command requires
# without running it.
bao policy write swarm-bootstrap - <<'EOF'
path "sys/policies/acl/swarm-controller" {
capabilities = ["create", "update"]
}
# Cert auth is a mount, and nothing has created it yet: reading `sys/auth` is
# how the unit checks, and `sudo` is what enabling one costs.
path "sys/auth" {
capabilities = ["read"]
}
path "sys/auth/cert" {
capabilities = ["create", "update", "sudo"]
}
path "auth/cert/certs/swarm-controller" {
capabilities = ["create", "update"]
}
EOF
# A token holding it. `-orphan` so it outlives the session that made it.
bao token create -policy=swarm-bootstrap -ttl=24h -orphan -display-name=swarm-bootstrap
```
Put the token's value at `services.hyperhive.deploy.bao.bootstrapTokenFile`
(all-local names that path for you), then rebuild. A one-shot unit inside the
store's container reads it, writes the `swarm-controller` policy, enables the
cert auth method, and creates the `swarm-controller` role that attaches the two.
⚠️ **None of this has been run against a live store.** Nothing in the tree has
ever authenticated to OpenBao, so treat the block above as derived rather than
exercised — the grants come from `-output-policy`, not from a swarm that came
up on them.
⚠️ **The role it creates has nothing to present a certificate for yet.** No
deployment shape mints a leaf whose CN reads `swarm-controller`, so until one
does, the role stays provisioning that waits for a consumer.
**Delete the file once that has run.** The unit skips when it's absent, so a
host that has finished bootstrapping stops carrying the credential — and the
TTL above means a forgotten one expires rather than lingering.
Whether anything more is needed depends on
`services.hyperhive.deploy.bao.seal`:
- **`pkcs11`** (the default) — the key is bound to the host's TPM and the
store unseals itself on every restart. `init` is the only manual step.
- **`shamir`** — no TPM, so `bao operator unseal` is needed again after
every restart, with the keys `init` printed.
The store serves TLS, and on a hive that deploys it you need do nothing: a
first-boot unit mints a CA of the store's own plus the two leaves it signs —
the store's server certificate and this host's client certificate — and points
`deploy.bao.serverCertFile`, `.serverKeyFile` and `.clientCaFile` at the store's
half, `.clientCertFile`, `.clientKeyFile` and `.serverCaFile` at the reader's.
Those are `mkDefault`s, so naming your own paths wins. Do that when your
certificates come from a real internal CA; the store has no opinion about
which. A hive that does **not** deploy the store names the reader's three
itself: that leaf is issued out of band, and it's the one credential the store
can't hand you, being what opens it. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is
**mTLS between services and the store**, a separate trust domain, because a
store that took its identity from an authority it will itself distribute could
never come up before that authority.
Making even the `init` unnecessary is tracked in issue #3768.
### 4 · Swarm SSO (only when `deploy.authelia`)
⚠️ **Required to finish the install, not optional.** Authelia treats an
empty user store as a fatal startup error, so until this runs the
container crash-loops and `auth.<swarm.domain>` answers `502 Bad
Gateway` — a working vhost in front of an upstream that refuses to
start. Skipping this step looks like a broken proxy.
```bash
# Runs as root on the host that RUNS authelia (not necessarily the
# controller host). Prints a generated password once — record it.
swarmctl user add mara --display-name Mara --email mara@example.com --group admins
```
⚠️ **Keep `--group admins`.** it's not decoration: operator-only
surfaces (the swarm UI below) are gated on that group, and an account
without it authenticates successfully and is then refused — which reads
like a broken login rather than a missing group.
If an account already exists without it, `user add` will refuse rather
than amend — adding the group afterwards is `swarmctl user update mara
--add-group admins`.
Detail, including what the password is and why this stays manual:
[`swarm/sso.md`](../swarm/sso.md).
### 5 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
Nothing to run — it's served on the swarm apex
(`https://<swarm.domain>/`) as soon as the host rebuilds. Two things
decide whether you can actually open it:
- **You are in `admins`** (step 3). The gateway asks authelia whether
you have a session; the rule that makes it mean _operator_ wants the
group. Without it you log in and still get bounced.
- **The name resolves to this host.** it's published to the hive's own
resolver and to `/etc/hosts` when `gateway.localHostsEntry` is on; from
anywhere else it needs a real DNS record like any other public name.
Detail, including why reachability is deliberately not the access
control: [`swarm/ui.md`](../swarm/ui.md).
### 6 · Matrix
```bash
# 5a. Ensure the hive-internal admin account exists first
hivectl matrix sync-admin
# 5b. Provision ruth's own matrix account — same bootstrap-bypass reasoning
# as forge above, still a required manual step.
hivectl matrix create-user ruth
# 5c. Invite the operator to the hive Space (and optionally to rooms)
hivectl matrix invite mara
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'
# 5d. Promote the operator to homeserver admin if needed
hivectl matrix promote-user mara
```
The human operator's own matrix account is created via swarm SSO instead of
a manual `hivectl` step — see step 3 (`swarmctl user add`).
### 7 · Spawn sub-agents
Sub-agent creation goes through the approval queue — ruth proposes, the
operator approves, the container builds. From ruth's own turn (inside
the container, via MCP tools):
```
# Step 1: initialise a new agent's config repo
request_init_config(name: "iris")
# → operator approves → config_ready event lands in the inbox
# Step 2: edit /agents/iris/config/agent.nix and commit it. Then the
# operator spawns iris (dashboard ◆ R3QU3ST SP4WN / Spawn approval),
# which builds + starts the container from that config.
# Later config changes: open a PR on agent-configs/iris (hive-forge);
# the operator reviews + approves it — no MCP tool call.
```
See [`approvals.md`](../agent-lifecycle/approvals.md) for the full flow.
### 8 · Useful host commands
```bash
# Roster: all agents, status, rev, parent, pending reminders
hivectl list-agents
# Restart a stuck container (no rebuild)
hivectl agent <agent> restart
# Open a Claude session inside an agent's container
hivectl agent <agent> choom
# Open hive web surfaces in a browser (or just print the URLs)
hivectl open # operator dashboard
hivectl open forge # Forgejo
hivectl open matrix # Matrix GUI (fluffychat)
```
See [`tools/hivectl.md`](../tools/hivectl.md) for every `hivectl` verb.
## Security notes
- **No forge admin token is stored in any agent state dir.** Agents
hold a regular agent token in their `forge-token` file; sensitive
creds (the core token, the matrix admin token) live on the host.
- All config changes (forge PRs on `agent-configs/<name>`) go through
operator approval — agents can't unilaterally rebuild containers, by design.
See [`boundary.md`](../trust-boundary/boundary.md) and [`security.md`](../trust-boundary/security.md).
- **Telemetry ingest is authenticated per hive**, and the `hive` label comes
from which hive authenticated rather than from the payload — so no hive can
report metrics as another. A first-run all-local hive gets this with nothing
to configure; joining a swarm you don't host needs one secret copied across.
See [`observability.md`](../scheduler/observability.md#authenticated-ingest).
Once the hive is running, ruth records anything it needs to remember
across restarts in `/agents/ruth/state/notes.md`.