docs: setup.md as a short all-local checklist; bao internals move to swarm/bao.md
This commit is contained in:
parent
3a0f74c1e7
commit
a7991c9242
3 changed files with 224 additions and 289 deletions
|
|
@ -11,8 +11,8 @@ declarations.
|
||||||
|
|
||||||
## Getting started
|
## Getting started
|
||||||
|
|
||||||
- **Bringing a fresh hive online?** → [`getting-started/setup.md`](getting-started/setup.md)
|
- **Bringing a fresh swarm online?** → [`getting-started/setup.md`](getting-started/setup.md)
|
||||||
(first-run `hivectl` bootstrap).
|
(the hand steps after the first switch).
|
||||||
- **What does the dashboard look like, and how do I use it?** →
|
- **What does the dashboard look like, and how do I use it?** →
|
||||||
[`web-ui/`](web-ui/README.md) — the operator-facing starting point;
|
[`web-ui/`](web-ui/README.md) — the operator-facing starting point;
|
||||||
its own sub-pages ([`shape`](web-ui/shape.md),
|
its own sub-pages ([`shape`](web-ui/shape.md),
|
||||||
|
|
@ -80,6 +80,8 @@ declarations.
|
||||||
where's that shape headed?** → [`swarm/credentials.md`](swarm/credentials.md)
|
where's that shape headed?** → [`swarm/credentials.md`](swarm/credentials.md)
|
||||||
(current state, target state, and the progressive-enhancement rule);
|
(current state, target state, and the progressive-enhancement rule);
|
||||||
[`swarm/secrets.md`](swarm/secrets.md) for where each file lives today.
|
[`swarm/secrets.md`](swarm/secrets.md) for where each file lives today.
|
||||||
|
- **How does the secret store come up, who writes its grants, how is it
|
||||||
|
sealed?** → [`swarm/bao.md`](swarm/bao.md).
|
||||||
|
|
||||||
## Scheduler, CI, observability
|
## Scheduler, CI, observability
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,102 +1,30 @@
|
||||||
# First-run setup (fresh-deploy bootstrap)
|
# First-run setup
|
||||||
|
|
||||||
How to bring a fresh hyperhive hive online: provision accounts, open
|
What's left to do by hand after the first `nixos-rebuild switch` of the
|
||||||
the gateway, bootstrap swarm SSO, make matrix reachable, and spawn the
|
[README's all-local quick start](../../README.md#quick-start-an-all-local-swarm)
|
||||||
first sub-agents.
|
— from a freshly built host to your first agent. Do the steps in order;
|
||||||
|
each one assumes the ones before it.
|
||||||
|
|
||||||
Aimed at `ruth` (the root/manager agent) on a fresh deploy, but it's a
|
Every command runs as **root on the host**.
|
||||||
plain reference doc — read it whenever you need the bootstrap command
|
|
||||||
sequence. All `hivectl` commands below run as **root on the host** (not
|
|
||||||
inside an agent container); the `request_*` steps run from ruth's own
|
|
||||||
turn via the MCP tools.
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
> **Not all-local?** Each step says which host it runs on. A split swarm
|
||||||
|
> also has credentials that can't be generated where they're read — each
|
||||||
|
> host's mTLS identity at the secret store, and each hive's telemetry
|
||||||
|
> ingest secret. Read [`swarm/secrets.md`](../swarm/secrets.md) first; it
|
||||||
|
> says which files you place, and where.
|
||||||
|
|
||||||
**Bringing up a hive that doesn't host its own swarm services?** Read
|
## 1 · Secret store
|
||||||
[`swarm/secrets.md`](../swarm/secrets.md) first. Everything below assumes
|
|
||||||
each credential is generated where it's read, which is true on an
|
|
||||||
all-local deploy and not otherwise — that page says which files an
|
|
||||||
operator has to place, and where.
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
_On the host running `swarm-bao`._ Everything else fetches its credentials
|
||||||
|
from the store, so it comes first.
|
||||||
## Step-by-step
|
|
||||||
|
|
||||||
### 1 · Forge
|
|
||||||
|
|
||||||
hive-c0re no longer creates agent forge users or mints agent tokens.
|
|
||||||
swarm-controller does, for every agent that holds a store identity: a pass
|
|
||||||
at start and every five minutes creates the forge user if it's missing,
|
|
||||||
then mints the token into the swarm secret store, where the agent fetches
|
|
||||||
it under that identity. An agent created with `swarmctl agent create` gets
|
|
||||||
its identity then. Ruth doesn't: hive-c0re creates her on its own at
|
|
||||||
startup, so she needs her identity minted by hand, once.
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# On the swarm-controller host: give ruth her store identity.
|
|
||||||
swarmctl agent mint-identity ruth
|
|
||||||
|
|
||||||
# On ruth's hive: re-apply her container config, which is when hive-c0re
|
|
||||||
# hands the new identity to the container.
|
|
||||||
hivectl agent ruth rebuild
|
|
||||||
```
|
|
||||||
|
|
||||||
You don't need to do anything else. The controller's next pass creates
|
|
||||||
ruth's forge user and mints her token, and her container fetches it within
|
|
||||||
about ten minutes. The same identity puts her in the matrix pass too (step
|
|
||||||
6). `swarmctl agent mint-forge-token ruth` skips the wait for
|
|
||||||
the pass. A hive without a swarm secret store has no path to a forge token
|
|
||||||
for ruth at all.
|
|
||||||
|
|
||||||
Swarm SSO creates the human operator's own forge account: the forge
|
|
||||||
makes it on their first login through authelia, and `swarmctl forge
|
|
||||||
make-admin <you>` then makes it a site admin — see _Swarm SSO_ below.
|
|
||||||
|
|
||||||
### 2 · Gateway (HTTP Basic auth)
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Add an operator login to the gateway (reads password from stdin)
|
|
||||||
echo "hunter2" | hivectl gateway create-user mara --password-stdin
|
|
||||||
|
|
||||||
# List existing users
|
|
||||||
hivectl gateway list-users
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3 · Secret store (only when `deploy.bao`)
|
|
||||||
|
|
||||||
⚠️ **A sealed store still answers.** OpenBao starts uninitialised and
|
|
||||||
sealed, so the container is up and the port responds while every read
|
|
||||||
times out — the failure looks like a hang, not like a store that was
|
|
||||||
never initialised. Do this before you point anything at it.
|
|
||||||
|
|
||||||
Run this **on the host**. The `bao` there is a wrapper carrying this store's
|
|
||||||
address, its CA, and the host's client certificate already, so nothing needs
|
|
||||||
exporting:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sudo bao operator init # keep the keys it prints and the root token OFF this host
|
sudo bao operator init # keep the keys it prints and the root token OFF this host
|
||||||
```
|
```
|
||||||
|
|
||||||
`sudo` because the client certificate and key sit under
|
Then bootstrap the **granter**, the one principal that writes every
|
||||||
`/var/lib/swarm-bao-pki`, which is mode `0700`.
|
`swarm-*` policy and role from then on. This is the only step that needs
|
||||||
|
the root token:
|
||||||
Inside the store's container (`nixos-container root-login swarm-bao`) the same
|
|
||||||
command needs two extra pieces, because the certificate carries no IP SAN and
|
|
||||||
its DNS name resolves to the bridge from in there: export
|
|
||||||
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
|
|
||||||
domain>` to verify the name while connecting on loopback. The host is the
|
|
||||||
shorter path.
|
|
||||||
|
|
||||||
While you still hold that root token, set up the **granter**: the one principal
|
|
||||||
that writes every `swarm-*` policy and role from then on. Cert auth answers a
|
|
||||||
_role_, so nothing can authenticate until some role exists. A short-lived
|
|
||||||
bootstrap token breaks that cycle once, and it's the only step that needs the
|
|
||||||
root token.
|
|
||||||
|
|
||||||
The policy it carries is `nix/host-modules/bao-bootstrap-policy.hcl`, shipped
|
|
||||||
on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It covers
|
|
||||||
the auth mounts and the granter's own policy and role, and nothing else. CI
|
|
||||||
fails when the unit using the token needs a path it lacks.
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sudo -i
|
sudo -i
|
||||||
|
|
@ -108,26 +36,9 @@ unset BAO_TOKEN
|
||||||
systemctl restart swarm-bao-granter-role
|
systemctl restart swarm-bao-granter-role
|
||||||
```
|
```
|
||||||
|
|
||||||
The token file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which
|
Check `systemctl status swarm-bao-granter-role` logs
|
||||||
all-local names for you. On a store host that isn't all-local, set it and
|
`Uploaded policy: bao-granter`, then restart the granting units that
|
||||||
rebuild first.
|
failed while they waited, and remove the token:
|
||||||
|
|
||||||
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
|
|
||||||
auth methods (and disables `approle` if an older store still has it mounted),
|
|
||||||
writes the `bao-granter` policy, and creates the
|
|
||||||
`bao-granter` role, which accepts the leaf
|
|
||||||
`/var/lib/swarm-bao-pki/granter.pem`. Every
|
|
||||||
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
|
|
||||||
mounts the KV and pki engines and writes the `swarm-controller` role, and each
|
|
||||||
sibling unit writes its own principal's policy and role. Every one runs on the
|
|
||||||
host, because every API listener but the loopback UI one demands a client
|
|
||||||
certificate and the host is the side that has one.
|
|
||||||
|
|
||||||
**Confirm with `systemctl status swarm-bao-granter-role`**, which should log
|
|
||||||
`Uploaded policy: bao-granter` and `Data written to: auth/cert/certs/bao-granter`.
|
|
||||||
Then restart the granting units that failed while they waited. These two
|
|
||||||
names cover every unit that logs in as the granter, and CI fails when one
|
|
||||||
doesn't:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
systemctl reset-failed 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
|
systemctl reset-failed 'swarm-bao-*-policy.service' swarm-bao-agent-pki.service
|
||||||
|
|
@ -136,223 +47,111 @@ systemctl status swarm-bao-controller-policy # Uploaded policy, Data written t
|
||||||
rm /var/lib/swarm-bao-bootstrap/grant.token
|
rm /var/lib/swarm-bao-bootstrap/grant.token
|
||||||
```
|
```
|
||||||
|
|
||||||
⚠️ Don't reach for `bao read auth/cert/…` to check. The host's `bao`
|
⏱️ A first attempt right after a rebuild may fail with `local node not
|
||||||
wrapper carries an address, a CA and a client certificate but deliberately
|
active` while the store comes up; the units retry every 30s for a day.
|
||||||
**no token**, so that read answers `403` whether or not the role exists.
|
|
||||||
|
|
||||||
⏱️ **Expect the first attempt to fail if you rebuilt into this.** A rebuild
|
With the default `pkcs11` seal the store unseals itself from here on. With
|
||||||
restarts the store, and the units race it: the store answers `local node not
|
`deploy.bao.seal = "shamir"`, run `bao operator unseal` after every restart.
|
||||||
active` until it finishes coming up. They retry every 30s for a day, so a
|
What the granter is, why it's root-equivalent, and the store's TLS:
|
||||||
sealed or late store heals itself.
|
[`swarm/bao.md`](../swarm/bao.md).
|
||||||
|
|
||||||
Until you set up the granter, each `swarm-bao-*-policy` unit **fails** and logs
|
## 2 · Your SSO account
|
||||||
the commands above. It never skips. Delete the token file only once
|
|
||||||
`swarm-bao-granter-role` has succeeded. That unit skips while the file is
|
|
||||||
absent, which is the steady state afterwards. The TTL above means a forgotten
|
|
||||||
token expires rather than lingering.
|
|
||||||
|
|
||||||
After that, a new or changed `swarm-*` grant needs no operator step: the unit
|
_On the host running authelia._ Authelia refuses to start with no users,
|
||||||
that writes it changes, and the deploy restarts it. A root step comes back only
|
so until this runs `auth.<swarm.domain>` answers `502 Bad Gateway`.
|
||||||
when the granter itself needs a path it lacks, such as a new mount.
|
|
||||||
|
|
||||||
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
|
|
||||||
them, not on every deploy. When a grant drifts in the store and its unit stays the
|
|
||||||
same, the next boot re-asserts it, not the next switch.
|
|
||||||
|
|
||||||
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
|
|
||||||
|
|
||||||
A store set up before the granter existed has every grant, but no `bao-granter`
|
|
||||||
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
|
|
||||||
unit fails and logs the one-time step. Run the two blocks above as they stand.
|
|
||||||
The old policy can go, with the root token again:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bao policy delete swarm-bootstrap
|
|
||||||
```
|
|
||||||
|
|
||||||
</details>
|
|
||||||
|
|
||||||
**Residual risk, stated plainly.** The granter is root-equivalent. It may write
|
|
||||||
any `swarm-*` policy with any content, and attach it to a role that accepts any
|
|
||||||
certificate; no bao ACL can constrain what a policy says. What bounds it:
|
|
||||||
|
|
||||||
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
|
|
||||||
module-eval pins each principal's grants. **Merging a change to that policy
|
|
||||||
text is granting it**: it takes effect on the next deploy with no bao step,
|
|
||||||
so code review is the only gate.
|
|
||||||
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
|
|
||||||
root in a `0700` directory, readable only by root units on the store's host.
|
|
||||||
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
|
|
||||||
and `controller-key.pem`, whose policy is already root-equivalent. Root on
|
|
||||||
that host gains nothing new.
|
|
||||||
- **Never copy `granter-key.pem` off the host** the way operators copy the
|
|
||||||
other leaves in that directory. That hands out root-equivalence.
|
|
||||||
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
|
|
||||||
subject. Rotate the store's CA, or have root point the `bao-granter` role at
|
|
||||||
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
|
|
||||||
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
|
|
||||||
one.
|
|
||||||
|
|
||||||
What else you need depends on
|
|
||||||
`services.hyperhive.deploy.bao.seal`:
|
|
||||||
|
|
||||||
- **`pkcs11`** (the default) — pkcs11 binds the key to the host's TPM, and
|
|
||||||
the store unseals itself on every restart. `init` is the only manual step.
|
|
||||||
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
|
|
||||||
restart, with the keys `init` printed.
|
|
||||||
|
|
||||||
The store serves TLS, and on a hive that deploys it you need do nothing: a
|
|
||||||
first-boot unit mints a CA of the store's own plus the leaves it signs — the
|
|
||||||
store's server certificate, this host's client certificate, and one per service
|
|
||||||
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
|
|
||||||
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
|
|
||||||
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
|
|
||||||
|
|
||||||
Those are `mkDefault`s, so naming your own paths wins. Do that when your
|
|
||||||
certificates come from a real internal CA; the store has no opinion about
|
|
||||||
which. A hive that does **not** deploy the store names the reader's three
|
|
||||||
itself, plus a pair for every principal it runs — see
|
|
||||||
[per-principal identities](../swarm/secrets.md#per-principal-identities) for the
|
|
||||||
list. The operator issues those leaves out of band; they're the credentials the
|
|
||||||
store can't hand you, being what opens it. ⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is
|
|
||||||
**mTLS between services and the store**, a separate trust domain, because a
|
|
||||||
store that took its identity from an authority it itself distributes could
|
|
||||||
never come up before that authority.
|
|
||||||
|
|
||||||
### 4 · Swarm SSO (only when `deploy.authelia`)
|
|
||||||
|
|
||||||
⚠️ **Required to finish the install, not optional.** Authelia treats an
|
|
||||||
empty user store as a fatal startup error, so until this runs the
|
|
||||||
container crash-loops and `auth.<swarm.domain>` answers `502 Bad
|
|
||||||
Gateway` — a working vhost in front of an upstream that refuses to
|
|
||||||
start. Skipping this step looks like a broken proxy.
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Runs as root on the host that RUNS authelia (not necessarily the
|
|
||||||
# controller host). Prints a generated password once — record it.
|
|
||||||
swarmctl user add mara --display-name Mara --email mara@example.com --group admins
|
swarmctl user add mara --display-name Mara --email mara@example.com --group admins
|
||||||
```
|
```
|
||||||
|
|
||||||
⚠️ **Keep `--group admins`.** it's not decoration: operator-only
|
It prints a generated password once — record it. Keep both flags:
|
||||||
surfaces (the swarm UI below) gate on that group, and an account
|
|
||||||
without it authenticates successfully and is then refused — which reads
|
|
||||||
like a broken login rather than a missing group.
|
|
||||||
|
|
||||||
If an account already exists without it, `user add` refuses rather
|
- **`--group admins`** — the swarm UI and other operator surfaces gate on
|
||||||
than amends — adding the group afterwards is `swarmctl user update mara
|
it. Without it you log in fine and are then refused.
|
||||||
--add-group admins`.
|
- **`--email`** — the forge won't create an account without one.
|
||||||
|
|
||||||
Then sign in to the forge once through authelia, with that account. That
|
Fix either afterwards with `swarmctl user update mara --add-group admins
|
||||||
first login creates your forge account, under the same username. Make it
|
--email …`. Details: [`swarm/sso.md`](../swarm/sso.md).
|
||||||
a site admin:
|
|
||||||
|
The swarm UI is now at `https://<swarm.domain>/`. The name resolves on the
|
||||||
|
box itself through `/etc/hosts`; from anywhere else it needs a real DNS
|
||||||
|
record. → [`swarm/ui.md`](../swarm/ui.md)
|
||||||
|
|
||||||
|
## 3 · Forge admin
|
||||||
|
|
||||||
|
Sign in to the forge once through SSO — that first login creates your
|
||||||
|
forge account under the same username. Then, _on the swarm-controller's
|
||||||
|
host_:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# On the swarm-controller's host. Fails until that first login has happened.
|
|
||||||
swarmctl forge make-admin mara
|
swarmctl forge make-admin mara
|
||||||
```
|
```
|
||||||
|
|
||||||
⚠️ **Keep `--email` too.** The forge won't create an account without an
|
|
||||||
email: a subject that has none gets the forge's link-account page and no
|
|
||||||
account. `swarmctl user update mara --email …` fixes it.
|
|
||||||
|
|
||||||
If the forge already has a local account with your username, the first
|
If the forge already has a local account with your username, the first
|
||||||
SSO login asks for that account's forge password once, to link the two.
|
SSO login asks for its forge password once, to link the two.
|
||||||
|
|
||||||
Detail, including what the password is and why this stays manual:
|
## 4 · Ruth's store identity
|
||||||
[`swarm/sso.md`](../swarm/sso.md).
|
|
||||||
|
|
||||||
### 5 · Swarm UI (only when `deploy.swarm-ui`, on by default with the controller)
|
hive-c0re creates the manager agent, ruth, on its own at startup — so
|
||||||
|
unlike agents made through the swarm, she has no store identity yet:
|
||||||
|
|
||||||
Nothing to run — it's served on the swarm apex
|
```bash
|
||||||
(`https://<swarm.domain>/`) as soon as the host rebuilds. Two things
|
swarmctl agent mint-identity ruth # on the swarm-controller's host
|
||||||
decide whether you can actually open it:
|
hivectl agent ruth rebuild # on ruth's hive: hands her the identity
|
||||||
|
```
|
||||||
|
|
||||||
- **You are in `admins`** (_Swarm SSO_ above). The gateway asks authelia whether
|
Within about ten minutes the controller creates her forge user and matrix
|
||||||
you have a session; the rule that makes it mean _operator_ wants the
|
account and mints their tokens into the store, where she fetches them.
|
||||||
group. Without it you log in and still get bounced.
|
`swarmctl agent mint-forge-token ruth` skips the wait for the forge token.
|
||||||
- **The name resolves to this host.** it's published to the hive's own
|
|
||||||
resolver and to `/etc/hosts` when `gateway.localHostsEntry` is on; from
|
|
||||||
anywhere else it needs a real DNS record like any other public name.
|
|
||||||
|
|
||||||
Detail, including why reachability is deliberately not the access
|
## 5 · Matrix
|
||||||
control: [`swarm/ui.md`](../swarm/ui.md).
|
|
||||||
|
|
||||||
### 6 · Matrix
|
Your matrix account comes from SSO. Invite it to the hive Space, and
|
||||||
|
optionally to rooms:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Invite the operator to the hive Space (and optionally to rooms)
|
|
||||||
hivectl matrix invite mara
|
hivectl matrix invite mara
|
||||||
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'
|
hivectl matrix invite @mara:yourserver --room '#hive-chat:yourserver'
|
||||||
```
|
```
|
||||||
|
|
||||||
The operator's own matrix account comes from SSO, not `hivectl` — matrix
|
## 6 · Your first agent
|
||||||
homeserver admin should eventually come from membership in authelia's
|
|
||||||
`admins` group; nobody has built that sync yet.
|
|
||||||
|
|
||||||
ruth's own matrix account comes from the swarm, like every agent's:
|
Create it from the swarm UI, or:
|
||||||
`swarm-controller` creates it within five minutes of her holding a store
|
|
||||||
identity (step 1), and her matrix daemon reads its token from the store.
|
|
||||||
Without that identity she has no matrix account, and this hive no longer
|
|
||||||
creates one.
|
|
||||||
|
|
||||||
Swarm SSO creates the human operator's own matrix account instead of
|
|
||||||
a manual `hivectl` step — see _Swarm SSO_ above (`swarmctl user add`).
|
|
||||||
|
|
||||||
### 7 · Spawn sub-agents
|
|
||||||
|
|
||||||
Sub-agent creation is a swarm-level operator action — agents have no
|
|
||||||
tool for it, and no hive can create one on its own:
|
|
||||||
|
|
||||||
```
|
|
||||||
# Create iris on hive pr1ma. The swarm controller seeds its config repo
|
|
||||||
# (agent-configs/iris) from the default template, then asks pr1ma to
|
|
||||||
# build + start the container from that config.
|
|
||||||
swarmctl agent create iris --hive pr1ma
|
|
||||||
|
|
||||||
# Later config changes: open a PR on agent-configs/iris (hive-forge);
|
|
||||||
# the operator reviews + approves it — no MCP tool call.
|
|
||||||
```
|
|
||||||
|
|
||||||
See [`approvals.md`](../agent-lifecycle/approvals.md) for the full flow.
|
|
||||||
|
|
||||||
### 8 · Useful host commands
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Roster: all agents, status, rev, pending reminders
|
swarmctl agent create iris --hive pr1ma
|
||||||
hivectl list-agents
|
|
||||||
|
|
||||||
# Restart a stuck container (no rebuild)
|
|
||||||
hivectl agent <agent> restart
|
|
||||||
|
|
||||||
# Open a Claude session inside an agent's container
|
|
||||||
hivectl agent <agent> choom
|
|
||||||
|
|
||||||
# Open hive web surfaces in a browser (or just print the URLs)
|
|
||||||
hivectl open # operator dashboard
|
|
||||||
hivectl open forge # Forgejo
|
|
||||||
hivectl open matrix # Matrix GUI (fluffychat)
|
|
||||||
```
|
```
|
||||||
|
|
||||||
See [`tools/hivectl.md`](../tools/hivectl.md) for every `hivectl` verb.
|
The controller provisions iris's identity, forge user and config repo
|
||||||
|
(`agent-configs/iris`, from the default template), then has the hive build
|
||||||
|
and start the container. It returns once the job is queued — watch the
|
||||||
|
swarm UI's job view for progress.
|
||||||
|
|
||||||
## Security notes
|
Later config changes are PRs on `agent-configs/iris`, approved by you. →
|
||||||
|
[`agent-lifecycle/approvals.md`](../agent-lifecycle/approvals.md)
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
## Optional · Lock the hive dashboard
|
||||||
|
|
||||||
- **No forge admin token is stored in any agent state dir.** Agents
|
The per-hive dashboard has no SSO in front of it. To put it behind HTTP
|
||||||
hold a regular agent token, fetched from the swarm secret store into
|
Basic auth, set `services.hyperhive.gateway.auth.enable = true` and add
|
||||||
`/run/hive-agent-forge-token/token` (or, for an agent without a store
|
logins:
|
||||||
identity, the `forge-token` file hive-c0re wrote before); sensitive
|
|
||||||
creds (the core token) live on the host.
|
|
||||||
- All config changes (forge PRs on `agent-configs/<name>`) go through
|
|
||||||
operator approval — agents can't unilaterally rebuild containers, by design.
|
|
||||||
See [`boundary.md`](../trust-boundary/boundary.md) and [`security.md`](../trust-boundary/security.md).
|
|
||||||
- **Each hive authenticates its own telemetry ingest**, and the `hive` label comes
|
|
||||||
from which hive authenticated rather than from the payload — so no hive can
|
|
||||||
report metrics as another. A first-run all-local hive gets this with nothing
|
|
||||||
to configure; joining a swarm you don't host needs one secret copied across.
|
|
||||||
See [`observability.md`](../scheduler/observability.md#authenticated-ingest).
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
Once the hive is running, ruth records anything it needs to remember
|
```bash
|
||||||
across restarts in `/agents/ruth/state/notes.md`.
|
echo "hunter2" | hivectl gateway create-user mara --password-stdin
|
||||||
|
```
|
||||||
|
|
||||||
|
→ [`networking/gateway.md`](../networking/gateway.md#http-basic-auth)
|
||||||
|
|
||||||
|
## Day to day
|
||||||
|
|
||||||
|
```bash
|
||||||
|
hivectl list-agents # this hive's agents, status, rev
|
||||||
|
hivectl agent <name> restart # stop + start, no rebuild
|
||||||
|
hivectl agent <name> watch # follow its live event stream
|
||||||
|
hivectl agent <name> choom # interactive claude session in its container
|
||||||
|
hivectl approvals pending # what's waiting on you
|
||||||
|
hivectl open # this hive's dashboard (or: forge, matrix)
|
||||||
|
```
|
||||||
|
|
||||||
|
Every verb: [`tools/hivectl.md`](../tools/hivectl.md) ·
|
||||||
|
[`tools/swarmctl-cli.md`](../tools/swarmctl-cli.md).
|
||||||
|
|
|
||||||
134
docs/swarm/bao.md
Normal file
134
docs/swarm/bao.md
Normal file
|
|
@ -0,0 +1,134 @@
|
||||||
|
# The swarm secret store (OpenBao)
|
||||||
|
|
||||||
|
The swarm runs one OpenBao store, `swarm-bao`. Every swarm-level
|
||||||
|
credential an agent or service needs is fetched from it under that
|
||||||
|
principal's own certificate identity. What lives in it and who reads what:
|
||||||
|
[`secrets.md`](secrets.md). This page covers the store itself — how it
|
||||||
|
comes up, who may write its grants, and how it's sealed and reached.
|
||||||
|
|
||||||
|
The operator steps (init, the one-time granter bootstrap) are in
|
||||||
|
[`../getting-started/setup.md`](../getting-started/setup.md#1--secret-store);
|
||||||
|
this page is the why behind them.
|
||||||
|
|
||||||
|
<!-- vale write-good.Passive = NO -->
|
||||||
|
|
||||||
|
## The granter
|
||||||
|
|
||||||
|
Cert auth answers a _role_, so nothing can authenticate until some role
|
||||||
|
exists. A short-lived bootstrap token breaks that cycle once, and it's the
|
||||||
|
only step that needs the root token.
|
||||||
|
|
||||||
|
The policy that token carries is `nix/host-modules/bao-bootstrap-policy.hcl`,
|
||||||
|
shipped on the store's host at `/etc/hyperhive/bao-bootstrap-policy.hcl`. It
|
||||||
|
covers the auth mounts and the granter's own policy and role, and nothing
|
||||||
|
else. CI fails when the unit using the token needs a path it lacks. The token
|
||||||
|
file is `services.hyperhive.deploy.bao.bootstrapTokenFile`, which all-local
|
||||||
|
names for you; on a store host that isn't all-local, set it and rebuild first.
|
||||||
|
|
||||||
|
`swarm-bao-granter-role` runs **on the host**. It enables the cert and oidc
|
||||||
|
auth methods (and disables `approle` if an older store still has it mounted),
|
||||||
|
writes the `bao-granter` policy, and creates the `bao-granter` role, which
|
||||||
|
accepts the leaf `/var/lib/swarm-bao-pki/granter.pem`. Every
|
||||||
|
`swarm-bao-*-policy` unit then logs in with that leaf. The controller's unit
|
||||||
|
mounts the KV and pki engines and writes the `swarm-controller` role, and each
|
||||||
|
sibling unit writes its own principal's policy and role. Every one runs on the
|
||||||
|
host, because every API listener but the loopback UI one demands a client
|
||||||
|
certificate and the host is the side that has one.
|
||||||
|
|
||||||
|
Until the granter exists, each `swarm-bao-*-policy` unit **fails** and logs
|
||||||
|
the bootstrap commands. It never skips. `swarm-bao-granter-role` skips while
|
||||||
|
the token file is absent, which is the steady state afterwards; the token's
|
||||||
|
TTL means a forgotten file expires rather than lingering.
|
||||||
|
|
||||||
|
After that, a new or changed `swarm-*` grant needs no operator step: the unit
|
||||||
|
that writes it changes, and the deploy restarts it. A root step comes back only
|
||||||
|
when the granter itself needs a path it lacks, such as a new mount.
|
||||||
|
|
||||||
|
⚠️ The granting units re-run on **boot** and whenever a deploy **changes**
|
||||||
|
them, not on every deploy. When a grant drifts in the store and its unit stays
|
||||||
|
the same, the next boot re-asserts it, not the next switch.
|
||||||
|
|
||||||
|
⚠️ Don't reach for `bao read auth/cert/…` to check a role. The host's `bao`
|
||||||
|
wrapper carries an address, a CA and a client certificate but deliberately
|
||||||
|
**no token**, so that read answers `403` whether or not the role exists. Read
|
||||||
|
the unit's journal instead.
|
||||||
|
|
||||||
|
<details><summary>Upgrading a swarm set up with the older swarm-bootstrap policy</summary>
|
||||||
|
|
||||||
|
A store set up before the granter existed has every grant, but no `bao-granter`
|
||||||
|
policy or role. After the deploy that introduces it, each `swarm-bao-*-policy`
|
||||||
|
unit fails and logs the one-time step. Run the setup step as it stands. The
|
||||||
|
old policy can go, with the root token again:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bao policy delete swarm-bootstrap
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
### Residual risk
|
||||||
|
|
||||||
|
The granter is root-equivalent. It may write any `swarm-*` policy with any
|
||||||
|
content, and attach it to a role that accepts any certificate; no bao ACL can
|
||||||
|
constrain what a policy says. What bounds it:
|
||||||
|
|
||||||
|
- `nix/host-modules/swarm-bao.nix` renders every policy it writes, and
|
||||||
|
module-eval pins each principal's grants. **Merging a change to that policy
|
||||||
|
text is granting it**: it takes effect on the next deploy with no bao step,
|
||||||
|
so code review is the only gate.
|
||||||
|
- Its key sits permanently at `/var/lib/swarm-bao-pki/granter-key.pem`, `0600`
|
||||||
|
root in a `0700` directory, readable only by root units on the store's host.
|
||||||
|
That host already holds `ca-key.pem`, which can mint a leaf with any subject,
|
||||||
|
and `controller-key.pem`, whose policy is already root-equivalent. Root on
|
||||||
|
that host gains nothing new.
|
||||||
|
- **Never copy `granter-key.pem` off the host** the way operators copy the
|
||||||
|
other leaves in that directory. That hands out root-equivalence.
|
||||||
|
- Nothing revokes a stolen leaf on its own: the role trusts the CA plus the
|
||||||
|
subject. Rotate the store's CA, or have root point the `bao-granter` role at
|
||||||
|
a new `deploy.bao.granterCommonName`. Deleting `granter{,-key}.pem` and
|
||||||
|
restarting `swarm-bao-pki` mints a new leaf, but doesn't invalidate the old
|
||||||
|
one.
|
||||||
|
|
||||||
|
## Sealing
|
||||||
|
|
||||||
|
`services.hyperhive.deploy.bao.seal`:
|
||||||
|
|
||||||
|
- **`pkcs11`** (the default) — binds the key to the host's TPM, and the store
|
||||||
|
unseals itself on every restart. `init` is the only manual step.
|
||||||
|
- **`shamir`** — no TPM, so run `bao operator unseal` again after every
|
||||||
|
restart, with the keys `init` printed.
|
||||||
|
|
||||||
|
⚠️ **A sealed store still answers.** The container is up and the port
|
||||||
|
responds while every read times out — the failure looks like a hang, not like
|
||||||
|
a store that was never initialised.
|
||||||
|
|
||||||
|
## TLS
|
||||||
|
|
||||||
|
The store serves TLS, and on a host that deploys it you need do nothing: a
|
||||||
|
first-boot unit mints a CA of the store's own plus the leaves it signs — the
|
||||||
|
store's server certificate, this host's client certificate, and one per service
|
||||||
|
principal — and points `deploy.bao.serverCertFile`, `.serverKeyFile` and
|
||||||
|
`.clientCaFile` at the store's half, `.clientCertFile`, `.clientKeyFile` and
|
||||||
|
`.serverCaFile` at the reader's, and each principal's own pair at its own leaf.
|
||||||
|
|
||||||
|
Those are `mkDefault`s, so naming your own paths wins. Do that when your
|
||||||
|
certificates come from a real internal CA; the store has no opinion about
|
||||||
|
which. A host that does **not** deploy the store names the reader's three
|
||||||
|
itself, plus a pair for every principal it runs — see
|
||||||
|
[per-principal identities](secrets.md#per-principal-identities) for the list.
|
||||||
|
The operator issues those leaves out of band; they're the credentials the
|
||||||
|
store can't hand you, being what opens it.
|
||||||
|
|
||||||
|
⚠️ Not the gateway's HTTPS certificates and not the hive CA — this is **mTLS
|
||||||
|
between services and the store**, a separate trust domain, because a store
|
||||||
|
that took its identity from an authority it itself distributes could never
|
||||||
|
come up before that authority.
|
||||||
|
|
||||||
|
Inside the store's container (`nixos-container root-login swarm-bao`) the
|
||||||
|
`bao` CLI needs two extra pieces, because the certificate carries no IP SAN
|
||||||
|
and its DNS name resolves to the bridge from in there: export
|
||||||
|
`BAO_ADDR=https://127.0.0.1:8200` alongside `BAO_TLS_SERVER_NAME=bao.<swarm
|
||||||
|
domain>`. The host's `bao` wrapper already carries the address, CA and client
|
||||||
|
certificate, so the host is the shorter path.
|
||||||
|
|
||||||
|
<!-- vale write-good.Passive = YES -->
|
||||||
Loading…
Reference in a new issue