Watch
0
0
Fork
You've already forked hyperhive
0

docs(swarm): facts + structure pass

swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
This commit is contained in:
atlas 2026-10-01 23:25:24 +02:00 • committed by mara
commit 270430a4b4
5 changed files with 408 additions and 402 deletions

View file

@ -9,15 +9,34 @@ deliberately **not** derived from
`services.hyperhive.deploy.hive-controller.enable`: turning it on is a statement
about swarm topology, not about whether this host runs a hive.
## What it does today
## What it does
Serves one `/health` endpoint and holds no state.
The swarm's control plane. The swarm UI and `swarmctl` are its clients.
That is the whole intent of the first slice. The point is to make the _unit_
real — service user, runtime and state directories, socket, nginx
reachability — so the swarm-level surfaces that follow have somewhere to land.
Inventing those surfaces before they are agreed would bake in a shape nobody
chose. See the `hyperhive.swarm` consolidation epic.
- **Hive directory** — serves `swarm.hives` (`GET /api/hives`) and what each
hive last published about itself (`GET /api/hives/status`).
- **Agent roster and wanted state** — every agent the swarm knows
(`GET /api/agents`, `/api/agents/status`), and the state it declares for
each one on its hive (`up`/`offline`/`paused`/`destroyed`,
`PUT /api/hives/{hive}/agents/{agent}/state`).
- **Job graph** — a `hive-jobq` scheduler, served at `GET /api/jobq/graph`.
Every provisioning step below runs as a node in it.
- **Agent creation** — `POST /api/agents` queues the SSO identity (through
`swarm-authelia-bridge`), forge user, config repo, store identity, forge
token and matrix account, declares the agent `paused`, then sends its hive
a deploy message.
- **Agent credentials** — at start and every five minutes it re-checks every
agent's forge token and matrix account, and renews store certificates and
queue secrets as they age.
- **Swarm-wide forge objects and webhooks** →
[`docs/swarm/README.md`](../docs/swarm/README.md#swarm-wide-forge-objects).
- **Relays** — each agent's terminal and turn-state header as SSE, agent
icons, the cross-repo issue report, and the UI's quick links.
It reads its configuration once at startup, from the environment the nix module
sets. The one file it persists is `webhook-secret` in its state directory.
The job graph lives in memory; hive status and wanted state live in the swarm
queue, so both survive a restart.
## Why a unix socket, not a port