Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/swarm-controller/README.md
atlas 270430a4b4 docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
2026-10-02 12:50:34 +02:00

79 lines
3.8 KiB
Markdown

# swarm-controller
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
owns what is true _across_ hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via `services.hyperhive.deploy.swarm-controller.enable`, which is
deliberately **not** derived from
`services.hyperhive.deploy.hive-controller.enable`: turning it on is a statement
about swarm topology, not about whether this host runs a hive.
## What it does
The swarm's control plane. The swarm UI and `swarmctl` are its clients.
- **Hive directory** — serves `swarm.hives` (`GET /api/hives`) and what each
hive last published about itself (`GET /api/hives/status`).
- **Agent roster and wanted state** — every agent the swarm knows
(`GET /api/agents`, `/api/agents/status`), and the state it declares for
each one on its hive (`up`/`offline`/`paused`/`destroyed`,
`PUT /api/hives/{hive}/agents/{agent}/state`).
- **Job graph** — a `hive-jobq` scheduler, served at `GET /api/jobq/graph`.
Every provisioning step below runs as a node in it.
- **Agent creation** — `POST /api/agents` queues the SSO identity (through
`swarm-authelia-bridge`), forge user, config repo, store identity, forge
token and matrix account, declares the agent `paused`, then sends its hive
a deploy message.
- **Agent credentials** — at start and every five minutes it re-checks every
agent's forge token and matrix account, and renews store certificates and
queue secrets as they age.
- **Swarm-wide forge objects and webhooks** →
[`docs/swarm/README.md`](../docs/swarm/README.md#swarm-wide-forge-objects).
- **Relays** — each agent's terminal and turn-state header as SSE, agent
icons, the cross-repo issue report, and the UI's quick links.
It reads its configuration once at startup, from the environment the nix module
sets. The one file it persists is `webhook-secret` in its state directory.
The job graph lives in memory; hive status and wanted state live in the swarm
queue, so both survive a restart.
## Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket
through a bind-mount. A listener that is never bound to an address cannot be
reached from off-host by mistake.
The socket path is `services.hyperhive.deploy.swarm-controller.socketPath`, default
`/run/swarm-controller/controller.sock`, exported to the process as
`SWARM_CONTROLLER_SOCKET`.
## ⚠️ The socket's directory is its access control
The socket is `0666`. It has to be: nginx runs as a different user and
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
sockets, and rests on the same argument — _"the bind source dir is per-agent on
host so blast radius is unchanged."_
What keeps that safe is that the directory holds **one** socket. So:
> **Never point `socketPath` at a directory that carries anything else.**
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
> Pointing nginx at that directory to reach this socket would put the admin
> socket within its reach too.
nginx is a host service, so nothing narrows what it can reach except the
directory itself — that is the whole of the access control. A unit test pins
the default path so a tidying edit fails instead of reviewing cleanly.
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
a **pair**: preserving the directory without the unlink means `bind` fails with
`EADDRINUSE` after a restart.
## Packaging
Built by the workspace derivation and extracted as its own package
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
`daemonBins` — that list is the core stack and drives the bundle
`services.hyperhive.c0re.package` points at, so folding this in would put a
swarm-scoped service into every hive's closure.