| Filename | Latest commit message | Latest commit date |
|---|---|---|
swarm-controller/README.md "What it does" was still missing two route
groups argus caught: linked external matrix/forge accounts
(PUT .../matrix-accounts/{account}, .../forge-accounts/{label}) and
config-PR status (GET /api/config-prs, /api/agents/{name}/config-pr).
Verified against the router at main.rs:2874-2899.
|
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-controller
The swarm-level daemon. Where hive-c0re owns the agents on one host, this
owns what is true across hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is
deliberately not derived from
services.hyperhive.deploy.hive-controller.enable: turning it on is a statement
about swarm topology, not about whether this host runs a hive.
What it does
The swarm's control plane. The swarm UI and swarmctl are its clients.
- Hive directory — serves
swarm.hives(GET /api/hives) and what each hive last published about itself (GET /api/hives/status). - Agent roster and wanted state — every agent the swarm knows
(
GET /api/agents,/api/agents/status), and the state it declares for each one on its hive (up/offline/paused/destroyed,PUT /api/hives/{hive}/agents/{agent}/state). - Job graph — a
hive-jobqscheduler, served atGET /api/jobq/graph. Every provisioning step below runs as a node in it. - Agent creation —
POST /api/agentsqueues the SSO identity (throughswarm-authelia-bridge), forge user, config repo, store identity, forge token and matrix account, declares the agentpaused, then sends its hive a deploy message. - Agent credentials — at start and every five minutes it re-checks every agent's forge token and matrix account, and renews store certificates and queue secrets as they age.
- Linked external accounts — an operator-supplied matrix or forge
account for one agent, stored in the swarm secret store
(
PUT /api/hives/{hive}/agents/{agent}/matrix-accounts/{account},.../forge-accounts/{label}); distinct from the agent's own swarm-minted accounts above. - Config PR status — each agent's open config-repo PR, cached from forge
webhooks (
GET /api/config-prs,/api/agents/{name}/config-pr). - Swarm-wide forge objects and webhooks →
docs/swarm/README.md. - Relays — each agent's terminal and turn-state header as SSE, agent icons, the cross-repo issue report, and the UI's quick links.
It reads its configuration once at startup, from the environment the nix module
sets. The one file it persists is webhook-secret in its state directory.
The job graph lives in memory; hive status and wanted state live in the swarm
queue, so both survive a restart.
Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.
The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default
/run/swarm-controller/controller.sock, exported to the process as
SWARM_CONTROLLER_SOCKET.
⚠️ The socket's directory is its access control
The socket is 0666. It has to be: nginx runs as a different user and
connect(2) needs write. This matches how hive-c0re publishes the per-agent
sockets, and rests on the same argument — "the bind source dir is per-agent on
host so blast radius is unchanged."
What keeps that safe is that the directory holds one socket. So:
Never point
socketPathat a directory that carries anything else./run/hyperhiveabove all — it holdshost.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.
nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.
RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are
a pair: preserving the directory without the unlink means bind fails with
EADDRINUSE after a restart.
Packaging
Built by the workspace derivation and extracted as its own package
(nix build .#swarm-controller). Deliberately not in nix/packages'
daemonBins — that list is the core stack and drives the bundle
services.hyperhive.c0re.package points at, so folding this in would put a
swarm-scoped service into every hive's closure.