Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 5971b51e65 swarm-controller: refuse a link over any stored object, decodable or not
The existence check read the path as the route's credential type, so an
object stored there that no longer decodes as one answered 500 instead of
409. It now reads the path untyped: anything stored holds the name.
2026-10-03 13:49:03 +02:00
..
src swarm-controller: refuse a link over any stored object, decodable or not 2026-10-03 13:49:03 +02:00
Cargo.toml swarm-controller: create every agent's subagent stream 2026-10-03 01:34:01 +02:00
README.md swarm UI: delete linked accounts — R2 fixes 2026-10-03 00:56:39 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is deliberately not derived from services.hyperhive.deploy.hive-controller.enable: turning it on is a statement about swarm topology, not about whether this host runs a hive.

What it does

The swarm's control plane. The swarm UI and swarmctl are its clients.

  • Hive directory — serves swarm.hives (GET /api/hives) and what each hive last published about itself (GET /api/hives/status).
  • Agent roster and wanted state — every agent the swarm knows (GET /api/agents, /api/agents/status), and the state it declares for each one on its hive (up/offline/paused/destroyed, PUT /api/hives/{hive}/agents/{agent}/state).
  • Job graph — a hive-jobq scheduler, served at GET /api/jobq/graph. Every provisioning step below runs as a node in it.
  • Agent creation — POST /api/agents queues the SSO identity (through swarm-authelia-bridge), forge user, config repo, store identity, forge token and matrix account, declares the agent paused, then sends its hive a deploy message.
  • Agent credentials — at start and every five minutes it re-checks every agent's forge token and matrix account, and renews store certificates and queue secrets as they age.
  • Linked external accounts — an operator-supplied matrix, forge or GitHub account for one agent, stored in the swarm secret store (PUT /api/hives/{hive}/agents/{agent}/matrix-accounts/{account}, .../forge-accounts/{label}, .../github-account); distinct from the agent's own swarm-minted accounts above. GET .../linked-accounts lists them, plus the agent's own main matrix account, by kind, name and host, never with a credential. Each of the three also takes a DELETE; matrix's also takes ?revoke=true, which logs the stored token out at its homeserver first and leaves the account in place if that fails.
  • Config PR status — each agent's open config-repo PR, cached from forge webhooks (GET /api/config-prs, /api/agents/{name}/config-pr).
  • Swarm-wide forge objects and webhooks → docs/swarm/README.md.
  • Relays — each agent's terminal and turn-state header as SSE, agent icons, the cross-repo issue report, and the UI's quick links.

It reads its configuration once at startup, from the environment the nix module sets. The one file it persists is webhook-secret in its state directory. The job graph lives in memory; hive status and wanted state live in the swarm queue, so both survive a restart.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.