Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 3380c1915f swarm UI: show the accounts linked to each agent
GET /api/hives/{hive}/agents/{agent}/linked-accounts returns one row per
account linked to the agent, as kind, name and host: each matrix account
under swarm/agents/<agent>/matrix (with its homeserver, and the agent's own
`main` marked reserved), each forge label under swarm/agents/<agent>/forge
(with its url), and github when swarm/agents/<agent>/github-token exists
(host github.com, which is not stored). No credential field is in the
response type.

Listing those two directories needs a new controller grant: `list` on
secret/metadata/swarm/agents/+/matrix and .../+/forge only, pinned in
bao-grants.nix as the only metadata stanzas under agents/ beside the queue
revocation. Checked against a dev OpenBao 2.6.3: the grant lists those two
directories and is refused on agents/, agents/<agent>/, and a leaf.

The swarm UI agent detail panel shows all rows under "accounts"; the table
view's matrix column shows the matrix rows. The link badges stay.

Refs #4855
2026-10-02 21:20:40 +02:00
..
src swarm UI: show the accounts linked to each agent 2026-10-02 21:20:40 +02:00
Cargo.toml swarm-controller: first credential-renewal pass waits for the queue connection 2026-09-30 15:46:56 +02:00
README.md swarm UI: show the accounts linked to each agent 2026-10-02 21:20:40 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is deliberately not derived from services.hyperhive.deploy.hive-controller.enable: turning it on is a statement about swarm topology, not about whether this host runs a hive.

What it does

The swarm's control plane. The swarm UI and swarmctl are its clients.

  • Hive directory — serves swarm.hives (GET /api/hives) and what each hive last published about itself (GET /api/hives/status).
  • Agent roster and wanted state — every agent the swarm knows (GET /api/agents, /api/agents/status), and the state it declares for each one on its hive (up/offline/paused/destroyed, PUT /api/hives/{hive}/agents/{agent}/state).
  • Job graph — a hive-jobq scheduler, served at GET /api/jobq/graph. Every provisioning step below runs as a node in it.
  • Agent creation — POST /api/agents queues the SSO identity (through swarm-authelia-bridge), forge user, config repo, store identity, forge token and matrix account, declares the agent paused, then sends its hive a deploy message.
  • Agent credentials — at start and every five minutes it re-checks every agent's forge token and matrix account, and renews store certificates and queue secrets as they age.
  • Linked external accounts — an operator-supplied matrix, forge or GitHub account for one agent, stored in the swarm secret store (PUT /api/hives/{hive}/agents/{agent}/matrix-accounts/{account}, .../forge-accounts/{label}, .../github-account); distinct from the agent's own swarm-minted accounts above. GET .../linked-accounts lists them, plus the agent's own main matrix account, by kind, name and host, never with a credential.
  • Config PR status — each agent's open config-repo PR, cached from forge webhooks (GET /api/config-prs, /api/agents/{name}/config-pr).
  • Swarm-wide forge objects and webhooks → docs/swarm/README.md.
  • Relays — each agent's terminal and turn-state header as SSE, agent icons, the cross-repo issue report, and the UI's quick links.

It reads its configuration once at startup, from the environment the nix module sets. The one file it persists is webhook-secret in its state directory. The job graph lives in memory; hive status and wanted state live in the swarm queue, so both survive a restart.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.