hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas e638db262e swarm-controller: keep a hive's read grant in step with its declaration
The controller writes an agent's credential; the hive fetches it back with
its own certificate. Nothing said which paths that certificate may read, so
the read half of a delivery answers 403 with no way to tell why.

The grant is derived from the declaration, so it is re-rendered at the one
place the declaration changes -- WantedWriter::set -- rather than at its
caller, which would work today and break on the second caller.

Emitted before the KV write: a grant that lands late is a 403 on an agent's
first fetch, while one that shrinks early only affects an agent already
being torn down. A failed write then leaves a superset the next declaration
re-renders.

Destroyed agents are filtered out. The declared set is a hive's whole
history -- a destroyed entry stays so that redeclaring it Up is refused as
the terminal transition it is -- so granting every declared agent would
leave a torn-down agent's credentials readable forever.

The sink is a trait because a missed emission is that same untraceable 403:
the double pins which agents were published, and the no-sink and refusing
arms pin the two deployments that are not a happy path. Not covered: the
call site inside set(), which needs a live queue.

The cert role moves to its own module on the way past. It is the
controller's identity at the store, not something the matrix route owns,
and the policy writer needs the same login.
2026-09-09 18:40:41 +02:00
..
src swarm-controller: keep a hive's read grant in step with its declaration 2026-09-09 18:40:41 +02:00
Cargo.toml swarm-controller: accept an agent's external matrix account and put it in the store 2026-09-08 15:53:50 +02:00
README.md deploy: move the controller's socket and credentials out of swarm.controller 2026-09-07 14:24:52 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is deliberately not derived from services.hyperhive.enable: turning it on is a statement about swarm topology, not about whether hyperhive is installed.

What it does today

Serves one /health endpoint and holds no state.

That is the whole intent of the first slice. The point is to make the unit real — service user, runtime and state directories, socket, nginx reachability — so the swarm-level surfaces that follow have somewhere to land. Inventing those surfaces before they are agreed would bake in a shape nobody chose. See #3066 and the hyperhive.swarm consolidation epic.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.