hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
iris 513554fe9a swarm: add a declared "paused" agent wanted state
mara (#4170): swarm-ui's wanted-state dropdown could only ever declare
up/offline/destroy, with no way to swarm-declare the existing hive-local
turn-loop pause (`hivectl agent pause|resume`).

`AgentState::Paused` is not a fifth peer of Up/Offline/Destroyed on the
power axis this enum otherwise answers — it's Up plus an orthogonal
turn-loop pause. `hive-c0re`'s `workers::wanted` reconcile loop now
decides the two axes independently (`decide` for power, the new
`decide_pause` for the marker), so a stopped agent declared Paused
converges with both a Start and a Pause in the same pass.

Known, deliberate limitation: a Paused declaration on an agent this
hive has never deployed only reaches Deploy this pass — writing the
pause marker into a harness dir that may not exist yet was judged not
worth the risk, so it converges on the next pass once the agent is
present instead.

swarm-ui's WantedMenu gains a fourth "paused" option (warning-tone
badge). No separate "resume" entry — selecting "up" from a paused row
already clears the marker via the same decide_pause path.

Pause/resume marker writes go through one shared
Coordinator::set_paused_by_name helper, used by both the interactive
dashboard pause/resume handlers and this reconcile loop, instead of
each duplicating the parse-name/write-marker/track-rescan shape.
swarm-ui's "offline" and "paused" confirm dialogs share one
confirmTarget state and one ConfirmDialog instead of two near-identical
copies.

Closes #4170
2026-09-11 01:35:40 +02:00
..
src swarm: add a declared "paused" agent wanted state 2026-09-11 01:35:40 +02:00
Cargo.toml swarm-controller: accept an agent's external matrix account and put it in the store 2026-09-08 15:53:50 +02:00
README.md check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is deliberately not derived from services.hyperhive.enable: turning it on is a statement about swarm topology, not about whether hyperhive is installed.

What it does today

Serves one /health endpoint and holds no state.

That is the whole intent of the first slice. The point is to make the unit real — service user, runtime and state directories, socket, nginx reachability — so the swarm-level surfaces that follow have somewhere to land. Inventing those surfaces before they are agreed would bake in a shape nobody chose. See the hyperhive.swarm consolidation epic.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.