hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 9b939f4626 refactor(#3255): one knowledge subject, single writer and many readers
Review call: the event was addressed per hive — `$SWARM.events.<hive>.knowledge`,
published in a loop over the roster, granted through a wildcard. It does not
need to be. The payload is empty and the event means the same thing to every
hive, so one publish to one subject delivers exactly what N publishes to N
subjects did, and core NATS already fans out to whoever is subscribed. A hive
that was down misses it either way and reconciles on its next periodic pull.

That deletes rather than reshuffles: the roster loop, the wildcard, and the
shared subject-building function whose entire purpose was keeping the grant and
the publish from drifting apart. With one literal there is nothing to disagree
about.

The per-hive shape was justified by the callout policy's rule that an extra
subject must contain the hive name. That rule governs `extra_hive_subjects` —
what a HIVE may publish. This subject lives in the controller's reader grant,
which the rule does not constrain, so a real rule was carried across into a
decision it had no authority over.

Knowledge becomes its own category rather than a leaf under a general event
namespace, since a namespace shaped for events that do not exist yet is a
decision made before there is anything to decide from. The empty config-PR match
arm goes with it: an arm with no body claims this is where the deploy path is
handled, and it is not.

The deny test stays and matters more, not less: with one shared subject a forged
event would reach the whole swarm where a per-hive one reached a single hive.
2026-08-19 21:05:52 +02:00
..
src refactor(#3255): one knowledge subject, single writer and many readers 2026-08-19 21:05:52 +02:00
Cargo.toml swarm-controller: extract jobq_metrics into its own hive-jobq-metrics crate 2026-08-19 18:10:22 +02:00
README.md docs(gateway): describe what is, not what changed 2026-08-11 18:09:51 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.swarm.controller.enable, which is deliberately not derived from services.hyperhive.enable: turning it on is a statement about swarm topology, not about whether hyperhive is installed.

What it does today

Serves one /health endpoint and holds no state.

That is the whole intent of the first slice. The point is to make the unit real — service user, runtime and state directories, socket, nginx reachability — so the swarm-level surfaces that follow have somewhere to land. Inventing those surfaces before they are agreed would bake in a shape nobody chose. See #3066 and the hyperhive.swarm consolidation epic.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.swarm.controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.