hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 395ecbdf41 swarm-secret-client: name the hive queue credential, and grant a hive its own kind
The agreement half of delivering the agent queue principal's client secret
through the store. No producer yet, so nothing writes this path — the unit
that does lands in the same PR, with the write grant it needs.

queue.rs is the sibling matrix.rs prescribes for a second kind of secret
rather than another field on a shared struct. Keyed per HIVE, not per agent:
the queue identity is minted once per hive at deploy time and says which hive
an agent belongs to, never which agent.

The client id rides with the secret for matrix.rs's stated reason — a
credential has to be reconstructable from the store alone, and deriving
`hive-<name>-agent` on the reading side is the split spelling the authelia
module warns denies every agent as a timeout.

policy.rs's render() takes the hive name now and emits a second, narrow
stanza for that hive's own path. The agent stanza is untouched: an agent's
path does not name its hive, so narrowing it still needs the enumeration
docs/trust-boundary/security.md rejects. A hive path does name its principal,
so scoping it costs nothing and drifts nowhere.

every_hive_gets_a_byte_identical_document is replaced rather than deleted.
Its surviving half is that the text is a function of the deploy-time name
alone, so a re-emission cannot drift; the new arms are that one hive's
document cannot reach another's path, and that a name which could close the
stanza is refused — live again now that a name reaches the document text.

Refs #3853
2026-09-12 10:56:50 +02:00
..
src swarm-secret-client: name the hive queue credential, and grant a hive its own kind 2026-09-12 10:56:50 +02:00
Cargo.toml convert hand-written enum as_str matches to strum derives workspace-wide 2026-09-12 00:06:31 +02:00
README.md check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is deliberately not derived from services.hyperhive.enable: turning it on is a statement about swarm topology, not about whether hyperhive is installed.

What it does today

Serves one /health endpoint and holds no state.

That is the whole intent of the first slice. The point is to make the unit real — service user, runtime and state directories, socket, nginx reachability — so the swarm-level surfaces that follow have somewhere to land. Inventing those surfaces before they are agreed would bake in a shape nobody chose. See the hyperhive.swarm consolidation epic.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.