One commit rather than two because they are not independent: the UI's `enable` had the controller's as its literal default, so moving the controller alone would leave the UI's default naming an option that no longer exists. The UI keeps that derivation in its new home — it is a view onto the controller's state and reaches it over that daemon's unix socket, so the host running the controller is the host that can serve it. Three spellings had to move together for the UI, not one: the `default`, the `defaultText` shown in the options doc, and the description prose that names the old path in words. A grep for the option path finds the first two. The sweep also reached outside nix: `swarm-controller`'s crate README and its `//!` module doc both named the option, as did this repo's own CLAUDE.md and four pages under docs/. An option's name is API, and its documentation lives wherever someone thought to write it down.
59 lines
2.6 KiB
Markdown
59 lines
2.6 KiB
Markdown
# swarm-controller
|
|
|
|
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
|
|
owns what is true *across* hives — so a swarm runs one of them and most hives
|
|
leave it off.
|
|
|
|
Opt-in per host via `services.hyperhive.deploy.controller`, which is
|
|
deliberately **not** derived from `services.hyperhive.enable`: turning it on is
|
|
a statement about swarm topology, not about whether hyperhive is installed.
|
|
|
|
## What it does today
|
|
|
|
Serves one `/health` endpoint and holds no state.
|
|
|
|
That is the whole intent of the first slice. The point is to make the *unit*
|
|
real — service user, runtime and state directories, socket, nginx
|
|
reachability — so the swarm-level surfaces that follow have somewhere to land.
|
|
Inventing those surfaces before they are agreed would bake in a shape nobody
|
|
chose. See #3066 and the `hyperhive.swarm` consolidation epic.
|
|
|
|
## Why a unix socket, not a port
|
|
|
|
The hive-gateway's nginx is the only intended client and reaches the socket
|
|
through a bind-mount. A listener that is never bound to an address cannot be
|
|
reached from off-host by mistake.
|
|
|
|
The socket path is `services.hyperhive.swarm.controller.socketPath`, default
|
|
`/run/swarm-controller/controller.sock`, exported to the process as
|
|
`SWARM_CONTROLLER_SOCKET`.
|
|
|
|
## ⚠️ The socket's directory is its access control
|
|
|
|
The socket is `0666`. It has to be: nginx runs as a different user and
|
|
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
|
|
sockets, and rests on the same argument — *"the bind source dir is per-agent on
|
|
host so blast radius is unchanged."*
|
|
|
|
What keeps that safe is that the directory holds **one** socket. So:
|
|
|
|
> **Never point `socketPath` at a directory that carries anything else.**
|
|
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
|
|
> Pointing nginx at that directory to reach this socket would put the admin
|
|
> socket within its reach too.
|
|
|
|
nginx is a host service, so nothing narrows what it can reach except the
|
|
directory itself — that is the whole of the access control. A unit test pins
|
|
the default path so a tidying edit fails instead of reviewing cleanly.
|
|
|
|
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
|
|
a **pair**: preserving the directory without the unlink means `bind` fails with
|
|
`EADDRINUSE` after a restart.
|
|
|
|
## Packaging
|
|
|
|
Built by the workspace derivation and extracted as its own package
|
|
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
|
|
`daemonBins` — that list is the core stack and drives the bundle
|
|
`services.hyperhive.c0re.package` points at, so folding this in would put a
|
|
swarm-scoped service into every hive's closure.
|