Two corrections from review, applied forward on this branch rather than
by rewriting it.
`deploy.<service>` was a bare bool, which makes
`deploy.forgejo = { enable; ci; }` unrepresentable -- the nested
CI-runner sub-option this namespace was designed around. Every entry is
now an attrset with an `enable`, so a second per-host deployment
decision becomes an ordinary addition rather than a migration.
`deploy.controller` is now `deploy.swarm-controller`, consistent with
`deploy.swarm-ui`, which was introduced in the same commit.
89 references rewritten across 24 files -- nix, Rust, docs, and the
repo's own CLAUDE.md.
The prefix-anchored sweep missed exactly one, and it was live code:
hive-tls.nix spells it `hyperhiveCfg.deploy.controller` -- the only
`hyperhiveCfg` prefix among 45 references. A suffix grep
(`\.deploy\.<name>`) finds it; a path-anchored one cannot, because the
head of a reference is whatever alias the reading file happens to bind.
59 lines
2.6 KiB
Markdown
59 lines
2.6 KiB
Markdown
# swarm-controller
|
|
|
|
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
|
|
owns what is true *across* hives — so a swarm runs one of them and most hives
|
|
leave it off.
|
|
|
|
Opt-in per host via `services.hyperhive.deploy.swarm-controller.enable`, which is
|
|
deliberately **not** derived from `services.hyperhive.enable`: turning it on is
|
|
a statement about swarm topology, not about whether hyperhive is installed.
|
|
|
|
## What it does today
|
|
|
|
Serves one `/health` endpoint and holds no state.
|
|
|
|
That is the whole intent of the first slice. The point is to make the *unit*
|
|
real — service user, runtime and state directories, socket, nginx
|
|
reachability — so the swarm-level surfaces that follow have somewhere to land.
|
|
Inventing those surfaces before they are agreed would bake in a shape nobody
|
|
chose. See #3066 and the `hyperhive.swarm` consolidation epic.
|
|
|
|
## Why a unix socket, not a port
|
|
|
|
The hive-gateway's nginx is the only intended client and reaches the socket
|
|
through a bind-mount. A listener that is never bound to an address cannot be
|
|
reached from off-host by mistake.
|
|
|
|
The socket path is `services.hyperhive.swarm.controller.socketPath`, default
|
|
`/run/swarm-controller/controller.sock`, exported to the process as
|
|
`SWARM_CONTROLLER_SOCKET`.
|
|
|
|
## ⚠️ The socket's directory is its access control
|
|
|
|
The socket is `0666`. It has to be: nginx runs as a different user and
|
|
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
|
|
sockets, and rests on the same argument — *"the bind source dir is per-agent on
|
|
host so blast radius is unchanged."*
|
|
|
|
What keeps that safe is that the directory holds **one** socket. So:
|
|
|
|
> **Never point `socketPath` at a directory that carries anything else.**
|
|
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
|
|
> Pointing nginx at that directory to reach this socket would put the admin
|
|
> socket within its reach too.
|
|
|
|
nginx is a host service, so nothing narrows what it can reach except the
|
|
directory itself — that is the whole of the access control. A unit test pins
|
|
the default path so a tidying edit fails instead of reviewing cleanly.
|
|
|
|
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
|
|
a **pair**: preserving the directory without the unlink means `bind` fails with
|
|
`EADDRINUSE` after a restart.
|
|
|
|
## Packaging
|
|
|
|
Built by the workspace derivation and extracted as its own package
|
|
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
|
|
`daemonBins` — that list is the core stack and drives the bundle
|
|
`services.hyperhive.c0re.package` points at, so folding this in would put a
|
|
swarm-scoped service into every hive's closure.
|