hyperhive/swarm-controller/README.md
atlas 07852cabc1 feat(3088): move the gateway's nginx + dnsmasq onto the host
The gateway's nginx + dnsmasq no longer run in their own nspawn container.
`nix/host-modules/hive-gateway/default.nix` loses the
`containers.hive-gateway` wrapper and everything that existed only to punch
holes in it: `privateNetwork = false`, `CAP_NET_ADMIN`, five bind mounts,
its own `stateVersion`, `networking.firewall.enable = false`,
`networking.resolvconf.enable = false`, and the `hive-gateway-resolv`
path+service pair. 465 -> 303 lines.

The container never bought isolation here. It shared the host netns by
necessity — nginx binds the host's :80/:443, dnsmasq answers on the bridge —
so each of those settings was undoing a boundary the gateway could not
afford in the first place.

Four things made it more than a deletion, none of them visible in the nix
diff:

- The self-signed cert service also imports the hive CA leaf, so removing it
  with the container would have left nginx naming a missing cert file, which
  it refuses to load at all.
- The nginx reload is a hive-priv verb. It still needs root, but no longer
  for the reason its doc gave, and `--machine=` was both transport and
  scope — so the unit name is now hard-coded in the helper as the
  containment.
- The lifecycle verb named a container that stops existing.
- `journalctl -M hive-gateway` had no machine to enter.

Per the operator's ruling, the operator verb keeps working and agents lose
it. `InfraContainer` answered three questions that used to share an answer;
it now splits into `name()` (identity), `target()` (Container vs HostUnit),
`service_unit()` (the systemd unit), and `agent_restartable()`, which the
MCP restart path checks before the capability so the refusal cannot read as
"ask for infra_admin". `SIBLING_CONTAINERS` drops the gateway — it gates the
requests that name a container as a string — while `FromStr` still accepts
it, because that answers what a name is, not who may act on it. The
dashboard's gateway journal reads host journald filtered to `nginx.service`.

Prose was corrected where it only named a location, and re-argued where the
container was doing security work: a `0666` per-agent socket was safe
because only the gateway container had the directory bind-mounted. There is
no mount now, so the directory permissions are the whole of the access
control — the constraint holds, its mechanism doesn't.

Gate: nix fmt / clippy --all-targets -D warnings / cargo test all clean (710
tests); hivectl-cli.md regenerated from the clap tree. The nix eval was run
in both TLS shapes at this commit: every delta in the rendered
virtualHosts is one of the three intended path moves, dnsmasq settings are
byte-identical, and the absence probe flips true -> false with bindMounts
emptied.
2026-08-11 18:01:03 +02:00

61 lines
2.7 KiB
Markdown

# swarm-controller
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
owns what is true *across* hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via `services.hyperhive.swarm.controller.enable`, which is
deliberately **not** derived from `services.hyperhive.enable`: turning it on is
a statement about swarm topology, not about whether hyperhive is installed.
## What it does today
Serves one `/health` endpoint and holds no state.
That is the whole intent of the first slice. The point is to make the *unit*
real — service user, runtime and state directories, socket, nginx
reachability — so the swarm-level surfaces that follow have somewhere to land.
Inventing those surfaces before they are agreed would bake in a shape nobody
chose. See #3066 and the `hyperhive.swarm` consolidation epic.
## Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket
through a bind-mount. A listener that is never bound to an address cannot be
reached from off-host by mistake.
The socket path is `services.hyperhive.swarm.controller.socketPath`, default
`/run/swarm-controller/controller.sock`, exported to the process as
`SWARM_CONTROLLER_SOCKET`.
## ⚠️ The socket's directory is its access control
The socket is `0666`. It has to be: nginx runs as a different user and
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
sockets, and rests on the same argument — *"the bind source dir is per-agent on
host so blast radius is unchanged."*
What keeps that safe is that the directory holds **one** socket. So:
> **Never point `socketPath` at a directory that carries anything else.**
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
> Pointing nginx at that directory to reach this socket would put the admin
> socket within its reach too.
This got *less* forgiving when nginx moved onto the host: the gateway used to
reach a unix upstream through a bind-mount, so the mount list was a second
bound on what it could touch. There is no mount now — the directory is the
whole of the access control. A unit test pins the default path so a tidying
edit fails instead of reviewing cleanly.
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
a **pair**: preserving the directory without the unlink means `bind` fails with
`EADDRINUSE` after a restart.
## Packaging
Built by the workspace derivation and extracted as its own package
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
`daemonBins` — that list is the core stack and drives the bundle
`services.hyperhive.c0re.package` points at, so folding this in would put a
swarm-scoped service into every hive's closure.