The gateway's nginx + dnsmasq no longer run in their own nspawn container. `nix/host-modules/hive-gateway/default.nix` loses the `containers.hive-gateway` wrapper and everything that existed only to punch holes in it: `privateNetwork = false`, `CAP_NET_ADMIN`, five bind mounts, its own `stateVersion`, `networking.firewall.enable = false`, `networking.resolvconf.enable = false`, and the `hive-gateway-resolv` path+service pair. 465 -> 303 lines. The container never bought isolation here. It shared the host netns by necessity — nginx binds the host's :80/:443, dnsmasq answers on the bridge — so each of those settings was undoing a boundary the gateway could not afford in the first place. Four things made it more than a deletion, none of them visible in the nix diff: - The self-signed cert service also imports the hive CA leaf, so removing it with the container would have left nginx naming a missing cert file, which it refuses to load at all. - The nginx reload is a hive-priv verb. It still needs root, but no longer for the reason its doc gave, and `--machine=` was both transport and scope — so the unit name is now hard-coded in the helper as the containment. - The lifecycle verb named a container that stops existing. - `journalctl -M hive-gateway` had no machine to enter. Per the operator's ruling, the operator verb keeps working and agents lose it. `InfraContainer` answered three questions that used to share an answer; it now splits into `name()` (identity), `target()` (Container vs HostUnit), `service_unit()` (the systemd unit), and `agent_restartable()`, which the MCP restart path checks before the capability so the refusal cannot read as "ask for infra_admin". `SIBLING_CONTAINERS` drops the gateway — it gates the requests that name a container as a string — while `FromStr` still accepts it, because that answers what a name is, not who may act on it. The dashboard's gateway journal reads host journald filtered to `nginx.service`. Prose was corrected where it only named a location, and re-argued where the container was doing security work: a `0666` per-agent socket was safe because only the gateway container had the directory bind-mounted. There is no mount now, so the directory permissions are the whole of the access control — the constraint holds, its mechanism doesn't. Gate: nix fmt / clippy --all-targets -D warnings / cargo test all clean (710 tests); hivectl-cli.md regenerated from the clap tree. The nix eval was run in both TLS shapes at this commit: every delta in the rendered virtualHosts is one of the three intended path moves, dnsmasq settings are byte-identical, and the absence probe flips true -> false with bindMounts emptied.
61 lines
2.7 KiB
Markdown
61 lines
2.7 KiB
Markdown
# swarm-controller
|
|
|
|
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
|
|
owns what is true *across* hives — so a swarm runs one of them and most hives
|
|
leave it off.
|
|
|
|
Opt-in per host via `services.hyperhive.swarm.controller.enable`, which is
|
|
deliberately **not** derived from `services.hyperhive.enable`: turning it on is
|
|
a statement about swarm topology, not about whether hyperhive is installed.
|
|
|
|
## What it does today
|
|
|
|
Serves one `/health` endpoint and holds no state.
|
|
|
|
That is the whole intent of the first slice. The point is to make the *unit*
|
|
real — service user, runtime and state directories, socket, nginx
|
|
reachability — so the swarm-level surfaces that follow have somewhere to land.
|
|
Inventing those surfaces before they are agreed would bake in a shape nobody
|
|
chose. See #3066 and the `hyperhive.swarm` consolidation epic.
|
|
|
|
## Why a unix socket, not a port
|
|
|
|
The hive-gateway's nginx is the only intended client and reaches the socket
|
|
through a bind-mount. A listener that is never bound to an address cannot be
|
|
reached from off-host by mistake.
|
|
|
|
The socket path is `services.hyperhive.swarm.controller.socketPath`, default
|
|
`/run/swarm-controller/controller.sock`, exported to the process as
|
|
`SWARM_CONTROLLER_SOCKET`.
|
|
|
|
## ⚠️ The socket's directory is its access control
|
|
|
|
The socket is `0666`. It has to be: nginx runs as a different user and
|
|
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
|
|
sockets, and rests on the same argument — *"the bind source dir is per-agent on
|
|
host so blast radius is unchanged."*
|
|
|
|
What keeps that safe is that the directory holds **one** socket. So:
|
|
|
|
> **Never point `socketPath` at a directory that carries anything else.**
|
|
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
|
|
> Pointing nginx at that directory to reach this socket would put the admin
|
|
> socket within its reach too.
|
|
|
|
This got *less* forgiving when nginx moved onto the host: the gateway used to
|
|
reach a unix upstream through a bind-mount, so the mount list was a second
|
|
bound on what it could touch. There is no mount now — the directory is the
|
|
whole of the access control. A unit test pins the default path so a tidying
|
|
edit fails instead of reviewing cleanly.
|
|
|
|
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
|
|
a **pair**: preserving the directory without the unlink means `bind` fails with
|
|
`EADDRINUSE` after a restart.
|
|
|
|
## Packaging
|
|
|
|
Built by the workspace derivation and extracted as its own package
|
|
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
|
|
`daemonBins` — that list is the core stack and drives the bundle
|
|
`services.hyperhive.c0re.package` points at, so folding this in would put a
|
|
swarm-scoped service into every hive's closure.
|