hyperhive/swarm-controller/README.md
atlas 3202cde704 nix: gate hive-c0re on deploy.hive-controller.enable, drop hyperhive.enable
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.

Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:

- Dropped: each swarm service and its glue keeps only its own deploy
  toggle (authelia, bao and its PKI glue, grafana, victorialogs,
  victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
  the controller/nats/matrix-ctl/publisher/services-issuer identities),
  the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
  its assertion, which feed hive-c0re and write under its state dir, plus
  their policy-order entry; the network identity assertions; hive-tls's
  two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
  (`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
  `deploy.matrix.enable` plus their client-identity checks. They read
  deploy.matrix's token file and registration script; their deploy.bao
  inputs are the client-half options a hive sets to read a store it does
  not run, so gating on deploy.bao.enable would drop the tested
  remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
  and fires wherever the hive, the store or the homeserver runs: each
  turns the name into an identifier with no fallback.

On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.

Refs #4500
2026-09-26 01:19:49 +02:00

60 lines
2.6 KiB
Markdown

# swarm-controller
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
owns what is true _across_ hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via `services.hyperhive.deploy.swarm-controller.enable`, which is
deliberately **not** derived from
`services.hyperhive.deploy.hive-controller.enable`: turning it on is a statement
about swarm topology, not about whether this host runs a hive.
## What it does today
Serves one `/health` endpoint and holds no state.
That is the whole intent of the first slice. The point is to make the _unit_
real — service user, runtime and state directories, socket, nginx
reachability — so the swarm-level surfaces that follow have somewhere to land.
Inventing those surfaces before they are agreed would bake in a shape nobody
chose. See the `hyperhive.swarm` consolidation epic.
## Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket
through a bind-mount. A listener that is never bound to an address cannot be
reached from off-host by mistake.
The socket path is `services.hyperhive.deploy.swarm-controller.socketPath`, default
`/run/swarm-controller/controller.sock`, exported to the process as
`SWARM_CONTROLLER_SOCKET`.
## ⚠️ The socket's directory is its access control
The socket is `0666`. It has to be: nginx runs as a different user and
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
sockets, and rests on the same argument — _"the bind source dir is per-agent on
host so blast radius is unchanged."_
What keeps that safe is that the directory holds **one** socket. So:
> **Never point `socketPath` at a directory that carries anything else.**
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
> Pointing nginx at that directory to reach this socket would put the admin
> socket within its reach too.
nginx is a host service, so nothing narrows what it can reach except the
directory itself — that is the whole of the access control. A unit test pins
the default path so a tidying edit fails instead of reviewing cleanly.
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
a **pair**: preserving the directory without the unlink means `bind` fails with
`EADDRINUSE` after a restart.
## Packaging
Built by the workspace derivation and extracted as its own package
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
`daemonBins` — that list is the core stack and drives the bundle
`services.hyperhive.c0re.package` points at, so folding this in would put a
swarm-scoped service into every hive's closure.