| Filename | Latest commit message | Latest commit date |
|---|---|---|
The read policy and cert-auth role for each hive were written once, at startup. On the deploy that surfaced this, the store was still coming up, the pass logged its warning and moved on, and no hive could log in until someone restarted the daemon — while cert auth answered "no chain matching all constraints", which reads like a certificate problem rather than a role that was never created. The bootstrap unit in swarm-bao.nix lost the same race and won on its retry 30s later. A daemon that boots alongside its store loses that race routinely; on a normal boot it is the ordinary case. The two passes fold into one `provision()` that logs in once instead of twice for two loops over the same list, keeping policy before role since the role names the policy. `ensure_hive_access` still awaits the first pass, so a store that is already up leaves nothing deferred, and only a pass that could not reach the store at all spawns the retry. The retry is `config_pr::spawn`'s idiom from this same crate: an interval task whose first tick is immediate. Its cadence and bound match the bootstrap unit's — 30s, ~a day — because the two halves of one race should not disagree about how long a wait is worth. `Error::MissingEnv` is what keeps it from spinning forever: no `BAO_*` set means a deployment that runs no store, where asking again changes nothing, so it returns Ok. Everything else is retryable, including an authority file that is not placed yet — the unit that writes it starts alongside this one. Both cases previously landed in the same "not managed here" line, so a store that was late looked exactly like one that was never configured. Per-hive failures keep their old behaviour: logged, skipped, Ok. A store that refuses one hive's write refuses it again, so the next start really is the right retry for those, and the module doc still says so. Closes #4176. |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-controller
The swarm-level daemon. Where hive-c0re owns the agents on one host, this
owns what is true across hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is
deliberately not derived from services.hyperhive.enable: turning it on is
a statement about swarm topology, not about whether hyperhive is installed.
What it does today
Serves one /health endpoint and holds no state.
That is the whole intent of the first slice. The point is to make the unit
real — service user, runtime and state directories, socket, nginx
reachability — so the swarm-level surfaces that follow have somewhere to land.
Inventing those surfaces before they are agreed would bake in a shape nobody
chose. See the hyperhive.swarm consolidation epic.
Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.
The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default
/run/swarm-controller/controller.sock, exported to the process as
SWARM_CONTROLLER_SOCKET.
⚠️ The socket's directory is its access control
The socket is 0666. It has to be: nginx runs as a different user and
connect(2) needs write. This matches how hive-c0re publishes the per-agent
sockets, and rests on the same argument — "the bind source dir is per-agent on
host so blast radius is unchanged."
What keeps that safe is that the directory holds one socket. So:
Never point
socketPathat a directory that carries anything else./run/hyperhiveabove all — it holdshost.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.
nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.
RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are
a pair: preserving the directory without the unlink means bind fails with
EADDRINUSE after a restart.
Packaging
Built by the workspace derivation and extracted as its own package
(nix build .#swarm-controller). Deliberately not in nix/packages'
daemonBins — that list is the core stack and drives the bundle
services.hyperhive.c0re.package points at, so folding this in would put a
swarm-scoped service into every hive's closure.