fix(nix): swarm-controller is not a core binary, and it needed a README
Per review: daemonBins is the core stack, and it drives the bundle that services.hyperhive.c0re.package points at -- so listing a swarm-scoped service there would put it in every hive's closure when one hive in a swarm runs it. It gets the same per-bin extractor, bound on its own, the way hivectl already is. The crate also had no README while every other one does. Both misses are the same shape: adding a thing without updating what describes the set of things.
This commit is contained in:
parent
0fe2babbee
commit
fb51006717
3 changed files with 71 additions and 1 deletions
|
|
@ -40,7 +40,6 @@ let
|
|||
hive-forge = "hyperhive Forgejo CLI";
|
||||
hive-forge-notify = "hyperhive per-agent Forgejo notification poller daemon";
|
||||
hive-github-notify = "hyperhive per-agent github.com notification poller daemon";
|
||||
swarm-controller = "hyperhive swarm-level controller daemon";
|
||||
};
|
||||
|
||||
# ONE compile of the whole workspace (every bin, sharing the
|
||||
|
|
@ -134,6 +133,16 @@ in
|
|||
}
|
||||
// binPkgs
|
||||
// {
|
||||
# Swarm-level controller daemon. Deliberately NOT in `daemonBins`:
|
||||
# that list is the core stack — the binaries hive-c0re and the agent
|
||||
# harness are made of — and it drives the `default` bundle that
|
||||
# `services.hyperhive.c0re.package` points at. This daemon is a
|
||||
# separate swarm-scoped service with its own module and its own
|
||||
# `package` option, and one hive in a swarm runs it, so folding it
|
||||
# into the core bundle would put it in every hive's closure to no end.
|
||||
# Uses the same per-bin extractor, just bound on its own.
|
||||
swarm-controller = mkBinPackage "swarm-controller" "hyperhive swarm-level controller daemon";
|
||||
|
||||
# Bundled browser assets — see ./frontend.nix. Output is
|
||||
# $out/{dashboard,agent}/ which the Rust binaries serve via
|
||||
# tower_http::ServeDir.
|
||||
|
|
|
|||
|
|
@ -1,6 +1,7 @@
|
|||
[package]
|
||||
name = "swarm-controller"
|
||||
version.workspace = true
|
||||
readme = "README.md"
|
||||
edition.workspace = true
|
||||
|
||||
[[bin]]
|
||||
|
|
|
|||
60
swarm-controller/README.md
Normal file
60
swarm-controller/README.md
Normal file
|
|
@ -0,0 +1,60 @@
|
|||
# swarm-controller
|
||||
|
||||
The **swarm-level** daemon. Where `hive-c0re` owns the agents on one host, this
|
||||
owns what is true *across* hives — so a swarm runs one of them and most hives
|
||||
leave it off.
|
||||
|
||||
Opt-in per host via `services.hyperhive.swarm.controller.enable`, which is
|
||||
deliberately **not** derived from `services.hyperhive.enable`: turning it on is
|
||||
a statement about swarm topology, not about whether hyperhive is installed.
|
||||
|
||||
## What it does today
|
||||
|
||||
Serves one `/health` endpoint and holds no state.
|
||||
|
||||
That is the whole intent of the first slice. The point is to make the *unit*
|
||||
real — service user, runtime and state directories, socket, nginx
|
||||
reachability — so the swarm-level surfaces that follow have somewhere to land.
|
||||
Inventing those surfaces before they are agreed would bake in a shape nobody
|
||||
chose. See #3066 and the `hyperhive.swarm` consolidation epic.
|
||||
|
||||
## Why a unix socket, not a port
|
||||
|
||||
The hive-gateway's nginx is the only intended client and reaches the socket
|
||||
through a bind-mount. A listener that is never bound to an address cannot be
|
||||
reached from off-host by mistake.
|
||||
|
||||
The socket path is `services.hyperhive.swarm.controller.socketPath`, default
|
||||
`/run/swarm-controller/controller.sock`, exported to the process as
|
||||
`SWARM_CONTROLLER_SOCKET`.
|
||||
|
||||
## ⚠️ The socket's directory is its access control
|
||||
|
||||
The socket is `0666`. It has to be: nginx runs as a different user and
|
||||
`connect(2)` needs write. This matches how `hive-c0re` publishes the per-agent
|
||||
sockets, and rests on the same argument — *"the bind source dir is per-agent on
|
||||
host so blast radius is unchanged."*
|
||||
|
||||
What keeps that safe is that the directory holds **one** socket and is
|
||||
bind-mounted into **one** container. So:
|
||||
|
||||
> **Never point `socketPath` at a directory that carries anything else.**
|
||||
> `/run/hyperhive` above all — it holds `host.sock`, the host **admin** socket.
|
||||
> Mounting that directory to reach this socket would hand the gateway container
|
||||
> the admin socket along with it.
|
||||
|
||||
Changing `socketPath` therefore means re-checking the gateway bind-mount, not
|
||||
just the daemon. A unit test pins the default path so a tidying edit fails
|
||||
instead of reviewing cleanly.
|
||||
|
||||
`RuntimeDirectoryPreserve=yes` and the daemon's stale-socket unlink on start are
|
||||
a **pair**: preserving the directory without the unlink means `bind` fails with
|
||||
`EADDRINUSE` after a restart.
|
||||
|
||||
## Packaging
|
||||
|
||||
Built by the workspace derivation and extracted as its own package
|
||||
(`nix build .#swarm-controller`). Deliberately **not** in `nix/packages`'
|
||||
`daemonBins` — that list is the core stack and drives the bundle
|
||||
`services.hyperhive.c0re.package` points at, so folding this in would put a
|
||||
swarm-scoped service into every hive's closure.
|
||||
Loading…
Reference in a new issue