| Filename | Latest commit message | Latest commit date |
|---|---|---|
`swarm/agents/<agent>/bao-mtls` did not exist, and neither did any per-agent identity at the secret store: `policy::agent_object_name`, `render_agent` and `render_agent_with_queue` had been written and never called outside their own tests. An agent's only "per-agent" secret today is read under the HIVE's certificate, through a wide grant on `swarm/agents/*` — so "per-agent" was presentational. The swarm now mints the certificate, so no hive ever needs the capability to mint one. `swarm-controller` is the service that does it: it already logs in to the store, and its existing grant already covers exactly the three objects written here (`create/update` on `secret/data/swarm/agents/*`, `sys/policies/acl/hive-*` and `auth/cert/certs/hive-*`). No new bao grant, and nothing co-located — a cert-auth role pins its authority by value, per role, so the controller issues from its own CA on its own host and pins that CA in the role it writes. No existing role changes. The mint node does not report success on a write. After publishing it connects again, with the leaf it just issued and under the role it just wrote, and reads the path back — so the policy, the role, the common name and the leaf are exercised in production on every agent creation. A certificate this code mints that the role this code writes will not accept turns the job node red at creation time instead of surfacing later as an agent container that cannot start. `TriggerDeploy` gains an `after_any` edge on the mint, not `after_ok`: a hive cannot pass down a certificate the swarm has not published, but a host with no authority configured must still create agents exactly as it does today. The private key is generated in memory and never written to disk on the controller — `SecretStore::connect_with_identity` takes the PEM the minter is already holding, so nothing is written out purely to be logged in with. Refs #4137 |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-controller
The swarm-level daemon. Where hive-c0re owns the agents on one host, this
owns what is true across hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is
deliberately not derived from services.hyperhive.enable: turning it on is
a statement about swarm topology, not about whether hyperhive is installed.
What it does today
Serves one /health endpoint and holds no state.
That is the whole intent of the first slice. The point is to make the unit
real — service user, runtime and state directories, socket, nginx
reachability — so the swarm-level surfaces that follow have somewhere to land.
Inventing those surfaces before they are agreed would bake in a shape nobody
chose. See the hyperhive.swarm consolidation epic.
Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.
The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default
/run/swarm-controller/controller.sock, exported to the process as
SWARM_CONTROLLER_SOCKET.
⚠️ The socket's directory is its access control
The socket is 0666. It has to be: nginx runs as a different user and
connect(2) needs write. This matches how hive-c0re publishes the per-agent
sockets, and rests on the same argument — "the bind source dir is per-agent on
host so blast radius is unchanged."
What keeps that safe is that the directory holds one socket. So:
Never point
socketPathat a directory that carries anything else./run/hyperhiveabove all — it holdshost.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.
nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.
RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are
a pair: preserving the directory without the unlink means bind fails with
EADDRINUSE after a restart.
Packaging
Built by the workspace derivation and extracted as its own package
(nix build .#swarm-controller). Deliberately not in nix/packages'
daemonBins — that list is the core stack and drives the bundle
services.hyperhive.c0re.package points at, so folding this in would put a
swarm-scoped service into every hive's closure.