| Filename | Latest commit message | Latest commit date |
|---|---|---|
The swarm UI had nowhere to POST an external matrix account to: this daemon had no matrix-account code at all and no `swarm-secret-client` dependency, so the last leg of #3726 — a credential reaching an agent — had no entry point. `PUT /api/hives/{hive}/agents/{agent}/matrix-accounts/{account}` writes the credential to the store under the agent's own path and publishes a `CredentialNotice` on that hive's credential subject. All three path names are load-bearing: agent + account locate the secret, hive routes the notice. The account is a path segment rather than a body field so that splitting the 1:1 account-to-agent mapping later is a new route, not a changed payload. Store first, notify second, and the order cannot be swapped: a notice that overtakes its own write reaches a hive that reads nothing, and the hive deliberately does not retry. The publish is followed by a flush for the reason `publish_deploy` flushes — `publish` hands the message to the connection's write buffer and returns, so the response could otherwise outrun the notice it reports as sent. The store client is built per request rather than held in `AppState`, matching what the hive side does inside `deliver`: a login that expires is not worth caching for a route this cold. `swarm_hive` is `declaration_target`'s two name checks, extracted so this handler makes them identically rather than in a second copy free to drift. `declaration_target` still tests the writer first, so a deployment with no queue answers 503 whatever the caller spelled. ## The nix half #4081 minted the controller's leaf and gave it `baoClientCertFile` / `baoClientKeyFile`, deliberately stopping there — the leaf is minted whether or not a controller runs on that host. Nothing consumed those options, so the identity never reached the process. Measured before writing: `git grep baoClientCertFile` returned 5 sites and zero consumers, against a control (`tokenEndpoint`, 4 hits in the same file) proving the search can see consumption where it exists. The unit now gets `BAO_ADDR` / `BAO_CLIENT_CERT` / `BAO_CLIENT_KEY` / `BAO_CACERT` and the matching `LoadCredential` entries, following `hive-c0re/environment.nix`'s `%d` credential shape. The gate is `deploy.swarm-controller.baoClientCertFile`, NOT `deploy.bao.clientCertFile`. The latter is the hive reader's identity and its policy scopes a hive's own secrets; wiring it here would evaluate, deploy, and fail only when the daemon tried to write an agent's credential. Two `module-eval` arms cover exactly that. The presence arm asserts the `LoadCredential` *source path* (`…:/var/lib/swarm-bao-pki/controller.pem`) and not just the `%d` name, because a `%d`-only assertion passes while the daemon holds the wrong policy. The absence arm (`controllerNoStore`) is what makes the presence arm mean anything. `RestrictAddressFamilies` already covers the store client; its own comment asks for the family to be added with the client, and AF_INET/AF_INET6 are present. Contributes to #3726 |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-controller
The swarm-level daemon. Where hive-c0re owns the agents on one host, this
owns what is true across hives — so a swarm runs one of them and most hives
leave it off.
Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is
deliberately not derived from services.hyperhive.enable: turning it on is
a statement about swarm topology, not about whether hyperhive is installed.
What it does today
Serves one /health endpoint and holds no state.
That is the whole intent of the first slice. The point is to make the unit
real — service user, runtime and state directories, socket, nginx
reachability — so the swarm-level surfaces that follow have somewhere to land.
Inventing those surfaces before they are agreed would bake in a shape nobody
chose. See #3066 and the hyperhive.swarm consolidation epic.
Why a unix socket, not a port
The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.
The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default
/run/swarm-controller/controller.sock, exported to the process as
SWARM_CONTROLLER_SOCKET.
⚠️ The socket's directory is its access control
The socket is 0666. It has to be: nginx runs as a different user and
connect(2) needs write. This matches how hive-c0re publishes the per-agent
sockets, and rests on the same argument — "the bind source dir is per-agent on
host so blast radius is unchanged."
What keeps that safe is that the directory holds one socket. So:
Never point
socketPathat a directory that carries anything else./run/hyperhiveabove all — it holdshost.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.
nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.
RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are
a pair: preserving the directory without the unlink means bind fails with
EADDRINUSE after a restart.
Packaging
Built by the workspace derivation and extracted as its own package
(nix build .#swarm-controller). Deliberately not in nix/packages'
daemonBins — that list is the core stack and drives the bundle
services.hyperhive.c0re.package points at, so folding this in would put a
swarm-scoped service into every hive's closure.