Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/swarm-controller
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 6170e74a31 swarm-bao: agent certificates issued by a store-generated agent CA
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.

- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
  tunes the mount, generates the root once (guarded on an empty issuer
  list, no replace branch), upserts the `swarm-agent` role (client
  certificates named `hive-agent-*` only, 90 days), caches the CA at
  /var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
  (client-ca.pem, then the agent CA). Host cert-auth roles still pin
  client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
  composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
  changed after openbao started, swarm-bao-agent-pki restarts
  openbao.service in the container; under `seal = "shamir"` it prints
  the step instead. Once swarm-bao-certs has a cached CA, later boots
  start openbao with it and do not restart.
- The controller policy gains exactly `update` on
  pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
  the leaf (the store generates the key), writes the agent's cert-auth
  role pinning the issuing CA bao returned, and writes the agent's
  policy as render_agent alone: the hive-shared queue credential stanza
  is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
  other pki role names); swarm-controller gets
  SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.

Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.

Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.

Closes #4756
2026-09-27 22:59:27 +02:00
..
src swarm-bao: agent certificates issued by a store-generated agent CA 2026-09-27 22:59:27 +02:00
Cargo.toml swarm-bao: agent certificates issued by a store-generated agent CA 2026-09-27 22:59:27 +02:00
README.md nix: gate hive-c0re on deploy.hive-controller.enable, drop hyperhive.enable 2026-09-26 01:19:49 +02:00

swarm-controller

The swarm-level daemon. Where hive-c0re owns the agents on one host, this owns what is true across hives — so a swarm runs one of them and most hives leave it off.

Opt-in per host via services.hyperhive.deploy.swarm-controller.enable, which is deliberately not derived from services.hyperhive.deploy.hive-controller.enable: turning it on is a statement about swarm topology, not about whether this host runs a hive.

What it does today

Serves one /health endpoint and holds no state.

That is the whole intent of the first slice. The point is to make the unit real — service user, runtime and state directories, socket, nginx reachability — so the swarm-level surfaces that follow have somewhere to land. Inventing those surfaces before they are agreed would bake in a shape nobody chose. See the hyperhive.swarm consolidation epic.

Why a unix socket, not a port

The hive-gateway's nginx is the only intended client and reaches the socket through a bind-mount. A listener that is never bound to an address cannot be reached from off-host by mistake.

The socket path is services.hyperhive.deploy.swarm-controller.socketPath, default /run/swarm-controller/controller.sock, exported to the process as SWARM_CONTROLLER_SOCKET.

⚠️ The socket's directory is its access control

The socket is 0666. It has to be: nginx runs as a different user and connect(2) needs write. This matches how hive-c0re publishes the per-agent sockets, and rests on the same argument — "the bind source dir is per-agent on host so blast radius is unchanged."

What keeps that safe is that the directory holds one socket. So:

Never point socketPath at a directory that carries anything else. /run/hyperhive above all — it holds host.sock, the host admin socket. Pointing nginx at that directory to reach this socket would put the admin socket within its reach too.

nginx is a host service, so nothing narrows what it can reach except the directory itself — that is the whole of the access control. A unit test pins the default path so a tidying edit fails instead of reviewing cleanly.

RuntimeDirectoryPreserve=yes and the daemon's stale-socket unlink on start are a pair: preserving the directory without the unlink means bind fails with EADDRINUSE after a restart.

Packaging

Built by the workspace derivation and extracted as its own package (nix build .#swarm-controller). Deliberately not in nix/packages' daemonBins — that list is the core stack and drives the bundle services.hyperhive.c0re.package points at, so folding this in would put a swarm-scoped service into every hive's closure.