hyperhive/docs/swarm/ca.md
atlas 433b294099 refactor(nix): swarm.peers becomes swarm.hives, a directory of every hive
One attrset describing every hive in the swarm including this one,
identical on every host, with hiveName selecting which entry is us.
"My peers" is derived (swarm.peerHives) rather than declared.

Every field in the old per-host peer list was intrinsic to the hive it
described, never to the pair -- so the list was a directory each host
kept its own copy of. Beyond the deduplication it removes a bug class:
two hosts could hold different endpoints for the same third hive with
nothing to detect the disagreement.

Drops the per-hive caCert. Trust inside a swarm derives from the swarm
root, which every hive chains to. What that genuinely removes is
trusting a hive whose root this swarm does not own -- a cross-swarm
problem that wants a mechanism of its own, not a field that happened to
work.

The matrix container's certificateFiles block goes with it and could
NOT be migrated: that list is read at build time and the swarm root is
a runtime file (its key must never enter the store), so there is no
build-time name to put there. caCert being a nix path was precisely
what made it the build-time distribution channel. Agents are unaffected
-- hive-tls folds the root into the hive trust bundle and the meta
renderer embeds that one file. Tracked separately.

Migration is an assertion plus warnings, not a rename: hives is peers
union {self}, and the set gains a member no existing config has written
down. A rename migrates a name and a default can re-root a meaning;
neither can conjure a new member. The warning explains, the self-entry
assertion stops the build.
2026-08-05 20:44:16 +02:00

7.1 KiB

Swarm CA

A hive's internal TLS chains to a swarm root CA: the root signs each hive's own CA, and that hive CA signs the gateway leaf. A peer that trusts the root once validates every hive in the swarm, present and future, instead of being pinned to each one by hand.

That is the whole point of the hierarchy — it turns per-peer trust from O(n²) hand-pinning into one anchor per swarm.

Two provisioning modes, one structure

What differs is who puts the artifacts on disk:

swarm root this hive's CA
default operator-provided operator-provided, else self-signed as before
autoConfigure = true (all on one host) generated by swarm-ca.service on first boot issued by hive-tls-ca.service under the root

services.hyperhive.swarm.ca.autoConfigure selects between them, and is off by default: a swarm's services and its hives can live on different hosts, and a host cannot tell whether it is the one holding the root, so setting the swarm CA up is an operator action rather than something a host assumes. Turn it on for an all-on-one-host deployment and the hierarchy costs no configuration.

It defaults from services.hyperhive.enableAllLocalDefaults, the single switch that says "this box is the whole deployment".

A hive given neither artifact keeps the self-signed CA it has always had. It serves TLS exactly as before and simply isn't part of a swarm's trust hierarchy — the right outcome for a hive nobody has federated yet. Only autoConfigure issues a hive sub-CA, because only that case can: signing one needs the root's private key.

Moving the swarm CA onto its own host is then a matter of moving services.hyperhive.swarm.ca.stateDir and leaving autoConfigure off — there is no second code path to switch to.

Constraints on the material

The root's private key never reaches the nix store: the store is world-readable and content-addressed, so a key committed to a flake is a key published to everyone who builds it. Only certificates are distributed.

Each hive CA is name-constrained (X.509 nameConstraints) to that hive's own domain, so a hive CA that leaks can only mint names inside its own subdomain — enforced by every verifier rather than by convention. The constraint excludes both IP families as well, since a permitted-DNS-only constraint says nothing about IP SANs.

The root is issued with pathlen:1: it may sign hive CAs, and those may sign leaves, and the chain stops there.

What to hand a peer

hivectl peer-config prints the cp line. The file is <tls.stateDir>/trust-bundle.pem — the hive CA plus the swarm root — not ca.pem.

The distinction is load-bearing rather than cosmetic: once a hive CA is an intermediate, it is no longer something a verifier can build a chain to. OpenSSL will not terminate a chain at a trusted non-self-signed certificate without -partial_chain, so handing a peer the bare intermediate produces a verification failure that reads like a bad cert rather than like a missing anchor. The bundle carries both, so the same file works whichever mode issued it.

On a hive that predates the swarm root the bundle is just that hive's self-signed CA, so the recipe does not change.

⚠️ Consumers must read trust-bundle.pem, never ca.pem directly. A consumer that reads ca.pem works fine on a hive that has always been self-signed and breaks the moment that hive adopts the hierarchy — so the failure is invisible on the deployment you are most likely to test on.

Adopting the hierarchy on an existing hive

A hive that predates the swarm root carries a self-signed ca.pem, and adopting the hierarchy means replacing it. That invalidates an anchor consumers already trust, and they refresh on their own schedule — agents only pick up new trust when their container restarts, peers only on their own rebuild. Who is allowed to decide that is what splits the two cases.

Where this host owns the root (autoConfigure)

Adoption happens by itself, once. hive-tls-ca.service notices that ca.pem does not chain to the root, keeps the old certificate as ca-previous.pem, and re-issues under the root; the next leaf is signed by the new CA.

It is safe to automate here precisely because this is the all-on-one-host shape: every consumer is on this box, so "when will they have refreshed" is knowable rather than guessed.

The old CA stays in trust-bundle.pem afterwards, so adoption is additive to the anchor set before it is subtractive — a container that has not restarted yet still validates. Removing ca-previous.pem is a deliberate later step: how long is long enough is a property of the deployment, not something the unit can know.

A marker file (.swarm-ca-adopted) records that this ran. Its absence is the trigger, so adoption fires once per hive rather than being re-decided on every activation.

Everywhere else

No automatic adoption. hive-tls-ca.service fails, loudly, naming both certificates and giving the two-command recipe:

rm <tls.stateDir>/ca.pem <tls.stateDir>/ca-key.pem
systemctl restart hive-tls-ca.service

Failing rather than warning is deliberate: a hive whose CA does not chain to the root it has been given is misconfigured, and a warning in a build log is not something anyone reads twice.

To keep the current CA on purpose — a hive that deliberately stays outside the hierarchy, or one mid-migration — touch the marker file named in the message. That is a decision, and it is recorded as one.

A hive with no root configured at all is not affected by any of this: it self-signs exactly as it always has.

Distributing the root

The root key is a runtime file for an obvious reason: the nix store is world-readable and content-addressed, so a key committed to a flake is a key published to every consumer of that flake.

The root certificate is a runtime file as a consequence — it lives beside the key under swarm.ca.stateDir — and that has a cost worth naming, because it is not obvious and it bites at a distance:

Nothing whose trust store is assembled at build time can reference the swarm root. security.pki.certificateFiles is read inside the derivation; the root does not exist there.

Two consumers, and only one of them is fine:

  • Agents are covered. hive-tls.nix folds the root into this hive's trust-bundle.pem, hive-c0re receives that path as HIVE_TLS_CA_PATH, and the meta-flake renderer embeds that one file next to every agent's flake. The bundle is the runtime-to-build-time bridge.
  • The Matrix container is not. It has no equivalent bridge, so it trusts no swarm-internal CA and federation with a self-signed peer does not validate. Giving it the root needs a runtime mechanism — bind-mount plus a bundle assembled at unit start, appending to the system bundle rather than replacing it — which is tracked separately.

To put the root on another host, copy the certificate to the same path there (scp <stateDir>/root.pem <host>:<stateDir>/root.pem). One anchor per host, not one per peer: a hive joining later needs no edit on the hives already running, which is the whole point of the hierarchy.