Review pass on top of the previous commit (85a94671) turned up three
real issues its own vale-only verification didn't catch:
- 'has not' -> 'hasn't' fired inside 'has nothing', corrupting
docs/swarm/README.md into 'hasn'thing to say upward'. Restored
'has nothing'.
- ~44 hits where the original text was genuinely sentence-initial
('It is', 'That is', 'They are', 'What is', 'Do not', ...) got
contracted in lowercase, silently de-capitalizing the start of a
sentence. Checked each one's real surrounding text and capitalized
the ones that needed it.
- Confirmed (by reading the actual text, not just trusting vale) that
the 3 remaining 'is not'/'has not'/'it is' hits after those fixes are
vale's own false positives: it matches these as a bare substring with
no trailing word boundary, so it also fires on 'is nothing', 'has
nothing', and the already-correct 'it isn't'. Left those alone --
there's no real issue in the text.
Re-verified with a fresh vale run: the remaining 17 hits are exactly
the union of the 14 legitimate ALL-CAPS/bold-emphasis skips from the
first commit and these 3 false positives. Full accounting in
/agents/iris/state/haiku-batches/contractions-report.md.
9.2 KiB
Swarm CA
A hive's internal TLS chains to a swarm root CA: the root signs each hive's own CA, and that hive CA signs the gateway leaf. A peer that trusts the root once validates every hive in the swarm, present and future, instead of being pinned to each one by hand.
That's the whole point of the hierarchy — it turns per-peer trust from O(n²) hand-pinning into one anchor per swarm.
Two provisioning modes, one structure
What differs is who puts the artifacts on disk:
| swarm root | this hive's CA | |
|---|---|---|
| default | operator-provided | operator-provided, else self-signed as before |
autoConfigure = true (all on one host) |
generated by swarm-ca.service on first boot |
issued by hive-tls-ca.service under the root |
services.hyperhive.swarm.ca.autoConfigure selects between them, and is
off by default: a swarm's services and its hives can live on
different hosts, and a host can't tell whether it's the one holding
the root, so setting the swarm CA up is an operator action rather than
something a host assumes. Turn it on for an all-on-one-host deployment
and the hierarchy costs no configuration.
It defaults from services.hyperhive.deploy.singleHostSwarm, the single
switch that says "this box is the whole deployment".
A hive given neither artifact keeps the self-signed CA it has always
had. It serves TLS exactly as before and isn't part of a
swarm's trust hierarchy — the right outcome for a hive nobody has
federated yet. Only autoConfigure issues a hive sub-CA, because only
that case can: signing one needs the root's private key.
Moving the swarm CA onto its own host is then a matter of moving
services.hyperhive.swarm.ca.stateDir and leaving autoConfigure off —
there is no second code path to switch to.
The warning about uncovered service names
The swarm's service names (swarm.forge.domain,
swarm.matrix.gatewayHost, swarm.authelia.domain) default to
siblings of the hive domain — forge.<swarm.domain>, not
forge.<hive domain>. The hive CA's leaf is a single-label wildcard
over the hive's own domain, so it can't cover them; only the
swarm-services leaf can, and this host signs that one only under
autoConfigure. A hive with neither serves the hive leaf on those names
and every client sees a name mismatch.
So hive-tls emits an eval-time warning naming the uncovered names.
It's deliberately not an assertion: this module can see what it is
able to issue, but not an operator-installed sub-CA in
swarm.ca.stateDir, an external ACME setup, or a certificate delivered
by any other means. If you have arranged one, the warning is expected
and can be ignored. Otherwise either install the sub-CA, or pin the
names back under the hive domain — a supported migration, since the
sub-CA is constrained to the configured names and the swarm root
carries no name constraints at all.
Constraints on the material
The root's private key never reaches the nix store: the store is world-readable and content-addressed, so a key committed to a flake is a key published to everyone who builds it. Only certificates are distributed.
Each hive CA is name-constrained (X.509 nameConstraints) to that
hive's own domain, so a hive CA that leaks can only mint names inside
its own subdomain — enforced by every verifier rather than by
convention. The constraint excludes both IP families as well, since a
permitted-DNS-only constraint says nothing about IP SANs.
The root is issued with pathlen:1: it may sign hive CAs, and those may
sign leaves, and the chain stops there.
What to hand a peer
hivectl peer-config prints the services.hyperhive.swarm.hives."<name>"
block a peer operator pastes into their own config. When this hive's
gateway serves a self-signed leaf under the hierarchy (detected by the
presence of <deploy.hive-controller.tls.stateDir>/trust-bundle.pem), it
also prints a one-time
scp line installing the swarm root —
<swarm.ca.stateDir>/root.pem, not this hive's own CA — on the peer's
host:
scp /var/lib/swarm-ca/root.pem <peer-host>:/var/lib/swarm-ca/root.pem
That's the point of the hierarchy: the root is installed once per
swarm host, not once per peer, so a hive joining later needs no edit on
the hives already running. A hive whose cert already chains to a public
CA has nothing to install — peer-config says so instead.
Handing a peer this hive's own ca.pem would not work even as a
one-off: once a hive CA is an intermediate under the swarm root, it's
no longer something a verifier can build a chain to — OpenSSL won't
terminate a chain at a trusted non-self-signed certificate without
-partial_chain. That's why the tool distributes the root, not a
per-hive file.
Adopting the hierarchy on an existing hive
A hive that predates the swarm root carries a self-signed ca.pem, and
adopting the hierarchy means replacing it. That invalidates an anchor
consumers already trust, and they refresh on their own schedule — agents
only pick up new trust when their container restarts, peers only on
their own rebuild. Who is allowed to decide that's what splits the
two cases.
Where this host owns the root (autoConfigure)
Adoption happens by itself, once. hive-tls-ca.service notices that
ca.pem doesn't chain to the root, keeps the old certificate as
ca-previous.pem, and re-issues under the root; the next leaf is signed
by the new CA.
It's safe to automate here precisely because this is the all-on-one-host shape: every consumer is on this box, so "when will they have refreshed" is knowable rather than guessed.
The old CA stays in trust-bundle.pem afterwards, so adoption is
additive to the anchor set before it's subtractive — a container
that hasn't restarted yet still validates. Removing ca-previous.pem
is a deliberate later step: how long is long enough is a property of the
deployment, not something the unit can know.
A marker file (.swarm-ca-adopted) records that this ran. Its absence
is the trigger, so adoption fires once per hive rather than being
re-decided on every activation.
Everywhere else
No automatic adoption. hive-tls-ca.service fails, loudly, naming both
certificates and giving the two-command recipe:
rm <deploy.hive-controller.tls.stateDir>/ca.pem <deploy.hive-controller.tls.stateDir>/ca-key.pem
systemctl restart hive-tls-ca.service
Failing rather than warning is deliberate: a hive whose CA doesn't chain to the root it has been given is misconfigured, and a warning in a build log isn't something anyone reads twice.
To keep the current CA on purpose — a hive that deliberately stays
outside the hierarchy, or one mid-migration — touch the marker file
named in the message. That's a decision, and it's recorded as one.
A hive with no root configured at all isn't affected by any of this: it self-signs exactly as it always has.
Distributing the root
The root key is a runtime file for the reason in Constraints on the material.
The root certificate is a runtime file as a consequence — it lives
beside the key under swarm.ca.stateDir — and that has a cost worth
naming, because it's subtle and it bites at a distance:
Nothing whose trust store is assembled at build time can reference the swarm root.
security.pki.certificateFilesis read inside the derivation; the root doesn't exist there.
Two consumers, and only one of them is fine:
-
Agents are covered.
hive-tls.nixfolds the root into this hive'strust-bundle.pem, hive-c0re receives that path asHIVE_TLS_CA_PATH, and the meta-flake renderer embeds that one file next to every agent's flake. The bundle is the runtime-to-build-time bridge. -
The Matrix container crosses the same bridge, via the shared
lib/hive-ca-trust.nixhelper thathive-ciandhive-forgealready use: the bundle is bind-mounted read-only into the container, and a oneshot concatenates it with the system CAs before tuwunel starts.The consumption differs per runtime and is the part worth knowing. tuwunel links no openssl, which makes
SSL_CERT_FILElook inapplicable — it isn't: its rustls-based TLS stack still resolves trust through the platform certificate store on Linux, and that store honorsSSL_CERT_FILE, so the env var takes effect the same way it would for an OpenSSL-linked binary.⚠️ Concatenate; never point
SSL_CERT_FILEat the anchor alone.openssl-probeuses it instead of the default store, so naming just the bundle would drop every public CA and break federation with the wider matrix network — a much bigger outage than the one being fixed. The same caveat applies tohive-forge(Go) for the same reason.
To put the root on another host, copy the certificate to the same path
there (scp <stateDir>/root.pem <host>:<stateDir>/root.pem). One anchor
per host, not one per peer: a hive joining later needs no edit on the
hives already running, which is the whole point of the hierarchy.