A swarm's services and its hives can live on different hosts, and a host
has no way to tell whether it is the one holding the root — so setting
the swarm CA up is an operator action, not something a host infers.
`swarm.ca.autoConfigure` is off by default and is the only thing that
issues a hive sub-CA, which is also the only case that *can*: signing
one needs the root's private key, and that key is deliberately not on a
hive host otherwise.
With the flag off, a hive self-signs its CA exactly as it always has.
That keeps a plain hive working out of the box; what it gives up is
membership of a swarm's trust hierarchy, which is the right thing to
give up for a hive nobody has federated.
Replaces a default derived from `swarm.peers == { }`. That read "no
peers declared" as "everything is local", which is not the same claim —
a hive can belong to a swarm it has not declared yet, or to one whose
services run elsewhere entirely.
163 lines
6.7 KiB
Nix
163 lines
6.7 KiB
Nix
# The swarm root CA: the anchor a whole swarm shares, and the issuer of
|
|
# each hive's own CA (which is where it gets used — see ./hive-tls.nix).
|
|
#
|
|
# Why a hierarchy at all: cross-hive trust is hand-pinned today
|
|
# (`swarm.peers.<d>.caCert`), so every hive must name every peer — O(n²)
|
|
# configuration that a new hive can only join by editing all the others.
|
|
# One root makes it O(1): trust the root once and every present *and
|
|
# future* peer validates.
|
|
#
|
|
# Two provisioning modes, ONE structure — what differs is who puts the
|
|
# artifacts on disk, never what the artifacts are:
|
|
#
|
|
# - autoconfigured (the default while the hive declares no peers): the
|
|
# unit below generates the root here on first boot.
|
|
# - operator-provided (multi-host): the operator installs the root cert
|
|
# — and, on a host that does not hold the root key, the hive CA too —
|
|
# into the state dirs, and this unit does nothing. Splitting a
|
|
# single-host swarm across hosts is then a matter of moving the
|
|
# service dirs and turning `autoConfigure` off, not of a second code
|
|
# path that has to be kept agreeing with the first.
|
|
#
|
|
# The root KEY is the reason this is a runtime file and not a nix option:
|
|
# the store is world-readable and content-addressed, so a key committed
|
|
# to a flake is a key *published* to every consumer of that flake. The
|
|
# root CERT has the opposite property, and is already distributed at
|
|
# build time via `swarm.peers.<d>.caCert`. That asymmetry is what makes
|
|
# the layering work at all.
|
|
{
|
|
lib,
|
|
config,
|
|
pkgs,
|
|
...
|
|
}:
|
|
let
|
|
cfg = config.services.hyperhive.swarm.ca;
|
|
hyperhiveCfg = config.services.hyperhive;
|
|
|
|
# The subject CN is a label for a human reading a chain, not an
|
|
# identity anything authenticates against. Fall through swarm name →
|
|
# hive domain → a constant so a hive that has set neither still
|
|
# evaluates; a missing `domain` is reported by its own assertion in
|
|
# hive-network.nix, and shouldn't also surface here as a null.
|
|
swarmLabel =
|
|
if hyperhiveCfg.swarm.name != null then
|
|
hyperhiveCfg.swarm.name
|
|
else if hyperhiveCfg.domain != null then
|
|
hyperhiveCfg.domain
|
|
else
|
|
"hyperhive";
|
|
in
|
|
{
|
|
options.services.hyperhive.swarm.ca = {
|
|
autoConfigure = lib.mkOption {
|
|
type = lib.types.bool;
|
|
default = false;
|
|
example = true;
|
|
description = ''
|
|
Run the whole swarm CA on this one host: generate the swarm
|
|
root when it is missing, and issue this hive's CA under it.
|
|
|
|
**Off by default, deliberately.** A swarm's services and its
|
|
hives can live on different hosts, and this host has no way to
|
|
tell whether it is the one holding the root — so the swarm CA
|
|
is something an operator sets up, not something a host decides
|
|
it is. Turn this on for an all-on-one-host deployment (dev
|
|
boxes, single-hive swarms) and get the hierarchy for free.
|
|
|
|
With it off, both artifacts are operator-provided: the root
|
|
under `stateDir`, and this hive's CA under
|
|
`services.hyperhive.tls.stateDir`. A hive given neither keeps
|
|
the self-signed CA it has always had — it simply isn't part of
|
|
a swarm's trust hierarchy, which is the correct outcome for a
|
|
hive nobody has federated yet.
|
|
'';
|
|
};
|
|
|
|
stateDir = lib.mkOption {
|
|
type = lib.types.str;
|
|
default = "/var/lib/swarm-ca";
|
|
description = ''
|
|
Host directory holding the swarm root CA: `root.pem` (the
|
|
anchor, safe to distribute — this is what a peer hive is
|
|
pointed at via `swarm.peers.<d>.caCert`) and `root-key.pem`
|
|
(0600, the one file that must never reach the nix store or
|
|
another host). The directory itself is 0700: nothing reads
|
|
out of it but the hive CA issuance in `hive-tls.nix`.
|
|
|
|
Moving the swarm CA to its own host is a matter of moving this
|
|
directory and setting `autoConfigure = false` here.
|
|
'';
|
|
};
|
|
|
|
validityDays = lib.mkOption {
|
|
type = lib.types.int;
|
|
default = 10950;
|
|
description = ''
|
|
Validity window of the swarm root CA in days (default ~30y).
|
|
Deliberately longer than `services.hyperhive.tls.caValidityDays`:
|
|
the root must outlive the hive CAs it issues, or those chains
|
|
expire out from under hives that are still perfectly happy with
|
|
their own intermediate. Rotating a root is the one operation in
|
|
this system with no partial-failure mode — it invalidates every
|
|
peer at once, paced by the slowest peer's rebuild — so it is
|
|
never automatic and this window is meant to be uneventful.
|
|
'';
|
|
};
|
|
};
|
|
|
|
config = lib.mkIf (hyperhiveCfg.enable && cfg.autoConfigure) {
|
|
systemd.services.swarm-ca = {
|
|
description = "Generate the swarm root CA when absent";
|
|
wantedBy = [ "multi-user.target" ];
|
|
path = [ pkgs.openssl ];
|
|
serviceConfig = {
|
|
Type = "oneshot";
|
|
RemainAfterExit = true;
|
|
UMask = "0077";
|
|
# Pin the journal identity (else it's the `script` store-path wrapper).
|
|
SyslogIdentifier = "swarm-ca";
|
|
};
|
|
script = ''
|
|
set -euo pipefail
|
|
d=${lib.escapeShellArg cfg.stateDir}
|
|
install -d -m 0700 "$d"
|
|
|
|
root="$d/root.pem"
|
|
rootk="$d/root-key.pem"
|
|
|
|
# Note the asymmetry with the hive CA in hive-tls.nix, which
|
|
# regenerates itself once expired: a root is never replaced
|
|
# automatically, not even an expired one. Consumers hold this
|
|
# cert, so replacing it is a swarm-wide flag day that wants an
|
|
# operator running it deliberately, with both roots trusted
|
|
# across the overlap.
|
|
if [ -s "$root" ] && [ -s "$rootk" ]; then
|
|
echo "swarm root CA already present at $root — leaving it alone"
|
|
exit 0
|
|
fi
|
|
|
|
# Half a root is not a root. Generating a fresh key beside an
|
|
# already-distributed cert (or the reverse) leaves every
|
|
# consumer trusting an anchor that no longer signs anything —
|
|
# and it would look like it worked.
|
|
if [ -e "$root" ] || [ -e "$rootk" ]; then
|
|
echo "swarm root CA half-provisioned ($root / $rootk) — refusing to generate over it" >&2
|
|
exit 1
|
|
fi
|
|
|
|
echo "generating swarm root CA at $root"
|
|
# pathlen:1 — the root signs hive CAs, which sign leaves. One
|
|
# intermediate below the root and no deeper.
|
|
openssl req -x509 -newkey rsa:4096 -nodes -sha256 \
|
|
-days ${toString cfg.validityDays} \
|
|
-keyout "$rootk" -out "$root" \
|
|
-subj "/CN=swarm-ca ${swarmLabel}" \
|
|
-addext "basicConstraints=critical,CA:TRUE,pathlen:1" \
|
|
-addext "keyUsage=critical,keyCertSign,cRLSign"
|
|
chmod 0600 "$rootk"
|
|
chmod 0644 "$root"
|
|
'';
|
|
};
|
|
};
|
|
}
|