swarm-nats: give the queue a name, a bao-issued leaf, and require TLS

The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.

- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
  `swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
  every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
  IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
  may only `update` `pki/issue/swarm-nats`, written by
  `swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
  paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
  as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
  the container, restarts nats when it rotates, and re-runs daily.
  It joins glue-bao-readers-policy-order, so it is ordered after its policy
  unit (`after` and `wants`, never `requires`) where the store is on the
  same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
  `allow_non_tls`. `validateConfig` is now off in every mode, because the
  build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
  host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
  `singleHostSwarm`, and the auth responder all dial
  `tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
  to the NATS connection too, so hive-c0re and the controller trust the
  leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
  swarm needs.

module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.

Closes #4626
This commit is contained in:
atlas 2026-09-24 15:43:25 +02:00 • committed by mara
commit 1d261b3fed
17 changed files with 748 additions and 108 deletions

View file

@ -201,16 +201,39 @@ let
userKey = "@USER_PUBKEY@";
issuerKey = "@ISSUER_PUBKEY@";
});
swarmDomain = config.services.hyperhive.swarm.domain;
# Total on a null swarm domain, for the reason ./swarm-bao.nix gives.
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
baoCfg = config.services.hyperhive.swarm.bao;
baoDeploy = deployCfg.bao;
# The queue's TLS leaf, bound read-only at the same path on both sides, the
# shape ./swarm-bao.nix's `tlsDir` uses. The key reaches the server through
# `LoadCredential`, as openbao's does: it is root-owned 0600 here, and the
# server runs as `nats`.
tlsDir = "/var/lib/swarm-nats-tls";
tlsCertPath = "${tlsDir}/cert.pem";
tlsKeyPath = "${tlsDir}/key.pem";
tlsKeyCredential = "tls-key";
tlsKeyCredentialPath = "/run/credentials/nats.service/${tlsKeyCredential}";
# Re-issue once the leaf is past half of the 720h `pki/roles/swarm-nats`
# grants it (./swarm-bao.nix), so the daily timer below has two weeks of
# retries before it lapses.
leafRenewSeconds = 360 * 3600;
in
{
# The swarm's message queue: one NATS server, reached by every hive.
# The swarm's message queue: one NATS server, reached by every hive at
# `tls://<swarm.nats.domain>:<port>`.
#
# ⚠️ There is deliberately NO gateway vhost here, and this is the first
# swarm service where that is true — the next reader will go looking for
# one. NATS speaks its own TCP protocol rather than HTTP, so nginx
# cannot front it the way it fronts the forge, matrix and authelia.
# Cross-hive reach is the wireguard mesh; `gateway.localNames` and the
# per-service vhost pattern do not apply.
# ⚠️ There is deliberately NO gateway vhost here. NATS speaks its own TCP
# protocol rather than HTTP, so nginx cannot front it the way it fronts the
# forge, matrix and authelia, and the server terminates TLS itself. The
# name is the half of the sibling pattern that applies, the way it is for
# bao: `gateway.localNames` on this host, the operator's DNS everywhere
# else.
options.services.hyperhive.swarm.nats = {
# `enable` moved to `services.hyperhive.deploy.nats.enable` — see
@ -226,6 +249,25 @@ in
# with `nixpkgs.overlays` if you need to. (`deploy.nats.authPackage`
# is a different thing: the callout responder, which IS ours.)
domain = lib.mkOption {
type = lib.types.str;
default = "nats.${domainBase}";
defaultText = lib.literalExpression ''"nats.''${services.hyperhive.swarm.domain}"'';
description = ''
Name every client reaches the queue on, as
`tls://<domain>:<port>`, and the only name its certificate carries.
A **sibling** of the swarm's other service names, for the reason
{option}`services.hyperhive.swarm.bao.domain` gives.
The host running the queue answers it through `gateway.localNames`.
Every other hive resolves it through the operator's upstream DNS,
which is where a multi-host swarm needs a record for it.
Not in {option}`services.hyperhive.swarm.serviceDomains`: that list
is the names a gateway vhost fronts, and nginx fronts nothing here.
'';
};
port = lib.mkOption {
type = lib.types.port;
default = 4222;
@ -234,12 +276,9 @@ in
sits outside hyperhive's claimed ranges (dashboard 7000, forge
3000, matrix 8008, every agent in 8100..8999 via FNV-1a hash).
Contributed to
`services.hyperhive.network.exposeHostPorts`, which opens it on
the bridge interface only — so it is reachable from agent
containers and not from the outside world. An agent connects to
`nats://<bridgeIp>:<port>`; loopback inside a container is the
agent itself, not this host.
Opened on the bridge interface, for agent containers, and on
`wg-hive` when this host is on the mesh, for other hives. Never
host-wide. TLS only: a client that does not speak it is refused.
Unlike {option}`monitorPort` and {option}`metricsPort`, the
address is not what bounds who may use this port: the queue's
@ -429,6 +468,36 @@ in
`calloutUserSeedFile`.
'';
};
baoClientCertFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-bao-pki/nats.pem";
description = ''
Client certificate `swarm-bao-nats-tls` presents to the swarm's
secret store when it asks for the queue's TLS leaf. Its subject must
be {option}`services.hyperhive.deploy.bao.natsCommonName`: cert auth
matches on the CN, and that role's policy may issue from
`pki/issue/swarm-nats` and nothing else.
No default. ./glue-nats-bao-identity.nix points it at the leaf
./glue-bao-tls.nix mints, where this host mints one. On a queue host
without the store, it is a file the operator copies from the store's
host.
A path, never a value.
'';
};
baoClientKeyFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-bao-pki/nats-key.pem";
description = ''
Private key for {option}`services.hyperhive.deploy.nats.baoClientCertFile`.
A path, never a value.
'';
};
};
config = lib.mkIf deployCfg.nats.enable {
@ -437,12 +506,17 @@ in
services.hyperhive.swarm.otel.journaldUnits = [
"nats"
"swarm-nats-auth"
"swarm-bao-nats-tls"
];
# Resolver only: NATS speaks its own protocol, so nginx fronts
# nothing here — but the auth responder introspects authelia by name.
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# The name every client dials, answered with the bridge IP on this host.
# Other hives resolve it through the operator's DNS, as they do bao's.
services.hyperhive.gateway.localNames = [ cfg.domain ];
assertions = [
{
# Fail at EVAL, not at boot: a queue that comes up unable to
@ -575,18 +649,27 @@ in
# by the firewall and looks exactly like every other NATS failure,
# a timeout.
#
# ⚠️ Bridge interface only — never the world. What makes it safe to
# open at all is that the queue is fail-closed: `auth_callout` admits
# nobody until the responder above answers for them, so an agent that
# reaches this port still has to present a token authelia vouches for.
# ⚠️ Never the world. What makes it safe to open at all is that the
# queue is fail-closed: `auth_callout` admits nobody until the responder
# above answers for them, so an agent that reaches this port still has to
# present a token authelia vouches for.
#
# An agent connects to `nats://<bridgeIp>:<port>`, NOT to
# `nats://127.0.0.1:<port>` — inside a container loopback is the
# *agent*. Every url in this repo today is the loopback one and each
# is correct for its reader, because those readers share the host
# netns; an agent does not.
# An agent reaches it by name, which dnsmasq answers with the bridge IP.
services.hyperhive.network.exposeHostPorts = [ cfg.port ];
# The other hives' way in, interface-scoped like
# ./swarm-snapshot-store.nix's receiver. The firewall is default-deny
# before a packet reaches the listener, so without this nothing in-tree
# lets another hive reach the queue at all.
networking.firewall.interfaces.wg-hive.allowedTCPPorts = lib.mkIf deployCfg.wireguard.enable [
cfg.port
];
# The bind source has to exist before the container starts, leaf or not:
# nixos-container refuses to start over a missing one, which would take
# the whole queue down rather than only its TLS.
systemd.tmpfiles.rules = [ "d ${tlsDir} 0755 root root -" ];
containers.swarm-nats = {
autoStart = true;
ephemeral = false;
@ -605,9 +688,14 @@ in
# Not agent containers, though: they have a netns of their own and
# the bridge firewall does not open this port.
privateNetwork = false;
# Binds only the public trust bundle, read-only. Empty when the gateway
# is not self-signed, so the whole trust path drops out cleanly.
bindMounts = caTrust.bindMount;
# The public trust bundle (empty when the gateway is not self-signed)
# and the queue's own TLS leaf, both read-only.
bindMounts = caTrust.bindMount // {
${tlsDir} = {
hostPath = tlsDir;
isReadOnly = true;
};
};
config =
{ ... }:
{
@ -652,13 +740,14 @@ in
serverName = "swarm-nats";
port = cfg.port;
# In auto mode the keys are empty until the generator runs,
# and `nats-server -t` rejects that ("Expected callout user to
# be a valid public account nkey, got \"\""), so leaving this
# on fails the BUILD of every all-local hive. Upstream's own
# description names the case: disable it when the config
# includes other files. The check moves to server start.
validateConfig = !deployCfg.nats.autoGenerateCallout;
# Off in every mode now. The `tls` block below names a leaf that
# exists only at runtime, and `nats-server -t` loads it, so the
# build-time check fails on every hive. In auto mode it failed
# already, on callout keys that are empty until the generator
# runs. Upstream's own description names the case: disable it
# when the config includes other files. The check moves to
# server start.
validateConfig = false;
settings =
calloutBlocks {
@ -687,9 +776,28 @@ in
# ⚠️ Loopback: unauthenticated, and `/connz` lists every
# connected client.
http = "127.0.0.1:${toString cfg.monitorPort}";
# TLS required on the client port: a server with a `tls`
# block advertises `tls_required` and refuses a client that
# does not upgrade. No `allow_non_tls`, which would keep the
# plaintext path open across the mesh. The 2s timeout is for
# a handshake that crosses the mesh; upstream's 0.5s is sized
# for a LAN.
tls = {
cert_file = tlsCertPath;
key_file = tlsKeyCredentialPath;
timeout = 2;
};
};
};
# The key as the `nats` user can read it; see `tlsDir`. Resolved at
# unit start only, which is why a renewed leaf restarts the server
# rather than reloading it.
systemd.services.nats.serviceConfig.LoadCredential = [
"${tlsKeyCredential}:${tlsKeyPath}"
];
# The translation layer. NATS has no Prometheus format of its
# own, so this reads the JSON monitoring endpoint above and
# re-serves it in the format the collector scrapes.
@ -754,7 +862,11 @@ in
serviceConfig = {
ExecStart = lib.concatStringsSep " " [
"${deployCfg.nats.authPackage}/bin/swarm-nats-auth"
"--nats-url nats://127.0.0.1:${toString cfg.port}"
# By name, like every other client: the leaf carries the name
# and no IP, so a loopback address fails verification. The
# container's resolver goes through the bridge, where dnsmasq
# answers it.
"--nats-url tls://${cfg.domain}:${toString cfg.port}"
"--user-seed-file \${CREDENTIALS_DIRECTORY}/callout-user.seed"
"--issuer-seed-file \${CREDENTIALS_DIRECTORY}/issuer.seed"
"--client-secret-file \${CREDENTIALS_DIRECTORY}/oidc-client.secret"
@ -966,5 +1078,148 @@ in
${lib.escapeShellArg hostRuntimeDir}/nats.conf
'';
};
# The queue's TLS leaf, issued by the secret store's `pki/roles/swarm-nats`
# under this host's `swarm-nats` identity. Same shape as ./hive-tls.nix's
# `swarm-services-cert`, whose comments carry the reasoning for the login,
# the single `issue` call split with `jq`, and the unseal-sized retry.
#
# Before the container, so a normal boot starts the server with its leaf.
# A store that comes up later fails this run; the retry lands the leaf and
# restarts the server, which until then refuses to start (TLS required,
# and its key credential is missing).
systemd.services.swarm-bao-nats-tls = {
description = "Issue the swarm queue's TLS leaf from the secret store's PKI";
wantedBy = [
"multi-user.target"
"container@${machine}.service"
];
before = [ "container@${machine}.service" ];
# Both absent where the store runs elsewhere, and ignored there; see
# `swarm-services-cert`. The ordering after this leaf's policy unit is
# ./glue-bao-readers-policy-order.nix's, because it holds only where the
# store is here too.
after = [
"container@${baoCfg.machine}.service"
"swarm-bao-pki.service"
];
wants = [ "container@${baoCfg.machine}.service" ];
path = [
baoDeploy.package
pkgs.jq
pkgs.openssl
pkgs.coreutils
pkgs.gnugrep
pkgs.systemd
];
startLimitBurst = 2880;
startLimitIntervalSec = 90000;
# No `RemainAfterExit`: the timer below starts this again, and starting
# an active unit is a no-op.
serviceConfig = {
Type = "oneshot";
UMask = "0077";
SyslogIdentifier = "swarm-bao-nats-tls";
Restart = "on-failure";
RestartSec = 30;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
}
// lib.optionalAttrs (deployCfg.nats.baoClientCertFile != null) {
BAO_CLIENT_CERT = deployCfg.nats.baoClientCertFile;
}
// lib.optionalAttrs (deployCfg.nats.baoClientKeyFile != null) {
BAO_CLIENT_KEY = deployCfg.nats.baoClientKeyFile;
}
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
BAO_CACERT = baoDeploy.serverCaFile;
};
script = ''
set -euo pipefail
d=${lib.escapeShellArg tlsDir}
name=${lib.escapeShellArg cfg.domain}
install -d -m 0755 "$d"
# Missing, past half its window, or naming something other than the
# configured domain.
reissue=0
{ [ -s "$d/cert.pem" ] && [ -s "$d/key.pem" ]; } || reissue=1
openssl x509 -in "$d/cert.pem" -noout -checkend ${toString leafRenewSeconds} >/dev/null 2>&1 || reissue=1
openssl x509 -in "$d/cert.pem" -noout -checkhost "$name" 2>/dev/null | grep -q ' does match ' || reissue=1
if [ "$reissue" = 0 ]; then
echo "queue leaf valid for $name — leaving it alone"
exit 0
fi
${
if deployCfg.nats.baoClientCertFile == null || deployCfg.nats.baoClientKeyFile == null then
''
echo "no bao client certificate configured for the queue, so its TLS leaf" >&2
echo "cannot be requested from the store." >&2
echo "Set services.hyperhive.deploy.nats.baoClient{Cert,Key}File" >&2
echo "to a leaf the store's CA signed with CN=${baoDeploy.natsCommonName}." >&2
exit 1''
else
""
}
err="$(mktemp)"
trap 'rm -f "$err"' EXIT
if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then
echo "could not log in to the swarm secret store with the queue's certificate." >&2
cat "$err" >&2
exit 1
fi
export BAO_TOKEN
echo "requesting the queue's TLS leaf for $name from the store"
resp="$(mktemp "$d/issue.json.XXXXXX")"
trap 'rm -f "$err" "$resp"' EXIT
if ! bao write -format=json \
${lib.escapeShellArg "${baoDeploy.servicesPkiMountPath}/issue/${baoDeploy.natsPkiRoleName}"} \
common_name="$name" > "$resp" 2>"$err"; then
echo "the store refused to issue the queue's certificate." >&2
cat "$err" >&2
exit 1
fi
jq -r '.data.private_key // empty' < "$resp" > "$d/key.pem.new"
jq -r '.data.certificate // empty' < "$resp" > "$d/cert.pem.new"
for f in "$d/key.pem.new" "$d/cert.pem.new"; do
if [ ! -s "$f" ]; then
echo "the store's response was missing a field: $f is empty" >&2
exit 1
fi
done
chmod 0600 "$d/key.pem.new"
chmod 0644 "$d/cert.pem.new"
mv -f "$d/key.pem.new" "$d/key.pem"
mv -f "$d/cert.pem.new" "$d/cert.pem"
# Renewal, or the late-store retry. On a normal boot the container is
# not up yet and starts the server with this leaf itself. A restart,
# not a reload: the key arrives by `LoadCredential`, which systemd
# resolves at start only. `reset-failed` because a server that found
# no key may have hit its start limit.
if systemctl is-active --quiet ${lib.escapeShellArg "container@${machine}.service"}; then
echo "queue leaf rotated — restarting the server in ${machine}"
systemctl --machine=${lib.escapeShellArg machine} reset-failed nats.service
systemctl --machine=${lib.escapeShellArg machine} restart nats.service
fi
'';
};
systemd.timers.swarm-bao-nats-tls = {
description = "Daily renewal check for the swarm queue's TLS leaf";
wantedBy = [ "timers.target" ];
timerConfig = {
OnCalendar = "daily";
Persistent = true;
};
};
};
}