hyperhive/nix/host-modules/swarm.nix
atlas 48f69fcdea feat(swarm): wire a hive's queue coordinates for status publishing
Three options, all three derived from ONE predicate — this host runs both
the queue and the IdP — so a defaulted set is all or nothing. Deriving
them per-service looks equivalent and is not: `enableRequiredServices`
turns on matrix and authelia but not nats, so an ordinary all-local hive
would resolve two of three and trip the assertion below. Making the
partial state unrepresentable is what keeps that assertion honest.

Deliberately not the shape swarm-controller uses. That module emits its
queue coordinates only when authelia and NATS are local, which is right
for a service that *is* a swarm-host service — but a hive is the one thing
in a swarm that routinely is not on the swarm host, so the same rule would
make status publishing work on exactly the deployment that needs it least.

There is no `enable`: three coordinates that are all set is the enable. An
extra flag would allow configured-but-off, which is one more state to
explain and one more way to be silently quiet.

A half-set trio is an eval error rather than a silent no-op, because its
runtime failure mode is the expensive kind — the daemon comes up fine,
never connects, and the hive reads never_reported on a dashboard nobody is
watching yet. With the defaults all-or-nothing, the assertion only ever
judges what an operator typed by hand.

The secret arrives by LoadCredential, not a copy: hive-c0re is a host
unit, so systemd hands it the file directly and the secret never gains a
second on-disk copy. The client id is not chosen here either — it is
`hive-<hiveName>`, the identity swarm-authelia.nix already declares for
every entry in the roster.
2026-08-16 13:14:03 +02:00

480 lines
22 KiB
Nix

# The swarm's directory: one entry per hive, **including this one**,
# identical on every host in the swarm. `services.hyperhive.hiveName`
# says which entry is us, and `peerHives` below derives the rest.
#
# Why a directory rather than a per-host peer list: every field here is
# intrinsic to the hive it describes — none of them says anything about
# the *pair*. A list where every field is intrinsic is a directory each
# host was keeping its own copy of, which is O(n²) duplication that
# deduplicates without loss. It is also a correctness gain: two hosts
# could hold different endpoints for the same third hive and nothing
# detected it. One entry per hive makes that unrepresentable.
#
# Consumed by swarm-controller's own hive directory (its `/api/hives`,
# swarm-controller.nix) and the mesh in ./swarm-wireguard.nix. The mesh
# lives there rather than here because bringing up an interface is
# host networking rather than swarm bookkeeping, and a host that runs
# no hive still needs it.
{
lib,
config,
...
}:
let
cfg = config.services.hyperhive;
swarmCfg = cfg.swarm;
# Public hostnames of the swarm's own services, in declaration order.
# `serviceDomains` below is this set sorted + deduplicated.
#
# ⚠️ These are NOT required to be under `swarm.domain`. An earlier
# revision asserted that, reasoning that the services sub-CA is
# constrained to the swarm's tree — but the sub-CA is constrained to
# the **configured names** (./swarm-ca.nix) and the swarm root carries
# no name constraints at all, so any configured name is issuable. The
# assertion encoded an intended shape, not a property of the code, and
# it rejected the supported migration path: a hive pinning its old
# `forge.<hive domain>` while joining a swarm at a different apex.
serviceDomains' = [
swarmCfg.forge.domain
swarmCfg.matrix.gatewayHost
swarmCfg.authelia.domain
]
# The swarm UI's name is a SIBLING of the other three, not a parent of
# them — the apex is as much a name needing a certificate as
# `forge.<apex>` is, and no CA in the hierarchy issues for it
# implicitly. Left out, its vhost falls back to the hive leaf and the
# swarm's front page opens with a name mismatch.
++ lib.optional swarmCfg.ui.enable swarmCfg.ui.domain;
# Hives whose entry still carries the removed `certFingerprint`. Scanned
# here, at top level, because that is the only place an assertion about a
# submodule field can actually live — see the option's own comment.
pinnedHives = lib.attrNames (
lib.filterAttrs (_: hive: hive.certFingerprint != null) swarmCfg.hives
);
# Whether this host can derive its own status-publishing coordinates —
# ONE condition for all three of them, deliberately.
#
# 🩸 They were three independent conditions first, and that was wrong in
# a way only an eval gate finds: `enableRequiredServices` turns on
# matrix and authelia but NOT nats (nats has no mode that enables it —
# see the auto-deploy question on the swarm-queue issue), so an ordinary
# all-local hive resolved authelia's two coordinates and not the queue
# URL. Two of three set is exactly what the assertion below rejects, so
# every `enableAllLocalDefaults` hive would have stopped evaluating.
#
# Deriving all three from one predicate makes the partial state
# unrepresentable rather than merely detected: a default set is all or
# nothing, and the assertion is then only ever about what an operator
# typed.
queueLocal = swarmCfg.nats.enable && swarmCfg.authelia.enable && cfg.hiveName != null;
in
{
options.services.hyperhive.swarm.hives = lib.mkOption {
type = lib.types.attrsOf (
lib.types.submodule (
{ name, ... }:
{
options = {
domain = lib.mkOption {
type = lib.types.str;
# `<name>.<swarm.domain>` is a derivation from two values an
# operator had to state explicitly (both are required), not a
# guess — and it is what makes this directory worth copying:
# a conventional swarm is `{ pr1ma = { }; umbra = { }; }`,
# names only, with a non-conventional hive saying so and only
# that. A shared file is read far more often than written.
#
# Total rather than a throw when `swarm.domain` is unset, for
# the reason ./hive-network.nix:155 gives in full: defaults
# that interpolate the domain are forced *while the assertion
# list evaluates*, so a throw here would replace the message
# naming the missing option with a coercion error naming this
# one. `.invalid` is reserved (RFC 2606) and fails loudly at
# resolution if it ever escaped — which the required-domain
# assertion is there to stop.
default = if swarmCfg.domain == null then "${name}.invalid" else "${name}.${swarmCfg.domain}";
defaultText = lib.literalExpression ''"''${name}.''${services.hyperhive.swarm.domain}"'';
example = "lab.example.com";
description = ''
Public DNS domain this hive occupies used for
swarm-controller's hive roster, agent identity
(qualified `agent@domain` labels), and Matrix federation
discovery.
Defaults to `<name>.<swarm.domain>`, the convention every
hive in a swarm follows, so a conventional directory is
names only. Set it for a hive that is addressed by
something else.
'';
};
wireguardPublicKey = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "base64pubkey=";
description = ''
WireGuard public key for this hive's host. Required when
`services.hyperhive.swarm.wireguard.enable = true` and
you want this hive reachable over the mesh. Null = TLS-
only peering (public internet, no mesh tunnel).
'';
};
wireguardEndpoint = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "203.0.113.1:51820";
description = ''
WireGuard endpoint for this hive in `host:port` form.
Null = this hive has no reachable endpoint, so the tunnel
is initiated from the other side.
Reads like a fact about the relationship and is not: it
says whether *this* hive can be dialled, which every other
hive in the swarm needs the same answer to.
'';
};
wireguardAddress = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "10.100.0.2/32";
description = ''
IP address (with prefix) of this hive's host on the
WireGuard mesh. Used as the `allowedIPs` for its
WireGuard config entry (`./swarm-wireguard.nix`), so
intra-swarm traffic can route over the mesh address
rather than the public domain. Required to include
a hive in the mesh (entries missing this field are
silently excluded from `wg-hive`).
'';
};
# Removed, and re-declared only so an existing definition
# produces an error that says *what happened*. Without it the
# module system says `The option ... does not exist`, which
# tells an operator nothing about why it went or what replaced
# it — and this field was set on real deployments.
#
# ⚠️ `lib.mkRemovedOptionModule` cannot do this job here, and
# neither of its two halves survives the move into a submodule:
# its `apply = throw` fires only when the value is READ, and
# nothing reads this any more (that being the point of removing
# it); its `config.assertions` half would land on a submodule
# that declares no `assertions` option at all. It is a
# top-level tool. The assertion below is the working shape, and
# it is the same one ./swarm-peers-removed.nix uses for the
# analogous `peers.<hive>.caCert`.
certFingerprint = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
visible = false;
internal = true;
description = ''
Removed. Trust inside a swarm comes from the swarm root CA
(`services.hyperhive.swarm.ca`, docs/swarm/ca.md), which
replaces per-hive leaf pinning entirely.
'';
};
};
}
)
);
default = { };
example = {
pr1ma = {
wireguardAddress = "10.100.0.1/32";
wireguardEndpoint = "203.0.113.1:51820";
};
edge = {
domain = "edge.elsewhere.example";
wireguardAddress = "10.100.0.2/32";
};
};
description = ''
Every hive in this swarm, keyed by `hiveName` **including this
host's own hive**. The same attrset is meant to be identical on
every host in the swarm, so it can be written once and shared;
`services.hyperhive.hiveName` is what makes a given host read it
as "me and four others" rather than "five peers".
Each entry needs no fields at all in the conventional case: a
hive's `domain` defaults to `<name>.<swarm.domain>`, so the whole
directory is usually a list of names.
It must contain an entry for `hiveName`, which is asserted this
host's own address is read out of it (it is where
`services.hyperhive.domain` derives from), and a hive that lists
everyone but itself would otherwise derive its own peer set as
*everything* and peer with itself.
'';
};
options.services.hyperhive.swarm.peerHives = lib.mkOption {
type = lib.types.attrsOf (lib.types.attrsOf lib.types.unspecified);
readOnly = true;
internal = true;
description = ''
Read-only: `hives` minus this host's own entry. Derived once here
rather than in each consumer, because "everything that isn't me"
is a filter four different modules were re-implementing and only
one of them has to be wrong for a hive to peer with itself.
'';
};
options.services.hyperhive.swarm.serviceDomains = lib.mkOption {
type = lib.types.listOf lib.types.str;
readOnly = true;
internal = true;
description = ''
Read-only: the public hostnames of the swarm's own services, in a
stable sorted order. Second derived set alongside `peerHives`, and
here for the same reason the CA that name-constrains these and
the leaf that carries them as SANs must agree exactly, and two
modules each assembling the list is how they stop agreeing.
Sorted and deduplicated deliberately: consumers compare this list
against what they issued last time to decide whether to re-issue,
so an unstable order would churn a certificate that other things
are meant to pin.
'';
};
config = {
services.hyperhive.swarm.peerHives = lib.filterAttrs (name: _: name != cfg.hiveName) swarmCfg.hives;
services.hyperhive.swarm.serviceDomains = lib.sort (a: b: a < b) (
lib.unique (lib.filter (d: d != null && d != "") serviceDomains')
);
assertions = [
{
# An EMPTY `hives` fires this too, deliberately: since
# `swarm.domain` became required, every hive is in a swarm — a
# swarm of one is still a swarm — so a directory with no entry
# for this host is missing one either way. It also has to fire
# here, because this host's own domain is now read out of the
# directory: without the entry `services.hyperhive.domain` is
# null and the generic required-domain assertion in
# ./hive-network.nix would fire instead, naming an option the
# operator should no longer be setting.
#
# Guarded on `hiveName != null` so the required-hiveName
# assertion in ./hyperhive.nix is what fires for that case —
# two assertions naming the same missing value is noise.
assertion = cfg.hiveName == null || swarmCfg.hives ? ${cfg.hiveName};
message = ''
services.hyperhive.swarm.hives has no entry for this hive
(services.hyperhive.hiveName = "${toString cfg.hiveName}").
`hives` describes every hive in the swarm including this one,
so that every host can share one identical attrset, and this
hive's own domain is derived from its entry. Add:
services.hyperhive.swarm.hives."${toString cfg.hiveName}" = { };
No fields are needed: `domain` defaults to
`<name>.<swarm.domain>`. Set it in the entry if this hive is
addressed by something else.
Declared hives: ${lib.concatStringsSep ", " (lib.attrNames swarmCfg.hives)}
'';
}
{
# An error rather than a warning, deliberately. Re-declaring the
# removed option above is what stops "option does not exist"; on
# its own it would also turn a config that used to FAIL into one
# that quietly evaluates with the setting ignored, which is a
# worse answer than the unhelpful error it replaces. The operator
# asked for a removal error, so this stays fatal until the field
# is gone from their config.
#
# Wording follows `lib.mkRemovedOptionModule`'s, so it reads like
# every other removal the module system reports.
assertion = pinnedHives == [ ];
message = ''
The option definition `services.hyperhive.swarm.hives.<hive>.certFingerprint'
no longer has any effect; please remove it.
Still set on: ${lib.concatStringsSep ", " pinnedHives}
It pinned a peer hive's TLS leaf for hive-c0re's own peer HTTPS
checks, and was removed along with the dashboard feature it
existed to serve nothing else ever consumed it.
There is no replacement, and none is needed for a hive inside
this swarm: trust comes from the swarm root CA
(services.hyperhive.swarm.ca see docs/swarm/ca.md), which
every hive chains to, so one anchor replaces per-hive pinning.
What that genuinely drops is trusting a hive whose root this
swarm does NOT own another swarm's, or one keeping its own
CA. That is a cross-swarm problem and wants a mechanism
designed for it.
'';
}
{
# Deliberately an assertion and not a silent "then publish
# nothing": a half-set trio is a config an operator believes is
# working, and its runtime failure mode is the expensive one —
# the daemon comes up fine, never connects, and the hive reads
# `never_reported` on a dashboard nobody is watching yet.
#
# Safe to add to an existing deployment: every `statusPublish`
# default is either all-local or all-null, so no config that
# evaluates today can be caught by this. It also encodes a
# property of the code rather than an intended shape — hive-c0re
# genuinely cannot publish with two of three coordinates — which
# is the distinction the `serviceDomains'` comment at the top of
# this file was written about.
assertion =
let
set = lib.filter (v: v != null) [
swarmCfg.statusPublish.natsUrl
swarmCfg.statusPublish.tokenEndpoint
swarmCfg.statusPublish.clientSecretFile
];
in
builtins.length set == 0 || builtins.length set == 3;
message = ''
services.hyperhive.swarm.statusPublish needs natsUrl,
tokenEndpoint and clientSecretFile set together or not at all
this hive has only some of them.
Currently:
natsUrl = ${toString swarmCfg.statusPublish.natsUrl}
tokenEndpoint = ${toString swarmCfg.statusPublish.tokenEndpoint}
clientSecretFile = ${toString swarmCfg.statusPublish.clientSecretFile}
Set the missing ones to publish this hive's status to the
swarm, or set all three to null to turn publishing off.
'';
}
];
};
# `enableRequiredServices` is declared in ./swarm-required-services.nix
# together with the per-service `enable`s it asserts — it is a
# deployment-shape switch rather than swarm bookkeeping, so it lives
# with its consequences instead of here.
options.services.hyperhive.swarm.snapshotStore = {
address = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "10.100.0.1";
description = ''
Mesh address of the swarm's snapshot store the single
`btrfs receive` endpoint every hive in this swarm pushes agent
snapshots to. Bare IP, no prefix.
There is exactly **one** store per swarm, not one per peer: the
receiver keys destinations by *agent*, so an agent that migrates
between hives keeps a single unbroken incremental chain. Per-hive
stores would split that chain in two, which is the case the store
exists to serve.
Null means this swarm has no store configured, and pushing fails
saying so rather than guessing an address. Set it on every hive
that pushes; the receiving host separately sets
`services.hyperhive.snapshotStore.enable`.
'';
};
port = lib.mkOption {
type = lib.types.port;
default = 51821;
description = ''
TCP port the swarm's snapshot store listens on. Must match the
receiving host's `services.hyperhive.snapshotStore.port`.
Defaulted (unlike `address`) because it is a shared convention
both sides read from the same option docs whereas an address
is deployment-specific and cannot be guessed.
'';
};
};
# How this hive reaches the swarm queue to offer its own status
# (hive-c0re's `swarm_status`). Three coordinates, defaulted from the
# local swarm services when this host runs them, and set by hand
# otherwise — one code path for both deployments.
#
# The alternative was the shape swarm-controller uses: emit the
# coordinates only when authelia and NATS are local, and nothing
# otherwise. That is right for the controller, which *is* a swarm-host
# service — but a hive is the one thing in a swarm that routinely is
# not on the swarm host, so the same rule would mean status publishing
# works on exactly the deployment that needs it least.
#
# There is no `enable`: three coordinates that are all set is the
# enable. An extra flag would let a hive be configured-but-off, which
# is one more state to explain and one more way to be silently quiet.
options.services.hyperhive.swarm.statusPublish = {
natsUrl = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = if queueLocal then "nats://127.0.0.1:${toString swarmCfg.nats.port}" else null;
defaultText = lib.literalExpression ''"nats://127.0.0.1:''${swarm.nats.port}" when this host runs the queue and the IdP, else null'';
example = "nats://10.100.0.1:4222";
description = ''
Where the swarm queue listens, as seen from *this* hive.
Defaults to loopback when this host runs the queue container
itself (it shares the host netns, so loopback is correct there
and is not the "localhost means the wrong thing" trap that
applies inside agent containers). A hive that is not the swarm
host has to name the swarm's mesh address.
Null disables status publishing: this hive computes its own
readiness as always, and simply offers it to nobody. The swarm
controller then reports it `never_reported`, which is the honest
reading.
'';
};
tokenEndpoint = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default =
if queueLocal && swarmCfg.authelia.url != null then
"${swarmCfg.authelia.url}/api/oidc/token"
else
null;
defaultText = lib.literalExpression ''"''${swarm.authelia.url}/api/oidc/token" when this host runs both the queue and the IdP, else null'';
example = "https://auth.example.com/api/oidc/token";
description = ''
The swarm IdP's OAuth2 token endpoint. This hive mints a
`client_credentials` access token there and presents it when
connecting to the queue, which authenticates it as
`hive-<hiveName>` the client
{file}`nix/host-modules/swarm-authelia.nix` already declares for
every entry in {option}`services.hyperhive.swarm.hives`.
'';
};
clientSecretFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default =
if queueLocal then "${swarmCfg.authelia.hostClientSecretDir}/hive-${cfg.hiveName}.secret" else null;
defaultText = lib.literalExpression ''"''${swarm.authelia.hostClientSecretDir}/hive-''${hiveName}.secret" when this host runs both the queue and the IdP, else null'';
example = "/var/lib/secrets/swarm-queue-client.secret";
description = ''
Path to a file holding the plaintext client secret for this
hive's `hive-<hiveName>` identity.
A path and not a value: a secret in the Nix store is world
readable, and one in the environment is readable by anything
that can open {file}`/proc/<pid>/environ`.
Defaults to authelia's own minted secret when the IdP runs on
this host. On any other hive the secret has to get here somehow,
and the swarm does not distribute it copy it out of the swarm
host's
{option}`services.hyperhive.swarm.authelia.hostClientSecretDir`
with whatever secret management this deployment already uses.
'';
};
};
}