Three options, all three derived from ONE predicate — this host runs both the queue and the IdP — so a defaulted set is all or nothing. Deriving them per-service looks equivalent and is not: `enableRequiredServices` turns on matrix and authelia but not nats, so an ordinary all-local hive would resolve two of three and trip the assertion below. Making the partial state unrepresentable is what keeps that assertion honest. Deliberately not the shape swarm-controller uses. That module emits its queue coordinates only when authelia and NATS are local, which is right for a service that *is* a swarm-host service — but a hive is the one thing in a swarm that routinely is not on the swarm host, so the same rule would make status publishing work on exactly the deployment that needs it least. There is no `enable`: three coordinates that are all set is the enable. An extra flag would allow configured-but-off, which is one more state to explain and one more way to be silently quiet. A half-set trio is an eval error rather than a silent no-op, because its runtime failure mode is the expensive kind — the daemon comes up fine, never connects, and the hive reads never_reported on a dashboard nobody is watching yet. With the defaults all-or-nothing, the assertion only ever judges what an operator typed by hand. The secret arrives by LoadCredential, not a copy: hive-c0re is a host unit, so systemd hands it the file directly and the secret never gains a second on-disk copy. The client id is not chosen here either — it is `hive-<hiveName>`, the identity swarm-authelia.nix already declares for every entry in the roster.
480 lines
22 KiB
Nix
480 lines
22 KiB
Nix
# The swarm's directory: one entry per hive, **including this one**,
|
|
# identical on every host in the swarm. `services.hyperhive.hiveName`
|
|
# says which entry is us, and `peerHives` below derives the rest.
|
|
#
|
|
# Why a directory rather than a per-host peer list: every field here is
|
|
# intrinsic to the hive it describes — none of them says anything about
|
|
# the *pair*. A list where every field is intrinsic is a directory each
|
|
# host was keeping its own copy of, which is O(n²) duplication that
|
|
# deduplicates without loss. It is also a correctness gain: two hosts
|
|
# could hold different endpoints for the same third hive and nothing
|
|
# detected it. One entry per hive makes that unrepresentable.
|
|
#
|
|
# Consumed by swarm-controller's own hive directory (its `/api/hives`,
|
|
# swarm-controller.nix) and the mesh in ./swarm-wireguard.nix. The mesh
|
|
# lives there rather than here because bringing up an interface is
|
|
# host networking rather than swarm bookkeeping, and a host that runs
|
|
# no hive still needs it.
|
|
{
|
|
lib,
|
|
config,
|
|
...
|
|
}:
|
|
let
|
|
cfg = config.services.hyperhive;
|
|
swarmCfg = cfg.swarm;
|
|
|
|
# Public hostnames of the swarm's own services, in declaration order.
|
|
# `serviceDomains` below is this set sorted + deduplicated.
|
|
#
|
|
# ⚠️ These are NOT required to be under `swarm.domain`. An earlier
|
|
# revision asserted that, reasoning that the services sub-CA is
|
|
# constrained to the swarm's tree — but the sub-CA is constrained to
|
|
# the **configured names** (./swarm-ca.nix) and the swarm root carries
|
|
# no name constraints at all, so any configured name is issuable. The
|
|
# assertion encoded an intended shape, not a property of the code, and
|
|
# it rejected the supported migration path: a hive pinning its old
|
|
# `forge.<hive domain>` while joining a swarm at a different apex.
|
|
serviceDomains' = [
|
|
swarmCfg.forge.domain
|
|
swarmCfg.matrix.gatewayHost
|
|
swarmCfg.authelia.domain
|
|
]
|
|
# The swarm UI's name is a SIBLING of the other three, not a parent of
|
|
# them — the apex is as much a name needing a certificate as
|
|
# `forge.<apex>` is, and no CA in the hierarchy issues for it
|
|
# implicitly. Left out, its vhost falls back to the hive leaf and the
|
|
# swarm's front page opens with a name mismatch.
|
|
++ lib.optional swarmCfg.ui.enable swarmCfg.ui.domain;
|
|
|
|
# Hives whose entry still carries the removed `certFingerprint`. Scanned
|
|
# here, at top level, because that is the only place an assertion about a
|
|
# submodule field can actually live — see the option's own comment.
|
|
pinnedHives = lib.attrNames (
|
|
lib.filterAttrs (_: hive: hive.certFingerprint != null) swarmCfg.hives
|
|
);
|
|
|
|
# Whether this host can derive its own status-publishing coordinates —
|
|
# ONE condition for all three of them, deliberately.
|
|
#
|
|
# 🩸 They were three independent conditions first, and that was wrong in
|
|
# a way only an eval gate finds: `enableRequiredServices` turns on
|
|
# matrix and authelia but NOT nats (nats has no mode that enables it —
|
|
# see the auto-deploy question on the swarm-queue issue), so an ordinary
|
|
# all-local hive resolved authelia's two coordinates and not the queue
|
|
# URL. Two of three set is exactly what the assertion below rejects, so
|
|
# every `enableAllLocalDefaults` hive would have stopped evaluating.
|
|
#
|
|
# Deriving all three from one predicate makes the partial state
|
|
# unrepresentable rather than merely detected: a default set is all or
|
|
# nothing, and the assertion is then only ever about what an operator
|
|
# typed.
|
|
queueLocal = swarmCfg.nats.enable && swarmCfg.authelia.enable && cfg.hiveName != null;
|
|
in
|
|
{
|
|
options.services.hyperhive.swarm.hives = lib.mkOption {
|
|
type = lib.types.attrsOf (
|
|
lib.types.submodule (
|
|
{ name, ... }:
|
|
{
|
|
options = {
|
|
domain = lib.mkOption {
|
|
type = lib.types.str;
|
|
# `<name>.<swarm.domain>` is a derivation from two values an
|
|
# operator had to state explicitly (both are required), not a
|
|
# guess — and it is what makes this directory worth copying:
|
|
# a conventional swarm is `{ pr1ma = { }; umbra = { }; }`,
|
|
# names only, with a non-conventional hive saying so and only
|
|
# that. A shared file is read far more often than written.
|
|
#
|
|
# Total rather than a throw when `swarm.domain` is unset, for
|
|
# the reason ./hive-network.nix:155 gives in full: defaults
|
|
# that interpolate the domain are forced *while the assertion
|
|
# list evaluates*, so a throw here would replace the message
|
|
# naming the missing option with a coercion error naming this
|
|
# one. `.invalid` is reserved (RFC 2606) and fails loudly at
|
|
# resolution if it ever escaped — which the required-domain
|
|
# assertion is there to stop.
|
|
default = if swarmCfg.domain == null then "${name}.invalid" else "${name}.${swarmCfg.domain}";
|
|
defaultText = lib.literalExpression ''"''${name}.''${services.hyperhive.swarm.domain}"'';
|
|
example = "lab.example.com";
|
|
description = ''
|
|
Public DNS domain this hive occupies — used for
|
|
swarm-controller's hive roster, agent identity
|
|
(qualified `agent@domain` labels), and Matrix federation
|
|
discovery.
|
|
|
|
Defaults to `<name>.<swarm.domain>`, the convention every
|
|
hive in a swarm follows, so a conventional directory is
|
|
names only. Set it for a hive that is addressed by
|
|
something else.
|
|
'';
|
|
};
|
|
|
|
wireguardPublicKey = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default = null;
|
|
example = "base64pubkey=";
|
|
description = ''
|
|
WireGuard public key for this hive's host. Required when
|
|
`services.hyperhive.swarm.wireguard.enable = true` and
|
|
you want this hive reachable over the mesh. Null = TLS-
|
|
only peering (public internet, no mesh tunnel).
|
|
'';
|
|
};
|
|
|
|
wireguardEndpoint = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default = null;
|
|
example = "203.0.113.1:51820";
|
|
description = ''
|
|
WireGuard endpoint for this hive in `host:port` form.
|
|
Null = this hive has no reachable endpoint, so the tunnel
|
|
is initiated from the other side.
|
|
|
|
Reads like a fact about the relationship and is not: it
|
|
says whether *this* hive can be dialled, which every other
|
|
hive in the swarm needs the same answer to.
|
|
'';
|
|
};
|
|
|
|
wireguardAddress = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default = null;
|
|
example = "10.100.0.2/32";
|
|
description = ''
|
|
IP address (with prefix) of this hive's host on the
|
|
WireGuard mesh. Used as the `allowedIPs` for its
|
|
WireGuard config entry (`./swarm-wireguard.nix`), so
|
|
intra-swarm traffic can route over the mesh address
|
|
rather than the public domain. Required to include
|
|
a hive in the mesh (entries missing this field are
|
|
silently excluded from `wg-hive`).
|
|
'';
|
|
};
|
|
|
|
# Removed, and re-declared only so an existing definition
|
|
# produces an error that says *what happened*. Without it the
|
|
# module system says `The option ... does not exist`, which
|
|
# tells an operator nothing about why it went or what replaced
|
|
# it — and this field was set on real deployments.
|
|
#
|
|
# ⚠️ `lib.mkRemovedOptionModule` cannot do this job here, and
|
|
# neither of its two halves survives the move into a submodule:
|
|
# its `apply = throw` fires only when the value is READ, and
|
|
# nothing reads this any more (that being the point of removing
|
|
# it); its `config.assertions` half would land on a submodule
|
|
# that declares no `assertions` option at all. It is a
|
|
# top-level tool. The assertion below is the working shape, and
|
|
# it is the same one ./swarm-peers-removed.nix uses for the
|
|
# analogous `peers.<hive>.caCert`.
|
|
certFingerprint = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default = null;
|
|
visible = false;
|
|
internal = true;
|
|
description = ''
|
|
Removed. Trust inside a swarm comes from the swarm root CA
|
|
(`services.hyperhive.swarm.ca`, docs/swarm/ca.md), which
|
|
replaces per-hive leaf pinning entirely.
|
|
'';
|
|
};
|
|
};
|
|
}
|
|
)
|
|
);
|
|
default = { };
|
|
example = {
|
|
pr1ma = {
|
|
wireguardAddress = "10.100.0.1/32";
|
|
wireguardEndpoint = "203.0.113.1:51820";
|
|
};
|
|
edge = {
|
|
domain = "edge.elsewhere.example";
|
|
wireguardAddress = "10.100.0.2/32";
|
|
};
|
|
};
|
|
description = ''
|
|
Every hive in this swarm, keyed by `hiveName` — **including this
|
|
host's own hive**. The same attrset is meant to be identical on
|
|
every host in the swarm, so it can be written once and shared;
|
|
`services.hyperhive.hiveName` is what makes a given host read it
|
|
as "me and four others" rather than "five peers".
|
|
|
|
Each entry needs no fields at all in the conventional case: a
|
|
hive's `domain` defaults to `<name>.<swarm.domain>`, so the whole
|
|
directory is usually a list of names.
|
|
|
|
It must contain an entry for `hiveName`, which is asserted — this
|
|
host's own address is read out of it (it is where
|
|
`services.hyperhive.domain` derives from), and a hive that lists
|
|
everyone but itself would otherwise derive its own peer set as
|
|
*everything* and peer with itself.
|
|
'';
|
|
};
|
|
|
|
options.services.hyperhive.swarm.peerHives = lib.mkOption {
|
|
type = lib.types.attrsOf (lib.types.attrsOf lib.types.unspecified);
|
|
readOnly = true;
|
|
internal = true;
|
|
description = ''
|
|
Read-only: `hives` minus this host's own entry. Derived once here
|
|
rather than in each consumer, because "everything that isn't me"
|
|
is a filter four different modules were re-implementing and only
|
|
one of them has to be wrong for a hive to peer with itself.
|
|
'';
|
|
};
|
|
|
|
options.services.hyperhive.swarm.serviceDomains = lib.mkOption {
|
|
type = lib.types.listOf lib.types.str;
|
|
readOnly = true;
|
|
internal = true;
|
|
description = ''
|
|
Read-only: the public hostnames of the swarm's own services, in a
|
|
stable sorted order. Second derived set alongside `peerHives`, and
|
|
here for the same reason — the CA that name-constrains these and
|
|
the leaf that carries them as SANs must agree exactly, and two
|
|
modules each assembling the list is how they stop agreeing.
|
|
|
|
Sorted and deduplicated deliberately: consumers compare this list
|
|
against what they issued last time to decide whether to re-issue,
|
|
so an unstable order would churn a certificate that other things
|
|
are meant to pin.
|
|
'';
|
|
};
|
|
|
|
config = {
|
|
services.hyperhive.swarm.peerHives = lib.filterAttrs (name: _: name != cfg.hiveName) swarmCfg.hives;
|
|
|
|
services.hyperhive.swarm.serviceDomains = lib.sort (a: b: a < b) (
|
|
lib.unique (lib.filter (d: d != null && d != "") serviceDomains')
|
|
);
|
|
|
|
assertions = [
|
|
{
|
|
# An EMPTY `hives` fires this too, deliberately: since
|
|
# `swarm.domain` became required, every hive is in a swarm — a
|
|
# swarm of one is still a swarm — so a directory with no entry
|
|
# for this host is missing one either way. It also has to fire
|
|
# here, because this host's own domain is now read out of the
|
|
# directory: without the entry `services.hyperhive.domain` is
|
|
# null and the generic required-domain assertion in
|
|
# ./hive-network.nix would fire instead, naming an option the
|
|
# operator should no longer be setting.
|
|
#
|
|
# Guarded on `hiveName != null` so the required-hiveName
|
|
# assertion in ./hyperhive.nix is what fires for that case —
|
|
# two assertions naming the same missing value is noise.
|
|
assertion = cfg.hiveName == null || swarmCfg.hives ? ${cfg.hiveName};
|
|
message = ''
|
|
services.hyperhive.swarm.hives has no entry for this hive
|
|
(services.hyperhive.hiveName = "${toString cfg.hiveName}").
|
|
|
|
`hives` describes every hive in the swarm including this one,
|
|
so that every host can share one identical attrset, and this
|
|
hive's own domain is derived from its entry. Add:
|
|
|
|
services.hyperhive.swarm.hives."${toString cfg.hiveName}" = { };
|
|
|
|
No fields are needed: `domain` defaults to
|
|
`<name>.<swarm.domain>`. Set it in the entry if this hive is
|
|
addressed by something else.
|
|
|
|
Declared hives: ${lib.concatStringsSep ", " (lib.attrNames swarmCfg.hives)}
|
|
'';
|
|
}
|
|
{
|
|
# An error rather than a warning, deliberately. Re-declaring the
|
|
# removed option above is what stops "option does not exist"; on
|
|
# its own it would also turn a config that used to FAIL into one
|
|
# that quietly evaluates with the setting ignored, which is a
|
|
# worse answer than the unhelpful error it replaces. The operator
|
|
# asked for a removal error, so this stays fatal until the field
|
|
# is gone from their config.
|
|
#
|
|
# Wording follows `lib.mkRemovedOptionModule`'s, so it reads like
|
|
# every other removal the module system reports.
|
|
assertion = pinnedHives == [ ];
|
|
message = ''
|
|
The option definition `services.hyperhive.swarm.hives.<hive>.certFingerprint'
|
|
no longer has any effect; please remove it.
|
|
|
|
Still set on: ${lib.concatStringsSep ", " pinnedHives}
|
|
|
|
It pinned a peer hive's TLS leaf for hive-c0re's own peer HTTPS
|
|
checks, and was removed along with the dashboard feature it
|
|
existed to serve — nothing else ever consumed it.
|
|
|
|
There is no replacement, and none is needed for a hive inside
|
|
this swarm: trust comes from the swarm root CA
|
|
(services.hyperhive.swarm.ca — see docs/swarm/ca.md), which
|
|
every hive chains to, so one anchor replaces per-hive pinning.
|
|
What that genuinely drops is trusting a hive whose root this
|
|
swarm does NOT own — another swarm's, or one keeping its own
|
|
CA. That is a cross-swarm problem and wants a mechanism
|
|
designed for it.
|
|
'';
|
|
}
|
|
{
|
|
# Deliberately an assertion and not a silent "then publish
|
|
# nothing": a half-set trio is a config an operator believes is
|
|
# working, and its runtime failure mode is the expensive one —
|
|
# the daemon comes up fine, never connects, and the hive reads
|
|
# `never_reported` on a dashboard nobody is watching yet.
|
|
#
|
|
# Safe to add to an existing deployment: every `statusPublish`
|
|
# default is either all-local or all-null, so no config that
|
|
# evaluates today can be caught by this. It also encodes a
|
|
# property of the code rather than an intended shape — hive-c0re
|
|
# genuinely cannot publish with two of three coordinates — which
|
|
# is the distinction the `serviceDomains'` comment at the top of
|
|
# this file was written about.
|
|
assertion =
|
|
let
|
|
set = lib.filter (v: v != null) [
|
|
swarmCfg.statusPublish.natsUrl
|
|
swarmCfg.statusPublish.tokenEndpoint
|
|
swarmCfg.statusPublish.clientSecretFile
|
|
];
|
|
in
|
|
builtins.length set == 0 || builtins.length set == 3;
|
|
message = ''
|
|
services.hyperhive.swarm.statusPublish needs natsUrl,
|
|
tokenEndpoint and clientSecretFile set together or not at all
|
|
— this hive has only some of them.
|
|
|
|
Currently:
|
|
natsUrl = ${toString swarmCfg.statusPublish.natsUrl}
|
|
tokenEndpoint = ${toString swarmCfg.statusPublish.tokenEndpoint}
|
|
clientSecretFile = ${toString swarmCfg.statusPublish.clientSecretFile}
|
|
|
|
Set the missing ones to publish this hive's status to the
|
|
swarm, or set all three to null to turn publishing off.
|
|
'';
|
|
}
|
|
];
|
|
};
|
|
|
|
# `enableRequiredServices` is declared in ./swarm-required-services.nix
|
|
# together with the per-service `enable`s it asserts — it is a
|
|
# deployment-shape switch rather than swarm bookkeeping, so it lives
|
|
# with its consequences instead of here.
|
|
|
|
options.services.hyperhive.swarm.snapshotStore = {
|
|
address = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default = null;
|
|
example = "10.100.0.1";
|
|
description = ''
|
|
Mesh address of the swarm's snapshot store — the single
|
|
`btrfs receive` endpoint every hive in this swarm pushes agent
|
|
snapshots to. Bare IP, no prefix.
|
|
|
|
There is exactly **one** store per swarm, not one per peer: the
|
|
receiver keys destinations by *agent*, so an agent that migrates
|
|
between hives keeps a single unbroken incremental chain. Per-hive
|
|
stores would split that chain in two, which is the case the store
|
|
exists to serve.
|
|
|
|
Null means this swarm has no store configured, and pushing fails
|
|
saying so rather than guessing an address. Set it on every hive
|
|
that pushes; the receiving host separately sets
|
|
`services.hyperhive.snapshotStore.enable`.
|
|
'';
|
|
};
|
|
|
|
port = lib.mkOption {
|
|
type = lib.types.port;
|
|
default = 51821;
|
|
description = ''
|
|
TCP port the swarm's snapshot store listens on. Must match the
|
|
receiving host's `services.hyperhive.snapshotStore.port`.
|
|
|
|
Defaulted (unlike `address`) because it is a shared convention
|
|
both sides read from the same option docs — whereas an address
|
|
is deployment-specific and cannot be guessed.
|
|
'';
|
|
};
|
|
};
|
|
|
|
# How this hive reaches the swarm queue to offer its own status
|
|
# (hive-c0re's `swarm_status`). Three coordinates, defaulted from the
|
|
# local swarm services when this host runs them, and set by hand
|
|
# otherwise — one code path for both deployments.
|
|
#
|
|
# The alternative was the shape swarm-controller uses: emit the
|
|
# coordinates only when authelia and NATS are local, and nothing
|
|
# otherwise. That is right for the controller, which *is* a swarm-host
|
|
# service — but a hive is the one thing in a swarm that routinely is
|
|
# not on the swarm host, so the same rule would mean status publishing
|
|
# works on exactly the deployment that needs it least.
|
|
#
|
|
# There is no `enable`: three coordinates that are all set is the
|
|
# enable. An extra flag would let a hive be configured-but-off, which
|
|
# is one more state to explain and one more way to be silently quiet.
|
|
options.services.hyperhive.swarm.statusPublish = {
|
|
natsUrl = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default = if queueLocal then "nats://127.0.0.1:${toString swarmCfg.nats.port}" else null;
|
|
defaultText = lib.literalExpression ''"nats://127.0.0.1:''${swarm.nats.port}" when this host runs the queue and the IdP, else null'';
|
|
example = "nats://10.100.0.1:4222";
|
|
description = ''
|
|
Where the swarm queue listens, as seen from *this* hive.
|
|
|
|
Defaults to loopback when this host runs the queue container
|
|
itself (it shares the host netns, so loopback is correct there
|
|
and is not the "localhost means the wrong thing" trap that
|
|
applies inside agent containers). A hive that is not the swarm
|
|
host has to name the swarm's mesh address.
|
|
|
|
Null disables status publishing: this hive computes its own
|
|
readiness as always, and simply offers it to nobody. The swarm
|
|
controller then reports it `never_reported`, which is the honest
|
|
reading.
|
|
'';
|
|
};
|
|
|
|
tokenEndpoint = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default =
|
|
if queueLocal && swarmCfg.authelia.url != null then
|
|
"${swarmCfg.authelia.url}/api/oidc/token"
|
|
else
|
|
null;
|
|
defaultText = lib.literalExpression ''"''${swarm.authelia.url}/api/oidc/token" when this host runs both the queue and the IdP, else null'';
|
|
example = "https://auth.example.com/api/oidc/token";
|
|
description = ''
|
|
The swarm IdP's OAuth2 token endpoint. This hive mints a
|
|
`client_credentials` access token there and presents it when
|
|
connecting to the queue, which authenticates it as
|
|
`hive-<hiveName>` — the client
|
|
{file}`nix/host-modules/swarm-authelia.nix` already declares for
|
|
every entry in {option}`services.hyperhive.swarm.hives`.
|
|
'';
|
|
};
|
|
|
|
clientSecretFile = lib.mkOption {
|
|
type = lib.types.nullOr lib.types.str;
|
|
default =
|
|
if queueLocal then "${swarmCfg.authelia.hostClientSecretDir}/hive-${cfg.hiveName}.secret" else null;
|
|
defaultText = lib.literalExpression ''"''${swarm.authelia.hostClientSecretDir}/hive-''${hiveName}.secret" when this host runs both the queue and the IdP, else null'';
|
|
example = "/var/lib/secrets/swarm-queue-client.secret";
|
|
description = ''
|
|
Path to a file holding the plaintext client secret for this
|
|
hive's `hive-<hiveName>` identity.
|
|
|
|
A path and not a value: a secret in the Nix store is world
|
|
readable, and one in the environment is readable by anything
|
|
that can open {file}`/proc/<pid>/environ`.
|
|
|
|
Defaults to authelia's own minted secret when the IdP runs on
|
|
this host. On any other hive the secret has to get here somehow,
|
|
and the swarm does not distribute it — copy it out of the swarm
|
|
host's
|
|
{option}`services.hyperhive.swarm.authelia.hostClientSecretDir`
|
|
with whatever secret management this deployment already uses.
|
|
'';
|
|
};
|
|
};
|
|
|
|
}
|