Watch
0
0
Fork
You've already forked hyperhive
0

nix: split swarm-nats into service and deploy-mode files

`swarm.nats` (what the queue is to every hive: domain, ports, client id)
moves to nix/host-modules/swarm-nats-service.nix, together with the only
two helpers it reads, `swarmDomain` and `domainBase`. Everything else --
the `deploy.nats` options, the whole `config` block including
`containers.swarm-nats`, and the helpers only they read -- stays in
nix/host-modules/swarm-nats.nix, which default.nix now imports alongside
the new file.

A pure move: option paths, option definitions and config are unchanged
apart from the transition comment above `deploy.nats`, which now names the
file `swarm.nats` lives in. The nats fixtures evaluate to the same host and
container toplevel derivations before and after.

Refs #3742
This commit is contained in:
atlas 2026-09-30 23:00:34 +02:00 • committed by mara
commit 07ca06dcca
3 changed files with 123 additions and 112 deletions

View file

@ -44,6 +44,7 @@
./swarm-bao.nix
./swarm-ca.nix
./swarm-secret-publisher.nix
./swarm-nats-service.nix
./swarm-nats.nix
./swarm-controller.nix
./swarm-grafana.nix

View file

@ -0,0 +1,117 @@
# The swarm's message queue as every hive sees it: where it is reached and
# what it is registered as, identical on every host. What the host running it
# decides, and the container itself, are in ./swarm-nats.nix.
{
lib,
config,
...
}:
let
swarmDomain = config.services.hyperhive.swarm.domain;
# Total on a null swarm domain, for the reason ./swarm-bao.nix gives.
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
in
{
options.services.hyperhive.swarm.nats = {
# `enable` moved to `services.hyperhive.deploy.nats.enable` — see
# ./deploy.nix. What the queue IS to every hive: its domain, ports
# and client id.
# ⚠️ Deliberately NO `package` option for the NATS server, anywhere —
# not here and not under `deploy.nats` either, where every other
# service's build now lives. `services.nats` upstream does not expose
# one; it resolves `pkgs.nats-server` itself, so an option would be
# ignored or need an overlay to mean anything, and an option that does
# not control what it names is worse than its absence. Pin the build
# with `nixpkgs.overlays` if you need to. (`deploy.nats.authPackage`
# is a different thing: the callout responder, which IS ours.)
domain = lib.mkOption {
type = lib.types.str;
default = "nats.${domainBase}";
defaultText = lib.literalExpression ''"nats.''${services.hyperhive.swarm.domain}"'';
description = ''
Name every client reaches the queue on, as
`tls://<domain>:<port>`, and the only name its certificate carries.
A **sibling** of the swarm's other service names, for the reason
{option}`services.hyperhive.swarm.bao.domain` gives.
The host running the queue answers it through `gateway.localNames`.
Every other hive resolves it through the operator's upstream DNS,
which is where a multi-host swarm needs a record for it.
Not in {option}`services.hyperhive.swarm.serviceDomains`: that list
is the names a gateway vhost fronts, and nginx fronts nothing here.
'';
};
port = lib.mkOption {
type = lib.types.port;
default = 4222;
description = ''
TCP port the queue listens on. 4222 is upstream's default and
sits outside hyperhive's claimed ranges (dashboard 7000, forge
3000, matrix 8008, every agent in 8100..8999 via FNV-1a hash).
Opened on the bridge interface, for agent containers, and on
`wg-hive` when this host is on the mesh, for other hives. Never
host-wide. TLS only: a client that does not speak it is refused.
Unlike {option}`monitorPort` and {option}`metricsPort`, the
address is not what bounds who may use this port: the queue's
`auth_callout` refuses every client it cannot identify.
'';
};
monitorPort = lib.mkOption {
type = lib.types.port;
default = 8222;
description = ''
Port NATS serves its **monitoring** endpoint on, bound to
loopback.
Not a metrics endpoint: the server has no Prometheus format of its
own. This serves `/varz`, `/connz`, `/routez` as JSON, and the
exporter below is what translates it — which is why enabling the
exporter without this produces a process that starts cleanly and
scrapes nothing.
⚠️ Loopback, and the exporter is the only intended reader. The
endpoint is unauthenticated and `/connz` names every connected
client, so the address it binds is the whole access control. Do
not widen it, and do not put it behind a gateway vhost expecting
that to add one.
'';
};
metricsPort = lib.mkOption {
type = lib.types.port;
default = 7777;
description = ''
Port the Prometheus exporter serves NATS's metrics on, bound to
loopback for the swarm collector to scrape.
⚠️ This and {option}`monitorPort` are two more claims on a port
space every swarm container shares — they run in this container but
`privateNetwork = false`, so a collision with any other hyperhive
service is a runtime coin toss over which process gets the port,
with nothing in any log saying so. Both defaults are upstream's
own (`nats-server` 8222, `prometheus-nats-exporter` 7777) and
neither is claimed elsewhere in this repo, checked when they were
added.
'';
};
clientId = lib.mkOption {
type = lib.types.str;
default = "swarm-nats";
description = ''
OAuth2 client id the queue's authentication path identifies
itself with. Must match the `id` of the corresponding entry in
`services.hyperhive.swarm.authelia.oidc.clients` — which this
module contributes for you when both run on this host.
'';
};
};
}

View file

@ -202,10 +202,6 @@ let
issuerKey = "@ISSUER_PUBKEY@";
});
swarmDomain = config.services.hyperhive.swarm.domain;
# Total on a null swarm domain, for the reason ./swarm-bao.nix gives.
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
baoCfg = config.services.hyperhive.swarm.bao;
baoDeploy = deployCfg.bao;
@ -267,114 +263,11 @@ in
# bao: `gateway.localNames` on this host, the operator's DNS everywhere
# else.
options.services.hyperhive.swarm.nats = {
# `enable` moved to `services.hyperhive.deploy.nats.enable` — see
# ./deploy.nix. What stays here is what the queue IS: its domain,
# ports, accounts and callout wiring.
# ⚠️ Deliberately NO `package` option for the NATS server, anywhere —
# not here and not under `deploy.nats` either, where every other
# service's build now lives. `services.nats` upstream does not expose
# one; it resolves `pkgs.nats-server` itself, so an option would be
# ignored or need an overlay to mean anything, and an option that does
# not control what it names is worse than its absence. Pin the build
# with `nixpkgs.overlays` if you need to. (`deploy.nats.authPackage`
# is a different thing: the callout responder, which IS ours.)
domain = lib.mkOption {
type = lib.types.str;
default = "nats.${domainBase}";
defaultText = lib.literalExpression ''"nats.''${services.hyperhive.swarm.domain}"'';
description = ''
Name every client reaches the queue on, as
`tls://<domain>:<port>`, and the only name its certificate carries.
A **sibling** of the swarm's other service names, for the reason
{option}`services.hyperhive.swarm.bao.domain` gives.
The host running the queue answers it through `gateway.localNames`.
Every other hive resolves it through the operator's upstream DNS,
which is where a multi-host swarm needs a record for it.
Not in {option}`services.hyperhive.swarm.serviceDomains`: that list
is the names a gateway vhost fronts, and nginx fronts nothing here.
'';
};
port = lib.mkOption {
type = lib.types.port;
default = 4222;
description = ''
TCP port the queue listens on. 4222 is upstream's default and
sits outside hyperhive's claimed ranges (dashboard 7000, forge
3000, matrix 8008, every agent in 8100..8999 via FNV-1a hash).
Opened on the bridge interface, for agent containers, and on
`wg-hive` when this host is on the mesh, for other hives. Never
host-wide. TLS only: a client that does not speak it is refused.
Unlike {option}`monitorPort` and {option}`metricsPort`, the
address is not what bounds who may use this port: the queue's
`auth_callout` refuses every client it cannot identify.
'';
};
monitorPort = lib.mkOption {
type = lib.types.port;
default = 8222;
description = ''
Port NATS serves its **monitoring** endpoint on, bound to
loopback.
Not a metrics endpoint: the server has no Prometheus format of its
own. This serves `/varz`, `/connz`, `/routez` as JSON, and the
exporter below is what translates it — which is why enabling the
exporter without this produces a process that starts cleanly and
scrapes nothing.
⚠️ Loopback, and the exporter is the only intended reader. The
endpoint is unauthenticated and `/connz` names every connected
client, so the address it binds is the whole access control. Do
not widen it, and do not put it behind a gateway vhost expecting
that to add one.
'';
};
metricsPort = lib.mkOption {
type = lib.types.port;
default = 7777;
description = ''
Port the Prometheus exporter serves NATS's metrics on, bound to
loopback for the swarm collector to scrape.
⚠️ This and {option}`monitorPort` are two more claims on a port
space every swarm container shares — they run in this container but
`privateNetwork = false`, so a collision with any other hyperhive
service is a runtime coin toss over which process gets the port,
with nothing in any log saying so. Both defaults are upstream's
own (`nats-server` 8222, `prometheus-nats-exporter` 7777) and
neither is claimed elsewhere in this repo, checked when they were
added.
'';
};
clientId = lib.mkOption {
type = lib.types.str;
default = "swarm-nats";
description = ''
OAuth2 client id the queue's authentication path identifies
itself with. Must match the `id` of the corresponding entry in
`services.hyperhive.swarm.authelia.oidc.clients` — which this
module contributes for you when both run on this host.
'';
};
};
# What stays above is what the queue IS to every hive: the ports it answers
# on and the client id it is registered under. What the host running it
# decides is below — which responder build it runs, whether it mints its own
# keypairs, and where the seeds sit. `enable` already lives in ./deploy.nix,
# which also carries the renames.
# What the queue IS to every hive, the ports it answers on and the client id
# it is registered under, is `swarm.nats` in ./swarm-nats-service.nix. What
# the host running it decides is below — which responder build it runs,
# whether it mints its own keypairs, and where the seeds sit. `enable`
# already lives in ./deploy.nix, which also carries the renames.
#
# ⚠️ The two PUBLIC keys move with their seeds rather than staying: peers
# receive the user key over the wire when they connect (docs/swarm/secrets.md),