swarm-nats: give the queue a name, a bao-issued leaf, and require TLS

The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.

- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
  `swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
  every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
  IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
  may only `update` `pki/issue/swarm-nats`, written by
  `swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
  paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
  as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
  the container, restarts nats when it rotates, and re-runs daily.
  It joins glue-bao-readers-policy-order, so it is ordered after its policy
  unit (`after` and `wants`, never `requires`) where the store is on the
  same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
  `allow_non_tls`. `validateConfig` is now off in every mode, because the
  build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
  host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
  `singleHostSwarm`, and the auth responder all dial
  `tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
  to the NATS connection too, so hive-c0re and the controller trust the
  leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
  swarm needs.

module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.

Closes #4626
This commit is contained in:
atlas 2026-09-24 15:43:25 +02:00 • committed by mara
commit 1d261b3fed
17 changed files with 748 additions and 108 deletions

View file

@ -343,11 +343,11 @@ half-configured hive is an eval error rather than one that quietly never
reports. They sit in two namespaces, because two of them are facts about
_this machine_ and one is the swarm's single address:
| option | what to set it to |
| ------------------------------------------------------- | ------------------------------------------------------ |
| `deploy.hive-controller.statusPublish.natsUrl` | where the swarm queue listens, as this hive reaches it |
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` |
| `deploy.hive-controller.statusPublish.clientSecretFile` | path to this hive's client secret, plaintext |
| option | what to set it to |
| ------------------------------------------------------- | --------------------------------------------- |
| `deploy.hive-controller.statusPublish.natsUrl` | `tls://<swarm.nats.domain>:<swarm.nats.port>` |
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` |
| `deploy.hive-controller.statusPublish.clientSecretFile` | path to this hive's client secret, plaintext |
On a host that runs the queue and the IdP itself, all three default to
the local ones and there is nothing to set. Any other hive needs them
@ -356,6 +356,14 @@ not distribute it. Copy `hive-<hiveName>.secret` out of the swarm host's
`deploy.authelia.hostClientSecretDir` with whatever secret management the
deployment already uses.
The queue URL is the same string on every hive. The queue accepts TLS only,
with a certificate for `swarm.nats.domain` (default `nats.<swarm.domain>`) and
no other name or address, so a URL with an IP address or `nats://` fails.
The queue's host resolves the name itself. **A multi-host swarm needs one
upstream DNS record**, `nats.<swarm.domain>` pointing at the queue host's mesh
address, the same contract as `bao.<swarm.domain>`. The queue's port is open
on `wg-hive` when that host is on the mesh.
The identity isn't a choice — a hive authenticates as `hive-<hiveName>`
and publishes under `hiveName`, the same name that keys `swarm.hives`.
@ -382,12 +390,11 @@ with the credential itself and aren't configurable.
| `deploy.hive-controller.queue.agentNatsUrl` | where the queue listens, as an agent **container** reaches it |
| `swarm.statusPublish.tokenEndpoint` | the swarm IdP's `/api/oidc/token` — the same one the hive uses |
On a host that runs the queue, `agentNatsUrl` defaults to
`nats://<network.bridgeIp>:<swarm.nats.port>`, which is the only address that
works from inside a container: the firewall opens the port on the bridge
interface and nowhere else. ⚠️ **Never a loopback address here** — the hive's own
`statusPublish.natsUrl` is loopback and correct, because `hive-c0re` shares the
host's network namespace. An agent doesn't, so `127.0.0.1` reaches the agent.
On a host that runs the queue, `agentNatsUrl` defaults to the same
`tls://<swarm.nats.domain>:<swarm.nats.port>` as the hive's own. Inside a
container the name resolves to the bridge address, where the firewall opens the
port. ⚠️ **Never a loopback address here**: an agent has its own network
namespace, so `127.0.0.1` reaches the agent.
The harness sees four variables, and treats them as all-or-none:
`HIVE_AGENT_NATS_URL` and `HIVE_AGENT_OIDC_TOKEN_ENDPOINT` from the two options

View file

@ -36,13 +36,13 @@ in
natsUrl = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "nats://10.42.0.1:4222";
example = "tls://nats.example.com:4222";
description = ''
Where the swarm queue listens, as this container reaches it.
Set by the generated meta flake from the host's
{option}`services.hyperhive.deploy.hive-controller.queue.agentNatsUrl`,
which is the bridge address rather than a loopback one — inside this
which names the queue rather than a loopback address — inside this
container `127.0.0.1` is the agent itself.
`null` means this hive has not been given the queue's address for its

View file

@ -39,6 +39,10 @@ in
inherit pkgs self nixosSystem;
inherit (pkgs) lib;
};
module-eval-nats-tls = import ./module-eval/nats-tls.nix {
inherit pkgs self nixosSystem;
inherit (pkgs) lib;
};
module-eval-matrix-core = import ./module-eval/matrix-core.nix {
inherit pkgs self nixosSystem;
inherit (pkgs) lib;

View file

@ -29,6 +29,7 @@
./glue-grafana-oidc-client.nix
./glue-matrix-bao-token.nix
./glue-matrix-ctl-bao-identity.nix
./glue-nats-bao-identity.nix
./glue-queue-agent-credential.nix
./glue-secret-publisher-bao-identity.nix
./glue-services-issuer-bao-identity.nix

View file

@ -1,7 +1,7 @@
# Glue: where the store and one of its readers share a host, the reader waits
# for the unit that writes the cert-auth role it logs in with.
#
# ONE PAIRING PER FILE — the store's four reader policy units ← the readers
# ONE PAIRING PER FILE — the store's reader policy units ← the readers
# they grant, and nothing else. Deleting this leaves every reader as it is on a
# host whose store is remote: it may log in before its role exists, and its own
# retries are what carry it past that.
@ -44,6 +44,8 @@ let
swarm-bao-otel-oidc =
deployCfg.swarm-otel.enable
&& havePair baoDeploy.otelOidcClientCertFile baoDeploy.otelOidcClientKeyFile;
# ./swarm-nats.nix: the queue's TLS leaf, not a secret, but the same wait.
swarm-bao-nats-tls = deployCfg.nats.enable;
};
orderAfterPolicy =

View file

@ -233,6 +233,13 @@ in
# credential would undo the separation the role was created for.
[ -s ${pkiDir}/services-issuer.pem ] || ${signLeaf} ${pkiDir} services-issuer \
${lib.escapeShellArg deployCfg.bao.servicesIssuerCommonName} "" clientAuth
# The queue's, which opens `pki/issue/swarm-nats` and nothing else.
# Minted here for the services issuer's reason: it opens the mint, so
# it cannot come out of it. On a queue host without the store this is
# the file an operator copies.
[ -s ${pkiDir}/nats.pem ] || ${signLeaf} ${pkiDir} nats \
${lib.escapeShellArg deployCfg.bao.natsCommonName} "" clientAuth
'';
};
};

View file

@ -0,0 +1,40 @@
# Glue: point the queue's TLS-leaf unit at the bao leaf minted for it.
#
# ONE PAIRING PER FILE — the swarm-nats principal ← bao, and nothing else.
# Deleting this leaves a queue host that takes an operator-provided path to
# that credential, which is what any deployment not minting its own already
# does.
#
# ⚠️ The minting is NOT here. ./glue-bao-tls.nix holds the CA and signs the
# leaf, because the thing that owns a private key owns issuing from it. What
# belongs here is the pairing: which paths `swarm-bao-nats-tls` presents.
#
# ⚠️ Gated on the leaf existing, not on the store being enabled — the same rule
# ./glue-matrix-ctl-bao-identity.nix states: a swarm runs ONE queue, and the
# hive hosting it need not be the hive hosting the store.
#
# Everything is `mkDefault`. An operator naming their own paths wins.
{
lib,
config,
...
}:
let
hyperhiveCfg = config.services.hyperhive;
deployCfg = hyperhiveCfg.deploy;
baoDeploy = deployCfg.bao;
# Where ./glue-bao-tls.nix puts the leaves, derived from the reader's own path
# rather than repeating that file's directory literal: an operator who moves
# the PKI moves both, and the two cannot drift apart.
haveMintedPki = baoDeploy.clientCertFile != null;
pkiDir = if haveMintedPki then builtins.dirOf baoDeploy.clientCertFile else null;
in
{
config = lib.mkIf (hyperhiveCfg.enable && deployCfg.nats.enable && haveMintedPki) {
services.hyperhive.deploy.nats = {
baoClientCertFile = lib.mkDefault "${pkiDir}/nats.pem";
baoClientKeyFile = lib.mkDefault "${pkiDir}/nats-key.pem";
};
};
}

View file

@ -289,10 +289,8 @@ in
# every container. A pair, gated together, because an agent that got one of
# them would report a half-configured queue instead of none.
#
# ⚠️ The url is the bridge address, NOT the loopback one beside it in
# `HIVE_C0RE_NATS_URL` above. Both are correct for their reader: this daemon
# shares the host netns, an agent does not, and inside a container
# `127.0.0.1` is the agent itself.
# ⚠️ By name, like `HIVE_C0RE_NATS_URL` above, and never a loopback
# address: inside a container `127.0.0.1` is the agent itself.
lib.optionalAttrs
(
config.services.hyperhive.deploy.hive-controller.queue.agentNatsUrl != null

View file

@ -88,11 +88,10 @@ in
config.services.hyperhive.swarm = {
ca.autoConfigure = lib.mkDefault cfg.deploy.singleHostSwarm;
# The controller's queue coordinates. Co-location is what makes these
# derivable at all — loopback only reaches the queue when the queue is
# here, and the minted client secret only exists on the host authelia
# ran its first boot on — so they belong to the mode that asserts this
# box is the whole deployment, not to the options' own `default`.
# The controller's queue coordinates. Kept with the mode, not in the
# options' own `default`, because the controller's minted client secret
# only exists on the host authelia ran its first boot on — so they belong
# to the mode that asserts this box is the whole deployment.
#
# Deriving them from `deploy.nats` / `deploy.authelia`
# inside those defaults is the mixing this file exists to prevent: the
@ -110,7 +109,7 @@ in
# needing a queue is the controller's own property in every topology,
# and only the convenience is local.
controller.queue.natsUrl = lib.mkIf cfg.deploy.singleHostSwarm (
lib.mkDefault "nats://127.0.0.1:${toString config.services.hyperhive.swarm.nats.port}"
lib.mkDefault "tls://${config.services.hyperhive.swarm.nats.domain}:${toString config.services.hyperhive.swarm.nats.port}"
);
};

View file

@ -384,6 +384,17 @@ let
}
'';
# The queue's principal: one `update` on its own role, the shape of the
# services issuer's grant above, narrowed to one name by the role itself.
natsPolicyName = "swarm-nats";
natsCn = baoDeploy.natsCommonName;
natsPkiRoleName = baoDeploy.natsPkiRoleName;
natsPolicyText = ''
path "${servicesPkiMountPath}/issue/${natsPkiRoleName}" {
capabilities = ["update"]
}
'';
# ── the four principals that used to share the hive's own leaf ─────────────
#
# 🩸 Each of the four reads exactly ONE path in the store, and until this
@ -1074,6 +1085,35 @@ in
'';
};
natsCommonName = lib.mkOption {
type = lib.types.str;
default = "swarm-nats";
description = ''
Subject the store's `swarm-nats` cert-auth role accepts: the identity
the queue's host presents to ask for the queue's TLS leaf.
Its own principal rather than
{option}`services.hyperhive.deploy.bao.servicesIssuerCommonName`,
whose role issues for every name in
{option}`services.hyperhive.swarm.serviceDomains`. This one may issue
from {option}`services.hyperhive.deploy.bao.natsPkiRoleName` alone,
and that role only for {option}`services.hyperhive.swarm.nats.domain`.
⚠️ Reserved as a hive name by ./swarm.nix, like its siblings.
'';
};
natsPkiRoleName = lib.mkOption {
type = lib.types.str;
default = "swarm-nats";
description = ''
Role on {option}`services.hyperhive.deploy.bao.servicesPkiMountPath`
the queue's TLS leaf is issued through. It allows exactly
{option}`services.hyperhive.swarm.nats.domain`: no subdomains, IPs or
localhost.
'';
};
matrixCtlHiveName = lib.mkOption {
type = lib.types.str;
default = toString hyperhiveCfg.hiveName;
@ -1518,6 +1558,7 @@ in
"swarm-bao-grafana-oidc-policy"
"swarm-bao-otel-oidc-policy"
"swarm-bao-services-issuer-policy"
"swarm-bao-nats-tls-policy"
];
# 🚫 No `swarm.otel.scrapeTargets.bao` entry any more, and its absence is
@ -2216,6 +2257,70 @@ in
'';
};
# A FIFTH sibling: the queue's issuing role, its policy and its login
# role, together. The `pki` role lives here rather than beside
# `swarm-services` in the controller's unit because it belongs to this
# principal; that unit creates the mount, hence the `after`.
#
# The role narrows exactly as `swarm-services` does, to one name. It is
# the queue's domain alone, since the same name reaches it from every
# hive; no IP SANs, since nothing dials an address.
systemd.services.swarm-bao-nats-tls-policy = lib.mkIf haveBootstrapToken {
description = "write the swarm queue's pki role, bao policy and cert-auth role";
after = [
"container@${cfg.machine}.service"
"swarm-bao-controller-policy.service"
];
wantedBy = [ "multi-user.target" ];
path = [
baoCli
pkgs.coreutils
];
unitConfig.ConditionPathExists = baoDeploy.bootstrapTokenFile;
# Same unseal wait as its siblings above.
startLimitBurst = 2880;
startLimitIntervalSec = 90000;
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
Restart = "on-failure";
RestartSec = 30;
};
script = ''
set -euo pipefail
BAO_TOKEN="$(cat ${lib.escapeShellArg baoDeploy.bootstrapTokenFile})"
export BAO_TOKEN
bao write ${lib.escapeShellArg "${servicesPkiMountPath}/roles/${natsPkiRoleName}"} \
allowed_domains=${lib.escapeShellArg hyperhiveCfg.swarm.nats.domain} \
allow_bare_domains=true \
allow_subdomains=false \
allow_glob_domains=false \
allow_localhost=false \
allow_any_name=false \
allow_ip_sans=false \
enforce_hostnames=true \
server_flag=true \
client_flag=false \
key_type=rsa \
key_bits=4096 \
ttl=${servicesPkiLeafTtl} \
max_ttl=${servicesPkiLeafTtl}
printf '%s' ${lib.escapeShellArg natsPolicyText} |
bao policy write ${lib.escapeShellArg natsPolicyName} -
''
+ lib.optionalString (baoDeploy.clientCaFile != null) ''
bao write auth/cert/certs/${lib.escapeShellArg natsPolicyName} \
certificate=@${tlsDir}/client-ca.pem \
allowed_common_names=${lib.escapeShellArg natsCn} \
token_policies=${lib.escapeShellArg natsPolicyName} \
display_name=${lib.escapeShellArg natsCn}
'';
};
# The CA bind source is written at runtime by a host unit, so the
# container has to start after it — otherwise nspawn sets up a mount
# over a file that does not exist yet.

View file

@ -365,18 +365,17 @@ in
natsUrl = lib.mkOption {
type = lib.types.str;
default = "";
example = "nats://queue.example.com:4222";
example = "tls://nats.example.com:4222";
description = ''
Where the controller reaches the swarm queue.
Empty means unset, which the assertion below refuses — a
controller with no queue is not a lighter controller.
`singleHostSwarm` fills this in with loopback, because
that address is only correct when the queue is on this host:
its container shares the host netns. That derivation lives with
the mode rather than here, so this option describes itself
rather than a deployment shape.
`singleHostSwarm` fills this in with
`tls://<swarm.nats.domain>:<port>`. That derivation lives with the
mode rather than here, so this option describes itself rather than
a deployment shape.
'';
};
@ -765,9 +764,8 @@ in
message = ''
services.hyperhive.swarm.controller.queue.natsUrl is unset.
It defaults to loopback only when this host also runs the queue
(`services.hyperhive.deploy.nats.enable`). A controller on its
own host has to be told where the queue is.
`singleHostSwarm` fills it in with the queue's name. A controller
in any other deployment has to be told the URL.
'';
}
{

View file

@ -201,16 +201,39 @@ let
userKey = "@USER_PUBKEY@";
issuerKey = "@ISSUER_PUBKEY@";
});
swarmDomain = config.services.hyperhive.swarm.domain;
# Total on a null swarm domain, for the reason ./swarm-bao.nix gives.
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
baoCfg = config.services.hyperhive.swarm.bao;
baoDeploy = deployCfg.bao;
# The queue's TLS leaf, bound read-only at the same path on both sides, the
# shape ./swarm-bao.nix's `tlsDir` uses. The key reaches the server through
# `LoadCredential`, as openbao's does: it is root-owned 0600 here, and the
# server runs as `nats`.
tlsDir = "/var/lib/swarm-nats-tls";
tlsCertPath = "${tlsDir}/cert.pem";
tlsKeyPath = "${tlsDir}/key.pem";
tlsKeyCredential = "tls-key";
tlsKeyCredentialPath = "/run/credentials/nats.service/${tlsKeyCredential}";
# Re-issue once the leaf is past half of the 720h `pki/roles/swarm-nats`
# grants it (./swarm-bao.nix), so the daily timer below has two weeks of
# retries before it lapses.
leafRenewSeconds = 360 * 3600;
in
{
# The swarm's message queue: one NATS server, reached by every hive.
# The swarm's message queue: one NATS server, reached by every hive at
# `tls://<swarm.nats.domain>:<port>`.
#
# ⚠️ There is deliberately NO gateway vhost here, and this is the first
# swarm service where that is true — the next reader will go looking for
# one. NATS speaks its own TCP protocol rather than HTTP, so nginx
# cannot front it the way it fronts the forge, matrix and authelia.
# Cross-hive reach is the wireguard mesh; `gateway.localNames` and the
# per-service vhost pattern do not apply.
# ⚠️ There is deliberately NO gateway vhost here. NATS speaks its own TCP
# protocol rather than HTTP, so nginx cannot front it the way it fronts the
# forge, matrix and authelia, and the server terminates TLS itself. The
# name is the half of the sibling pattern that applies, the way it is for
# bao: `gateway.localNames` on this host, the operator's DNS everywhere
# else.
options.services.hyperhive.swarm.nats = {
# `enable` moved to `services.hyperhive.deploy.nats.enable` — see
@ -226,6 +249,25 @@ in
# with `nixpkgs.overlays` if you need to. (`deploy.nats.authPackage`
# is a different thing: the callout responder, which IS ours.)
domain = lib.mkOption {
type = lib.types.str;
default = "nats.${domainBase}";
defaultText = lib.literalExpression ''"nats.''${services.hyperhive.swarm.domain}"'';
description = ''
Name every client reaches the queue on, as
`tls://<domain>:<port>`, and the only name its certificate carries.
A **sibling** of the swarm's other service names, for the reason
{option}`services.hyperhive.swarm.bao.domain` gives.
The host running the queue answers it through `gateway.localNames`.
Every other hive resolves it through the operator's upstream DNS,
which is where a multi-host swarm needs a record for it.
Not in {option}`services.hyperhive.swarm.serviceDomains`: that list
is the names a gateway vhost fronts, and nginx fronts nothing here.
'';
};
port = lib.mkOption {
type = lib.types.port;
default = 4222;
@ -234,12 +276,9 @@ in
sits outside hyperhive's claimed ranges (dashboard 7000, forge
3000, matrix 8008, every agent in 8100..8999 via FNV-1a hash).
Contributed to
`services.hyperhive.network.exposeHostPorts`, which opens it on
the bridge interface only — so it is reachable from agent
containers and not from the outside world. An agent connects to
`nats://<bridgeIp>:<port>`; loopback inside a container is the
agent itself, not this host.
Opened on the bridge interface, for agent containers, and on
`wg-hive` when this host is on the mesh, for other hives. Never
host-wide. TLS only: a client that does not speak it is refused.
Unlike {option}`monitorPort` and {option}`metricsPort`, the
address is not what bounds who may use this port: the queue's
@ -429,6 +468,36 @@ in
`calloutUserSeedFile`.
'';
};
baoClientCertFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-bao-pki/nats.pem";
description = ''
Client certificate `swarm-bao-nats-tls` presents to the swarm's
secret store when it asks for the queue's TLS leaf. Its subject must
be {option}`services.hyperhive.deploy.bao.natsCommonName`: cert auth
matches on the CN, and that role's policy may issue from
`pki/issue/swarm-nats` and nothing else.
No default. ./glue-nats-bao-identity.nix points it at the leaf
./glue-bao-tls.nix mints, where this host mints one. On a queue host
without the store, it is a file the operator copies from the store's
host.
A path, never a value.
'';
};
baoClientKeyFile = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "/var/lib/swarm-bao-pki/nats-key.pem";
description = ''
Private key for {option}`services.hyperhive.deploy.nats.baoClientCertFile`.
A path, never a value.
'';
};
};
config = lib.mkIf deployCfg.nats.enable {
@ -437,12 +506,17 @@ in
services.hyperhive.swarm.otel.journaldUnits = [
"nats"
"swarm-nats-auth"
"swarm-bao-nats-tls"
];
# Resolver only: NATS speaks its own protocol, so nginx fronts
# nothing here — but the auth responder introspects authelia by name.
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# The name every client dials, answered with the bridge IP on this host.
# Other hives resolve it through the operator's DNS, as they do bao's.
services.hyperhive.gateway.localNames = [ cfg.domain ];
assertions = [
{
# Fail at EVAL, not at boot: a queue that comes up unable to
@ -575,18 +649,27 @@ in
# by the firewall and looks exactly like every other NATS failure,
# a timeout.
#
# ⚠️ Bridge interface only — never the world. What makes it safe to
# open at all is that the queue is fail-closed: `auth_callout` admits
# nobody until the responder above answers for them, so an agent that
# reaches this port still has to present a token authelia vouches for.
# ⚠️ Never the world. What makes it safe to open at all is that the
# queue is fail-closed: `auth_callout` admits nobody until the responder
# above answers for them, so an agent that reaches this port still has to
# present a token authelia vouches for.
#
# An agent connects to `nats://<bridgeIp>:<port>`, NOT to
# `nats://127.0.0.1:<port>` — inside a container loopback is the
# *agent*. Every url in this repo today is the loopback one and each
# is correct for its reader, because those readers share the host
# netns; an agent does not.
# An agent reaches it by name, which dnsmasq answers with the bridge IP.
services.hyperhive.network.exposeHostPorts = [ cfg.port ];
# The other hives' way in, interface-scoped like
# ./swarm-snapshot-store.nix's receiver. The firewall is default-deny
# before a packet reaches the listener, so without this nothing in-tree
# lets another hive reach the queue at all.
networking.firewall.interfaces.wg-hive.allowedTCPPorts = lib.mkIf deployCfg.wireguard.enable [
cfg.port
];
# The bind source has to exist before the container starts, leaf or not:
# nixos-container refuses to start over a missing one, which would take
# the whole queue down rather than only its TLS.
systemd.tmpfiles.rules = [ "d ${tlsDir} 0755 root root -" ];
containers.swarm-nats = {
autoStart = true;
ephemeral = false;
@ -605,9 +688,14 @@ in
# Not agent containers, though: they have a netns of their own and
# the bridge firewall does not open this port.
privateNetwork = false;
# Binds only the public trust bundle, read-only. Empty when the gateway
# is not self-signed, so the whole trust path drops out cleanly.
bindMounts = caTrust.bindMount;
# The public trust bundle (empty when the gateway is not self-signed)
# and the queue's own TLS leaf, both read-only.
bindMounts = caTrust.bindMount // {
${tlsDir} = {
hostPath = tlsDir;
isReadOnly = true;
};
};
config =
{ ... }:
{
@ -652,13 +740,14 @@ in
serverName = "swarm-nats";
port = cfg.port;
# In auto mode the keys are empty until the generator runs,
# and `nats-server -t` rejects that ("Expected callout user to
# be a valid public account nkey, got \"\""), so leaving this
# on fails the BUILD of every all-local hive. Upstream's own
# description names the case: disable it when the config
# includes other files. The check moves to server start.
validateConfig = !deployCfg.nats.autoGenerateCallout;
# Off in every mode now. The `tls` block below names a leaf that
# exists only at runtime, and `nats-server -t` loads it, so the
# build-time check fails on every hive. In auto mode it failed
# already, on callout keys that are empty until the generator
# runs. Upstream's own description names the case: disable it
# when the config includes other files. The check moves to
# server start.
validateConfig = false;
settings =
calloutBlocks {
@ -687,9 +776,28 @@ in
# ⚠️ Loopback: unauthenticated, and `/connz` lists every
# connected client.
http = "127.0.0.1:${toString cfg.monitorPort}";
# TLS required on the client port: a server with a `tls`
# block advertises `tls_required` and refuses a client that
# does not upgrade. No `allow_non_tls`, which would keep the
# plaintext path open across the mesh. The 2s timeout is for
# a handshake that crosses the mesh; upstream's 0.5s is sized
# for a LAN.
tls = {
cert_file = tlsCertPath;
key_file = tlsKeyCredentialPath;
timeout = 2;
};
};
};
# The key as the `nats` user can read it; see `tlsDir`. Resolved at
# unit start only, which is why a renewed leaf restarts the server
# rather than reloading it.
systemd.services.nats.serviceConfig.LoadCredential = [
"${tlsKeyCredential}:${tlsKeyPath}"
];
# The translation layer. NATS has no Prometheus format of its
# own, so this reads the JSON monitoring endpoint above and
# re-serves it in the format the collector scrapes.
@ -754,7 +862,11 @@ in
serviceConfig = {
ExecStart = lib.concatStringsSep " " [
"${deployCfg.nats.authPackage}/bin/swarm-nats-auth"
"--nats-url nats://127.0.0.1:${toString cfg.port}"
# By name, like every other client: the leaf carries the name
# and no IP, so a loopback address fails verification. The
# container's resolver goes through the bridge, where dnsmasq
# answers it.
"--nats-url tls://${cfg.domain}:${toString cfg.port}"
"--user-seed-file \${CREDENTIALS_DIRECTORY}/callout-user.seed"
"--issuer-seed-file \${CREDENTIALS_DIRECTORY}/issuer.seed"
"--client-secret-file \${CREDENTIALS_DIRECTORY}/oidc-client.secret"
@ -966,5 +1078,148 @@ in
${lib.escapeShellArg hostRuntimeDir}/nats.conf
'';
};
# The queue's TLS leaf, issued by the secret store's `pki/roles/swarm-nats`
# under this host's `swarm-nats` identity. Same shape as ./hive-tls.nix's
# `swarm-services-cert`, whose comments carry the reasoning for the login,
# the single `issue` call split with `jq`, and the unseal-sized retry.
#
# Before the container, so a normal boot starts the server with its leaf.
# A store that comes up later fails this run; the retry lands the leaf and
# restarts the server, which until then refuses to start (TLS required,
# and its key credential is missing).
systemd.services.swarm-bao-nats-tls = {
description = "Issue the swarm queue's TLS leaf from the secret store's PKI";
wantedBy = [
"multi-user.target"
"container@${machine}.service"
];
before = [ "container@${machine}.service" ];
# Both absent where the store runs elsewhere, and ignored there; see
# `swarm-services-cert`. The ordering after this leaf's policy unit is
# ./glue-bao-readers-policy-order.nix's, because it holds only where the
# store is here too.
after = [
"container@${baoCfg.machine}.service"
"swarm-bao-pki.service"
];
wants = [ "container@${baoCfg.machine}.service" ];
path = [
baoDeploy.package
pkgs.jq
pkgs.openssl
pkgs.coreutils
pkgs.gnugrep
pkgs.systemd
];
startLimitBurst = 2880;
startLimitIntervalSec = 90000;
# No `RemainAfterExit`: the timer below starts this again, and starting
# an active unit is a no-op.
serviceConfig = {
Type = "oneshot";
UMask = "0077";
SyslogIdentifier = "swarm-bao-nats-tls";
Restart = "on-failure";
RestartSec = 30;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
}
// lib.optionalAttrs (deployCfg.nats.baoClientCertFile != null) {
BAO_CLIENT_CERT = deployCfg.nats.baoClientCertFile;
}
// lib.optionalAttrs (deployCfg.nats.baoClientKeyFile != null) {
BAO_CLIENT_KEY = deployCfg.nats.baoClientKeyFile;
}
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
BAO_CACERT = baoDeploy.serverCaFile;
};
script = ''
set -euo pipefail
d=${lib.escapeShellArg tlsDir}
name=${lib.escapeShellArg cfg.domain}
install -d -m 0755 "$d"
# Missing, past half its window, or naming something other than the
# configured domain.
reissue=0
{ [ -s "$d/cert.pem" ] && [ -s "$d/key.pem" ]; } || reissue=1
openssl x509 -in "$d/cert.pem" -noout -checkend ${toString leafRenewSeconds} >/dev/null 2>&1 || reissue=1
openssl x509 -in "$d/cert.pem" -noout -checkhost "$name" 2>/dev/null | grep -q ' does match ' || reissue=1
if [ "$reissue" = 0 ]; then
echo "queue leaf valid for $name — leaving it alone"
exit 0
fi
${
if deployCfg.nats.baoClientCertFile == null || deployCfg.nats.baoClientKeyFile == null then
''
echo "no bao client certificate configured for the queue, so its TLS leaf" >&2
echo "cannot be requested from the store." >&2
echo "Set services.hyperhive.deploy.nats.baoClient{Cert,Key}File" >&2
echo "to a leaf the store's CA signed with CN=${baoDeploy.natsCommonName}." >&2
exit 1''
else
""
}
err="$(mktemp)"
trap 'rm -f "$err"' EXIT
if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then
echo "could not log in to the swarm secret store with the queue's certificate." >&2
cat "$err" >&2
exit 1
fi
export BAO_TOKEN
echo "requesting the queue's TLS leaf for $name from the store"
resp="$(mktemp "$d/issue.json.XXXXXX")"
trap 'rm -f "$err" "$resp"' EXIT
if ! bao write -format=json \
${lib.escapeShellArg "${baoDeploy.servicesPkiMountPath}/issue/${baoDeploy.natsPkiRoleName}"} \
common_name="$name" > "$resp" 2>"$err"; then
echo "the store refused to issue the queue's certificate." >&2
cat "$err" >&2
exit 1
fi
jq -r '.data.private_key // empty' < "$resp" > "$d/key.pem.new"
jq -r '.data.certificate // empty' < "$resp" > "$d/cert.pem.new"
for f in "$d/key.pem.new" "$d/cert.pem.new"; do
if [ ! -s "$f" ]; then
echo "the store's response was missing a field: $f is empty" >&2
exit 1
fi
done
chmod 0600 "$d/key.pem.new"
chmod 0644 "$d/cert.pem.new"
mv -f "$d/key.pem.new" "$d/key.pem"
mv -f "$d/cert.pem.new" "$d/cert.pem"
# Renewal, or the late-store retry. On a normal boot the container is
# not up yet and starts the server with this leaf itself. A restart,
# not a reload: the key arrives by `LoadCredential`, which systemd
# resolves at start only. `reset-failed` because a server that found
# no key may have hit its start limit.
if systemctl is-active --quiet ${lib.escapeShellArg "container@${machine}.service"}; then
echo "queue leaf rotated — restarting the server in ${machine}"
systemctl --machine=${lib.escapeShellArg machine} reset-failed nats.service
systemctl --machine=${lib.escapeShellArg machine} restart nats.service
fi
'';
};
systemd.timers.swarm-bao-nats-tls = {
description = "Daily renewal check for the swarm queue's TLS leaf";
wantedBy = [ "timers.target" ];
timerConfig = {
OnCalendar = "daily";
Persistent = true;
};
};
};
}

View file

@ -475,10 +475,8 @@ in
collector is a plain producer against this name exactly like every
other client of a swarm service, resolved locally by dnsmasq on a
co-located host and over the real network otherwise. There is no
separate loopback-vs-remote knob to get wrong: `swarm-nats` is the
deliberate exception to this pattern (its cross-hive reach is the
wireguard mesh, not the gateway), everything else in this swarm
addresses its siblings by name.
separate loopback-vs-remote knob to get wrong: everything in this
swarm addresses its siblings by name.
'';
};

View file

@ -50,6 +50,7 @@ let
deployCfg.bao.grafanaOidcCommonName
deployCfg.bao.otelOidcCommonName
deployCfg.bao.servicesIssuerCommonName
deployCfg.bao.natsCommonName
]
# The two per-hive readers' subjects, spelled out per hive rather than as the
# prefix. The prefix alone would reserve the wrong string: the role for hive
@ -542,17 +543,17 @@ in
options.services.hyperhive.deploy.hive-controller.statusPublish = {
natsUrl = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = if queueLocal then "nats://127.0.0.1:${toString swarmCfg.nats.port}" else null;
defaultText = lib.literalExpression ''"nats://127.0.0.1:''${swarm.nats.port}" when this host runs the queue and the IdP, else null'';
example = "nats://10.100.0.1:4222";
default =
if queueLocal then "tls://${swarmCfg.nats.domain}:${toString swarmCfg.nats.port}" else null;
defaultText = lib.literalExpression ''"tls://''${swarm.nats.domain}:''${swarm.nats.port}" when this host runs the queue and the IdP, else null'';
example = "tls://nats.example.com:4222";
description = ''
Where the swarm queue listens, as seen from *this* hive.
Defaults to loopback when this host runs the queue container
itself (it shares the host netns, so loopback is correct there
and is not the "localhost means the wrong thing" trap that
applies inside agent containers). A hive that is not the swarm
host has to name the swarm's mesh address.
Defaults to the queue's name when this host runs the queue and the
IdP. A hive that is not the swarm host sets the same URL: the name
resolves through the operator's DNS there. TLS only, and by name,
since the queue's certificate carries the name and no address.
Null disables status publishing: this hive computes its own
readiness as always, and simply offers it to nobody. The swarm
@ -589,9 +590,8 @@ in
};
# The same queue, reached from one layer further in. An agent container has
# its own network namespace, so it needs an address of its own rather than
# the one beside it in `statusPublish.natsUrl` — sharing that option would
# hand every agent a loopback address that resolves to the agent.
# its own network namespace and resolver, so this is its own option, even
# though by name it defaults to the same URL as `statusPublish.natsUrl`.
#
# Only the address lives here. The credential does not: it is published per
# hive and read out of the store by ./glue-queue-agent-credential.nix, which
@ -599,18 +599,18 @@ in
options.services.hyperhive.deploy.hive-controller.queue.agentNatsUrl = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default =
if queueLocal then "nats://${cfg.network.bridgeIp}:${toString swarmCfg.nats.port}" else null;
defaultText = lib.literalExpression ''"nats://''${network.bridgeIp}:''${swarm.nats.port}" when this host runs the queue and the IdP, else null'';
example = "nats://10.100.0.1:4222";
if queueLocal then "tls://${swarmCfg.nats.domain}:${toString swarmCfg.nats.port}" else null;
defaultText = lib.literalExpression ''"tls://''${swarm.nats.domain}:''${swarm.nats.port}" when this host runs the queue and the IdP, else null'';
example = "tls://nats.example.com:4222";
description = ''
Where the swarm queue listens, as an agent *container* on this host
reaches it.
Defaults to the bridge address when this host runs the queue, because
that is the only address it is reachable at from a container:
{option}`services.hyperhive.swarm.nats.port` is opened on the bridge
interface alone. ⚠️ Never a loopback address — inside an agent's network
namespace `127.0.0.1` is the agent, not this host.
Defaults to the queue's name when this host runs the queue. The agent
resolves it through the bridge, where dnsmasq answers it with the
bridge address. ⚠️ Never a loopback address: inside an agent's network
namespace `127.0.0.1` is the agent, not this host, and the queue's
certificate names no address anyway.
Null means this hive's agents have not been given the queue's address.
Together with

View file

@ -123,19 +123,17 @@ let
== toString allLocal.services.hyperhive.deploy.hive-controller.queue.agentCredentialDir;
}
{
# The one address in this file that must NOT be loopback. Both spellings
# sit in the same unit's environment and are correct for their own
# reader: hive-c0re shares the host netns, an agent container does not,
# so a copy-paste between them reaches the agent itself and the symptom
# is a connect that hangs.
name = "the agents' queue address is the bridge, not the loopback one the hive itself uses";
# Inside an agent container loopback is the agent itself, so the agents'
# address must never be one. By name it is the same string the hive
# dials, which ./nats-tls.nix pins for every client.
name = "the agents' queue address is the queue's name, never loopback";
ok =
let
e = allLocal.systemd.services.hive-c0re.environment;
in
e.HIVE_AGENT_NATS_URL == "nats://${allLocal.services.hyperhive.network.bridgeIp}:4222"
e.HIVE_AGENT_NATS_URL == "tls://nats.t.local:4222"
&& !(lib.hasInfix "127.0.0.1" e.HIVE_AGENT_NATS_URL)
&& e.HIVE_AGENT_NATS_URL != e.HIVE_C0RE_NATS_URL;
&& e.HIVE_AGENT_NATS_URL == e.HIVE_C0RE_NATS_URL;
}
{
# The agents mint against the swarm's IdP, the same endpoint the hive's

View file

@ -0,0 +1,217 @@
# `checks.module-eval-nats-tls` — see ./lib.nix for the shared rationale (why
# this suite exists, naming convention, "evaluates not executes").
#
# The queue's name, its bao-issued leaf, and the clients that dial it.
{
pkgs,
lib,
self,
nixosSystem,
}:
let
inherit
(import ./lib.nix {
inherit
pkgs
lib
self
nixosSystem
;
})
hive
runGroup
bridgePorts
;
natsName = "nats.t.local";
natsUrl = "tls://${natsName}:4222";
# Every service on one host, with a bootstrap token so the store's granting
# units render. The queue, the store and every in-tree client of the queue
# are all here, so the scan below reads each of them.
allLocal = hive {
deploy.singleHostSwarm = true;
deploy.bao.bootstrapTokenFile = "/run/secrets/bao-bootstrap.token";
};
# The same host on the mesh.
allLocalMesh = hive {
deploy.singleHostSwarm = true;
deploy.wireguard.enable = true;
deploy.wireguard.address = "10.100.0.1/24";
};
# The queue on a host whose store is elsewhere: no local policy unit.
queueNoStore = hive {
deploy.nats.enable = true;
deploy.nats.autoGenerateCallout = true;
};
policyScript = allLocal.systemd.services.swarm-bao-nats-tls-policy.script;
leafUnit = allLocal.systemd.services.swarm-bao-nats-tls;
natsContainer = allLocal.containers.swarm-nats.config;
natsTls = natsContainer.services.nats.settings.tls;
# Every `(nats|tls)://…` in a string.
urlsIn =
s: map builtins.head (builtins.filter builtins.isList (builtins.split "((nats|tls)://[^ '\"]+)" s));
# Every in-tree queue client, found rather than listed. The Rust client reads
# its address from `<PREFIX>_NATS_URL` (`swarm_queue_client::QueueConfig::
# from_env`), so any unit on the host or in a container that is handed one
# carries a variable of that shape. The responder takes a flag instead.
# Keyed by where each came from, so the control below can name them.
clientUrls =
machine:
let
fromUnits =
where: services:
lib.concatLists (
lib.mapAttrsToList (
unit: s:
lib.mapAttrsToList (var: v: {
name = "${where}/${unit}/${var}";
value = v;
}) (lib.filterAttrs (var: v: lib.hasSuffix "_NATS_URL" var && v != null) (s.environment or { }))
) services
);
containerUnits = lib.concatLists (
lib.mapAttrsToList (c: cc: fromUnits c cc.config.systemd.services) machine.containers
);
responder =
map
(u: {
name = "swarm-nats/swarm-nats-auth/--nats-url";
value = u;
})
(
urlsIn machine.containers.swarm-nats.config.systemd.services.swarm-nats-auth.serviceConfig.ExecStart
);
in
lib.listToAttrs (fromUnits "host" machine.systemd.services ++ containerUnits ++ responder);
scanned = clientUrls allLocal;
cases = [
{
# Control first: a scan that found nothing would pass the next case
# vacuously. These are the four clients in the tree today.
name = "the client scan finds hive-c0re, the agents, the controller and the responder";
ok = lib.all (k: scanned ? ${k}) [
"host/hive-c0re/HIVE_C0RE_NATS_URL"
"host/hive-c0re/HIVE_AGENT_NATS_URL"
"host/swarm-controller/SWARM_CONTROLLER_NATS_URL"
"swarm-nats/swarm-nats-auth/--nats-url"
];
}
{
# The server requires TLS and its leaf carries the name alone, so a
# `nats://` URL or an address is a client that cannot connect. Every one
# found, not the four above: a client added later is held to it too.
name = "every in-tree queue client dials tls://<the queue's name>:4222";
ok = lib.all (u: u == natsUrl) (lib.attrValues scanned);
}
{
# The option defaults the scan reads through, so a client that stops
# reading them does not also escape the property above.
name = "the queue URL options default to the name on the queue's host";
ok =
let
d = allLocal.services.hyperhive.deploy;
in
d.hive-controller.statusPublish.natsUrl == natsUrl
&& d.hive-controller.queue.agentNatsUrl == natsUrl
&& allLocal.services.hyperhive.swarm.controller.queue.natsUrl == natsUrl;
}
{
name = "the queue's name is served by this host's resolver, at the bridge address";
ok =
lib.elem natsName allLocal.services.hyperhive.gateway.localNames
&& lib.elem "/${natsName}/${allLocal.services.hyperhive.network.bridgeIp}" allLocal.services.dnsmasq.settings.address;
}
{
name = "the queue's pki role issues for its name alone";
ok = lib.all (arg: lib.hasInfix arg policyScript) [
"roles/swarm-nats \\"
"allowed_domains=${lib.escapeShellArg natsName} \\"
"allow_bare_domains=true \\"
"allow_subdomains=false \\"
"allow_glob_domains=false \\"
"allow_localhost=false \\"
"allow_any_name=false \\"
"allow_ip_sans=false \\"
"server_flag=true \\"
"client_flag=false \\"
];
}
{
# Every `path` the policy names, not a search for the one expected: a
# second grant added later fails here.
name = "the queue's policy grants pki/issue/swarm-nats and nothing else";
ok =
let
paths = map builtins.head (
builtins.filter builtins.isList (builtins.split "path \"([^\"]*)\"" policyScript)
);
in
paths == [ "pki/issue/swarm-nats" ]
&& lib.hasInfix ''capabilities = ["update"]'' policyScript
&& lib.hasInfix "allowed_common_names=swarm-nats" policyScript
&& lib.hasInfix "token_policies=swarm-nats" policyScript;
}
{
# Mint to consume: the leaf the unit writes is the one the server reads,
# and the login leaf glue-bao-tls signs is the one the unit presents.
name = "the server serves the leaf the host unit issues, and the unit logs in as swarm-nats";
ok =
natsTls.cert_file == "/var/lib/swarm-nats-tls/cert.pem"
&& natsTls.key_file == "/run/credentials/nats.service/tls-key"
&& lib.elem "tls-key:/var/lib/swarm-nats-tls/key.pem" natsContainer.systemd.services.nats.serviceConfig.LoadCredential
&& allLocal.containers.swarm-nats.bindMounts ? "/var/lib/swarm-nats-tls"
&& lib.hasInfix "d=/var/lib/swarm-nats-tls" leafUnit.script
&& lib.hasInfix "pki/issue/swarm-nats" leafUnit.script
&& leafUnit.environment.BAO_CLIENT_CERT == "/var/lib/swarm-bao-pki/nats.pem"
&& lib.hasInfix "[ -s /var/lib/swarm-bao-pki/nats.pem ]" allLocal.systemd.services.swarm-bao-pki.script
&& lib.hasInfix "swarm-nats \"\" clientAuth" allLocal.systemd.services.swarm-bao-pki.script;
}
{
name = "the server requires TLS: no allow_non_tls";
ok = !(natsContainer.services.nats.settings ? allow_non_tls);
}
{
name = "4222 is open on wg-hive when this host is on the mesh, and never host-wide";
ok =
lib.elem 4222 allLocalMesh.networking.firewall.interfaces.wg-hive.allowedTCPPorts
&& !(lib.elem 4222 allLocalMesh.networking.firewall.allowedTCPPorts)
&& lib.elem 4222 (bridgePorts allLocalMesh);
}
{
name = "4222 is not opened on wg-hive when this host is not on the mesh";
ok =
!(lib.elem 4222
(allLocal.networking.firewall.interfaces.wg-hive or { allowedTCPPorts = [ ]; }).allowedTCPPorts
)
&& !(lib.elem 4222 allLocal.networking.firewall.allowedTCPPorts);
}
{
# Ordering, never a requirement: the policy unit skips once the bootstrap
# token is gone, and a skipped unit counts as done.
name = "the leaf unit is ordered after its policy unit, with no requires";
ok =
let
p = "swarm-bao-nats-tls-policy.service";
in
lib.elem p leafUnit.after && lib.elem p leafUnit.wants && !(lib.elem p (leafUnit.requires or [ ]));
}
{
name = "a queue host whose store is elsewhere orders its leaf unit after no policy unit";
ok =
let
u = queueNoStore.systemd.services.swarm-bao-nats-tls;
in
!(lib.elem "swarm-bao-nats-tls-policy.service" u.after)
&& !(lib.elem "swarm-bao-nats-tls-policy.service" u.wants);
}
];
in
runGroup "nats-tls" cases

View file

@ -323,7 +323,7 @@ const MAX_RECONNECT_DELAY: std::time::Duration = std::time::Duration::from_mins(
/// this process reads no file it was not pointed at.
#[derive(Debug, Clone)]
pub struct QueueConfig {
/// `nats://host:port` for the swarm queue.
/// `tls://<name>:<port>` for the swarm queue.
pub url: String,
/// Authelia's token endpoint, e.g. `https://auth.<swarm>/api/oidc/token`.
pub token_endpoint: String,
@ -335,8 +335,8 @@ pub struct QueueConfig {
/// read here, and putting it in the environment would publish it to
/// anything that can read `/proc/<pid>/environ`.
pub client_secret_file: PathBuf,
/// Extra trust anchor for the token endpoint, when it is not signed by
/// a publicly-trusted CA.
/// Extra trust anchor for the token endpoint and the queue itself, when
/// they are not signed by a publicly-trusted CA.
///
/// Optional, and deliberately NOT part of the all-or-none group below: a
/// swarm fronted by a public certificate needs no extra anchor, and
@ -673,6 +673,7 @@ pub async fn connect(cfg: QueueConfig) -> Result<async_nats::Client, Error> {
// `async-nats` do what it already does well — back off and try again.
let http = build_http_client(&cfg)?;
let url = cfg.url.clone();
let ca_file = cfg.ca_file.clone();
// Shared across every invocation of the callback below, which is the
// whole point: the callback fires per connection ATTEMPT, so without
@ -681,7 +682,7 @@ pub async fn connect(cfg: QueueConfig) -> Result<async_nats::Client, Error> {
let cache: std::sync::Arc<tokio::sync::Mutex<Option<CachedToken>>> =
std::sync::Arc::new(tokio::sync::Mutex::new(None));
let client = async_nats::ConnectOptions::with_auth_callback(move |_nonce| {
let mut options = async_nats::ConnectOptions::with_auth_callback(move |_nonce| {
let http = http.clone();
let cfg = cfg.clone();
let cache = cache.clone();
@ -745,13 +746,23 @@ pub async fn connect(cfg: QueueConfig) -> Result<async_nats::Client, Error> {
// near expiry. (Before the cache, "runs the callback" meant "mints",
// and this retry loop was a token-request loop. See
// `MAX_RECONNECT_DELAY`.)
.retry_on_initial_connect()
.connect(&url)
.await
.map_err(|source| Error::Connect {
url: url.clone(),
source,
})?;
.retry_on_initial_connect();
// The queue's leaf chains to the swarm's own PKI root, which a host
// process's platform store does not hold: the same anchor the token
// endpoint needs. async-nats then trusts this file INSTEAD of the platform
// roots, which is right here, because the queue has no public certificate.
if let Some(path) = ca_file {
options = options.add_root_certificates(path);
}
let client = options
.connect(&url)
.await
.map_err(|source| Error::Connect {
url: url.clone(),
source,
})?;
Ok(client)
}