swarm-bao: refuse a remote reader that named seven of the eight leaves

The four-way client-cert split gives each store reader its own leaf, and
three of the four readers render only where their own leaf exists. On a
host that mints its own PKI glue-bao-tls.nix defaults all eight, so there
is nothing to do; on a hand-configured remote-store hive, omitting one
pair used to mean that unit silently did not render — a privilege-
narrowing unit absent from a green build, with the missing unit as the
only evidence.

Each of the three now asserts its own pair, shaped after
swarm-grafana.nix's haveClientIdentity assertion and named to the pair it
needs. What differs from Grafana's is the gate: these fire only where the
host demonstrably reads the store (it holds deploy.bao.clientCertFile and
clientKeyFile) and the consumer is on. A host with no store identity is
the supported no-store deployment and still evaluates; the collector's
no-secret degrade is untouched, because that host holds no clientCertFile
either.

Also rewords three passive-voice sentences in docs/swarm/secrets.md that
vale flagged, and documents what the refusal costs and where it stays
silent.
This commit is contained in:
atlas 2026-09-23 09:56:43 +02:00 committed by mara
commit d3e4951cc8
6 changed files with 627 additions and 276 deletions

View file

@ -49,6 +49,14 @@ let
haveClientIdentity =
baoDeploy.matrixTokenClientCertFile != null && baoDeploy.matrixTokenClientKeyFile != null;
# Does this host read the store at all — the hive's own leaf, which is the
# one thing a remote-store deployment has always had to place by hand. Only
# used to decide whether a missing per-principal leaf is a mistake or a
# deployment that has no store: a host holding neither is the supported
# no-store shape, and one holding this pair but not the pair above named
# seven of the eight options and stopped.
hiveReaderIdentity = baoDeploy.clientCertFile != null && baoDeploy.clientKeyFile != null;
# Where the token lives in the store. A path, not a convention to guess at:
# whoever writes it and whoever reads it must agree, and the agreement
# belongs in one visible place.
@ -73,134 +81,178 @@ let
matrixMachine = "hive-matrix";
in
{
config = lib.mkIf (hyperhiveCfg.enable && haveClientIdentity && deployCfg.matrix.enable) {
# Same rule as the unit's own gate: this reader exists on a host that has a
# client identity and a homeserver, which is not every host that runs the
# store, so the store's module cannot name it.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-matrix-token" ];
config = lib.mkMerge [
# ⚠️ A SEPARATE arm from the unit below, and that separation is the whole
# mechanism: the unit's arm is gated on `haveClientIdentity`, so an
# assertion written inside it could never be reached in the state it
# exists to report.
#
# Shaped after ./swarm-grafana.nix's `haveClientIdentity` assertion — the
# same refusal, named to this principal's own pair. What differs is the
# gate. Grafana asserts wherever Grafana runs, because a Grafana with no
# store identity has no way in at all; a homeserver with no store identity
# is a hive that has no store, which is supported. So this one additionally
# requires `hiveReaderIdentity`: the host demonstrably reads the store, and
# named every option but this pair.
(lib.mkIf (hyperhiveCfg.enable && deployCfg.matrix.enable && hiveReaderIdentity) {
assertions = [
{
assertion = haveClientIdentity;
message = ''
This host reads the swarm secret store (services.hyperhive.deploy.bao.clientCertFile
is set) and runs a homeserver, so it needs the matrix appservice
token reader's own client identity: set both
systemd.services.swarm-bao-matrix-token = {
description = "fetch the matrix appservice token from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.
# `Requires=` on an absent unit fails the job outright, so the ordering is
# conditional even though the read is not: off-host there is nothing local
# to wait for, and the timeout below is what bounds the attempt instead.
after = lib.optionals baoDeploy.enable [
"swarm-bao-pki.service"
"container@${baoCfg.machine}.service"
services.hyperhive.deploy.bao.matrixTokenClientCertFile
services.hyperhive.deploy.bao.matrixTokenClientKeyFile
swarm-bao-matrix-token.service fetches this hive's appservice token
out of the store, and without these it is not rendered at all
leaving the homeserver authenticating hive-c0re against whatever is
already on disk, which is a 401 on every request naming nothing.
This reader's OWN leaf, not deploy.bao.clientCertFile. That one is
the hive's, and its grant reads every secret in the store; this role
reads the one appservice-token path. Pointing this option at the
hive's leaf would evaluate, deploy and log in and undo the split.
On a hive that runs the store, glue-bao-tls.nix supplies both as
defaults and there is nothing to do. Elsewhere the leaf is issued
from that CA out of band and named here see docs/swarm/secrets.md.
'';
}
];
wants = lib.optionals baoDeploy.enable [ "container@${baoCfg.machine}.service" ];
requires = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" ];
before = [ "container@${matrixMachine}.service" ];
wantedBy = [ "container@${matrixMachine}.service" ];
path = [
deployCfg.bao.package
pkgs.coreutils
];
# Sized for the race this loses, not for an unseal. `swarm-bao` comes up
# seconds before this unit asks, and the cert-auth role it logs in
# against is written seconds after — so a few short attempts cover it.
# ⚠️ `swarm-bao-controller-policy`'s 2880 × 30s is NOT the model to copy.
# That unit blocks nothing; this one is `Before=` the homeserver's
# container, and whether that ordering waits across an auto-restart is
# unverified — so an hours-long window would be a bet on an unknown,
# where a minute is not. A store still sealed after it keeps the degrade
# below, as today.
#
# `StartLimit*` are `[Unit]` settings, so they go here and not in
# `serviceConfig` — systemd ignores them under `[Service]`. The window
# has to exceed `RestartSec × burst`.
startLimitBurst = 4;
startLimitIntervalSec = 300;
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
# What actually bounds the read below. Stated here rather than
# left to systemd's default, so the number a boot waits on is in
# the file that waits.
TimeoutStartSec = 30;
Restart = "on-failure";
RestartSec = 15;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
BAO_CLIENT_CERT = baoDeploy.matrixTokenClientCertFile;
BAO_CLIENT_KEY = baoDeploy.matrixTokenClientKeyFile;
}
# Absent means the system trust store, which is what a deployment with a
# real CA wants and what a self-signed one must not be left with.
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
BAO_CACERT = baoDeploy.serverCaFile;
};
script = ''
set -euo pipefail
})
# A sealed or uninitialised store answers on the port and never
# answers the read, so "the store is up" is not the same as "the
# store can answer". `TimeoutStartSec` above is the bound; the
# homeserver only `Wants=` this unit, so hitting it degrades to
# keeping the local token rather than holding up the container.
# `bao`'s own message is the only thing separating a missing value
# from a refused identity from an unreachable host. This unit's
# degraded mode is correct for all three, so it reports which one
# rather than asserting all three in a sentence of ours — a reader
# that cannot say why it read nothing is indistinguishable from a
# broken one.
err="$(mktemp)"
trap 'rm -f "$err"' EXIT
(lib.mkIf (hyperhiveCfg.enable && haveClientIdentity && deployCfg.matrix.enable) {
# Same rule as the unit's own gate: this reader exists on a host that has a
# client identity and a homeserver, which is not every host that runs the
# store, so the store's module cannot name it.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-matrix-token" ];
# Cert auth is a login, not a transport setting. The `BAO_CLIENT_*`
# variables above only decide which certificate the TLS handshake
# presents; without a token `bao` asks its token helper instead, and
# that is a `sh` this unit's `path` does not carry. `-token-only`
# answers on stdout and skips the helper on both sides.
systemd.services.swarm-bao-matrix-token = {
description = "fetch the matrix appservice token from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.
# `Requires=` on an absent unit fails the job outright, so the ordering is
# conditional even though the read is not: off-host there is nothing local
# to wait for, and the timeout below is what bounds the attempt instead.
after = lib.optionals baoDeploy.enable [
"swarm-bao-pki.service"
"container@${baoCfg.machine}.service"
];
wants = lib.optionals baoDeploy.enable [ "container@${baoCfg.machine}.service" ];
requires = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" ];
before = [ "container@${matrixMachine}.service" ];
wantedBy = [ "container@${matrixMachine}.service" ];
path = [
deployCfg.bao.package
pkgs.coreutils
];
# Sized for the race this loses, not for an unseal. `swarm-bao` comes up
# seconds before this unit asks, and the cert-auth role it logs in
# against is written seconds after — so a few short attempts cover it.
# ⚠️ `swarm-bao-controller-policy`'s 2880 × 30s is NOT the model to copy.
# That unit blocks nothing; this one is `Before=` the homeserver's
# container, and whether that ordering waits across an auto-restart is
# unverified — so an hours-long window would be a bet on an unknown,
# where a minute is not. A store still sealed after it keeps the degrade
# below, as today.
#
# Fails LOUDLY, unlike the read below: the three states a login failure
# covers — store not up, sealed, role not written yet — are all things
# a retry fixes, and `Restart=on-failure` above is what retries. Exiting
# 0 here spends the whole boot on a condition that was seconds old.
if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then
echo "could not log in to swarm-bao with this host's certificate; keeping the token hive-matrix already has." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
# `StartLimit*` are `[Unit]` settings, so they go here and not in
# `serviceConfig` — systemd ignores them under `[Service]`. The window
# has to exceed `RestartSec × burst`.
startLimitBurst = 4;
startLimitIntervalSec = 300;
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
# What actually bounds the read below. Stated here rather than
# left to systemd's default, so the number a boot waits on is in
# the file that waits.
TimeoutStartSec = 30;
Restart = "on-failure";
RestartSec = 15;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
BAO_CLIENT_CERT = baoDeploy.matrixTokenClientCertFile;
BAO_CLIENT_KEY = baoDeploy.matrixTokenClientKeyFile;
}
# Absent means the system trust store, which is what a deployment with a
# real CA wants and what a self-signed one must not be left with.
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
BAO_CACERT = baoDeploy.serverCaFile;
};
script = ''
set -euo pipefail
# A sealed or uninitialised store answers on the port and never
# answers the read, so "the store is up" is not the same as "the
# store can answer". `TimeoutStartSec` above is the bound; the
# homeserver only `Wants=` this unit, so hitting it degrades to
# keeping the local token rather than holding up the container.
# `bao`'s own message is the only thing separating a missing value
# from a refused identity from an unreachable host. This unit's
# degraded mode is correct for all three, so it reports which one
# rather than asserting all three in a sentence of ours — a reader
# that cannot say why it read nothing is indistinguishable from a
# broken one.
err="$(mktemp)"
trap 'rm -f "$err"' EXIT
# Cert auth is a login, not a transport setting. The `BAO_CLIENT_*`
# variables above only decide which certificate the TLS handshake
# presents; without a token `bao` asks its token helper instead, and
# that is a `sh` this unit's `path` does not carry. `-token-only`
# answers on stdout and skips the helper on both sides.
#
# Fails LOUDLY, unlike the read below: the three states a login failure
# covers — store not up, sealed, role not written yet — are all things
# a retry fixes, and `Restart=on-failure` above is what retries. Exiting
# 0 here spends the whole boot on a condition that was seconds old.
if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then
echo "could not log in to swarm-bao with this host's certificate; keeping the token hive-matrix already has." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
fi
exit 1
fi
exit 1
fi
export BAO_TOKEN
export BAO_TOKEN
if ! token="$(bao kv get -field=value ${lib.escapeShellArg tokenPath} 2>"$err")"; then
echo "swarm-bao did not return ${tokenPath}; keeping the token hive-matrix already has." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
if ! token="$(bao kv get -field=value ${lib.escapeShellArg tokenPath} 2>"$err")"; then
echo "swarm-bao did not return ${tokenPath}; keeping the token hive-matrix already has." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
fi
exit 0
fi
exit 0
fi
if [ -z "$token" ]; then
echo "swarm-bao returned an empty ${tokenPath}; keeping the local token." >&2
exit 0
fi
if [ -z "$token" ]; then
echo "swarm-bao returned an empty ${tokenPath}; keeping the local token." >&2
exit 0
fi
umask 077
printf '%s\n' "$token" > ${lib.escapeShellArg (toString deployCfg.matrix.appserviceTokenFile)}
chmod 0600 ${lib.escapeShellArg (toString deployCfg.matrix.appserviceTokenFile)}
umask 077
printf '%s\n' "$token" > ${lib.escapeShellArg (toString deployCfg.matrix.appserviceTokenFile)}
chmod 0600 ${lib.escapeShellArg (toString deployCfg.matrix.appserviceTokenFile)}
# Re-stamp the registration file from the token just written. The
# token is half an agreement — the registration the homeserver loads
# has to carry the same value — so writing the file and stopping
# would leave the homeserver authenticating hive-c0re against
# whatever activation put there: a 401 on every request, naming
# nothing. Unconditional rather than on-change, because this unit
# has no way to know what the registration currently says.
#
# hive-matrix's own renderer rather than a `printf` here, so the
# registration's shape has one home.
${deployCfg.matrix.appserviceRegistrationScript}
'';
};
};
# Re-stamp the registration file from the token just written. The
# token is half an agreement — the registration the homeserver loads
# has to carry the same value — so writing the file and stopping
# would leave the homeserver authenticating hive-c0re against
# whatever activation put there: a 401 on every request, naming
# nothing. Unconditional rather than on-change, because this unit
# has no way to know what the registration currently says.
#
# hive-matrix's own renderer rather than a `printf` here, so the
# registration's shape has one home.
${deployCfg.matrix.appserviceRegistrationScript}
'';
};
})
];
}

View file

@ -51,6 +51,14 @@ let
haveClientIdentity =
baoDeploy.queueAgentClientCertFile != null && baoDeploy.queueAgentClientKeyFile != null;
# Does this host read the store at all — the hive's own leaf, which is the
# one thing a remote-store deployment has always had to place by hand. Only
# used to decide whether a missing per-principal leaf is a mistake or a
# deployment that has no store: a host holding neither is the supported
# no-store shape, and one holding this pair but not the pair above named
# seven of the eight options and stopped.
hiveReaderIdentity = baoDeploy.clientCertFile != null && baoDeploy.clientKeyFile != null;
credentialDir = toString deployCfg.hive-controller.queue.agentCredentialDir;
secretFile = "${credentialDir}/secret";
clientIdFile = "${credentialDir}/client_id";
@ -94,144 +102,192 @@ in
};
};
config = lib.mkIf (hyperhiveCfg.enable && haveClientIdentity) {
# Same rule as the unit's own gate: this reader exists on any host holding
# a client identity, which is not every host that runs the store, so the
# store's module cannot name it.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-queue-agent" ];
config = lib.mkMerge [
# ⚠️ A SEPARATE arm from the unit below, and that separation is the whole
# mechanism: the unit's arm is gated on `haveClientIdentity`, so an
# assertion written inside it could never be reached in the state it
# exists to report.
#
# Shaped after ./swarm-grafana.nix's `haveClientIdentity` assertion — the
# same refusal, named to this principal's own pair. Where Grafana's gate
# is `deploy.grafana.enable`, this reader has no toggle of its own to
# check: every hive runs one for its own agents, so reading the store at
# all is what asks for it. Hence `hiveReaderIdentity` alone — a host
# holding no hive leaf has no store to read and nothing is missing.
(lib.mkIf (hyperhiveCfg.enable && hiveReaderIdentity) {
assertions = [
{
assertion = haveClientIdentity;
message = ''
This host reads the swarm secret store (services.hyperhive.deploy.bao.clientCertFile
is set), so it needs the agent queue credential reader's own client
identity: set both
systemd.services.swarm-bao-queue-agent = {
description = "fetch this hive's agent queue credential from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.
# `Requires=` on an absent unit fails the job outright, so the ordering is
# conditional even though the read is not: off-host there is nothing local
# to wait for, and the timeout below is what bounds the attempt instead.
after = lib.optionals baoDeploy.enable [
"swarm-bao-pki.service"
"container@${baoCfg.machine}.service"
];
wants = lib.optionals baoDeploy.enable [ "container@${baoCfg.machine}.service" ];
requires = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" ];
# Ordered before hive-c0re, so no agent container renders ahead of an
# attempt at its credential. `Wants=`, not `Requires=`: a store this
# unit can't reach delays hive-c0re's start by its own start-limit
# window (`TimeoutStartSec`, retried up to `startLimitBurst` times
# below) rather than failing it — hive-c0re starts once that window
# elapses, whatever credential is or isn't on disk by then.
before = [ "hive-c0re.service" ];
wantedBy = [
"multi-user.target"
"hive-c0re.service"
];
path = [
baoDeploy.package
pkgs.coreutils
];
# Sized for the race this loses, not for an unseal: `swarm-bao` comes up
# seconds before this unit asks, and the cert-auth role it logs in
# against is written seconds after, so a few short attempts cover it.
# An hours-long window would be a bet on a store that is sealed, and the
# degrade below is already correct for that.
#
# `StartLimit*` are `[Unit]` settings, so they go here and not in
# `serviceConfig` — systemd ignores them under `[Service]`. The window
# has to exceed `RestartSec × burst`.
startLimitBurst = 4;
startLimitIntervalSec = 300;
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
# What actually bounds the reads below. Stated here rather than left to
# systemd's default, so the number a boot waits on is in the file that
# waits.
TimeoutStartSec = 30;
Restart = "on-failure";
RestartSec = 15;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
BAO_CLIENT_CERT = baoDeploy.queueAgentClientCertFile;
BAO_CLIENT_KEY = baoDeploy.queueAgentClientKeyFile;
}
# Absent means the system trust store, which is what a deployment with a
# real CA wants and what a self-signed one must not be left with.
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
BAO_CACERT = baoDeploy.serverCaFile;
};
script = ''
set -euo pipefail
services.hyperhive.deploy.bao.queueAgentClientCertFile
services.hyperhive.deploy.bao.queueAgentClientKeyFile
# `bao`'s own message is the only thing separating a missing value from
# a refused identity from an unreachable host. This unit's degraded
# mode is correct for all three, so it reports which one rather than
# asserting all three in a sentence of ours.
err="$(mktemp)"
trap 'rm -f "$err"' EXIT
swarm-bao-queue-agent.service fetches this hive's agent queue
credential out of the store, and without these it is not rendered
at all leaving hive-c0re with no queue credential and no unit
that would ever have written one.
# Cert auth is a login, not a transport setting. The `BAO_CLIENT_*`
# variables above only decide which certificate the TLS handshake
# presents; without a token `bao` asks its token helper instead, and
# that is a `sh` this unit's `path` does not carry. `-token-only`
# answers on stdout and skips the helper on both sides.
Every hive runs this reader for its own agents, so unlike the other
three principals there is no per-service toggle that turns it off:
a hive that reads the store at all is a hive that needs this pair.
This reader's OWN leaf, not deploy.bao.clientCertFile. That one
is the hive's, and its grant reads every secret in the store; this
role reads the one queue-credential path. Pointing this option at
the hive's leaf would evaluate, deploy and log in and undo the
split.
On a hive that runs the store, glue-bao-tls.nix supplies both as
defaults and there is nothing to do. Elsewhere the leaf is issued
from that CA out of band and named here see docs/swarm/secrets.md.
'';
}
];
})
(lib.mkIf (hyperhiveCfg.enable && haveClientIdentity) {
# Same rule as the unit's own gate: this reader exists on any host holding
# a client identity, which is not every host that runs the store, so the
# store's module cannot name it.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-queue-agent" ];
systemd.services.swarm-bao-queue-agent = {
description = "fetch this hive's agent queue credential from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.
# `Requires=` on an absent unit fails the job outright, so the ordering is
# conditional even though the read is not: off-host there is nothing local
# to wait for, and the timeout below is what bounds the attempt instead.
after = lib.optionals baoDeploy.enable [
"swarm-bao-pki.service"
"container@${baoCfg.machine}.service"
];
wants = lib.optionals baoDeploy.enable [ "container@${baoCfg.machine}.service" ];
requires = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" ];
# Ordered before hive-c0re, so no agent container renders ahead of an
# attempt at its credential. `Wants=`, not `Requires=`: a store this
# unit can't reach delays hive-c0re's start by its own start-limit
# window (`TimeoutStartSec`, retried up to `startLimitBurst` times
# below) rather than failing it — hive-c0re starts once that window
# elapses, whatever credential is or isn't on disk by then.
before = [ "hive-c0re.service" ];
wantedBy = [
"multi-user.target"
"hive-c0re.service"
];
path = [
baoDeploy.package
pkgs.coreutils
];
# Sized for the race this loses, not for an unseal: `swarm-bao` comes up
# seconds before this unit asks, and the cert-auth role it logs in
# against is written seconds after, so a few short attempts cover it.
# An hours-long window would be a bet on a store that is sealed, and the
# degrade below is already correct for that.
#
# Fails LOUDLY, unlike the reads below: the three states a login
# failure covers — store not up, sealed, role not written yet — are all
# things a retry fixes, and `Restart=on-failure` above is what retries.
if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then
echo "could not log in to swarm-bao with this host's certificate; leaving the queue credential in ${credentialDir} as it is." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
# `StartLimit*` are `[Unit]` settings, so they go here and not in
# `serviceConfig` — systemd ignores them under `[Service]`. The window
# has to exceed `RestartSec × burst`.
startLimitBurst = 4;
startLimitIntervalSec = 300;
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
# What actually bounds the reads below. Stated here rather than left to
# systemd's default, so the number a boot waits on is in the file that
# waits.
TimeoutStartSec = 30;
Restart = "on-failure";
RestartSec = 15;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";
BAO_CLIENT_CERT = baoDeploy.queueAgentClientCertFile;
BAO_CLIENT_KEY = baoDeploy.queueAgentClientKeyFile;
}
# Absent means the system trust store, which is what a deployment with a
# real CA wants and what a self-signed one must not be left with.
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
BAO_CACERT = baoDeploy.serverCaFile;
};
script = ''
set -euo pipefail
# `bao`'s own message is the only thing separating a missing value from
# a refused identity from an unreachable host. This unit's degraded
# mode is correct for all three, so it reports which one rather than
# asserting all three in a sentence of ours.
err="$(mktemp)"
trap 'rm -f "$err"' EXIT
# Cert auth is a login, not a transport setting. The `BAO_CLIENT_*`
# variables above only decide which certificate the TLS handshake
# presents; without a token `bao` asks its token helper instead, and
# that is a `sh` this unit's `path` does not carry. `-token-only`
# answers on stdout and skips the helper on both sides.
#
# Fails LOUDLY, unlike the reads below: the three states a login
# failure covers — store not up, sealed, role not written yet — are all
# things a retry fixes, and `Restart=on-failure` above is what retries.
if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then
echo "could not log in to swarm-bao with this host's certificate; leaving the queue credential in ${credentialDir} as it is." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
fi
exit 1
fi
exit 1
fi
export BAO_TOKEN
export BAO_TOKEN
# Two reads of one object rather than one `-format=json` parsed with
# `jq`: no sibling unit carries `jq` on its `path`, and the pair cannot
# actually disagree — the client id is derived from this hive's name, so
# a rotation landing between these two calls changes the secret and
# rewrites the same id.
if ! secret="$(bao kv get -field=value ${lib.escapeShellArg credentialPath} 2>"$err")"; then
echo "swarm-bao did not return ${credentialPath}; this hive's agents have no queue credential yet." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
# Two reads of one object rather than one `-format=json` parsed with
# `jq`: no sibling unit carries `jq` on its `path`, and the pair cannot
# actually disagree — the client id is derived from this hive's name, so
# a rotation landing between these two calls changes the secret and
# rewrites the same id.
if ! secret="$(bao kv get -field=value ${lib.escapeShellArg credentialPath} 2>"$err")"; then
echo "swarm-bao did not return ${credentialPath}; this hive's agents have no queue credential yet." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
fi
exit 0
fi
exit 0
fi
if ! client_id="$(bao kv get -field=client_id ${lib.escapeShellArg credentialPath} 2>"$err")"; then
echo "${credentialPath} holds no client_id; the secret alone is not a usable credential, so nothing is written." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
if ! client_id="$(bao kv get -field=client_id ${lib.escapeShellArg credentialPath} 2>"$err")"; then
echo "${credentialPath} holds no client_id; the secret alone is not a usable credential, so nothing is written." >&2
if [ -s "$err" ]; then
cat "$err" >&2
else
echo "bao failed without writing a diagnostic." >&2
fi
exit 0
fi
exit 0
fi
# Both or neither, for the reason `QueueConfig::from_env` refuses a
# half-set environment: a client that finds one of the two comes up
# "fine" and never connects.
if [ -z "$secret" ] || [ -z "$client_id" ]; then
echo "swarm-bao returned an empty field of ${credentialPath}; leaving the files as they are." >&2
exit 0
fi
# Both or neither, for the reason `QueueConfig::from_env` refuses a
# half-set environment: a client that finds one of the two comes up
# "fine" and never connects.
if [ -z "$secret" ] || [ -z "$client_id" ]; then
echo "swarm-bao returned an empty field of ${credentialPath}; leaving the files as they are." >&2
exit 0
fi
install -d -m 0755 ${lib.escapeShellArg credentialDir}
install -d -m 0755 ${lib.escapeShellArg credentialDir}
umask 077
printf '%s\n' "$secret" > ${lib.escapeShellArg secretFile}
chmod 0600 ${lib.escapeShellArg secretFile}
umask 077
printf '%s\n' "$secret" > ${lib.escapeShellArg secretFile}
chmod 0600 ${lib.escapeShellArg secretFile}
# `0644` on purpose: an OIDC client id is presented to the token
# endpoint on every connection and is public by construction.
printf '%s\n' "$client_id" > ${lib.escapeShellArg clientIdFile}
chmod 0644 ${lib.escapeShellArg clientIdFile}
'';
};
};
# `0644` on purpose: an OIDC client id is presented to the token
# endpoint on every connection and is public by construction.
printf '%s\n' "$client_id" > ${lib.escapeShellArg clientIdFile}
chmod 0644 ${lib.escapeShellArg clientIdFile}
'';
};
})
];
}

View file

@ -242,12 +242,14 @@ let
# A reader of the store is defined by holding a certificate the store
# accepts, never by standing next to it — the rule
# ./glue-matrix-bao-token.nix states in full. Unlike Grafana's identical-
# looking flag, this one stays outside `assertions`: the store-reading unit
# below simply does not render without it, the same choice
# looking flag, this one does not gate an assertion on its own: the
# store-reading unit below simply does not render without it, the same choice
# ./glue-matrix-bao-token.nix and ./glue-queue-agent-credential.nix make for
# their own optional readers, because a collector with no client identity is
# their own readers, because a collector with no client identity is
# `haveCollectorSecret = false` above, and that is already a supported,
# merely degraded shape rather than a service with no way in at all.
# merely degraded shape rather than a service with no way in at all. What
# IS asserted is narrower and lives in `assertions` below — see
# `hiveReaderIdentity`.
#
# 🩸 The collector's OWN leaf, not `clientCertFile` — the hive's, which four
# units used to share. Bao matches a cert-auth role on the CN, so one leaf for
@ -257,6 +259,16 @@ let
haveClientIdentity =
baoDeploy.otelOidcClientCertFile != null && baoDeploy.otelOidcClientKeyFile != null;
# Does this host read the store at all — the hive's own leaf, which is the
# one thing a remote-store deployment has always had to place by hand. It is
# what separates the degrade the comment above describes from the mistake the
# assertion below reports: a collector on a host holding NO store identity is
# the supported shape, and one on a host that demonstrably reads the store
# named seven of the eight options and stopped. Before the four-way split
# that second host rendered this unit off `clientCertFile`, so it is a silent
# regression rather than a choice anyone made.
hiveReaderIdentity = baoDeploy.clientCertFile != null && baoDeploy.clientKeyFile != null;
# Where the publisher on authelia's host leaves this client's secret —
# composed from the same swarm-wide `clientId` the registration in
# ./glue-swarm-otel-oidc-client.nix uses, so a rename cannot leave one of
@ -732,11 +744,14 @@ in
# value.
#
# ⚠️ Renders only where `haveClientIdentity` holds, unlike
# ./swarm-grafana.nix's equivalent unit. That module asserts the identity
# because a Grafana with none has no way in at all; this collector without
# one is `haveCollectorSecret = false` above — already a supported,
# merely degraded shape, so the unit that would fetch a credential simply
# does not exist rather than refusing the build for want of one.
# ./swarm-grafana.nix's equivalent unit. That module refuses any host that
# runs Grafana without the identity, because a Grafana with none has no way
# in at all; this collector without one is `haveCollectorSecret = false`
# above — already a supported, merely degraded shape, so the unit that would
# fetch a credential simply does not exist rather than refusing the build
# for want of one. The one case that IS refused is the host that already
# holds the hive's leaf and is missing only this pair, which is a silent
# regression rather than that degrade — `hiveReaderIdentity` above.
systemd.services.swarm-bao-otel-oidc = lib.mkIf haveClientIdentity {
description = "fetch the swarm collector's OIDC client secret from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.
@ -860,6 +875,45 @@ in
systemd.services."container@${cfg.machine}" = caTrust.containerOrdering;
assertions = [
{
# Shaped after ./swarm-grafana.nix's `haveClientIdentity` assertion —
# the same refusal, named to this principal's own pair. What differs is
# the `!hiveReaderIdentity ||` guard, and it is what keeps the degrade
# the flag's own comment describes intact: Grafana refuses any host
# that runs it without a leaf, because a Grafana with no SSO has no way
# in at all, while a collector with none still receives telemetry. So
# this fires only where the host already reads the store and is missing
# this one pair — which before the four-way split rendered the unit off
# `clientCertFile`, and now silently does not.
assertion = !hiveReaderIdentity || haveClientIdentity;
message = ''
This host reads the swarm secret store (services.hyperhive.deploy.bao.clientCertFile
is set) and runs the swarm collector, so it needs the collector's own
client identity: set both
services.hyperhive.deploy.bao.otelOidcClientCertFile
services.hyperhive.deploy.bao.otelOidcClientKeyFile
swarm-bao-otel-oidc.service fetches the collector's OIDC client
secret out of the store, and without these it is not rendered at all
so this collector would authenticate to nothing, quietly, on a host
that has everything it needs to fetch the secret.
A collector on a host holding NO store identity is a different and
supported shape: it runs without a client secret and still receives
telemetry. That is not this host.
The collector's OWN leaf, not deploy.bao.clientCertFile. That one
is the hive's, and its grant reads every secret in the store; this
role reads the one path this collector's client secret lives at.
Pointing this option at the hive's leaf would evaluate, deploy and
log in and undo the split.
On a hive that runs the store, glue-bao-tls.nix supplies both as
defaults and there is nothing to do. Elsewhere the leaf is issued
from that CA out of band and named here see docs/swarm/secrets.md.
'';
}
# ⚠️ An assertion that this collector has "somewhere to send" was REMOVED
# rather than relaxed: it read the store's *per-host* enable, so it
# rejected at eval the very deployment the stores are reached by domain