Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/nix/host-modules/lib/refresh-consumer.nix
atlas 7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00

57 lines
2.5 KiB
Nix

# Shared shell steps for a host oneshot that fetches a credential into a
# file a service inside a container loads at start (`LoadCredential`, or a
# config file expanded once while parsing). Such a service never sees a
# value that lands after it started — a late fetch or a rotation — until it
# is restarted, so the fetch restarts it, and only when the value changed:
# a routine re-fetch of the same value leaves it alone.
#
# secret_differs <path> <value>
# True when <path> is missing or holds something other than <value>.
# Call it BEFORE `atomic_write_secret` writes <path>. Compares with
# `$(< path)`, a bash builtin, so the value never becomes an argument in
# /proc; both sides lose their trailing newlines, which is how
# `atomic_write_secret` writes and `$(bao …)` reads.
#
# refresh_consumer <machine> <unit>
# Nothing while `container@<machine>` is not active: a container that
# starts later loads the file as it starts. Otherwise a running
# consumer is restarted, and a failed one —
# a consumer that found no credential may have hit its start limit — is
# reset and started. A consumer stopped on purpose stays stopped.
# `--no-block` throughout: a caller ordered `Before=` its consumer must
# not wait on a job that waits on the caller (the deadlock stated at
# `swarm-services-cert`'s propagation in ../hive-tls.nix).
#
# Pure function — NOT a NixOS module. Call it from a module's `let`:
#
# refreshConsumer = import ./lib/refresh-consumer.nix { };
# script = ''
# ${refreshConsumer}
# changed=0
# if secret_differs "$path" "$secret"; then changed=1; fi
# atomic_write_secret 0400 root:root "$path" "$secret"
# if [ "$changed" = 1 ]; then refresh_consumer my-machine my.service; fi
# '';
#
# Requires `systemd` on the caller's `path`.
{ }:
''
secret_differs() {
[ ! -f "$1" ] || [ "$(< "$1")" != "$2" ]
}
refresh_consumer() {
local machine="$1" unit="$2"
if ! systemctl is-active --quiet "container@$machine.service"; then
return 0
fi
if systemctl --machine="$machine" is-failed --quiet "$unit"; then
echo "the credential for $unit in $machine changed and $unit had failed — starting it"
systemctl --machine="$machine" reset-failed "$unit"
systemctl --machine="$machine" start --no-block "$unit"
else
echo "the credential for $unit in $machine changed — restarting it if it runs"
systemctl --machine="$machine" try-restart --no-block "$unit"
fi
}
''