Watch
0
0
Fork
You've already forked hyperhive
0

credential units: restart consumers on a changed credential; fix the ordering claim

The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
This commit is contained in:
atlas 2026-09-30 00:00:28 +02:00 • committed by mara
commit 7eb966fe2b
9 changed files with 178 additions and 22 deletions

View file

@ -167,6 +167,7 @@ let
collectorSecretPath = "/run/credentials/opentelemetry-collector.service/${collectorCredentialId}";
atomicWriteSecret = import ./lib/atomic-write-secret.nix { };
refreshConsumer = import ./lib/refresh-consumer.nix { };
# `attrNames` is sorted, so this is a function of the hive SET and not of
# the order anyone wrote it in.
@ -790,9 +791,11 @@ in
path = [
baoDeploy.package
pkgs.coreutils
pkgs.systemd
];
# ./lib/store-retry.nix. The collector's container is ordered after
# this unit, so it waits while the login below keeps failing.
# ./lib/store-retry.nix. The collector's container is ordered after this
# unit and waits for one attempt; a secret a later attempt lands restarts
# the collector (below).
inherit (storeRetry) startLimitBurst startLimitIntervalSec;
serviceConfig = storeRetry.serviceConfig // {
Type = "oneshot";
@ -817,6 +820,7 @@ in
set -euo pipefail
${atomicWriteSecret}
${refreshConsumer}
# `bao`'s own message is the only thing separating a missing value
# from a refused identity from an unreachable host. This unit's
@ -863,7 +867,15 @@ in
# no uid to give this to — `LoadCredential` reads it as root before
# the sandbox exists and re-exposes it to whichever uid the unit got.
install -d -m 0755 ${lib.escapeShellArg collectorHostSecretDir}
changed=0
if secret_differs ${lib.escapeShellArg collectorHostSecretPath} "$secret"; then changed=1; fi
atomic_write_secret 0400 root:root ${lib.escapeShellArg collectorHostSecretPath} "$secret"
# `LoadCredential` copies the file at start only, and a collector that
# started without it is in its start limit.
if [ "$changed" = 1 ]; then
refresh_consumer ${lib.escapeShellArg cfg.machine} opentelemetry-collector.service
fi
'';
};