Watch
0
0
Fork
You've already forked hyperhive
0

credential units: restart consumers on a changed credential; fix the ordering claim

The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
This commit is contained in:
atlas 2026-09-30 00:00:28 +02:00 • committed by mara
commit 7eb966fe2b
9 changed files with 178 additions and 22 deletions

View file

@ -239,6 +239,7 @@ let
grafanaUid = config.ids.uids.grafana;
atomicWriteSecret = import ./lib/atomic-write-secret.nix { };
refreshConsumer = import ./lib/refresh-consumer.nix { };
# Shared host netns, like every sibling swarm container: the gateway
# reaches this at 127.0.0.1:<port>.
@ -673,9 +674,11 @@ in
path = [
baoDeploy.package
pkgs.coreutils
pkgs.systemd
];
# ./lib/store-retry.nix. Grafana's container is ordered after this unit,
# so it waits while the login below keeps failing.
# ./lib/store-retry.nix. Grafana's container is ordered after this unit
# and waits for one attempt; a secret a later attempt lands restarts
# Grafana (below).
inherit (storeRetry) startLimitBurst startLimitIntervalSec;
serviceConfig = storeRetry.serviceConfig // {
Type = "oneshot";
@ -700,6 +703,7 @@ in
set -euo pipefail
${atomicWriteSecret}
${refreshConsumer}
# `bao`'s own message is the only thing separating a missing value from
# a refused identity from an unreachable host. This unit's degraded
@ -757,7 +761,14 @@ in
# grants the group nothing; if this mode ever widens, the gid has to be
# discovered at runtime rather than assumed.
install -d -m 0755 ${lib.escapeShellArg hostSecretDir}
changed=0
if secret_differs ${lib.escapeShellArg hostSecretPath} "$secret"; then changed=1; fi
atomic_write_secret 0400 ${lib.escapeShellArg "${toString config.ids.uids.grafana}:0"} ${lib.escapeShellArg hostSecretPath} "$secret"
# `$__file{}` is expanded once, while Grafana parses its config.
if [ "$changed" = 1 ]; then
refresh_consumer ${lib.escapeShellArg cfg.machine} grafana.service
fi
'';
};