The previous commit's comments said a unit in auto-restart keeps its start job, so anything ordered after it waits for the whole 24h retry window. That is wrong under the default RestartMode=normal: each failed attempt passes through `failed`, which ends that start job. `After=` dependents proceed after one attempt, `Requires=` dependents fail with `dependency`, and the retries continue as fresh start jobs. The 2026-09-24 journal shows it with the already-2880 swarm-services-cert: nginx got "Dependency failed" 1ms after the first failure, and switch-to-configuration exited before the first restart was scheduled. The comments in lib/store-retry.nix, glue-matrix-bao-token.nix, glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix now say that, and so does docs/swarm/credentials.md. Because dependents start after one attempt, a consumer that loads its credential at start never sees a value a later attempt lands, or a rotated one. nix/host-modules/lib/refresh-consumer.nix adds `secret_differs` and `refresh_consumer`, and the four fetch units whose consumers take a start-time copy call them after the write, only when the value changed: - swarm-bao-matrix-token -> tuwunel.service in hive-matrix - swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel - swarm-bao-grafana-oidc -> grafana.service in the grafana container - swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao A running consumer is try-restarted, a failed one is reset and started, all with --no-block. Inline in the fetch script rather than a PathChanged path unit because the fetch script is the only writer and already knows whether the value changed, and it is the same shape as this PR's nginx hook and swarm-bao-nats-tls's restart of nats. module-eval-bao-grants gains one case per consumer. Refs #4662
35 lines
1.5 KiB
Nix
35 lines
1.5 KiB
Nix
# Retry shape for a oneshot that fetches a credential or certificate from
|
||
# the secret store: every 30s for 24h. Under `seal = "shamir"` an operator
|
||
# unseals BY HAND, and a fetch fails for as long as that takes, or for as
|
||
# long as the store or the gateway in front of it is down. `start-limit-hit`
|
||
# does not clear when the store comes back, so a budget shorter than the
|
||
# outage leaves the unit failed until something starts it again.
|
||
#
|
||
# `StartLimit*` are `[Unit]` settings — systemd ignores them under
|
||
# `[Service]` — and the window must exceed `RestartSec × burst` or it closes
|
||
# between attempts and the burst is never reached: 2880 × 30s is 24h inside
|
||
# a 25h window.
|
||
#
|
||
# Under the default `RestartMode=normal` each failed attempt passes through
|
||
# `failed`, which ends that attempt's start job: a unit ordered `After=` it
|
||
# waits for ONE attempt, a unit that `Requires=` it fails with `dependency`,
|
||
# and the retries continue in the background as fresh start jobs. A consumer
|
||
# that loaded the credential before it landed does not see it without a
|
||
# restart — ./refresh-consumer.nix.
|
||
#
|
||
# Pure attrset — NOT a NixOS module. Use from a unit definition:
|
||
#
|
||
# storeRetry = import ./lib/store-retry.nix { };
|
||
# systemd.services.foo = {
|
||
# inherit (storeRetry) startLimitBurst startLimitIntervalSec;
|
||
# serviceConfig = storeRetry.serviceConfig // { Type = "oneshot"; };
|
||
# };
|
||
{ }:
|
||
{
|
||
startLimitBurst = 2880;
|
||
startLimitIntervalSec = 90000;
|
||
serviceConfig = {
|
||
Restart = "on-failure";
|
||
RestartSec = 30;
|
||
};
|
||
}
|