credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its start job, so anything ordered after it waits for the whole 24h retry window. That is wrong under the default RestartMode=normal: each failed attempt passes through `failed`, which ends that start job. `After=` dependents proceed after one attempt, `Requires=` dependents fail with `dependency`, and the retries continue as fresh start jobs. The 2026-09-24 journal shows it with the already-2880 swarm-services-cert: nginx got "Dependency failed" 1ms after the first failure, and switch-to-configuration exited before the first restart was scheduled. The comments in lib/store-retry.nix, glue-matrix-bao-token.nix, glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix now say that, and so does docs/swarm/credentials.md. Because dependents start after one attempt, a consumer that loads its credential at start never sees a value a later attempt lands, or a rotated one. nix/host-modules/lib/refresh-consumer.nix adds `secret_differs` and `refresh_consumer`, and the four fetch units whose consumers take a start-time copy call them after the write, only when the value changed: - swarm-bao-matrix-token -> tuwunel.service in hive-matrix - swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel - swarm-bao-grafana-oidc -> grafana.service in the grafana container - swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao A running consumer is try-restarted, a failed one is reset and started, all with --no-block. Inline in the fetch script rather than a PathChanged path unit because the fetch script is the only writer and already knows whether the value changed, and it is the same shape as this PR's nginx hook and swarm-bao-nats-tls's restart of nats. module-eval-bao-grants gains one case per consumer. Refs #4662
This commit is contained in:
parent
b3b42d3279
commit
7eb966fe2b
9 changed files with 178 additions and 22 deletions
|
|
@ -174,9 +174,17 @@ needed the store first. The unit exists whenever
|
|||
sets from its own store address — the same all-or-nothing gate the per-hive
|
||||
queue credential beside it uses, and the reason the delivery above never lands
|
||||
in a container with nothing to read it. It fails loudly where the hive-side
|
||||
readers degrade quietly, which is deliberate: a missing queue secret means a
|
||||
swarm whose publisher has yet to run, while a refused certificate means an
|
||||
agent that believes it reaches the store and never does.
|
||||
readers treat a missing value as a normal state, which is deliberate: a
|
||||
missing queue secret means a swarm whose publisher has yet to run, while a
|
||||
refused certificate means an agent that believes it reaches the store and
|
||||
never does.
|
||||
|
||||
Both kinds of unit retry a failed login every 30 seconds for a day
|
||||
(`nix/host-modules/lib/store-retry.nix`). A unit ordered after one waits for a
|
||||
single attempt, not for the retries: an attempt that fails ends that start, so
|
||||
the dependent starts with whatever is already on disk. When a later attempt
|
||||
lands a changed value, the homeserver, Grafana and collector readers restart
|
||||
the service that loaded it (`nix/host-modules/lib/refresh-consumer.nix`).
|
||||
|
||||
## Progressive enhancement
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue