Watch
0
0
Fork
You've already forked hyperhive
0

credential units: restart consumers on a changed credential; fix the ordering claim

The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
This commit is contained in:
atlas 2026-09-30 00:00:28 +02:00 • committed by mara
commit 7eb966fe2b
9 changed files with 178 additions and 22 deletions

View file

@ -174,9 +174,17 @@ needed the store first. The unit exists whenever
sets from its own store address — the same all-or-nothing gate the per-hive
queue credential beside it uses, and the reason the delivery above never lands
in a container with nothing to read it. It fails loudly where the hive-side
readers degrade quietly, which is deliberate: a missing queue secret means a
swarm whose publisher has yet to run, while a refused certificate means an
agent that believes it reaches the store and never does.
readers treat a missing value as a normal state, which is deliberate: a
missing queue secret means a swarm whose publisher has yet to run, while a
refused certificate means an agent that believes it reaches the store and
never does.
Both kinds of unit retry a failed login every 30 seconds for a day
(`nix/host-modules/lib/store-retry.nix`). A unit ordered after one waits for a
single attempt, not for the retries: an attempt that fails ends that start, so
the dependent starts with whatever is already on disk. When a later attempt
lands a changed value, the homeserver, Grafana and collector readers restart
the service that loaded it (`nix/host-modules/lib/refresh-consumer.nix`).
## Progressive enhancement