swarm-otel: bound the collector's restart backoff
The collector's oidc/* authenticators call out to authelia at startup, so a restart that races authelia's own (a redeploy that touches both, a store outage) can fail immediately. nixpkgs' upstream opentelemetry-collector module sets Restart=always with no RestartSec, so systemd's defaults (100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well under a second and leave the unit in start-limit-hit, dead until someone resets it by hand. Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec window (12/120s) sized so the burst can never trip while authelia comes back — same values host-modules/otel.nix already uses for the sibling host-tier collector, which depends on authelia the same way. Pins the [Unit]-vs-[Service] placement and the window relation in module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing case for the host tier.
This commit is contained in:
parent
375123f5d2
commit
916c441e82
2 changed files with 57 additions and 10 deletions
|
|
@ -223,6 +223,27 @@ let
|
|||
name = "the swarm collector's journald receiver carries the shared PRIORITY mapping";
|
||||
ok = carriesJournaldSeverity (otelSettings otelNoStores).receivers.journald;
|
||||
}
|
||||
{
|
||||
# Sibling of ./hive-otel.nix's identical case for the host-tier
|
||||
# collector: this one's `oidc/*` extensions call out to authelia at
|
||||
# startup, so it fails the same way when that restart races its own.
|
||||
# nixpkgs ships `Restart = "always"` with no `RestartSec`, so without
|
||||
# this the unit burns its five default attempts inside two seconds and
|
||||
# lands in `start-limit-hit`, where it stops retrying. The third clause
|
||||
# is the one that has to hold — `StartLimit*` are `[Unit]` settings
|
||||
# that systemd ignores under `[Service]`, so a bound written into
|
||||
# `serviceConfig` renders, deploys and does nothing.
|
||||
name = "the swarm collector backs off a failed bind from [Unit], not [Service]";
|
||||
ok =
|
||||
let
|
||||
u = otelNoStores.containers.swarm-otel.config.systemd.services.opentelemetry-collector;
|
||||
in
|
||||
u.serviceConfig.RestartSec or 0 > 0
|
||||
&& toString u.unitConfig.StartLimitBurst == "12"
|
||||
&& !(u.serviceConfig ? StartLimitBurst)
|
||||
# The window has to outlast every attempt, or the burst is unreachable.
|
||||
&& u.startLimitIntervalSec or 0 > u.serviceConfig.RestartSec * u.startLimitBurst;
|
||||
}
|
||||
];
|
||||
in
|
||||
runGroup "swarm-otel-core" cases
|
||||
|
|
|
|||
Loading…
Reference in a new issue