Watch
0
0
Fork
You've already forked hyperhive
0

swarm-otel: bound the collector's restart backoff

The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.

Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
This commit is contained in:
atlas 2026-09-29 00:02:24 +02:00 • committed by mara
commit 916c441e82
2 changed files with 57 additions and 10 deletions

View file

@ -223,6 +223,27 @@ let
name = "the swarm collector's journald receiver carries the shared PRIORITY mapping";
ok = carriesJournaldSeverity (otelSettings otelNoStores).receivers.journald;
}
{
# Sibling of ./hive-otel.nix's identical case for the host-tier
# collector: this one's `oidc/*` extensions call out to authelia at
# startup, so it fails the same way when that restart races its own.
# nixpkgs ships `Restart = "always"` with no `RestartSec`, so without
# this the unit burns its five default attempts inside two seconds and
# lands in `start-limit-hit`, where it stops retrying. The third clause
# is the one that has to hold — `StartLimit*` are `[Unit]` settings
# that systemd ignores under `[Service]`, so a bound written into
# `serviceConfig` renders, deploys and does nothing.
name = "the swarm collector backs off a failed bind from [Unit], not [Service]";
ok =
let
u = otelNoStores.containers.swarm-otel.config.systemd.services.opentelemetry-collector;
in
u.serviceConfig.RestartSec or 0 > 0
&& toString u.unitConfig.StartLimitBurst == "12"
&& !(u.serviceConfig ? StartLimitBurst)
# The window has to outlast every attempt, or the burst is unreachable.
&& u.startLimitIntervalSec or 0 > u.serviceConfig.RestartSec * u.startLimitBurst;
}
];
in
runGroup "swarm-otel-core" cases