Watch
0
0
Fork
You've already forked hyperhive
0

credential units: 24h retry shape; start a failed nginx when the cert lands

Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
This commit is contained in:
atlas 2026-09-29 23:15:29 +02:00 • committed by mara
commit b3b42d3279
10 changed files with 137 additions and 125 deletions

View file

@ -866,6 +866,34 @@ let
&& u.serviceConfig.Restart == "on-failure"
) grantingUnitNames;
}
{
# A fetch from the store outlives a store or gateway that is down or
# sealed for hours: `start-limit-hit` does not clear when the store comes
# back, so a shorter budget leaves the credential missing until someone
# starts the unit by hand.
name = "every unit fetching a credential or certificate from the store retries 2880 times at 30s";
ok =
lib.all
(
unit:
let
s = baoGrantWithConsumers.systemd.services;
u = s.${unit};
in
s ? ${unit}
&& toString u.unitConfig.StartLimitBurst == "2880"
&& toString u.unitConfig.StartLimitIntervalSec == "90000"
&& toString u.serviceConfig.RestartSec == "30"
&& u.serviceConfig.Restart == "on-failure"
)
(
[
"swarm-services-cert"
"swarm-bao-forwarder-oidc"
]
++ policyReaders
);
}
{
# The only unit left acting with the token, so the only one that may
# skip on it.