Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.
- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
hive-agent-bao-identity and hive-agent-forge-token use it.
queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
forge-token.nix is a fetch unit with the same short budget that was
added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
(--no-block) a loaded nginx that is not active; an active nginx keeps
the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
fetch unit, swarm-services-cert included.
Refs #4662
32 lines
1.3 KiB
Nix
32 lines
1.3 KiB
Nix
# Retry shape for a oneshot that fetches a credential or certificate from
|
||
# the secret store: every 30s for 24h. Under `seal = "shamir"` an operator
|
||
# unseals BY HAND, and a fetch fails for as long as that takes, or for as
|
||
# long as the store or the gateway in front of it is down. `start-limit-hit`
|
||
# does not clear when the store comes back, so a budget shorter than the
|
||
# outage leaves the unit failed until something starts it again.
|
||
#
|
||
# `StartLimit*` are `[Unit]` settings — systemd ignores them under
|
||
# `[Service]` — and the window must exceed `RestartSec × burst` or it closes
|
||
# between attempts and the burst is never reached: 2880 × 30s is 24h inside
|
||
# a 25h window.
|
||
#
|
||
# ⚠️ A unit in auto-restart is still `activating`, so its start job stays
|
||
# queued across attempts: anything ordered `After=` it waits for as long as
|
||
# it retries, up to the full 24h.
|
||
#
|
||
# Pure attrset — NOT a NixOS module. Use from a unit definition:
|
||
#
|
||
# storeRetry = import ./lib/store-retry.nix { };
|
||
# systemd.services.foo = {
|
||
# inherit (storeRetry) startLimitBurst startLimitIntervalSec;
|
||
# serviceConfig = storeRetry.serviceConfig // { Type = "oneshot"; };
|
||
# };
|
||
{ }:
|
||
{
|
||
startLimitBurst = 2880;
|
||
startLimitIntervalSec = 90000;
|
||
serviceConfig = {
|
||
Restart = "on-failure";
|
||
RestartSec = 30;
|
||
};
|
||
}
|