Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/nix/host-modules/lib/store-retry.nix
atlas b3b42d3279 credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
2026-09-30 07:45:47 +02:00

32 lines
1.3 KiB
Nix
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Retry shape for a oneshot that fetches a credential or certificate from
# the secret store: every 30s for 24h. Under `seal = "shamir"` an operator
# unseals BY HAND, and a fetch fails for as long as that takes, or for as
# long as the store or the gateway in front of it is down. `start-limit-hit`
# does not clear when the store comes back, so a budget shorter than the
# outage leaves the unit failed until something starts it again.
#
# `StartLimit*` are `[Unit]` settings — systemd ignores them under
# `[Service]` — and the window must exceed `RestartSec × burst` or it closes
# between attempts and the burst is never reached: 2880 × 30s is 24h inside
# a 25h window.
#
# ⚠️ A unit in auto-restart is still `activating`, so its start job stays
# queued across attempts: anything ordered `After=` it waits for as long as
# it retries, up to the full 24h.
#
# Pure attrset — NOT a NixOS module. Use from a unit definition:
#
# storeRetry = import ./lib/store-retry.nix { };
# systemd.services.foo = {
# inherit (storeRetry) startLimitBurst startLimitIntervalSec;
# serviceConfig = storeRetry.serviceConfig // { Type = "oneshot"; };
# };
{ }:
{
startLimitBurst = 2880;
startLimitIntervalSec = 90000;
serviceConfig = {
Restart = "on-failure";
RestartSec = 30;
};
}