credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.
- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
hive-agent-bao-identity and hive-agent-forge-token use it.
queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
forge-token.nix is a fetch unit with the same short budget that was
added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
(--no-block) a loaded nginx that is not active; an active nginx keeps
the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
fetch unit, swarm-services-cert included.
Refs #4662
This commit is contained in:
parent
3db1233da0
commit
b3b42d3279
10 changed files with 137 additions and 125 deletions
|
|
@ -49,6 +49,8 @@ let
|
|||
# The store's address is the whole switch, as in ./bao.nix and
|
||||
# ./queue-identity.nix.
|
||||
configured = cfg.addr != null;
|
||||
|
||||
storeRetry = import ../host-modules/lib/store-retry.nix { };
|
||||
in
|
||||
{
|
||||
options.services.hyperhive.agent.forge.tokenFile = lib.mkOption {
|
||||
|
|
@ -90,9 +92,9 @@ in
|
|||
pkgs.coreutils
|
||||
pkgs.diffutils
|
||||
];
|
||||
startLimitBurst = 4;
|
||||
startLimitIntervalSec = 300;
|
||||
serviceConfig = {
|
||||
# ../host-modules/lib/store-retry.nix.
|
||||
inherit (storeRetry) startLimitBurst startLimitIntervalSec;
|
||||
serviceConfig = storeRetry.serviceConfig // {
|
||||
Type = "oneshot";
|
||||
# Not `RemainAfterExit`: the timer below has to be able to start this
|
||||
# unit again, and an active unit cannot be started.
|
||||
|
|
@ -100,8 +102,6 @@ in
|
|||
# in it, alive between runs instead.
|
||||
RemainAfterExit = false;
|
||||
TimeoutStartSec = 30;
|
||||
Restart = "on-failure";
|
||||
RestartSec = 15;
|
||||
User = agentName;
|
||||
Group = agentName;
|
||||
RuntimeDirectory = runtimeDir;
|
||||
|
|
|
|||
Loading…
Reference in a new issue