Watch
0
0
Fork
You've already forked hyperhive
0

credential units: 24h retry shape; start a failed nginx when the cert lands

Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
This commit is contained in:
atlas 2026-09-29 23:15:29 +02:00 • committed by mara
commit b3b42d3279
10 changed files with 137 additions and 125 deletions

View file

@ -79,6 +79,7 @@ let
credentialPath = "secret/swarm/hives/${hyperhiveCfg.hiveName}/queue/agent";
atomicWriteSecret = import ./lib/atomic-write-secret.nix { };
storeRetry = import ./lib/store-retry.nix { };
in
{
options.services.hyperhive.deploy.hive-controller.queue = {
@ -172,10 +173,10 @@ in
requires = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" ];
# Ordered before hive-c0re, so no agent container renders ahead of an
# attempt at its credential. `Wants=`, not `Requires=`: a store this
# unit can't reach delays hive-c0re's start by its own start-limit
# window (`TimeoutStartSec`, retried up to `startLimitBurst` times
# below) rather than failing it — hive-c0re starts once that window
# elapses, whatever credential is or isn't on disk by then.
# unit can't reach delays hive-c0re's start for as long as this unit
# retries (up to the 24h of ./lib/store-retry.nix) rather than failing
# it — hive-c0re starts once the retries end, whatever credential is
# or isn't on disk by then.
before = [ "hive-c0re.service" ];
wantedBy = [
"multi-user.target"
@ -185,26 +186,15 @@ in
baoDeploy.package
pkgs.coreutils
];
# Sized for the race this loses, not for an unseal: `swarm-bao` comes up
# seconds before this unit asks, and the cert-auth role it logs in
# against is written seconds after, so a few short attempts cover it.
# An hours-long window would be a bet on a store that is sealed, and the
# degrade below is already correct for that.
#
# `StartLimit*` are `[Unit]` settings, so they go here and not in
# `serviceConfig` — systemd ignores them under `[Service]`. The window
# has to exceed `RestartSec × burst`.
startLimitBurst = 4;
startLimitIntervalSec = 300;
serviceConfig = {
# ./lib/store-retry.nix.
inherit (storeRetry) startLimitBurst startLimitIntervalSec;
serviceConfig = storeRetry.serviceConfig // {
Type = "oneshot";
RemainAfterExit = true;
# What actually bounds the reads below. Stated here rather than left to
# systemd's default, so the number a boot waits on is in the file that
# waits.
# What actually bounds each attempt at the reads below. Stated here
# rather than left to systemd's default, so the number a boot waits on
# is in the file that waits.
TimeoutStartSec = 30;
Restart = "on-failure";
RestartSec = 15;
};
environment = {
BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}";