credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.
- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
hive-agent-bao-identity and hive-agent-forge-token use it.
queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
forge-token.nix is a fetch unit with the same short budget that was
added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
(--no-block) a loaded nginx that is not active; an active nginx keeps
the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
fetch unit, swarm-services-cert included.
Refs #4662
This commit is contained in:
parent
3db1233da0
commit
b3b42d3279
10 changed files with 137 additions and 125 deletions
|
|
@ -299,6 +299,7 @@ let
|
|||
// lib.optionalAttrs (baoDeploy.serverCaFile != null) {
|
||||
BAO_CACERT = baoDeploy.serverCaFile;
|
||||
};
|
||||
storeRetry = import ./lib/store-retry.nix { };
|
||||
in
|
||||
{
|
||||
# Host-side TLS trust root for the self-signed gateway mode.
|
||||
|
|
@ -751,21 +752,14 @@ in
|
|||
];
|
||||
wants = [ "container@${baoCfg.machine}.service" ];
|
||||
path = servicesCertPath;
|
||||
# Sized like the store's own granting units, and for the same
|
||||
# reason: under `seal = "shamir"` an operator unseals BY HAND, and
|
||||
# the login below fails for as long as that takes. 2880 × 30s is
|
||||
# 24h inside a 25h window — `StartLimit*` are `[Unit]` settings, so
|
||||
# the window must exceed `RestartSec × burst` or it closes between
|
||||
# attempts and the burst is never reached.
|
||||
startLimitBurst = 2880;
|
||||
startLimitIntervalSec = 90000;
|
||||
serviceConfig = {
|
||||
# ./lib/store-retry.nix: the login below fails for as long as the
|
||||
# store is sealed or down.
|
||||
inherit (storeRetry) startLimitBurst startLimitIntervalSec;
|
||||
serviceConfig = storeRetry.serviceConfig // {
|
||||
Type = "oneshot";
|
||||
RemainAfterExit = true;
|
||||
UMask = "0077";
|
||||
SyslogIdentifier = "swarm-services-cert";
|
||||
Restart = "on-failure";
|
||||
RestartSec = 30;
|
||||
};
|
||||
environment = servicesCertEnvironment;
|
||||
script = ''
|
||||
|
|
@ -926,18 +920,18 @@ in
|
|||
mv -f "$svcroot.new" "$svcroot"
|
||||
chmod 0644 "$svcroot"
|
||||
|
||||
# Propagation, for the retry and renewal paths only. Ordered before
|
||||
# both of these, so on a normal boot they have not run yet,
|
||||
# `is-active` is false, and ordering alone does the work. What this
|
||||
# covers is the store coming up hours after the gateway did, and
|
||||
# `swarm-services-cert-renew` rotating the leaf under a running
|
||||
# gateway: nginx serves a *copy* of the leaf, so re-issuing the
|
||||
# source changes nothing until the copy is remade.
|
||||
# Propagation. nginx serves a *copy* of the leaf, so re-issuing the
|
||||
# source changes nothing until the copy is remade. A running
|
||||
# gateway (`swarm-services-cert-renew` rotating the leaf) gets the
|
||||
# copy remade and a reload. A gateway that is not running gets
|
||||
# started: a start job still waiting on this unit absorbs the
|
||||
# request, and a start that already ended on this unit's
|
||||
# `requiredBy` has nothing else that would start it again.
|
||||
#
|
||||
# ⚠️ `--no-block`, and it is not a preference. This unit declares
|
||||
# `Before=` both of these, so a blocking `systemctl restart`
|
||||
# enqueues a job that systemd will not start until this unit is
|
||||
# active — and this unit is not active until its ExecStart
|
||||
# `Before=` both of these, so a blocking `systemctl restart` or
|
||||
# `start` enqueues a job that systemd will not start until this
|
||||
# unit is active — and this unit is not active until its ExecStart
|
||||
# returns, which is waiting on that job. A deadlock, held until
|
||||
# the 24h retry window's `TimeoutStartSec` fires. Queueing the
|
||||
# job and letting it run once we exit is the only ordering that
|
||||
|
|
@ -952,6 +946,10 @@ in
|
|||
echo "swarm-services leaf rotated — re-importing and reloading nginx"
|
||||
systemctl restart --no-block hive-gateway-self-signed-cert.service
|
||||
systemctl reload --no-block nginx.service
|
||||
elif [ "$(systemctl show -P LoadState nginx.service)" = loaded ]; then
|
||||
echo "swarm-services leaf landed with nginx not running — starting it"
|
||||
systemctl reset-failed nginx.service
|
||||
systemctl start --no-block nginx.service
|
||||
fi
|
||||
|
||||
# The bundle is assembled by `hive-tls-ca`, which runs BEFORE this
|
||||
|
|
|
|||
Loading…
Reference in a new issue