nix(ci): gate the actions runner on nix-daemon readiness
After a hive-ci restart the act_runner re-registers and immediately claims any queued jobs, which can beat the in-container nix daemon coming up: its nix-daemon.socket carries ConditionPathIsReadWrite=/nix/var/nix/daemon-socket and is skipped until /nix/var is read-write, so the first nix-dependent build dispatches into a cold daemon and hangs/retries (observed: a 55m48s nix flake check vs the normal ~30s — a build-offload stall, not a code failure; a non-nix step on the same runner passed in 1s, masking it). Ordering the runner after/wants nix-daemon.socket does not help — a condition-skipped unit satisfies systemd ordering immediately. Instead add a blocking ExecStartPre that polls until the daemon actually answers (nix store ping), placed with mkBefore so it runs ahead of the upstream module's runner-registration ExecStartPre. The runner therefore cannot register or claim jobs until nix is usable. Topology-agnostic (works whether the daemon is in-container or a shared host socket); ~180s ceiling then the unit fails cleanly rather than claiming jobs into a dead daemon. Eval-proven: gitea-runner-hive's ExecStartPre is [wait-nix-daemon, gitea-register-runner-hive] in that order.
This commit is contained in:
parent
1a3f82a459
commit
58cc720173
1 changed files with 37 additions and 1 deletions
|
|
@ -406,7 +406,43 @@ in
|
|||
pkgs.nix
|
||||
];
|
||||
|
||||
systemd.services."gitea-runner-hive".path = [ pkgs.nix ];
|
||||
systemd.services."gitea-runner-hive" = {
|
||||
path = [ pkgs.nix ];
|
||||
# Gate runner start (and therefore job registration/claiming) on
|
||||
# the in-container nix daemon being reachable. After a hive-ci
|
||||
# restart the runner re-registers and immediately claims any
|
||||
# queued jobs — which can beat the nix daemon coming up: its
|
||||
# `nix-daemon.socket` carries
|
||||
# `ConditionPathIsReadWrite=/nix/var/nix/daemon-socket` and is
|
||||
# skipped until /nix/var is read-write, so the first
|
||||
# nix-dependent build dispatches into a cold daemon and
|
||||
# hangs/retries (observed: a 55m48s `nix flake check` vs the
|
||||
# normal ~30s, a build-offload stall, not a code failure).
|
||||
#
|
||||
# Ordering `after`/`wants` the socket unit does NOT fix this — a
|
||||
# condition-skipped unit satisfies systemd ordering immediately,
|
||||
# so the runner would still start before the daemon is live.
|
||||
# Instead block in ExecStartPre by polling the daemon until it
|
||||
# actually answers; this is topology-agnostic (works whether the
|
||||
# daemon is in-container or a shared host socket). `mkBefore` so
|
||||
# this runs ahead of any pre-steps the upstream module adds.
|
||||
serviceConfig.ExecStartPre = lib.mkBefore [
|
||||
(pkgs.writeShellScript "wait-nix-daemon" ''
|
||||
# Up to ~180s; the daemon is normally up within seconds, this
|
||||
# only bites in the post-restart cold window. Non-fatal shape:
|
||||
# if it never comes up the unit fails cleanly (recoverable)
|
||||
# rather than the runner claiming jobs into a dead daemon.
|
||||
for _ in $(seq 1 90); do
|
||||
if ${pkgs.nix}/bin/nix store ping >/dev/null 2>&1; then
|
||||
exit 0
|
||||
fi
|
||||
sleep 2
|
||||
done
|
||||
echo "nix daemon not reachable after 180s" >&2
|
||||
exit 1
|
||||
'')
|
||||
];
|
||||
};
|
||||
};
|
||||
};
|
||||
};
|
||||
|
|
|
|||
Loading…
Reference in a new issue