nix: ship the journals of the units an apply can leave failed

The units on the path a deploy takes to TLS, the store's grants and the
swarm collector itself were not on the host collector's journald
allowlist, so an ingest outage one of them caused showed in the store
only as every source going quiet at once.

Each module names its own units, per the option's rule:
- hive-tls.nix: hive-tls-ca, swarm-services-cert
- hive-gateway: hive-gateway-self-signed-cert (self-signed mode only)
- swarm-bao.nix: the seven grant units beside
  swarm-bao-services-issuer-policy
- swarm-otel.nix: container@<machine>, and
  nixos-rebuild-switch-to-configuration, the transient unit nixos-rebuild
  runs the activation in and whose syslog lines carry its status

The module-eval arm pins each unit as both listed and defined, since a
listed name that matches nothing is silent.
This commit is contained in:
atlas 2026-09-24 13:07:44 +02:00 • committed by mara
commit aa719da571
5 changed files with 77 additions and 3 deletions

View file

@ -129,8 +129,13 @@ in
# Every request to every hyperhive service passes through here, so this
# is the one unit that can say a service was unreachable rather than
# merely quiet. Named even on hives that run no swarm collector: the
# option is inert unless one is collecting on this host.
services.hyperhive.swarm.otel.journaldUnits = [ "nginx" ];
# option is inert unless one is collecting on this host. nginx
# `Requires=` the self-signed copy, so its journal is the other half of
# why nginx did not start.
services.hyperhive.swarm.otel.journaldUnits = [
"nginx"
]
++ lib.optional useSelfSigned "hive-gateway-self-signed-cert";
assertions = [
{

View file

@ -367,6 +367,13 @@ in
# rebuild of every hive that is not the CA host, about a fallback
# that no longer happens.
# Both run on the deploy and sit on the gateway's start path, so a
# failure in either takes TLS down for every service behind it.
services.hyperhive.swarm.otel.journaldUnits = [
"hive-tls-ca"
"swarm-services-cert"
];
# Generate (and rotate) the hive CA + gateway leaf before anything
# serves it. Idempotent: the CA is created once and reused; the leaf
# is re-signed on expiry under the same CA so the anchor is stable.

View file

@ -1500,7 +1500,7 @@ in
# addresses on every command.
environment.systemPackages = [ baoCli ];
# The in-container unit plus the two host-side ones this module defines.
# The in-container unit plus the host-side ones this module defines.
# `swarm-bao-pki` and `swarm-bao-matrix-token` are declared by the glue
# modules that create them, per the option's own rule — and a name
# nothing defines is silently ignored, so naming them from here would
@ -1510,6 +1510,13 @@ in
"swarm-bao-certs"
"swarm-bao-token"
"swarm-bao-forwarder-oidc"
"swarm-bao-controller-policy"
"swarm-bao-secret-publisher-policy"
"swarm-bao-matrix-ctl-policy"
"swarm-bao-matrix-token-policy"
"swarm-bao-queue-agent-policy"
"swarm-bao-grafana-oidc-policy"
"swarm-bao-otel-oidc-policy"
"swarm-bao-services-issuer-policy"
];

View file

@ -660,9 +660,17 @@ in
# delivering says so in the store it stopped delivering to. That is
# less circular than it sounds: the failure that matters here is a
# single hive's receiver refusing pushes, not the process dying.
#
# The container's own host unit is what says the collector never started
# — a failed bind source or a `Requires=` that did not come up — which
# the process inside cannot report. The activation has no module to name
# it: nixos-rebuild runs switch-to-configuration as this transient unit,
# and its syslog lines are the deploy's start, finish or failure status.
services.hyperhive.swarm.otel.journaldUnits = [
"opentelemetry-collector"
"swarm-bao-otel-oidc"
"container@${cfg.machine}"
"nixos-rebuild-switch-to-configuration"
];
# The metrics counterpart to the journal line above, closing the same