Watch
0
0
Fork
You've already forked hyperhive
0

swarm-otel: ship the whole host journal, drop user sessions after it

The swarm collector's journald receiver read only the units listed in
`services.hyperhive.swarm.otel.journaldUnits`. A unit nobody listed
never reached the store, and a misspelt entry shipped nothing without
an error. The list existed to keep an operator's desktop session out of
a store every swarm operator can read, but the receiver can only match
positively, so the only way to express "not user sessions" was to name
every service instead.

The receiver now reads the whole host journal, and a new
`filter/exclude-user-sessions` processor in the `logs/<swarm>` pipeline
drops records whose `_SYSTEMD_SLICE` is `user-<uid>.slice` (session
scopes and `user@<uid>.service`). The per-hive `logs/<hive>` pipelines
carry agent-container journals only and get no filter.

`journaldUnits` is removed with `mkRemovedOptionModule`, together with
its non-empty assertion and the entry each host module added. The four
module-eval membership checks go with it, replaced by one structural
case in swarm-otel-core.

Closes #3646
This commit is contained in:
atlas 2026-09-30 22:51:52 +02:00 • committed by mara
commit 6b1e825c0a
25 changed files with 72 additions and 324 deletions

View file

@ -259,12 +259,10 @@ in
# forwards nothing.
directory = "/var/log/journal";
storage = "file_storage";
# No `units` allowlist, unlike the swarm tier's receiver. That
# one needs one because the host's journal also holds an
# operator's own session; a container's journal is the harness
# and what the harness spawns, so there is no foreign traffic to
# filter out and an allowlist would only be a list to forget to
# update.
# No `units` allowlist, and no user-session filter like the swarm
# tier's: a container's journal is the harness and what the
# harness spawns, so there is no foreign traffic to filter out and
# an allowlist would only be a list to forget to update.
# journald's PRIORITY carries the level every line already has;
# without this the record reaches VictoriaLogs with

View file

@ -120,11 +120,6 @@ in
# every client certificate already trusting it, so a rebuild that
# "refreshed" it would lock every reader in the swarm out at once — the
# same rule the store's TPM PIN unit follows, for a sharper reason.
# Declared beside the unit it names, not in the store's module: an entry
# exists only where the unit does, and this one is minted by glue that not
# every hive runs.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-pki" ];
systemd.services.swarm-bao-pki = {
description = "mint the swarm secret store's own CA and leaves";
before = [ "swarm-bao-certs.service" ];

View file

@ -130,11 +130,6 @@ in
})
(lib.mkIf (haveClientIdentity && deployCfg.matrix.enable) {
# Same rule as the unit's own gate: this reader exists on a host that has a
# client identity and a homeserver, which is not every host that runs the
# store, so the store's module cannot name it.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-matrix-token" ];
systemd.services.swarm-bao-matrix-token = {
description = "fetch the matrix appservice token from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.

View file

@ -154,11 +154,6 @@ in
})
(lib.mkIf (deployCfg.hive-controller.enable && haveClientIdentity) {
# Same rule as the unit's own gate: this reader exists on any host holding
# a client identity, which is not every host that runs the store, so the
# store's module cannot name it.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-bao-queue-agent" ];
systemd.services.swarm-bao-queue-agent = {
description = "fetch this hive's agent queue credential from the swarm secret store";
# Every one of these names a unit that exists only where the store runs.

View file

@ -159,26 +159,6 @@ in
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
services.hyperhive.network.enable = lib.mkDefault true;
# The daemon that owns every container on this hive, and the helper it
# delegates its root operations to. An agent asking why a container did
# not come up is asking about one of these two.
#
# The four agent-side units are named here rather than by the
# `agent-modules/` that define them, which is the one case where "a
# module names its own units" cannot hold: those modules are evaluated
# inside the guest, and this option belongs to the host. This module is
# the host's only knowledge that agent containers exist at all. They are
# also exactly the units the dashboard offers as journal filters, so
# without them the store cannot answer a question the UI can ask.
services.hyperhive.swarm.otel.journaldUnits = [
"hive-c0re"
"hive-priv"
"hive-agent"
"hive-mcp-http"
"hive-bash-daemon"
"hive-matrix-daemon"
];
assertions = [
{
# The pinned claude reaches agents as a bare path, so this

View file

@ -193,11 +193,6 @@ in
}
];
# The runner's journal, named as the unit is *inside* the container —
# nixpkgs derives `gitea-runner-<instance>` from the attr below, so this
# name follows that attr rather than `cfg.name`.
services.hyperhive.swarm.otel.journaldUnits = [ "gitea-runner-hive" ];
# The runner container has a private netns of its own, so it is
# attached to the bridge and reaches the forge by name through the
# gateway vhost the assertion above already insists on.

View file

@ -466,14 +466,6 @@ in
};
config = lib.mkIf deployCfg.forgejo.enable {
# Same principle as the vhost below — this service's own surface lives
# with the service. The SSO source is named alongside forgejo because
# its failure mode is a login that silently falls back, not an error.
services.hyperhive.swarm.otel.journaldUnits = [
"forgejo"
"forgejo-sso-source"
];
# This service's own gateway surface: the vhost that fronts it and
# the name the hive resolver answers for. Declared here rather than
# in the gateway so the forge's public face lives with the forge —

View file

@ -126,17 +126,6 @@ in
# file's `let` and are not option surface.
services.hyperhive.gateway.lib = vhostLib;
# Every request to every hyperhive service passes through here, so this
# is the one unit that can say a service was unreachable rather than
# merely quiet. Named even on hives that run no swarm collector: the
# option is inert unless one is collecting on this host. nginx
# `Requires=` the self-signed copy, so its journal is the other half of
# why nginx did not start.
services.hyperhive.swarm.otel.journaldUnits = [
"nginx"
]
++ lib.optional useSelfSigned "hive-gateway-self-signed-cert";
assertions = [
{
assertion = !(cfg.tls.acme.enable && cfg.tls.certDir != null);

View file

@ -30,10 +30,6 @@ in
# address, so the resolver is itself a consumer one layer down.
services.hyperhive.network.enable = lib.mkDefault true;
# A name that stops resolving presents as every client timing out at
# once, so this unit's journal is worth reading swarm-wide.
services.hyperhive.swarm.otel.journaldUnits = [ "dnsmasq" ];
# The host asks the hive's own resolver, at the BRIDGE IP.
#
# Every container inherits a COPY of this host's `/etc/resolv.conf`

View file

@ -842,16 +842,6 @@ in
# through the hive's dnsmasq whether or not the gateway fronts it.
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# The homeserver's own journal (`tuwunel` is the unit name inside the
# container, whatever the nixpkgs option is called), plus the host-side
# oneshot that mints its OIDC secret — carrying the same `ssoLocal`
# guard the unit itself is declared under, so the list never names a
# unit this deployment does not define.
services.hyperhive.swarm.otel.journaldUnits = [
"tuwunel"
]
++ lib.optional ssoLocal "hive-matrix-oidc-secret";
# This swarm-ui quick-links entry. Gated on `gui.enable` too, not just
# `gatewayHost != null`: `/` on that vhost only serves fluffychat
# (below) when the GUI is on — otherwise the link would 404. Same

View file

@ -420,16 +420,6 @@ in
# rebuild of every hive that is not the CA host, about a fallback
# that no longer happens.
# The first two run on the deploy and sit on the gateway's start path,
# so a failure in either takes TLS down for every service behind it.
# The renewal sits on no start path, and a failure there is a leaf that
# lapses days later unless someone reads why it failed now.
services.hyperhive.swarm.otel.journaldUnits = [
"hive-tls-ca"
"swarm-services-cert"
"swarm-services-cert-renew"
];
# Generate (and rotate) the hive CA + gateway leaf before anything
# serves it. Idempotent: the CA is created once and reused; the leaf
# is re-signed on expiry under the same CA so the anchor is stable.

View file

@ -1093,19 +1093,6 @@ in
services.hyperhive.gateway.enable = lib.mkDefault true;
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# The bridge as well as authelia: it is the half that writes the identity
# store, and its refusals are returned to callers as a bare 401.
#
# `unitName`, never a literal: upstream derives the unit from the instance
# name, so the running unit is `authelia-<instance>` and a hardcoded
# "authelia" matches nothing. Nothing reports that — the receiver's `units`
# is an allowlist, so a name that matches no journal entry is silently
# absent and reads exactly like a service with nothing to say.
services.hyperhive.swarm.otel.journaldUnits = [
unitName
"swarm-authelia-bridge"
];
# Declared here rather than in the collector's module, per the option's
# own rule: an entry exists only where the service that named it runs.
#

View file

@ -2233,32 +2233,6 @@ in
# granter step, on the host where that step runs.
environment.etc."hyperhive/bao-bootstrap-policy.hcl".source = ./bao-bootstrap-policy.hcl;
# The in-container unit plus the host-side ones this module defines.
# `swarm-bao-pki` and `swarm-bao-matrix-token` are declared by the glue
# modules that create them, per the option's own rule — and a name
# nothing defines is silently ignored, so naming them from here would
# read as coverage on hives that have neither.
services.hyperhive.swarm.otel.journaldUnits = [
"openbao"
"swarm-bao-certs"
"swarm-bao-token"
"swarm-bao-forwarder-oidc"
"swarm-bao-granter-role"
"swarm-bao-controller-policy"
"swarm-bao-secret-publisher-policy"
"swarm-bao-matrix-ctl-policy"
"swarm-bao-matrix-token-policy"
"swarm-bao-queue-agent-policy"
"swarm-bao-grafana-oidc-policy"
"swarm-bao-otel-oidc-policy"
"swarm-bao-forwarder-oidc-policy"
"swarm-bao-services-issuer-policy"
"swarm-bao-nats-tls-policy"
"swarm-bao-agent-pki"
"swarm-bao-nats-auth-policy"
"swarm-bao-operator-viewer-policy"
];
# 🚫 No `swarm.otel.scrapeTargets.bao` entry any more, and its absence is
# the deliverable rather than a tidy-up. That option is read only by the
# SWARM collector, over loopback, so declaring the store there meant its
@ -2266,14 +2240,6 @@ in
# nowhere else — silently, since a store with no entry looks identical to
# one nothing scrapes. The scrape moved into this container's own
# collector (see `receivers.prometheus` below), which travels with the
# store.
#
# Logs are a different story from that move: `journaldUnits` above still
# names this store's units, but bao's journal now lives only inside the
# container (see the forwarder's own comment below) — nothing links it
# into the host tree any more, so the shared collector can never see
# these units. The list stays untouched for the sibling containers that
# still ride it; bao's own entries come out once delivery through the
# container's own collector is confirmed.
# ⚠️ The one nginx exception to this file's header, and it is one because
@ -3701,18 +3667,14 @@ in
# is supposed to ship its own logs to the next hop rather than
# leave a *different* container's collector to find them.
#
# The journal lives inside this container now, the same shape as
# every agent container: nothing links it into the host's
# /var/log/journal tree, so `swarm.otel.journaldUnits` near the
# top of this file can no longer reach any of bao's units through
# the shared collector — see the comment there.
# The journal lives inside this container, the same shape as every
# agent container: nothing links it into the host's
# /var/log/journal tree, so the shared collector's journal
# receiver never sees any of bao's units.
#
# ⚠️ NO `units` allowlist here, and that is the point rather than a
# simplification. A shared collector needs one because the journal
# it reads holds six containers' units plus the host's own; this
# one reads only openbao, the UI's nginx and the two oneshots, so
# there is nothing foreign to separate out — and a list of unit names
# is a thing to get wrong, which ships nothing while looking healthy.
# ⚠️ NO `units` allowlist here: this journal holds only openbao,
# the UI's nginx and the two oneshots, and a list of unit names is
# a thing to get wrong, which ships nothing while looking healthy.
assertions = [
{
# Sibling of the agent forwarder's identical assertion, and it

View file

@ -107,10 +107,6 @@ in
};
config = lib.mkIf cfg.autoConfigure {
# A CA that fails to issue is invisible until something makes a TLS call
# hours later, so this oneshot is worth more than most services.
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-ca" ];
systemd.services.swarm-ca = {
description = "Generate the swarm root CA when absent";
wantedBy = [ "multi-user.target" ];

View file

@ -643,8 +643,6 @@ in
};
config = lib.mkIf deployCfg.swarm-controller.enable {
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-controller" ];
users.users.swarm-controller = {
isSystemUser = true;
group = "swarm-controller";

View file

@ -525,20 +525,6 @@ in
services.hyperhive.gateway.enable = lib.mkDefault true;
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# The secret oneshots as well as grafana itself: each runs before it and
# fails in ways grafana then reports only as a login that does not work.
#
# Unconditional, because every unit named here now renders in every
# deployment. This list used to be assembled with `lib.optional` per
# delivery route, which was the right shape while there were two — a unit
# name that never renders is a journald scrape target matching nothing,
# which reads as a quiet unit rather than an absent one.
services.hyperhive.swarm.otel.journaldUnits = [
"grafana"
"swarm-grafana-secret-key"
"swarm-bao-grafana-oidc"
];
services.hyperhive.swarm.controller.links = [
{
label = "Grafana";

View file

@ -561,14 +561,6 @@ in
};
config = lib.mkIf deployCfg.nats.enable {
# The responder as well as the server: a denial reaches the client as a
# timeout, so the server's own log is the only place it is an error.
services.hyperhive.swarm.otel.journaldUnits = [
"nats"
"swarm-nats-auth"
"swarm-bao-nats-tls"
];
# Resolver only: NATS speaks its own protocol, so nginx fronts
# nothing here — but the auth responder introspects authelia by name.
services.hyperhive.gateway.dns.enable = lib.mkDefault true;

View file

@ -314,6 +314,15 @@ let
storeRetry = import ./lib/store-retry.nix { };
in
{
imports = [
(lib.mkRemovedOptionModule [ "services" "hyperhive" "swarm" "otel" "journaldUnits" ] ''
The swarm collector no longer takes a unit allowlist: it ships the
journal of every system unit on the host and drops only user
sessions (records under a `user-<uid>.slice`). Delete the setting;
nothing replaces it.
'')
];
# `enable` moved to `services.hyperhive.deploy.swarm-otel.enable` — see ./deploy.nix.
# ⚠️ That is the SWARM collector. The per-hive one keeps its own
# `services.hyperhive.otel.enable` (./otel.nix) and is a different
@ -572,32 +581,6 @@ in
'';
};
journaldUnits = lib.mkOption {
type = lib.types.listOf lib.types.str;
default = [ ];
example = [ "nginx" ];
description = ''
systemd units whose journal this collector ships to the swarm's
log store. Nothing outside this list is collected.
Every module that defines a unit worth reading swarm-wide adds
its own names here, rather than one list naming them all: a
service that is not running contributes nothing, and a service
added later arrives with its units already declared. A central
list would be a second place that has to know which services
exist, and it would go stale in the direction that hides a
service's logs rather than the one that shows too many.
Names are matched as journald `_SYSTEMD_UNIT` values, so an
in-container unit is named exactly as it is inside its container
— `authelia`, not `container@swarm-authelia`.
⚠️ A name that matches nothing is not an error anywhere. The unit
simply never appears in the store, which looks the same as a
service that had nothing to say.
'';
};
clientId = lib.mkOption {
type = lib.types.str;
readOnly = true;
@ -667,28 +650,12 @@ in
services.hyperhive.gateway.enable = lib.mkDefault true;
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# This collector reads its own journal, so a pipeline that stops
# delivering says so in the store it stopped delivering to. That is
# less circular than it sounds: the failure that matters here is a
# single hive's receiver refusing pushes, not the process dying.
#
# The container's own host unit is what says the collector never started
# — a failed bind source or a `Requires=` that did not come up — which
# the process inside cannot report. The activation has no module to name
# it: nixos-rebuild runs switch-to-configuration as this transient unit,
# and its syslog lines are the deploy's start, finish or failure status.
services.hyperhive.swarm.otel.journaldUnits = [
"opentelemetry-collector"
"swarm-bao-otel-oidc"
"container@${cfg.machine}"
"nixos-rebuild-switch-to-configuration"
];
# The metrics counterpart to the journal line above, closing the same
# gap from the other side: the journal says the process is alive, these
# counters say whether it is dropping what it receives. Loopback works
# here because `privateNetwork = false` — this container shares the
# host's netns, so `127.0.0.1` is where `telemetryPort` is bound.
# The metrics counterpart to this collector's own journal, which its
# receiver ships like any other unit's: the journal says the process is
# alive, these counters say whether it is dropping what it receives.
# Loopback works here because `privateNetwork = false` — this container
# shares the host's netns, so `127.0.0.1` is where `telemetryPort` is
# bound.
services.hyperhive.swarm.otel.scrapeTargets.collector = "127.0.0.1:${toString cfg.telemetryPort}";
# OTLP/HTTP, not a browsable UI, but the same reverse-proxy shape as
@ -936,26 +903,7 @@ in
# rejected at eval the very deployment the stores are reached by domain
# for — a collector on its own host never built.
{
# An empty allowlist is not an empty collection: the receiver
# renders no unit filter at all and reads the host's entire
# journal, an operator's desktop session included. Fail-open, and
# silent — so the one configuration that must not be reachable is
# "logs are being shipped and nothing said which".
assertion = !collectLogs || cfg.journaldUnits != [ ];
message = ''
The swarm collector has a log destination but
services.hyperhive.swarm.otel.journaldUnits is empty. An empty
list is not "collect nothing" — it is no filter at all, and this
collector would ship every unit on the host, including any
operator user session, to a store the whole swarm can read.
Name the units worth collecting, or turn off log collection by
leaving both log destinations unset.
'';
}
{
# The assertion above covers *which* units to collect; this one
# covers whether there is a journal on disk to collect them from.
# Whether there is a journal on disk to collect from.
# A `bindMounts` entry never creates its `hostPath`, and unlike
# the CA bind source above there is no unit to order after — the
# directory exists because journald was told to store
@ -1332,8 +1280,7 @@ in
}
) cfg.publishedScrapeTargets;
}
# The host's whole journal directory, read through a unit
# allowlist.
# The host's whole journal directory, every unit in it.
#
# The directory is the host's rather than the swarm
# containers', because the host units are where the incidents
@ -1341,13 +1288,11 @@ in
# none of them is a swarm container. A source restricted to
# `swarm-*` would exclude the most-needed one.
#
# But that directory also holds every other unit on the host,
# including an operator's `user@.service` session on a hive
# that runs its services on a workstation. The allowlist is
# what separates "this host's services" from "this host", and
# only a per-unit one can: the receiver has no system-only
# switch, and its `matches` field is an allowlist too, so
# "everything except user sessions" is not expressible.
# ⚠️ That directory also holds an operator's desktop session
# on a hive that runs its services on a workstation. The
# receiver cannot exclude it — `units`, `matches` and the rest
# are all positive matches — so `filter/exclude-user-sessions`
# in this receiver's pipeline is what keeps it out of the store.
#
# Attribution is journald's either way — these fields are
# written by it rather than by the logging process. ⚠️ Use
@ -1357,7 +1302,6 @@ in
// lib.optionalAttrs collectLogs {
journald = {
directory = hostJournalDir;
units = cfg.journaldUnits;
# The SAME operator list the agent tier attaches to its own
# receiver (nix/agent-modules/otel.nix), and it transfers
# without adjustment: the entry shape is the receiver's,
@ -1365,8 +1309,8 @@ in
# `journalctl -o json`, so `PRIORITY` is spelled and typed
# identically whether the directory it was pointed at holds
# a host's journal or a container's. What differs between
# the two receivers is which journal and which units —
# neither of which the parser reads.
# the two receivers is which journal, and the parser does
# not read that.
operators = import ../journald-severity.nix;
# Persists the read cursor, so a restart resumes where
# the last run stopped. Without it the receiver starts
@ -1577,7 +1521,10 @@ in
"journald"
"otlp/${storeProducerName}"
];
processors = [ "resource/${swarmTierName}" ];
processors = [
"filter/exclude-user-sessions"
"resource/${swarmTierName}"
];
exporters = logExporterNames;
};
};
@ -1754,6 +1701,23 @@ in
action = "upsert";
}
];
}
# Drops an operator's desktop session from the host journal:
# journald stamps `_SYSTEMD_SLICE=user-<uid>.slice` on every
# record from a logind session scope and from `user@<uid>.service`.
# System units, containers and `user.slice` itself do not match.
# The store's forwarder shares the pipeline, and its container
# runs no login session, so nothing of it is dropped.
#
# `ignore`: a record whose body is not a map cannot be indexed,
# and is kept (with a warning) rather than dropped.
// lib.optionalAttrs collectLogs {
"filter/exclude-user-sessions" = {
error_mode = "ignore";
logs.log_record = [
''IsMatch(body["_SYSTEMD_SLICE"], "^user-[0-9]+\\.slice$")''
];
};
};
};
};

View file

@ -163,8 +163,6 @@ in
};
config = lib.mkIf active {
services.hyperhive.swarm.otel.journaldUnits = [ "swarm-secret-publish" ];
# Re-publish when authelia rotates a secret. The mint writes the file, so
# the file is the event — there is no signal from authelia to subscribe to.
systemd.paths.swarm-secret-publish = {

View file

@ -162,10 +162,6 @@ in
services.hyperhive.gateway.enable = lib.mkDefault true;
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# Declared here rather than in the collector, so this store's logs are
# collected because it runs, not because a list elsewhere remembered it.
services.hyperhive.swarm.otel.journaldUnits = [ "victorialogs" ];
services.hyperhive.swarm.controller.links = [
{
label = "Logs";

View file

@ -112,10 +112,6 @@ in
services.hyperhive.gateway.enable = lib.mkDefault true;
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
# Declared here rather than in the collector, so this store's logs are
# collected because it runs, not because a list elsewhere remembered it.
services.hyperhive.swarm.otel.journaldUnits = [ "victoriametrics" ];
services.hyperhive.swarm.controller.links = [
{
label = "Metrics";

View file

@ -160,14 +160,6 @@ let
&& lib.hasInfix "${dir}/secret" s
&& lib.hasInfix "${dir}/client_id" s;
}
{
# A reader off the store's host is a reader whose journal is the only
# record of why a hive's agents never connected, so the collector has to
# be told the unit exists. Nothing else can say it: the store's module
# does not know who holds a certificate.
name = "the queue credential reader's journal reaches the collector";
ok = builtins.elem "swarm-bao-queue-agent" baoRemoteReader.services.hyperhive.swarm.otel.journaldUnits;
}
{
# No agent container may render before this unit has had its attempts,
# and the edge that guarantees it must delay hive-c0re rather than sink

View file

@ -282,10 +282,9 @@ let
ok = !(lib.elem "--link-journal=host" (baoWithCollector.containers.swarm-bao.extraFlags or [ ]));
}
{
# The whole journal, which is what the shared collector's unit allowlist
# is not. A `units` list here would render and deploy perfectly while
# shipping only the units someone remembered to name — the failure this
# forwarder exists to end.
# The whole journal. A `units` list here would render and deploy
# perfectly while shipping only the units someone remembered to name —
# the failure this forwarder exists to end.
name = "the store's forwarder filters no units";
ok = !((baoForwarder baoWithCollector).settings.receivers.journald ? units);
}

View file

@ -92,10 +92,6 @@ let
lib.attrValues (removeAttrs units [ triggeredName ])
));
}
{
name = "the renewal's journal ships beside the boot issuance's";
ok = lib.elem triggeredName allLocal.services.hyperhive.swarm.otel.journaldUnits;
}
{
# Moved with the role, in both places: the threshold is half of what the
# store grants, and the store grants what the option says.

View file

@ -62,53 +62,24 @@ let
deploy.bao.otelOidcClientKeyFile = "/etc/pki/bao-otel-oidc-key.pem";
};
# The collector beside the store holding a bootstrap token: the one shape in
# which every unit an apply can leave failed renders on the same host.
otelApplyPath = hive {
deploy.swarm-otel.enable = true;
deploy.authelia.enable = true;
deploy.bao.enable = true;
deploy.bao.bootstrapTokenFile = "/run/secrets/bao-bootstrap.token";
};
# Units on the path a deploy takes to TLS, the store's grants and the
# collector itself. A failure among them silences ingest, so without their
# journals the store can show that ingest stopped but not which unit
# stopped it.
applyPathUnits = [
"container@swarm-otel"
"hive-tls-ca"
"swarm-services-cert"
"hive-gateway-self-signed-cert"
"swarm-bao-granter-role"
"swarm-bao-controller-policy"
"swarm-bao-secret-publisher-policy"
"swarm-bao-matrix-ctl-policy"
"swarm-bao-matrix-token-policy"
"swarm-bao-queue-agent-policy"
"swarm-bao-grafana-oidc-policy"
"swarm-bao-otel-oidc-policy"
"swarm-bao-services-issuer-policy"
];
cases = [
{
# Listed AND defined, because a name that matches nothing is not an
# error anywhere: a unit renamed out from under its entry would pass a
# membership check and still never reach the store.
name = "every unit on the apply path is defined and ships its journal";
# The host journal is read whole, so the filter is the only thing
# keeping an operator's desktop session out of the store. A processor a
# pipeline names but the config does not define is a collector that
# refuses to start, hence both halves.
name = "the host journal ships every unit, through the user-session filter";
ok =
let
m = otelApplyPath;
s = otelSettings otelNoStores;
journalPipelines = lib.filter (p: builtins.elem "journald" p.receivers) (
lib.attrValues s.service.pipelines
);
in
lib.all (
u: builtins.elem u m.services.hyperhive.swarm.otel.journaldUnits && m.systemd.services ? ${u}
) applyPathUnits;
}
{
# Transient, so nothing here defines it and only membership can be
# pinned: nixos-rebuild names the unit it runs the activation in.
name = "the activation's journal ships beside the units it starts";
ok = builtins.elem "nixos-rebuild-switch-to-configuration" otelApplyPath.services.hyperhive.swarm.otel.journaldUnits;
!(s.receivers.journald ? units)
&& journalPipelines != [ ]
&& lib.all (p: builtins.elem "filter/exclude-user-sessions" p.processors) journalPipelines
&& s.processors."filter/exclude-user-sessions".logs.log_record or [ ] != [ ];
}
{
# 🩸 The arm that guards the ruling this slice landed under, the