# The swarm's telemetry collector: one per swarm, in a `swarm-otel` # nixos-container beside the swarm's other shared services. This is the # **swarm** tier; `otel.nix` is the hive tier that forwards into it. # `docs/scheduler/observability.md` owns the tier model and why the two stay separate # when co-located. # # A container rather than a second host unit, for the same reason every # sibling swarm service is one — and because `services.opentelemetry-collector` # is a singleton NixOS option, already spoken for on the host by the hive # tier. A container gets its own evaluation and therefore its own collector. { pkgs, lib, config, ... }: let cfg = config.services.hyperhive.swarm.otel; deployCfg = config.services.hyperhive.deploy; otelCfg = config.services.hyperhive.otel; vmCfg = config.services.hyperhive.swarm.victoriametrics; vlCfg = config.services.hyperhive.swarm.victorialogs; hyperhiveCfg = config.services.hyperhive; gatewayCfg = hyperhiveCfg.gateway; baoCfg = hyperhiveCfg.swarm.bao; baoDeploy = deployCfg.bao; # The collector names its components `/`, where `` is a # hive name for the per-hive pipelines and this literal for the swarm tier's # own. The two share one namespace and are merged with `//`, so a hive named # this would silently REPLACE the swarm tier's parts — and lose its own # pipeline in the process, while its receiver keeps accepting pushes. # # Bound once and interpolated at every swarm-tier use below, so the assertion # that reserves it is checking the same string the config emits. A literal # repeated at each site would let the guard and the config drift apart, which # is the failure this guard exists to prevent. # Read-only option below, not a bare literal — `swarm-controller.nix` needs # the identical string to build the same audience/endpoint, and a value # bound once here (rather than copy-pasted at both sites) is the only way # the two can't drift apart. swarmTierName = cfg.producerName; # The SECOND swarm-tier producer: the collector inside the secret store's # container (`swarm-bao.nix`). It gets an owner of its own rather than # pushing into `swarmTierName`'s route, because the audience that route # checks is `swarm-controller`'s own client id — one identity per principal, # so a second principal brings its own client, its own audience and # therefore its own authenticator and receiver. # # Named after the store's container, so `otlp/swarm-bao` reads as the thing # that pushes into it. No `reserved-names.nix` entry is needed for it the # way `swarm` has one: the name CONTAINS `swarm`, which # `nix/reserved-hive-fragments.nix` forbids as a substring of any hive name, # so no hive can ever own these components. The assertion below checks that # fragment is still listed rather than trusting it. storeProducerName = baoCfg.machine; storeClientId = baoCfg.otel.clientId; reservedHiveFragments = import ../reserved-hive-fragments.nix; # Every `` no hive may take. Read from `nix/reserved-names.nix`, the # same file the daemons are handed as `HIVE_RESERVED_NAMES`, because agent # names and hive names are ONE namespace going forward — a locally-owned # list here would be a second copy to keep in step, which is the failure a # single blacklist exists to prevent. # # The assertion below still checks the string this module emits: the file # is asserted to CONTAIN `swarmTierName`, so a rename that dropped it from # the file (or a `producerName` override the file was never updated for) # would be an eval error rather than a silently missing guard. # ⚠️ The guards themselves live in ./swarm.nix, which declares # `swarm.hives` and is unconditional. They were here and gated on this # module's own `enable`, so a swarm without the collector had no hive-name # check at all. What stays here is the assertion below: this module's own # entry must still be in that shared list. reservedOwners = import ../reserved-names.nix; # A published target is declared as ONE url, because that url is also the # audience its token is minted for — but prometheus wants the same fact in # three fields. Split it here rather than asking a service to state it # twice: two spellings of one address is a mismatch waiting to happen, and # the failure is a valid token refused at the target. # # `null` when the shape is wrong, which the assertion below reports by name. # ⚠️ The pattern requires `https`. A published endpoint reached over plain # http would ship a bearer token in clear text, and that is worth an eval # error rather than a warning nobody reads. parsePublished = url: builtins.match "https://([^/]+)(/.*)" url; # A loopback target may carry a path and a query as well as `host:port`, # because not every exporter serves `/metrics`: openbao 404s there and # answers on `/v1/sys/metrics?format=prometheus`. Split for the same reason # the published one is — one declaration, three prometheus fields. # # ⚠️ The query may not ride along in `metrics_path`: prometheus percent- # encodes the `?`, so the request goes to a path that does not exist. It has # to become `params`, whose values are lists. # # `null` when the shape is wrong, which the assertion below reports by name. parseScrape = t: builtins.match "([^/?]+)(/[^?]*)?(\\?(.+))?" t; queryParams = query: lib.listToAttrs ( map ( kv: let parts = lib.splitString "=" kv; in lib.nameValuePair (lib.head parts) [ (lib.concatStringsSep "=" (lib.tail parts)) ] ) (lib.splitString "&" query) ); # Both optional fields are omitted rather than defaulted, so a plain # `host:port` renders the config it rendered before this grammar existed. loopbackScrapeConfig = job: target: let parts = parseScrape target; path = lib.elemAt parts 1; query = lib.elemAt parts 3; in { job_name = job; static_configs = [ { targets = [ (lib.elemAt parts 0) ]; } ]; } // lib.optionalAttrs (path != null) { metrics_path = path; } // lib.optionalAttrs (query != null) { params = queryParams query; }; # The client secret takes three names, and the reason is `DynamicUser`. # # Upstream's collector unit runs with `DynamicUser = true`, and the # prometheus receiver opens `client_secret_file` ITSELF, at runtime, as that # user — so there is no stable uid to hand a file to, and the root-owned # 0400 shape `hive-matrix-oidc-secret` delivers to would be unreadable. # (Matrix gets away with it because `LoadCredential` reads the file as root # before the sandbox exists, and tuwunel never opens that path itself.) # # `LoadCredential` solves both halves: systemd reads the file as root and # re-exposes it to the dynamic user under a path that does not depend on # which uid it turned out to be. # # At rest in the container's tree — written by the host oneshot below. # Under /var/lib and not /run because the collector may start before the # delivery unit on a later boot, and a secret that evaporates on reboot # turns a working scrape into an intermittent one. collectorSecretInContainer = "/var/lib/swarm-otel-oidc/${cfg.clientId}.secret"; # The same file seen from the host, which is where the delivery unit # writes it — spelled once for the same reason ./swarm-grafana.nix spells # its own pair once: the unit that writes it and the option that names it # are a few hundred lines apart, and a collector reading a path nothing # writes is an export that fails with nothing in any log about the file. collectorHostSecretPath = "/var/lib/nixos-containers/${cfg.machine}${collectorSecretInContainer}"; collectorHostSecretDir = builtins.dirOf collectorHostSecretPath; collectorCredentialId = "oidc-client-secret"; # What the scrape config points at. ⚠️ This path and the `LoadCredential` # id below are one fact spelled twice by systemd's design — both derive from # `collectorCredentialId` so they cannot drift; a mismatch is a file the # collector cannot open, discovered at runtime and nowhere else. collectorSecretPath = "/run/credentials/opentelemetry-collector.service/${collectorCredentialId}"; atomicWriteSecret = import ./lib/atomic-write-secret.nix { }; refreshConsumer = import ./lib/refresh-consumer.nix { }; # `attrNames` is sorted, so this is a function of the hive SET and not of # the order anyone wrote it in. # # These ports are internal and appear in no URL: a hive addresses its own # receiver as a PATH on this collector's single gateway name, and nginx — # rendered from this same evaluation — is the only thing that ever names # the port. That is what makes deriving them safe here and unsafe in the # obvious other place: were a hive told a port, inserting a hive would # renumber the ones after it and silently move a port a running hive was # already sending to. hivePorts = lib.listToAttrs ( lib.imap0 (i: h: lib.nameValuePair h (cfg.port + i)) (lib.attrNames hyperhiveCfg.swarm.hives) ); # The swarm's authelia is reached by its gateway name, whose leaf is # issued by the swarm services sub-CA — so this container needs the same # runtime CA trust every other consumer of a swarm-service name needs. # The CA is generated at runtime and cannot be baked into a derivation, # which is why it arrives as a bind mount rather than # `security.pki.certificateFiles`. caTrust = import ./lib/hive-ca-trust.nix { inherit lib; tlsCfg = deployCfg.hive-controller.tls; inherit gatewayCfg; }; # `unknown` rather than omitting the label, copying `agent-modules/otel.nix` # deliberately: a producer that cannot name its swarm should say so in the # same vocabulary as every other producer, so a query never has to handle # both "the label is absent" and "the label says unknown". swarmDisplayName = if hyperhiveCfg.swarm.name == null then "unknown" else hyperhiveCfg.swarm.name; # Where the journal lives on BOTH sides of the bind mount below — one string # because the receiver reads the path it is mounted at, and two spellings of # one path is a mount that succeeds and a receiver that finds nothing. # # 🔑 Why the host's directory is enough to see a container — and why it is # only enough because that container asks for it. nixpkgs hardcodes # `--link-journal=try-guest`, which puts the journal inside the container and # leaves the host with a symlink into that container's transient root: a # reader here cannot follow it, and it dangles the moment the container # stops. Every container block except `swarm-bao` therefore sets # `extraFlags = [ "--link-journal=host" ]`, which the invocation expands # AFTER the hardcoded flag, so the files land here under their own # machine-id subdirectory and are bind-mounted into the guest instead. # # ⚠️ Drop that flag from one of THOSE containers and this receiver silently # stops seeing it — no error, just a unit that never appears in the store. # `swarm-bao` is the one exception on purpose: its journal stays inside the # container for its own collector instead (see `swarm-bao.nix`), so it was # never one of the units this receiver could see either. hostJournalDir = "/var/log/journal"; # Each store's OTLP route, by domain. One binding because the same string is # both the address requested and the audience the token is minted for — two # spellings present as a valid token refused at the store. # ⚠️ ./swarm-otel-service.nix binds both again for `swarm.otel.audience`, # the audiences the client is registered for; the two copies must agree. metricsPushUrl = "https://${vmCfg.domain}/opentelemetry/api/v1/push"; logsPushUrl = "https://${vlCfg.domain}/insert/opentelemetry/v1/logs"; # Read by the extensions block, `service.extensions` and each exporter's # authenticator. Derived rather than repeated: an authenticator naming an # unlisted extension starts clean and authenticates nothing. pushAudiences = { victoriametrics = metricsPushUrl; victorialogs = logsPushUrl; }; pushAuthenticator = name: "oauth2client/${name}"; # Whether this collector holds a credential — a property of the credential, # not of where any other service runs. Not an assertion: a collector that # pushes nowhere authenticated still receives from every hive, so refusing # to build one would make a supported shape illegal for want of a value # that only degrades what it can push. haveCollectorSecret = deployCfg.swarm-otel.clientSecretFile != null; # A reader of the store is defined by holding a certificate the store # accepts, never by standing next to it — the rule # ./glue-matrix-bao-token.nix states in full. Unlike Grafana's identical- # looking flag, this one does not gate an assertion on its own: the # store-reading unit below simply does not render without it, the same choice # ./glue-matrix-bao-token.nix and ./glue-queue-agent-credential.nix make for # their own readers, because a collector with no client identity is # `haveCollectorSecret = false` above, and that is already a supported, # merely degraded shape rather than a service with no way in at all. What # IS asserted is narrower and lives in `assertions` below — see # `hiveReaderIdentity`. # # 🩸 The collector's OWN leaf, not `clientCertFile` — the hive's, which four # units used to share. Bao matches a cert-auth role on the CN, so one leaf for # four readers was ONE principal holding the union of four grants: this unit # could read every agent credential in the swarm and Grafana's OIDC client # secret, when what it needs is the one path `storeSecretPath` names. haveClientIdentity = baoDeploy.otelOidcClientCertFile != null && baoDeploy.otelOidcClientKeyFile != null; # Does this host read the store at all — the hive's own leaf, which is the # one thing a remote-store deployment has always had to place by hand. It is # what separates the degrade the comment above describes from the mistake the # assertion below reports: a collector on a host holding NO store identity is # the supported shape, and one on a host that demonstrably reads the store # named seven of the eight options and stopped. Before the four-way split # that second host rendered this unit off `clientCertFile`, so it is a silent # regression rather than a choice anyone made. hiveReaderIdentity = baoDeploy.clientCertFile != null && baoDeploy.clientKeyFile != null; # Where the publisher on authelia's host leaves this client's secret — # composed from the same swarm-wide `clientId` the registration in # ./glue-swarm-otel-oidc-client.nix uses, so a rename cannot leave one of # them behind. The `services` segment is `swarm-secret-client`'s # `path::Kind::Service`, the same prefix ./swarm-grafana.nix reads its own # client secret under — one write grant in ./swarm-bao.nix and one hive read # grant in `policy::render` already cover it. storeSecretPath = "secret/swarm/services/${cfg.clientId}/oidc/client"; # The operator-configured upstream, named once: the same exporter carries # every signal, so metrics and logs both reach it without a second # definition. upstreamExporterName = if otelCfg.protocol == "grpc" then "otlp" else "otlphttp"; upstreamExporters = lib.optional (otelCfg.endpoint != "") upstreamExporterName; # One list, read by every pipeline: the per-hive pipelines fan out to # exactly the same destinations as the single pipeline they replace. # Written once because "which exporters" is a property of this tier, not # of which hive a sample came from. # ⚠️ Unconditional: `deploy.victoriametrics.enable` says "this host RUNS the # store", and a swarm has one either way. Gating on it left a collector # elsewhere with no exporter at all, dropping everything silently. exporterNames = upstreamExporters ++ [ "otlphttp/victoriametrics" ]; # Same fan-out for logs, unconditional for the same reason. logExporterNames = upstreamExporters ++ [ "otlphttp/victorialogs" ]; # An empty exporter list is not a quiet no-op — the collector rejects it. # Now always satisfied, as a consequence of the log store always existing. collectLogs = logExporterNames != [ ]; # Shared host netns, like every sibling swarm service: the hive tier # reaches this collector, and this collector reaches the metrics # store, without either crossing a network boundary that would need # its own trust material. privateNetwork = false; storeRetry = import ./lib/store-retry.nix { }; in { # What the collector IS to the swarm — the client it is registered as, where # it exports — is `swarm.otel` in ./swarm-otel-service.nix. The secret is a # path on the machine that runs it, so it hangs off the deployment. `enable` # already lives in ./deploy.nix, which also carries the rename. options.services.hyperhive.deploy.swarm-otel = { clientSecretFile = lib.mkOption { type = lib.types.nullOr lib.types.str; default = null; example = "/var/lib/swarm-otel-oidc/swarm-collector.secret"; description = '' Path, inside this collector's container, to its OIDC client secret. `null` means it holds no credential: it still exports, and the authenticated destinations refuse it. Set automatically where authelia is co-located; a deployment that places authelia elsewhere points this at a file it delivers itself. ⚠️ A path, never a value — a secret interpolated into a nix expression renders world-readable into the store. ''; }; }; config = lib.mkIf deployCfg.swarm-otel.enable { # The gateway name, inside `deployCfg.swarm-otel.enable` — that guard is the load-bearing # part. Every hive in a swarm may know this collector exists, but only # the host that RUNS it may claim the name; a client hive declaring the # vhost would answer for a service it does not have. services.hyperhive.gateway.localNames = [ cfg.domain ]; # This host serves the vhost, and the container behind it resolves # through the hive's dnsmasq. services.hyperhive.gateway.enable = lib.mkDefault true; services.hyperhive.gateway.dns.enable = lib.mkDefault true; # The metrics counterpart to this collector's own journal, which its # receiver ships like any other unit's: the journal says the process is # alive, these counters say whether it is dropping what it receives. # Loopback works here because `privateNetwork = false` — this container # shares the host's netns, so `127.0.0.1` is where `telemetryPort` is # bound. services.hyperhive.swarm.otel.scrapeTargets.collector = "127.0.0.1:${toString cfg.telemetryPort}"; # OTLP/HTTP, not a browsable UI, but the same reverse-proxy shape as # every sibling swarm service: TLS terminates here, then plain http to # the co-located container over loopback (shared netns, like the store # this collector writes to). services.nginx.virtualHosts."${cfg.domain}" = (gatewayCfg.lib.tlsFor cfg.domain) // { listen = gatewayCfg.lib.listen; extraConfig = gatewayCfg.lib.securityHeaders; locations = # One name for the whole collector, and the hive is a path under # it. The alternative — a vhost per hive — needs a certificate, # a DNS name and a `localNames` entry per hive to express the # same routing the gateway already does for free. # # ⚠️ The trailing slash on both sides is load-bearing: it is what # strips `/` before the request reaches the receiver, which # serves `/v1/metrics` and knows nothing about hives. Without it # the receiver sees `//v1/metrics` and answers 404 to a # request that authenticated perfectly. lib.mapAttrs' ( h: p: lib.nameValuePair "/${h}/" { proxyPass = "http://127.0.0.1:${toString p}/"; } ) hivePorts // { # The swarm-tier producer's own route — same trailing-slash # shape as the per-hive locations above, and load-bearing for # the same reason: it strips `/${producerName}` before the # request reaches the receiver, which knows nothing about the # path it was found at. "/${swarmTierName}/".proxyPass = "http://127.0.0.1:${toString cfg.producerPort}/"; # The secret store's forwarder, on its own route for its own # audience — same trailing slashes, same reason. "/${storeProducerName}/".proxyPass = "http://127.0.0.1:${toString cfg.storeProducerPort}/"; # There is no swarm-wide inbox, and a closed door is the honest # description of that. Every other route into this collector # belongs to exactly one hive or the swarm-tier producer above. "/".return = "404"; }; }; # The identities this tier authenticates against. Stated rather than # assumed: the option is an operator's to turn off, and this tier does # not work without it. services.hyperhive.swarm.authelia.oidc.hiveIdentities = true; # The collector's own identity, for the other direction: the hive # identities above are how this collector authenticates its *callers*, # this is how it authenticates *itself* to a service published behind # the gateway. # # Registering the client is NOT here any more: it has to happen on the # host that runs authelia, and this whole block is gated on the host that # runs the collector. ./glue-swarm-otel-oidc-client.nix is where it moved # to, the same split ./swarm-grafana.nix made for its own client. # THE delivery unit — one route, in every deployment. The secret authelia # minted arrives out of the swarm secret store, which the publisher on # authelia's host wrote it into, whether authelia is a network away or in # the container next door. # # 🩸 A unit here used to copy the plaintext directly out of authelia's # host tree, reachable only because they share this host's network # namespace — which stopped working the moment authelia moved to another # host, which is the defect this whole change exists to fix. The ruling # that deleted it rather than gave it a remote sibling: the store exists # so a host holds ONE out-of-band secret — its client certificate — and # reads everything else with it. Recorded in docs/swarm/secrets.md. # # Shaped after ./glue-queue-agent-credential.nix: a cert login that fails # LOUDLY, since every state it fails on is one a retry fixes, then a read # that degrades QUIETLY, since no retry turns "no value there" into a # value. # # ⚠️ Renders only where `haveClientIdentity` holds, unlike # ./swarm-grafana.nix's equivalent unit. That module refuses any host that # runs Grafana without the identity, because a Grafana with none has no way # in at all; this collector without one is `haveCollectorSecret = false` # above — already a supported, merely degraded shape, so the unit that would # fetch a credential simply does not exist rather than refusing the build # for want of one. The one case that IS refused is the host that already # holds the hive's leaf and is missing only this pair, which is a silent # regression rather than that degrade — `hiveReaderIdentity` above. systemd.services.swarm-bao-otel-oidc = lib.mkIf haveClientIdentity { description = "fetch the swarm collector's OIDC client secret from the swarm secret store"; # Every one of these names a unit that exists only where the store runs. # `Requires=` on an absent unit fails the job outright, so the ordering # is conditional even though the read is not: off-host there is nothing # local to wait for, and the timeout below bounds the attempt instead. after = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" "container@${baoCfg.machine}.service" ]; wants = lib.optionals baoDeploy.enable [ "container@${baoCfg.machine}.service" ]; requires = lib.optionals baoDeploy.enable [ "swarm-bao-pki.service" ]; before = [ "container@${cfg.machine}.service" ]; wantedBy = [ "multi-user.target" "container@${cfg.machine}.service" ]; path = [ baoDeploy.package pkgs.coreutils pkgs.systemd ]; # ./lib/store-retry.nix. The collector's container is ordered after this # unit and waits for one attempt; a secret a later attempt lands restarts # the collector (below). inherit (storeRetry) startLimitBurst startLimitIntervalSec; serviceConfig = storeRetry.serviceConfig // { Type = "oneshot"; RemainAfterExit = true; SyslogIdentifier = "swarm-bao-otel-oidc"; # What actually bounds each attempt below. Stated here rather than left # to systemd's default, so the number a boot waits on is in the file # that waits. TimeoutStartSec = 30; }; environment = { BAO_ADDR = "https://${baoCfg.domain}:${toString baoCfg.port}"; BAO_CLIENT_CERT = baoDeploy.otelOidcClientCertFile; BAO_CLIENT_KEY = baoDeploy.otelOidcClientKeyFile; } # Absent means the system trust store, which is what a deployment with a # real CA wants and what a self-signed one must not be left with. // lib.optionalAttrs (baoDeploy.serverCaFile != null) { BAO_CACERT = baoDeploy.serverCaFile; }; script = '' set -euo pipefail ${atomicWriteSecret} ${refreshConsumer} # `bao`'s own message is the only thing separating a missing value # from a refused identity from an unreachable host. This unit's # degraded mode is correct for all three, so it reports which one # rather than asserting all three in a sentence of ours. err="$(mktemp)" trap 'rm -f "$err"' EXIT # Cert auth is a login, not a transport setting. The `BAO_CLIENT_*` # variables above only decide which certificate the TLS handshake # presents; without a token `bao` asks its token helper instead, and # that is a `sh` this unit's `path` does not carry. `-token-only` # answers on stdout and skips the helper on both sides. if ! BAO_TOKEN="$(bao login -method=cert -token-only 2>"$err")"; then echo "could not log in to swarm-bao with this host's certificate; leaving the collector's OIDC client secret as it is." >&2 if [ -s "$err" ]; then cat "$err" >&2 else echo "bao failed without writing a diagnostic." >&2 fi exit 1 fi export BAO_TOKEN if ! secret="$(bao kv get -field=value ${lib.escapeShellArg storeSecretPath} 2>"$err")"; then echo "swarm-bao did not return ${storeSecretPath}; the collector has no OIDC client secret yet." >&2 if [ -s "$err" ]; then cat "$err" >&2 else echo "bao failed without writing a diagnostic." >&2 fi exit 0 fi if [ -z "$secret" ]; then echo "swarm-bao returned an empty ${storeSecretPath}; leaving the file as it is." >&2 exit 0 fi # root-owned 0400, written with a shell builtin and never handed to a # program: `printf` is bash's own, so the plaintext never becomes an # argument in /proc the way `install <<<"$secret"` or an `echo` from # `path` would. The collector runs under `DynamicUser`, so there is # no uid to give this to — `LoadCredential` reads it as root before # the sandbox exists and re-exposes it to whichever uid the unit got. install -d -m 0755 ${lib.escapeShellArg collectorHostSecretDir} if secret_differs ${lib.escapeShellArg collectorHostSecretPath} "$secret"; then atomic_write_secret 0400 root:root ${lib.escapeShellArg collectorHostSecretPath} "$secret" fi # `LoadCredential` copies the file at start only, and a collector that # started without it is in its start limit. refresh_consumer ${lib.escapeShellArg cfg.machine} opentelemetry-collector.service ${lib.escapeShellArg collectorHostSecretPath} ''; }; # Where the delivery unit above lands the secret. `mkDefault`, so a # deployment delivering it some other way just sets the option — and # `mkIf haveClientIdentity` so a host with no store identity is left with # `clientSecretFile == null`, the already-supported degrade rather than a # path nothing ever writes. services.hyperhive.deploy.swarm-otel.clientSecretFile = lib.mkIf haveClientIdentity ( lib.mkDefault collectorSecretInContainer ); # The CA bind source is written at runtime by a host unit, so the # container has to start after it — otherwise nspawn sets up a mount # over a file that does not exist yet. systemd.services."container@${cfg.machine}" = caTrust.containerOrdering; assertions = [ { # Shaped after ./swarm-grafana.nix's `haveClientIdentity` assertion — # the same refusal, named to this principal's own pair. What differs is # the `!hiveReaderIdentity ||` guard, and it is what keeps the degrade # the flag's own comment describes intact: Grafana refuses any host # that runs it without a leaf, because a Grafana with no SSO has no way # in at all, while a collector with none still receives telemetry. So # this fires only where the host already reads the store and is missing # this one pair — which before the four-way split rendered the unit off # `clientCertFile`, and now silently does not. assertion = !hiveReaderIdentity || haveClientIdentity; message = '' This host reads the swarm secret store (services.hyperhive.deploy.bao.clientCertFile is set) and runs the swarm collector, so it needs the collector's own client identity: set both services.hyperhive.deploy.bao.otelOidcClientCertFile services.hyperhive.deploy.bao.otelOidcClientKeyFile swarm-bao-otel-oidc.service fetches the collector's OIDC client secret out of the store, and without these it is not rendered at all — so this collector would authenticate to nothing, quietly, on a host that has everything it needs to fetch the secret. A collector on a host holding NO store identity is a different and supported shape: it runs without a client secret and still receives telemetry. That is not this host. ⚠️ The collector's OWN leaf, not deploy.bao.clientCertFile. That one is the hive's, and its grant reads every secret in the store; this role reads the one path this collector's client secret lives at. Pointing this option at the hive's leaf would evaluate, deploy and log in — and undo the split. On a hive that runs the store, glue-bao-tls.nix supplies both as defaults and there is nothing to do. Elsewhere the leaf is issued from that CA out of band and named here — see docs/swarm/secrets.md. ''; } # ⚠️ An assertion that this collector has "somewhere to send" was REMOVED # rather than relaxed: it read the store's *per-host* enable, so it # rejected at eval the very deployment the stores are reached by domain # for — a collector on its own host never built. { # Whether there is a journal on disk to collect from. # A `bindMounts` entry never creates its `hostPath`, and unlike # the CA bind source above there is no unit to order after — the # directory exists because journald was told to store # persistently, or not at all. `auto` is deliberately not # rejected: it uses the directory when it exists, and eval cannot # see whether it does. assertion = !collectLogs || !(lib.elem config.services.journald.storage [ "volatile" "none" ]); message = '' The swarm collector is configured to ship journal logs, but services.journald.storage is "${config.services.journald.storage}" on this host. journald only writes ${hostJournalDir} when it stores persistently: with "volatile" the journal lives in /run/log/journal, and with "none" there is none at all. This collector's journald receiver reads ${hostJournalDir} and the container bind-mounts that path, so the collector would not start at all — nixos-container refuses to start when a bind source is missing. Set services.journald.storage = "persistent" (the NixOS default), or turn off log collection by leaving both log destinations unset. ''; } { # Without a roster there are no receivers at all, so this # collector would listen on nothing while looking configured. assertion = hyperhiveCfg.swarm.hives != { }; message = '' services.hyperhive.deploy.swarm-otel.enable is true but services.hyperhive.swarm.hives is empty: ingest is authenticated per hive, so an empty roster means this collector accepts nothing from anyone. List the swarm's hives. ''; } { # The blacklist now lives in a shared file, so this module no longer # controls its contents — and a guard whose subject can be edited # elsewhere has to assert that its own case is still in there. Without # this, deleting one line from `reserved-names.nix` would silently # retire the check below rather than fail anything. assertion = lib.elem swarmTierName reservedOwners; message = '' nix/reserved-names.nix no longer contains '${swarmTierName}', which the swarm collector needs reserved: it names components `/` and uses the hive name as the owner, so a hive called '${swarmTierName}' would replace the swarm tier's own pipelines and lose its own. Put it back, or give this module a different swarmTierName. ''; } { # The store forwarder's components are named after a container whose # name merely CONTAINS the reserved word, so what protects them is # the substring list rather than the equality one. Same reason the # assertion above exists: the list lives in another file, and a guard # whose subject can be edited elsewhere has to assert its own case is # still covered. assertion = lib.any (f: lib.hasInfix f storeProducerName) reservedHiveFragments; message = '' nix/reserved-hive-fragments.nix forbids no substring of '${storeProducerName}', which the swarm collector needs kept out of the hive namespace: it names components `/`, so a hive called '${storeProducerName}' would replace the secret store forwarder's own receiver and pipeline entry. Put the fragment back, or give the store's container a name one of them covers. ''; } { # A published target that is not an `https://host/path` url. Without # this the split returns null and the failure surfaces as # `attempt to call elemAt on null` from inside the renderer, naming # neither the option nor the offending value. # # It also enforces the scheme: a published endpoint scraped over plain # http would put a bearer token on the wire in clear text. assertion = lib.all (u: parsePublished u != null) (lib.attrValues cfg.publishedScrapeTargets); message = '' services.hyperhive.swarm.otel.publishedScrapeTargets has ${ lib.concatMapStringsSep ", " (kv: "${kv.name} = \"${kv.value}\"") ( lib.filter (kv: parsePublished kv.value == null) (lib.attrsToList cfg.publishedScrapeTargets) ) }, which is not of the form https:///. The value is both the address scraped and the audience the collector's token is minted for, so it has to be the exact url a request goes to. https is required: this hop carries a bearer token, and plain http would put it on the wire in clear text. ''; } { # Same reason the published one is asserted: a value the grammar # rejects surfaces as `attempt to call elemAt on null` from inside # the renderer, naming neither the option nor the offending value. assertion = lib.all (t: parseScrape t != null) (lib.attrValues cfg.scrapeTargets); message = '' services.hyperhive.swarm.otel.scrapeTargets has ${ lib.concatMapStringsSep ", " (kv: "${kv.name} = \"${kv.value}\"") ( lib.filter (kv: parseScrape kv.value == null) (lib.attrsToList cfg.scrapeTargets) ) }, which is not of the form :[/][?=]. A trailing `?` with nothing after it is the usual cause. This option is loopback-and-unauthenticated by contract, so it takes no scheme and no credential. ''; } { # A job name used by BOTH scrape options. Within one option this # cannot happen — the module system refuses two definitions of the # same key with different values — but the two options # are separate, so nothing arbitrates between them, and they # render into a single `scrape_configs` LIST where nothing # overwrites anything: both entries ship under one `job_name`. # # The message names which side each collision came from, which is # the part a rendered-config error could not tell an operator. assertion = lib.intersectLists (lib.attrNames cfg.scrapeTargets) (lib.attrNames cfg.publishedScrapeTargets) == [ ]; message = '' services.hyperhive.swarm.otel: ${ lib.concatMapStringsSep ", " (j: "'${j}'") ( lib.intersectLists (lib.attrNames cfg.scrapeTargets) (lib.attrNames cfg.publishedScrapeTargets) ) } is declared as both a loopback scrapeTargets job and a publishedScrapeTargets job. They render into one prometheus scrape_configs list, so both entries would ship under the same job_name — nothing overwrites anything, and the two are scraped by different rules with different trust. Rename one. A target is either reachable on loopback because it shares this host, or published and reached with a credential; it should not be described as both. ''; } { # A port collision between two listeners on one host is a runtime # coin toss with nothing in any log — the failure this whole # comment budget exists to prevent. Checked against every port # reachable from here; a port some other module picks is not. # # ⚠️ `cfg.port` is deliberately absent from `others`: it is the # FIRST element of the derived range, so listing it would make this # assertion fire on every config. assertion = let derived = lib.attrValues hivePorts; others = [ cfg.telemetryPort cfg.producerPort cfg.storeProducerPort otelCfg.collector.port ] ++ lib.optional deployCfg.victoriametrics.enable vmCfg.port; all = derived ++ others; in lib.length (lib.unique all) == lib.length all; message = '' services.hyperhive.swarm.otel: the receiver range starting at port (${toString cfg.port}, one port per hive in services.hyperhive.swarm.hives) overlaps another port on this host. Every swarm container shares the host's network namespace, so two listeners claiming one port is not a build failure — it is whichever process started first, silently. Move services.hyperhive.swarm.otel.port to a free range. ''; } ]; containers.${cfg.machine} = { autoStart = true; ephemeral = false; # Journal files on the host, not inside the container: nixpkgs hardcodes # --link-journal=try-guest, and EXTRA_NSPAWN_FLAGS expands after it. extraFlags = [ "--link-journal=host" ]; inherit privateNetwork; # The upstream credential is operator-provided and lives on the host. # Read-only, and only when one is configured — binding a path that # does not exist makes nixos-container refuse to start the container, # which is a stall several layers from its cause. bindMounts = lib.optionalAttrs (otelCfg.headersCredential != null) { ${otelCfg.headersCredential} = { hostPath = otelCfg.headersCredential; isReadOnly = true; }; } # The public hive CA, read-only — only when something in here # actually verifies a swarm-service name. // caTrust.bindMount # The host's journal, read-only, and only when something will read it. # Read-only is the whole security posture of this mount: the collector # has no business writing to a journal, and one that cannot write # cannot corrupt the record it is reporting on. # # The mount is the only half we supply. Journal files are # `0640 root:systemd-journal` and the unit runs `DynamicUser`, so # reading them needs that group — which upstream's own collector unit # already grants unconditionally. Adding it here again is not # harmless: systemd list options CONCATENATE, so a second copy renders # `[ "systemd-journal" "systemd-journal" ]` and quietly becomes a # second owner of a fact upstream may later change. # # It works across the mount because that gid is FIXED at 62 in # nixpkgs' `ids.nix`. A per-host allocation would leave the host's # ownership naming a different group inside the container, and the # failure would be a receiver that starts cleanly and reads nothing. // lib.optionalAttrs collectLogs { ${hostJournalDir} = { hostPath = hostJournalDir; isReadOnly = true; }; }; config = { ... }: { # This tier is the one that resolves an operator-configured # hostname: `otel.endpoint` is an external URL, and reaching it # is the entire reason this container holds a credential. The # `/etc/resolv.conf` nixos-containers copies in is a snapshot # taken once at boot, so without this the upstream export # depends on the host's file having been right at that instant. imports = [ ../container-modules/swarm-container.nix (import ../container-modules/swarm-container-resolver.nix { inherit (config.services.hyperhive.network) bridgeIp; dnsConsumers = [ "opentelemetry-collector.service" ]; }) ] # `SSL_CERT_FILE` REPLACES the trust store rather than adding to # it, so a failed assembly yields an empty pool and every TLS # call fails while the unit looks healthy. That is why this is # the shared helper — it carries the `Requires` and the # non-empty check — and not a local `cat`. ++ [ (caTrust.trustBundle { inherit pkgs; name = cfg.machine; consumers = [ "opentelemetry-collector" ]; }) ]; system.stateVersion = config.system.stateVersion; services.hyperhive.swarmContainer = { inherit privateNetwork; }; services.opentelemetry-collector = { enable = true; package = pkgs.opentelemetry-collector-contrib; # Runs `otelcol validate` at build time. ⚠️ A parser, not a # wiring check: it accepts a receiver naming an absent # extension, and the collector then dies at startup. A green # build does not prove this config starts, never mind that a # sample arrives — which is why this module's gate pushes a # real sample through both tiers into the store. validateConfigFile = true; settings = { # One receiver per hive, and that multiplicity is forced # rather than chosen. The `hive` # label has to come from something the sender cannot write, # and the only such thing here is WHICH RECEIVER accepted # the sample: a processor cannot read the token's claims # (`from_context` reads request metadata, and asking it for # an auth claim yields nothing — silently, with a healthy # startup), and one receiver holding many credentials never # reveals which one matched. receivers = lib.mapAttrs' ( h: p: lib.nameValuePair "otlp/${h}" { protocols.http = { endpoint = "127.0.0.1:${toString p}"; auth.authenticator = "oidc/${h}"; }; } ) hivePorts # The swarm tier's OWN receiver — authenticated exactly # like a hive's, just against a different identity (the # swarm-level producer is not a hive; see # `producerPort`'s description for what its audience # checks and why). // { "otlp/${swarmTierName}".protocols.http = { endpoint = "127.0.0.1:${toString cfg.producerPort}"; auth.authenticator = "oidc/${swarmTierName}"; }; } # The secret store's forwarder, which is a swarm-tier # producer with a principal of its own — a receiver per # audience, because an `oidc` authenticator checks one. // { "otlp/${storeProducerName}".protocols.http = { endpoint = "127.0.0.1:${toString cfg.storeProducerPort}"; auth.authenticator = "oidc/${storeProducerName}"; }; } # MERGED with the per-hive receivers, never assigned over # them. A plain assignment here would drop every hive's # receiver and still render a valid config that starts # cleanly — the collector has no opinion about how many # pipelines it was supposed to have. # # Only emitted when a service has actually declared a # target: a `prometheus` receiver with nothing to scrape is # the shape this whole issue is about, a config that renders # and deploys perfectly while adding no data. // lib.optionalAttrs (cfg.scrapeTargets != { } || cfg.publishedScrapeTargets != { }) { prometheus.config.scrape_configs = lib.mapAttrsToList loopbackScrapeConfig cfg.scrapeTargets # ⚠️ `++`, so the two kinds of target land in ONE list — # which is exactly why a job name may not appear in both # options. A list concatenation does not resolve a # collision the way an attrset would: both entries ship # under the same `job_name`. The assertion above is what # stands between that and a deploy. ++ lib.mapAttrsToList ( job: url: let parts = parsePublished url; in { job_name = job; # Split from the declared url — see `parsePublished`. scheme = "https"; static_configs = [ { targets = [ (lib.elemAt parts 0) ]; } ]; metrics_path = lib.elemAt parts 1; # Prometheus-native oauth2, not the collector's # `oauth2client` extension: the prometheus receiver # takes upstream's scrape config verbatim and does not # accept an `auth` block naming a collector extension. oauth2 = { client_id = cfg.clientId; # A path, never a value — nothing here may read the # secret, or it lands in the store world-readable. client_secret_file = collectorSecretPath; token_url = "${hyperhiveCfg.swarm.authelia.url}/api/oidc/token"; # The scope authelia's bearer-authz check looks for. # The client is REGISTERED for it (the collector's # entry sets `bearerAuthz`), but prometheus asks for # no scopes unless told to, so the token came back # carrying none and every scrape was refused at # introspection with "the requested scope is # invalid, unknown, or malformed". # # Which is the rule stated directly below for the # other field, and it was written here before this # line existed: the two travel together, and a # config read cannot see the one that is MISSING. scopes = [ "authelia.bearer.authz" ]; # The audience is the target's own url, and authelia # checks it against the address being requested. # Registered ≠ requested: a client that does not ASK # for an audience gets a token with `aud: []` however # complete its registration looks. endpoint_params.audience = url; }; } ) cfg.publishedScrapeTargets; } # The host's whole journal directory, every unit in it. # # The directory is the host's rather than the swarm # containers', because the host units are where the incidents # live — the gateway's nginx, the core daemon, dnsmasq — and # none of them is a swarm container. A source restricted to # `swarm-*` would exclude the most-needed one. # # ⚠️ That directory also holds an operator's desktop session # on a hive that runs its services on a workstation. The # receiver cannot exclude it — `units`, `matches` and the rest # are all positive matches — so `filter/exclude-user-sessions` # in this receiver's pipeline is what keeps it out of the store. # # Attribution is journald's either way — these fields are # written by it rather than by the logging process. ⚠️ Use # `_MACHINE_ID` to tell machines apart, not `_HOSTNAME`: a # hostname is a config value two machines can share, where a # machine id is per-machine by construction. // lib.optionalAttrs collectLogs { journald = { directory = hostJournalDir; # The SAME operator list the agent tier attaches to its own # receiver (nix/agent-modules/otel.nix), and it transfers # without adjustment: the entry shape is the receiver's, # not the journal's. Both are this one component running # `journalctl -o json`, so `PRIORITY` is spelled and typed # identically whether the directory it was pointed at holds # a host's journal or a container's. What differs between # the two receivers is which journal, and the parser does # not read that. operators = import ../journald-severity.nix; # Persists the read cursor, so a restart resumes where # the last run stopped. Without it the receiver starts # from the journal's tail and ships only what is written # AFTER it starts: every restart silently drops whatever # landed while the collector was down — no error, no # replay, just a hole. # # Cursor and journal share a lifetime here, which is what # makes the pointer meaningful: the journal is the host's, # bind-mounted in above, and this container is # `ephemeral = false`, so the cursor under the unit's # `StateDirectory` outlives a container restart exactly as # the journal it indexes does. storage = "file_storage"; }; }; exporters = { # `metrics_endpoint`, NOT `endpoint`: the latter is a # base that otlphttp appends `/v1/metrics` to, while # VictoriaMetrics serves OTLP at # `/opentelemetry/api/v1/push`. With `endpoint` the # collector answers 200 to its own clients and posts the # samples to a path that does not exist. "otlphttp/victoriametrics" = { metrics_endpoint = metricsPushUrl; } // lib.optionalAttrs haveCollectorSecret { auth.authenticator = pushAuthenticator "victoriametrics"; }; } // lib.optionalAttrs (otelCfg.endpoint != "") { ${if otelCfg.protocol == "grpc" then "otlp" else "otlphttp"} = { endpoint = otelCfg.endpoint; } // lib.optionalAttrs (otelCfg.headersCredential != null) { # Interpolated by the collector from its environment at # runtime, never by nix: `EnvironmentFile` below is what # puts it there, so the value is not read into the store. headers.${otelCfg.collector.upstreamHeaderName} = "\${env:${otelCfg.collector.upstreamHeaderName}}"; } // lib.optionalAttrs (otelCfg.protocol == "http/json") { encoding = "json"; }; } // { # `logs_endpoint`, NOT `endpoint`: the latter is a base the # exporter appends `/v1/logs` to, and VictoriaLogs serves # `/insert/opentelemetry/v1/logs` — a different route from the # metrics store's. Both spellings pass `otelcol validate`, and # the store answers 400 to every wrong path, so neither end # tells you which one you sent. # # 🔑 The query parameters are not tuning. The journald receiver # leaves the OTLP body empty and carries the entry as journal # fields, so without `_msg_field` every record stores the # literal `missing _msg field` as its message and text search # finds nothing, while ingest keeps answering 200. # `_stream_fields` is one stream per unit per machine. # `_MACHINE_ID` is what makes that true — it is per-machine by # construction, where `_HOSTNAME` is only as distinct as the # hostnames happen to be. Keying on the hostname alone merges # the streams of any two machines sharing one, which is what # the host and containerised `opentelemetry-collector.service` # would do. "otlphttp/victorialogs" = { logs_endpoint = logsPushUrl + "?_msg_field=MESSAGE&_stream_fields=_HOSTNAME,_MACHINE_ID,_SYSTEMD_UNIT"; } // lib.optionalAttrs haveCollectorSecret { auth.authenticator = pushAuthenticator "victorialogs"; }; }; # Moves this collector's self-metrics off the built-in # default of `localhost:8888`, which the hive tier holds. # # ⚠️ `metrics.address` is the spelling that looks right and # is REJECTED by this collector version: # `'migration.MetricsConfigV030' has invalid keys: # address`. `readers` is the schema it accepts, and the # difference is a startup failure rather than a warning. service.telemetry.metrics.readers = [ { pull.exporter.prometheus = { host = "127.0.0.1"; port = cfg.telemetryPort; }; } ]; # ⚠️ An extension that is configured but not listed here is # INERT — the collector starts clean and the receiver # naming it authenticates nothing. Derived from the same # attrset as the receivers so the two cannot disagree. service.extensions = # Same rule, and the journald receiver's `storage:` above # would name a component the collector never starts. [ "file_storage" ] ++ map (h: "oidc/${h}") (lib.attrNames hivePorts) # The swarm-tier producer's own authenticator — unconditional, # same reasoning as its receiver above: `otlp/${swarmTierName}` # always exists, so the authenticator it names must too, or an # extension-not-listed startup failure follows every deploy. ++ [ "oidc/${swarmTierName}" # Unconditional for the same reason, and one level sharper: # the store runs in every swarm, so its forwarder's # receiver is as permanent as the swarm tier's own. "oidc/${storeProducerName}" ] # The push side's authenticators. Same rule as above: one an # exporter names but this list omits is INERT — the collector # starts clean and pushes unauthenticated. ++ lib.optionals haveCollectorSecret (map pushAuthenticator (lib.attrNames pushAudiences)); # Fan-out, not a choice: with both configured the same # samples go upstream AND into the swarm's store. The store # is for looking at this swarm; the upstream is for whoever # aggregates across swarms, and neither replaces the other. # `exporterNames` is shared by every pipeline — where a # sample goes is a property of this tier, not of the hive # that sent it. service.pipelines = lib.mapAttrs' ( h: _: lib.nameValuePair "metrics/${h}" { receivers = [ "otlp/${h}" ]; processors = [ "resource/${h}" ]; exporters = exporterNames; } ) hivePorts # Its OWN pipeline, and that separation is the ruling, not a # tidiness choice: every `resource/` above UPSERTS a # `hive` key, so a scraped swarm sample routed through any of # them would acquire the one label a swarm-level service must # not have. Keeping it out of them makes the absence # structural rather than something to remember to strip. // { # `otlp/${swarmTierName}` is unconditional (see the # receiver above), so this pipeline is too — a swarm-tier # producer must always have somewhere to land, unlike # `prometheus`, which only joins the receiver list once # something has actually declared a scrape target. # # ⚠️ That gate is the SAME condition the receiver is # defined under, and the two must stay spelled the same # way: a receiver defined here and attached to no # pipeline is scrape configs that render, deploy and # deliver nothing — the silent shape this file keeps # arguing against, one level up. "metrics/${swarmTierName}" = { receivers = [ "otlp/${swarmTierName}" # The store forwarder's receiver joins the swarm # tier's pipeline rather than bringing one of its own: # the stamp it needs is `resource/${swarmTierName}`'s # exact stamp — the swarm's name and deliberately no # `hive` — and a second pipeline would be that same # processor under a second name to keep in step. What # a producer is separated FOR is the credential it # proves, and that is the receiver's job, not the # pipeline's. "otlp/${storeProducerName}" ] ++ lib.optional (cfg.scrapeTargets != { } || cfg.publishedScrapeTargets != { }) "prometheus"; processors = [ "resource/${swarmTierName}" ]; exporters = exporterNames; }; } # A hive's logs, and the receiver is the SAME one its metrics # arrive on: a receiver is not per-signal, so `otlp/${h}` # feeds this pipeline and `metrics/${h}` above with no second # port, credential or audience. Which is also why the `hive` # label stays trustworthy for logs without any new mechanism # — it still comes from which receiver accepted the sample. # # The senders are agent containers, one forwarder each # (nix/agent-modules/otel.nix), via their own hive's collector. # Without this pipeline that whole path terminates here: the # receiver accepts the push, answers 200, and the records go # nowhere. // lib.optionalAttrs collectLogs ( lib.mapAttrs' ( h: _: lib.nameValuePair "logs/${h}" { receivers = [ "otlp/${h}" ]; processors = [ "resource/${h}" ]; exporters = logExporterNames; } ) hivePorts ) # Logs fan out exactly as metrics do: the swarm's store when it # runs, the operator's upstream when one is configured, both # when both. The local store is a destination rather than the # reason to collect, so turning it off leaves a hive that still # ships its journal to whoever aggregates across swarms. # # Same stamp as the scraped swarm services, for the same reason: # a host's logs belong to the swarm, and there is no honest # `hive` value to put on them. // lib.optionalAttrs collectLogs { "logs/${swarmTierName}" = { # This host's own journal, and the store container's, # which arrives over the receiver above rather than off # a disk — the store's forwarder reads a journal no # reader on another host can see. receivers = [ "journald" "otlp/${storeProducerName}" ]; processors = [ "filter/exclude-user-sessions" "resource/${swarmTierName}" ]; exporters = logExporterNames; }; }; } // { extensions = # The journald receiver's cursor store. No `StateDirectory` # is set alongside it because nixpkgs' own # `opentelemetry-collector` module already declares # `StateDirectory = "opentelemetry-collector"` (and points # `WorkingDirectory` at the same `%S` path) — systemd # therefore creates this directory, owned by the unit's # `DynamicUser`, before every start. A second declaration # here would only be a second owner of a fact upstream # states. # # `create_directory` is still load-bearing, and for a # build-time reason rather than a runtime one: # `validateConfigFile` above runs `otelcol validate` in the # nix build sandbox, where systemd has never run and the # path does not exist. Without it the validator refuses # with "directory must exist" and the build fails. { file_storage = { directory = "/var/lib/opentelemetry-collector"; create_directory = true; }; } // lib.mapAttrs' ( h: _: lib.nameValuePair "oidc/${h}" { issuer_url = hyperhiveCfg.swarm.authelia.url; # The audience this hive's client is registered to # request, and the reason one hive's token is refused by # another hive's receiver. Same expression authelia # registers it under — a second spelling here would deny # every hive, as a 401 that blames the token. audience = "${hyperhiveCfg.swarm.authelia.hiveClientPrefix}${h}"; # ⛔ DO NOT ADD `issuer_ca_path` HERE. The failure it # causes is invisible to every check we have. # # It loads only the FIRST certificate in the file it # names. The bundle assembled for this container is # `system CAs ++ hive trust bundle`, so the anchor sits # ~123rd and is never in the pool: the extension then # cannot verify authelia and the whole collector exits # `x509: certificate signed by unknown authority`, on # every start, with 125 valid certificates in the file. # # Leaving it unset makes the extension use the process # trust store, which `trustBundle` already populates via # `SSL_CERT_FILE` — and *that* consumer reads every # certificate regardless of order. One file, two # consumers, opposite parsing: the fix is to stop naming # it twice, not to reorder the bundle. # # ⚠️ Nor is pointing it at the hive trust bundle a fix: # `hive-tls.nix` writes that leading with the *hive* CA # (`nameConstraints` = this hive's domain), which cannot # issue a swarm-level name at all. # # (Kept from the original note, still true and still # worth not re-deriving: `issuer_ca_file`, `ca_file` and # `tls.ca_file` are INVALID KEYS for this extension.) } ) hivePorts # The swarm-tier producer's own authenticator. Same extension # type and same `issuer_url`/no-`issuer_ca_path` shape as every # `oidc/${h}` above — only the audience differs, and # deliberately does not reuse `hiveClientPrefix`: # `swarm-controller` is explicitly not a hive (see its own # module's `queueClientId` doc comment), so its audience is # its own client id rather than a value from the per-hive # namespace. Read via the option rather than a literal so a # rename of `queueClientId`'s default cannot silently # desync the two ends — this authenticator and the client # requesting the audience it checks both have to agree, or # a real swarm-controller push gets a healthy-looking 401. // { "oidc/${swarmTierName}" = { issuer_url = hyperhiveCfg.swarm.authelia.url; audience = config.services.hyperhive.swarm.controller.queueClientId; }; } # The secret store forwarder's own authenticator. Same shape # again, and the audience is that principal's own client id # — the self-referential form `swarm-authelia.nix` registers # a hive client under and `swarm-controller` uses above, # read from the option `glue-swarm-bao-otel-oidc-client.nix` # registers so the two ends cannot spell it differently. // { "oidc/${storeProducerName}" = { issuer_url = hyperhiveCfg.swarm.authelia.url; audience = storeClientId; }; } # The push side. Opposite direction to every `oidc/*` above — # those VALIDATE a token arriving; these OBTAIN one to send. # Hence `oauth2client` rather than `oidc`, and hence a client # id and secret rather than an issuer and an audience to check. // lib.optionalAttrs haveCollectorSecret ( lib.mapAttrs' ( name: url: lib.nameValuePair (pushAuthenticator name) { client_id = cfg.clientId; # A path, never a value. Same file the scrape side reads. client_secret_file = collectorSecretPath; token_url = "${hyperhiveCfg.swarm.authelia.url}/api/oidc/token"; # Registered is not requested: a client that does not ASK # for the scope gets a token carrying none, and one that # does not ask for an audience gets `aud: []`. scopes = [ "authelia.bearer.authz" ]; endpoint_params.audience = url; } ) pushAudiences ); # `upsert`, not `insert`: a sender that stamps its own # `hive` must be OVERWRITTEN, not deferred to. This # processor is the whole attribution boundary — the value # is a constant per receiver, so it says which hive # authenticated, not which hive claimed to be sending. processors = lib.mapAttrs' ( h: _: lib.nameValuePair "resource/${h}" { attributes = [ { key = "hive"; value = h; action = "upsert"; } # Stamped here rather than on a separate upstream-only # pipeline, which would double the pipeline count to # withhold one constant label from the local store. It # is redundant there — one VictoriaMetrics per swarm, so # every series in it already belongs to this swarm — but # a constant label multiplies no series, and it means # what LEAVES and what STAYS have the same shape. # # Upstream is where it stops being redundant: that is the # one hop where several swarms can land in one store, and # samples that cannot name their swarm collide there # exactly as hives collided here before per-hive # receivers existed. { key = "swarm"; value = swarmDisplayName; action = "upsert"; } ]; } ) hivePorts # The swarm tier's own stamp: `swarm` and deliberately NO # `hive`. A scraped swarm service belongs to the swarm, not to # any one hive, so there is no honest value to put there — and # an invented one (a sentinel, the local hive's name) would be # queried as though it meant something. # # ⚠️ Unconditional, not gated on `scrapeTargets != {} || # collectLogs`: a processor a pipeline names # but the config does not define is a collector that refuses # to start, and `metrics/${swarmTierName}` now ALWAYS exists # (its `otlp/${swarmTierName}` receiver is unconditional too, # for a swarm-tier producer to push to) — so this processor # has to exist unconditionally right alongside it. // { "resource/${swarmTierName}".attributes = [ { # The metric LABEL, a different namespace from the # component name above — deliberately not interpolated. key = "swarm"; value = swarmDisplayName; action = "upsert"; } ]; } # Drops an operator's desktop session from the host journal: # journald stamps `_SYSTEMD_SLICE=user-.slice` on every # record from a logind session scope and from `user@.service`. # System units, containers and `user.slice` itself do not match. # The store's forwarder shares the pipeline, and its container # runs no login session, so nothing of it is dropped. # # `ignore`: a record whose body is not a map cannot be indexed, # and is kept (with a warning) rather than dropped. // lib.optionalAttrs collectLogs { "filter/exclude-user-sessions" = { error_mode = "ignore"; logs.log_record = [ ''IsMatch(body["_SYSTEMD_SLICE"], "^user-[0-9]+\\.slice$")'' ]; }; }; }; }; systemd.services.opentelemetry-collector = lib.mkMerge [ # ⚠️ `StartLimit*` are `[Unit]` settings; systemd IGNORES them # under `[Service]` — in `serviceConfig` they render, deploy and # do nothing. `checks.module-eval-swarm-otel-core` pins them in # `unitConfig`, same shape as `checks.module-eval-hive-otel` for # the sibling host-tier collector. `Restart` is deliberately NOT # set: nixpkgs' own module defines it at normal priority, so a # second definition is a module-system conflict. # # This collector's `oidc/*` extensions call out to authelia at # startup, so a start that races authelia's own restart fails — # and with systemd's defaults (100ms `RestartSec`, 5-in-10s # start limit) that burns the whole allowance before authelia is # back, landing in `start-limit-hit`. The window here (12 × 5s = # 60s) has to exceed `RestartSec × burst` with margin for that # drift, same values as `otel.nix`'s host-tier collector, which # shares the same dependency. { startLimitBurst = 12; startLimitIntervalSec = 120; serviceConfig.RestartSec = 5; } # The credential file is already `NAME=value`, systemd's # EnvironmentFile format — so the secret reaches the process as # an environment variable without being read by nix, written to # the store, or passed in argv. (lib.optionalAttrs (otelCfg.headersCredential != null) { serviceConfig.EnvironmentFile = otelCfg.headersCredential; }) # The other credential, and the other direction: the one above # authenticates this collector's export onward, this one # authenticates it to a service it scrapes. # # ⚠️ Gated on the secret existing, not on what it is used for. # `LoadCredential` on a missing source is a unit that refuses to # start, and the option that names this file is only set where # something delivers it — so this follows that condition # exactly rather than restating a narrower one. Where it is false # no authenticator is rendered either, so the collector starts and # is refused by the stores rather than failing to start. (lib.optionalAttrs haveCollectorSecret { serviceConfig.LoadCredential = [ "${collectorCredentialId}:${deployCfg.swarm-otel.clientSecretFile}" ]; }) ]; }; }; }; }