Nothing in this deployment read a Prometheus endpoint, so enabling
/metrics on a managed service added an endpoint and no data. The swarm
collector receives OTLP pushes and does not scrape; VictoriaMetrics
stores what is pushed and has no scrape config. The pipeline was entirely
push-based and every such service is pull-based.
Adds a prometheus receiver, a resource/swarm processor and a
metrics/swarm pipeline alongside the per-hive ones.
The pipeline is separate because that is the ruling, not for tidiness:
every resource/<hive> processor UPSERTS a hive key, so a scraped swarm
sample routed through one would acquire the single label a swarm-level
service must not have. Keeping it out makes the absence structural rather
than something to remember to strip, the same way hive stays a property
of which authenticated receiver accepted a push.
Targets come from an option each service fills in from its own module,
under its own enable, rather than a list assembled here. That is what
puts the scraper and the target on the same host by construction: an
entry exists only where the service that named it runs. Co-location is
true of the all-local deployment and is not a guarantee, and that is
exactly the case where assuming it is invisible.
All three additions MERGE with the per-hive attrsets rather than
replacing them. A plain assignment would drop every hive's receiver,
processor and pipeline and still render a config the collector starts
cleanly on.
Empty target set emits no receiver, no processor and no pipeline — an
enabled scraper with nothing to scrape is the inert configuration this
issue is about, and the target set ships empty here because the targets
themselves are separate issues.
Formatting verified with nix fmt. The evaluation gate is not written yet;
nothing here has been evaluated against a fixture.
The swarm collector could never verify its OIDC issuer, so it exited at
startup on every boot and the hive tier dropped every metric.
`issuer_ca_path` loads only the FIRST certificate in the file it names.
The bundle assembled for this container is `system CAs ++ hive anchors`,
so the swarm CA sits ~123rd and was never in the pool: the extension got
whichever public CA sorts first, could verify nothing of ours, and failed
`x509: certificate signed by unknown authority` — with 125 valid
certificates in the file.
Leaving the option unset makes the extension use the process trust store,
which `trustBundle` already populates via `SSL_CERT_FILE`, and that
consumer reads every certificate regardless of order. One file, two
consumers, opposite parsing; the fix is to stop naming it twice rather
than to reorder the bundle.
Measured with the deployed binary against the deployed config, varying
only the CA source: root-only OK, root-first OK, root-last FAILS,
root-second FAILS, and unset-with-SSL_CERT_FILE OK against a control that
correctly fails when the anchor is absent.
mara on the tracking issue: "we need the swarm label for upstream otel at
least (the out of swarm one)."
Stamped in the per-hive `resource` processor rather than on a separate
upstream-only pipeline, which would double the pipeline count to withhold one
constant label from the local store. It is redundant there — one metrics store
per swarm, so every series in it already belongs to this swarm — but a constant
label multiplies no series, and it means what leaves and what stays have the
same shape.
Upstream is where it stops being redundant: that is the one hop where several
swarms can land in one store, and samples that cannot name their swarm collide
there exactly as hives collided here before per-hive receivers existed.
`unknown` when unnamed rather than an absent label, copying the agent path so
a query never has to handle both "the label is missing" and "the label says
unknown".
mara, reviewing this PR: "hives always require an identity, swarm controller
and auth is not optional."
So `requireHiveIdentity` is gone rather than defaulted, and with it every
branch that had to describe an unauthenticated collector. The swarm tier now
serves per-hive receivers only, and `/` answers 404 because there is no
swarm-wide inbox to route to. A hive with no credential is a build error, not
a quieter mode.
`hivePortBase` goes too: with per-hive receivers unconditional, `port` IS the
base of the range. That keeps one documented knob instead of adding a second,
and its advice ("move it if something else claims that range") still holds.
Two assertions replace the toggle — an empty hive roster, and a null
`authelia.url`. The second matters because a guessed issuer URL evaluates
cleanly, deploys cleanly, and then refuses every hive at runtime.
⚠️ `cfg.port` is deliberately no longer compared against the derived range in
the collision assertion: it is now the range's first element, so listing it
would make that assertion fire on every config.
This also retires the asymmetry guard added earlier in review — the state it
protected against (auth off on one side, credential still set on the other)
is no longer representable.
The swarm collector accepted OTLP from anyone who could reach it, and took
the `hive` resource attribute from the payload. So any writer on the swarm
network could attribute metrics to any hive, and nothing downstream could
tell.
The label now comes from which receiver accepted the sample: one receiver
per hive, each behind an `oidc` extension verifying a token minted for that
hive's audience, each feeding a pipeline whose `resource` processor upserts
a constant. A sender cannot influence it, because the only input is which
authenticated port the bytes arrived on.
That multiplicity is forced rather than preferred. A processor cannot read
the token's claims — `from_context` reads request metadata, and asking it
for an auth claim yields nothing, silently, with a healthy startup — and
one receiver holding several credentials never reveals which one matched.
The per-hive ports are internal: a hive reaches its receiver as a path
under this collector's existing gateway name, so nginx (rendered from this
same evaluation) is the only thing that names a port. Fronting each hive
with its own vhost would need a certificate, a DNS name and a gateway entry
per hive to express routing the gateway already does.
Turning this on removes the unauthenticated receiver. While an open port
still accepts samples the per-hive receivers are decoration, so this is the
switch itself rather than a hardening layer beside it; a swarm that wants
the open receiver says so.
`hive-ca-trust.nix` grows `bundlePathFor`, because a consumer taking its own
CA argument has to name the bundle rather than just have `SSL_CERT_FILE`
exported at it.
Drops swarm.otel.url (a loopback default an operator had to override on a
split host) in favor of swarm.otel.domain -- the same
gateway.localNames + nginx-vhost-through-the-gateway shape every other
swarm service (authelia, grafana, victoriametrics, ui) already uses. The
hive tier's exporter now reaches it as https://<domain> unconditionally,
resolved locally by dnsmasq on a co-located host and over the real network
otherwise, instead of a config knob nobody sets until they hit the silent
drop.
Costs CA trust on the hive tier: otel.nix wires
lib/hive-ca-trust.nix's trustBundle with hostUnit = true on the
opentelemetry-collector host unit, the same flag #3441/#3442 added for
swarm-controller and hive-c0re.
mara, #3125 comment 58363: "go c".
All four sibling swarm containers import swarm-container-resolver.nix;
this one did not. It matters more here than most: otel.endpoint is an
operator-configured external hostname, and reaching it is the entire
reason this container holds a credential.
Also aligns two details with those siblings - the enable default is
asserted from swarm-required-services.nix with the metrics pair it
feeds, so that file remains the one place a service host is declared,
and machine is readOnly since its description already calls it a fact
rather than a knob.
The pre-push lint refuses them, and rightly: a comment that names an
issue number ages into a pointer at a closed thread. The constraint each
one carried is stated directly instead.
A collector serves its own metrics on localhost:8888 unless told
otherwise, and co-located tiers share a network namespace, so the second
one to start dies with 'bind: address already in use'.
The port appears in neither config - it is a default inside the binary -
so comparing the ports the configs name reports them distinct. A
behavioural probe found it by being unable to start the chain.
metrics.address is the spelling that looks right and is rejected by this
version ('migration.MetricsConfigV030' has invalid keys: address);
readers is the schema it accepts.
The swarm tier is the only holder of the upstream credential, the only
writer to the swarm's metrics store, and (once #3283 lands) the place
that stamps hive= from the authenticated connection rather than from
anything a sender can choose. Today one collector does both tiers' jobs,
which works only because they land on one box.
A container rather than a second host unit, for two reasons that agree:
every sibling swarm service is one, and `services.opentelemetry-collector`
is a singleton NixOS option already spoken for on the host by the hive
tier. A container gets its own evaluation and therefore its own
collector.
Port defaults to 4319, deliberately not the OTLP default 4318 the hive
tier uses: swarm containers share the host netns, and two listeners
claiming one port is not a build failure but a runtime coin toss with
nothing in any log saying so -- the same collision grafana and the forge
hit on 3000.
`url` is an option with a co-located default rather than a loopback
literal in the exporter, so a split-host deployment is a config change
instead of a code change.
Wires nothing yet: the hive tier still exports directly, and switching it
over is the next commit.