Commit graph hyperhive/nix/host-modules
Author SHA1 Message Date
atlas
792d7f503f forge, matrix: SSO is not optional
Both services carried an `sso.enable` defaulting to false, so a swarm's
own forge and homeserver shipped with their identity provider switched
off unless an operator remembered two lines. Grafana never had the
toggle and is the shape the other two now match.

Behaves as if the setting were true: `ssoLocal` loses one conjunct, the
three assertions become unconditional, and the login source and
identity_provider render always.

The option is removed rather than defaulted, so a config that turned SSO
OFF fails where that line is instead of silently gaining a login
provider on the next rebuild.
2026-08-24 23:06:25 +02:00
atlas
4336436457 swarm-otel: collect only the units the swarm's services declare
The journald receiver was configured with a directory and no filter, so
the swarm's log store received every unit on the host that runs the
collector. On a hive whose services live on a workstation that includes
the operator's desktop session, in a store every swarm operator can read.

The receiver has no system-only switch and its `matches` field is an
allowlist too, so what the swarm collects has to be stated rather than
excluded. Each service module names its own units: a service that is not
running contributes nothing, and one added later arrives declared.

An empty list is fail-open — the receiver renders no filter at all and
reads everything — so it is asserted against.
2026-08-24 22:05:45 +02:00
atlas
923638ab98 fix(swarm-grafana): drop the Loki-only Logs Drilldown app
grafana-lokiexplore-app cannot query VictoriaLogs. Its volume views call
/loki/api/v1/index/volume, which VictoriaLogs does not implement and
answers "unsupported path requested", and upstream's position is that
the datasource plugin must provide drilldown support. No configuration
here changes that, so the app only ever offered a menu entry that looks
like a broken feature rather than an absent one.

The option's own description already ruled its siblings out for fronting
backends this swarm does not run. That premise expired when the swarm
gained a log store, which is what made adding this one look reasonable —
so the reason is rewritten rather than the list alone: Logs Drilldown is
excluded on compatibility, not on absence.

Explore with the VictoriaLogs datasource is the log browser, and
docs/swarm/services.md now says so where an operator looks for it.
2026-08-24 21:29:08 +02:00
atlas
228842b9d8 feat(swarm-grafana): scrape Grafana's own metrics over a loopback listener
Grafana was the one swarm service that exports Prometheus metrics and had no
scrape target, so the swarm's metrics UI was invisible to the metrics it
displays.

It serves on a unix socket and claims no TCP port, which is deliberate and
stays that way; a prometheus target has to be a host:port, so nginx re-serves
the one endpoint a scraper needs on loopback. An exact-match location, not a
prefix: widening it would re-serve the whole UI without authorization.

The metrics section is now stated explicitly rather than inherited, because a
scrape target depends on it and a changed upstream default would take the
endpoint away while nginx kept answering.
2026-08-24 19:59:46 +02:00
iris
ed95e692e1 swarm: publish an authenticated gateway vhost for VictoriaLogs
Adds services.hyperhive.swarm.victorialogs.domain and a gateway vhost
gated by the same auth_request check against authelia that swarm-uis
own vhost uses (same swarmAuthRequest shape, copied not shared - see
the file top comment for why). The stores own listener stays
loopback-only and unauthenticated exactly as before; the collector
still writes to it directly, never through this vhost, so this only
adds a new authenticated read path.

Registers the new domain in swarm.nixs serviceDomains so the
swarm-services sub-CA issues for it (a missed entry silently falls
back to the hive leaf, which cannot cover a name under a different
apex - swarm.nixs own comment on that list documents the incident
this caused before).

Verified with a throwaway module-eval (same technique as the flake
module-eval check): vhost only exists when victorialogs.enable is
set, forceSSL/no addSSL, both locations present, domain correctly
registered/absent from serviceDomains, links entry present.
2026-08-24 18:43:26 +02:00
atlas
17288589ef feat(swarm-grafana): provision the log store as a datasource, with logs drilldown
Grafana could reach the swarm's metrics and not its logs, so the store that
landed with the collector pipeline had no reader.

Adds the VictoriaLogs datasource plugin and the Logs Drilldown app, and
provisions the datasource beside the metrics one. The drilldown matters as much
as the connection: an unfamiliar log stream is explorable without writing a
LogsQL query first, which is the difference between a store you can query and
one you can use.
2026-08-24 18:27:44 +02:00
atlas
012ab5bb37 fix(swarm-otel): stop re-granting the journal group upstream already grants
nixpkgs' collector unit already sets SupplementaryGroups to systemd-journal
unconditionally, with a comment saying why. Systemd list options concatenate,
so this module's copy rendered ["systemd-journal" "systemd-journal"] and made
this a second owner of a fact upstream may later change.

The bind mount stays, since that half is genuinely ours.
2026-08-24 17:27:47 +02:00
atlas
f46ef39ef3 feat(swarm-otel): fan logs out like metrics, so the local store is optional
The logs half was gated on the swarm's log store being enabled, so turning
that store off stopped collection entirely rather than leaving the upstream
export. Logs now fan out exactly as metrics do: the store when it runs, the
operator's upstream when one is configured, both when both.

The receiver, the journal mount and the group grant follow whether there is
anywhere to send logs, not whether the local store exists. Tying the mount to
the store instead would render a receiver that can read nothing.
2026-08-24 17:21:26 +02:00
atlas
bfa25c6419 feat(swarm): start the log store with the other required services
Deferred until the collector had a logs pipeline writing to it: a store
nothing writes to starts, answers queries and returns nothing, so the first
person to look concludes there were no logs rather than that nothing was
collecting them.
2026-08-24 17:00:38 +02:00
atlas
1af0138928 fix(swarm-otel): tell the log store which field carries the message
The journald receiver leaves the OTLP body empty and carries the entry as a
map of journal fields, so VictoriaLogs had no message to index and wrote a
placeholder into _msg on every record. Ingest returned 200, every field was
present, and a plain search for a line sitting in the store found nothing.

_msg_field names the field that holds the text. _stream_fields is the
difference between one stream for the whole host and one per unit per
machine; both are set by journald itself and both are low-cardinality.
2026-08-24 16:59:01 +02:00
atlas
0c755e04c8 feat(swarm-otel): collect the host journal into the swarm's log store
A journald receiver reading the host's journal directory, a logs pipeline
stamped with the swarm tier's own resource processor, and an otlphttp
exporter pointed at VictoriaLogs. All four parts are conditional on the log
store being enabled, so a swarm without one renders exactly as before.

The host's directory is enough to see every container: nspawn links a
non-ephemeral container's journal guest-side, so the files live on the host
under the container's machine-id, and journalctl descends into those
subdirectories. Measured, along with the fixed systemd-journal gid that makes
the group grant meaningful across the bind mount.
2026-08-24 16:48:41 +02:00
atlas
527e07c5e0 feat(swarm-victorialogs): a log store for the swarm
An agent can verify that a unit was launched and never that it is
working: container journals are not reachable, so a diagnosis stops at
the first broken component -- which is precisely the component whose own
instrument is least likely to be legible. This is the store half of
collecting logs centrally so the question becomes answerable.

Mirrors swarm-victoriametrics deliberately: same container shape, same
loopback pin, same self-scrape arrangement. Two differences, both
intentional.

No gateway vhost. VictoriaLogs' ingest and query endpoints carry no
authentication of their own, exactly like the metrics store's -- and the
metrics store IS published under a resolvable name, which is an open
question rather than a settled design. Publishing this one the same way
would repeat that before the first instance is decided.

Retention defaults to 30d against the metrics store's 5y. Logs are
orders of magnitude larger per unit of time and their value decays much
faster: a log line answers what happened during an incident, a metric
answers whether this is worse than last quarter.

Not wired into enableRequiredServices yet -- that lands with the
collector pipeline, so we do not start a store nothing writes to.
2026-08-24 16:36:18 +02:00
atlas
79f4132b4f fix(swarm-nats): the metrics exporter needs a collector flag or it exits
prometheus-nats-exporter has never served a metric. Upstream's module
renders `-addr … -port … ${extraFlags} ${url}` and defaults extraFlags
to the empty list, but the binary refuses to start without at least one
collector: it exits 1 with "no Collectors specified". So the unit logged
Started, the process was gone milliseconds later, and every scrape was
refused -- up=0 continuously, scrape_duration 0.6ms, zero samples.

Verified by running the exact argv both ways: without a collector it
exits 1, with -varz it stays up and logs the listener. A bogus flag is
rejected with exit 2, so the check measures acceptance rather than
tolerance.

varz is the server itself, connz makes a client that will not stay
connected visible, and jsz covers JetStream, which this swarm uses for
the status KV. The rest describe a clustered deployment we do not have.

An empty list is not a neutral default when the program requires a flag,
and nothing about it is visible to evaluation -- which is why the gate
now reads the rendered ExecStart rather than only asserting enable.
2026-08-24 16:06:27 +02:00
damocles
7e253a3421 hive-forge: timestamp-suffix the swarm-controller token name to avoid a re-mint collision 2026-08-24 16:02:42 +02:00
atlas
2d0d8c686a feat(swarm-grafana): provision an authelia dashboard
The swarm-services board was mostly authelia, so it becomes its own page
and is trimmed out of that one (mara, #3494). Eleven panels: uptime,
authentications and failures, authorization decisions, requests and
verdicts by status code, request and OIDC latency, and the three Go
process signals.

OIDC latency gets its own panel rather than being folded into the
general one because authelia keeps a separate duration family for it,
and every machine-to-machine credential in this swarm is minted through
those endpoints -- averaging the two produces a number describing
neither.

Counters are counted over the dashboard range rather than rated. At this
volume a rate window contains no requests, so rate() returns zero and
draws a flat line, which is indistinguishable from a broken query; the
quantile panels are worse, because a quantile over all-zero buckets is
NaN and renders empty rather than zero.
2026-08-24 15:59:28 +02:00
damocles
215a747b13 hive-forge: grant the swarm-controller token write:admin for CreateForgeUser 2026-08-24 15:42:59 +02:00
atlas
d2de3e8a2a docs(swarm-otel): scrapeTargets constrains the target, not the scraper
The option's description claimed declaring an entry from the service's
own module put 'the scraper and the target on the same host by
construction rather than by luck'. It does not. It constrains where the
target is; nothing in it places the collector, and the two enable flags
are co-located by a shared lib.mkDefault rather than by construction.

Split across hosts, a target is silently never scraped — the service's
host declares an entry no local collector reads, the collector's host
never enabled the service. No error surfaces, and no assertion can
catch it: separate hosts are separate evaluations with no shared
context, so the doc telling the truth is the only mechanism there is.

The same paragraph already warned co-location was not a guarantee, four
lines below the sentence claiming it was; a reader arriving for
permission stopped at the permission. This one did.

swarm-nats carries the concrete caveat for its own contribution.
2026-08-24 15:22:57 +02:00
atlas
e5c44f835a feat(#3518): expose NATS broker metrics via prometheus-nats-exporter
NATS has no Prometheus format of its own. It serves a JSON monitoring
endpoint, and prometheus-nats-exporter translates that — so this is two
changes in order, not one: without the monitoring endpoint the exporter
starts cleanly and scrapes nothing, which is the inert-config shape the
scrape work exists to avoid.

Both listeners are loopback and the exporter is the monitoring
endpoint's only intended reader: it is unauthenticated and /connz names
every connected client, so the address it binds is the whole access
control.

The scrape target is declared here rather than in the collector's
module, gated on a collector existing to read it — an entry exists only
where the service that named it runs.
2026-08-24 15:22:57 +02:00
iris
10a294a2f8 swarm-ui: swap colors.css via a plain nginx location, not a package-copy derivation
mara, on review: "i thought we just swap a css file via nginx
config?" -- right instinct. The package-copy overlay (cp -r + install)
only made sense for hive-c0re's servedFrontend, which backs multiple
serve points (dashboard root + every per-agent gateway route) from one
swapped tree. swarm-ui has exactly one location serving cfg.package, so
an = /static/colors.css exact-match override -- the same idiom every
other single-path override on this vhost already uses (/api/whoami,
/api/docs) -- replaces the whole derivation with one location block.

Verified with the same throwaway nixosSystem eval as the previous
commit: unthemed case has no colors.css location and the / root is
cfg.package unchanged.
2026-08-24 14:28:25 +02:00
iris
2ecf842a1b swarm-ui: apply the operator's stylix theme, same as the dashboard already does
mara: "swarm dash is in my own colors, but tab colors on swarm ui are
the default catpucchin one." Root cause: hive-c0re/theme.nix's stylix
overlay only ever wrote a themed colors.css onto the dashboard/agent
frontend subtrees -- swarm-ui, served from its own separate package,
was never in scope.

Extracts the stylix-detection + colors.css-generation logic (previously
inline in theme.nix) into a shared nix/host-modules/stylix-theme.nix,
imported by both theme.nix and swarm-ui.nix -- one source instead of a
second copy that has to agree by inspection. swarm-ui.nix gains its own
themedPackage overlay (same shape as theme.nix's themedFrontend: copy
the package, overwrite static/colors.css) and serves that instead of
cfg.package directly when stylix is active; a clean passthrough
otherwise.

Verified with a throwaway nixosSystem eval (this repo's own flake
checks do not exercise gateway-module wiring) confirming the unthemed
path resolves cfg.package unchanged.
2026-08-24 14:28:25 +02:00
atlas
2e06957b32 feat(swarm-otel): deliver the collector's client secret from authelia
The scrape config names a client_secret_file; this is what puts a file
there. A host oneshot copies authelia's minted secret between the two
container trees — a copy and not a bind mount, because the secret does
not exist until authelia's first boot and nixos-container refuses to
start on a missing bind source, which on a fresh swarm is a permanent
stall presenting as broken metrics.

The collector runs under DynamicUser and the prometheus receiver opens
client_secret_file itself, as that user, so there is no uid to hand the
file to. LoadCredential reads it as root before the sandbox exists and
re-exposes it under a path that does not depend on which uid the unit
got; the scrape config points there. Both spellings derive from one
binding, since a mismatch is a file the collector cannot open and
nothing but a runtime 401 would say so.
2026-08-24 13:46:48 +02:00
atlas
97b104b54f feat(swarm-otel): scrape a published target with an authelia bearer token
The published targets now render as prometheus scrape jobs, so the
option's name is true: the collector reaches them by name over https,
using prometheus-native oauth2 with the audience set to the target's own
url. Registered is not the same as requested — a client that does not
ask for an audience gets a token with an empty one however complete its
registration looks.

The url is split into scheme, target and metrics_path rather than asked
for three times: two spellings of one address is a mismatch waiting to
happen, and the failure it produces is a valid token refused at the
target. A malformed or non-https url is an assertion rather than a null
dereference from inside the renderer.
2026-08-24 13:42:14 +02:00
atlas
94ef8428e0 refactor(swarm-otel): one declaration for a published scrape target
`collectorAudiences` was a list of URLs services contributed so the
collector's client could be registered for them. Slice C needs the same
URLs as scrape jobs, and a job needs a name the audience list has no
room for — so services would have contributed to two options that must
agree.

They now contribute `publishedScrapeTargets` once, as `<job> = "<url>"`,
and the client's audiences derive from it. The drift that would have
needed maintaining is gone, and its failure mode was the quiet one: a
target whose audience was forgotten authenticates against nothing and
reads as a broken scrape rather than a missing registration.

Adds an assertion for the one collision the module system cannot catch.
Two definitions of the same key within one option are already refused
(measured); across the two scrape options nothing arbitrates, and both
entries would render into a single scrape_configs list under one
job_name.
2026-08-24 13:36:32 +02:00
atlas
94204ac82b fix(swarm-authelia): guard bearerAuthz against a silent no-op on kind
Review found that nothing stopped `bearerAuthz = true` on an interactive
client. `renderClient` reads the flag only in the machine branch, so the
scope is never emitted: the client authenticates, is authorised for
nothing, and the build is green. `kind` defaults to `interactive`, so it
is reached by forgetting a field rather than by writing a wrong one —
this module's own failure mode one level up.

The other half of the review asked to relax the method assertion to
accept `null`, on the strength of the option's doc calling `null`
authelia's default. Measured instead: under `authelia.bearer.authz`
authelia refuses the omission outright, so the assertion was right and
the DOC was wrong. The doc now carries the exception, and the assertion
message says null is refused rather than leaving a reader to infer it.
2026-08-24 13:35:14 +02:00
atlas
3895a1e21d feat(#3517): register the swarm collector as an audienced oauth2 client
The forge's `/metrics` is published behind the gateway and denied to
everyone, waiting on a client to allow. This is that client.

An audience is a URL: authelia validates a bearer token against the
address being requested, and a client may only request an audience it is
registered for, so registration is the authorisation. The URLs are owned
by the services that publish them while audiences attach to one client,
so services contribute to a list and this module builds the single entry
— the `gateway.localNames` split, forced here by client definitions
concatenating rather than merging into a shared entry.

The access-control rule asks the client list whether the collector is
registered rather than re-deriving the conditions that register it. The
two drifting is not a build failure: authelia refuses a rule naming an
unknown client in its startup validator, so SSO fails to restart.
2026-08-24 13:35:14 +02:00
atlas
da88d450dd feat(#3517): a named capability for authelia's bearer-authz scope
A scraper reaches a service published behind the gateway by presenting an
access token to authelia's authz endpoint, which requires the client to
carry authelia.bearer.authz. Nothing could express that: scopes are
derived from kind, and a machine client rendered an empty list.

A named capability rather than a free-form scopes list, for the reason
the derivation exists — authelia refuses some scope/grant combinations
outright, openid with client_credentials among them, and a list makes
those expressible again.

The two assertions carry their weight: authelia checks the same
obligations, but in its preStart validator, so a violation builds and
deploys cleanly and then fails to restart with swarm SSO attached to it.
2026-08-24 13:35:14 +02:00
atlas
4de4878e74 fix(swarm-otel): reserve the swarm tier's component names from hive names
The collector names components `<kind>/<owner>` — a hive name for the
per-hive pipelines, the literal `swarm` for the swarm tier's own. Both
land in one attrset via `//`, so a hive named `swarm` replaced the swarm
tier's parts and lost its own: its receiver kept accepting pushes into a
pipeline that routed nowhere, and its samples lost the `hive` stamp that
makes attribution unforgeable. Zero failed assertions.

The reserved name is now bound once and interpolated at each swarm-tier
use, so the guard checks the same string the config emits rather than a
copy of it. A second swarm-tier pipeline joins the list and inherits the
check without touching the assertion.
2026-08-24 13:21:34 +02:00
atlas
8178b0b55c refactor(swarm-authelia): name the collector option in full at the gate
There are two otel options one word apart: swarm.otel is the swarm's
collector, hyperhive.otel is the per-hive tier that ships upstream and
never reads scrapeTargets. The gate went through a let binding declared
800 lines from its use, so which one it referred to was not visible where
it mattered — mara had to ask.

Gating the wrong one is not a build error. It is a target that is either
always declared or never declared, and both look like working config.
2026-08-24 12:37:39 +02:00
atlas
541dd29820 fix(swarm-authelia): only declare the scrape target where a collector reads it
The comment claimed the option's rule keeps scraper and target on one host
by construction. It does not. Both services default from
enableRequiredServices via mkDefault, which is an invitation to override
rather than a guarantee, so co-location is a property of the auto-deployed
topology and not of the module.

Gating the target on the collector's own enable makes the loopback address
honest: a host running authelia without a collector no longer declares a
target nothing can read. That absence was the part worth fixing, because it
is silent — no error, no metrics, nothing in a log to notice.

This does not make authelia scrapeable from another host. That needs the
endpoint published under a name with a certificate and an audience, which
is separate work; the option's docs now say so where someone splitting the
two would read it.
2026-08-24 12:36:54 +02:00
atlas
31fa891a4e feat(swarm-authelia): expose prometheus metrics and declare the scrape target
The endpoint was off, so nothing reported on the swarm's own SSO. Enabling
it alone would have added no data — the scraper that reads it only landed
with the swarm-tier prometheus receiver.

Loopback only, like the main listener and for a stronger reason: this
endpoint authenticates nothing and reports request volumes and outcomes for
every login on the swarm.

metricsPort is an option rather than a literal because every swarm container
shares the host netns, so two services picking the same port do not conflict
at build time — one loses at runtime with nothing in any log. 9959 is
upstream's default and is unclaimed across nix/.
2026-08-24 12:36:54 +02:00
atlas
b2ee674415 fix(#3517): metrics are not optional, and the endpoint is denied by default
Two review findings, folded together.

mara: a swarm integrated auto deployed forge always has metrics, so the
toggle is gone. The endpoint follows behindGateway instead, which is the
swarm-integrated shape and the condition the protected location lives
under. Serving it without that location would put it on a listener
openFirewall can expose with nothing in front.

argus: /metrics matched no access_control rule, so default_policy
one_factor governed it. That is any authenticated subject, which today
means any operator and tomorrow any agent. My audience argument covered
the Bearer path only; the same endpoint also accepts CookieSession, and a
cookie carries no audience at all, so the audience was never what stood
between a browser session and this data.

The rule is deny rather than a client-scoped allow because the collector's
client does not exist yet. Authelia refuses a subject naming an
unregistered client, and does so in a preStart validator rather than at
build time, so naming one early yields a green nixos-rebuild and dead
swarm SSO on the next restart. Denying until the client is registered
makes publishing the endpoint safe on its own; registering it is a
one-line change from deny to that allow.

Rule order is load-bearing: authelia takes the first match.
2026-08-24 12:06:51 +02:00
atlas
4e17deada5 feat(#3517): let a machine authenticate to the authz endpoint with a bearer token
The gateway authenticates scrapers so services do not each grow a static
bearer of their own, but authelia's auth_request endpoint ran its default
strategies, which are cookie-only. A scraper's OAuth2 access token was
refused no matter how it was minted.

CookieSession is listed explicitly because authn_strategies replaces the
defaults rather than extending them. Omitting it evaluates, renders and
starts, and silently ends every operator session on the swarm UI, which
uses this same endpoint.

Unconditional rather than keyed to whichever service is scraped today:
this makes a scheme available, not an authorisation. Authelia refuses a
token carrying no audience for the requested URL and only issues a client
audiences it is registered for, so nothing passes until a client is
registered against a specific URL.
2026-08-24 12:06:51 +02:00
atlas
ae129835ae feat(#3517): publish forgejo's metrics behind the gateway
Forgejo can serve prometheus metrics but nothing turned them on, and
turning them on alone would have published them: forgejo serves
`/metrics` on its normal listener and the gateway vhost proxies `/` to
that listener, so the existing catch-all would have carried the endpoint
to anyone. The option is therefore one switch for both halves, and the
`= /metrics` location is an exact match so it outranks that prefix.

Authentication is the gateway's rather than forgejo's own `[metrics]
TOKEN`: a scraper presents an audience-scoped authelia token which nginx
checks via auth_request, so the swarm keeps one identity system instead
of gaining a static bearer per service.

The subrequest deliberately omits the `error_page 401 =302` that
swarm-ui uses. That redirect sends a browser to a login page; a scraper
would follow it and parse HTML as metrics.
2026-08-24 12:06:51 +02:00
iris
dabd0fc823 swarm-ui: header profile menu — initials avatar, authelia settings + logout (#3570)
Adds an "/api/whoami" same-origin nginx proxy to authelia's own
GET /api/user/info (session-cookie authenticated, no swarm-controller
code needed) and a new UserMenu header component: a generated initials
avatar (first letter of display name, coloured from the same seven
base16 chromatic slots the nav accent already cycles through) opening a
popover with the signed-in name, a link to authelia settings, and log
out — both reusing the existing "Authelia" entry from GET /api/links
rather than a second source of the domain.

Per mara's call on the open avatar-mechanism question: initials now,
a real uploaded photo (authelia's settings UI implies pics are
settable) is an explicit future item, not blocking this.
2026-08-24 00:01:30 +02:00
atlas
9ac037b21a feat(#3494): provision the claude-usage dashboard
Dashboard 2 of the set the operator asked for: how Claude is being used
rather than which agent is using it, so the axes are model, effort, token
type and query source. The only deliberate overlap with the agents page
is the cost/token headline.

Every panel was run against the live store before this landed. Three
panels were dropped rather than shipped, because their label has exactly
one live value today and a page of single-bar charts is the same silent
failure as an empty one.

Turn count and turn length are PROXIES and say so in their descriptions:
nothing exports turn stats, so a session record stands in for a turn,
which holds because each turn runs a new claude process.
2026-08-23 23:21:50 +02:00
atlas
0ddb504a98 feat(#3494): filter the agents dashboard by hive
The operator asked for this when the dashboard was first reviewed and it was
deferred, not declined: the container metric family carried no hive label, so
selecting a hive emptied every container panel while the "All" default hid the
problem completely -- a `.*` matcher matches series where the label is absent.

That family now carries hive and swarm, verified against the live store rather
than inferred from the fix having merged: a hive selection returns the same 7
agents as the All default, at every window out to 168h, with a nonexistent
hive returning zero.

The agent list chains off the hive selection, so picking a hive narrows the
agent dropdown rather than leaving entries in it that resolve to nothing.
2026-08-23 22:51:25 +02:00
atlas
ece21b0f7f feat(#3554): let a hive-owned service declare a scrape target 2026-08-23 22:45:56 +02:00
atlas
9fec1e1dca feat(#3494): provision the agents dashboard from the repo
Grafana served no dashboards: only datasources were provisioned, while
the module header already claimed dashboards were. This ships the agents
dashboard as a file provider and makes that sentence true.

The datasource uid is bound once and substituted into the dashboard at
build time. Committing the literal would make the dashboard a second
speller of a name the datasource already owns, and the drift failure is
silent -- panels render empty rather than erroring.

The shipped copy drops the `DS` datasource variable: it exists so an
operator can pick a store on manual import, and a provisioned dashboard
must not ask.
2026-08-23 22:45:25 +02:00
atlas
1f24193609 feat(swarm-victoriametrics): scrape the store's own prometheus endpoint
The scraper shipped with no targets, so nothing exercised it. This is its
first user, and the one with the least new surface: victoriametrics
publishes prometheus metrics on the listener it already serves queries on,
so there is no exporter, no extra port and no new reach — the collector's
otlphttp exporter already writes to that same loopback address.

Declared from this module rather than the collector's, per the option's own
rule: an entry exists only where the service that named it runs.
2026-08-19 21:40:40 +02:00
atlas
ff5c76c42f feat(#3520): scrape swarm-service prometheus endpoints at the swarm tier
Nothing in this deployment read a Prometheus endpoint, so enabling
/metrics on a managed service added an endpoint and no data. The swarm
collector receives OTLP pushes and does not scrape; VictoriaMetrics
stores what is pushed and has no scrape config. The pipeline was entirely
push-based and every such service is pull-based.

Adds a prometheus receiver, a resource/swarm processor and a
metrics/swarm pipeline alongside the per-hive ones.

The pipeline is separate because that is the ruling, not for tidiness:
every resource/<hive> processor UPSERTS a hive key, so a scraped swarm
sample routed through one would acquire the single label a swarm-level
service must not have. Keeping it out makes the absence structural rather
than something to remember to strip, the same way hive stays a property
of which authenticated receiver accepted a push.

Targets come from an option each service fills in from its own module,
under its own enable, rather than a list assembled here. That is what
puts the scraper and the target on the same host by construction: an
entry exists only where the service that named it runs. Co-location is
true of the all-local deployment and is not a guarantee, and that is
exactly the case where assuming it is invisible.

All three additions MERGE with the per-hive attrsets rather than
replacing them. A plain assignment would drop every hive's receiver,
processor and pipeline and still render a config the collector starts
cleanly on.

Empty target set emits no receiver, no processor and no pipeline — an
enabled scraper with nothing to scrape is the inert configuration this
issue is about, and the target set ships empty here because the targets
themselves are separate issues.

Formatting verified with nix fmt. The evaluation gate is not written yet;
nothing here has been evaluated against a fixture.
2026-08-19 21:40:40 +02:00
atlas
893b686c15 fix(#3527): a missing source must report as zero, not as empty
argus, reviewing, ran the guard against a source path that does not
exist rather than one that is empty. grep writes nothing to stdout in
that case, so `|| true` left the count variable empty and the -eq test
died with "integer expected" instead of reporting.

The unit still failed — cat hits the same missing file and set -e stops
it — but with a generic "no such file" rather than the message naming
which half is absent, which is the only thing this guard is for.

`|| echo 0` on all three counts. The gate gained the arm that was
missing: an absent source, asserted to fail THROUGH the guard rather
than merely to fail.
2026-08-19 20:27:00 +02:00
atlas
40557ffb74 fix(#3527): count the hive half, not the assembled bundle
The guard inspected the assembled file for any certificate. The system
store always holds certificates, so it passed unconditionally — including
in the one case it was written to catch, where the hive CA half
contributed nothing.

That half is the only one that matters here: every name these consumers
verify is issued by our own CA, so a bundle of nothing but public CAs is,
for this purpose, an empty bundle that measures as full. The failure is
silent and total — the unit reports success and every egress TLS call to
a swarm service then fails.

Counts the source on its own before assembling, and checks the result
carries what both halves brought, so a source truncated between the count
and the copy is caught too.

Scope is stated at the guard: it proves the anchor was contributed, not
that it is usable. A consumer reading only the first certificate ignores
it regardless, which is what took the swarm collector down, and no check
on this file can see that. Only a handshake can.
2026-08-19 20:16:16 +02:00
atlas
b8b571fea0 fix(#3524): stop naming the trust bundle as the oidc issuer anchor
The swarm collector could never verify its OIDC issuer, so it exited at
startup on every boot and the hive tier dropped every metric.

`issuer_ca_path` loads only the FIRST certificate in the file it names.
The bundle assembled for this container is `system CAs ++ hive anchors`,
so the swarm CA sits ~123rd and was never in the pool: the extension got
whichever public CA sorts first, could verify nothing of ours, and failed
`x509: certificate signed by unknown authority` — with 125 valid
certificates in the file.

Leaving the option unset makes the extension use the process trust store,
which `trustBundle` already populates via `SSL_CERT_FILE`, and that
consumer reads every certificate regardless of order. One file, two
consumers, opposite parsing; the fix is to stop naming it twice rather
than to reorder the bundle.

Measured with the deployed binary against the deployed config, varying
only the CA source: root-only OK, root-first OK, root-last FAILS,
root-second FAILS, and unset-with-SSL_CERT_FILE OK against a control that
correctly fails when the anchor is absent.
2026-08-19 18:23:42 +02:00
damocles
582b5cf83d fix(#3500): make the swarm-controller forge account a site admin 2026-08-19 17:13:54 +02:00
atlas
7880483b51 swarm-otel: stamp the swarm on every hive's pipeline
mara on the tracking issue: "we need the swarm label for upstream otel at
least (the out of swarm one)."

Stamped in the per-hive `resource` processor rather than on a separate
upstream-only pipeline, which would double the pipeline count to withhold one
constant label from the local store. It is redundant there — one metrics store
per swarm, so every series in it already belongs to this swarm — but a constant
label multiplies no series, and it means what leaves and what stays have the
same shape.

Upstream is where it stops being redundant: that is the one hop where several
swarms can land in one store, and samples that cannot name their swarm collide
there exactly as hives collided here before per-hive receivers existed.

`unknown` when unnamed rather than an absent label, copying the agent path so
a query never has to handle both "the label is missing" and "the label says
unknown".
2026-08-19 17:06:47 +02:00
atlas
738cc413e7 otel: wait for the telemetry client secret instead of racing it
`LoadCredential` naming a missing path is fatal at unit start, and this hive's
secret is minted by authelia's first-boot generator inside its own container —
nothing orders a host unit against that.

nixpkgs sets `Restart = "always"` on the collector with no `RestartSec`, so
that failure is instant: the unit burns systemd's 5-starts-in-10s allowance in
well under a second, lands in `start-limit-hit`, and stops retrying entirely.
`Restart = always` reads like it makes this self-healing and does the opposite
— a slow-failing unit retries until the secret appears, a fast-failing one
exhausts its limit before the thing it waits for can exist.

A oneshot converts the fast failure into a slow one, which is what that restart
policy is actually good at. Copied from `hive-forge-oidc-secret.service`, which
already solves this for the forge: bounded wait, then fail loudly naming the
file — never skip, because a skip yields a collector that starts and ships
nothing.

`TimeoutStartSec` exceeds the wait on purpose: `DefaultTimeoutStartSec` is 90s
and would kill the unit before it could emit that message.

The ordering against authelia's container is conditional — on a hive that does
not host the provider the secret is operator-provided, and naming a unit that
does not exist orders nothing, silently. The wait itself still applies there,
so a file that arrives late is tolerated rather than fatal.
2026-08-19 15:34:40 +02:00
atlas
9bd2b9e9e6 otel: a hive always authenticates — drop the unauthenticated mode
mara, reviewing this PR: "hives always require an identity, swarm controller
and auth is not optional."

So `requireHiveIdentity` is gone rather than defaulted, and with it every
branch that had to describe an unauthenticated collector. The swarm tier now
serves per-hive receivers only, and `/` answers 404 because there is no
swarm-wide inbox to route to. A hive with no credential is a build error, not
a quieter mode.

`hivePortBase` goes too: with per-hive receivers unconditional, `port` IS the
base of the range. That keeps one documented knob instead of adding a second,
and its advice ("move it if something else claims that range") still holds.

Two assertions replace the toggle — an empty hive roster, and a null
`authelia.url`. The second matters because a guessed issuer URL evaluates
cleanly, deploys cleanly, and then refuses every hive at runtime.

⚠️ `cfg.port` is deliberately no longer compared against the derived range in
the collision assertion: it is now the range's first element, so listing it
would make that assertion fire on every config.

This also retires the asymmetry guard added earlier in review — the state it
protected against (auth off on one side, credential still set on the other)
is no longer representable.
2026-08-19 15:27:09 +02:00
atlas
7da7915150 otel: refuse a half-configured escape hatch instead of 404ing silently
Turning ingest auth off without clearing a hive's credential leaves that
hive's collector authenticating and addressing its own path, while an
unauthenticated swarm tier serves one catch-all and forwards the URI
unchanged. The receiver is asked for a path it does not serve, so telemetry
stops with 404s and retries — no 401, no assertion, nothing in any log
naming auth.

Only reachable by overriding one side without the other, since both defaults
derive from the same flag. That is what makes it worth a build error rather
than a caveat: an operator who flips the documented escape hatch has no
reason to suspect the sending half.

Found in review by argus.
2026-08-19 15:27:09 +02:00
atlas
cb787997bd otel: the hive tier presents its own identity to the swarm collector
The receiving half authenticates per hive, so this half has to prove which
hive it is. It mints a token against the swarm's authelia with this hive's
client and posts to that hive's path on the collector's gateway name.

Holding a credential is what decides whether this tier authenticates —
`clientSecretFile` non-null — rather than a second switch that could
disagree with it. The default is the secret this host's own authelia
minted, which is right exactly when the IdP runs here; a hive that is not
that host names wherever the file landed, the same manual-copy shape the
identities option already documents as unsolved.

Two things that a diff will not explain:

`endpoint_params.audience` is not redundant with the client's registered
audience. Registering only makes an audience permissible; a token minted
without asking for one carries `aud: []` and every receiver refuses it,
with a config that reads correctly at both ends.

`client_secret_file` keeps the secret out of nix altogether — the
collector opens the file itself. It is a real key of this extension,
checked against the shipped binary with a deliberate typo rejected in the
same run, so "accepted" is distinguishable from "ignores everything". The
path comes from systemd's `CREDENTIALS_DIRECTORY`, so nothing hardcodes a
`/run/credentials` layout.

An assertion covers the one deployment where this can go wrong silently:
a host running both tiers with ingest authenticated and no credential to
present would 401 against a collector on the same machine.
2026-08-19 15:27:09 +02:00
atlas
08faa0970e swarm-otel: authenticate ingest per hive, and stamp the hive from the receiver
The swarm collector accepted OTLP from anyone who could reach it, and took
the `hive` resource attribute from the payload. So any writer on the swarm
network could attribute metrics to any hive, and nothing downstream could
tell.

The label now comes from which receiver accepted the sample: one receiver
per hive, each behind an `oidc` extension verifying a token minted for that
hive's audience, each feeding a pipeline whose `resource` processor upserts
a constant. A sender cannot influence it, because the only input is which
authenticated port the bytes arrived on.

That multiplicity is forced rather than preferred. A processor cannot read
the token's claims — `from_context` reads request metadata, and asking it
for an auth claim yields nothing, silently, with a healthy startup — and
one receiver holding several credentials never reveals which one matched.

The per-hive ports are internal: a hive reaches its receiver as a path
under this collector's existing gateway name, so nginx (rendered from this
same evaluation) is the only thing that names a port. Fronting each hive
with its own vhost would need a certificate, a DNS name and a gateway entry
per hive to express routing the gateway already does.

Turning this on removes the unauthenticated receiver. While an open port
still accepts samples the per-hive receivers are decoration, so this is the
switch itself rather than a hardening layer beside it; a swarm that wants
the open receiver says so.

`hive-ca-trust.nix` grows `bundlePathFor`, because a consumer taking its own
CA argument has to name the bundle rather than just have `SSL_CERT_FILE`
exported at it.
2026-08-19 15:27:09 +02:00