Grafana was the one swarm service that exports Prometheus metrics and had no
scrape target, so the swarm's metrics UI was invisible to the metrics it
displays.
It serves on a unix socket and claims no TCP port, which is deliberate and
stays that way; a prometheus target has to be a host:port, so nginx re-serves
the one endpoint a scraper needs on loopback. An exact-match location, not a
prefix: widening it would re-serve the whole UI without authorization.
The metrics section is now stated explicitly rather than inherited, because a
scrape target depends on it and a changed upstream default would take the
endpoint away while nginx kept answering.
Every hive-labelled series in the store also carries an agent label, so a
hive is only ever visible as the sum of its agents — and a hive whose c0re
has stopped is indistinguishable from one that simply hosts none.
Adds three instruments to the exporter hive-c0re already runs, each a
projection of a value the process computes anyway: process.uptime (the
semconv name — the spec defines it as a double gauge in seconds, which is
exactly this instrument), hyperhive.hive.degraded, and
hyperhive.hive.warnings split by level. None carries an agent attribute;
that absence is what makes them selectable as hive-scoped.
The health pair reads warnings::readiness() rather than deriving its own
verdict, and degraded ships as a series instead of being left for a
dashboard query to compute from warnings{level="crit"} — either would put
the "what counts as unhealthy" rule in a second place that disagrees
silently the first time a degrading condition is added.
mara, on review: forge and matrix are swarm-level too (both are
swarm-wide singletons per their own module comments - one Forgejo,
one Matrix homeserver, not one per hive), and the single combined tree
made the host separation unclear. Split into two trees: everything a
hive host always runs, and everything that runs once per swarm on
whichever host opts in (can be the same host or a different one).
The diagram was single-hive only - no mention of swarm-controller,
swarm-ui, authelia SSO, the metrics/log stores, Grafana, NATS or the
swarm-level otel collector, despite all of them being real, shipped
services (several on by default via enableAllLocalDefaults). Adds a
swarm section alongside agent containers, same terse one-line-per-
service style the existing optional-containers section already uses.
swarm-ui gets a note that it has no own container (served straight
from the gateway's nginx, unlike every sibling swarm service).
Adds services.hyperhive.swarm.victorialogs.domain and a gateway vhost
gated by the same auth_request check against authelia that swarm-uis
own vhost uses (same swarmAuthRequest shape, copied not shared - see
the file top comment for why). The stores own listener stays
loopback-only and unauthenticated exactly as before; the collector
still writes to it directly, never through this vhost, so this only
adds a new authenticated read path.
Registers the new domain in swarm.nixs serviceDomains so the
swarm-services sub-CA issues for it (a missed entry silently falls
back to the hive leaf, which cannot cover a name under a different
apex - swarm.nixs own comment on that list documents the incident
this caused before).
Verified with a throwaway module-eval (same technique as the flake
module-eval check): vhost only exists when victorialogs.enable is
set, forceSSL/no addSSL, both locations present, domain correctly
registered/absent from serviceDomains, links entry present.
LinksMenu and SettingsMenu triggers mixed a full-colour emoji (link)
with a plain text glyph (gear) - different rendering paths mean
different, unfixable sizes/styles. Replace both with matching inline
SVG icons (feather/lucide gear + link glyphs), same viewBox/stroke/
size, so the two buttons finally share one rendering path.
Also: shared base.css never zeroed the default UA body margin, which
showed as a bg-coloured strip around the whole viewport edge on any
full-bleed header (swarm-uis .shell-header among them). Zeroed it in
the shared file so every consumer (dashboard, agent UI, swarm-ui) gets
the fix, not just swarm-ui.
Screenshot-verified at both desktop and phone widths.
Grafana could reach the swarm's metrics and not its logs, so the store that
landed with the collector pipeline had no reader.
Adds the VictoriaLogs datasource plugin and the Logs Drilldown app, and
provisions the datasource beside the metrics one. The drilldown matters as much
as the connection: an unfamiliar log stream is explorable without writing a
LogsQL query first, which is the difference between a store you can query and
one you can use.
The secrets page discussed 'the telemetry collector' as a reader needing no
delivery, but there are two: the hive's is a host unit and reads authelia's
file in place, while the swarm's runs in a container and gets a copy placed by
a host oneshot.
States plainly that the container one has no operator-provided variant, which
is a consequence of it running beside authelia rather than a gap.
nixpkgs' collector unit already sets SupplementaryGroups to systemd-journal
unconditionally, with a comment saying why. Systemd list options concatenate,
so this module's copy rendered ["systemd-journal" "systemd-journal"] and made
this a second owner of a fact upstream may later change.
The bind mount stays, since that half is genuinely ours.
The logs half was gated on the swarm's log store being enabled, so turning
that store off stopped collection entirely rather than leaving the upstream
export. Logs now fan out exactly as metrics do: the store when it runs, the
operator's upstream when one is configured, both when both.
The receiver, the journal mount and the group grant follow whether there is
anywhere to send logs, not whether the local store exists. Tying the mount to
the store instead would render a receiver that can read nothing.
Deferred until the collector had a logs pipeline writing to it: a store
nothing writes to starts, answers queries and returns nothing, so the first
person to look concludes there were no logs rather than that nothing was
collecting them.
The journald receiver leaves the OTLP body empty and carries the entry as a
map of journal fields, so VictoriaLogs had no message to index and wrote a
placeholder into _msg on every record. Ingest returned 200, every field was
present, and a plain search for a line sitting in the store found nothing.
_msg_field names the field that holds the text. _stream_fields is the
difference between one stream for the whole host and one per unit per
machine; both are set by journald itself and both are low-cardinality.
A journald receiver reading the host's journal directory, a logs pipeline
stamped with the swarm tier's own resource processor, and an otlphttp
exporter pointed at VictoriaLogs. All four parts are conditional on the log
store being enabled, so a swarm without one renders exactly as before.
The host's directory is enough to see every container: nspawn links a
non-ephemeral container's journal guest-side, so the files live on the host
under the container's machine-id, and journalctl descends into those
subdirectories. Measured, along with the fixed systemd-journal gid that makes
the group grant meaningful across the bind mount.
An agent can verify that a unit was launched and never that it is
working: container journals are not reachable, so a diagnosis stops at
the first broken component -- which is precisely the component whose own
instrument is least likely to be legible. This is the store half of
collecting logs centrally so the question becomes answerable.
Mirrors swarm-victoriametrics deliberately: same container shape, same
loopback pin, same self-scrape arrangement. Two differences, both
intentional.
No gateway vhost. VictoriaLogs' ingest and query endpoints carry no
authentication of their own, exactly like the metrics store's -- and the
metrics store IS published under a resolvable name, which is an open
question rather than a settled design. Publishing this one the same way
would repeat that before the first instance is decided.
Retention defaults to 30d against the metrics store's 5y. Logs are
orders of magnitude larger per unit of time and their value decays much
faster: a log line answers what happened during an incident, a metric
answers whether this is worse than last quarter.
Not wired into enableRequiredServices yet -- that lands with the
collector pipeline, so we do not start a store nothing writes to.
prometheus-nats-exporter has never served a metric. Upstream's module
renders `-addr … -port … ${extraFlags} ${url}` and defaults extraFlags
to the empty list, but the binary refuses to start without at least one
collector: it exits 1 with "no Collectors specified". So the unit logged
Started, the process was gone milliseconds later, and every scrape was
refused -- up=0 continuously, scrape_duration 0.6ms, zero samples.
Verified by running the exact argv both ways: without a collector it
exits 1, with -varz it stays up and logs the listener. A bogus flag is
rejected with exit 2, so the check measures acceptance rather than
tolerance.
varz is the server itself, connz makes a client that will not stay
connected visible, and jsz covers JetStream, which this swarm uses for
the status KV. The rest describe a clustered deployment we do not have.
An empty list is not a neutral default when the program requires a flag,
and nothing about it is visible to evaluation -- which is why the gate
now reads the rendered ExecStart rather than only asserting enable.
The swarm-services board was mostly authelia, so it becomes its own page
and is trimmed out of that one (mara, #3494). Eleven panels: uptime,
authentications and failures, authorization decisions, requests and
verdicts by status code, request and OIDC latency, and the three Go
process signals.
OIDC latency gets its own panel rather than being folded into the
general one because authelia keeps a separate duration family for it,
and every machine-to-machine credential in this swarm is minted through
those endpoints -- averaging the two produces a number describing
neither.
Counters are counted over the dashboard range rather than rated. At this
volume a rate window contains no requests, so rate() returns zero and
draws a flat line, which is indistinguishable from a broken query; the
quantile panels are worse, because a quantile over all-zero buckets is
NaN and renders empty rather than zero.
Part of #3591 (mara: "pop ups and menus appearing should animate").
Each popover mounts fresh on open ({open ? <div> : null}, not a state
transition), so a keyframe animation on the popover element itself is
the right tool -- same shape as Shell.css's own shell-page-enter
(fade + a slight translate/scale settle), including the identical
three-rule motion-guard (base rule, prefers-reduced-motion media
query, data-motion=reduce/allow explicit overrides).
Scope: LinksMenu, SettingsMenu, UserMenu -- the three header popovers.
Not included here (posted findings on the issue instead of guessing):
the refresh-interval picker's dropdown (native <select>, whose open
popup is OS/browser chrome outside CSS reach in current browsers --
"not themed" is a platform limitation, not a bug in this component's
own styling) and the jobs graph's node animations (JobqGraph is a
@hive/shared component consumed by both swarm-ui and the per-hive
dashboard, real design/implementation work on shared infra, not a
same-shape mechanical extension of an existing pattern).
Screenshot-verified the settled (post-animation) state renders
correctly; a static screenshot cannot show an in-flight CSS animation,
so this leans on exact structural parity with the already-shipped
Shell.css pattern for the animation's own correctness.
The option's description claimed declaring an entry from the service's
own module put 'the scraper and the target on the same host by
construction rather than by luck'. It does not. It constrains where the
target is; nothing in it places the collector, and the two enable flags
are co-located by a shared lib.mkDefault rather than by construction.
Split across hosts, a target is silently never scraped — the service's
host declares an entry no local collector reads, the collector's host
never enabled the service. No error surfaces, and no assertion can
catch it: separate hosts are separate evaluations with no shared
context, so the doc telling the truth is the only mechanism there is.
The same paragraph already warned co-location was not a guarantee, four
lines below the sentence claiming it was; a reader arriving for
permission stopped at the permission. This one did.
swarm-nats carries the concrete caveat for its own contribution.
NATS has no Prometheus format of its own. It serves a JSON monitoring
endpoint, and prometheus-nats-exporter translates that — so this is two
changes in order, not one: without the monitoring endpoint the exporter
starts cleanly and scrapes nothing, which is the inert-config shape the
scrape work exists to avoid.
Both listeners are loopback and the exporter is the monitoring
endpoint's only intended reader: it is unauthenticated and /connz names
every connected client, so the address it binds is the whole access
control.
The scrape target is declared here rather than in the collector's
module, gated on a collector existing to read it — an entry exists only
where the service that named it runs.
mara, on review: "i thought we just swap a css file via nginx
config?" -- right instinct. The package-copy overlay (cp -r + install)
only made sense for hive-c0re's servedFrontend, which backs multiple
serve points (dashboard root + every per-agent gateway route) from one
swapped tree. swarm-ui has exactly one location serving cfg.package, so
an = /static/colors.css exact-match override -- the same idiom every
other single-path override on this vhost already uses (/api/whoami,
/api/docs) -- replaces the whole derivation with one location block.
Verified with the same throwaway nixosSystem eval as the previous
commit: unthemed case has no colors.css location and the / root is
cfg.package unchanged.
mara: "swarm dash is in my own colors, but tab colors on swarm ui are
the default catpucchin one." Root cause: hive-c0re/theme.nix's stylix
overlay only ever wrote a themed colors.css onto the dashboard/agent
frontend subtrees -- swarm-ui, served from its own separate package,
was never in scope.
Extracts the stylix-detection + colors.css-generation logic (previously
inline in theme.nix) into a shared nix/host-modules/stylix-theme.nix,
imported by both theme.nix and swarm-ui.nix -- one source instead of a
second copy that has to agree by inspection. swarm-ui.nix gains its own
themedPackage overlay (same shape as theme.nix's themedFrontend: copy
the package, overwrite static/colors.css) and serves that instead of
cfg.package directly when stylix is active; a clean passthrough
otherwise.
Verified with a throwaway nixosSystem eval (this repo's own flake
checks do not exercise gateway-module wiring) confirming the unthemed
path resolves cfg.package unchanged.
The scrape config names a client_secret_file; this is what puts a file
there. A host oneshot copies authelia's minted secret between the two
container trees — a copy and not a bind mount, because the secret does
not exist until authelia's first boot and nixos-container refuses to
start on a missing bind source, which on a fresh swarm is a permanent
stall presenting as broken metrics.
The collector runs under DynamicUser and the prometheus receiver opens
client_secret_file itself, as that user, so there is no uid to hand the
file to. LoadCredential reads it as root before the sandbox exists and
re-exposes it under a path that does not depend on which uid the unit
got; the scrape config points there. Both spellings derive from one
binding, since a mismatch is a file the collector cannot open and
nothing but a runtime 401 would say so.
The published targets now render as prometheus scrape jobs, so the
option's name is true: the collector reaches them by name over https,
using prometheus-native oauth2 with the audience set to the target's own
url. Registered is not the same as requested — a client that does not
ask for an audience gets a token with an empty one however complete its
registration looks.
The url is split into scheme, target and metrics_path rather than asked
for three times: two spellings of one address is a mismatch waiting to
happen, and the failure it produces is a valid token refused at the
target. A malformed or non-https url is an assertion rather than a null
dereference from inside the renderer.
`collectorAudiences` was a list of URLs services contributed so the
collector's client could be registered for them. Slice C needs the same
URLs as scrape jobs, and a job needs a name the audience list has no
room for — so services would have contributed to two options that must
agree.
They now contribute `publishedScrapeTargets` once, as `<job> = "<url>"`,
and the client's audiences derive from it. The drift that would have
needed maintaining is gone, and its failure mode was the quiet one: a
target whose audience was forgotten authenticates against nothing and
reads as a broken scrape rather than a missing registration.
Adds an assertion for the one collision the module system cannot catch.
Two definitions of the same key within one option are already refused
(measured); across the two scrape options nothing arbitrates, and both
entries would render into a single scrape_configs list under one
job_name.
Review found that nothing stopped `bearerAuthz = true` on an interactive
client. `renderClient` reads the flag only in the machine branch, so the
scope is never emitted: the client authenticates, is authorised for
nothing, and the build is green. `kind` defaults to `interactive`, so it
is reached by forgetting a field rather than by writing a wrong one —
this module's own failure mode one level up.
The other half of the review asked to relax the method assertion to
accept `null`, on the strength of the option's doc calling `null`
authelia's default. Measured instead: under `authelia.bearer.authz`
authelia refuses the omission outright, so the assertion was right and
the DOC was wrong. The doc now carries the exception, and the assertion
message says null is refused rather than leaving a reader to infer it.
The forge's `/metrics` is published behind the gateway and denied to
everyone, waiting on a client to allow. This is that client.
An audience is a URL: authelia validates a bearer token against the
address being requested, and a client may only request an audience it is
registered for, so registration is the authorisation. The URLs are owned
by the services that publish them while audiences attach to one client,
so services contribute to a list and this module builds the single entry
— the `gateway.localNames` split, forced here by client definitions
concatenating rather than merging into a shared entry.
The access-control rule asks the client list whether the collector is
registered rather than re-deriving the conditions that register it. The
two drifting is not a build failure: authelia refuses a rule naming an
unknown client in its startup validator, so SSO fails to restart.
A scraper reaches a service published behind the gateway by presenting an
access token to authelia's authz endpoint, which requires the client to
carry authelia.bearer.authz. Nothing could express that: scopes are
derived from kind, and a machine client rendered an empty list.
A named capability rather than a free-form scopes list, for the reason
the derivation exists — authelia refuses some scope/grant combinations
outright, openid with client_credentials among them, and a list makes
those expressible again.
The two assertions carry their weight: authelia checks the same
obligations, but in its preStart validator, so a violation builds and
deploys cleanly and then fails to restart with swarm SSO attached to it.
The collector names components `<kind>/<owner>` — a hive name for the
per-hive pipelines, the literal `swarm` for the swarm tier's own. Both
land in one attrset via `//`, so a hive named `swarm` replaced the swarm
tier's parts and lost its own: its receiver kept accepting pushes into a
pipeline that routed nowhere, and its samples lost the `hive` stamp that
makes attribution unforgeable. Zero failed assertions.
The reserved name is now bound once and interpolated at each swarm-tier
use, so the guard checks the same string the config emits rather than a
copy of it. A second swarm-tier pipeline joins the list and inherits the
check without touching the assertion.
There are two otel options one word apart: swarm.otel is the swarm's
collector, hyperhive.otel is the per-hive tier that ships upstream and
never reads scrapeTargets. The gate went through a let binding declared
800 lines from its use, so which one it referred to was not visible where
it mattered — mara had to ask.
Gating the wrong one is not a build error. It is a target that is either
always declared or never declared, and both look like working config.
The comment claimed the option's rule keeps scraper and target on one host
by construction. It does not. Both services default from
enableRequiredServices via mkDefault, which is an invitation to override
rather than a guarantee, so co-location is a property of the auto-deployed
topology and not of the module.
Gating the target on the collector's own enable makes the loopback address
honest: a host running authelia without a collector no longer declares a
target nothing can read. That absence was the part worth fixing, because it
is silent — no error, no metrics, nothing in a log to notice.
This does not make authelia scrapeable from another host. That needs the
endpoint published under a name with a certificate and an audience, which
is separate work; the option's docs now say so where someone splitting the
two would read it.
The endpoint was off, so nothing reported on the swarm's own SSO. Enabling
it alone would have added no data — the scraper that reads it only landed
with the swarm-tier prometheus receiver.
Loopback only, like the main listener and for a stronger reason: this
endpoint authenticates nothing and reports request volumes and outcomes for
every login on the swarm.
metricsPort is an option rather than a literal because every swarm container
shares the host netns, so two services picking the same port do not conflict
at build time — one loses at runtime with nothing in any log. 9959 is
upstream's default and is unclaimed across nix/.
mara: "should be https://auth.constellation.darkest.space/settings in
profile pic menu" -- the UserMenu link was pointing at the plain
authelia domain root, which lands on the portal rather than the
account settings page. Appends /settings client-side, same base-URL
source as before (GET /api/links Authelia entry).