Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
16d578e692 | ||
|
|
c364d262e5 | ||
|
|
f6870c6a85 | ||
|
|
e01ecef18e | ||
|
|
bedfa786a7 |
7 changed files with 722 additions and 11 deletions
|
|
@ -16,8 +16,11 @@ services.hyperhive.otel = {
|
|||
};
|
||||
```
|
||||
|
||||
`enable` is the single gate. `endpoint` (required when enabled) is where
|
||||
telemetry ultimately goes.
|
||||
`enable` is the single gate. `endpoint` is where telemetry goes upstream —
|
||||
required when enabled *unless* this host runs the swarm's own metrics store
|
||||
(`swarm.victoriametrics.enable`), which is a destination in its own right. With
|
||||
both, telemetry goes to both. See
|
||||
[`swarm/services.md`](swarm/services.md#metrics-victoriametrics--grafana).
|
||||
|
||||
**There is exactly one way telemetry leaves a hive: through the collector that
|
||||
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
|
||||
|
|
@ -58,11 +61,16 @@ closed by this design, and nothing here should be read as closing it.
|
|||
|
||||
Master switch. When true, all other options below take effect.
|
||||
|
||||
### `services.hyperhive.otel.endpoint` — string, required when enabled
|
||||
### `services.hyperhive.otel.endpoint` — string, required when enabled unless the swarm store runs here
|
||||
|
||||
OTLP collector endpoint URL. Set as `OTEL_EXPORTER_OTLP_ENDPOINT` for every
|
||||
Upstream OTLP endpoint URL. Set as `OTEL_EXPORTER_OTLP_ENDPOINT` for every
|
||||
agent. Example: `"https://collector.example.com/otel"`.
|
||||
|
||||
Leave it empty **only** on a host running `swarm.victoriametrics.enable` — the
|
||||
local store is then the destination and the collector writes there instead.
|
||||
With neither, `enable` is refused at eval: telemetry with nowhere to go is a
|
||||
misconfiguration, not a quiet no-op.
|
||||
|
||||
### `services.hyperhive.otel.protocol` — enum, default `"http/protobuf"`
|
||||
|
||||
OTLP wire protocol, passed as `OTEL_EXPORTER_OTLP_PROTOCOL`. Accepted values:
|
||||
|
|
@ -142,12 +150,16 @@ protects it from other containers, not from the agent itself. As long as the
|
|||
direct path stays *selectable*, that hole stays selectable; an option that can
|
||||
reintroduce it is a hole with extra steps.
|
||||
|
||||
**`endpoint` keeps meaning "where telemetry ultimately goes."** The collector
|
||||
**`endpoint` keeps meaning "where telemetry goes upstream."** The collector
|
||||
does not redefine it — the agent-facing value is *derived*
|
||||
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
|
||||
keeps working unchanged. The bridge port is contributed to `exposeHostPorts`
|
||||
automatically; there is nothing to open by hand.
|
||||
|
||||
What the collector *added* is a second destination: on a host running the
|
||||
swarm's metrics store it writes there too, so `endpoint` is no longer the only
|
||||
place telemetry can land — and no longer the only way to have one.
|
||||
|
||||
### `services.hyperhive.otel.collector.port` — port, default `4318`
|
||||
|
||||
The OTLP/HTTP port the collector listens on, bound to the bridge IP only.
|
||||
|
|
|
|||
|
|
@ -61,3 +61,47 @@ See [`sso.md`](sso.md) for bootstrapping the first user and the OIDC
|
|||
relying-party flow, and [`secrets.md`](secrets.md) for where each of
|
||||
authelia's keys is generated and read.
|
||||
|
||||
### Metrics (VictoriaMetrics + Grafana)
|
||||
|
||||
The swarm's telemetry lands in one VictoriaMetrics and is read through
|
||||
one Grafana, in two containers at `metrics.<swarm-domain>` and
|
||||
`grafana.<swarm-domain>`. Two containers rather than one so Grafana can
|
||||
be restarted or broken without taking the time-series database with it.
|
||||
|
||||
Both follow `swarm.enableRequiredServices` like authelia and matrix, so
|
||||
the swarm's service host gets them with everything else. They derive
|
||||
together: a store with no UI is unreadable and a UI with no store is
|
||||
empty. To run one without the other, set it directly:
|
||||
|
||||
```nix
|
||||
services.hyperhive.swarm.victoriametrics.enable = true;
|
||||
services.hyperhive.swarm.grafana.enable = false;
|
||||
```
|
||||
|
||||
⚠️ **This starts a database that grows for as long as the swarm runs.**
|
||||
See `retentionPeriod` below before leaving it at its default.
|
||||
|
||||
| Option | When you'd touch it |
|
||||
|---|---|
|
||||
| `swarm.victoriametrics.retentionPeriod` | Default `5y`. Lower it once you have measured how fast this swarm actually fills a disk — the default is deliberately generous because too-short silently discards history you cannot get back. |
|
||||
| `swarm.grafana.oidc.role` | Default `Admin` for everyone who logs in. Lower to `Viewer`/`Editor` if the swarm grows operators who should not be able to reconfigure Grafana. |
|
||||
| `swarm.grafana.datasourceUrl` | Only if you front VictoriaMetrics with something else. It defaults to the store on this host, which is the only thing it can reach. |
|
||||
|
||||
**Logging in.** Grafana is behind swarm SSO, so the accounts are the
|
||||
authelia ones — there is no separate Grafana password, and the local
|
||||
login form is switched off whenever SSO is configured. If you enable
|
||||
Grafana on a host with no authelia, the form stays on and Grafana's
|
||||
default `admin`/`admin` applies; change it before exposing that host.
|
||||
|
||||
**Where the data comes from.** With `otel.enable` on, the hive's OTEL
|
||||
collector writes into this store as well as to any upstream endpoint —
|
||||
both, not one or the other, since a local store is for looking at this
|
||||
swarm and an upstream is for whoever aggregates across swarms. That also
|
||||
means `otel.endpoint` is no longer required when the store runs here: a
|
||||
hive with a local store already has somewhere for telemetry to go. See
|
||||
[`../observability.md`](../observability.md).
|
||||
|
||||
Neither container is reachable except through the gateway: both bind
|
||||
loopback, and VictoriaMetrics' write endpoint takes no credential, so
|
||||
the collector is the only intended writer.
|
||||
|
||||
|
|
|
|||
|
|
@ -26,8 +26,10 @@
|
|||
./swarm-ca.nix
|
||||
./swarm-nats.nix
|
||||
./swarm-controller.nix
|
||||
./swarm-grafana.nix
|
||||
./swarm-snapshot-store.nix
|
||||
./swarm-ui.nix
|
||||
./swarm-victoriametrics.nix
|
||||
./swarm-wireguard.nix
|
||||
./swarm.nix
|
||||
./swarm-peers-removed.nix
|
||||
|
|
|
|||
|
|
@ -36,8 +36,15 @@
|
|||
default = "";
|
||||
example = "https://collector.example.com/otel";
|
||||
description = ''
|
||||
OTLP collector endpoint, set as `OTEL_EXPORTER_OTLP_ENDPOINT`
|
||||
for every agent. Required when `enable` is true.
|
||||
Upstream OTLP endpoint, set as `OTEL_EXPORTER_OTLP_ENDPOINT` for
|
||||
every agent.
|
||||
|
||||
Required when `enable` is true, **unless** this host runs the
|
||||
swarm's metrics store
|
||||
({option}`services.hyperhive.swarm.victoriametrics.enable`) — that
|
||||
store is a destination in its own right, and with both configured
|
||||
telemetry goes to both. With neither, `enable` is refused rather
|
||||
than silently exporting nowhere.
|
||||
'';
|
||||
};
|
||||
|
||||
|
|
@ -158,8 +165,25 @@
|
|||
(lib.mkIf config.services.hyperhive.c0re.enable {
|
||||
assertions = lib.optionals config.services.hyperhive.otel.enable [
|
||||
{
|
||||
assertion = config.services.hyperhive.otel.endpoint != "";
|
||||
message = "services.hyperhive.otel.enable is true but services.hyperhive.otel.endpoint is empty.";
|
||||
# Telemetry has to go SOMEWHERE, but "somewhere" stopped meaning
|
||||
# "an upstream endpoint" once the swarm grew its own store: a hive
|
||||
# running `swarm.victoriametrics` is a complete destination on its
|
||||
# own, and requiring an external endpoint as well would make the
|
||||
# all-local mode impossible to express.
|
||||
#
|
||||
# This only ever relaxes the old rule — every config that passed
|
||||
# before still passes.
|
||||
assertion =
|
||||
config.services.hyperhive.otel.endpoint != ""
|
||||
|| config.services.hyperhive.swarm.victoriametrics.enable;
|
||||
message = ''
|
||||
services.hyperhive.otel.enable is true but telemetry has nowhere
|
||||
to go: services.hyperhive.otel.endpoint is empty and
|
||||
services.hyperhive.swarm.victoriametrics.enable is false.
|
||||
|
||||
Set the endpoint to export upstream, or enable the swarm's
|
||||
metrics store to keep telemetry on this host.
|
||||
'';
|
||||
}
|
||||
];
|
||||
})
|
||||
|
|
@ -176,6 +200,29 @@
|
|||
# becoming "how agents talk to the collector". The agent half
|
||||
# is pinned to OTLP/HTTP by the receiver below (derived in
|
||||
# hive-c0re/environment.nix).
|
||||
vmCfg = config.services.hyperhive.swarm.victoriametrics;
|
||||
# Two independent destinations, either of which may be absent: an
|
||||
# upstream the operator named, and the swarm's own store when this
|
||||
# host runs it. The assertion above guarantees at least one.
|
||||
upstreamConfigured = otel.endpoint != "";
|
||||
localStore = vmCfg.enable;
|
||||
|
||||
storeName = "otlphttp/victoriametrics";
|
||||
store = {
|
||||
# ⚠️ `metrics_endpoint`, NOT `endpoint`, and the difference is
|
||||
# invisible until you read the far end: `endpoint` is a BASE that
|
||||
# otlphttp appends `/v1/metrics` to, while VictoriaMetrics serves
|
||||
# OTLP at `/opentelemetry/api/v1/push`. With `endpoint` the
|
||||
# collector still answers 200 to its own clients and the samples
|
||||
# are silently posted to a path that does not exist.
|
||||
# `metrics_endpoint` is used verbatim.
|
||||
#
|
||||
# Measured end-to-end rather than read: a real sample crossed a
|
||||
# real collector into a real store, and the same probe with
|
||||
# `endpoint` never arrived — see `state/probe-3265-collector-to-vm.sh`.
|
||||
metrics_endpoint = "http://127.0.0.1:${toString vmCfg.port}/opentelemetry/api/v1/push";
|
||||
};
|
||||
|
||||
grpcUpstream = otel.protocol == "grpc";
|
||||
upstreamName = if grpcUpstream then "otlp" else "otlphttp";
|
||||
upstream = {
|
||||
|
|
@ -213,10 +260,17 @@
|
|||
validateConfigFile = true;
|
||||
settings = {
|
||||
receivers.otlp.protocols.http.endpoint = listen;
|
||||
exporters.${upstreamName} = upstream;
|
||||
exporters =
|
||||
lib.optionalAttrs upstreamConfigured { ${upstreamName} = upstream; }
|
||||
// lib.optionalAttrs localStore { ${storeName} = store; };
|
||||
service.pipelines.metrics = {
|
||||
receivers = [ "otlp" ];
|
||||
exporters = [ upstreamName ];
|
||||
# Fan-out, not a choice: with both configured the same
|
||||
# samples go upstream AND into the swarm's store. A local
|
||||
# store is for looking at this swarm; an upstream is for
|
||||
# whoever aggregates across swarms, and neither replaces
|
||||
# the other.
|
||||
exporters = lib.optional upstreamConfigured upstreamName ++ lib.optional localStore storeName;
|
||||
};
|
||||
};
|
||||
};
|
||||
|
|
|
|||
416
nix/host-modules/swarm-grafana.nix
Normal file
416
nix/host-modules/swarm-grafana.nix
Normal file
|
|
@ -0,0 +1,416 @@
|
|||
# The swarm's metrics UI: one Grafana for the whole swarm, in a
|
||||
# `swarm-grafana` nixos-container next to the VictoriaMetrics store it reads.
|
||||
#
|
||||
# Two containers rather than one, on the operator's call: Grafana can be
|
||||
# restarted, reconfigured or broken without taking the TSDB down with it.
|
||||
# They are a pair, not a unit.
|
||||
#
|
||||
# Dashboards are provisioned from config, deliberately not deployed from git.
|
||||
{
|
||||
pkgs,
|
||||
lib,
|
||||
config,
|
||||
...
|
||||
}:
|
||||
let
|
||||
cfg = config.services.hyperhive.swarm.grafana;
|
||||
hyperhiveCfg = config.services.hyperhive;
|
||||
gatewayCfg = hyperhiveCfg.gateway;
|
||||
tlsCfg = hyperhiveCfg.tls;
|
||||
autheliaCfg = hyperhiveCfg.swarm.authelia;
|
||||
vmCfg = hyperhiveCfg.swarm.victoriametrics;
|
||||
swarmDomain = hyperhiveCfg.swarm.domain;
|
||||
|
||||
caTrust = import ./lib/hive-ca-trust.nix { inherit lib tlsCfg gatewayCfg; };
|
||||
useSelfSigned = caTrust.useSelfSigned;
|
||||
|
||||
# Total on a null swarm domain for the same reason every sibling module is:
|
||||
# the required-domain assertion in hive-network.nix should be what an
|
||||
# operator sees, not a coercion error from here.
|
||||
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
|
||||
|
||||
# The all-local case: this host runs BOTH Grafana and the swarm's authelia,
|
||||
# so the minted secret can be moved without an operator. Same split the
|
||||
# forge and matrix modules document.
|
||||
ssoLocal = cfg.enable && autheliaCfg.enable;
|
||||
autheliaUrl = toString autheliaCfg.url;
|
||||
|
||||
# Where the plaintext lands inside the container. Under /var/lib rather
|
||||
# than /run: Grafana may start before the delivery unit on a later boot,
|
||||
# and a secret that evaporates on reboot turns a working login into an
|
||||
# intermittent one.
|
||||
secretPath = "/var/lib/grafana-oidc/${cfg.oidc.clientId}.secret";
|
||||
|
||||
# Format-locked by Grafana: the generic OAuth callback is always
|
||||
# `<root_url>/login/generic_oauth`. Declared once here and read by both
|
||||
# the authelia client and Grafana itself.
|
||||
redirectUri = "https://${cfg.domain}/login/generic_oauth";
|
||||
|
||||
grafanaCaBundle = "/run/swarm-grafana-ca/ca-bundle.crt";
|
||||
in
|
||||
{
|
||||
options.services.hyperhive.swarm.grafana = {
|
||||
enable = lib.mkOption {
|
||||
type = lib.types.bool;
|
||||
default = false;
|
||||
description = ''
|
||||
Run the swarm's metrics UI on this host. Off by default and not
|
||||
derived from {option}`services.hyperhive.enable`: a swarm has one
|
||||
Grafana, so enabling it is a decision about swarm topology rather
|
||||
than about whether hyperhive is installed.
|
||||
'';
|
||||
};
|
||||
|
||||
package = lib.mkOption {
|
||||
type = lib.types.package;
|
||||
default = pkgs.grafana;
|
||||
defaultText = lib.literalExpression "pkgs.grafana";
|
||||
description = "Grafana package to run.";
|
||||
};
|
||||
|
||||
machine = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
readOnly = true;
|
||||
default = "swarm-grafana";
|
||||
description = ''
|
||||
Container name. Read-only: the name appears in host paths and in
|
||||
`machinectl`, so it is a fact other modules may read rather than a
|
||||
knob.
|
||||
'';
|
||||
};
|
||||
|
||||
domain = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "grafana.${domainBase}";
|
||||
defaultText = lib.literalExpression ''"grafana.''${services.hyperhive.swarm.domain}"'';
|
||||
description = ''
|
||||
Name the gateway serves this on. A sibling of the swarm's other
|
||||
service names, so the swarm-services sub-CA can issue for it — see
|
||||
`hive-tls.nix` for why a service name being a sibling rather than a
|
||||
child decides which CA may sign it.
|
||||
|
||||
⚠️ Changing this changes the OAuth redirect URI, which authelia
|
||||
matches exactly. Both sides move together because both derive from
|
||||
this option; an operator who pins one by hand breaks the login.
|
||||
'';
|
||||
};
|
||||
|
||||
port = lib.mkOption {
|
||||
type = lib.types.port;
|
||||
default = 3000;
|
||||
description = ''
|
||||
Port Grafana listens on, bound to loopback only. Upstream's own
|
||||
default, kept so an operator reading Grafana documentation finds
|
||||
what they expect.
|
||||
'';
|
||||
};
|
||||
|
||||
datasourceUrl = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "http://127.0.0.1:${toString vmCfg.port}";
|
||||
defaultText = lib.literalExpression ''"http://127.0.0.1:''${toString services.hyperhive.swarm.victoriametrics.port}"'';
|
||||
description = ''
|
||||
Where the provisioned datasource points. Defaults to the metrics
|
||||
store on this host, which is the only place it can be: that store
|
||||
binds loopback, so a Grafana somewhere else could not reach it
|
||||
anyway. Set explicitly if a deployment fronts VictoriaMetrics with
|
||||
something that does listen wider.
|
||||
'';
|
||||
};
|
||||
|
||||
oidc = {
|
||||
clientId = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "swarm-grafana";
|
||||
description = ''
|
||||
The authelia OIDC client id. Names the application rather than
|
||||
the protocol, per the convention in
|
||||
{option}`services.hyperhive.swarm.authelia.oidc.clients`.
|
||||
'';
|
||||
};
|
||||
|
||||
role = lib.mkOption {
|
||||
type = lib.types.enum [
|
||||
"Viewer"
|
||||
"Editor"
|
||||
"Admin"
|
||||
];
|
||||
default = "Admin";
|
||||
example = "Editor";
|
||||
description = ''
|
||||
Grafana org role every SSO user is assigned.
|
||||
|
||||
`Admin` by default, and that is a considered default rather than
|
||||
a permissive one: the login form is disabled whenever SSO is
|
||||
configured, so this is the *only* way anyone reaches Grafana —
|
||||
a `Viewer` default would produce a swarm nobody can administer.
|
||||
Passing authelia already means being an operator of this swarm;
|
||||
its user store is the small, `swarmctl`-managed one.
|
||||
|
||||
Lower it if a swarm ever grows read-only operators, which is a
|
||||
one-line change here.
|
||||
'';
|
||||
};
|
||||
};
|
||||
};
|
||||
|
||||
config = lib.mkIf (hyperhiveCfg.enable && cfg.enable) {
|
||||
# The gateway name and the quick-link, both inside `cfg.enable` — that
|
||||
# guard is the load-bearing part. Every hive in a swarm may know this UI
|
||||
# exists, but only the host that RUNS it may claim the name; a client
|
||||
# hive declaring the vhost would answer for a service it does not have.
|
||||
services.hyperhive.gateway.localNames = [ cfg.domain ];
|
||||
|
||||
services.hyperhive.swarm.controller.links = [
|
||||
{
|
||||
label = "Grafana";
|
||||
icon = "📊";
|
||||
url = "https://${cfg.domain}/";
|
||||
}
|
||||
];
|
||||
|
||||
# One declaration, two readers. Grafana's callback URL is format-locked
|
||||
# to its own root URL; making the operator restate it in authelia's
|
||||
# client list would be a second source of truth for a string whose
|
||||
# mismatch is a silently rejected login.
|
||||
#
|
||||
# `kind` is left at its `interactive` default: a person logs in here.
|
||||
services.hyperhive.swarm.authelia.oidc.clients = lib.mkIf ssoLocal [
|
||||
{
|
||||
id = cfg.oidc.clientId;
|
||||
description = "HyperHive swarm metrics";
|
||||
redirectUris = [ redirectUri ];
|
||||
}
|
||||
];
|
||||
|
||||
assertions = [
|
||||
{
|
||||
# Grafana reaches the token endpoint server-to-server, so a null URL
|
||||
# would become a request to `null/api/oidc/token` — a DNS failure
|
||||
# several layers from its cause. Only reachable by enabling authelia
|
||||
# and clearing its `url`, which is why it is an assertion and not a
|
||||
# fallback.
|
||||
assertion = !ssoLocal || autheliaCfg.url != null;
|
||||
message = ''
|
||||
services.hyperhive.swarm.grafana requires
|
||||
services.hyperhive.swarm.authelia.url when authelia is enabled.
|
||||
|
||||
Grafana exchanges its authorization code at
|
||||
`''${url}/api/oidc/token` from inside its container. With the URL
|
||||
null there is no endpoint to name.
|
||||
'';
|
||||
}
|
||||
];
|
||||
|
||||
# Websockets: Grafana Live streams panel updates over one, and without
|
||||
# the upgrade headers dashboards load and then never refresh — which
|
||||
# reads as stale data rather than as a proxy fault.
|
||||
services.nginx.virtualHosts."${cfg.domain}" = (gatewayCfg.lib.tlsFor cfg.domain) // {
|
||||
listen = gatewayCfg.lib.listen;
|
||||
extraConfig = gatewayCfg.lib.securityHeaders;
|
||||
locations."/" = {
|
||||
proxyPass = "http://127.0.0.1:${toString cfg.port}/";
|
||||
proxyWebsockets = true;
|
||||
extraConfig = ''
|
||||
# Grafana builds its OAuth redirect from the ORIGINAL request.
|
||||
# Without these every request looks like it arrived at 127.0.0.1
|
||||
# over plain http, and the redirect sent to authelia names a
|
||||
# host the browser cannot reach.
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
proxy_set_header X-Forwarded-Host $host;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
'';
|
||||
};
|
||||
};
|
||||
|
||||
# Order the container after the host CA service so the bind source below
|
||||
# exists before nspawn sets the mount up.
|
||||
systemd.services."container@${cfg.machine}" = caTrust.containerOrdering;
|
||||
|
||||
# The secret delivery. It runs on the HOST because that is the only place
|
||||
# both container trees are addressable: they share this host's network
|
||||
# namespace, which makes them feel co-located, but their filesystem roots
|
||||
# are separate — Grafana cannot open a path inside authelia's tree
|
||||
# however local the port looks.
|
||||
#
|
||||
# ⚠️ Deliberately a copy and not a `bindMounts` entry. nixos-container
|
||||
# refuses to start when a bind source is missing, and this secret does
|
||||
# not exist until authelia's first boot has minted it — so binding it
|
||||
# would make Grafana wait on a file that waits on a container that starts
|
||||
# after it.
|
||||
systemd.services.swarm-grafana-oidc-secret = lib.mkIf ssoLocal {
|
||||
description = "deliver Grafana's OIDC client secret from authelia";
|
||||
after = [ "container@${autheliaCfg.machine}.service" ];
|
||||
requires = [ "container@${autheliaCfg.machine}.service" ];
|
||||
before = [ "container@${cfg.machine}.service" ];
|
||||
wantedBy = [ "container@${cfg.machine}.service" ];
|
||||
serviceConfig = {
|
||||
Type = "oneshot";
|
||||
RemainAfterExit = true;
|
||||
SyslogIdentifier = "swarm-grafana-oidc-secret";
|
||||
# ⚠️ Longer than the wait below, and that is the whole point:
|
||||
# `DefaultTimeoutStartSec` is 90s, so a 120s bounded wait is killed
|
||||
# by systemd at 90 — before it can emit the error naming the file it
|
||||
# waited for. The timeout has to outlive the thing it is timing.
|
||||
TimeoutStartSec = "180s";
|
||||
};
|
||||
path = [ pkgs.coreutils ];
|
||||
script = ''
|
||||
set -euo pipefail
|
||||
|
||||
src=${lib.escapeShellArg "${autheliaCfg.hostClientSecretDir}/${cfg.oidc.clientId}.secret"}
|
||||
dst=${lib.escapeShellArg "/var/lib/nixos-containers/${cfg.machine}${secretPath}"}
|
||||
|
||||
# authelia's container is up, but its first-boot generator may still
|
||||
# be minting. Bounded wait, then fail: a silent skip here produces a
|
||||
# Grafana whose only login path dead-ends.
|
||||
deadline=$(( SECONDS + 120 ))
|
||||
while [ ! -s "$src" ]; do
|
||||
if [ "$SECONDS" -ge "$deadline" ]; then
|
||||
echo "authelia has not minted $src after 120s" >&2
|
||||
exit 1
|
||||
fi
|
||||
sleep 2
|
||||
done
|
||||
|
||||
# Owned by Grafana's own uid, unlike the matrix sibling which lands
|
||||
# root-owned: tuwunel's secret is read by `LoadCredential` as root
|
||||
# before the sandbox exists, whereas Grafana expands `$__file{}`
|
||||
# itself, as itself, while parsing its config. These containers set
|
||||
# no `privateUsers`, so the host uid is the container uid, and both
|
||||
# sides take it from the same static NixOS id.
|
||||
#
|
||||
# Group is root, not grafana, and that is forced rather than chosen:
|
||||
# `ids.uids.grafana` is a static id but there is no `ids.gids.grafana`
|
||||
# — the group's gid is allocated at activation inside the container,
|
||||
# so the host cannot know it at eval time. Harmless here because 0400
|
||||
# grants the group nothing; if this mode ever widens, the gid has to
|
||||
# be discovered at runtime rather than assumed.
|
||||
install -D -m 0400 -o ${toString config.ids.uids.grafana} -g 0 "$src" "$dst"
|
||||
'';
|
||||
};
|
||||
|
||||
containers.${cfg.machine} = {
|
||||
autoStart = true;
|
||||
ephemeral = false;
|
||||
# Shared host netns, like every sibling swarm container: the gateway
|
||||
# reaches this at 127.0.0.1:<port>.
|
||||
privateNetwork = false;
|
||||
|
||||
bindMounts = { } // caTrust.bindMount;
|
||||
|
||||
config =
|
||||
{ ... }:
|
||||
{
|
||||
system.stateVersion = "26.05";
|
||||
|
||||
# This container shares the host netns, so its own firewall.service
|
||||
# would rewrite the HOST ruleset at every boot. The host firewall
|
||||
# owns all filtering.
|
||||
networking.firewall.enable = false;
|
||||
# Keep the host-copied /etc/resolv.conf intact — resolvconf's
|
||||
# host-tracking would regenerate it to an empty file, since the
|
||||
# host's copy doesn't cross the boundary after start.
|
||||
networking.resolvconf.enable = lib.mkForce false;
|
||||
|
||||
# Self-signed mode: Grafana is Go, and Go's `SSL_CERT_FILE`
|
||||
# *replaces* the default bundle rather than adding to it — so
|
||||
# concatenate the system CAs with the bind-mounted hive CA instead
|
||||
# of pointing at the CA alone, which would lose every public
|
||||
# anchor. /run is tmpfs, so this is rebuilt from the current CA
|
||||
# each boot rather than going stale.
|
||||
#
|
||||
# Without it the browser half of the login succeeds and the
|
||||
# server-to-server token exchange fails with an x509 "unknown
|
||||
# authority" — the same shape of failure the swarm queue hit.
|
||||
systemd.services.swarm-grafana-ca-bundle = lib.mkIf useSelfSigned {
|
||||
description = "assemble Grafana TLS trust bundle (system CAs + hive CA)";
|
||||
wantedBy = [ "grafana.service" ];
|
||||
before = [ "grafana.service" ];
|
||||
serviceConfig = {
|
||||
Type = "oneshot";
|
||||
RemainAfterExit = true;
|
||||
SyslogIdentifier = "swarm-grafana-ca-bundle";
|
||||
};
|
||||
path = [ pkgs.coreutils ];
|
||||
script = ''
|
||||
set -euo pipefail
|
||||
install -d -m 0755 /run/swarm-grafana-ca
|
||||
cat /etc/ssl/certs/ca-certificates.crt ${caTrust.caContainerPath} \
|
||||
> ${grafanaCaBundle}
|
||||
chmod 0644 ${grafanaCaBundle}
|
||||
'';
|
||||
};
|
||||
|
||||
systemd.services.grafana.environment.SSL_CERT_FILE = lib.mkIf useSelfSigned grafanaCaBundle;
|
||||
|
||||
services.grafana = {
|
||||
enable = true;
|
||||
package = cfg.package;
|
||||
|
||||
settings = {
|
||||
server = {
|
||||
# Already upstream's default (measured), but pinned rather
|
||||
# than inherited: the gateway is the only intended client,
|
||||
# and this being loopback is what keeps the UI from being
|
||||
# published on whatever else the host is reachable on.
|
||||
http_addr = "127.0.0.1";
|
||||
http_port = cfg.port;
|
||||
domain = cfg.domain;
|
||||
# Grafana builds its own OAuth redirect from this. Left at
|
||||
# upstream's `%(protocol)s://%(domain)s:%(http_port)s/` it
|
||||
# would name `http://<domain>:3000/`, which authelia has
|
||||
# never heard of.
|
||||
root_url = "https://${cfg.domain}/";
|
||||
};
|
||||
|
||||
analytics = {
|
||||
reporting_enabled = false;
|
||||
check_for_updates = false;
|
||||
};
|
||||
|
||||
users.auto_assign_org_role = cfg.oidc.role;
|
||||
|
||||
# No local password path at all when SSO is configured. This
|
||||
# is not tidiness: Grafana ships an `admin`/`admin` account,
|
||||
# and this vhost is on the public gateway.
|
||||
auth.disable_login_form = ssoLocal;
|
||||
}
|
||||
// lib.optionalAttrs ssoLocal {
|
||||
"auth.generic_oauth" = {
|
||||
enabled = true;
|
||||
name = "HyperHive";
|
||||
client_id = cfg.oidc.clientId;
|
||||
# ⚠️ `$__file{}`, never the secret itself — anything else
|
||||
# here is world-readable in the nix store. Note the module's
|
||||
# own leak assertion does NOT cover this key (it checks
|
||||
# `database.password`, `security.admin_password` and
|
||||
# datasource `secureJsonData`), so nothing but this comment
|
||||
# stands between a literal and the store.
|
||||
client_secret = "$__file{${secretPath}}";
|
||||
scopes = "openid profile email groups";
|
||||
auth_url = "${autheliaUrl}/api/oidc/authorization";
|
||||
token_url = "${autheliaUrl}/api/oidc/token";
|
||||
api_url = "${autheliaUrl}/api/oidc/userinfo";
|
||||
use_pkce = true;
|
||||
};
|
||||
};
|
||||
|
||||
provision.datasources.settings = {
|
||||
apiVersion = 1;
|
||||
datasources = [
|
||||
{
|
||||
name = "VictoriaMetrics";
|
||||
type = "prometheus";
|
||||
uid = "swarm-victoriametrics";
|
||||
url = cfg.datasourceUrl;
|
||||
access = "proxy";
|
||||
isDefault = true;
|
||||
}
|
||||
];
|
||||
};
|
||||
};
|
||||
};
|
||||
};
|
||||
};
|
||||
}
|
||||
|
|
@ -58,5 +58,17 @@ in
|
|||
# mode was minting the queue's callout nkeys and then never starting
|
||||
# the queue they authenticate against.
|
||||
nats.enable = lib.mkDefault swarmCfg.enableRequiredServices;
|
||||
|
||||
# The metrics pair. Once per swarm and optional, so they meet the rule
|
||||
# in the option's description the same way the three above do — a hive
|
||||
# that is not the service host is a *client* of this Grafana, not a
|
||||
# second one.
|
||||
#
|
||||
# They derive together on purpose: a store with no UI is unreadable and
|
||||
# a UI with no store is empty, so there is no sensible deployment that
|
||||
# takes one and not the other from this switch. An operator who wants
|
||||
# exactly one still sets it directly, which `mkDefault` allows.
|
||||
victoriametrics.enable = lib.mkDefault swarmCfg.enableRequiredServices;
|
||||
grafana.enable = lib.mkDefault swarmCfg.enableRequiredServices;
|
||||
};
|
||||
}
|
||||
|
|
|
|||
171
nix/host-modules/swarm-victoriametrics.nix
Normal file
171
nix/host-modules/swarm-victoriametrics.nix
Normal file
|
|
@ -0,0 +1,171 @@
|
|||
# The swarm's metrics store: one VictoriaMetrics for the whole swarm, in a
|
||||
# `swarm-victoriametrics` nixos-container.
|
||||
#
|
||||
# A LOCAL time-series database rather than only an external sink, and that is
|
||||
# the point rather than a convenience: the swarm dashboard has to stay
|
||||
# readable when the outside world is unreachable. Same failure-domain property
|
||||
# the hive-status KV has — a view of the system must not depend on the system
|
||||
# it is viewing being healthy.
|
||||
#
|
||||
# It is the collector that feeds this (the gateway OTEL collector), not the
|
||||
# agents directly: one ingest point per swarm, authenticated there.
|
||||
{
|
||||
pkgs,
|
||||
lib,
|
||||
config,
|
||||
...
|
||||
}:
|
||||
let
|
||||
cfg = config.services.hyperhive.swarm.victoriametrics;
|
||||
hyperhiveCfg = config.services.hyperhive;
|
||||
gatewayCfg = hyperhiveCfg.gateway;
|
||||
swarmDomain = hyperhiveCfg.swarm.domain;
|
||||
|
||||
# Total on a null swarm domain for the same reason every sibling module is:
|
||||
# the required-domain assertion in hive-network.nix should be what an
|
||||
# operator sees, not a coercion error from here.
|
||||
domainBase = if swarmDomain == null then "invalid" else swarmDomain;
|
||||
in
|
||||
{
|
||||
options.services.hyperhive.swarm.victoriametrics = {
|
||||
enable = lib.mkOption {
|
||||
type = lib.types.bool;
|
||||
default = false;
|
||||
description = ''
|
||||
Run the swarm's metrics store on this host. Off by default and not
|
||||
derived from {option}`services.hyperhive.enable`: a swarm has one
|
||||
metrics store, so enabling it is a decision about swarm topology
|
||||
rather than about whether hyperhive is installed.
|
||||
'';
|
||||
};
|
||||
|
||||
package = lib.mkOption {
|
||||
type = lib.types.package;
|
||||
default = pkgs.victoriametrics;
|
||||
defaultText = lib.literalExpression "pkgs.victoriametrics";
|
||||
description = "VictoriaMetrics package to run.";
|
||||
};
|
||||
|
||||
machine = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
readOnly = true;
|
||||
default = "swarm-victoriametrics";
|
||||
description = ''
|
||||
Container name. Read-only: the name appears in host paths and in
|
||||
`machinectl`, so it is a fact other modules may read rather than a
|
||||
knob.
|
||||
'';
|
||||
};
|
||||
|
||||
domain = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "metrics.${domainBase}";
|
||||
defaultText = lib.literalExpression ''"metrics.''${services.hyperhive.swarm.domain}"'';
|
||||
description = ''
|
||||
Name the gateway serves this on. A sibling of the swarm's other
|
||||
service names, so the swarm-services sub-CA can issue for it — see
|
||||
`hive-tls.nix` for why a service name being a sibling rather than a
|
||||
child decides which CA may sign it.
|
||||
'';
|
||||
};
|
||||
|
||||
port = lib.mkOption {
|
||||
type = lib.types.port;
|
||||
default = 8428;
|
||||
description = ''
|
||||
Port VictoriaMetrics listens on, bound to loopback only (see
|
||||
below). Upstream's own default, kept so an operator reading
|
||||
VictoriaMetrics documentation finds what they expect.
|
||||
'';
|
||||
};
|
||||
|
||||
retentionPeriod = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "5y";
|
||||
example = "90d";
|
||||
description = ''
|
||||
How long samples are kept.
|
||||
|
||||
Deliberately a high default rather than a required option: the two
|
||||
failure directions are not symmetric. Too long fills a disk, which
|
||||
is visible and recoverable by lowering this; too short **destroys
|
||||
history**, silently and permanently. So the safe default is generous
|
||||
and an operator lowers it once they have measured how fast this swarm
|
||||
actually accumulates data.
|
||||
'';
|
||||
};
|
||||
};
|
||||
|
||||
config = lib.mkIf (hyperhiveCfg.enable && cfg.enable) {
|
||||
# The gateway name and the quick-link, both inside `cfg.enable` — that
|
||||
# guard is the load-bearing part. Every hive in a swarm may know this
|
||||
# store exists, but only the host that RUNS it may claim the name; a
|
||||
# client hive declaring the vhost would answer for a service it does not
|
||||
# have.
|
||||
services.hyperhive.gateway.localNames = [ cfg.domain ];
|
||||
|
||||
services.hyperhive.swarm.controller.links = [
|
||||
{
|
||||
label = "Metrics";
|
||||
icon = "📈";
|
||||
url = "https://${cfg.domain}/";
|
||||
}
|
||||
];
|
||||
|
||||
services.nginx.virtualHosts."${cfg.domain}" = (gatewayCfg.lib.tlsFor cfg.domain) // {
|
||||
listen = gatewayCfg.lib.listen;
|
||||
extraConfig = gatewayCfg.lib.securityHeaders;
|
||||
locations."/" = {
|
||||
proxyPass = "http://127.0.0.1:${toString cfg.port}/";
|
||||
};
|
||||
};
|
||||
|
||||
containers.${cfg.machine} = {
|
||||
autoStart = true;
|
||||
ephemeral = false;
|
||||
# Shared host netns, like every sibling swarm container: the gateway
|
||||
# reaches this at 127.0.0.1:<port>.
|
||||
privateNetwork = false;
|
||||
|
||||
config =
|
||||
{ ... }:
|
||||
{
|
||||
system.stateVersion = "26.05";
|
||||
|
||||
# This container shares the host netns, so its own firewall.service
|
||||
# would rewrite the HOST ruleset at every boot. The host firewall
|
||||
# owns all filtering.
|
||||
networking.firewall.enable = false;
|
||||
# Keep the host-copied /etc/resolv.conf intact — resolvconf's
|
||||
# host-tracking would regenerate it to an empty file, since the
|
||||
# host's copy doesn't cross the boundary after start.
|
||||
networking.resolvconf.enable = lib.mkForce false;
|
||||
|
||||
services.victoriametrics = {
|
||||
enable = true;
|
||||
package = cfg.package;
|
||||
retentionPeriod = cfg.retentionPeriod;
|
||||
|
||||
# ⚠️ PINNED TO LOOPBACK, and this is a correction rather than a
|
||||
# preference: upstream's default is `:8428`, i.e. every
|
||||
# interface. The gateway is the only intended client and it is on
|
||||
# this host, so binding wider would publish an unauthenticated
|
||||
# write endpoint (see the OTLP note below) to whatever the host
|
||||
# is reachable on.
|
||||
listenAddress = "127.0.0.1:${toString cfg.port}";
|
||||
};
|
||||
};
|
||||
};
|
||||
};
|
||||
|
||||
# 🔑 OTLP ingest needs no flag. Measured against the pinned 1.146.0 rather
|
||||
# than inferred from the module's option list, which has no OTLP switch and
|
||||
# so reads as though the feature were absent: the running server answers
|
||||
# `POST /opentelemetry/api/v1/push` with 200 (a nonexistent path answers
|
||||
# 400, so that 200 means the route exists). The `-opentelemetry.*` flags
|
||||
# only tune naming and limits, and are reachable via `extraOptions` if a
|
||||
# deployment ever needs them.
|
||||
#
|
||||
# ⚠️ That endpoint is unauthenticated, which is why `listenAddress` above is
|
||||
# loopback and why the collector — not agents — is the writer.
|
||||
}
|
||||
Loading…
Reference in a new issue