`swarm.victorialogs` (what the log store is to every hive: container name, domain, port) moves to nix/host-modules/swarm-victorialogs-service.nix, together with the only two helpers it reads, `swarmDomain` and `domainBase`. Everything else -- the `deploy.victorialogs` options, the whole `config` block including `containers.swarm-victorialogs`, the file header and the helpers only they read (`swarmAuthRequest` among them) -- stays in nix/host-modules/swarm-victorialogs.nix, which default.nix now imports alongside the new file. `hyperhiveCfg` (an alias for `config.services.hyperhive`, not an option) is read by both halves, so it is duplicated into the service file rather than shared. A pure move: option paths, option definitions and config are unchanged apart from the comment above the `deploy.victorialogs` options, which now names the file `swarm.victorialogs` lives in. Fixtures enabling the store evaluate to the same host and container toplevel derivations before and after. Refs #3742
275 lines
13 KiB
Nix
275 lines
13 KiB
Nix
# The swarm's log store: one VictoriaLogs for the whole swarm, in a
|
|
# `swarm-victorialogs` nixos-container beside the metrics store it mirrors.
|
|
#
|
|
# Why a store at all rather than reading journals directly: an agent can
|
|
# verify that a unit was *launched* and never that it is *working*. The
|
|
# container journals are not reachable from an agent, so a diagnosis stops at
|
|
# the first component that is broken — which is precisely the component whose
|
|
# own instrument is least likely to be legible. Collecting logs centrally
|
|
# makes the question answerable without host access.
|
|
#
|
|
# It is the collector that feeds this, not the services directly: one ingest
|
|
# point per swarm, same shape as the metrics path.
|
|
#
|
|
# Gateway vhost is authenticated, unlike the metrics store's: an operator
|
|
# reaches this at `https://${cfg.domain}/` gated by the same `auth_request`
|
|
# check against authelia that swarm-ui's own vhost uses (`swarmAuthRequest`
|
|
# below, same shape as `swarm-ui.nix`'s) — see that file's copy for the full
|
|
# rationale (forceSSL is load-bearing there too: authelia answers a plain-http
|
|
# auth subrequest with 400, which `auth_request` cannot interpret as anything
|
|
# but a broken check). The store's own listener stays loopback-only and
|
|
# unauthenticated, and the vhost is now the ONLY way in from outside: the
|
|
# collector pushes through it too, at its own `= /insert/...` location, because
|
|
# a swarm has one log store and the collector need not share a host with it.
|
|
# ⚠️ That ingest location deliberately does not carry `swarmAuthRequest` — a
|
|
# pusher handed its `error_page 401 =302` follows the redirect and POSTs at a
|
|
# login page, which answers 200. See the location itself.
|
|
#
|
|
# The query API has a machine route of the same shape at `^~ /select/logsql/`,
|
|
# for the same reason, and with no filtering of what an authenticated caller
|
|
# may read — see that location.
|
|
{
|
|
pkgs,
|
|
lib,
|
|
config,
|
|
...
|
|
}:
|
|
let
|
|
cfg = config.services.hyperhive.swarm.victorialogs;
|
|
deployCfg = config.services.hyperhive.deploy;
|
|
networkCfg = config.services.hyperhive.network;
|
|
hyperhiveCfg = config.services.hyperhive;
|
|
gatewayCfg = hyperhiveCfg.gateway;
|
|
|
|
# Copy of `swarm-ui.nix`'s own `swarmAuthRequest` — not shared code because
|
|
# this vhost needs exactly the two locations that reference it (`/` and the
|
|
# internal auth-request target below) and pulling in swarm-ui.nix for one
|
|
# string would couple this module to swarm-ui existing at all, which it
|
|
# need not. Same three lines, same reasoning as that file's own comment.
|
|
swarmAuthRequest = ''
|
|
auth_request /__hive_authelia;
|
|
auth_request_set $target_url $scheme://$http_host$request_uri;
|
|
error_page 401 =302 https://${hyperhiveCfg.swarm.authelia.domain}/?rd=$target_url;
|
|
'';
|
|
|
|
# Shared host netns, like every sibling swarm container: the collector
|
|
# and Grafana reach this at 127.0.0.1:<port>.
|
|
privateNetwork = false;
|
|
in
|
|
{
|
|
# What the store IS from any hive's point of view, the name it answers on and
|
|
# the port, is `swarm.victorialogs` in ./swarm-victorialogs-service.nix.
|
|
# `enable`, `package` and `retentionPeriod` are decisions of the host that
|
|
# runs it and live under `deploy.*`.
|
|
|
|
# Retention is a property of the store this host runs, not something the
|
|
# swarm has to agree on: it is read only where the container is defined,
|
|
# and a hive that is a *client* of the log store never consults it. That
|
|
# makes it a `deploy.*` value by the same rule as the seal on the secret
|
|
# store — options on the auto-deployed service itself.
|
|
options.services.hyperhive.deploy.victorialogs.package = lib.mkOption {
|
|
type = lib.types.package;
|
|
default = pkgs.victorialogs;
|
|
defaultText = lib.literalExpression "pkgs.victorialogs";
|
|
description = "VictoriaLogs package to run.";
|
|
};
|
|
|
|
options.services.hyperhive.deploy.victorialogs.retentionPeriod = lib.mkOption {
|
|
type = lib.types.str;
|
|
default = "30d";
|
|
example = "90d";
|
|
description = ''
|
|
How long log data is kept.
|
|
|
|
Deliberately far shorter than the metrics store's retention: logs
|
|
are orders of magnitude larger per unit of time, and their value
|
|
decays much faster. A log line answers "what happened during that
|
|
incident"; a metric answers "is this worse than last quarter".
|
|
'';
|
|
};
|
|
|
|
config = lib.mkIf deployCfg.victorialogs.enable {
|
|
# This store publishes its own health as prometheus metrics on the same
|
|
# listener it serves queries on, so the swarm's collector scrapes it with
|
|
# no exporter and no extra port — same arrangement as the metrics store.
|
|
#
|
|
# Declared here rather than in the collector's module because that is the
|
|
# rule the option carries: an entry exists only where the service that
|
|
# named it runs. ⚠️ That constrains the TARGET, not the scraper — a
|
|
# deployment that splits this container away from the collector's host
|
|
# silently drops the entry, and no assertion can see it, because separate
|
|
# hosts are separate evaluations.
|
|
services.hyperhive.swarm.otel.scrapeTargets.victorialogs = "127.0.0.1:${toString cfg.port}";
|
|
|
|
# The gateway name and the quick-link, both inside `deployCfg.victorialogs.enable` — same
|
|
# "only the host that runs the service may claim the name" guard every
|
|
# sibling swarm-service module uses (`swarm-grafana.nix`,
|
|
# `swarm-victoriametrics.nix`).
|
|
services.hyperhive.gateway.localNames = [ cfg.domain ];
|
|
|
|
# This host serves the vhost, and the container behind it resolves
|
|
# through the hive's dnsmasq.
|
|
services.hyperhive.gateway.enable = lib.mkDefault true;
|
|
services.hyperhive.gateway.dns.enable = lib.mkDefault true;
|
|
|
|
services.hyperhive.swarm.controller.links = [
|
|
{
|
|
label = "Logs";
|
|
icon = "📜";
|
|
url = "https://${cfg.domain}/";
|
|
}
|
|
];
|
|
|
|
# Authenticated front door onto the loopback-only store — see the
|
|
# file-top comment for why this is safe to add without touching the
|
|
# store's own (still unauthenticated, still loopback) listener at all.
|
|
# `removeAttrs`/`forceSSL`: same asymmetry `swarm-ui.nix` documents —
|
|
# authelia answers a plain-http auth subrequest with 400, which
|
|
# `auth_request` cannot read as anything but a broken check, so this
|
|
# vhost needs `forceSSL` rather than the `addSSL` every unauthenticated
|
|
# sibling vhost uses.
|
|
services.nginx.virtualHosts."${cfg.domain}" =
|
|
(builtins.removeAttrs (gatewayCfg.lib.tlsFor cfg.domain) [ "addSSL" ])
|
|
// {
|
|
forceSSL = true;
|
|
listen = gatewayCfg.lib.listen;
|
|
extraConfig = gatewayCfg.lib.securityHeaders;
|
|
locations = {
|
|
"/" = {
|
|
proxyPass = "http://127.0.0.1:${toString cfg.port}/";
|
|
extraConfig = swarmAuthRequest;
|
|
};
|
|
|
|
# The swarm collector's ingest route. `=` so it outranks the `/`
|
|
# prefix above — without a more specific location it would ride that
|
|
# catch-all, and that is the whole hazard here.
|
|
#
|
|
# ⚠️ Deliberately NOT `swarmAuthRequest`. That block ends in
|
|
# `error_page 401 =302`, which is right for a browser and wrong for a
|
|
# pusher: handed a redirect it follows the redirect and POSTs its
|
|
# batch at a login page, which answers 200. Ingest then reports
|
|
# healthy while storing nothing. A machine route lets the 401 reach
|
|
# the client unchanged.
|
|
#
|
|
# `auth_request` does not inherit across sibling locations (see the
|
|
# note on `swarmAuthRequest` above), so naming it here is required
|
|
# rather than redundant.
|
|
"= /insert/opentelemetry/v1/logs" = {
|
|
proxyPass = "http://127.0.0.1:${toString cfg.port}/insert/opentelemetry/v1/logs";
|
|
extraConfig = ''
|
|
auth_request /__hive_authelia;
|
|
'';
|
|
};
|
|
|
|
# The read side of that same split: the route a caller holding a
|
|
# bearer token queries, as opposed to the `/` above which is the
|
|
# operator's browser. Bare `auth_request` for the ingest route's
|
|
# exact reason — `error_page 401 =302` hands an unauthenticated
|
|
# machine caller authelia's login page as a **200 with an HTML
|
|
# body**, so a client that reads the status code records a
|
|
# successful query that returned no rows. 401 is the only answer
|
|
# here that a caller cannot mistake for an empty result.
|
|
#
|
|
# `^~` so it outranks the `/` catch-all and stays ahead of any
|
|
# regex location a later change adds. The prefix is the whole
|
|
# LogsQL query API and nothing else — `query`, `tail`, `hits`,
|
|
# `facets`, the `stats_query*` and the `field_*`/`stream_*`
|
|
# routes — while the browser UI sits beside it under
|
|
# `/select/vmui`. Both read off the pinned build, not the docs.
|
|
#
|
|
# The query is forwarded unmodified — no scoping parameter is
|
|
# injected, so any authenticated caller reads the whole swarm's
|
|
# logs. That is the rule in force, not an omission: read
|
|
# permissions are a later thing, and this location is where one
|
|
# attaches when it exists.
|
|
"^~ /select/logsql/" = {
|
|
# No URI part, so the request path and its query string reach
|
|
# the store as the caller sent them.
|
|
proxyPass = "http://127.0.0.1:${toString cfg.port}";
|
|
extraConfig = ''
|
|
auth_request /__hive_authelia;
|
|
'';
|
|
};
|
|
|
|
# The subrequest itself — same target, same header set, same
|
|
# reasoning as `swarm-ui.nix`'s own copy (measured against the
|
|
# pinned authelia binary, not copied from an example).
|
|
"= /__hive_authelia" = {
|
|
proxyPass = "https://${hyperhiveCfg.swarm.authelia.domain}/api/authz/auth-request";
|
|
# nixpkgs appends its OWN `Host $host` after extraConfig,
|
|
# which would override verifiedProxyTo's — see the comment
|
|
# on verifiedProxyTo in hive-gateway/vhost-lib.nix.
|
|
recommendedProxySettings = false;
|
|
extraConfig = ''
|
|
internal;
|
|
${gatewayCfg.lib.verifiedProxyTo hyperhiveCfg.swarm.authelia.domain}
|
|
proxy_pass_request_body off;
|
|
proxy_set_header Content-Length "";
|
|
proxy_set_header X-Original-Method $request_method;
|
|
proxy_set_header X-Original-URL $scheme://$http_host$request_uri;
|
|
proxy_set_header X-Forwarded-Proto $scheme;
|
|
proxy_set_header X-Forwarded-Host $http_host;
|
|
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
|
'';
|
|
};
|
|
};
|
|
};
|
|
|
|
containers.${cfg.machine} = {
|
|
autoStart = true;
|
|
ephemeral = false;
|
|
# Journal files on the host, not inside the container: nixpkgs hardcodes
|
|
# --link-journal=try-guest, and EXTRA_NSPAWN_FLAGS expands after it.
|
|
extraFlags = [ "--link-journal=host" ];
|
|
inherit privateNetwork;
|
|
|
|
config =
|
|
{ ... }:
|
|
{
|
|
imports = [
|
|
../container-modules/swarm-container.nix
|
|
(import ../container-modules/swarm-container-resolver.nix {
|
|
inherit (networkCfg) bridgeIp;
|
|
dnsConsumers = [ "victorialogs.service" ];
|
|
})
|
|
];
|
|
|
|
services.hyperhive.swarmContainer = { inherit privateNetwork; };
|
|
|
|
services.victorialogs = {
|
|
enable = true;
|
|
package = deployCfg.victorialogs.package;
|
|
|
|
# ⚠️ PINNED TO LOOPBACK for the same reason the metrics store is,
|
|
# and it matters more here: upstream's default listens on every
|
|
# interface, and this endpoint accepts writes as well as reads.
|
|
# Nothing authenticates it either — upstream offers one basic-auth
|
|
# pair (`-httpAuth.username` / `-httpAuth.password`) and this
|
|
# module sets neither. The bind address is the boundary.
|
|
listenAddress = "127.0.0.1:${toString cfg.port}";
|
|
|
|
extraOptions = [ "-retentionPeriod=${deployCfg.victorialogs.retentionPeriod}" ];
|
|
};
|
|
};
|
|
};
|
|
};
|
|
|
|
# 🔑 OTLP ingest is served at `/insert/opentelemetry/v1/logs`, measured
|
|
# against the pinned build rather than read from docs. The path differs
|
|
# from the metrics store's `/opentelemetry/api/v1/push`, so an exporter
|
|
# configured by analogy with the metrics one is silently wrong.
|
|
#
|
|
# ⚠️ And the status code cannot tell you which is which: a wrong path, a
|
|
# right path with a wrong body, and a nonsense path ALL answer 400. Only
|
|
# the server log distinguishes them — the real route complains about the
|
|
# encoding ("json encoding isn't supported ... use protobuf"), while
|
|
# anything else logs "unsupported path requested". Verified with a
|
|
# nonsense-path control, because a probe where every arm returns the same
|
|
# code is not evidence.
|
|
#
|
|
# ⚠️ Like the metrics store's, that endpoint is unauthenticated — which is
|
|
# why `listenAddress` above is loopback and why nothing reaches it except
|
|
# through the gateway, where the ingest route is authenticated. (This module
|
|
# does declare a vhost; the line that used to say otherwise was already
|
|
# wrong before the collector started pushing through it.)
|
|
}
|