Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/services.md
atlas 5ec658e0fa docs(networking): drop remaining stale authelia-empty-user claims
error-pages.nix paragraph (gateway.md:433) blamed a dead authelia
upstream on an empty user set; the real reason the route earns a
custom page is that a bare 502 there blames the proxy while the
gateway itself is fine. gateway.md:38 dropped 'yet' from the
placeholder-while-empty phrasing. services.md:109 corrected
'seeds an empty users database' to the disabled placeholder subject
swarm-authelia.nix actually seeds (swarm-authelia.nix:873).

Refs #3902
2026-10-02 17:25:37 +02:00

12 KiB

Swarm-wide services

Some things exist once per swarm rather than once per hive: the forge, the matrix homeserver, SSO, the secret store, the queue, and the metrics and log stack. This page says which host runs them and what a hive that runs none of them configures instead. The all-local quick start sets everything with one line → README.

Two options say where they live, and everything else derives:

services.hyperhive.deploy.singleHostSwarm = true;   # everything on this box
# or, for a dedicated services host with hives elsewhere:
services.hyperhive.deploy.allSwarmServices = true;

deploy.allSwarmServices is what "the swarm's shared services run here" means: every once-per-swarm service takes its enable from it. That's the whole rule, stated once — the per-service sections below don't repeat it, so a service that stops deriving is a visible difference rather than one more paragraph saying the same thing.

singleHostSwarm is the all-on-one-box mode above it. It defaults deploy.allSwarmServices, the swarm CA (swarm.ca.autoConfigure, generated on this host), the swarm controller (deploy.swarm-controller.enable), the host's /etc/hosts entries for the names it serves (gateway.localHostsEntry), the queue's auth-callout keys (deploy.nats.autoGenerateCallout) and where the secret store's bootstrap token goes (deploy.bao.bootstrapTokenFile). You can still set each derived toggle on its own, and that setting wins.

Both default to off, and that's deliberate: a host can't tell whether it's meant to be the swarm's service host, so this is an operator saying so rather than something inferred. With them off, a hive is a client of those services — it configures how to reach them and runs none of them.

That includes the forge (deploy.forgejo.enable). Every swarm needs one, being the canonical store for the meta flake and every agent's config repo, so exactly one host must turn it on: singleHostSwarm, allSwarmServices, or deploy.forgejo.enable = true set by hand. A host that enables its services one by one without any of those runs no forge.

A hive that runs none of these reaches each one by name — the forge at forge.<swarm-domain>, for example. Only the host running a service answers its name from its own resolver, so on a swarm spread over more than one host, the operator's DNS has to resolve those names to that host.

Moving an existing hive's forge to the swarm's

A hive that stops running the forge keeps the old container's state at /var/lib/nixos-containers/hive-forge/. Nothing moves it to the swarm's forge: push anything worth keeping there by hand. Its /var/lib/hyperhive/forge-core-token came from that old forge and fails against the swarm's one.

Deployment shapes

Those two options are what makes the difference between deployments, so the shapes worth naming are the ones they produce:

  • All-local. Everything on one machine: singleHostSwarm = true, plus deploy.hive-controller.enable = true for a hive to run agents on. After the first switch, the steps in setup.md remain.
  • Services on the swarm controller host. Set deploy.allSwarmServices and deploy.swarm-controller.enable there, with hives elsewhere. The controller doesn't derive from allSwarmServices.
  • Fully spread out. One container / VM / machine per service, somewhere.

These are a set, not a ladder with a correct top, and hyperhive supports in-between shapes. You can set each derived toggle on its own (see above), which is what makes "all local except X" a configuration rather than an unsupported edge case. Nothing in hyperhive prescribes a deployment model, so a doc that treats one shape as the real one and the others as compromises is wrong about the system rather than opinionated.

The practical consequence is that co-location is normal: in the first two shapes services share a host by design. Where a specific service has something to say about sharing a host --- extra requirements, or a cost worth weighing --- that belongs in the doc for that service, not here.

Services

SSO (authelia)

One authelia per swarm, in a swarm-authelia container, at auth.<swarm-domain>. Operator and agents are both subjects of the same provider, differentiated by roles and claims rather than by mechanism — there is one IdP and one auth path.

  • deploy.authelia — run the container here.
  • swarm.authelia.url — where clients go to authenticate. Present on every hive and the same value on all of them: https://<swarm.authelia.domain>, whether or not this host runs the container. Set it explicitly when joining a swarm whose IdP is under another name.

Agent subjects come from swarm-controller's agent-creation job, written into the users database by swarm-authelia-bridge; human ones come from swarmctl user add → setup.md § 2. On first boot this module seeds the users database with a disabled placeholder subject, so authelia has a non-empty store to start against before any real account exists (swarm-authelia.nix:873). Authelia generates session and storage keys in the container on first boot and never rotates them automatically; replacing one invalidates data already written (sessions, the encrypted store), so that's an operator action.

Storage is local sqlite and the notifier writes to a file. Both are small-deployment choices, and the scope is the justification: redis buys shared session state across replicas and there is one instance; SMTP exists to mail humans, and provisioning here is programmatic.

See sso.md for bootstrapping the first user and the OIDC relying-party flow, and secrets.md for where authelia generates and reads each of its keys.

Metrics (VictoriaMetrics + Grafana)

The swarm's telemetry lands in one VictoriaMetrics, and one Grafana reads it, in two containers at metrics.<swarm-domain> and grafana.<swarm-domain>. Two containers rather than one so you can restart or break Grafana without taking the time-series database with it.

They derive together: a store with no UI is unreadable and a UI with no store is empty. To run one without the other, set it directly:

services.hyperhive.deploy.victoriametrics.enable = true;
services.hyperhive.deploy.grafana.enable = false;

⚠️ This starts a database that grows for as long as the swarm runs. See retentionPeriod below before leaving it at its default.

Option When you'd touch it
deploy.victoriametrics.retentionPeriod Default 5y. Lower it once you have measured how fast this swarm actually fills a disk — the default is deliberately generous because too-short silently discards history you can't get back.
swarm.grafana.oidc.role Default Admin for everyone who logs in. Lower to Viewer/Editor if the swarm grows operators who shouldn't be able to reconfigure Grafana.
deploy.grafana.datasourceUrl Only if you front VictoriaMetrics with something else. It defaults to the store on this host, which is the only thing it can reach.

Logging in. Grafana is behind swarm SSO, so the accounts are the authelia ones — there is no separate Grafana password, and this module switches off the local login form whenever SSO is configured. If you enable Grafana on a host with no authelia, the form stays on and Grafana's default admin/admin applies; change it before exposing that host.

Where the data comes from. The swarm's OTEL collector, below.

Neither container is reachable except through the gateway: both bind loopback, and VictoriaMetrics' write endpoint takes no credential, so the collector is the only intended writer.

Logs (VictoriaLogs)

Each hive ships its journals to one VictoriaLogs at logs.<swarm-domain>, behind the same SSO as everything else: the swarm's own service containers, the hive's daemons and infra containers, and the harness units inside every agent container. The collector below is what writes to it.

Reading them. Open Grafana, pick Explore, and choose the VictoriaLogs datasource — it's provisioned for you. Grafana's Logs Drilldown app is deliberately not installed: it only supports Loki, and no setting here changes that, so Explore is the log browser for this swarm.

Option When you'd touch it
deploy.victorialogs.retentionPeriod Default 30d, far shorter than the metrics store's — logs are bulkier per unit of value and are typically read within days of ingestion. Raise it if you need to answer questions about last quarter.
swarm.victorialogs.domain Only to rename it.
swarm.victorialogs.port Only if something else on the services host already claims 9428.

Like the metrics store, it binds loopback and takes no credential of its own: the gateway vhost is the only way in, and the collector is the only intended writer.

Telemetry collector (OTEL)

The swarm's collector receives from every hive's own collector and is the only process that decides where telemetry goes: it always writes the stores above, and exports to otel.endpoint as well when that's set. It holds the upstream credential, which is why no hive and no agent needs one.

It runs in a swarm-otel container. Its swarm.otel.port defaults to 4319 rather than OTLP's usual 4318, which the hive tier uses — swarm containers share the host's network namespace, so two collectors on one port is a coin toss at runtime rather than an error at build time.

Every hive's own collector reaches this one by its gateway name, swarm.otel.domain (default otel.<swarm domain>) — the same by-domain-through-the-gateway shape every other swarm service uses, not a loopback URL an operator has to redirect. Nothing needs setting on a hive that doesn't run the swarm's services; the name resolves through the gateway either way.

Option When you'd touch it
swarm.otel.domain Only to rename it — the default already resolves correctly for every hive in the swarm.
swarm.otel.port Only if something else on the services host already claims 4319.

Both store exporters are unconditional, and deploy.victoriametrics.enable doesn't gate them: that option says this host runs the store, while the swarm has one either way, reached by its swarm name through the gateway.

Agent-side configuration, and what a hive's own collector does, are in ../scheduler/observability.md.