hyperhive/docs/swarm/services.md
iris fab2a0dedc docs: fix genuine passive-voice hits in docs/swarm
Read all 94 write-good.Passive hits across docs/swarm/ (ca.md,
README.md, secrets.md, services.md, sso.md, ui.md) in context. 44 are
genuine catches with a nameable, usually already-established actor
(swarm-controller, authelia, swarmctl, the controller, the gateway,
this module, hyperhive itself, or 'the operator' for manual actions) —
rewritten to active. 50 are legitimate passives or false catches, left
alone: predicate-adjective state descriptions (is expected/misconfigured/
broken), negative-capability idioms (no X is needed/placed, can't be
Yed/listed/fetched), config-state conditionals (whenever/when X is
enabled/configured/set), requirement-list labels (is required),
'is tracked as' idiom, backward-looking changelog facts with no actor
(was removed/verified/introduced), ambiguous-actor statements left
conservatively alone (agents are created and destroyed — could be
hive-c0re or swarm-controller, doc doesn't say), and a couple of
deliberately-parallel idiom pairs.

Several sibling-inconsistency fixes: a passive clause sitting next to
an already-active sibling describing the same fact/mechanism (ca.md's
two-bullet consumer list, README's 4-item WireGuard-mesh bullet list,
README's controller-registers-hooks paragraph, sso.md's followed-a-302
sentence).

Verified via vale on the whole directory, diffed against main's exact
baseline (not just the Passive count): write-good.Passive 94 -> 50
exactly, every other category unchanged (1 pre-existing
Microsoft.Contractions error at services... at secrets.md:182,
8 TooWordy, 1 Microsoft.We, 1 Microsoft.FirstPerson — same counts,
same locations).
2026-09-08 15:54:16 +02:00

186 lines
10 KiB
Markdown

# Swarm-wide services
Some things exist once per **swarm** rather than once per hive. Two
options say where the optional ones live, and everything else derives:
```nix
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
# or, for a dedicated services host with hives elsewhere:
services.hyperhive.deploy.allSwarmServices = true;
```
**`deploy.allSwarmServices` is what "the swarm's shared services run
here" means: every once-per-swarm service that's _optional_ takes its
`enable` from it.** That's the whole rule, stated once — the per-service
sections below don't repeat it, so a service that stops deriving is a
visible difference rather than one more paragraph saying the same thing.
`singleHostSwarm` is the all-on-one-box switch above it: it defaults
both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host
generates the swarm CA here). Each derived toggle can still be set on its own,
which wins, so "all local except X" needs no further option.
**Both default to off**, and that's deliberate: a host can't tell
whether it's meant to be the swarm's service host, so this is an
operator saying so rather than something inferred. With them off, a hive
is a _client_ of those services — it configures how to reach them and
runs none of them.
The forge is the exception, and not because it's per-hive: it's
swarm-wide but **not optional**, being the canonical store for the meta
flake and every agent's config repo, so it deploys with hyperhive itself
and has no `enable` to derive from anything.
## Deployment shapes
Those two options are what makes the difference between deployments, so
the shapes worth naming are the ones they produce:
- **All-local.** Everything on one machine:
`singleHostSwarm = true`. Setup is automatic apart from
choosing a domain and creating the first user.
- **Services on the swarm controller host.**
`deploy.allSwarmServices = true` there; the required services
deploy together on that host, with hives elsewhere.
- **Fully spread out.** One container / VM / machine per service,
somewhere.
**These are a set, not a ladder with a correct top, and hyperhive
supports in-between shapes.** Each derived toggle can be set on its own (see
above), which is what makes "all local except X" a configuration rather
than an unsupported edge case. Nothing in hyperhive prescribes a
deployment model, so a doc that treats one shape as the real one and
the others as compromises is wrong about the system rather than
opinionated.
The practical consequence is that **co-location is normal**: in the
first two shapes services share a host by design. Where a specific
service has something to say about sharing a host --- extra
requirements, or a cost worth weighing --- that belongs in the doc for
that service, not here.
## Services
### SSO (authelia)
One authelia per swarm, in a `swarm-authelia` container, at
`auth.<swarm-domain>`. Operator and agents are both subjects of the same
provider, differentiated by roles and claims rather than by mechanism —
there is one IdP and one auth path.
- **`deploy.authelia`** — run the container here.
- **`swarm.authelia.url`** — where clients go to authenticate.
Present on **every** hive, defaulting to this host's own instance only
when this module is the thing running it; otherwise `null`, and a hive
joining someone else's swarm sets it explicitly. Null means "no SSO
configured", and consumers say so rather than guessing an address.
swarm-controller writes the users database, not by hand: agents
are created and destroyed continuously, so the subject set is dynamic.
This module only guarantees the file exists and parses, so authelia
starts with nobody in it rather than failing to start — a provider with
no subjects yet is the correct state before anything has provisioned
them. Authelia generates session and storage keys in the container on
first boot and never rotates them automatically; replacing one
invalidates data already written (sessions, the encrypted store), so
that's an operator action.
Storage is local sqlite and the notifier writes to a file. Both are
small-deployment choices, and the scope is the justification: redis
buys shared session state across replicas and there is one instance;
SMTP exists to mail humans, and provisioning here is programmatic.
See [`sso.md`](sso.md) for bootstrapping the first user and the OIDC
relying-party flow, and [`secrets.md`](secrets.md) for where authelia
generates and reads each of its keys.
### Metrics (VictoriaMetrics + Grafana)
The swarm's telemetry lands in one VictoriaMetrics, and one Grafana
reads it, in two containers at `metrics.<swarm-domain>` and
`grafana.<swarm-domain>`. Two containers rather than one so Grafana can
be restarted or broken without taking the time-series database with it.
They derive together: a store with no UI is unreadable and a UI with no
store is empty. To run one without the other, set it directly:
```nix
services.hyperhive.deploy.victoriametrics.enable = true;
services.hyperhive.deploy.grafana.enable = false;
```
⚠️ **This starts a database that grows for as long as the swarm runs.**
See `retentionPeriod` below before leaving it at its default.
| Option | When you'd touch it |
| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `deploy.victoriametrics.retentionPeriod` | Default `5y`. Lower it once you have measured how fast this swarm actually fills a disk — the default is deliberately generous because too-short silently discards history you can't get back. |
| `swarm.grafana.oidc.role` | Default `Admin` for everyone who logs in. Lower to `Viewer`/`Editor` if the swarm grows operators who shouldn't be able to reconfigure Grafana. |
| `deploy.grafana.datasourceUrl` | Only if you front VictoriaMetrics with something else. It defaults to the store on this host, which is the only thing it can reach. |
**Logging in.** Grafana is behind swarm SSO, so the accounts are the
authelia ones — there is no separate Grafana password, and this module
switches off the local login form whenever SSO is configured. If you enable
Grafana on a host with no authelia, the form stays on and Grafana's
default `admin`/`admin` applies; change it before exposing that host.
**Where the data comes from.** The swarm's OTEL collector, below.
Neither container is reachable except through the gateway: both bind
loopback, and VictoriaMetrics' write endpoint takes no credential, so
the collector is the only intended writer.
### Logs (VictoriaLogs)
Each hive ships its journals to one VictoriaLogs at `logs.<swarm-domain>`,
behind the same SSO as everything else: the swarm's own service containers,
the hive's daemons and infra containers, and the harness units inside every
agent container. The collector below is what writes to it.
**Reading them.** Open Grafana, pick **Explore**, and choose the
`VictoriaLogs` datasource — it's provisioned for you. Grafana's _Logs
Drilldown_ app is deliberately not installed: it only supports Loki, and
no setting here changes that, so Explore is the log browser for this
swarm.
| Option | When you'd touch it |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `deploy.victorialogs.retentionPeriod` | Default `30d`, far shorter than the metrics store's — logs are bulkier per unit of value and are typically read within days of being written. Raise it if you need to answer questions about last quarter. |
| `swarm.victorialogs.domain` | Only to rename it. |
| `swarm.victorialogs.port` | Only if something else on the services host already claims `9428`. |
Like the metrics store, it binds loopback and takes no credential of its
own: the gateway vhost is the only way in, and the collector is the only
intended writer.
### Telemetry collector (OTEL)
The swarm's collector receives from every hive's own collector and is the
only process that decides where telemetry goes: it writes the store above
and exports to `otel.endpoint`, doing both when both are configured. It
also holds the upstream credential, which is why no hive and no agent
needs one.
It runs in a `swarm-otel` container. Its `swarm.otel.port` defaults to
`4319` rather than OTLP's usual `4318`, which the hive tier uses — swarm
containers share the host's network namespace, so two collectors on one
port is a coin toss at runtime rather than an error at build time.
Every hive's own collector reaches this one by its gateway name,
`swarm.otel.domain` (default `otel.<swarm domain>`) — the same
by-domain-through-the-gateway shape every other swarm service uses, not a
loopback URL an operator has to redirect. Nothing needs setting on a hive
that doesn't run the swarm's services; the name resolves through the
gateway either way.
| Option | When you'd touch it |
| ------------------- | --------------------------------------------------------------------------------------- |
| `swarm.otel.domain` | Only to rename it — the default already resolves correctly for every hive in the swarm. |
| `swarm.otel.port` | Only if something else on the services host already claims `4319`. |
With neither `otel.endpoint` nor the store enabled, this module refuses
the collector at eval — a tier that receives samples and drops them
looks healthy while losing data.
Agent-side configuration, and what a hive's own collector does, are in
[`../scheduler/observability.md`](../scheduler/observability.md).