Read all 94 write-good.Passive hits across docs/swarm/ (ca.md, README.md, secrets.md, services.md, sso.md, ui.md) in context. 44 are genuine catches with a nameable, usually already-established actor (swarm-controller, authelia, swarmctl, the controller, the gateway, this module, hyperhive itself, or 'the operator' for manual actions) — rewritten to active. 50 are legitimate passives or false catches, left alone: predicate-adjective state descriptions (is expected/misconfigured/ broken), negative-capability idioms (no X is needed/placed, can't be Yed/listed/fetched), config-state conditionals (whenever/when X is enabled/configured/set), requirement-list labels (is required), 'is tracked as' idiom, backward-looking changelog facts with no actor (was removed/verified/introduced), ambiguous-actor statements left conservatively alone (agents are created and destroyed — could be hive-c0re or swarm-controller, doc doesn't say), and a couple of deliberately-parallel idiom pairs. Several sibling-inconsistency fixes: a passive clause sitting next to an already-active sibling describing the same fact/mechanism (ca.md's two-bullet consumer list, README's 4-item WireGuard-mesh bullet list, README's controller-registers-hooks paragraph, sso.md's followed-a-302 sentence). Verified via vale on the whole directory, diffed against main's exact baseline (not just the Passive count): write-good.Passive 94 -> 50 exactly, every other category unchanged (1 pre-existing Microsoft.Contractions error at services... at secrets.md:182, 8 TooWordy, 1 Microsoft.We, 1 Microsoft.FirstPerson — same counts, same locations).
186 lines
10 KiB
Markdown
186 lines
10 KiB
Markdown
# Swarm-wide services
|
|
|
|
Some things exist once per **swarm** rather than once per hive. Two
|
|
options say where the optional ones live, and everything else derives:
|
|
|
|
```nix
|
|
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
|
|
# or, for a dedicated services host with hives elsewhere:
|
|
services.hyperhive.deploy.allSwarmServices = true;
|
|
```
|
|
|
|
**`deploy.allSwarmServices` is what "the swarm's shared services run
|
|
here" means: every once-per-swarm service that's _optional_ takes its
|
|
`enable` from it.** That's the whole rule, stated once — the per-service
|
|
sections below don't repeat it, so a service that stops deriving is a
|
|
visible difference rather than one more paragraph saying the same thing.
|
|
|
|
`singleHostSwarm` is the all-on-one-box switch above it: it defaults
|
|
both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host
|
|
generates the swarm CA here). Each derived toggle can still be set on its own,
|
|
which wins, so "all local except X" needs no further option.
|
|
|
|
**Both default to off**, and that's deliberate: a host can't tell
|
|
whether it's meant to be the swarm's service host, so this is an
|
|
operator saying so rather than something inferred. With them off, a hive
|
|
is a _client_ of those services — it configures how to reach them and
|
|
runs none of them.
|
|
|
|
The forge is the exception, and not because it's per-hive: it's
|
|
swarm-wide but **not optional**, being the canonical store for the meta
|
|
flake and every agent's config repo, so it deploys with hyperhive itself
|
|
and has no `enable` to derive from anything.
|
|
|
|
## Deployment shapes
|
|
|
|
Those two options are what makes the difference between deployments, so
|
|
the shapes worth naming are the ones they produce:
|
|
|
|
- **All-local.** Everything on one machine:
|
|
`singleHostSwarm = true`. Setup is automatic apart from
|
|
choosing a domain and creating the first user.
|
|
- **Services on the swarm controller host.**
|
|
`deploy.allSwarmServices = true` there; the required services
|
|
deploy together on that host, with hives elsewhere.
|
|
- **Fully spread out.** One container / VM / machine per service,
|
|
somewhere.
|
|
|
|
**These are a set, not a ladder with a correct top, and hyperhive
|
|
supports in-between shapes.** Each derived toggle can be set on its own (see
|
|
above), which is what makes "all local except X" a configuration rather
|
|
than an unsupported edge case. Nothing in hyperhive prescribes a
|
|
deployment model, so a doc that treats one shape as the real one and
|
|
the others as compromises is wrong about the system rather than
|
|
opinionated.
|
|
|
|
The practical consequence is that **co-location is normal**: in the
|
|
first two shapes services share a host by design. Where a specific
|
|
service has something to say about sharing a host --- extra
|
|
requirements, or a cost worth weighing --- that belongs in the doc for
|
|
that service, not here.
|
|
|
|
## Services
|
|
|
|
### SSO (authelia)
|
|
|
|
One authelia per swarm, in a `swarm-authelia` container, at
|
|
`auth.<swarm-domain>`. Operator and agents are both subjects of the same
|
|
provider, differentiated by roles and claims rather than by mechanism —
|
|
there is one IdP and one auth path.
|
|
|
|
- **`deploy.authelia`** — run the container here.
|
|
- **`swarm.authelia.url`** — where clients go to authenticate.
|
|
Present on **every** hive, defaulting to this host's own instance only
|
|
when this module is the thing running it; otherwise `null`, and a hive
|
|
joining someone else's swarm sets it explicitly. Null means "no SSO
|
|
configured", and consumers say so rather than guessing an address.
|
|
|
|
swarm-controller writes the users database, not by hand: agents
|
|
are created and destroyed continuously, so the subject set is dynamic.
|
|
This module only guarantees the file exists and parses, so authelia
|
|
starts with nobody in it rather than failing to start — a provider with
|
|
no subjects yet is the correct state before anything has provisioned
|
|
them. Authelia generates session and storage keys in the container on
|
|
first boot and never rotates them automatically; replacing one
|
|
invalidates data already written (sessions, the encrypted store), so
|
|
that's an operator action.
|
|
|
|
Storage is local sqlite and the notifier writes to a file. Both are
|
|
small-deployment choices, and the scope is the justification: redis
|
|
buys shared session state across replicas and there is one instance;
|
|
SMTP exists to mail humans, and provisioning here is programmatic.
|
|
|
|
See [`sso.md`](sso.md) for bootstrapping the first user and the OIDC
|
|
relying-party flow, and [`secrets.md`](secrets.md) for where authelia
|
|
generates and reads each of its keys.
|
|
|
|
### Metrics (VictoriaMetrics + Grafana)
|
|
|
|
The swarm's telemetry lands in one VictoriaMetrics, and one Grafana
|
|
reads it, in two containers at `metrics.<swarm-domain>` and
|
|
`grafana.<swarm-domain>`. Two containers rather than one so Grafana can
|
|
be restarted or broken without taking the time-series database with it.
|
|
|
|
They derive together: a store with no UI is unreadable and a UI with no
|
|
store is empty. To run one without the other, set it directly:
|
|
|
|
```nix
|
|
services.hyperhive.deploy.victoriametrics.enable = true;
|
|
services.hyperhive.deploy.grafana.enable = false;
|
|
```
|
|
|
|
⚠️ **This starts a database that grows for as long as the swarm runs.**
|
|
See `retentionPeriod` below before leaving it at its default.
|
|
|
|
| Option | When you'd touch it |
|
|
| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
| `deploy.victoriametrics.retentionPeriod` | Default `5y`. Lower it once you have measured how fast this swarm actually fills a disk — the default is deliberately generous because too-short silently discards history you can't get back. |
|
|
| `swarm.grafana.oidc.role` | Default `Admin` for everyone who logs in. Lower to `Viewer`/`Editor` if the swarm grows operators who shouldn't be able to reconfigure Grafana. |
|
|
| `deploy.grafana.datasourceUrl` | Only if you front VictoriaMetrics with something else. It defaults to the store on this host, which is the only thing it can reach. |
|
|
|
|
**Logging in.** Grafana is behind swarm SSO, so the accounts are the
|
|
authelia ones — there is no separate Grafana password, and this module
|
|
switches off the local login form whenever SSO is configured. If you enable
|
|
Grafana on a host with no authelia, the form stays on and Grafana's
|
|
default `admin`/`admin` applies; change it before exposing that host.
|
|
|
|
**Where the data comes from.** The swarm's OTEL collector, below.
|
|
|
|
Neither container is reachable except through the gateway: both bind
|
|
loopback, and VictoriaMetrics' write endpoint takes no credential, so
|
|
the collector is the only intended writer.
|
|
|
|
### Logs (VictoriaLogs)
|
|
|
|
Each hive ships its journals to one VictoriaLogs at `logs.<swarm-domain>`,
|
|
behind the same SSO as everything else: the swarm's own service containers,
|
|
the hive's daemons and infra containers, and the harness units inside every
|
|
agent container. The collector below is what writes to it.
|
|
|
|
**Reading them.** Open Grafana, pick **Explore**, and choose the
|
|
`VictoriaLogs` datasource — it's provisioned for you. Grafana's _Logs
|
|
Drilldown_ app is deliberately not installed: it only supports Loki, and
|
|
no setting here changes that, so Explore is the log browser for this
|
|
swarm.
|
|
|
|
| Option | When you'd touch it |
|
|
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
| `deploy.victorialogs.retentionPeriod` | Default `30d`, far shorter than the metrics store's — logs are bulkier per unit of value and are typically read within days of being written. Raise it if you need to answer questions about last quarter. |
|
|
| `swarm.victorialogs.domain` | Only to rename it. |
|
|
| `swarm.victorialogs.port` | Only if something else on the services host already claims `9428`. |
|
|
|
|
Like the metrics store, it binds loopback and takes no credential of its
|
|
own: the gateway vhost is the only way in, and the collector is the only
|
|
intended writer.
|
|
|
|
### Telemetry collector (OTEL)
|
|
|
|
The swarm's collector receives from every hive's own collector and is the
|
|
only process that decides where telemetry goes: it writes the store above
|
|
and exports to `otel.endpoint`, doing both when both are configured. It
|
|
also holds the upstream credential, which is why no hive and no agent
|
|
needs one.
|
|
|
|
It runs in a `swarm-otel` container. Its `swarm.otel.port` defaults to
|
|
`4319` rather than OTLP's usual `4318`, which the hive tier uses — swarm
|
|
containers share the host's network namespace, so two collectors on one
|
|
port is a coin toss at runtime rather than an error at build time.
|
|
|
|
Every hive's own collector reaches this one by its gateway name,
|
|
`swarm.otel.domain` (default `otel.<swarm domain>`) — the same
|
|
by-domain-through-the-gateway shape every other swarm service uses, not a
|
|
loopback URL an operator has to redirect. Nothing needs setting on a hive
|
|
that doesn't run the swarm's services; the name resolves through the
|
|
gateway either way.
|
|
|
|
| Option | When you'd touch it |
|
|
| ------------------- | --------------------------------------------------------------------------------------- |
|
|
| `swarm.otel.domain` | Only to rename it — the default already resolves correctly for every hive in the swarm. |
|
|
| `swarm.otel.port` | Only if something else on the services host already claims `4319`. |
|
|
|
|
With neither `otel.endpoint` nor the store enabled, this module refuses
|
|
the collector at eval — a tier that receives samples and drops them
|
|
looks healthy while losing data.
|
|
|
|
Agent-side configuration, and what a hive's own collector does, are in
|
|
[`../scheduler/observability.md`](../scheduler/observability.md).
|