Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/services.md
atlas 270430a4b4 docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
2026-10-02 12:50:34 +02:00

218 lines
12 KiB
Markdown

# Swarm-wide services
Some things exist once per **swarm** rather than once per hive: the forge,
the matrix homeserver, SSO, the secret store, the queue, and the metrics and
log stack. This page says which host runs them and what a hive that runs none
of them configures instead. The all-local quick start sets everything with one
line → [README](../../README.md#quick-start-an-all-local-swarm).
Two options say where they live, and everything else derives:
```nix
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
# or, for a dedicated services host with hives elsewhere:
services.hyperhive.deploy.allSwarmServices = true;
```
**`deploy.allSwarmServices` is what "the swarm's shared services run
here" means: every once-per-swarm service takes its `enable` from it.** That's the whole rule, stated once — the per-service
sections below don't repeat it, so a service that stops deriving is a
visible difference rather than one more paragraph saying the same thing.
`singleHostSwarm` is the all-on-one-box mode above it. It defaults
`deploy.allSwarmServices`, the swarm CA (`swarm.ca.autoConfigure`, generated
on this host), the swarm controller (`deploy.swarm-controller.enable`), the
host's `/etc/hosts` entries for the names it serves
(`gateway.localHostsEntry`), the queue's auth-callout keys
(`deploy.nats.autoGenerateCallout`) and where the secret store's bootstrap
token goes (`deploy.bao.bootstrapTokenFile`). You can still set each derived
toggle on its own, which wins, so "all local except X" needs no further
option.
**Both default to off**, and that's deliberate: a host can't tell
whether it's meant to be the swarm's service host, so this is an
operator saying so rather than something inferred. With them off, a hive
is a _client_ of those services — it configures how to reach them and
runs none of them.
That includes the forge (`deploy.forgejo.enable`). Every swarm needs
one, being the canonical store for the meta flake and every agent's
config repo, so **exactly one host must turn it on**: `singleHostSwarm`,
`allSwarmServices`, or `deploy.forgejo.enable = true` set by hand. A
host that enables its services one by one without any of those runs no
forge.
A hive that runs none of these reaches each one by name — the forge at
`forge.<swarm-domain>`, for example. Only the host running a service
answers its name from its own resolver, so on a swarm spread over
more than one host, the operator's DNS has to resolve those names to that
host.
<details><summary>Moving an existing hive's forge to the swarm's</summary>
A hive that stops running the forge keeps the old container's state at
`/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's
forge: push anything worth keeping there by hand. Its
`/var/lib/hyperhive/forge-core-token` came from that old forge and
fails against the swarm's one.
</details>
## Deployment shapes
Those two options are what makes the difference between deployments, so
the shapes worth naming are the ones they produce:
- **All-local.** Everything on one machine:
`singleHostSwarm = true`, plus `deploy.hive-controller.enable = true` for
a hive to run agents on. After the first switch, the steps in
[`setup.md`](../getting-started/setup.md) remain.
- **Services on the swarm controller host.**
Set `deploy.allSwarmServices` and `deploy.swarm-controller.enable` there,
with hives elsewhere. The controller doesn't derive from
`allSwarmServices`.
- **Fully spread out.** One container / VM / machine per service,
somewhere.
**These are a set, not a ladder with a correct top, and hyperhive
supports in-between shapes.** You can set each derived toggle on its own (see
above), which is what makes "all local except X" a configuration rather
than an unsupported edge case. Nothing in hyperhive prescribes a
deployment model, so a doc that treats one shape as the real one and
the others as compromises is wrong about the system rather than
opinionated.
The practical consequence is that **co-location is normal**: in the
first two shapes services share a host by design. Where a specific
service has something to say about sharing a host --- extra
requirements, or a cost worth weighing --- that belongs in the doc for
that service, not here.
## Services
### SSO (authelia)
One authelia per swarm, in a `swarm-authelia` container, at
`auth.<swarm-domain>`. Operator and agents are both subjects of the same
provider, differentiated by roles and claims rather than by mechanism —
there is one IdP and one auth path.
- **`deploy.authelia`** — run the container here.
- **`swarm.authelia.url`** — where clients go to authenticate. Present
on **every** hive and the same value on all of them:
`https://<swarm.authelia.domain>`, whether or not this host runs the
container. Set it explicitly when joining a swarm whose IdP is under
another name.
Agent subjects come from swarm-controller's agent-creation job, written
into the users database by `swarm-authelia-bridge`; human ones come from
`swarmctl user add` → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
On first boot this module seeds an empty users database. Authelia generates session and storage keys in the container on
first boot and never rotates them automatically; replacing one
invalidates data already written (sessions, the encrypted store), so
that's an operator action.
Storage is local sqlite and the notifier writes to a file. Both are
small-deployment choices, and the scope is the justification: redis
buys shared session state across replicas and there is one instance;
SMTP exists to mail humans, and provisioning here is programmatic.
See [`sso.md`](sso.md) for bootstrapping the first user and the OIDC
relying-party flow, and [`secrets.md`](secrets.md) for where authelia
generates and reads each of its keys.
### Metrics (VictoriaMetrics + Grafana)
The swarm's telemetry lands in one VictoriaMetrics, and one Grafana
reads it, in two containers at `metrics.<swarm-domain>` and
`grafana.<swarm-domain>`. Two containers rather than one so you can
restart or break Grafana without taking the time-series database with it.
They derive together: a store with no UI is unreadable and a UI with no
store is empty. To run one without the other, set it directly:
```nix
services.hyperhive.deploy.victoriametrics.enable = true;
services.hyperhive.deploy.grafana.enable = false;
```
⚠️ **This starts a database that grows for as long as the swarm runs.**
See `retentionPeriod` below before leaving it at its default.
| Option | When you'd touch it |
| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `deploy.victoriametrics.retentionPeriod` | Default `5y`. Lower it once you have measured how fast this swarm actually fills a disk — the default is deliberately generous because too-short silently discards history you can't get back. |
| `swarm.grafana.oidc.role` | Default `Admin` for everyone who logs in. Lower to `Viewer`/`Editor` if the swarm grows operators who shouldn't be able to reconfigure Grafana. |
| `deploy.grafana.datasourceUrl` | Only if you front VictoriaMetrics with something else. It defaults to the store on this host, which is the only thing it can reach. |
<!-- vale write-good.Passive = NO -->
**Logging in.** Grafana is behind swarm SSO, so the accounts are the
authelia ones — there is no separate Grafana password, and this module
switches off the local login form whenever SSO is configured. If you enable
Grafana on a host with no authelia, the form stays on and Grafana's
default `admin`/`admin` applies; change it before exposing that host.
<!-- vale write-good.Passive = YES -->
**Where the data comes from.** The swarm's OTEL collector, below.
Neither container is reachable except through the gateway: both bind
loopback, and VictoriaMetrics' write endpoint takes no credential, so
the collector is the only intended writer.
### Logs (VictoriaLogs)
Each hive ships its journals to one VictoriaLogs at `logs.<swarm-domain>`,
behind the same SSO as everything else: the swarm's own service containers,
the hive's daemons and infra containers, and the harness units inside every
agent container. The collector below is what writes to it.
**Reading them.** Open Grafana, pick **Explore**, and choose the
`VictoriaLogs` datasource — it's provisioned for you. Grafana's _Logs
Drilldown_ app is deliberately not installed: it only supports Loki, and
no setting here changes that, so Explore is the log browser for this
swarm.
| Option | When you'd touch it |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `deploy.victorialogs.retentionPeriod` | Default `30d`, far shorter than the metrics store's — logs are bulkier per unit of value and are typically read within days of ingestion. Raise it if you need to answer questions about last quarter. |
| `swarm.victorialogs.domain` | Only to rename it. |
| `swarm.victorialogs.port` | Only if something else on the services host already claims `9428`. |
Like the metrics store, it binds loopback and takes no credential of its
own: the gateway vhost is the only way in, and the collector is the only
intended writer.
### Telemetry collector (OTEL)
The swarm's collector receives from every hive's own collector and is the
only process that decides where telemetry goes: it always writes the stores
above, and exports to `otel.endpoint` as well when that's set. It holds
the upstream credential, which is why no hive and no agent needs one.
It runs in a `swarm-otel` container. Its `swarm.otel.port` defaults to
`4319` rather than OTLP's usual `4318`, which the hive tier uses — swarm
containers share the host's network namespace, so two collectors on one
port is a coin toss at runtime rather than an error at build time.
Every hive's own collector reaches this one by its gateway name,
`swarm.otel.domain` (default `otel.<swarm domain>`) — the same
by-domain-through-the-gateway shape every other swarm service uses, not a
loopback URL an operator has to redirect. Nothing needs setting on a hive
that doesn't run the swarm's services; the name resolves through the
gateway either way.
| Option | When you'd touch it |
| ------------------- | --------------------------------------------------------------------------------------- |
| `swarm.otel.domain` | Only to rename it — the default already resolves correctly for every hive in the swarm. |
| `swarm.otel.port` | Only if something else on the services host already claims `4319`. |
**Both store exporters are unconditional**, and `deploy.victoriametrics.enable`
doesn't gate them: that option says this host _runs_ the store, while the swarm
has one either way, reached by its swarm name through the gateway. A collector
with no exporter would receive from every hive and drop it silently, because an
absent exporter isn't an error.
Agent-side configuration, and what a hive's own collector does, are in
[`../scheduler/observability.md`](../scheduler/observability.md).