Watch
0
0
Fork
You've already forked hyperhive
0

docs(swarm): facts + structure pass

swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
This commit is contained in:
atlas 2026-10-01 23:25:24 +02:00 • committed by mara
commit 270430a4b4
5 changed files with 408 additions and 402 deletions

View file

@ -1,7 +1,12 @@
# Swarm-wide services
Some things exist once per **swarm** rather than once per hive. Two
options say where the optional ones live, and everything else derives:
Some things exist once per **swarm** rather than once per hive: the forge,
the matrix homeserver, SSO, the secret store, the queue, and the metrics and
log stack. This page says which host runs them and what a hive that runs none
of them configures instead. The all-local quick start sets everything with one
line → [README](../../README.md#quick-start-an-all-local-swarm).
Two options say where they live, and everything else derives:
```nix
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
@ -14,10 +19,15 @@ here" means: every once-per-swarm service takes its `enable` from it.** That's t
sections below don't repeat it, so a service that stops deriving is a
visible difference rather than one more paragraph saying the same thing.
`singleHostSwarm` is the all-on-one-box switch above it: it defaults
both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host
generates the swarm CA here). You can still set each derived toggle on its own,
which wins, so "all local except X" needs no further option.
`singleHostSwarm` is the all-on-one-box mode above it. It defaults
`deploy.allSwarmServices`, the swarm CA (`swarm.ca.autoConfigure`, generated
on this host), the swarm controller (`deploy.swarm-controller.enable`), the
host's `/etc/hosts` entries for the names it serves
(`gateway.localHostsEntry`), the queue's auth-callout keys
(`deploy.nats.autoGenerateCallout`) and where the secret store's bootstrap
token goes (`deploy.bao.bootstrapTokenFile`). You can still set each derived
toggle on its own, which wins, so "all local except X" needs no further
option.
**Both default to off**, and that's deliberate: a host can't tell
whether it's meant to be the swarm's service host, so this is an
@ -38,23 +48,29 @@ answers its name from its own resolver, so on a swarm spread over
more than one host, the operator's DNS has to resolve those names to that
host.
<details><summary>Moving an existing hive's forge to the swarm's</summary>
A hive that stops running the forge keeps the old container's state at
`/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's
forge: push anything worth keeping there by hand. Its
`/var/lib/hyperhive/forge-core-token` came from that old forge and
fails against the swarm's one.
</details>
## Deployment shapes
Those two options are what makes the difference between deployments, so
the shapes worth naming are the ones they produce:
- **All-local.** Everything on one machine:
`singleHostSwarm = true`. Setup is automatic apart from
choosing a domain and creating the first user.
`singleHostSwarm = true`, plus `deploy.hive-controller.enable = true` for
a hive to run agents on. After the first switch, the steps in
[`setup.md`](../getting-started/setup.md) remain.
- **Services on the swarm controller host.**
`deploy.allSwarmServices = true` there; the required services
deploy together on that host, with hives elsewhere.
Set `deploy.allSwarmServices` and `deploy.swarm-controller.enable` there,
with hives elsewhere. The controller doesn't derive from
`allSwarmServices`.
- **Fully spread out.** One container / VM / machine per service,
somewhere.
@ -88,12 +104,10 @@ there is one IdP and one auth path.
container. Set it explicitly when joining a swarm whose IdP is under
another name.
swarm-controller writes the users database, not by hand: hive-c0re
creates and destroys agents continuously, so the subject set is dynamic.
This module only guarantees the file exists and parses, so authelia
starts with nobody in it rather than failing to start — a provider with
no subjects yet is the correct state before anything has provisioned
them. Authelia generates session and storage keys in the container on
Agent subjects come from swarm-controller's agent-creation job, written
into the users database by `swarm-authelia-bridge`; human ones come from
`swarmctl user add` → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
On first boot this module seeds an empty users database. Authelia generates session and storage keys in the container on
first boot and never rotates them automatically; replacing one
invalidates data already written (sessions, the encrypted store), so
that's an operator action.
@ -196,9 +210,9 @@ gateway either way.
**Both store exporters are unconditional**, and `deploy.victoriametrics.enable`
doesn't gate them: that option says this host _runs_ the store, while the swarm
has one either way, reached by its swarm name through the gateway. Gating on it
once left a collector on any other host with no exporter at all — receiving from
every hive and dropping it, silently, because an absent exporter isn't an error.
has one either way, reached by its swarm name through the gateway. A collector
with no exporter would receive from every hive and drop it silently, because an
absent exporter isn't an error.
Agent-side configuration, and what a hive's own collector does, are in
[`../scheduler/observability.md`](../scheduler/observability.md).