diff --git a/docs/swarm/README.md b/docs/swarm/README.md index 9c233eb6..454cdebf 100644 --- a/docs/swarm/README.md +++ b/docs/swarm/README.md @@ -1,308 +1,55 @@ -# Multi-hive swarms +# The swarm -A **swarm** is a collection of agents that share an identity and -coordinate across one or more hives. A single hyperhive instance -running on one host is already a swarm (one hive). This doc covers -the additional config needed when the swarm spans multiple hosts. +The **swarm** is where things live: agent identities and accounts, secrets, +the job graph that creates and places agents, telemetry, and the UI you +drive it all from. **Hives are the substrate** — NixOS hosts that run agent +containers on the swarm's behalf. Every hive belongs to a swarm; a single +host is a swarm of one. -For the full option reference rather than prose: `services.hyperhive.swarm.*` -(swarm-wide facts, identical on every host) and `services.hyperhive.deploy.*` -(this host's own deployment decisions — does _this_ machine run grafana, -the swarm controller, authelia, …) are separate generated pages, `nix -build .#docs-swarm` / `.#docs-deploy` or the website's `/options/swarm.html` -/ `/options/deploy.html`. +This page is for the operator. It covers the control plane, how a hive +joins the directory, and how hives and agents report upward. The steps for a +fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local +config is the [README quick start](../../README.md#quick-start-an-all-local-swarm). -## Terminology +Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical +on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs +grafana, the controller, authelia, …) → +[options reference](https://hyperhive.darkest.space/options/), or +`nix build .#docs-swarm` / `.#docs-deploy`. -- **hive** — a single hyperhive installation on one host. Has its - own `services.hyperhive.domain` DNS name and its own set of agent - containers. -- **swarm** — one or more hives whose operators have declared them - as peers. You can qualify an agent as `agent@hive-domain`. -- **peer hive** — any hive in `services.hyperhive.swarm.hives` other - than this one. Peers are _derived_, not declared: the directory lists - every hive including yourself, and `hiveName` says which one you are. +## Where each piece lives -## Hive identity config - -```nix -services.hyperhive = { - swarm.domain = "example.com"; # required — the swarm's DNS domain - hiveName = "pr1ma"; # required — this hive's label in it - swarm.name = "constellat1on"; # shared swarm display name (optional) - - # required — the directory, identical on every host in the swarm. - # Names only: each entry's `domain` defaults to .. - swarm.hives = { - pr1ma = { }; - edge = { }; - }; -}; -``` - - - -`swarm.domain` and `hiveName` are **required** whenever hyperhive is -enabled; eval fails with a hint naming each. Neither defaults, -because a guessed value here is a wrong hostname that evaluates cleanly -and deploys — an eval failure asking the operator to write the address -down is the cheaper outcome. **Upgrading past this release means setting -both once.** - - - -You must still set `domain` too, but you no longer _write_ it: it's read from -this hive's own entry in the directory, whose `domain` defaults to -`.`. A conventional swarm states no addresses at -all, and a hive addressed by something else states it in the one place -the other hives read — `swarm.hives.edge.domain = "edge.elsewhere.example";`. - -Setting `services.hyperhive.domain` directly still works and still wins, -with a **deprecation warning**. The reason it's deprecated isn't tidiness: -that option is local to one host, and the operator copies the directory -to every host, so a value written only there leaves every peer pointing -somewhere else with nothing detecting the disagreement. - -⚠️ **Upgrading:** a hive that has been running on `swarm.domain` + -`hiveName` alone now needs its own directory entry — -`services.hyperhive.swarm.hives. = { };`, one line, no value. -Eval fails naming it if you forget. - -`domain` drives `HYPERHIVE_HIVE_DOMAIN` in every container so agents can -form qualified labels (`iris@pr1ma.example.com`). - -`swarm.name` is purely display — it surfaces in the dashboard chrome -header and per-agent system prompts, and federated hives at different -domains can share one. `hiveName` surfaces in the same places but is -_not_ only display: it's the leftmost label of the hive's domain. That -`swarm.name` sits under `swarm` and `hiveName` doesn't is the whole -distinction — one names this hive, the other names the group it belongs -to. - -See `docs/process/conventions.md` § Hive identity for the env-var chain -and `qualify()` / `qualified_label()` semantics. - -## Swarm CA - -A hive's internal TLS chains to a **swarm root CA**, so a peer that -trusts the root validates every hive in the swarm rather than pinning -to each one by hand. Provisioning modes, what to hand a peer -(`trust-bundle.pem`, never `ca.pem`), the name constraints on a hive -CA, and how an existing hive adopts the hierarchy: [`ca.md`](ca.md). - -## Running the swarm's shared services - -One authelia, one matrix, one forge per swarm — which host runs them, -and what a hive that runs none of them configures instead: -[`services.md`](services.md). - -## Single sign-on - -Which secrets the SSO provider generates, which one has a reader in -another container, and the three ways that one gets delivered: -[`sso.md`](sso.md). - -## Secrets - -Every credential the swarm holds, who mints it, where it must live, and -which of the three topologies makes it the operator's job to place: -[`secrets.md`](secrets.md). - -Where that shape is **going** — the per-secret minter/reader/renewal -contract, the target of one mTLS identity per host and everything else -through the store, and the test a change has to pass to count as movement -toward it: [`credentials.md`](credentials.md). It supersedes `secrets.md` -when the migration completes. - -## Swarm UI - -The operator-only web surface on the swarm apex, why reaching it needs -the `admins` group rather than just a session, and the four sites you -wire a swarm service name into: [`ui.md`](ui.md). - -## The swarm's hive directory - -```nix -services.hyperhive.swarm.hives = { - pr1ma = { }; # this host, per hiveName - lab = { }; # a second hive in the swarm - edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention -}; -``` - -One attrset describing **every** hive in the swarm, **including this -one**, keyed by that hive's `hiveName`. It's meant to be _identical on -every host_ — write it once, share it, and each host reads it correctly -because `services.hyperhive.hiveName` says which entry is itself. - -Empty (the default) means this host isn't in a swarm. Once non-empty it -**must** contain an entry for `hiveName`; eval fails naming the missing -hive. That assertion is load-bearing rather than pedantic — "my peers" -comes from _everything that isn't me_, so a directory that doesn't -contain you derives every hive as a peer and you peer with yourself. - -`domain` defaults to `.`, the convention every hive -follows, so a conventional directory is names only. The default is a -derivation from two values the operator already had to state — the swarm's -domain and the entry's own name — rather than a guess, which is what makes -it safe here when a guessed hostname wouldn't be. Set it only for a hive -addressed by something else. - -> **No per-hive CA field exists, and no per-hive cert pinning.** Trust -> inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every -> hive chains to it, so one anchor replaces per-hive pinning entirely. -> What that genuinely drops is trusting a hive whose root this swarm -> does _not_ own — another swarm's, or one keeping its own CA. That's -> a cross-swarm problem and wants a mechanism designed for it. (An -> earlier `certFingerprint` field existed for exactly that gap, pinning -> a peer's TLS leaf for hive-c0re's own peer HTTPS checks — removed -> along with the dashboard feature it existed to serve, since nothing -> else ever consumed it.) - -## What the config does at runtime - -1. **Swarm-wide hive roster** — swarm-controller reads this same - directory and serves it at `GET /api/hives`; `swarm-ui`'s overview - page renders it (`docs/swarm/ui.md`). This is the operator-facing - "what hives exist" surface — a per-hive dashboard "peer hives" - display existed here once; it no longer exists, in favour of this. - -2. **Matrix federation** — when `matrix.enable` is on, tuwunel - federates with the peer's matrix server (discovered via the peer's - `.well-known/matrix/server` delegation, which the gateway serves). - Federation validates the peer's TLS certificate against the matrix - **container's** trust bundle, independent of this directory. - - ⚠️ **That container currently trusts no swarm-internal CA**, so a - self-signed gateway certificate doesn't federate. You can't list the - swarm root there: `security.pki.certificateFiles` is - read when the system is _built_, and the root is a runtime file (its - key must never enter the store), so there is no build-time name for - it. Bridging that needs a runtime mechanism; a separate issue tracks - it. Until then, federation needs CA-issued certs (ACME). See - `docs/integrations/matrix.md` for federation firewall + TLS requirements. - -3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each - entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress` - to configure `wg-hive`. See "WireGuard inter-hive mesh" below. - -## One directory, not a bilateral declaration - -Both hives hold the **same** `hives` attrset; neither declares the -other. What differs between the two hosts is only `hiveName`: - -``` -# hive A # hive B -hiveName = "pr1ma"; hiveName = "edge"; -swarm.hives = { … }; swarm.hives = { … }; # byte-identical -``` - -That's the point of the shape, and it removes a class of bug rather -than saving typing: a per-host peer list let two hosts hold _different_ -facts about the same third hive — a stale endpoint, a rotated -fingerprint — with nothing to detect the disagreement. One entry per -hive makes it unrepresentable. - -## WireGuard inter-hive mesh (optional) - -The peer config above uses public HTTPS for all inter-hive traffic. -For private deployments — or to reduce latency and TLS overhead on -intra-swarm traffic — hive-c0re can configure a host-to-host -WireGuard mesh. - -### Generating keys - -On each hive host: - -```bash -wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key -wg pubkey < /etc/wireguard/hive.key # → share this with peer operators -``` - -### Config example (two hives) - -```nix -# hive A (pr1ma.example.com, mesh IP 10.100.0.1) -services.hyperhive = { - deploy.wireguard = { - enable = true; - privateKeyFile = "/etc/wireguard/hive.key"; - address = "10.100.0.1/24"; - listenPort = 51820; # optional, default 51820 - }; - - # The same `hives` attrset both hosts hold — mesh fields included, - # since "where this hive can be dialled" is a fact about that hive. - swarm.hives = { - pr1ma = { - domain = "pr1ma.example.com"; - wireguardPublicKey = "base64keyA="; - wireguardEndpoint = "198.51.100.1:51820"; - wireguardAddress = "10.100.0.1/32"; - }; - edge = { - domain = "edge.corp"; - wireguardPublicKey = "base64keyB="; - wireguardEndpoint = "203.0.113.42:51820"; - wireguardAddress = "10.100.0.2/32"; - }; - }; -}; - -# hive B (edge.corp, mesh IP 10.100.0.2) -services.hyperhive = { - deploy.wireguard = { - enable = true; - privateKeyFile = "/etc/wireguard/hive.key"; - address = "10.100.0.2/24"; - }; - - swarm.hives = { /* … identical to hive A's … */ }; -}; -``` - -### What the mesh does - -- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the - host (not inside agent containers; containers reach peers via the - host's routing table). -- It opens UDP port 51820 (or `listenPort`) on the host firewall. -- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly - from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s - `allowedIPs`, so intra-swarm traffic can route over the mesh address - rather than the public domain. -- It sets `persistentKeepalive = 25` by default; override or null to - disable (not needed when both sides have public IPs and no NAT). - -### NAT / one-sided endpoints - -If one host is behind NAT and can't accept incoming connections, only -that host needs a null `wireguardEndpoint` on the peer config — the -other side initiates. With keepalive on, the NAT hole stays open. - -If both hosts are behind NAT, you need a STUN relay or a third host -(exit node). Out of scope for v0. - -## Snapshot store - -One further option lives in this namespace but its docs live with the -service it points at: `services.hyperhive.swarm.snapshotStore.{address, -port}` tells this hive where the swarm's `btrfs receive` endpoint is, so -`hivectl agent subvol snapshot push` has somewhere to stream to. - -It's genuinely swarm-scoped rather than per-peer — a swarm has exactly -one store, because the receiver keys destinations by _agent_ so a -migrating agent keeps one unbroken incremental chain. See -[snapshot-store.md](../networking/snapshot-store.md). +- **control plane** — `swarm-controller`: hive directory, agent roster, job + graph, agent creation. → [below](#swarm-controller) +- **swarm UI** — the operator's day-to-day surface, on the swarm apex, + `admins` only. → [`ui.md`](ui.md) +- **shared services** — one forge, homeserver, SSO, queue and metrics/logs + stack, each on whichever host you put it. → [`services.md`](services.md) +- **SSO** — the secrets authelia generates and how each reaches its reader. + → [`sso.md`](sso.md) +- **secrets** — every credential the swarm holds, who mints it and where it + lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal + contract → [`credentials.md`](credentials.md) · how the store comes up, + who writes its grants and how it unseals → [`bao.md`](bao.md) +- **swarm CA** — the root every hive's internal TLS chains to, and what to + hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md) +- **snapshot store** — the swarm's one `btrfs receive` endpoint, + `swarm.snapshotStore.{address,port}`. → + [`snapshot-store.md`](../networking/snapshot-store.md) ## Swarm controller -`services.hyperhive.deploy.swarm-controller.enable` runs the `swarm-controller` -daemon on this host. **Off by default and deliberately not derived from -`services.hyperhive.deploy.hive-controller.enable`**: a swarm has one +`swarm-controller` is the swarm's control plane: it holds the hive directory, +the agent roster and their wanted state, and the job graph. Creating an agent — +from the swarm UI or `swarmctl agent create --hive ` — queues its SSO +identity, forge user, config repo, store identity and matrix account, then +sends the hive a deploy message. A new agent starts `paused`. + +`services.hyperhive.deploy.swarm-controller.enable` runs it on this host; +`singleHostSwarm` turns it on. **Otherwise off by default and deliberately +not derived from `deploy.hive-controller.enable`**: a swarm has one controller, so enabling it states a fact about swarm topology, not about -whether this host runs a hive. Every hive runs `hive-c0re` (the agents on that host); one -hive additionally runs this (what's true across hives). +whether this host runs a hive. What it serves, why it's a unix socket rather than a port, and the socket-directory constraint that governs where `socketPath` may point: @@ -409,6 +156,8 @@ how it gets there. A hive lacking the queue's address for its agents sets none of the four and each agent logs that it has none; a half-set environment logs an error and the harness keeps serving. +
What an agent publishes over the queue + What an agent does with that connection is publish its terminal. Every row its own web UI renders also goes to `$SWARM.term.`, one subject per agent, so a swarm-level terminal can follow one agent without subscribing to the swarm's @@ -468,6 +217,8 @@ key. Swarm-side, `GET /api/agents//icon` serves the stored bytes, and 404 means the agent has no icon. swarm-ui's agent cards load it as an `` and show the dimmed hyperhive mark for an agent without one, as the hive dashboard does. +
+ ### Swarm-wide forge objects The controller also keeps the forge objects that are one per swarm, not @@ -484,7 +235,7 @@ one per hive. It ensures them at start and every five minutes after - the `agent-configs` org avatar (`deploy.swarm-controller.configOrgAvatarPng`). -hive-c0re no longer creates any of them. A pass that can't finish logs a +A pass that can't finish logs a `warn` line per object plus `swarm forge objects: pass incomplete` in `journalctl -u swarm-controller`, and retries on the next tick. While the controller is down the objects stay as they are. @@ -503,14 +254,14 @@ re-derive. Approval happens once, at the swarm level: a hive receives a decision, not an event to adjudicate. **`internal/knowledge` is on that path.** The controller's is the only -hook on it: hives no longer register their own (see -`docs/integrations/knowledge.md` for clearing a leftover). A webhook has exactly one target URL, so per-hive -registration never added a recipient — it took delivery away from -whichever hive registered before it. +hook on it; hives register none of their own +([`knowledge.md`](../integrations/knowledge.md) covers clearing a leftover). +A webhook has exactly one target URL, so a second registration would take +delivery away from the first rather than add a recipient. -**The `agent-configs` org isn't yet.** Each hive still registers its own +**The `agent-configs` org isn't.** Each hive registers its own `pull_request` hook there, so that repo has two — the hive's and the controller's — and **both are expected; don't delete either.** Removing a hive's stops it acting on config PRs; removing the controller's just @@ -529,15 +280,206 @@ To check it's working, push to `internal/knowledge` and look for `webhook: verified delivery` in `journalctl -u swarm-controller`. A refused delivery logs `webhook: refused delivery` with the reason. +## Hives: the substrate + +- **hive** — one host running `hive-c0re` and its agent containers + (`deploy.hive-controller.enable`). Addressed as `.`. +- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the + shared services and controller. You can qualify an agent as + `agent@hive-domain`. +- **peer hive** — any hive in the directory other than this one. Peers are + _derived_, not declared: the directory lists every hive including + yourself, and `hiveName` says which one you are. + +### Hive identity config + +```nix +services.hyperhive = { + swarm.domain = "example.com"; # required — the swarm's DNS domain + hiveName = "pr1ma"; # required — this hive's label in it + swarm.name = "constellat1on"; # shared swarm display name (optional) + + # required — the directory, identical on every host in the swarm. + # Names only: each entry's `domain` defaults to .. + swarm.hives = { + pr1ma = { }; # this host, per hiveName + lab = { }; # a second hive + edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention + }; +}; +``` + + + +`swarm.domain` is **required** on a host that runs a hive, and `hiveName` on +a host that runs a hive, the secret store or the homeserver; eval fails with +a hint naming each. Neither defaults, because a guessed value here is a +wrong hostname that evaluates cleanly and deploys. + + + +**The directory is one attrset, identical on every host.** It describes +every hive in the swarm, **including this one**, keyed by `hiveName`; what +differs between hosts is only `hiveName`. It **must** contain an entry for +this host's `hiveName`, and an empty directory fails that check too — eval +names the missing hive. Peers are _every entry but this host's_, so a +directory without this host would make every hive a peer, itself included. + +``` +# hive A # hive B +hiveName = "pr1ma"; hiveName = "edge"; +swarm.hives = { … }; swarm.hives = { … }; # byte-identical +``` + +One entry per hive means two hosts can't hold _different_ facts about the +same third hive, such as a stale endpoint. + +`domain` defaults to `.`, so a conventional directory +is names only. Set it only for a hive addressed by something else. This +hive's own `services.hyperhive.domain` comes from its entry; it drives +`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified +labels (`iris@pr1ma.example.com`). + +`swarm.name` is display only — the dashboard chrome header and per-agent +system prompts — and federated hives at different domains can share one. +`hiveName` surfaces in the same places but is also the leftmost label of the +hive's domain. `swarm.name` names the group; `hiveName` names this hive. + +The env-var chain and `qualify()` / `qualified_label()` semantics: +[`conventions.md`](../process/conventions.md) § Hive identity. + +Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive +chains to it, so there is no per-hive CA field and no per-hive cert pinning. +Trusting a hive whose root this swarm doesn't own — another swarm's — has no +mechanism. + +
Upgrading an existing hive + +- Set `swarm.domain` and `hiveName` once; eval fails naming each until you do. +- Add the hive's own directory entry, + `services.hyperhive.swarm.hives. = { };` — one line, no value. + Eval fails naming it if you forget. +- Setting `services.hyperhive.domain` directly still works and still wins, + with a **deprecation warning**: the option is local to one host while + every host holds the directory, so a value written only there leaves every + peer pointing somewhere else. Move it into the hive's directory entry, or + drop it if it's the conventional `.`. +- Setting `swarm.hives..certFingerprint` fails eval; the field no + longer exists. Trust comes from the swarm root instead. + +
+ +### What the directory feeds + +1. **Swarm-wide hive roster** — swarm-controller reads this same + directory and serves it at `GET /api/hives`; the swarm UI's front page + renders it ([`ui.md`](ui.md)). + +2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel + federates with the peer's matrix server (discovered via the peer's + `.well-known/matrix/server` delegation, which the gateway serves). + Federation validates the peer's TLS certificate against the matrix + **container's** trust bundle, independent of this directory. + + ⚠️ On a hive with self-signed gateway certificates, the container + also trusts this hive's `trust-bundle.pem`, bound in at runtime and + added to the public CAs; it ends at the swarm root, so a peer whose + certificate chains to the same root validates. A peer outside this + swarm's root needs a CA-issued certificate (ACME). Federation firewall + and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md). + +3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each + entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress` + to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional). + +### WireGuard inter-hive mesh (optional) + +The peer config above uses public HTTPS for all inter-hive traffic. +For private deployments — or to reduce latency and TLS overhead on +intra-swarm traffic — hyperhive can configure a host-to-host +WireGuard mesh. + +#### Generating keys + +On each hive host: + +```bash +wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key +wg pubkey < /etc/wireguard/hive.key # → share this with peer operators +``` + +#### Config example (two hives) + +```nix +# hive A (pr1ma.example.com, mesh IP 10.100.0.1) +services.hyperhive = { + deploy.wireguard = { + enable = true; + privateKeyFile = "/etc/wireguard/hive.key"; + address = "10.100.0.1/24"; + listenPort = 51820; # optional, default 51820 + }; + + # The same `hives` attrset both hosts hold — mesh fields included, + # since "where this hive can be dialled" is a fact about that hive. + swarm.hives = { + pr1ma = { + domain = "pr1ma.example.com"; + wireguardPublicKey = "base64keyA="; + wireguardEndpoint = "198.51.100.1:51820"; + wireguardAddress = "10.100.0.1/32"; + }; + edge = { + domain = "edge.corp"; + wireguardPublicKey = "base64keyB="; + wireguardEndpoint = "203.0.113.42:51820"; + wireguardAddress = "10.100.0.2/32"; + }; + }; +}; + +# hive B (edge.corp, mesh IP 10.100.0.2) +services.hyperhive = { + deploy.wireguard = { + enable = true; + privateKeyFile = "/etc/wireguard/hive.key"; + address = "10.100.0.2/24"; + }; + + swarm.hives = { /* … identical to hive A's … */ }; +}; +``` + +#### What the mesh does + +- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the + host (not inside agent containers; containers reach peers via the + host's routing table). +- It opens UDP port 51820 (or `listenPort`) on the host firewall. +- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly + from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s + `allowedIPs`, so intra-swarm traffic can route over the mesh address + rather than the public domain. +- It sets `persistentKeepalive = 25` by default; override or null to + disable (not needed when both sides have public IPs and no NAT). + +#### NAT / one-sided endpoints + +If one host is behind NAT and can't accept incoming connections, only +that host needs a null `wireguardEndpoint` on the peer config — the +other side initiates. With keepalive on, the NAT hole stays open. + +If both hosts are behind NAT, you need a STUN relay or a third host +(exit node); hyperhive sets up neither. + ## Cross-references -- `docs/networking/snapshot-store.md` — the swarm's `btrfs receive` endpoint, and - the `swarm.snapshotStore` option that points a hive at it -- `docs/process/conventions.md` § Hive identity — env vars, qualified labels -- `docs/integrations/matrix.md` — matrix federation, TLS cert autogeneration, - firewall posture -- `docs/swarm/ui.md` — the swarm-wide hive roster, now the operator - surface for "what hives exist" (superseded the per-hive dashboard's - old "peer hives" display) -- `docs/networking/gateway.md` — nginx vhosts and the `.well-known/matrix/` - autodiscovery scheme +- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist" +- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the + swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option +- [`../process/conventions.md`](../process/conventions.md) § Hive identity — + env vars, qualified labels +- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix + federation, TLS cert autogeneration, firewall posture +- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and + the `.well-known/matrix/` autodiscovery scheme diff --git a/docs/swarm/services.md b/docs/swarm/services.md index c140e1fd..3c1fd084 100644 --- a/docs/swarm/services.md +++ b/docs/swarm/services.md @@ -1,7 +1,12 @@ # Swarm-wide services -Some things exist once per **swarm** rather than once per hive. Two -options say where the optional ones live, and everything else derives: +Some things exist once per **swarm** rather than once per hive: the forge, +the matrix homeserver, SSO, the secret store, the queue, and the metrics and +log stack. This page says which host runs them and what a hive that runs none +of them configures instead. The all-local quick start sets everything with one +line → [README](../../README.md#quick-start-an-all-local-swarm). + +Two options say where they live, and everything else derives: ```nix services.hyperhive.deploy.singleHostSwarm = true; # everything on this box @@ -14,10 +19,15 @@ here" means: every once-per-swarm service takes its `enable` from it.** That's t sections below don't repeat it, so a service that stops deriving is a visible difference rather than one more paragraph saying the same thing. -`singleHostSwarm` is the all-on-one-box switch above it: it defaults -both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host -generates the swarm CA here). You can still set each derived toggle on its own, -which wins, so "all local except X" needs no further option. +`singleHostSwarm` is the all-on-one-box mode above it. It defaults +`deploy.allSwarmServices`, the swarm CA (`swarm.ca.autoConfigure`, generated +on this host), the swarm controller (`deploy.swarm-controller.enable`), the +host's `/etc/hosts` entries for the names it serves +(`gateway.localHostsEntry`), the queue's auth-callout keys +(`deploy.nats.autoGenerateCallout`) and where the secret store's bootstrap +token goes (`deploy.bao.bootstrapTokenFile`). You can still set each derived +toggle on its own, which wins, so "all local except X" needs no further +option. **Both default to off**, and that's deliberate: a host can't tell whether it's meant to be the swarm's service host, so this is an @@ -38,23 +48,29 @@ answers its name from its own resolver, so on a swarm spread over more than one host, the operator's DNS has to resolve those names to that host. +
Moving an existing hive's forge to the swarm's + A hive that stops running the forge keeps the old container's state at `/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's forge: push anything worth keeping there by hand. Its `/var/lib/hyperhive/forge-core-token` came from that old forge and fails against the swarm's one. +
+ ## Deployment shapes Those two options are what makes the difference between deployments, so the shapes worth naming are the ones they produce: - **All-local.** Everything on one machine: - `singleHostSwarm = true`. Setup is automatic apart from - choosing a domain and creating the first user. + `singleHostSwarm = true`, plus `deploy.hive-controller.enable = true` for + a hive to run agents on. After the first switch, the steps in + [`setup.md`](../getting-started/setup.md) remain. - **Services on the swarm controller host.** - `deploy.allSwarmServices = true` there; the required services - deploy together on that host, with hives elsewhere. + Set `deploy.allSwarmServices` and `deploy.swarm-controller.enable` there, + with hives elsewhere. The controller doesn't derive from + `allSwarmServices`. - **Fully spread out.** One container / VM / machine per service, somewhere. @@ -88,12 +104,10 @@ there is one IdP and one auth path. container. Set it explicitly when joining a swarm whose IdP is under another name. -swarm-controller writes the users database, not by hand: hive-c0re -creates and destroys agents continuously, so the subject set is dynamic. -This module only guarantees the file exists and parses, so authelia -starts with nobody in it rather than failing to start — a provider with -no subjects yet is the correct state before anything has provisioned -them. Authelia generates session and storage keys in the container on +Agent subjects come from swarm-controller's agent-creation job, written +into the users database by `swarm-authelia-bridge`; human ones come from +`swarmctl user add` → [setup.md § 2](../getting-started/setup.md#2--your-sso-account). +On first boot this module seeds an empty users database. Authelia generates session and storage keys in the container on first boot and never rotates them automatically; replacing one invalidates data already written (sessions, the encrypted store), so that's an operator action. @@ -196,9 +210,9 @@ gateway either way. **Both store exporters are unconditional**, and `deploy.victoriametrics.enable` doesn't gate them: that option says this host _runs_ the store, while the swarm -has one either way, reached by its swarm name through the gateway. Gating on it -once left a collector on any other host with no exporter at all — receiving from -every hive and dropping it, silently, because an absent exporter isn't an error. +has one either way, reached by its swarm name through the gateway. A collector +with no exporter would receive from every hive and drop it silently, because an +absent exporter isn't an error. Agent-side configuration, and what a hive's own collector does, are in [`../scheduler/observability.md`](../scheduler/observability.md). diff --git a/docs/swarm/ui.md b/docs/swarm/ui.md index f5c826eb..6c8d5f58 100644 --- a/docs/swarm/ui.md +++ b/docs/swarm/ui.md @@ -1,10 +1,23 @@ # Swarm UI -The swarm's own web surface, served by the gateway on the **swarm apex** -(`services.hyperhive.swarm.domain`) and readable only by operators. +The swarm's own web surface and the operator's day-to-day view: served by +the gateway on the **swarm apex** (`services.hyperhive.swarm.domain`), +readable only by operators. The per-hive dashboard, on each hive's own +domain, covers host-level detail for one hive. -Distinct from the per-hive dashboard, which lives on the hive domain and -answers for one host. This one is the view _across_ hives. +## What it shows + +| route | what | +| ------------------------- | ------------------------------------------------------------------------------------------- | +| `/` | the hive directory, each hive with its last reported status | +| `/agents` | every agent: status, config PR, wanted state; create agents, link forge and matrix accounts | +| `/agents//terminal` | one agent's live terminal | +| `/jobs` | the controller's job graph — where agent creation and credential mints show progress | +| `/issues` | a cross-repo issue report | + +Everything it shows comes from [`swarm-controller`](../../swarm-controller/README.md). +An agent created here or with `swarmctl agent create` starts `paused`; set it +`up` from its card. ## Enabling @@ -35,28 +48,16 @@ requiring `group:admins`. An account without that group authenticates fine and still gets bounced. ```sh -swarmctl user add --group admins +swarmctl user add --email @example.com --group admins +swarmctl user update --add-group admins # an account that already exists ``` - +`--email` isn't needed for the UI, but the forge won't create your account +without one → [setup.md § 2](../getting-started/setup.md#2--your-sso-account). -`admins` deliberately, not a new word: [`../getting-started/setup.md`](../getting-started/setup.md) has -told every operator to create exactly that group since the bootstrap step -existed, so an account made by following the guide already passes. This -is the first rule that _consumes_ a group name — inventing a second one -would have meant those accounts silently failing a check they were -supposed to pass. - - - -An account created without any group needs re-adding with the flag — -`swarmctl` reads the existing entry out of `users.yml,` so the group is -what changes. - -Why a group and not a list of usernames: agents are getting authelia -accounts of their own (matrix SSO), and _authenticated_ would then -include every agent in the hive. The group is the only thing standing -between "an operator's page" and "anyone with a session." +Why a group and not "any session": agents are authelia subjects too, so +_authenticated_ includes every agent in the swarm. The group is the only +thing standing between "an operator's page" and "anyone with a session." ## What it costs to be reachable @@ -66,7 +67,22 @@ not a hole: **reachability isn't the access control here.** An agent that resolves the name and connects still has no operator session, and the subrequest denies it. -## Two wiring sites +## Quick links + +The swarm UI's header carries a single 🔗 button, visible on every route, +opening a popover of links to other swarm-wide services. Backed by +`GET /api/links` (swarm-controller), which serves +`services.hyperhive.swarm.controller.links` (a `listOf { label, icon, url }`, +same shape as the per-agent `services.hyperhive.agent.dashboardLinks`). + +Each service's own module contributes its entry when it's enabled on the +controller's host — `swarm-authelia.nix`, `hive-matrix.nix`, +`hive-forge/default.nix`, `swarm-grafana.nix`, `swarm-victorialogs.nix`, and +`swarm-ui.nix` for this UI's own API docs. Adding a link for a new service is +a nix-only change to that service's module, or an operator adding an entry +directly. An empty list hides the button. + +
Adding a swarm service name: the two wiring sites Adding a swarm service name means touching two things. Missing the second ships as a different flavour of "works from the host, broken from @@ -99,23 +115,7 @@ and the apex is a **sibling** of `forge.` / `chat.` / implicitly. Left out, the vhost falls back to the hive leaf and the swarm's front page opens with a name mismatch. -## Quick links - -The swarm UI's header carries a single 🔗 button, visible on every route, -opening a popover of links to other swarm-wide services — authelia, -matrix, forge, this UI's own swagger docs. Backed by `GET /api/links` -(swarm-controller), which serves `services.hyperhive.swarm.controller.links` -(a `listOf { label, icon, url }`, same shape as the per-agent -`services.hyperhive.agent.dashboardLinks`). - -Rather than one central hardcoded list, each service's own module -contributes its own entry when it's actually enabled on the controller's -host — `swarm-authelia.nix`, `hive-matrix.nix` and `hive-forge/default.nix` -all do, the same list-merge idiom `gateway.localNames` uses above. Adding a -link for a new service is a nix-only change to that service's own module -(or an operator adding an entry directly); no swarm-controller or swarm-ui -change needed. Empty list hides the button rather than showing an empty -popover. +
## Cross-references diff --git a/swarm-controller/README.md b/swarm-controller/README.md index 05a3d19f..6e015250 100644 --- a/swarm-controller/README.md +++ b/swarm-controller/README.md @@ -9,15 +9,34 @@ deliberately **not** derived from `services.hyperhive.deploy.hive-controller.enable`: turning it on is a statement about swarm topology, not about whether this host runs a hive. -## What it does today +## What it does -Serves one `/health` endpoint and holds no state. +The swarm's control plane. The swarm UI and `swarmctl` are its clients. -That is the whole intent of the first slice. The point is to make the _unit_ -real — service user, runtime and state directories, socket, nginx -reachability — so the swarm-level surfaces that follow have somewhere to land. -Inventing those surfaces before they are agreed would bake in a shape nobody -chose. See the `hyperhive.swarm` consolidation epic. +- **Hive directory** — serves `swarm.hives` (`GET /api/hives`) and what each + hive last published about itself (`GET /api/hives/status`). +- **Agent roster and wanted state** — every agent the swarm knows + (`GET /api/agents`, `/api/agents/status`), and the state it declares for + each one on its hive (`up`/`offline`/`paused`/`destroyed`, + `PUT /api/hives/{hive}/agents/{agent}/state`). +- **Job graph** — a `hive-jobq` scheduler, served at `GET /api/jobq/graph`. + Every provisioning step below runs as a node in it. +- **Agent creation** — `POST /api/agents` queues the SSO identity (through + `swarm-authelia-bridge`), forge user, config repo, store identity, forge + token and matrix account, declares the agent `paused`, then sends its hive + a deploy message. +- **Agent credentials** — at start and every five minutes it re-checks every + agent's forge token and matrix account, and renews store certificates and + queue secrets as they age. +- **Swarm-wide forge objects and webhooks** → + [`docs/swarm/README.md`](../docs/swarm/README.md#swarm-wide-forge-objects). +- **Relays** — each agent's terminal and turn-state header as SSE, agent + icons, the cross-repo issue report, and the UI's quick links. + +It reads its configuration once at startup, from the environment the nix module +sets. The one file it persists is `webhook-secret` in its state directory. +The job graph lives in memory; hive status and wanted state live in the swarm +queue, so both survive a restart. ## Why a unix socket, not a port diff --git a/swarmctl/README.md b/swarmctl/README.md index e526fe7b..ca97da99 100644 --- a/swarmctl/README.md +++ b/swarmctl/README.md @@ -31,18 +31,9 @@ that. ## One file, two writers -`users.yml` — authelia's own users database — is read and written -directly. There is no second store. - -There used to be: a private `users.json` here, canonical, with `users.yml` -rendered from it, while `swarm-authelia-bridge` kept its own pair against -the _same_ physical file. Two canonical stores for one file is a seam, and -it bit — a writer whose own JSON was missing could not tell "nothing here -yet" from "someone else's users", and refused to write. - -The argument for the split was that it let this crate work without a YAML -parser. It didn't: the JSON was read back on every run, so the round-trip -was already being paid — the two files differed only in _format_. +`swarmctl` reads and writes `users.yml` — authelia's own users database — +directly, with no second store. `swarm-authelia-bridge` writes agent +subjects into the same file. ⚠️ The file is round-tripped, so **comments and hand-formatting do not survive a write**. Values do, and so do keys this binary does not model. @@ -62,18 +53,30 @@ an error. | `SWARMCTL_AUTHELIA_USERS_FILE` | host-side path of the users database | | `SWARMCTL_AUTHELIA_MACHINE` | container name, for `systemctl -M` | | `SWARMCTL_AUTHELIA_UNIT` | authelia's unit inside that container | -| `SWARM_CONTROLLER_SOCKET` | the controller's unix socket, for `agent create` | +| `SWARM_CONTROLLER_SOCKET` | the controller's unix socket, for `agent` and `forge` verbs | + +Every verb and flag: [`docs/tools/swarmctl-cli.md`](../docs/tools/swarmctl-cli.md). +The sections below cover why each verb behaves as it does. ## `user add` ```console -# swarmctl user add mara --display-name "Mara" --group admins +# swarmctl user add mara --display-name "Mara" --email mara@example.com --group admins ``` +Keep both flags: + +- **`--group admins`** — the swarm UI and other operator surfaces gate on it. +- **`--email`** — the forge won't create an account without one. + +`user add` refuses a username that already exists; fix an existing account +with `swarmctl user update mara --add-group admins --email …`. + The password is **generated by authelia** (`crypto hash generate argon2 --random`) and printed once. It is never passed on a command line: `/proc//cmdline` is world-readable, so a password in argv is readable -by any local process for the lifetime of the call. +by any local process for the lifetime of the call. `user reset-password` +generates a new one the same way. ## `agent create` @@ -89,13 +92,17 @@ module sets from the daemon's own `socketPath`). It prints the queued job's node id **and stops there**. It deliberately does not wait. The endpoint queues a DAG — SSO identity, -forge user, forge repo, repo membership, config-repo seed, then a deploy +forge user, forge repo, repo membership, config-repo seed, store identity, +forge token, matrix account, a `paused` wanted state, then a deploy message — and the last of those _publishes_: the hive's `hive-c0re` picks it up and converges on its own clock, out of the controller's sight. So even a fully settled graph would not mean the agent is up, and there is nothing this CLI could wait for that would let it say so honestly. Watch the swarm UI's job view for the rest. +The new agent starts `paused`: it doesn't drive turns until you set it `up` +in the swarm UI. + No approval gate, for the same reason nothing else here has one: running this binary already means being root on the controller's host. @@ -110,6 +117,30 @@ crate does not link it, and there is no wire-type crate between them. Two fields out, two in, both ends validating — a drift shows up as a 400 naming the field. +## `agent mint-identity` + +```console +# swarmctl agent mint-identity ruth +``` + +`POST /api/agents/{name}/identity` on the swarm-controller. Re-runs the +store-identity mint for one agent that already exists — the backfill for an +agent the swarm never created, such as the manager agent `hive-c0re` makes at +startup. It re-mints the agent's store certificate, which the agent picks up +the next time its container boots, and leaves an existing queue secret alone. +Queues and returns, like `agent create`. + +## `agent mint-forge-token` + +```console +# swarmctl agent mint-forge-token ruth +``` + +`POST /api/agents/{name}/forge-token` on the swarm-controller. Checks one +agent's forge token and mints it if it's missing or stale. The controller +already does this for every agent with a store identity at start and every +five minutes; this verb skips the wait. Queues and returns. + ## `forge make-admin` ```console