docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity and the directory follow as the substrate. Upgrade notes move into a <details> block, the per-agent queue publishing detail into another, and the one-paragraph pointer sections collapse into a link list. Fact fixes, checked against origin/main: - an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean "not in a swarm" - swarm.domain is required with a hive (hive-network.nix:156,188), hiveName with a hive, store or homeserver (hyperhive.nix:161-166) - the matrix container trusts the hive's trust-bundle.pem at runtime under self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85) - singleHostSwarm also defaults the controller, localHostsEntry, the nats callout keys and the bao bootstrap token path (local-defaults.nix:72-129) - swarm-controller serves far more than /health: roster, wanted state, job graph, agent creation and credential mints (main.rs:2874-2899) - swarmctl user add needs --email for the forge account and refuses an existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document agent mint-identity and mint-forge-token - agent creation also mints store identity, forge token and matrix account, and declares the agent paused (main.rs:1822-1920, 247-248) Refs #3902
This commit is contained in:
parent
f688cfcdf0
commit
270430a4b4
5 changed files with 408 additions and 402 deletions
|
|
@ -1,308 +1,55 @@
|
||||||
# Multi-hive swarms
|
# The swarm
|
||||||
|
|
||||||
A **swarm** is a collection of agents that share an identity and
|
The **swarm** is where things live: agent identities and accounts, secrets,
|
||||||
coordinate across one or more hives. A single hyperhive instance
|
the job graph that creates and places agents, telemetry, and the UI you
|
||||||
running on one host is already a swarm (one hive). This doc covers
|
drive it all from. **Hives are the substrate** — NixOS hosts that run agent
|
||||||
the additional config needed when the swarm spans multiple hosts.
|
containers on the swarm's behalf. Every hive belongs to a swarm; a single
|
||||||
|
host is a swarm of one.
|
||||||
|
|
||||||
For the full option reference rather than prose: `services.hyperhive.swarm.*`
|
This page is for the operator. It covers the control plane, how a hive
|
||||||
(swarm-wide facts, identical on every host) and `services.hyperhive.deploy.*`
|
joins the directory, and how hives and agents report upward. The steps for a
|
||||||
(this host's own deployment decisions — does _this_ machine run grafana,
|
fresh swarm are in [`setup.md`](../getting-started/setup.md); the all-local
|
||||||
the swarm controller, authelia, …) are separate generated pages, `nix
|
config is the [README quick start](../../README.md#quick-start-an-all-local-swarm).
|
||||||
build .#docs-swarm` / `.#docs-deploy` or the website's `/options/swarm.html`
|
|
||||||
/ `/options/deploy.html`.
|
|
||||||
|
|
||||||
## Terminology
|
Option reference: `services.hyperhive.swarm.*` (swarm-wide facts, identical
|
||||||
|
on every host) and `services.hyperhive.deploy.*` (whether _this_ host runs
|
||||||
|
grafana, the controller, authelia, …) →
|
||||||
|
[options reference](https://hyperhive.darkest.space/options/), or
|
||||||
|
`nix build .#docs-swarm` / `.#docs-deploy`.
|
||||||
|
|
||||||
- **hive** — a single hyperhive installation on one host. Has its
|
## Where each piece lives
|
||||||
own `services.hyperhive.domain` DNS name and its own set of agent
|
|
||||||
containers.
|
|
||||||
- **swarm** — one or more hives whose operators have declared them
|
|
||||||
as peers. You can qualify an agent as `agent@hive-domain`.
|
|
||||||
- **peer hive** — any hive in `services.hyperhive.swarm.hives` other
|
|
||||||
than this one. Peers are _derived_, not declared: the directory lists
|
|
||||||
every hive including yourself, and `hiveName` says which one you are.
|
|
||||||
|
|
||||||
## Hive identity config
|
- **control plane** — `swarm-controller`: hive directory, agent roster, job
|
||||||
|
graph, agent creation. → [below](#swarm-controller)
|
||||||
```nix
|
- **swarm UI** — the operator's day-to-day surface, on the swarm apex,
|
||||||
services.hyperhive = {
|
`admins` only. → [`ui.md`](ui.md)
|
||||||
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
- **shared services** — one forge, homeserver, SSO, queue and metrics/logs
|
||||||
hiveName = "pr1ma"; # required — this hive's label in it
|
stack, each on whichever host you put it. → [`services.md`](services.md)
|
||||||
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
- **SSO** — the secrets authelia generates and how each reaches its reader.
|
||||||
|
→ [`sso.md`](sso.md)
|
||||||
# required — the directory, identical on every host in the swarm.
|
- **secrets** — every credential the swarm holds, who mints it and where it
|
||||||
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
lives → [`secrets.md`](secrets.md) · the per-secret minter/reader/renewal
|
||||||
swarm.hives = {
|
contract → [`credentials.md`](credentials.md) · how the store comes up,
|
||||||
pr1ma = { };
|
who writes its grants and how it unseals → [`bao.md`](bao.md)
|
||||||
edge = { };
|
- **swarm CA** — the root every hive's internal TLS chains to, and what to
|
||||||
};
|
hand a peer (`trust-bundle.pem`, never `ca.pem`). → [`ca.md`](ca.md)
|
||||||
};
|
- **snapshot store** — the swarm's one `btrfs receive` endpoint,
|
||||||
```
|
`swarm.snapshotStore.{address,port}`. →
|
||||||
|
[`snapshot-store.md`](../networking/snapshot-store.md)
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
|
|
||||||
`swarm.domain` and `hiveName` are **required** whenever hyperhive is
|
|
||||||
enabled; eval fails with a hint naming each. Neither defaults,
|
|
||||||
because a guessed value here is a wrong hostname that evaluates cleanly
|
|
||||||
and deploys — an eval failure asking the operator to write the address
|
|
||||||
down is the cheaper outcome. **Upgrading past this release means setting
|
|
||||||
both once.**
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
You must still set `domain` too, but you no longer _write_ it: it's read from
|
|
||||||
this hive's own entry in the directory, whose `domain` defaults to
|
|
||||||
`<name>.<swarm.domain>`. A conventional swarm states no addresses at
|
|
||||||
all, and a hive addressed by something else states it in the one place
|
|
||||||
the other hives read — `swarm.hives.edge.domain = "edge.elsewhere.example";`.
|
|
||||||
|
|
||||||
Setting `services.hyperhive.domain` directly still works and still wins,
|
|
||||||
with a **deprecation warning**. The reason it's deprecated isn't tidiness:
|
|
||||||
that option is local to one host, and the operator copies the directory
|
|
||||||
to every host, so a value written only there leaves every peer pointing
|
|
||||||
somewhere else with nothing detecting the disagreement.
|
|
||||||
|
|
||||||
⚠️ **Upgrading:** a hive that has been running on `swarm.domain` +
|
|
||||||
`hiveName` alone now needs its own directory entry —
|
|
||||||
`services.hyperhive.swarm.hives.<hiveName> = { };`, one line, no value.
|
|
||||||
Eval fails naming it if you forget.
|
|
||||||
|
|
||||||
`domain` drives `HYPERHIVE_HIVE_DOMAIN` in every container so agents can
|
|
||||||
form qualified labels (`iris@pr1ma.example.com`).
|
|
||||||
|
|
||||||
`swarm.name` is purely display — it surfaces in the dashboard chrome
|
|
||||||
header and per-agent system prompts, and federated hives at different
|
|
||||||
domains can share one. `hiveName` surfaces in the same places but is
|
|
||||||
_not_ only display: it's the leftmost label of the hive's domain. That
|
|
||||||
`swarm.name` sits under `swarm` and `hiveName` doesn't is the whole
|
|
||||||
distinction — one names this hive, the other names the group it belongs
|
|
||||||
to.
|
|
||||||
|
|
||||||
See `docs/process/conventions.md` § Hive identity for the env-var chain
|
|
||||||
and `qualify()` / `qualified_label()` semantics.
|
|
||||||
|
|
||||||
## Swarm CA
|
|
||||||
|
|
||||||
A hive's internal TLS chains to a **swarm root CA**, so a peer that
|
|
||||||
trusts the root validates every hive in the swarm rather than pinning
|
|
||||||
to each one by hand. Provisioning modes, what to hand a peer
|
|
||||||
(`trust-bundle.pem`, never `ca.pem`), the name constraints on a hive
|
|
||||||
CA, and how an existing hive adopts the hierarchy: [`ca.md`](ca.md).
|
|
||||||
|
|
||||||
## Running the swarm's shared services
|
|
||||||
|
|
||||||
One authelia, one matrix, one forge per swarm — which host runs them,
|
|
||||||
and what a hive that runs none of them configures instead:
|
|
||||||
[`services.md`](services.md).
|
|
||||||
|
|
||||||
## Single sign-on
|
|
||||||
|
|
||||||
Which secrets the SSO provider generates, which one has a reader in
|
|
||||||
another container, and the three ways that one gets delivered:
|
|
||||||
[`sso.md`](sso.md).
|
|
||||||
|
|
||||||
## Secrets
|
|
||||||
|
|
||||||
Every credential the swarm holds, who mints it, where it must live, and
|
|
||||||
which of the three topologies makes it the operator's job to place:
|
|
||||||
[`secrets.md`](secrets.md).
|
|
||||||
|
|
||||||
Where that shape is **going** — the per-secret minter/reader/renewal
|
|
||||||
contract, the target of one mTLS identity per host and everything else
|
|
||||||
through the store, and the test a change has to pass to count as movement
|
|
||||||
toward it: [`credentials.md`](credentials.md). It supersedes `secrets.md`
|
|
||||||
when the migration completes.
|
|
||||||
|
|
||||||
## Swarm UI
|
|
||||||
|
|
||||||
The operator-only web surface on the swarm apex, why reaching it needs
|
|
||||||
the `admins` group rather than just a session, and the four sites you
|
|
||||||
wire a swarm service name into: [`ui.md`](ui.md).
|
|
||||||
|
|
||||||
## The swarm's hive directory
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.swarm.hives = {
|
|
||||||
pr1ma = { }; # this host, per hiveName
|
|
||||||
lab = { }; # a second hive in the swarm
|
|
||||||
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
One attrset describing **every** hive in the swarm, **including this
|
|
||||||
one**, keyed by that hive's `hiveName`. It's meant to be _identical on
|
|
||||||
every host_ — write it once, share it, and each host reads it correctly
|
|
||||||
because `services.hyperhive.hiveName` says which entry is itself.
|
|
||||||
|
|
||||||
Empty (the default) means this host isn't in a swarm. Once non-empty it
|
|
||||||
**must** contain an entry for `hiveName`; eval fails naming the missing
|
|
||||||
hive. That assertion is load-bearing rather than pedantic — "my peers"
|
|
||||||
comes from _everything that isn't me_, so a directory that doesn't
|
|
||||||
contain you derives every hive as a peer and you peer with yourself.
|
|
||||||
|
|
||||||
`domain` defaults to `<name>.<swarm.domain>`, the convention every hive
|
|
||||||
follows, so a conventional directory is names only. The default is a
|
|
||||||
derivation from two values the operator already had to state — the swarm's
|
|
||||||
domain and the entry's own name — rather than a guess, which is what makes
|
|
||||||
it safe here when a guessed hostname wouldn't be. Set it only for a hive
|
|
||||||
addressed by something else.
|
|
||||||
|
|
||||||
> **No per-hive CA field exists, and no per-hive cert pinning.** Trust
|
|
||||||
> inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every
|
|
||||||
> hive chains to it, so one anchor replaces per-hive pinning entirely.
|
|
||||||
> What that genuinely drops is trusting a hive whose root this swarm
|
|
||||||
> does _not_ own — another swarm's, or one keeping its own CA. That's
|
|
||||||
> a cross-swarm problem and wants a mechanism designed for it. (An
|
|
||||||
> earlier `certFingerprint` field existed for exactly that gap, pinning
|
|
||||||
> a peer's TLS leaf for hive-c0re's own peer HTTPS checks — removed
|
|
||||||
> along with the dashboard feature it existed to serve, since nothing
|
|
||||||
> else ever consumed it.)
|
|
||||||
|
|
||||||
## What the config does at runtime
|
|
||||||
|
|
||||||
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
|
||||||
directory and serves it at `GET /api/hives`; `swarm-ui`'s overview
|
|
||||||
page renders it (`docs/swarm/ui.md`). This is the operator-facing
|
|
||||||
"what hives exist" surface — a per-hive dashboard "peer hives"
|
|
||||||
display existed here once; it no longer exists, in favour of this.
|
|
||||||
|
|
||||||
2. **Matrix federation** — when `matrix.enable` is on, tuwunel
|
|
||||||
federates with the peer's matrix server (discovered via the peer's
|
|
||||||
`.well-known/matrix/server` delegation, which the gateway serves).
|
|
||||||
Federation validates the peer's TLS certificate against the matrix
|
|
||||||
**container's** trust bundle, independent of this directory.
|
|
||||||
|
|
||||||
⚠️ **That container currently trusts no swarm-internal CA**, so a
|
|
||||||
self-signed gateway certificate doesn't federate. You can't list the
|
|
||||||
swarm root there: `security.pki.certificateFiles` is
|
|
||||||
read when the system is _built_, and the root is a runtime file (its
|
|
||||||
key must never enter the store), so there is no build-time name for
|
|
||||||
it. Bridging that needs a runtime mechanism; a separate issue tracks
|
|
||||||
it. Until then, federation needs CA-issued certs (ACME). See
|
|
||||||
`docs/integrations/matrix.md` for federation firewall + TLS requirements.
|
|
||||||
|
|
||||||
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
|
||||||
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
|
||||||
to configure `wg-hive`. See "WireGuard inter-hive mesh" below.
|
|
||||||
|
|
||||||
## One directory, not a bilateral declaration
|
|
||||||
|
|
||||||
Both hives hold the **same** `hives` attrset; neither declares the
|
|
||||||
other. What differs between the two hosts is only `hiveName`:
|
|
||||||
|
|
||||||
```
|
|
||||||
# hive A # hive B
|
|
||||||
hiveName = "pr1ma"; hiveName = "edge";
|
|
||||||
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
|
||||||
```
|
|
||||||
|
|
||||||
That's the point of the shape, and it removes a class of bug rather
|
|
||||||
than saving typing: a per-host peer list let two hosts hold _different_
|
|
||||||
facts about the same third hive — a stale endpoint, a rotated
|
|
||||||
fingerprint — with nothing to detect the disagreement. One entry per
|
|
||||||
hive makes it unrepresentable.
|
|
||||||
|
|
||||||
## WireGuard inter-hive mesh (optional)
|
|
||||||
|
|
||||||
The peer config above uses public HTTPS for all inter-hive traffic.
|
|
||||||
For private deployments — or to reduce latency and TLS overhead on
|
|
||||||
intra-swarm traffic — hive-c0re can configure a host-to-host
|
|
||||||
WireGuard mesh.
|
|
||||||
|
|
||||||
### Generating keys
|
|
||||||
|
|
||||||
On each hive host:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
|
||||||
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
|
||||||
```
|
|
||||||
|
|
||||||
### Config example (two hives)
|
|
||||||
|
|
||||||
```nix
|
|
||||||
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
|
||||||
services.hyperhive = {
|
|
||||||
deploy.wireguard = {
|
|
||||||
enable = true;
|
|
||||||
privateKeyFile = "/etc/wireguard/hive.key";
|
|
||||||
address = "10.100.0.1/24";
|
|
||||||
listenPort = 51820; # optional, default 51820
|
|
||||||
};
|
|
||||||
|
|
||||||
# The same `hives` attrset both hosts hold — mesh fields included,
|
|
||||||
# since "where this hive can be dialled" is a fact about that hive.
|
|
||||||
swarm.hives = {
|
|
||||||
pr1ma = {
|
|
||||||
domain = "pr1ma.example.com";
|
|
||||||
wireguardPublicKey = "base64keyA=";
|
|
||||||
wireguardEndpoint = "198.51.100.1:51820";
|
|
||||||
wireguardAddress = "10.100.0.1/32";
|
|
||||||
};
|
|
||||||
edge = {
|
|
||||||
domain = "edge.corp";
|
|
||||||
wireguardPublicKey = "base64keyB=";
|
|
||||||
wireguardEndpoint = "203.0.113.42:51820";
|
|
||||||
wireguardAddress = "10.100.0.2/32";
|
|
||||||
};
|
|
||||||
};
|
|
||||||
};
|
|
||||||
|
|
||||||
# hive B (edge.corp, mesh IP 10.100.0.2)
|
|
||||||
services.hyperhive = {
|
|
||||||
deploy.wireguard = {
|
|
||||||
enable = true;
|
|
||||||
privateKeyFile = "/etc/wireguard/hive.key";
|
|
||||||
address = "10.100.0.2/24";
|
|
||||||
};
|
|
||||||
|
|
||||||
swarm.hives = { /* … identical to hive A's … */ };
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
### What the mesh does
|
|
||||||
|
|
||||||
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
|
||||||
host (not inside agent containers; containers reach peers via the
|
|
||||||
host's routing table).
|
|
||||||
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
|
||||||
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
|
||||||
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
|
||||||
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
|
||||||
rather than the public domain.
|
|
||||||
- It sets `persistentKeepalive = 25` by default; override or null to
|
|
||||||
disable (not needed when both sides have public IPs and no NAT).
|
|
||||||
|
|
||||||
### NAT / one-sided endpoints
|
|
||||||
|
|
||||||
If one host is behind NAT and can't accept incoming connections, only
|
|
||||||
that host needs a null `wireguardEndpoint` on the peer config — the
|
|
||||||
other side initiates. With keepalive on, the NAT hole stays open.
|
|
||||||
|
|
||||||
If both hosts are behind NAT, you need a STUN relay or a third host
|
|
||||||
(exit node). Out of scope for v0.
|
|
||||||
|
|
||||||
## Snapshot store
|
|
||||||
|
|
||||||
One further option lives in this namespace but its docs live with the
|
|
||||||
service it points at: `services.hyperhive.swarm.snapshotStore.{address,
|
|
||||||
port}` tells this hive where the swarm's `btrfs receive` endpoint is, so
|
|
||||||
`hivectl agent <name> subvol snapshot push` has somewhere to stream to.
|
|
||||||
|
|
||||||
It's genuinely swarm-scoped rather than per-peer — a swarm has exactly
|
|
||||||
one store, because the receiver keys destinations by _agent_ so a
|
|
||||||
migrating agent keeps one unbroken incremental chain. See
|
|
||||||
[snapshot-store.md](../networking/snapshot-store.md).
|
|
||||||
|
|
||||||
## Swarm controller
|
## Swarm controller
|
||||||
|
|
||||||
`services.hyperhive.deploy.swarm-controller.enable` runs the `swarm-controller`
|
`swarm-controller` is the swarm's control plane: it holds the hive directory,
|
||||||
daemon on this host. **Off by default and deliberately not derived from
|
the agent roster and their wanted state, and the job graph. Creating an agent —
|
||||||
`services.hyperhive.deploy.hive-controller.enable`**: a swarm has one
|
from the swarm UI or `swarmctl agent create --hive <h>` — queues its SSO
|
||||||
|
identity, forge user, config repo, store identity and matrix account, then
|
||||||
|
sends the hive a deploy message. A new agent starts `paused`.
|
||||||
|
|
||||||
|
`services.hyperhive.deploy.swarm-controller.enable` runs it on this host;
|
||||||
|
`singleHostSwarm` turns it on. **Otherwise off by default and deliberately
|
||||||
|
not derived from `deploy.hive-controller.enable`**: a swarm has one
|
||||||
controller, so enabling it states a fact about swarm topology, not about
|
controller, so enabling it states a fact about swarm topology, not about
|
||||||
whether this host runs a hive. Every hive runs `hive-c0re` (the agents on that host); one
|
whether this host runs a hive.
|
||||||
hive additionally runs this (what's true across hives).
|
|
||||||
|
|
||||||
What it serves, why it's a unix socket rather than a port, and the
|
What it serves, why it's a unix socket rather than a port, and the
|
||||||
socket-directory constraint that governs where `socketPath` may point:
|
socket-directory constraint that governs where `socketPath` may point:
|
||||||
|
|
@ -409,6 +156,8 @@ how it gets there. A hive lacking the queue's address for its
|
||||||
agents sets none of the four and each agent logs that it has none; a half-set
|
agents sets none of the four and each agent logs that it has none; a half-set
|
||||||
environment logs an error and the harness keeps serving.
|
environment logs an error and the harness keeps serving.
|
||||||
|
|
||||||
|
<details><summary>What an agent publishes over the queue</summary>
|
||||||
|
|
||||||
What an agent does with that connection is publish its terminal. Every row its
|
What an agent does with that connection is publish its terminal. Every row its
|
||||||
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
|
own web UI renders also goes to `$SWARM.term.<agent>`, one subject per agent, so
|
||||||
a swarm-level terminal can follow one agent without subscribing to the swarm's
|
a swarm-level terminal can follow one agent without subscribing to the swarm's
|
||||||
|
|
@ -468,6 +217,8 @@ key. Swarm-side, `GET /api/agents/<name>/icon` serves the stored bytes, and 404
|
||||||
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
|
means the agent has no icon. swarm-ui's agent cards load it as an `<img>` and show the
|
||||||
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
|
dimmed hyperhive mark for an agent without one, as the hive dashboard does.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
### Swarm-wide forge objects
|
### Swarm-wide forge objects
|
||||||
|
|
||||||
The controller also keeps the forge objects that are one per swarm, not
|
The controller also keeps the forge objects that are one per swarm, not
|
||||||
|
|
@ -484,7 +235,7 @@ one per hive. It ensures them at start and every five minutes after
|
||||||
- the `agent-configs` org avatar
|
- the `agent-configs` org avatar
|
||||||
(`deploy.swarm-controller.configOrgAvatarPng`).
|
(`deploy.swarm-controller.configOrgAvatarPng`).
|
||||||
|
|
||||||
hive-c0re no longer creates any of them. A pass that can't finish logs a
|
A pass that can't finish logs a
|
||||||
`warn` line per object plus `swarm forge objects: pass incomplete` in
|
`warn` line per object plus `swarm forge objects: pass incomplete` in
|
||||||
`journalctl -u swarm-controller`, and retries on the next tick. While the
|
`journalctl -u swarm-controller`, and retries on the next tick. While the
|
||||||
controller is down the objects stay as they are.
|
controller is down the objects stay as they are.
|
||||||
|
|
@ -503,14 +254,14 @@ re-derive. Approval happens once, at the swarm level: a hive receives a
|
||||||
decision, not an event to adjudicate.
|
decision, not an event to adjudicate.
|
||||||
|
|
||||||
**`internal/knowledge` is on that path.** The controller's is the only
|
**`internal/knowledge` is on that path.** The controller's is the only
|
||||||
hook on it: hives no longer register their own (see
|
hook on it; hives register none of their own
|
||||||
`docs/integrations/knowledge.md` for clearing a leftover). A webhook has exactly one target URL, so per-hive
|
([`knowledge.md`](../integrations/knowledge.md) covers clearing a leftover).
|
||||||
registration never added a recipient — it took delivery away from
|
A webhook has exactly one target URL, so a second registration would take
|
||||||
whichever hive registered before it.
|
delivery away from the first rather than add a recipient.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
|
|
||||||
**The `agent-configs` org isn't yet.** Each hive still registers its own
|
**The `agent-configs` org isn't.** Each hive registers its own
|
||||||
`pull_request` hook there, so that repo has two — the hive's and the
|
`pull_request` hook there, so that repo has two — the hive's and the
|
||||||
controller's — and **both are expected; don't delete either.** Removing
|
controller's — and **both are expected; don't delete either.** Removing
|
||||||
a hive's stops it acting on config PRs; removing the controller's just
|
a hive's stops it acting on config PRs; removing the controller's just
|
||||||
|
|
@ -529,15 +280,206 @@ To check it's working, push to `internal/knowledge` and look for
|
||||||
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
|
`webhook: verified delivery` in `journalctl -u swarm-controller`. A
|
||||||
refused delivery logs `webhook: refused delivery` with the reason.
|
refused delivery logs `webhook: refused delivery` with the reason.
|
||||||
|
|
||||||
|
## Hives: the substrate
|
||||||
|
|
||||||
|
- **hive** — one host running `hive-c0re` and its agent containers
|
||||||
|
(`deploy.hive-controller.enable`). Addressed as `<hiveName>.<swarm.domain>`.
|
||||||
|
- **swarm** — every hive in `services.hyperhive.swarm.hives`, plus the
|
||||||
|
shared services and controller. You can qualify an agent as
|
||||||
|
`agent@hive-domain`.
|
||||||
|
- **peer hive** — any hive in the directory other than this one. Peers are
|
||||||
|
_derived_, not declared: the directory lists every hive including
|
||||||
|
yourself, and `hiveName` says which one you are.
|
||||||
|
|
||||||
|
### Hive identity config
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive = {
|
||||||
|
swarm.domain = "example.com"; # required — the swarm's DNS domain
|
||||||
|
hiveName = "pr1ma"; # required — this hive's label in it
|
||||||
|
swarm.name = "constellat1on"; # shared swarm display name (optional)
|
||||||
|
|
||||||
|
# required — the directory, identical on every host in the swarm.
|
||||||
|
# Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
|
||||||
|
swarm.hives = {
|
||||||
|
pr1ma = { }; # this host, per hiveName
|
||||||
|
lab = { }; # a second hive
|
||||||
|
edge = { domain = "edge.elsewhere.example"; }; # addressed off-convention
|
||||||
|
};
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
<!-- vale write-good.Passive = NO -->
|
||||||
|
|
||||||
|
`swarm.domain` is **required** on a host that runs a hive, and `hiveName` on
|
||||||
|
a host that runs a hive, the secret store or the homeserver; eval fails with
|
||||||
|
a hint naming each. Neither defaults, because a guessed value here is a
|
||||||
|
wrong hostname that evaluates cleanly and deploys.
|
||||||
|
|
||||||
|
<!-- vale write-good.Passive = YES -->
|
||||||
|
|
||||||
|
**The directory is one attrset, identical on every host.** It describes
|
||||||
|
every hive in the swarm, **including this one**, keyed by `hiveName`; what
|
||||||
|
differs between hosts is only `hiveName`. It **must** contain an entry for
|
||||||
|
this host's `hiveName`, and an empty directory fails that check too — eval
|
||||||
|
names the missing hive. Peers are _every entry but this host's_, so a
|
||||||
|
directory without this host would make every hive a peer, itself included.
|
||||||
|
|
||||||
|
```
|
||||||
|
# hive A # hive B
|
||||||
|
hiveName = "pr1ma"; hiveName = "edge";
|
||||||
|
swarm.hives = { … }; swarm.hives = { … }; # byte-identical
|
||||||
|
```
|
||||||
|
|
||||||
|
One entry per hive means two hosts can't hold _different_ facts about the
|
||||||
|
same third hive, such as a stale endpoint.
|
||||||
|
|
||||||
|
`domain` defaults to `<name>.<swarm.domain>`, so a conventional directory
|
||||||
|
is names only. Set it only for a hive addressed by something else. This
|
||||||
|
hive's own `services.hyperhive.domain` comes from its entry; it drives
|
||||||
|
`HYPERHIVE_HIVE_DOMAIN` in every container so agents can form qualified
|
||||||
|
labels (`iris@pr1ma.example.com`).
|
||||||
|
|
||||||
|
`swarm.name` is display only — the dashboard chrome header and per-agent
|
||||||
|
system prompts — and federated hives at different domains can share one.
|
||||||
|
`hiveName` surfaces in the same places but is also the leftmost label of the
|
||||||
|
hive's domain. `swarm.name` names the group; `hiveName` names this hive.
|
||||||
|
|
||||||
|
The env-var chain and `qualify()` / `qualified_label()` semantics:
|
||||||
|
[`conventions.md`](../process/conventions.md) § Hive identity.
|
||||||
|
|
||||||
|
Trust inside a swarm comes from the swarm root ([`ca.md`](ca.md)): every hive
|
||||||
|
chains to it, so there is no per-hive CA field and no per-hive cert pinning.
|
||||||
|
Trusting a hive whose root this swarm doesn't own — another swarm's — has no
|
||||||
|
mechanism.
|
||||||
|
|
||||||
|
<details><summary>Upgrading an existing hive</summary>
|
||||||
|
|
||||||
|
- Set `swarm.domain` and `hiveName` once; eval fails naming each until you do.
|
||||||
|
- Add the hive's own directory entry,
|
||||||
|
`services.hyperhive.swarm.hives.<hiveName> = { };` — one line, no value.
|
||||||
|
Eval fails naming it if you forget.
|
||||||
|
- Setting `services.hyperhive.domain` directly still works and still wins,
|
||||||
|
with a **deprecation warning**: the option is local to one host while
|
||||||
|
every host holds the directory, so a value written only there leaves every
|
||||||
|
peer pointing somewhere else. Move it into the hive's directory entry, or
|
||||||
|
drop it if it's the conventional `<hiveName>.<swarm.domain>`.
|
||||||
|
- Setting `swarm.hives.<name>.certFingerprint` fails eval; the field no
|
||||||
|
longer exists. Trust comes from the swarm root instead.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
### What the directory feeds
|
||||||
|
|
||||||
|
1. **Swarm-wide hive roster** — swarm-controller reads this same
|
||||||
|
directory and serves it at `GET /api/hives`; the swarm UI's front page
|
||||||
|
renders it ([`ui.md`](ui.md)).
|
||||||
|
|
||||||
|
2. **Matrix federation** — when this host runs the homeserver (`deploy.matrix.enable`), tuwunel
|
||||||
|
federates with the peer's matrix server (discovered via the peer's
|
||||||
|
`.well-known/matrix/server` delegation, which the gateway serves).
|
||||||
|
Federation validates the peer's TLS certificate against the matrix
|
||||||
|
**container's** trust bundle, independent of this directory.
|
||||||
|
|
||||||
|
⚠️ On a hive with self-signed gateway certificates, the container
|
||||||
|
also trusts this hive's `trust-bundle.pem`, bound in at runtime and
|
||||||
|
added to the public CAs; it ends at the swarm root, so a peer whose
|
||||||
|
certificate chains to the same root validates. A peer outside this
|
||||||
|
swarm's root needs a CA-issued certificate (ACME). Federation firewall
|
||||||
|
and TLS requirements: [`integrations/matrix.md`](../integrations/matrix.md).
|
||||||
|
|
||||||
|
3. **WireGuard mesh** (optional) — `deploy.wireguard.enable` reads each
|
||||||
|
entry's `wireguardPublicKey`/`wireguardEndpoint`/`wireguardAddress`
|
||||||
|
to configure `wg-hive`. See [below](#wireguard-inter-hive-mesh-optional).
|
||||||
|
|
||||||
|
### WireGuard inter-hive mesh (optional)
|
||||||
|
|
||||||
|
The peer config above uses public HTTPS for all inter-hive traffic.
|
||||||
|
For private deployments — or to reduce latency and TLS overhead on
|
||||||
|
intra-swarm traffic — hyperhive can configure a host-to-host
|
||||||
|
WireGuard mesh.
|
||||||
|
|
||||||
|
#### Generating keys
|
||||||
|
|
||||||
|
On each hive host:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
|
||||||
|
wg pubkey < /etc/wireguard/hive.key # → share this with peer operators
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Config example (two hives)
|
||||||
|
|
||||||
|
```nix
|
||||||
|
# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
|
||||||
|
services.hyperhive = {
|
||||||
|
deploy.wireguard = {
|
||||||
|
enable = true;
|
||||||
|
privateKeyFile = "/etc/wireguard/hive.key";
|
||||||
|
address = "10.100.0.1/24";
|
||||||
|
listenPort = 51820; # optional, default 51820
|
||||||
|
};
|
||||||
|
|
||||||
|
# The same `hives` attrset both hosts hold — mesh fields included,
|
||||||
|
# since "where this hive can be dialled" is a fact about that hive.
|
||||||
|
swarm.hives = {
|
||||||
|
pr1ma = {
|
||||||
|
domain = "pr1ma.example.com";
|
||||||
|
wireguardPublicKey = "base64keyA=";
|
||||||
|
wireguardEndpoint = "198.51.100.1:51820";
|
||||||
|
wireguardAddress = "10.100.0.1/32";
|
||||||
|
};
|
||||||
|
edge = {
|
||||||
|
domain = "edge.corp";
|
||||||
|
wireguardPublicKey = "base64keyB=";
|
||||||
|
wireguardEndpoint = "203.0.113.42:51820";
|
||||||
|
wireguardAddress = "10.100.0.2/32";
|
||||||
|
};
|
||||||
|
};
|
||||||
|
};
|
||||||
|
|
||||||
|
# hive B (edge.corp, mesh IP 10.100.0.2)
|
||||||
|
services.hyperhive = {
|
||||||
|
deploy.wireguard = {
|
||||||
|
enable = true;
|
||||||
|
privateKeyFile = "/etc/wireguard/hive.key";
|
||||||
|
address = "10.100.0.2/24";
|
||||||
|
};
|
||||||
|
|
||||||
|
swarm.hives = { /* … identical to hive A's … */ };
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
#### What the mesh does
|
||||||
|
|
||||||
|
- hyperhive configures `networking.wireguard.interfaces.wg-hive` on the
|
||||||
|
host (not inside agent containers; containers reach peers via the
|
||||||
|
host's routing table).
|
||||||
|
- It opens UDP port 51820 (or `listenPort`) on the host firewall.
|
||||||
|
- `swarm-wireguard.nix` reads each entry's `wireguardAddress` directly
|
||||||
|
from `services.hyperhive.swarm.peerHives` to build `wg-hive`'s
|
||||||
|
`allowedIPs`, so intra-swarm traffic can route over the mesh address
|
||||||
|
rather than the public domain.
|
||||||
|
- It sets `persistentKeepalive = 25` by default; override or null to
|
||||||
|
disable (not needed when both sides have public IPs and no NAT).
|
||||||
|
|
||||||
|
#### NAT / one-sided endpoints
|
||||||
|
|
||||||
|
If one host is behind NAT and can't accept incoming connections, only
|
||||||
|
that host needs a null `wireguardEndpoint` on the peer config — the
|
||||||
|
other side initiates. With keepalive on, the NAT hole stays open.
|
||||||
|
|
||||||
|
If both hosts are behind NAT, you need a STUN relay or a third host
|
||||||
|
(exit node); hyperhive sets up neither.
|
||||||
|
|
||||||
## Cross-references
|
## Cross-references
|
||||||
|
|
||||||
- `docs/networking/snapshot-store.md` — the swarm's `btrfs receive` endpoint, and
|
- [`ui.md`](ui.md) — the swarm UI, the operator surface for "what hives exist"
|
||||||
the `swarm.snapshotStore` option that points a hive at it
|
- [`../networking/snapshot-store.md`](../networking/snapshot-store.md) — the
|
||||||
- `docs/process/conventions.md` § Hive identity — env vars, qualified labels
|
swarm's `btrfs receive` endpoint and the `swarm.snapshotStore` option
|
||||||
- `docs/integrations/matrix.md` — matrix federation, TLS cert autogeneration,
|
- [`../process/conventions.md`](../process/conventions.md) § Hive identity —
|
||||||
firewall posture
|
env vars, qualified labels
|
||||||
- `docs/swarm/ui.md` — the swarm-wide hive roster, now the operator
|
- [`../integrations/matrix.md`](../integrations/matrix.md) — matrix
|
||||||
surface for "what hives exist" (superseded the per-hive dashboard's
|
federation, TLS cert autogeneration, firewall posture
|
||||||
old "peer hives" display)
|
- [`../networking/gateway.md`](../networking/gateway.md) — nginx vhosts and
|
||||||
- `docs/networking/gateway.md` — nginx vhosts and the `.well-known/matrix/`
|
the `.well-known/matrix/` autodiscovery scheme
|
||||||
autodiscovery scheme
|
|
||||||
|
|
|
||||||
|
|
@ -1,7 +1,12 @@
|
||||||
# Swarm-wide services
|
# Swarm-wide services
|
||||||
|
|
||||||
Some things exist once per **swarm** rather than once per hive. Two
|
Some things exist once per **swarm** rather than once per hive: the forge,
|
||||||
options say where the optional ones live, and everything else derives:
|
the matrix homeserver, SSO, the secret store, the queue, and the metrics and
|
||||||
|
log stack. This page says which host runs them and what a hive that runs none
|
||||||
|
of them configures instead. The all-local quick start sets everything with one
|
||||||
|
line → [README](../../README.md#quick-start-an-all-local-swarm).
|
||||||
|
|
||||||
|
Two options say where they live, and everything else derives:
|
||||||
|
|
||||||
```nix
|
```nix
|
||||||
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
|
services.hyperhive.deploy.singleHostSwarm = true; # everything on this box
|
||||||
|
|
@ -14,10 +19,15 @@ here" means: every once-per-swarm service takes its `enable` from it.** That's t
|
||||||
sections below don't repeat it, so a service that stops deriving is a
|
sections below don't repeat it, so a service that stops deriving is a
|
||||||
visible difference rather than one more paragraph saying the same thing.
|
visible difference rather than one more paragraph saying the same thing.
|
||||||
|
|
||||||
`singleHostSwarm` is the all-on-one-box switch above it: it defaults
|
`singleHostSwarm` is the all-on-one-box mode above it. It defaults
|
||||||
both `deploy.allSwarmServices` and `swarm.ca.autoConfigure` (this host
|
`deploy.allSwarmServices`, the swarm CA (`swarm.ca.autoConfigure`, generated
|
||||||
generates the swarm CA here). You can still set each derived toggle on its own,
|
on this host), the swarm controller (`deploy.swarm-controller.enable`), the
|
||||||
which wins, so "all local except X" needs no further option.
|
host's `/etc/hosts` entries for the names it serves
|
||||||
|
(`gateway.localHostsEntry`), the queue's auth-callout keys
|
||||||
|
(`deploy.nats.autoGenerateCallout`) and where the secret store's bootstrap
|
||||||
|
token goes (`deploy.bao.bootstrapTokenFile`). You can still set each derived
|
||||||
|
toggle on its own, which wins, so "all local except X" needs no further
|
||||||
|
option.
|
||||||
|
|
||||||
**Both default to off**, and that's deliberate: a host can't tell
|
**Both default to off**, and that's deliberate: a host can't tell
|
||||||
whether it's meant to be the swarm's service host, so this is an
|
whether it's meant to be the swarm's service host, so this is an
|
||||||
|
|
@ -38,23 +48,29 @@ answers its name from its own resolver, so on a swarm spread over
|
||||||
more than one host, the operator's DNS has to resolve those names to that
|
more than one host, the operator's DNS has to resolve those names to that
|
||||||
host.
|
host.
|
||||||
|
|
||||||
|
<details><summary>Moving an existing hive's forge to the swarm's</summary>
|
||||||
|
|
||||||
A hive that stops running the forge keeps the old container's state at
|
A hive that stops running the forge keeps the old container's state at
|
||||||
`/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's
|
`/var/lib/nixos-containers/hive-forge/`. Nothing moves it to the swarm's
|
||||||
forge: push anything worth keeping there by hand. Its
|
forge: push anything worth keeping there by hand. Its
|
||||||
`/var/lib/hyperhive/forge-core-token` came from that old forge and
|
`/var/lib/hyperhive/forge-core-token` came from that old forge and
|
||||||
fails against the swarm's one.
|
fails against the swarm's one.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
## Deployment shapes
|
## Deployment shapes
|
||||||
|
|
||||||
Those two options are what makes the difference between deployments, so
|
Those two options are what makes the difference between deployments, so
|
||||||
the shapes worth naming are the ones they produce:
|
the shapes worth naming are the ones they produce:
|
||||||
|
|
||||||
- **All-local.** Everything on one machine:
|
- **All-local.** Everything on one machine:
|
||||||
`singleHostSwarm = true`. Setup is automatic apart from
|
`singleHostSwarm = true`, plus `deploy.hive-controller.enable = true` for
|
||||||
choosing a domain and creating the first user.
|
a hive to run agents on. After the first switch, the steps in
|
||||||
|
[`setup.md`](../getting-started/setup.md) remain.
|
||||||
- **Services on the swarm controller host.**
|
- **Services on the swarm controller host.**
|
||||||
`deploy.allSwarmServices = true` there; the required services
|
Set `deploy.allSwarmServices` and `deploy.swarm-controller.enable` there,
|
||||||
deploy together on that host, with hives elsewhere.
|
with hives elsewhere. The controller doesn't derive from
|
||||||
|
`allSwarmServices`.
|
||||||
- **Fully spread out.** One container / VM / machine per service,
|
- **Fully spread out.** One container / VM / machine per service,
|
||||||
somewhere.
|
somewhere.
|
||||||
|
|
||||||
|
|
@ -88,12 +104,10 @@ there is one IdP and one auth path.
|
||||||
container. Set it explicitly when joining a swarm whose IdP is under
|
container. Set it explicitly when joining a swarm whose IdP is under
|
||||||
another name.
|
another name.
|
||||||
|
|
||||||
swarm-controller writes the users database, not by hand: hive-c0re
|
Agent subjects come from swarm-controller's agent-creation job, written
|
||||||
creates and destroys agents continuously, so the subject set is dynamic.
|
into the users database by `swarm-authelia-bridge`; human ones come from
|
||||||
This module only guarantees the file exists and parses, so authelia
|
`swarmctl user add` → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
|
||||||
starts with nobody in it rather than failing to start — a provider with
|
On first boot this module seeds an empty users database. Authelia generates session and storage keys in the container on
|
||||||
no subjects yet is the correct state before anything has provisioned
|
|
||||||
them. Authelia generates session and storage keys in the container on
|
|
||||||
first boot and never rotates them automatically; replacing one
|
first boot and never rotates them automatically; replacing one
|
||||||
invalidates data already written (sessions, the encrypted store), so
|
invalidates data already written (sessions, the encrypted store), so
|
||||||
that's an operator action.
|
that's an operator action.
|
||||||
|
|
@ -196,9 +210,9 @@ gateway either way.
|
||||||
|
|
||||||
**Both store exporters are unconditional**, and `deploy.victoriametrics.enable`
|
**Both store exporters are unconditional**, and `deploy.victoriametrics.enable`
|
||||||
doesn't gate them: that option says this host _runs_ the store, while the swarm
|
doesn't gate them: that option says this host _runs_ the store, while the swarm
|
||||||
has one either way, reached by its swarm name through the gateway. Gating on it
|
has one either way, reached by its swarm name through the gateway. A collector
|
||||||
once left a collector on any other host with no exporter at all — receiving from
|
with no exporter would receive from every hive and drop it silently, because an
|
||||||
every hive and dropping it, silently, because an absent exporter isn't an error.
|
absent exporter isn't an error.
|
||||||
|
|
||||||
Agent-side configuration, and what a hive's own collector does, are in
|
Agent-side configuration, and what a hive's own collector does, are in
|
||||||
[`../scheduler/observability.md`](../scheduler/observability.md).
|
[`../scheduler/observability.md`](../scheduler/observability.md).
|
||||||
|
|
|
||||||
|
|
@ -1,10 +1,23 @@
|
||||||
# Swarm UI
|
# Swarm UI
|
||||||
|
|
||||||
The swarm's own web surface, served by the gateway on the **swarm apex**
|
The swarm's own web surface and the operator's day-to-day view: served by
|
||||||
(`services.hyperhive.swarm.domain`) and readable only by operators.
|
the gateway on the **swarm apex** (`services.hyperhive.swarm.domain`),
|
||||||
|
readable only by operators. The per-hive dashboard, on each hive's own
|
||||||
|
domain, covers host-level detail for one hive.
|
||||||
|
|
||||||
Distinct from the per-hive dashboard, which lives on the hive domain and
|
## What it shows
|
||||||
answers for one host. This one is the view _across_ hives.
|
|
||||||
|
| route | what |
|
||||||
|
| ------------------------- | ------------------------------------------------------------------------------------------- |
|
||||||
|
| `/` | the hive directory, each hive with its last reported status |
|
||||||
|
| `/agents` | every agent: status, config PR, wanted state; create agents, link forge and matrix accounts |
|
||||||
|
| `/agents/<name>/terminal` | one agent's live terminal |
|
||||||
|
| `/jobs` | the controller's job graph — where agent creation and credential mints show progress |
|
||||||
|
| `/issues` | a cross-repo issue report |
|
||||||
|
|
||||||
|
Everything it shows comes from [`swarm-controller`](../../swarm-controller/README.md).
|
||||||
|
An agent created here or with `swarmctl agent create` starts `paused`; set it
|
||||||
|
`up` from its card.
|
||||||
|
|
||||||
## Enabling
|
## Enabling
|
||||||
|
|
||||||
|
|
@ -35,28 +48,16 @@ requiring `group:admins`. An account without that group authenticates
|
||||||
fine and still gets bounced.
|
fine and still gets bounced.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
swarmctl user add <you> --group admins
|
swarmctl user add <you> --email <you>@example.com --group admins
|
||||||
|
swarmctl user update <you> --add-group admins # an account that already exists
|
||||||
```
|
```
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
`--email` isn't needed for the UI, but the forge won't create your account
|
||||||
|
without one → [setup.md § 2](../getting-started/setup.md#2--your-sso-account).
|
||||||
|
|
||||||
`admins` deliberately, not a new word: [`../getting-started/setup.md`](../getting-started/setup.md) has
|
Why a group and not "any session": agents are authelia subjects too, so
|
||||||
told every operator to create exactly that group since the bootstrap step
|
_authenticated_ includes every agent in the swarm. The group is the only
|
||||||
existed, so an account made by following the guide already passes. This
|
thing standing between "an operator's page" and "anyone with a session."
|
||||||
is the first rule that _consumes_ a group name — inventing a second one
|
|
||||||
would have meant those accounts silently failing a check they were
|
|
||||||
supposed to pass.
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
An account created without any group needs re-adding with the flag —
|
|
||||||
`swarmctl` reads the existing entry out of `users.yml,` so the group is
|
|
||||||
what changes.
|
|
||||||
|
|
||||||
Why a group and not a list of usernames: agents are getting authelia
|
|
||||||
accounts of their own (matrix SSO), and _authenticated_ would then
|
|
||||||
include every agent in the hive. The group is the only thing standing
|
|
||||||
between "an operator's page" and "anyone with a session."
|
|
||||||
|
|
||||||
## What it costs to be reachable
|
## What it costs to be reachable
|
||||||
|
|
||||||
|
|
@ -66,7 +67,22 @@ not a hole: **reachability isn't the access control here.** An agent
|
||||||
that resolves the name and connects still has no operator session, and
|
that resolves the name and connects still has no operator session, and
|
||||||
the subrequest denies it.
|
the subrequest denies it.
|
||||||
|
|
||||||
## Two wiring sites
|
## Quick links
|
||||||
|
|
||||||
|
The swarm UI's header carries a single 🔗 button, visible on every route,
|
||||||
|
opening a popover of links to other swarm-wide services. Backed by
|
||||||
|
`GET /api/links` (swarm-controller), which serves
|
||||||
|
`services.hyperhive.swarm.controller.links` (a `listOf { label, icon, url }`,
|
||||||
|
same shape as the per-agent `services.hyperhive.agent.dashboardLinks`).
|
||||||
|
|
||||||
|
Each service's own module contributes its entry when it's enabled on the
|
||||||
|
controller's host — `swarm-authelia.nix`, `hive-matrix.nix`,
|
||||||
|
`hive-forge/default.nix`, `swarm-grafana.nix`, `swarm-victorialogs.nix`, and
|
||||||
|
`swarm-ui.nix` for this UI's own API docs. Adding a link for a new service is
|
||||||
|
a nix-only change to that service's module, or an operator adding an entry
|
||||||
|
directly. An empty list hides the button.
|
||||||
|
|
||||||
|
<details><summary>Adding a swarm service name: the two wiring sites</summary>
|
||||||
|
|
||||||
Adding a swarm service name means touching two things. Missing the
|
Adding a swarm service name means touching two things. Missing the
|
||||||
second ships as a different flavour of "works from the host, broken from
|
second ships as a different flavour of "works from the host, broken from
|
||||||
|
|
@ -99,23 +115,7 @@ and the apex is a **sibling** of `forge.<swarm>` / `chat.<swarm>` /
|
||||||
implicitly. Left out, the vhost falls back to the hive leaf and the
|
implicitly. Left out, the vhost falls back to the hive leaf and the
|
||||||
swarm's front page opens with a name mismatch.
|
swarm's front page opens with a name mismatch.
|
||||||
|
|
||||||
## Quick links
|
</details>
|
||||||
|
|
||||||
The swarm UI's header carries a single 🔗 button, visible on every route,
|
|
||||||
opening a popover of links to other swarm-wide services — authelia,
|
|
||||||
matrix, forge, this UI's own swagger docs. Backed by `GET /api/links`
|
|
||||||
(swarm-controller), which serves `services.hyperhive.swarm.controller.links`
|
|
||||||
(a `listOf { label, icon, url }`, same shape as the per-agent
|
|
||||||
`services.hyperhive.agent.dashboardLinks`).
|
|
||||||
|
|
||||||
Rather than one central hardcoded list, each service's own module
|
|
||||||
contributes its own entry when it's actually enabled on the controller's
|
|
||||||
host — `swarm-authelia.nix`, `hive-matrix.nix` and `hive-forge/default.nix`
|
|
||||||
all do, the same list-merge idiom `gateway.localNames` uses above. Adding a
|
|
||||||
link for a new service is a nix-only change to that service's own module
|
|
||||||
(or an operator adding an entry directly); no swarm-controller or swarm-ui
|
|
||||||
change needed. Empty list hides the button rather than showing an empty
|
|
||||||
popover.
|
|
||||||
|
|
||||||
## Cross-references
|
## Cross-references
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -9,15 +9,34 @@ deliberately **not** derived from
|
||||||
`services.hyperhive.deploy.hive-controller.enable`: turning it on is a statement
|
`services.hyperhive.deploy.hive-controller.enable`: turning it on is a statement
|
||||||
about swarm topology, not about whether this host runs a hive.
|
about swarm topology, not about whether this host runs a hive.
|
||||||
|
|
||||||
## What it does today
|
## What it does
|
||||||
|
|
||||||
Serves one `/health` endpoint and holds no state.
|
The swarm's control plane. The swarm UI and `swarmctl` are its clients.
|
||||||
|
|
||||||
That is the whole intent of the first slice. The point is to make the _unit_
|
- **Hive directory** — serves `swarm.hives` (`GET /api/hives`) and what each
|
||||||
real — service user, runtime and state directories, socket, nginx
|
hive last published about itself (`GET /api/hives/status`).
|
||||||
reachability — so the swarm-level surfaces that follow have somewhere to land.
|
- **Agent roster and wanted state** — every agent the swarm knows
|
||||||
Inventing those surfaces before they are agreed would bake in a shape nobody
|
(`GET /api/agents`, `/api/agents/status`), and the state it declares for
|
||||||
chose. See the `hyperhive.swarm` consolidation epic.
|
each one on its hive (`up`/`offline`/`paused`/`destroyed`,
|
||||||
|
`PUT /api/hives/{hive}/agents/{agent}/state`).
|
||||||
|
- **Job graph** — a `hive-jobq` scheduler, served at `GET /api/jobq/graph`.
|
||||||
|
Every provisioning step below runs as a node in it.
|
||||||
|
- **Agent creation** — `POST /api/agents` queues the SSO identity (through
|
||||||
|
`swarm-authelia-bridge`), forge user, config repo, store identity, forge
|
||||||
|
token and matrix account, declares the agent `paused`, then sends its hive
|
||||||
|
a deploy message.
|
||||||
|
- **Agent credentials** — at start and every five minutes it re-checks every
|
||||||
|
agent's forge token and matrix account, and renews store certificates and
|
||||||
|
queue secrets as they age.
|
||||||
|
- **Swarm-wide forge objects and webhooks** →
|
||||||
|
[`docs/swarm/README.md`](../docs/swarm/README.md#swarm-wide-forge-objects).
|
||||||
|
- **Relays** — each agent's terminal and turn-state header as SSE, agent
|
||||||
|
icons, the cross-repo issue report, and the UI's quick links.
|
||||||
|
|
||||||
|
It reads its configuration once at startup, from the environment the nix module
|
||||||
|
sets. The one file it persists is `webhook-secret` in its state directory.
|
||||||
|
The job graph lives in memory; hive status and wanted state live in the swarm
|
||||||
|
queue, so both survive a restart.
|
||||||
|
|
||||||
## Why a unix socket, not a port
|
## Why a unix socket, not a port
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -31,18 +31,9 @@ that.
|
||||||
|
|
||||||
## One file, two writers
|
## One file, two writers
|
||||||
|
|
||||||
`users.yml` — authelia's own users database — is read and written
|
`swarmctl` reads and writes `users.yml` — authelia's own users database —
|
||||||
directly. There is no second store.
|
directly, with no second store. `swarm-authelia-bridge` writes agent
|
||||||
|
subjects into the same file.
|
||||||
There used to be: a private `users.json` here, canonical, with `users.yml`
|
|
||||||
rendered from it, while `swarm-authelia-bridge` kept its own pair against
|
|
||||||
the _same_ physical file. Two canonical stores for one file is a seam, and
|
|
||||||
it bit — a writer whose own JSON was missing could not tell "nothing here
|
|
||||||
yet" from "someone else's users", and refused to write.
|
|
||||||
|
|
||||||
The argument for the split was that it let this crate work without a YAML
|
|
||||||
parser. It didn't: the JSON was read back on every run, so the round-trip
|
|
||||||
was already being paid — the two files differed only in _format_.
|
|
||||||
|
|
||||||
⚠️ The file is round-tripped, so **comments and hand-formatting do not
|
⚠️ The file is round-tripped, so **comments and hand-formatting do not
|
||||||
survive a write**. Values do, and so do keys this binary does not model.
|
survive a write**. Values do, and so do keys this binary does not model.
|
||||||
|
|
@ -62,18 +53,30 @@ an error.
|
||||||
| `SWARMCTL_AUTHELIA_USERS_FILE` | host-side path of the users database |
|
| `SWARMCTL_AUTHELIA_USERS_FILE` | host-side path of the users database |
|
||||||
| `SWARMCTL_AUTHELIA_MACHINE` | container name, for `systemctl -M` |
|
| `SWARMCTL_AUTHELIA_MACHINE` | container name, for `systemctl -M` |
|
||||||
| `SWARMCTL_AUTHELIA_UNIT` | authelia's unit inside that container |
|
| `SWARMCTL_AUTHELIA_UNIT` | authelia's unit inside that container |
|
||||||
| `SWARM_CONTROLLER_SOCKET` | the controller's unix socket, for `agent create` |
|
| `SWARM_CONTROLLER_SOCKET` | the controller's unix socket, for `agent` and `forge` verbs |
|
||||||
|
|
||||||
|
Every verb and flag: [`docs/tools/swarmctl-cli.md`](../docs/tools/swarmctl-cli.md).
|
||||||
|
The sections below cover why each verb behaves as it does.
|
||||||
|
|
||||||
## `user add`
|
## `user add`
|
||||||
|
|
||||||
```console
|
```console
|
||||||
# swarmctl user add mara --display-name "Mara" --group admins
|
# swarmctl user add mara --display-name "Mara" --email mara@example.com --group admins
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Keep both flags:
|
||||||
|
|
||||||
|
- **`--group admins`** — the swarm UI and other operator surfaces gate on it.
|
||||||
|
- **`--email`** — the forge won't create an account without one.
|
||||||
|
|
||||||
|
`user add` refuses a username that already exists; fix an existing account
|
||||||
|
with `swarmctl user update mara --add-group admins --email …`.
|
||||||
|
|
||||||
The password is **generated by authelia** (`crypto hash generate argon2
|
The password is **generated by authelia** (`crypto hash generate argon2
|
||||||
--random`) and printed once. It is never passed on a command line:
|
--random`) and printed once. It is never passed on a command line:
|
||||||
`/proc/<pid>/cmdline` is world-readable, so a password in argv is readable
|
`/proc/<pid>/cmdline` is world-readable, so a password in argv is readable
|
||||||
by any local process for the lifetime of the call.
|
by any local process for the lifetime of the call. `user reset-password`
|
||||||
|
generates a new one the same way.
|
||||||
|
|
||||||
## `agent create`
|
## `agent create`
|
||||||
|
|
||||||
|
|
@ -89,13 +92,17 @@ module sets from the daemon's own `socketPath`). It prints the queued
|
||||||
job's node id **and stops there**.
|
job's node id **and stops there**.
|
||||||
|
|
||||||
It deliberately does not wait. The endpoint queues a DAG — SSO identity,
|
It deliberately does not wait. The endpoint queues a DAG — SSO identity,
|
||||||
forge user, forge repo, repo membership, config-repo seed, then a deploy
|
forge user, forge repo, repo membership, config-repo seed, store identity,
|
||||||
|
forge token, matrix account, a `paused` wanted state, then a deploy
|
||||||
message — and the last of those _publishes_: the hive's `hive-c0re` picks
|
message — and the last of those _publishes_: the hive's `hive-c0re` picks
|
||||||
it up and converges on its own clock, out of the controller's sight. So
|
it up and converges on its own clock, out of the controller's sight. So
|
||||||
even a fully settled graph would not mean the agent is up, and there is
|
even a fully settled graph would not mean the agent is up, and there is
|
||||||
nothing this CLI could wait for that would let it say so honestly. Watch
|
nothing this CLI could wait for that would let it say so honestly. Watch
|
||||||
the swarm UI's job view for the rest.
|
the swarm UI's job view for the rest.
|
||||||
|
|
||||||
|
The new agent starts `paused`: it doesn't drive turns until you set it `up`
|
||||||
|
in the swarm UI.
|
||||||
|
|
||||||
No approval gate, for the same reason nothing else here has one: running
|
No approval gate, for the same reason nothing else here has one: running
|
||||||
this binary already means being root on the controller's host.
|
this binary already means being root on the controller's host.
|
||||||
|
|
||||||
|
|
@ -110,6 +117,30 @@ crate does not link it, and there is no wire-type crate between them.
|
||||||
Two fields out, two in, both ends validating — a drift shows up as a
|
Two fields out, two in, both ends validating — a drift shows up as a
|
||||||
400 naming the field.
|
400 naming the field.
|
||||||
|
|
||||||
|
## `agent mint-identity`
|
||||||
|
|
||||||
|
```console
|
||||||
|
# swarmctl agent mint-identity ruth
|
||||||
|
```
|
||||||
|
|
||||||
|
`POST /api/agents/{name}/identity` on the swarm-controller. Re-runs the
|
||||||
|
store-identity mint for one agent that already exists — the backfill for an
|
||||||
|
agent the swarm never created, such as the manager agent `hive-c0re` makes at
|
||||||
|
startup. It re-mints the agent's store certificate, which the agent picks up
|
||||||
|
the next time its container boots, and leaves an existing queue secret alone.
|
||||||
|
Queues and returns, like `agent create`.
|
||||||
|
|
||||||
|
## `agent mint-forge-token`
|
||||||
|
|
||||||
|
```console
|
||||||
|
# swarmctl agent mint-forge-token ruth
|
||||||
|
```
|
||||||
|
|
||||||
|
`POST /api/agents/{name}/forge-token` on the swarm-controller. Checks one
|
||||||
|
agent's forge token and mints it if it's missing or stale. The controller
|
||||||
|
already does this for every agent with a store identity at start and every
|
||||||
|
five minutes; this verb skips the wait. Queues and returns.
|
||||||
|
|
||||||
## `forge make-admin`
|
## `forge make-admin`
|
||||||
|
|
||||||
```console
|
```console
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue