From b83fdc60f324127761fc1a0bf8d3c795ff88d84a Mon Sep 17 00:00:00 2001 From: atlas Date: Fri, 2 Oct 2026 09:26:41 +0200 Subject: [PATCH] docs: drop statements about absent things, state current behaviour --- docs/integrations/matrix.md | 19 +++++++------------ docs/networking/gateway.md | 12 ++++++------ docs/networking/network.md | 4 +--- docs/scheduler/observability.md | 17 ++++++++--------- 4 files changed, 22 insertions(+), 30 deletions(-) diff --git a/docs/integrations/matrix.md b/docs/integrations/matrix.md index 32db5285..aad3babc 100644 --- a/docs/integrations/matrix.md +++ b/docs/integrations/matrix.md @@ -207,17 +207,14 @@ tuwunel has none. Promoting a user to homeserver admin and resetting a password both need an admin **sender**: `!admin …` messages into `#admins:`, and tuwunel only treats a message as a command when its sender is already an -admin. `@hive-:` has no admin sender to make that call with. Both are -swarm-level operations. +admin. Only `@swarm` is a homeserver admin, so both are swarm-level +operations.
Store precedence for a hive's sender account -`swarm/services/matrix/sender-token` is a shared path `hive-c0re` and -`swarm-controller` both build from the same function; nothing reads it. `ensure_hive_user` reads the per-hive path `swarm/hives//matrix/sender-token` **first, on every sweep**, not -just when that file is missing (`sender_source`'s decision), so a value -at the shared path never takes effect once the per-hive path has one. +just when that file is missing (`sender_source`'s decision). `swarm-controller` mints a per-hive token within five minutes of a hive appearing; the sweep then overwrites the per-hive file, no boot required. @@ -246,9 +243,7 @@ halves. `registrationTokenFile` is a removed option: a config that still sets it fails to evaluate with a message naming the appservice. -- **Access tokens are independent of the registration token.** An - access token lives on the device that minted it, so accounts and - sessions are unaffected by the registration token's presence. +- **Access tokens live on the device that minted them.** `login_with_password` stays on, so the password fallback works too. - **The homeserver's `admin_execute` promotes only `@swarm` at boot** @@ -271,9 +266,9 @@ Initial rollout settings: until you list peers. - `allow_registration = false`. tuwunel checks this flag only for requests that arrive **without** an appservice token, so the appservices - still create accounts and tuwunel refuses everyone else. It's not a - hardening afterthought: with no registration token configured, - `allow_registration = true` makes tuwunel refuse to start unless + still create accounts and tuwunel refuses everyone else. With no + registration token configured, `allow_registration = true` makes + tuwunel refuse to start unless `yes_i_am_very_very_sure_…_open_registration_…` is also set. - `allow_encryption` — server-side E2EE switch, sourced from `services.hyperhive.swarm.matrix.allowEncryption` (**default `false`**, opt-in). diff --git a/docs/networking/gateway.md b/docs/networking/gateway.md index 6d9caa36..8a9a3adf 100644 --- a/docs/networking/gateway.md +++ b/docs/networking/gateway.md @@ -2,7 +2,7 @@ Every host's nginx: the one front door for whatever this host serves. A swarm service running here (forge, matrix, SSO, the swarm UI, the metrics and log stores) declares its own vhost through the gateway; the gateway itself adds the hive's own surface — dashboard, per-agent UIs, matrix discovery. -_For the operator configuring `services.hyperhive.gateway.*` on a host._ nginx and the hive resolver (dnsmasq) run on the host next to hive-c0re, not in a container: they bind `:80`/`:443` and the bridge address, so a network namespace of their own would isolate nothing. +_For the operator configuring `services.hyperhive.gateway.*` on a host._ nginx and the hive resolver (dnsmasq) run on the host next to hive-c0re, not in a container: they bind `:80`/`:443` and the bridge address. You rarely switch it on yourself. `gateway.enable` defaults to off, and every module that serves a vhost or needs hive names to resolve sets `gateway.enable` / `gateway.dns.enable` with `mkDefault true` — the hive controller, each swarm service, CI. @@ -41,7 +41,7 @@ Per-agent UIs stay sub-path; forge and matrix get sub-domains → [Sub-domain sh ## TLS modes -The gateway always terminates TLS: there is no http-only mode. Which certificate it serves depends on what you configure: +The gateway always terminates TLS. Which certificate it serves depends on what you configure: | mode | config | cert source | `.well-known` scheme | |---|---|---|---| @@ -77,7 +77,7 @@ The issuer is a **host-held hive CA**, not a bare self-signed leaf. `hive-tls-ca ⚠️ **Keep the import unit.** It does two jobs nginx needs: it re-modes the key to `0640 root:nginx` (nginx's pre-start `nginx -t` runs as the nginx user and fails on the CA's `0600 root:root` key), and it makes sure **every cert path the config names exists** — when the swarm-services leaf is missing it installs the hive leaf in its place. nginx refuses a config naming a missing cert file, so without that fallback one missing leaf takes down every vhost, not just one. -**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust, and nothing can wire a runtime-generated leaf into an agent's build-time trust store. Agents and federation peers trust the stable CA once; leaf rotation never breaks them. +**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust, and an agent's build sets its trust store once, before any leaf exists. Agents and federation peers trust the stable CA once; leaf rotation never breaks them. **What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is an intermediate under the swarm root ([`swarm/ca.md`](../swarm/ca.md) has the hierarchy), and a verifier can't stop at an intermediate — so the bundle carries the hive CA plus its root. nginx serves the leaf with the hive CA appended for the same reason. Agents (via `security.pki.certificateFiles`), the CI and forge containers and federating peers all read the bundle. @@ -107,7 +107,7 @@ nginx reads the directory directly. Keep the key readable by nginx: ### Fronting with an external TLS terminator -The gateway has no plain-http upstream mode. Either give the gateway the real cert (`tls.certDir` or `tls.acme`) so it serves proper TLS itself, or front it over a unix socket rather than a plain-http TCP port. `.well-known/matrix/*` responses always advertise `https` ([Discovery flow](#discovery-flow-matrix)). +Give the gateway the real cert (`tls.certDir` or `tls.acme`) so it serves proper TLS itself, or front it over a unix socket rather than a plain-http TCP port. `.well-known/matrix/*` responses always advertise `https` ([Discovery flow](#discovery-flow-matrix)). ## HTTP Basic auth @@ -234,7 +234,7 @@ matrix-dart-sdk (FluffyChat and others) always fetches the well-known over `http Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.:"}`. The port is always explicit, even 443: a delegated host without a port means the federation default 8448, not 443. Peers then reach `/_matrix/` on the chat vhost through the gateway, so the gateway must be reachable from them (`gateway.openFirewall`). -⚠️ The gateway serves both `.well-known` routes on the **hive** vhost. Matrix looks them up at the `serverName`, which defaults to the bare swarm domain, and the swarm UI's apex vhost serves no `.well-known/matrix/*` route. +⚠️ The gateway serves both `.well-known` routes on the **hive** vhost only. Matrix looks them up at the `serverName`, which defaults to the bare swarm domain, so a lookup against the swarm domain reaches the swarm UI's apex vhost and finds nothing — pin `serverName` to the hive domain, or serve `.well-known` at the swarm domain yourself. ## Sub-domain shape (rationale) @@ -266,7 +266,7 @@ The `chat.` vhost (fluffychat) serves a flutter/SPA bundle via the Accept - hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap) - a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type -Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $uri/ $matrix_spa_target =404` falls through to a plain `404` — a missing asset is just missing. No extension allowlist, no `if` block, no regex heuristics. +Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $uri/ $matrix_spa_target =404` falls through to a plain `404` — a missing asset is just missing. #### Dashboard: path-based routing (not Accept-header) diff --git a/docs/networking/network.md b/docs/networking/network.md index 60c34647..d399fe85 100644 --- a/docs/networking/network.md +++ b/docs/networking/network.md @@ -7,9 +7,7 @@ need the bridge set it with `mkDefault true` — the hive controller, CI, the hive collector, and the gateway's resolver, which every swarm service turns on. Configured via `services.hyperhive.network.*`. -Isolation is the only mode; agent containers never share the host netns. -`services.hyperhive.network.isolateContainers` and -`services.hyperhive.network.upstreamDns` don't exist. +Agent containers always run in a private netns. ## Network map diff --git a/docs/scheduler/observability.md b/docs/scheduler/observability.md index fcaaf4ce..1825aa56 100644 --- a/docs/scheduler/observability.md +++ b/docs/scheduler/observability.md @@ -13,12 +13,12 @@ catalogued below. Telemetry crosses two collectors, and which one you configure depends on what the host is: -| | runs where | receives from | does | -| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- | -| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's stores and exports upstream | -| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination | +| | runs where | receives from | does | +| ------------------------------------ | ---------------------- | --------------------------------- | ----------------------------------------------- | +| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's stores and exports upstream | +| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential | -An all-local host runs both, and needs nothing said about the hop between them. +An all-local host runs both; the hop between them configures itself. ```nix services.hyperhive.otel = { @@ -31,7 +31,7 @@ services.hyperhive.otel = { ## Enabling export `otel.enable` is the single gate on a hive: one switch in the host config -covers every agent container on it, with no per-agent opt-in or opt-out. +covers every agent container on it. `endpoint` is where telemetry ends up after it leaves the swarm — optional, because the swarm's own metrics store (`deploy.victoriametrics`) is a destination in its own right. With both, telemetry goes to both. See @@ -50,7 +50,7 @@ The hive collector reaches the swarm collector by its gateway name (`swarm.otel.domain`, default `otel.`) — the same DNS-and-CA-trust shape every hive-to-swarm-service hop uses. On the host running the swarm collector the hive's dnsmasq answers that name; elsewhere it resolves through -ordinary DNS. Nothing here needs setting for the split-host case. +ordinary DNS. ⚠️ **The hive collector carries every bit of that hive's telemetry.** It runs on the same host as the agents and restarts on failure. Telemetry isn't @@ -62,8 +62,7 @@ the control plane, so degraded telemetry isn't degraded operation. every agent needs the credential — and the only place to hand it to an agent container is somewhere the agent itself can read, its own claude settings among them. `0600` protects a secret from other containers, not from the -agent it belongs to. An option that could select that path would reopen the -hole. +agent it belongs to. **The tiers stay separate on one box.** An all-local hive is a statement about _where_ processes run, not about the shape of the deployment. A boundary that