From 067f4e5699e8c89bb91772a8e6a8fa7a23da4faf Mon Sep 17 00:00:00 2001 From: atlas Date: Thu, 1 Oct 2026 23:31:46 +0200 Subject: [PATCH] docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix gateway.md: split the opener into what/audience/enable; vhost map in two tables (swarm-service vhosts declared by their own modules, then the hive vhost) matching vhosts.nix and the service modules; gateway.enable exists and is set with mkDefault by the modules that need it; Basic auth scope, dashboard /health/ prefix, error-page rendering, matrix body limit and forge link source corrected; nginx internals grouped under one Internals section with their headings unchanged. network.md: gateway and dnsmasq run on the host, not in a container; network.enable is set by the modules that need it; shared-netns firewall rule covers every swarm service container; hive-priv writes the nspawn conf; domain sentence rewritten; removed options moved into
. jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS show different graphs drawn by the same component. observability.md: swarm tier first; history narration cut; network access deduplicated into a link to network.md; options link made absolute. matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and SSO options under deploy.matrix; .well-known is served on the hive domain; roadmap sentence deleted; stale hive-c0re provisioning claims fixed; serverName upgrade note moved into
. Refs #3902 --- docs/integrations/matrix.md | 57 ++- docs/networking/gateway.md | 791 ++++++++++++-------------------- docs/networking/network.md | 62 +-- docs/scheduler/jobq.md | 32 +- docs/scheduler/observability.md | 162 +++---- 5 files changed, 452 insertions(+), 652 deletions(-) diff --git a/docs/integrations/matrix.md b/docs/integrations/matrix.md index d46abf7a..5658a245 100644 --- a/docs/integrations/matrix.md +++ b/docs/integrations/matrix.md @@ -2,9 +2,16 @@ Private Matrix homeserver (matrix-tuwunel — the conduwuit successor) wrapped in a nixos-container, plus optional fluffychat-web -client at `chat./` (the `gatewayHost` vhost). Configured via -`services.hyperhive.swarm.matrix.*`; vhost routing lives in -[`gateway.md`](../networking/gateway.md). +client at `chat./` (the `gatewayHost` vhost). A swarm runs one +homeserver, on one host. Two namespaces configure it: + +- `services.hyperhive.swarm.matrix.*` — what the homeserver **is**, as + every hive sees it: `serverName`, `gatewayHost`, ports, `allowEncryption`. +- `services.hyperhive.deploy.matrix.*` — what the host running it decides: + `enable`, `gui.enable`, `openFirewall`, `trustedServers`, + `maxRequestSize`, `sso.clientSecretFile`. + +Vhost routing lives in [`gateway.md`](../networking/gateway.md). ## Container shape @@ -33,9 +40,10 @@ Two distinct hostnames: *irrevocably* in every `@user:` and `!room:` identifier minted on this homeserver. You can't change it later without abandoning every account and chat history. Defaults to the - bare `services.hyperhive.swarm.domain`; clients autodiscover the - actual API endpoint via the `.well-known/matrix/{client,server}` - routes the gateway serves at that domain. + bare `services.hyperhive.swarm.domain`. The gateway serves the + `.well-known/matrix/{client,server}` discovery routes on the matrix + host's **hive** domain, not the swarm domain + ([Discovery flow](../networking/gateway.md#discovery-flow-matrix)). - **`gatewayHost`** — the API listener hostname, where the gateway's matrix vhost proxies `/_matrix/*` to tuwunel. Defaults to `chat.`. Set to `null` to skip the @@ -54,12 +62,12 @@ room id, so adopting a new one does **not** rename the existing users and rooms it strands them, because their ids still name a homeserver that no longer answers. -### Upgrading a homeserver that already has ids +
Upgrading a homeserver that already has ids -`serverName`'s default has changed across releases. A homeserver that -has already minted ids under an older default must **pin the value it -actually minted them under**, not adopt the new default — see above -for why adopting a new one strands existing users and rooms: +A homeserver that minted ids under an older `serverName` default must +**pin the value it actually minted them under**, not adopt the current +default — see above for why adopting a new one strands existing users +and rooms: ```nix services.hyperhive.swarm.matrix = { @@ -71,13 +79,14 @@ services.hyperhive.swarm.matrix = { A rebuild on a host that already has a homeserver prints a `hive-matrix: WARNING — … serverName is unset` line when this is missing, -naming the value it's about to default to. That warning is why this -section exists; it never fails the rebuild, so it's on you to act on it -before the homeserver mints the ids. +naming the value it's about to default to. It never fails the rebuild, so act on +it before the homeserver mints the ids. + +
## Default-closed firewall -`openFirewall` defaults to `false` (secure-by-default): the host +`deploy.matrix.openFirewall` defaults to `false`: the host reaches the homeserver on loopback, and agent containers reach it at `chat.` via the gateway — so the firewall hole only matters for access from *outside* the host. Flip to `true` when @@ -198,9 +207,8 @@ tuwunel has none. Promoting a user to homeserver admin and resetting a password both need an admin **sender**: `!admin …` messages into `#admins:`, and tuwunel only treats a message as a command when its sender is already an -admin. `@hive-:` has no admin sender to make that call with. They're -swarm-level operations: matrix admin should eventually come from -membership in authelia's `admins` group; nobody has built that sync yet. +admin. `@hive-:` has no admin sender to make that call with. Both are +swarm-level operations.
Upgrading a hive that shared one sender account with every other hive @@ -274,8 +282,8 @@ Initial rollout settings: restart. `trusted_servers = []` keeps it effectively closed until you list peers. - `allow_registration = false`. tuwunel checks this flag only for - requests that arrive **without** an appservice token, so hive-c0re - provisions exactly as before and tuwunel refuses everyone else. It's not a + requests that arrive **without** an appservice token, so the appservices + still create accounts and tuwunel refuses everyone else. It's not a hardening afterthought: with no registration token configured, `allow_registration = true` makes tuwunel refuse to start unless `yes_i_am_very_very_sure_…_open_registration_…` is also set. @@ -298,8 +306,7 @@ Initial rollout settings: ## Hive Matrix Space -On first boot, after hive-c0re provisions all agent accounts, it -creates a private **Matrix Space** named `"hive"` using its own hive +On first boot hive-c0re creates a private **Matrix Space** named `"hive"` using its own hive account (`@hive-:`) and invites every provisioned agent into it. This gives the operator a single Space in FluffyChat or any Matrix client that groups all agent-to-agent + operator rooms in one @@ -331,7 +338,7 @@ re-creation (for example after a homeserver wipe). ## Configuration tuning ```nix -services.hyperhive.swarm.matrix = { +services.hyperhive.deploy.matrix = { trustedServers = [ "matrix.org" "example.com" ]; # default: [] maxRequestSize = 20000000; # default: 20 MB }; @@ -364,14 +371,14 @@ surprising behaviour: SSO is unconditional, so the three below are requirements of running a homeserver at all rather than of a setting: -- **Set `sso.clientSecretFile`** — fails at eval, not at boot: +- **Set `deploy.matrix.sso.clientSecretFile`** — fails at eval, not at boot: tuwunel reads its identity providers from the config file, so a half-configured one can stop the homeserver from starting outright rather than merely hiding a login button. On a host that also runs the swarm's authelia it's wired up for you. - **Set `swarm.authelia.url`** — without a provider URL there is nothing to discover against. -- **Set `gatewayHost != null`** — the SSO callback URL is +- **Set `swarm.matrix.gatewayHost != null`** — the SSO callback URL is format-locked to `/_matrix/client/unstable/login/sso/callback/`, and the identity provider needs a public name to redirect the browser to. diff --git a/docs/networking/gateway.md b/docs/networking/gateway.md index c500b4a8..80adc2db 100644 --- a/docs/networking/gateway.md +++ b/docs/networking/gateway.md @@ -1,103 +1,290 @@ # hive-gateway -This host's nginx fronts the hyperhive web surfaces running on it — next to hive-c0re, not in its own container: it shares the host netns anyway (see [Vhost map](#vhost-map) below), so containerizing it would buy no network isolation while costing a resolv.conf sync, a machine-bus reload, and three bind mounts. System-config (not meta-flake managed). Configured via `services.hyperhive.gateway.*` + per-subsystem opt-in flags in `services.hyperhive.{forge,matrix,...}`. the modules that need them assert `gateway.enable` and `gateway.dns.enable`, so a host serving a vhost or resolving hive names gets them without an opt-in. +Every host's nginx: the one front door for whatever this host serves. A swarm service running here (forge, matrix, SSO, the swarm UI, the metrics and log stores) declares its own vhost through the gateway; the gateway itself adds the hive's own surface — dashboard, per-agent UIs, matrix discovery. + +_For the operator configuring `services.hyperhive.gateway.*` on a host._ nginx and the hive resolver (dnsmasq) run on the host next to hive-c0re, not in a container: they bind `:80`/`:443` and the bridge address, so a network namespace of their own would isolate nothing. + +You rarely switch it on yourself. `gateway.enable` defaults to off, and every module that serves a vhost or needs hive names to resolve sets `gateway.enable` / `gateway.dns.enable` with `mkDefault true` — the hive controller, each swarm service, CI. ## Vhost map -| URL | vhost | upstream | source | -| --- | --- | --- | --- | -| `/` | `_` (catch-all) | dashboard dist (static, from `servedFrontend`); `/api/` + `/webhook/` → hive-c0re (`7000`) | always | -| `/agent//` | `_` | per-agent harness (UDS or TCP) | `agents.conf` (runtime-generated) | -| `/.well-known/matrix/{client,server}` | `_` | inline JSON (no upstream) | `matrix.enable && domain != null` | -| `/matrix/` (deprecated) | `_` | 301 → `chat./` | `matrix.gui.enable` | -| `forge./` | `forge.` | forgejo (`3000`) | `deploy.forgejo.behindGateway` | -| `chat./_matrix/*` | `chat.` | tuwunel (`8008`) | `matrix.gatewayHost != null` | -| `chat./` | `chat.` | fluffychat-web static | `matrix.gui.enable` | -| `chat./config.json` | `chat.` | inline JSON (FluffyChat boot config) | `matrix.gui.enable && domain != null` | -| `auth./` | `auth.` | authelia (`9091`) | `deploy.authelia` | -| `/` | `` | swarm-ui dist (static), behind an authelia subrequest | `deploy.swarm-ui` | +**Swarm services.** Each module declares its vhost on the host that runs the service, under the swarm domain: -Only the host that **runs** authelia declares the authelia vhost, not every hive that uses it — a client hive knows the swarm's `authelia.url` but must not answer for a name it doesn't serve. Its server name is exactly `swarm.authelia.domain`: authelia validates `authelia_url ⊂ session cookie domain` at startup, so a near-miss is a container that refuses to boot. It carries no `auth_basic` — the login page must not sit behind the login mechanism it replaces — and sets the four `X-Forwarded-{Proto,Host,Uri,For}` headers, since authelia decides by the *original* request rather than the hop it sees. +| URL | upstream | declared by, when | +| --- | --- | --- | +| `/` | swarm-ui dist (static), behind an authelia subrequest | `swarm-ui.nix`, `deploy.swarm-ui.enable` | +| `auth./` | authelia (`9091`) | `swarm-authelia.nix`, `deploy.authelia.enable` | +| `forge./` | forgejo (`3000`) | `hive-forge/`, `deploy.forgejo.behindGateway` | +| `chat./_matrix/*` | tuwunel (`8008`) | `hive-matrix.nix`, `swarm.matrix.gatewayHost != null` | +| `chat./` | fluffychat-web static (404 with the GUI off) | `hive-matrix.nix`, `deploy.matrix.gui.enable` | +| `chat./config.json` | inline JSON (FluffyChat boot config) | `hive-matrix.nix`, `deploy.matrix.gui.enable` | +| `grafana.`, `metrics.`, `logs.`, `otel.`, `bao.` | the matching swarm service | that service's module → [`swarm/services.md`](../swarm/services.md) | - -⚠️ **A `502` from this vhost means authelia itself isn't answering, not that the proxy is misconfigured.** Check `journalctl -M swarm-authelia -u authelia-swarm` before suspecting anything here. A swarm with no users of its own answers with a login page that refuses everyone; the first account is created in [`swarm/sso.md`](../swarm/sso.md). - +**The hive's own vhost**, named for the hive domain: -Per-agent UIs stay sub-path, forge and matrix get sub-domains — see -[Sub-domain shape (rationale)](#sub-domain-shape-rationale) below for why. +| URL | upstream | when | +| --- | --- | --- | +| `/` | dashboard dist (static, from `servedFrontend`) | always | +| `/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always | +| `/api/docs/` | themed Swagger UI dist (static) | always | +| `/agent//` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) | +| `/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` | +| `/matrix/` (deprecated) | 301 → `chat./` | `deploy.matrix.gui.enable` and `gatewayHost` set | + +The catch-all `_` vhost answers any other `Host` with `444` (connection closed, no response). It's `mkDefault`, so to make your own vhost the default server, set `services.nginx.virtualHosts."_".default = false;` — an eval assertion names both when two claim it. + +Only the host that **runs** authelia declares `auth.`; a hive that merely uses SSO knows `swarm.authelia.url` but doesn't answer for that name. The server name must be exactly `swarm.authelia.domain` — authelia checks that `authelia_url` sits inside its session cookie domain at startup, and refuses to boot otherwise. The vhost carries no Basic auth (that would put the login page behind the login it replaces) and passes `X-Forwarded-{Proto,Host,Uri,For}`, because authelia decides by the original request. + +⚠️ `auth.` answering with the **sso unavailable** page means authelia isn't answering at all — check `journalctl -M swarm-authelia -u authelia-swarm`. A swarm with no real accounts yet boots fine (a disabled placeholder user keeps authelia's user store non-empty) and serves a login page that refuses everyone; add the first account with `swarmctl user add …` ([`setup.md`](../getting-started/setup.md)). + +Per-agent UIs stay sub-path; forge and matrix get sub-domains → [Sub-domain shape](#sub-domain-shape-rationale). + +## TLS modes + +The gateway always terminates TLS: there is no http-only mode. Which certificate it serves depends on what you configure: + +| mode | config | cert source | `.well-known` scheme | +|---|---|---|---| +| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | the host's hive CA signs a gateway leaf (RSA-4096) | `https` | +| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` | +| operator cert | `tls.certDir` set | read from the operator's dir | `https` | + +`tls.certDir` together with `tls.acme.enable` fails an assertion — pick one. + +⚠️ One vhost class has no TLS: a vhost bound to loopback for a local consumer. `grafana-metrics` (`nix/host-modules/swarm-grafana.nix`) is the one in the tree: it listens on `127.0.0.1` only and serves `= /metrics` from grafana's unix socket for this host's collector. Everything with a routable name follows the table above. + +### ACME / Let's Encrypt (`tls.acme`) + +The simplest production path for a public domain: + +```nix +services.hyperhive.gateway = { + openFirewall = true; + tls.acme = { + enable = true; + email = "admin@example.com"; # required + }; +}; +``` + +nginx obtains and renews certs via the HTTP-01 challenge on `port` (default 80); they land in `/var/lib/acme/`, managed by nixpkgs's `security.acme`. Every name this host serves needs a public DNS record pointing here — the hive domain and each swarm-service name in the [vhost map](#vhost-map) — and `openFirewall = true` so Let's Encrypt reaches `/.well-known/acme-challenge/`. + +### Self-signed TLS (default) + +On by default, listening on `httpsPort` (default 443) beside the plain-http `port` (default 80). + +The issuer is a **host-held hive CA**, not a bare self-signed leaf. `hive-tls-ca.service` (from `hive-tls.nix`) generates a long-lived CA (`services.hyperhive.deploy.hive-controller.tls.caValidityDays`, default 7300 days) under `…tls.stateDir` (default `/var/lib/hive-tls`) and signs a gateway **leaf** with it (`leafValidityDays`, default 30). `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`). + +⚠️ **Keep the import unit.** It does two jobs nginx needs: it re-modes the key to `0640 root:nginx` (nginx's pre-start `nginx -t` runs as the nginx user and fails on the CA's `0600 root:root` key), and it makes sure **every cert path the config names exists** — when the swarm-services leaf is missing it installs the hive leaf in its place. nginx refuses a config naming a missing cert file, so without that fallback one missing leaf takes down every vhost, not just one. + +**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust, and nothing can wire a runtime-generated leaf into an agent's build-time trust store. Agents and federation peers trust the stable CA once; leaf rotation never breaks them. + +**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is an intermediate under the swarm root ([`swarm/ca.md`](../swarm/ca.md) has the hierarchy), and a verifier can't stop at an intermediate — so the bundle carries the hive CA plus its root. nginx serves the leaf with the hive CA appended for the same reason. Agents (via `security.pki.certificateFiles`), the CI and forge containers and federating peers all read the bundle. + +**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) fetches `https:///.well-known/matrix/client` and never falls back to plain http, so without TLS the browser client can't bootstrap. + +**Cert shape**: the leaf's CN is the bare hive domain; its SANs are `` and `*.`. The hive CA is name-constrained to ``, so it can't sign a swarm-service name outside it — a violating SAN would invalidate the whole leaf. Those names get the swarm-services leaf instead ([`swarm/ca.md`](../swarm/ca.md)). + +**Rotation**: `hive-tls-ca.service` re-signs the leaf when it's missing, within 30 days of expiry, or no longer covers the configured names, always under the same CA. It regenerates the CA only if missing or expired. To force a leaf rotation, delete `gateway.pem` under the state dir, restart the unit, then reload `nginx`. + +**Cert prompts**: browsers warn once per host until you add the hive's `trust-bundle.pem` (the anchor, not the leaf) to the browser or OS trust store. + +### Operator-provided cert (`tls.certDir`) + +For a cert from a real CA (Let's Encrypt via your own `security.acme`, a corporate CA): + +```nix +services.hyperhive.gateway = { + tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir + # tls.certName = "cert.pem"; # default — matches security.acme layout + # tls.keyName = "key.pem"; # default — matches security.acme layout +}; +``` + +nginx reads the directory directly. Keep the key readable by nginx: + +- `security.acme` writes keys `0640 root:acme`, which the `nginx` user can't read. Set `security.acme.certs."example.com".group = "nginx";` (or make the key `0644` if your threat model allows). Otherwise nginx fails at startup with the reason in the journal. + +### Fronting with an external TLS terminator + +The gateway has no plain-http upstream mode. Either give the gateway the real cert (`tls.certDir` or `tls.acme`) so it serves proper TLS itself, or front it over a unix socket rather than a plain-http TCP port. `.well-known/matrix/*` responses always advertise `https` ([Discovery flow](#discovery-flow-matrix)). + +## HTTP Basic auth + +Optional: Basic auth on the hive's dashboard. Add a login first, then enable it — the htpasswd file exists from first boot, and an empty one refuses everyone. + +_On the hive host:_ + +```sh +hivectl gateway create-user alice --password-stdin # password on stdin +hivectl gateway delete-user bob +hivectl gateway list-users +``` + +```nix +services.hyperhive.gateway.auth = { + enable = true; + # realm = "hyperhive"; # default; must not contain `"` or `$` +}; +``` + +`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password ` also works but lands in shell history. + +**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent//` routes, which come from `agents.conf` and inherit no auth from `/`. + +A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first. + +## Firewall posture (host-level) + +`services.hyperhive.gateway.openFirewall = true` opens `port` and `httpsPort` on the host firewall — both, since the gateway always serves TLS. It defaults to off; set it for any reach from outside the host. + +nginx is the one external entry point. The per-agent web-port range (`8100`–`8999`) stays closed: agents serve their UI on a unix socket ([Per-agent unix-socket upstream](#per-agent-unix-socket-upstream)), and the hashed TCP port (`lifecycle::agent_web_port`) is only a fallback bind for an agent missing `HIVE_WEB_SOCKET` — the gateway never proxies through it. + +The dashboard port (`services.hyperhive.c0re.dashboardPort`, default 7000) binds `127.0.0.1` only, so remote dashboard access goes through the gateway. The dashboard can approve, deny and destroy; don't expose it without a reverse proxy in front. + +## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain + +Agents poll `HIVE_FORGE_URL` for Forgejo notifications and run every `hive-forge` call against it. `nix/host-modules/hive-c0re/environment.nix` sets it to `http://` (default `forge.` — a swarm runs one forge). Agents run in a private netns and can't reach the host's loopback, so they resolve that name through the bridge dnsmasq to the bridge IP and reach nginx on port 80 (the bridge firewall opens 80 and 443) — the same path an operator's browser takes. + +## hive-forge container shape + +Private Forgejo in a nixos-container named `hive-forge` (not `h-*`, so c0re's lifecycle scanner leaves it alone; manage it with the standard `nixos-container` CLI). The container keeps it from colliding with any `services.forgejo` the operator already runs on the host: separate systemd namespace, separate state dir, separate port. + +It shares the host network namespace (`privateNetwork = false`): the container is for state and unit isolation, not network isolation. Agent containers, by contrast, are network-isolated and reach the forge through the gateway ([`HIVE_FORGE_URL`](#hive_forge_url-agents-reach-the-forge-via-the-gateway-by-domain)). + +State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/` and survives restarts and reboots; destroy the container to wipe it. + +Only the swarm's forge host runs it (`services.hyperhive.deploy.forgejo.enable`, see [`swarm/services.md`](../swarm/services.md)). Every other hive reaches that host's gateway by `forge.`. + +### Network and port configuration + +```nix +services.hyperhive.swarm.forge = { + httpPort = 3000; # default — outside hyperhive's 7000 / 8100-8999 + sshPort = 2222; # default — git-over-SSH, off the host's openssh on 22 +}; +# Which ports the forge answers on is swarm-wide; whether THIS host opens +# them in its firewall is a deployment decision, so it lives under deploy.*. +services.hyperhive.deploy.forgejo.openFirewall = false; # default +``` + +`sshPort` serves `git clone/push/pull` over SSH (`git@:owner/repo.git` with `-p 2222`). SSH goes straight to Forgejo, not through nginx. + +`openFirewall` (default `false`) opens `httpPort` and `sshPort` on the host. Agents don't need it — they come through the gateway. Set it for a browser reaching `http://:/` directly, or for external git clients pushing over SSH. The forge vhost behind the gateway (`deploy.forgejo.behindGateway`, default `true`) needs only the gateway's own `openFirewall`. + +### `rootUrl` override + +```nix +services.hyperhive.swarm.forge.rootUrl = "https://forge.example.com/"; +``` + +`rootUrl` (default `null`) overrides the Forgejo `ROOT_URL` derived from `forge.domain` and the gateway: + +| Shape | Derived `ROOT_URL` | +|---|---| +| `deploy.forgejo.behindGateway = true` | `https:///` (no port suffix when `gateway.httpsPort == 443`) | +| `deploy.forgejo.behindGateway = false` | `http://:/` | + +Set it when `forge.domain` differs from the public URL, or for a bespoke shape such as an external reverse proxy on another host or path. It must end with `/` (an assertion enforces this). + +## Security headers + +Every named vhost sets these at server scope (the `_` catch-all only closes connections): + +| Header | Value | +|--------|-------| +| `X-Frame-Options` | `SAMEORIGIN` | +| `X-Content-Type-Options` | `nosniff` | +| `Referrer-Policy` | `strict-origin-when-cross-origin` | + +nginx doesn't merge `add_header`: a `location` that sets its own (the CORS locations `/.well-known/matrix/client` and `/_matrix/`) inherits none of the server-scope headers, so those locations repeat them. Locations with no `add_header` of their own pick them up. + +### HSTS (`gateway.hsts`) + +Opt-in, off by default: + +```nix +services.hyperhive.gateway.hsts = { + enable = true; # default: false + maxAge = 31536000; # default: 1 year (required for preload list) + includeSubDomains = true; # default: true +}; +``` + +When enabled, every vhost adds `Strict-Transport-Security: max-age=…[; includeSubDomains]`. HSTS pins https in the browser: a deployment that later loses TLS locks browsers out until `max-age` expires, so enable it only when TLS is permanent. + +## Local dev (`localHostsEntry`) + +`services.hyperhive.gateway.localHostsEntry = true` maps to `127.0.0.1` in the host's `/etc/hosts`: + +- the hive domain; +- every name a module on this host contributes to `gateway.localNames` — each swarm service this host runs adds its own (`forge.` when behind the gateway, `chat.`, `auth.`, the swarm UI's apex, …). + +`services.hyperhive.deploy.singleHostSwarm` turns it on. Leave it off with real DNS. ## Discovery flow (matrix) -Operator points client at ``. Sequence: +The operator points a client at ``: -1. Client fetches `https:///.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat."}}` (no port suffix when gateway listens on 443). The gateway always terminates TLS, so the scheme is always `https`; a non-default `httpsPort` shows up as the port suffix. -2. Client connects to `chat./_matrix/client/...`. -3. Gateway routes `/_matrix/*` → tuwunel at `127.0.0.1:8008`. +1. The client fetches `https:///.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat."}}` (no port suffix when `httpsPort` is 443). +2. It connects to `chat./_matrix/client/...`. +3. The gateway proxies `/_matrix/*` to tuwunel at `127.0.0.1:8008`. -matrix-dart-sdk (FluffyChat etc.) hardcodes `https` for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path. +matrix-dart-sdk (FluffyChat and others) always fetches the well-known over `https`, which is why [self-signed TLS](#self-signed-tls-default) is on by default. - -Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.:"}` (the federation delegation always carries an explicit port, even the HTTPS default 443 — the https-implies-443 elision only applies to the client base_url above). Gateway only listens on configured `port` (+ `httpsPort` when TLS on); cross-hive federation needs either an SRV record (`_matrix._tcp.chat.` → port 80 / 443) OR `matrix.openFirewall = true` so peers reach tuwunel's federation port directly. Hyperhive is closed/internal in most deployments, so this rarely bites. - +Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.:"}`. The port is always explicit, even 443: a delegated host without a port means the federation default 8448, not 443. Peers then reach `/_matrix/` on the chat vhost through the gateway, so the gateway must be reachable from them (`gateway.openFirewall`). -## SPA fallback (Accept-header pattern) +⚠️ The gateway serves both `.well-known` routes on the **hive** vhost. Matrix looks them up at the `serverName`, which defaults to the bare swarm domain, and the swarm UI's apex vhost serves no `.well-known/matrix/*` route. -The per-agent UIs and the `chat.` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard vhost instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide: +## Sub-domain shape (rationale) + +Sub-domain for forge and matrix, sub-path for per-agent UIs: + +- forgejo's default `ROOT_URL = http:///` works without `X-Forwarded-Prefix` handling; sub-domain hosting is the canonical Forgejo shape. +- matrix deployments put the API listener on its own name, and federation expects that. `gatewayHost` defaults to `chat.` rather than the spec-conventional `matrix.`; set it explicitly for the conventional label. +- per-agent UIs are hyperhive-internal and base-path-aware for `/agent//`. A sub-domain per agent would multiply DNS and TLS cost for nothing. +- cookie and storage isolation: a forge XSS can't reach the dashboard session, because they're different origins. + +`services.hyperhive.swarm.forge.domain` and `services.hyperhive.swarm.matrix.gatewayHost` take a full hostname (`forge.darkest.space`, `git.example.com`), not a label glued to a domain. + +## Tuning knobs + +Per-vhost timeouts and body-size limits live in the location blocks: + +- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), websockets on. +- matrix `/_matrix/` (tuwunel): `client_max_body_size` = `deploy.matrix.maxRequestSize` + 1 MiB, so tuwunel stays the tighter limit and returns a matrix error a client can act on; `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*`, websockets on. +- per-agent `/agent//`: `proxy_read_timeout 1d` (long-lived SSE and WebSocket UIs), websockets on, `X-Forwarded-Prefix` set so the harness can build absolute URLs. + +## Internals + +How the gateway does what the sections above describe. Read before changing `hive-gateway/`, `gateway_nginx.rs` or `agent_sockets.rs`. + +### SPA fallback (Accept-header pattern) + +The per-agent UIs and the `chat.` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide: - hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap) - a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type Solution: an `nginx http`-context `map $http_accept $_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $_spa_target ` falls through to ``. No extension allowlist, no `if` block, no regex heuristics. -For matrix / per-agent static assets, `` is `=404` (a missing asset is just missing). +For matrix and per-agent static assets, `` is `=404` (a missing asset is just missing). -### Dashboard: path-based routing (not Accept-header) +#### Dashboard: path-based routing (not Accept-header) -Every hive-c0re route lives under `/api/` plus the single `/webhook/knowledge` endpoint, so the dashboard vhost routes by **path**, not Accept header — deterministic, unlike a content-type split where the same URL could resolve differently depending on the caller's `Accept` header: +hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header: -- `location /api/` → hive-c0re (`7000`): all dashboard data, actions/mutations, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). Carries `proxy_buffering off` + a 1d read timeout for the streams. -- `location /webhook/` → hive-c0re: the knowledge webhook. -- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html` (SPA fallback). +- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live. +- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded. +- `location /health/` → hive-c0re: liveness and readiness. +- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`. -Each location carries a duplicated `auth_basic` block (separate locations don't inherit it). This keeps the gateway static-serving the dashboard dist while hive-c0re stays API-only — a frontend-only change doesn't rebuild or restart the core daemon. A new top-level c0re route prefix (beyond `/api` + `/webhook`) needs a matching `location` added to the dashboard vhost. +The gateway static-serves the dist while hive-c0re stays API-only, so a frontend-only change doesn't restart the core daemon. A new top-level c0re route prefix needs a matching `location` on the hive vhost (`hive-gateway/vhosts.nix`). -## Local dev (`localHostsEntry`) +### Per-agent unix-socket upstream -`services.hyperhive.gateway.localHostsEntry = true` adds entries to the host's `/etc/hosts`: - -- `` → `127.0.0.1` -- `forge.` → `127.0.0.1` (when deploy.forgejo.behindGateway) -- `chat.` → `127.0.0.1` (when matrix.gatewayHost set) -- `auth.` → `127.0.0.1` (when deploy.authelia) - -`lib.unique` de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off. - -## Sub-domain shape (rationale) - -Operator decision: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs. - -- forgejo's default `ROOT_URL = http:///` works without any `X-Forwarded-Prefix` gymnastics — sub-domain hosting is the canonical Forgejo deploy shape. -- matrix-spec deployments universally use `matrix.` for the actual API listener — federation already expects this. (`gatewayHost`'s own default departs from that convention — `chat.`, not `matrix.` — since `serverName` is swarm-wide but `gatewayHost` is per-hive; operators who want the spec-conventional label can still set it explicitly.) -- per-agent UIs are hyperhive-internal and base-path-aware specifically for `/agent//`. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins. -- cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins. - -`services.hyperhive.{forge.domain,matrix.gatewayHost}` take the full hostname (`forge.darkest.space`, `git.example.com`) rather than a label that gets concatenated with hive-domain — operators want control over the full shape, not a forced `