Watch
0
0
Fork
You've already forked hyperhive
0

docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix

gateway.md: split the opener into what/audience/enable; vhost map in two
tables (swarm-service vhosts declared by their own modules, then the hive
vhost) matching vhosts.nix and the service modules; gateway.enable exists
and is set with mkDefault by the modules that need it; Basic auth scope,
dashboard /health/ prefix, error-page rendering, matrix body limit and
forge link source corrected; nginx internals grouped under one Internals
section with their headings unchanged.

network.md: gateway and dnsmasq run on the host, not in a container;
network.enable is set by the modules that need it; shared-netns firewall
rule covers every swarm service container; hive-priv writes the nspawn
conf; domain sentence rewritten; removed options moved into <details>.

jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS
show different graphs drawn by the same component.

observability.md: swarm tier first; history narration cut; network access
deduplicated into a link to network.md; options link made absolute.

matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and
SSO options under deploy.matrix; .well-known is served on the hive domain;
roadmap sentence deleted; stale hive-c0re provisioning claims fixed;
serverName upgrade note moved into <details>.

Refs #3902
This commit is contained in:
atlas 2026-10-01 23:31:46 +02:00
commit 067f4e5699
5 changed files with 433 additions and 633 deletions

View file

@ -1,103 +1,290 @@
# hive-gateway
This host's nginx fronts the hyperhive web surfaces running on it — next to hive-c0re, not in its own container: it shares the host netns anyway (see [Vhost map](#vhost-map) below), so containerizing it would buy no network isolation while costing a resolv.conf sync, a machine-bus reload, and three bind mounts. System-config (not meta-flake managed). Configured via `services.hyperhive.gateway.*` + per-subsystem opt-in flags in `services.hyperhive.{forge,matrix,...}`. the modules that need them assert `gateway.enable` and `gateway.dns.enable`, so a host serving a vhost or resolving hive names gets them without an opt-in.
Every host's nginx: the one front door for whatever this host serves. A swarm service running here (forge, matrix, SSO, the swarm UI, the metrics and log stores) declares its own vhost through the gateway; the gateway itself adds the hive's own surface — dashboard, per-agent UIs, matrix discovery.
_For the operator configuring `services.hyperhive.gateway.*` on a host._ nginx and the hive resolver (dnsmasq) run on the host next to hive-c0re, not in a container: they bind `:80`/`:443` and the bridge address, so a network namespace of their own would isolate nothing.
You rarely switch it on yourself. `gateway.enable` defaults to off, and every module that serves a vhost or needs hive names to resolve sets `gateway.enable` / `gateway.dns.enable` with `mkDefault true` — the hive controller, each swarm service, CI.
## Vhost map
| URL | vhost | upstream | source |
| --- | --- | --- | --- |
| `<hive>/` | `_` (catch-all) | dashboard dist (static, from `servedFrontend`); `/api/` + `/webhook/` → hive-c0re (`7000`) | always |
| `<hive>/agent/<name>/` | `_` | per-agent harness (UDS or TCP) | `agents.conf` (runtime-generated) |
| `<hive>/.well-known/matrix/{client,server}` | `_` | inline JSON (no upstream) | `matrix.enable && domain != null` |
| `<hive>/matrix/` (deprecated) | `_` | 301 → `chat.<swarm>/` | `matrix.gui.enable` |
| `forge.<swarm>/` | `forge.<swarm>` | forgejo (`3000`) | `deploy.forgejo.behindGateway` |
| `chat.<swarm>/_matrix/*` | `chat.<swarm>` | tuwunel (`8008`) | `matrix.gatewayHost != null` |
| `chat.<swarm>/` | `chat.<swarm>` | fluffychat-web static | `matrix.gui.enable` |
| `chat.<swarm>/config.json` | `chat.<swarm>` | inline JSON (FluffyChat boot config) | `matrix.gui.enable && domain != null` |
| `auth.<swarm>/` | `auth.<swarm>` | authelia (`9091`) | `deploy.authelia` |
| `<swarm>/` | `<swarm>` | swarm-ui dist (static), behind an authelia subrequest | `deploy.swarm-ui` |
**Swarm services.** Each module declares its vhost on the host that runs the service, under the swarm domain:
Only the host that **runs** authelia declares the authelia vhost, not every hive that uses it — a client hive knows the swarm's `authelia.url` but must not answer for a name it doesn't serve. Its server name is exactly `swarm.authelia.domain`: authelia validates `authelia_url ⊂ session cookie domain` at startup, so a near-miss is a container that refuses to boot. It carries no `auth_basic` — the login page must not sit behind the login mechanism it replaces — and sets the four `X-Forwarded-{Proto,Host,Uri,For}` headers, since authelia decides by the *original* request rather than the hop it sees.
| URL | upstream | declared by, when |
| --- | --- | --- |
| `<swarm>/` | swarm-ui dist (static), behind an authelia subrequest | `swarm-ui.nix`, `deploy.swarm-ui.enable` |
| `auth.<swarm>/` | authelia (`9091`) | `swarm-authelia.nix`, `deploy.authelia.enable` |
| `forge.<swarm>/` | forgejo (`3000`) | `hive-forge/`, `deploy.forgejo.behindGateway` |
| `chat.<swarm>/_matrix/*` | tuwunel (`8008`) | `hive-matrix.nix`, `swarm.matrix.gatewayHost != null` |
| `chat.<swarm>/` | fluffychat-web static (404 with the GUI off) | `hive-matrix.nix`, `deploy.matrix.gui.enable` |
| `chat.<swarm>/config.json` | inline JSON (FluffyChat boot config) | `hive-matrix.nix`, `deploy.matrix.gui.enable` |
| `grafana.<swarm>`, `metrics.<swarm>`, `logs.<swarm>`, `otel.<swarm>`, `bao.<swarm>` | the matching swarm service | that service's module → [`swarm/services.md`](../swarm/services.md) |
<!-- vale write-good.Passive = NO -->
⚠️ **A `502` from this vhost means authelia itself isn't answering, not that the proxy is misconfigured.** Check `journalctl -M swarm-authelia -u authelia-swarm` before suspecting anything here. A swarm with no users of its own answers with a login page that refuses everyone; the first account is created in [`swarm/sso.md`](../swarm/sso.md).
<!-- vale write-good.Passive = YES -->
**The hive's own vhost**, named for the hive domain:
Per-agent UIs stay sub-path, forge and matrix get sub-domains — see
[Sub-domain shape (rationale)](#sub-domain-shape-rationale) below for why.
| URL | upstream | when |
| --- | --- | --- |
| `<hive>/` | dashboard dist (static, from `servedFrontend`) | always |
| `<hive>/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always |
| `<hive>/api/docs/` | themed Swagger UI dist (static) | always |
| `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) |
| `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` |
| `<hive>/matrix/` (deprecated) | 301 → `chat.<swarm>/` | `deploy.matrix.gui.enable` and `gatewayHost` set |
The catch-all `_` vhost answers any other `Host` with `444` (connection closed, no response). It's `mkDefault`, so to make your own vhost the default server, set `services.nginx.virtualHosts."_".default = false;` — an eval assertion names both when two claim it.
Only the host that **runs** authelia declares `auth.<swarm>`; a hive that merely uses SSO knows `swarm.authelia.url` but doesn't answer for that name. The server name must be exactly `swarm.authelia.domain` — authelia checks that `authelia_url` sits inside its session cookie domain at startup, and refuses to boot otherwise. The vhost carries no Basic auth (that would put the login page behind the login it replaces) and passes `X-Forwarded-{Proto,Host,Uri,For}`, because authelia decides by the original request.
⚠️ `auth.<swarm>` answering with the **sso unavailable** page means authelia isn't answering at all — check `journalctl -M swarm-authelia -u authelia-swarm`. A swarm with no real accounts yet boots fine (a disabled placeholder user keeps authelia's user store non-empty) and serves a login page that refuses everyone; add the first account with `swarmctl user add …` ([`setup.md`](../getting-started/setup.md)).
Per-agent UIs stay sub-path; forge and matrix get sub-domains → [Sub-domain shape](#sub-domain-shape-rationale).
## TLS modes
The gateway always terminates TLS: there is no http-only mode. Which certificate it serves depends on what you configure:
| mode | config | cert source | `.well-known` scheme |
|---|---|---|---|
| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | the host's hive CA signs a gateway leaf (RSA-4096) | `https` |
| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` |
| operator cert | `tls.certDir` set | read from the operator's dir | `https` |
`tls.certDir` together with `tls.acme.enable` fails an assertion — pick one.
⚠️ One vhost class has no TLS: a vhost bound to loopback for a local consumer. `grafana-metrics` (`nix/host-modules/swarm-grafana.nix`) is the one in the tree: it listens on `127.0.0.1` only and serves `= /metrics` from grafana's unix socket for this host's collector. Everything with a routable name follows the table above.
### ACME / Let's Encrypt (`tls.acme`)
The simplest production path for a public domain:
```nix
services.hyperhive.gateway = {
openFirewall = true;
tls.acme = {
enable = true;
email = "admin@example.com"; # required
};
};
```
nginx obtains and renews certs via the HTTP-01 challenge on `port` (default 80); they land in `/var/lib/acme/`, managed by nixpkgs's `security.acme`. Every name this host serves needs a public DNS record pointing here — the hive domain and each swarm-service name in the [vhost map](#vhost-map) — and `openFirewall = true` so Let's Encrypt reaches `/.well-known/acme-challenge/`.
### Self-signed TLS (default)
On by default, listening on `httpsPort` (default 443) beside the plain-http `port` (default 80).
The issuer is a **host-held hive CA**, not a bare self-signed leaf. `hive-tls-ca.service` (from `hive-tls.nix`) generates a long-lived CA (`services.hyperhive.deploy.hive-controller.tls.caValidityDays`, default 7300 days) under `…tls.stateDir` (default `/var/lib/hive-tls`) and signs a gateway **leaf** with it (`leafValidityDays`, default 30). `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`).
⚠️ **Keep the import unit.** It does two jobs nginx needs: it re-modes the key to `0640 root:nginx` (nginx's pre-start `nginx -t` runs as the nginx user and fails on the CA's `0600 root:root` key), and it makes sure **every cert path the config names exists** — when the swarm-services leaf is missing it installs the hive leaf in its place. nginx refuses a config naming a missing cert file, so without that fallback one missing leaf takes down every vhost, not just one.
**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust, and nothing can wire a runtime-generated leaf into an agent's build-time trust store. Agents and federation peers trust the stable CA once; leaf rotation never breaks them.
**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is an intermediate under the swarm root ([`swarm/ca.md`](../swarm/ca.md) has the hierarchy), and a verifier can't stop at an intermediate — so the bundle carries the hive CA plus its root. nginx serves the leaf with the hive CA appended for the same reason. Agents (via `security.pki.certificateFiles`), the CI and forge containers and federating peers all read the bundle.
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) fetches `https://<host>/.well-known/matrix/client` and never falls back to plain http, so without TLS the browser client can't bootstrap.
**Cert shape**: the leaf's CN is the bare hive domain; its SANs are `<hive>` and `*.<hive>`. The hive CA is name-constrained to `<hive>`, so it can't sign a swarm-service name outside it — a violating SAN would invalidate the whole leaf. Those names get the swarm-services leaf instead ([`swarm/ca.md`](../swarm/ca.md)).
**Rotation**: `hive-tls-ca.service` re-signs the leaf when it's missing, within 30 days of expiry, or no longer covers the configured names, always under the same CA. It regenerates the CA only if missing or expired. To force a leaf rotation, delete `gateway.pem` under the state dir, restart the unit, then reload `nginx`.
**Cert prompts**: browsers warn once per host until you add the hive's `trust-bundle.pem` (the anchor, not the leaf) to the browser or OS trust store.
### Operator-provided cert (`tls.certDir`)
For a cert from a real CA (Let's Encrypt via your own `security.acme`, a corporate CA):
```nix
services.hyperhive.gateway = {
tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir
# tls.certName = "cert.pem"; # default — matches security.acme layout
# tls.keyName = "key.pem"; # default — matches security.acme layout
};
```
nginx reads the directory directly. Keep the key readable by nginx:
- `security.acme` writes keys `0640 root:acme`, which the `nginx` user can't read. Set `security.acme.certs."example.com".group = "nginx";` (or make the key `0644` if your threat model allows). Otherwise nginx fails at startup with the reason in the journal.
### Fronting with an external TLS terminator
The gateway has no plain-http upstream mode. Either give the gateway the real cert (`tls.certDir` or `tls.acme`) so it serves proper TLS itself, or front it over a unix socket rather than a plain-http TCP port. `.well-known/matrix/*` responses always advertise `https` ([Discovery flow](#discovery-flow-matrix)).
## HTTP Basic auth
Optional: Basic auth on the hive's dashboard. Add a login first, then enable it — the htpasswd file exists from first boot, and an empty one refuses everyone.
_On the hive host:_
```sh
hivectl gateway create-user alice --password-stdin # password on stdin
hivectl gateway delete-user bob
hivectl gateway list-users
```
```nix
services.hyperhive.gateway.auth = {
enable = true;
# realm = "hyperhive"; # default; must not contain `"` or `$`
};
```
`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history.
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first.
## Firewall posture (host-level)
`services.hyperhive.gateway.openFirewall = true` opens `port` and `httpsPort` on the host firewall — both, since the gateway always serves TLS. It defaults to off; set it for any reach from outside the host.
nginx is the one external entry point. The per-agent web-port range (`8100`–`8999`) stays closed: agents serve their UI on a unix socket ([Per-agent unix-socket upstream](#per-agent-unix-socket-upstream)), and the hashed TCP port (`lifecycle::agent_web_port`) is only a fallback bind for an agent missing `HIVE_WEB_SOCKET` — the gateway never proxies through it.
The dashboard port (`services.hyperhive.c0re.dashboardPort`, default 7000) binds `127.0.0.1` only, so remote dashboard access goes through the gateway. The dashboard can approve, deny and destroy; don't expose it without a reverse proxy in front.
## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain
Agents poll `HIVE_FORGE_URL` for Forgejo notifications and run every `hive-forge` call against it. `nix/host-modules/hive-c0re/environment.nix` sets it to `http://<swarm.forge.domain>` (default `forge.<swarm-domain>` — a swarm runs one forge). Agents run in a private netns and can't reach the host's loopback, so they resolve that name through the bridge dnsmasq to the bridge IP and reach nginx on port 80 (the bridge firewall opens 80 and 443) — the same path an operator's browser takes.
## hive-forge container shape
Private Forgejo in a nixos-container named `hive-forge` (not `h-*`, so c0re's lifecycle scanner leaves it alone; manage it with the standard `nixos-container` CLI). The container keeps it from colliding with any `services.forgejo` the operator already runs on the host: separate systemd namespace, separate state dir, separate port.
It shares the host network namespace (`privateNetwork = false`): the container is for state and unit isolation, not network isolation. Agent containers, by contrast, are network-isolated and reach the forge through the gateway ([`HIVE_FORGE_URL`](#hive_forge_url-agents-reach-the-forge-via-the-gateway-by-domain)).
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/` and survives restarts and reboots; destroy the container to wipe it.
Only the swarm's forge host runs it (`services.hyperhive.deploy.forgejo.enable`, see [`swarm/services.md`](../swarm/services.md)). Every other hive reaches that host's gateway by `forge.<swarm-domain>`.
### Network and port configuration
```nix
services.hyperhive.swarm.forge = {
httpPort = 3000; # default — outside hyperhive's 7000 / 8100-8999
sshPort = 2222; # default — git-over-SSH, off the host's openssh on 22
};
# Which ports the forge answers on is swarm-wide; whether THIS host opens
# them in its firewall is a deployment decision, so it lives under deploy.*.
services.hyperhive.deploy.forgejo.openFirewall = false; # default
```
`sshPort` serves `git clone/push/pull` over SSH (`git@<domain>:owner/repo.git` with `-p 2222`). SSH goes straight to Forgejo, not through nginx.
`openFirewall` (default `false`) opens `httpPort` and `sshPort` on the host. Agents don't need it — they come through the gateway. Set it for a browser reaching `http://<host>:<httpPort>/` directly, or for external git clients pushing over SSH. The forge vhost behind the gateway (`deploy.forgejo.behindGateway`, default `true`) needs only the gateway's own `openFirewall`.
### `rootUrl` override
```nix
services.hyperhive.swarm.forge.rootUrl = "https://forge.example.com/";
```
`rootUrl` (default `null`) overrides the Forgejo `ROOT_URL` derived from `forge.domain` and the gateway:
| Shape | Derived `ROOT_URL` |
|---|---|
| `deploy.forgejo.behindGateway = true` | `https://<forge.domain>/` (no port suffix when `gateway.httpsPort == 443`) |
| `deploy.forgejo.behindGateway = false` | `http://<forge.domain>:<httpPort>/` |
Set it when `forge.domain` differs from the public URL, or for a bespoke shape such as an external reverse proxy on another host or path. It must end with `/` (an assertion enforces this).
## Security headers
Every named vhost sets these at server scope (the `_` catch-all only closes connections):
| Header | Value |
|--------|-------|
| `X-Frame-Options` | `SAMEORIGIN` |
| `X-Content-Type-Options` | `nosniff` |
| `Referrer-Policy` | `strict-origin-when-cross-origin` |
nginx doesn't merge `add_header`: a `location` that sets its own (the CORS locations `/.well-known/matrix/client` and `/_matrix/`) inherits none of the server-scope headers, so those locations repeat them. Locations with no `add_header` of their own pick them up.
### HSTS (`gateway.hsts`)
Opt-in, off by default:
```nix
services.hyperhive.gateway.hsts = {
enable = true; # default: false
maxAge = 31536000; # default: 1 year (required for preload list)
includeSubDomains = true; # default: true
};
```
When enabled, every vhost adds `Strict-Transport-Security: max-age=…[; includeSubDomains]`. HSTS pins https in the browser: a deployment that later loses TLS locks browsers out until `max-age` expires, so enable it only when TLS is permanent.
## Local dev (`localHostsEntry`)
`services.hyperhive.gateway.localHostsEntry = true` maps to `127.0.0.1` in the host's `/etc/hosts`:
- the hive domain;
- every name a module on this host contributes to `gateway.localNames` — each swarm service this host runs adds its own (`forge.<swarm>` when behind the gateway, `chat.<swarm>`, `auth.<swarm>`, the swarm UI's apex, …).
`services.hyperhive.deploy.singleHostSwarm` turns it on. Leave it off with real DNS.
## Discovery flow (matrix)
Operator points client at `<hive>`. Sequence:
The operator points a client at `<hive>`:
1. Client fetches `https://<hive>/.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat.<swarm>"}}` (no port suffix when gateway listens on 443). The gateway always terminates TLS, so the scheme is always `https`; a non-default `httpsPort` shows up as the port suffix.
2. Client connects to `chat.<swarm>/_matrix/client/...`.
3. Gateway routes `/_matrix/*` → tuwunel at `127.0.0.1:8008`.
1. The client fetches `https://<hive>/.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat.<swarm>"}}` (no port suffix when `httpsPort` is 443).
2. It connects to `chat.<swarm>/_matrix/client/...`.
3. The gateway proxies `/_matrix/*` to tuwunel at `127.0.0.1:8008`.
matrix-dart-sdk (FluffyChat etc.) hardcodes `https` for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path.
matrix-dart-sdk (FluffyChat and others) always fetches the well-known over `https`, which is why [self-signed TLS](#self-signed-tls-default) is on by default.
<!-- vale write-good.Passive = NO -->
Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.<swarm>:<httpsPort>"}` (the federation delegation always carries an explicit port, even the HTTPS default 443 — the https-implies-443 elision only applies to the client base_url above). Gateway only listens on configured `port` (+ `httpsPort` when TLS on); cross-hive federation needs either an SRV record (`_matrix._tcp.chat.<swarm>` → port 80 / 443) OR `matrix.openFirewall = true` so peers reach tuwunel's federation port directly. Hyperhive is closed/internal in most deployments, so this rarely bites.
<!-- vale write-good.Passive = YES -->
Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.<swarm>:<httpsPort>"}`. The port is always explicit, even 443: a delegated host without a port means the federation default 8448, not 443. Peers then reach `/_matrix/` on the chat vhost through the gateway, so the gateway must be reachable from them (`gateway.openFirewall`).
## SPA fallback (Accept-header pattern)
⚠️ The gateway serves both `.well-known` routes on the **hive** vhost. Matrix looks them up at the `serverName`, which defaults to the bare swarm domain, and the swarm UI's apex vhost serves no `.well-known/matrix/*` route.
The per-agent UIs and the `chat.<swarm>` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard vhost instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide:
## Sub-domain shape (rationale)
Sub-domain for forge and matrix, sub-path for per-agent UIs:
- forgejo's default `ROOT_URL = http://<host>/` works without `X-Forwarded-Prefix` handling; sub-domain hosting is the canonical Forgejo shape.
- matrix deployments put the API listener on its own name, and federation expects that. `gatewayHost` defaults to `chat.<swarm-domain>` rather than the spec-conventional `matrix.<server_name>`; set it explicitly for the conventional label.
- per-agent UIs are hyperhive-internal and base-path-aware for `/agent/<name>/`. A sub-domain per agent would multiply DNS and TLS cost for nothing.
- cookie and storage isolation: a forge XSS can't reach the dashboard session, because they're different origins.
`services.hyperhive.swarm.forge.domain` and `services.hyperhive.swarm.matrix.gatewayHost` take a full hostname (`forge.darkest.space`, `git.example.com`), not a label glued to a domain.
## Tuning knobs
Per-vhost timeouts and body-size limits live in the location blocks:
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), websockets on.
- matrix `/_matrix/` (tuwunel): `client_max_body_size` = `deploy.matrix.maxRequestSize` + 1 MiB, so tuwunel stays the tighter limit and returns a matrix error a client can act on; `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*`, websockets on.
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE and WebSocket UIs), websockets on, `X-Forwarded-Prefix` set so the harness can build absolute URLs.
## Internals
How the gateway does what the sections above describe. Read before changing `hive-gateway/`, `gateway_nginx.rs` or `agent_sockets.rs`.
### SPA fallback (Accept-header pattern)
The per-agent UIs and the `chat.<swarm>` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide:
- hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap)
- a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type
Solution: an `nginx http`-context `map $http_accept $<name>_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $<name>_spa_target <final>` falls through to `<final>`. No extension allowlist, no `if` block, no regex heuristics.
For matrix / per-agent static assets, `<final>` is `=404` (a missing asset is just missing).
For matrix and per-agent static assets, `<final>` is `=404` (a missing asset is just missing).
### Dashboard: path-based routing (not Accept-header)
#### Dashboard: path-based routing (not Accept-header)
Every hive-c0re route lives under `/api/` plus the single `/webhook/knowledge` endpoint, so the dashboard vhost routes by **path**, not Accept header — deterministic, unlike a content-type split where the same URL could resolve differently depending on the caller's `Accept` header:
hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions/mutations, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). Carries `proxy_buffering off` + a 1d read timeout for the streams.
- `location /webhook/` → hive-c0re: the knowledge webhook.
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html` (SPA fallback).
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live.
- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded.
- `location /health/` → hive-c0re: liveness and readiness.
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`.
Each location carries a duplicated `auth_basic` block (separate locations don't inherit it). This keeps the gateway static-serving the dashboard dist while hive-c0re stays API-only — a frontend-only change doesn't rebuild or restart the core daemon. A new top-level c0re route prefix (beyond `/api` + `/webhook`) needs a matching `location` added to the dashboard vhost.
The gateway static-serves the dist while hive-c0re stays API-only, so a frontend-only change doesn't restart the core daemon. A new top-level c0re route prefix needs a matching `location` on the hive vhost (`hive-gateway/vhosts.nix`).
## Local dev (`localHostsEntry`)
### Per-agent unix-socket upstream
`services.hyperhive.gateway.localHostsEntry = true` adds entries to the host's `/etc/hosts`:
- `<hive-domain>` → `127.0.0.1`
- `forge.<swarm>` → `127.0.0.1` (when deploy.forgejo.behindGateway)
- `chat.<swarm>` → `127.0.0.1` (when matrix.gatewayHost set)
- `auth.<swarm>` → `127.0.0.1` (when deploy.authelia)
`lib.unique` de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
## Sub-domain shape (rationale)
Operator decision: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
- forgejo's default `ROOT_URL = http://<host>/` works without any `X-Forwarded-Prefix` gymnastics — sub-domain hosting is the canonical Forgejo deploy shape.
- matrix-spec deployments universally use `matrix.<server_name>` for the actual API listener — federation already expects this. (`gatewayHost`'s own default departs from that convention — `chat.<swarm-domain>`, not `matrix.<server_name>` — since `serverName` is swarm-wide but `gatewayHost` is per-hive; operators who want the spec-conventional label can still set it explicitly.)
- per-agent UIs are hyperhive-internal and base-path-aware specifically for `/agent/<name>/`. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins.
- cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
`services.hyperhive.{forge.domain,matrix.gatewayHost}` take the full hostname (`forge.darkest.space`, `git.example.com`) rather than a label that gets concatenated with hive-domain — operators want control over the full shape, not a forced `<label>.<hive-domain>` pattern.
## Tuning knobs
Per-vhost timeouts + body-size limits live in the location blocks:
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), `proxyWebsockets = true` (live-update endpoints).
- matrix `/_matrix/` (tuwunel): `client_max_body_size 50M` (media uploads), `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*` (federation + cross-origin clients), `proxyWebsockets = true`.
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE / WebSocket dashboards), `proxyWebsockets = true`, `X-Forwarded-Prefix` set so the harness can build absolute URLs when relative isn't enough.
SSH for forge stays direct on `services.hyperhive.swarm.forge.sshPort` — separate listener protocol, not HTTP-over-nginx.
## Per-agent unix-socket upstream
All agents bind their web UI on a unix-domain socket at
`/run/hive-agent/<name>/web.sock` — the `HIVE_WEB_SOCKET` env var is
now set unconditionally for every agent. The mechanism:
Every agent binds its web UI on a unix-domain socket at
`/run/hive-agent/<name>/web.sock`. The mechanism:
1. **Agent side**. The nix module sets `HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock`
on every harness service env; `web_ui::serve` binds a
@ -106,7 +293,7 @@ now set unconditionally for every agent. The mechanism:
(`/run/hive-agent/<name>/`) into the agent's container. Dir
bind, not file bind — file bind-mounts don't survive the
harness's `unlink + bind(2)` cycle on socket replace. Per-agent
subdir keeps each agent's container blind to siblings' sockets.
subdir keeps each agent's container from seeing siblings' sockets.
The dir is `0751`, owned by the agent's container uid/gid (hive-priv
creates it `0751 root` before each start when missing, and the
@ -115,11 +302,11 @@ now set unconditionally for every agent. The mechanism:
own `0666`. The gateway is one of three principals sharing that dir
and doesn't own its ownership rules — see
[`docs/trust-boundary/boundary.md`](../trust-boundary/boundary.md#the-per-agent-socket-dir).
3. **Marker gate**. After successful `bind_unix`, the harness drops
3. **Marker gate**. After a successful `bind_unix`, the harness drops
`<dir>/hyperhive-socket-bound` next to the socket. c0re's
`agent_sockets::write` filters its JSON map by marker presence —
only agents whose harness has actually bound the socket appear there.
(Legacy name `.bound` also accepted during the transition window.)
only agents whose harness has bound the socket appear there.
It also accepts the older `.bound` name.
4. **Gateway side**. `gateway_nginx::write` generates
`/var/lib/hive-gateway/conf/agents.conf` — a plain nginx include
file with one `location /agent/<name>/` block per agent. Always
@ -152,275 +339,13 @@ on subsequent poll ticks.
`agents.conf` uses atomic `<path>.tmp` + `rename()` writes so a crashing
c0re process never leaves a partial or unparseable file behind.
## Dashboard link shape (gateway vs direct)
### Per-agent static frontend split
When the gateway is in front, the SW4RM tab builds per-agent links
as same-origin `/agent/<name>/…` URLs instead of the legacy direct
`http://<host>:<container.port>/` TCP shape. The signal comes from
`StateSnapshot.gateway_enabled`, sourced from the
`HIVE_GATEWAY_ENABLED` env the c0re NixOS module now always sets
(`services.hyperhive.gateway.enable` no longer exists — the gateway runs
unconditionally alongside hyperhive), so this is effectively always
true; the `false` branch stays as a defensive fallback for the
env being unset. Three render sites
flip together: the primary agent-name link, the favicon fetch
(`<url>/icon`), and the nav-strip `container`-kind links from
`DashboardState.links` (`GET /api/dashboard-state`). `forge`-kind nav-strip links still
resolve against `http://<host>:3000` (separate sub-domain transition
tracked by `deploy.forgejo.behindGateway`); `external`-kind links are
already absolute. See `docs/web-ui/dashboard.md::Container row` for the
frontend-side derivation.
## TLS modes
The gateway always terminates TLS — self-signed is the implicit floor when
the operator configures nothing else, so there is no http-only mode. Three
modes, selected by which (if any) external TLS source the operator sets:
⚠️ **One vhost class is exempt, so expect it during a TLS audit.** A vhost bound to loopback for a
local consumer carries a single plain-HTTP listen and no TLS. `grafana-metrics`
(`nix/host-modules/swarm-grafana.nix`) is the only one in the tree today: it listens on `127.0.0.1`
alone and serves one `= /metrics` location from grafana's unix socket, for the collector on this
host to scrape. Nothing off-host can reach it, so TLS there protects nothing. The modes below
cover every vhost with a routable name.
| mode | config | cert source | `.well-known` scheme |
|---|---|---|---|
| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | host hive-CA signs a gateway leaf (RSA-4096) | `https` |
| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` |
| operator cert | `tls.certDir` set | read from the operator's dir | `https` |
The `gateway.selfSignedTls` option has been **removed** — self-signed
is now derived from the absence of `tls.certDir` / `tls.acme`. A config
that still sets it fails eval with a removal message; use `tls.certDir`
/ `tls.acme` to override the default.
### ACME / Let's Encrypt (`tls.acme`)
Simplest production path for operators with a public domain:
```nix
services.hyperhive.gateway = {
openFirewall = true;
tls.acme = {
enable = true;
email = "admin@example.com";
};
};
```
nginx obtains and autorenews certs via the ACME HTTP-01 challenge on `port` (default 80). Certs land in `/var/lib/acme/` on the host, managed by nixpkgs's `security.acme` in the ordinary way.
**Requirements**: `services.hyperhive.domain` must be publicly DNS-resolvable to this host, and `openFirewall = true` so Let's Encrypt can reach `/.well-known/acme-challenge/`. Each active vhost (main domain, `forge.<swarm-domain>`, `chat.<swarm-domain>`) gets its own cert via separate ACME challenges — the swarm services default to names under `services.hyperhive.swarm.domain`, so **every one of those names must resolve to this host too**, not just the hive's own.
Mutual exclusion: `tls.certDir` set together with `tls.acme.enable = true` fails an assertion — pick one external TLS source (or neither, for the self-signed default).
### Self-signed TLS (default)
On by default, and listens on `httpsPort` (default 443) on every routable vhost beside the plain-http `port` (default 80). (See the loopback-only exception under *TLS modes* above.)
The issuer is a **host-held hive CA**, not a bare self-signed leaf. A host service (`hive-tls-ca.service`, from the `hive-tls` module) generates a long-lived CA (`services.hyperhive.deploy.hive-controller.tls.caValidityDays`, default ~20y) under `services.hyperhive.deploy.hive-controller.tls.stateDir` (default `/var/lib/hive-tls`), then signs a gateway **leaf** (`leafValidityDays`, default 30d) with it. `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`).
⚠️ **Don't collapse that import unit into pointing nginx at the CA dir.**
It does two jobs, and skipping it has taken the gateway down in production
before. It re-modes the leaf (`hive-tls-ca` writes the key `0600
root:root`; nginx's pre-start `nginx -t` runs as the *nginx user*, so a
`0600` key fails the config test and blocks the unit), and it guarantees
**every cert path the nginx config names exists** — which is what the
swarm-services fallback below is for.
**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust — and nothing can wire a runtime-generated leaf into an agent's build-time trust store at all. With a stable CA, agents and federation peers trust it *once*; leaf rotation never re-breaks them.
**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is itself issued under the swarm root ([`swarm/ca.md`](../swarm/ca.md) has the hierarchy), and an intermediate isn't a chain a verifier can terminate at — so the bundle carries the hive CA plus whatever it's rooted at. `hive-gateway-self-signed-cert` hands nginx the leaf with the hive CA appended for the same reason. Everything that trusts the hive's TLS reads the bundle: agents (via `security.pki.certificateFiles`), the CI and forge containers, and a federating peer.
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) hardcodes `https://<host>/.well-known/matrix/client` for homeserver discovery and refuses to fall back to plain http. Without TLS the browser client can't bootstrap.
**Cert shape**: leaf subject CN = bare hive domain; subjectAltName is `<hive>` plus wildcard `*.<hive>`, so all current and future sub-domain vhosts validate under the same leaf + the hive CA. You can't add a swarm service whose name is *not* under this hive's domain here — the hive CA is name-constrained to `<hive>`, and a violating SAN invalidates the whole leaf, not just that name. Those names get the swarm-services leaf instead ([`swarm/ca.md`](../swarm/ca.md)).
<!-- vale write-good.Passive = NO -->
**Rotation**: `hive-tls-ca.service` is idempotent — it re-signs the leaf when it's missing or within 30 days of expiry, always under the same CA (so consumer trust is undisturbed). It regenerates the CA itself only if missing or already expired. To force a leaf rotation, delete `gateway.pem` under the state dir and restart the unit, then reload `nginx`.
<!-- vale write-good.Passive = YES -->
**Cert prompts**: browsers still warn once per host until the operator adds the hive's `trust-bundle.pem` to the browser/OS trust store (an anchor, not the leaf, is the thing to trust). A separate mechanism wires agent trust (see the agent-trust work for `/run/hive-ca`).
### Operator-provided cert (`tls.certDir`)
For operators with a real CA cert (Let's Encrypt, corporate CA, etc.):
```nix
services.hyperhive.gateway = {
tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir
# tls.certName = "cert.pem"; # default — matches security.acme layout
# tls.keyName = "key.pem"; # default — matches security.acme layout
};
```
nginx reads the directory directly and uses `cert.pem` + `key.pem` (override `tls.certName`/`tls.keyName` for different filenames). Both modes listen on `httpsPort` (default 443) and emit `https://` in `.well-known` responses.
`tls.certDir` and `tls.acme.enable` set together is an assertion error.
**Key file permissions**: nixpkgs's `security.acme` outputs private keys as `0640 root:acme` by default. nginx runs as the `nginx` user and can't read a key with that ownership. Fix with:
```nix
security.acme.certs."example.com".group = "nginx";
```
or make the key world-readable (`0644`) if your threat model allows it. nginx errors out at startup on a key it can't read — the error is explicit in the journal, not a silent failure.
### Fronting with an external TLS terminator
No http-only mode exists (see [TLS modes](#tls-modes) above). Two paths
for an operator who wants their own TLS terminator:
- give the gateway the real cert via `tls.certDir` (or `tls.acme`) so it
serves proper TLS directly — no separate proxy needed; or
- front it over a **unix socket** rather than a plain-http TCP port (the
intended direction for "bring your own proxy" — the gateway isn't meant
to expose an unencrypted TCP upstream).
Because of this, `.well-known/matrix/{client,server}` discovery responses
always advertise `https` (see [Discovery flow](#discovery-flow-matrix) above).
## Firewall posture (host-level)
The gateway is unconditional — `services.hyperhive.gateway.enable` no
longer exists, there is no gateway-off mode. nginx is always the sole
external entry point and routes to agents over the UDS upstream
described above (see [Per-agent unix-socket
upstream](#per-agent-unix-socket-upstream)), so the per-agent web-port
range `8100..8999` stays closed on the host firewall
unconditionally — opening it would defeat the single-front-door story.
The hashed TCP port (`lifecycle::agent_web_port`) still exists as a
fallback bind for an agent whose `HIVE_WEB_SOCKET` env somehow ends up
unset, but nothing opens a matching firewall hole for it and the
gateway itself never proxies through it.
`services.hyperhive.gateway.openFirewall = true` opens both `port` and
`httpsPort` — both are always served, since the gateway always terminates
TLS (see [TLS modes](#tls-modes) above).
Every agent hashes into the same port range (no special case), so
one range opening covers every container.
<!-- vale write-good.Passive = NO -->
The dashboard port (`services.hyperhive.c0re.dashboardPort`, default 7000) is *not*
listed in either case — it binds `127.0.0.1` only, so a firewall
hole would be a no-op. Remote dashboard access flows through the
gateway. Operators who opt out of the gateway lose external
dashboard reach by design — the surface is privileged (approve /
deny / destroy), and operators must not expose it without a real reverse
proxy in front.
<!-- vale write-good.Passive = YES -->
## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain
Agents poll `HIVE_FORGE_URL` for Forgejo notifications + run all
`hive-forge` calls against it. Network isolation is always on (the
shared-netns mode no longer exists), so agents run in a private netns and
can never reach the host's loopback.
`nix/host-modules/hive-c0re/environment.nix` sets `HIVE_FORGE_URL` to
`http://<forge.domain>` (default `forge.<swarm-domain>` — a swarm runs
one forge; you must set `services.hyperhive.domain`). Agents
get the bridge dnsmasq as their resolver, resolve the hostname →
bridge IP, then reach nginx on port 80 (the bridge firewall opens
80+443). nginx proxies to forgejo — the same path an operator browser
takes, no raw port exposure needed.
## hive-forge container shape
Private Forgejo wrapped in a nixos-container (`hive-forge`, not
`h-*` — keeps c0re's lifecycle scanner out of the picture; the
operator manages it via the standard `nixos-container` CLI). The
container also keeps hive-forge from fighting any `services.forgejo`
the operator already runs on the host — separate systemd namespace,
separate state dir, separate port unless the operator deliberately
collides.
The forge container shares the host network namespace
(`privateNetwork = false`), so forgejo's listeners look like a
host-side service — nixos-container is here for state + systemd-unit
isolation, not network isolation. Note this is the FORGE container;
agent containers are network-isolated and reach the forge through the
gateway by `forge.<swarm-domain>` (see `HIVE_FORGE_URL` above), not via
the host's loopback.
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/`
and survives container restart / host reboot. To wipe, destroy the
container.
The container runs only on the swarm's forge host
(`services.hyperhive.deploy.forgejo.enable`, see
[`../swarm/services.md`](../swarm/services.md)). Every other hive
reaches that host's gateway by `forge.<swarm-domain>`.
### Network and port configuration
```nix
services.hyperhive.swarm.forge = {
httpPort = 3000; # default — HTTP listener; outside hyperhive's 7000/8100-8999 range
sshPort = 2222; # default — git-over-SSH; kept off 22 so it doesn't collide with the host openssh
};
# Which ports the forge answers on is swarm-wide; whether THIS host opens
# them in its firewall is a deployment decision, so it lives under deploy.*.
services.hyperhive.deploy.forgejo.openFirewall = false; # default
```
`httpPort` (default **3000**) is the port Forgejo's HTTP server binds to.
It sits outside hyperhive's reserved ranges (dashboard 7000,
agents 8100–8999) so a default install has no port fights. Change it
only if you already have another process bound to 3000.
`sshPort` (default **2222**) is the port Forgejo's built-in SSH server
uses for `git clone/push/pull` over SSH (`git@<domain>:owner/repo.git`
via `-p 2222`). Port 22 stays alone on the host for openssh.
<!-- vale write-good.Passive = NO -->
`openFirewall` (default **false**) controls whether the host firewall
opens `httpPort` and `sshPort`. Off by default (secure by
default): agents reach Forgejo through the gateway (`forge.<swarm-domain>` on
the bridge), not the raw port, so no firewall hole is needed. Flip to
`true` when you need:
- The operator's browser to reach `http://<host>:<httpPort>/` directly
(not behind the gateway).
- External git clients that push/pull via SSH directly to the host.
<!-- vale write-good.Passive = YES -->
Forgejo served through the gateway (`deploy.forgejo.behindGateway = true`) does
not need `openFirewall` — the gateway's own `openFirewall` option covers
that path.
### `rootUrl` override
```nix
services.hyperhive.swarm.forge.rootUrl = "https://forge.example.com/";
```
`rootUrl` (default **null**) overrides the Forgejo `ROOT_URL` that's
autoderived from `forge.domain` + gateway state. The autoderivation
covers most cases:
| Shape | Autoderived `ROOT_URL` |
|---|---|
| `deploy.forgejo.behindGateway = true` | `https://<forge.domain>/` (port suffix omitted when `gateway.httpsPort == 443`) |
| `deploy.forgejo.behindGateway = false` | `http://<forge.domain>:<httpPort>/` |
The gateway always terminates TLS, so the `behindGateway = true` case is
always advertised over `https://`; only the direct (`behindGateway =
false`) shape stays `http://`. Set `rootUrl` explicitly when
`forge.domain` resolves differently from the public URL, or for a
genuinely bespoke shape (for example an external reverse proxy on a different
host/path). Must end with `/` (Forgejo requirement; an assertion
enforces this).
## Per-agent static frontend split
hive-c0re injects `HIVE_AGENT_FRONTEND_DIR` into its service
environment, pointing at the `agent/` subdirectory of the frontend dist
set by `services.hyperhive.c0re.frontend`. That option has a default, so
the variable is always set.
hive-c0re's environment sets `HIVE_AGENT_FRONTEND_DIR` to the `agent/`
subdirectory of `services.hyperhive.c0re.servedFrontend`.
The nginx include generator (`gateway_nginx::write`) reads
this variable and, when set, emits split location blocks per agent
instead of the legacy single-proxy block.
instead of a single proxy block.
**Location priority for `/agent/<name>/...`:**
@ -461,8 +386,7 @@ checking file existence):
| `/agent/iris/events/live` | no file match → `@iris_dynamic` | proxied (SSE) |
Adding a new HTML page to the frontend dist (`dist/<page>.html`)
automatically makes it reachable at `/agent/<name>/<page>` — no
generator change needed.
makes it reachable at `/agent/<name>/<page>` with no generator change.
**Why `^~` for `/static/`**: the `^~` prefix gives this block higher
priority than the plain prefix `location /agent/<name>/`, so compiled
@ -475,169 +399,59 @@ store path baked in at hive-c0re build time, and c0re (writing
`agents.conf`) and nginx (serving files from it) are on the same machine,
so they see the same store.
**Graceful degradation**: if `HIVE_AGENT_FRONTEND_DIR` is empty or
unset (for example a build that predates `services.hyperhive.c0re.frontend`), each agent gets the
legacy single-proxy block and nginx forwards all traffic to the agent
daemon as before.
**Unset variable**: if `HIVE_AGENT_FRONTEND_DIR` is empty or unset, each
agent gets a single proxy block and nginx forwards all traffic to the
agent daemon.
**`extraFiles`**: per-agent `services.hyperhive.agent.frontend.extraFiles` are in
**`extraFiles`**: per-agent `services.hyperhive.agent.frontend.extraFiles` sit in
`mergedDist`, not in the base `services.hyperhive.c0re.frontend` dist. They're not under
the nix-store `alias` path, so requests for them fall through
`try_files` to `@<name>_dynamic`, and the agent daemon serves them
as before.
`try_files` to `@<name>_dynamic`, and the agent daemon serves them.
## Per-agent error pages
### Per-agent error pages
`/agent/<name>/` requests hit two failure modes; both get static
HTML pages instead of nginx's default error chrome:
- **Agent not found** (`/agent/<unknown>/...`) — name isn't in
`agentPortsTable`. nginx's prefix match falls back to the bare
- **Agent not found** (`/agent/<unknown>/...`) — no `location` for that
name in `agents.conf`. nginx's prefix match falls back to the bare
`/agent/` catch-all, which `return 404`s and `error_page 404` rewrites
to `/__hive_agent_not_found` → serves `not-found.html` with a link
back to the dashboard.
- **Agent unreachable** (`502 / 503 / 504` from `proxy_pass`) — the
per-agent harness isn't responding (container restarting, crash
recovery, etc.). `proxy_intercept_errors on` + `error_page 502 503
recovery). `proxy_intercept_errors on` + `error_page 502 503
504 = /__hive_agent_unreachable` rewrites to `unreachable.html`.
`pkgs.runCommand` builds both pages at deploy time (one nix
derivation `hyperhive-agent-error-pages` with `not-found.html` +
`unreachable.html` inside), served via two `internal` nginx
locations with `alias` to the exact file. `internal` keeps the
files from being directly request-able by operators — only nginx's
own error-handling can reach them.
`hive-gateway/error-pages.nix` renders every page (`notFound`, `unreachable`,
`unauthorized`, `ssoUnavailable`) from one template with `pkgs.writeText`;
nginx serves each through an `internal` location with `alias` to the file,
so only nginx's own error handling can reach them. They depend on nothing
from the frontend dist and render even when hive-c0re is down. A service
module reaches them through `gateway.lib.errorPages`.
Page styling: minimal inline CSS matching the dashboard's catppuccin
palette (`#1e1e2e` bg, `#cdd6f4` text, `#cba6f7` heading). No
dependencies on the frontend dist — these pages render even when
hive-c0re itself is down.
A route earns a custom page when the default status code would blame the
wrong component. The per-agent routes qualify (a 502 there means the
harness is restarting, not a gateway fault), and so does
`auth.<swarm>` — a dead authelia upstream almost always means no users
yet, and a bare 502 blames the proxy, the one part that works.
Forge, matrix and fluffychat keep nginx's defaults: a dead upstream there
means what the status code says.
<!-- vale write-good.Passive = NO -->
Scope is intentionally narrow: a route earns a custom page when the
default status code would point at the wrong component. The per-agent
routes qualify (a 502 there means the harness is restarting, not that
the gateway is broken), and so does `auth.<swarm>` — a dead authelia
upstream almost always means the user store was never bootstrapped, and
a bare 502 blames the proxy, which is the one part that's working.
<!-- vale write-good.Passive = YES -->
### Dashboard link shape (gateway vs direct)
Forge / matrix / fluffychat still get nginx defaults: their upstreams
being down means what the status code says, so a themed page would add
styling and no information.
The dashboard builds per-agent links as same-origin `/agent/<name>/…` URLs
when `StateSnapshot.gateway_enabled` is true, and direct
`http://<host>:<port>/` links otherwise. The flag comes from
`HIVE_GATEWAY_ENABLED`, which `hive-c0re/environment.nix` sets to `1` on
every hive, so the direct shape is a fallback for the variable being unset.
Three render sites follow it: the agent-name link, the favicon fetch
(`<url>/icon`), and the nav-strip `container`-kind links from
`DashboardState.links` (`GET /api/dashboard-state`). Forge links come from
`swarm.forge.publicUrl`; with it unset the dashboard hides them rather than
guess. See `docs/web-ui/dashboard.md::Container row` for the frontend side.
## HTTP Basic auth
<!-- vale write-good.Passive = NO -->
`services.hyperhive.gateway.auth.enable = true` gates every request to
the main vhost (`_`) behind HTTP Basic auth. nginx's built-in `auth_basic`
module validates credentials; no extra service or host-side daemon is
required.
<!-- vale write-good.Passive = YES -->
**Setup:**
```nix
services.hyperhive.gateway.auth = {
enable = true;
# realm = "hyperhive"; # optional, default shown
};
```
<!-- vale write-good.Passive = NO -->
The credential store lives at the fixed path
`/var/lib/hive-gateway/conf/gateway.htpasswd` on the host. A tmpfiles
rule pre-creates the file on first boot; no manual path configuration
is required. nginx reads it at that path directly.
<!-- vale write-good.Passive = YES -->
Manage users with `hivectl gateway`. `hivectl` sends the request over the
host admin socket and the `hive-c0re` daemon performs the write at its
canonical path — the daemon never exposes a path to the CLI:
```sh
# Add or update a user (prompted for password):
hivectl gateway create-user alice --password-stdin
# Add with inline password (visible in shell history — avoid for sensitive creds):
hivectl gateway create-user bob --password hunter2
# Remove a user:
hivectl gateway delete-user bob
# List current usernames:
hivectl gateway list-users
```
<!-- vale write-good.Passive = NO -->
The daemon hashes passwords with BCrypt (cost 12) and writes
`$2y$`-prefixed hashes that nginx accepts natively. No external
`htpasswd` binary is required.
<!-- vale write-good.Passive = YES -->
**What's not gated:** per-agent UI routes emitted into `agents.conf`
(served under `/agent/<name>/`) inherit no auth from `/` — nginx
applies `auth_basic` per-location. Full per-agent coverage is a
follow-up.
**Realm:** the `WWW-Authenticate: Basic realm="..."` string browsers
display in the credential dialog. Defaults to `"hyperhive"`. Must not
contain `"` or `$`.
**Custom 401 page:** when credentials are absent or wrong, nginx serves
a Catppuccin-styled `unauthorized.html` page (built into the same Nix
derivation as the agent error pages) that tells the operator which
`hivectl` command to run to create a user. The response status is still
`401` (`error_page 401 =401 /__hive_auth_unauthorized`) so browsers
present the login dialog on the first visit — users who dismiss the
dialog see the human-readable hint. The internal exact-match location
(`= /__hive_auth_unauthorized`) beats `location /` in nginx's prefix
ordering, preventing the subrequest from looping back through
`auth_basic`.
## Security headers
The gateway emits the following headers at server scope on every
vhost (`_`, `forge.<swarm-domain>`, `chat.<swarm-domain>`):
| Header | Value |
|--------|-------|
| `X-Frame-Options` | `SAMEORIGIN` |
| `X-Content-Type-Options` | `nosniff` |
| `Referrer-Policy` | `strict-origin-when-cross-origin` |
nginx's `add_header` inheritance rule: a `location` block that sets its
own `add_header` does **not** inherit server-scope headers. API locations
that carry their own CORS headers (for example `/.well-known/matrix/client`,
`/_matrix/`) are therefore unaffected. HTML-serving and proxy locations
with no `add_header` of their own pick the security headers up
automatically.
### HSTS (`gateway.hsts`)
HSTS is **opt-in** and disabled by default:
```nix
services.hyperhive.gateway.hsts = {
enable = true; # default: false
maxAge = 31536000; # default: 1 year (required for preload list)
includeSubDomains = true; # default: true
};
```
When enabled, the gateway adds a `Strict-Transport-Security: max-age=...[; includeSubDomains]`
header alongside the other security headers.
**Opt-in rationale**: HSTS pins HTTPS in the browser's preload cache;
enabling it on a deployment that later loses TLS locks browsers out
until `max-age` expires. Only enable when TLS is permanent.
Since the gateway always terminates TLS (see [TLS modes](#tls-modes)
above), an enabled HSTS header is always served over https — there is no
TLS-less mode that could violate it.
## Dialing another vhost by name (`verifiedProxyTo`)
### Dialing another vhost by name (`verifiedProxyTo`)
<!-- vale write-good.Passive = NO -->
`vhost-lib.nix`'s `verifiedProxyTo` builds the `proxy_ssl_*` /
@ -682,4 +496,3 @@ recursing on its own `auth_request` until nginx's subrequest-depth limit
turns it into a plain 500. Every call site sets `recommendedProxySettings
= false` on the location for exactly this reason — nixpkgs' version
would still clobber this one.

View file

@ -1,14 +1,22 @@
# hive-network
Host-side bridge + per-agent private-netns isolation, up on a host
where something attaches to it (`services.hyperhive.network.enable`,
asserted by the modules that need it rather than set by hand).
Configured via `services.hyperhive.network.*`.
The bridge every container on a host attaches to, and the private-netns
isolation of agent and CI containers behind it. It comes up on any host that runs agents or swarm services:
`services.hyperhive.network.enable` defaults to off, and the modules that
need the bridge set it with `mkDefault true` — the hive controller, CI,
the hive collector, and the gateway's resolver, which every swarm service
turns on. Configured via `services.hyperhive.network.*`.
> Isolation is the only mode — there is no shared-netns fallback. The
> former `services.hyperhive.network.isolateContainers` and
> `services.hyperhive.network.upstreamDns` options no longer exist; a
> config that still sets one fails eval with a removal message.
Isolation is the only mode; agent containers never share the host netns.
<details><summary>Upgrading a config that sets isolateContainers or upstreamDns</summary>
Neither option exists. A config that still sets
`services.hyperhive.network.isolateContainers` or
`services.hyperhive.network.upstreamDns` fails eval with a removal
message; drop the line.
</details>
## Network map
@ -45,7 +53,7 @@ that touches it.
| container | netns | IPv4 | listens / reached via |
| -------------- | ----------------------- | -------------- | ----------------------------------------------------------------------------------------------- |
| `hive-gateway` | host (shared) | host addresses | nginx `:80`/`:443` (every vhost); dnsmasq `bridgeIp:53` + DHCP `:67` on the bridge |
| gateway (host) | host | host addresses | nginx `:80`/`:443` (every vhost); dnsmasq `bridgeIp:53` + DHCP `:67` on the bridge |
| `hive-forge` | host (shared) | host addresses | forgejo `:3000` http, `:2222` git-ssh; fronted by the `forge.<swarm-domain>` vhost |
| `hive-matrix` | host (shared) | host addresses | tuwunel `:8008` (+ optional federation port); fronted by the matrix vhost |
| `hive-ci` | private, veth on bridge | DHCP pool | outbound only (runner → forge); no inbound surface |
@ -79,11 +87,9 @@ The flows, end to end:
## Container shape (where dnsmasq lives)
Co-located in the existing `hive-gateway` container — single
front-door for both DNS and HTTP, saves a sibling container, single
systemd-unit / state surface to monitor. The gateway shares host
netns (`privateNetwork = false`) so dnsmasq's `bind-interfaces`
listener on `bridgeIp` is on the host's bridge interface.
dnsmasq runs on the host itself, next to nginx — both come from the
gateway module, one front door for DNS and HTTP. Its `bind-interfaces`
listener sits on `bridgeIp`, on the host's bridge interface.
## Configuration
@ -99,10 +105,11 @@ listener on `bridgeIp` is on the host's bridge interface.
}
```
You must set `services.hyperhive.domain` — the dnsmasq resolver
is authoritative for `<hive-domain>` and its sub-domains. You don't
write it: it's read from this hive's entry in the swarm directory
(`docs/swarm/README.md` § Hive identity config).
The hive domain (`services.hyperhive.domain`, which the resolver answers
for) comes from this hive's entry in the swarm directory:
`swarm.hives.<hiveName>.domain`, default `<hiveName>.<swarm.domain>`.
Eval fails until `swarm.domain` and that entry exist
([`swarm/README.md`](../swarm/README.md) § Hive identity config).
## Bridge addressing
@ -162,13 +169,13 @@ agent containers.
receives DHCP via a regular UDP socket (it doesn't use a
netfilter-bypassing raw socket), so the hole is mandatory — without
it containers never get a lease and fall back to 169.254.x.x.
- Ports 80 and 443 let isolated agents reach nginx (gateway
container, shared host netns) for the forge sub-domain, per-agent
- Ports 80 and 443 let isolated agents reach nginx (on the host) for the forge sub-domain, per-agent
UI proxies, and any other HTTP services.
The **host** firewall is the only firewall. The shared-netns infra
containers (gateway, forge, matrix) set
`networking.firewall.enable = false`: a NixOS firewall inside a
The **host** firewall is the only firewall. The swarm service
containers that share the host netns (forge, matrix, authelia, bao, the
metrics and log stores) run with `networking.firewall.enable = false`
(`nix/container-modules/swarm-container.nix`): a NixOS firewall inside a
shared-netns container runs against the _host_ ruleset — at container
boot its `firewall-start` flushes the `nixos-fw` chains, rebuilds them
from the container's (empty) port list, and deletes the host's
@ -228,15 +235,16 @@ address arithmetic.
`hive-c0re` reads `HIVE_NETWORK_BRIDGE` + `HIVE_NETWORK_SUBNET` and passes
`PRIVATE_NETWORK=1`, `LOCAL_ADDRESS=` (empty), `HOST_ADDRESS=<bridge-ip>`,
and `HOST_BRIDGE=<bridgeName>` via `lifecycle::set_nspawn_flags` when
creating or updating containers. hive-c0re validates both variables **once at
and `HOST_BRIDGE=<bridgeName>` when creating or updating containers:
`lifecycle::set_nspawn_flags` hands them to hive-priv, which rewrites
`/etc/nixos-containers/<container>.conf`. hive-c0re validates both variables **once at
daemon startup**, not per container: they're process-global, so a
missing or malformed value is a misconfigured daemon rather than one bad
container, and failing at boot gives a single diagnostic instead of one
per agent. No non-isolated mode exists to fall back to. hive-c0re leaves `LOCAL_ADDRESS` empty so the
container's dhcpcd acquires an address from the bridge dnsmasq pool
(`networking.useDHCP = true` in `nix/agent-modules/network.nix`). This applies uniformly
to all containers — agents and service containers alike.
(`networking.useDHCP = true` in `nix/agent-modules/network.nix`). This applies to
every agent container, manager included.
`HOST_ADDRESS` is the bridge gateway IP (the address part of
`HIVE_NETWORK_SUBNET`, via `lifecycle::bridge_gateway_ip` — taken verbatim