hyperhive/docs/gateway.md

602 lines
33 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# hive-gateway
Single nginx in front of every hyperhive web surface. Runs on the **host**, next to hive-c0re, rather than in its own container: it shares the host netns anyway (see [Vhost map](#vhost-map) below), so containerizing it would buy no network isolation while costing a resolv.conf sync, a machine-bus reload, and three bind mounts. System-config (not meta-flake managed). Configured via `services.hyperhive.gateway.*` + per-subsystem opt-in flags in `services.hyperhive.{forge,matrix,...}`.
## Vhost map
| URL | vhost | upstream | source |
| --- | --- | --- | --- |
| `<hive>/` | `_` (catch-all) | dashboard dist (static, from `servedFrontend`); `/api/` + `/webhook/` → hive-c0re (`7000`) | always |
| `<hive>/agent/<name>/` | `_` | per-agent harness (UDS or TCP) | `agents.conf` (runtime-generated) |
| `<hive>/.well-known/matrix/{client,server}` | `_` | inline JSON (no upstream) | `matrix.enable && domain != null` |
| `<hive>/matrix/` (deprecated) | `_` | 301 → `matrix.<hive>/` | `matrix.gui.enable` |
| `forge.<hive>/` | `forge.<hive>` | forgejo (`3000`) | `forge.behindGateway` |
| `matrix.<hive>/_matrix/*` | `matrix.<hive>` | tuwunel (`8008`) | `matrix.gatewayHost != null` |
| `matrix.<hive>/` | `matrix.<hive>` | fluffychat-web static | `matrix.gui.enable` |
| `matrix.<hive>/config.json` | `matrix.<hive>` | inline JSON (FluffyChat boot config) | `matrix.gui.enable && domain != null` |
| `auth.<swarm>/` | `auth.<swarm>` | authelia (`9091`) | `swarm.authelia.enable` |
| `<swarm>/` | `<swarm>` | swarm-ui dist (static), behind an authelia subrequest | `swarm.ui.enable` |
The authelia vhost is declared only by the host that **runs** authelia, not by every hive that uses it — a client hive knows the swarm's `authelia.url` but must not answer for a name it doesn't serve. Its server name is exactly `swarm.authelia.domain`: authelia validates `authelia_url ⊂ session cookie domain` at startup, so a near-miss is a container that refuses to boot. It carries no `auth_basic` — the login page must not sit behind the login mechanism it replaces — and sets the four `X-Forwarded-{Proto,Host,Uri,For}` headers, since authelia decides by the *original* request rather than the hop it sees.
⚠️ **A `502` from this vhost usually means authelia has no users yet, not that the proxy is misconfigured.** Authelia treats an empty user store as a fatal startup error, so an enabled-but-unbootstrapped swarm crash-loops the container while the vhost in front of it works perfectly. Check `journalctl -M swarm-authelia -u authelia-swarm` before suspecting anything here; the bootstrap step is in [`swarm/sso.md`](swarm/sso.md).
Per-agent UIs stay sub-path, forge and matrix get sub-domains — see
[Sub-domain shape (rationale)](#sub-domain-shape-rationale) below for why.
## Discovery flow (matrix)
Operator points client at `<hive>`. Sequence:
1. Client fetches `https://<hive>/.well-known/matrix/client``{"m.homeserver":{"base_url":"https://matrix.<hive>"}}` (no port suffix when gateway listens on 443). The gateway always terminates TLS, so the scheme is always `https`; a non-default `httpsPort` is reflected as the port suffix.
2. Client connects to `matrix.<hive>/_matrix/client/...`.
3. Gateway routes `/_matrix/*` → tuwunel at `127.0.0.1:8008`.
matrix-dart-sdk (FluffyChat etc.) hardcodes `https` for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path.
Federation peers fetch `.well-known/matrix/server``{"m.server":"matrix.<hive>"}` and connect to `matrix.<hive>:8448` per spec default. Gateway only listens on configured `port` (+ `httpsPort` when TLS on); cross-hive federation needs either an SRV record (`_matrix._tcp.matrix.<hive>` → port 80 / 443) OR `matrix.openFirewall = true` so peers reach tuwunel's federation port directly. Hyperhive is mostly closed/internal, so this rarely bites.
## SPA fallback (Accept-header pattern)
The per-agent UIs and the `matrix.<hive>` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard vhost instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide:
- hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap)
- a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type
Solution: an `nginx http`-context `map $http_accept $<name>_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $<name>_spa_target <final>` falls through to `<final>`. No extension allowlist, no `if` block, no regex heuristics.
For matrix / per-agent static assets, `<final>` is `=404` (a missing asset is just missing).
### Dashboard: path-based routing (not Accept-header)
Every hive-c0re backend route lives under `/api/` plus the single `/webhook/knowledge` endpoint, so the dashboard vhost routes by **path**, not Accept header — deterministic, unlike a content-type split where the same URL could resolve differently depending on the caller's `Accept` header:
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions/mutations, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). Carries `proxy_buffering off` + a 1d read timeout for the streams.
- `location /webhook/` → hive-c0re: the knowledge webhook.
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html` (SPA fallback).
Each location carries a duplicated `auth_basic` block (separate locations don't inherit it). This keeps the gateway static-serving the dashboard dist while hive-c0re stays API-only — a frontend-only change doesn't rebuild or restart the core daemon. A new top-level c0re route prefix (beyond `/api` + `/webhook`) needs a matching `location` added to the dashboard vhost.
## Local dev (`localHostsEntry`)
`services.hyperhive.gateway.localHostsEntry = true` adds entries to the host's `/etc/hosts`:
- `<hive-domain>``127.0.0.1`
- `forge.<hive>``127.0.0.1` (when forge.behindGateway)
- `matrix.<hive>``127.0.0.1` (when matrix.gatewayHost set)
- `auth.<swarm>``127.0.0.1` (when swarm.authelia.enable)
`lib.unique` de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
## Sub-domain shape (rationale)
Operator decision: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
- forgejo's default `ROOT_URL = http://<host>/` works without any `X-Forwarded-Prefix` gymnastics — sub-domain hosting is the canonical Forgejo deploy shape.
- matrix-spec deployments universally use `matrix.<server_name>` for the actual API listener — federation already expects this.
- per-agent UIs are hyperhive-internal and base-path-aware specifically for `/agent/<name>/`. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins.
- cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
`services.hyperhive.{forge.domain,matrix.gatewayHost}` take the full hostname (`forge.darkest.space`, `git.example.com`) rather than a label that gets concatenated with hive-domain — operators want control over the full shape, not a forced `<label>.<hive-domain>` pattern.
## Tuning knobs
Per-vhost timeouts + body-size limits live in the location blocks:
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), `proxyWebsockets = true` (live-update endpoints).
- matrix `/_matrix/` (tuwunel): `client_max_body_size 50M` (media uploads), `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*` (federation + cross-origin clients), `proxyWebsockets = true`.
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE / WebSocket dashboards), `proxyWebsockets = true`, `X-Forwarded-Prefix` set so the harness can build absolute URLs when relative isn't enough.
SSH for forge stays direct on `cfg.sshPort` — separate listener protocol, not HTTP-over-nginx.
## Per-agent unix-socket upstream
All agents bind their web UI on a unix-domain socket at
`/run/hive-agent/<name>/web.sock` — the `HIVE_WEB_SOCKET` env var is
now set unconditionally for every agent. The mechanism:
1. **Agent side**. `HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock`
is set on every harness service env; `web_ui::serve` binds a
`UnixListener` at that path.
2. **Host side**. `hive-c0re` bind-mounts the per-agent subdir
(`/run/hive-agent/<name>/`) into the agent's container. Dir
bind, not file bind — file bind-mounts don't survive the
harness's `unlink + bind(2)` cycle on socket replace. Per-agent
subdir keeps each agent's container blind to siblings' sockets.
The dir is `0751`, owned by the agent's container uid/gid, so
nginx reaches `web.sock` through `o=--x` (traverse) and the socket's
own `0666`. The gateway is one of three principals sharing that dir
and does not own its ownership rules — see
[`docs/boundary.md`](boundary.md#the-per-agent-socket-dir).
3. **Marker gate**. After successful `bind_unix`, the harness drops
`<dir>/hyperhive-socket-bound` next to the socket. c0re's
`agent_sockets::write` filters its JSON map by marker presence —
only agents whose harness has actually bound the socket appear there.
(Legacy name `.bound` also accepted during the transition window.)
4. **Gateway side**. `gateway_nginx::write` generates
`/var/lib/hive-gateway/conf/agents.conf` — a plain nginx include
file with one `location /agent/<name>/` block per agent. Always
a UDS upstream (`http://unix:/run/hive-agent/<name>/web.sock:/`);
if the socket is not yet bound, nginx returns 502 caught by the
`error_page 502 503 504 = /__hive_agent_unreachable` directive.
nginx includes `/var/lib/hive-gateway/conf/agents.conf` — the same
path c0re writes, since both run on the host.
After each write, c0re triggers the appropriate nginx action via
`hive-priv` (which is root; hive-c0re runs as the unprivileged
`hive-core` user and cannot act on a system unit).
`hive-priv` queries `ActiveState` and dispatches:
- active → `systemctl reload nginx` (SIGHUP, zero-downtime)
- failed → `systemctl reset-failed nginx` + `systemctl start nginx`
- otherwise → `systemctl start nginx`
This is an explicit trigger rather than a path unit watching the
file: the write and the reload belong in one causal chain c0re can
retry and report on (`RELOAD_PENDING`), not two units racing on an
inotify event.
c0re regenerates `agents.conf` (and triggers a reload) on two
triggers: every topology change (new/removed agents) and every 10s
marker poll tick (`agent_sockets::spawn_poll`). `write()` is
idempotent — skips the rename when content is unchanged. Failed reloads
are retried automatically on subsequent poll ticks via
`gateway_nginx::reload_if_pending`.
`agents.conf` uses atomic `<path>.tmp` + `rename()` writes so a crashing
c0re process never leaves a partial or unparseable file behind.
## Dashboard link shape (gateway vs direct)
When the gateway is in front, the SW4RM tab builds per-agent links
as same-origin `/agent/<name>/…` URLs instead of the legacy direct
`http://<host>:<container.port>/` TCP shape. The signal comes from
`StateSnapshot.gateway_enabled`, sourced from the
`HIVE_GATEWAY_ENABLED` env the c0re NixOS module sets when
`services.hyperhive.gateway.enable = true`. Three render sites
flip together: the primary agent-name link, the favicon fetch
(`<url>/icon`), and the nav-strip `container`-kind links from
`DashboardState.links` (`GET /api/dashboard-state`). `forge`-kind nav-strip links still
resolve against `http://<host>:3000` (separate sub-domain transition
tracked by `forge.behindGateway`); `external`-kind links are
already absolute. See `docs/web-ui/dashboard.md::Container row` for the
frontend-side derivation.
## TLS modes
The gateway always terminates TLS — self-signed is the implicit floor when
nothing else is configured, so there is no http-only mode. Three modes,
selected by which (if any) external TLS source is set:
| mode | config | cert source | `.well-known` scheme |
|---|---|---|---|
| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | host hive-CA signs a gateway leaf (RSA-4096) | `https` |
| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` |
| operator cert | `tls.certDir` set | read from the operator's dir | `https` |
The `gateway.selfSignedTls` option is **deprecated and ignored** — self-signed
is now derived from the absence of `tls.certDir` / `tls.acme`. Setting it to
`false` (which used to select http-only or force an external cert) warns and
has no effect; use `tls.certDir` / `tls.acme` to override the default.
### ACME / Let's Encrypt (`tls.acme`)
Simplest production path for operators with a public domain:
```nix
services.hyperhive.gateway = {
openFirewall = true;
tls.acme = {
enable = true;
email = "admin@example.com";
};
};
```
nginx obtains and auto-renews certs via the ACME HTTP-01 challenge on `port` (default 80). Certs land in `/var/lib/acme/` on the host, managed by nixpkgs's `security.acme` in the ordinary way.
**Requirements**: `services.hyperhive.domain` must be publicly DNS-resolvable to this host, and `openFirewall = true` so Let's Encrypt can reach `/.well-known/acme-challenge/`. Each active vhost (main domain, `forge.<swarm-domain>`, `chat.<swarm-domain>`) gets its own cert via separate ACME challenges — the swarm services default to names under `services.hyperhive.swarm.domain`, so **every one of those names must resolve to this host too**, not just the hive's own.
Mutual exclusion: `tls.certDir` set together with `tls.acme.enable = true` fails an assertion — pick one external TLS source (or neither, for the self-signed default).
**Swarm peers**: CA-signed certs are trusted by default — this hive's entry in `swarm.hives` needs no `certFingerprint`.
### Self-signed TLS (default)
On by default, and listens on `httpsPort` (default 443) on every vhost beside the plain-http `port` (default 80).
The issuer is a **host-held hive CA**, not a bare self-signed leaf. A host service (`hive-tls-ca.service`, from the `hive-tls` module) generates a long-lived CA (`services.hyperhive.tls.caValidityDays`, default ~20y) under `services.hyperhive.tls.stateDir` (default `/var/lib/hive-tls`), then signs a gateway **leaf** (`leafValidityDays`, default 30d) with it. `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`).
⚠️ **Do not collapse that import unit into pointing nginx at the CA dir.**
It does two jobs, and skipping it has taken the gateway down in production
before. It re-modes the leaf (`hive-tls-ca` writes the key `0600
root:root`; nginx's pre-start `nginx -t` runs as the *nginx user*, so a
`0600` key fails the config test and blocks the unit), and it guarantees
**every cert path the nginx config names exists** — which is what the
swarm-services fallback below is for.
**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust — and a runtime-generated leaf can't be wired into an agent's build-time trust store at all. With a stable CA, agents and federation peers trust it *once*; leaf rotation never re-breaks them.
**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is itself issued under the swarm root ([`swarm/ca.md`](swarm/ca.md) has the hierarchy), and an intermediate is not a chain a verifier can terminate at — so the bundle carries the hive CA plus whatever it is rooted at. nginx is handed the leaf with the hive CA appended for the same reason. Everything that trusts the hive's TLS reads the bundle: agents (via `security.pki.certificateFiles`), the CI and forge containers, and a federating peer.
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) hardcodes `https://<host>/.well-known/matrix/client` for homeserver discovery and refuses to fall back to plain http. Without TLS the browser client cannot bootstrap.
**Cert shape**: leaf subject CN = bare hive domain; subjectAltName is `<hive>` plus wildcard `*.<hive>`, so all current and future sub-domain vhosts validate under the same leaf + the hive CA. A swarm service whose name is *not* under this hive's domain cannot be added here — the hive CA is name-constrained to `<hive>`, and a violating SAN invalidates the whole leaf, not just that name. Those names get the swarm-services leaf instead ([`swarm/ca.md`](swarm/ca.md)).
**Rotation**: `hive-tls-ca.service` is idempotent — it re-signs the leaf when it is missing or within 30 days of expiry, always under the same CA (so consumer trust is undisturbed). The CA itself is regenerated only if missing or already expired. To force a leaf rotation, delete `gateway.pem` under the state dir and restart the unit, then reload `nginx`.
**Cert prompts**: browsers still warn once per host until the hive's `trust-bundle.pem` is added to the browser/OS trust store (an anchor, not the leaf, is the thing to trust). Agent trust is wired separately (see the agent-trust work for `/run/hive-ca`).
### Operator-provided cert (`tls.certDir`)
For operators with a real CA cert (Let's Encrypt, corporate CA, etc.):
```nix
services.hyperhive.gateway = {
tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir
# tls.certName = "cert.pem"; # default — matches security.acme layout
# tls.keyName = "key.pem"; # default — matches security.acme layout
};
```
nginx reads the directory directly and uses `cert.pem` + `key.pem` (override `tls.certName`/`tls.keyName` for different filenames). Both modes listen on `httpsPort` (default 443) and emit `https://` in `.well-known` responses.
`tls.certDir` and `tls.acme.enable` set together is an assertion error.
**Key file permissions**: nixpkgs's `security.acme` outputs private keys as `0640 root:acme` by default. nginx runs as the `nginx` user and cannot read a key with that ownership. Fix with:
```nix
security.acme.certs."example.com".group = "nginx";
```
or make the key world-readable (`0644`) if your threat model allows it. nginx errors out at startup on a key it can't read — the error is explicit in the journal, not a silent failure.
**Swarm directory entry**: when using a CA-signed cert, this hive's entry needs no `certFingerprint` — the standard CA bundle validates:
```nix
services.hyperhive.swarm.hives.example = { domain = "example.com"; }; # no certFingerprint needed
```
### Fronting with an external TLS terminator
There is no http-only mode (see [TLS modes](#tls-modes) above). Two paths
for an operator who wants their own TLS terminator:
- give the gateway the real cert via `tls.certDir` (or `tls.acme`) so it
serves proper TLS directly — no separate proxy needed; or
- front it over a **unix socket** rather than a plain-http TCP port (the
intended direction for "bring your own proxy" — the gateway is not meant
to expose an unencrypted TCP upstream).
Because of this, `.well-known/matrix/{client,server}` discovery responses
always advertise `https` (see [Discovery flow](#discovery-flow-matrix) above).
## Firewall posture (host-level)
`hive-c0re.nix` opens the per-agent web-port range
`8100..8999` in the host firewall **only when
`services.hyperhive.gateway.enable = false`**. With the gateway on
(default) it's the sole external entry point and routes to agents over
the UDS upstream described above (see [Per-agent unix-socket
upstream](#per-agent-unix-socket-upstream)) — leaving the per-agent
ports firewall-open would defeat the single-front-door story. The
hashed TCP port (`lifecycle::agent_web_port`) still exists as a direct
host-loopback fallback for the pre-UDS/gateway-disabled case, but isn't
what the gateway itself proxies through.
`services.hyperhive.gateway.openFirewall = true` opens both `port` and
`httpsPort` — both are always served, since the gateway always terminates
TLS (see [TLS modes](#tls-modes) above).
Every agent hashes into the same port range (no special case), so
one range opening covers every container.
The dashboard port (`cfg.dashboardPort`, default 7000) is *not*
listed in either case — it binds `127.0.0.1` only, so a firewall
hole would be a no-op. Remote dashboard access flows through the
gateway. Operators who opt out of the gateway lose external
dashboard reach by design — the surface is privileged (approve /
deny / destroy) and must not be exposed without a real reverse
proxy in front.
## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain
Agents poll `HIVE_FORGE_URL` for Forgejo notifications + run all
`hive-forge` calls against it. Network isolation is always on (the
shared-netns mode was removed), so agents run in a private netns and
can never reach the host's loopback. `hive-c0re.nix` sets
`HIVE_FORGE_URL` to `http://<forge.domain>` (default
`forge.<hive-domain>`; `services.hyperhive.domain` is required). Agents
get the bridge dnsmasq as their resolver, resolve the hostname →
bridge IP, then reach nginx on port 80 (the bridge firewall opens
80+443). nginx proxies to forgejo — the same path an operator browser
takes, no raw port exposure needed.
## hive-forge container shape
Private Forgejo wrapped in a nixos-container (`hive-forge`, not
`h-*` — keeps c0re's lifecycle scanner out of the picture; the
operator manages it via the standard `nixos-container` CLI). The
container also keeps hive-forge from fighting any `services.forgejo`
the operator already runs on the host — separate systemd namespace,
separate state dir, separate port unless the operator deliberately
collides.
The forge container shares the host network namespace
(`privateNetwork = false`), so forgejo's listeners look like a
host-side service — nixos-container is here for state + systemd-unit
isolation, not network isolation. Note this is the FORGE container;
agent containers are network-isolated and reach the forge through the
gateway by `forge.<swarm-domain>` (see `HIVE_FORGE_URL` above), not via
the host's loopback.
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/`
and survives container restart / host reboot. To wipe, destroy the
container.
### Network and port configuration
```nix
services.hyperhive.forge = {
httpPort = 3000; # default — HTTP listener; outside hyperhive's 7000/8100-8999 range
sshPort = 2222; # default — git-over-SSH; kept off 22 so it doesn't collide with the host openssh
openFirewall = false; # default — expose httpPort + sshPort to the host firewall
};
```
`httpPort` (default **3000**) is the port Forgejo's HTTP server binds to.
It sits outside hyperhive's reserved ranges (dashboard 7000,
agents 81008999) so a default install has no port fights. Change it
only if you already have another process bound to 3000.
`sshPort` (default **2222**) is the port Forgejo's built-in SSH server
uses for `git clone/push/pull` over SSH (`git@<domain>:owner/repo.git`
via `-p 2222`). Port 22 is left alone on the host for openssh.
`openFirewall` (default **false**) controls whether `httpPort` and
`sshPort` are opened in the host firewall. Off by default (secure by
default): agents reach Forgejo through the gateway (`forge.<swarm-domain>` on
the bridge), not the raw port, so no firewall hole is needed. Flip to
`true` when you need:
- The operator's browser to reach `http://<host>:<httpPort>/` directly
(not behind the gateway).
- External git clients that push/pull via SSH directly to the host.
Forgejo served through the gateway (`forge.behindGateway = true`) does
not need `openFirewall` — the gateway's own `openFirewall` option covers
that path.
### `rootUrl` override
```nix
services.hyperhive.forge.rootUrl = "https://forge.example.com/";
```
`rootUrl` (default **null**) overrides the Forgejo `ROOT_URL` that is
auto-derived from `forge.domain` + gateway state. The auto-derivation
covers most cases:
| Shape | Auto-derived `ROOT_URL` |
|---|---|
| `behindGateway = true` | `http://<forge.domain>/` (port suffix omitted when `gateway.port == 80`) |
| `behindGateway = false` | `http://<forge.domain>:<httpPort>/` |
The auto-derivation always uses `http://`. Set `rootUrl` explicitly when
you need `https://` (e.g. behind a TLS-terminating reverse proxy, or when
clone URLs must carry `https://` because the gateway terminates TLS), or
when `forge.domain` resolves differently from the public URL. Must end with
`/` (Forgejo requirement; an assertion enforces this).
## Per-agent static frontend split
When `services.hyperhive.frontend` is configured, hive-c0re injects
`HIVE_AGENT_FRONTEND_DIR = "${cfg.frontend}/agent"` into its service
environment. The nginx include generator (`gateway_nginx::write`) reads
this variable and, when set, emits split location blocks per agent
instead of the legacy single-proxy block.
**Location priority for `/agent/<name>/...`:**
```nginx
# 1. Compiled assets — content-addressed nix store path, cache forever
location ^~ /agent/<name>/static/ {
alias <frontend>/static/;
expires 1y;
add_header Cache-Control "public, immutable, max-age=31536000";
}
# 2. Static dist + proxy fallback for dynamic paths
location /agent/<name>/ {
alias <frontend>/;
try_files $uri $uri.html $uri/index.html @<name>_dynamic;
}
# 3. Proxy catchall — API, events, icon, send, login, …
location @<name>_dynamic {
proxy_pass <upstream>;
proxy_set_header X-Forwarded-Prefix /agent/<name>;
proxy_intercept_errors on;
error_page 502 503 504 = /__hive_agent_unreachable;
# … (full proxy header block)
}
```
**`try_files` resolution** (nginx applies the `alias` mapping before
checking file existence):
| request | resolved | outcome |
| --- | --- | --- |
| `/agent/iris/` | `<frontend>/index.html` | main agent page |
| `/agent/iris/stats` | `<frontend>/stats.html` | stats page |
| `/agent/iris/screen` | `<frontend>/screen.html` | screen page |
| `/agent/iris/static/app.js` | caught by `^~` block first | served with immutable cache |
| `/agent/iris/api/state` | no file match → `@iris_dynamic` | proxied to agent daemon |
| `/agent/iris/events/live` | no file match → `@iris_dynamic` | proxied (SSE) |
Adding a new HTML page to the frontend dist (`dist/<page>.html`)
automatically makes it reachable at `/agent/<name>/<page>` — no
generator change needed.
**Why `^~` for `/static/`**: the `^~` prefix gives this block higher
priority than the plain prefix `location /agent/<name>/`, so compiled
JS/CSS assets skip `try_files` entirely and get the immutable cache
headers. Nix store paths are content-addressed — the hash changes on
any content change — so `max-age=31536000` is safe.
**Why the nix store path resolves**: `HIVE_AGENT_FRONTEND_DIR` is a nix
store path baked in at hive-c0re build time, and c0re (writing
`agents.conf`) and nginx (serving files from it) are on the same machine,
so they see the same store.
**Graceful degradation**: if `HIVE_AGENT_FRONTEND_DIR` is empty or
unset (e.g. a build that predates `cfg.frontend`), each agent gets the
legacy single-proxy block and all traffic is forwarded to the agent
daemon as before.
**`extraFiles`**: per-agent `hyperhive.frontend.extraFiles` are in
`mergedDist`, not in the base `cfg.frontend` dist. They are not under
the nix-store `alias` path, so requests for them fall through
`try_files` to `@<name>_dynamic` and are served by the agent daemon
as before.
## Per-agent error pages
`/agent/<name>/` requests hit two failure modes; both get static
HTML pages instead of nginx's default error chrome:
- **Agent not found** (`/agent/<unknown>/...`) — name isn't in
`agentPortsTable`. nginx's prefix match falls back to the bare
`/agent/` catch-all, which `return 404`s and `error_page 404` rewrites
to `/__hive_agent_not_found` → serves `not-found.html` with a link
back to the dashboard.
- **Agent unreachable** (`502 / 503 / 504` from `proxy_pass`) — the
per-agent harness isn't responding (container restarting, crash
recovery, etc.). `proxy_intercept_errors on` + `error_page 502 503
504 = /__hive_agent_unreachable` rewrites to `unreachable.html`.
Both pages are built at deploy time via `pkgs.runCommand` (one nix
derivation `hyperhive-agent-error-pages` with `not-found.html` +
`unreachable.html` inside) and served via two `internal` nginx
locations with `alias` to the exact file. `internal` keeps the
files from being directly request-able by operators — only nginx's
own error-handling can reach them.
Page styling: minimal inline CSS matching the dashboard's catppuccin
palette (`#1e1e2e` bg, `#cdd6f4` text, `#cba6f7` heading). No
dependencies on the frontend dist — these pages render even when
hive-c0re itself is down.
Scope is intentionally narrow: a route earns a custom page when the
default status code would point at the wrong component. The per-agent
routes qualify (a 502 there means the harness is restarting, not that
the gateway is broken), and so does `auth.<swarm>` — a dead authelia
upstream almost always means the user store was never bootstrapped, and
a bare 502 blames the proxy, which is the one part that is working.
Forge / matrix / fluffychat still get nginx defaults: their upstreams
being down means what the status code says, so a themed page would add
styling and no information.
## HTTP Basic auth
`services.hyperhive.gateway.auth.enable = true` gates every request to
the main vhost (`_`) behind HTTP Basic auth. nginx's built-in `auth_basic`
module validates credentials; no extra service or host-side daemon is
required.
**Setup:**
```nix
services.hyperhive.gateway.auth = {
enable = true;
# realm = "hyperhive"; # optional, default shown
};
```
The credential store lives at the fixed path
`/var/lib/hive-gateway/conf/gateway.htpasswd` on the host. A tmpfiles
rule pre-creates the file on first boot; no manual path configuration
is required. nginx reads it at that path directly.
Manage users with `hivectl gateway`. `hivectl` sends the request over the
host admin socket and the `hive-c0re` daemon performs the write at its
canonical path — no path is exposed to the CLI:
```sh
# Add or update a user (prompted for password):
hivectl gateway create-user alice --password-stdin
# Add with inline password (visible in shell history — avoid for sensitive creds):
hivectl gateway create-user bob --password hunter2
# Remove a user:
hivectl gateway delete-user bob
# List current usernames:
hivectl gateway list-users
```
The daemon hashes passwords with BCrypt (cost 12) and writes
`$2y$`-prefixed hashes that nginx accepts natively. No external
`htpasswd` binary is required.
**What is not gated:** per-agent UI routes emitted into `agents.conf`
(served under `/agent/<name>/`) inherit no auth from `/` — nginx
applies `auth_basic` per-location. Full per-agent coverage is a
follow-up.
**Realm:** the `WWW-Authenticate: Basic realm="..."` string browsers
display in the credential dialog. Defaults to `"hyperhive"`. Must not
contain `"` or `$`.
**Custom 401 page:** when credentials are absent or wrong, nginx serves
a Catppuccin-styled `unauthorized.html` page (built into the same Nix
derivation as the agent error pages) that tells the operator which
`hivectl` command to run to create a user. The response status is still
`401` (`error_page 401 =401 /__hive_auth_unauthorized`) so browsers
present the login dialog on the first visit — users who dismiss the
dialog see the human-readable hint. The internal exact-match location
(`= /__hive_auth_unauthorized`) beats `location /` in nginx's prefix
ordering, preventing the subrequest from looping back through
`auth_basic`.
## Security headers
The following headers are emitted at server scope on every gateway
vhost (`_`, `forge.<swarm-domain>`, `chat.<swarm-domain>`):
| Header | Value |
|--------|-------|
| `X-Frame-Options` | `SAMEORIGIN` |
| `X-Content-Type-Options` | `nosniff` |
| `Referrer-Policy` | `strict-origin-when-cross-origin` |
nginx's `add_header` inheritance rule: a `location` block that sets its
own `add_header` does **not** inherit server-scope headers. API locations
that carry their own CORS headers (e.g. `/.well-known/matrix/client`,
`/_matrix/`) are therefore unaffected. HTML-serving and proxy locations
with no `add_header` of their own pick the security headers up
automatically.
### HSTS (`gateway.hsts`)
HSTS is **opt-in** and disabled by default:
```nix
services.hyperhive.gateway.hsts = {
enable = true; # default: false
maxAge = 31536000; # default: 1 year (required for preload list)
includeSubDomains = true; # default: true
};
```
When enabled, a `Strict-Transport-Security: max-age=...[; includeSubDomains]`
header is added alongside the other security headers.
**Opt-in rationale**: HSTS pins HTTPS in the browser's preload cache;
enabling it on a deployment that later loses TLS locks browsers out
until `max-age` expires. Only enable when TLS is permanent.
Since the gateway always terminates TLS (see [TLS modes](#tls-modes)
above), an enabled HSTS header is always served over https — there is no
TLS-less mode that could violate it.