Follow-up to the plumbing removal: docs/swarm/README.md gets the biggest rewrite (drops the whole "Fingerprint format" section, fixes the runtime-effects list, the WireGuard config example + "what the mesh does" bullet), docs/gateway.md and hive-gateway/options.nix drop 4 "needs no certFingerprint" mentions, swarm-peers-removed.nix's migration-warning text no longer tells an upgrading operator to carry a field over that no longer exists, swarm.nix/swarm-wireguard.nix/ swarm-controller.nix/swarm-controller's main.rs get comment fixes where they described the now-removed HYPERHIVE_PEERS shape. Also caught one more stale "peer hives" mention in docs/web-ui/README.md's SW4RM tab description that the first pass on this issue missed.
601 lines
33 KiB
Markdown
601 lines
33 KiB
Markdown
# hive-gateway
|
||
|
||
Single nginx in front of every hyperhive web surface. Runs on the **host**, next to hive-c0re, rather than in its own container: it shares the host netns anyway (see [Vhost map](#vhost-map) below), so containerizing it would buy no network isolation while costing a resolv.conf sync, a machine-bus reload, and three bind mounts. System-config (not meta-flake managed). Configured via `services.hyperhive.gateway.*` + per-subsystem opt-in flags in `services.hyperhive.{forge,matrix,...}`.
|
||
|
||
## Vhost map
|
||
|
||
| URL | vhost | upstream | source |
|
||
| --- | --- | --- | --- |
|
||
| `<hive>/` | `_` (catch-all) | dashboard dist (static, from `servedFrontend`); `/api/` + `/webhook/` → hive-c0re (`7000`) | always |
|
||
| `<hive>/agent/<name>/` | `_` | per-agent harness (UDS or TCP) | `agents.conf` (runtime-generated) |
|
||
| `<hive>/.well-known/matrix/{client,server}` | `_` | inline JSON (no upstream) | `matrix.enable && domain != null` |
|
||
| `<hive>/matrix/` (deprecated) | `_` | 301 → `chat.<swarm>/` | `matrix.gui.enable` |
|
||
| `forge.<swarm>/` | `forge.<swarm>` | forgejo (`3000`) | `forge.behindGateway` |
|
||
| `chat.<swarm>/_matrix/*` | `chat.<swarm>` | tuwunel (`8008`) | `matrix.gatewayHost != null` |
|
||
| `chat.<swarm>/` | `chat.<swarm>` | fluffychat-web static | `matrix.gui.enable` |
|
||
| `chat.<swarm>/config.json` | `chat.<swarm>` | inline JSON (FluffyChat boot config) | `matrix.gui.enable && domain != null` |
|
||
| `auth.<swarm>/` | `auth.<swarm>` | authelia (`9091`) | `swarm.authelia.enable` |
|
||
| `<swarm>/` | `<swarm>` | swarm-ui dist (static), behind an authelia subrequest | `swarm.ui.enable` |
|
||
|
||
The authelia vhost is declared only by the host that **runs** authelia, not by every hive that uses it — a client hive knows the swarm's `authelia.url` but must not answer for a name it doesn't serve. Its server name is exactly `swarm.authelia.domain`: authelia validates `authelia_url ⊂ session cookie domain` at startup, so a near-miss is a container that refuses to boot. It carries no `auth_basic` — the login page must not sit behind the login mechanism it replaces — and sets the four `X-Forwarded-{Proto,Host,Uri,For}` headers, since authelia decides by the *original* request rather than the hop it sees.
|
||
|
||
⚠️ **A `502` from this vhost usually means authelia has no users yet, not that the proxy is misconfigured.** Authelia treats an empty user store as a fatal startup error, so an enabled-but-unbootstrapped swarm crash-loops the container while the vhost in front of it works perfectly. Check `journalctl -M swarm-authelia -u authelia-swarm` before suspecting anything here; the bootstrap step is in [`swarm/sso.md`](swarm/sso.md).
|
||
|
||
Per-agent UIs stay sub-path, forge and matrix get sub-domains — see
|
||
[Sub-domain shape (rationale)](#sub-domain-shape-rationale) below for why.
|
||
|
||
## Discovery flow (matrix)
|
||
|
||
Operator points client at `<hive>`. Sequence:
|
||
|
||
1. Client fetches `https://<hive>/.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat.<swarm>"}}` (no port suffix when gateway listens on 443). The gateway always terminates TLS, so the scheme is always `https`; a non-default `httpsPort` is reflected as the port suffix.
|
||
2. Client connects to `chat.<swarm>/_matrix/client/...`.
|
||
3. Gateway routes `/_matrix/*` → tuwunel at `127.0.0.1:8008`.
|
||
|
||
matrix-dart-sdk (FluffyChat etc.) hardcodes `https` for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path.
|
||
|
||
Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.<swarm>:<httpsPort>"}` (the federation delegation always carries an explicit port, even the HTTPS default 443 — the https-implies-443 elision only applies to the client base_url above). Gateway only listens on configured `port` (+ `httpsPort` when TLS on); cross-hive federation needs either an SRV record (`_matrix._tcp.chat.<swarm>` → port 80 / 443) OR `matrix.openFirewall = true` so peers reach tuwunel's federation port directly. Hyperhive is mostly closed/internal, so this rarely bites.
|
||
|
||
## SPA fallback (Accept-header pattern)
|
||
|
||
The per-agent UIs and the `chat.<swarm>` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard vhost instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide:
|
||
|
||
- hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap)
|
||
- a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type
|
||
|
||
Solution: an `nginx http`-context `map $http_accept $<name>_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $<name>_spa_target <final>` falls through to `<final>`. No extension allowlist, no `if` block, no regex heuristics.
|
||
|
||
For matrix / per-agent static assets, `<final>` is `=404` (a missing asset is just missing).
|
||
|
||
### Dashboard: path-based routing (not Accept-header)
|
||
|
||
Every hive-c0re backend route lives under `/api/` plus the single `/webhook/knowledge` endpoint, so the dashboard vhost routes by **path**, not Accept header — deterministic, unlike a content-type split where the same URL could resolve differently depending on the caller's `Accept` header:
|
||
|
||
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions/mutations, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). Carries `proxy_buffering off` + a 1d read timeout for the streams.
|
||
- `location /webhook/` → hive-c0re: the knowledge webhook.
|
||
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html` (SPA fallback).
|
||
|
||
Each location carries a duplicated `auth_basic` block (separate locations don't inherit it). This keeps the gateway static-serving the dashboard dist while hive-c0re stays API-only — a frontend-only change doesn't rebuild or restart the core daemon. A new top-level c0re route prefix (beyond `/api` + `/webhook`) needs a matching `location` added to the dashboard vhost.
|
||
|
||
## Local dev (`localHostsEntry`)
|
||
|
||
`services.hyperhive.gateway.localHostsEntry = true` adds entries to the host's `/etc/hosts`:
|
||
|
||
- `<hive-domain>` → `127.0.0.1`
|
||
- `forge.<swarm>` → `127.0.0.1` (when forge.behindGateway)
|
||
- `chat.<swarm>` → `127.0.0.1` (when matrix.gatewayHost set)
|
||
- `auth.<swarm>` → `127.0.0.1` (when swarm.authelia.enable)
|
||
|
||
`lib.unique` de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
|
||
|
||
## Sub-domain shape (rationale)
|
||
|
||
Operator decision: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
|
||
|
||
- forgejo's default `ROOT_URL = http://<host>/` works without any `X-Forwarded-Prefix` gymnastics — sub-domain hosting is the canonical Forgejo deploy shape.
|
||
- matrix-spec deployments universally use `matrix.<server_name>` for the actual API listener — federation already expects this.
|
||
- per-agent UIs are hyperhive-internal and base-path-aware specifically for `/agent/<name>/`. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins.
|
||
- cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
|
||
|
||
`services.hyperhive.{forge.domain,matrix.gatewayHost}` take the full hostname (`forge.darkest.space`, `git.example.com`) rather than a label that gets concatenated with hive-domain — operators want control over the full shape, not a forced `<label>.<hive-domain>` pattern.
|
||
|
||
## Tuning knobs
|
||
|
||
Per-vhost timeouts + body-size limits live in the location blocks:
|
||
|
||
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), `proxyWebsockets = true` (live-update endpoints).
|
||
- matrix `/_matrix/` (tuwunel): `client_max_body_size 50M` (media uploads), `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*` (federation + cross-origin clients), `proxyWebsockets = true`.
|
||
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE / WebSocket dashboards), `proxyWebsockets = true`, `X-Forwarded-Prefix` set so the harness can build absolute URLs when relative isn't enough.
|
||
|
||
SSH for forge stays direct on `cfg.sshPort` — separate listener protocol, not HTTP-over-nginx.
|
||
|
||
## Per-agent unix-socket upstream
|
||
|
||
All agents bind their web UI on a unix-domain socket at
|
||
`/run/hive-agent/<name>/web.sock` — the `HIVE_WEB_SOCKET` env var is
|
||
now set unconditionally for every agent. The mechanism:
|
||
|
||
1. **Agent side**. `HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock`
|
||
is set on every harness service env; `web_ui::serve` binds a
|
||
`UnixListener` at that path.
|
||
2. **Host side**. `hive-c0re` bind-mounts the per-agent subdir
|
||
(`/run/hive-agent/<name>/`) into the agent's container. Dir
|
||
bind, not file bind — file bind-mounts don't survive the
|
||
harness's `unlink + bind(2)` cycle on socket replace. Per-agent
|
||
subdir keeps each agent's container blind to siblings' sockets.
|
||
|
||
The dir is `0751`, owned by the agent's container uid/gid, so
|
||
nginx reaches `web.sock` through `o=--x` (traverse) and the socket's
|
||
own `0666`. The gateway is one of three principals sharing that dir
|
||
and does not own its ownership rules — see
|
||
[`docs/boundary.md`](boundary.md#the-per-agent-socket-dir).
|
||
3. **Marker gate**. After successful `bind_unix`, the harness drops
|
||
`<dir>/hyperhive-socket-bound` next to the socket. c0re's
|
||
`agent_sockets::write` filters its JSON map by marker presence —
|
||
only agents whose harness has actually bound the socket appear there.
|
||
(Legacy name `.bound` also accepted during the transition window.)
|
||
4. **Gateway side**. `gateway_nginx::write` generates
|
||
`/var/lib/hive-gateway/conf/agents.conf` — a plain nginx include
|
||
file with one `location /agent/<name>/` block per agent. Always
|
||
a UDS upstream (`http://unix:/run/hive-agent/<name>/web.sock:/`);
|
||
if the socket is not yet bound, nginx returns 502 caught by the
|
||
`error_page 502 503 504 = /__hive_agent_unreachable` directive.
|
||
nginx includes `/var/lib/hive-gateway/conf/agents.conf` — the same
|
||
path c0re writes, since both run on the host.
|
||
After each write, c0re triggers the appropriate nginx action via
|
||
`hive-priv` (which is root; hive-c0re runs as the unprivileged
|
||
`hive-core` user and cannot act on a system unit).
|
||
`hive-priv` queries `ActiveState` and dispatches:
|
||
- active → `systemctl reload nginx` (SIGHUP, zero-downtime)
|
||
- failed → `systemctl reset-failed nginx` + `systemctl start nginx`
|
||
- otherwise → `systemctl start nginx`
|
||
This is an explicit trigger rather than a path unit watching the
|
||
file: the write and the reload belong in one causal chain c0re can
|
||
retry and report on (`RELOAD_PENDING`), not two units racing on an
|
||
inotify event.
|
||
|
||
c0re regenerates `agents.conf` (and triggers a reload) on two
|
||
triggers: every topology change (new/removed agents) and every 10s
|
||
marker poll tick (`agent_sockets::spawn_poll`). `write()` is
|
||
idempotent — skips the rename when content is unchanged. Failed reloads
|
||
are retried automatically on subsequent poll ticks via
|
||
`gateway_nginx::reload_if_pending`.
|
||
|
||
`agents.conf` uses atomic `<path>.tmp` + `rename()` writes so a crashing
|
||
c0re process never leaves a partial or unparseable file behind.
|
||
|
||
## Dashboard link shape (gateway vs direct)
|
||
|
||
When the gateway is in front, the SW4RM tab builds per-agent links
|
||
as same-origin `/agent/<name>/…` URLs instead of the legacy direct
|
||
`http://<host>:<container.port>/` TCP shape. The signal comes from
|
||
`StateSnapshot.gateway_enabled`, sourced from the
|
||
`HIVE_GATEWAY_ENABLED` env the c0re NixOS module now always sets
|
||
(`services.hyperhive.gateway.enable` was removed — the gateway runs
|
||
unconditionally alongside hyperhive), so this is effectively always
|
||
true; the `false` branch is retained as a defensive fallback for the
|
||
env being unset. Three render sites
|
||
flip together: the primary agent-name link, the favicon fetch
|
||
(`<url>/icon`), and the nav-strip `container`-kind links from
|
||
`DashboardState.links` (`GET /api/dashboard-state`). `forge`-kind nav-strip links still
|
||
resolve against `http://<host>:3000` (separate sub-domain transition
|
||
tracked by `forge.behindGateway`); `external`-kind links are
|
||
already absolute. See `docs/web-ui/dashboard.md::Container row` for the
|
||
frontend-side derivation.
|
||
|
||
## TLS modes
|
||
|
||
The gateway always terminates TLS — self-signed is the implicit floor when
|
||
nothing else is configured, so there is no http-only mode. Three modes,
|
||
selected by which (if any) external TLS source is set:
|
||
|
||
| mode | config | cert source | `.well-known` scheme |
|
||
|---|---|---|---|
|
||
| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | host hive-CA signs a gateway leaf (RSA-4096) | `https` |
|
||
| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` |
|
||
| operator cert | `tls.certDir` set | read from the operator's dir | `https` |
|
||
|
||
The `gateway.selfSignedTls` option has been **removed** — self-signed
|
||
is now derived from the absence of `tls.certDir` / `tls.acme`. A config
|
||
that still sets it fails eval with a removal message; use `tls.certDir`
|
||
/ `tls.acme` to override the default.
|
||
|
||
### ACME / Let's Encrypt (`tls.acme`)
|
||
|
||
Simplest production path for operators with a public domain:
|
||
|
||
```nix
|
||
services.hyperhive.gateway = {
|
||
openFirewall = true;
|
||
tls.acme = {
|
||
enable = true;
|
||
email = "admin@example.com";
|
||
};
|
||
};
|
||
```
|
||
|
||
nginx obtains and auto-renews certs via the ACME HTTP-01 challenge on `port` (default 80). Certs land in `/var/lib/acme/` on the host, managed by nixpkgs's `security.acme` in the ordinary way.
|
||
|
||
**Requirements**: `services.hyperhive.domain` must be publicly DNS-resolvable to this host, and `openFirewall = true` so Let's Encrypt can reach `/.well-known/acme-challenge/`. Each active vhost (main domain, `forge.<swarm-domain>`, `chat.<swarm-domain>`) gets its own cert via separate ACME challenges — the swarm services default to names under `services.hyperhive.swarm.domain`, so **every one of those names must resolve to this host too**, not just the hive's own.
|
||
|
||
Mutual exclusion: `tls.certDir` set together with `tls.acme.enable = true` fails an assertion — pick one external TLS source (or neither, for the self-signed default).
|
||
|
||
### Self-signed TLS (default)
|
||
|
||
On by default, and listens on `httpsPort` (default 443) on every vhost beside the plain-http `port` (default 80).
|
||
|
||
The issuer is a **host-held hive CA**, not a bare self-signed leaf. A host service (`hive-tls-ca.service`, from the `hive-tls` module) generates a long-lived CA (`services.hyperhive.tls.caValidityDays`, default ~20y) under `services.hyperhive.tls.stateDir` (default `/var/lib/hive-tls`), then signs a gateway **leaf** (`leafValidityDays`, default 30d) with it. `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`).
|
||
|
||
⚠️ **Do not collapse that import unit into pointing nginx at the CA dir.**
|
||
It does two jobs, and skipping it has taken the gateway down in production
|
||
before. It re-modes the leaf (`hive-tls-ca` writes the key `0600
|
||
root:root`; nginx's pre-start `nginx -t` runs as the *nginx user*, so a
|
||
`0600` key fails the config test and blocks the unit), and it guarantees
|
||
**every cert path the nginx config names exists** — which is what the
|
||
swarm-services fallback below is for.
|
||
|
||
**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust — and a runtime-generated leaf can't be wired into an agent's build-time trust store at all. With a stable CA, agents and federation peers trust it *once*; leaf rotation never re-breaks them.
|
||
|
||
**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is itself issued under the swarm root ([`swarm/ca.md`](swarm/ca.md) has the hierarchy), and an intermediate is not a chain a verifier can terminate at — so the bundle carries the hive CA plus whatever it is rooted at. nginx is handed the leaf with the hive CA appended for the same reason. Everything that trusts the hive's TLS reads the bundle: agents (via `security.pki.certificateFiles`), the CI and forge containers, and a federating peer.
|
||
|
||
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) hardcodes `https://<host>/.well-known/matrix/client` for homeserver discovery and refuses to fall back to plain http. Without TLS the browser client cannot bootstrap.
|
||
|
||
**Cert shape**: leaf subject CN = bare hive domain; subjectAltName is `<hive>` plus wildcard `*.<hive>`, so all current and future sub-domain vhosts validate under the same leaf + the hive CA. A swarm service whose name is *not* under this hive's domain cannot be added here — the hive CA is name-constrained to `<hive>`, and a violating SAN invalidates the whole leaf, not just that name. Those names get the swarm-services leaf instead ([`swarm/ca.md`](swarm/ca.md)).
|
||
|
||
**Rotation**: `hive-tls-ca.service` is idempotent — it re-signs the leaf when it is missing or within 30 days of expiry, always under the same CA (so consumer trust is undisturbed). The CA itself is regenerated only if missing or already expired. To force a leaf rotation, delete `gateway.pem` under the state dir and restart the unit, then reload `nginx`.
|
||
|
||
**Cert prompts**: browsers still warn once per host until the hive's `trust-bundle.pem` is added to the browser/OS trust store (an anchor, not the leaf, is the thing to trust). Agent trust is wired separately (see the agent-trust work for `/run/hive-ca`).
|
||
|
||
### Operator-provided cert (`tls.certDir`)
|
||
|
||
For operators with a real CA cert (Let's Encrypt, corporate CA, etc.):
|
||
|
||
```nix
|
||
services.hyperhive.gateway = {
|
||
tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir
|
||
# tls.certName = "cert.pem"; # default — matches security.acme layout
|
||
# tls.keyName = "key.pem"; # default — matches security.acme layout
|
||
};
|
||
```
|
||
|
||
nginx reads the directory directly and uses `cert.pem` + `key.pem` (override `tls.certName`/`tls.keyName` for different filenames). Both modes listen on `httpsPort` (default 443) and emit `https://` in `.well-known` responses.
|
||
|
||
`tls.certDir` and `tls.acme.enable` set together is an assertion error.
|
||
|
||
**Key file permissions**: nixpkgs's `security.acme` outputs private keys as `0640 root:acme` by default. nginx runs as the `nginx` user and cannot read a key with that ownership. Fix with:
|
||
|
||
```nix
|
||
security.acme.certs."example.com".group = "nginx";
|
||
```
|
||
|
||
or make the key world-readable (`0644`) if your threat model allows it. nginx errors out at startup on a key it can't read — the error is explicit in the journal, not a silent failure.
|
||
|
||
### Fronting with an external TLS terminator
|
||
|
||
There is no http-only mode (see [TLS modes](#tls-modes) above). Two paths
|
||
for an operator who wants their own TLS terminator:
|
||
|
||
- give the gateway the real cert via `tls.certDir` (or `tls.acme`) so it
|
||
serves proper TLS directly — no separate proxy needed; or
|
||
- front it over a **unix socket** rather than a plain-http TCP port (the
|
||
intended direction for "bring your own proxy" — the gateway is not meant
|
||
to expose an unencrypted TCP upstream).
|
||
|
||
Because of this, `.well-known/matrix/{client,server}` discovery responses
|
||
always advertise `https` (see [Discovery flow](#discovery-flow-matrix) above).
|
||
|
||
## Firewall posture (host-level)
|
||
|
||
The gateway is unconditional — `services.hyperhive.gateway.enable` was
|
||
removed, there is no gateway-off mode. nginx is always the sole
|
||
external entry point and routes to agents over the UDS upstream
|
||
described above (see [Per-agent unix-socket
|
||
upstream](#per-agent-unix-socket-upstream)), so the per-agent web-port
|
||
range `8100..8999` stays closed on the host firewall
|
||
unconditionally — opening it would defeat the single-front-door story.
|
||
The hashed TCP port (`lifecycle::agent_web_port`) still exists as a
|
||
fallback bind for an agent whose `HIVE_WEB_SOCKET` env somehow ends up
|
||
unset, but nothing opens a matching firewall hole for it and the
|
||
gateway itself never proxies through it.
|
||
|
||
`services.hyperhive.gateway.openFirewall = true` opens both `port` and
|
||
`httpsPort` — both are always served, since the gateway always terminates
|
||
TLS (see [TLS modes](#tls-modes) above).
|
||
|
||
Every agent hashes into the same port range (no special case), so
|
||
one range opening covers every container.
|
||
|
||
The dashboard port (`cfg.dashboardPort`, default 7000) is *not*
|
||
listed in either case — it binds `127.0.0.1` only, so a firewall
|
||
hole would be a no-op. Remote dashboard access flows through the
|
||
gateway. Operators who opt out of the gateway lose external
|
||
dashboard reach by design — the surface is privileged (approve /
|
||
deny / destroy) and must not be exposed without a real reverse
|
||
proxy in front.
|
||
|
||
## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain
|
||
|
||
Agents poll `HIVE_FORGE_URL` for Forgejo notifications + run all
|
||
`hive-forge` calls against it. Network isolation is always on (the
|
||
shared-netns mode was removed), so agents run in a private netns and
|
||
can never reach the host's loopback.
|
||
`nix/host-modules/hive-c0re/environment.nix` sets `HIVE_FORGE_URL` to
|
||
`http://<forge.domain>` (default `forge.<swarm-domain>` — a swarm runs
|
||
one forge; `services.hyperhive.domain` is required). Agents
|
||
get the bridge dnsmasq as their resolver, resolve the hostname →
|
||
bridge IP, then reach nginx on port 80 (the bridge firewall opens
|
||
80+443). nginx proxies to forgejo — the same path an operator browser
|
||
takes, no raw port exposure needed.
|
||
|
||
## hive-forge container shape
|
||
|
||
Private Forgejo wrapped in a nixos-container (`hive-forge`, not
|
||
`h-*` — keeps c0re's lifecycle scanner out of the picture; the
|
||
operator manages it via the standard `nixos-container` CLI). The
|
||
container also keeps hive-forge from fighting any `services.forgejo`
|
||
the operator already runs on the host — separate systemd namespace,
|
||
separate state dir, separate port unless the operator deliberately
|
||
collides.
|
||
|
||
The forge container shares the host network namespace
|
||
(`privateNetwork = false`), so forgejo's listeners look like a
|
||
host-side service — nixos-container is here for state + systemd-unit
|
||
isolation, not network isolation. Note this is the FORGE container;
|
||
agent containers are network-isolated and reach the forge through the
|
||
gateway by `forge.<swarm-domain>` (see `HIVE_FORGE_URL` above), not via
|
||
the host's loopback.
|
||
|
||
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/`
|
||
and survives container restart / host reboot. To wipe, destroy the
|
||
container.
|
||
|
||
### Network and port configuration
|
||
|
||
```nix
|
||
services.hyperhive.forge = {
|
||
httpPort = 3000; # default — HTTP listener; outside hyperhive's 7000/8100-8999 range
|
||
sshPort = 2222; # default — git-over-SSH; kept off 22 so it doesn't collide with the host openssh
|
||
openFirewall = false; # default — expose httpPort + sshPort to the host firewall
|
||
};
|
||
```
|
||
|
||
`httpPort` (default **3000**) is the port Forgejo's HTTP server binds to.
|
||
It sits outside hyperhive's reserved ranges (dashboard 7000,
|
||
agents 8100–8999) so a default install has no port fights. Change it
|
||
only if you already have another process bound to 3000.
|
||
|
||
`sshPort` (default **2222**) is the port Forgejo's built-in SSH server
|
||
uses for `git clone/push/pull` over SSH (`git@<domain>:owner/repo.git`
|
||
via `-p 2222`). Port 22 is left alone on the host for openssh.
|
||
|
||
`openFirewall` (default **false**) controls whether `httpPort` and
|
||
`sshPort` are opened in the host firewall. Off by default (secure by
|
||
default): agents reach Forgejo through the gateway (`forge.<swarm-domain>` on
|
||
the bridge), not the raw port, so no firewall hole is needed. Flip to
|
||
`true` when you need:
|
||
- The operator's browser to reach `http://<host>:<httpPort>/` directly
|
||
(not behind the gateway).
|
||
- External git clients that push/pull via SSH directly to the host.
|
||
|
||
Forgejo served through the gateway (`forge.behindGateway = true`) does
|
||
not need `openFirewall` — the gateway's own `openFirewall` option covers
|
||
that path.
|
||
|
||
### `rootUrl` override
|
||
|
||
```nix
|
||
services.hyperhive.forge.rootUrl = "https://forge.example.com/";
|
||
```
|
||
|
||
`rootUrl` (default **null**) overrides the Forgejo `ROOT_URL` that is
|
||
auto-derived from `forge.domain` + gateway state. The auto-derivation
|
||
covers most cases:
|
||
|
||
| Shape | Auto-derived `ROOT_URL` |
|
||
|---|---|
|
||
| `behindGateway = true` | `https://<forge.domain>/` (port suffix omitted when `gateway.httpsPort == 443`) |
|
||
| `behindGateway = false` | `http://<forge.domain>:<httpPort>/` |
|
||
|
||
The gateway always terminates TLS, so the `behindGateway = true` case is
|
||
always advertised over `https://`; only the direct (`behindGateway =
|
||
false`) shape stays `http://`. Set `rootUrl` explicitly when
|
||
`forge.domain` resolves differently from the public URL, or for a
|
||
genuinely bespoke shape (e.g. an external reverse proxy on a different
|
||
host/path). Must end with `/` (Forgejo requirement; an assertion
|
||
enforces this).
|
||
|
||
## Per-agent static frontend split
|
||
|
||
When `services.hyperhive.frontend` is configured, hive-c0re injects
|
||
`HIVE_AGENT_FRONTEND_DIR = "${cfg.frontend}/agent"` into its service
|
||
environment. The nginx include generator (`gateway_nginx::write`) reads
|
||
this variable and, when set, emits split location blocks per agent
|
||
instead of the legacy single-proxy block.
|
||
|
||
**Location priority for `/agent/<name>/...`:**
|
||
|
||
```nginx
|
||
# 1. Compiled assets — content-addressed nix store path, cache forever
|
||
location ^~ /agent/<name>/static/ {
|
||
alias <frontend>/static/;
|
||
expires 1y;
|
||
add_header Cache-Control "public, immutable, max-age=31536000";
|
||
}
|
||
|
||
# 2. Static dist + proxy fallback for dynamic paths
|
||
location /agent/<name>/ {
|
||
alias <frontend>/;
|
||
try_files $uri $uri.html $uri/index.html @<name>_dynamic;
|
||
}
|
||
|
||
# 3. Proxy catchall — API, events, icon, send, login, …
|
||
location @<name>_dynamic {
|
||
proxy_pass <upstream>;
|
||
proxy_set_header X-Forwarded-Prefix /agent/<name>;
|
||
proxy_intercept_errors on;
|
||
error_page 502 503 504 = /__hive_agent_unreachable;
|
||
# … (full proxy header block)
|
||
}
|
||
```
|
||
|
||
**`try_files` resolution** (nginx applies the `alias` mapping before
|
||
checking file existence):
|
||
|
||
| request | resolved | outcome |
|
||
| --- | --- | --- |
|
||
| `/agent/iris/` | `<frontend>/index.html` | main agent page |
|
||
| `/agent/iris/stats` | `<frontend>/stats.html` | stats page |
|
||
| `/agent/iris/screen` | `<frontend>/screen.html` | screen page |
|
||
| `/agent/iris/static/app.js` | caught by `^~` block first | served with immutable cache |
|
||
| `/agent/iris/api/state` | no file match → `@iris_dynamic` | proxied to agent daemon |
|
||
| `/agent/iris/events/live` | no file match → `@iris_dynamic` | proxied (SSE) |
|
||
|
||
Adding a new HTML page to the frontend dist (`dist/<page>.html`)
|
||
automatically makes it reachable at `/agent/<name>/<page>` — no
|
||
generator change needed.
|
||
|
||
**Why `^~` for `/static/`**: the `^~` prefix gives this block higher
|
||
priority than the plain prefix `location /agent/<name>/`, so compiled
|
||
JS/CSS assets skip `try_files` entirely and get the immutable cache
|
||
headers. Nix store paths are content-addressed — the hash changes on
|
||
any content change — so `max-age=31536000` is safe.
|
||
|
||
**Why the nix store path resolves**: `HIVE_AGENT_FRONTEND_DIR` is a nix
|
||
store path baked in at hive-c0re build time, and c0re (writing
|
||
`agents.conf`) and nginx (serving files from it) are on the same machine,
|
||
so they see the same store.
|
||
|
||
**Graceful degradation**: if `HIVE_AGENT_FRONTEND_DIR` is empty or
|
||
unset (e.g. a build that predates `cfg.frontend`), each agent gets the
|
||
legacy single-proxy block and all traffic is forwarded to the agent
|
||
daemon as before.
|
||
|
||
**`extraFiles`**: per-agent `hyperhive.frontend.extraFiles` are in
|
||
`mergedDist`, not in the base `cfg.frontend` dist. They are not under
|
||
the nix-store `alias` path, so requests for them fall through
|
||
`try_files` to `@<name>_dynamic` and are served by the agent daemon
|
||
as before.
|
||
|
||
## Per-agent error pages
|
||
|
||
`/agent/<name>/` requests hit two failure modes; both get static
|
||
HTML pages instead of nginx's default error chrome:
|
||
|
||
- **Agent not found** (`/agent/<unknown>/...`) — name isn't in
|
||
`agentPortsTable`. nginx's prefix match falls back to the bare
|
||
`/agent/` catch-all, which `return 404`s and `error_page 404` rewrites
|
||
to `/__hive_agent_not_found` → serves `not-found.html` with a link
|
||
back to the dashboard.
|
||
|
||
- **Agent unreachable** (`502 / 503 / 504` from `proxy_pass`) — the
|
||
per-agent harness isn't responding (container restarting, crash
|
||
recovery, etc.). `proxy_intercept_errors on` + `error_page 502 503
|
||
504 = /__hive_agent_unreachable` rewrites to `unreachable.html`.
|
||
|
||
Both pages are built at deploy time via `pkgs.runCommand` (one nix
|
||
derivation `hyperhive-agent-error-pages` with `not-found.html` +
|
||
`unreachable.html` inside) and served via two `internal` nginx
|
||
locations with `alias` to the exact file. `internal` keeps the
|
||
files from being directly request-able by operators — only nginx's
|
||
own error-handling can reach them.
|
||
|
||
Page styling: minimal inline CSS matching the dashboard's catppuccin
|
||
palette (`#1e1e2e` bg, `#cdd6f4` text, `#cba6f7` heading). No
|
||
dependencies on the frontend dist — these pages render even when
|
||
hive-c0re itself is down.
|
||
|
||
Scope is intentionally narrow: a route earns a custom page when the
|
||
default status code would point at the wrong component. The per-agent
|
||
routes qualify (a 502 there means the harness is restarting, not that
|
||
the gateway is broken), and so does `auth.<swarm>` — a dead authelia
|
||
upstream almost always means the user store was never bootstrapped, and
|
||
a bare 502 blames the proxy, which is the one part that is working.
|
||
|
||
Forge / matrix / fluffychat still get nginx defaults: their upstreams
|
||
being down means what the status code says, so a themed page would add
|
||
styling and no information.
|
||
|
||
## HTTP Basic auth
|
||
|
||
`services.hyperhive.gateway.auth.enable = true` gates every request to
|
||
the main vhost (`_`) behind HTTP Basic auth. nginx's built-in `auth_basic`
|
||
module validates credentials; no extra service or host-side daemon is
|
||
required.
|
||
|
||
**Setup:**
|
||
|
||
```nix
|
||
services.hyperhive.gateway.auth = {
|
||
enable = true;
|
||
# realm = "hyperhive"; # optional, default shown
|
||
};
|
||
```
|
||
|
||
The credential store lives at the fixed path
|
||
`/var/lib/hive-gateway/conf/gateway.htpasswd` on the host. A tmpfiles
|
||
rule pre-creates the file on first boot; no manual path configuration
|
||
is required. nginx reads it at that path directly.
|
||
|
||
Manage users with `hivectl gateway`. `hivectl` sends the request over the
|
||
host admin socket and the `hive-c0re` daemon performs the write at its
|
||
canonical path — no path is exposed to the CLI:
|
||
|
||
```sh
|
||
# Add or update a user (prompted for password):
|
||
hivectl gateway create-user alice --password-stdin
|
||
|
||
# Add with inline password (visible in shell history — avoid for sensitive creds):
|
||
hivectl gateway create-user bob --password hunter2
|
||
|
||
# Remove a user:
|
||
hivectl gateway delete-user bob
|
||
|
||
# List current usernames:
|
||
hivectl gateway list-users
|
||
```
|
||
|
||
The daemon hashes passwords with BCrypt (cost 12) and writes
|
||
`$2y$`-prefixed hashes that nginx accepts natively. No external
|
||
`htpasswd` binary is required.
|
||
|
||
**What is not gated:** per-agent UI routes emitted into `agents.conf`
|
||
(served under `/agent/<name>/`) inherit no auth from `/` — nginx
|
||
applies `auth_basic` per-location. Full per-agent coverage is a
|
||
follow-up.
|
||
|
||
**Realm:** the `WWW-Authenticate: Basic realm="..."` string browsers
|
||
display in the credential dialog. Defaults to `"hyperhive"`. Must not
|
||
contain `"` or `$`.
|
||
|
||
**Custom 401 page:** when credentials are absent or wrong, nginx serves
|
||
a Catppuccin-styled `unauthorized.html` page (built into the same Nix
|
||
derivation as the agent error pages) that tells the operator which
|
||
`hivectl` command to run to create a user. The response status is still
|
||
`401` (`error_page 401 =401 /__hive_auth_unauthorized`) so browsers
|
||
present the login dialog on the first visit — users who dismiss the
|
||
dialog see the human-readable hint. The internal exact-match location
|
||
(`= /__hive_auth_unauthorized`) beats `location /` in nginx's prefix
|
||
ordering, preventing the subrequest from looping back through
|
||
`auth_basic`.
|
||
|
||
## Security headers
|
||
|
||
The following headers are emitted at server scope on every gateway
|
||
vhost (`_`, `forge.<swarm-domain>`, `chat.<swarm-domain>`):
|
||
|
||
| Header | Value |
|
||
|--------|-------|
|
||
| `X-Frame-Options` | `SAMEORIGIN` |
|
||
| `X-Content-Type-Options` | `nosniff` |
|
||
| `Referrer-Policy` | `strict-origin-when-cross-origin` |
|
||
|
||
nginx's `add_header` inheritance rule: a `location` block that sets its
|
||
own `add_header` does **not** inherit server-scope headers. API locations
|
||
that carry their own CORS headers (e.g. `/.well-known/matrix/client`,
|
||
`/_matrix/`) are therefore unaffected. HTML-serving and proxy locations
|
||
with no `add_header` of their own pick the security headers up
|
||
automatically.
|
||
|
||
### HSTS (`gateway.hsts`)
|
||
|
||
HSTS is **opt-in** and disabled by default:
|
||
|
||
```nix
|
||
services.hyperhive.gateway.hsts = {
|
||
enable = true; # default: false
|
||
maxAge = 31536000; # default: 1 year (required for preload list)
|
||
includeSubDomains = true; # default: true
|
||
};
|
||
```
|
||
|
||
When enabled, a `Strict-Transport-Security: max-age=...[; includeSubDomains]`
|
||
header is added alongside the other security headers.
|
||
|
||
**Opt-in rationale**: HSTS pins HTTPS in the browser's preload cache;
|
||
enabling it on a deployment that later loses TLS locks browsers out
|
||
until `max-age` expires. Only enable when TLS is permanent.
|
||
|
||
Since the gateway always terminates TLS (see [TLS modes](#tls-modes)
|
||
above), an enabled HSTS header is always served over https — there is no
|
||
TLS-less mode that could violate it.
|
||
|