Last pass of the docs-from-code → docs/ epic (#708). Drops every attribution cookie from docs/ + README.md + CLAUDE.md so the source-tree files no longer reference the issue tracker. Issue threads + commit history retain the references — those are the canonical record. - README.md: drop #701 / #660×2 / #551 from matrix + display-name sections, rephrase to convey the semantics directly - CLAUDE.md: scrub 18 cookies from the file map (#655, #15, #784, #832, #444, #425, #361, #548, #598, #539, #544, #589, #701, #658, #280, #660, #551, #764, #772, #793, #14, #805) - docs/agent-hierarchy.md: drop #658 ×3 (per-agent user is the current shape, not a transition) - docs/conventions.md: drop #571 (replaced with a docs xref to persistence.md::matrix-avatar-sync) - docs/gateway.md: scrub vhost-map table cookies + Sub-domain rationale + Per-agent unix-socket upstream + Self-signed TLS + Firewall posture + HIVE_FORGE_URL + Per-agent error pages sections; drop the trailing 'Sequencing history' issue list + the 'Next-up' issue-link footnote - docs/matrix.md: scrub serverName/gatewayHost + Default-closed firewall + Provisioning flow + Initial rollout + Assertion rationale + fluffychat-web build fixes; drop the trailing 'Sequencing history' issue list - docs/network.md: drop 'Why ship before #14' #805 quote + Container shape #805 attribution + trailing 'Sequencing history' + Cross-references issue links; rename v2 column to 'after netns isolation' - docs/web-ui.md: drop #784 from Container row, replace with a docs xref to docs/gateway.md::Per-agent unix-socket upstream Only remaining #NNN in docs/ is the literal markdown-heading example in docs/forge.md (`#tag`, `#123`, `#!/bin/bash`) which demonstrates the renderer's behaviour — not an attribution cookie.
231 lines
14 KiB
Markdown
231 lines
14 KiB
Markdown
# hive-gateway
|
|
|
|
Single nginx in front of every hyperhive web surface. Container `hive-gateway`, shared host netns, system-config (not meta-flake managed). Configured via `services.hyperhive.gateway.*` + per-subsystem opt-in flags in `services.hyperhive.{forge,matrix,...}`.
|
|
|
|
## Vhost map
|
|
|
|
| URL | vhost | upstream | source |
|
|
| --- | --- | --- | --- |
|
|
| `<hive>/` | `_` (catch-all) | hive-c0re dashboard (`7000`) | always |
|
|
| `<hive>/agent/<name>/` | `_` | per-agent harness on `agent_web_port(name)` | `agentPortsFile` JSON |
|
|
| `<hive>/.well-known/matrix/{client,server}` | `_` | inline JSON (no upstream) | `matrix.enable && domain != null` |
|
|
| `<hive>/matrix/` (deprecated) | `_` | 301 → `matrix.<hive>/` | `matrix.gui.enable` |
|
|
| `forge.<hive>/` | `forge.<hive>` | forgejo (`3000`) | `forge.behindGateway` |
|
|
| `matrix.<hive>/_matrix/*` | `matrix.<hive>` | tuwunel (`8008`) | `matrix.gatewayHost != null` |
|
|
| `matrix.<hive>/` | `matrix.<hive>` | fluffychat-web static | `matrix.gui.enable` |
|
|
| `matrix.<hive>/config.json` | `matrix.<hive>` | inline JSON (FluffyChat boot config) | `matrix.gui.enable && domain != null` |
|
|
|
|
Per-agent UIs stay sub-path because they're hyperhive-internal and base-path-aware. External standard apps (forge / matrix) get sub-domains because their defaults work cleanly at sub-domain root + per-origin cookies / storage isolation matters.
|
|
|
|
## Discovery flow (matrix)
|
|
|
|
Operator points client at `<hive>`. Sequence:
|
|
|
|
1. Client fetches `https://<hive>/.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://matrix.<hive>"}}` (no port suffix when gateway listens on 443). With `selfSignedTls = false` the scheme drops to http and the port suffix reflects the bare `port` instead.
|
|
2. Client connects to `matrix.<hive>/_matrix/client/...`.
|
|
3. Gateway routes `/_matrix/*` → tuwunel at `127.0.0.1:8008`.
|
|
|
|
matrix-dart-sdk (FluffyChat etc.) hardcodes `https` for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path.
|
|
|
|
Federation peers fetch `.well-known/matrix/server` → `{"m.server":"matrix.<hive>"}` and connect to `matrix.<hive>:8448` per spec default. Gateway only listens on configured `port` (+ `httpsPort` when TLS on); cross-hive federation needs either an SRV record (`_matrix._tcp.matrix.<hive>` → port 80 / 443) OR `matrix.openFirewall = true` so peers reach tuwunel's federation port directly. Hyperhive is mostly closed/internal, so this rarely bites.
|
|
|
|
## SPA fallback (Accept-header pattern)
|
|
|
|
The `<hive>` catch-all and the `matrix.<hive>` vhost both serve a flutter SPA (per-agent UI, fluffychat). Two requirements collide:
|
|
|
|
- hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap)
|
|
- missing assets must surface as 404, not as HTML with wrong content-type
|
|
|
|
Solution: an `nginx http`-context `map $http_accept $matrix_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; asset fetches (`Accept: image/*`, `*/*`, etc.) get a sentinel nonexistent path → `try_files` falls through to `=404`. No extension allowlist, no `if` block, no regex heuristics.
|
|
|
|
## Local dev (`localHostsEntry`)
|
|
|
|
`services.hyperhive.gateway.localHostsEntry = true` adds entries to the host's `/etc/hosts`:
|
|
|
|
- `<hive-domain>` → `127.0.0.1`
|
|
- `forge.<hive>` → `127.0.0.1` (when forge.behindGateway)
|
|
- `matrix.<hive>` → `127.0.0.1` (when matrix.gatewayHost set)
|
|
|
|
`lib.unique` de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
|
|
|
|
## Sub-domain shape (rationale)
|
|
|
|
Operator decision: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
|
|
|
|
- forgejo's default `ROOT_URL = http://<host>/` works without any `X-Forwarded-Prefix` gymnastics — sub-domain hosting is the canonical Forgejo deploy shape.
|
|
- matrix-spec deployments universally use `matrix.<server_name>` for the actual API listener — federation already expects this.
|
|
- per-agent UIs are hyperhive-internal and base-path-aware specifically for `/agent/<name>/`. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins.
|
|
- cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
|
|
|
|
`services.hyperhive.{forge.domain,matrix.gatewayHost}` take the full hostname (`forge.darkest.space`, `git.example.com`) rather than a label that gets concatenated with hive-domain — operators want control over the full shape, not a forced `<label>.<hive-domain>` pattern.
|
|
|
|
## Tuning knobs
|
|
|
|
Per-vhost timeouts + body-size limits live in the location blocks:
|
|
|
|
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), `proxyWebsockets = true` (live-update endpoints).
|
|
- matrix `/_matrix/` (tuwunel): `client_max_body_size 50M` (media uploads), `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*` (federation + cross-origin clients), `proxyWebsockets = true`.
|
|
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE / WebSocket dashboards), `proxyWebsockets = true`, `X-Forwarded-Prefix` set so the harness can build absolute URLs when relative isn't enough.
|
|
|
|
SSH for forge stays direct on `cfg.sshPort` — separate listener protocol, not HTTP-over-nginx.
|
|
|
|
## Per-agent unix-socket upstream
|
|
|
|
Sub-agent `/agent/<name>/` upstreams flip from TCP loopback to a
|
|
unix-domain socket as each agent opts in. The mechanism:
|
|
|
|
1. **Agent side** (`hyperhive.web.useUnixSocket = true` in
|
|
`agent.nix`). Sets `HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock`
|
|
on the harness service env; `web_ui::serve` binds a `UnixListener`
|
|
at that path instead of TCP.
|
|
2. **Host side**. `hive-c0re` bind-mounts the per-agent subdir
|
|
(`/run/hive-agent/<name>/`) into the agent's container. Dir
|
|
bind, not file bind — file bind-mounts don't survive the
|
|
harness's `unlink + bind(2)` cycle on socket replace. Per-agent
|
|
subdir keeps each agent's container blind to siblings' sockets.
|
|
3. **Marker gate**. After successful `bind_unix`, the harness drops
|
|
`<dir>/.bound` next to the socket. c0re's `agent_sockets::write`
|
|
filters its JSON map by marker presence — only agents whose
|
|
harness has actually bound the socket appear there. Without this
|
|
filter, the gateway would `proxy_pass` to a non-existent socket
|
|
for every sub-agent that hasn't opted in yet.
|
|
4. **Gateway side**. Reads `agent-sockets.json` at request-handling
|
|
time and routes `/agent/<name>/` to
|
|
`http://unix:/run/hive-agent/<name>/web.sock:/`. Whole
|
|
`/run/hive-agent/` is bind-mounted read-only into the gateway
|
|
container so it can reach every published socket.
|
|
|
|
c0re re-fires `agent_sockets::write` every 10s so newly-bound
|
|
markers get picked up without needing a container-start hook in
|
|
every lifecycle path. `write()` is idempotent: steady-state cost is
|
|
one stat per agent per tick.
|
|
|
|
Transition: agents that haven't flipped `useUnixSocket = true` still
|
|
appear in `agent-ports.json` (the legacy TCP map) and the gateway
|
|
falls back to TCP for them. A future cleanup will drop the TCP map +
|
|
the harness's TCP bind once every agent's flipped.
|
|
|
|
## Dashboard link shape (gateway vs direct)
|
|
|
|
When the gateway is in front, the SW4RM tab builds per-agent links
|
|
as same-origin `/agent/<name>/…` URLs instead of the legacy direct
|
|
`http://<host>:<container.port>/` TCP shape. The signal comes from
|
|
`StateSnapshot.gateway_enabled`, sourced from the
|
|
`HIVE_GATEWAY_ENABLED` env the c0re NixOS module sets when
|
|
`services.hyperhive.gateway.enable = true`. Three render sites
|
|
flip together: the primary agent-name link, the favicon fetch
|
|
(`<url>/icon`), and the nav-strip `container`-kind links from
|
|
`/api/agent/<name>/links`. `forge`-kind nav-strip links still
|
|
resolve against `http://<host>:3000` (separate sub-domain transition
|
|
tracked by `forge.behindGateway`); `external`-kind links are
|
|
already absolute. See `docs/web-ui.md::Container row` for the
|
|
frontend-side derivation.
|
|
|
|
## Self-signed TLS (`selfSignedTls`)
|
|
|
|
On by default. The gateway generates a self-signed RSA-4096 cert at first boot (10-year validity) and listens on `httpsPort` (default 443) on every vhost beside the plain-http `port` (default 80).
|
|
|
|
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) hardcodes `https://<host>/.well-known/matrix/client` for homeserver discovery and refuses to fall back to plain http. Without TLS the browser client cannot bootstrap. For headless agent traffic and operator dashboard reach plain http is fine, so the http listen stays in parallel — operators can keep using `http://<hive>/` from the dashboard if they don't care about the cert prompt.
|
|
|
|
**Cert shape**: subject CN = bare hive domain; subjectAltName covers `<hive>` + wildcard `*.<hive>` so all current and future sub-domain vhosts (matrix, forge, ...) validate under the same cert. Stored at `/var/lib/hive-gateway/tls/{cert,key}.pem` inside the gateway container (`ephemeral = false`, so persisted across container restart).
|
|
|
|
**Regeneration**: the generator unit (`hive-gateway-self-signed-cert.service`) is gated by `ConditionPathExists=!cert.pem`, so it's a one-shot. To rotate (e.g. cert leak, expiry approaching), delete `cert.pem` inside the gateway container and restart `nginx.service` — the generator runs as a `Before=` dependency. No automatic rotation; the cert is purely a workaround for the well-known fetch.
|
|
|
|
**Production**: operators fronting hyperhive with a real reverse proxy (caddy + ACME, traefik + Let's Encrypt, etc.) set `selfSignedTls = false`. The proxy handles termination upstream, the gateway listens on `port` only.
|
|
|
|
**Cert prompts**: browsers warn once per host on first visit. With the wildcard SAN, `https://<hive>/`, `https://matrix.<hive>/`, and `https://forge.<hive>/` are covered by the same cert, but the browser still prompts per origin (per-host security state). FluffyChat needs the user to accept both `matrix.<hive>` (the SPA itself) and `<hive>` (the well-known fetch endpoint).
|
|
|
|
**`.well-known/matrix/{client,server}` scheme**: switches to `https` when `selfSignedTls` is on, so matrix-dart-sdk doesn't downgrade and tuwunel federation peers get a TLS-fronted base_url. The legacy `/matrix/*` 301 redirect also flips to https.
|
|
|
|
## Firewall posture (host-level)
|
|
|
|
`hive-c0re.nix` opens the per-agent web-port range
|
|
`8100..8999` in the host firewall **only when
|
|
`services.hyperhive.gateway.enable = false`**. With the gateway on
|
|
(default), it's the sole external entry point and proxies to
|
|
`127.0.0.1:<port>` internally — leaving the per-agent ports
|
|
firewall-open would defeat the single-front-door story.
|
|
|
|
`services.hyperhive.gateway.openFirewall = true` opens `port` plus
|
|
`httpsPort` when `selfSignedTls = true` (default). Operators who
|
|
flip `selfSignedTls = false` to front the gateway with a real
|
|
TLS-terminating reverse proxy on the host get only `port` opened.
|
|
|
|
The manager hashes into the same port range as sub-agents (no
|
|
"manager pinned at 8000" special case), so one range opening covers
|
|
every container.
|
|
|
|
The dashboard port (`cfg.dashboardPort`, default 7000) is *not*
|
|
listed in either case — it binds `127.0.0.1` only, so a firewall
|
|
hole would be a no-op. Remote dashboard access flows through the
|
|
gateway. Operators who opt out of the gateway lose external
|
|
dashboard reach by design — the surface is privileged (approve /
|
|
deny / destroy) and must not be exposed without a real reverse
|
|
proxy in front.
|
|
|
|
## `HIVE_FORGE_URL`: loopback for in-cluster, sub-domain for the operator
|
|
|
|
Agents poll `HIVE_FORGE_URL` for Forgejo notifications + run all
|
|
`hive-forge` calls against it. `hive-c0re.nix` pins this to
|
|
`http://127.0.0.1:<forge.httpPort>` for the in-cluster path: every
|
|
agent container shares the host's network namespace, so loopback
|
|
reaches the forge container directly with no DNS lookup needed.
|
|
|
|
The sub-domain default (`forge.<hive-domain>`) is for **operator
|
|
browsers + cross-host clients**, not in-cluster traffic. Using the
|
|
sub-domain URL inside agent containers would fail every `hive-forge`
|
|
invocation with "Name or service not known" — the agent's nspawn
|
|
doesn't have DNS for the external hostname.
|
|
|
|
## hive-forge container shape
|
|
|
|
Private Forgejo wrapped in a nixos-container (`hive-forge`, not
|
|
`h-*` — keeps c0re's lifecycle scanner out of the picture; the
|
|
operator manages it via the standard `nixos-container` CLI). The
|
|
container also keeps hive-forge from fighting any `services.forgejo`
|
|
the operator already runs on the host — separate systemd namespace,
|
|
separate state dir, separate port unless the operator deliberately
|
|
collides.
|
|
|
|
Container shares the host network namespace
|
|
(`privateNetwork = false`) so agents reach the forge at
|
|
`http://localhost:<httpPort>` without extra plumbing — nixos-container
|
|
is here for state + systemd-unit isolation, not network isolation.
|
|
|
|
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/`
|
|
and survives container restart / host reboot. To wipe, destroy the
|
|
container.
|
|
|
|
## Per-agent error pages
|
|
|
|
`/agent/<name>/` requests hit two failure modes; both get static
|
|
HTML pages instead of nginx's default error chrome:
|
|
|
|
- **Agent not found** (`/agent/<unknown>/...`) — name isn't in
|
|
`agentPortsTable`. nginx's prefix match falls back to the bare
|
|
`/agent/` catch-all, which `return 404`s and `error_page 404` rewrites
|
|
to `/__hive_agent_not_found` → serves `not-found.html` with a link
|
|
back to the dashboard.
|
|
|
|
- **Agent unreachable** (`502 / 503 / 504` from `proxy_pass`) — the
|
|
per-agent harness isn't responding (container restarting, crash
|
|
recovery, etc.). `proxy_intercept_errors on` + `error_page 502 503
|
|
504 = /__hive_agent_unreachable` rewrites to `unreachable.html`.
|
|
|
|
Both pages are built at deploy time via `pkgs.runCommand` (one nix
|
|
derivation `hyperhive-agent-error-pages` with `not-found.html` +
|
|
`unreachable.html` inside) and served via two `internal` nginx
|
|
locations with `alias` to the exact file. `internal` keeps the
|
|
files from being directly request-able by operators — only nginx's
|
|
own error-handling can reach them.
|
|
|
|
Page styling: minimal inline CSS matching the dashboard's catppuccin
|
|
palette (`#1e1e2e` bg, `#cdd6f4` text, `#cba6f7` heading). No
|
|
dependencies on the frontend dist — these pages render even when
|
|
hive-c0re itself is down.
|
|
|
|
Scope is intentionally narrow: only routes already special-cased in
|
|
the nginx config get custom error pages. Other gateway routes
|
|
(forge / matrix / fluffychat) get nginx defaults — extending the
|
|
custom-error pattern there is a separate follow-up.
|
|
|