15 KiB
hive-gateway
Single nginx in front of every hyperhive web surface. Container hive-gateway, shared host netns, system-config (not meta-flake managed). Configured via services.hyperhive.gateway.* + per-subsystem opt-in flags in services.hyperhive.{forge,matrix,...}.
Vhost map
| URL | vhost | upstream | source |
|---|---|---|---|
<hive>/ |
_ (catch-all) |
hive-c0re dashboard (7000) |
always |
<hive>/agent/<name>/ |
_ |
per-agent harness on agent_web_port(name) |
agentPortsFile JSON, #15 |
<hive>/.well-known/matrix/{client,server} |
_ |
inline JSON (no upstream) | matrix.enable && domain != null, #660 / #747 |
<hive>/matrix/ (deprecated) |
_ |
301 → matrix.<hive>/ |
matrix.gui.enable, #772 |
forge.<hive>/ |
forge.<hive> |
forgejo (3000) |
forge.behindGateway, #754 |
matrix.<hive>/_matrix/* |
matrix.<hive> |
tuwunel (8008) |
matrix.gatewayHost != null, #764 |
matrix.<hive>/ |
matrix.<hive> |
fluffychat-web static | matrix.gui.enable, #772 |
matrix.<hive>/config.json |
matrix.<hive> |
inline JSON (FluffyChat boot config) | matrix.gui.enable && domain != null, #736 |
Per-agent UIs stay sub-path because they're hyperhive-internal and base-path-aware (iris #731). External standard apps (forge / matrix) get sub-domains because their defaults work cleanly at sub-domain root + per-origin cookies / storage isolation matters.
Discovery flow (matrix)
Operator points client at <hive>. Sequence:
- Client fetches
https://<hive>/.well-known/matrix/client→{"m.homeserver":{"base_url":"https://matrix.<hive>"}}(no port suffix when gateway listens on 443). WithselfSignedTls = falsethe scheme drops to http and the port suffix reflects the bareportinstead. - Client connects to
matrix.<hive>/_matrix/client/.... - Gateway routes
/_matrix/*→ tuwunel at127.0.0.1:8008.
matrix-dart-sdk (FluffyChat etc.) hardcodes https for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path.
Federation peers fetch .well-known/matrix/server → {"m.server":"matrix.<hive>"} and connect to matrix.<hive>:8448 per spec default. Gateway only listens on configured port (+ httpsPort when TLS on); cross-hive federation needs either an SRV record (_matrix._tcp.matrix.<hive> → port 80 / 443) OR matrix.openFirewall = true so peers reach tuwunel's federation port directly. Hyperhive is mostly closed/internal, so this rarely bites.
SPA fallback (Accept-header pattern)
The <hive> catch-all and the matrix.<hive> vhost both serve a flutter SPA (per-agent UI, fluffychat). Two requirements collide:
- hard-refresh on a sub-route must serve
index.html(SPA's client-side router takes over after JS bootstrap) - missing assets must surface as 404, not as HTML with wrong content-type (the original #643 bug)
Solution: an nginx http-context map $http_accept $matrix_spa_target { ... } keyed on the request's Accept header. Browser navigations (Accept: text/html,...) get index.html; asset fetches (Accept: image/*, */*, etc.) get a sentinel nonexistent path → try_files falls through to =404. No extension allowlist, no if block, no regex heuristics. #686 + #729 thread for the design history.
Local dev (localHostsEntry)
services.hyperhive.gateway.localHostsEntry = true adds entries to the host's /etc/hosts:
<hive-domain>→127.0.0.1forge.<hive>→127.0.0.1(when forge.behindGateway)matrix.<hive>→127.0.0.1(when matrix.gatewayHost set)
lib.unique de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
Sub-domain shape (rationale)
mara verdict at #749:9609 + #747:9722: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
- forgejo's default
ROOT_URL = http://<host>/works without anyX-Forwarded-Prefixgymnastics — sub-domain hosting is the canonical Forgejo deploy shape. - matrix-spec deployments universally use
matrix.<server_name>for the actual API listener — federation already expects this. - per-agent UIs are hyperhive-internal; iris's #731 made them base-path-aware specifically for
/agent/<name>/. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins. - cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
services.hyperhive.{forge.domain,matrix.gatewayHost} take the full hostname (forge.darkest.space, git.example.com) rather than a label that gets concatenated with hive-domain — mara on #754:9684 wanted operator control over the full shape, not a forced <label>.<hive-domain> pattern.
Tuning knobs
Per-vhost timeouts + body-size limits live in the location blocks:
- forge
/(forgejo):client_max_body_size 1G(LFS),proxy_read_timeout 1h(multi-GB clones),proxyWebsockets = true(live-update endpoints). - matrix
/_matrix/(tuwunel):client_max_body_size 50M(media uploads),proxy_read_timeout 1h(long-poll/sync), CORS*(federation + cross-origin clients),proxyWebsockets = true. - per-agent
/agent/<name>/:proxy_read_timeout 1d(long-lived SSE / WebSocket dashboards),proxyWebsockets = true,X-Forwarded-Prefixset so the harness can build absolute URLs when relative isn't enough.
SSH for forge stays direct on cfg.sshPort — separate listener protocol, not HTTP-over-nginx.
Per-agent unix-socket upstream (#784)
Sub-agent /agent/<name>/ upstreams flip from TCP loopback to a
unix-domain socket as each agent opts in. The mechanism:
- Agent side (
hyperhive.web.useUnixSocket = trueinagent.nix, #815). SetsHIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sockon the harness service env;web_ui::servebinds aUnixListenerat that path instead of TCP. - Host side.
hive-c0rebind-mounts the per-agent subdir (/run/hive-agent/<name>/) into the agent's container (#813). Dir bind, not file bind — file bind-mounts don't survive the harness'sunlink + bind(2)cycle on socket replace. Per-agent subdir keeps each agent's container blind to siblings' sockets (mara on #800). - Marker gate. After successful
bind_unix, the harness drops<dir>/.boundnext to the socket. c0re'sagent_sockets::writefilters its JSON map by marker presence — only agents whose harness has actually bound the socket appear there (#784 atlas gate). Without this filter, the gateway wouldproxy_passto a non-existent socket for every sub-agent that hasn't opted in yet. - Gateway side (#829). Reads
agent-sockets.jsonat request-handling time and routes/agent/<name>/tohttp://unix:/run/hive-agent/<name>/web.sock:/. Whole/run/hive-agent/is bind-mounted read-only into the gateway container so it can reach every published socket.
c0re re-fires agent_sockets::write every 10s so newly-bound
markers get picked up without needing a container-start hook in
every lifecycle path. write() is idempotent: steady-state cost is
one stat per agent per tick.
Transition: agents that haven't flipped useUnixSocket = true still
appear in agent-ports.json (the legacy TCP map) and the gateway
falls back to TCP for them. Step 4 of #784 will drop the TCP map +
the harness's TCP bind once every agent's flipped.
Dashboard link shape (gateway vs direct)
When the gateway is in front, the SW4RM tab builds per-agent links
as same-origin /agent/<name>/… URLs instead of the legacy direct
http://<host>:<container.port>/ TCP shape. The signal comes from
StateSnapshot.gateway_enabled, sourced from the
HIVE_GATEWAY_ENABLED env the c0re NixOS module sets when
services.hyperhive.gateway.enable = true. Three render sites
flip together: the primary agent-name link, the favicon fetch
(<url>/icon), and the nav-strip container-kind links from
/api/agent/<name>/links. forge-kind nav-strip links still
resolve against http://<host>:3000 (separate sub-domain transition
tracked by forge.behindGateway); external-kind links are
already absolute. See docs/web-ui.md::Container row for the
frontend-side derivation.
Sequencing history
- #15 v0 (per-agent routing, #740) — first sub-app behind the gateway, JSON port table from c0re.
- #686 / #729 — Accept-header SPA fallback pattern.
- #749 / #754 — forge to sub-domain (mara: sub-domain over sub-path).
- #747 / #764 — matrix sub-domain vhost +
.well-knowndelegation. - #772 / #775 — fluffychat hops from
<hive>/matrix/tomatrix.<hive>/. - #784 / #800 / #813 / #815 / #822 / #829 — sub-agent UI flips to unix-domain socket upstream, opt-in per agent.
Next-up tracked separately: #14 (container netns isolation), TLS (#594).
Self-signed TLS (selfSignedTls)
On by default since #837. The gateway generates a self-signed RSA-4096 cert at first boot (10-year validity) and listens on httpsPort (default 443) on every vhost beside the plain-http port (default 80).
Why on by default: matrix-dart-sdk (FluffyChat's SDK) hardcodes https://<host>/.well-known/matrix/client for homeserver discovery and refuses to fall back to plain http. Without TLS the browser client cannot bootstrap (#837). For headless agent traffic and operator dashboard reach plain http is fine, so the http listen stays in parallel — operators can keep using http://<hive>/ from the dashboard if they don't care about the cert prompt.
Cert shape: subject CN = bare hive domain; subjectAltName covers <hive> + wildcard *.<hive> so all current and future sub-domain vhosts (matrix, forge, ...) validate under the same cert. Stored at /var/lib/hive-gateway/tls/{cert,key}.pem inside the gateway container (ephemeral = false, so persisted across container restart).
Regeneration: the generator unit (hive-gateway-self-signed-cert.service) is gated by ConditionPathExists=!cert.pem, so it's a one-shot. To rotate (e.g. cert leak, expiry approaching), delete cert.pem inside the gateway container and restart nginx.service — the generator runs as a Before= dependency. No automatic rotation; the cert is purely a workaround for the well-known fetch.
Production: operators fronting hyperhive with a real reverse proxy (caddy + ACME, traefik + Let's Encrypt, etc.) set selfSignedTls = false. The proxy handles termination upstream, the gateway listens on port only.
Cert prompts: browsers warn once per host on first visit. With the wildcard SAN, https://<hive>/, https://matrix.<hive>/, and https://forge.<hive>/ are covered by the same cert, but the browser still prompts per origin (per-host security state). FluffyChat needs the user to accept both matrix.<hive> (the SPA itself) and <hive> (the well-known fetch endpoint).
.well-known/matrix/{client,server} scheme: switches to https when selfSignedTls is on, so matrix-dart-sdk doesn't downgrade and tuwunel federation peers get a TLS-fronted base_url. The legacy /matrix/* 301 redirect also flips to https.
Firewall posture (host-level)
hive-c0re.nix opens the per-agent web-port range
8100..8999 in the host firewall only when
services.hyperhive.gateway.enable = false. With the gateway on
(default), it's the sole external entry point and proxies to
127.0.0.1:<port> internally — leaving the per-agent ports
firewall-open would defeat the single-front-door story (closes
#621).
services.hyperhive.gateway.openFirewall = true opens port plus
httpsPort when selfSignedTls = true (default). Operators who
flip selfSignedTls = false to front the gateway with a real
TLS-terminating reverse proxy on the host get only port opened.
Manager hashes into the same range since #753 (no more "manager pinned at 8000" special case), so one range opening covers every container.
The dashboard port (cfg.dashboardPort, default 7000) is not
listed in either case — since #652 it binds 127.0.0.1 only, so a
firewall hole would be a no-op. Remote dashboard access flows
through the gateway. Operators who opt out of the gateway lose
external dashboard reach by design — the surface is privileged
(approve / deny / destroy) and must not be exposed without a real
reverse proxy in front.
HIVE_FORGE_URL: loopback for in-cluster, sub-domain for the operator
Agents poll HIVE_FORGE_URL for Forgejo notifications + run all
hive-forge calls against it. hive-c0re.nix pins this to
http://127.0.0.1:<forge.httpPort> for the in-cluster path: every
agent container shares the host's network namespace, so loopback
reaches the forge container directly with no DNS lookup needed
(closes #761).
The post-#754 sub-domain default (forge.<hive-domain>) is for
operator browsers + cross-host clients, not in-cluster traffic.
Using the sub-domain URL inside agent containers would fail every
hive-forge invocation with "Name or service not known" — the
agent's nspawn doesn't have DNS for the external hostname.
hive-forge container shape
Private Forgejo wrapped in a nixos-container (hive-forge, not
h-* — keeps c0re's lifecycle scanner out of the picture; the
operator manages it via the standard nixos-container CLI). The
container also keeps hive-forge from fighting any services.forgejo
the operator already runs on the host — separate systemd namespace,
separate state dir, separate port unless the operator deliberately
collides.
Container shares the host network namespace
(privateNetwork = false) so agents reach the forge at
http://localhost:<httpPort> without extra plumbing — nixos-container
is here for state + systemd-unit isolation, not network isolation.
State lives at /var/lib/nixos-containers/hive-forge/var/lib/forgejo/
and survives container restart / host reboot. To wipe, destroy the
container.
Per-agent error pages
/agent/<name>/ requests hit two failure modes; both get static
HTML pages instead of nginx's default error chrome (#755):
-
Agent not found (
/agent/<unknown>/...) — name isn't inagentPortsTable. nginx's prefix match falls back to the bare/agent/catch-all, whichreturn 404s anderror_page 404rewrites to/__hive_agent_not_found→ servesnot-found.htmlwith a link back to the dashboard. -
Agent unreachable (
502 / 503 / 504fromproxy_pass) — the per-agent harness isn't responding (container restarting, crash recovery, etc.).proxy_intercept_errors on+error_page 502 503 504 = /__hive_agent_unreachablerewrites tounreachable.html.
Both pages are built at deploy time via pkgs.runCommand (one nix
derivation hyperhive-agent-error-pages with not-found.html +
unreachable.html inside) and served via two internal nginx
locations with alias to the exact file. internal keeps the
files from being directly request-able by operators — only nginx's
own error-handling can reach them.
Page styling: minimal inline CSS matching the dashboard's catppuccin
palette (#1e1e2e bg, #cdd6f4 text, #cba6f7 heading). No
dependencies on the frontend dist — these pages render even when
hive-c0re itself is down.
Scope is intentionally narrow per mara on #755: "only for routes already special cased in the nginx config". Other gateway routes (forge / matrix / fluffychat) get nginx defaults — extending the custom-error pattern there is a separate follow-up.