Switch per-agent gateway upstreams from TCP loopback to unix-domain socket when the agent has opted in via `hyperhive.web.useUnixSocket` (#822). Coexists with the TCP path during rollout. Changes: - New `agentSocketsFile` option (default `/var/lib/hyperhive/agent-sockets.json`) — c0re writes the map there via `hive_c0re::agent_sockets::write` (#809). - `agentSocketsTable = lib.importJSON ...` (graceful empty when file missing). - `agentUpstreamFor name port` picks `http://unix:<path>:/` when the socket has a JSON entry AND the file exists at eval time; else `http://127.0.0.1:<port>/`. Path-exists gate guards against c0re's blanket-emit shape during the canary window (agents in `agent-sockets.json` who haven't actually flipped have no bound socket on disk → fall back to TCP). Damocles will ship a `.bound` marker filter on the c0re side (#784 step 2d follow-up); once that's in, the path-exists check is redundant but harmless. Step 4 drops it entirely along with the TCP fallback. - `containers.hive-gateway.bindMounts."/run/hive-agent"` — read-only, unconditional. Inert when no agents have opted in. Required so nginx inside the gateway container can `connect(2)` to the per-agent sockets damocles's #813 bind-mounts into agent containers at the same paths. Docs: - `docs/gateway.md::Per-agent UDS upstream (#784)` — full rollout flow, subdir-bind rationale (damocles #813), eval-time gate explainer, step 4 drop plan. `nix flake check` clean; `nix fmt` clean. Canary plan: once #822 (`useUnixSocket` option) lands + this PR merges, manager flips atlas's agent.nix to `useUnixSocket = true` via the config-update flow. End-to-end validation against atlas before broader rollout.
12 KiB
hive-gateway
Single nginx in front of every hyperhive web surface. Container hive-gateway, shared host netns, system-config (not meta-flake managed). Configured via services.hyperhive.gateway.* + per-subsystem opt-in flags in services.hyperhive.{forge,matrix,...}.
Vhost map
| URL | vhost | upstream | source |
|---|---|---|---|
<hive>/ |
_ (catch-all) |
hive-c0re dashboard (7000) |
always |
<hive>/agent/<name>/ |
_ |
per-agent harness on agent_web_port(name) |
agentPortsFile JSON, #15 |
<hive>/.well-known/matrix/{client,server} |
_ |
inline JSON (no upstream) | matrix.enable && domain != null, #660 / #747 |
<hive>/matrix/ (deprecated) |
_ |
301 → matrix.<hive>/ |
matrix.gui.enable, #772 |
forge.<hive>/ |
forge.<hive> |
forgejo (3000) |
forge.behindGateway, #754 |
matrix.<hive>/_matrix/* |
matrix.<hive> |
tuwunel (8008) |
matrix.gatewayHost != null, #764 |
matrix.<hive>/ |
matrix.<hive> |
fluffychat-web static | matrix.gui.enable, #772 |
matrix.<hive>/config.json |
matrix.<hive> |
inline JSON (FluffyChat boot config) | matrix.gui.enable && domain != null, #736 |
Per-agent UIs stay sub-path because they're hyperhive-internal and base-path-aware (iris #731). External standard apps (forge / matrix) get sub-domains because their defaults work cleanly at sub-domain root + per-origin cookies / storage isolation matters.
Discovery flow (matrix)
Operator points client at <hive>. Sequence:
- Client fetches
http://<hive>/.well-known/matrix/client→{"m.homeserver":{"base_url":"http://matrix.<hive>"}}(no port suffix when gateway listens on 80). - Client connects to
matrix.<hive>/_matrix/client/.... - Gateway routes
/_matrix/*→ tuwunel at127.0.0.1:8008.
Federation peers fetch .well-known/matrix/server → {"m.server":"matrix.<hive>"} and connect to matrix.<hive>:8448 per spec default. Gateway only listens on configured port; cross-hive federation needs either an SRV record (_matrix._tcp.matrix.<hive> → port 80) OR matrix.openFirewall = true so peers reach tuwunel's federation port directly. Hyperhive is mostly closed/internal, so this rarely bites.
SPA fallback (Accept-header pattern)
The <hive> catch-all and the matrix.<hive> vhost both serve a flutter SPA (per-agent UI, fluffychat). Two requirements collide:
- hard-refresh on a sub-route must serve
index.html(SPA's client-side router takes over after JS bootstrap) - missing assets must surface as 404, not as HTML with wrong content-type (the original #643 bug)
Solution: an nginx http-context map $http_accept $matrix_spa_target { ... } keyed on the request's Accept header. Browser navigations (Accept: text/html,...) get index.html; asset fetches (Accept: image/*, */*, etc.) get a sentinel nonexistent path → try_files falls through to =404. No extension allowlist, no if block, no regex heuristics. #686 + #729 thread for the design history.
Local dev (localHostsEntry)
services.hyperhive.gateway.localHostsEntry = true adds entries to the host's /etc/hosts:
<hive-domain>→127.0.0.1forge.<hive>→127.0.0.1(when forge.behindGateway)matrix.<hive>→127.0.0.1(when matrix.gatewayHost set)
lib.unique de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
Sub-domain shape (rationale)
mara verdict at #749:9609 + #747:9722: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
- forgejo's default
ROOT_URL = http://<host>/works without anyX-Forwarded-Prefixgymnastics — sub-domain hosting is the canonical Forgejo deploy shape. - matrix-spec deployments universally use
matrix.<server_name>for the actual API listener — federation already expects this. - per-agent UIs are hyperhive-internal; iris's #731 made them base-path-aware specifically for
/agent/<name>/. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins. - cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
services.hyperhive.{forge.domain,matrix.gatewayHost} take the full hostname (forge.darkest.space, git.example.com) rather than a label that gets concatenated with hive-domain — mara on #754:9684 wanted operator control over the full shape, not a forced <label>.<hive-domain> pattern.
Tuning knobs
Per-vhost timeouts + body-size limits live in the location blocks:
- forge
/(forgejo):client_max_body_size 1G(LFS),proxy_read_timeout 1h(multi-GB clones),proxyWebsockets = true(live-update endpoints). - matrix
/_matrix/(tuwunel):client_max_body_size 50M(media uploads),proxy_read_timeout 1h(long-poll/sync), CORS*(federation + cross-origin clients),proxyWebsockets = true. - per-agent
/agent/<name>/:proxy_read_timeout 1d(long-lived SSE / WebSocket dashboards),proxyWebsockets = true,X-Forwarded-Prefixset so the harness can build absolute URLs when relative isn't enough.
SSH for forge stays direct on cfg.sshPort — separate listener protocol, not HTTP-over-nginx.
Sequencing history
- #15 v0 (per-agent routing, #740) — first sub-app behind the gateway, JSON port table from c0re.
- #686 / #729 — Accept-header SPA fallback pattern.
- #749 / #754 — forge to sub-domain (mara: sub-domain over sub-path).
- #747 / #764 — matrix sub-domain vhost +
.well-knowndelegation. - #772 / #775 — fluffychat hops from
<hive>/matrix/tomatrix.<hive>/.
Next-up tracked separately: #14 (container netns isolation), TLS (#594).
Firewall posture (host-level)
hive-c0re.nix opens the per-agent web-port range
8100..8999 in the host firewall only when
services.hyperhive.gateway.enable = false. With the gateway on
(default), it's the sole external entry point and proxies to
127.0.0.1:<port> internally — leaving the per-agent ports
firewall-open would defeat the single-front-door story (closes
#621).
Manager hashes into the same range since #753 (no more "manager pinned at 8000" special case), so one range opening covers every container.
The dashboard port (cfg.dashboardPort, default 7000) is not
listed in either case — since #652 it binds 127.0.0.1 only, so a
firewall hole would be a no-op. Remote dashboard access flows
through the gateway. Operators who opt out of the gateway lose
external dashboard reach by design — the surface is privileged
(approve / deny / destroy) and must not be exposed without a real
reverse proxy in front.
HIVE_FORGE_URL: loopback for in-cluster, sub-domain for the operator
Agents poll HIVE_FORGE_URL for Forgejo notifications + run all
hive-forge calls against it. hive-c0re.nix pins this to
http://127.0.0.1:<forge.httpPort> for the in-cluster path: every
agent container shares the host's network namespace, so loopback
reaches the forge container directly with no DNS lookup needed
(closes #761).
The post-#754 sub-domain default (forge.<hive-domain>) is for
operator browsers + cross-host clients, not in-cluster traffic.
Using the sub-domain URL inside agent containers would fail every
hive-forge invocation with "Name or service not known" — the
agent's nspawn doesn't have DNS for the external hostname.
hive-forge container shape
Private Forgejo wrapped in a nixos-container (hive-forge, not
h-* — keeps c0re's lifecycle scanner out of the picture; the
operator manages it via the standard nixos-container CLI). The
container also keeps hive-forge from fighting any services.forgejo
the operator already runs on the host — separate systemd namespace,
separate state dir, separate port unless the operator deliberately
collides.
Container shares the host network namespace
(privateNetwork = false) so agents reach the forge at
http://localhost:<httpPort> without extra plumbing — nixos-container
is here for state + systemd-unit isolation, not network isolation.
State lives at /var/lib/nixos-containers/hive-forge/var/lib/forgejo/
and survives container restart / host reboot. To wipe, destroy the
container.
Per-agent error pages
/agent/<name>/ requests hit two failure modes; both get static
HTML pages instead of nginx's default error chrome (#755):
-
Agent not found (
/agent/<unknown>/...) — name isn't inagentPortsTable. nginx's prefix match falls back to the bare/agent/catch-all, whichreturn 404s anderror_page 404rewrites to/__hive_agent_not_found→ servesnot-found.htmlwith a link back to the dashboard. -
Agent unreachable (
502 / 503 / 504fromproxy_pass) — the per-agent harness isn't responding (container restarting, crash recovery, etc.).proxy_intercept_errors on+error_page 502 503 504 = /__hive_agent_unreachablerewrites tounreachable.html.
Both pages are built at deploy time via pkgs.runCommand (one nix
derivation hyperhive-agent-error-pages with not-found.html +
unreachable.html inside) and served via two internal nginx
locations with alias to the exact file. internal keeps the
files from being directly request-able by operators — only nginx's
own error-handling can reach them.
Page styling: minimal inline CSS matching the dashboard's catppuccin
palette (#1e1e2e bg, #cdd6f4 text, #cba6f7 heading). No
dependencies on the frontend dist — these pages render even when
hive-c0re itself is down.
Scope is intentionally narrow per mara on #755: "only for routes already special cased in the nginx config". Other gateway routes (forge / matrix / fluffychat) get nginx defaults — extending the custom-error pattern there is a separate follow-up.
Per-agent UDS upstream (#784)
Per-agent /agent/<name>/ upstreams default to TCP loopback
(http://127.0.0.1:<port>/) but each agent can opt in to unix-
domain socket upstream by flipping hyperhive.web.useUnixSocket = true in its agent.nix. Rollout flow:
- Harness binds a
UnixListenerat/run/hive-agent/<name>/web.sockwhenHIVE_WEB_SOCKETis set (PR #800). The env var is set byharness-base.nixfrom theuseUnixSocketoption (#822). - hive-c0re writes a sibling
agent-sockets.jsonnext toagent-ports.json(PR #809) and bind-mounts the per-agent subdir/run/hive-agent/<name>/into each sub-agent container viaset_nspawn_flags(PR #813). Path-shape lives inhive_c0re::agent_sockets::socket_path_for(name)— one canonical derivation, no triangulation across the c0re / harness / gateway boundaries. - Gateway reads both
agentPortsFile+agentSocketsFileat deploy time. Per agent: a socket entry beats the TCP port. The gateway container bind-mounts/run/hive-agent/read-only so nginx inside canconnect(2)to the per-agent sockets.
Mixed state during rollout: agents flip per-agent. Agents that
haven't opted in keep the TCP path; agents that have flipped use
the UDS path. The two coexist on the same gateway with zero
per-agent special-casing in the nginx config (agentUpstreamFor
resolves the right shape from the JSON maps).
Eval-time gate during the rollout window: agentUpstreamFor
checks builtins.pathExists on the socket path before picking the
UDS upstream. c0re's agent_sockets::write emits an entry for every
sub-agent regardless of whether they've actually flipped, so the
gateway has no other signal that a given agent is or isn't actually
binding the socket. The path-exists check works because a flipped
agent's harness binds the socket on container start, and the
gateway-container rebuild (which re-runs nix eval) happens on every
topology change — so a freshly-flipped agent flips through TCP →
UDS over one rebuild cycle. Once c0re's .bound marker filter
ships (#784 step 2d follow-up), agent-sockets.json only contains
agents that have actually bound, and the path-exists check is
redundant but harmless. Step 4 drops it.
Why per-agent subdir (not a flat /run/hive-agent/<name>.sock):
the harness's bind_unix helper unlinks any stale socket before
calling bind(2), and a file bind-mount loses its host-side anchor
on unlink. Dir bind-mount keeps the same dir inode visible on both
sides, so the new web.sock shows up on the host the moment the
harness binds it (damocles #813 design note).
Step 4 plan: once every agent has flipped + soaked, the
agentPortsFile fallback drops + the harness's TCP bind goes away
entirely. Tracked at #784 step 4.