Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/networking/gateway.md
atlas 54556e4661 docs(networking): fix round — argus review + audit E1/E2
observability.md: "Why two tiers" claimed the harness currently writes an
upstream token into the agent's own claude settings; that path was removed
with the direct-export mode it served (nix/agent-modules/otel.nix:56-63).
Reworded as the hypothetical the paragraph is actually making.

gateway.md: restored the agent-trust pointer to
/run/hive-ca/trust-bundle.pem in "Cert prompts" (hive-ca-trust.nix:41),
dropped by the earlier rewrite. Corrected the SPA-fallback section: only
the chat.<swarm> vhost uses the Accept-header map
(hive-matrix.nix:693-696); per-agent split mode uses file-existence
try_files (gateway_nginx.rs:93-134), not the same mechanism.

Refs #3902
2026-10-02 17:25:36 +02:00

32 KiB
Raw Blame History

hive-gateway

Every host's nginx: the one front door for whatever this host serves. A swarm service running here (forge, matrix, SSO, the swarm UI, the metrics and log stores) declares its own vhost through the gateway; the gateway itself adds the hive's own surface — dashboard, per-agent UIs, matrix discovery.

For the operator configuring services.hyperhive.gateway.* on a host. nginx and the hive resolver (dnsmasq) run on the host next to hive-c0re, not in a container: they bind :80/:443 and the bridge address, so a network namespace of their own would isolate nothing.

You rarely switch it on yourself. gateway.enable defaults to off, and every module that serves a vhost or needs hive names to resolve sets gateway.enable / gateway.dns.enable with mkDefault true — the hive controller, each swarm service, CI.

Vhost map

Swarm services. Each module declares its vhost on the host that runs the service, under the swarm domain:

URL upstream declared by, when
<swarm>/ swarm-ui dist (static), behind an authelia subrequest swarm-ui.nix, deploy.swarm-ui.enable
auth.<swarm>/ authelia (9091) swarm-authelia.nix, deploy.authelia.enable
forge.<swarm>/ forgejo (3000) hive-forge/, deploy.forgejo.behindGateway
chat.<swarm>/_matrix/* tuwunel (8008) hive-matrix.nix, swarm.matrix.gatewayHost != null
chat.<swarm>/ fluffychat-web static (404 with the GUI off) hive-matrix.nix, deploy.matrix.gui.enable
chat.<swarm>/config.json inline JSON (FluffyChat boot config) hive-matrix.nix, deploy.matrix.gui.enable
grafana.<swarm>, metrics.<swarm>, logs.<swarm>, otel.<swarm>, bao.<swarm> the matching swarm service that service's module → swarm/services.md

The hive's own vhost, named for the hive domain:

URL upstream when
<hive>/ dashboard dist (static, from servedFrontend) always
<hive>/api/, /webhook/, /health/ hive-c0re (7000) always
<hive>/api/docs/ themed Swagger UI dist (static) always
<hive>/agent/<name>/ per-agent harness over its unix socket agents.conf (runtime-generated)
<hive>/.well-known/matrix/{client,server} inline JSON deploy.matrix.enable
<hive>/matrix/ (deprecated) 301 → chat.<swarm>/ deploy.matrix.gui.enable and gatewayHost set

The catch-all _ vhost answers any other Host with 444 (connection closed, no response). It's mkDefault, so to make your own vhost the default server, set services.nginx.virtualHosts."_".default = false; — an eval assertion names both when two claim it.

Only the host that runs authelia declares auth.<swarm>; a hive that merely uses SSO knows swarm.authelia.url but doesn't answer for that name. The server name must be exactly swarm.authelia.domain — authelia checks that authelia_url sits inside its session cookie domain at startup, and refuses to boot otherwise. The vhost carries no Basic auth (that would put the login page behind the login it replaces) and passes X-Forwarded-{Proto,Host,Uri,For}, because authelia decides by the original request.

⚠️ auth.<swarm> answering with the sso unavailable page means authelia isn't answering at all — check journalctl -M swarm-authelia -u authelia-swarm. A swarm with no real accounts yet boots fine (a disabled placeholder user keeps authelia's user store non-empty) and serves a login page that refuses everyone; add the first account with swarmctl user add … (setup.md).

Per-agent UIs stay sub-path; forge and matrix get sub-domains → Sub-domain shape.

TLS modes

The gateway always terminates TLS: there is no http-only mode. Which certificate it serves depends on what you configure:

mode config cert source .well-known scheme
self-signed (default) neither tls.certDir nor tls.acme set the host's hive CA signs a gateway leaf (RSA-4096) https
ACME (Let's Encrypt) tls.acme.enable = true nginx via HTTP-01 https
operator cert tls.certDir set read from the operator's dir https

tls.certDir together with tls.acme.enable fails an assertion — pick one.

⚠️ One vhost class has no TLS: a vhost bound to loopback for a local consumer. grafana-metrics (nix/host-modules/swarm-grafana.nix) is the one in the tree: it listens on 127.0.0.1 only and serves = /metrics from grafana's unix socket for this host's collector. Everything with a routable name follows the table above.

ACME / Let's Encrypt (tls.acme)

The simplest production path for a public domain:

services.hyperhive.gateway = {
  openFirewall  = true;
  tls.acme = {
    enable = true;
    email  = "admin@example.com";   # required
  };
};

nginx obtains and renews certs via the HTTP-01 challenge on port (default 80); they land in /var/lib/acme/, managed by nixpkgs's security.acme. Every name this host serves needs a public DNS record pointing here — the hive domain and each swarm-service name in the vhost map — and openFirewall = true so Let's Encrypt reaches /.well-known/acme-challenge/.

Self-signed TLS (default)

On by default, listening on httpsPort (default 443) beside the plain-http port (default 80).

The issuer is a host-held hive CA, not a bare self-signed leaf. hive-tls-ca.service (from hive-tls.nix) generates a long-lived CA (services.hyperhive.deploy.hive-controller.tls.caValidityDays, default 7300 days) under …tls.stateDir (default /var/lib/hive-tls) and signs a gateway leaf with it (leafValidityDays, default 30). hive-gateway-self-signed-cert then imports the leaf into nginx's state dir (/var/lib/hive-gateway/tls/{cert,key}.pem).

⚠️ Keep the import unit. It does two jobs nginx needs: it re-modes the key to 0640 root:nginx (nginx's pre-start nginx -t runs as the nginx user and fails on the CA's 0600 root:root key), and it makes sure every cert path the config names exists — when the swarm-services leaf is missing it installs the hive leaf in its place. nginx refuses a config naming a missing cert file, so without that fallback one missing leaf takes down every vhost, not just one.

Why a CA, not a bare leaf: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust, and nothing can wire a runtime-generated leaf into an agent's build-time trust store. Agents and federation peers trust the stable CA once; leaf rotation never breaks them.

What consumers trust: trust-bundle.pem in the same state dir, not ca.pem. The hive CA is an intermediate under the swarm root (swarm/ca.md has the hierarchy), and a verifier can't stop at an intermediate — so the bundle carries the hive CA plus its root. nginx serves the leaf with the hive CA appended for the same reason. Agents (via security.pki.certificateFiles), the CI and forge containers and federating peers all read the bundle.

Why on by default: matrix-dart-sdk (FluffyChat's SDK) fetches https://<host>/.well-known/matrix/client and never falls back to plain http, so without TLS the browser client can't bootstrap.

Cert shape: the leaf's CN is the bare hive domain; its SANs are <hive> and *.<hive>. The hive CA is name-constrained to <hive>, so it can't sign a swarm-service name outside it — a violating SAN would invalidate the whole leaf. Those names get the swarm-services leaf instead (swarm/ca.md).

Rotation: hive-tls-ca.service re-signs the leaf when it's missing, within 30 days of expiry, or no longer covers the configured names, always under the same CA. It regenerates the CA only if missing or expired. To force a leaf rotation, delete gateway.pem under the state dir, restart the unit, then reload nginx.

Cert prompts: browsers warn once per host until you add the hive's trust-bundle.pem (the anchor, not the leaf) to the browser or OS trust store. A container that needs to trust it for its own outbound TLS gets the bundle a different way, bind-mounted read-only at /run/hive-ca/trust-bundle.pem (nix/host-modules/lib/hive-ca-trust.nix).

Operator-provided cert (tls.certDir)

For a cert from a real CA (Let's Encrypt via your own security.acme, a corporate CA):

services.hyperhive.gateway = {
  tls.certDir = "/var/lib/acme/example.com";  # nixpkgs security.acme output dir
  # tls.certName = "cert.pem";   # default — matches security.acme layout
  # tls.keyName  = "key.pem";    # default — matches security.acme layout
};

nginx reads the directory directly. Keep the key readable by nginx:

  • security.acme writes keys 0640 root:acme, which the nginx user can't read. Set security.acme.certs."example.com".group = "nginx"; (or make the key 0644 if your threat model allows). Otherwise nginx fails at startup with the reason in the journal.

Fronting with an external TLS terminator

The gateway has no plain-http upstream mode. Either give the gateway the real cert (tls.certDir or tls.acme) so it serves proper TLS itself, or front it over a unix socket rather than a plain-http TCP port. .well-known/matrix/* responses always advertise https (Discovery flow).

HTTP Basic auth

Optional: Basic auth on the hive's dashboard. Add a login first, then enable it — the htpasswd file exists from first boot, and an empty one refuses everyone.

On the hive host:

hivectl gateway create-user alice --password-stdin   # password on stdin
hivectl gateway delete-user bob
hivectl gateway list-users
services.hyperhive.gateway.auth = {
  enable = true;
  # realm = "hyperhive";  # default; must not contain `"` or `$`
};

hivectl asks hive-c0re over the host admin socket, and the daemon writes /var/lib/hive-gateway/conf/gateway.htpasswd itself, bcrypt (cost 12) with $2y$ hashes nginx reads natively. --password <pw> also works but lands in shell history.

What it gates: /, /api/ and /api/docs/ on the hive vhost. Not gated: /webhook/ (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), /health/ (for uptime monitors; status only), /.well-known/matrix/*, and the per-agent /agent/<name>/ routes, which come from agents.conf and inherit no auth from /.

A failed or missing login gets 401 with a styled unauthorized.html naming the hivectl command to run, so browsers still show the login dialog first.

Firewall posture (host-level)

services.hyperhive.gateway.openFirewall = true opens port and httpsPort on the host firewall — both, since the gateway always serves TLS. It defaults to off; set it for any reach from outside the host.

nginx is the one external entry point. The per-agent web-port range (8100–8999) stays closed: agents serve their UI on a unix socket (Per-agent unix-socket upstream), and the hashed TCP port (lifecycle::agent_web_port) is only a fallback bind for an agent missing HIVE_WEB_SOCKET — the gateway never proxies through it.

The dashboard port (services.hyperhive.c0re.dashboardPort, default 7000) binds 127.0.0.1 only, so remote dashboard access goes through the gateway. The dashboard can approve, deny and destroy; don't expose it without a reverse proxy in front.

HIVE_FORGE_URL: agents reach the forge via the gateway by domain

Agents poll HIVE_FORGE_URL for Forgejo notifications and run every hive-forge call against it. nix/host-modules/hive-c0re/environment.nix sets it to http://<swarm.forge.domain> (default forge.<swarm-domain> — a swarm runs one forge). Agents run in a private netns and can't reach the host's loopback, so they resolve that name through the bridge dnsmasq to the bridge IP and reach nginx on port 80 (the bridge firewall opens 80 and 443) — the same path an operator's browser takes.

hive-forge container shape

Private Forgejo in a nixos-container named hive-forge (not h-*, so c0re's lifecycle scanner leaves it alone; manage it with the standard nixos-container CLI). The container keeps it from colliding with any services.forgejo the operator already runs on the host: separate systemd namespace, separate state dir, separate port.

It shares the host network namespace (privateNetwork = false): the container is for state and unit isolation, not network isolation. Agent containers, by contrast, are network-isolated and reach the forge through the gateway (HIVE_FORGE_URL).

State lives at /var/lib/nixos-containers/hive-forge/var/lib/forgejo/ and survives restarts and reboots; destroy the container to wipe it.

Only the swarm's forge host runs it (services.hyperhive.deploy.forgejo.enable, see swarm/services.md). Every other hive reaches that host's gateway by forge.<swarm-domain>.

Network and port configuration

services.hyperhive.swarm.forge = {
  httpPort = 3000;   # default — outside hyperhive's 7000 / 8100-8999
  sshPort  = 2222;   # default — git-over-SSH, off the host's openssh on 22
};
# Which ports the forge answers on is swarm-wide; whether THIS host opens
# them in its firewall is a deployment decision, so it lives under deploy.*.
services.hyperhive.deploy.forgejo.openFirewall = false; # default

sshPort serves git clone/push/pull over SSH (git@<domain>:owner/repo.git with -p 2222). SSH goes straight to Forgejo, not through nginx.

openFirewall (default false) opens httpPort and sshPort on the host. Agents don't need it — they come through the gateway. Set it for a browser reaching http://<host>:<httpPort>/ directly, or for external git clients pushing over SSH. The forge vhost behind the gateway (deploy.forgejo.behindGateway, default true) needs only the gateway's own openFirewall.

rootUrl override

services.hyperhive.swarm.forge.rootUrl = "https://forge.example.com/";

rootUrl (default null) overrides the Forgejo ROOT_URL derived from forge.domain and the gateway:

Shape Derived ROOT_URL
deploy.forgejo.behindGateway = true https://<forge.domain>/ (no port suffix when gateway.httpsPort == 443)
deploy.forgejo.behindGateway = false http://<forge.domain>:<httpPort>/

Set it when forge.domain differs from the public URL, or for a bespoke shape such as an external reverse proxy on another host or path. It must end with / (an assertion enforces this).

Security headers

Every named vhost sets these at server scope (the _ catch-all only closes connections):

Header Value
X-Frame-Options SAMEORIGIN
X-Content-Type-Options nosniff
Referrer-Policy strict-origin-when-cross-origin

nginx doesn't merge add_header: a location that sets its own (the CORS locations /.well-known/matrix/client and /_matrix/) inherits none of the server-scope headers, so those locations repeat them. Locations with no add_header of their own pick them up.

HSTS (gateway.hsts)

Opt-in, off by default:

services.hyperhive.gateway.hsts = {
  enable = true;             # default: false
  maxAge = 31536000;         # default: 1 year (required for preload list)
  includeSubDomains = true;  # default: true
};

When enabled, every vhost adds Strict-Transport-Security: max-age=…[; includeSubDomains]. HSTS pins https in the browser: a deployment that later loses TLS locks browsers out until max-age expires, so enable it only when TLS is permanent.

Local dev (localHostsEntry)

services.hyperhive.gateway.localHostsEntry = true maps to 127.0.0.1 in the host's /etc/hosts:

  • the hive domain;
  • every name a module on this host contributes to gateway.localNames — each swarm service this host runs adds its own (forge.<swarm> when behind the gateway, chat.<swarm>, auth.<swarm>, the swarm UI's apex, …).

services.hyperhive.deploy.singleHostSwarm turns it on. Leave it off with real DNS.

Discovery flow (matrix)

The operator points a client at <hive>:

  1. The client fetches https://<hive>/.well-known/matrix/client → {"m.homeserver":{"base_url":"https://chat.<swarm>"}} (no port suffix when httpsPort is 443).
  2. It connects to chat.<swarm>/_matrix/client/....
  3. The gateway proxies /_matrix/* to tuwunel at 127.0.0.1:8008.

matrix-dart-sdk (FluffyChat and others) always fetches the well-known over https, which is why self-signed TLS is on by default.

Federation peers fetch .well-known/matrix/server → {"m.server":"chat.<swarm>:<httpsPort>"}. The port is always explicit, even 443: a delegated host without a port means the federation default 8448, not 443. Peers then reach /_matrix/ on the chat vhost through the gateway, so the gateway must be reachable from them (gateway.openFirewall).

⚠️ The gateway serves both .well-known routes on the hive vhost. Matrix looks them up at the serverName, which defaults to the bare swarm domain, and the swarm UI's apex vhost serves no .well-known/matrix/* route.

Sub-domain shape (rationale)

Sub-domain for forge and matrix, sub-path for per-agent UIs:

  • forgejo's default ROOT_URL = http://<host>/ works without X-Forwarded-Prefix handling; sub-domain hosting is the canonical Forgejo shape.
  • matrix deployments put the API listener on its own name, and federation expects that. gatewayHost defaults to chat.<swarm-domain> rather than the spec-conventional matrix.<server_name>; set it explicitly for the conventional label.
  • per-agent UIs are hyperhive-internal and base-path-aware for /agent/<name>/. A sub-domain per agent would multiply DNS and TLS cost for nothing.
  • cookie and storage isolation: a forge XSS can't reach the dashboard session, because they're different origins.

services.hyperhive.swarm.forge.domain and services.hyperhive.swarm.matrix.gatewayHost take a full hostname (forge.darkest.space, git.example.com), not a label glued to a domain.

Tuning knobs

Per-vhost timeouts and body-size limits live in the location blocks:

  • forge / (forgejo): client_max_body_size 1G (LFS), proxy_read_timeout 1h (multi-GB clones), websockets on.
  • matrix /_matrix/ (tuwunel): client_max_body_size = deploy.matrix.maxRequestSize + 1 MiB, so tuwunel stays the tighter limit and returns a matrix error a client can act on; proxy_read_timeout 1h (long-poll /sync), CORS *, websockets on.
  • per-agent /agent/<name>/: proxy_read_timeout 1d (long-lived SSE and WebSocket UIs), websockets on, X-Forwarded-Prefix set so the harness can build absolute URLs.

Internals

How the gateway does what the sections above describe. Read before changing hive-gateway/, gateway_nginx.rs or agent_sockets.rs.

SPA fallback (Accept-header pattern)

The chat.<swarm> vhost (fluffychat) serves a flutter/SPA bundle via the Accept-header pattern below. Per-agent UIs solve the same two requirements differently, with file-existence try_files — see Per-agent static frontend split. The dashboard routes by path — see Dashboard: path-based routing below. Two requirements collide for a vhost whose client-side router owns every route under /:

  • hard-refresh on a sub-route must serve index.html (SPA's client-side router takes over after JS bootstrap)
  • a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type

Solution: an nginx http-context map $http_accept $matrix_spa_target { ... } keyed on the request's Accept header. Browser navigations (Accept: text/html,...) get index.html; everything else (Accept: image/*, */*, application/json, text/event-stream, …) gets a sentinel nonexistent path, so try_files $uri $uri/ $matrix_spa_target =404 falls through to a plain 404 — a missing asset is just missing. No extension allowlist, no if block, no regex heuristics.

Dashboard: path-based routing (not Accept-header)

hive-c0re serves exactly three prefixes, so the dashboard routes by path — deterministic, where a content-type split would let one URL resolve differently by the caller's Accept header:

  • location /api/ → hive-c0re (7000): all dashboard data, actions, and the two SSE streams (/api/dashboard/stream, /api/build-logs/id/{id}/stream). proxy_buffering off and a 1d read timeout keep the streams live.
  • location /webhook/ → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded.
  • location /health/ → hive-c0re: liveness and readiness.
  • location / → the dashboard dist (from the servedFrontend nix-store path) with try_files $uri /index.html.

The gateway static-serves the dist while hive-c0re stays API-only, so a frontend-only change doesn't restart the core daemon. A new top-level c0re route prefix needs a matching location on the hive vhost (hive-gateway/vhosts.nix).

Per-agent unix-socket upstream

Every agent binds its web UI on a unix-domain socket at /run/hive-agent/<name>/web.sock. The mechanism:

  1. Agent side. The nix module sets HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock on every harness service env; web_ui::serve binds a UnixListener at that path.

  2. Host side. hive-c0re bind-mounts the per-agent subdir (/run/hive-agent/<name>/) into the agent's container. Dir bind, not file bind — file bind-mounts don't survive the harness's unlink + bind(2) cycle on socket replace. Per-agent subdir keeps each agent's container from seeing siblings' sockets.

    The dir is 0751, owned by the agent's container uid/gid (hive-priv creates it 0751 root before each start when missing, and the container's activation hands it to the agent user), so nginx reaches web.sock through o=--x (traverse) and the socket's own 0666. The gateway is one of three principals sharing that dir and doesn't own its ownership rules — see docs/trust-boundary/boundary.md.

  3. Marker gate. After a successful bind_unix, the harness drops <dir>/hyperhive-socket-bound next to the socket. c0re's agent_sockets::write filters its JSON map by marker presence — only agents whose harness has bound the socket appear there. It also accepts the older .bound name.

  4. Gateway side. gateway_nginx::write generates /var/lib/hive-gateway/conf/agents.conf — a plain nginx include file with one location /agent/<name>/ block per agent. Always a UDS upstream (http://unix:/run/hive-agent/<name>/web.sock:/); if the socket isn't yet bound, nginx returns 502 caught by the error_page 502 503 504 = /__hive_agent_unreachable directive. nginx includes /var/lib/hive-gateway/conf/agents.conf — the same path c0re writes, since both run on the host. After each write, c0re triggers the appropriate nginx action via hive-priv (which is root; hive-c0re runs as the unprivileged hive-core user and can't act on a system unit). hive-priv queries ActiveState and dispatches:

    • active → systemctl reload nginx (SIGHUP, zero-downtime)
    • failed → systemctl reset-failed nginx + systemctl start nginx
    • otherwise → systemctl start nginx This is an explicit trigger rather than a path unit watching the file: the write and the reload belong in one causal chain c0re can retry and report on (RELOAD_PENDING), not two units racing on an inotify event.

c0re regenerates agents.conf (and triggers a reload) on two triggers: every topology change (new/removed agents) and every 10s marker poll tick (agent_sockets::spawn_poll). write() is idempotent — skips the rename when content is unchanged. gateway_nginx::reload_if_pending automatically retries failed reloads on subsequent poll ticks.

agents.conf uses atomic <path>.tmp + rename() writes so a crashing c0re process never leaves a partial or unparseable file behind.

Per-agent static frontend split

hive-c0re's environment sets HIVE_AGENT_FRONTEND_DIR to the agent/ subdirectory of services.hyperhive.c0re.servedFrontend. The nginx include generator (gateway_nginx::write) reads this variable and, when set, emits split location blocks per agent instead of a single proxy block.

Location priority for /agent/<name>/...:

# 1. Compiled assets — content-addressed nix store path, cache forever
location ^~ /agent/<name>/static/ {
    alias <frontend>/static/;
    expires 1y;
    add_header Cache-Control "public, immutable, max-age=31536000";
}

# 2. Static dist + proxy fallback for dynamic paths
location /agent/<name>/ {
    alias <frontend>/;
    try_files $uri $uri.html $uri/index.html @<name>_dynamic;
}

# 3. Proxy catchall — API, events, icon, send, login, …
location @<name>_dynamic {
    proxy_pass <upstream>;
    proxy_set_header X-Forwarded-Prefix /agent/<name>;
    proxy_intercept_errors on;
    error_page 502 503 504 = /__hive_agent_unreachable;
    # … (full proxy header block)
}

try_files resolution (nginx applies the alias mapping before checking file existence):

request resolved outcome
/agent/iris/ <frontend>/index.html main agent page
/agent/iris/stats <frontend>/stats.html stats page
/agent/iris/screen <frontend>/screen.html screen page
/agent/iris/static/app.js caught by ^~ block first served with immutable cache
/agent/iris/api/state no file match → @iris_dynamic proxied to agent daemon
/agent/iris/events/live no file match → @iris_dynamic proxied (SSE)

Adding a new HTML page to the frontend dist (dist/<page>.html) makes it reachable at /agent/<name>/<page> with no generator change.

Why ^~ for /static/: the ^~ prefix gives this block higher priority than the plain prefix location /agent/<name>/, so compiled JS/CSS assets skip try_files entirely and get the immutable cache headers. Nix store paths are content-addressed — the hash changes on any content change — so max-age=31536000 is safe.

Why the nix store path resolves: HIVE_AGENT_FRONTEND_DIR is a nix store path baked in at hive-c0re build time, and c0re (writing agents.conf) and nginx (serving files from it) are on the same machine, so they see the same store.

Unset variable: if HIVE_AGENT_FRONTEND_DIR is empty or unset, each agent gets a single proxy block and nginx forwards all traffic to the agent daemon.

extraFiles: per-agent services.hyperhive.agent.frontend.extraFiles sit in mergedDist, not in the base services.hyperhive.c0re.frontend dist. They're not under the nix-store alias path, so requests for them fall through try_files to @<name>_dynamic, and the agent daemon serves them.

Per-agent error pages

/agent/<name>/ requests hit two failure modes; both get static HTML pages instead of nginx's default error chrome:

  • Agent not found (/agent/<unknown>/...) — no location for that name in agents.conf. nginx's prefix match falls back to the bare /agent/ catch-all, which return 404s and error_page 404 rewrites to /__hive_agent_not_found → serves not-found.html with a link back to the dashboard.
  • Agent unreachable (502 / 503 / 504 from proxy_pass) — the per-agent harness isn't responding (container restarting, crash recovery). proxy_intercept_errors on + error_page 502 503 504 = /__hive_agent_unreachable rewrites to unreachable.html.

hive-gateway/error-pages.nix renders every page (notFound, unreachable, unauthorized, ssoUnavailable) from one template with pkgs.writeText; nginx serves each through an internal location with alias to the file, so only nginx's own error handling can reach them. They depend on nothing from the frontend dist and render even when hive-c0re is down. A service module reaches them through gateway.lib.errorPages.

A route earns a custom page when the default status code would blame the wrong component. The per-agent routes qualify (a 502 there means the harness is restarting, not a gateway fault), and so does auth.<swarm> — a dead authelia upstream almost always means no users yet, and a bare 502 blames the proxy, the one part that works. Forge, matrix and fluffychat keep nginx's defaults: a dead upstream there means what the status code says.

The dashboard builds per-agent links as same-origin /agent/<name>/… URLs when StateSnapshot.gateway_enabled is true, and direct http://<host>:<port>/ links otherwise. The flag comes from HIVE_GATEWAY_ENABLED, which hive-c0re/environment.nix sets to 1 on every hive, so the direct shape is a fallback for the variable being unset. Three render sites follow it: the agent-name link, the favicon fetch (<url>/icon), and the nav-strip container-kind links from DashboardState.links (GET /api/dashboard-state). Forge links come from swarm.forge.publicUrl; with it unset the dashboard hides them rather than guess. See docs/web-ui/dashboard.md::Container row for the frontend side.

Dialing another vhost by name (verifiedProxyTo)

vhost-lib.nix's verifiedProxyTo builds the proxy_ssl_* / proxy_set_header block a module uses to dial another service on this same gateway BY NAME over https, verified. One definition rather than a copy per module: nginx verifies nothing by default (proxy_ssl_verify is off), so a proxy_pass https://… without these lines is encrypted and unauthenticated. That failure is invisible — it works, and keeps working, against any certificate at all.

Every line earns its place, each confirmed against a real nginx with the opposite arm run as a control:

  • verify + depth — the chain is leaf -> intermediate -> root.
  • trusted_cert — the bundle; nginx reads ALL certs in the file, which the bundle's own doc warns isn't true of every consumer.
  • ssl_name — checks the HOSTNAME too. Without it a chain-only check accepts any certificate this CA ever signed, and for an internal CA that's every service on the hive.
  • server_name on — sends SNI, or the far end can't pick a cert.

⚠️ Session-cache footgun: the module leaves proxy_ssl_session_reuse at its default (on), deliberately — nginx uses this on per-request auth subrequests, so every request pays the handshake it avoids. Worth knowing when testing though: nginx keys the session cache by upstream address and NOT by trust config, so two locations pointing at one upstream with different trust don't verify independently.

⚠️ Host-header clobber footgun: verifiedProxyTo also pins Host (and reinstates the rest of nginx's recommendedProxySettings header set) to the target name rather than letting nginx fill it in later. name here resolves back to THIS gateway — every consumer dials another vhost on the same nginx, not a separate host — and nginx picks the vhost to answer an HTTPS request from the Host header, not from the TLS SNI that proxy_ssl_name sends. recommendedProxySettings's own Host $host (the CALLER's host, not the target) is textually appended by nixpkgs AFTER a location's extraConfig — so it always wins over a proxy_set_header Host written in the location body, and the subrequest loops back into the ORIGINAL vhost instead of reaching the target, recursing on its own auth_request until nginx's subrequest-depth limit turns it into a plain 500. Every call site sets recommendedProxySettings = false on the location for exactly this reason — nixpkgs' version would still clobber this one.