docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix
gateway.md: split the opener into what/audience/enable; vhost map in two tables (swarm-service vhosts declared by their own modules, then the hive vhost) matching vhosts.nix and the service modules; gateway.enable exists and is set with mkDefault by the modules that need it; Basic auth scope, dashboard /health/ prefix, error-page rendering, matrix body limit and forge link source corrected; nginx internals grouped under one Internals section with their headings unchanged. network.md: gateway and dnsmasq run on the host, not in a container; network.enable is set by the modules that need it; shared-netns firewall rule covers every swarm service container; hive-priv writes the nspawn conf; domain sentence rewritten; removed options moved into <details>. jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS show different graphs drawn by the same component. observability.md: swarm tier first; history narration cut; network access deduplicated into a link to network.md; options link made absolute. matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and SSO options under deploy.matrix; .well-known is served on the hive domain; roadmap sentence deleted; stale hive-c0re provisioning claims fixed; serverName upgrade note moved into <details>. Refs #3902
This commit is contained in:
parent
2b2608a491
commit
067f4e5699
5 changed files with 433 additions and 633 deletions
|
|
@ -2,9 +2,16 @@
|
||||||
|
|
||||||
Private Matrix homeserver (matrix-tuwunel — the conduwuit
|
Private Matrix homeserver (matrix-tuwunel — the conduwuit
|
||||||
successor) wrapped in a nixos-container, plus optional fluffychat-web
|
successor) wrapped in a nixos-container, plus optional fluffychat-web
|
||||||
client at `chat.<swarm-domain>/` (the `gatewayHost` vhost). Configured via
|
client at `chat.<swarm-domain>/` (the `gatewayHost` vhost). A swarm runs one
|
||||||
`services.hyperhive.swarm.matrix.*`; vhost routing lives in
|
homeserver, on one host. Two namespaces configure it:
|
||||||
[`gateway.md`](../networking/gateway.md).
|
|
||||||
|
- `services.hyperhive.swarm.matrix.*` — what the homeserver **is**, as
|
||||||
|
every hive sees it: `serverName`, `gatewayHost`, ports, `allowEncryption`.
|
||||||
|
- `services.hyperhive.deploy.matrix.*` — what the host running it decides:
|
||||||
|
`enable`, `gui.enable`, `openFirewall`, `trustedServers`,
|
||||||
|
`maxRequestSize`, `sso.clientSecretFile`.
|
||||||
|
|
||||||
|
Vhost routing lives in [`gateway.md`](../networking/gateway.md).
|
||||||
|
|
||||||
## Container shape
|
## Container shape
|
||||||
|
|
||||||
|
|
@ -33,9 +40,10 @@ Two distinct hostnames:
|
||||||
*irrevocably* in every `@user:<server_name>` and `!room:<server_name>`
|
*irrevocably* in every `@user:<server_name>` and `!room:<server_name>`
|
||||||
identifier minted on this homeserver. You can't change it later
|
identifier minted on this homeserver. You can't change it later
|
||||||
without abandoning every account and chat history. Defaults to the
|
without abandoning every account and chat history. Defaults to the
|
||||||
bare `services.hyperhive.swarm.domain`; clients autodiscover the
|
bare `services.hyperhive.swarm.domain`. The gateway serves the
|
||||||
actual API endpoint via the `.well-known/matrix/{client,server}`
|
`.well-known/matrix/{client,server}` discovery routes on the matrix
|
||||||
routes the gateway serves at that domain.
|
host's **hive** domain, not the swarm domain
|
||||||
|
([Discovery flow](../networking/gateway.md#discovery-flow-matrix)).
|
||||||
- **`gatewayHost`** — the API listener hostname, where the gateway's
|
- **`gatewayHost`** — the API listener hostname, where the gateway's
|
||||||
matrix vhost proxies `/_matrix/*` to tuwunel. Defaults to
|
matrix vhost proxies `/_matrix/*` to tuwunel. Defaults to
|
||||||
`chat.<services.hyperhive.swarm.domain>`. Set to `null` to skip the
|
`chat.<services.hyperhive.swarm.domain>`. Set to `null` to skip the
|
||||||
|
|
@ -54,12 +62,12 @@ room id, so adopting a new one does **not** rename the existing users and rooms
|
||||||
it strands them, because their ids still name a homeserver that no
|
it strands them, because their ids still name a homeserver that no
|
||||||
longer answers.
|
longer answers.
|
||||||
|
|
||||||
### Upgrading a homeserver that already has ids
|
<details><summary>Upgrading a homeserver that already has ids</summary>
|
||||||
|
|
||||||
`serverName`'s default has changed across releases. A homeserver that
|
A homeserver that minted ids under an older `serverName` default must
|
||||||
has already minted ids under an older default must **pin the value it
|
**pin the value it actually minted them under**, not adopt the current
|
||||||
actually minted them under**, not adopt the new default — see above
|
default — see above for why adopting a new one strands existing users
|
||||||
for why adopting a new one strands existing users and rooms:
|
and rooms:
|
||||||
|
|
||||||
```nix
|
```nix
|
||||||
services.hyperhive.swarm.matrix = {
|
services.hyperhive.swarm.matrix = {
|
||||||
|
|
@ -71,13 +79,14 @@ services.hyperhive.swarm.matrix = {
|
||||||
|
|
||||||
A rebuild on a host that already has a homeserver prints a
|
A rebuild on a host that already has a homeserver prints a
|
||||||
`hive-matrix: WARNING — … serverName is unset` line when this is missing,
|
`hive-matrix: WARNING — … serverName is unset` line when this is missing,
|
||||||
naming the value it's about to default to. That warning is why this
|
naming the value it's about to default to. It never fails the rebuild, so act on
|
||||||
section exists; it never fails the rebuild, so it's on you to act on it
|
it before the homeserver mints the ids.
|
||||||
before the homeserver mints the ids.
|
|
||||||
|
</details>
|
||||||
|
|
||||||
## Default-closed firewall
|
## Default-closed firewall
|
||||||
|
|
||||||
`openFirewall` defaults to `false` (secure-by-default): the host
|
`deploy.matrix.openFirewall` defaults to `false`: the host
|
||||||
reaches the homeserver on loopback, and agent containers reach it
|
reaches the homeserver on loopback, and agent containers reach it
|
||||||
at `chat.<swarm-domain>` via the gateway — so the firewall hole only
|
at `chat.<swarm-domain>` via the gateway — so the firewall hole only
|
||||||
matters for access from *outside* the host. Flip to `true` when
|
matters for access from *outside* the host. Flip to `true` when
|
||||||
|
|
@ -198,9 +207,8 @@ tuwunel has none.
|
||||||
Promoting a user to homeserver admin and resetting a password both need an
|
Promoting a user to homeserver admin and resetting a password both need an
|
||||||
admin **sender**: `!admin …` messages into `#admins:<server_name>`, and
|
admin **sender**: `!admin …` messages into `#admins:<server_name>`, and
|
||||||
tuwunel only treats a message as a command when its sender is already an
|
tuwunel only treats a message as a command when its sender is already an
|
||||||
admin. `@hive-<hive>:` has no admin sender to make that call with. They're
|
admin. `@hive-<hive>:` has no admin sender to make that call with. Both are
|
||||||
swarm-level operations: matrix admin should eventually come from
|
swarm-level operations.
|
||||||
membership in authelia's `admins` group; nobody has built that sync yet.
|
|
||||||
|
|
||||||
<details><summary>Upgrading a hive that shared one sender account with every other hive</summary>
|
<details><summary>Upgrading a hive that shared one sender account with every other hive</summary>
|
||||||
|
|
||||||
|
|
@ -274,8 +282,8 @@ Initial rollout settings:
|
||||||
restart. `trusted_servers = []` keeps it effectively closed
|
restart. `trusted_servers = []` keeps it effectively closed
|
||||||
until you list peers.
|
until you list peers.
|
||||||
- `allow_registration = false`. tuwunel checks this flag only for
|
- `allow_registration = false`. tuwunel checks this flag only for
|
||||||
requests that arrive **without** an appservice token, so hive-c0re
|
requests that arrive **without** an appservice token, so the appservices
|
||||||
provisions exactly as before and tuwunel refuses everyone else. It's not a
|
still create accounts and tuwunel refuses everyone else. It's not a
|
||||||
hardening afterthought: with no registration token configured,
|
hardening afterthought: with no registration token configured,
|
||||||
`allow_registration = true` makes tuwunel refuse to start unless
|
`allow_registration = true` makes tuwunel refuse to start unless
|
||||||
`yes_i_am_very_very_sure_…_open_registration_…` is also set.
|
`yes_i_am_very_very_sure_…_open_registration_…` is also set.
|
||||||
|
|
@ -298,8 +306,7 @@ Initial rollout settings:
|
||||||
|
|
||||||
## Hive Matrix Space
|
## Hive Matrix Space
|
||||||
|
|
||||||
On first boot, after hive-c0re provisions all agent accounts, it
|
On first boot hive-c0re creates a private **Matrix Space** named `"hive"` using its own hive
|
||||||
creates a private **Matrix Space** named `"hive"` using its own hive
|
|
||||||
account (`@hive-<hive>:<server_name>`) and invites every provisioned agent
|
account (`@hive-<hive>:<server_name>`) and invites every provisioned agent
|
||||||
into it. This gives the operator a single Space in FluffyChat or any
|
into it. This gives the operator a single Space in FluffyChat or any
|
||||||
Matrix client that groups all agent-to-agent + operator rooms in one
|
Matrix client that groups all agent-to-agent + operator rooms in one
|
||||||
|
|
@ -331,7 +338,7 @@ re-creation (for example after a homeserver wipe).
|
||||||
## Configuration tuning
|
## Configuration tuning
|
||||||
|
|
||||||
```nix
|
```nix
|
||||||
services.hyperhive.swarm.matrix = {
|
services.hyperhive.deploy.matrix = {
|
||||||
trustedServers = [ "matrix.org" "example.com" ]; # default: []
|
trustedServers = [ "matrix.org" "example.com" ]; # default: []
|
||||||
maxRequestSize = 20000000; # default: 20 MB
|
maxRequestSize = 20000000; # default: 20 MB
|
||||||
};
|
};
|
||||||
|
|
@ -364,14 +371,14 @@ surprising behaviour:
|
||||||
SSO is unconditional, so the three below are requirements of running a
|
SSO is unconditional, so the three below are requirements of running a
|
||||||
homeserver at all rather than of a setting:
|
homeserver at all rather than of a setting:
|
||||||
|
|
||||||
- **Set `sso.clientSecretFile`** — fails at eval, not at boot:
|
- **Set `deploy.matrix.sso.clientSecretFile`** — fails at eval, not at boot:
|
||||||
tuwunel reads its identity providers from the config file, so a
|
tuwunel reads its identity providers from the config file, so a
|
||||||
half-configured one can stop the homeserver from starting outright
|
half-configured one can stop the homeserver from starting outright
|
||||||
rather than merely hiding a login button. On a host that also runs
|
rather than merely hiding a login button. On a host that also runs
|
||||||
the swarm's authelia it's wired up for you.
|
the swarm's authelia it's wired up for you.
|
||||||
- **Set `swarm.authelia.url`** — without a provider URL there
|
- **Set `swarm.authelia.url`** — without a provider URL there
|
||||||
is nothing to discover against.
|
is nothing to discover against.
|
||||||
- **Set `gatewayHost != null`** — the SSO callback URL is
|
- **Set `swarm.matrix.gatewayHost != null`** — the SSO callback URL is
|
||||||
format-locked to `<homeserver>/_matrix/client/unstable/login/sso/callback/<client_id>`,
|
format-locked to `<homeserver>/_matrix/client/unstable/login/sso/callback/<client_id>`,
|
||||||
and the identity provider needs a public name to redirect the
|
and the identity provider needs a public name to redirect the
|
||||||
browser to.
|
browser to.
|
||||||
|
|
|
||||||
|
|
@ -1,103 +1,290 @@
|
||||||
# hive-gateway
|
# hive-gateway
|
||||||
|
|
||||||
This host's nginx fronts the hyperhive web surfaces running on it — next to hive-c0re, not in its own container: it shares the host netns anyway (see [Vhost map](#vhost-map) below), so containerizing it would buy no network isolation while costing a resolv.conf sync, a machine-bus reload, and three bind mounts. System-config (not meta-flake managed). Configured via `services.hyperhive.gateway.*` + per-subsystem opt-in flags in `services.hyperhive.{forge,matrix,...}`. the modules that need them assert `gateway.enable` and `gateway.dns.enable`, so a host serving a vhost or resolving hive names gets them without an opt-in.
|
Every host's nginx: the one front door for whatever this host serves. A swarm service running here (forge, matrix, SSO, the swarm UI, the metrics and log stores) declares its own vhost through the gateway; the gateway itself adds the hive's own surface — dashboard, per-agent UIs, matrix discovery.
|
||||||
|
|
||||||
|
_For the operator configuring `services.hyperhive.gateway.*` on a host._ nginx and the hive resolver (dnsmasq) run on the host next to hive-c0re, not in a container: they bind `:80`/`:443` and the bridge address, so a network namespace of their own would isolate nothing.
|
||||||
|
|
||||||
|
You rarely switch it on yourself. `gateway.enable` defaults to off, and every module that serves a vhost or needs hive names to resolve sets `gateway.enable` / `gateway.dns.enable` with `mkDefault true` — the hive controller, each swarm service, CI.
|
||||||
|
|
||||||
## Vhost map
|
## Vhost map
|
||||||
|
|
||||||
| URL | vhost | upstream | source |
|
**Swarm services.** Each module declares its vhost on the host that runs the service, under the swarm domain:
|
||||||
|
|
||||||
|
| URL | upstream | declared by, when |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `<swarm>/` | swarm-ui dist (static), behind an authelia subrequest | `swarm-ui.nix`, `deploy.swarm-ui.enable` |
|
||||||
|
| `auth.<swarm>/` | authelia (`9091`) | `swarm-authelia.nix`, `deploy.authelia.enable` |
|
||||||
|
| `forge.<swarm>/` | forgejo (`3000`) | `hive-forge/`, `deploy.forgejo.behindGateway` |
|
||||||
|
| `chat.<swarm>/_matrix/*` | tuwunel (`8008`) | `hive-matrix.nix`, `swarm.matrix.gatewayHost != null` |
|
||||||
|
| `chat.<swarm>/` | fluffychat-web static (404 with the GUI off) | `hive-matrix.nix`, `deploy.matrix.gui.enable` |
|
||||||
|
| `chat.<swarm>/config.json` | inline JSON (FluffyChat boot config) | `hive-matrix.nix`, `deploy.matrix.gui.enable` |
|
||||||
|
| `grafana.<swarm>`, `metrics.<swarm>`, `logs.<swarm>`, `otel.<swarm>`, `bao.<swarm>` | the matching swarm service | that service's module → [`swarm/services.md`](../swarm/services.md) |
|
||||||
|
|
||||||
|
**The hive's own vhost**, named for the hive domain:
|
||||||
|
|
||||||
|
| URL | upstream | when |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `<hive>/` | dashboard dist (static, from `servedFrontend`) | always |
|
||||||
|
| `<hive>/api/`, `/webhook/`, `/health/` | hive-c0re (`7000`) | always |
|
||||||
|
| `<hive>/api/docs/` | themed Swagger UI dist (static) | always |
|
||||||
|
| `<hive>/agent/<name>/` | per-agent harness over its unix socket | `agents.conf` (runtime-generated) |
|
||||||
|
| `<hive>/.well-known/matrix/{client,server}` | inline JSON | `deploy.matrix.enable` |
|
||||||
|
| `<hive>/matrix/` (deprecated) | 301 → `chat.<swarm>/` | `deploy.matrix.gui.enable` and `gatewayHost` set |
|
||||||
|
|
||||||
|
The catch-all `_` vhost answers any other `Host` with `444` (connection closed, no response). It's `mkDefault`, so to make your own vhost the default server, set `services.nginx.virtualHosts."_".default = false;` — an eval assertion names both when two claim it.
|
||||||
|
|
||||||
|
Only the host that **runs** authelia declares `auth.<swarm>`; a hive that merely uses SSO knows `swarm.authelia.url` but doesn't answer for that name. The server name must be exactly `swarm.authelia.domain` — authelia checks that `authelia_url` sits inside its session cookie domain at startup, and refuses to boot otherwise. The vhost carries no Basic auth (that would put the login page behind the login it replaces) and passes `X-Forwarded-{Proto,Host,Uri,For}`, because authelia decides by the original request.
|
||||||
|
|
||||||
|
⚠️ `auth.<swarm>` answering with the **sso unavailable** page means authelia isn't answering at all — check `journalctl -M swarm-authelia -u authelia-swarm`. A swarm with no real accounts yet boots fine (a disabled placeholder user keeps authelia's user store non-empty) and serves a login page that refuses everyone; add the first account with `swarmctl user add …` ([`setup.md`](../getting-started/setup.md)).
|
||||||
|
|
||||||
|
Per-agent UIs stay sub-path; forge and matrix get sub-domains → [Sub-domain shape](#sub-domain-shape-rationale).
|
||||||
|
|
||||||
|
## TLS modes
|
||||||
|
|
||||||
|
The gateway always terminates TLS: there is no http-only mode. Which certificate it serves depends on what you configure:
|
||||||
|
|
||||||
|
| mode | config | cert source | `.well-known` scheme |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| `<hive>/` | `_` (catch-all) | dashboard dist (static, from `servedFrontend`); `/api/` + `/webhook/` → hive-c0re (`7000`) | always |
|
| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | the host's hive CA signs a gateway leaf (RSA-4096) | `https` |
|
||||||
| `<hive>/agent/<name>/` | `_` | per-agent harness (UDS or TCP) | `agents.conf` (runtime-generated) |
|
| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` |
|
||||||
| `<hive>/.well-known/matrix/{client,server}` | `_` | inline JSON (no upstream) | `matrix.enable && domain != null` |
|
| operator cert | `tls.certDir` set | read from the operator's dir | `https` |
|
||||||
| `<hive>/matrix/` (deprecated) | `_` | 301 → `chat.<swarm>/` | `matrix.gui.enable` |
|
|
||||||
| `forge.<swarm>/` | `forge.<swarm>` | forgejo (`3000`) | `deploy.forgejo.behindGateway` |
|
|
||||||
| `chat.<swarm>/_matrix/*` | `chat.<swarm>` | tuwunel (`8008`) | `matrix.gatewayHost != null` |
|
|
||||||
| `chat.<swarm>/` | `chat.<swarm>` | fluffychat-web static | `matrix.gui.enable` |
|
|
||||||
| `chat.<swarm>/config.json` | `chat.<swarm>` | inline JSON (FluffyChat boot config) | `matrix.gui.enable && domain != null` |
|
|
||||||
| `auth.<swarm>/` | `auth.<swarm>` | authelia (`9091`) | `deploy.authelia` |
|
|
||||||
| `<swarm>/` | `<swarm>` | swarm-ui dist (static), behind an authelia subrequest | `deploy.swarm-ui` |
|
|
||||||
|
|
||||||
Only the host that **runs** authelia declares the authelia vhost, not every hive that uses it — a client hive knows the swarm's `authelia.url` but must not answer for a name it doesn't serve. Its server name is exactly `swarm.authelia.domain`: authelia validates `authelia_url ⊂ session cookie domain` at startup, so a near-miss is a container that refuses to boot. It carries no `auth_basic` — the login page must not sit behind the login mechanism it replaces — and sets the four `X-Forwarded-{Proto,Host,Uri,For}` headers, since authelia decides by the *original* request rather than the hop it sees.
|
`tls.certDir` together with `tls.acme.enable` fails an assertion — pick one.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
⚠️ One vhost class has no TLS: a vhost bound to loopback for a local consumer. `grafana-metrics` (`nix/host-modules/swarm-grafana.nix`) is the one in the tree: it listens on `127.0.0.1` only and serves `= /metrics` from grafana's unix socket for this host's collector. Everything with a routable name follows the table above.
|
||||||
⚠️ **A `502` from this vhost means authelia itself isn't answering, not that the proxy is misconfigured.** Check `journalctl -M swarm-authelia -u authelia-swarm` before suspecting anything here. A swarm with no users of its own answers with a login page that refuses everyone; the first account is created in [`swarm/sso.md`](../swarm/sso.md).
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
Per-agent UIs stay sub-path, forge and matrix get sub-domains — see
|
### ACME / Let's Encrypt (`tls.acme`)
|
||||||
[Sub-domain shape (rationale)](#sub-domain-shape-rationale) below for why.
|
|
||||||
|
The simplest production path for a public domain:
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive.gateway = {
|
||||||
|
openFirewall = true;
|
||||||
|
tls.acme = {
|
||||||
|
enable = true;
|
||||||
|
email = "admin@example.com"; # required
|
||||||
|
};
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
nginx obtains and renews certs via the HTTP-01 challenge on `port` (default 80); they land in `/var/lib/acme/`, managed by nixpkgs's `security.acme`. Every name this host serves needs a public DNS record pointing here — the hive domain and each swarm-service name in the [vhost map](#vhost-map) — and `openFirewall = true` so Let's Encrypt reaches `/.well-known/acme-challenge/`.
|
||||||
|
|
||||||
|
### Self-signed TLS (default)
|
||||||
|
|
||||||
|
On by default, listening on `httpsPort` (default 443) beside the plain-http `port` (default 80).
|
||||||
|
|
||||||
|
The issuer is a **host-held hive CA**, not a bare self-signed leaf. `hive-tls-ca.service` (from `hive-tls.nix`) generates a long-lived CA (`services.hyperhive.deploy.hive-controller.tls.caValidityDays`, default 7300 days) under `…tls.stateDir` (default `/var/lib/hive-tls`) and signs a gateway **leaf** with it (`leafValidityDays`, default 30). `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`).
|
||||||
|
|
||||||
|
⚠️ **Keep the import unit.** It does two jobs nginx needs: it re-modes the key to `0640 root:nginx` (nginx's pre-start `nginx -t` runs as the nginx user and fails on the CA's `0600 root:root` key), and it makes sure **every cert path the config names exists** — when the swarm-services leaf is missing it installs the hive leaf in its place. nginx refuses a config naming a missing cert file, so without that fallback one missing leaf takes down every vhost, not just one.
|
||||||
|
|
||||||
|
**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust, and nothing can wire a runtime-generated leaf into an agent's build-time trust store. Agents and federation peers trust the stable CA once; leaf rotation never breaks them.
|
||||||
|
|
||||||
|
**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is an intermediate under the swarm root ([`swarm/ca.md`](../swarm/ca.md) has the hierarchy), and a verifier can't stop at an intermediate — so the bundle carries the hive CA plus its root. nginx serves the leaf with the hive CA appended for the same reason. Agents (via `security.pki.certificateFiles`), the CI and forge containers and federating peers all read the bundle.
|
||||||
|
|
||||||
|
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) fetches `https://<host>/.well-known/matrix/client` and never falls back to plain http, so without TLS the browser client can't bootstrap.
|
||||||
|
|
||||||
|
**Cert shape**: the leaf's CN is the bare hive domain; its SANs are `<hive>` and `*.<hive>`. The hive CA is name-constrained to `<hive>`, so it can't sign a swarm-service name outside it — a violating SAN would invalidate the whole leaf. Those names get the swarm-services leaf instead ([`swarm/ca.md`](../swarm/ca.md)).
|
||||||
|
|
||||||
|
**Rotation**: `hive-tls-ca.service` re-signs the leaf when it's missing, within 30 days of expiry, or no longer covers the configured names, always under the same CA. It regenerates the CA only if missing or expired. To force a leaf rotation, delete `gateway.pem` under the state dir, restart the unit, then reload `nginx`.
|
||||||
|
|
||||||
|
**Cert prompts**: browsers warn once per host until you add the hive's `trust-bundle.pem` (the anchor, not the leaf) to the browser or OS trust store.
|
||||||
|
|
||||||
|
### Operator-provided cert (`tls.certDir`)
|
||||||
|
|
||||||
|
For a cert from a real CA (Let's Encrypt via your own `security.acme`, a corporate CA):
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive.gateway = {
|
||||||
|
tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir
|
||||||
|
# tls.certName = "cert.pem"; # default — matches security.acme layout
|
||||||
|
# tls.keyName = "key.pem"; # default — matches security.acme layout
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
nginx reads the directory directly. Keep the key readable by nginx:
|
||||||
|
|
||||||
|
- `security.acme` writes keys `0640 root:acme`, which the `nginx` user can't read. Set `security.acme.certs."example.com".group = "nginx";` (or make the key `0644` if your threat model allows). Otherwise nginx fails at startup with the reason in the journal.
|
||||||
|
|
||||||
|
### Fronting with an external TLS terminator
|
||||||
|
|
||||||
|
The gateway has no plain-http upstream mode. Either give the gateway the real cert (`tls.certDir` or `tls.acme`) so it serves proper TLS itself, or front it over a unix socket rather than a plain-http TCP port. `.well-known/matrix/*` responses always advertise `https` ([Discovery flow](#discovery-flow-matrix)).
|
||||||
|
|
||||||
|
## HTTP Basic auth
|
||||||
|
|
||||||
|
Optional: Basic auth on the hive's dashboard. Add a login first, then enable it — the htpasswd file exists from first boot, and an empty one refuses everyone.
|
||||||
|
|
||||||
|
_On the hive host:_
|
||||||
|
|
||||||
|
```sh
|
||||||
|
hivectl gateway create-user alice --password-stdin # password on stdin
|
||||||
|
hivectl gateway delete-user bob
|
||||||
|
hivectl gateway list-users
|
||||||
|
```
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive.gateway.auth = {
|
||||||
|
enable = true;
|
||||||
|
# realm = "hyperhive"; # default; must not contain `"` or `$`
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
`hivectl` asks hive-c0re over the host admin socket, and the daemon writes `/var/lib/hive-gateway/conf/gateway.htpasswd` itself, bcrypt (cost 12) with `$2y$` hashes nginx reads natively. `--password <pw>` also works but lands in shell history.
|
||||||
|
|
||||||
|
**What it gates:** `/`, `/api/` and `/api/docs/` on the hive vhost. **Not gated:** `/webhook/` (Forgejo can't send Basic credentials; the handler checks the HMAC signature instead), `/health/` (for uptime monitors; status only), `/.well-known/matrix/*`, and the per-agent `/agent/<name>/` routes, which come from `agents.conf` and inherit no auth from `/`.
|
||||||
|
|
||||||
|
A failed or missing login gets `401` with a styled `unauthorized.html` naming the `hivectl` command to run, so browsers still show the login dialog first.
|
||||||
|
|
||||||
|
## Firewall posture (host-level)
|
||||||
|
|
||||||
|
`services.hyperhive.gateway.openFirewall = true` opens `port` and `httpsPort` on the host firewall — both, since the gateway always serves TLS. It defaults to off; set it for any reach from outside the host.
|
||||||
|
|
||||||
|
nginx is the one external entry point. The per-agent web-port range (`8100`–`8999`) stays closed: agents serve their UI on a unix socket ([Per-agent unix-socket upstream](#per-agent-unix-socket-upstream)), and the hashed TCP port (`lifecycle::agent_web_port`) is only a fallback bind for an agent missing `HIVE_WEB_SOCKET` — the gateway never proxies through it.
|
||||||
|
|
||||||
|
The dashboard port (`services.hyperhive.c0re.dashboardPort`, default 7000) binds `127.0.0.1` only, so remote dashboard access goes through the gateway. The dashboard can approve, deny and destroy; don't expose it without a reverse proxy in front.
|
||||||
|
|
||||||
|
## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain
|
||||||
|
|
||||||
|
Agents poll `HIVE_FORGE_URL` for Forgejo notifications and run every `hive-forge` call against it. `nix/host-modules/hive-c0re/environment.nix` sets it to `http://<swarm.forge.domain>` (default `forge.<swarm-domain>` — a swarm runs one forge). Agents run in a private netns and can't reach the host's loopback, so they resolve that name through the bridge dnsmasq to the bridge IP and reach nginx on port 80 (the bridge firewall opens 80 and 443) — the same path an operator's browser takes.
|
||||||
|
|
||||||
|
## hive-forge container shape
|
||||||
|
|
||||||
|
Private Forgejo in a nixos-container named `hive-forge` (not `h-*`, so c0re's lifecycle scanner leaves it alone; manage it with the standard `nixos-container` CLI). The container keeps it from colliding with any `services.forgejo` the operator already runs on the host: separate systemd namespace, separate state dir, separate port.
|
||||||
|
|
||||||
|
It shares the host network namespace (`privateNetwork = false`): the container is for state and unit isolation, not network isolation. Agent containers, by contrast, are network-isolated and reach the forge through the gateway ([`HIVE_FORGE_URL`](#hive_forge_url-agents-reach-the-forge-via-the-gateway-by-domain)).
|
||||||
|
|
||||||
|
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/` and survives restarts and reboots; destroy the container to wipe it.
|
||||||
|
|
||||||
|
Only the swarm's forge host runs it (`services.hyperhive.deploy.forgejo.enable`, see [`swarm/services.md`](../swarm/services.md)). Every other hive reaches that host's gateway by `forge.<swarm-domain>`.
|
||||||
|
|
||||||
|
### Network and port configuration
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive.swarm.forge = {
|
||||||
|
httpPort = 3000; # default — outside hyperhive's 7000 / 8100-8999
|
||||||
|
sshPort = 2222; # default — git-over-SSH, off the host's openssh on 22
|
||||||
|
};
|
||||||
|
# Which ports the forge answers on is swarm-wide; whether THIS host opens
|
||||||
|
# them in its firewall is a deployment decision, so it lives under deploy.*.
|
||||||
|
services.hyperhive.deploy.forgejo.openFirewall = false; # default
|
||||||
|
```
|
||||||
|
|
||||||
|
`sshPort` serves `git clone/push/pull` over SSH (`git@<domain>:owner/repo.git` with `-p 2222`). SSH goes straight to Forgejo, not through nginx.
|
||||||
|
|
||||||
|
`openFirewall` (default `false`) opens `httpPort` and `sshPort` on the host. Agents don't need it — they come through the gateway. Set it for a browser reaching `http://<host>:<httpPort>/` directly, or for external git clients pushing over SSH. The forge vhost behind the gateway (`deploy.forgejo.behindGateway`, default `true`) needs only the gateway's own `openFirewall`.
|
||||||
|
|
||||||
|
### `rootUrl` override
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive.swarm.forge.rootUrl = "https://forge.example.com/";
|
||||||
|
```
|
||||||
|
|
||||||
|
`rootUrl` (default `null`) overrides the Forgejo `ROOT_URL` derived from `forge.domain` and the gateway:
|
||||||
|
|
||||||
|
| Shape | Derived `ROOT_URL` |
|
||||||
|
|---|---|
|
||||||
|
| `deploy.forgejo.behindGateway = true` | `https://<forge.domain>/` (no port suffix when `gateway.httpsPort == 443`) |
|
||||||
|
| `deploy.forgejo.behindGateway = false` | `http://<forge.domain>:<httpPort>/` |
|
||||||
|
|
||||||
|
Set it when `forge.domain` differs from the public URL, or for a bespoke shape such as an external reverse proxy on another host or path. It must end with `/` (an assertion enforces this).
|
||||||
|
|
||||||
|
## Security headers
|
||||||
|
|
||||||
|
Every named vhost sets these at server scope (the `_` catch-all only closes connections):
|
||||||
|
|
||||||
|
| Header | Value |
|
||||||
|
|--------|-------|
|
||||||
|
| `X-Frame-Options` | `SAMEORIGIN` |
|
||||||
|
| `X-Content-Type-Options` | `nosniff` |
|
||||||
|
| `Referrer-Policy` | `strict-origin-when-cross-origin` |
|
||||||
|
|
||||||
|
nginx doesn't merge `add_header`: a `location` that sets its own (the CORS locations `/.well-known/matrix/client` and `/_matrix/`) inherits none of the server-scope headers, so those locations repeat them. Locations with no `add_header` of their own pick them up.
|
||||||
|
|
||||||
|
### HSTS (`gateway.hsts`)
|
||||||
|
|
||||||
|
Opt-in, off by default:
|
||||||
|
|
||||||
|
```nix
|
||||||
|
services.hyperhive.gateway.hsts = {
|
||||||
|
enable = true; # default: false
|
||||||
|
maxAge = 31536000; # default: 1 year (required for preload list)
|
||||||
|
includeSubDomains = true; # default: true
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
When enabled, every vhost adds `Strict-Transport-Security: max-age=…[; includeSubDomains]`. HSTS pins https in the browser: a deployment that later loses TLS locks browsers out until `max-age` expires, so enable it only when TLS is permanent.
|
||||||
|
|
||||||
|
## Local dev (`localHostsEntry`)
|
||||||
|
|
||||||
|
`services.hyperhive.gateway.localHostsEntry = true` maps to `127.0.0.1` in the host's `/etc/hosts`:
|
||||||
|
|
||||||
|
- the hive domain;
|
||||||
|
- every name a module on this host contributes to `gateway.localNames` — each swarm service this host runs adds its own (`forge.<swarm>` when behind the gateway, `chat.<swarm>`, `auth.<swarm>`, the swarm UI's apex, …).
|
||||||
|
|
||||||
|
`services.hyperhive.deploy.singleHostSwarm` turns it on. Leave it off with real DNS.
|
||||||
|
|
||||||
## Discovery flow (matrix)
|
## Discovery flow (matrix)
|
||||||
|
|
||||||
Operator points client at `<hive>`. Sequence:
|
The operator points a client at `<hive>`:
|
||||||
|
|
||||||
1. Client fetches `https://<hive>/.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat.<swarm>"}}` (no port suffix when gateway listens on 443). The gateway always terminates TLS, so the scheme is always `https`; a non-default `httpsPort` shows up as the port suffix.
|
1. The client fetches `https://<hive>/.well-known/matrix/client` → `{"m.homeserver":{"base_url":"https://chat.<swarm>"}}` (no port suffix when `httpsPort` is 443).
|
||||||
2. Client connects to `chat.<swarm>/_matrix/client/...`.
|
2. It connects to `chat.<swarm>/_matrix/client/...`.
|
||||||
3. Gateway routes `/_matrix/*` → tuwunel at `127.0.0.1:8008`.
|
3. The gateway proxies `/_matrix/*` to tuwunel at `127.0.0.1:8008`.
|
||||||
|
|
||||||
matrix-dart-sdk (FluffyChat etc.) hardcodes `https` for the well-known fetch regardless of input scheme, so the discovery endpoint MUST be https — see "Self-signed TLS" below for the cert generation that backs the default-on path.
|
matrix-dart-sdk (FluffyChat and others) always fetches the well-known over `https`, which is why [self-signed TLS](#self-signed-tls-default) is on by default.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.<swarm>:<httpsPort>"}`. The port is always explicit, even 443: a delegated host without a port means the federation default 8448, not 443. Peers then reach `/_matrix/` on the chat vhost through the gateway, so the gateway must be reachable from them (`gateway.openFirewall`).
|
||||||
Federation peers fetch `.well-known/matrix/server` → `{"m.server":"chat.<swarm>:<httpsPort>"}` (the federation delegation always carries an explicit port, even the HTTPS default 443 — the https-implies-443 elision only applies to the client base_url above). Gateway only listens on configured `port` (+ `httpsPort` when TLS on); cross-hive federation needs either an SRV record (`_matrix._tcp.chat.<swarm>` → port 80 / 443) OR `matrix.openFirewall = true` so peers reach tuwunel's federation port directly. Hyperhive is closed/internal in most deployments, so this rarely bites.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
## SPA fallback (Accept-header pattern)
|
⚠️ The gateway serves both `.well-known` routes on the **hive** vhost. Matrix looks them up at the `serverName`, which defaults to the bare swarm domain, and the swarm UI's apex vhost serves no `.well-known/matrix/*` route.
|
||||||
|
|
||||||
The per-agent UIs and the `chat.<swarm>` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard vhost instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide:
|
## Sub-domain shape (rationale)
|
||||||
|
|
||||||
|
Sub-domain for forge and matrix, sub-path for per-agent UIs:
|
||||||
|
|
||||||
|
- forgejo's default `ROOT_URL = http://<host>/` works without `X-Forwarded-Prefix` handling; sub-domain hosting is the canonical Forgejo shape.
|
||||||
|
- matrix deployments put the API listener on its own name, and federation expects that. `gatewayHost` defaults to `chat.<swarm-domain>` rather than the spec-conventional `matrix.<server_name>`; set it explicitly for the conventional label.
|
||||||
|
- per-agent UIs are hyperhive-internal and base-path-aware for `/agent/<name>/`. A sub-domain per agent would multiply DNS and TLS cost for nothing.
|
||||||
|
- cookie and storage isolation: a forge XSS can't reach the dashboard session, because they're different origins.
|
||||||
|
|
||||||
|
`services.hyperhive.swarm.forge.domain` and `services.hyperhive.swarm.matrix.gatewayHost` take a full hostname (`forge.darkest.space`, `git.example.com`), not a label glued to a domain.
|
||||||
|
|
||||||
|
## Tuning knobs
|
||||||
|
|
||||||
|
Per-vhost timeouts and body-size limits live in the location blocks:
|
||||||
|
|
||||||
|
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), websockets on.
|
||||||
|
- matrix `/_matrix/` (tuwunel): `client_max_body_size` = `deploy.matrix.maxRequestSize` + 1 MiB, so tuwunel stays the tighter limit and returns a matrix error a client can act on; `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*`, websockets on.
|
||||||
|
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE and WebSocket UIs), websockets on, `X-Forwarded-Prefix` set so the harness can build absolute URLs.
|
||||||
|
|
||||||
|
## Internals
|
||||||
|
|
||||||
|
How the gateway does what the sections above describe. Read before changing `hive-gateway/`, `gateway_nginx.rs` or `agent_sockets.rs`.
|
||||||
|
|
||||||
|
### SPA fallback (Accept-header pattern)
|
||||||
|
|
||||||
|
The per-agent UIs and the `chat.<swarm>` vhost serve a flutter/SPA bundle via the Accept-header pattern below. The dashboard instead routes by **path** — see [Dashboard: path-based routing](#dashboard-path-based-routing-not-accept-header) below. Two requirements collide:
|
||||||
|
|
||||||
- hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap)
|
- hard-refresh on a sub-route must serve `index.html` (SPA's client-side router takes over after JS bootstrap)
|
||||||
- a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type
|
- a non-navigation request that isn't an on-disk asset must NOT get HTML with the wrong content-type
|
||||||
|
|
||||||
Solution: an `nginx http`-context `map $http_accept $<name>_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $<name>_spa_target <final>` falls through to `<final>`. No extension allowlist, no `if` block, no regex heuristics.
|
Solution: an `nginx http`-context `map $http_accept $<name>_spa_target { ... }` keyed on the request's Accept header. Browser navigations (`Accept: text/html,...`) get `index.html`; everything else (`Accept: image/*`, `*/*`, `application/json`, `text/event-stream`, …) gets a sentinel nonexistent path, so `try_files $uri $<name>_spa_target <final>` falls through to `<final>`. No extension allowlist, no `if` block, no regex heuristics.
|
||||||
|
|
||||||
For matrix / per-agent static assets, `<final>` is `=404` (a missing asset is just missing).
|
For matrix and per-agent static assets, `<final>` is `=404` (a missing asset is just missing).
|
||||||
|
|
||||||
### Dashboard: path-based routing (not Accept-header)
|
#### Dashboard: path-based routing (not Accept-header)
|
||||||
|
|
||||||
Every hive-c0re route lives under `/api/` plus the single `/webhook/knowledge` endpoint, so the dashboard vhost routes by **path**, not Accept header — deterministic, unlike a content-type split where the same URL could resolve differently depending on the caller's `Accept` header:
|
hive-c0re serves exactly three prefixes, so the dashboard routes by **path** — deterministic, where a content-type split would let one URL resolve differently by the caller's `Accept` header:
|
||||||
|
|
||||||
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions/mutations, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). Carries `proxy_buffering off` + a 1d read timeout for the streams.
|
- `location /api/` → hive-c0re (`7000`): all dashboard data, actions, and the two SSE streams (`/api/dashboard/stream`, `/api/build-logs/id/{id}/stream`). `proxy_buffering off` and a 1d read timeout keep the streams live.
|
||||||
- `location /webhook/` → hive-c0re: the knowledge webhook.
|
- `location /webhook/` → hive-c0re: knowledge push and config-PR approval triggers, HMAC-guarded.
|
||||||
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html` (SPA fallback).
|
- `location /health/` → hive-c0re: liveness and readiness.
|
||||||
|
- `location /` → the dashboard dist (from the `servedFrontend` nix-store path) with `try_files $uri /index.html`.
|
||||||
|
|
||||||
Each location carries a duplicated `auth_basic` block (separate locations don't inherit it). This keeps the gateway static-serving the dashboard dist while hive-c0re stays API-only — a frontend-only change doesn't rebuild or restart the core daemon. A new top-level c0re route prefix (beyond `/api` + `/webhook`) needs a matching `location` added to the dashboard vhost.
|
The gateway static-serves the dist while hive-c0re stays API-only, so a frontend-only change doesn't restart the core daemon. A new top-level c0re route prefix needs a matching `location` on the hive vhost (`hive-gateway/vhosts.nix`).
|
||||||
|
|
||||||
## Local dev (`localHostsEntry`)
|
### Per-agent unix-socket upstream
|
||||||
|
|
||||||
`services.hyperhive.gateway.localHostsEntry = true` adds entries to the host's `/etc/hosts`:
|
Every agent binds its web UI on a unix-domain socket at
|
||||||
|
`/run/hive-agent/<name>/web.sock`. The mechanism:
|
||||||
- `<hive-domain>` → `127.0.0.1`
|
|
||||||
- `forge.<swarm>` → `127.0.0.1` (when deploy.forgejo.behindGateway)
|
|
||||||
- `chat.<swarm>` → `127.0.0.1` (when matrix.gatewayHost set)
|
|
||||||
- `auth.<swarm>` → `127.0.0.1` (when deploy.authelia)
|
|
||||||
|
|
||||||
`lib.unique` de-dupes if any sub-domain happens to equal another entry. Operators with real DNS leave it off.
|
|
||||||
|
|
||||||
## Sub-domain shape (rationale)
|
|
||||||
|
|
||||||
Operator decision: sub-domain over sub-path for forge + matrix, sub-path for per-agent UIs.
|
|
||||||
|
|
||||||
- forgejo's default `ROOT_URL = http://<host>/` works without any `X-Forwarded-Prefix` gymnastics — sub-domain hosting is the canonical Forgejo deploy shape.
|
|
||||||
- matrix-spec deployments universally use `matrix.<server_name>` for the actual API listener — federation already expects this. (`gatewayHost`'s own default departs from that convention — `chat.<swarm-domain>`, not `matrix.<server_name>` — since `serverName` is swarm-wide but `gatewayHost` is per-hive; operators who want the spec-conventional label can still set it explicitly.)
|
|
||||||
- per-agent UIs are hyperhive-internal and base-path-aware specifically for `/agent/<name>/`. Sub-domain per agent would multiply DNS + TLS-per-subdomain cost without per-app config wins.
|
|
||||||
- cookie / storage isolation: a future forge XSS can't reach the dashboard session because they're different origins.
|
|
||||||
|
|
||||||
`services.hyperhive.{forge.domain,matrix.gatewayHost}` take the full hostname (`forge.darkest.space`, `git.example.com`) rather than a label that gets concatenated with hive-domain — operators want control over the full shape, not a forced `<label>.<hive-domain>` pattern.
|
|
||||||
|
|
||||||
## Tuning knobs
|
|
||||||
|
|
||||||
Per-vhost timeouts + body-size limits live in the location blocks:
|
|
||||||
|
|
||||||
- forge `/` (forgejo): `client_max_body_size 1G` (LFS), `proxy_read_timeout 1h` (multi-GB clones), `proxyWebsockets = true` (live-update endpoints).
|
|
||||||
- matrix `/_matrix/` (tuwunel): `client_max_body_size 50M` (media uploads), `proxy_read_timeout 1h` (long-poll `/sync`), CORS `*` (federation + cross-origin clients), `proxyWebsockets = true`.
|
|
||||||
- per-agent `/agent/<name>/`: `proxy_read_timeout 1d` (long-lived SSE / WebSocket dashboards), `proxyWebsockets = true`, `X-Forwarded-Prefix` set so the harness can build absolute URLs when relative isn't enough.
|
|
||||||
|
|
||||||
SSH for forge stays direct on `services.hyperhive.swarm.forge.sshPort` — separate listener protocol, not HTTP-over-nginx.
|
|
||||||
|
|
||||||
## Per-agent unix-socket upstream
|
|
||||||
|
|
||||||
All agents bind their web UI on a unix-domain socket at
|
|
||||||
`/run/hive-agent/<name>/web.sock` — the `HIVE_WEB_SOCKET` env var is
|
|
||||||
now set unconditionally for every agent. The mechanism:
|
|
||||||
|
|
||||||
1. **Agent side**. The nix module sets `HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock`
|
1. **Agent side**. The nix module sets `HIVE_WEB_SOCKET=/run/hive-agent/<name>/web.sock`
|
||||||
on every harness service env; `web_ui::serve` binds a
|
on every harness service env; `web_ui::serve` binds a
|
||||||
|
|
@ -106,7 +293,7 @@ now set unconditionally for every agent. The mechanism:
|
||||||
(`/run/hive-agent/<name>/`) into the agent's container. Dir
|
(`/run/hive-agent/<name>/`) into the agent's container. Dir
|
||||||
bind, not file bind — file bind-mounts don't survive the
|
bind, not file bind — file bind-mounts don't survive the
|
||||||
harness's `unlink + bind(2)` cycle on socket replace. Per-agent
|
harness's `unlink + bind(2)` cycle on socket replace. Per-agent
|
||||||
subdir keeps each agent's container blind to siblings' sockets.
|
subdir keeps each agent's container from seeing siblings' sockets.
|
||||||
|
|
||||||
The dir is `0751`, owned by the agent's container uid/gid (hive-priv
|
The dir is `0751`, owned by the agent's container uid/gid (hive-priv
|
||||||
creates it `0751 root` before each start when missing, and the
|
creates it `0751 root` before each start when missing, and the
|
||||||
|
|
@ -115,11 +302,11 @@ now set unconditionally for every agent. The mechanism:
|
||||||
own `0666`. The gateway is one of three principals sharing that dir
|
own `0666`. The gateway is one of three principals sharing that dir
|
||||||
and doesn't own its ownership rules — see
|
and doesn't own its ownership rules — see
|
||||||
[`docs/trust-boundary/boundary.md`](../trust-boundary/boundary.md#the-per-agent-socket-dir).
|
[`docs/trust-boundary/boundary.md`](../trust-boundary/boundary.md#the-per-agent-socket-dir).
|
||||||
3. **Marker gate**. After successful `bind_unix`, the harness drops
|
3. **Marker gate**. After a successful `bind_unix`, the harness drops
|
||||||
`<dir>/hyperhive-socket-bound` next to the socket. c0re's
|
`<dir>/hyperhive-socket-bound` next to the socket. c0re's
|
||||||
`agent_sockets::write` filters its JSON map by marker presence —
|
`agent_sockets::write` filters its JSON map by marker presence —
|
||||||
only agents whose harness has actually bound the socket appear there.
|
only agents whose harness has bound the socket appear there.
|
||||||
(Legacy name `.bound` also accepted during the transition window.)
|
It also accepts the older `.bound` name.
|
||||||
4. **Gateway side**. `gateway_nginx::write` generates
|
4. **Gateway side**. `gateway_nginx::write` generates
|
||||||
`/var/lib/hive-gateway/conf/agents.conf` — a plain nginx include
|
`/var/lib/hive-gateway/conf/agents.conf` — a plain nginx include
|
||||||
file with one `location /agent/<name>/` block per agent. Always
|
file with one `location /agent/<name>/` block per agent. Always
|
||||||
|
|
@ -152,275 +339,13 @@ on subsequent poll ticks.
|
||||||
`agents.conf` uses atomic `<path>.tmp` + `rename()` writes so a crashing
|
`agents.conf` uses atomic `<path>.tmp` + `rename()` writes so a crashing
|
||||||
c0re process never leaves a partial or unparseable file behind.
|
c0re process never leaves a partial or unparseable file behind.
|
||||||
|
|
||||||
## Dashboard link shape (gateway vs direct)
|
### Per-agent static frontend split
|
||||||
|
|
||||||
When the gateway is in front, the SW4RM tab builds per-agent links
|
hive-c0re's environment sets `HIVE_AGENT_FRONTEND_DIR` to the `agent/`
|
||||||
as same-origin `/agent/<name>/…` URLs instead of the legacy direct
|
subdirectory of `services.hyperhive.c0re.servedFrontend`.
|
||||||
`http://<host>:<container.port>/` TCP shape. The signal comes from
|
|
||||||
`StateSnapshot.gateway_enabled`, sourced from the
|
|
||||||
`HIVE_GATEWAY_ENABLED` env the c0re NixOS module now always sets
|
|
||||||
(`services.hyperhive.gateway.enable` no longer exists — the gateway runs
|
|
||||||
unconditionally alongside hyperhive), so this is effectively always
|
|
||||||
true; the `false` branch stays as a defensive fallback for the
|
|
||||||
env being unset. Three render sites
|
|
||||||
flip together: the primary agent-name link, the favicon fetch
|
|
||||||
(`<url>/icon`), and the nav-strip `container`-kind links from
|
|
||||||
`DashboardState.links` (`GET /api/dashboard-state`). `forge`-kind nav-strip links still
|
|
||||||
resolve against `http://<host>:3000` (separate sub-domain transition
|
|
||||||
tracked by `deploy.forgejo.behindGateway`); `external`-kind links are
|
|
||||||
already absolute. See `docs/web-ui/dashboard.md::Container row` for the
|
|
||||||
frontend-side derivation.
|
|
||||||
|
|
||||||
## TLS modes
|
|
||||||
|
|
||||||
The gateway always terminates TLS — self-signed is the implicit floor when
|
|
||||||
the operator configures nothing else, so there is no http-only mode. Three
|
|
||||||
modes, selected by which (if any) external TLS source the operator sets:
|
|
||||||
|
|
||||||
⚠️ **One vhost class is exempt, so expect it during a TLS audit.** A vhost bound to loopback for a
|
|
||||||
local consumer carries a single plain-HTTP listen and no TLS. `grafana-metrics`
|
|
||||||
(`nix/host-modules/swarm-grafana.nix`) is the only one in the tree today: it listens on `127.0.0.1`
|
|
||||||
alone and serves one `= /metrics` location from grafana's unix socket, for the collector on this
|
|
||||||
host to scrape. Nothing off-host can reach it, so TLS there protects nothing. The modes below
|
|
||||||
cover every vhost with a routable name.
|
|
||||||
|
|
||||||
| mode | config | cert source | `.well-known` scheme |
|
|
||||||
|---|---|---|---|
|
|
||||||
| self-signed (default) | neither `tls.certDir` nor `tls.acme` set | host hive-CA signs a gateway leaf (RSA-4096) | `https` |
|
|
||||||
| ACME (Let's Encrypt) | `tls.acme.enable = true` | nginx via HTTP-01 | `https` |
|
|
||||||
| operator cert | `tls.certDir` set | read from the operator's dir | `https` |
|
|
||||||
|
|
||||||
The `gateway.selfSignedTls` option has been **removed** — self-signed
|
|
||||||
is now derived from the absence of `tls.certDir` / `tls.acme`. A config
|
|
||||||
that still sets it fails eval with a removal message; use `tls.certDir`
|
|
||||||
/ `tls.acme` to override the default.
|
|
||||||
|
|
||||||
### ACME / Let's Encrypt (`tls.acme`)
|
|
||||||
|
|
||||||
Simplest production path for operators with a public domain:
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.gateway = {
|
|
||||||
openFirewall = true;
|
|
||||||
tls.acme = {
|
|
||||||
enable = true;
|
|
||||||
email = "admin@example.com";
|
|
||||||
};
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
nginx obtains and autorenews certs via the ACME HTTP-01 challenge on `port` (default 80). Certs land in `/var/lib/acme/` on the host, managed by nixpkgs's `security.acme` in the ordinary way.
|
|
||||||
|
|
||||||
**Requirements**: `services.hyperhive.domain` must be publicly DNS-resolvable to this host, and `openFirewall = true` so Let's Encrypt can reach `/.well-known/acme-challenge/`. Each active vhost (main domain, `forge.<swarm-domain>`, `chat.<swarm-domain>`) gets its own cert via separate ACME challenges — the swarm services default to names under `services.hyperhive.swarm.domain`, so **every one of those names must resolve to this host too**, not just the hive's own.
|
|
||||||
|
|
||||||
Mutual exclusion: `tls.certDir` set together with `tls.acme.enable = true` fails an assertion — pick one external TLS source (or neither, for the self-signed default).
|
|
||||||
|
|
||||||
### Self-signed TLS (default)
|
|
||||||
|
|
||||||
On by default, and listens on `httpsPort` (default 443) on every routable vhost beside the plain-http `port` (default 80). (See the loopback-only exception under *TLS modes* above.)
|
|
||||||
|
|
||||||
The issuer is a **host-held hive CA**, not a bare self-signed leaf. A host service (`hive-tls-ca.service`, from the `hive-tls` module) generates a long-lived CA (`services.hyperhive.deploy.hive-controller.tls.caValidityDays`, default ~20y) under `services.hyperhive.deploy.hive-controller.tls.stateDir` (default `/var/lib/hive-tls`), then signs a gateway **leaf** (`leafValidityDays`, default 30d) with it. `hive-gateway-self-signed-cert` then imports the leaf into nginx's state dir (`/var/lib/hive-gateway/tls/{cert,key}.pem`).
|
|
||||||
|
|
||||||
⚠️ **Don't collapse that import unit into pointing nginx at the CA dir.**
|
|
||||||
It does two jobs, and skipping it has taken the gateway down in production
|
|
||||||
before. It re-modes the leaf (`hive-tls-ca` writes the key `0600
|
|
||||||
root:root`; nginx's pre-start `nginx -t` runs as the *nginx user*, so a
|
|
||||||
`0600` key fails the config test and blocks the unit), and it guarantees
|
|
||||||
**every cert path the nginx config names exists** — which is what the
|
|
||||||
swarm-services fallback below is for.
|
|
||||||
|
|
||||||
**Why a CA, not a bare leaf**: a bare self-signed leaf is its own trust anchor, so every regeneration is a new anchor every consumer must re-trust — and nothing can wire a runtime-generated leaf into an agent's build-time trust store at all. With a stable CA, agents and federation peers trust it *once*; leaf rotation never re-breaks them.
|
|
||||||
|
|
||||||
**What consumers trust**: `trust-bundle.pem` in the same state dir, not `ca.pem`. The hive CA is itself issued under the swarm root ([`swarm/ca.md`](../swarm/ca.md) has the hierarchy), and an intermediate isn't a chain a verifier can terminate at — so the bundle carries the hive CA plus whatever it's rooted at. `hive-gateway-self-signed-cert` hands nginx the leaf with the hive CA appended for the same reason. Everything that trusts the hive's TLS reads the bundle: agents (via `security.pki.certificateFiles`), the CI and forge containers, and a federating peer.
|
|
||||||
|
|
||||||
**Why on by default**: matrix-dart-sdk (FluffyChat's SDK) hardcodes `https://<host>/.well-known/matrix/client` for homeserver discovery and refuses to fall back to plain http. Without TLS the browser client can't bootstrap.
|
|
||||||
|
|
||||||
**Cert shape**: leaf subject CN = bare hive domain; subjectAltName is `<hive>` plus wildcard `*.<hive>`, so all current and future sub-domain vhosts validate under the same leaf + the hive CA. You can't add a swarm service whose name is *not* under this hive's domain here — the hive CA is name-constrained to `<hive>`, and a violating SAN invalidates the whole leaf, not just that name. Those names get the swarm-services leaf instead ([`swarm/ca.md`](../swarm/ca.md)).
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
**Rotation**: `hive-tls-ca.service` is idempotent — it re-signs the leaf when it's missing or within 30 days of expiry, always under the same CA (so consumer trust is undisturbed). It regenerates the CA itself only if missing or already expired. To force a leaf rotation, delete `gateway.pem` under the state dir and restart the unit, then reload `nginx`.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
**Cert prompts**: browsers still warn once per host until the operator adds the hive's `trust-bundle.pem` to the browser/OS trust store (an anchor, not the leaf, is the thing to trust). A separate mechanism wires agent trust (see the agent-trust work for `/run/hive-ca`).
|
|
||||||
|
|
||||||
### Operator-provided cert (`tls.certDir`)
|
|
||||||
|
|
||||||
For operators with a real CA cert (Let's Encrypt, corporate CA, etc.):
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.gateway = {
|
|
||||||
tls.certDir = "/var/lib/acme/example.com"; # nixpkgs security.acme output dir
|
|
||||||
# tls.certName = "cert.pem"; # default — matches security.acme layout
|
|
||||||
# tls.keyName = "key.pem"; # default — matches security.acme layout
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
nginx reads the directory directly and uses `cert.pem` + `key.pem` (override `tls.certName`/`tls.keyName` for different filenames). Both modes listen on `httpsPort` (default 443) and emit `https://` in `.well-known` responses.
|
|
||||||
|
|
||||||
`tls.certDir` and `tls.acme.enable` set together is an assertion error.
|
|
||||||
|
|
||||||
**Key file permissions**: nixpkgs's `security.acme` outputs private keys as `0640 root:acme` by default. nginx runs as the `nginx` user and can't read a key with that ownership. Fix with:
|
|
||||||
|
|
||||||
```nix
|
|
||||||
security.acme.certs."example.com".group = "nginx";
|
|
||||||
```
|
|
||||||
|
|
||||||
or make the key world-readable (`0644`) if your threat model allows it. nginx errors out at startup on a key it can't read — the error is explicit in the journal, not a silent failure.
|
|
||||||
|
|
||||||
### Fronting with an external TLS terminator
|
|
||||||
|
|
||||||
No http-only mode exists (see [TLS modes](#tls-modes) above). Two paths
|
|
||||||
for an operator who wants their own TLS terminator:
|
|
||||||
|
|
||||||
- give the gateway the real cert via `tls.certDir` (or `tls.acme`) so it
|
|
||||||
serves proper TLS directly — no separate proxy needed; or
|
|
||||||
- front it over a **unix socket** rather than a plain-http TCP port (the
|
|
||||||
intended direction for "bring your own proxy" — the gateway isn't meant
|
|
||||||
to expose an unencrypted TCP upstream).
|
|
||||||
|
|
||||||
Because of this, `.well-known/matrix/{client,server}` discovery responses
|
|
||||||
always advertise `https` (see [Discovery flow](#discovery-flow-matrix) above).
|
|
||||||
|
|
||||||
## Firewall posture (host-level)
|
|
||||||
|
|
||||||
The gateway is unconditional — `services.hyperhive.gateway.enable` no
|
|
||||||
longer exists, there is no gateway-off mode. nginx is always the sole
|
|
||||||
external entry point and routes to agents over the UDS upstream
|
|
||||||
described above (see [Per-agent unix-socket
|
|
||||||
upstream](#per-agent-unix-socket-upstream)), so the per-agent web-port
|
|
||||||
range `8100..8999` stays closed on the host firewall
|
|
||||||
unconditionally — opening it would defeat the single-front-door story.
|
|
||||||
The hashed TCP port (`lifecycle::agent_web_port`) still exists as a
|
|
||||||
fallback bind for an agent whose `HIVE_WEB_SOCKET` env somehow ends up
|
|
||||||
unset, but nothing opens a matching firewall hole for it and the
|
|
||||||
gateway itself never proxies through it.
|
|
||||||
|
|
||||||
`services.hyperhive.gateway.openFirewall = true` opens both `port` and
|
|
||||||
`httpsPort` — both are always served, since the gateway always terminates
|
|
||||||
TLS (see [TLS modes](#tls-modes) above).
|
|
||||||
|
|
||||||
Every agent hashes into the same port range (no special case), so
|
|
||||||
one range opening covers every container.
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
The dashboard port (`services.hyperhive.c0re.dashboardPort`, default 7000) is *not*
|
|
||||||
listed in either case — it binds `127.0.0.1` only, so a firewall
|
|
||||||
hole would be a no-op. Remote dashboard access flows through the
|
|
||||||
gateway. Operators who opt out of the gateway lose external
|
|
||||||
dashboard reach by design — the surface is privileged (approve /
|
|
||||||
deny / destroy), and operators must not expose it without a real reverse
|
|
||||||
proxy in front.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
## `HIVE_FORGE_URL`: agents reach the forge via the gateway by domain
|
|
||||||
|
|
||||||
Agents poll `HIVE_FORGE_URL` for Forgejo notifications + run all
|
|
||||||
`hive-forge` calls against it. Network isolation is always on (the
|
|
||||||
shared-netns mode no longer exists), so agents run in a private netns and
|
|
||||||
can never reach the host's loopback.
|
|
||||||
`nix/host-modules/hive-c0re/environment.nix` sets `HIVE_FORGE_URL` to
|
|
||||||
`http://<forge.domain>` (default `forge.<swarm-domain>` — a swarm runs
|
|
||||||
one forge; you must set `services.hyperhive.domain`). Agents
|
|
||||||
get the bridge dnsmasq as their resolver, resolve the hostname →
|
|
||||||
bridge IP, then reach nginx on port 80 (the bridge firewall opens
|
|
||||||
80+443). nginx proxies to forgejo — the same path an operator browser
|
|
||||||
takes, no raw port exposure needed.
|
|
||||||
|
|
||||||
## hive-forge container shape
|
|
||||||
|
|
||||||
Private Forgejo wrapped in a nixos-container (`hive-forge`, not
|
|
||||||
`h-*` — keeps c0re's lifecycle scanner out of the picture; the
|
|
||||||
operator manages it via the standard `nixos-container` CLI). The
|
|
||||||
container also keeps hive-forge from fighting any `services.forgejo`
|
|
||||||
the operator already runs on the host — separate systemd namespace,
|
|
||||||
separate state dir, separate port unless the operator deliberately
|
|
||||||
collides.
|
|
||||||
|
|
||||||
The forge container shares the host network namespace
|
|
||||||
(`privateNetwork = false`), so forgejo's listeners look like a
|
|
||||||
host-side service — nixos-container is here for state + systemd-unit
|
|
||||||
isolation, not network isolation. Note this is the FORGE container;
|
|
||||||
agent containers are network-isolated and reach the forge through the
|
|
||||||
gateway by `forge.<swarm-domain>` (see `HIVE_FORGE_URL` above), not via
|
|
||||||
the host's loopback.
|
|
||||||
|
|
||||||
State lives at `/var/lib/nixos-containers/hive-forge/var/lib/forgejo/`
|
|
||||||
and survives container restart / host reboot. To wipe, destroy the
|
|
||||||
container.
|
|
||||||
|
|
||||||
The container runs only on the swarm's forge host
|
|
||||||
(`services.hyperhive.deploy.forgejo.enable`, see
|
|
||||||
[`../swarm/services.md`](../swarm/services.md)). Every other hive
|
|
||||||
reaches that host's gateway by `forge.<swarm-domain>`.
|
|
||||||
|
|
||||||
### Network and port configuration
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.swarm.forge = {
|
|
||||||
httpPort = 3000; # default — HTTP listener; outside hyperhive's 7000/8100-8999 range
|
|
||||||
sshPort = 2222; # default — git-over-SSH; kept off 22 so it doesn't collide with the host openssh
|
|
||||||
};
|
|
||||||
# Which ports the forge answers on is swarm-wide; whether THIS host opens
|
|
||||||
# them in its firewall is a deployment decision, so it lives under deploy.*.
|
|
||||||
services.hyperhive.deploy.forgejo.openFirewall = false; # default
|
|
||||||
```
|
|
||||||
|
|
||||||
`httpPort` (default **3000**) is the port Forgejo's HTTP server binds to.
|
|
||||||
It sits outside hyperhive's reserved ranges (dashboard 7000,
|
|
||||||
agents 8100–8999) so a default install has no port fights. Change it
|
|
||||||
only if you already have another process bound to 3000.
|
|
||||||
|
|
||||||
`sshPort` (default **2222**) is the port Forgejo's built-in SSH server
|
|
||||||
uses for `git clone/push/pull` over SSH (`git@<domain>:owner/repo.git`
|
|
||||||
via `-p 2222`). Port 22 stays alone on the host for openssh.
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
`openFirewall` (default **false**) controls whether the host firewall
|
|
||||||
opens `httpPort` and `sshPort`. Off by default (secure by
|
|
||||||
default): agents reach Forgejo through the gateway (`forge.<swarm-domain>` on
|
|
||||||
the bridge), not the raw port, so no firewall hole is needed. Flip to
|
|
||||||
`true` when you need:
|
|
||||||
- The operator's browser to reach `http://<host>:<httpPort>/` directly
|
|
||||||
(not behind the gateway).
|
|
||||||
- External git clients that push/pull via SSH directly to the host.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
Forgejo served through the gateway (`deploy.forgejo.behindGateway = true`) does
|
|
||||||
not need `openFirewall` — the gateway's own `openFirewall` option covers
|
|
||||||
that path.
|
|
||||||
|
|
||||||
### `rootUrl` override
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.swarm.forge.rootUrl = "https://forge.example.com/";
|
|
||||||
```
|
|
||||||
|
|
||||||
`rootUrl` (default **null**) overrides the Forgejo `ROOT_URL` that's
|
|
||||||
autoderived from `forge.domain` + gateway state. The autoderivation
|
|
||||||
covers most cases:
|
|
||||||
|
|
||||||
| Shape | Autoderived `ROOT_URL` |
|
|
||||||
|---|---|
|
|
||||||
| `deploy.forgejo.behindGateway = true` | `https://<forge.domain>/` (port suffix omitted when `gateway.httpsPort == 443`) |
|
|
||||||
| `deploy.forgejo.behindGateway = false` | `http://<forge.domain>:<httpPort>/` |
|
|
||||||
|
|
||||||
The gateway always terminates TLS, so the `behindGateway = true` case is
|
|
||||||
always advertised over `https://`; only the direct (`behindGateway =
|
|
||||||
false`) shape stays `http://`. Set `rootUrl` explicitly when
|
|
||||||
`forge.domain` resolves differently from the public URL, or for a
|
|
||||||
genuinely bespoke shape (for example an external reverse proxy on a different
|
|
||||||
host/path). Must end with `/` (Forgejo requirement; an assertion
|
|
||||||
enforces this).
|
|
||||||
|
|
||||||
## Per-agent static frontend split
|
|
||||||
|
|
||||||
hive-c0re injects `HIVE_AGENT_FRONTEND_DIR` into its service
|
|
||||||
environment, pointing at the `agent/` subdirectory of the frontend dist
|
|
||||||
set by `services.hyperhive.c0re.frontend`. That option has a default, so
|
|
||||||
the variable is always set.
|
|
||||||
The nginx include generator (`gateway_nginx::write`) reads
|
The nginx include generator (`gateway_nginx::write`) reads
|
||||||
this variable and, when set, emits split location blocks per agent
|
this variable and, when set, emits split location blocks per agent
|
||||||
instead of the legacy single-proxy block.
|
instead of a single proxy block.
|
||||||
|
|
||||||
**Location priority for `/agent/<name>/...`:**
|
**Location priority for `/agent/<name>/...`:**
|
||||||
|
|
||||||
|
|
@ -461,8 +386,7 @@ checking file existence):
|
||||||
| `/agent/iris/events/live` | no file match → `@iris_dynamic` | proxied (SSE) |
|
| `/agent/iris/events/live` | no file match → `@iris_dynamic` | proxied (SSE) |
|
||||||
|
|
||||||
Adding a new HTML page to the frontend dist (`dist/<page>.html`)
|
Adding a new HTML page to the frontend dist (`dist/<page>.html`)
|
||||||
automatically makes it reachable at `/agent/<name>/<page>` — no
|
makes it reachable at `/agent/<name>/<page>` with no generator change.
|
||||||
generator change needed.
|
|
||||||
|
|
||||||
**Why `^~` for `/static/`**: the `^~` prefix gives this block higher
|
**Why `^~` for `/static/`**: the `^~` prefix gives this block higher
|
||||||
priority than the plain prefix `location /agent/<name>/`, so compiled
|
priority than the plain prefix `location /agent/<name>/`, so compiled
|
||||||
|
|
@ -475,169 +399,59 @@ store path baked in at hive-c0re build time, and c0re (writing
|
||||||
`agents.conf`) and nginx (serving files from it) are on the same machine,
|
`agents.conf`) and nginx (serving files from it) are on the same machine,
|
||||||
so they see the same store.
|
so they see the same store.
|
||||||
|
|
||||||
**Graceful degradation**: if `HIVE_AGENT_FRONTEND_DIR` is empty or
|
**Unset variable**: if `HIVE_AGENT_FRONTEND_DIR` is empty or unset, each
|
||||||
unset (for example a build that predates `services.hyperhive.c0re.frontend`), each agent gets the
|
agent gets a single proxy block and nginx forwards all traffic to the
|
||||||
legacy single-proxy block and nginx forwards all traffic to the agent
|
agent daemon.
|
||||||
daemon as before.
|
|
||||||
|
|
||||||
**`extraFiles`**: per-agent `services.hyperhive.agent.frontend.extraFiles` are in
|
**`extraFiles`**: per-agent `services.hyperhive.agent.frontend.extraFiles` sit in
|
||||||
`mergedDist`, not in the base `services.hyperhive.c0re.frontend` dist. They're not under
|
`mergedDist`, not in the base `services.hyperhive.c0re.frontend` dist. They're not under
|
||||||
the nix-store `alias` path, so requests for them fall through
|
the nix-store `alias` path, so requests for them fall through
|
||||||
`try_files` to `@<name>_dynamic`, and the agent daemon serves them
|
`try_files` to `@<name>_dynamic`, and the agent daemon serves them.
|
||||||
as before.
|
|
||||||
|
|
||||||
## Per-agent error pages
|
### Per-agent error pages
|
||||||
|
|
||||||
`/agent/<name>/` requests hit two failure modes; both get static
|
`/agent/<name>/` requests hit two failure modes; both get static
|
||||||
HTML pages instead of nginx's default error chrome:
|
HTML pages instead of nginx's default error chrome:
|
||||||
|
|
||||||
- **Agent not found** (`/agent/<unknown>/...`) — name isn't in
|
- **Agent not found** (`/agent/<unknown>/...`) — no `location` for that
|
||||||
`agentPortsTable`. nginx's prefix match falls back to the bare
|
name in `agents.conf`. nginx's prefix match falls back to the bare
|
||||||
`/agent/` catch-all, which `return 404`s and `error_page 404` rewrites
|
`/agent/` catch-all, which `return 404`s and `error_page 404` rewrites
|
||||||
to `/__hive_agent_not_found` → serves `not-found.html` with a link
|
to `/__hive_agent_not_found` → serves `not-found.html` with a link
|
||||||
back to the dashboard.
|
back to the dashboard.
|
||||||
|
|
||||||
- **Agent unreachable** (`502 / 503 / 504` from `proxy_pass`) — the
|
- **Agent unreachable** (`502 / 503 / 504` from `proxy_pass`) — the
|
||||||
per-agent harness isn't responding (container restarting, crash
|
per-agent harness isn't responding (container restarting, crash
|
||||||
recovery, etc.). `proxy_intercept_errors on` + `error_page 502 503
|
recovery). `proxy_intercept_errors on` + `error_page 502 503
|
||||||
504 = /__hive_agent_unreachable` rewrites to `unreachable.html`.
|
504 = /__hive_agent_unreachable` rewrites to `unreachable.html`.
|
||||||
|
|
||||||
`pkgs.runCommand` builds both pages at deploy time (one nix
|
`hive-gateway/error-pages.nix` renders every page (`notFound`, `unreachable`,
|
||||||
derivation `hyperhive-agent-error-pages` with `not-found.html` +
|
`unauthorized`, `ssoUnavailable`) from one template with `pkgs.writeText`;
|
||||||
`unreachable.html` inside), served via two `internal` nginx
|
nginx serves each through an `internal` location with `alias` to the file,
|
||||||
locations with `alias` to the exact file. `internal` keeps the
|
so only nginx's own error handling can reach them. They depend on nothing
|
||||||
files from being directly request-able by operators — only nginx's
|
from the frontend dist and render even when hive-c0re is down. A service
|
||||||
own error-handling can reach them.
|
module reaches them through `gateway.lib.errorPages`.
|
||||||
|
|
||||||
Page styling: minimal inline CSS matching the dashboard's catppuccin
|
A route earns a custom page when the default status code would blame the
|
||||||
palette (`#1e1e2e` bg, `#cdd6f4` text, `#cba6f7` heading). No
|
wrong component. The per-agent routes qualify (a 502 there means the
|
||||||
dependencies on the frontend dist — these pages render even when
|
harness is restarting, not a gateway fault), and so does
|
||||||
hive-c0re itself is down.
|
`auth.<swarm>` — a dead authelia upstream almost always means no users
|
||||||
|
yet, and a bare 502 blames the proxy, the one part that works.
|
||||||
|
Forge, matrix and fluffychat keep nginx's defaults: a dead upstream there
|
||||||
|
means what the status code says.
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
### Dashboard link shape (gateway vs direct)
|
||||||
Scope is intentionally narrow: a route earns a custom page when the
|
|
||||||
default status code would point at the wrong component. The per-agent
|
|
||||||
routes qualify (a 502 there means the harness is restarting, not that
|
|
||||||
the gateway is broken), and so does `auth.<swarm>` — a dead authelia
|
|
||||||
upstream almost always means the user store was never bootstrapped, and
|
|
||||||
a bare 502 blames the proxy, which is the one part that's working.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
Forge / matrix / fluffychat still get nginx defaults: their upstreams
|
The dashboard builds per-agent links as same-origin `/agent/<name>/…` URLs
|
||||||
being down means what the status code says, so a themed page would add
|
when `StateSnapshot.gateway_enabled` is true, and direct
|
||||||
styling and no information.
|
`http://<host>:<port>/` links otherwise. The flag comes from
|
||||||
|
`HIVE_GATEWAY_ENABLED`, which `hive-c0re/environment.nix` sets to `1` on
|
||||||
|
every hive, so the direct shape is a fallback for the variable being unset.
|
||||||
|
Three render sites follow it: the agent-name link, the favicon fetch
|
||||||
|
(`<url>/icon`), and the nav-strip `container`-kind links from
|
||||||
|
`DashboardState.links` (`GET /api/dashboard-state`). Forge links come from
|
||||||
|
`swarm.forge.publicUrl`; with it unset the dashboard hides them rather than
|
||||||
|
guess. See `docs/web-ui/dashboard.md::Container row` for the frontend side.
|
||||||
|
|
||||||
## HTTP Basic auth
|
### Dialing another vhost by name (`verifiedProxyTo`)
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
`services.hyperhive.gateway.auth.enable = true` gates every request to
|
|
||||||
the main vhost (`_`) behind HTTP Basic auth. nginx's built-in `auth_basic`
|
|
||||||
module validates credentials; no extra service or host-side daemon is
|
|
||||||
required.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
**Setup:**
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.gateway.auth = {
|
|
||||||
enable = true;
|
|
||||||
# realm = "hyperhive"; # optional, default shown
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
The credential store lives at the fixed path
|
|
||||||
`/var/lib/hive-gateway/conf/gateway.htpasswd` on the host. A tmpfiles
|
|
||||||
rule pre-creates the file on first boot; no manual path configuration
|
|
||||||
is required. nginx reads it at that path directly.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
Manage users with `hivectl gateway`. `hivectl` sends the request over the
|
|
||||||
host admin socket and the `hive-c0re` daemon performs the write at its
|
|
||||||
canonical path — the daemon never exposes a path to the CLI:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
# Add or update a user (prompted for password):
|
|
||||||
hivectl gateway create-user alice --password-stdin
|
|
||||||
|
|
||||||
# Add with inline password (visible in shell history — avoid for sensitive creds):
|
|
||||||
hivectl gateway create-user bob --password hunter2
|
|
||||||
|
|
||||||
# Remove a user:
|
|
||||||
hivectl gateway delete-user bob
|
|
||||||
|
|
||||||
# List current usernames:
|
|
||||||
hivectl gateway list-users
|
|
||||||
```
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
The daemon hashes passwords with BCrypt (cost 12) and writes
|
|
||||||
`$2y$`-prefixed hashes that nginx accepts natively. No external
|
|
||||||
`htpasswd` binary is required.
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
**What's not gated:** per-agent UI routes emitted into `agents.conf`
|
|
||||||
(served under `/agent/<name>/`) inherit no auth from `/` — nginx
|
|
||||||
applies `auth_basic` per-location. Full per-agent coverage is a
|
|
||||||
follow-up.
|
|
||||||
|
|
||||||
**Realm:** the `WWW-Authenticate: Basic realm="..."` string browsers
|
|
||||||
display in the credential dialog. Defaults to `"hyperhive"`. Must not
|
|
||||||
contain `"` or `$`.
|
|
||||||
|
|
||||||
**Custom 401 page:** when credentials are absent or wrong, nginx serves
|
|
||||||
a Catppuccin-styled `unauthorized.html` page (built into the same Nix
|
|
||||||
derivation as the agent error pages) that tells the operator which
|
|
||||||
`hivectl` command to run to create a user. The response status is still
|
|
||||||
`401` (`error_page 401 =401 /__hive_auth_unauthorized`) so browsers
|
|
||||||
present the login dialog on the first visit — users who dismiss the
|
|
||||||
dialog see the human-readable hint. The internal exact-match location
|
|
||||||
(`= /__hive_auth_unauthorized`) beats `location /` in nginx's prefix
|
|
||||||
ordering, preventing the subrequest from looping back through
|
|
||||||
`auth_basic`.
|
|
||||||
|
|
||||||
## Security headers
|
|
||||||
|
|
||||||
The gateway emits the following headers at server scope on every
|
|
||||||
vhost (`_`, `forge.<swarm-domain>`, `chat.<swarm-domain>`):
|
|
||||||
|
|
||||||
| Header | Value |
|
|
||||||
|--------|-------|
|
|
||||||
| `X-Frame-Options` | `SAMEORIGIN` |
|
|
||||||
| `X-Content-Type-Options` | `nosniff` |
|
|
||||||
| `Referrer-Policy` | `strict-origin-when-cross-origin` |
|
|
||||||
|
|
||||||
nginx's `add_header` inheritance rule: a `location` block that sets its
|
|
||||||
own `add_header` does **not** inherit server-scope headers. API locations
|
|
||||||
that carry their own CORS headers (for example `/.well-known/matrix/client`,
|
|
||||||
`/_matrix/`) are therefore unaffected. HTML-serving and proxy locations
|
|
||||||
with no `add_header` of their own pick the security headers up
|
|
||||||
automatically.
|
|
||||||
|
|
||||||
### HSTS (`gateway.hsts`)
|
|
||||||
|
|
||||||
HSTS is **opt-in** and disabled by default:
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.gateway.hsts = {
|
|
||||||
enable = true; # default: false
|
|
||||||
maxAge = 31536000; # default: 1 year (required for preload list)
|
|
||||||
includeSubDomains = true; # default: true
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
When enabled, the gateway adds a `Strict-Transport-Security: max-age=...[; includeSubDomains]`
|
|
||||||
header alongside the other security headers.
|
|
||||||
|
|
||||||
**Opt-in rationale**: HSTS pins HTTPS in the browser's preload cache;
|
|
||||||
enabling it on a deployment that later loses TLS locks browsers out
|
|
||||||
until `max-age` expires. Only enable when TLS is permanent.
|
|
||||||
|
|
||||||
Since the gateway always terminates TLS (see [TLS modes](#tls-modes)
|
|
||||||
above), an enabled HSTS header is always served over https — there is no
|
|
||||||
TLS-less mode that could violate it.
|
|
||||||
|
|
||||||
## Dialing another vhost by name (`verifiedProxyTo`)
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
`vhost-lib.nix`'s `verifiedProxyTo` builds the `proxy_ssl_*` /
|
`vhost-lib.nix`'s `verifiedProxyTo` builds the `proxy_ssl_*` /
|
||||||
|
|
@ -682,4 +496,3 @@ recursing on its own `auth_request` until nginx's subrequest-depth limit
|
||||||
turns it into a plain 500. Every call site sets `recommendedProxySettings
|
turns it into a plain 500. Every call site sets `recommendedProxySettings
|
||||||
= false` on the location for exactly this reason — nixpkgs' version
|
= false` on the location for exactly this reason — nixpkgs' version
|
||||||
would still clobber this one.
|
would still clobber this one.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,14 +1,22 @@
|
||||||
# hive-network
|
# hive-network
|
||||||
|
|
||||||
Host-side bridge + per-agent private-netns isolation, up on a host
|
The bridge every container on a host attaches to, and the private-netns
|
||||||
where something attaches to it (`services.hyperhive.network.enable`,
|
isolation of agent and CI containers behind it. It comes up on any host that runs agents or swarm services:
|
||||||
asserted by the modules that need it rather than set by hand).
|
`services.hyperhive.network.enable` defaults to off, and the modules that
|
||||||
Configured via `services.hyperhive.network.*`.
|
need the bridge set it with `mkDefault true` — the hive controller, CI,
|
||||||
|
the hive collector, and the gateway's resolver, which every swarm service
|
||||||
|
turns on. Configured via `services.hyperhive.network.*`.
|
||||||
|
|
||||||
> Isolation is the only mode — there is no shared-netns fallback. The
|
Isolation is the only mode; agent containers never share the host netns.
|
||||||
> former `services.hyperhive.network.isolateContainers` and
|
|
||||||
> `services.hyperhive.network.upstreamDns` options no longer exist; a
|
<details><summary>Upgrading a config that sets isolateContainers or upstreamDns</summary>
|
||||||
> config that still sets one fails eval with a removal message.
|
|
||||||
|
Neither option exists. A config that still sets
|
||||||
|
`services.hyperhive.network.isolateContainers` or
|
||||||
|
`services.hyperhive.network.upstreamDns` fails eval with a removal
|
||||||
|
message; drop the line.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
## Network map
|
## Network map
|
||||||
|
|
||||||
|
|
@ -45,7 +53,7 @@ that touches it.
|
||||||
|
|
||||||
| container | netns | IPv4 | listens / reached via |
|
| container | netns | IPv4 | listens / reached via |
|
||||||
| -------------- | ----------------------- | -------------- | ----------------------------------------------------------------------------------------------- |
|
| -------------- | ----------------------- | -------------- | ----------------------------------------------------------------------------------------------- |
|
||||||
| `hive-gateway` | host (shared) | host addresses | nginx `:80`/`:443` (every vhost); dnsmasq `bridgeIp:53` + DHCP `:67` on the bridge |
|
| gateway (host) | host | host addresses | nginx `:80`/`:443` (every vhost); dnsmasq `bridgeIp:53` + DHCP `:67` on the bridge |
|
||||||
| `hive-forge` | host (shared) | host addresses | forgejo `:3000` http, `:2222` git-ssh; fronted by the `forge.<swarm-domain>` vhost |
|
| `hive-forge` | host (shared) | host addresses | forgejo `:3000` http, `:2222` git-ssh; fronted by the `forge.<swarm-domain>` vhost |
|
||||||
| `hive-matrix` | host (shared) | host addresses | tuwunel `:8008` (+ optional federation port); fronted by the matrix vhost |
|
| `hive-matrix` | host (shared) | host addresses | tuwunel `:8008` (+ optional federation port); fronted by the matrix vhost |
|
||||||
| `hive-ci` | private, veth on bridge | DHCP pool | outbound only (runner → forge); no inbound surface |
|
| `hive-ci` | private, veth on bridge | DHCP pool | outbound only (runner → forge); no inbound surface |
|
||||||
|
|
@ -79,11 +87,9 @@ The flows, end to end:
|
||||||
|
|
||||||
## Container shape (where dnsmasq lives)
|
## Container shape (where dnsmasq lives)
|
||||||
|
|
||||||
Co-located in the existing `hive-gateway` container — single
|
dnsmasq runs on the host itself, next to nginx — both come from the
|
||||||
front-door for both DNS and HTTP, saves a sibling container, single
|
gateway module, one front door for DNS and HTTP. Its `bind-interfaces`
|
||||||
systemd-unit / state surface to monitor. The gateway shares host
|
listener sits on `bridgeIp`, on the host's bridge interface.
|
||||||
netns (`privateNetwork = false`) so dnsmasq's `bind-interfaces`
|
|
||||||
listener on `bridgeIp` is on the host's bridge interface.
|
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
|
|
@ -99,10 +105,11 @@ listener on `bridgeIp` is on the host's bridge interface.
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
You must set `services.hyperhive.domain` — the dnsmasq resolver
|
The hive domain (`services.hyperhive.domain`, which the resolver answers
|
||||||
is authoritative for `<hive-domain>` and its sub-domains. You don't
|
for) comes from this hive's entry in the swarm directory:
|
||||||
write it: it's read from this hive's entry in the swarm directory
|
`swarm.hives.<hiveName>.domain`, default `<hiveName>.<swarm.domain>`.
|
||||||
(`docs/swarm/README.md` § Hive identity config).
|
Eval fails until `swarm.domain` and that entry exist
|
||||||
|
([`swarm/README.md`](../swarm/README.md) § Hive identity config).
|
||||||
|
|
||||||
## Bridge addressing
|
## Bridge addressing
|
||||||
|
|
||||||
|
|
@ -162,13 +169,13 @@ agent containers.
|
||||||
receives DHCP via a regular UDP socket (it doesn't use a
|
receives DHCP via a regular UDP socket (it doesn't use a
|
||||||
netfilter-bypassing raw socket), so the hole is mandatory — without
|
netfilter-bypassing raw socket), so the hole is mandatory — without
|
||||||
it containers never get a lease and fall back to 169.254.x.x.
|
it containers never get a lease and fall back to 169.254.x.x.
|
||||||
- Ports 80 and 443 let isolated agents reach nginx (gateway
|
- Ports 80 and 443 let isolated agents reach nginx (on the host) for the forge sub-domain, per-agent
|
||||||
container, shared host netns) for the forge sub-domain, per-agent
|
|
||||||
UI proxies, and any other HTTP services.
|
UI proxies, and any other HTTP services.
|
||||||
|
|
||||||
The **host** firewall is the only firewall. The shared-netns infra
|
The **host** firewall is the only firewall. The swarm service
|
||||||
containers (gateway, forge, matrix) set
|
containers that share the host netns (forge, matrix, authelia, bao, the
|
||||||
`networking.firewall.enable = false`: a NixOS firewall inside a
|
metrics and log stores) run with `networking.firewall.enable = false`
|
||||||
|
(`nix/container-modules/swarm-container.nix`): a NixOS firewall inside a
|
||||||
shared-netns container runs against the _host_ ruleset — at container
|
shared-netns container runs against the _host_ ruleset — at container
|
||||||
boot its `firewall-start` flushes the `nixos-fw` chains, rebuilds them
|
boot its `firewall-start` flushes the `nixos-fw` chains, rebuilds them
|
||||||
from the container's (empty) port list, and deletes the host's
|
from the container's (empty) port list, and deletes the host's
|
||||||
|
|
@ -228,15 +235,16 @@ address arithmetic.
|
||||||
|
|
||||||
`hive-c0re` reads `HIVE_NETWORK_BRIDGE` + `HIVE_NETWORK_SUBNET` and passes
|
`hive-c0re` reads `HIVE_NETWORK_BRIDGE` + `HIVE_NETWORK_SUBNET` and passes
|
||||||
`PRIVATE_NETWORK=1`, `LOCAL_ADDRESS=` (empty), `HOST_ADDRESS=<bridge-ip>`,
|
`PRIVATE_NETWORK=1`, `LOCAL_ADDRESS=` (empty), `HOST_ADDRESS=<bridge-ip>`,
|
||||||
and `HOST_BRIDGE=<bridgeName>` via `lifecycle::set_nspawn_flags` when
|
and `HOST_BRIDGE=<bridgeName>` when creating or updating containers:
|
||||||
creating or updating containers. hive-c0re validates both variables **once at
|
`lifecycle::set_nspawn_flags` hands them to hive-priv, which rewrites
|
||||||
|
`/etc/nixos-containers/<container>.conf`. hive-c0re validates both variables **once at
|
||||||
daemon startup**, not per container: they're process-global, so a
|
daemon startup**, not per container: they're process-global, so a
|
||||||
missing or malformed value is a misconfigured daemon rather than one bad
|
missing or malformed value is a misconfigured daemon rather than one bad
|
||||||
container, and failing at boot gives a single diagnostic instead of one
|
container, and failing at boot gives a single diagnostic instead of one
|
||||||
per agent. No non-isolated mode exists to fall back to. hive-c0re leaves `LOCAL_ADDRESS` empty so the
|
per agent. No non-isolated mode exists to fall back to. hive-c0re leaves `LOCAL_ADDRESS` empty so the
|
||||||
container's dhcpcd acquires an address from the bridge dnsmasq pool
|
container's dhcpcd acquires an address from the bridge dnsmasq pool
|
||||||
(`networking.useDHCP = true` in `nix/agent-modules/network.nix`). This applies uniformly
|
(`networking.useDHCP = true` in `nix/agent-modules/network.nix`). This applies to
|
||||||
to all containers — agents and service containers alike.
|
every agent container, manager included.
|
||||||
|
|
||||||
`HOST_ADDRESS` is the bridge gateway IP (the address part of
|
`HOST_ADDRESS` is the bridge gateway IP (the address part of
|
||||||
`HIVE_NETWORK_SUBNET`, via `lifecycle::bridge_gateway_ip` — taken verbatim
|
`HIVE_NETWORK_SUBNET`, via `lifecycle::bridge_gateway_ip` — taken verbatim
|
||||||
|
|
|
||||||
|
|
@ -1,11 +1,13 @@
|
||||||
# The job queue, for operators
|
# The job queue, for operators
|
||||||
|
|
||||||
Every container operation — rebuild, first-spawn, a config-PR deploy,
|
Long-running work runs through a job graph. The swarm controller keeps one
|
||||||
power changes — runs through one shared job queue. This page explains
|
for swarm-level work — creating an agent's identity, forge user and config
|
||||||
what the job queue _is_, as a general idea, independent of what any one
|
repo. Each hive's hive-c0re keeps its own for container operations —
|
||||||
subsystem uses it for. For the hive-c0re-specific step catalogue and the
|
rebuild, first-spawn, a config-PR deploy, power changes. This page explains
|
||||||
engineering internals (scheduler, leases, resource windows) see
|
what the job queue _is_, as a general idea, independent of what either uses
|
||||||
[`coordinator.md`](coordinator.md) instead.
|
it for. For the hive-c0re step catalogue and the engineering internals
|
||||||
|
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)
|
||||||
|
instead.
|
||||||
|
|
||||||
## What the job queue is, in the abstract
|
## What the job queue is, in the abstract
|
||||||
|
|
||||||
|
|
@ -24,18 +26,18 @@ Two ideas are all there is to it:
|
||||||
|
|
||||||
The engine's whole job is: whenever a step's ordering and resource needs
|
The engine's whole job is: whenever a step's ordering and resource needs
|
||||||
are both satisfied, run it. It has no opinion on what the steps _do_ —
|
are both satisfied, run it. It has no opinion on what the steps _do_ —
|
||||||
that's supplied by whoever builds the graph. hive-c0re is the one thing
|
that's supplied by whoever builds the graph. The swarm controller and
|
||||||
building graphs on it today, but nothing about the engine is specific to
|
hive-c0re each build their own graph on it, and nothing about the engine
|
||||||
containers or rebuilds; there's nothing stopping another subsystem from
|
is specific to either.
|
||||||
using the same engine for its own unrelated queue.
|
|
||||||
|
|
||||||
## Watching it happen
|
## Watching it happen
|
||||||
|
|
||||||
Each **row** you see in a queue view (the **BU1LDS** page's R3BU1LD QU3U3
|
Two views, one per graph, drawn by the same component: the swarm UI's
|
||||||
— see [`web-ui/dashboard.md`](../web-ui/dashboard.md) — and swarm-ui's
|
`/jobs` page shows the swarm controller's graph, and the hive dashboard's
|
||||||
`/jobs` page both render the same underlying graph) is one job; the rows
|
**BU1LDS** page (R3BU1LD QU3U3 — see
|
||||||
nested under it are that job's steps, in order (occasionally a couple run
|
[`web-ui/dashboard.md`](../web-ui/dashboard.md)) shows that hive's. Each
|
||||||
side by side). A step shows one of:
|
**row** is one job; the rows nested under it are that job's steps, in
|
||||||
|
order (occasionally a couple run side by side). A step shows one of:
|
||||||
|
|
||||||
<!-- vale write-good.Passive = NO -->
|
<!-- vale write-good.Passive = NO -->
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,48 +1,74 @@
|
||||||
# Observability (OpenTelemetry)
|
# Observability (OpenTelemetry)
|
||||||
|
|
||||||
hyperhive has built-in support for exporting each agent container's telemetry to
|
The swarm has one telemetry pipeline. The swarm's collector writes the
|
||||||
any OTLP-compatible collector: Claude Code statistics — token usage, cost, tool
|
swarm's metrics and log stores and, if you name one, exports upstream; each
|
||||||
call counts — via Claude Code's built-in OpenTelemetry integration, and the
|
hive runs a collector of its own that forwards its agents' telemetry there.
|
||||||
container's journal, forwarded by a collector running inside it.
|
What flows through it: each agent's Claude Code statistics (token usage,
|
||||||
|
cost, tool call counts) via Claude Code's built-in OpenTelemetry
|
||||||
|
integration, each agent container's journal, and the hyperhive metrics
|
||||||
|
catalogued below.
|
||||||
|
|
||||||
This is a **hive-wide** setting: one switch in the host NixOS config enables it
|
## The two collectors
|
||||||
for every agent container simultaneously. No per-agent opt-in or opt-out exists.
|
|
||||||
|
|
||||||
## Enabling export
|
Telemetry crosses two collectors, and which one you configure depends on what
|
||||||
|
the host is:
|
||||||
|
|
||||||
|
| | runs where | receives from | does |
|
||||||
|
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
|
||||||
|
| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's stores and exports upstream |
|
||||||
|
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
|
||||||
|
|
||||||
|
An all-local host runs both, and needs nothing said about the hop between them.
|
||||||
|
|
||||||
```nix
|
```nix
|
||||||
services.hyperhive.otel = {
|
services.hyperhive.otel = {
|
||||||
enable = true;
|
enable = true;
|
||||||
endpoint = "https://collector.example.com/otel";
|
endpoint = "https://collector.example.com/otel"; # the upstream
|
||||||
|
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
|
||||||
};
|
};
|
||||||
```
|
```
|
||||||
|
|
||||||
`enable` is the single gate. `endpoint` is where telemetry ends up after it
|
## Enabling export
|
||||||
leaves the swarm — optional, because the swarm's own metrics store
|
|
||||||
(`deploy.victoriametrics`) is a destination in its own right. With both,
|
`otel.enable` is the single gate on a hive: one switch in the host config
|
||||||
telemetry goes to both. See
|
covers every agent container on it, with no per-agent opt-in or opt-out.
|
||||||
|
`endpoint` is where telemetry ends up after it leaves the swarm — optional,
|
||||||
|
because the swarm's own metrics store (`deploy.victoriametrics`) is a
|
||||||
|
destination in its own right. With both, telemetry goes to both. See
|
||||||
[`swarm/services.md`](../swarm/services.md#metrics-victoriametrics--grafana).
|
[`swarm/services.md`](../swarm/services.md#metrics-victoriametrics--grafana).
|
||||||
|
|
||||||
**Telemetry leaves a hive exactly one way: through the collector that
|
**Telemetry leaves a hive exactly one way: through the collector that
|
||||||
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
|
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
|
||||||
they export unauthenticated to a bridge address only their own containers can
|
they export unauthenticated to a bridge address only their own containers can
|
||||||
reach. That collector forwards to the swarm's
|
reach (`http://<bridgeIp>:<collector port>`; the otel module opens that port
|
||||||
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), which is
|
on the bridge itself). That collector forwards to the swarm's
|
||||||
the single process holding the upstream credential and the only writer to the
|
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), the
|
||||||
swarm's store. No agent holds a copy, and neither does this hive.
|
single process holding the upstream credential and the only writer to the
|
||||||
|
swarm's stores. No agent holds a copy, and neither does the hive.
|
||||||
|
|
||||||
The hive collector reaches the swarm collector by its gateway name
|
The hive collector reaches the swarm collector by its gateway name
|
||||||
(`swarm.otel.domain`, default `otel.<swarm domain>`) — the same DNS-and-CA-trust
|
(`swarm.otel.domain`, default `otel.<swarm domain>`) — the same DNS-and-CA-trust
|
||||||
shape every hive-to-swarm-service hop uses, not a URL an operator has to point
|
shape every hive-to-swarm-service hop uses. On the host running the swarm
|
||||||
anywhere. A hive that doesn't run the swarm's services still resolves that
|
collector the hive's dnsmasq answers that name; elsewhere it resolves through
|
||||||
name through the gateway; nothing here needs setting for the split-host case.
|
ordinary DNS. Nothing here needs setting for the split-host case.
|
||||||
|
|
||||||
⚠️ **The collector is therefore in the path of all telemetry.** It runs on the
|
⚠️ **The hive collector carries every bit of that hive's telemetry.** It
|
||||||
same host as the agents and restarts on failure, and telemetry isn't the
|
runs on the same host as the agents and restarts on failure. Telemetry isn't
|
||||||
control plane — degraded telemetry isn't degraded operation — but the export
|
the control plane, so degraded telemetry isn't degraded operation.
|
||||||
no longer survives independently of anything host-side.
|
|
||||||
|
|
||||||
### what the agent→collector hop is and isn't
|
### Why two tiers
|
||||||
|
|
||||||
|
**The hive tier isn't optional.** Exporting straight to `endpoint` would mean
|
||||||
|
every agent needs the credential — and the harness delivers that token into
|
||||||
|
the agent's own `~/.claude/settings.json`, a file the agent can read. `0600`
|
||||||
|
protects it from other containers, not from the agent itself. An option that
|
||||||
|
could select the direct path would reopen that hole.
|
||||||
|
|
||||||
|
**The tiers stay separate on one box.** An all-local hive is a statement about
|
||||||
|
_where_ processes run, not about the shape of the deployment. A boundary that
|
||||||
|
disappears locally is one the local deployment stops testing.
|
||||||
|
|
||||||
|
### What the agent→collector hop is and isn't
|
||||||
|
|
||||||
**It has no application-level auth.** The receiver takes any OTLP that reaches
|
**It has no application-level auth.** The receiver takes any OTLP that reaches
|
||||||
it; what bounds who can reach it's the firewall — `exposeHostPorts` opens the
|
it; what bounds who can reach it's the firewall — `exposeHostPorts` opens the
|
||||||
|
|
@ -55,10 +81,9 @@ credential.** Neither tier can tell a container's genuine Claude Code stats
|
||||||
from anything else shaped like OTLP arriving on that port — including data
|
from anything else shaped like OTLP arriving on that port — including data
|
||||||
smuggled out in resource attributes on an otherwise-legitimate export.
|
smuggled out in resource attributes on an otherwise-legitimate export.
|
||||||
|
|
||||||
That's a **different risk from the one the collector fixes**, and strictly
|
That's a **different risk from the one the collector fixes**: an agent that
|
||||||
smaller than what preceded it: before, every agent held the upstream credential
|
held the upstream credential could do everything above _and_ use the token
|
||||||
itself, so it could do all of the above _and_ use the token anywhere else. The
|
anywhere else. The collector removes the token and keeps the pipe. Agents are inside the trust
|
||||||
collector removes the token and keeps the pipe. Agents are inside the trust
|
|
||||||
boundary (`docs/trust-boundary/security.md`: capability = accepted risk), so an agent being
|
boundary (`docs/trust-boundary/security.md`: capability = accepted risk), so an agent being
|
||||||
able to _send_ is an accepted extension of that boundary — but it's not
|
able to _send_ is an accepted extension of that boundary — but it's not
|
||||||
closed by this design, and don't read anything here as closing it.
|
closed by this design, and don't read anything here as closing it.
|
||||||
|
|
@ -74,7 +99,7 @@ hive's collector can label its data as any other agent.
|
||||||
agent container forwards its own journal through this port — every unit in it at
|
agent container forwards its own journal through this port — every unit in it at
|
||||||
`info` and above, not an allowlist. That's the harness, the MCP daemons and
|
`info` and above, not an allowlist. That's the harness, the MCP daemons and
|
||||||
whatever a tool call spawned, so command lines and error text now leave the
|
whatever a tool call spawned, so command lines and error text now leave the
|
||||||
container where before only counts did. The trust boundary is unchanged (same
|
container, not just counts. The trust boundary is unchanged (same
|
||||||
destination, same credential, and an agent could already send arbitrary OTLP);
|
destination, same credential, and an agent could already send arbitrary OTLP);
|
||||||
what changes is how much detail leaves by default.
|
what changes is how much detail leaves by default.
|
||||||
|
|
||||||
|
|
@ -98,54 +123,6 @@ rather than the claims it authenticated with.
|
||||||
If you need per-agent numbers you can act on, take them from the agent's own
|
If you need per-agent numbers you can act on, take them from the agent's own
|
||||||
turn-stats rather than from a metric label.
|
turn-stats rather than from a metric label.
|
||||||
|
|
||||||
## Options reference
|
|
||||||
|
|
||||||
The nix module (`nix/host-modules/otel.nix`) generates every
|
|
||||||
`services.hyperhive.otel.*` option's full type/default/description/
|
|
||||||
example straight into [`/options/`](/options/) (host options — `nix build
|
|
||||||
.#docs-host` for a local render). The build keeps that page honest in a
|
|
||||||
way a hand-copied version here can't be, so it's the reference, not this
|
|
||||||
doc. What follows is what a flat per-option listing can't express: the
|
|
||||||
two-tier architecture, the security model, and how the options interact.
|
|
||||||
|
|
||||||
## The two collectors
|
|
||||||
|
|
||||||
Telemetry crosses two collectors, and which one you configure depends on what
|
|
||||||
the host is:
|
|
||||||
|
|
||||||
| | runs where | receives from | does |
|
|
||||||
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
|
|
||||||
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
|
|
||||||
| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's store and exports upstream |
|
|
||||||
|
|
||||||
An all-local host runs both, and needs nothing said about the hop between them.
|
|
||||||
|
|
||||||
```nix
|
|
||||||
services.hyperhive.otel = {
|
|
||||||
enable = true;
|
|
||||||
endpoint = "https://collector.example.com/otel"; # the upstream
|
|
||||||
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
**Why the hive tier isn't optional.** Exporting straight to `endpoint` means
|
|
||||||
every agent needs the credential to authenticate — and the harness delivers
|
|
||||||
that token into the agent's own `~/.claude/settings.json`, a file the agent can
|
|
||||||
read. `0600` protects it from other containers, not from the agent itself. As
|
|
||||||
long as the direct path stays _selectable_, that hole stays selectable; an
|
|
||||||
option that can reintroduce it's a hole with extra steps.
|
|
||||||
|
|
||||||
**Why the tiers stay separate on one box.** They're not collapsed when
|
|
||||||
co-located: an all-local hive is a statement about _where_ processes run, not
|
|
||||||
about the shape of the deployment. A boundary that disappears locally is one
|
|
||||||
the local deployment stops testing.
|
|
||||||
|
|
||||||
**`endpoint` keeps meaning "where telemetry goes upstream."** Neither tier
|
|
||||||
redefines it — the agent-facing value is _derived_
|
|
||||||
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
|
|
||||||
keeps working unchanged. The otel module contributes the bridge port to `exposeHostPorts`
|
|
||||||
automatically; there is nothing to open by hand.
|
|
||||||
|
|
||||||
### Authenticated ingest
|
### Authenticated ingest
|
||||||
|
|
||||||
The swarm tier gives **each hive its own receiver**, and stamps the `hive` label
|
The swarm tier gives **each hive its own receiver**, and stamps the `hive` label
|
||||||
|
|
@ -180,26 +157,19 @@ a build error naming the reason rather than telemetry silently going nowhere.
|
||||||
|
|
||||||
## Network access
|
## Network access
|
||||||
|
|
||||||
Agent containers can only reach the host on ports 80 and 443 by default. To let
|
Agent containers reach the host only on ports 80 and 443 (plus DNS). The
|
||||||
them reach some other host-local service you run yourself — a database, a
|
collector port needs nothing from you: `otel.enable` opens it on the bridge.
|
||||||
scratch HTTP endpoint — open its port on the bridge:
|
To let agents reach some other host-local service of your own →
|
||||||
|
[`networking/network.md`](../networking/network.md#reaching-host-services-exposehostports).
|
||||||
|
|
||||||
```nix
|
## Options reference
|
||||||
services.hyperhive.network.exposeHostPorts = [ 5432 ];
|
|
||||||
```
|
|
||||||
|
|
||||||
and point whatever consumes it at `10.42.0.1:5432` rather than loopback: inside
|
The nix module (`nix/host-modules/otel.nix`) generates every
|
||||||
a container, loopback is the _container_. The bridge IP is the host's address on
|
`services.hyperhive.otel.*` option's full type, default, description and
|
||||||
the `hive-br0` bridge. The service must also bind an address the bridge can
|
example into the [options reference](https://hyperhive.darkest.space/options/)
|
||||||
reach — a `127.0.0.1`-only listener stays unreachable no matter what the
|
(host options — `nix build .#docs-host` for a local render). That's the
|
||||||
firewall allows. See `docs/networking/network.md::Reaching host services` for details.
|
reference; this page covers what a per-option listing can't: the two-tier
|
||||||
|
architecture, the security model, and how the options interact.
|
||||||
<!-- vale write-good.Passive = NO -->
|
|
||||||
|
|
||||||
⚠️ **None of this is needed for hyperhive's own telemetry** — `otel.enable`
|
|
||||||
contributes the collector's port and derives the agent-facing endpoint itself.
|
|
||||||
|
|
||||||
<!-- vale write-good.Passive = YES -->
|
|
||||||
|
|
||||||
## Built-in resource labels
|
## Built-in resource labels
|
||||||
|
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue