argus on #775 v3: "the `gatewayHost` option description's
server_name-vs-gatewayHost essay + federation SRV note are also
candidates for [docs/gateway.md] section."
Cuts the gatewayHost option's description from ~40 lines (with
inline duplication of the discovery flow, when-to-set-which, and
federation port caveat) down to ~8 lines pointing at
`docs/gateway.md`. The brief `server_name vs gatewayHost` clarifier
stays in code because it disambiguates two SIMILAR-LOOKING options
on the same module — operators reading option docs need the
distinction inline, not behind a doc link.
Also trimmed `matrix.gui.enable` + `matrix.gui.package` descriptions
to similar shapes — point at docs/gateway.md for the architecture,
keep the override-shape hints in code.
Push includes the rebase onto current main (#764 + 0af6ea1 + others
landed since #775 was opened; cherry-picked commits get skipped
cleanly).
Net: matrix.nix loses ~70 lines of inline prose. No behavioral
change (verified gatewayHost still resolves to `matrix.<hive>`).
mara on PR #775: "this is too much docs in code - move bigger picture
stuff to md files and put refs in code"
New `docs/gateway.md` consolidates the gateway architecture story
that was spreading across long inline comments in `hive-gateway.nix`,
`hive-matrix.nix`, and `hive-forge.nix`:
- vhost map (which URL serves what, which upstream, which option)
- matrix discovery flow (.well-known → sub-domain delegation
sequence)
- Accept-header SPA fallback pattern (#686 / #729 design history)
- local-dev `localHostsEntry` story
- sub-domain rationale (mara verdict tracking) + when sub-path is
right (hyperhive-internal apps)
- per-vhost tuning knobs (forge LFS, matrix long-poll, agent SSE)
- sequencing history (which PR added which routing piece)
In-code comments in the two nix modules get trimmed to short refs
into the doc — keeps the *why* in the markdown while the *what*
stays alongside the code:
- hive-gateway.nix: top-of-file comment, `agentPortsTable`,
`appendHttpConfig`, every location block + vhost
- hive-matrix.nix: `fluffychat-web-fixed`, `fluffychat-web-imaging`,
the dart compile postInstall
README.md gets a new row in the docs table pointing at gateway.md.
Verified `nix eval` still resolves the same vhost + location layout
after the comment trim — no behavioral change, just less in-code
prose.
mara on #764:9897: "host the fluffy chat app at / as follow up?"
Moves fluffychat-web from the bare-domain sub-path
(`<hive>/matrix/`) to the matrix sub-domain root
(`matrix.<hive>/`). Follow-up to #764 (matrix vhost itself), per
mara's gateway-architecture verdict (sub-domain for external standard
apps, sub-path for hyperhive-internal). Stacked on
`atlas/747-matrix-behind-gateway` — depends on #764 landing first.
## Mechanics
**hive-matrix.nix:**
- Drop `flutterBuildFlags = [ "--base-href" "/matrix/" ]` from
`fluffychat-web-fixed`. Upstream default `--base-href "/"` is correct
at sub-domain root.
- Update option docs to reflect new mount point.
**hive-gateway.nix:**
- `$matrix_spa_target` map target flips from `/matrix/index.html` →
`/index.html` (sub-domain root now).
- New `<hive>/matrix/*` location: `rewrite ^/matrix/(.*)$
matrix.<hive>/$1 permanent;` — 301 redirect preserves bookmark +
deep-link compatibility for `<hive>/matrix/#/rooms/...` URLs during
the transition.
- `<hive>/matrix/config.json` location removed (moved to `/config.json`
on the matrix vhost).
- Matrix vhost (#764) gains `/` location: serves fluffychat dist as
static files with the Accept-header SPA fallback (`/_matrix/`
proxying to tuwunel keeps working via nginx longer-prefix-wins
precedence). When `gui.enable = false`, `/` returns 404 cleanly.
- Matrix vhost gains `= /config.json` for the FluffyChat boot-config
pre-fill (#736).
## Verified
```
vhosts: ["_", "forge.test.local", "matrix.test.local"]
bare locations: ["/", "/matrix/", "= /.well-known/matrix/client",
"= /.well-known/matrix/server"]
matrix vhost locations: ["/", "/_matrix/", "= /config.json"]
/matrix/ extraConfig: "rewrite ^/matrix/(.*)$ http://matrix.test.local/$1 permanent;"
matrix vhost / alias: /nix/store/...fluffychat-web-2.6.0/
```
Full container toplevel builds clean.
## Risk
Medium. Two breaking changes for operators:
1. **Bookmark migration**: `http://<hive>/matrix/#/rooms/...` 301s
to `http://matrix.<hive>/#/rooms/...`. Browser bookmarks +
shared links keep working via the redirect; can be cleaned up
once it's been in the wild long enough.
2. **fluffychat-web dist hash changes**: dropping the
`--base-href "/matrix/"` flag changes the derivation hash, so
`gui.package` rebuilds even though the source is the same.
Operators on substitute caches will fetch the new dist; building
from source takes the same time as before.
The `.well-known/matrix/{client,server}` delegation (already
advertising `matrix.<hive>` per #764) means matrix clients
auto-discover the new location — no client config change needed.
## Sequencing
**Depends on #764** — needs the matrix vhost to host the new `/`
location. Merge after #764 lands + soaks. If #764 changes shape
during review I'll rebase + force-push.
Closes#772.
Folds both 🟡 notes from argus's #764 review:
1. **Empty-string assertion on `cfg.gatewayHost`**: same footgun as
the forge.domain rejection from #754 — empty would render `.<hive>`
shaped garbage in both nginx server_name (wildcard catch-all,
surprising) and /etc/hosts (invalid entry). Fail loud at toplevel
build with a message pointing at `null` as the right opt-out.
2. **Federation port-8448 caveat in `gatewayHost` docs**: when the
gateway listens on 80, `.well-known/matrix/server` advertises
`${gatewayHost}` with no port suffix → matrix federation spec
falls back to port 8448 → no listener on 8448 → cross-hive
federation requires either `_matrix._tcp.${gatewayHost}` SRV
record OR `services.hyperhive.matrix.openFirewall = true`.
Hyperhive is mostly closed/internal so this rarely bites, but
the option docs now flag it for the federation-curious operator.
Verified: `gatewayHost = ""` triggers the new assertion at toplevel
build with the expected message; default still resolves to
`matrix.<hive>` cleanly.
mara on #747:9722: "this still seems to be an issue in current version"
(after #751 closed without merge). Mirroring the forge sub-domain
pattern just merged as #754 for matrix per mara's #749:9609 verdict
(sub-domain over sub-path for forge + matrix, "not user-visible for
matrix because the .well-known/matrix/{client,server} redirect routes
clients through automatically").
## Mechanics
**New `services.hyperhive.matrix.gatewayHost`** — nullable str, defaults
to `matrix.<services.hyperhive.domain>` when hive-domain set, else
null. Full hostname (`matrix.darkest.space`, `homeserver.internal.lan`)
for bespoke shapes per mara's #754:9684 "specify full domain in
options instead" pattern.
**Gateway:** new `server { server_name = matrixCfg.gatewayHost; }`
block proxying `/_matrix/...` → `http://127.0.0.1:<httpPort>/_matrix/...`
with matrix-spec CORS + tuned for long-poll `/sync` (1h timeout) +
typical media uploads (50M body cap). `/` returns 404 — nothing
else lives at the matrix vhost. Matches the forge vhost shape from #754.
**`.well-known/matrix/{client,server}`** (already served at bare hive-
domain since #660): now points at `matrixCfg.gatewayHost` (no port
suffix when gateway is on the canonical port 80) instead of the
direct `<hive-domain>:<httpPort>` shape. Falls back to direct shape
when `gatewayHost = null` (no hive-domain, or operator nulled it).
**`localHostsEntry` extension**: `/etc/hosts` (when set) now adds the
matrix sub-domain → 127.0.0.1 alongside hive-domain + forge.domain.
`lib.unique` collapses any duplicate (edge case if operator sets
gatewayHost equal to hive-domain).
## Verified via `nix eval`
```
vhosts: ["_", "forge.test.local", "matrix.test.local"]
gatewayHost: "matrix.test.local"
client wellknown: m.homeserver.base_url = "http://matrix.test.local"
server wellknown: m.server = "matrix.test.local"
/etc/hosts: ["test.local", "forge.test.local", "matrix.test.local"]
```
## What this fixes for #747
mara's HAR showed `GET /.well-known/matrix/client` and
`GET /_matrix/client/versions` both failing on `pr1ma.darkest.space`:
1. **`.well-known/matrix/client`** was advertising
`http://pr1ma.darkest.space:8008` — that URL only works if tuwunel's
port 8008 is firewall-open to the operator's browser (it isn't by
default — `services.hyperhive.matrix.openFirewall` defaults to false
since #651). Now advertises `http://matrix.pr1ma.darkest.space/`
which goes through the gateway on the (already-open) port 80.
2. **`/_matrix/client/versions`** was hitting the bare-domain `"_"`
vhost, which has no `/_matrix/` location — fell through to `/` →
c0re's dashboard upstream → 404. Now hits the new `matrix.<hive>`
vhost which proxies the request to tuwunel cleanly.
server_name + serverName unaffected — matrix identifiers (`@alice:<hive>`)
still embed the bare hive-domain per #660; only the wire-level transport
URL moves to the sub-domain.
## Risk
Medium. Existing matrix tokens / sessions stay valid because:
- `serverName` (the identifier domain) doesn't change
- tuwunel's `/_matrix/` endpoints serve the same requests, just reached
via the new sub-domain instead of the direct port
Operators with `services.hyperhive.matrix.openFirewall = true` and
external clients reaching `:8008` directly keep working too — the
sub-domain vhost is additive, doesn't take away the direct port.
## Sequencing
This is a parallel matrix-side mirror of #754 (forge). Both follow
the same mara-verdict pattern; once both have soaked, the gateway-
behind-everything story is done for v0.
Closes#747.
mara on PR #754: "would it be better to specify full forge domain in
options instead?"
Drops the awkward `cfg.subdomain` label option. Now `cfg.domain` is
the single source of truth for both the forgejo `DOMAIN` setting
(existing semantics) AND the gateway vhost server-name (new).
## Before / after
```nix
# before: separate label + cfg.domain juggling
services.hyperhive.forge.subdomain = "forge"; # → forge.<hive>
services.hyperhive.forge.domain = "localhost"; # unused for vhost
# after: full domain, single option
services.hyperhive.forge.domain = "forge.darkest.space"; # ← used for ROOT_URL + vhost
```
## Default
`cfg.domain` default auto-derives:
- `forge.<services.hyperhive.domain>` when hive-domain is set
- `"localhost"` otherwise (pre-#749 direct-on-port shape)
So the common case (hive-domain set) gets `forge.<hive>` for free,
operators with a bespoke shape (`git.example.com`) set the full
hostname directly.
## Assertions
- `cfg.domain != ""` — empty would render `.<hive>` shaped garbage
in both server_name + /etc/hosts.
- `cfg.behindGateway → gateway.enable` — can't route through a
gateway that isn't running.
(The previous "subdomain = empty" assertion is dropped — that
edge case is gone with the rename.)
## Verified
- default with `hyperhive.domain = "test.local"` → `forge.test.local`,
`ROOT_URL = http://forge.test.local/`, vhost present
- `forge.domain = "git.example.com"` → `git.example.com`,
`ROOT_URL = http://git.example.com/`, vhost = `["_", "git.example.com"]`
- `gateway.enable = false` → `forge.domain` falls back to `localhost`,
`ROOT_URL = http://localhost:3000/`, no gateway vhost
(`behindGateway = false`)
- `/etc/hosts` (when `localHostsEntry = true`) → unique entries for
hive-domain + forge.domain (de-duped via `lib.unique` for the
edge case where forge.domain = hive-domain)
- full container toplevel builds clean
## PR title
(Will fix the PR title separately — still says "/forge/" which is
wrong since the rewrite to sub-domain shape.)
argus on PR #754 v2 review:
> `subdomain = ""` edge case: when `cfg.subdomain = ""`, the
> `localHostsEntry` appends ".${domain}" (invalid hostname; bare
> domain is already covered) and the virtualHosts key becomes
> ".${domain}" (nginx treats this as a wildcard catch-all, not a
> bare-domain server block). docs call this "advanced: collides
> with dashboard server block" — the actual nginx behavior is
> more surprising than that.
Fix: reject `""` at assertion time rather than ship the surprising
behaviour. Bare-domain landing is what the dashboard already
serves; there's no use case for `""` that null doesn't already
cover. Updated option description + dropped the now-dead branch
from the `subdomain` let-binding.
Verified: `services.hyperhive.forge.subdomain = ""` triggers the
new assertion at toplevel build with a clear message pointing at
`null` as the right opt-out. Default + `null` paths still build
clean.
mara on #749:9609: "we will go with sub domains for forge and matrix
(redirected in well known in the latter case, not user visible). close /
fix PRs you have open that dont match this."
Reshapes the v1 sub-path (`<host>/forge/`) approach into a sub-domain
vhost (`forge.<host>/`) per the mara verdict. matrix gets the same
treatment in damocles's #751 follow-up.
## Why sub-domain
- forgejo's default `ROOT_URL = http://<host>/` works without any
`X-Forwarded-Prefix` gymnastics — sub-domain hosting is the
canonical Forgejo deploy shape, matches every upstream-doc example.
- Cookie / storage isolation between the dashboard and forge (XSS blast
radius shrinks; a future forge XSS can't reach dashboard session).
- matches the matrix-spec pattern that #751 wires up for the
homeserver.
## Mechanics
**forge options:**
- `services.hyperhive.forge.subdomain` — nullable str, default `"forge"`
→ rendered sub-domain is `forge.<hive-domain>`. Set to `null` to opt
out (forge stays direct on `httpPort`); set to `""` for bare-domain
landing (advanced, collides with dashboard).
- `services.hyperhive.forge.rootUrl` — nullable str override. When
null, auto-derived: `http://<subdomain>.<hive>/` when gateway is on
+ subdomain set, else `http://<domain>:<httpPort>/` (direct).
- **Asserts** rootUrl ends with `/` (argus 🟡 on #754: forgejo's
ROOT_URL contract requires trailing slash, else emits
`https://forge.example.com.user.id` shaped garbage). Asserts
`subdomain != null` requires `hyperhive.domain` set.
**gateway:**
- New `virtualHosts."<subdomain>.<hive-domain>"` server block —
separate from the `"_"` catch-all. Proxies all `/` →
`http://127.0.0.1:<forge.httpPort>/` so forgejo handles requests at
root (no prefix translation needed; matches the upstream-default
ROOT_URL shape).
- Git-tuned: `client_max_body_size 1G`, `proxy_read_timeout 1h`,
`proxy_send_timeout 1h`, `proxy_buffering off`,
`proxyWebsockets = true`. SSH stays direct on `cfg.sshPort`.
- `networking.hosts` (when `localHostsEntry = true`) now also adds
`forge.<hive-domain> -> 127.0.0.1` for the dev loop.
## Verified
- `nix eval ROOT_URL` → `http://forge.test.local/` (default with
gateway on)
- `nix eval ROOT_URL` with `gateway.enable = false` → `http://localhost:3000/`
(current direct shape preserved)
- `nix eval virtualHosts attrs` → `["_", "forge.test.local"]`
- `nix eval networking.hosts` with `localHostsEntry = true` →
`{"127.0.0.1": ["test.local", "forge.test.local"], ...}`
- bad rootUrl (no trailing /) triggers assertion at toplevel build
with the spelled-out forgejo failure mode
- full container toplevel builds clean
(`nixos-system-hive-gateway-26.05pre-git`)
## Migration
ROOT_URL change is a one-way migration on rebuild:
- Existing agent `git remote origin` URLs (`http://localhost:3000/...`)
**keep working** — forgejo accepts any inbound URL; the URL on the
agent side is unchanged.
- New clone-link copy-paste from forge UI uses `forge.<hive>/...` —
operators copying clones after this lands need to go through the
new sub-domain.
- Direct browsing on `:3000` shows pages with `forge.<hive>` links →
works if hosts entry / DNS resolves, broken otherwise. Operators
should switch to `http://forge.<hive>/`.
## Out of scope
- TLS termination (mara explicit on #15: no TLS v0)
- SSH-over-HTTPS / wildcard cert provisioning
- matrix sub-domain (damocles's #751, sibling work)
Closes#749. Addresses argus 🟡 on #754.
mara on PR #740 comment 9295: "we decided to go with the json" (issue #15 comment 9270:
"nginx container lives in system config, so it cannot be just rebuilt
from meta flake. go for the json file the c0re writes").
Drops:
- `cfg.agents` listOf str option
- Replicated FNV-1a hash + char-code table + manager-port special case
- Drift-hazard comment (no more rust↔nix constant sync)
Adds:
- `cfg.agentPortsFile = "/var/lib/hyperhive/agent-ports.json"` (default,
nullable to disable) — path to a JSON map of `{ "<name>": <port> }`
written by hive-c0re on every topology change.
- `agentPortsTable` reads the file at eval time via
`builtins.fromJSON (builtins.readFile path)`, guarded by
`builtins.pathExists` so a missing file gracefully defaults to `{}`.
- Per-agent locations generated via `lib.mapAttrs'` over the table —
one location block per entry; empty table → empty attrset → no
per-agent blocks, pre-#15 shape.
Rust-side dependency: hive-c0re needs to emit the JSON file on every
topology change. Coordinating with damocles via a separate ping — the
nix side ships now with safe defaults (missing file = no routes, no
behavior change vs main).
Verified:
- nix eval with `/tmp/test-agent-ports.json` → 4 per-agent blocks at
correct ports (8178 iris, 8267 argus, 8304 atlas, 8549 damocles)
- nix eval with nonexistent file → only `/` location (graceful default)
- full container toplevel builds clean with matrix on
Empty file case mirrors the previous empty-list default — purely
additive, old `<host>:<port>/` direct reach untouched, no per-agent
blocks until c0re writes the JSON. Operator can also `null` the
option to disable entirely.
Per mara on #14 (comment 9081): focused, purely additive to what's
there, no TLS / no manager special cases, old `<host>:<port>/` path
keeps working. Builds on iris's #731 (agent UI now serves
document-relative URLs so it works under any nginx prefix).
Mechanics:
- New `services.hyperhive.gateway.agents` option (`listOf str`,
default `[]`) lists sub-agent names to expose at
`/agent/<name>/` through the gateway.
- For each name, generate one `location /agent/<name>/` block that
`proxy_pass`es to `http://127.0.0.1:<port>/`, where `<port>`
is computed from the same FNV-1a hash hive-c0re uses internally
(`lifecycle::agent_web_port`).
- Trailing-slash pair on location + proxy_pass strips the
`/agent/<name>` prefix on the upstream side — agent server
receives `GET /`, `GET /api/state`, `GET /screen/ws`, etc. as if
reached directly on its port.
- `X-Forwarded-Prefix` set so the harness can build correct absolute
URLs for cases where document-relative isn't enough.
- `proxyWebsockets = true` + `proxy_buffering off` keeps SSE
+ WS endpoints working transparently.
- Empty `cfg.agents` (default) → no per-agent blocks generated.
- Manager not included — already gets `/` via the c0re upstream.
FNV-1a hash replicated in nix to match `lifecycle::agent_web_port`
line-for-line. Verified against rust output for 8 representative
agent names:
agent | nix | rust | match
iris | 8178 | 8178 | ✓
atlas | 8304 | 8304 | ✓
argus | 8267 | 8267 | ✓
damocles | 8549 | 8549 | ✓
manager | 8000 | 8000 | ✓ (special case)
dmatrix | 8266 | 8266 | ✓
triage | 8737 | 8737 | ✓
bitburner | 8658 | 8658 | ✓
Drift hazard documented in the let-block comment: if the rust
constants change (MANAGER_PORT, WEB_PORT_BASE, WEB_PORT_RANGE, or
the FNV-1a parameters), the nix copy needs a lockstep bump or
gateway will proxy to wrong ports. Tracked in the option's
description as a follow-up to single-source via
`/var/lib/hyperhive/meta/topology.json` lib.importJSON OR runtime
nginx-include written by c0re.
Char-code lookup table covers `[a-z0-9_-]` — the current
`hyperhive.user.name` alphabet. Names with other chars produce an
eval-time error rather than a silent wrong hash.
Verified:
- `nix eval` on the locations attrset for [iris atlas argus damocles]
→ correct ports (matching rust impl) on each `/agent/<name>/` block
- empty `cfg.agents` default → no per-agent blocks (`[ "/" ]` only)
- full container toplevel builds cleanly with 7 agents + matrix on
(`nixos-system-hive-gateway-26.05pre-git`)
Sequencing per mara: this is #15 v0 (gateway-side per-agent routing,
purely additive). #14 netns isolation follows once this soaks.
Out of scope: TLS, manager special-case routing, per-agent unix
sockets (mara: "at some point the agent servers will be domain
sockets"), CORS workaround removal at `POST /answer-question/{id}`,
gateway auth.
Closes#15 v0.
mara on PR #729: "this still feels hacky - is there a proper way to do this?"
damocles: agreed, "Accept-header map is meaningfully better than the
allowlist [...] one map definition that encodes browser semantics directly,
vs ~20 extensions to keep synced with whatever fluffychat (and any future
hyperhive-served SPA) decides to ship".
The previous shape (#684 catch-all regex, then this PR v1's
extension allowlist) leaned on heuristics to distinguish "missing
asset → 404" from "unknown SPA route → fall back to index.html".
Both shapes were fragile against a SPA shipping a new extension,
and the allowlist became dead code the moment a route ended in
`.html-ish-suffix`.
The proper distinction lives at the HTTP layer: top-frame browser
navigations send `Accept: text/html,...` (chrome/firefox/safari are
consistent on this). Asset fetches from script tags / img / fetch() /
XHR send asset-typed Accepts (`image/*`, `application/javascript`,
`*/*`) without `text/html`.
Mechanics: an `nginx http`-context `map` keyed on `$http_accept`
emits either `/matrix/index.html` (navigation) or a sentinel
nonexistent path (`/__matrix_spa_no_html_fallback`); the location's
`try_files $uri $uri/ $matrix_spa_target =404;` does the right thing
for both cases. No extension list, no regex narrowing, no `if` block,
no named-location fallback.
The `map` lives in `services.nginx.appendHttpConfig` (only added
when the matrix GUI is on, otherwise no `map` directive at all).
The location's `extraConfig` is now a single `try_files` line.
Verified via `nix eval` on both the rendered `appendHttpConfig` and
the location's `extraConfig`. Full closure build pending operator
deploy.
Closes#686.
mara's first deploy hit:
Error: Couldn't resolve the package 'matrix' in 'package:matrix/matrix.dart'.
/nix/store/k9j8ns45fz7rpjp6rzk33ydjng67pgm0-source/web/native_executor.dart:1:8:
Error: Not found: 'package:matrix/matrix.dart'
Root cause: `dart compile js` walks up from the source file's dir to
find `.dart_tool/package_config.json`. My previous postInstall passed
`$src/web/native_executor.dart` — pointing dart at the unpacked nix
source, which has no `.dart_tool/` (pub-get wrote it to the build CWD,
not the read-only store path).
Fix: use a relative path `web/native_executor.dart`. nixpkgs's
buildFlutterApplication leaves CWD at the source root in postInstall
(its installPhase is just `cp -r build/web "$out"` with no `cd`
first — see `pkgs/development/compilers/flutter/build-support/
build-flutter-application.nix`), so the relative path walks up from
`web/` to the build CWD where pub-get's package_config lives.
Verified by `nix eval`; full closure build pending operator deploy.
Followup to #697 (the original fix; merged but mara's deploy then
surfaced this regression).
mara on PR #697: "this still puts us in the position of having to update
that dependency in sync with upstream. cant we use the one from the
nixpkgs build directly somehow?"
Drops the parallel `fetchurl` + sha256 pin in `fluffychat-web-imaging`.
Source now comes from
`pkgs.fluffychat-web.passthru.pubspecLock.dependencySources.native_imaging`
— the exact derivation that fluffychat-web's flutter build already pulls
into its pub-cache for the dart-side bindings. Version likewise pulled
from `passthru.pubspecLock.dependencyVersions.native_imaging`.
Result: when nixpkgs bumps `pkgs.fluffychat-web` (and with it the
pubspec.lock-resolved native_imaging version), our build automatically
picks up the matching source. No parallel hash to bump, no risk of drift
between the dart-side bindings and the wasm-side C compile.
Verified the build still works against the pub-cache-sourced derivation
(same Makefile, same emscripten flow):
$ nix-build test-passthru.nix
...
buildPhase completed in 52 seconds
$ ls /nix/store/.../fluffychat-web-imaging-0.4.0/
Imaging.js (9956 bytes)
Imaging.wasm (67363 bytes)
Byte-for-byte identical to the previous v2 output, just sourced from
the same store path fluffychat-web itself uses.
Follow-up to argus's v2 🟢 review of #697. No regression on the prior
review feedback — `make -C js` + explicit installPhase paths still in place.
Switches from `cd js && make ...` to `make -C js ...` so buildPhase
leaves pwd at the source root. installPhase's `js/Imaging.{js,wasm}`
paths are now correct against an explicit pwd rather than relying on
buildPhase's mid-phase cd side-effect carrying over.
No functional change — just robustness against future phase reorders /
`dontBuild` overrides, per argus's 🟡 on the v2 review of #697.
mara on PR #697: "dont use the prebuilt binary, fix the compile of the
one in nixpkgs (however its easiest: overlay, own derivation based on
it, hell if you want to you can fix buildFlutterApplication, idk)."
Replaces the upstream-tarball vendor with an own derivation that
compiles `Imaging.{js,wasm}` from the `native_imaging` dart package's
C source via `pkgs.emscripten`. Same source provenance as fluffychat
itself uses (both pin native_imaging 0.4.0 from pub.dev), now actually
exercised at build time.
Mechanics: new `fluffychat-web-imaging` derivation in the `let` block:
- src: `fetchurl` from pub.dev's `native_imaging-0.4.0.tar.gz`
(hash sha256-ztessYuApDFXjJBo65w+51+N85SR6K2vRxY1usKC1lE=)
- nativeBuildInputs: emscripten + cmake + gnumake + jq
- buildPhase: `cd js && make Imaging.js Imaging.wasm`
(`HOME` + `EM_CACHE` set in TMPDIR so emscripten's sysroot
builds work in the sandbox — standard nixpkgs pattern for
emcc-using derivations, see pkgs/top-level/emscripten-packages.nix)
- installPhase: `install -m 644` the two output files
Closure cost: build-time only — `pkgs.emscripten` is ~3.6 GiB
(LLVM + toolchain). Runtime closure is just the two produced files,
nothing emscripten-shaped survives into the deployed dist.
`postInstall` in `fluffychat-web-fixed` now references
`${fluffychat-web-imaging}` for the install copies, replacing the
previous reference to the deleted `fluffychat-web-imaging-prebuilt`
runCommandLocal.
Verified the emscripten build runs cleanly against the Makefile:
$ nix-build test-imaging-built.nix
...
emcc -s MODULARIZE=1 -s ALLOW_MEMORY_GROWTH=1 -O3 --closure 1 ...
cache:INFO: generating system library: sysroot/lib/.../libstubs.a ...
cache:INFO: generating system library: sysroot/lib/.../libc.a ...
cache:INFO: generating system library: sysroot/lib/.../libc++-noexcept.a ...
cache:INFO: generating system library: sysroot/lib/.../libc++abi-noexcept.a ...
make: Nothing to be done for 'Imaging.wasm'.
buildPhase completed in 52 seconds
/nix/store/x5ds7rkrgfgyy3gb1lk44ak8mkdvdx0p-fluffychat-web-imaging-0.4.0
$ ls /nix/store/.../fluffychat-web-imaging-0.4.0/
Imaging.js Imaging.wasm
$ stat -c '%s' .../Imaging.js .../Imaging.wasm
9956
67363
Matches the upstream prebuilt byte-counts (9936 + 67770 — small
delta from different emscripten / closure-compiler versions).
mara on #685: "do the post build step. check if there are any more file
that should have been built." Investigated the full diff between
`pkgs.fluffychat-web` (nixpkgs's nix-built dist) and the upstream
prebuilt release tarball. **Three** files missing from the nix build:
1. `native_executor.js` — flutter web worker entry. Source is
`web/native_executor.dart` in fluffychat. `flutter341.buildFlutterApplication`
skips `web/*.dart` worker entries; needs a separate `dart compile js`
pass. Solved by adding `pkgs.flutter341.dart` to nativeBuildInputs +
`dart compile js` in postInstall.
2. `Imaging.js` (~10 KB) + `Imaging.wasm` (~67 KB) — emscripten-compiled
C library from the `native_imaging` dart package (vendored by
Famedly). The package ships only C source + a Makefile that builds
them via `emcc`; the package does NOT ship prebuilt versions —
they're expected to be built at install time. nixpkgs's flutter
builder doesn't run that pipeline. Two paths considered:
- run emcc at build time: +~600 MB of `pkgs.emscripten` closure
for two files
- vendor the prebuilt files from the upstream release tarball:
same fluffychat release version → byte-identical output
Chose vendoring (cheaper closure, same result). Pinned to
`pkgs.fluffychat-web.version`-templated URL with sha256, so a
version bump auto-fetches the matching prebuilt.
The single other missing file (`native_executor.js.deps`) is a Dart
build-metadata artefact, not used at runtime — ignored.
Mechanics: two new `let`-bindings in `nix/modules/hive-matrix.nix`:
- `fluffychat-web-imaging-prebuilt` — small `runCommandLocal` that
fetches the upstream `fluffychat-web.tar.gz` and extracts just the
two Imaging files. Hash pinned, URL templated on the nixpkgs
fluffychat-web version.
- `fluffychat-web-fixed` — `pkgs.fluffychat-web.overrideAttrs` carrying
forward the existing `--base-href "/matrix/"` override (#634) plus
the new postInstall that runs `dart compile js` on
`web/native_executor.dart` and installs the two Imaging files from
the prebuilt derivation.
Then `services.hyperhive.matrix.gui.package`'s default flips from the
inline overrideAttrs to `fluffychat-web-fixed`.
Symptom this resolves: fluffychat-web's blank-page-after-load (#643)
caused by `main.dart.js` requesting `native_executor.js` and the SPA
runtime never booting. With `native_executor.js` present + the
gateway-side SPA-fallback fix (#684 making missing assets visible),
flutter's bootstrap completes and the login form is usable.
Verified `nix eval` produces a different derivation hash than the
unpatched `pkgs.fluffychat-web` (aag55wgh... vs ldgcxy50...),
confirming the override takes effect. Full closure build pending
operator deploy — local sandbox networking flaky.
Closes#685.
iris's diagnosis on #643 (mara's fluffychat-web login attempt): the gateway's
`/matrix/` location used
try_files $uri $uri/ /matrix/index.html;
which silently returned `index.html` (Content-Type: text/html, status 200) for
ANY missing path under `/matrix/`, including static assets like
`native_executor.js`. flutter's bootstrap requested that JS file, got HTML back,
failed to load the JS runtime, and the page rendered blank without any visible
error in the browser console.
(Confirmed root cause for the missing file itself: the upstream `fluffychat-web`
dist in nixpkgs ships `native_executor.dart` but no compiled `native_executor.js`,
even though `main.dart.js` references the latter. That's a separate
fluffychat-web packaging issue — tracked separately; this PR fixes only the
gateway-side masking that hides such failures.)
Replaces the inline `try_files` fallback with a named-location fallback that
distinguishes between route-shaped URIs (no extension) and asset-shaped URIs
(any `.<ext>` suffix):
location /matrix/ {
alias <pkg>/;
try_files $uri $uri/ @matrix_spa_fallback;
}
location @matrix_spa_fallback {
if ($uri ~ "\.[A-Za-z0-9]+$") {
return 404;
}
rewrite ^ /matrix/index.html last;
}
Routes still fall back to `index.html` so SPA client-side routing keeps
working; missing assets now surface a real 404 so flutter (and the operator's
devtools) can see the failure.
Verified the rendered nginx location attr via
`nix eval .#nixosConfigurations.* .... locations."@matrix_spa_fallback".extraConfig`.
argus on #661🟡: "matrix IDs embed server_name irrevocably — anyone
already running with the old default would be broken."
Append a `**Breaking change as of #660**` paragraph to the
`serverName` option description with the exact opt-back-in string,
matching the pattern from #651's openFirewall flip. PR body + commit
already documented the breakage; this surfaces it in the option's
own description so it shows up in `nix flake show` + the auto-
generated options docs right next to the option.
mara on #660: "Matrix domain should default to hive domain if not
set otherwise / redirect matrix clients with .well-known"
Two coupled changes:
1. `services.hyperhive.matrix.serverName` default flipped from
`matrix.${services.hyperhive.domain}` (subdomain) to just
`${services.hyperhive.domain}` (bare hive domain).
This is a "for new deploys only" change — `server_name` is
embedded irrevocably in every user/room ID, so existing
homeservers must set `serverName` explicitly to preserve the
subdomain shape if that's where their identifiers were minted.
Description updated to point at the .well-known piece below.
2. `hive-gateway` nginx now serves matrix-spec `.well-known`
auto-discovery JSON at the canonical location when matrix is
enabled + hive domain set:
GET /.well-known/matrix/client
{"m.homeserver":{"base_url":"http://<domain>:<httpPort>"}}
+ Access-Control-Allow-Origin: * (per matrix spec)
GET /.well-known/matrix/server
{"m.server":"<domain>:<httpPort>"}
tuwunel serves both client + federation on the same `httpPort`
(see hive-matrix.nix), so both records point at the same
endpoint. No-op when matrix isn't enabled or hive domain isn't
set — nothing to advertise.
Combined effect: with `services.hyperhive.domain = "darkest.space"` +
matrix enabled, a matrix client pointed at `darkest.space` resolves
through `.well-known` to the actual `:8008` endpoint, no subdomain
needed. MXIDs become `@atlas:darkest.space` (was: `@atlas:matrix.darkest.space`).
Verified via `nix eval`:
- server_name = "darkest.space" (was "matrix.darkest.space")
- gateway locations include `= /.well-known/matrix/client` + `= /.well-known/matrix/server`
- well-known/matrix/client returns the spec-shaped JSON
Caveat: `m.homeserver.base_url` advertises HTTP (no TLS yet —
follow-up). matrix clients increasingly require HTTPS for new
account creation, so the v0 setup works for local-network testing
but won't satisfy public clients until the gateway TLS story lands.
Closes#660.
Per iris's recommendation on #644 [comment 8043](http://localhost:3000/hyperhive/hyperhive/issues/644#issuecomment-8043):
swap the `chown root:tuwunel + chmod 0640 + pinned GID 10042` shape
(shipped via #649) for systemd's `LoadCredential=` mechanism.
How it works: systemd reads the host-side file at service start,
copies it into a per-service credentials dir
(`/run/credentials/tuwunel.service/registration_token`) owned by
the dynamic user with mode 0400. Service reads from there. All the
namespace mapping happens transparently inside systemd — keeps
`DynamicUser=true` + `PrivateUsers=true` intact.
Net diff from current shape:
- DROP `users.groups.tuwunel.gid = 10042;` from BOTH host AND container
- DROP `chown root:tuwunel "$tokenFile"; chmod 0640 "$tokenFile"`
from activation script; replace with `chmod 0600` (root:root)
- DROP `[ "var" "users" ]` activation dep on `users` (no longer
needs the group to exist before chown)
- ADD `systemd.services.tuwunel.serviceConfig.LoadCredential = [...]`
inside the container config
- CHANGE `registration_token_file` from the bind-mount path to
`/run/credentials/tuwunel.service/registration_token`
- KEEP the bind mount + activation-script token generation (load
credential reads the bind-mounted host file at service start)
Verified via `nix eval`:
- host: no `users.groups.tuwunel` (was: gid = 10042)
- container: tuwunel group exists with `gid = null` (auto-allocated;
no longer pinned to match host since it doesn't need to)
- container: tuwunel.service.serviceConfig.LoadCredential =
`["registration_token:/var/lib/hyperhive/matrix-register-token"]`
- container: services.matrix-tuwunel.settings.global.registration_token_file =
`/run/credentials/tuwunel.service/registration_token`
`/run/credentials/<service>/<id>` is a systemd-stable path
(documented in `man systemd.exec` → LoadCredential); safe to
hardcode.
argus picked option (a) on #653: put the upgrade note in each option's
`description` so it shows up in `nix flake show` + the rendered
options docs, right next to the option itself. cheapest option, no
eval-time noise (a `warnings` block would fire on every new
deployment that wants false — the normal case now).
Appended a `**Breaking change as of #651**` paragraph to each of the
three `openFirewall` descriptions, naming the exact option string the
operator needs to set to restore the old behaviour.
Gateway's note specifically calls out that external reach is the
common case (operator's primary entry point), so the upgrade hint
is most likely needed there.
mara on #651: "Dont default openFirewall to true."
Flip the `openFirewall` default from `true` to `false` for all three
modules that expose host-side ports:
- `services.hyperhive.forge.openFirewall` (httpPort 3000 + sshPort 2222)
- `services.hyperhive.gateway.openFirewall` (port 80)
- `services.hyperhive.matrix.openFirewall` (httpPort 8008)
Rationale: secure-by-default. With shared host netns, the host +
every agent container reach these services via `localhost` regardless
of the firewall — the open only matters for access from outside the
host. Operators who want external reach now flip the bool explicitly:
services.hyperhive.gateway.openFirewall = true;
Each description updated to explain the new default + when to flip
it (operator's browser, external git clients, federation announcement,
etc.). Behind a host-level reverse proxy that handles TLS, leave off.
Verified via `nix eval` on a clean stub config:
- forge openFirewall = false
- gateway openFirewall = false
- matrix openFirewall = false
- networking.firewall.allowedTCPPorts = [] (was: [80 2222 3000 8008])
Note: c0re's direct ports (7000/8000/8100-8999) are gated separately
via #621 on `gateway.enable` — that gate stays; this PR only touches
the per-module `openFirewall` knobs.
Closes#651.
Per #621 (filed as follow-up to #620 v0): when the gateway is on
(now the default), the c0re dashboard / manager / sub-agent direct
ports should NOT be open in the host firewall — the gateway nginx
is the sole external entry point, proxying to `127.0.0.1:7000` etc.
internally. Leaving them open in the firewall defeats the "single
front door" story.
Wraps the existing `allowedTCPPorts` + `allowedTCPPortRanges` blocks
in `lib.mkIf (!config.services.hyperhive.gateway.enable)`. Operators
who opt out of the gateway still get the direct ports opened so the
legacy `http://<host>:7000/` flow keeps working.
Verified via `nix eval`:
| gateway | allowedTCPPorts (host firewall) | allowedTCPPortRanges |
| --- | --- | --- |
| on | `[80 2222 3000]` (gateway + forge) | `[]` |
| off | `[2222 3000 7000 8000]` (forge + c0re + manager) | `[{from=8100; to=8999}]` (agents) |
Forge ports stay direct in both modes — `hive-forge.nix` opens them
independently and they're not proxied through the gateway (that's a
separate follow-up if wanted).
Closes#621.
Per mara's directive on #609: stand up a single nginx in its own
nixos-container, serve the matrix GUI static dist there, proxy
everything else to hive-c0re. v0 is HTTP-only; TLS / public-domain
shape lands in follow-ups.
New `nix/modules/hive-gateway.nix` declaring `containers.hive-gateway`
modelled on `hive-forge`:
- nixos-container running nginx, shares host netns
- `location /matrix/` → static-serves `hyperhive.matrix.gui.package`
(fluffychat-web by default) when `matrix.gui.enable` is true
- `location /` → proxy_pass to `127.0.0.1:${dashboardPort}` with
websocket + SSE upgrade headers + 1d read timeout
Options (`hyperhive.gateway.*`):
- `enable` (default `true`) — gateway on by default, opt out to bypass
- `port` (default `80`) — nginx listen port on the host
- `upstreamHost` / `upstreamPort` — c0re target, defaults to
`127.0.0.1:${services.hive-c0re.dashboardPort}`
- `openFirewall` (default `true`) — open the listen port
- `localHostsEntry` (default `false`) — when true, adds an
`/etc/hosts` entry mapping `hyperhive.domain` → `127.0.0.1` for
local-dev / test loops without real DNS (per mara's spec)
`hive-c0re.nix` updates: when gateway is enabled, skip wiring
`HIVE_MATRIX_GUI_DIR` (gateway owns `/matrix/` now). When gateway is
off, c0re's pre-existing matrix mount stays as the fallback.
README: short "Optional" block introducing the gateway + the
`localHostsEntry` knob.
```sh
nix flake check --no-build
nix build .#docs-host
```
End-to-end eval matrix:
| gateway.enable | matrix.gui.enable | c0re HIVE_MATRIX_GUI_DIR | gateway container |
| --- | --- | --- | --- |
| true (default) | true | unset (gateway serves) | present |
| true | false | unset | present, no /matrix |
| false | true | set (c0re serves) | absent |
| false | false | unset | absent |
- TLS termination — separate follow-up once mara picks a story
(self-signed-mkcert vs operator-provided certs)
- Per-agent UI routing (`/agent/<name>/`) — depends on agent base-path
support which is a frontend lift
- Subdomain routing for `matrix.${hyperhive.domain}` — same-origin
`/matrix/` is the v0 shape per mara ("leave everything else as is")
Closes part of #609 (matrix GUI re-rooting onto nginx); leaves the
issue open for the subdomain re-root + `.well-known/matrix/client`
piece once the multi-host story matures.
Per [mara on PR #615 comment 7349](http://localhost:3000/hyperhive/hyperhive/pulls/615#issuecomment-7349):
> follow nix conventions, services.hyperhive it is. the earlier we
> change this, the less breakage.
Renames the entire host-side option tree under `services.hyperhive.*`:
- `services.hive-c0re.*` → `services.hyperhive.c0re.*`
- `hyperhive.enable` → `services.hyperhive.enable`
- `hyperhive.domain` → `services.hyperhive.domain`
- `hyperhive.forge.*` → `services.hyperhive.forge.*`
- `hyperhive.matrix.*` → `services.hyperhive.matrix.*`
Per mara's "earlier = less breakage", the previous `services.hive-c0re.enable`
deprecation alias is dropped. Operators get a clear eval error on the
old paths pointing at the rename. Single migration moment.
Per-agent options in `nix/templates/harness-base.nix` (`hyperhive.model`,
`hyperhive.allowedRecipients`, etc.) stay at `hyperhive.*` — they're
container-level config, not services in the NixOS sense.
Verified via `nix flake check --no-build` + an end-to-end NixOS eval
exercising every renamed path.
Follow-up needed: rust source comments referencing the old NixOS
option names (`hive-c0re/src/{meta,coordinator,main,dashboard}.rs`)
should be updated in a separate pure-rust PR to keep this one
strictly nix-only.
- Move options.services.hive-c0re → options.hyperhive.c0re
- Add options.hyperhive.enable to auto-enable c0re + subsystems
- Add deprecation alias for services.hive-c0re.enable (backward compat)
- Update doc references in README, flake.nix, docs, harness-base.nix
- Simplifies config: 'hyperhive.enable = true' now enables everything
Existing operator configs using services.hive-c0re.enable will
continue to work but emit a deprecation warning. Aligns the option
namespace with the existing hyperhive.* family (matrix, forge, domain).
fixes#612
Cuts every `include_bytes!`/`include_str!` of a non-rust path in
the workspace over to runtime file loads from `$HIVE_ASSETS_DIR`
(the `hyperhive-assets` derivation introduced in the previous
commit). After this commit the rust derivation has no compile-time
dependency on `branding/*` or `hive-ag3nt/prompts/*` anymore.
Call-site flips:
- `hive-c0re/src/forge.rs::CORE_AVATAR_PNG` /
`CONFIG_ORG_AVATAR_PNG`: were `include_bytes!` of
`branding/hyperhive.png` and `$OUT_DIR/agent-configs.png`. Now
`ensure_core_avatar` / `ensure_config_org_avatar` `tokio::fs::read`
via `hive_sh4re::assets::{core_avatar_png, config_org_avatar_png}`
at startup. The `agent-configs.png` is now rendered by the
`hyperhive-assets` derivation's rsvg-convert step (was
`hive-c0re/build.rs` + librsvg on the rust derivation's
nativeBuildInputs — both gone in the next commit).
- `hive-ag3nt/src/prompt.rs::TEMPLATE`: `render` now takes the
template as an argument; `write_system_prompt` reads it once from
`$HIVE_ASSETS_DIR/prompts/system.md` before calling render. The
test module still `include_str!`s the production template so
`cargo test --workspace` doesn't need `HIVE_ASSETS_DIR` set —
this is the only remaining compile-time reference to the file
from the rust workspace, gated to `#[cfg(test)]`.
- `hive-ag3nt/src/turn.rs::CLAUDE_SETTINGS`: was `include_str!`'d
and written via `tokio::fs::write`; now `tokio::fs::copy` from
`$HIVE_ASSETS_DIR/prompts/claude-settings.json` into the
per-agent socket dir.
- `hive-ag3nt/src/web_ui.rs::DEFAULT_ICON`: was `include_str!`'d;
now read on-demand from `$HIVE_ASSETS_DIR/branding/hyperhive.svg`
inside `serve_icon`. Falls back to an empty body if missing so
the endpoint never panics on a misconfigured container (matches
the existing "per-agent icon.svg override" fallthrough).
`HIVE_ASSETS_DIR` wiring:
- Inside containers: `nix/templates/harness-base.nix`
`environment.variables` sets it to
`${pkgs.hyperhive-assets}/share/hyperhive` (resolved through
the default overlay applied in `mkContainer`). Verified by
building `agent-base-toplevel` and grepping the resulting
`/etc/set-environment`.
- Host-side: `nix/modules/hive-c0re.nix` adds an `assets` option
defaulting to `hyperhive.packages.${system}.assets`, threaded
in from the flake's nixosModules wiring, and sets the same env
var on the `hive-c0re` systemd unit so the daemon's
`forge::ensure_*_avatar` startup hooks find the PNGs.
`hive-c0re/build.rs` deleted entirely; `[package].build` removed
from `hive-c0re/Cargo.toml`; rsvg-convert dependency lives in the
assets derivation only.
Validated: `nix build .#default .#checks.x86_64-linux.clippy
.#agent-base-toplevel .#manager-toplevel --fallback` all succeed.
`/etc/set-environment` in the toplevel shows
`HIVE_ASSETS_DIR="/nix/store/.../hyperhive-assets-0.1.0/share/hyperhive"`.