docs(gotchas): group entries by area, fix misplaced nix-fmt section

This commit is contained in:
iris 2026-08-15 11:51:18 +02:00 committed by mara
commit e9e4408af1

View file

@ -2,8 +2,12 @@
NixOS + nspawn quirks and lessons we hit the hard way. If something
here looks unmotivated in the code, there's usually a story underneath.
Grouped by area — jump to the section that matches what you're
touching.
## `nixos-container` doesn't expose `--bind` on the CLI
## NixOS / nspawn containers
### `nixos-container` doesn't expose `--bind` on the CLI
The CLI doesn't accept `--bind`. Path is via `EXTRA_NSPAWN_FLAGS` in
`/etc/nixos-containers/<NAME>.conf` — the start script
@ -11,18 +15,18 @@ The CLI doesn't accept `--bind`. Path is via `EXTRA_NSPAWN_FLAGS` in
`systemd-nspawn` invocation. `lifecycle::set_nspawn_flags()` rewrites
this line.
## `/run/systemd/nspawn/*.nspawn` overrides are ignored
### `/run/systemd/nspawn/*.nspawn` overrides are ignored
`nixos-container`'s start script builds the nspawn command line
directly. Dropping a `.nspawn` file under `/run/systemd/nspawn/`
looks like the obvious extension point and does nothing. Use
`EXTRA_NSPAWN_FLAGS` (above).
## `boot.isNspawnContainer = true`
### `boot.isNspawnContainer = true`
Not `boot.isContainer = true`. Renamed in nixos-25.11+.
## `nixos-container create` auto-assigns `HOST_ADDRESS` / `LOCAL_ADDRESS`
### `nixos-container create` auto-assigns `HOST_ADDRESS` / `LOCAL_ADDRESS`
…in the `.conf`. The start script's `if HOST_ADDRESS set →
--network-veth` branch then forces a private netns — silently fatal
@ -30,7 +34,7 @@ for our web UIs (the bind is invisible from the host). We
force-clear `HOST_ADDRESS` / `LOCAL_ADDRESS` / `HOST_ADDRESS6` /
`LOCAL_ADDRESS6` / `HOST_BRIDGE` and set `PRIVATE_NETWORK=0`.
## systemd service PATH ≠ host PATH
### systemd service PATH ≠ host PATH
The hive-c0re service sets `path = [ pkgs.git "/run/current-system/sw" ]`.
In-container harness services do the same so anything an agent adds
@ -40,7 +44,7 @@ editing the service definition.
`environment.HYPERHIVE_GIT` bakes git's absolute path in (read by
`lifecycle::git_command()`) for the host.
## `systemd.services.*.path` appends `/bin` to every entry
### `systemd.services.*.path` appends `/bin` to every entry
NixOS's `systemd.services.<unit>.path` list feeds every entry through
`lib.makeBinPath`, which **appends `/bin` unconditionally**. That's
@ -62,19 +66,21 @@ contains a non-existent directory. The first symptom is usually
the setuid sudo wrapper lives at `/run/wrappers/bin/sudo` and
the path entry resolves to `/run/wrappers/bin/bin` instead.
## `RuntimeDirectoryPreserve = "yes"`
### `RuntimeDirectoryPreserve = "yes"`
…keeps `/run/hyperhive/` (and the per-agent sub-dirs) across
hive-c0re restarts. Without it, every restart wipes bind sources and
existing containers can't be started.
## `register_agent` is idempotent
### `register_agent` is idempotent
Drops any prior socket task before rebinding. Required so a
hive-c0re restart followed by `rebuild alice` recreates the agent's
socket without needing a clean reinstall.
## `claude-code` is unfree
## Claude Code packaging & credentials
### `claude-code` is unfree
`claude-code` comes from the flake's main `nixpkgs` (nixos-26.05).
It's unfree, so the agent modules set `config.allowUnfreePredicate`
@ -125,7 +131,7 @@ the hive's `claude` out from under it. The price of the root is that an
old `claude-code` can't be reclaimed until every agent has rebuilt past
it and the old generations are gone.
## Claude credentials are per-agent
### Claude credentials are per-agent
`/var/lib/hyperhive/agents/<name>/claude/` bind-mounts to
`/home/<name>/.claude` (RW). Sharing one dir across agents is NOT viable —
@ -133,7 +139,7 @@ OAuth refresh tokens rotate, so any sibling refresh invalidates all
the others. Login flow runs from the per-agent web UI; creds persist
across `destroy`/recreate (`--purge` wipes them).
## Persistent notes dir per agent
### Persistent notes dir per agent
`/var/lib/hyperhive/agents/<name>/state/` bind-mounts to
`/agents/<name>/state` (RW; uniform for all agents).
@ -143,7 +149,9 @@ durable knowledge here (`notes.md`, anything else). The harness also
writes its events log here (`hyperhive-events.sqlite`).
Survives `destroy`/recreate alongside the claude dir.
## Web UI ports collide on hash
## Networking & ports
### Web UI ports collide on hash
Sub-agent web UI ports are deterministic FNV-1a of the agent name
modulo 900 (range 8100..8999). With ~30 agents the birthday-paradox
@ -155,7 +163,7 @@ reproducible from just the name. Every agent hashes into
8100..8999 via the same FNV-1a; dashboard
at `cfg.dashboardPort` (default 7000).
## Restart races on TCP bind
### Restart races on TCP bind
Both the dashboard and per-agent web UI use `tokio::net::TcpSocket`
with `SO_REUSEADDR` plus a retry-on-`AddrInUse` loop (12 tries,
@ -166,22 +174,18 @@ overlap" case. REUSEADDR does **not** allow two simultaneous
`LISTEN` sockets on the same port (that would be `SO_REUSEPORT`,
which we don't use) — exclusivity is preserved.
## Orphan approvals
## Approvals
### Orphan approvals
If state dirs are wiped out from under a pending approval (test
scripts, manual `rm -rf`), the dashboard's next render marks them
`failed` with note `"agent state dir missing"` so they fall out of
`pending`. They stay in sqlite for audit.
## Nix store `cp -r` preserves read-only bits
## Gateway / SPA serving
Copying a nix store path with `cp -r src/. $out/` inside a
`pkgs.runCommand` derivation preserves the read-only permissions of
store files. Any subsequent write into the copied tree (adding new
files in subdirectories) fails with `EPERM`. Fix: pass
`--no-preserve=mode,ownership` so the output tree is writable.
## SPA fallback: use `Accept` header map, not `try_files ... /index.html`
### SPA fallback: use `Accept` header map, not `try_files ... /index.html`
The naive nginx pattern for a path-prefix SPA (`try_files $uri $uri/
/matrix/index.html`) silently swallows asset 404s — a missing JS file
@ -210,7 +214,17 @@ firefox / safari are consistent). Asset fetches (`image/*`,
fall through to the trailing `=404`. No extension list to maintain;
no named-location indirection needed.
## `nix build flake#name` does not walk into `nixosConfigurations`
## Build & dev workflow
### Nix store `cp -r` preserves read-only bits
Copying a nix store path with `cp -r src/. $out/` inside a
`pkgs.runCommand` derivation preserves the read-only permissions of
store files. Any subsequent write into the copied tree (adding new
files in subdirectories) fails with `EPERM`. Fix: pass
`--no-preserve=mode,ownership` so the output tree is writable.
### `nix build flake#name` does not walk into `nixosConfigurations`
`nix build` resolves the fragment (`#name`) against the flake's
**top-level output attrs** — not against `nixosConfigurations`
@ -233,12 +247,7 @@ instead of `meta#nixosConfigurations.argus.config…`. The fix:
`split_once('#')` to separate flake path from name, then template
`{path}#nixosConfigurations.{name}.config.system.build.toplevel`.
## `hive-forge`: prefer over raw curl pipelines
Full CLI reference: [`docs/tools/forge.md`](tools/forge.md).
Never use raw `curl` for forge access.
## Containerized nix-daemon needs `sandbox-fallback = true`
### Containerized nix-daemon needs `sandbox-fallback = true`
Agent containers bind-mount the host's nix-daemon socket. nspawn
containers don't get user-namespaces by default, so `nix build`
@ -249,7 +258,7 @@ and fail outright if the host daemon's
fall back to unsandboxed local builds rather than failing. Security
implications: `docs/security.md`.
## Linking workspace binaries locally needs `nix develop`
### Linking workspace binaries locally needs `nix develop`
The Rust workspace links `libsqlite3-sys` (rusqlite) against the
system `libsqlite3`. Agent containers carry no system libsqlite3 on
@ -271,7 +280,7 @@ e.g. `docs/tools/hivectl-cli.md` via the `hivectl markdown-docs`
subcommand (its `hivectl-docs` flake check otherwise only fails in
CI on drift).
## Split asset derivations away from the rust workspace
### Split asset derivations away from the rust workspace
`nix/packages/assets.nix` builds the branding SVG/PNG family + claude
system-prompt template + claude-settings JSON as its own derivation,
@ -285,7 +294,46 @@ The agent-configs PNG is rendered from the SVG via `rsvg-convert` at
build time; librsvg dependency lives here, not in the rust
derivation's `nativeBuildInputs`.
## Weston VNC compositor (per-agent `hyperhive.gui.enable`)
### `nix fmt` fails in a git worktree with "object not found"
`nix fmt` (and any `nix` command that fetches a `git+file://` flake
URL) uses libgit2 internally to compute `revCount` — the number of
commits reachable from HEAD. This walk fails with:
```
error: getting Git object '<hash>': object not found (libgit2 error code = 9)
```
when a commit that was reachable at some earlier evaluation is now gone
(GC'd, rebased away, or pruned). The failure is persistent: clearing
`~/.cache/nix/{eval-cache-v6,gitv3,fetcher-cache-v4.sqlite}` does not
help because the missing object is a structural gap in the git object
graph itself, not in nix's caches.
**Workaround: use a plain clone, not a git worktree.**
```bash
git clone http://<forge>/hyperhive/hyperhive.git ~/hh-work
cd ~/hh-work && nix fmt
```
The root cause is specific to worktrees: a worktree shares the object
store with its parent repo. If the parent repo's history was rewritten
(rebase, force-push, `git gc --prune`) while the worktree was checked
out at a branch tip that references the pruned commits via its reflog or
history, libgit2's rev-walk encounters the gap. A plain clone has its
own self-consistent object store and is immune to the issue.
## Tooling
### `hive-forge`: prefer over raw curl pipelines
Full CLI reference: [`docs/tools/forge.md`](tools/forge.md).
Never use raw `curl` for forge access.
## GUI (weston/VNC)
### Weston VNC compositor (per-agent `hyperhive.gui.enable`)
`nix/agent-modules/weston-vnc.nix` adds an optional Weston Wayland
compositor with the VNC backend, surfaced as
@ -361,7 +409,9 @@ connects to the compositor at `127.0.0.1:<vnc_port>`.
`wl_event_source_timer_update` treats as "disarm", so the
compositor never goes idle and never locks.
## Nix options reference (`nix/docs/default.nix`)
## Nix docs pipeline
### Nix options reference (`nix/docs/default.nix`)
`pkgs.nixosOptionsDoc` over two evaluated module trees:
`hostEval` (a stub NixOS system loading the `nix/host-modules/` aggregator with every
@ -397,7 +447,7 @@ options tree picks up everything under that root — picking against
stray roots produces an empty tree and renders the host page as
template chrome with no `<h2>` headers.
### Docs drv stability: `nixSrc`
#### Docs drv stability: `nixSrc`
Naively, the docs evaluation depends on `self` (the flake's store path),
so every commit — even Rust-only or frontend-only changes — produces new
@ -426,33 +476,3 @@ Why `builtins.unsafeDiscardStringContext`? The path string
make `builtins.path` include `self` as a build dependency even after
content-addressing the directory. Discarding the context makes the
resulting `nixSrc` truly independent of `self`'s store path.
### `nix fmt` fails in a git worktree with "object not found"
`nix fmt` (and any `nix` command that fetches a `git+file://` flake
URL) uses libgit2 internally to compute `revCount` — the number of
commits reachable from HEAD. This walk fails with:
```
error: getting Git object '<hash>': object not found (libgit2 error code = 9)
```
when a commit that was reachable at some earlier evaluation is now gone
(GC'd, rebased away, or pruned). The failure is persistent: clearing
`~/.cache/nix/{eval-cache-v6,gitv3,fetcher-cache-v4.sqlite}` does not
help because the missing object is a structural gap in the git object
graph itself, not in nix's caches.
**Workaround: use a plain clone, not a git worktree.**
```bash
git clone http://<forge>/hyperhive/hyperhive.git ~/hh-work
cd ~/hh-work && nix fmt
```
The root cause is specific to worktrees: a worktree shares the object
store with its parent repo. If the parent repo's history was rewritten
(rebase, force-push, `git gc --prune`) while the worktree was checked
out at a branch tip that references the pruned commits via its reflog or
history, libgit2's rev-walk encounters the gap. A plain clone has its
own self-consistent object store and is immune to the issue.