diff --git a/docs/gotchas.md b/docs/gotchas.md index 6cbc7f9c..b5bf6782 100644 --- a/docs/gotchas.md +++ b/docs/gotchas.md @@ -2,8 +2,12 @@ NixOS + nspawn quirks and lessons we hit the hard way. If something here looks unmotivated in the code, there's usually a story underneath. +Grouped by area — jump to the section that matches what you're +touching. -## `nixos-container` doesn't expose `--bind` on the CLI +## NixOS / nspawn containers + +### `nixos-container` doesn't expose `--bind` on the CLI The CLI doesn't accept `--bind`. Path is via `EXTRA_NSPAWN_FLAGS` in `/etc/nixos-containers/.conf` — the start script @@ -11,18 +15,18 @@ The CLI doesn't accept `--bind`. Path is via `EXTRA_NSPAWN_FLAGS` in `systemd-nspawn` invocation. `lifecycle::set_nspawn_flags()` rewrites this line. -## `/run/systemd/nspawn/*.nspawn` overrides are ignored +### `/run/systemd/nspawn/*.nspawn` overrides are ignored `nixos-container`'s start script builds the nspawn command line directly. Dropping a `.nspawn` file under `/run/systemd/nspawn/` looks like the obvious extension point and does nothing. Use `EXTRA_NSPAWN_FLAGS` (above). -## `boot.isNspawnContainer = true` +### `boot.isNspawnContainer = true` Not `boot.isContainer = true`. Renamed in nixos-25.11+. -## `nixos-container create` auto-assigns `HOST_ADDRESS` / `LOCAL_ADDRESS` +### `nixos-container create` auto-assigns `HOST_ADDRESS` / `LOCAL_ADDRESS` …in the `.conf`. The start script's `if HOST_ADDRESS set → --network-veth` branch then forces a private netns — silently fatal @@ -30,7 +34,7 @@ for our web UIs (the bind is invisible from the host). We force-clear `HOST_ADDRESS` / `LOCAL_ADDRESS` / `HOST_ADDRESS6` / `LOCAL_ADDRESS6` / `HOST_BRIDGE` and set `PRIVATE_NETWORK=0`. -## systemd service PATH ≠ host PATH +### systemd service PATH ≠ host PATH The hive-c0re service sets `path = [ pkgs.git "/run/current-system/sw" ]`. In-container harness services do the same so anything an agent adds @@ -40,7 +44,7 @@ editing the service definition. `environment.HYPERHIVE_GIT` bakes git's absolute path in (read by `lifecycle::git_command()`) for the host. -## `systemd.services.*.path` appends `/bin` to every entry +### `systemd.services.*.path` appends `/bin` to every entry NixOS's `systemd.services..path` list feeds every entry through `lib.makeBinPath`, which **appends `/bin` unconditionally**. That's @@ -62,19 +66,21 @@ contains a non-existent directory. The first symptom is usually the setuid sudo wrapper lives at `/run/wrappers/bin/sudo` and the path entry resolves to `/run/wrappers/bin/bin` instead. -## `RuntimeDirectoryPreserve = "yes"` +### `RuntimeDirectoryPreserve = "yes"` …keeps `/run/hyperhive/` (and the per-agent sub-dirs) across hive-c0re restarts. Without it, every restart wipes bind sources and existing containers can't be started. -## `register_agent` is idempotent +### `register_agent` is idempotent Drops any prior socket task before rebinding. Required so a hive-c0re restart followed by `rebuild alice` recreates the agent's socket without needing a clean reinstall. -## `claude-code` is unfree +## Claude Code packaging & credentials + +### `claude-code` is unfree `claude-code` comes from the flake's main `nixpkgs` (nixos-26.05). It's unfree, so the agent modules set `config.allowUnfreePredicate` @@ -125,7 +131,7 @@ the hive's `claude` out from under it. The price of the root is that an old `claude-code` can't be reclaimed until every agent has rebuilt past it and the old generations are gone. -## Claude credentials are per-agent +### Claude credentials are per-agent `/var/lib/hyperhive/agents//claude/` bind-mounts to `/home//.claude` (RW). Sharing one dir across agents is NOT viable — @@ -133,7 +139,7 @@ OAuth refresh tokens rotate, so any sibling refresh invalidates all the others. Login flow runs from the per-agent web UI; creds persist across `destroy`/recreate (`--purge` wipes them). -## Persistent notes dir per agent +### Persistent notes dir per agent `/var/lib/hyperhive/agents//state/` bind-mounts to `/agents//state` (RW; uniform for all agents). @@ -143,7 +149,9 @@ durable knowledge here (`notes.md`, anything else). The harness also writes its events log here (`hyperhive-events.sqlite`). Survives `destroy`/recreate alongside the claude dir. -## Web UI ports collide on hash +## Networking & ports + +### Web UI ports collide on hash Sub-agent web UI ports are deterministic FNV-1a of the agent name modulo 900 (range 8100..8999). With ~30 agents the birthday-paradox @@ -155,7 +163,7 @@ reproducible from just the name. Every agent hashes into 8100..8999 via the same FNV-1a; dashboard at `cfg.dashboardPort` (default 7000). -## Restart races on TCP bind +### Restart races on TCP bind Both the dashboard and per-agent web UI use `tokio::net::TcpSocket` with `SO_REUSEADDR` plus a retry-on-`AddrInUse` loop (12 tries, @@ -166,22 +174,18 @@ overlap" case. REUSEADDR does **not** allow two simultaneous `LISTEN` sockets on the same port (that would be `SO_REUSEPORT`, which we don't use) — exclusivity is preserved. -## Orphan approvals +## Approvals + +### Orphan approvals If state dirs are wiped out from under a pending approval (test scripts, manual `rm -rf`), the dashboard's next render marks them `failed` with note `"agent state dir missing"` so they fall out of `pending`. They stay in sqlite for audit. -## Nix store `cp -r` preserves read-only bits +## Gateway / SPA serving -Copying a nix store path with `cp -r src/. $out/` inside a -`pkgs.runCommand` derivation preserves the read-only permissions of -store files. Any subsequent write into the copied tree (adding new -files in subdirectories) fails with `EPERM`. Fix: pass -`--no-preserve=mode,ownership` so the output tree is writable. - -## SPA fallback: use `Accept` header map, not `try_files ... /index.html` +### SPA fallback: use `Accept` header map, not `try_files ... /index.html` The naive nginx pattern for a path-prefix SPA (`try_files $uri $uri/ /matrix/index.html`) silently swallows asset 404s — a missing JS file @@ -210,7 +214,17 @@ firefox / safari are consistent). Asset fetches (`image/*`, fall through to the trailing `=404`. No extension list to maintain; no named-location indirection needed. -## `nix build flake#name` does not walk into `nixosConfigurations` +## Build & dev workflow + +### Nix store `cp -r` preserves read-only bits + +Copying a nix store path with `cp -r src/. $out/` inside a +`pkgs.runCommand` derivation preserves the read-only permissions of +store files. Any subsequent write into the copied tree (adding new +files in subdirectories) fails with `EPERM`. Fix: pass +`--no-preserve=mode,ownership` so the output tree is writable. + +### `nix build flake#name` does not walk into `nixosConfigurations` `nix build` resolves the fragment (`#name`) against the flake's **top-level output attrs** — not against `nixosConfigurations` @@ -233,12 +247,7 @@ instead of `meta#nixosConfigurations.argus.config…`. The fix: `split_once('#')` to separate flake path from name, then template `{path}#nixosConfigurations.{name}.config.system.build.toplevel`. -## `hive-forge`: prefer over raw curl pipelines - -Full CLI reference: [`docs/tools/forge.md`](tools/forge.md). -Never use raw `curl` for forge access. - -## Containerized nix-daemon needs `sandbox-fallback = true` +### Containerized nix-daemon needs `sandbox-fallback = true` Agent containers bind-mount the host's nix-daemon socket. nspawn containers don't get user-namespaces by default, so `nix build` @@ -249,7 +258,7 @@ and fail outright if the host daemon's fall back to unsandboxed local builds rather than failing. Security implications: `docs/security.md`. -## Linking workspace binaries locally needs `nix develop` +### Linking workspace binaries locally needs `nix develop` The Rust workspace links `libsqlite3-sys` (rusqlite) against the system `libsqlite3`. Agent containers carry no system libsqlite3 on @@ -271,7 +280,7 @@ e.g. `docs/tools/hivectl-cli.md` via the `hivectl markdown-docs` subcommand (its `hivectl-docs` flake check otherwise only fails in CI on drift). -## Split asset derivations away from the rust workspace +### Split asset derivations away from the rust workspace `nix/packages/assets.nix` builds the branding SVG/PNG family + claude system-prompt template + claude-settings JSON as its own derivation, @@ -285,7 +294,46 @@ The agent-configs PNG is rendered from the SVG via `rsvg-convert` at build time; librsvg dependency lives here, not in the rust derivation's `nativeBuildInputs`. -## Weston VNC compositor (per-agent `hyperhive.gui.enable`) +### `nix fmt` fails in a git worktree with "object not found" + +`nix fmt` (and any `nix` command that fetches a `git+file://` flake +URL) uses libgit2 internally to compute `revCount` — the number of +commits reachable from HEAD. This walk fails with: + +``` +error: getting Git object '': object not found (libgit2 error code = 9) +``` + +when a commit that was reachable at some earlier evaluation is now gone +(GC'd, rebased away, or pruned). The failure is persistent: clearing +`~/.cache/nix/{eval-cache-v6,gitv3,fetcher-cache-v4.sqlite}` does not +help because the missing object is a structural gap in the git object +graph itself, not in nix's caches. + +**Workaround: use a plain clone, not a git worktree.** + +```bash +git clone http:///hyperhive/hyperhive.git ~/hh-work +cd ~/hh-work && nix fmt +``` + +The root cause is specific to worktrees: a worktree shares the object +store with its parent repo. If the parent repo's history was rewritten +(rebase, force-push, `git gc --prune`) while the worktree was checked +out at a branch tip that references the pruned commits via its reflog or +history, libgit2's rev-walk encounters the gap. A plain clone has its +own self-consistent object store and is immune to the issue. + +## Tooling + +### `hive-forge`: prefer over raw curl pipelines + +Full CLI reference: [`docs/tools/forge.md`](tools/forge.md). +Never use raw `curl` for forge access. + +## GUI (weston/VNC) + +### Weston VNC compositor (per-agent `hyperhive.gui.enable`) `nix/agent-modules/weston-vnc.nix` adds an optional Weston Wayland compositor with the VNC backend, surfaced as @@ -361,7 +409,9 @@ connects to the compositor at `127.0.0.1:`. `wl_event_source_timer_update` treats as "disarm", so the compositor never goes idle and never locks. -## Nix options reference (`nix/docs/default.nix`) +## Nix docs pipeline + +### Nix options reference (`nix/docs/default.nix`) `pkgs.nixosOptionsDoc` over two evaluated module trees: `hostEval` (a stub NixOS system loading the `nix/host-modules/` aggregator with every @@ -397,7 +447,7 @@ options tree picks up everything under that root — picking against stray roots produces an empty tree and renders the host page as template chrome with no `

` headers. -### Docs drv stability: `nixSrc` +#### Docs drv stability: `nixSrc` Naively, the docs evaluation depends on `self` (the flake's store path), so every commit — even Rust-only or frontend-only changes — produces new @@ -426,33 +476,3 @@ Why `builtins.unsafeDiscardStringContext`? The path string make `builtins.path` include `self` as a build dependency even after content-addressing the directory. Discarding the context makes the resulting `nixSrc` truly independent of `self`'s store path. - -### `nix fmt` fails in a git worktree with "object not found" - -`nix fmt` (and any `nix` command that fetches a `git+file://` flake -URL) uses libgit2 internally to compute `revCount` — the number of -commits reachable from HEAD. This walk fails with: - -``` -error: getting Git object '': object not found (libgit2 error code = 9) -``` - -when a commit that was reachable at some earlier evaluation is now gone -(GC'd, rebased away, or pruned). The failure is persistent: clearing -`~/.cache/nix/{eval-cache-v6,gitv3,fetcher-cache-v4.sqlite}` does not -help because the missing object is a structural gap in the git object -graph itself, not in nix's caches. - -**Workaround: use a plain clone, not a git worktree.** - -```bash -git clone http:///hyperhive/hyperhive.git ~/hh-work -cd ~/hh-work && nix fmt -``` - -The root cause is specific to worktrees: a worktree shares the object -store with its parent repo. If the parent repo's history was rewritten -(rebase, force-push, `git gc --prune`) while the worktree was checked -out at a branch tip that references the pruned commits via its reflog or -history, libgit2's rev-walk encounters the gap. A plain clone has its -own self-consistent object store and is immune to the issue.