hyperhive/hive-priv
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-08-28 18:37:08 +02:00
..
src hive-priv: filter/trim before counting output-path lines, fix log level 2026-08-28 18:37:08 +02:00
Cargo.toml hive-priv: replace json! with typed structs for account sidecar files 2026-08-13 23:16:17 +02:00
README.md hive-priv: filter/trim before counting output-path lines, fix log level 2026-08-28 18:37:08 +02:00

hive-priv

The minimal root privileged-helper for hive-c0re. It runs as root and exposes a narrow unix socket at /run/hive/priv.sock that accepts PrivRequest JSON lines and performs only the handful of operations that genuinely require root — bind-mount edits, nsenter into a container, btrfs subvolume ops. All coordination logic (broker, HTTP, scheduling) stays in the unprivileged hive-c0re process, which delegates here.

Why it exists

Privsep. hive-c0re runs as the unprivileged hive-core user so a bug or a prompt-injection in the large daemon can't directly wield root. The few root operations it needs are funnelled through this small, auditable helper instead. See docs/boundary.md and docs/security.md for the privilege boundary.

Security model

  • Strict allowlist. Every request is validated against a container-name allowlist before any filesystem or process operation — only names matching the hive convention (h-*, the manager container, known sibling service containers) are accepted.
  • No pass-through. Every PrivRequest variant maps to a single known operation; there is no arbitrary-command escape hatch.
  • Socket-activated, always. systemd binds /run/hive/priv.sock (SocketGroup=hive-core, 0660) and passes the listener as fd 3 (LISTEN_FDS); the helper requires this and has no self-bind fallback, so dev and prod take the identical path and the group grant always holds.

The wire contract (PrivRequest / response types) lives in the separate hive-priv-sock crate so this root binary depends on just the protocol shapes, not the whole daemon-shared crate.

Implementation notes

Container toplevel builds (create/update)

container_flake_action (in src/main.rs) builds nixosConfigurations.<name>.config.system.build.toplevel itself (nix_build_toplevel) and passes the resolved store path to nixos-container create/update via --system-path, for both verbs.

Why not let nixos-container build it (the old create-only behaviour, update used its own --flake path): nixos-container's own buildFlake() — invoked whenever --system-path isn't passed — builds to a hardcoded relative path. $systemPath is only ever assigned from the CLI flag or from buildFlake()'s own result, so with no flag it stays undef and "$systemPath.tmp" interpolates to the bare string .tmp in whatever the caller's cwd happens to be. buildFlake() itself takes no lock at all: create wraps its whole action in an exclusive flock before calling it, but update used to call it with no lock whatsoever — so create's lock never protected against a concurrent update clobbering the same .tmp. hive-priv never sets a per-call cwd, so with services.hyperhive.c0re.buildSlots > 1, two concurrent calls (any mix of create/update) could share that one .tmp: one's readlink(".tmp") resolving to the other's build output, handing an agent's container the wrong agent's closure — the "agent container gets closure of other agent" bug (hyperhive#3312).

Building the toplevel here and passing the resolved store path via --system-path for every call means buildFlake() never runs at all, for either verb — no shared .tmp left to race on, no locking invariant of a script we don't own to keep track of. --no-link avoids a competing out-link race of our own; we only need the store path, not a GC root (it's safe from collection for as long as it takes nixos-container to register it against the container's own profile, same window every other --print-out-paths consumer already relies on).

Forwards stderr live, captures stdout silently — deliberately not symmetric. This build is the multi-minute phase of a create/ update, and it used to run inside nixos-container's own --flake invocation, which streams every line to the caller in real time. Buffering it instead (Command::output(), as this function first shipped) regressed that: nothing on the wire — dashboard or journalctl -f alike — until the whole build finishes, then everything at once. So both pipes are drained concurrently (needed to avoid deadlocking if either pipe fills while the other is being read), but only stderr — where nix's own progress goes — is forwarded live, same shape as container_run_streaming. stdout is different: --print-out-paths writes only the final store path there, once, at the end — forwarding it the same way would risk interleaving a progress line into the value this function hands back as --system-path, trading a closure-mixup bug for a corrupted-argument one. So stdout lines are accumulated silently and only consulted after the exit status is known to be success — and even then, exactly one non-empty, trimmed line is required (nix build --print-out-paths prints one line per output, not one line total; config.system.build.toplevel is single-output today, but a bare whole-buffer .trim() would silently hand a multi-line string to --system-path the day that ever changes, and a bare untrimmed/unfiltered .lines() turns a lone "\n" into a bogus empty-string "path" — a wrong line count, or a blank/whitespace one, bail!s instead).