| Filename | Latest commit message | Latest commit date |
|---|---|---|
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290) that pre-created every agent's bind sources. The start preamble already creates them for every c0re-driven start, and on this host only hive-c0re starts agent containers. The file was also the reason the socket dir's owner had to be declared there, which is how it spent its life at `0777 root root` whenever the uid could not be resolved (#4742). - hive-priv gains `EnsureAgentSocketDir { name }`, called from `set_nspawn_flags` in every start path. It creates `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a directory is left alone. hive-c0re's own create_dir_all went: its /run is read-only under ProtectSystem=strict. - The container's `hive-agent-user-migrate` activation chowns that dir to the agent user and sets 0751, the same way it already handles state/ and harness/. It refuses a symlink or non-directory there, since `test -d` and chmod follow links. No host-side passwd parse, and no window where the dir is world-writable. - `/run/hyperhive/agents/<name>` stays created by hive-c0re itself (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re binds as hive-core, so it must not become root- or agent-owned. - The `/run/hive-agent` parent is declared in hive-priv.nix, `0755 root:root`, instead of hive-gateway's hive-core rule. hive-priv is its only writer now, and hive-priv's ReadWritePaths needs it to exist. - The manager start in `ensure_root_agent` now goes through `converge_start_preamble` + `start_with_fallback`. It was a bare start, so after a reboot the manager's bind sources existed only because of the tmpfiles file, and its limits drop-in did not exist at all. - Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`, `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles body builder and their tests, plus the three call sites. - Legacy: hive-priv unlinks the file at every start, ignoring ENOENT. `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that does the same unlink and returns Ok, for an older hive-c0re. Salvaged from #4752: the boundary.md correction that nginx only dials, because ProtectSystem=strict makes its /run read-only. Behaviour change: a manual `nixos-container start h-<name>` right after a reboot, before hive-c0re has started that agent, now fails on a missing bind source instead of starting. Closes #4742 |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
hive-priv
The minimal root privileged-helper for hive-c0re. It runs as root and
exposes a narrow unix socket at /run/hive/priv.sock that accepts PrivRequest
JSON lines and performs only the handful of operations that genuinely require
root — bind-mount edits, nsenter into a container, btrfs subvolume ops. All
coordination logic (broker, HTTP, scheduling) stays in the unprivileged
hive-c0re process, which delegates here.
Why it exists
Privsep. hive-c0re runs as the unprivileged hive-core user so a bug or a
prompt-injection in the large daemon can't directly wield root. The few root
operations it needs are funnelled through this small, auditable helper instead.
See docs/trust-boundary/boundary.md and docs/trust-boundary/security.md for the privilege boundary.
Security model
- Strict allowlist. Every request is validated against a container-name
allowlist before any filesystem or process operation — only names matching the
hive convention (
h-*, the manager container, known sibling service containers) are accepted. - No pass-through. Every
PrivRequestvariant maps to a single known operation; there is no arbitrary-command escape hatch. - Socket-activated, always. systemd binds
/run/hive/priv.sock(SocketGroup=hive-core,0660) and passes the listener as fd 3 (LISTEN_FDS); the helper requires this and has no self-bind fallback, so dev and prod take the identical path and the group grant always holds.
The wire contract (PrivRequest / response types) lives in the separate
hive-priv-sock crate so this root binary depends on just the protocol shapes,
not the whole daemon-shared crate.
Implementation notes
Container toplevel builds (create/update)
container_flake_action (in src/main.rs) builds
nixosConfigurations.<name>.config.system.build.toplevel itself
(nix_build_toplevel) and passes the resolved store path to
nixos-container create/update via --system-path, for both verbs.
Why not let nixos-container build it (the old create-only
behaviour, update used its own --flake path): nixos-container's
own buildFlake() — invoked whenever --system-path isn't passed —
builds to a hardcoded relative path. $systemPath is only ever
assigned from the CLI flag or from buildFlake()'s own result, so with
no flag it stays undef and "$systemPath.tmp" interpolates to the
bare string .tmp in whatever the caller's cwd happens to be.
buildFlake() itself takes no lock at all: create wraps its whole
action in an exclusive flock before calling it, but update used to
call it with no lock whatsoever — so create's lock never protected
against a concurrent update clobbering the same .tmp. hive-priv never
sets a per-call cwd, so with services.hyperhive.c0re.buildSlots > 1,
two concurrent calls (any mix of create/update) could share that one
.tmp: one's readlink(".tmp") resolving to the other's build
output, handing an agent's container the wrong agent's closure — the
"agent container gets closure of other agent" bug.
Building the toplevel here and passing the resolved store path via
--system-path for every call means buildFlake() never runs at all,
for either verb — no shared .tmp left to race on, no locking invariant
of a script we don't own to keep track of. --no-link avoids a
competing out-link race of our own; we only need the store path, not a
GC root (it's safe from collection for as long as it takes
nixos-container to register it against the container's own profile,
same window every other --print-out-paths consumer already relies on).
Forwards stderr live, captures stdout silently — deliberately not
symmetric. This build is the multi-minute phase of a create/
update, and it used to run inside nixos-container's own --flake
invocation, which streams every line to the caller in real time.
Buffering it instead (Command::output(), as this function first
shipped) regressed that: nothing on the wire — dashboard or
journalctl -f alike — until the whole build finishes, then everything
at once. So both pipes are drained concurrently (needed to avoid
deadlocking if either pipe fills while the other is being read), but
only stderr — where nix's own progress goes — is forwarded live, same
shape as container_run_streaming. stdout is different:
--print-out-paths writes only the final store path there, once, at
the end — forwarding it the same way would risk interleaving a progress
line into the value this function hands back as --system-path, trading
a closure-mixup bug for a corrupted-argument one. So stdout lines are
accumulated silently and only consulted after the exit status is known
to be success — and even then, exactly one non-empty, trimmed line is
required (nix build --print-out-paths prints one line per output,
not one line total; config.system.build.toplevel is single-output
today, but a bare whole-buffer .trim() would silently hand a
multi-line string to --system-path the day that ever changes, and a
bare untrimmed/unfiltered .lines() turns a lone "\n" into a bogus
empty-string "path" — a wrong line count, or a blank/whitespace one,
bail!s instead).