hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf

/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
This commit is contained in:
atlas 2026-09-27 04:51:07 +02:00 • committed by mara
commit 2252c55df8
21 changed files with 369 additions and 421 deletions

View file

@ -73,7 +73,7 @@ Cheap — no build slot:
| `PauseDrain` | await the harness reporting `PauseAcknowledged`, bounded timeout; best-effort like `Drain` |
| `DestroyContainer` | `nixos-container destroy` + un-registration (drop from the roster, clear the ephemeral runtime dir). Runs downstream of a `Stop`, so deliberately excluded from `takes_container_down` — the container is already down by the time it claims |
| `PurgeState` | the `purge = true` half of a destroy: delete the agent's state subvolume (via hive-priv) plus its state/applied dirs. Own node because it's conditional and the irreversible step |
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone, resync tmpfiles. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
| `DestroyBookkeeping` | the post-destroy tail — meta sync, fail pending approvals, drop the power intent, notify the manager, rescan, re-emit the tombstone. Same split rationale as `RebuildBookkeeping`/`Swap`. Its `purge` flag only selects the wording of the approval-failure reason and the manager notification — the destructive work is `PurgeState`'s |
| `SetWanted` | write the durable power intent (`wanted = Up`/`Offline`) as the head node of a power-op DAG, replacing the old pre-submit side effect. Takes the agent lease even though it's a store write, so the intent write and the tail `Reconcile` are atomic per-agent — two racing power ops can't clobber each other's intent before either reconciles |
| `FinalizeDeploy` | deploy phase 3 — drop the rollback ref, plant `deployed/<id>`, commit the staged `flake.lock`. The first two git steps are fatal on purpose, so a confirmed-good deploy's outcome and the repo's state can't disagree |
| `ResolveApproval` | tail of an approval-carrying DAG — resolve the approval row from how the work ended (`AfterAny`, one node emitted per outcome). Agentless: the approval row already names its agent |