hyperhive/docs/trust-boundary/boundary.md
atlas 2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00

170 lines
8.5 KiB
Markdown

# The operator/agent boundary
<!-- vale write-good.Passive = NO -->
Design rationale for hyperhive's two-principal trust model. The
_implementation_ work — container network isolation, the unifying
gateway, core-daemon privsep — is tracked as `area:ops` issues on
the forge.
<!-- vale write-good.Passive = YES -->
The operator/agent boundary is technically enforced, not just a
convention: containers run in private netns (network isolation is
always on), the gateway proxies all operator-facing traffic, and
`hive-c0re` runs as the unprivileged `hive-core` user.
## Two principals, two paths
- **Operator** — reaches every UI (the dashboard + every
per-agent page) through the gateway, on one origin.
The core daemon serves operator-authority actions (approve / deny,
answer-as-operator, lifecycle POSTs), reachable only via the
gateway.
- **Agent** — speaks only for itself, only over its per-agent
unix socket. The socket's identity _is_ the agent (see
`docs/process/conventions.md`, "identity = socket"). An agent must not
be able to reach the core daemon's HTTP surface, another
agent's socket, or another agent's web UI.
## Design rule
**Operator-authority actions never get a per-agent-socket entry
point.** They live in hive-c0re.
Worked example — destroying or rebuilding a container is a
`POST /api/{destroy,rebuild}/{name}` on the core dashboard, _never_ a
per-agent-socket `Request` variant. If it were a per-agent-socket
request, a compromised agent could `curl` its own socket and destroy
or rebuild itself (or, if the variant took an arbitrary target, another
agent) without ever touching the core's own authenticated surface.
## Why network isolation is the load-bearing step
Without network isolation, containers share the host network namespace
and can reach `localhost:<core-port>`, the dashboard, and every other
agent's web port — the operator/agent split is on the honour system and
every boundary claim above is aspirational. Network isolation is what
makes the boundary _real_; the gateway and privsep are ergonomics and
defence-in-depth layered on top.
Network isolation is complete and always on: every agent container
runs in a private netns behind the hive bridge, and there is no
shared-netns mode. See `docs/networking/network.md`.
Concretely, the core daemon's dashboard `/api` carries **no
application-layer authentication** — the core daemon serves operator-authority
routes unauthenticated at the HTTP layer. Their protection is entirely (a) the
gateway, which fronts all operator traffic and is where operator auth lives,
and (b) network isolation, which keeps agents — and `hive-ci`'s untrusted PR
builds — off host-loopback so nothing can reach `127.0.0.1:<dashboard_port>`
directly. This is deliberate given the load-bearing role of network isolation
above, but it's a standing invariant: the `/api` must never bind to a
non-loopback address or get exposed outside the gateway, and every new
operator-authority route inherits that assumption. hyperhive treats `hive-ci`
like an agent for this purpose — it runs untrusted PR code and is
netns-isolated for the same reason.
The boundary rests on three layers:
1. **Gateway** — fronts all surfaces (dashboard + every per-agent UI)
on one origin. An nginx nixos-container proxies per-agent UIs under
`/agent/<name>/`, which is what lets each agent page's inbox panel
POST `mark-all-read` to the core dashboard's
`/api/agent/{name}/mark-all-read` go same-origin instead of needing
a cross-origin CORS shim. Pure ergonomics — no behavioural risk on
its own.
2. **Network isolation** — the load-bearing layer: every agent
container runs in a private netns behind the hive bridge, always
on and unconditional. This is what turns the operator/agent split
from an honour-system convention into an enforced boundary.
3. **Privsep** — defence in depth on the core process; `hive-c0re`
runs as the unprivileged `hive-core` user and delegates root
operations to `hive-priv`, a narrow socket-activated helper. See
[`docs/trust-boundary/security.md`](security.md) for the privilege boundary table.
### hive-priv socket activation
`hive-priv` is **always** socket-activated by the `hive-priv.socket`
systemd unit. The unit binds `/run/hive/priv.sock` with
`SocketGroup=hive-core` and mode `0660` and passes the ready listener
to the helper as fd 3 (`LISTEN_FDS`). The helper requires this and
bails if it isn't socket-activated.
⚠️ Intentionally, no self-bind fallback exists: if `hive-priv` bound
the socket itself, it would create the file owned by root's primary
group rather than `hive-core`, and a `hive-core` client couldn't
connect the way the socket unit's `SocketGroup` grant intends.
Requiring socket activation everywhere keeps dev and prod on the
exact same path, so the group grant always holds.
### the per-agent socket dir
**Three principals that share no group** reach `/run/hive-agent/<name>/`,
which is why its mode is what it's:
| principal | reaches | needs |
| ------------------- | ---------------------------------------- | ------------ |
| the agent's harness | binds + unlinks `agent.sock`, `web.sock` | owner, `rwx` |
| `hive-c0re` | dials `agent.sock` (todo wakes) | traverse |
| the gateway's nginx | dials `web.sock` | traverse |
The last two land in "other," so the dir is **`0751`, owned by the
agent's container uid/gid** — `o=--x` is traverse without listing, and
both sockets are `0666`, which is all a dialer needs.
<!-- vale write-good.Passive = NO -->
**One mechanism creates it, one sets its owner.** Before every start,
hive-priv creates the dir when missing (`EnsureAgentSocketDir`, `0751
root:root`, no symlink followed) and leaves an existing one alone. The
container's own activation (`hive-agent-user-migrate`) then chowns it to
the agent user and sets `0751`. The container has no user namespace, so
that uid is the host inode's owner. Don't add a host-side chown or chmod:
two owners of one path revert each other. Until the container activates,
the dir is `0751 root`: nothing but root can plant a socket in it, and a
legacy root-run harness can still bind.
<!-- vale write-good.Passive = YES -->
The mode is load-bearing, not cosmetic. Write permission on a
_directory_ is what confers the right to unlink its entries, whoever owns
them, and the sticky bit is the only thing that would restrain that (it
isn't set here). A world-writable socket dir therefore lets anything
able to reach the path delete an agent's socket and bind its own. On the
host, any non-root process whose `/run` is writable can do that, such as
a login session or dnsmasq. nginx only dials: `ProtectSystem=strict` makes its view
of `/run` read-only. Dropping `o=w` removes that permission rather than
qualifying it.
⚠️ **The gateway's nginx and dnsmasq are host services, next to
`hive-c0re`** (see `docs/networking/gateway.md`) — there is no namespace between
them and the rest of the host. That costs no network isolation: nginx
binds the host's `:80`/`:443` and reaches `localhost` upstreams, which
would require opening a netns anyway.
🔑 It does mean nothing _implicitly_ scopes the privileged reload verb —
see [`docs/trust-boundary/security.md`](security.md#hive-c0re-privilege-separation) for
how `PrivRequest::ReloadGatewayNginx`'s containment works.
⚠️ Contrast `/shared`, which _is_ sticky world-writable (`1777`): it has
many legitimate writers, so sticky is the best available answer there.
This dir has exactly one writer, so it needs no world write at all.
### host admin socket access (`hivectl`)
`hivectl` drives the whole hive — spawn / kill / destroy / rebuild /
deploy — over the **host admin socket** `/run/hyperhive/host.sock`,
socket-activated by the `hive-c0re.socket` unit. That socket _is_ the
full-control surface, so who can connect to it's a real trust
boundary.
By default the socket is `0660` group-owned by **`hive-admin`**, an
empty group — so it's effectively **root-only** until an operator is
explicitly granted access. Grant sudoless `hivectl` by listing login
users in `services.hyperhive.c0re.adminUsers`; the option adds each to
`hive-admin`, and members connect without `sudo`. The runtime dir
`/run/hyperhive` is `0751` (traverse-only, no listing) so the group can
reach the socket path; the socket's own `0660 hive-admin` mode gates
the connection, and the per-agent subdirs under it keep their own
restrictive perms. Keep `adminUsers` to trusted operators — membership
is equivalent to root over the hive.