Per review: docs represent current state. Every "used to" / "no longer" clause this branch introduced is gone — including the History section in network.md, which was a whole subsection about a sync mechanism that doesn't exist. Where the removed clause was carrying a real constraint, the constraint stays and is stated in the present tense instead of as a delta: nothing narrows what the gateway's nginx can reach except the directory permissions in front of a socket, and nothing bounds `ReloadGatewayNginx` except the hard-coded unit name. Those read as rules now rather than as the story of how they came to be rules.
158 lines
7.8 KiB
Markdown
158 lines
7.8 KiB
Markdown
# The operator/agent boundary
|
|
|
|
Design rationale for hyperhive's two-principal trust model. The
|
|
_implementation_ work — container network isolation, the unifying
|
|
gateway, core-daemon privsep — is tracked as `area:ops` issues on
|
|
the forge.
|
|
|
|
The operator/agent boundary is now technically enforced, not just a
|
|
convention. Containers run in private netns (network isolation is
|
|
always on), the gateway proxies all operator-facing traffic, and
|
|
`hive-c0re` runs as the unprivileged `hive-core` user. All three
|
|
`area:ops` pillars — network isolation, the gateway, and privsep —
|
|
are complete and active.
|
|
|
|
## Two principals, two paths
|
|
|
|
- **Operator** — reaches every UI (the dashboard + every
|
|
per-agent page) through the gateway, on one origin.
|
|
Operator-authority actions (approve / deny, answer-as-operator,
|
|
lifecycle POSTs) are served by the core daemon and only
|
|
reachable via the gateway.
|
|
- **Agent** — speaks only for itself, only over its per-agent
|
|
unix socket. The socket's identity _is_ the agent (see
|
|
`docs/conventions.md`, "identity = socket"). An agent must not
|
|
be able to reach the core daemon's HTTP surface, another
|
|
agent's socket, or another agent's web UI.
|
|
|
|
## Design rule
|
|
|
|
**Operator-authority actions never get a per-agent-socket entry
|
|
point.** They live on the core backend.
|
|
|
|
Worked example — answering an operator-targeted question is a
|
|
`POST /answer-question/{id}` on the core dashboard, _never_ an
|
|
`AgentRequest` variant. If it were a per-agent-socket request, an
|
|
agent could `curl` its own socket and spoof an operator answer.
|
|
The per-agent web UI POSTs cross-origin to the core for these
|
|
(see the inline-answer feature — the loose-ends section on each
|
|
agent page).
|
|
|
|
## Why network isolation is the load-bearing step
|
|
|
|
Without network isolation, containers share the host network namespace
|
|
and can reach `localhost:<core-port>`, the dashboard, and every other
|
|
agent's web port — the operator/agent split is on the honour system and
|
|
every boundary claim above is aspirational. Network isolation is what
|
|
makes the boundary _real_; the gateway and privsep are ergonomics and
|
|
defence-in-depth layered on top.
|
|
|
|
Network isolation is now complete and always on: every agent container
|
|
runs in a private netns behind the hive bridge. The shared-netns mode
|
|
was removed. See `docs/network.md`.
|
|
|
|
Concretely, the core daemon's dashboard `/api` carries **no
|
|
application-layer authentication** — operator-authority routes are served
|
|
unauthenticated at the HTTP layer. Their protection is entirely (a) the
|
|
gateway, which fronts all operator traffic and is where operator auth lives,
|
|
and (b) network isolation, which keeps agents — and `hive-ci`'s untrusted PR
|
|
builds — off host-loopback so nothing can reach `127.0.0.1:<dashboard_port>`
|
|
directly. This is deliberate given the load-bearing role of network isolation
|
|
above, but it is a standing invariant: the `/api` must never be bound to a
|
|
non-loopback address or exposed outside the gateway, and every new
|
|
operator-authority route inherits that assumption. `hive-ci` is treated like an
|
|
agent for this purpose — it runs untrusted PR code and is netns-isolated for
|
|
the same reason.
|
|
|
|
The `area:ops` issues followed this sequencing:
|
|
|
|
1. **Gateway** — pure ergonomics win, unblocks same-origin (lets the
|
|
cross-origin CORS shim on `/answer-question/{id}` go away), no
|
|
behavioural risk. An nginx nixos-container now sits in front of all
|
|
surfaces; per-agent UIs are proxied under `/agent/<name>/`.
|
|
2. **Network isolation** — the load-bearing step that turns the
|
|
honour-system split into an enforced boundary. **Complete** —
|
|
always-on, unconditional; the shared-netns mode was removed.
|
|
3. **Privsep** — defence in depth on the core process; `hive-c0re`
|
|
runs as the unprivileged `hive-core` user and delegates root
|
|
operations to `hive-priv`, a narrow socket-activated helper. See
|
|
[`docs/security.md`](security.md) for the privilege boundary table.
|
|
|
|
### hive-priv socket activation
|
|
|
|
`hive-priv` is **always** socket-activated by the `hive-priv.socket`
|
|
systemd unit. The unit binds `/run/hive/priv.sock` with
|
|
`SocketGroup=hive-core` and mode `0660` and passes the ready listener
|
|
to the helper as fd 3 (`LISTEN_FDS`). The helper requires this and
|
|
bails if it isn't socket-activated — there is intentionally no
|
|
self-bind fallback.
|
|
|
|
Dropping the old fallback removed a dev/prod divergence: when
|
|
`hive-priv` bound the socket itself it created the file owned by
|
|
root's primary group rather than `hive-core`, so a `hive-core` client
|
|
couldn't connect the way the socket unit's `SocketGroup` grant
|
|
intends. Requiring socket activation everywhere means dev and prod
|
|
take the exact same path and the group grant always holds.
|
|
|
|
### the per-agent socket dir
|
|
|
|
`/run/hive-agent/<name>/` is shared by **three principals that share no
|
|
group**, which is why its mode is what it is:
|
|
|
|
| principal | reaches | needs |
|
|
|---|---|---|
|
|
| the agent's harness | binds + unlinks `agent.sock`, `web.sock` | owner, `rwx` |
|
|
| `hive-c0re` | dials `agent.sock` (todo wakes) | traverse |
|
|
| the gateway's nginx | dials `web.sock` | traverse |
|
|
|
|
The last two land in "other", so the dir is **`0751`, owned by the
|
|
agent's container uid/gid** — `o=--x` is traverse without listing, and
|
|
both sockets are `0666`, which is all a dialer needs.
|
|
|
|
**Ownership is declared, not repaired.** The tmpfiles.d entry written by
|
|
`SyncAgentTmpfiles` names the uid/gid directly. Do not add a chown
|
|
alongside it: `d` re-applies on every boot *and* every agent
|
|
spawn/destroy, so ownership set afterwards is reverted the next time any
|
|
agent changes — which is exactly how this dir spent a long time at
|
|
`0777 root root` while a privileged chown appeared to be fixing it.
|
|
|
|
The mode is load-bearing, not cosmetic. Write permission on a
|
|
*directory* is what confers the right to unlink its entries, whoever owns
|
|
them, and the sticky bit is the only thing that would restrain that (it
|
|
is not set here). A world-writable socket dir therefore lets anything
|
|
able to reach the path delete an agent's socket and bind its own — and
|
|
nginx reaches all of `/run/hive-agent` as a plain host path. Dropping
|
|
`o=w` removes that permission rather than qualifying it.
|
|
|
|
⚠️ **The gateway's nginx and dnsmasq are host services, next to
|
|
`hive-c0re`** (see `docs/gateway.md`) — there is no namespace between
|
|
them and the rest of the host. That costs no network isolation: nginx
|
|
binds the host's `:80`/`:443` and reaches `localhost` upstreams, which a
|
|
netns would have to be opened up for anyway.
|
|
🔑 It does mean nothing *implicitly* scopes the privileged reload verb,
|
|
so the scope is explicit: the unit name is hard-coded in `hive-priv` —
|
|
see `PrivRequest::ReloadGatewayNginx`. **A caller cannot name the unit,
|
|
so the verb cannot be steered at another service.**
|
|
|
|
⚠️ Contrast `/shared`, which *is* sticky world-writable (`1777`): it has
|
|
many legitimate writers, so sticky is the best available answer there.
|
|
This dir has exactly one writer, so it needs no world write at all.
|
|
|
|
### host admin socket access (`hivectl`)
|
|
|
|
`hivectl` drives the whole hive — spawn / kill / destroy / rebuild /
|
|
deploy — over the **host admin socket** `/run/hyperhive/host.sock`,
|
|
socket-activated by the `hive-c0re.socket` unit. That socket *is* the
|
|
full-control surface, so who can connect to it is a real trust
|
|
boundary.
|
|
|
|
By default the socket is `0660` group-owned by **`hive-admin`**, an
|
|
empty group — so it is effectively **root-only** until an operator is
|
|
explicitly granted access. Grant sudoless `hivectl` by listing login
|
|
users in `services.hyperhive.c0re.adminUsers`; each is added to
|
|
`hive-admin`, and members connect without `sudo`. The runtime dir
|
|
`/run/hyperhive` is `0751` (traverse-only, no listing) so the group can
|
|
reach the socket path; the socket's own `0660 hive-admin` mode gates
|
|
the connection, and the per-agent subdirs under it keep their own
|
|
restrictive perms. Keep `adminUsers` to trusted operators — membership
|
|
is equivalent to root over the hive.
|