177 lines
8.6 KiB
Markdown
177 lines
8.6 KiB
Markdown
# hive-network
|
|
|
|
Host-side bridge + per-agent private-netns isolation — always on
|
|
whenever hyperhive is enabled. Configured via
|
|
`services.hyperhive.network.*`.
|
|
|
|
> **Historical note:** the bridge and private-netns isolation landed in
|
|
> two separate phases. `services.hyperhive.network.enable` and
|
|
> `services.hyperhive.network.isolateContainers` are retained as
|
|
> deprecated no-op options so existing configs eval without change; both
|
|
> are ignored — isolation is the only mode.
|
|
|
|
## Container shape (where dnsmasq lives)
|
|
|
|
Co-located in the existing `hive-gateway` container — single
|
|
front-door for both DNS and HTTP, saves a sibling container, single
|
|
systemd-unit / state surface to monitor. The gateway shares host
|
|
netns (`privateNetwork = false`) so dnsmasq's `bind-interfaces`
|
|
listener on `bridgeIp` is on the host's bridge interface.
|
|
|
|
## Configuration
|
|
|
|
```nix
|
|
{
|
|
services.hyperhive = {
|
|
enable = true;
|
|
domain = "darkest.space";
|
|
# network.bridgeIp = "10.42.0.1"; # default
|
|
# network.upstreamDns = [ "1.1.1.1" "9.9.9.9" ]; # default
|
|
};
|
|
}
|
|
```
|
|
|
|
Requires `services.hyperhive.domain` to be set — the dnsmasq resolver
|
|
is authoritative for `<hive-domain>` and its sub-domains.
|
|
|
|
## Bridge addressing
|
|
|
|
Default subnet is `10.42.0.0/24`, host-side gateway at `10.42.0.1`.
|
|
RFC 1918 space, unlikely to clash with operator's existing setup;
|
|
override `bridgeIp` + `bridgePrefixLength` if a different range is
|
|
already in use. `/24` gives 254 usable per-agent addresses — enough
|
|
for any single-host hive; bigger swarms or tighter addressing
|
|
schemes pick their own.
|
|
|
|
## Resolver behaviour
|
|
|
|
dnsmasq is **authoritative** for the hive's own zones — answers
|
|
`<hive-domain>`, `forge.<hive-domain>`, `matrix.<hive-domain>`
|
|
queries with the bridge IP (where nginx is reachable). Everything
|
|
else gets forwarded to `upstreamDns`. Containers don't need to know
|
|
the upstream — they query the bridge IP and dnsmasq does the right
|
|
thing per-name.
|
|
|
|
`bind-interfaces` + `interface = [ bridgeName "lo" ]` means the
|
|
listener only accepts queries from the bridge interface (plus lo for
|
|
container health-checks). External hosts can't reach it — no
|
|
DNS-amplification surface even when the operator opens port 80 for
|
|
gateway HTTP.
|
|
|
|
`resolveLocalQueries = false` keeps dnsmasq out of the host's own
|
|
resolution stack — the host's resolver (systemd-resolved, plain
|
|
glibc nss, dnscrypt-proxy, etc.) keeps doing whatever the operator
|
|
configured. The hive resolver is purely for inbound queries from
|
|
agent containers.
|
|
|
|
## Firewall posture
|
|
|
|
`networking.firewall.interfaces.<bridge>.allowedUDPPorts = [ 53 67 ]`
|
|
`networking.firewall.interfaces.<bridge>.allowedTCPPorts = [ 53 80 443 ]`
|
|
|
|
- Port 53 opens the resolver on the bridge interface only. Other
|
|
interfaces stay closed. The hive resolver isn't an external-facing
|
|
service.
|
|
- Port 67 (UDP) admits DHCP requests to the dnsmasq pool. dnsmasq
|
|
receives DHCP via a regular UDP socket (it does not use a
|
|
netfilter-bypassing raw socket), so the hole is mandatory — without
|
|
it containers never get a lease and fall back to 169.254.x.x.
|
|
- Ports 80 and 443 let isolated agents reach nginx (gateway
|
|
container, shared host netns) for the forge sub-domain, per-agent
|
|
UI proxies, and any other HTTP services.
|
|
|
|
### Reaching host services (`exposeHostPorts`)
|
|
|
|
By default agents can only reach the host on 80/443 (+53 DNS), so a
|
|
host-side service on another port — e.g. a dev OTEL collector for
|
|
`services.hyperhive.otel.endpoint` (see `docs/observability.md`) — is unreachable.
|
|
|
|
`services.hyperhive.network.exposeHostPorts = [ 4318 ];` opens each
|
|
listed TCP port `P` on the bridge-interface `allowedTCPPorts`, so an
|
|
agent can connect to `<bridgeIp>:P` (point the collector endpoint at
|
|
`http://<bridgeIp>:4318`, default `http://10.42.0.1:4318`).
|
|
|
|
This is **firewall-only**: the host service must bind an address
|
|
reachable from the bridge — `0.0.0.0` or the bridge IP — not loopback
|
|
only. The bridge→`127.0.0.0/8` DROP rule (below) is unchanged, so a
|
|
service bound to `127.0.0.1` only stays unreachable; rebind it to
|
|
`0.0.0.0`. (An earlier revision shipped a per-port
|
|
`systemd-socket-proxyd` bridge→loopback forwarder, but that collides
|
|
EADDRINUSE with any collector already bound to `0.0.0.0` — which is the
|
|
common case — so the proxy was dropped in favour of opening the port.)
|
|
|
|
The port is reachable by **every** agent on the bridge subnet (like
|
|
DNS/gateway), so only expose services safe for any agent to reach.
|
|
|
|
## Container isolation
|
|
|
|
Each agent container runs in a private network namespace with a dedicated
|
|
veth pair attached to the bridge. The following table summarises what
|
|
the nix side sets up unconditionally:
|
|
|
|
| effect | mechanism |
|
|
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
|
|
| IP forwarding | `boot.kernel.sysctl."net.ipv4.ip_forward" = 1` |
|
|
| Internet NAT | `networking.nat { enable = true; internalInterfaces = [ bridgeName ]; }` — MASQUERADE on packets leaving via any external NIC |
|
|
| Loopback DROP | `networking.firewall.extraInputRules` — drops bridge-subnet → `127.0.0.0/8` traffic; defence-in-depth against routing table leaks |
|
|
| Gateway access | `networking.firewall.interfaces.<bridge>.allowedTCPPorts = [ 80 443 ]` — lets isolated agents (private netns, veth on bridge) reach nginx on the host |
|
|
| c0re signal | `HIVE_NETWORK_ISOLATION=1`, `HIVE_NETWORK_BRIDGE`, `HIVE_NETWORK_SUBNET` in `systemd.services.hive-c0re.environment` |
|
|
|
|
`HIVE_NETWORK_SUBNET` is the host-side bridge IP + prefix (e.g.
|
|
`10.42.0.1/24`), **not** the canonical network address. The Rust side
|
|
must normalise (bitwise-AND with mask) before subnet membership checks or
|
|
address arithmetic.
|
|
|
|
### What the Rust side does
|
|
|
|
`hive-c0re` reads `HIVE_NETWORK_ISOLATION` and passes
|
|
`PRIVATE_NETWORK=1`, `LOCAL_ADDRESS=` (empty), `HOST_ADDRESS=<bridge-ip>`,
|
|
and `HOST_BRIDGE=<bridgeName>` via `lifecycle::set_nspawn_flags` when
|
|
creating or updating containers. `LOCAL_ADDRESS` is left empty so the
|
|
container's dhcpcd acquires an address from the bridge dnsmasq pool
|
|
(`networking.useDHCP = true` in `harness-base.nix`). This applies uniformly
|
|
to all containers — agents and service containers alike.
|
|
|
|
`HOST_ADDRESS` is the bridge gateway IP (the address part of
|
|
`HIVE_NETWORK_SUBNET`, via `lifecycle::bridge_gateway_ip` — taken verbatim
|
|
so a non-`.1` operator override still resolves to wherever the bridge
|
|
actually lives). It is **load-bearing**: nixos-container's container-side
|
|
network setup only installs a default route (`ip route add default via
|
|
$HOST_ADDRESS`) when `HOST_ADDRESS` is non-empty. In bridge mode the
|
|
host-side address/route setup is skipped, so writing it only affects the
|
|
container's default route — without it the container comes up with an IP
|
|
but no path off the bridge subnet (no internet, no `api.anthropic.com`).
|
|
|
|
### How the isolated container gets its resolver
|
|
|
|
nixos-container copies the **host's** `/etc/resolv.conf` into the container
|
|
at every start. The host resolver (e.g. `127.0.0.53` from systemd-resolved,
|
|
or a LAN router) is unreachable from a private netns and isn't
|
|
authoritative for the hive's own zones, so it is replaced with the
|
|
bridge dnsmasq at boot. Because the copy happens on every start, a
|
|
declarative `environment.etc."resolv.conf"` would be clobbered — so the
|
|
wiring is runtime:
|
|
|
|
- `hive-priv` drops a marker file (`/etc/hyperhive-bridge-dns`, carrying the
|
|
gateway IP) into each container's `/etc`.
|
|
- the `hyperhive-isolated-dns` oneshot (harness-base.nix), gated on that
|
|
marker, rewrites `/etc/resolv.conf` to `nameserver <gateway-ip>` at boot.
|
|
It is ordered `before` the harness (`hive-ag3nt`), the matrix daemon, and
|
|
`tea-login` so the resolver is correct before the first DNS lookup.
|
|
|
|
**Why isolation is safe**: all hive-c0re communication goes
|
|
through unix domain sockets (`/run/hive/mcp.sock` for agent requests,
|
|
`/run/hive/priv.sock` for privileged ops).
|
|
These are bind-mounted into containers via the nspawn conf. UDS paths
|
|
traverse the VFS, not the network stack, so `PRIVATE_NETWORK=1` does not
|
|
affect them.
|
|
|
|
The nix side also enables IP forwarding + NAT (agents reach the internet
|
|
through the host) and drops bridge-subnet → loopback traffic (defence-in-depth
|
|
against a compromised agent reaching the c0re dashboard HTTP at
|
|
`127.0.0.1`). Agents have no legitimate reason to reach the dashboard over
|
|
loopback — the hive-c0re admin socket is a UDS, not TCP.
|
|
|
|
## Cross-references
|
|
|
|
- `docs/gateway.md` — vhost map + the gateway container's other duties
|