# hive-network Host-side bridge + per-agent DNS resolver — the foundation that makes container netns isolation safe to land. Configured via `services.hyperhive.network.*`; off by default during rollout. ## Why ship before netns isolation If netns isolation lands first, agent containers lose `/etc/resolv.conf` propagation from the host and DNS breaks until a separate resolver is up. Inverting the sequence — bridge + dnsmasq first, netns flip second — makes the flag day boring: the resolver endpoint is already live, agents just discover it via veth instead of shared netns. ## v1 vs v2 | feature | v1 (this PR) | v2 (after netns isolation) | | ------------------------ | ------------------------------------------------------- | ----------------------------------------------- | | bridge interface | created on host, no slave NICs | per-agent veth pairs attach | | dnsmasq binding | bridge IP (reachable via host loopback in shared netns) | bridge IP (reachable via veth in private netns) | | agent container netns | shared host | private | | agent `/etc/resolv.conf` | unchanged (host DNS) | `nameserver ` | | `address` rules target | `` (works in both modes) | unchanged from v1 | The `address` rules ship pointing at the bridge IP from v1 so the DNS contract is fixed before any container actually depends on it — minimises the things that flip on netns day. ## Container shape (where dnsmasq lives) Co-located in the existing `hive-gateway` container — single front-door for both DNS and HTTP, saves a sibling container, single systemd-unit / state surface to monitor. The gateway shares host netns (`privateNetwork = false`) so dnsmasq's `bind-interfaces` listener on `bridgeIp` works without any veth gymnastics today; when agent containers flip to private netns the binding doesn't change (it's still on the host's bridge interface). ## Configuration ```nix { services.hyperhive = { enable = true; domain = "darkest.space"; network.enable = true; # opt in to bridge + DNS network.bridgeIp = "10.42.0.1"; # default network.upstreamDns = [ # default Cloudflare + Quad9 "1.1.1.1" "9.9.9.9" ]; }; } ``` Asserts `services.hyperhive.domain != null` (resolver needs a domain to be authoritative for) + `services.hyperhive.gateway.enable = true` (resolver lives in the gateway container). ## Bridge addressing Default subnet is `10.42.0.0/24`, host-side gateway at `10.42.0.1`. RFC 1918 space, unlikely to clash with operator's existing setup; override `bridgeIp` + `bridgePrefixLength` if a different range is already in use. `/24` gives 254 usable per-agent addresses — enough for any single-host hive; bigger swarms or tighter addressing schemes pick their own. ## Resolver behaviour dnsmasq is **authoritative** for the hive's own zones — answers ``, `forge.`, `matrix.` queries with the bridge IP (where nginx is reachable). Everything else gets forwarded to `upstreamDns`. Containers don't need to know the upstream — they query the bridge IP and dnsmasq does the right thing per-name. `bind-interfaces` + `interface = [ bridgeName "lo" ]` means the listener only accepts queries from the bridge interface (plus lo for container health-checks). External hosts can't reach it — no DNS-amplification surface even when the operator opens port 80 for gateway HTTP. `resolveLocalQueries = false` keeps dnsmasq out of the host's own resolution stack — the host's resolver (systemd-resolved, plain glibc nss, dnscrypt-proxy, etc.) keeps doing whatever the operator configured. The hive resolver is purely for inbound queries from agent containers. ## Firewall posture `networking.firewall.interfaces..allowedUDPPorts = [ 53 ]` - `allowedTCPPorts = [ 53 ]` opens the resolver on the bridge interface only. Other interfaces stay closed. The hive resolver isn't an external-facing service. When `isolateContainers = true`, `allowedTCPPorts` is extended with `[ 80 443 ]` so isolated agents can reach nginx (gateway container, shared host netns) for the forge sub-domain, per-agent UI proxies, and any other HTTP services. ### Reaching host services (`exposeHostPorts`) By default agents can only reach the host on 80/443 (+53 DNS), so a host-side service on another port — e.g. a dev OTEL collector for `services.hyperhive.otel.endpoint` — is unreachable. `services.hyperhive.network.exposeHostPorts = [ 4318 ];` opens each listed TCP port `P` on the bridge-interface `allowedTCPPorts`, so an agent can connect to `:P` (point the collector endpoint at `http://:4318`, default `http://10.42.0.1:4318`). This is **firewall-only**: the host service must bind an address reachable from the bridge — `0.0.0.0` or the bridge IP — not loopback only. The bridge→`127.0.0.0/8` DROP rule (below) is unchanged, so a service bound to `127.0.0.1` only stays unreachable; rebind it to `0.0.0.0`. (An earlier revision shipped a per-port `systemd-socket-proxyd` bridge→loopback forwarder, but that collides EADDRINUSE with any collector already bound to `0.0.0.0` — which is the common case — so the proxy was dropped in favour of opening the port.) The port is reachable by **every** agent on the bridge subnet (like DNS/gateway), so only expose services safe for any agent to reach. ## Container isolation `services.hyperhive.network.isolateContainers` (default `false`) flips agent containers from shared host netns to private netns. Set only after `enable = true` is stable in production — an assertion blocks the reverse. ### What the nix side does when `isolateContainers = true` | effect | mechanism | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | IP forwarding | `boot.kernel.sysctl."net.ipv4.ip_forward" = 1` | | Internet NAT | `networking.nat { enable = true; internalInterfaces = [ bridgeName ]; }` — MASQUERADE on packets leaving via any external NIC | | Loopback DROP | `networking.firewall.extraInputRules` — drops bridge-subnet → `127.0.0.0/8` traffic; defence-in-depth against routing table leaks | | Gateway access | `networking.firewall.interfaces..allowedTCPPorts = [ 80 443 ]` — lets isolated agents reach nginx on the host (shared netns) | | Forge URL | `HIVE_FORGE_URL` flips from `http://127.0.0.1:3000` to `http://forge.` — agents resolve via dnsmasq, nginx proxies to forgejo | | c0re signal | `HIVE_NETWORK_ISOLATION=1`, `HIVE_NETWORK_BRIDGE`, `HIVE_NETWORK_SUBNET` in `systemd.services.hive-c0re.environment` | `HIVE_NETWORK_SUBNET` is the host-side bridge IP + prefix (e.g. `10.42.0.1/24`), **not** the canonical network address. The Rust side must normalise (bitwise-AND with mask) before subnet membership checks or address arithmetic. ### What the Rust side does `hive-c0re` reads `HIVE_NETWORK_ISOLATION` and, when set, passes `PRIVATE_NETWORK=1`, `LOCAL_ADDRESS=`, `HOST_ADDRESS=`, and `HOST_BRIDGE=` via `lifecycle::set_nspawn_flags` when creating or updating containers. Each agent gets a deterministic IP derived from its name so the address is reproducible across destroy/recreate. This applies uniformly to all containers — no special case. `HOST_ADDRESS` is the bridge gateway IP (the address part of `HIVE_NETWORK_SUBNET`, via `lifecycle::bridge_gateway_ip` — taken verbatim so a non-`.1` operator override still resolves to wherever the bridge actually lives). It is **load-bearing**: nixos-container's container-side network setup only installs a default route (`ip route add default via $HOST_ADDRESS`) when `HOST_ADDRESS` is non-empty. In bridge mode the host-side address/route setup is skipped, so writing it only affects the container's default route — without it the container comes up with an IP but no path off the bridge subnet (no internet, no `api.anthropic.com`). ### How the isolated container gets its resolver nixos-container copies the **host's** `/etc/resolv.conf` into the container at every start. The host resolver (e.g. `127.0.0.53` from systemd-resolved, or a LAN router) is unreachable from a private netns and isn't authoritative for the hive's own zones, so it must be replaced with the bridge dnsmasq (the gateway IP). Because the copy happens on every start, a declarative `environment.etc."resolv.conf"` would be clobbered — so the wiring is runtime: - `hive-priv` drops a marker file (`/etc/hyperhive-bridge-dns`, carrying the gateway IP) into the container's `/etc` **only when isolated**, removing it otherwise — so one shared container toplevel behaves correctly in both netns modes. - the `hyperhive-isolated-dns` oneshot (harness-base.nix), gated on that marker, rewrites `/etc/resolv.conf` to `nameserver ` at boot. It is ordered `before` the harness (`hive-ag3nt`), the matrix daemon, and `tea-login` so the resolver is correct before the first DNS lookup; it's an instant no-op in shared-netns mode (the marker is absent, so `ConditionPathExists` skips it). **Why isolation is safe**: all hive-c0re communication goes through unix domain sockets (`/run/hive/mcp.sock` for agent requests, `/run/hive/priv.sock` for privileged ops). These are bind-mounted into containers via the nspawn conf. UDS paths traverse the VFS, not the network stack, so `PRIVATE_NETWORK=1` does not affect them. The nix side also enables IP forwarding + NAT (agents reach the internet through the host) and drops bridge-subnet → loopback traffic (defence-in-depth against a compromised agent reaching the c0re dashboard HTTP at `127.0.0.1`). Agents have no legitimate reason to reach the dashboard over loopback — the hive-c0re admin socket is a UDS, not TCP. ### Prerequisites before flipping on - All agents must have `hyperhive.web.useUnixSocket = true`. Agents that still bind TCP on `0.0.0.0:` will be reachable at their bridge IP from other agents on the same subnet — defeating the isolation goal. The gateway routes via unix sockets so gateway reach is unaffected. ### Migration behaviour Containers are destroyed and re-created when the flag flips. Agent state under `/agents//state/` is bind-mounted and survives; the container rootfs is recreated cleanly from the nix store. ## Cross-references - `docs/gateway.md` — vhost map + the gateway container's other duties