hyperhive/docs/network.md
iris 174876094e docs(observability): document OTEL configuration options
Add docs/observability.md covering all services.hyperhive.otel.*
options: enable, endpoint, protocol, headersCredential,
extraResourceAttributes, debug (new in cb0a66147a), and
metricIntervalMs.

Includes:
- Built-in OTEL_RESOURCE_ATTRIBUTES labels (service.name, agent, hive, swarm)
- Cumulative temporality note (avoids Prometheus DELTA drop)
- Network note for host-side collectors on non-standard ports,
  cross-referencing docs/network.md exposeHostPorts

Also:
- CLAUDE.md: add reading-path entry for the new doc
- docs/network.md: link the OTEL mention to observability.md

Closes no issue — gap found during doc sweep.
2026-07-04 13:17:42 +02:00

11 KiB

hive-network

Host-side bridge + per-agent DNS resolver — the foundation that makes container netns isolation safe to land. Configured via services.hyperhive.network.*; off by default during rollout.

Why ship before netns isolation

If netns isolation lands first, agent containers lose /etc/resolv.conf propagation from the host and DNS breaks until a separate resolver is up. Inverting the sequence — bridge + dnsmasq first, netns flip second — makes the flag day boring: the resolver endpoint is already live, agents just discover it via veth instead of shared netns.

v1 vs v2

feature v1 (this PR) v2 (after netns isolation)
bridge interface created on host, no slave NICs per-agent veth pairs attach
dnsmasq binding bridge IP (reachable via host loopback in shared netns) bridge IP (reachable via veth in private netns)
agent container netns shared host private
agent /etc/resolv.conf unchanged (host DNS) nameserver <bridge-ip>
address rules target <bridge-ip> (works in both modes) unchanged from v1

The address rules ship pointing at the bridge IP from v1 so the DNS contract is fixed before any container actually depends on it — minimises the things that flip on netns day.

Container shape (where dnsmasq lives)

Co-located in the existing hive-gateway container — single front-door for both DNS and HTTP, saves a sibling container, single systemd-unit / state surface to monitor. The gateway shares host netns (privateNetwork = false) so dnsmasq's bind-interfaces listener on bridgeIp works without any veth gymnastics today; when agent containers flip to private netns the binding doesn't change (it's still on the host's bridge interface).

Configuration

{
  services.hyperhive = {
    enable = true;
    domain = "darkest.space";
    network.enable = true;            # opt in to bridge + DNS
    network.bridgeIp = "10.42.0.1";   # default
    network.upstreamDns = [            # default Cloudflare + Quad9
      "1.1.1.1"
      "9.9.9.9"
    ];
  };
}

Asserts services.hyperhive.domain != null (resolver needs a domain to be authoritative for) + services.hyperhive.gateway.enable = true (resolver lives in the gateway container).

Bridge addressing

Default subnet is 10.42.0.0/24, host-side gateway at 10.42.0.1. RFC 1918 space, unlikely to clash with operator's existing setup; override bridgeIp + bridgePrefixLength if a different range is already in use. /24 gives 254 usable per-agent addresses — enough for any single-host hive; bigger swarms or tighter addressing schemes pick their own.

Resolver behaviour

dnsmasq is authoritative for the hive's own zones — answers <hive-domain>, forge.<hive-domain>, matrix.<hive-domain> queries with the bridge IP (where nginx is reachable). Everything else gets forwarded to upstreamDns. Containers don't need to know the upstream — they query the bridge IP and dnsmasq does the right thing per-name.

bind-interfaces + interface = [ bridgeName "lo" ] means the listener only accepts queries from the bridge interface (plus lo for container health-checks). External hosts can't reach it — no DNS-amplification surface even when the operator opens port 80 for gateway HTTP.

resolveLocalQueries = false keeps dnsmasq out of the host's own resolution stack — the host's resolver (systemd-resolved, plain glibc nss, dnscrypt-proxy, etc.) keeps doing whatever the operator configured. The hive resolver is purely for inbound queries from agent containers.

Firewall posture

networking.firewall.interfaces.<bridge>.allowedUDPPorts = [ 53 ]

  • allowedTCPPorts = [ 53 ] opens the resolver on the bridge interface only. Other interfaces stay closed. The hive resolver isn't an external-facing service.

When isolateContainers = true, allowedTCPPorts is extended with [ 80 443 ] so isolated agents can reach nginx (gateway container, shared host netns) for the forge sub-domain, per-agent UI proxies, and any other HTTP services.

Reaching host services (exposeHostPorts)

By default agents can only reach the host on 80/443 (+53 DNS), so a host-side service on another port — e.g. a dev OTEL collector for services.hyperhive.otel.endpoint (see docs/observability.md) — is unreachable.

services.hyperhive.network.exposeHostPorts = [ 4318 ]; opens each listed TCP port P on the bridge-interface allowedTCPPorts, so an agent can connect to <bridgeIp>:P (point the collector endpoint at http://<bridgeIp>:4318, default http://10.42.0.1:4318).

This is firewall-only: the host service must bind an address reachable from the bridge — 0.0.0.0 or the bridge IP — not loopback only. The bridge→127.0.0.0/8 DROP rule (below) is unchanged, so a service bound to 127.0.0.1 only stays unreachable; rebind it to 0.0.0.0. (An earlier revision shipped a per-port systemd-socket-proxyd bridge→loopback forwarder, but that collides EADDRINUSE with any collector already bound to 0.0.0.0 — which is the common case — so the proxy was dropped in favour of opening the port.)

The port is reachable by every agent on the bridge subnet (like DNS/gateway), so only expose services safe for any agent to reach.

Container isolation

services.hyperhive.network.isolateContainers (default false) flips agent containers from shared host netns to private netns. Set only after enable = true is stable in production — an assertion blocks the reverse.

What the nix side does when isolateContainers = true

effect mechanism
IP forwarding boot.kernel.sysctl."net.ipv4.ip_forward" = 1
Internet NAT networking.nat { enable = true; internalInterfaces = [ bridgeName ]; } — MASQUERADE on packets leaving via any external NIC
Loopback DROP networking.firewall.extraInputRules — drops bridge-subnet → 127.0.0.0/8 traffic; defence-in-depth against routing table leaks
Gateway access networking.firewall.interfaces.<bridge>.allowedTCPPorts = [ 80 443 ] — lets isolated agents reach nginx on the host (shared netns)
Forge URL HIVE_FORGE_URL flips from http://127.0.0.1:3000 to http://forge.<domain> — agents resolve via dnsmasq, nginx proxies to forgejo
c0re signal HIVE_NETWORK_ISOLATION=1, HIVE_NETWORK_BRIDGE, HIVE_NETWORK_SUBNET in systemd.services.hive-c0re.environment

HIVE_NETWORK_SUBNET is the host-side bridge IP + prefix (e.g. 10.42.0.1/24), not the canonical network address. The Rust side must normalise (bitwise-AND with mask) before subnet membership checks or address arithmetic.

What the Rust side does

hive-c0re reads HIVE_NETWORK_ISOLATION and, when set, passes PRIVATE_NETWORK=1, LOCAL_ADDRESS=<deterministic-ip>, HOST_ADDRESS=<bridge-ip>, and HOST_BRIDGE=<bridgeName> via lifecycle::set_nspawn_flags when creating or updating containers. Each agent gets a deterministic IP derived from its name so the address is reproducible across destroy/recreate. This applies uniformly to all containers — no special case.

HOST_ADDRESS is the bridge gateway IP (the address part of HIVE_NETWORK_SUBNET, via lifecycle::bridge_gateway_ip — taken verbatim so a non-.1 operator override still resolves to wherever the bridge actually lives). It is load-bearing: nixos-container's container-side network setup only installs a default route (ip route add default via $HOST_ADDRESS) when HOST_ADDRESS is non-empty. In bridge mode the host-side address/route setup is skipped, so writing it only affects the container's default route — without it the container comes up with an IP but no path off the bridge subnet (no internet, no api.anthropic.com).

How the isolated container gets its resolver

nixos-container copies the host's /etc/resolv.conf into the container at every start. The host resolver (e.g. 127.0.0.53 from systemd-resolved, or a LAN router) is unreachable from a private netns and isn't authoritative for the hive's own zones, so it must be replaced with the bridge dnsmasq (the gateway IP). Because the copy happens on every start, a declarative environment.etc."resolv.conf" would be clobbered — so the wiring is runtime:

  • hive-priv drops a marker file (/etc/hyperhive-bridge-dns, carrying the gateway IP) into the container's /etc only when isolated, removing it otherwise — so one shared container toplevel behaves correctly in both netns modes.
  • the hyperhive-isolated-dns oneshot (harness-base.nix), gated on that marker, rewrites /etc/resolv.conf to nameserver <gateway-ip> at boot. It is ordered before the harness (hive-ag3nt), the matrix daemon, and tea-login so the resolver is correct before the first DNS lookup; it's an instant no-op in shared-netns mode (the marker is absent, so ConditionPathExists skips it).

Why isolation is safe: all hive-c0re communication goes through unix domain sockets (/run/hive/mcp.sock for agent requests, /run/hive/priv.sock for privileged ops). These are bind-mounted into containers via the nspawn conf. UDS paths traverse the VFS, not the network stack, so PRIVATE_NETWORK=1 does not affect them.

The nix side also enables IP forwarding + NAT (agents reach the internet through the host) and drops bridge-subnet → loopback traffic (defence-in-depth against a compromised agent reaching the c0re dashboard HTTP at 127.0.0.1). Agents have no legitimate reason to reach the dashboard over loopback — the hive-c0re admin socket is a UDS, not TCP.

Prerequisites before flipping on

  • All agents must have hyperhive.web.useUnixSocket = true. Agents that still bind TCP on 0.0.0.0:<port> will be reachable at their bridge IP from other agents on the same subnet — defeating the isolation goal. The gateway routes via unix sockets so gateway reach is unaffected.

Migration behaviour

Containers are destroyed and re-created when the flag flips. Agent state under /agents/<name>/state/ is bind-mounted and survives; the container rootfs is recreated cleanly from the nix store.

Cross-references

  • docs/gateway.md — vhost map + the gateway container's other duties