Watch
0
0
Fork
You've already forked hyperhive
0

docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix

gateway.md: split the opener into what/audience/enable; vhost map in two
tables (swarm-service vhosts declared by their own modules, then the hive
vhost) matching vhosts.nix and the service modules; gateway.enable exists
and is set with mkDefault by the modules that need it; Basic auth scope,
dashboard /health/ prefix, error-page rendering, matrix body limit and
forge link source corrected; nginx internals grouped under one Internals
section with their headings unchanged.

network.md: gateway and dnsmasq run on the host, not in a container;
network.enable is set by the modules that need it; shared-netns firewall
rule covers every swarm service container; hive-priv writes the nspawn
conf; domain sentence rewritten; removed options moved into <details>.

jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS
show different graphs drawn by the same component.

observability.md: swarm tier first; history narration cut; network access
deduplicated into a link to network.md; options link made absolute.

matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and
SSO options under deploy.matrix; .well-known is served on the hive domain;
roadmap sentence deleted; stale hive-c0re provisioning claims fixed;
serverName upgrade note moved into <details>.

Refs #3902
This commit is contained in:
atlas 2026-10-01 23:31:46 +02:00
commit 067f4e5699
5 changed files with 433 additions and 633 deletions

View file

@ -1,14 +1,22 @@
# hive-network
Host-side bridge + per-agent private-netns isolation, up on a host
where something attaches to it (`services.hyperhive.network.enable`,
asserted by the modules that need it rather than set by hand).
Configured via `services.hyperhive.network.*`.
The bridge every container on a host attaches to, and the private-netns
isolation of agent and CI containers behind it. It comes up on any host that runs agents or swarm services:
`services.hyperhive.network.enable` defaults to off, and the modules that
need the bridge set it with `mkDefault true` — the hive controller, CI,
the hive collector, and the gateway's resolver, which every swarm service
turns on. Configured via `services.hyperhive.network.*`.
> Isolation is the only mode — there is no shared-netns fallback. The
> former `services.hyperhive.network.isolateContainers` and
> `services.hyperhive.network.upstreamDns` options no longer exist; a
> config that still sets one fails eval with a removal message.
Isolation is the only mode; agent containers never share the host netns.
<details><summary>Upgrading a config that sets isolateContainers or upstreamDns</summary>
Neither option exists. A config that still sets
`services.hyperhive.network.isolateContainers` or
`services.hyperhive.network.upstreamDns` fails eval with a removal
message; drop the line.
</details>
## Network map
@ -45,7 +53,7 @@ that touches it.
| container | netns | IPv4 | listens / reached via |
| -------------- | ----------------------- | -------------- | ----------------------------------------------------------------------------------------------- |
| `hive-gateway` | host (shared) | host addresses | nginx `:80`/`:443` (every vhost); dnsmasq `bridgeIp:53` + DHCP `:67` on the bridge |
| gateway (host) | host | host addresses | nginx `:80`/`:443` (every vhost); dnsmasq `bridgeIp:53` + DHCP `:67` on the bridge |
| `hive-forge` | host (shared) | host addresses | forgejo `:3000` http, `:2222` git-ssh; fronted by the `forge.<swarm-domain>` vhost |
| `hive-matrix` | host (shared) | host addresses | tuwunel `:8008` (+ optional federation port); fronted by the matrix vhost |
| `hive-ci` | private, veth on bridge | DHCP pool | outbound only (runner → forge); no inbound surface |
@ -79,11 +87,9 @@ The flows, end to end:
## Container shape (where dnsmasq lives)
Co-located in the existing `hive-gateway` container — single
front-door for both DNS and HTTP, saves a sibling container, single
systemd-unit / state surface to monitor. The gateway shares host
netns (`privateNetwork = false`) so dnsmasq's `bind-interfaces`
listener on `bridgeIp` is on the host's bridge interface.
dnsmasq runs on the host itself, next to nginx — both come from the
gateway module, one front door for DNS and HTTP. Its `bind-interfaces`
listener sits on `bridgeIp`, on the host's bridge interface.
## Configuration
@ -99,10 +105,11 @@ listener on `bridgeIp` is on the host's bridge interface.
}
```
You must set `services.hyperhive.domain` — the dnsmasq resolver
is authoritative for `<hive-domain>` and its sub-domains. You don't
write it: it's read from this hive's entry in the swarm directory
(`docs/swarm/README.md` § Hive identity config).
The hive domain (`services.hyperhive.domain`, which the resolver answers
for) comes from this hive's entry in the swarm directory:
`swarm.hives.<hiveName>.domain`, default `<hiveName>.<swarm.domain>`.
Eval fails until `swarm.domain` and that entry exist
([`swarm/README.md`](../swarm/README.md) § Hive identity config).
## Bridge addressing
@ -162,13 +169,13 @@ agent containers.
receives DHCP via a regular UDP socket (it doesn't use a
netfilter-bypassing raw socket), so the hole is mandatory — without
it containers never get a lease and fall back to 169.254.x.x.
- Ports 80 and 443 let isolated agents reach nginx (gateway
container, shared host netns) for the forge sub-domain, per-agent
- Ports 80 and 443 let isolated agents reach nginx (on the host) for the forge sub-domain, per-agent
UI proxies, and any other HTTP services.
The **host** firewall is the only firewall. The shared-netns infra
containers (gateway, forge, matrix) set
`networking.firewall.enable = false`: a NixOS firewall inside a
The **host** firewall is the only firewall. The swarm service
containers that share the host netns (forge, matrix, authelia, bao, the
metrics and log stores) run with `networking.firewall.enable = false`
(`nix/container-modules/swarm-container.nix`): a NixOS firewall inside a
shared-netns container runs against the _host_ ruleset — at container
boot its `firewall-start` flushes the `nixos-fw` chains, rebuilds them
from the container's (empty) port list, and deletes the host's
@ -228,15 +235,16 @@ address arithmetic.
`hive-c0re` reads `HIVE_NETWORK_BRIDGE` + `HIVE_NETWORK_SUBNET` and passes
`PRIVATE_NETWORK=1`, `LOCAL_ADDRESS=` (empty), `HOST_ADDRESS=<bridge-ip>`,
and `HOST_BRIDGE=<bridgeName>` via `lifecycle::set_nspawn_flags` when
creating or updating containers. hive-c0re validates both variables **once at
and `HOST_BRIDGE=<bridgeName>` when creating or updating containers:
`lifecycle::set_nspawn_flags` hands them to hive-priv, which rewrites
`/etc/nixos-containers/<container>.conf`. hive-c0re validates both variables **once at
daemon startup**, not per container: they're process-global, so a
missing or malformed value is a misconfigured daemon rather than one bad
container, and failing at boot gives a single diagnostic instead of one
per agent. No non-isolated mode exists to fall back to. hive-c0re leaves `LOCAL_ADDRESS` empty so the
container's dhcpcd acquires an address from the bridge dnsmasq pool
(`networking.useDHCP = true` in `nix/agent-modules/network.nix`). This applies uniformly
to all containers — agents and service containers alike.
(`networking.useDHCP = true` in `nix/agent-modules/network.nix`). This applies to
every agent container, manager included.
`HOST_ADDRESS` is the bridge gateway IP (the address part of
`HIVE_NETWORK_SUBNET`, via `lifecycle::bridge_gateway_ip` — taken verbatim