docs(3191): drop the migration history from the gateway comments

Per mara on the PR: how a thing used to be deployed stops being
relevant the moment this is merged and the hives are rebuilt, so the
comments state the current constraint and nothing about the container
that used to hold it.

Removes the two ⚠️ blocks this PR added (hive-tls-ca's ordering, the
resign propagation), the matrix ordering's, dnsmasq's resolv.conf
paragraph, and the pre-existing 'REMOVED WITH THE CONTAINER' inventory
in the gateway module. The 🚨 do-not-simplify warning on the cert-copy
unit keeps both its reasons and loses the container framing.

Kept deliberately: the two operator-facing 'this used to default to
true' option notes (migration information for someone upgrading), and
the SupplementaryGroups block, which documents why an absence is
load-bearing rather than how a deployment used to look.
This commit is contained in:
atlas 2026-08-12 13:26:58 +02:00
commit b12be26f20
4 changed files with 11 additions and 62 deletions

View file

@ -155,32 +155,6 @@ in
"f /var/lib/hive-gateway/conf/gateway.htpasswd 0644 hive-core hive-core - -" "f /var/lib/hive-gateway/conf/gateway.htpasswd 0644 hive-core hive-core - -"
]; ];
# ⚠️ REMOVED WITH THE CONTAINER, and each one was a workaround for the
# boundary rather than a thing nginx or dnsmasq needed:
#
# - `privateNetwork = false` — the container already shared the host
# netns, which is why nginx bound host ports and `localhost`
# upstreams reached hive-c0re. On the host that is simply true.
# - `additionalCapabilities = [ "CAP_NET_ADMIN" ]` — dnsmasq refuses
# to start with a `dhcp-range` unless it holds NET_ADMIN, and
# nspawn's bounding set dropped it for a host-netns container.
# Host root has it.
# - `networking.firewall.enable = false` — a container sharing the
# host netns would run *its* firewall.service against the HOST
# ruleset, flushing nixos-fw and deleting the nixos-nat-* chains on
# every boot. With one machine there is one firewall (below).
# - `networking.resolvconf.enable = false` + the `hive-gateway-resolv`
# path/service pair — the container's /etc/resolv.conf was a
# one-shot copy frozen at start, so a host network change left
# dnsmasq forwarding to a resolver that was gone. The whole
# watch-copy-reload machine existed to bridge two files. There is
# now one.
# - three bind mounts — /run/hive-agent, /run/hive-state, and either
# /run/hive-tls (operator cert) or /run/hive-ca (self-signed);
# those last two are mutually exclusive mkIfs, so it was never
# four. All plain host paths now.
#
# See `docs/network.md::Resolver behaviour` for the resolver history.
# ACME (Let's Encrypt) integration. nginx vhosts set # ACME (Let's Encrypt) integration. nginx vhosts set
# `enableACME = true` via the vhost builder; this provides the # `enableACME = true` via the vhost builder; this provides the
# shared ACME config (acceptTerms + email). # shared ACME config (acceptTerms + email).
@ -191,15 +165,14 @@ in
# Import the hive-CA leaf into nginx's state dir before nginx starts. # Import the hive-CA leaf into nginx's state dir before nginx starts.
# #
# 🚨 THIS LOOKS LIKE A LEFTOVER OF THE CONTAINER AND IS NOT — do not # 🚨 DO NOT "simplify" this into pointing nginx at the CA dir. It
# "simplify" it into pointing nginx at the CA dir. It does TWO jobs: # does TWO jobs:
# #
# (1) It re-modes the leaf. `hive-tls-ca` writes the key 0600 # (1) It re-modes the leaf. `hive-tls-ca` writes the key 0600
# root:root; nginx's pre-start `nginx -t` runs as the nginx # root:root; nginx's pre-start `nginx -t` runs as the nginx
# *user*, so a 0600 key fails the config test with # *user*, so a 0600 key fails the config test with
# `BIO_new_file() … Permission denied` and blocks the unit — # `BIO_new_file() … Permission denied` and blocks the unit —
# hence the 0640 root:nginx copy below. That is a file-mode fact, # hence the 0640 root:nginx copy below.
# not a namespace one, and it did not go away with the container.
# (2) It guarantees that **every cert path the nginx config names # (2) It guarantees that **every cert path the nginx config names
# exists** — which is what the swarm-services fallback at the # exists** — which is what the swarm-services fallback at the
# bottom of the script is for. nginx refuses to load a config # bottom of the script is for. nginx refuses to load a config

View file

@ -69,11 +69,8 @@
dhcp-leasefile = "/var/lib/dnsmasq/dnsmasq.leases"; dhcp-leasefile = "/var/lib/dnsmasq/dnsmasq.leases";
# No explicit upstream: non-hive queries follow dnsmasq's # No explicit upstream: non-hive queries follow dnsmasq's
# resolv.conf default — the host's own `/etc/resolv.conf`, so the # resolv.conf default — the host's own `/etc/resolv.conf`, so the
# hive always uses the host's resolvers and follows them live. This # hive always uses the host's resolvers and follows them live with
# used to be a *copy* nixos-container made at container start, kept # no copy to go stale. Deliberately no fallback
# fresh by a host-side path unit that pushed in a new one and
# reloaded dnsmasq; running on the host deleted both the copy and
# the machinery that watched it. Deliberately no fallback
# `server=`: dnsmasq queries # `server=`: dnsmasq queries
# all known upstreams in parallel, so a hardcoded public resolver # all known upstreams in parallel, so a hardcoded public resolver
# would take a share of *normal* traffic, not just fill in when the # would take a share of *normal* traffic, not just fill in when the

View file

@ -591,14 +591,6 @@ in
# boot requirement. Soft `after` ordering (not `requires`) keeps the # boot requirement. Soft `after` ordering (not `requires`) keeps the
# matrix container's lifecycle decoupled from the resolver's. # matrix container's lifecycle decoupled from the resolver's.
# #
# ⚠️ This named `container@hive-gateway.service` until the gateway
# moved onto the host: dnsmasq lived in that container, so ordering
# after the container was how you ordered after the resolver. The
# container is gone and dnsmasq is a plain host `services.dnsmasq`,
# so the ordering now names the resolver directly — which is what it
# always meant. Naming the *container* was already indirection; it
# just happened to be correct while the container existed.
#
# `mkMerge`, not a bare assignment: `caTrust.containerOrdering` also # `mkMerge`, not a bare assignment: `caTrust.containerOrdering` also
# sets `after`/`requires` (so the bound trust bundle exists before # sets `after`/`requires` (so the bound trust bundle exists before
# nspawn wires the mount up), and two plain assignments to the same # nspawn wires the mount up), and two plain assignments to the same

View file

@ -316,15 +316,8 @@ in
# The consumer is `hive-gateway-self-signed-cert`, which copies the # The consumer is `hive-gateway-self-signed-cert`, which copies the
# leaf into the gateway's state dir at the mode nginx can read, and # leaf into the gateway's state dir at the mode nginx can read, and
# which nginx in turn `Requires=`. So this must run first or that # which nginx in turn `Requires=`. So this must run first or that
# copy fails under `set -eu` and takes nginx down with it. # copy fails under `set -eu` and takes nginx down with it — order
# # against the unit that reads the file.
# ⚠️ This used to name `container@hive-gateway.service` — ordering
# ran through the container, because the copy happened *at container
# start*. Moving the gateway onto the host retired the container
# without retiring the dependency: both units became plain host
# units with nothing sequencing them, and the copy could win the
# race on a fast disk. Order against the unit that reads the file,
# not against the thing that used to host it.
before = [ "hive-gateway-self-signed-cert.service" ]; before = [ "hive-gateway-self-signed-cert.service" ];
requiredBy = [ "hive-gateway-self-signed-cert.service" ]; requiredBy = [ "hive-gateway-self-signed-cert.service" ];
# The issuance below needs the swarm root key on disk, and (for the # The issuance below needs the swarm root key on disk, and (for the
@ -482,16 +475,10 @@ in
# the source therefore changes nothing on its own — the copy has to be # the source therefore changes nothing on its own — the copy has to be
# remade and nginx reloaded, which is what the two calls below do. # remade and nginx reloaded, which is what the two calls below do.
# #
# ⚠️ Both calls used to be `systemctl -M hive-gateway … || true`, from # ⚠️ Neither call is `|| true`: a swallowed propagation failure means
# when nginx lived in a container: `-M` entered the machine, and the # the leaf rotates on disk while nginx keeps serving the old copy
# `|| true` covered a *stopped* container, which was a normal state # until it expires, with this unit reporting success the whole time.
# ("its next boot imports the rotated leaf anyway"). The container is # A failure here must fail the timer.
# gone, so `-M hive-gateway` names a machine nothing creates — both
# calls failed, both failures were swallowed, and the unit logged
# "propagating" and exited 0 while nginx kept serving the stale copy
# until it expired. A host unit failing is NOT a normal state: no
# `-M`, and no `|| true` either, so a broken propagation fails the
# timer loudly instead of reporting success.
systemd.services.hive-tls-resign = { systemd.services.hive-tls-resign = {
description = "Re-sign the gateway TLS leaf and reload nginx"; description = "Re-sign the gateway TLS leaf and reload nginx";
# hive-tls-ca must have run first so the CA key exists before we try # hive-tls-ca must have run first so the CA key exists before we try