feat(3088): move the gateway's nginx + dnsmasq onto the host
The gateway's nginx + dnsmasq no longer run in their own nspawn container. `nix/host-modules/hive-gateway/default.nix` loses the `containers.hive-gateway` wrapper and everything that existed only to punch holes in it: `privateNetwork = false`, `CAP_NET_ADMIN`, five bind mounts, its own `stateVersion`, `networking.firewall.enable = false`, `networking.resolvconf.enable = false`, and the `hive-gateway-resolv` path+service pair. 465 -> 303 lines. The container never bought isolation here. It shared the host netns by necessity — nginx binds the host's :80/:443, dnsmasq answers on the bridge — so each of those settings was undoing a boundary the gateway could not afford in the first place. Four things made it more than a deletion, none of them visible in the nix diff: - The self-signed cert service also imports the hive CA leaf, so removing it with the container would have left nginx naming a missing cert file, which it refuses to load at all. - The nginx reload is a hive-priv verb. It still needs root, but no longer for the reason its doc gave, and `--machine=` was both transport and scope — so the unit name is now hard-coded in the helper as the containment. - The lifecycle verb named a container that stops existing. - `journalctl -M hive-gateway` had no machine to enter. Per the operator's ruling, the operator verb keeps working and agents lose it. `InfraContainer` answered three questions that used to share an answer; it now splits into `name()` (identity), `target()` (Container vs HostUnit), `service_unit()` (the systemd unit), and `agent_restartable()`, which the MCP restart path checks before the capability so the refusal cannot read as "ask for infra_admin". `SIBLING_CONTAINERS` drops the gateway — it gates the requests that name a container as a string — while `FromStr` still accepts it, because that answers what a name is, not who may act on it. The dashboard's gateway journal reads host journald filtered to `nginx.service`. Prose was corrected where it only named a location, and re-argued where the container was doing security work: a `0666` per-agent socket was safe because only the gateway container had the directory bind-mounted. There is no mount now, so the directory permissions are the whole of the access control — the constraint holds, its mechanism doesn't. Gate: nix fmt / clippy --all-targets -D warnings / cargo test all clean (710 tests); hivectl-cli.md regenerated from the clap tree. The nix eval was run in both TLS shapes at this commit: every delta in the rendered virtualHosts is one of the three intended path moves, dnsmasq settings are byte-identical, and the absence probe flips true -> false with bindMounts emptied.
This commit is contained in:
parent
cae2cf8df6
commit
07852cabc1
34 changed files with 704 additions and 618 deletions
|
|
@ -31,12 +31,13 @@ pub(super) async fn handle_start(coord: &Arc<Coordinator>, agent: &str, name: &s
|
|||
/// orthogonal: it is gated on the `infra_admin` capability and audited, so it
|
||||
/// stays ahead of the topology guard.
|
||||
pub(super) async fn handle_restart(coord: &Arc<Coordinator>, agent: &str, name: &str) -> Response {
|
||||
// Infra-container restart: an agent holding the `infra_admin`
|
||||
// capability can restart a hive infrastructure container (hive-ci /
|
||||
// hive-gateway / hive-forge / hive-matrix) by passing its name to the
|
||||
// same restart tool. The `InfraContainer` enum parse both recognises
|
||||
// these (never agent children, so disjoint from the child path below)
|
||||
// and yields the typed value the restart path needs.
|
||||
// Infra restart: an agent holding the `infra_admin` capability can
|
||||
// restart a hive infrastructure service (hive-ci / hive-forge /
|
||||
// hive-matrix) by passing its name to the same restart tool. The
|
||||
// `InfraContainer` enum parse both recognises these (never agent
|
||||
// children, so disjoint from the child path below) and yields the typed
|
||||
// value the restart path needs. It recognises `hive-gateway` too, which
|
||||
// is then refused — a name the agent surface knows but may not act on.
|
||||
if let Ok(container) = name.parse::<hive_priv_sock::InfraContainer>() {
|
||||
return handle_restart_infra(coord, agent, container).await;
|
||||
}
|
||||
|
|
@ -60,7 +61,7 @@ async fn handle_restart_infra(
|
|||
agent: &str,
|
||||
container: hive_priv_sock::InfraContainer,
|
||||
) -> Response {
|
||||
let name = container.unit_name();
|
||||
let name = container.name();
|
||||
// Record the attempt in the operator-visible privileged-action audit
|
||||
// trail, then emit a live `AuditEntryAdded` so the dashboard audit view
|
||||
// appends it off `/dashboard/stream`. Best-effort: `record` returns the
|
||||
|
|
@ -75,6 +76,22 @@ async fn handle_restart_infra(
|
|||
coord.emit_audit_entry(entry);
|
||||
}
|
||||
};
|
||||
// Some targets are off-limits to agents regardless of capability — the
|
||||
// gateway, because nginx fronts every hive service from the host and an
|
||||
// agent bouncing it takes out the forge, the dashboard and matrix at
|
||||
// once, including the route its own fix would have to travel. Checked
|
||||
// before the capability so the refusal doesn't read as "ask for
|
||||
// infra_admin"; no capability grants this.
|
||||
if !container.agent_restartable() {
|
||||
tracing::warn!(%agent, %name, "agent: infra restart denied (not agent-restartable)");
|
||||
audit(
|
||||
crate::audit_log::AuditOutcome::Err,
|
||||
Some("denied: target is not agent-restartable"),
|
||||
);
|
||||
return Response::Err {
|
||||
message: format!("`{name}` cannot be restarted by an agent; ask the operator"),
|
||||
};
|
||||
}
|
||||
if !crate::capabilities::has_cap(agent, hive_sh4re::permissions::Capability::InfraAdmin) {
|
||||
tracing::warn!(%agent, %name, "agent: infra restart denied (no infra_admin capability)");
|
||||
audit(
|
||||
|
|
|
|||
Loading…
Reference in a new issue