feat(job-queue): promote the meta-repo deploy window to a queue resource

The two-phase approval deploy keeps a bumped `flake.lock` staged
uncommitted for the whole container build, so no other meta mutation may
land inside that span — until now enforced by a process-global
`meta::exclusive()` mutex held inside each executor fn.

A `MutexGuard` cannot outlive the fn that takes it, which is what blocks
decomposing the opaque `ApprovalDeploy` node into scheduler-visible
sub-nodes: the window has to span them. Replace the mutex with
`Resource::MetaWindow`, a global capacity-1 queue resource declared by
every meta-mutating node kind (`NodeKind::needs_meta_window`). Resources
are held by a subtree root across its whole subtree, so a later increment
can hang the deploy's phases under one window-holding parent.

Same global serialisation as before, and the scheduler now blocks a node
from being claimed rather than parking a worker on a mutex.

Split the rebuild's meta preamble out of `Prebuild` into a new `MetaSync`
node. `Prebuild` must NOT hold the window: the old mutex was deliberately
scoped to drop before the multi-minute toplevel build, which only reads
the store, and a cap-1 global held across it would serialise every
agent's rebuild behind every other's. `MetaSync` is a sibling root that
`Prebuild` deps `AfterOk` on — not its parent, since a parent's resource
covers its whole subtree and would reintroduce exactly that problem.

Queue tests: shape assertions gain the extra node, which is the point of
the change (phases become nodes). The concurrency invariants are intact
but observed one step later — the `MetaSync` heads take turns on the
window, exactly as the runtime mutex made them, so those tests now
complete the heads before asserting that the prebuilds overlap.
This commit is contained in:
atlas 2026-07-25 20:00:29 +02:00 committed by mara
commit dfadacd45f
8 changed files with 311 additions and 148 deletions

View file

@ -22,6 +22,17 @@ pub enum Resource {
/// the rest of that agent's subtree via the crate's recursive lock, so two
/// DAGs never interleave container ops on one agent.
Agent(String),
/// The meta-repo mutation window — a global singleton (default capacity 1)
/// held by any node that mutates the meta repo, so two meta mutations never
/// interleave. Replaces the former runtime `meta::exclusive()` mutex: a
/// `MutexGuard` cannot span scheduler nodes, but a resource held by a
/// subtree root *can* — which is what lets the two-phase deploy
/// (`prepare_deploy` stages `flake.lock` uncommitted across the whole
/// container build, `finalize_deploy`/`abort_deploy` resolve it) be
/// decomposed into sub-nodes instead of one opaque node. Descendants of a
/// holder re-enter it through the crate's recursive lock, exactly like
/// [`Resource::Agent`].
MetaWindow,
}
impl NodeKind {
@ -31,7 +42,14 @@ impl NodeKind {
/// container-affecting kinds ([`NodeKind::needs_lease`]). Lease-exempt
/// container ops (`Start` / `Stop`, fanned out by a lease-holding
/// `Reconcile`) hold no lease of their own — they re-enter the ancestor's
/// `Agent` lock through the crate's recursive re-entrancy.
/// `Agent` lock through the crate's recursive re-entrancy. Meta-mutating
/// kinds ([`NodeKind::needs_meta_window`]) additionally take the global
/// [`Resource::MetaWindow`].
///
/// All of a node's resource edges are acquired **atomically**
/// (`try_acquire_all`) — a node never holds one resource while waiting on
/// another, so the multi-resource kinds (a `MetaLock` wants a build slot
/// *and* the meta window) cannot deadlock against each other.
pub fn resource_deps(&self) -> Vec<Dep<Resource>> {
let mut deps = Vec::new();
if self.needs_build_slot() {
@ -46,6 +64,12 @@ impl NodeKind {
count: 1,
});
}
if self.needs_meta_window() {
deps.push(Dep::Resource {
name: Resource::MetaWindow,
count: 1,
});
}
deps
}
}