jobq: delete DagView/NodeView, the second projection of one graph

Two views of the same graph existed: the typed `DagView`/`NodeView`
(`/api/state.rebuild_queue`, the `QueueDag` socket request, and the
`RebuildQueueChanged` payload) and `hive-jobq-wire`'s generic
`GraphNode` (`/api/jobq/graph`, `QueueNodes`). Every consumer has moved
to the generic one, so the typed pair is deleted rather than kept in
agreement with it.

What that removes, beyond the types: the `QueueDag` request and
`HostResponse::dags`; `Queue::snapshot`; `dag_view`, `visible_dags`,
`shown_on_wire`, `dag_finished_at` and `containers`; and the
`rebuild_queue` field on `/api/state`. `RebuildQueueChanged` keeps its
seq and loses its payload — nothing read it, and shipping the graph
both on an event and on an endpoint is the duplication this issue is
about. It stays an event rather than becoming a poll because
push-on-change is what every other live surface here does.

Two behaviours came out simpler for a structural reason. `await_dags`
needed two rules — settled means "gone from the snapshot" *or* "present
with every node terminal" — because the typed view evicted finished
groups; the generic view doesn't, so pending is just "some node isn't
terminal". And `state_of` in the tests no longer derives a roll-up at
all: a group root's own state is the scheduler's answer.

That second one found a bug. `cancelled_dag_still_runs_its_approval
tail` asserted the group reads `Cancelled` while the tail it exists to
protect was still pending — `rollup_state` flattened the surviving
child away and called the group settled. The root reads `Finishing`,
which is what the scheduler documents: own logic done, children still
running. The test now asserts that, with the reasoning inline so it
doesn't get "fixed" back.

Kept: `Source`, `State`, `PermPayload` and the `NodeId` alias in
`hive-host-sock::jobs` — shared vocabulary, still used by hivectl.
This commit is contained in:
atlas 2026-08-03 21:07:58 +02:00 committed by mara
commit f707c60f90
10 changed files with 131 additions and 615 deletions

View file

@ -150,16 +150,6 @@ async fn dispatch(req: &HostRequest, coord: Arc<Coordinator>) -> HostResponse {
HostRequest::Rebuild { name } => {
submit_single(&coord, name.as_str(), Verb::Rebuild).await
}
HostRequest::QueueDag { id } => {
// A multi-step op is one DAG now (no fan-out children to gather).
let dags = coord
.job_queue
.snapshot()
.into_iter()
.filter(|d| d.id == *id)
.collect();
HostResponse::dags(dags)
}
HostRequest::QueueNodes { ids } => {
HostResponse::nodes(coord.job_queue.node_subtrees(ids))
}
@ -901,15 +891,15 @@ async fn handle_stop(
async fn await_dags(coord: &Arc<Coordinator>, ids: &[u64], timeout: std::time::Duration) {
let deadline = std::time::Instant::now() + timeout;
loop {
let snap = coord.job_queue.snapshot();
// A DAG has settled when it's either gone from the snapshot (fully
// `Done` DAGs drop out) or still present but with every node terminal
// (a `Failed`/`Cancelled` DAG lingers). It's pending only while it has
// a non-terminal node.
let pending = ids.iter().any(|id| {
snap.iter()
.any(|d| d.id == *id && d.nodes.iter().any(|n| !n.state.is_terminal()))
});
// Every node under the named roots, terminal ones included — unlike the
// old typed snapshot, a settled group does not drop out of this view.
// So "pending" is simply "some node hasn't finished", with no second
// rule for the disappeared case.
let pending = coord
.job_queue
.node_subtrees(ids)
.iter()
.any(|n| !n.state.is_terminal());
if !pending {
return;
}