| Filename | Latest commit message | Latest commit date |
|---|---|---|
Splits what was one `Cancelled` outcome into two, because they were two
different facts wearing one name:
- `Skipped` — the node's own edges ruled it out. Expected; the failure
branch of a run that succeeded is `Skipped`. A parent's roll-up
**ignores** it.
- `Cancelled` — the work was dropped before it could start. Still
not-success for the roll-up, as before.
Without that split, branching on outcome defeats itself: exactly one
branch is always ruled out, `any_child_failed` counted it, and every DAG
containing a branch would have rolled up failed no matter how the run
went. Caught in review before it was written, not after.
`AFTER_ANY` becomes `{Done, Failed, Skipped}` — "anything except the work
being dropped". That is what it always meant; it only swept in
cancellation because cancellation wasn't distinguishable from
elimination. Audited every user rather than assuming, which is how the
one regression in my own proposal surfaced: `{Done, Failed}` would have
refused to run rebuild's recovery `Reconcile` after a failed `MetaSync`
(that eliminates `Prebuild`, so the tail's dep is `Skipped`, not
`Failed`) and left the container down.
With that, the templates stop computing outcomes and let the graph pick:
- `ResolveApproval { approval_id, outcome }` — one tail per outcome, each
edged to accept only its own, so exactly one is ever runnable.
- `EmitRebuilt { agent, ok }` — a pair. `ok` is not derived, it is which
of the two the graph let run.
Edges are conjunctive, so "any of these roots failed" is not directly
sayable. The composition: the success branch is `AFTER_OK` on every root
(so it is itself eliminated the moment one doesn't succeed), and the
failure branch keys off *that* elimination. The failure branch also
waits on every root — without it, a failed `Prebuild` eliminates the
success branch immediately and the failure would be announced while the
recovery `Reconcile` was still running. The tests caught that one.
Deletes, all of them #2770's host-side debt:
- `Claim.deps`, `DepOutcome`, `Claim::deps_state`, `Claim::deps_error`
and the dep-snapshotting loop in `claim_ready`. Executors read their
own variant now; nothing inspects anything.
- `NodeKind::is_tail()` and the `cancel` exemption built on it. Sparing
is derived from the edges: `cancel` keeps a node iff one of its edges
accepts `Cancelled`. An approval tail names it and survives to resolve
the row; `Reconcile` doesn't and is cancelled with the rest. My earlier
claim that this couldn't dissolve was only true while `AFTER_ANY`
accepted cancellation.
`resolve_approval_dag` / `deploy_terminal_tag` now take `TerminalState`
rather than the wire `State`, so both matches are exhaustive instead of
ending in a catch-all.
Skipped nodes are filtered off the wire alongside `Done` ones. That costs
some dashboard detail on a failed rebuild — which steps were skipped —
and the tests say so with a pointer to the follow-up. Surfacing them as
`Cancelled` instead would be worse: the client roll-up ranks `Cancelled`
above `Running`, so a successful DAG with a not-taken branch would read
as cancelled.
|
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
hive-jobq
A persistent job-DAG scheduler, extracted from hive-c0re's in-tree job_queue
as a domain-agnostic library. It schedules a single persistent graph of
nodes over named resources; it knows nothing about containers, rebuilds, or any
hyperhive type — the node payload N and resource name R are both generic, so
the caller supplies its own domain.
When to use it
Reach for this crate whenever you need to run a DAG of interdependent work items under bounded, named concurrency — the hive-c0re rebuild/lifecycle queue is the first consumer, but nothing here is specific to it. The caller defines the node kinds, wires deps, and supplies a runner; the scheduler decides what can start.
Model
One persistent graph for the whole system, not a DAG per job. Enqueuing inserts a self-contained sub-DAG and returns the new node ids; the scheduler runs a continuous loop, starting every node whose deps are satisfied:
- Resource deps are named counting semaphores over a caller-chosen type
R— e.g.build-slot(capacity N),agent/<name>(capacity 1), or any unconfigured name (capacity 1, created on use). A node acquires all its resource deps atomically at start (all-or-nothing) — no hold-and-wait, so no deadlock. - Node deps wait on another node per
DepWhen:AfterOkneeds success (a failed dep cancels the dependent),AfterAnyonly needs terminal.
A node carries two independent axes: its Deps (ordering + resource needs) and
its parent (structural grouping). The parent chain, not the node edges, is
what the scheduler consults for resource re-entrancy: a resource unit is held
for the acquiring node plus its whole parent subtree, and a descendant needing
a resource an ancestor already holds re-uses that grant (a re-entrant borrow,
one branch at a time) rather than taking a fresh unit.
A NodeId is opaque, stable, and monotonic (safe to persist). The scheduler is
single-threaded — it owns the resource table and mutates it directly.
Shape
Graph<N, R>— the persistent node store.insertmints ids and validates dep/parent references;set_stateis the single state-transition choke point (and where each node's lifecycle timestamps —started_at/finished_at,DateTime<Utc>— are stamped).Node<N, R>—{ id, parent, payload, deps, state, started_at, finished_at, error }. All fields public; derives serde for persistence + the wire.Scheduler<N, R>— drives the graph:settle()starts every ready node (acquiring resources atomically),complete(id, outcome)reports a finished node's result and rolls terminality up the parent chain, releasing grants once a subtree is done.Outcome::{Done, Failed(String)}— the failure reason ridesFailedonto the node'serror.ResourceTable<R>— per-name capacities; unconfigured names default to capacity 1.
See the crate-root and scheduler module //! docs for the full borrow/release
model.