hyperhive/hive-jobq
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas e646656c92 c0re's queue tests no longer drive the scheduler
The last two claim-driven tests were both arranging node states to observe
something that never needed a run:

`settled_dag_leaves_the_snapshot_despite_its_skipped_branch` completed all
seven nodes of a rebuild to assert the DAG left the snapshot. That is one
predicate over a list of states. `shown_on_wire` is it, split out of
`dag_view`, and the cases can now be named rather than arranged — including
the empty set, the one input where "any" and "all" disagree. It takes
states rather than projected nodes so the caller skips projecting what it
is about to discard; a `NodeView` costs a `build_logs` lookup.

`failed_node_cancels_downstream_but_afterany_reconcile_runs` asserted three
unrelated things from one arranged failure: the cascade (hive_jobq's, and
already tested there), the wire filter (now `shown_on_wire`), and the
roll-up. `DagView::rollup_state` lives in hive-host-sock, which had no
tests at all — it does now, next to the invariant, covering the ordering
its own doc comment says has silently disagreed with the frontend before.

With nothing left claiming, `Claimed` / `ClaimReady` / `CompleteNode` /
`claim_one` / `settle_rebuild_tail` are deleted. Claim/complete sites in
`job_queue/tests.rs`: 109 -> 0.

jobq narrows to match: `settle` is gone (it was a `claim_one` loop
returning a Vec, and its only callers were tests — it lives in the test
module now), `claim_one` is private, and `complete_growing` is
`pub(crate)`. `claim_next` is the whole run-loop surface.

`complete` stays `pub` for one caller, noted at the definition: `submit`
completes a group root with no logic of its own so it parks in `Finishing`
and its children unblock. That is a statement about the node, not an event
to report, and it wants to be expressible at insert time.
2026-08-02 22:00:34 +02:00
..
src c0re's queue tests no longer drive the scheduler 2026-08-02 22:00:34 +02:00
Cargo.toml jobq: make NodeGuid an actual guid 2026-08-02 15:32:05 +02:00
README.md jobq: a job asks for the ids it wants back 2026-08-02 15:32:05 +02:00

hive-jobq

A persistent job-DAG scheduler, extracted from hive-c0re's in-tree job_queue as a domain-agnostic library. It schedules a single persistent graph of nodes over named resources; it knows nothing about containers, rebuilds, or any hyperhive type — the node payload N and resource name R are both generic, so the caller supplies its own domain.

When to use it

Reach for this crate whenever you need to run a DAG of interdependent work items under bounded, named concurrency — the hive-c0re rebuild/lifecycle queue is the first consumer, but nothing here is specific to it. The caller defines the node kinds, wires deps, and supplies a runner; the scheduler decides what can start.

Model

One persistent graph for the whole system, not a DAG per job. Enqueuing inserts a self-contained sub-DAG and returns the ids of the nodes the job asked for, in the order it named them; the scheduler runs a continuous loop, starting every node whose deps are satisfied:

  • Resource deps are named counting semaphores over a caller-chosen type R — e.g. build-slot (capacity N), agent/<name> (capacity 1), or any unconfigured name (capacity 1, created on use). A node acquires all its resource deps atomically at start (all-or-nothing) — no hold-and-wait, so no deadlock.
  • Node deps wait on another node per DepWhen: AfterOk needs success (a failed dep cancels the dependent), AfterAny only needs terminal.

A node carries two independent axes: its Deps (ordering + resource needs) and its parent (structural grouping). The parent chain, not the node edges, is what the scheduler consults for resource re-entrancy: a resource unit is held for the acquiring node plus its whole parent subtree, and a descendant needing a resource an ancestor already holds re-uses that grant (a re-entrant borrow, one branch at a time) rather than taking a fresh unit.

A NodeId is opaque, stable, and monotonic (safe to persist). The scheduler is single-threaded — it owns the resource table and mutates it directly.

Shape

  • Graph<N, R> — the persistent node store. insert mints ids and validates dep/parent references; set_state is the single state-transition choke point (and where each node's lifecycle timestamps — started_at / finished_at, DateTime<Utc> — are stamped).
  • Node<N, R>{ id, parent, payload, deps, state, started_at, finished_at, error }. All fields public; derives serde for persistence + the wire.
  • Scheduler<N, R> — drives the graph: settle() starts every ready node (acquiring resources atomically), complete(id, outcome) reports a finished node's result and rolls terminality up the parent chain, releasing grants once a subtree is done. Outcome::{Done, Failed(String)} — the failure reason rides Failed onto the node's error.
  • ResourceTable<R> — per-name capacities; unconfigured names default to capacity 1.

See the crate-root and scheduler module //! docs for the full borrow/release model.