hyperhive/hive-jobq
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 3ebfed1226 jobq: pin the fairness guarantee where it is made
fifo_fairness_for_the_slot lived in hive-c0re and submitted three
rebuilds, driving one to completion to watch the freed build slot go to
the earlier waiter. The guarantee it was checking is this crate's:
claim_one scans nodes in insertion order and takes the first satisfiable
one. Nothing here tested it -- the property hive-jobq provides was
asserted only downstream, through a host's templates.

a_contended_resource_goes_to_the_oldest_waiter tests it directly.
Mutation-checked: reversing the scan order fails it.

What is hive-c0re's is which nodes contend for the slot at all, and that
is a declaration, so its half is now a declared_resources table with
nothing running. The measured shape corrected an assumption on the way:
MetaSync takes the meta window only, and the agent lease starts at
StopForUpdate -- the first node that touches the container -- not at the
head of the chain. Prebuild deliberately holds no lease, which is what
lets it overlap another DAG on the same agent while the container is
still up.

"Uniform hold across the chain" needs no test of its own: a resource is
held for the acquirer's whole subtree, and the parent nesting is already
asserted in rebuild_chain_is_declared_serial.
2026-08-02 22:00:34 +02:00
..
src jobq: pin the fairness guarantee where it is made 2026-08-02 22:00:34 +02:00
Cargo.toml jobq: make NodeGuid an actual guid 2026-08-02 15:32:05 +02:00
README.md jobq: a job asks for the ids it wants back 2026-08-02 15:32:05 +02:00

hive-jobq

A persistent job-DAG scheduler, extracted from hive-c0re's in-tree job_queue as a domain-agnostic library. It schedules a single persistent graph of nodes over named resources; it knows nothing about containers, rebuilds, or any hyperhive type — the node payload N and resource name R are both generic, so the caller supplies its own domain.

When to use it

Reach for this crate whenever you need to run a DAG of interdependent work items under bounded, named concurrency — the hive-c0re rebuild/lifecycle queue is the first consumer, but nothing here is specific to it. The caller defines the node kinds, wires deps, and supplies a runner; the scheduler decides what can start.

Model

One persistent graph for the whole system, not a DAG per job. Enqueuing inserts a self-contained sub-DAG and returns the ids of the nodes the job asked for, in the order it named them; the scheduler runs a continuous loop, starting every node whose deps are satisfied:

  • Resource deps are named counting semaphores over a caller-chosen type R — e.g. build-slot (capacity N), agent/<name> (capacity 1), or any unconfigured name (capacity 1, created on use). A node acquires all its resource deps atomically at start (all-or-nothing) — no hold-and-wait, so no deadlock.
  • Node deps wait on another node per DepWhen: AfterOk needs success (a failed dep cancels the dependent), AfterAny only needs terminal.

A node carries two independent axes: its Deps (ordering + resource needs) and its parent (structural grouping). The parent chain, not the node edges, is what the scheduler consults for resource re-entrancy: a resource unit is held for the acquiring node plus its whole parent subtree, and a descendant needing a resource an ancestor already holds re-uses that grant (a re-entrant borrow, one branch at a time) rather than taking a fresh unit.

A NodeId is opaque, stable, and monotonic (safe to persist). The scheduler is single-threaded — it owns the resource table and mutates it directly.

Shape

  • Graph<N, R> — the persistent node store. insert mints ids and validates dep/parent references; set_state is the single state-transition choke point (and where each node's lifecycle timestamps — started_at / finished_at, DateTime<Utc> — are stamped).
  • Node<N, R>{ id, parent, payload, deps, state, started_at, finished_at, error }. All fields public; derives serde for persistence + the wire.
  • Scheduler<N, R> — drives the graph: settle() starts every ready node (acquiring resources atomically), complete(id, outcome) reports a finished node's result and rolls terminality up the parent chain, releasing grants once a subtree is done. Outcome::{Done, Failed(String)} — the failure reason rides Failed onto the node's error.
  • ResourceTable<R> — per-name capacities; unconfigured names default to capacity 1.

See the crate-root and scheduler module //! docs for the full borrow/release model.