hyperhive/hive-jobq/README.md
atlas bf138ae79a jobq: a job asks for the ids it wants back
The operator's instruction on the issue was "the closure returns an array
of guids, and enqueue_job returns the node ids in that order". What was
here instead returned a HashMap of everything inserted, and no caller used
the keys: submit dropped the return, insert_group did into_values(), and
the scheduler ignored what append_subgraph handed back. The guid-keyed
lookup was dead weight, and into_values() made that Vec arbitrarily
ordered -- harmless only because nothing read it.

insert_job now takes FnOnce(&JobBuilder) -> Vec<NodeGuid> and returns the
matching ids positionally. A handle from another job is UnknownNode rather
than a silent omission: the return is positional, so a short vector would
misalign every id after it.

c0re's Declare stays FnOnce(&Job) and the wrapper names no handles in one
place, rather than ending seven templates in an empty vector -- a DAG is
addressed by its container node, which submit inserts itself. That frees
insert_group from needing every id, so the node_rt pre-seeding goes too:
NodeRuntime is one Option field and every reader already tolerated a
missing entry (entry().or_default(), get().and_then(), iter().find()).

The tests are the argument for the shape: capturing a handle through a
mutable binding to look it up in the map afterwards collapses into
returning it and destructuring the result.
2026-08-02 15:32:05 +02:00

59 lines
3 KiB
Markdown

# hive-jobq
A persistent job-DAG scheduler, extracted from hive-c0re's in-tree `job_queue`
as a **domain-agnostic** library. It schedules a single persistent graph of
nodes over named resources; it knows nothing about containers, rebuilds, or any
hyperhive type — the node payload `N` and resource name `R` are both generic, so
the caller supplies its own domain.
## When to use it
Reach for this crate whenever you need to run a DAG of interdependent work items
under bounded, named concurrency — the hive-c0re rebuild/lifecycle queue is the
first consumer, but nothing here is specific to it. The caller defines the node
kinds, wires deps, and supplies a runner; the scheduler decides what can start.
## Model
One **persistent graph** for the whole system, not a DAG per job. Enqueuing
inserts a self-contained sub-DAG and returns the ids of the nodes the job
*asked* for, in the order it named them; the scheduler runs a continuous loop,
starting every node whose deps are satisfied:
- **Resource deps** are named counting semaphores over a caller-chosen type `R`
— e.g. `build-slot` (capacity N), `agent/<name>` (capacity 1), or any
unconfigured name (capacity 1, created on use). A node acquires *all* its
resource deps atomically at start (all-or-nothing) — no hold-and-wait, so no
deadlock.
- **Node deps** wait on another node per `DepWhen`: `AfterOk` needs success (a
failed dep cancels the dependent), `AfterAny` only needs terminal.
A node carries two independent axes: its `Dep`s (ordering + resource needs) and
its `parent` (structural grouping). The **parent chain**, not the node edges, is
what the scheduler consults for resource re-entrancy: a resource unit is held
for the acquiring node *plus its whole parent subtree*, and a descendant needing
a resource an ancestor already holds re-uses that grant (a re-entrant borrow,
one branch at a time) rather than taking a fresh unit.
A `NodeId` is opaque, stable, and monotonic (safe to persist). The scheduler is
single-threaded — it owns the resource table and mutates it directly.
## Shape
- **`Graph<N, R>`** — the persistent node store. `insert` mints ids and
validates dep/parent references; `set_state` is the single state-transition
choke point (and where each node's lifecycle timestamps —
`started_at` / `finished_at`, `DateTime<Utc>` — are stamped).
- **`Node<N, R>`** — `{ id, parent, payload, deps, state, started_at,
finished_at, error }`. All fields public; derives serde for persistence + the
wire.
- **`Scheduler<N, R>`** — drives the graph: `settle()` starts every ready node
(acquiring resources atomically), `complete(id, outcome)` reports a finished
node's result and rolls terminality up the parent chain, releasing grants once
a subtree is done. `Outcome::{Done, Failed(String)}` — the failure reason
rides `Failed` onto the node's `error`.
- **`ResourceTable<R>`** — per-name capacities; unconfigured names default to
capacity 1.
See the crate-root and `scheduler` module `//!` docs for the full borrow/release
model.