hyperhive/hive-jobq/README.md
atlas 39b95c2ede treefmt: apply prettier
Pure `nix fmt` output from the commit before this one — no hand edits.
203 files: 52 md, 42 tsx, 32 js, 32 css, 21 ts, 13 html, 8 json, 3 mjs.

Reproduce with `nix develop -c nix fmt` on the parent commit; the result
should be byte-identical to this tree.

None of the 13 `.prettierignore` entries appears here — verified by
intersecting the changed-file list against the ignore file, with a
control proving the intersection finds a match when one exists.
2026-09-02 15:25:07 +02:00

67 lines
3.5 KiB
Markdown

# hive-jobq
A job-DAG scheduler, extracted from hive-c0re's in-tree `job_queue`
as a **domain-agnostic** library. It schedules a single in-memory graph of
nodes over named resources; it knows nothing about containers, rebuilds, or any
hyperhive type — the node payload `N` and resource name `R` are both generic, so
the caller supplies its own domain.
## When to use it
Reach for this crate whenever you need to run a DAG of interdependent work items
under bounded, named concurrency — the hive-c0re rebuild/lifecycle queue is the
first consumer, but nothing here is specific to it. The caller defines the node
kinds, wires deps, and supplies a runner; the scheduler decides what can start.
## Model
**Runtime-only: nothing writes this graph to disk.** The serde impls exist for
the wire projection (`hive-jobq-wire`) and a possible future store; no caller
loads one, so ids and timestamps are stable within a run, not across restarts.
hive-c0re constructs an empty graph every boot and re-derives desired state with
its reconcile sweep.
One **shared graph** for the whole system, not a DAG per job. Enqueuing
inserts a self-contained sub-DAG and returns the ids of the nodes the job
_asked_ for, in the order it named them; the scheduler runs a continuous loop,
starting every node whose deps are satisfied:
- **Resource deps** are named counting semaphores over a caller-chosen type `R`
— e.g. `build-slot` (capacity N), `agent/<name>` (capacity 1), or any
unconfigured name (capacity 1, created on use). A node acquires _all_ its
resource deps atomically at start (all-or-nothing) — no hold-and-wait, so no
deadlock.
- **Node deps** wait on another node per `DepWhen`: `AfterOk` needs success (a
failed dep cancels the dependent), `AfterAny` only needs terminal.
A node carries two independent axes: its `Dep`s (ordering + resource needs) and
its `parent` (structural grouping). The **parent chain**, not the node edges, is
what the scheduler consults for resource re-entrancy: a resource unit is held
for the acquiring node _plus its whole parent subtree_, and a descendant needing
a resource an ancestor already holds re-uses that grant (a re-entrant borrow,
one branch at a time) rather than taking a fresh unit.
A `NodeId` is opaque, stable and monotonic **within a run** — a fresh process
mints ids from zero, so an id stored outside it is a historical record, not a
handle that will resolve later. The scheduler is
single-threaded — it owns the resource table and mutates it directly.
## Shape
- **`Graph<N, R>`** — the in-memory node store. `insert` mints ids and
validates dep/parent references; `set_state` is the single state-transition
choke point (and where each node's lifecycle timestamps —
`started_at` / `finished_at`, `DateTime<Utc>` — are stamped).
- **`Node<N, R>`** — `{ id, parent, payload, deps, state, started_at,
finished_at, error }`. All fields public; derives serde for the wire
projection (and so a store could be added — nothing calls one today).
- **`Scheduler<N, R>`** — drives the graph: `settle()` starts every ready node
(acquiring resources atomically), `complete(id, outcome)` reports a finished
node's result and rolls terminality up the parent chain, releasing grants once
a subtree is done. `Outcome::{Done, Failed(String)}` — the failure reason
rides `Failed` onto the node's `error`.
- **`ResourceTable<R>`** — per-name capacities; unconfigured names default to
capacity 1.
See the crate-root and `scheduler` module `//!` docs for the full borrow/release
model.