Seventh batch of the ongoing write-good.Passive pass (hyperhive#4042):
read all 65 hits across jobq.md/ci.md/observability.md/coordinator.md
in context and rewrote 41 with a clearly nameable actor -- mostly
hive-c0re, nix/the nix module, the harness, or a specific fn/type
named right there or nearby (coordinator.md's node-inventory table
and DAG-shape descriptions name concrete Rust items constantly, so
the actor is almost always sitting in the same sentence).
Left 24 alone: predicate-adjective-copula state descriptions ("is
stuck", "is gone", "is unaffected", "is done", etc. -- the largest
recurring bucket this batch, especially in observability.md's
scope/status descriptions), negative-capability idioms ("no X is
needed/left", "X can't be written down"), the established "is
tracked as a follow-up" idiom, a firewall-shorthand notation
("bridge->127.0.0.0/8 is dropped") where rewriting would break the
compact rule-like format, a CLI-flag "(repeatable)" annotation ("May
be repeated"), a Rust type-signature fact ("`moves` is typed ..."),
a hypothetical/counterfactual maintenance-burden clause, a
readiness-condition list ("a node is ready when ... every dep is
satisfied"), and one deliberately-parallel idiom pair
("When OTEL is enabled" used identically twice as a section-opening
convention -- fixing one would break the parallelism, not the
opposite).
One self-caught regression: an early attempt to fix "used by every
`Reconcile` node's start action" (a reduced participial clause, not
flagged) into "is used by every `Reconcile` node's start action"
introduced a brand-new flagged passive. Caught by the post-edit vale
count (expected 65->24, got 65->25) not matching, same discipline as
the docs/turn-loop batch's tail-truncation catch -- re-ran with
active voice instead ("Every `Reconcile` node's start action uses
this fallback").
Verified via vale before/after: 65 -> 24 write-good.Passive hits,
exactly the 24 left alone above; error count and other warning
categories unchanged. Re-read every changed line in full surrounding
context after editing before running the final vale check.
52 lines
2.7 KiB
Markdown
52 lines
2.7 KiB
Markdown
# The job queue, for operators
|
|
|
|
Every container operation — rebuild, first-spawn, a config-PR deploy,
|
|
power changes — runs through one shared job queue. This page explains
|
|
what the job queue _is_, as a general idea, independent of what any one
|
|
subsystem uses it for. For the hive-c0re-specific step catalogue and the
|
|
engineering internals (scheduler, leases, resource windows) see
|
|
[`coordinator.md`](coordinator.md) instead.
|
|
|
|
## What the job queue is, in the abstract
|
|
|
|
"jobq" is a generic engine for running many interdependent jobs under
|
|
limited concurrency — it has no idea what a "container" or a "rebuild" is.
|
|
Two ideas are all there is to it:
|
|
|
|
- **A job is a small graph of steps**, not one opaque blob. Steps can
|
|
depend on each other (this step only starts once that one finishes), so
|
|
a big operation is really a short, ordered sequence — not a single
|
|
black box that's either "done" or "not done."
|
|
- **A step can need a shared resource**, which only so many steps can hold
|
|
at once (a "slot"). If every currently running step already holds the
|
|
slots it needs, a new step that wants the same one waits its turn —
|
|
that's the whole reason things queue instead of all firing at once.
|
|
|
|
The engine's whole job is: whenever a step's ordering and resource needs
|
|
are both satisfied, run it. It has no opinion on what the steps _do_ —
|
|
that's supplied by whoever builds the graph. hive-c0re is the one thing
|
|
building graphs on it today, but nothing about the engine is specific to
|
|
containers or rebuilds; there's nothing stopping another subsystem from
|
|
using the same engine for its own unrelated queue.
|
|
|
|
## Watching it happen
|
|
|
|
Each **row** you see in a queue view (the **BU1LDS** page's R3BU1LD QU3U3
|
|
— see [`web-ui/dashboard.md`](../web-ui/dashboard.md) — and swarm-ui's
|
|
`/jobs` page both render the same underlying graph) is one job; the rows
|
|
nested under it are that job's steps, in order (occasionally a couple run
|
|
side by side). A step shows one of:
|
|
|
|
| Glyph | Meaning |
|
|
| ----- | ------------------------------------------------------- |
|
|
| `⏸` | queued, waiting its turn |
|
|
| `▶` | running |
|
|
| `◐` | its own work is done, waiting on a step nested under it |
|
|
| `✔` | finished successfully |
|
|
| `✖` | failed |
|
|
| `⊘` | cancelled |
|
|
| `·` | skipped (not needed for this run) |
|
|
|
|
A step that isn't needed for a given run shows as `·` rather than
|
|
dropping out of the tree entirely, so the same kind of operation keeps a
|
|
recognizable shape run to run, whichever steps it actually needed.
|