Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/scheduler/jobq.md
atlas a88ed9f24e docs: config changes are operator merges on the forge
Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.

Refs #4850
2026-10-02 23:13:04 +02:00

3.1 KiB

The job queue, for operators

Long-running work runs through a job graph. The swarm controller keeps one for swarm-level work — creating an agent's identity, forge user and config repo. Each hive's hive-c0re keeps its own for container operations — rebuild, first-spawn, power changes. This page explains what the job queue is, as a general idea, independent of what either uses it for. For the hive-c0re step catalogue and the engineering internals (scheduler, leases, resource windows) see coordinator.md instead.

What the job queue is, in the abstract

"jobq" is a generic engine for running many interdependent jobs under limited concurrency — it has no idea what a "container" or a "rebuild" is. Two ideas are all there is to it:

  • A job is a small graph of steps, not one opaque blob. Steps can depend on each other (this step only starts once that one finishes), so a big operation is really a short, ordered sequence — not a single black box that's either "done" or "not done."
  • A step can need a shared resource, which only so many steps can hold at once (a "slot"). If every currently running step already holds the slots it needs, a new step that wants the same one waits its turn — that's the whole reason things queue instead of all firing at once.

The engine's whole job is: whenever a step's ordering and resource needs are both satisfied, run it. It has no opinion on what the steps do — that's supplied by whoever builds the graph. The engine is generic: the swarm controller and hive-c0re each build their own graph on it.

Watching it happen

Two views, one per graph, drawn by the same component: the swarm UI's /jobs page shows the swarm controller's graph, and the hive dashboard's BU1LDS page (R3BU1LD QU3U3 — see web-ui/dashboard.md) shows that hive's. Each row is one job; the rows nested under it are that job's steps, in order (occasionally a couple run side by side). A step shows one of:

Glyph Meaning
⏸ queued, waiting its turn
▶ running
◐ its own work is done, waiting on a step nested under it
✔ finished successfully
✖ failed
⊘ cancelled
· skipped (not needed for this run)

A step that isn't needed for a given run stays in the graph as · rather than being absent from it, so the same kind of operation keeps a recognizable shape run to run, whichever steps it actually needed.

⚠️ The view hides ✔ and · by default, so that full shape isn't what you see first — tick them back on in the state filter. The selection is part of the request, not a display toggle: hidden states are never fetched.