522 lines
29 KiB
Rust
522 lines
29 KiB
Rust
//! Data model for the generic job-DAG queue: the node kinds — the primitive
|
|
//! operations — and what each one carries. The `State` / `PermPayload` wire
|
|
//! enums live in `hive_host_sock::jobs` (they travel on the host admin
|
|
//! socket) and are re-exported here for the queue's internal use. The graph
|
|
//! itself is served through `hive_jobq_wire`'s generic projection — there is
|
|
//! no second, typed view of it any more.
|
|
//!
|
|
//! **One level, not two.** The node is the unit of everything: scheduling,
|
|
//! execution, build-log, cancel, and the dashboard group (a group root's
|
|
//! subtree *is* the group). A DAG used to be a second level above it, with
|
|
//! its own store and its own id; there is no container node any more, so a
|
|
//! job is exactly the nodes it declared. See `docs/coordinator.md::Job queue`
|
|
//! for the full design.
|
|
|
|
pub use hive_host_sock::jobs::{PermPayload, State};
|
|
use serde::Serialize;
|
|
|
|
use hive_jobq::TerminalState;
|
|
|
|
/// The primitive operations — each kind maps to one executor fn in
|
|
/// `exec.rs`, a thin wrapper over existing `lifecycle.rs` / `meta.rs`
|
|
/// code. Concurrency is gated by two resource classes (see
|
|
/// [`Resource`](super::resource::Resource), declared per node where the node
|
|
/// is constructed rather than derived from its kind); the
|
|
/// meta *repo* is serialized by `meta::META_LOCK` inside the wrapped
|
|
/// functions themselves, which is why there is no `GitCommit` node —
|
|
/// a standalone commit node would open a dirty-working-tree window
|
|
/// between nodes that the fused `meta.rs` ops deliberately close.
|
|
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
|
|
#[serde(tag = "kind", rename_all = "snake_case")]
|
|
pub enum NodeKind {
|
|
/// The rebuild's meta-repo preamble: `lifecycle::prepare_rebuild_dirs`,
|
|
/// an idempotent meta `sync_agents`, and an optional per-agent relock.
|
|
/// `relock = false` only for meta-update cascade rebuilds (re-locking
|
|
/// would revert the bump the cascade just committed).
|
|
///
|
|
/// Its own node — ahead of, and *not* an ancestor of, [`NodeKind::Prebuild`]
|
|
/// — precisely because it is the only part of the rebuild that mutates the
|
|
/// meta repo and so holds the global
|
|
/// [`Resource::MetaWindow`](super::resource::Resource::MetaWindow). Fusing
|
|
/// it into `Prebuild` (or making it `Prebuild`'s parent, which holds a
|
|
/// resource across the whole subtree) would extend that global window over
|
|
/// the multi-minute toplevel build and serialize rebuilds hive-wide.
|
|
/// Store/meta work only — build-slot- and lease-exempt.
|
|
MetaSync { agent: String, relock: bool },
|
|
/// Out-of-band toplevel build while the container keeps serving:
|
|
/// `lifecycle::prebuild_toplevel`, reading a meta repo the upstream
|
|
/// [`NodeKind::MetaSync`] has already synced. The warm build is skipped
|
|
/// when the container is already down — it only exists to shrink the
|
|
/// swap's downtime, which a stopped agent doesn't need (`Swap` builds
|
|
/// inline instead).
|
|
Prebuild { agent: String },
|
|
/// `nixos-container update` profile-swap (requires the container
|
|
/// stopped). Re-applies nspawn flags + resource limits first —
|
|
/// rebuild is the reconcile verb. The post-rebuild bookkeeping tail
|
|
/// lives in the sibling `RebuildBookkeeping` node.
|
|
Swap { agent: String },
|
|
/// The post-`Swap` bookkeeping tail as a first-class node: rev marker,
|
|
/// forge + matrix sync, manager kick, container rescan, meta-inputs
|
|
/// snapshot. Split out of `Swap` for dashboard visibility + retry
|
|
/// granularity. Deps `AfterOk(Swap)`, so it runs only when the profile
|
|
/// swap succeeded; the tail `Reconcile` deps `AfterAny(RebuildBookkeeping)`, so on
|
|
/// swap failure this node is cancel-cascaded (a terminal state) and
|
|
/// recovery still runs. Store/forge/matrix work only — no nix build.
|
|
///
|
|
/// Declares **no resources**: it is a coordinated child of
|
|
/// [`NodeKind::AgentWindow`], which holds the agent lease (and the build
|
|
/// slot) for the whole rebuild subtree. See `templates.rs`'s module doc for
|
|
/// why the brace is the one place a resource is declared on behalf of
|
|
/// others.
|
|
RebuildBookkeeping { agent: String },
|
|
/// First-spawn pre-create provisioning: proposed/applied repos,
|
|
/// state subvolume, and meta registration (`sync_agents`). Runs
|
|
/// ahead of `Create` so the `nixos-container create --flake
|
|
/// meta#<name>` ref resolves. Store/meta-only — no container yet —
|
|
/// so it's lease- and build-slot-exempt like `Prebuild`.
|
|
Provision { agent: String },
|
|
/// First-spawn `nixos-container create` proper. Assumes the
|
|
/// upstream `Provision` node already registered the agent in meta.
|
|
Create { agent: String },
|
|
/// `nixos-container destroy` plus the un-registration that follows it:
|
|
/// drop the agent from the coordinator's roster and clear its ephemeral
|
|
/// runtime dir.
|
|
///
|
|
/// **Deliberately not in [`NodeKind::takes_container_down`]**, and that
|
|
/// is the design rather than an oversight. This node runs *downstream of
|
|
/// a `Stop`*, which already carries the flag honestly, so by the time it
|
|
/// claims there is nothing left to take down. A container still live here
|
|
/// is a real bug and must page someone — a `true` would absorb exactly
|
|
/// that signal, and the flag's whole asymmetry (see that method) is that
|
|
/// a wrong `true` silently swallows a crash.
|
|
DestroyContainer { agent: String },
|
|
/// The `purge = true` half of a destroy: delete the agent's state
|
|
/// subvolume (via hive-priv, since a subvolume root defeats
|
|
/// `remove_dir_all`) plus its state and applied dirs. Its own node
|
|
/// because it is conditional — a plain destroy never inserts it — and
|
|
/// because it is the irreversible step, so it earns a distinct row in
|
|
/// the graph rather than hiding inside a bookkeeping tail.
|
|
PurgeState { agent: String },
|
|
/// The post-destroy bookkeeping tail: meta sync, fail the agent's pending
|
|
/// approvals, drop the durable power intent, notify the manager, rescan
|
|
/// containers, re-emit the tombstone + schedule snapshots, resync
|
|
/// tmpfiles. Split from [`NodeKind::DestroyContainer`] for the same
|
|
/// reason [`NodeKind::RebuildBookkeeping`] is split from `Swap`:
|
|
/// dashboard visibility and retry granularity for work that is pure
|
|
/// store/meta bookkeeping and touches no container.
|
|
///
|
|
/// `purge` only selects the wording of the approval-failure reason and
|
|
/// the manager notification; the destructive work is `PurgeState`'s.
|
|
DestroyBookkeeping { agent: String, purge: bool },
|
|
/// Meta flake lock bump. `sweep = false`: `meta::lock_update`
|
|
/// (commit fused, under `META_LOCK`) with this node's own `inputs`;
|
|
/// `sweep = true`: `meta::lock_update_hyperhive`, *non-fatal* (a
|
|
/// failed boot-time bump must not cancel the fan-out rebuilds).
|
|
/// On success the scheduler appends child `Rebuild` DAGs: the
|
|
/// precomputed `fanout` list when present (boot sweep), else the
|
|
/// post-bump affected set (`meta_update_cascade_agents`).
|
|
///
|
|
/// `inputs` are the flake inputs to bump — empty means "all", and the
|
|
/// boot sweep leaves them empty since it bumps `hyperhive` alone. They
|
|
/// ride this node because it is the only thing that reads them.
|
|
MetaLock {
|
|
sweep: bool,
|
|
fanout: Option<Vec<String>>,
|
|
inputs: Vec<String>,
|
|
},
|
|
/// Idempotent power converge *planner*: read `wanted` + observed
|
|
/// state and decide the action (start if `Up` & down, stop if
|
|
/// `Offline` & up, else noop). The mechanical work is not done in
|
|
/// this node — it fans a child [`NodeKind::Start`] / [`NodeKind::Stop`]
|
|
/// DAG out at runtime so the sub-step is a first-class DAG node.
|
|
Reconcile { agent: String },
|
|
/// Mechanical container start: the start preamble (runtime dir +
|
|
/// drop-ins), `start_with_fallback`, MCP listener registration, and
|
|
/// the manager kick. Fanned out by a [`NodeKind::Reconcile`] that
|
|
/// observed `wanted = Up` and the container down.
|
|
Start { agent: String },
|
|
/// Mechanical container stop: `nixos-container` kill, MCP listener
|
|
/// unregister, and the `Killed` manager notify. Fanned out by a
|
|
/// [`NodeKind::Reconcile`] that observed `wanted = Offline` and up.
|
|
Stop { agent: String },
|
|
/// Mechanical `nixos-container stop` for the profile swap. Never
|
|
/// touches `wanted`. Noop if already stopped.
|
|
StopForUpdate { agent: String },
|
|
/// Set the graceful-stop fence + kick the harness so it runs one
|
|
/// stop-checkpoint turn.
|
|
Signal { agent: String },
|
|
/// Await the harness clearing the fence, bounded by
|
|
/// `GRACEFUL_STOP_TIMEOUT`. Resolves ok either way — the
|
|
/// downstream `Reconcile` performs the actual stop.
|
|
Drain { agent: String },
|
|
/// Write the pause marker (`Coordinator::set_paused`) + mark
|
|
/// `pause_pending`. No kick, unlike `Signal` — the harness's own
|
|
/// between-turns poll (`PAUSE_POLL`, 1s default) is already
|
|
/// responsive enough, and `Signal`'s kick-message body ("you were
|
|
/// just (re)started") would be actively misleading here.
|
|
PauseSignal { agent: String },
|
|
/// Await the harness reporting `PauseAcknowledged`, bounded by a
|
|
/// timeout. Resolves ok either way — pausing is best-effort from
|
|
/// the queue's perspective, same as `Drain`.
|
|
PauseDrain { agent: String },
|
|
/// `set_nspawn_flags` + `set_resource_limits` + daemon-reload.
|
|
WriteDropin { agent: String },
|
|
/// Commit `tool-groups.json` / `capabilities.json` per its `payload`
|
|
/// (commit fused under `META_LOCK`). The payload rides this node — the only
|
|
/// consumer — rather than the generic DAG container.
|
|
WritePermFile { agent: String, payload: PermPayload },
|
|
/// Topology move(s) — `set-parent` (len 1) or `set-parent-bulk` (len N) —
|
|
/// as a single queue node. Agentless like [`NodeKind::MetaLock`]: a
|
|
/// reparent touches the meta repo, not any one container, and a bulk
|
|
/// move spans multiple agents anyway. Declares the meta window, same
|
|
/// precedent as [`NodeKind::WritePermFile`] (also a small
|
|
/// git-commit-under-`META_LOCK` op) — a reparent's commit must not land
|
|
/// inside another node's staged deploy `prepare_deploy`→`finalize_deploy`
|
|
/// window. `(child, new_parent)` pairs, applied in order under one
|
|
/// `META_LOCK` acquisition / one git commit (`meta::bulk_commit_topology`
|
|
/// handles both the single- and multi-move case uniformly).
|
|
Reparent {
|
|
moves: Vec<(hive_types::Ident, Option<hive_types::Ident>)>,
|
|
},
|
|
/// Group root of the approval-deploy (`MergeConfigPr`) subtree, and the
|
|
/// node that **owns the deploy window**. It performs no work of its own —
|
|
/// it exists so the resources it declares (the global
|
|
/// [`Resource::MetaWindow`](super::resource::Resource::MetaWindow), the
|
|
/// agent lease, a build slot) are held continuously across every child
|
|
/// phase, which a per-node acquisition could not guarantee.
|
|
///
|
|
/// All three resources are declared *here*, on one node, on purpose. The
|
|
/// queue acquires a node's resources atomically (all-or-nothing), so a
|
|
/// single multi-resource root can never hold one and block on another —
|
|
/// whereas letting a child take the build slot while its parent held the
|
|
/// meta window would introduce exactly that pattern, and with it a
|
|
/// lock-ordering argument that has to be re-verified on every future edit.
|
|
/// Cheap, too: the window has to span the container build regardless (see
|
|
/// [`NodeKind::DeployApply`]), so nothing is over-serialised by hoisting
|
|
/// the slot and the lease up alongside it.
|
|
///
|
|
/// Carries the approval row every phase below it re-reads, like each of
|
|
/// those phases does — the id is the node's own payload, not something a
|
|
/// DAG-level catch-all hands down.
|
|
DeployWindow { agent: String, approval_id: i64 },
|
|
/// Group root of a rebuild subtree — the **brace**: declares the agent
|
|
/// lease and the build slot, holds both for the whole subtree, and performs
|
|
/// no work of its own. Same pure-resource-holder shape as
|
|
/// [`NodeKind::DeployWindow`], scoped to one agent.
|
|
///
|
|
/// Its children declare no resources and borrow these grants, which is what
|
|
/// lets `Prebuild` run beside the `Signal` → `Drain` window. Why braces
|
|
/// exist and what they cost: `docs/coordinator.md`, _Braces_.
|
|
AgentWindow { agent: String },
|
|
/// Deploy phase 1 — **verify only, mutates nothing.** Drift-gate the
|
|
/// approval's PR head, fetch it into the applied repo, and eval-verify the
|
|
/// merge head. Any failure here aborts the deploy with the forge state
|
|
/// untouched, so it is safely retryable and cancel-safe: nothing downstream
|
|
/// has happened yet.
|
|
MergeVerify { agent: String, approval_id: i64 },
|
|
/// Deploy phase 2 — the irreversible fast-forward plus the *opening* half of
|
|
/// the two-phase meta deploy: park the rollback ref, ff-merge the reviewed
|
|
/// head to `main` via the forge API, ff `applied/main`, and
|
|
/// `meta::prepare_deploy` (which stages `flake.lock` uncommitted).
|
|
///
|
|
/// It does **not** run the container rebuild itself. It grows the ordinary
|
|
/// rebuild subgraph into this DAG as its own children
|
|
/// ([`super::templates::deploy_rebuild_nodes`], `relock = false` — the lock
|
|
/// is already staged), so the multi-minute build renders as the same real
|
|
/// nodes every other rebuild does instead of one opaque box. Closing the
|
|
/// staged-lock window is likewise its own node
|
|
/// ([`NodeKind::FinalizeDeploy`]), and the compensation path is
|
|
/// [`NodeKind::DeployTail`].
|
|
DeployApply { agent: String, approval_id: i64 },
|
|
/// Deploy phase 3 — close the two-phase meta deploy once the rebuild
|
|
/// subgraph under [`NodeKind::DeployApply`] has come up clean: drop the
|
|
/// rollback ref, plant the `deployed/<id>` tag, commit the staged
|
|
/// `flake.lock` (`meta::finalize_deploy`).
|
|
///
|
|
/// Its two git steps are **fatal**, deliberately. They are the writes that
|
|
/// tell [`NodeKind::DeployTail`] a deploy confirmed good, so a node that
|
|
/// merely warned on them could report success while leaving the tail
|
|
/// looking at the git state of a failure — and the tail would then roll a
|
|
/// *good* deploy back. Failing loudly keeps the node's outcome and the
|
|
/// repo's state saying the same thing.
|
|
///
|
|
/// The trailing `meta::finalize_deploy` stays warn-only: by then the
|
|
/// container already runs the new config, and an uncommitted staged lock is
|
|
/// something the operator can commit by hand.
|
|
FinalizeDeploy { agent: String, approval_id: i64 },
|
|
/// Deploy compensation **and bookkeeping** tail — `AfterAny`
|
|
/// [`NodeKind::DeployApply`], so it runs on success, failure, and cancel
|
|
/// alike, in the same spirit as the rebuild template's tail `Reconcile`
|
|
/// ("always runs, decides internally"). It:
|
|
/// 1. compensates a merge that landed but was never finalized — roll `main`
|
|
/// back, reset the tree, `meta::abort_deploy`, plant `failed/<id>`;
|
|
/// 2. mirrors whichever deploy tag got planted to the forge config repo
|
|
/// (`forge::push_config`), always, best-effort;
|
|
/// 3. posts the failing build log back onto the config PR when the deploy
|
|
/// failed, so the manager sees the rejection without leaving the forge.
|
|
///
|
|
/// Steps 2 and 3 are why this is `DeployTail` and not `AbortDeploy`: it has
|
|
/// work to do on the success path too, and a node name that claims
|
|
/// otherwise would be a lie on the dashboard.
|
|
///
|
|
/// For (1) it needs no knowledge of how far the deploy got, because that state is
|
|
/// parked in the applied repo rather than passed between nodes:
|
|
/// `DeployApply` writes the pre-merge `main` sha to
|
|
/// `refs/hyperhive/rollback/<approval-id>` before the fast-forward and
|
|
/// [`NodeKind::FinalizeDeploy`] deletes it. So the ref existing *is*
|
|
/// the "a merge landed and was not finalized" signal, and its absence makes
|
|
/// this node a no-op. Parking it in git rather than in a node payload also
|
|
/// means it survives a `hive-c0re` restart mid-deploy, which an in-memory
|
|
/// queue does not.
|
|
DeployTail { agent: String, approval_id: i64 },
|
|
/// Tail node of an approval-carrying DAG (spawn / opaque deploy / config-PR
|
|
/// merge): resolve the approval row from how the work actually ended.
|
|
///
|
|
/// Weak-edged (`DepWhen::AFTER_ANY`) like [`NodeKind::DeployTail`], so it runs on
|
|
/// success, failure **and cancel** alike and decides internally. It reads its
|
|
/// dependencies' terminal states from the node's dependency list rather than
|
|
/// re-deriving them from the world the way `DeployTail` reads git: a node is
|
|
/// *told* how the work it follows ended, it does not go back out and ask.
|
|
///
|
|
/// Agentless on purpose: the approval row already names its agent, so
|
|
/// carrying one here would be a second copy free to drift. Like
|
|
/// [`NodeKind::MetaLock`] it reports `""` from [`NodeKind::agent`] and takes
|
|
/// no lease — which is also what lets one close a multi-agent DAG.
|
|
ResolveApproval {
|
|
approval_id: i64,
|
|
/// Which outcome this node reports. A template emits **one per outcome**,
|
|
/// each edged to accept only that one, so exactly one is ever runnable
|
|
/// and the executor has nothing to decide — it resolves the row the way
|
|
/// its own variant says. The `Cancelled` one is also the node that
|
|
/// [`super::JobQueue::cancel`] spares, since its edge is the only one
|
|
/// that accepts a dropped dependency.
|
|
outcome: TerminalState,
|
|
},
|
|
/// Tail node of a rebuild / perm-change: emit this agent's `Rebuilt` manager
|
|
/// event — `ok` when its deps are `Done`, `!ok` carrying the failure note when
|
|
/// they `Failed`, and **nothing at all** when they `Cancelled` (a cancelled DAG
|
|
/// never ran, so there is no rebuild to report).
|
|
///
|
|
/// One node per **agent** — a multi-agent DAG reports each agent's own
|
|
/// outcome rather than painting all of them with the whole DAG's roll-up —
|
|
/// and one per **outcome**: `ok` isn't computed here, it's which of the pair
|
|
/// the graph let run.
|
|
///
|
|
/// No cancel variant, deliberately: a DAG dropped before it started has no
|
|
/// rebuild to report, and neither tail's edge accepts `Cancelled`, so both
|
|
/// are cancelled with the rest and nothing is emitted.
|
|
EmitRebuilt { agent: String, ok: bool },
|
|
/// Write the agent's durable power intent (`wanted = Up` when `up`, else
|
|
/// `Offline`) as a first-class DAG node, at the head of a power-op
|
|
/// template so the downstream `Reconcile` reads it. Replaces the old
|
|
/// pre-submit `set_wanted` side effect: the intent write is now part of
|
|
/// the atomic DAG (crash-safe, per-agent — a multi-agent DAG carries one
|
|
/// `SetWanted` per agent). Build-slot-exempt (a store write), but
|
|
/// **lease-needing**: it takes the agent's lifecycle lease so the whole
|
|
/// power-op DAG (intent write → reconcile) is atomic per-agent — two
|
|
/// racing ops (e.g. restart vs stop) can't clobber each other's intent
|
|
/// before either reconciles, which is the point of moving the write into
|
|
/// the DAG. (In `stale_start` the lease is thus held across the head
|
|
/// `Prebuild`, but that's a no-op there — the agent is down, so prebuild
|
|
/// is skipped.)
|
|
SetWanted { agent: String, up: bool },
|
|
/// One-shot boot-time forge user/token sweep for every existing
|
|
/// container (`forge::ensure_all`) — moved out of a bare `tokio::spawn`
|
|
/// so it shows as real work on the dashboard instead of running
|
|
/// invisibly until it fails. Agentless: it sweeps every container, not
|
|
/// one. Store/network-I/O only — build-slot- and lease-exempt, same
|
|
/// class as [`NodeKind::MetaSync`]/[`NodeKind::Provision`].
|
|
ForgeSweep,
|
|
/// One-shot boot-time matrix user/space sweep (`matrix::ensure_all`),
|
|
/// for the same dashboard-visibility reason as [`NodeKind::ForgeSweep`].
|
|
/// The *periodic* re-sweep (every 30 min, recovering token files
|
|
/// `hive-matrix-daemon` deleted) stays a background loop in
|
|
/// `main.rs` — only the boot-time instance is a DAG node.
|
|
MatrixSweep,
|
|
/// One-shot boot-time Forgejo webhook registration, for
|
|
/// `internal/knowledge` (push → git pull) and the `agent-configs` org
|
|
/// (`pull_request` → config-PR approval). No-op when the core token, hive
|
|
/// domain, or HMAC secret aren't available yet — mirrors the guard the
|
|
/// `tokio::spawn` block it replaced already used. Agentless, build-slot-
|
|
/// and lease-exempt.
|
|
WebhookRegister,
|
|
/// One-shot boot-time `/knowledge` pull (`knowledge::pull`), reconciling
|
|
/// any commits that landed while `hive-c0re` was down. Same rationale as
|
|
/// [`NodeKind::MatrixSweep`]: the periodic hourly re-pull stays a
|
|
/// background loop, only the boot-time instance is a DAG node.
|
|
KnowledgePull,
|
|
}
|
|
|
|
/// How a hive-c0re node describes itself to a generic graph viewer.
|
|
///
|
|
/// Every field in the wire node's `data` representation here used to be a named column on
|
|
/// `NodeView`, meaningful for one node kind and `null` on all the others. As
|
|
/// free-form data it costs the wire type nothing, and a generic consumer
|
|
/// renders it without knowing what any of it means.
|
|
impl hive_jobq_wire::WireNode for NodeKind {
|
|
fn label(&self) -> String {
|
|
self.as_str().to_owned()
|
|
}
|
|
|
|
fn data(&self, id: hive_jobq_wire::WireId) -> serde_json::Value {
|
|
let mut data = serde_json::Map::new();
|
|
let agent = self.agent();
|
|
if !agent.is_empty() {
|
|
data.insert("agent".to_owned(), agent.into());
|
|
}
|
|
if let NodeKind::DeployWindow { approval_id, .. } = self {
|
|
data.insert("approval_id".to_owned(), (*approval_id).into());
|
|
}
|
|
if let NodeKind::MetaLock { inputs, .. } = self
|
|
&& !inputs.is_empty()
|
|
{
|
|
data.insert("inputs".to_owned(), inputs.clone().into());
|
|
}
|
|
// Not in the payload at all — the build log is keyed on node identity
|
|
// in a side table, which is why `data` is handed the id.
|
|
if let Some(log) = crate::build_logs::global().and_then(|h| h.id_for_node(id)) {
|
|
data.insert("build_log_id".to_owned(), log.into());
|
|
}
|
|
if data.is_empty() {
|
|
serde_json::Value::Null
|
|
} else {
|
|
serde_json::Value::Object(data)
|
|
}
|
|
}
|
|
}
|
|
|
|
impl NodeKind {
|
|
/// Wire string for the node's label on the graph wire
|
|
/// ([`hive_jobq_wire::WireNode::label`]).
|
|
pub fn as_str(&self) -> &'static str {
|
|
match self {
|
|
NodeKind::MetaSync { .. } => "meta_sync",
|
|
NodeKind::Prebuild { .. } => "prebuild",
|
|
NodeKind::Swap { .. } => "swap",
|
|
NodeKind::RebuildBookkeeping { .. } => "rebuild_bookkeeping",
|
|
NodeKind::Provision { .. } => "provision",
|
|
NodeKind::Create { .. } => "create",
|
|
NodeKind::DestroyContainer { .. } => "destroy_container",
|
|
NodeKind::PurgeState { .. } => "purge_state",
|
|
NodeKind::DestroyBookkeeping { .. } => "destroy_bookkeeping",
|
|
NodeKind::MetaLock { .. } => "meta_lock",
|
|
NodeKind::Reconcile { .. } => "reconcile",
|
|
NodeKind::Start { .. } => "start",
|
|
NodeKind::Stop { .. } => "stop",
|
|
NodeKind::StopForUpdate { .. } => "stop_for_update",
|
|
NodeKind::Signal { .. } => "signal",
|
|
NodeKind::Drain { .. } => "drain",
|
|
NodeKind::PauseSignal { .. } => "pause_signal",
|
|
NodeKind::PauseDrain { .. } => "pause_drain",
|
|
NodeKind::WriteDropin { .. } => "write_dropin",
|
|
NodeKind::WritePermFile { .. } => "write_perm_file",
|
|
NodeKind::Reparent { .. } => "reparent",
|
|
NodeKind::DeployWindow { .. } => "deploy_window",
|
|
NodeKind::AgentWindow { .. } => "agent_window",
|
|
NodeKind::MergeVerify { .. } => "merge_verify",
|
|
NodeKind::DeployApply { .. } => "deploy_apply",
|
|
NodeKind::FinalizeDeploy { .. } => "finalize_deploy",
|
|
NodeKind::DeployTail { .. } => "deploy_tail",
|
|
NodeKind::ResolveApproval { .. } => "resolve_approval",
|
|
NodeKind::EmitRebuilt { .. } => "emit_rebuilt",
|
|
NodeKind::SetWanted { .. } => "set_wanted",
|
|
NodeKind::ForgeSweep => "forge_sweep",
|
|
NodeKind::MatrixSweep => "matrix_sweep",
|
|
NodeKind::WebhookRegister => "webhook_register",
|
|
NodeKind::KnowledgePull => "knowledge_pull",
|
|
}
|
|
}
|
|
|
|
/// The agent this node targets, or `""` for agentless kinds
|
|
/// ([`NodeKind::MetaLock`] on the `hyperhive` pseudo-agent,
|
|
/// [`NodeKind::Reparent`] which can span multiple agents, and
|
|
/// [`NodeKind::ResolveApproval`] which acts on an approval row).
|
|
#[must_use]
|
|
pub fn agent(&self) -> &str {
|
|
match self {
|
|
NodeKind::MetaSync { agent, .. }
|
|
| NodeKind::Prebuild { agent }
|
|
| NodeKind::Swap { agent }
|
|
| NodeKind::RebuildBookkeeping { agent }
|
|
| NodeKind::Provision { agent }
|
|
| NodeKind::Create { agent }
|
|
| NodeKind::DestroyContainer { agent }
|
|
| NodeKind::PurgeState { agent }
|
|
| NodeKind::DestroyBookkeeping { agent, .. }
|
|
| NodeKind::Reconcile { agent }
|
|
| NodeKind::Start { agent }
|
|
| NodeKind::Stop { agent }
|
|
| NodeKind::StopForUpdate { agent }
|
|
| NodeKind::Signal { agent }
|
|
| NodeKind::Drain { agent }
|
|
| NodeKind::PauseSignal { agent }
|
|
| NodeKind::PauseDrain { agent }
|
|
| NodeKind::WriteDropin { agent }
|
|
| NodeKind::WritePermFile { agent, .. }
|
|
| NodeKind::DeployWindow { agent, .. }
|
|
| NodeKind::AgentWindow { agent }
|
|
| NodeKind::MergeVerify { agent, .. }
|
|
| NodeKind::DeployApply { agent, .. }
|
|
| NodeKind::FinalizeDeploy { agent, .. }
|
|
| NodeKind::DeployTail { agent, .. }
|
|
| NodeKind::EmitRebuilt { agent, .. }
|
|
| NodeKind::SetWanted { agent, .. } => agent,
|
|
NodeKind::MetaLock { .. }
|
|
| NodeKind::Reparent { .. }
|
|
| NodeKind::ResolveApproval { .. }
|
|
| NodeKind::ForgeSweep
|
|
| NodeKind::MatrixSweep
|
|
| NodeKind::WebhookRegister
|
|
| NodeKind::KnowledgePull => "",
|
|
}
|
|
}
|
|
|
|
/// Whether running this node is *expected* to take the agent's container
|
|
/// down. Feeds `RunningTransient::takes_container_down`, which the crash watcher
|
|
/// reads to tell an intentional stop from a crash.
|
|
///
|
|
/// This is a **safety** question, not a display one — it decides whether a
|
|
/// vanished container raises an alert. It is deliberately not derived from
|
|
/// the pill label: a label is free to be renamed or added without moving
|
|
/// the alerting boundary, and only the operation itself knows its intent.
|
|
///
|
|
/// Default is `false`, and that asymmetry is the point. A wrong `false`
|
|
/// costs a spurious crash event; a wrong `true` **swallows a real crash**
|
|
/// silently. So a kind earns `true` by being listed here, and anything new
|
|
/// is noisy-but-safe until someone decides otherwise.
|
|
#[must_use]
|
|
pub fn takes_container_down(&self) -> bool {
|
|
matches!(
|
|
self,
|
|
// Explicit stops, and the quiesce steps that precede one.
|
|
NodeKind::Stop { .. }
|
|
| NodeKind::StopForUpdate { .. }
|
|
| NodeKind::Signal { .. }
|
|
| NodeKind::Drain { .. }
|
|
| NodeKind::SetWanted { up: false, .. }
|
|
// The rebuild's own machinery: the container is down across the
|
|
// swap and the drop-in write that reconfigures it.
|
|
| NodeKind::Swap { .. }
|
|
| NodeKind::WriteDropin { .. }
|
|
)
|
|
// Everything else is `false` on purpose, including the ones that would
|
|
// be easy to wave through:
|
|
// - `Create` / `Start` / `SetWanted{up}` bring a container UP. A
|
|
// container disappearing *while starting* is a genuine crash and has
|
|
// to keep reporting as one.
|
|
// - `DestroyContainer` looks like the most obvious `true` on this list
|
|
// and is the one that must stay `false`. It is edged downstream of a
|
|
// `Stop`, so the container is already down when it claims; the stop
|
|
// that the operator asked for is accounted for by the node that
|
|
// performs it. A container found alive at destroy time is a genuine
|
|
// bug, and a `true` here would suppress the alert that says so.
|
|
// - `Reconcile` is a planner; it fans out `Start` / `Stop`, which carry
|
|
// their own answer.
|
|
// - `DeployWindow` brackets a deploy without itself stopping anything.
|
|
// - `AgentWindow` likewise, though it *parents* nodes that answer
|
|
// `true`. This is per-node, not per-subtree, and each of those
|
|
// children reports for itself — so `true` here would only widen
|
|
// suppression over the build and tail, where a vanished container is
|
|
// still a real crash.
|
|
}
|
|
}
|