diff --git a/docs/approvals.md b/docs/approvals.md index d54b590e..ca84dc77 100644 --- a/docs/approvals.md +++ b/docs/approvals.md @@ -38,33 +38,23 @@ request. comments, CI status) and sees a matching card on the dashboard with a "review PR on forge" deep link. They click ◆ APPR0VE (or `hivectl approvals approve ` on the CLI) once satisfied. -4. On approve, a deploy DAG runs three phases under a resource-holding - `DeployWindow` root (see *Queue templates* below): - - `MergeVerify` re-reads the live PR head and **aborts if it drifted** - from the reviewed `fetched_sha` (the submitter must push again, - which queues a fresh approval); then fetches that head into the - applied repo and **eval-verifies** it — a flake eval on a throwaway - checkout. This is the trust gate: it relies on c0re's own eval, not - on any in-repo (agent-forgeable) signal like a CI status. Nothing is - mutated in this phase, so a rejection here leaves the forge and the - applied repo exactly as they were. - - `DeployApply` parks the pre-merge `applied/main` in - `refs/hyperhive/rollback/`, then fast-forward-merges - the reviewed head to the forge config repo's `main` (this IS the - merge — a `core`-authenticated ff-merge pinned to the reviewed sha, - so a moved PR head can't substitute bytes), and runs the deploy - proper (`deploy_applied_target`): ff `applied/main`, two-phase meta - deploy, container rebuild. On success it drops the rollback ref and - plants `deployed/`. - - `DeployTail` runs on **every** outcome, including a cancel-cascade. - If the rollback ref survived, the deploy never confirmed good: it - rolls `applied/main` back, resyncs the working tree, and aborts the - staged meta lock, so the agent stays on its last-good tree. Then it - mirrors the config repo (and its new deploy tag) to the forge. - - The rollback state lives in a **git ref, not a local variable**, on - purpose: hive-c0re can restart between the apply and the tail, and the - tail still has to know what to undo when it does. +4. On approve, `run_merge_config_pr`: + - re-reads the live PR head and **aborts if it drifted** from the + reviewed `fetched_sha` (the submitter must push again, which queues + a fresh approval); + - fetches that head into the applied repo and **eval-verifies** it — + a flake eval on a throwaway checkout. This is the trust gate: it + relies on c0re's own eval, not on any in-repo (agent-forgeable) + signal like a CI status; + - fast-forwards the forge config repo's `main` to the reviewed head + (this IS the merge — a `core`-authenticated ff-push, so a moved PR + head can't substitute bytes) and marks the PR merged (best-effort; + `main` is already there); + - runs the shared deploy tail (`deploy_applied_target`): builds under + the tag sequence below, and on success `applied/main` + fast-forwards to the reviewed commit + the meta lock bumps. On + failure, main stays put and the working tree resets to the previous + deployed commit. 5. `HelperEvent::ApprovalResolved` (and `Rebuilt`) land in the **submitting agent's** inbox via `notify_submitter`, carrying both the canonical sha and the terminal tag (the approval row carries a @@ -110,12 +100,13 @@ kind-specific payload carrier. approval row. No MCP tool call needed — the forge PR IS the request. `commit_ref` stores the **PR number** (decimal), and `fetched_sha` is the PR **head sha at queue time** (the "reviewed" sha). On approve, - the deploy DAG's `MergeVerify` phase re-reads the live PR head and - aborts if it drifted from `fetched_sha` (submitter must push again to - re-trigger), then fetches that head into the applied repo and - eval-verifies it; `DeployApply` fast-forward-merges the forge config - repo's `main` to it (the merge) and runs `deploy_applied_target`; - `DeployTail` compensates on failure. Never a first spawn. + `run_merge_config_pr` re-reads the live PR head and aborts if it + drifted from `fetched_sha` (submitter must push again to + re-trigger), then fetches that head into the applied repo, + eval-verifies it, fast-forwards the forge config repo's `main` to + it (the merge), marks the PR merged (best-effort — `main` is + already there), and runs the shared deploy tail + (`deploy_applied_target`). Never a first spawn. - `Spawn` — direct container creation from the agent's config repo. `commit_ref` is empty. Submitted via `HostRequest::RequestSpawn` (operator-gated, the `◆ R3QU3ST SP4WN` dashboard button + @@ -248,9 +239,8 @@ place (it only re-locks when the declared url itself changes), so the forge-declared / applied-deployed split is stable. Per-deploy lock flow (two-phase, owned by -`actions::run_deploy_apply` → `deploy_applied_target` → -`meta::{prepare,finalize,abort}_deploy`, with the abort half moved out -into `actions::run_deploy_tail`): +`actions::run_merge_config_pr` → `deploy_applied_target` → +`meta::{prepare,finalize,abort}_deploy`): 1. `meta::prepare_deploy(name)` runs `nix flake lock --update-input agent-` without @@ -264,13 +254,10 @@ into `actions::run_deploy_tail`): ")` stages `flake.lock` and commits with `deploy deployed/ `. Meta's git log gains one entry per successful deploy. -4. On failure — the `DeployTail` node runs `meta::abort_deploy()` - (`git restore flake.lock`) so the meta history shows only +4. On failure — `meta::abort_deploy()` runs + `git restore flake.lock` so the meta history shows only successes; the failure stays as an annotated `failed/` - tag in `applied/`. The tail runs on every outcome, so this - also covers a hive-c0re restart mid-build: the staged lock is - dropped and `applied/main` rolled back from the parked - `refs/hyperhive/rollback/`. + tag in `applied/`. Single-phase variants exist for paths without rollback semantics: `meta::lock_update_for_rebuild(name)` for @@ -359,22 +346,18 @@ the approval handler submits a DAG to the global job queue | `ApprovalKind` | DAG submitted | source | |---|---|---| -| `MergeConfigPr` | `rebuild` (`DeployWindow` root + `MergeVerify → DeployApply` + `DeployTail`) | `approval` | +| `MergeConfigPr` | `rebuild` (single opaque `ApprovalDeploy` node) | `approval` | | `UpdateMetaInputs` | `meta_update` (`MetaLock` + rebuild fan-out) | `approval` | | `Spawn` | `spawn` (`Create → WriteDropin → Reconcile`) | `approval` | | `InitConfig` | — runs inline (sub-second git seed) | — | | `SchedulePrompt` | — runs inline (single sqlite insert) | — | -The DAG carries the originating `approval_id`, surfaced on the node that -owns it — for a deploy that's the `DeployWindow` root, so the dashboard -renders one approval card, not four. **Every** queued kind resolves -through `actions::resolve_approval_dag` when its DAG settles terminal: -the deploy's phases are ordinary queue nodes, so the DAG's own terminal -state is the authoritative outcome. That hook fires the matching -`HelperEvent::*` via `finish_approval`, derives the `Rebuilt` event's -terminal tag (verifying the tag actually resolves in the applied repo — -a pre-merge rejection plants none), posts the failing build log back to -the config PR, and for a spawn runs the post-spawn forge bookkeeping. +The DAG carries the originating `approval_id`. The `ApprovalDeploy` +node runs `run_approval_merge_config_pr` (the two-phase meta deploy +stays inside `actions.rs`) and fires the matching `HelperEvent::*` via +`finish_approval` itself; `Spawn` and `UpdateMetaInputs` DAGs resolve +through `actions::resolve_approval_dag` when the DAG settles terminal +(a spawn additionally runs the post-spawn forge bookkeeping there). Two visible consequences: diff --git a/docs/coordinator.md b/docs/coordinator.md index 1ed1fbf1..a16f11eb 100644 --- a/docs/coordinator.md +++ b/docs/coordinator.md @@ -37,15 +37,12 @@ Nix-heavy — hold one of the `buildSlots` permits for the node's duration: | `Swap` | drop-in rewrite + `nixos-container update` profile-swap (requires the container stopped); the post-swap bookkeeping tail lives in the sibling `PostSwap` node | | `Create` | first-spawn provisioning + `nixos-container create` (atomic build+create) | | `MetaLock` | meta flake lock bump (`lock_update` / boot-sweep `lock_update_hyperhive`, commit fused — see below); fans out child `Rebuild` DAGs on completion | -| `DeployWindow` | resource-holding root of the merge-config-PR deploy subtree — declares the build slot, the lease and the meta window, then completes immediately so its children run under them (see _Approvals_ below) | -| `DeployApply` | the deploy's irreversible half: ff-merge the reviewed PR head, two-phase meta deploy, container rebuild | +| `ApprovalDeploy` | the opaque apply-commit / merge-config-PR pipeline (see _Approvals_ below) | Cheap — no build slot: | Node | Behavior | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------ | -| `MergeVerify` | the deploy's pre-merge gate — PR-head drift check, fetch, `verify_commit` eval. Mutates nothing, so a rejection here needs no compensation | -| `DeployTail` | the deploy's `AfterAny` compensation + bookkeeping tail — rolls `applied/main` back from the parked `refs/hyperhive/rollback/` and aborts the staged meta lock when the deploy never confirmed good, then mirrors the config repo to the forge. Infallible by construction | | `MetaSync` | the rebuild's meta preamble — rebuild-dir prep, idempotent meta `sync_agents`, optional per-agent relock. Holds the `MetaWindow` resource (below); deliberately its own node so the window never covers `Prebuild`'s multi-minute build | | `Reconcile` | idempotent power converge: read `wanted` (below) + observed state; start if `Up` & down (cold-start fallback included), stop if `Offline` & up, else noop | | `StopForUpdate` | mechanical `nixos-container stop` for the profile swap; never touches `wanted`; noop if already stopped | @@ -67,8 +64,7 @@ container build: - **The deploy window** (`Resource::MetaWindow`): a global, capacity-1 queue resource declared by every node kind that mutates the meta repo — `MetaSync`, `MetaLock`, `WritePermFile`, `Provision`'s agent registration, and - `DeployWindow` — the deploy subtree's root, which holds it across every - phase below it (`NodeKind::needs_meta_window`). Two meta + `ApprovalDeploy` for its whole span (`NodeKind::needs_meta_window`). Two meta mutations can therefore never interleave, so no commit lands inside another node's staged window. It is a queue resource rather than a runtime mutex because a resource is held by a subtree root across its whole subtree, which @@ -184,7 +180,7 @@ resources are free. Resources: per-node; a DAG can span agents) and globally exclusive per agent across all DAGs: acquired at a container-affecting node (`SetWanted`, `StopForUpdate`, `Swap`, `Signal`, `Drain`, `Reconcile`, `WriteDropin`, - `Create`, `DeployWindow`), held by the owning DAG until it's terminal, + `Create`, `ApprovalDeploy`), held by the owning DAG until it's terminal, so two DAGs never interleave container ops on the same agent. A DAG touching several agents holds one lease per agent. (`SetWanted` is a store write, not a container op, but takes the lease anyway so a power-op DAG's @@ -217,30 +213,15 @@ per template. ### Approvals -`MergeConfigPr` approvals ride as a four-node deploy subtree: - -``` -DeployWindow (root — build slot + lease + meta window, no work of its own) -├── MergeVerify drift gate, fetch, verify_commit -├── DeployApply AfterOk(verify) park rollback ref, ff-merge, deploy -└── DeployTail AfterAny(apply) compensate, mirror to forge -``` - -The root's resources are held across the whole subtree, so the two-phase -`prepare_deploy` / `finalize_deploy` span keeps its staged `flake.lock` -protected even though the phases are separate nodes. Splitting them buys -three things a single opaque node couldn't have: per-phase visibility on the -dashboard, a `MergeVerify` failure that provably mutated nothing, and a -compensation step that survives a hive-c0re restart — the pre-merge -`applied/main` is parked in `refs/hyperhive/rollback/`, not in a -local variable, so `DeployTail` can still undo a half-finished deploy after a -crash. - -`Spawn` and `UpdateMetaInputs` approvals map onto the ordinary `spawn` / -`meta-update` shapes. The scheduler fires `actions::resolve_approval_dag` -exactly once when **any** approval-carrying DAG settles terminal — deploys -included, since their outcome is now the DAG's own state (including -cancelled-while-queued, which fails the approval instead of dangling it). +`MergeConfigPr` approvals ride as single-node +`ApprovalDeploy` DAGs: the two-phase `prepare_deploy` / `finalize_deploy` / +`abort_deploy` meta orchestration stays inside `actions.rs` in v1 +(deliberately not modeled as scheduler nodes) and resolves the approval +itself. `Spawn` and `UpdateMetaInputs` approvals map onto the ordinary +`spawn` / `meta-update` shapes; the scheduler fires +`actions::resolve_approval_dag` exactly once when such a DAG settles +terminal (including cancelled-while-queued, which fails the approval instead +of dangling it). ### Wire shape @@ -357,9 +338,8 @@ Sequence for a rebuild DAG (each step is its own queue node): The approval apply-commit pipeline still drives `lifecycle::rebuild_no_meta` (the fused stop/update/start path with an inline start) inside its -`DeployApply` node, because it verifies the agent comes back up before -finalizing the deploy tag. Breaking that fused path apart into the -`Prebuild → Swap → Reconcile` chain above is increment 2b of #2664, not 2a. +`ApprovalDeploy` node, because it verifies the agent comes back up before +finalizing the deploy tag. ### Cold-start fallback diff --git a/hive-c0re/src/actions.rs b/hive-c0re/src/actions.rs index 3db5de87..13e3992e 100644 --- a/hive-c0re/src/actions.rs +++ b/hive-c0re/src/actions.rs @@ -18,14 +18,15 @@ use crate::lifecycle; /// (operator no longer blocks on a 30-90s spinner for `MergeConfigPr`). /// /// Dispatch: -/// - `MergeConfigPr` → a `DeployWindow` DAG (`MergeVerify → DeployApply → -/// DeployTail` under a resource-holding root; ~30-90s) +/// - `MergeConfigPr` → a single-node `ApprovalDeploy` +/// DAG (the two-phase meta deploy stays opaque in v1; ~30-90s) /// - `UpdateMetaInputs` → a `MetaUpdate` DAG (fan-out on completion) /// - `Spawn` → a `Spawn` DAG (`Create → WriteDropin → Reconcile`) /// - `InitConfig` → inline (<1s; queue card would be noise) /// -/// Every queued kind — deploys included — resolves its approval row via -/// [`resolve_approval_dag`] when the DAG settles terminal. +/// `ApprovalDeploy` resolves the approval inside its pipeline; the +/// `MetaUpdate` / `Spawn` DAGs resolve via [`resolve_approval_dag`] +/// when their DAG settles terminal. pub async fn approve(coord: Arc, id: i64) -> Result<()> { let approval = coord.approvals.mark_approved(id)?; tracing::info!( @@ -103,10 +104,10 @@ pub async fn approve(coord: Arc, id: i64) -> Result<()> { } ApprovalKind::MergeConfigPr => { // The work ends in a container rebuild, so route it through the - // rebuild queue. The queue worker dispatches the deploy DAG's - // nodes to `run_deploy_merge_verify` (drift gate + eval), - // `run_deploy_apply` (ff-merge + rebuild) and `run_deploy_tail` - // (compensation + forge mirror). + // rebuild queue. The queue worker dispatches MergeConfigPr + // approvals to `run_merge_config_pr` (verify the reviewed PR head, + // ff the forge config repo's main to it, mark merged, then the + // deploy tail). enqueue_approval_rebuild( &coord, approval.agent.as_str(), @@ -118,10 +119,10 @@ pub async fn approve(coord: Arc, id: i64) -> Result<()> { } } -/// Submit the deploy DAG tied to an approval id. Used by the `MergeConfigPr` -/// dispatch arm — the work ends in a container rebuild routed through the -/// queue. See [`crate::job_queue::templates::approval_deploy`] for the node -/// shape; the executor dispatches each node to the `run_deploy_*` bodies below. +/// Submit the single-node `ApprovalDeploy` DAG tied to an approval id. +/// Used by the `MergeConfigPr` dispatch arm — the work ends in a container +/// rebuild routed through the queue; the node executor runs +/// `run_merge_config_pr`. fn enqueue_approval_rebuild( coord: &Arc, agent: &str, @@ -141,226 +142,45 @@ fn enqueue_approval_rebuild( coord.emit_rebuild_queue_snapshot(); } -/// Ref under which [`run_deploy_apply`] parks the pre-merge `applied/main` -/// sha, for [`run_deploy_tail`] to compensate with. -/// -/// Deliberately a *git ref in the applied repo* rather than an in-memory value -/// handed between nodes: hive-c0re can restart between the apply and the tail, -/// and the whole point of splitting the deploy is that the tail still knows -/// what to undo when it does. The ref's existence IS the "a merge landed but -/// hasn't been confirmed good yet" flag — [`run_deploy_apply`] drops it the -/// moment the rebuild succeeds. -fn rollback_ref(approval_id: i64) -> String { - format!("refs/hyperhive/rollback/{approval_id}") -} - -/// Everything a deploy node needs, re-derived from sqlite on each node rather -/// than cached across the DAG. Nothing here is *computed* by an earlier node — -/// `pr` and `reviewed` are fields of the approval row the operator signed off -/// on — so re-reading is both cheap and the authoritative source of truth. -struct DeployCtx { - approval: hive_sh4re::Approval, - /// PR number, parsed from `approval.commit_ref`. - pr: u64, - /// The PR head sha the operator reviewed (`approval.fetched_sha`). - reviewed: String, - agent_dir: std::path::PathBuf, - applied_dir: std::path::PathBuf, - /// The agent's forge config repo (`/`). - repo: String, -} - -fn deploy_ctx(coord: &Coordinator, approval_id: i64) -> Result { +/// Worker entry point for `ApprovalKind::MergeConfigPr` queue entries — the +/// config-change flow's deploy worker. Re-fetches the approval row, runs the +/// merge pipeline, and fires +/// `ApprovalResolved` + the `Rebuilt` lifecycle event via `finish_approval`. +/// `run_merge_config_pr` already fast-forwarded the forge repo's `main` to the +/// reviewed head (that IS the merge), so `push_config`'s `main` refspec is a +/// no-op — but it still mirrors the `deployed/` / `failed/` tag the +/// deploy tail plants onto the merged sha, giving the merged commit a +/// forge-visible deploy marker. A `MergeConfigPr` is never a first spawn (the +/// agent already exists). +pub async fn run_approval_merge_config_pr( + coord: &Arc, + queue_entry_id: Option, + approval_id: i64, +) -> Result<()> { let approval = fetch_approval_for_worker(coord, approval_id, ApprovalKind::MergeConfigPr)?; - let pr: u64 = approval.commit_ref.parse().map_err(|e| { - anyhow::anyhow!( - "parse PR number from commit_ref {:?}: {e}", - approval.commit_ref - ) - })?; - let reviewed = approval.fetched_sha.clone().ok_or_else(|| { - anyhow::anyhow!("merge config pr approval {approval_id} has no reviewed head sha") - })?; - Ok(DeployCtx { - pr, - reviewed, - agent_dir: crate::paths::agent_runtime_dir(approval.agent.as_str()), - applied_dir: crate::paths::applied_dir(approval.agent.as_str()), - repo: crate::forge::config_repo(approval.agent.as_str()), - approval, - }) -} - -/// `MergeVerify` node body — everything that can say "no" before anything is -/// mutated. `approval.commit_ref` is the PR number; `approval.fetched_sha` is -/// the PR head the operator reviewed. Steps: -/// 1. drift gate — re-read the live PR head; if it moved since review, abort -/// (the operator must re-review the new head); -/// 2. fetch the reviewed head into the applied repo so later git ops resolve -/// it locally; -/// 3. eval-verify the reviewed commit against the meta flake. -/// -/// Nothing here needs undoing on failure: the fetch only adds objects, and -/// `main` doesn't move. That's the whole reason this is its own node — a -/// failure at this stage leaves [`run_deploy_tail`] with no ref to compensate. -/// -/// # Errors -/// -/// Returns an error if the approval can't be loaded, if the live PR head has -/// drifted from the reviewed sha, if fetching that head into the applied repo -/// fails, or if the eval-verify of the reviewed commit fails. Every one of -/// these leaves the forge and `main` untouched, so the node is safely -/// retryable. -pub async fn run_deploy_merge_verify( - coord: &Arc, - queue_entry_id: Option, - approval_id: i64, -) -> Result<()> { - let ctx = deploy_ctx(coord, approval_id)?; - let pr = ctx.pr; - let reviewed = ctx.reviewed.as_str(); - - // 1. Drift gate: the live PR head must still equal what was reviewed. - coord.set_queue_step(queue_entry_id, "verify PR head"); - let head = crate::forge::pr_head_sha(&ctx.repo, pr) - .await - .map_err(|e| anyhow::anyhow!("read PR #{pr} head: {e}"))?; - if head != reviewed { - bail!( - "PR #{pr} head drifted since review (reviewed {reviewed}, now {head}); re-review before merging" - ); - } - - // 2. Fetch the reviewed head into applied so ff/verify/deploy resolve it. - coord.set_queue_step(queue_entry_id, "fetch PR head"); - crate::forge::fetch_pr_head_into_applied(&ctx.repo, pr) - .await - .map_err(|e| anyhow::anyhow!("fetch PR #{pr} head into applied: {e}"))?; - - // 3. Eval-verify BEFORE the irreversible merge (bad nix fails fast here). - coord.set_queue_step(queue_entry_id, "verify proposal (eval)"); - crate::meta::verify_commit(ctx.approval.agent.as_str(), &ctx.applied_dir, reviewed) - .await - .map_err(|e| anyhow::anyhow!("verify merge head {reviewed}: {e:#}"))?; - Ok(()) -} - -/// `DeployApply` node body — the irreversible half. Parks the rollback ref, -/// fast-forward-merges the PR (THE merge), then runs the deploy proper. -/// -/// The ref is parked *before* the merge, so a hive-c0re crash anywhere from -/// here on still leaves [`run_deploy_tail`] enough to undo. If the merge itself -/// fails, `main` never moved and the tail's compensation is a no-op against the -/// same sha — harmless, and cheaper than trying to be clever about it. -/// -/// # Errors -/// -/// Returns an error if the approval can't be loaded, if reading or parking the -/// pre-merge `main` sha fails, if the forge refuses the fast-forward merge -/// (including a head that drifted between verify and merge), or if the deploy -/// of the merged target fails. From the merge onward a failure is *not* -/// retryable on its own — [`run_deploy_tail`] runs `AfterAny` to compensate. -pub async fn run_deploy_apply( - coord: &Arc, - queue_entry_id: Option, - approval_id: i64, -) -> Result<()> { - let ctx = deploy_ctx(coord, approval_id)?; - let agent = ctx.approval.agent.as_str(); - let pr = ctx.pr; - - let prev_main = lifecycle::git_rev_parse(&ctx.applied_dir, "refs/heads/main") - .await - .map_err(|e| anyhow::anyhow!("read applied/main: {e:#}"))?; - lifecycle::git_update_ref(&ctx.applied_dir, &rollback_ref(approval_id), &prev_main) - .await - .map_err(|e| anyhow::anyhow!("park rollback ref for approval {approval_id}: {e:#}"))?; - - // THE merge: fast-forward-only merge the reviewed head to `main` via the - // forge API, pinned to the reviewed sha (`head_commit_id`). This one call - // both advances `main` to the reviewed head and marks the PR merged — no - // direct push to the protected branch. A failure here means `main` was NOT - // advanced, so it's fatal: we must not deploy a head the forge didn't merge. - coord.set_queue_step(queue_entry_id, "fast-forward-merge PR"); - match crate::forge::merge_config_pr_ff(&ctx.repo, pr, &ctx.reviewed).await { - Ok(()) => {} - Err(crate::forge::ForgeMergeError::HeadDrift { expected, actual }) => bail!( - "PR #{pr} head drifted before merge (reviewed {expected}, now {actual}); re-review before merging" - ), - Err(e) => bail!("ff-merge PR #{pr}: {e}"), - } - - deploy_applied_target( - coord, - agent, - &ctx.agent_dir, - &ctx.applied_dir, - &ctx.reviewed, - approval_id, - queue_entry_id, - ) - .await -} - -/// `DeployTail` node body — compensation + bookkeeping, `AfterAny` the apply -/// node so it runs on every outcome including a cancel-cascade. Infallible by -/// construction: it is the recovery step, so it has nothing to hand a failure -/// to. Every fallible call inside warns and continues. -/// -/// 1. If the rollback ref survived, the deploy did not confirm good: roll -/// `applied/main` back to the parked sha, resync the working tree, and drop -/// the staged meta lock so the deploy log only ever shows successes. -/// 2. Mirror the agent's config repo to the forge. `main` is already ff'd by -/// the merge, so only the `deployed/` / `failed/` tag refspec -/// actually lands — that's what gives the merged commit a forge-visible -/// deploy marker. -/// -/// Takes `agent` from the node payload rather than the approval row so it still -/// works if the row vanished underneath the DAG (deny race, purge). -pub async fn run_deploy_tail( - coord: &Arc, - queue_entry_id: Option, - agent: &str, - approval_id: i64, -) { - let applied_dir = crate::paths::applied_dir(agent); - let rollback = rollback_ref(approval_id); - if let Ok(prev_main) = lifecycle::git_rev_parse(&applied_dir, &rollback).await { - // Belt and braces: `run_deploy_apply` drops the ref before it plants - // `deployed/`, so seeing both means the *delete* failed on an - // otherwise-successful deploy. Rolling back there would be the worst - // outcome this node can produce, so the tag wins. - if lifecycle::git_rev_parse(&applied_dir, &format!("deployed/{approval_id}")) - .await - .is_ok() - { - tracing::warn!( - %agent, approval_id, - "deploy tail: rollback ref outlived a successful deploy; dropping it without compensating" - ); - } else { - coord.set_queue_step(queue_entry_id, "roll back applied/main"); - if let Err(e) = - lifecycle::git_update_ref(&applied_dir, "refs/heads/main", &prev_main).await - { - tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: main rollback failed"); - } - if let Err(e) = lifecycle::git_read_tree_reset(&applied_dir, "refs/heads/main").await { - tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: rollback read-tree failed"); - } - if let Err(e) = crate::meta::abort_deploy().await { - tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: meta abort_deploy failed"); - } - } - if let Err(e) = lifecycle::git_delete_ref(&applied_dir, &rollback).await { - tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: drop rollback ref failed"); - } - } - + let agent_dir = crate::paths::agent_runtime_dir(approval.agent.as_str()); + let applied_dir = crate::paths::applied_dir(approval.agent.as_str()); + // Captured up front to scope the failure-comment's build-log lookup to + // rows this deploy produced (see `post_merge_failure_to_pr`). + let since_ts = hive_sh4re::wire_time::now_unix(); + coord.set_queue_step(queue_entry_id, "merge config pr"); + let (result, terminal_tag) = + run_merge_config_pr(coord, &approval, &agent_dir, &applied_dir, queue_entry_id).await; + // Mirror the deploy bookkeeping tag (`deployed/` or `failed/`) the + // deploy tail planted onto the merged sha to the forge config repo, so the + // merged commit carries a forge-visible deploy marker. `main` is already + // ff'd by the merge, so only the tag refspec actually lands; best-effort, + // never fails the approval. coord.set_queue_step(queue_entry_id, "forge push"); - if let Err(e) = crate::forge::push_config(agent).await { - tracing::warn!(%agent, error = ?e, "forge: push_config after merge failed"); + if let Err(e) = crate::forge::push_config(approval.agent.as_str()).await { + tracing::warn!(agent = %approval.agent, error = ?e, "forge: push_config after merge failed"); } + // On a failed deploy, surface the failing build log back onto the PR so + // the manager sees why it was rejected without leaving the forge. + if let Err(e) = &result { + post_merge_failure_to_pr(coord, &approval, since_ts, e).await; + } + finish_approval(coord, &approval, result, terminal_tag) } /// Max stderr bytes to inline in a PR failure comment. Keeps the comment @@ -374,25 +194,21 @@ const PR_FAIL_LOG_TAIL_BYTES: usize = 4000; /// approval-resolution path. /// /// The failing `build_log` row is located heuristically: the most recent `fail` -/// row for this agent that started at/after the approval was decided (i.e. when -/// its deploy DAG was submitted). Because deploys are serialised per agent -/// through the queue, that is the step which just failed — `verify`, -/// `prepare-deploy`, `prebuild`, or the container rebuild. Pre-build failures -/// (drift gate, fetch) create no `build_log` row, so the comment then carries -/// only the error text. +/// row for this agent that started at/after `since_ts` (the caller's function +/// entry). Because deploys are serialised per agent through the queue, that is +/// the step which just failed — `verify`, `prepare-deploy`, `prebuild`, or the +/// container rebuild. Pre-build failures (drift gate, fetch) create no `build_log` +/// row, so the comment then carries only the error text. async fn post_merge_failure_to_pr( coord: &Arc, approval: &hive_sh4re::Approval, + since_ts: i64, err: &anyhow::Error, ) { let Ok(pr) = approval.commit_ref.parse::() else { return; }; let repo = crate::forge::config_repo(approval.agent.as_str()); - let since_ts = approval - .resolved_at - .unwrap_or(approval.requested_at) - .timestamp(); let log_section = coord .build_logs @@ -437,6 +253,131 @@ fn tail_bytes(s: &str, max_bytes: usize) -> String { format!("[… truncated …]\n{}", &s[start..]) } +/// PR-merge config pipeline. `approval.commit_ref` is the PR number; +/// `approval.fetched_sha` is the PR head sha the operator reviewed. Steps: +/// 1. drift gate — re-read the live PR head; if it moved since review, abort +/// WITHOUT mutating anything (the operator must re-review the new head); +/// 2. fetch the reviewed head into the applied repo so later git ops resolve +/// it locally; +/// 3. eval-verify the reviewed commit against the meta flake BEFORE the +/// irreversible push; +/// 4. fast-forward the forge repo's `main` to the reviewed head — THE merge; +/// 5. mark the PR merged (best-effort: `main` is already at the head, so a +/// failure here is logged, not fatal); +/// 6. run the shared deploy tail (`deploy_applied_target`): ff applied/main, +/// meta deploy, container rebuild, finalize/rollback. +/// +/// Returns `(build result, terminal tag)` like `deploy_applied_target`. Any +/// pre-merge abort returns `Err` with no mutation; the operator re-reviews. +async fn run_merge_config_pr( + coord: &Arc, + approval: &hive_sh4re::Approval, + agent_dir: &std::path::Path, + applied_dir: &std::path::Path, + queue_entry_id: Option, +) -> (Result<()>, Option) { + let id = approval.id; + let pr: u64 = match approval.commit_ref.parse() { + Ok(n) => n, + Err(e) => { + return ( + Err(anyhow::anyhow!( + "parse PR number from commit_ref {:?}: {e}", + approval.commit_ref + )), + None, + ); + } + }; + let reviewed = match approval.fetched_sha.as_deref() { + Some(s) => s.to_owned(), + None => { + return ( + Err(anyhow::anyhow!( + "merge config pr approval {id} has no reviewed head sha" + )), + None, + ); + } + }; + let repo = crate::forge::config_repo(approval.agent.as_str()); + + // 1. Drift gate: the live PR head must still equal what was reviewed. + coord.set_queue_step(queue_entry_id, "verify PR head"); + let head = match crate::forge::pr_head_sha(&repo, pr).await { + Ok(h) => h, + Err(e) => return (Err(anyhow::anyhow!("read PR #{pr} head: {e}")), None), + }; + if head != reviewed { + return ( + Err(anyhow::anyhow!( + "PR #{pr} head drifted since review (reviewed {reviewed}, now {head}); re-review before merging" + )), + None, + ); + } + + // 2. Fetch the reviewed head into applied so ff/verify/deploy resolve it. + coord.set_queue_step(queue_entry_id, "fetch PR head"); + if let Err(e) = crate::forge::fetch_pr_head_into_applied(&repo, pr).await { + return ( + Err(anyhow::anyhow!("fetch PR #{pr} head into applied: {e}")), + None, + ); + } + + // 3. Eval-verify BEFORE the irreversible push (bad nix fails fast here). + coord.set_queue_step(queue_entry_id, "verify proposal (eval)"); + if let Err(e) = + crate::meta::verify_commit(approval.agent.as_str(), applied_dir, &reviewed).await + { + return ( + Err(anyhow::anyhow!("verify merge head {reviewed}: {e:#}")), + None, + ); + } + + // Capture the currently-deployed sha for the deploy tail's rollback. + let prev_main_sha = match lifecycle::git_rev_parse(applied_dir, "refs/heads/main").await { + Ok(s) => s, + Err(e) => return (Err(anyhow::anyhow!("read applied/main: {e:#}")), None), + }; + + // 4. THE merge: fast-forward-only merge the reviewed head to `main` via the + // forge API, pinned to the reviewed sha (`head_commit_id`). This one call + // both advances `main` to the reviewed head and marks the PR merged — no + // direct push to the protected branch. Unlike the old push-then-mark split, + // a failure here means `main` was NOT advanced, so it's fatal: we must not + // deploy a head the forge didn't merge. + coord.set_queue_step(queue_entry_id, "fast-forward-merge PR"); + match crate::forge::merge_config_pr_ff(&repo, pr, &reviewed).await { + Ok(()) => {} + Err(crate::forge::ForgeMergeError::HeadDrift { expected, actual }) => { + return ( + Err(anyhow::anyhow!( + "PR #{pr} head drifted before merge (reviewed {expected}, now {actual}); re-review before merging" + )), + None, + ); + } + Err(e) => return (Err(anyhow::anyhow!("ff-merge PR #{pr}: {e}")), None), + } + + // 5. Deploy tail. target == finalize == the reviewed head. + deploy_applied_target( + coord, + approval.agent.as_str(), + agent_dir, + applied_dir, + &reviewed, + &reviewed, + id, + &prev_main_sha, + queue_entry_id, + ) + .await +} + /// Inline (non-queued) handler for `ApprovalKind::SchedulePrompt`. /// On approve, decode the `SchedulePromptPayload` JSON from the /// approval's `commit_ref`, insert a row into `scheduled_prompts` @@ -467,19 +408,24 @@ async fn run_approval_schedule_prompt( finish_approval(coord, &approval, result, None) } -/// Terminal hook for approval-carrying DAGs — the job queue's scheduler calls -/// this exactly once when such a DAG settles terminal. Every approval-carrying -/// template resolves here, deploys included: the deploy pipeline is ordinary -/// queue nodes now, so the DAG's own terminal state is the authoritative -/// outcome and there's no in-node resolution to skip around. +/// Terminal hook for approval-carrying DAGs — the job queue's +/// scheduler calls this exactly once when such a DAG settles terminal. +/// `MetaUpdate` and `Spawn` approval DAGs resolve here (their work is +/// ordinary queue nodes); the opaque `ApprovalDeploy` pipeline resolves +/// *inside* its node, so its DAG is skipped — unless it was cancelled +/// while still queued, in which case the node never ran and the row +/// would otherwise dangle forever. pub(crate) async fn resolve_approval_dag( coord: &Arc, terminal: &crate::job_queue::TerminalDag, ) { - use crate::job_queue::State; + use crate::job_queue::{State, Template}; let Some(approval_id) = terminal.approval_id else { return; }; + if terminal.template == Template::Rebuild && terminal.state != State::Cancelled { + return; // ApprovalDeploy resolved inside the node. + } let approval = match coord.approvals.get(approval_id) { Ok(Some(a)) => a, Ok(None) => { @@ -502,61 +448,22 @@ pub(crate) async fn resolve_approval_dag( .unwrap_or_else(|| "job dag failed".to_owned()) )), }; - let mut terminal_tag = None; - match approval.kind { - ApprovalKind::Spawn => { - // Post-spawn forge bookkeeping (user, config repo mirror, meta - // access) — warn-only, then the resolution events + a rescan so - // the dashboard reflects the post-spawn state either way. - if result.is_ok() { - forge_after_first_spawn(coord, approval.agent.as_str()).await; - } else { - coord.rescan_containers_and_emit().await; - crate::dashboard::emit_tombstones_snapshot(coord).await; - } + if approval.kind == ApprovalKind::Spawn { + // Post-spawn forge bookkeeping (user, config repo mirror, meta + // access) — warn-only, then the resolution events + a rescan so + // the dashboard reflects the post-spawn state either way. + if result.is_ok() { + forge_after_first_spawn(coord, approval.agent.as_str()).await; + } else { + coord.rescan_containers_and_emit().await; + crate::dashboard::emit_tombstones_snapshot(coord).await; } - ApprovalKind::MergeConfigPr => { - terminal_tag = - deploy_terminal_tag(approval.agent.as_str(), approval_id, terminal.state).await; - // On a failed deploy, surface the failing build log back onto the - // PR so the manager sees why it was rejected without leaving the - // forge. Posted here rather than inside a node because this is the - // one place that holds the DAG's definitive error — a `MergeVerify` - // rejection and a `DeployApply` build failure both land here. - if let Err(e) = &result { - post_merge_failure_to_pr(coord, &approval, e).await; - } - } - _ => {} } - if let Err(e) = finish_approval(coord, &approval, result, terminal_tag) { + if let Err(e) = finish_approval(coord, &approval, result, None) { tracing::warn!(approval_id, error = ?e, "approval dag resolved with failure"); } } -/// Which bookkeeping tag a settled deploy DAG actually planted, for the -/// `Rebuilt` event's `tag` field. The state picks the candidate name, but the -/// applied repo has the final say: a pre-merge rejection (`MergeVerify` drift -/// gate, eval failure) fails the DAG without ever planting `failed/`, and -/// tag plants are best-effort. Reporting a tag that isn't there would send the -/// manager looking for a ref that doesn't exist. -async fn deploy_terminal_tag( - agent: &str, - approval_id: i64, - state: crate::job_queue::State, -) -> Option { - use crate::job_queue::State; - let candidate = match state { - State::Done => format!("deployed/{approval_id}"), - State::Cancelled => return None, - _ => format!("failed/{approval_id}"), - }; - lifecycle::git_rev_parse(&crate::paths::applied_dir(agent), &candidate) - .await - .ok() - .map(|_| candidate) -} - /// Re-fetch an approval row from sqlite for a queue-worker dispatch. /// Bails if the row is gone (deny race), if its kind doesn't match, /// or if the lookup itself fails. The kind check is defensive — the @@ -740,47 +647,62 @@ fn finish_approval( result } -/// Post-merge deploy for the config-PR flow. Fast-forwards `applied/main` to -/// `target`, syncs the working tree, runs the meta two-phase deploy + container -/// rebuild, and plants the `deployed/` / `failed/` bookkeeping tag. -/// -/// **Undo is not this function's job.** Every early return here leaves the -/// applied repo dirty on purpose — [`run_deploy_tail`] owns compensation, and -/// it runs whether this returns `Err`, panics, or never returns at all because -/// hive-c0re was restarted underneath it. That's the whole point of parking the -/// pre-merge sha in a git ref instead of a local variable. +/// Deploy tail for the config-PR merge flow. Fast-forwards `applied/main` to +/// `target_ref`, syncs the working tree, runs the meta two-phase deploy + +/// container rebuild, and plants the `deployed/` / +/// `failed/` bookkeeping tags. On build failure it rolls +/// `applied/main` back to `prev_main_sha` and aborts the staged meta lock so +/// the agent stays on its last-good tree. Returns the build result + the +/// terminal tag name. /// /// Caller-specific bits stay OUT of here: fetching the PR head, the -/// `verify_commit` gate, and the ff-merge. `target` is both what `applied/main` -/// fast-forwards to and the sha `meta::finalize_deploy` records — for a merge -/// they are always the same reviewed head. The agent always already exists here -/// (a merge is never a first spawn), so there's no `sync_agents` step — the +/// `verify_commit` gate, the ff-merge, and forge mark-merged. `finalize_sha` +/// is the sha recorded by `meta::finalize_deploy`; `target_ref` is what +/// `applied/main` fast-forwards to. The agent always already exists here (a +/// merge is never a first spawn), so there's no `sync_agents` step — the /// operator `Spawn` flow owns first-time meta registration. +#[allow( + clippy::too_many_arguments, + clippy::too_many_lines, + reason = "one sequential ff/deploy/rebuild/finalize pipeline; splitting it \ + would obscure the linear flow" +)] async fn deploy_applied_target( coord: &Arc, agent: &str, agent_dir: &std::path::Path, applied_dir: &std::path::Path, - target: &str, - id: i64, + target_ref: &str, + finalize_sha: &str, + tag_base: i64, + prev_main_sha: &str, queue_entry_id: Option, -) -> Result<()> { +) -> (Result<()>, Option) { + let id = tag_base; + coord.set_queue_step(queue_entry_id, "fast-forward applied/main"); - // Fast-forward applied/main to target + sync the working tree. Meta input - // pins `?ref=main`, so this is what makes nix re-lock to the target commit - // on the prepare_deploy step below. - lifecycle::git_update_ref(applied_dir, "refs/heads/main", target) - .await - .map_err(|e| anyhow::anyhow!("ff main to {target}: {e:#}"))?; - lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main") - .await - .map_err(|e| anyhow::anyhow!("read-tree to main: {e:#}"))?; + // Fast-forward applied/main to target_ref + sync the working tree. + // Meta input pins `?ref=main`, so this is what makes nix re-lock to + // the target commit on the prepare_deploy step below. On build + // failure we roll main back to prev_main_sha so a crash leaves the + // agent on its last-good tree. + if let Err(e) = lifecycle::git_update_ref(applied_dir, "refs/heads/main", target_ref).await { + return (Err(anyhow::anyhow!("ff main to {target_ref}: {e:#}")), None); + } + if let Err(e) = lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main").await { + // main is ahead; working tree didn't sync. Roll main back to + // keep the two consistent before bailing. + let _ = lifecycle::git_update_ref(applied_dir, "refs/heads/main", prev_main_sha).await; + return (Err(anyhow::anyhow!("read-tree to main: {e:#}")), None); + } coord.set_queue_step(queue_entry_id, "meta prepare_deploy"); // Phase 1 of the meta two-phase deploy: relock without committing. - crate::meta::prepare_deploy(agent) - .await - .map_err(|e| anyhow::anyhow!("meta prepare_deploy: {e:#}"))?; + if let Err(e) = crate::meta::prepare_deploy(agent).await { + let _ = lifecycle::git_update_ref(applied_dir, "refs/heads/main", prev_main_sha).await; + let _ = lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main").await; + return (Err(anyhow::anyhow!("meta prepare_deploy: {e:#}")), None); + } // Container-level rebuild (or first-time create) against meta#. // Step labels are emitted inside rebuild_no_meta via the callback so @@ -804,18 +726,11 @@ async fn deploy_applied_target( match build_result { Ok(_) => { coord.set_queue_step(queue_entry_id, "finalize deploy"); - // Drop the compensation ref FIRST: from here the deploy is good and - // the tail must not roll `main` back. Ordering it ahead of the tag - // plant is what makes the tail's `deployed/` cross-check a - // second line of defence rather than the only one. - if let Err(e) = lifecycle::git_delete_ref(applied_dir, &rollback_ref(id)).await { - tracing::warn!(%agent, %id, error = ?e, "drop rollback ref after successful deploy failed"); - } let tag = format!("deployed/{id}"); - if let Err(e) = lifecycle::git_tag(applied_dir, &tag, target).await { + if let Err(e) = lifecycle::git_tag(applied_dir, &tag, target_ref).await { tracing::warn!(%agent, %id, error = ?e, "plant deployed tag failed"); } - if let Err(e) = crate::meta::finalize_deploy(agent, target, &tag).await { + if let Err(e) = crate::meta::finalize_deploy(agent, finalize_sha, &tag).await { // The build itself succeeded — meta lock landed but // couldn't be committed. Surface as a soft warn so the // operator can git-commit by hand if they care. @@ -827,18 +742,32 @@ async fn deploy_applied_target( // proposal, agent picks up where it left off with the // new env / packages. coord.kick_agent(agent, "config update applied"); - Ok(()) + (Ok(()), Some(tag)) } Err(e) => { - // Plant the failure marker here rather than in the tail: this is - // the only place that holds the build error to annotate it with. - // The repo-state rollback is the tail's, via the parked ref. let tag = format!("failed/{id}"); let body = format!("{e:#}"); - if let Err(te) = lifecycle::git_tag_annotated(applied_dir, &tag, target, &body).await { + if let Err(te) = + lifecycle::git_tag_annotated(applied_dir, &tag, target_ref, &body).await + { tracing::warn!(%agent, %id, error = ?te, "annotate failed tag failed"); } - Err(e) + // Roll main back to last known-good so the on-disk state + // matches what nixos-container last successfully built. + if let Err(re) = + lifecycle::git_update_ref(applied_dir, "refs/heads/main", prev_main_sha).await + { + tracing::warn!(%agent, %id, error = ?re, "main rollback failed"); + } + if let Err(re) = lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main").await { + tracing::warn!(%agent, %id, error = ?re, "rollback read-tree failed"); + } + // Drop the staged meta lock change so the deploy log + // only ever shows successes. + if let Err(ae) = crate::meta::abort_deploy().await { + tracing::warn!(%agent, %id, error = ?ae, "meta abort_deploy failed"); + } + (Err(e), Some(tag)) } } } diff --git a/hive-c0re/src/coordinator.rs b/hive-c0re/src/coordinator.rs index 0e78a4d7..c24247f9 100644 --- a/hive-c0re/src/coordinator.rs +++ b/hive-c0re/src/coordinator.rs @@ -668,14 +668,12 @@ impl Coordinator { /// Update the `step` label on the currently-running node of DAG /// `id` and (if it actually changed) re-emit the queue snapshot so - /// the dashboard renders the new phase. DAG-id-only surface for the - /// approval-deploy bodies in `actions.rs`, which don't know their node id. - /// The lookup is still exact for them: a deploy DAG is a strictly - /// sequential chain whose resource-holding root sits in `Finishing` while - /// the phases run, so at most one node is ever `Running`. Queue executors - /// that do know their node id use the precise per-node sink in - /// `job_queue::exec` instead. No-op when `id` is `None` (callers not - /// running from the queue) or when nothing is `Running`. + /// the dashboard renders the new phase. DAG-id-only surface for + /// the opaque approval pipeline in `actions.rs`, whose callbacks + /// don't know node ids (its DAGs are single-node, so the lookup is + /// exact); queue executors use the precise per-node sink in + /// `job_queue::exec` instead. No-op when `id` is `None` (callers + /// not running from the queue) or when nothing is `Running`. pub fn set_queue_step(self: &Arc, id: Option, step: &str) { let Some(id) = id else { return }; if self.job_queue.set_step_running(id, step) { diff --git a/hive-c0re/src/forge/pr_merge.rs b/hive-c0re/src/forge/pr_merge.rs index 6b685736..bd259e6e 100644 --- a/hive-c0re/src/forge/pr_merge.rs +++ b/hive-c0re/src/forge/pr_merge.rs @@ -1,5 +1,5 @@ //! PR-based config-flow merge primitives — the forge-side mechanics -//! hive-c0re's deploy apply node (`actions::run_deploy_apply`) orchestrates to +//! hive-c0re's approve-handler (`run_merge_config_pr`) orchestrates to //! land an operator-approved config PR. Part of the operator trust //! boundary; moved verbatim from the `forge` module root. @@ -15,7 +15,7 @@ use super::{CONFIG_ORG, api, core_token, forge_git_url}; // The dashboard-approve-driven flow has hive-c0re verify an operator-approved // config PR, then land it: fast-forward-merge the verified sha into the // protected default branch via the forge merge API (= the merge). These fns are -// the forge-side mechanics the c0re deploy apply node (`run_deploy_apply`) +// the forge-side mechanics the c0re approve-handler (`run_merge_config_pr`) // orchestrates; the orchestration fetches the verified sha into the agent's // applied repo (for the eval-verify) before calling `merge_config_pr_ff`. The // core token is sourced internally (`core_token`), never passed in. `repo` is diff --git a/hive-c0re/src/forge/repos.rs b/hive-c0re/src/forge/repos.rs index 5a6957a8..a9fe91d8 100644 --- a/hive-c0re/src/forge/repos.rs +++ b/hive-c0re/src/forge/repos.rs @@ -907,7 +907,7 @@ async fn apply_operator_branch_protection(repo: &str, token: &str) -> Result<()> /// - **`main` is never directly pushable** — no push is enabled on the /// protected branch, so neither the agent (a write collaborator) nor /// hive-c0re can `git push` it. It only advances via the config-PR merge -/// node (`actions::run_deploy_apply`), which fast-forward-*merges* the reviewed +/// handler (`run_merge_config_pr`), which fast-forward-*merges* the reviewed /// head through the forge merge API (`Do=fast-forward-only`, /// `head_commit_id` pinned to the reviewed sha). /// - **merge is whitelisted to `core`** — only hive-c0re can merge a config PR; diff --git a/hive-c0re/src/job_queue/exec.rs b/hive-c0re/src/job_queue/exec.rs index 6fc5eb36..58f3ef1c 100644 --- a/hive-c0re/src/job_queue/exec.rs +++ b/hive-c0re/src/job_queue/exec.rs @@ -99,10 +99,7 @@ pub(super) async fn run_node(coord: &Arc, claim: &Claim) -> Result< NodeKind::Drain { .. } => run_drain(coord, claim, &ctx).await, NodeKind::WriteDropin { .. } => run_write_dropin(coord, claim).await, NodeKind::WritePermFile { .. } => run_write_perm_file(coord, claim, &ctx).await, - NodeKind::DeployWindow { .. } => run_deploy_window(claim), - NodeKind::MergeVerify { .. } => run_merge_verify(coord, claim).await, - NodeKind::DeployApply { .. } => run_deploy_apply(coord, claim).await, - NodeKind::DeployTail { .. } => run_deploy_tail(coord, claim).await, + NodeKind::ApprovalDeploy { .. } => run_approval_deploy(coord, claim).await, NodeKind::SetWanted { up, .. } => run_set_wanted(coord, claim, *up), // Pure grouping container — no work; completing it lets it reach // `Finishing` so its child template nodes start. The DAG's terminal @@ -581,66 +578,27 @@ async fn run_write_perm_file( Ok(NodeOutput::default()) } -/// The approval id every deploy phase re-reads its approval row by. Fails the -/// node when the DAG carries none, which would mean a `MergeConfigPr` DAG was -/// built without going through `templates::approval_deploy`. -fn deploy_approval_id(claim: &Claim) -> Result { - claim +/// Opaque approval deploy pipeline for `MergeConfigPr`: verify + ff-merge the +/// reviewed PR head, then the container rebuild. The two-phase +/// prepare/finalize/abort meta deploy — and the approval resolution — stay +/// inside `actions.rs` in v1 (design doc §9). +async fn run_approval_deploy(coord: &Arc, claim: &Claim) -> Result { + let approval_id = claim .approval_id - .with_context(|| format!("approval deploy dag {} has no approval_id", claim.dag_id)) -} - -/// The deploy subtree's root: pure resource holder, no work of its own. -/// -/// It exists so the global meta window (plus the agent lease and a build slot) -/// is held continuously across every phase below it. `prepare_deploy` leaves -/// `flake.lock` staged-uncommitted for the whole container build, and any other -/// meta mutation landing inside that span would sweep the staged lock into its -/// own commit and neuter `abort_deploy` — so the window has to outlive any one -/// node, which the `MutexGuard` this replaced could not do. -/// -/// Completing immediately moves it to `Finishing`, which is what starts the -/// children; the resources stay held until the whole subtree settles. -fn run_deploy_window(claim: &Claim) -> Result { - deploy_approval_id(claim)?; - Ok(NodeOutput::default()) -} - -/// Deploy phase 1 — drift gate, fetch, eval-verify. Mutates nothing, so a -/// failure here cancel-cascades the rest of the subtree with the forge and the -/// applied repo exactly as they were. -async fn run_merge_verify(coord: &Arc, claim: &Claim) -> Result { - crate::actions::run_deploy_merge_verify(coord, Some(claim.dag_id), deploy_approval_id(claim)?) + .with_context(|| format!("approval_deploy dag {} has no approval_id", claim.dag_id))?; + // The deploy window covers the whole prepare→finalize span + // (`NodeKind::needs_meta_window`, held by the scheduler for this + // node): `prepare_deploy` stages `flake.lock` uncommitted for the + // entire container build, and no other meta mutation may land inside + // that window (it would sweep the staged lock and neuter + // `abort_deploy`). Holding it as a queue resource — rather than a + // `MutexGuard` that cannot outlive this fn — is what lets increment 2 + // decompose this node into sub-nodes under a window-holding parent. + crate::actions::run_approval_merge_config_pr(coord, Some(claim.dag_id), approval_id) .await .map(|()| NodeOutput::default()) } -/// Deploy phase 2 — the irreversible half: ff-merge, two-phase meta deploy, -/// container rebuild, finalize. -async fn run_deploy_apply(coord: &Arc, claim: &Claim) -> Result { - crate::actions::run_deploy_apply(coord, Some(claim.dag_id), deploy_approval_id(claim)?) - .await - .map(|()| NodeOutput::default()) -} - -/// Deploy compensation + bookkeeping tail. `AfterAny` the apply node, so it -/// runs on every outcome; it is deliberately infallible (see -/// [`crate::actions::run_deploy_tail`]) — a failing tail must not flip an -/// otherwise-successful deploy's DAG state. -/// -/// Takes the agent from the node payload so the tail can still compensate when -/// the approval row is gone (deny race, purge). -async fn run_deploy_tail(coord: &Arc, claim: &Claim) -> Result { - crate::actions::run_deploy_tail( - coord, - Some(claim.dag_id), - claim.kind.agent(), - deploy_approval_id(claim)?, - ) - .await; - Ok(NodeOutput::default()) -} - /// Compute which agents a `nix flake update ` on the meta /// flake affects — the fan-out set for `MetaUpdate` DAGs. Empty /// `inputs` or any input under `hyperhive` → every container; diff --git a/hive-c0re/src/job_queue/mod.rs b/hive-c0re/src/job_queue/mod.rs index 09d90c24..b549d8d2 100644 --- a/hive-c0re/src/job_queue/mod.rs +++ b/hive-c0re/src/job_queue/mod.rs @@ -789,11 +789,8 @@ impl QueueInner { Dep::Resource { .. } => None, }) .collect(); - // Non-derivable per-node payload rides the node that owns it. For a - // deploy that's the subtree root: the phases below it are ordinary - // nodes, and hanging the approval link off all four would render the - // same card four times. - let approval_id = matches!(node.payload, NodeKind::DeployWindow { .. }) + // Non-derivable per-node payload rides the node that owns it. + let approval_id = matches!(node.payload, NodeKind::ApprovalDeploy { .. }) .then_some(meta.approval_id) .flatten(); let inputs = if matches!(node.payload, NodeKind::MetaLock { .. }) { diff --git a/hive-c0re/src/job_queue/model.rs b/hive-c0re/src/job_queue/model.rs index e69c2bd5..d0ba1157 100644 --- a/hive-c0re/src/job_queue/model.rs +++ b/hive-c0re/src/job_queue/model.rs @@ -186,64 +186,11 @@ pub enum NodeKind { /// (commit fused under `META_LOCK`). The payload rides this node — the only /// consumer — rather than the generic DAG container. WritePermFile { agent: String, payload: PermPayload }, - /// Group root of the approval-deploy (`MergeConfigPr`) subtree, and the - /// node that **owns the deploy window**. It performs no work of its own — - /// it exists so the resources it declares (the global - /// [`Resource::MetaWindow`](super::resource::Resource::MetaWindow), the - /// agent lease, a build slot) are held continuously across every child - /// phase, which a per-node acquisition could not guarantee. - /// - /// All three resources are declared *here*, on one node, on purpose. The - /// queue acquires a node's resources atomically (all-or-nothing), so a - /// single multi-resource root can never hold one and block on another — - /// whereas letting a child take the build slot while its parent held the - /// meta window would introduce exactly that pattern, and with it a - /// lock-ordering argument that has to be re-verified on every future edit. - /// Cheap, too: the window has to span the container build regardless (see - /// [`NodeKind::DeployApply`]), so nothing is over-serialised by hoisting - /// the slot and the lease up alongside it. - DeployWindow { agent: String }, - /// Deploy phase 1 — **verify only, mutates nothing.** Drift-gate the - /// approval's PR head, fetch it into the applied repo, and eval-verify the - /// merge head. Any failure here aborts the deploy with the forge state - /// untouched, so it is safely retryable and cancel-safe: nothing downstream - /// has happened yet. - MergeVerify { agent: String }, - /// Deploy phase 2 — everything from the irreversible fast-forward onward: - /// ff-merge the reviewed head to `main` via the forge API, two-phase meta - /// `prepare_deploy`, the container rebuild, then on success the - /// `deployed/` tag + `finalize_deploy`. - /// - /// Still one node in this increment: splitting the tail into - /// `FfMain`/`PrepareDeploy`/rebuild/`FinalizeDeploy` children is the next - /// one. What *is* already split out is the compensation path — see - /// [`NodeKind::DeployTail`]. - DeployApply { agent: String }, - /// Deploy compensation **and bookkeeping** tail — `AfterAny` - /// [`NodeKind::DeployApply`], so it runs on success, failure, and cancel - /// alike, in the same spirit as the rebuild template's tail `Reconcile` - /// ("always runs, decides internally"). It: - /// 1. compensates a merge that landed but was never finalized — roll `main` - /// back, reset the tree, `meta::abort_deploy`, plant `failed/`; - /// 2. mirrors whichever deploy tag got planted to the forge config repo - /// (`forge::push_config`), always, best-effort; - /// 3. posts the failing build log back onto the config PR when the deploy - /// failed, so the manager sees the rejection without leaving the forge. - /// - /// Steps 2 and 3 are why this is `DeployTail` and not `AbortDeploy`: it has - /// work to do on the success path too, and a node name that claims - /// otherwise would be a lie on the dashboard. - /// - /// For (1) it needs no knowledge of how far the deploy got, because that state is - /// parked in the applied repo rather than passed between nodes: - /// `DeployApply` writes the pre-merge `main` sha to - /// `refs/hyperhive/rollback/` before the fast-forward and - /// deletes it once the deploy has been finalized. So the ref existing *is* - /// the "a merge landed and was not finalized" signal, and its absence makes - /// this node a no-op. Parking it in git rather than in a node payload also - /// means it survives a `hive-c0re` restart mid-deploy, which an in-memory - /// queue does not. - DeployTail { agent: String }, + /// Opaque approval deploy pipeline (`MergeConfigPr`): the two-phase + /// prepare/finalize/abort meta deploy stays inside `actions.rs` in v1 — + /// deliberately not + /// modeled as scheduler nodes (see the design doc §9). + ApprovalDeploy { agent: String }, /// Write the agent's durable power intent (`wanted = Up` when `up`, else /// `Offline`) as a first-class DAG node, at the head of a power-op /// template so the downstream `Reconcile` reads it. Replaces the old @@ -296,10 +243,7 @@ impl NodeKind { NodeKind::Drain { .. } => "drain", NodeKind::WriteDropin { .. } => "write_dropin", NodeKind::WritePermFile { .. } => "write_perm_file", - NodeKind::DeployWindow { .. } => "deploy_window", - NodeKind::MergeVerify { .. } => "merge_verify", - NodeKind::DeployApply { .. } => "deploy_apply", - NodeKind::DeployTail { .. } => "deploy_tail", + NodeKind::ApprovalDeploy { .. } => "approval_deploy", NodeKind::SetWanted { .. } => "set_wanted", NodeKind::Dag { .. } => "dag", } @@ -325,10 +269,7 @@ impl NodeKind { | NodeKind::Drain { agent } | NodeKind::WriteDropin { agent } | NodeKind::WritePermFile { agent, .. } - | NodeKind::DeployWindow { agent } - | NodeKind::MergeVerify { agent } - | NodeKind::DeployApply { agent } - | NodeKind::DeployTail { agent } + | NodeKind::ApprovalDeploy { agent } | NodeKind::SetWanted { agent, .. } => agent, NodeKind::MetaLock { .. } | NodeKind::Dag { .. } => "", } @@ -343,7 +284,7 @@ impl NodeKind { | NodeKind::Swap { .. } | NodeKind::Create { .. } | NodeKind::MetaLock { .. } - | NodeKind::DeployWindow { .. } + | NodeKind::ApprovalDeploy { .. } ) } @@ -365,7 +306,7 @@ impl NodeKind { | NodeKind::Signal { .. } | NodeKind::Drain { .. } | NodeKind::WriteDropin { .. } - | NodeKind::DeployWindow { .. } + | NodeKind::ApprovalDeploy { .. } | NodeKind::SetWanted { .. } ) } @@ -396,7 +337,7 @@ impl NodeKind { | NodeKind::Provision { .. } | NodeKind::MetaLock { .. } | NodeKind::WritePermFile { .. } - | NodeKind::DeployWindow { .. } + | NodeKind::ApprovalDeploy { .. } ) } } diff --git a/hive-c0re/src/job_queue/templates.rs b/hive-c0re/src/job_queue/templates.rs index 0fd03da8..f2b3747c 100644 --- a/hive-c0re/src/job_queue/templates.rs +++ b/hive-c0re/src/job_queue/templates.rs @@ -140,24 +140,10 @@ pub fn rebuild(agent: &str, source: Source, reason: String, relock: bool) -> Dag } } -/// Approval-driven deploy (`MergeConfigPr`) as a phase subtree rather than the -/// single opaque node it used to be. Structure: -/// - `DeployWindow` (0, **root**): the resource holder — global meta window, -/// agent lease, build slot — held across every child below. No work of its -/// own; it reaches `Finishing` immediately and the children run inside it. -/// - `MergeVerify` (1, child): drift-gate + fetch + eval-verify. Mutates -/// nothing, so a failure here cancel-cascades its siblings with the forge and -/// the applied repo exactly as they were. -/// - `DeployApply` (2, child, `AfterOk` `MergeVerify`): the irreversible half — -/// ff-merge, `prepare_deploy`, rebuild, `finalize_deploy`. -/// - `DeployTail` (3, child, `AfterAny` `DeployApply`): the compensation + -/// bookkeeping tail — rollback when a merge landed unfinalized, forge tag -/// mirror, PR failure comment (see [`NodeKind::DeployTail`]). -/// -/// The window still spans the container build, as it must: `prepare_deploy` -/// leaves `flake.lock` staged-uncommitted for the build's whole duration. +/// Approval-driven deploy (`MergeConfigPr`): the whole two-phase pipeline +/// stays one opaque node in v1 (design doc §9) — wire-visible as a `rebuild` +/// card like today. pub fn approval_deploy(agent: &str, approval_id: i64, reason: String) -> DagSpec { - let a = || agent.to_owned(); DagSpec { template: Template::Rebuild, source: Source::Approval, @@ -165,19 +151,12 @@ pub fn approval_deploy(agent: &str, approval_id: i64, reason: String) -> DagSpec approval_id: Some(approval_id), inputs: Vec::new(), transient: Some(TransientKind::Rebuilding), - nodes: vec![ - node(NodeKind::DeployWindow { agent: a() }, Vec::new()), - child(0, NodeKind::MergeVerify { agent: a() }, Vec::new()), - child(0, NodeKind::DeployApply { agent: a() }, after_ok(1)), - child( - 0, - NodeKind::DeployTail { agent: a() }, - vec![Dep { - on: 2, - when: DepWhen::AfterAny, - }], - ), - ], + nodes: vec![node( + NodeKind::ApprovalDeploy { + agent: agent.to_owned(), + }, + Vec::new(), + )], } } diff --git a/hive-c0re/src/job_queue/tests.rs b/hive-c0re/src/job_queue/tests.rs index 53d2ee55..ba46586c 100644 --- a/hive-c0re/src/job_queue/tests.rs +++ b/hive-c0re/src/job_queue/tests.rs @@ -926,88 +926,6 @@ fn cancelled_dag_finalizes_with_terminal_rollup() { ); } -// ---- approval deploy subtree ---- - -/// The config-PR deploy is a subtree, not one opaque node. The -/// resource-holding root completes immediately (its `Finishing` state is the -/// parent gate that releases the children), then the phases run strictly in -/// order — and the `AfterAny` tail still runs when the irreversible half fails, -/// because it's the node that compensates for it. -#[test] -fn deploy_dag_runs_phases_in_order_and_tails_a_failed_apply() { - let q = JobQueue::new(1); - let id = submit( - &q, - templates::approval_deploy("agent-a", 7, "approval #7".to_owned()), - ); - - let root = claim_one(&q); - assert!( - matches!(root.kind, NodeKind::DeployWindow { .. }), - "root claims first: it holds the meta window for the whole subtree" - ); - q.complete_node(id, root.node_id, Ok(())); - - let verify = claim_one(&q); - assert!(matches!(verify.kind, NodeKind::MergeVerify { .. })); - q.complete_node(id, verify.node_id, Ok(())); - - let apply = claim_one(&q); - assert!(matches!(apply.kind, NodeKind::DeployApply { .. })); - q.complete_node( - id, - apply.node_id, - Err("nixos-container update blew up".into()), - ); - - let tail = claim_one(&q); - assert!( - matches!(tail.kind, NodeKind::DeployTail { .. }), - "AfterAny tail runs on a failed apply — that's the whole point of it" - ); - q.complete_node(id, tail.node_id, Ok(())); - - let summary = q.terminal_summary(id).expect("dag terminal"); - assert_eq!( - summary.state, - State::Failed, - "an Ok tail must not launder a failed deploy into a success" - ); - assert_eq!(summary.approval_id, Some(7)); -} - -/// A pre-merge rejection (drift gate, eval failure) cancel-cascades the -/// irreversible half via its `AfterOk` edge, but the tail is still reached — -/// it owns the forge mirror, not just compensation. -#[test] -fn deploy_dag_skips_apply_but_still_runs_tail_when_verify_fails() { - let q = JobQueue::new(1); - let id = submit( - &q, - templates::approval_deploy("agent-a", 9, "approval #9".to_owned()), - ); - - let root = claim_one(&q); - q.complete_node(id, root.node_id, Ok(())); - let verify = claim_one(&q); - q.complete_node( - id, - verify.node_id, - Err("PR head drifted since review".into()), - ); - - let tail = claim_one(&q); - assert!( - matches!(tail.kind, NodeKind::DeployTail { .. }), - "apply is cancel-cascaded, so the tail is the next claimable node" - ); - q.complete_node(id, tail.node_id, Ok(())); - - let summary = q.terminal_summary(id).expect("dag terminal"); - assert_eq!(summary.state, State::Failed); - assert_eq!(summary.approval_id, Some(9)); -} - // ---- steps, build logs, history ---- #[test] diff --git a/hive-c0re/src/lifecycle/git.rs b/hive-c0re/src/lifecycle/git.rs index 3461c956..a54299d9 100644 --- a/hive-c0re/src/lifecycle/git.rs +++ b/hive-c0re/src/lifecycle/git.rs @@ -167,13 +167,3 @@ pub async fn git_read_tree_reset(dir: &Path, target: &str) -> Result<()> { pub async fn git_update_ref(dir: &Path, refname: &str, target: &str) -> Result<()> { git(dir, &["update-ref", refname, target]).await } - -/// Delete a ref. The counterpart to [`git_update_ref`] for the bookkeeping -/// refs a deploy parks in the applied repo (`refs/hyperhive/rollback/`, -/// which records the pre-merge `main` so the deploy tail can compensate a -/// merge that landed but never finalized). `update-ref -d` is a no-op-free -/// delete: it errors if the ref does not exist, so callers that treat absence -/// as "nothing to undo" should check with [`git_rev_parse`] first. -pub async fn git_delete_ref(dir: &Path, refname: &str) -> Result<()> { - git(dir, &["update-ref", "-d", refname]).await -} diff --git a/hive-c0re/src/lifecycle/mod.rs b/hive-c0re/src/lifecycle/mod.rs index a1a2347f..1419f354 100644 --- a/hive-c0re/src/lifecycle/mod.rs +++ b/hive-c0re/src/lifecycle/mod.rs @@ -7,8 +7,8 @@ mod setup; mod tests; pub use git::{ - git, git_command, git_delete_ref, git_read_tree_reset, git_rev_parse, git_tag, - git_tag_annotated, git_update_ref, + git, git_command, git_read_tree_reset, git_rev_parse, git_tag, git_tag_annotated, + git_update_ref, }; pub use host_config::write_dropins; pub use setup::{ diff --git a/hive-c0re/src/lifecycle/setup.rs b/hive-c0re/src/lifecycle/setup.rs index eef1f6ec..39d037db 100644 --- a/hive-c0re/src/lifecycle/setup.rs +++ b/hive-c0re/src/lifecycle/setup.rs @@ -71,7 +71,7 @@ async fn ensure_applied_remote(proposed_dir: &Path, name: &str) -> Result<()> { /// proposed's initial commit in via `git fetch`, tag it `deployed/0`. /// This is the *only* time hive-c0re reads from `proposed` for an /// agent — subsequent config changes are fetched from the reviewed -/// forge PR head at merge time (see `actions::run_deploy_apply`). +/// forge PR head at merge time (see `actions::run_merge_config_pr`). /// /// `proposed_dir` is `None` on rebuild paths where the repo already /// exists — we just verify it's the right shape and bail otherwise. diff --git a/hive-c0re/src/socket_server/config_approvals.rs b/hive-c0re/src/socket_server/config_approvals.rs index 6cf7617b..95f1caaa 100644 --- a/hive-c0re/src/socket_server/config_approvals.rs +++ b/hive-c0re/src/socket_server/config_approvals.rs @@ -89,7 +89,7 @@ pub(super) fn handle_request_update_meta_inputs( /// dashboard shows the pending card immediately. /// /// The PR head sha is stored as `fetched_sha` on the approval row — the -/// "reviewed sha" the deploy's `MergeVerify` node drift-gates +/// "reviewed sha" the approve handler (`run_merge_config_pr`) drift-gates /// against before doing anything irreversible. This does NOT fetch the commit /// into the applied repo at submission time (that happens inside the approve /// handler, step 2, after the drift check). No flake pre-flight either —