actions: split the config-PR deploy into verify / apply / tail
`run_approval_merge_config_pr` and `run_merge_config_pr` are gone; the three phases are `run_deploy_merge_verify` (drift gate, fetch, verify — mutates nothing), `run_deploy_apply` (merge + build) and `run_deploy_tail` (compensation + push). The rollback state is a git ref in the applied repo (`refs/hyperhive/rollback/<approval-id>`) rather than a value handed between nodes, because hive-c0re can restart between the apply and the tail and the tail still has to know what to undo. Rolling `main` back on a *successful* deploy is the worst thing the tail can do, so it is guarded twice: the apply drops the rollback ref before it plants `deployed/<id>`, and the tail refuses to compensate at all if `deployed/<id>` resolves. It takes two independent git failures to get there. `run_deploy_tail` returns nothing and warns on every error — a failing compensation must not mask the deploy's own verdict, which the terminal hook takes from the DAG's roll-up.
This commit is contained in:
parent
8899c9f355
commit
27ecda7b13
2 changed files with 393 additions and 280 deletions
|
|
@ -18,15 +18,14 @@ use crate::lifecycle;
|
||||||
/// (operator no longer blocks on a 30-90s spinner for `MergeConfigPr`).
|
/// (operator no longer blocks on a 30-90s spinner for `MergeConfigPr`).
|
||||||
///
|
///
|
||||||
/// Dispatch:
|
/// Dispatch:
|
||||||
/// - `MergeConfigPr` → a single-node `ApprovalDeploy`
|
/// - `MergeConfigPr` → a `DeployWindow` DAG (`MergeVerify → DeployApply →
|
||||||
/// DAG (the two-phase meta deploy stays opaque in v1; ~30-90s)
|
/// DeployTail` under a resource-holding root; ~30-90s)
|
||||||
/// - `UpdateMetaInputs` → a `MetaUpdate` DAG (fan-out on completion)
|
/// - `UpdateMetaInputs` → a `MetaUpdate` DAG (fan-out on completion)
|
||||||
/// - `Spawn` → a `Spawn` DAG (`Create → WriteDropin → Reconcile`)
|
/// - `Spawn` → a `Spawn` DAG (`Create → WriteDropin → Reconcile`)
|
||||||
/// - `InitConfig` → inline (<1s; queue card would be noise)
|
/// - `InitConfig` → inline (<1s; queue card would be noise)
|
||||||
///
|
///
|
||||||
/// `ApprovalDeploy` resolves the approval inside its pipeline; the
|
/// Every queued kind — deploys included — resolves its approval row via
|
||||||
/// `MetaUpdate` / `Spawn` DAGs resolve via [`resolve_approval_dag`]
|
/// [`resolve_approval_dag`] when the DAG settles terminal.
|
||||||
/// when their DAG settles terminal.
|
|
||||||
pub async fn approve(coord: Arc<Coordinator>, id: i64) -> Result<()> {
|
pub async fn approve(coord: Arc<Coordinator>, id: i64) -> Result<()> {
|
||||||
let approval = coord.approvals.mark_approved(id)?;
|
let approval = coord.approvals.mark_approved(id)?;
|
||||||
tracing::info!(
|
tracing::info!(
|
||||||
|
|
@ -104,10 +103,10 @@ pub async fn approve(coord: Arc<Coordinator>, id: i64) -> Result<()> {
|
||||||
}
|
}
|
||||||
ApprovalKind::MergeConfigPr => {
|
ApprovalKind::MergeConfigPr => {
|
||||||
// The work ends in a container rebuild, so route it through the
|
// The work ends in a container rebuild, so route it through the
|
||||||
// rebuild queue. The queue worker dispatches MergeConfigPr
|
// rebuild queue. The queue worker dispatches the deploy DAG's
|
||||||
// approvals to `run_merge_config_pr` (verify the reviewed PR head,
|
// nodes to `run_deploy_merge_verify` (drift gate + eval),
|
||||||
// ff the forge config repo's main to it, mark merged, then the
|
// `run_deploy_apply` (ff-merge + rebuild) and `run_deploy_tail`
|
||||||
// deploy tail).
|
// (compensation + forge mirror).
|
||||||
enqueue_approval_rebuild(
|
enqueue_approval_rebuild(
|
||||||
&coord,
|
&coord,
|
||||||
approval.agent.as_str(),
|
approval.agent.as_str(),
|
||||||
|
|
@ -119,10 +118,10 @@ pub async fn approve(coord: Arc<Coordinator>, id: i64) -> Result<()> {
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Submit the single-node `ApprovalDeploy` DAG tied to an approval id.
|
/// Submit the deploy DAG tied to an approval id. Used by the `MergeConfigPr`
|
||||||
/// Used by the `MergeConfigPr` dispatch arm — the work ends in a container
|
/// dispatch arm — the work ends in a container rebuild routed through the
|
||||||
/// rebuild routed through the queue; the node executor runs
|
/// queue. See [`crate::job_queue::templates::approval_deploy`] for the node
|
||||||
/// `run_merge_config_pr`.
|
/// shape; the executor dispatches each node to the `run_deploy_*` bodies below.
|
||||||
fn enqueue_approval_rebuild(
|
fn enqueue_approval_rebuild(
|
||||||
coord: &Arc<Coordinator>,
|
coord: &Arc<Coordinator>,
|
||||||
agent: &str,
|
agent: &str,
|
||||||
|
|
@ -142,45 +141,226 @@ fn enqueue_approval_rebuild(
|
||||||
coord.emit_rebuild_queue_snapshot();
|
coord.emit_rebuild_queue_snapshot();
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Worker entry point for `ApprovalKind::MergeConfigPr` queue entries — the
|
/// Ref under which [`run_deploy_apply`] parks the pre-merge `applied/main`
|
||||||
/// config-change flow's deploy worker. Re-fetches the approval row, runs the
|
/// sha, for [`run_deploy_tail`] to compensate with.
|
||||||
/// merge pipeline, and fires
|
///
|
||||||
/// `ApprovalResolved` + the `Rebuilt` lifecycle event via `finish_approval`.
|
/// Deliberately a *git ref in the applied repo* rather than an in-memory value
|
||||||
/// `run_merge_config_pr` already fast-forwarded the forge repo's `main` to the
|
/// handed between nodes: hive-c0re can restart between the apply and the tail,
|
||||||
/// reviewed head (that IS the merge), so `push_config`'s `main` refspec is a
|
/// and the whole point of splitting the deploy is that the tail still knows
|
||||||
/// no-op — but it still mirrors the `deployed/<id>` / `failed/<id>` tag the
|
/// what to undo when it does. The ref's existence IS the "a merge landed but
|
||||||
/// deploy tail plants onto the merged sha, giving the merged commit a
|
/// hasn't been confirmed good yet" flag — [`run_deploy_apply`] drops it the
|
||||||
/// forge-visible deploy marker. A `MergeConfigPr` is never a first spawn (the
|
/// moment the rebuild succeeds.
|
||||||
/// agent already exists).
|
fn rollback_ref(approval_id: i64) -> String {
|
||||||
pub async fn run_approval_merge_config_pr(
|
format!("refs/hyperhive/rollback/{approval_id}")
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Everything a deploy node needs, re-derived from sqlite on each node rather
|
||||||
|
/// than cached across the DAG. Nothing here is *computed* by an earlier node —
|
||||||
|
/// `pr` and `reviewed` are fields of the approval row the operator signed off
|
||||||
|
/// on — so re-reading is both cheap and the authoritative source of truth.
|
||||||
|
struct DeployCtx {
|
||||||
|
approval: hive_sh4re::Approval,
|
||||||
|
/// PR number, parsed from `approval.commit_ref`.
|
||||||
|
pr: u64,
|
||||||
|
/// The PR head sha the operator reviewed (`approval.fetched_sha`).
|
||||||
|
reviewed: String,
|
||||||
|
agent_dir: std::path::PathBuf,
|
||||||
|
applied_dir: std::path::PathBuf,
|
||||||
|
/// The agent's forge config repo (`<owner>/<name>`).
|
||||||
|
repo: String,
|
||||||
|
}
|
||||||
|
|
||||||
|
fn deploy_ctx(coord: &Coordinator, approval_id: i64) -> Result<DeployCtx> {
|
||||||
|
let approval = fetch_approval_for_worker(coord, approval_id, ApprovalKind::MergeConfigPr)?;
|
||||||
|
let pr: u64 = approval.commit_ref.parse().map_err(|e| {
|
||||||
|
anyhow::anyhow!(
|
||||||
|
"parse PR number from commit_ref {:?}: {e}",
|
||||||
|
approval.commit_ref
|
||||||
|
)
|
||||||
|
})?;
|
||||||
|
let reviewed = approval.fetched_sha.clone().ok_or_else(|| {
|
||||||
|
anyhow::anyhow!("merge config pr approval {approval_id} has no reviewed head sha")
|
||||||
|
})?;
|
||||||
|
Ok(DeployCtx {
|
||||||
|
pr,
|
||||||
|
reviewed,
|
||||||
|
agent_dir: crate::paths::agent_runtime_dir(approval.agent.as_str()),
|
||||||
|
applied_dir: crate::paths::applied_dir(approval.agent.as_str()),
|
||||||
|
repo: crate::forge::config_repo(approval.agent.as_str()),
|
||||||
|
approval,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `MergeVerify` node body — everything that can say "no" before anything is
|
||||||
|
/// mutated. `approval.commit_ref` is the PR number; `approval.fetched_sha` is
|
||||||
|
/// the PR head the operator reviewed. Steps:
|
||||||
|
/// 1. drift gate — re-read the live PR head; if it moved since review, abort
|
||||||
|
/// (the operator must re-review the new head);
|
||||||
|
/// 2. fetch the reviewed head into the applied repo so later git ops resolve
|
||||||
|
/// it locally;
|
||||||
|
/// 3. eval-verify the reviewed commit against the meta flake.
|
||||||
|
///
|
||||||
|
/// Nothing here needs undoing on failure: the fetch only adds objects, and
|
||||||
|
/// `main` doesn't move. That's the whole reason this is its own node — a
|
||||||
|
/// failure at this stage leaves [`run_deploy_tail`] with no ref to compensate.
|
||||||
|
///
|
||||||
|
/// # Errors
|
||||||
|
///
|
||||||
|
/// Returns an error if the approval can't be loaded, if the live PR head has
|
||||||
|
/// drifted from the reviewed sha, if fetching that head into the applied repo
|
||||||
|
/// fails, or if the eval-verify of the reviewed commit fails. Every one of
|
||||||
|
/// these leaves the forge and `main` untouched, so the node is safely
|
||||||
|
/// retryable.
|
||||||
|
pub async fn run_deploy_merge_verify(
|
||||||
coord: &Arc<Coordinator>,
|
coord: &Arc<Coordinator>,
|
||||||
queue_entry_id: Option<u64>,
|
queue_entry_id: Option<u64>,
|
||||||
approval_id: i64,
|
approval_id: i64,
|
||||||
) -> Result<()> {
|
) -> Result<()> {
|
||||||
let approval = fetch_approval_for_worker(coord, approval_id, ApprovalKind::MergeConfigPr)?;
|
let ctx = deploy_ctx(coord, approval_id)?;
|
||||||
let agent_dir = crate::paths::agent_runtime_dir(approval.agent.as_str());
|
let pr = ctx.pr;
|
||||||
let applied_dir = crate::paths::applied_dir(approval.agent.as_str());
|
let reviewed = ctx.reviewed.as_str();
|
||||||
// Captured up front to scope the failure-comment's build-log lookup to
|
|
||||||
// rows this deploy produced (see `post_merge_failure_to_pr`).
|
// 1. Drift gate: the live PR head must still equal what was reviewed.
|
||||||
let since_ts = hive_sh4re::wire_time::now_unix();
|
coord.set_queue_step(queue_entry_id, "verify PR head");
|
||||||
coord.set_queue_step(queue_entry_id, "merge config pr");
|
let head = crate::forge::pr_head_sha(&ctx.repo, pr)
|
||||||
let (result, terminal_tag) =
|
.await
|
||||||
run_merge_config_pr(coord, &approval, &agent_dir, &applied_dir, queue_entry_id).await;
|
.map_err(|e| anyhow::anyhow!("read PR #{pr} head: {e}"))?;
|
||||||
// Mirror the deploy bookkeeping tag (`deployed/<id>` or `failed/<id>`) the
|
if head != reviewed {
|
||||||
// deploy tail planted onto the merged sha to the forge config repo, so the
|
bail!(
|
||||||
// merged commit carries a forge-visible deploy marker. `main` is already
|
"PR #{pr} head drifted since review (reviewed {reviewed}, now {head}); re-review before merging"
|
||||||
// ff'd by the merge, so only the tag refspec actually lands; best-effort,
|
);
|
||||||
// never fails the approval.
|
}
|
||||||
|
|
||||||
|
// 2. Fetch the reviewed head into applied so ff/verify/deploy resolve it.
|
||||||
|
coord.set_queue_step(queue_entry_id, "fetch PR head");
|
||||||
|
crate::forge::fetch_pr_head_into_applied(&ctx.repo, pr)
|
||||||
|
.await
|
||||||
|
.map_err(|e| anyhow::anyhow!("fetch PR #{pr} head into applied: {e}"))?;
|
||||||
|
|
||||||
|
// 3. Eval-verify BEFORE the irreversible merge (bad nix fails fast here).
|
||||||
|
coord.set_queue_step(queue_entry_id, "verify proposal (eval)");
|
||||||
|
crate::meta::verify_commit(ctx.approval.agent.as_str(), &ctx.applied_dir, reviewed)
|
||||||
|
.await
|
||||||
|
.map_err(|e| anyhow::anyhow!("verify merge head {reviewed}: {e:#}"))?;
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `DeployApply` node body — the irreversible half. Parks the rollback ref,
|
||||||
|
/// fast-forward-merges the PR (THE merge), then runs the deploy proper.
|
||||||
|
///
|
||||||
|
/// The ref is parked *before* the merge, so a hive-c0re crash anywhere from
|
||||||
|
/// here on still leaves [`run_deploy_tail`] enough to undo. If the merge itself
|
||||||
|
/// fails, `main` never moved and the tail's compensation is a no-op against the
|
||||||
|
/// same sha — harmless, and cheaper than trying to be clever about it.
|
||||||
|
///
|
||||||
|
/// # Errors
|
||||||
|
///
|
||||||
|
/// Returns an error if the approval can't be loaded, if reading or parking the
|
||||||
|
/// pre-merge `main` sha fails, if the forge refuses the fast-forward merge
|
||||||
|
/// (including a head that drifted between verify and merge), or if the deploy
|
||||||
|
/// of the merged target fails. From the merge onward a failure is *not*
|
||||||
|
/// retryable on its own — [`run_deploy_tail`] runs `AfterAny` to compensate.
|
||||||
|
pub async fn run_deploy_apply(
|
||||||
|
coord: &Arc<Coordinator>,
|
||||||
|
queue_entry_id: Option<u64>,
|
||||||
|
approval_id: i64,
|
||||||
|
) -> Result<()> {
|
||||||
|
let ctx = deploy_ctx(coord, approval_id)?;
|
||||||
|
let agent = ctx.approval.agent.as_str();
|
||||||
|
let pr = ctx.pr;
|
||||||
|
|
||||||
|
let prev_main = lifecycle::git_rev_parse(&ctx.applied_dir, "refs/heads/main")
|
||||||
|
.await
|
||||||
|
.map_err(|e| anyhow::anyhow!("read applied/main: {e:#}"))?;
|
||||||
|
lifecycle::git_update_ref(&ctx.applied_dir, &rollback_ref(approval_id), &prev_main)
|
||||||
|
.await
|
||||||
|
.map_err(|e| anyhow::anyhow!("park rollback ref for approval {approval_id}: {e:#}"))?;
|
||||||
|
|
||||||
|
// THE merge: fast-forward-only merge the reviewed head to `main` via the
|
||||||
|
// forge API, pinned to the reviewed sha (`head_commit_id`). This one call
|
||||||
|
// both advances `main` to the reviewed head and marks the PR merged — no
|
||||||
|
// direct push to the protected branch. A failure here means `main` was NOT
|
||||||
|
// advanced, so it's fatal: we must not deploy a head the forge didn't merge.
|
||||||
|
coord.set_queue_step(queue_entry_id, "fast-forward-merge PR");
|
||||||
|
match crate::forge::merge_config_pr_ff(&ctx.repo, pr, &ctx.reviewed).await {
|
||||||
|
Ok(()) => {}
|
||||||
|
Err(crate::forge::ForgeMergeError::HeadDrift { expected, actual }) => bail!(
|
||||||
|
"PR #{pr} head drifted before merge (reviewed {expected}, now {actual}); re-review before merging"
|
||||||
|
),
|
||||||
|
Err(e) => bail!("ff-merge PR #{pr}: {e}"),
|
||||||
|
}
|
||||||
|
|
||||||
|
deploy_applied_target(
|
||||||
|
coord,
|
||||||
|
agent,
|
||||||
|
&ctx.agent_dir,
|
||||||
|
&ctx.applied_dir,
|
||||||
|
&ctx.reviewed,
|
||||||
|
approval_id,
|
||||||
|
queue_entry_id,
|
||||||
|
)
|
||||||
|
.await
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `DeployTail` node body — compensation + bookkeeping, `AfterAny` the apply
|
||||||
|
/// node so it runs on every outcome including a cancel-cascade. Infallible by
|
||||||
|
/// construction: it is the recovery step, so it has nothing to hand a failure
|
||||||
|
/// to. Every fallible call inside warns and continues.
|
||||||
|
///
|
||||||
|
/// 1. If the rollback ref survived, the deploy did not confirm good: roll
|
||||||
|
/// `applied/main` back to the parked sha, resync the working tree, and drop
|
||||||
|
/// the staged meta lock so the deploy log only ever shows successes.
|
||||||
|
/// 2. Mirror the agent's config repo to the forge. `main` is already ff'd by
|
||||||
|
/// the merge, so only the `deployed/<id>` / `failed/<id>` tag refspec
|
||||||
|
/// actually lands — that's what gives the merged commit a forge-visible
|
||||||
|
/// deploy marker.
|
||||||
|
///
|
||||||
|
/// Takes `agent` from the node payload rather than the approval row so it still
|
||||||
|
/// works if the row vanished underneath the DAG (deny race, purge).
|
||||||
|
pub async fn run_deploy_tail(
|
||||||
|
coord: &Arc<Coordinator>,
|
||||||
|
queue_entry_id: Option<u64>,
|
||||||
|
agent: &str,
|
||||||
|
approval_id: i64,
|
||||||
|
) {
|
||||||
|
let applied_dir = crate::paths::applied_dir(agent);
|
||||||
|
let rollback = rollback_ref(approval_id);
|
||||||
|
if let Ok(prev_main) = lifecycle::git_rev_parse(&applied_dir, &rollback).await {
|
||||||
|
// Belt and braces: `run_deploy_apply` drops the ref before it plants
|
||||||
|
// `deployed/<id>`, so seeing both means the *delete* failed on an
|
||||||
|
// otherwise-successful deploy. Rolling back there would be the worst
|
||||||
|
// outcome this node can produce, so the tag wins.
|
||||||
|
if lifecycle::git_rev_parse(&applied_dir, &format!("deployed/{approval_id}"))
|
||||||
|
.await
|
||||||
|
.is_ok()
|
||||||
|
{
|
||||||
|
tracing::warn!(
|
||||||
|
%agent, approval_id,
|
||||||
|
"deploy tail: rollback ref outlived a successful deploy; dropping it without compensating"
|
||||||
|
);
|
||||||
|
} else {
|
||||||
|
coord.set_queue_step(queue_entry_id, "roll back applied/main");
|
||||||
|
if let Err(e) =
|
||||||
|
lifecycle::git_update_ref(&applied_dir, "refs/heads/main", &prev_main).await
|
||||||
|
{
|
||||||
|
tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: main rollback failed");
|
||||||
|
}
|
||||||
|
if let Err(e) = lifecycle::git_read_tree_reset(&applied_dir, "refs/heads/main").await {
|
||||||
|
tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: rollback read-tree failed");
|
||||||
|
}
|
||||||
|
if let Err(e) = crate::meta::abort_deploy().await {
|
||||||
|
tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: meta abort_deploy failed");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if let Err(e) = lifecycle::git_delete_ref(&applied_dir, &rollback).await {
|
||||||
|
tracing::warn!(%agent, approval_id, error = ?e, "deploy tail: drop rollback ref failed");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
coord.set_queue_step(queue_entry_id, "forge push");
|
coord.set_queue_step(queue_entry_id, "forge push");
|
||||||
if let Err(e) = crate::forge::push_config(approval.agent.as_str()).await {
|
if let Err(e) = crate::forge::push_config(agent).await {
|
||||||
tracing::warn!(agent = %approval.agent, error = ?e, "forge: push_config after merge failed");
|
tracing::warn!(%agent, error = ?e, "forge: push_config after merge failed");
|
||||||
}
|
}
|
||||||
// On a failed deploy, surface the failing build log back onto the PR so
|
|
||||||
// the manager sees why it was rejected without leaving the forge.
|
|
||||||
if let Err(e) = &result {
|
|
||||||
post_merge_failure_to_pr(coord, &approval, since_ts, e).await;
|
|
||||||
}
|
|
||||||
finish_approval(coord, &approval, result, terminal_tag)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Max stderr bytes to inline in a PR failure comment. Keeps the comment
|
/// Max stderr bytes to inline in a PR failure comment. Keeps the comment
|
||||||
|
|
@ -194,21 +374,25 @@ const PR_FAIL_LOG_TAIL_BYTES: usize = 4000;
|
||||||
/// approval-resolution path.
|
/// approval-resolution path.
|
||||||
///
|
///
|
||||||
/// The failing `build_log` row is located heuristically: the most recent `fail`
|
/// The failing `build_log` row is located heuristically: the most recent `fail`
|
||||||
/// row for this agent that started at/after `since_ts` (the caller's function
|
/// row for this agent that started at/after the approval was decided (i.e. when
|
||||||
/// entry). Because deploys are serialised per agent through the queue, that is
|
/// its deploy DAG was submitted). Because deploys are serialised per agent
|
||||||
/// the step which just failed — `verify`, `prepare-deploy`, `prebuild`, or the
|
/// through the queue, that is the step which just failed — `verify`,
|
||||||
/// container rebuild. Pre-build failures (drift gate, fetch) create no `build_log`
|
/// `prepare-deploy`, `prebuild`, or the container rebuild. Pre-build failures
|
||||||
/// row, so the comment then carries only the error text.
|
/// (drift gate, fetch) create no `build_log` row, so the comment then carries
|
||||||
|
/// only the error text.
|
||||||
async fn post_merge_failure_to_pr(
|
async fn post_merge_failure_to_pr(
|
||||||
coord: &Arc<Coordinator>,
|
coord: &Arc<Coordinator>,
|
||||||
approval: &hive_sh4re::Approval,
|
approval: &hive_sh4re::Approval,
|
||||||
since_ts: i64,
|
|
||||||
err: &anyhow::Error,
|
err: &anyhow::Error,
|
||||||
) {
|
) {
|
||||||
let Ok(pr) = approval.commit_ref.parse::<u64>() else {
|
let Ok(pr) = approval.commit_ref.parse::<u64>() else {
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
let repo = crate::forge::config_repo(approval.agent.as_str());
|
let repo = crate::forge::config_repo(approval.agent.as_str());
|
||||||
|
let since_ts = approval
|
||||||
|
.resolved_at
|
||||||
|
.unwrap_or(approval.requested_at)
|
||||||
|
.timestamp();
|
||||||
|
|
||||||
let log_section = coord
|
let log_section = coord
|
||||||
.build_logs
|
.build_logs
|
||||||
|
|
@ -253,131 +437,6 @@ fn tail_bytes(s: &str, max_bytes: usize) -> String {
|
||||||
format!("[… truncated …]\n{}", &s[start..])
|
format!("[… truncated …]\n{}", &s[start..])
|
||||||
}
|
}
|
||||||
|
|
||||||
/// PR-merge config pipeline. `approval.commit_ref` is the PR number;
|
|
||||||
/// `approval.fetched_sha` is the PR head sha the operator reviewed. Steps:
|
|
||||||
/// 1. drift gate — re-read the live PR head; if it moved since review, abort
|
|
||||||
/// WITHOUT mutating anything (the operator must re-review the new head);
|
|
||||||
/// 2. fetch the reviewed head into the applied repo so later git ops resolve
|
|
||||||
/// it locally;
|
|
||||||
/// 3. eval-verify the reviewed commit against the meta flake BEFORE the
|
|
||||||
/// irreversible push;
|
|
||||||
/// 4. fast-forward the forge repo's `main` to the reviewed head — THE merge;
|
|
||||||
/// 5. mark the PR merged (best-effort: `main` is already at the head, so a
|
|
||||||
/// failure here is logged, not fatal);
|
|
||||||
/// 6. run the shared deploy tail (`deploy_applied_target`): ff applied/main,
|
|
||||||
/// meta deploy, container rebuild, finalize/rollback.
|
|
||||||
///
|
|
||||||
/// Returns `(build result, terminal tag)` like `deploy_applied_target`. Any
|
|
||||||
/// pre-merge abort returns `Err` with no mutation; the operator re-reviews.
|
|
||||||
async fn run_merge_config_pr(
|
|
||||||
coord: &Arc<Coordinator>,
|
|
||||||
approval: &hive_sh4re::Approval,
|
|
||||||
agent_dir: &std::path::Path,
|
|
||||||
applied_dir: &std::path::Path,
|
|
||||||
queue_entry_id: Option<u64>,
|
|
||||||
) -> (Result<()>, Option<String>) {
|
|
||||||
let id = approval.id;
|
|
||||||
let pr: u64 = match approval.commit_ref.parse() {
|
|
||||||
Ok(n) => n,
|
|
||||||
Err(e) => {
|
|
||||||
return (
|
|
||||||
Err(anyhow::anyhow!(
|
|
||||||
"parse PR number from commit_ref {:?}: {e}",
|
|
||||||
approval.commit_ref
|
|
||||||
)),
|
|
||||||
None,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
};
|
|
||||||
let reviewed = match approval.fetched_sha.as_deref() {
|
|
||||||
Some(s) => s.to_owned(),
|
|
||||||
None => {
|
|
||||||
return (
|
|
||||||
Err(anyhow::anyhow!(
|
|
||||||
"merge config pr approval {id} has no reviewed head sha"
|
|
||||||
)),
|
|
||||||
None,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
};
|
|
||||||
let repo = crate::forge::config_repo(approval.agent.as_str());
|
|
||||||
|
|
||||||
// 1. Drift gate: the live PR head must still equal what was reviewed.
|
|
||||||
coord.set_queue_step(queue_entry_id, "verify PR head");
|
|
||||||
let head = match crate::forge::pr_head_sha(&repo, pr).await {
|
|
||||||
Ok(h) => h,
|
|
||||||
Err(e) => return (Err(anyhow::anyhow!("read PR #{pr} head: {e}")), None),
|
|
||||||
};
|
|
||||||
if head != reviewed {
|
|
||||||
return (
|
|
||||||
Err(anyhow::anyhow!(
|
|
||||||
"PR #{pr} head drifted since review (reviewed {reviewed}, now {head}); re-review before merging"
|
|
||||||
)),
|
|
||||||
None,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
|
|
||||||
// 2. Fetch the reviewed head into applied so ff/verify/deploy resolve it.
|
|
||||||
coord.set_queue_step(queue_entry_id, "fetch PR head");
|
|
||||||
if let Err(e) = crate::forge::fetch_pr_head_into_applied(&repo, pr).await {
|
|
||||||
return (
|
|
||||||
Err(anyhow::anyhow!("fetch PR #{pr} head into applied: {e}")),
|
|
||||||
None,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
|
|
||||||
// 3. Eval-verify BEFORE the irreversible push (bad nix fails fast here).
|
|
||||||
coord.set_queue_step(queue_entry_id, "verify proposal (eval)");
|
|
||||||
if let Err(e) =
|
|
||||||
crate::meta::verify_commit(approval.agent.as_str(), applied_dir, &reviewed).await
|
|
||||||
{
|
|
||||||
return (
|
|
||||||
Err(anyhow::anyhow!("verify merge head {reviewed}: {e:#}")),
|
|
||||||
None,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
|
|
||||||
// Capture the currently-deployed sha for the deploy tail's rollback.
|
|
||||||
let prev_main_sha = match lifecycle::git_rev_parse(applied_dir, "refs/heads/main").await {
|
|
||||||
Ok(s) => s,
|
|
||||||
Err(e) => return (Err(anyhow::anyhow!("read applied/main: {e:#}")), None),
|
|
||||||
};
|
|
||||||
|
|
||||||
// 4. THE merge: fast-forward-only merge the reviewed head to `main` via the
|
|
||||||
// forge API, pinned to the reviewed sha (`head_commit_id`). This one call
|
|
||||||
// both advances `main` to the reviewed head and marks the PR merged — no
|
|
||||||
// direct push to the protected branch. Unlike the old push-then-mark split,
|
|
||||||
// a failure here means `main` was NOT advanced, so it's fatal: we must not
|
|
||||||
// deploy a head the forge didn't merge.
|
|
||||||
coord.set_queue_step(queue_entry_id, "fast-forward-merge PR");
|
|
||||||
match crate::forge::merge_config_pr_ff(&repo, pr, &reviewed).await {
|
|
||||||
Ok(()) => {}
|
|
||||||
Err(crate::forge::ForgeMergeError::HeadDrift { expected, actual }) => {
|
|
||||||
return (
|
|
||||||
Err(anyhow::anyhow!(
|
|
||||||
"PR #{pr} head drifted before merge (reviewed {expected}, now {actual}); re-review before merging"
|
|
||||||
)),
|
|
||||||
None,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
Err(e) => return (Err(anyhow::anyhow!("ff-merge PR #{pr}: {e}")), None),
|
|
||||||
}
|
|
||||||
|
|
||||||
// 5. Deploy tail. target == finalize == the reviewed head.
|
|
||||||
deploy_applied_target(
|
|
||||||
coord,
|
|
||||||
approval.agent.as_str(),
|
|
||||||
agent_dir,
|
|
||||||
applied_dir,
|
|
||||||
&reviewed,
|
|
||||||
&reviewed,
|
|
||||||
id,
|
|
||||||
&prev_main_sha,
|
|
||||||
queue_entry_id,
|
|
||||||
)
|
|
||||||
.await
|
|
||||||
}
|
|
||||||
|
|
||||||
/// Inline (non-queued) handler for `ApprovalKind::SchedulePrompt`.
|
/// Inline (non-queued) handler for `ApprovalKind::SchedulePrompt`.
|
||||||
/// On approve, decode the `SchedulePromptPayload` JSON from the
|
/// On approve, decode the `SchedulePromptPayload` JSON from the
|
||||||
/// approval's `commit_ref`, insert a row into `scheduled_prompts`
|
/// approval's `commit_ref`, insert a row into `scheduled_prompts`
|
||||||
|
|
@ -408,24 +467,19 @@ async fn run_approval_schedule_prompt(
|
||||||
finish_approval(coord, &approval, result, None)
|
finish_approval(coord, &approval, result, None)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Terminal hook for approval-carrying DAGs — the job queue's
|
/// Terminal hook for approval-carrying DAGs — the job queue's scheduler calls
|
||||||
/// scheduler calls this exactly once when such a DAG settles terminal.
|
/// this exactly once when such a DAG settles terminal. Every approval-carrying
|
||||||
/// `MetaUpdate` and `Spawn` approval DAGs resolve here (their work is
|
/// template resolves here, deploys included: the deploy pipeline is ordinary
|
||||||
/// ordinary queue nodes); the opaque `ApprovalDeploy` pipeline resolves
|
/// queue nodes now, so the DAG's own terminal state is the authoritative
|
||||||
/// *inside* its node, so its DAG is skipped — unless it was cancelled
|
/// outcome and there's no in-node resolution to skip around.
|
||||||
/// while still queued, in which case the node never ran and the row
|
|
||||||
/// would otherwise dangle forever.
|
|
||||||
pub(crate) async fn resolve_approval_dag(
|
pub(crate) async fn resolve_approval_dag(
|
||||||
coord: &Arc<Coordinator>,
|
coord: &Arc<Coordinator>,
|
||||||
terminal: &crate::job_queue::TerminalDag,
|
terminal: &crate::job_queue::TerminalDag,
|
||||||
) {
|
) {
|
||||||
use crate::job_queue::{State, Template};
|
use crate::job_queue::State;
|
||||||
let Some(approval_id) = terminal.approval_id else {
|
let Some(approval_id) = terminal.approval_id else {
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
if terminal.template == Template::Rebuild && terminal.state != State::Cancelled {
|
|
||||||
return; // ApprovalDeploy resolved inside the node.
|
|
||||||
}
|
|
||||||
let approval = match coord.approvals.get(approval_id) {
|
let approval = match coord.approvals.get(approval_id) {
|
||||||
Ok(Some(a)) => a,
|
Ok(Some(a)) => a,
|
||||||
Ok(None) => {
|
Ok(None) => {
|
||||||
|
|
@ -448,22 +502,61 @@ pub(crate) async fn resolve_approval_dag(
|
||||||
.unwrap_or_else(|| "job dag failed".to_owned())
|
.unwrap_or_else(|| "job dag failed".to_owned())
|
||||||
)),
|
)),
|
||||||
};
|
};
|
||||||
if approval.kind == ApprovalKind::Spawn {
|
let mut terminal_tag = None;
|
||||||
// Post-spawn forge bookkeeping (user, config repo mirror, meta
|
match approval.kind {
|
||||||
// access) — warn-only, then the resolution events + a rescan so
|
ApprovalKind::Spawn => {
|
||||||
// the dashboard reflects the post-spawn state either way.
|
// Post-spawn forge bookkeeping (user, config repo mirror, meta
|
||||||
if result.is_ok() {
|
// access) — warn-only, then the resolution events + a rescan so
|
||||||
forge_after_first_spawn(coord, approval.agent.as_str()).await;
|
// the dashboard reflects the post-spawn state either way.
|
||||||
} else {
|
if result.is_ok() {
|
||||||
coord.rescan_containers_and_emit().await;
|
forge_after_first_spawn(coord, approval.agent.as_str()).await;
|
||||||
crate::dashboard::emit_tombstones_snapshot(coord).await;
|
} else {
|
||||||
|
coord.rescan_containers_and_emit().await;
|
||||||
|
crate::dashboard::emit_tombstones_snapshot(coord).await;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
ApprovalKind::MergeConfigPr => {
|
||||||
|
terminal_tag =
|
||||||
|
deploy_terminal_tag(approval.agent.as_str(), approval_id, terminal.state).await;
|
||||||
|
// On a failed deploy, surface the failing build log back onto the
|
||||||
|
// PR so the manager sees why it was rejected without leaving the
|
||||||
|
// forge. Posted here rather than inside a node because this is the
|
||||||
|
// one place that holds the DAG's definitive error — a `MergeVerify`
|
||||||
|
// rejection and a `DeployApply` build failure both land here.
|
||||||
|
if let Err(e) = &result {
|
||||||
|
post_merge_failure_to_pr(coord, &approval, e).await;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
_ => {}
|
||||||
}
|
}
|
||||||
if let Err(e) = finish_approval(coord, &approval, result, None) {
|
if let Err(e) = finish_approval(coord, &approval, result, terminal_tag) {
|
||||||
tracing::warn!(approval_id, error = ?e, "approval dag resolved with failure");
|
tracing::warn!(approval_id, error = ?e, "approval dag resolved with failure");
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Which bookkeeping tag a settled deploy DAG actually planted, for the
|
||||||
|
/// `Rebuilt` event's `tag` field. The state picks the candidate name, but the
|
||||||
|
/// applied repo has the final say: a pre-merge rejection (`MergeVerify` drift
|
||||||
|
/// gate, eval failure) fails the DAG without ever planting `failed/<id>`, and
|
||||||
|
/// tag plants are best-effort. Reporting a tag that isn't there would send the
|
||||||
|
/// manager looking for a ref that doesn't exist.
|
||||||
|
async fn deploy_terminal_tag(
|
||||||
|
agent: &str,
|
||||||
|
approval_id: i64,
|
||||||
|
state: crate::job_queue::State,
|
||||||
|
) -> Option<String> {
|
||||||
|
use crate::job_queue::State;
|
||||||
|
let candidate = match state {
|
||||||
|
State::Done => format!("deployed/{approval_id}"),
|
||||||
|
State::Cancelled => return None,
|
||||||
|
_ => format!("failed/{approval_id}"),
|
||||||
|
};
|
||||||
|
lifecycle::git_rev_parse(&crate::paths::applied_dir(agent), &candidate)
|
||||||
|
.await
|
||||||
|
.ok()
|
||||||
|
.map(|_| candidate)
|
||||||
|
}
|
||||||
|
|
||||||
/// Re-fetch an approval row from sqlite for a queue-worker dispatch.
|
/// Re-fetch an approval row from sqlite for a queue-worker dispatch.
|
||||||
/// Bails if the row is gone (deny race), if its kind doesn't match,
|
/// Bails if the row is gone (deny race), if its kind doesn't match,
|
||||||
/// or if the lookup itself fails. The kind check is defensive — the
|
/// or if the lookup itself fails. The kind check is defensive — the
|
||||||
|
|
@ -647,62 +740,47 @@ fn finish_approval(
|
||||||
result
|
result
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Deploy tail for the config-PR merge flow. Fast-forwards `applied/main` to
|
/// Post-merge deploy for the config-PR flow. Fast-forwards `applied/main` to
|
||||||
/// `target_ref`, syncs the working tree, runs the meta two-phase deploy +
|
/// `target`, syncs the working tree, runs the meta two-phase deploy + container
|
||||||
/// container rebuild, and plants the `deployed/<tag_base>` /
|
/// rebuild, and plants the `deployed/<id>` / `failed/<id>` bookkeeping tag.
|
||||||
/// `failed/<tag_base>` bookkeeping tags. On build failure it rolls
|
///
|
||||||
/// `applied/main` back to `prev_main_sha` and aborts the staged meta lock so
|
/// **Undo is not this function's job.** Every early return here leaves the
|
||||||
/// the agent stays on its last-good tree. Returns the build result + the
|
/// applied repo dirty on purpose — [`run_deploy_tail`] owns compensation, and
|
||||||
/// terminal tag name.
|
/// it runs whether this returns `Err`, panics, or never returns at all because
|
||||||
|
/// hive-c0re was restarted underneath it. That's the whole point of parking the
|
||||||
|
/// pre-merge sha in a git ref instead of a local variable.
|
||||||
///
|
///
|
||||||
/// Caller-specific bits stay OUT of here: fetching the PR head, the
|
/// Caller-specific bits stay OUT of here: fetching the PR head, the
|
||||||
/// `verify_commit` gate, the ff-merge, and forge mark-merged. `finalize_sha`
|
/// `verify_commit` gate, and the ff-merge. `target` is both what `applied/main`
|
||||||
/// is the sha recorded by `meta::finalize_deploy`; `target_ref` is what
|
/// fast-forwards to and the sha `meta::finalize_deploy` records — for a merge
|
||||||
/// `applied/main` fast-forwards to. The agent always already exists here (a
|
/// they are always the same reviewed head. The agent always already exists here
|
||||||
/// merge is never a first spawn), so there's no `sync_agents` step — the
|
/// (a merge is never a first spawn), so there's no `sync_agents` step — the
|
||||||
/// operator `Spawn` flow owns first-time meta registration.
|
/// operator `Spawn` flow owns first-time meta registration.
|
||||||
#[allow(
|
|
||||||
clippy::too_many_arguments,
|
|
||||||
clippy::too_many_lines,
|
|
||||||
reason = "one sequential ff/deploy/rebuild/finalize pipeline; splitting it \
|
|
||||||
would obscure the linear flow"
|
|
||||||
)]
|
|
||||||
async fn deploy_applied_target(
|
async fn deploy_applied_target(
|
||||||
coord: &Arc<Coordinator>,
|
coord: &Arc<Coordinator>,
|
||||||
agent: &str,
|
agent: &str,
|
||||||
agent_dir: &std::path::Path,
|
agent_dir: &std::path::Path,
|
||||||
applied_dir: &std::path::Path,
|
applied_dir: &std::path::Path,
|
||||||
target_ref: &str,
|
target: &str,
|
||||||
finalize_sha: &str,
|
id: i64,
|
||||||
tag_base: i64,
|
|
||||||
prev_main_sha: &str,
|
|
||||||
queue_entry_id: Option<u64>,
|
queue_entry_id: Option<u64>,
|
||||||
) -> (Result<()>, Option<String>) {
|
) -> Result<()> {
|
||||||
let id = tag_base;
|
|
||||||
|
|
||||||
coord.set_queue_step(queue_entry_id, "fast-forward applied/main");
|
coord.set_queue_step(queue_entry_id, "fast-forward applied/main");
|
||||||
// Fast-forward applied/main to target_ref + sync the working tree.
|
// Fast-forward applied/main to target + sync the working tree. Meta input
|
||||||
// Meta input pins `?ref=main`, so this is what makes nix re-lock to
|
// pins `?ref=main`, so this is what makes nix re-lock to the target commit
|
||||||
// the target commit on the prepare_deploy step below. On build
|
// on the prepare_deploy step below.
|
||||||
// failure we roll main back to prev_main_sha so a crash leaves the
|
lifecycle::git_update_ref(applied_dir, "refs/heads/main", target)
|
||||||
// agent on its last-good tree.
|
.await
|
||||||
if let Err(e) = lifecycle::git_update_ref(applied_dir, "refs/heads/main", target_ref).await {
|
.map_err(|e| anyhow::anyhow!("ff main to {target}: {e:#}"))?;
|
||||||
return (Err(anyhow::anyhow!("ff main to {target_ref}: {e:#}")), None);
|
lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main")
|
||||||
}
|
.await
|
||||||
if let Err(e) = lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main").await {
|
.map_err(|e| anyhow::anyhow!("read-tree to main: {e:#}"))?;
|
||||||
// main is ahead; working tree didn't sync. Roll main back to
|
|
||||||
// keep the two consistent before bailing.
|
|
||||||
let _ = lifecycle::git_update_ref(applied_dir, "refs/heads/main", prev_main_sha).await;
|
|
||||||
return (Err(anyhow::anyhow!("read-tree to main: {e:#}")), None);
|
|
||||||
}
|
|
||||||
|
|
||||||
coord.set_queue_step(queue_entry_id, "meta prepare_deploy");
|
coord.set_queue_step(queue_entry_id, "meta prepare_deploy");
|
||||||
// Phase 1 of the meta two-phase deploy: relock without committing.
|
// Phase 1 of the meta two-phase deploy: relock without committing.
|
||||||
if let Err(e) = crate::meta::prepare_deploy(agent).await {
|
crate::meta::prepare_deploy(agent)
|
||||||
let _ = lifecycle::git_update_ref(applied_dir, "refs/heads/main", prev_main_sha).await;
|
.await
|
||||||
let _ = lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main").await;
|
.map_err(|e| anyhow::anyhow!("meta prepare_deploy: {e:#}"))?;
|
||||||
return (Err(anyhow::anyhow!("meta prepare_deploy: {e:#}")), None);
|
|
||||||
}
|
|
||||||
|
|
||||||
// Container-level rebuild (or first-time create) against meta#<name>.
|
// Container-level rebuild (or first-time create) against meta#<name>.
|
||||||
// Step labels are emitted inside rebuild_no_meta via the callback so
|
// Step labels are emitted inside rebuild_no_meta via the callback so
|
||||||
|
|
@ -726,11 +804,18 @@ async fn deploy_applied_target(
|
||||||
match build_result {
|
match build_result {
|
||||||
Ok(_) => {
|
Ok(_) => {
|
||||||
coord.set_queue_step(queue_entry_id, "finalize deploy");
|
coord.set_queue_step(queue_entry_id, "finalize deploy");
|
||||||
|
// Drop the compensation ref FIRST: from here the deploy is good and
|
||||||
|
// the tail must not roll `main` back. Ordering it ahead of the tag
|
||||||
|
// plant is what makes the tail's `deployed/<id>` cross-check a
|
||||||
|
// second line of defence rather than the only one.
|
||||||
|
if let Err(e) = lifecycle::git_delete_ref(applied_dir, &rollback_ref(id)).await {
|
||||||
|
tracing::warn!(%agent, %id, error = ?e, "drop rollback ref after successful deploy failed");
|
||||||
|
}
|
||||||
let tag = format!("deployed/{id}");
|
let tag = format!("deployed/{id}");
|
||||||
if let Err(e) = lifecycle::git_tag(applied_dir, &tag, target_ref).await {
|
if let Err(e) = lifecycle::git_tag(applied_dir, &tag, target).await {
|
||||||
tracing::warn!(%agent, %id, error = ?e, "plant deployed tag failed");
|
tracing::warn!(%agent, %id, error = ?e, "plant deployed tag failed");
|
||||||
}
|
}
|
||||||
if let Err(e) = crate::meta::finalize_deploy(agent, finalize_sha, &tag).await {
|
if let Err(e) = crate::meta::finalize_deploy(agent, target, &tag).await {
|
||||||
// The build itself succeeded — meta lock landed but
|
// The build itself succeeded — meta lock landed but
|
||||||
// couldn't be committed. Surface as a soft warn so the
|
// couldn't be committed. Surface as a soft warn so the
|
||||||
// operator can git-commit by hand if they care.
|
// operator can git-commit by hand if they care.
|
||||||
|
|
@ -742,32 +827,18 @@ async fn deploy_applied_target(
|
||||||
// proposal, agent picks up where it left off with the
|
// proposal, agent picks up where it left off with the
|
||||||
// new env / packages.
|
// new env / packages.
|
||||||
coord.kick_agent(agent, "config update applied");
|
coord.kick_agent(agent, "config update applied");
|
||||||
(Ok(()), Some(tag))
|
Ok(())
|
||||||
}
|
}
|
||||||
Err(e) => {
|
Err(e) => {
|
||||||
|
// Plant the failure marker here rather than in the tail: this is
|
||||||
|
// the only place that holds the build error to annotate it with.
|
||||||
|
// The repo-state rollback is the tail's, via the parked ref.
|
||||||
let tag = format!("failed/{id}");
|
let tag = format!("failed/{id}");
|
||||||
let body = format!("{e:#}");
|
let body = format!("{e:#}");
|
||||||
if let Err(te) =
|
if let Err(te) = lifecycle::git_tag_annotated(applied_dir, &tag, target, &body).await {
|
||||||
lifecycle::git_tag_annotated(applied_dir, &tag, target_ref, &body).await
|
|
||||||
{
|
|
||||||
tracing::warn!(%agent, %id, error = ?te, "annotate failed tag failed");
|
tracing::warn!(%agent, %id, error = ?te, "annotate failed tag failed");
|
||||||
}
|
}
|
||||||
// Roll main back to last known-good so the on-disk state
|
Err(e)
|
||||||
// matches what nixos-container last successfully built.
|
|
||||||
if let Err(re) =
|
|
||||||
lifecycle::git_update_ref(applied_dir, "refs/heads/main", prev_main_sha).await
|
|
||||||
{
|
|
||||||
tracing::warn!(%agent, %id, error = ?re, "main rollback failed");
|
|
||||||
}
|
|
||||||
if let Err(re) = lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main").await {
|
|
||||||
tracing::warn!(%agent, %id, error = ?re, "rollback read-tree failed");
|
|
||||||
}
|
|
||||||
// Drop the staged meta lock change so the deploy log
|
|
||||||
// only ever shows successes.
|
|
||||||
if let Err(ae) = crate::meta::abort_deploy().await {
|
|
||||||
tracing::warn!(%agent, %id, error = ?ae, "meta abort_deploy failed");
|
|
||||||
}
|
|
||||||
(Err(e), Some(tag))
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -99,7 +99,10 @@ pub(super) async fn run_node(coord: &Arc<Coordinator>, claim: &Claim) -> Result<
|
||||||
NodeKind::Drain { .. } => run_drain(coord, claim, &ctx).await,
|
NodeKind::Drain { .. } => run_drain(coord, claim, &ctx).await,
|
||||||
NodeKind::WriteDropin { .. } => run_write_dropin(coord, claim).await,
|
NodeKind::WriteDropin { .. } => run_write_dropin(coord, claim).await,
|
||||||
NodeKind::WritePermFile { .. } => run_write_perm_file(coord, claim, &ctx).await,
|
NodeKind::WritePermFile { .. } => run_write_perm_file(coord, claim, &ctx).await,
|
||||||
NodeKind::ApprovalDeploy { .. } => run_approval_deploy(coord, claim).await,
|
NodeKind::DeployWindow { .. } => run_deploy_window(claim),
|
||||||
|
NodeKind::MergeVerify { .. } => run_merge_verify(coord, claim).await,
|
||||||
|
NodeKind::DeployApply { .. } => run_deploy_apply(coord, claim).await,
|
||||||
|
NodeKind::DeployTail { .. } => run_deploy_tail(coord, claim).await,
|
||||||
NodeKind::SetWanted { up, .. } => run_set_wanted(coord, claim, *up),
|
NodeKind::SetWanted { up, .. } => run_set_wanted(coord, claim, *up),
|
||||||
// Pure grouping container — no work; completing it lets it reach
|
// Pure grouping container — no work; completing it lets it reach
|
||||||
// `Finishing` so its child template nodes start. The DAG's terminal
|
// `Finishing` so its child template nodes start. The DAG's terminal
|
||||||
|
|
@ -578,27 +581,66 @@ async fn run_write_perm_file(
|
||||||
Ok(NodeOutput::default())
|
Ok(NodeOutput::default())
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Opaque approval deploy pipeline for `MergeConfigPr`: verify + ff-merge the
|
/// The approval id every deploy phase re-reads its approval row by. Fails the
|
||||||
/// reviewed PR head, then the container rebuild. The two-phase
|
/// node when the DAG carries none, which would mean a `MergeConfigPr` DAG was
|
||||||
/// prepare/finalize/abort meta deploy — and the approval resolution — stay
|
/// built without going through `templates::approval_deploy`.
|
||||||
/// inside `actions.rs` in v1 (design doc §9).
|
fn deploy_approval_id(claim: &Claim) -> Result<i64> {
|
||||||
async fn run_approval_deploy(coord: &Arc<Coordinator>, claim: &Claim) -> Result<NodeOutput> {
|
claim
|
||||||
let approval_id = claim
|
|
||||||
.approval_id
|
.approval_id
|
||||||
.with_context(|| format!("approval_deploy dag {} has no approval_id", claim.dag_id))?;
|
.with_context(|| format!("approval deploy dag {} has no approval_id", claim.dag_id))
|
||||||
// The deploy window covers the whole prepare→finalize span
|
}
|
||||||
// (`NodeKind::needs_meta_window`, held by the scheduler for this
|
|
||||||
// node): `prepare_deploy` stages `flake.lock` uncommitted for the
|
/// The deploy subtree's root: pure resource holder, no work of its own.
|
||||||
// entire container build, and no other meta mutation may land inside
|
///
|
||||||
// that window (it would sweep the staged lock and neuter
|
/// It exists so the global meta window (plus the agent lease and a build slot)
|
||||||
// `abort_deploy`). Holding it as a queue resource — rather than a
|
/// is held continuously across every phase below it. `prepare_deploy` leaves
|
||||||
// `MutexGuard` that cannot outlive this fn — is what lets increment 2
|
/// `flake.lock` staged-uncommitted for the whole container build, and any other
|
||||||
// decompose this node into sub-nodes under a window-holding parent.
|
/// meta mutation landing inside that span would sweep the staged lock into its
|
||||||
crate::actions::run_approval_merge_config_pr(coord, Some(claim.dag_id), approval_id)
|
/// own commit and neuter `abort_deploy` — so the window has to outlive any one
|
||||||
|
/// node, which the `MutexGuard` this replaced could not do.
|
||||||
|
///
|
||||||
|
/// Completing immediately moves it to `Finishing`, which is what starts the
|
||||||
|
/// children; the resources stay held until the whole subtree settles.
|
||||||
|
fn run_deploy_window(claim: &Claim) -> Result<NodeOutput> {
|
||||||
|
deploy_approval_id(claim)?;
|
||||||
|
Ok(NodeOutput::default())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Deploy phase 1 — drift gate, fetch, eval-verify. Mutates nothing, so a
|
||||||
|
/// failure here cancel-cascades the rest of the subtree with the forge and the
|
||||||
|
/// applied repo exactly as they were.
|
||||||
|
async fn run_merge_verify(coord: &Arc<Coordinator>, claim: &Claim) -> Result<NodeOutput> {
|
||||||
|
crate::actions::run_deploy_merge_verify(coord, Some(claim.dag_id), deploy_approval_id(claim)?)
|
||||||
.await
|
.await
|
||||||
.map(|()| NodeOutput::default())
|
.map(|()| NodeOutput::default())
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Deploy phase 2 — the irreversible half: ff-merge, two-phase meta deploy,
|
||||||
|
/// container rebuild, finalize.
|
||||||
|
async fn run_deploy_apply(coord: &Arc<Coordinator>, claim: &Claim) -> Result<NodeOutput> {
|
||||||
|
crate::actions::run_deploy_apply(coord, Some(claim.dag_id), deploy_approval_id(claim)?)
|
||||||
|
.await
|
||||||
|
.map(|()| NodeOutput::default())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Deploy compensation + bookkeeping tail. `AfterAny` the apply node, so it
|
||||||
|
/// runs on every outcome; it is deliberately infallible (see
|
||||||
|
/// [`crate::actions::run_deploy_tail`]) — a failing tail must not flip an
|
||||||
|
/// otherwise-successful deploy's DAG state.
|
||||||
|
///
|
||||||
|
/// Takes the agent from the node payload so the tail can still compensate when
|
||||||
|
/// the approval row is gone (deny race, purge).
|
||||||
|
async fn run_deploy_tail(coord: &Arc<Coordinator>, claim: &Claim) -> Result<NodeOutput> {
|
||||||
|
crate::actions::run_deploy_tail(
|
||||||
|
coord,
|
||||||
|
Some(claim.dag_id),
|
||||||
|
claim.kind.agent(),
|
||||||
|
deploy_approval_id(claim)?,
|
||||||
|
)
|
||||||
|
.await;
|
||||||
|
Ok(NodeOutput::default())
|
||||||
|
}
|
||||||
|
|
||||||
/// Compute which agents a `nix flake update <inputs>` on the meta
|
/// Compute which agents a `nix flake update <inputs>` on the meta
|
||||||
/// flake affects — the fan-out set for `MetaUpdate` DAGs. Empty
|
/// flake affects — the fan-out set for `MetaUpdate` DAGs. Empty
|
||||||
/// `inputs` or any input under `hyperhive` → every container;
|
/// `inputs` or any input under `hyperhive` → every container;
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue