Watch
0
0
Fork
You've already forked hyperhive
0

config PRs: an operator's Forgejo merge deploys the merged rev

A config PR merged in the Forgejo UI changed nothing on the hive: the
hive's webhook ignores `closed`, its poll then cancels the dashboard
card, and `applied/main` stays where it was.

swarm-controller reads `merged`/`merge_commit_sha` off the
`pull_request` delivery it already receives for `agent-configs`, finds
the hive placing the agent by scanning every hive's wanted state (the
scan `declarations_elsewhere` already ran, factored out), and queues a
`TriggerDeploy` carrying the rev. Zero or several claimants deploy
nothing and log the claimants.

`DeployRequest` gains `rev: Option<String>` with `serde(default)`, so
rev-less payloads from either side keep decoding.

hive-c0re, given a rev for an agent it runs: a no-op when
`applied/main` already is the rev (a dashboard merge deploys its own
PR); otherwise it fetches the forge `main` with the core token,
requires the rev to descend from `applied/main` (the ancestry gate,
factored out of `run_deploy_merge_verify`), fast-forwards by CAS and
queues the usual relocking rebuild. No eval-verify on this path, per
mara (#4850 c90075). A refusal is commented on the PR that merged the
rev, found by commit.

swarm-controller's forge-objects pass converges every config repo's
`main` rule to merge whitelist `operators` + `core` and approval
whitelist `operators`. The hive's boot PATCH stops forcing
`enable_approvals_whitelist` off, so the two do not fight.

Refs #4850
This commit is contained in:
atlas 2026-10-02 18:45:18 +02:00
commit f1c695c212
11 changed files with 732 additions and 76 deletions

View file

@ -196,13 +196,8 @@ pub async fn run_deploy_merge_verify(
// 3. Ancestry gate: `main` must be reachable from the reviewed head, or the
// "fast-forward" in prepare_applied_target is really a rewind that drops
// every commit between the PR's base and where `main` actually is now.
let current_main = lifecycle::git_rev_parse(&ctx.applied_dir, "refs/heads/main")
.await
.map_err(|e| anyhow::anyhow!("read applied/main: {e:#}"))?;
if !lifecycle::git_is_ancestor(&ctx.applied_dir, &current_main, reviewed)
.await
.map_err(|e| anyhow::anyhow!("ancestry check {current_main}..{reviewed}: {e:#}"))?
{
let (current_main, descends) = applied_main_under(&ctx.applied_dir, reviewed).await?;
if !descends {
bail!(
"PR #{pr} does not descend from applied/main (main {current_main}, reviewed {reviewed}); \
merging it would discard commits — rebase the PR onto main and re-review"
@ -221,6 +216,80 @@ pub async fn run_deploy_merge_verify(
Ok(())
}
/// `applied/main`, and whether `target` descends from it. Fast-forwarding
/// `main` to a `target` that does not is a rewind: it drops every commit on
/// `main` that `target` lacks.
async fn applied_main_under(applied_dir: &std::path::Path, target: &str) -> Result<(String, bool)> {
let current_main = lifecycle::git_rev_parse(applied_dir, "refs/heads/main")
.await
.map_err(|e| anyhow::anyhow!("read applied/main: {e:#}"))?;
let descends = lifecycle::git_is_ancestor(applied_dir, &current_main, target)
.await
.map_err(|e| anyhow::anyhow!("ancestry check {current_main}..{target}: {e:#}"))?;
Ok((current_main, descends))
}
/// What [`advance_applied_to_rev`] did to `applied/main`.
#[derive(Debug, PartialEq, Eq)]
pub(crate) enum RevAdvance {
/// `applied/main` already was the rev, so there is nothing to deploy.
AlreadyApplied,
/// `applied/main` was fast-forwarded to the rev.
Advanced,
}
/// Fast-forward `agent`'s `applied/main` to `rev`, a commit the swarm asks
/// this hive to deploy, so the next relocking rebuild builds it.
///
/// No eval-verify on this path: a config that does not evaluate fails the
/// rebuild instead, with `applied/main` already at `rev`.
///
/// # Errors
///
/// Returns an error if the forge fetch fails, if `rev` does not descend from
/// `applied/main`, or if `main` moved during the fast-forward. `main` does not
/// move in any of these.
pub(crate) async fn advance_applied_to_rev(agent: &str, rev: &str) -> Result<RevAdvance> {
advance_applied_main(
&crate::paths::applied_dir(agent),
rev,
crate::forge::fetch_forge_main(agent),
)
.await
}
/// [`advance_applied_to_rev`] with the forge fetch passed in, so it is
/// testable against a local repo.
///
/// `applied/main` is compared to `rev` twice: before the fetch, so an applied
/// rev costs no forge round-trip, and after it, because a dashboard deploy of
/// the same merge fast-forwards `main` itself and may do so meanwhile.
async fn advance_applied_main<T>(
applied_dir: &std::path::Path,
rev: &str,
fetch: impl std::future::Future<Output = Result<T>>,
) -> Result<RevAdvance> {
let main = lifecycle::git_rev_parse(applied_dir, "refs/heads/main")
.await
.map_err(|e| anyhow::anyhow!("read applied/main: {e:#}"))?;
if main == rev {
return Ok(RevAdvance::AlreadyApplied);
}
fetch.await.context("fetch the forge config main")?;
let (main, descends) = applied_main_under(applied_dir, rev).await?;
if main == rev {
return Ok(RevAdvance::AlreadyApplied);
}
if !descends {
bail!(
"{rev} does not descend from applied/main ({main}); fast-forwarding to it would \
discard commits"
);
}
ff_applied_main(applied_dir, rev, &main).await?;
Ok(RevAdvance::Advanced)
}
/// `DeployApply` node body — the irreversible half. Parks the rollback ref,
/// fast-forward-merges the PR (THE merge), then opens the deploy.
///
@ -422,6 +491,28 @@ async fn post_merge_failure_to_pr(
}
}
/// Post why the merged commit `rev` of `agent`'s config repo was not deployed
/// onto the PR whose merge produced it. Best-effort, like
/// [`post_merge_failure_to_pr`]: a failure here is logged only.
pub(crate) async fn post_rev_deploy_failure_to_pr(agent: &str, rev: &str, err: &anyhow::Error) {
let repo = crate::forge::config_repo(agent);
let pr = match crate::forge::merged_pr_for_commit(&repo, rev).await {
Ok(pr) => pr,
Err(e) => {
tracing::warn!(%agent, %rev, error = %e, "find the PR that merged a failed deploy's rev failed");
return;
}
};
let body = format!(
"## ⚠️ config deploy failed\n\n\
The merged commit `{rev}` was not deployed:\n\n\
```\n{err:#}\n```"
);
if let Err(e) = crate::forge::post_pr_comment(&repo, pr, &body).await {
tracing::warn!(%agent, %pr, error = ?e, "post rev deploy-failure comment to PR failed");
}
}
/// Return the last `max_bytes` of `s`, snapped to a char boundary, prefixed
/// with an elision marker when truncated.
fn tail_bytes(s: &str, max_bytes: usize) -> String {
@ -672,24 +763,9 @@ async fn prepare_applied_target(
expected_main: &str,
node_id: Option<u64>,
) -> Result<()> {
// Fast-forward applied/main to target + sync the working tree. Meta input
// pins `?ref=main`, so this is what makes nix re-lock to the target commit
// on the prepare_deploy step below.
//
// Compare-and-swap, not a bare set: `main` must still be the sha the caller
// read before the merge. A plain `update-ref` here moves the branch to
// `target` whatever it currently points at, which turns "fast-forward" into
// "discard anything that landed in the meantime" — the ancestry gate in
// run_deploy_merge_verify only proves the target is safe against the `main`
// observed *then*, so this is what makes that proof still true *now*.
lifecycle::git_update_ref_cas(applied_dir, "refs/heads/main", target, expected_main)
.await
.map_err(|e| {
anyhow::anyhow!("ff main {expected_main} -> {target} (concurrent move?): {e:#}")
})?;
lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main")
.await
.map_err(|e| anyhow::anyhow!("read-tree to main: {e:#}"))?;
// Meta input pins `?ref=main`, so this is what makes nix re-lock to the
// target commit on the prepare_deploy step below.
ff_applied_main(applied_dir, target, expected_main).await?;
// Phase 1 of the meta two-phase deploy: relock without committing. The
// staged lock then stays uncommitted across the whole appended rebuild —
@ -699,6 +775,29 @@ async fn prepare_applied_target(
.map_err(|e| anyhow::anyhow!("meta prepare_deploy: {e:#}"))
}
/// Fast-forward `applied/main` to `target` and sync the working tree.
///
/// Compare-and-swap, not a bare set: `main` must still be `expected_main`, the
/// sha the caller's ancestry gate ([`applied_main_under`]) ran against. A plain
/// `update-ref` moves the branch to `target` whatever it currently points at,
/// which turns "fast-forward" into "discard anything that landed in the
/// meantime" — the gate only proves `target` safe against the `main` observed
/// *then*, so this is what makes that proof still true *now*.
async fn ff_applied_main(
applied_dir: &std::path::Path,
target: &str,
expected_main: &str,
) -> Result<()> {
lifecycle::git_update_ref_cas(applied_dir, "refs/heads/main", target, expected_main)
.await
.map_err(|e| {
anyhow::anyhow!("ff main {expected_main} -> {target} (concurrent move?): {e:#}")
})?;
lifecycle::git_read_tree_reset(applied_dir, "refs/heads/main")
.await
.map_err(|e| anyhow::anyhow!("read-tree to main: {e:#}"))
}
/// `FinalizeDeploy` node body — phase 2 of the meta two-phase deploy, run once
/// the appended rebuild subgraph has built, swapped, and brought the container
/// back up.
@ -802,3 +901,105 @@ pub fn deny(coord: &Coordinator, id: i64, note: Option<&str>) -> Result<()> {
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::{RevAdvance, advance_applied_main};
use crate::lifecycle;
/// A repo with `main` at `b`, child of `a`, and `c`, a sibling of `b` off
/// `a`. Returns the dir and the shas `[a, b, c]`.
async fn repo() -> (tempfile::TempDir, [String; 3]) {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path();
let commit = |msg: &'static str| async move {
lifecycle::git(
path,
&[
"-c",
"user.name=t",
"-c",
"user.email=t@t",
"commit",
"-q",
"--allow-empty",
"-m",
msg,
],
)
.await
.expect("commit");
lifecycle::git_rev_parse(path, "HEAD").await.expect("HEAD")
};
lifecycle::git(path, &["init", "-q", "--initial-branch=main"])
.await
.expect("init");
let a = commit("a").await;
let b = commit("b").await;
lifecycle::git(path, &["checkout", "-q", "-b", "side", &a])
.await
.expect("branch");
let c = commit("c").await;
lifecycle::git(path, &["checkout", "-q", "main"])
.await
.expect("checkout main");
(dir, [a, b, c])
}
async fn main_of(dir: &tempfile::TempDir) -> String {
lifecycle::git_rev_parse(dir.path(), "refs/heads/main")
.await
.expect("main")
}
#[tokio::test]
async fn the_applied_rev_is_a_no_op_without_a_fetch() {
let (dir, [_, b, _]) = repo().await;
let fetched = std::cell::Cell::new(false);
let fetch = async {
fetched.set(true);
anyhow::Ok(())
};
let outcome = advance_applied_main(dir.path(), &b, fetch).await;
assert_eq!(outcome.expect("no-op"), RevAdvance::AlreadyApplied);
assert!(!fetched.get(), "an applied rev must not reach the forge");
assert_eq!(main_of(&dir).await, b);
}
#[tokio::test]
async fn a_rev_not_descending_from_main_is_refused_and_main_stays() {
let (dir, [_, b, c]) = repo().await;
let outcome = advance_applied_main(dir.path(), &c, async { anyhow::Ok(()) }).await;
let err = outcome.expect_err("a sibling of main is not a fast-forward");
assert!(
format!("{err:#}").contains("does not descend from applied/main"),
"{err:#}"
);
assert_eq!(main_of(&dir).await, b);
}
#[tokio::test]
async fn a_descendant_rev_fast_forwards_main() {
let (dir, [a, b, _]) = repo().await;
lifecycle::git_update_ref(dir.path(), "refs/heads/main", &a)
.await
.expect("rewind main");
let outcome = advance_applied_main(dir.path(), &b, async { anyhow::Ok(()) }).await;
assert_eq!(outcome.expect("fast-forward"), RevAdvance::Advanced);
assert_eq!(main_of(&dir).await, b);
}
#[tokio::test]
async fn a_failed_fetch_leaves_main_alone() {
let (dir, [a, b, _]) = repo().await;
lifecycle::git_update_ref(dir.path(), "refs/heads/main", &a)
.await
.expect("rewind main");
let outcome = advance_applied_main(dir.path(), &b, async {
Err::<(), _>(anyhow::anyhow!("forge down"))
})
.await;
assert!(outcome.is_err());
assert_eq!(main_of(&dir).await, a);
}
}

View file

@ -12,10 +12,10 @@ mod repos;
mod users;
pub use pr_merge::{
ForgeMergeError, config_repo, fetch_pr_head_into_applied, merge_config_pr_ff, post_pr_comment,
pr_head_sha, pr_is_open,
ForgeMergeError, config_repo, fetch_pr_head_into_applied, merge_config_pr_ff,
merged_pr_for_commit, post_pr_comment, pr_head_sha, pr_is_open,
};
pub use reconcile::{reconcile_config_apply, reconcile_config_status};
pub use reconcile::{fetch_forge_main, reconcile_config_apply, reconcile_config_status};
pub use repos::{
clone_config_into_proposed, ensure_config_repo, ensure_meta_remote, ensure_repo,
fast_forward_applied_main, fetch_config_main_into_applied, meta_read_access, push_config,

View file

@ -218,6 +218,33 @@ pub async fn merge_config_pr_ff(repo: &str, pr: u64, sha: &str) -> Result<(), Fo
}
}
/// The number of the merged PR that put commit `sha` on `repo`'s base branch.
///
/// # Errors
/// `Other` on absent core token, malformed repo, no such PR, or
/// transport/API failure.
pub async fn merged_pr_for_commit(repo: &str, sha: &str) -> Result<u64, ForgeMergeError> {
let token = core_token()
.ok_or_else(|| ForgeMergeError::Other(anyhow::anyhow!("forge core token absent")))?;
let (owner, name) = repo.split_once('/').ok_or_else(|| {
ForgeMergeError::Other(anyhow::anyhow!("forge repo `{repo}` is not owner/name"))
})?;
let client = api(&token).map_err(ForgeMergeError::Other)?;
let pull = client
.repo_get_commit_pull_request(owner, name, sha)
.await
.map_err(|e| {
ForgeMergeError::Other(anyhow::Error::from(e).context("GET pull request of commit"))
})?;
pull.number
.and_then(|n| u64::try_from(n).ok())
.ok_or_else(|| {
ForgeMergeError::Other(anyhow::anyhow!(
"pull request of {sha} in {repo} carries no number"
))
})
}
/// Post a comment to PR (= issue) `pr` on `repo` as the core forge user.
/// PRs are issues in Forgejo, so the PR number is the issue index. Used to
/// surface a failed config-approval deploy's build log back onto the PR so

View file

@ -25,7 +25,8 @@ const FORGE_MAIN_REF: &str = "refs/hyperhive/forge-config-main";
/// Fetch `agent-configs/<agent>` `main` into the applied repo's scratch
/// ref (read-only, no working-tree change) and return the applied dir.
async fn fetch_forge_main(agent: &str) -> Result<PathBuf> {
/// Also how a swarm deploy of a merged commit gets that commit's objects.
pub async fn fetch_forge_main(agent: &str) -> Result<PathBuf> {
if !is_present().await {
anyhow::bail!("forge is not running");
}

View file

@ -667,14 +667,14 @@ fn main_branch_protection_option() -> CreateBranchProtectionOption {
/// node (`actions::run_deploy_apply`), which fast-forward-*merges* the reviewed
/// head through the forge merge API (`Do=fast-forward-only`,
/// `head_commit_id` pinned to the reviewed sha).
/// - **merge is whitelisted to `core`** — only hive-c0re can merge a config PR;
/// the agent can push feature branches + open PRs but can't land them.
/// - **the operator's dashboard approval is the gate** — approval happens on
/// the `MergeConfigPr` card and hive-c0re only merges an approved PR. There's
/// deliberately no Forgejo `required_approvals` review requirement: the flow
/// never does an in-forge review, so requiring one would only dead-block the
/// `core` merge. The dashboard approval + the `core`-only merge whitelist are
/// the real gate.
/// - **merge is whitelisted to `core`** — the agent can push feature branches +
/// open PRs but can't land them. swarm-controller adds the `operators` team
/// to the whitelist, so an operator can also merge in the forge UI.
/// - **the operator's dashboard approval is the gate for `core`** — approval
/// happens on the `MergeConfigPr` card and hive-c0re only merges an approved
/// PR. There's deliberately no Forgejo `required_approvals` review
/// requirement: that flow never does an in-forge review, so requiring one
/// would only dead-block the `core` merge.
/// - **fast-forward-only** — `main` only ever advances by fast-forward; a raced
/// non-ff `main` is refused by the merge API rather than force-moved.
///
@ -736,6 +736,11 @@ async fn apply_config_repo_branch_protection(repo: &str, token: &str) -> Result<
/// left `None`, so a repo carrying an older push-based or approval-gated rule is
/// actually *converged* rather than merely re-asserted — a PATCH leaves unset
/// fields untouched. Every field unrelated to this policy stays `None`.
///
/// The approval and merge whitelists' **teams** stay `None` too, and so does
/// `enable_approvals_whitelist`: swarm-controller puts the `operators` team on
/// both, so operators can merge config PRs in the forge UI, and this PATCH
/// runs on every boot.
fn config_repo_protection_edit() -> EditBranchProtectionOption {
EditBranchProtectionOption {
apply_to_admins: None,
@ -745,7 +750,7 @@ fn config_repo_protection_edit() -> EditBranchProtectionOption {
block_on_outdated_branch: None,
block_on_rejected_reviews: None,
dismiss_stale_approvals: None,
enable_approvals_whitelist: Some(false),
enable_approvals_whitelist: None,
enable_merge_whitelist: Some(true),
enable_push: Some(false),
enable_push_whitelist: Some(false),

View file

@ -144,14 +144,14 @@ pub fn spawn(
/// Listen on the swarm's event subjects and act on what arrives.
///
/// The controller decides *what a forge delivery means* and addresses the
/// result here; this end does not know a forge exists. Two events today:
/// result here; this end reads no forge delivery. Two events today:
/// **the knowledge repository changed**, answered by the pull this daemon
/// already runs at boot, and **deploy this agent**, answered by the same
/// rebuild insert the operator's own verb makes.
///
/// Only the deploy event carries a payload, and only the agent name: its
/// subject already names the hive, so it listens on its own rather than a
/// swarm-wide feed. Even then it is a trigger, never the config git owns.
/// Only the deploy event carries a payload: the agent name, and the config
/// commit to deploy when the swarm names one. Its subject already names the
/// hive, so it listens on its own rather than a swarm-wide feed.
///
/// # A missed message costs the two events very differently
///
@ -264,7 +264,7 @@ async fn handle_deploy_request(
return;
}
};
let agent = request.agent;
let swarm_queue_client::DeployRequest { agent, rev } = request;
// Which of the two meanings this request has is decided here, and the
// predicate is "does a container exist", not "is one running":
@ -283,6 +283,15 @@ async fn handle_deploy_request(
}
};
// Only a rebuild applies `rev`. A first deploy ignores it and builds the
// `applied` repo it seeds from the forge's `main`, or the one it finds.
if known
&& let Some(rev) = rev.as_deref()
&& !apply_rev(&agent, rev).await
{
return;
}
let inserted = if known {
// The same insert the operator's own `rebuild` verb makes, relock and
// all: "deploy this agent" means here exactly what it already meant, and
@ -303,6 +312,29 @@ async fn handle_deploy_request(
}
}
/// Move `agent`'s `applied/main` to `rev` ahead of its rebuild, and say
/// whether that rebuild should be queued. A `rev` already applied is taken as
/// deployed or deploying — a dashboard merge deploys the PR it merges — so
/// no second rebuild is queued. That also skips a rev put in place by
/// `hivectl forge reconcile-config`, which does not deploy.
///
/// A refusal is commented on the PR that merged `rev`. A failed rebuild is
/// not: it surfaces where every rebuild failure does, on this hive.
async fn apply_rev(agent: &str, rev: &str) -> bool {
match crate::actions::advance_applied_to_rev(agent, rev).await {
Ok(crate::actions::RevAdvance::Advanced) => true,
Ok(crate::actions::RevAdvance::AlreadyApplied) => {
tracing::info!(%agent, %rev, "swarm events: deploy requested for the applied rev; nothing queued");
false
}
Err(e) => {
tracing::warn!(%agent, %rev, error = %format!("{e:#}"), "swarm events: rev not applied; nothing queued");
crate::actions::post_rev_deploy_failure_to_pr(agent, rev, &e).await;
false
}
}
}
/// Queue the first deploy of an agent this hive does not have yet: seed its
/// power intent, then insert the DAG.
///