Watch
0
0
Fork
You've already forked hyperhive
0

config PRs: remove the hive's config-PR webhook, poll and core merge

An operator's merge on the forge deploys a config PR through
swarm-controller's DeployRequest{rev}. The hive-side path that queued a
MergeConfigPr approval and merged the PR as `core` goes:

- the `/webhook/config-pr` receiver, its HMAC secret, the WebhookRegister
  boot node and the org-hook registration; the hive vhost's `/webhook/`
  location
- the 5-minute config-PR poll
- ApprovalKind::MergeConfigPr, its dashboard card, and the deploy DAG it
  drove (DeployWindow, MergeVerify, DeployApply, FinalizeDeploy,
  DeployTail), with verify_commit, the two-phase meta deploy, rollback
  refs, the PR-failure comment and forge/pr_merge.rs
- `fetched_sha`, `sha_short`/`pr_number` on approval events, and
  `sha`/`tag` on HelperEvent::ApprovalResolved: only the merge path set
  them

`config_repo`, `merged_pr_for_commit` and `post_pr_comment` move to
forge/pr_comment.rs for the merged-rev deploy's refusal comment.
Approvals v5 drops stored `merge_config_pr` rows; a test reopens a v4
database holding them.

Closes #4850
This commit is contained in:
atlas 2026-10-02 22:32:17 +02:00
commit efbfec6d01
39 changed files with 286 additions and 3199 deletions

View file

@ -1,18 +1,15 @@
//! Swarm-level config-PR status, kept current by both a webhook nudge and a
//! periodic poll.
//!
//! `hive-c0re::forge::config_pr_poll` already scans `agent-configs/*` for
//! open PRs — but it does that once per hive, to queue that hive's own
//! `MergeConfigPr` approval, and nothing at swarm level reads the result.
//! swarm-ui's config-PR panel needs a *swarm*-level answer to "does agent X
//! have an open config PR" that does not depend on which hive currently
//! hosts X being reachable.
//!
//! Both paths write the same [`ConfigPrCache`]:
//!
//! - [`spawn`] — a periodic full rescan, mirroring `hive-c0re`'s own poll
//! shape. The backstop: catches anything a missed delivery loses, and is
//! what populates the cache before the first delivery ever arrives.
//! - [`spawn`] — a periodic full rescan. The backstop: catches anything a
//! missed delivery loses, and is what populates the cache before the first
//! delivery ever arrives.
//! - [`ConfigPrCache::apply_webhook_delivery`] — called from
//! `crate::webhook::post_webhook_forge` on a verified `ConfigPr` delivery.
//! The low-latency path: a PR opening or closing shows up immediately
@ -21,11 +18,6 @@
//! [`merged`] reads the same delivery for a merge, which the webhook handler
//! turns into a deploy. The poll has no counterpart: it lists open PRs only,
//! so a merge whose delivery is lost deploys nothing.
//!
//! Per mara's review call: ship both from the start rather than the poll
//! alone — the eventual swarm-level replacement for `hive-c0re`'s own
//! poll+webhook pair needs both anyway, so building only half here would be
//! work redone rather than work reused.
use std::collections::HashMap;
use std::sync::{Arc, Mutex};
@ -119,9 +111,7 @@ struct WebhookRepository {
name: String,
}
/// How often to rescan `agent-configs/*`. Matches the interval named in
/// `hive-c0re::forge::config_pr_poll`'s own doc comment — same org, same
/// staleness tolerance, no reason for the two to disagree.
/// How often to rescan `agent-configs/*`.
const POLL_INTERVAL: std::time::Duration = std::time::Duration::from_mins(5);
/// The latest full scan, replaced atomically each cycle.

View file

@ -603,10 +603,7 @@ impl Client {
}
/// Every agent in [`CONFIG_ORG`] with an open config PR, keyed by agent
/// name. Mirrors `hive-c0re::forge::config_pr_poll::poll_open_config_prs`'s
/// scan shape (list repos in the org, list open PRs per repo) but returns
/// data instead of side-effecting an approval queue — this daemon has no
/// approval system of its own; it exists so `GET
/// name: list repos in the org, then open PRs per repo. It exists so `GET
/// /api/agents/{name}/config-pr` has something to answer from,
/// independent of any one hive being up.
///
@ -649,9 +646,7 @@ impl Client {
}
};
// Only the first open PR matters for the panel — a config repo
// is meant to carry at most one live proposal at a time (the
// same assumption `hive-c0re`'s poller and the `MergeConfigPr`
// approval flow both make).
// is meant to carry at most one live proposal at a time.
if let Some(pr) = prs.into_iter().next() {
let Some(pr_number) = pr.number.and_then(|n| u64::try_from(n).ok()) else {
continue;
@ -900,19 +895,9 @@ impl Client {
/// forge event reaches [`crate::webhook`] instead of the endpoint only
/// being reachable by hand.
///
/// **These are registered ALONGSIDE the per-hive hooks, not instead of
/// them.** Every hive keeps receiving and acting on its own deliveries
/// exactly as today; the controller receives a copy and (for now) logs
/// it. Taking the hive-side registration away is a later step, and it
/// has to be later: fan-out swarm→hive does not exist yet, so a hook
/// moved now would point at a receiver that forwards nowhere — silent on
/// both sides, indistinguishable from no activity.
///
/// ⛔ **No stale-hook deletion arm, unlike the two per-hive registrars
/// this otherwise mirrors.** Their arm deletes hooks matching their own
/// path with a foreign base; copying it here would delete the hives'
/// live hooks, which are not stale — they are the path still in
/// production. The controller only ever adds its own.
/// Adds its own hooks only. A hive's `/webhook/config-pr` hook left on
/// the `agent-configs` org by an older release stays registered and
/// delivers to a route no hive serves.
///
/// Idempotent: an existing hook with the same `target_url` is left
/// alone, so this is safe on every boot.
@ -921,8 +906,7 @@ impl Client {
///
/// Returns the first failure. A listing failure is not fatal — it falls
/// through to the create attempt, which is idempotent server-side by way
/// of the already-exists fold, the same best-effort shape `hive-c0re`
/// uses.
/// of the already-exists fold.
pub async fn ensure_swarm_webhooks(&self, public_base: &str, secret: &str) -> Result<()> {
for kind in DeliveryKind::ALL {
let target_url = kind.target_url(public_base);
@ -989,7 +973,7 @@ impl Client {
);
}
Err(e) => {
// Best-effort, same as the per-hive registrars: a forge that
// Best-effort: a forge that
// cannot be listed may still accept a create, and a
// duplicate create is folded into success below. `warn!`
// because that fold is an assumption about the forge, not a
@ -1054,8 +1038,7 @@ fn transitive_reach(start: i64, adj: &HashMap<i64, Vec<i64>>) -> usize {
}
/// Where a hook lives. The knowledge hook is repo-scoped and the config-PR
/// hook is org-scoped, mirroring exactly where the per-hive registrars put
/// theirs — a hook on the wrong scope would never fire, and forgejo would
/// hook is org-scoped — a hook on the wrong scope would never fire, and forgejo would
/// report that as a perfectly healthy hook with no deliveries.
enum HookScope<'a> {
Repo {

View file

@ -10,16 +10,6 @@
//! the payload and emits a semantic message — *the knowledge repo changed*,
//! *deploy agent X at rev Y* — addressed to the hives that need it. One place
//! decides what a forge payload means, so no hive re-derives it.
//!
//! Receipt is all that is wired today: the payload is opaque bytes keyed by a
//! [`DeliveryKind`] from the URL path, and nothing consumes it until the
//! swarm→hive channel exists.
//!
//! The HMAC code is deliberately **not** shared with `hive-c0re`: that copy is
//! leaving, and a shared crate is right only when a second consumer arrives.
//!
//! **These hooks are registered ALONGSIDE the per-hive ones** — see
//! [`crate::forge::Client::ensure_swarm_webhooks`].
use anyhow::{Context as _, Result};
use axum::{
@ -199,15 +189,11 @@ pub(super) enum DeliveryKind {
/// The route prefix a registered `target_url` must point at.
///
/// ⚠️ Deliberately **not** `/webhook/knowledge` or `/webhook/config-pr`, the
/// paths the per-hive receivers use. A hive-side registrar deletes any hook
/// whose URL ends with *its* path but has a different base — see
/// `hive-c0re`'s `forge::ensure_config_pr_webhook`. A swarm-level hook under
/// such a path would therefore be deleted by every hive on every boot, and
/// the symptom is a hook that silently stops existing.
/// `webhook_urls_survive_the_hive_side_reapers` pins that, and its own doc
/// records why `/webhook/knowledge` stays in the check even though the
/// knowledge registrar no longer reaps.
/// ⚠️ Deliberately **not** `/webhook/knowledge` or `/webhook/config-pr`.
/// Older `hive-c0re` releases delete, on every boot, any hook whose URL ends
/// with one of those paths on a base other than their own; a swarm-level
/// hook there would silently stop existing while such a hive runs.
/// `webhook_urls_survive_the_hive_side_reapers` pins that.
const ROUTE_PREFIX: &str = "/webhook/forge/";
impl DeliveryKind {
@ -661,22 +647,18 @@ mod tests {
}
}
/// A cross-daemon invariant with nothing else to enforce it: a per-hive
/// registrar in `hive-c0re` **deletes** hooks whose URL ends with its
/// own path but carries a different base. A controller URL matching such
/// a suffix would be deleted by every hive on every boot — the swarm hook
/// would simply cease to exist, with the cause in a different daemon's
/// startup sweep. Serving these under `/webhook/forge/` is what avoids
/// it, and this is the only place that says so in a form that fails.
/// A cross-daemon invariant with nothing else to enforce it: older
/// `hive-c0re` releases **delete** hooks whose URL ends with
/// `/webhook/knowledge` or `/webhook/config-pr` but carries a base other
/// than their own. A controller URL matching such a suffix would be
/// deleted by every such hive on every boot, with the cause in a
/// different daemon's startup sweep. Serving these under
/// `/webhook/forge/` is what avoids it, and this is the only place that
/// says so in a form that fails.
///
/// `forge::ensure_config_pr_webhook` still reaps that way. The knowledge
/// registrar no longer does — it was replaced by a removal that matches
/// the full URL, so it cannot touch another hive's hook. **The
/// `/webhook/knowledge` arm is kept anyway**, because the hazard is not
/// this repository's current code: it is whatever is *deployed*, and a
/// hive still running the previous version reaps by suffix until it is
/// upgraded. Drop that arm once no such hive can exist, not when the
/// source stops mentioning it.
/// Current `hive-c0re` removes only hooks matching its own full URL. The
/// hazard is what is *deployed*, not this tree: drop this test once no
/// hive running an older release can exist.
#[test]
fn webhook_urls_survive_the_hive_side_reapers() {
for suffix in ["/webhook/knowledge", "/webhook/config-pr"] {