swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the operators team in agents and agent-configs, the pull-mirrors, internal/docs, internal/knowledge (public, README-seeded) and the agent-configs org avatar are one set per forge. hive-c0re ensured them in its boot sweep, as the core admin, and only on the hive co-located with the forge container. swarm-controller now reconciles them at start and every 5 minutes (forge/objects.rs: observe -> pure plan -> apply). A failed object logs a warn line plus a pass summary and is retried next tick. create_repo ensures the agent-configs org and its operators team first, so a config repo's merge gate never depends on the periodic pass having run. hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo, ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/ set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot knowledge::remove_webhook cleanup, with their now-unused helpers. nix: the mirror list moves from the hive-c0re unit (HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit (SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are declared on a host that runs no controller. c0re.orgAvatarPng is renamed to deploy.swarm-controller.configOrgAvatarPng. Refs #3782
This commit is contained in:
parent
85ba45b2de
commit
20419ccd41
18 changed files with 1456 additions and 822 deletions
|
|
@ -34,6 +34,7 @@ use utoipa::ToSchema;
|
|||
use crate::webhook::DeliveryKind;
|
||||
|
||||
pub mod agent_token;
|
||||
pub mod objects;
|
||||
pub mod site_admin;
|
||||
|
||||
/// An agent's open config-PR, as [`Client::list_open_config_prs`] reports it
|
||||
|
|
@ -81,12 +82,9 @@ pub struct IssueReportRow {
|
|||
}
|
||||
|
||||
/// The `operators` team, whitelisted for the merge gate on every repo
|
||||
/// this client protects — provisioned by `hive-c0re::forge::repos`
|
||||
/// already (`ensure_operators_team`), not re-provisioned here. If that
|
||||
/// assumption ever breaks (this becomes reachable before any hive's
|
||||
/// `hive-c0re` has run its startup sweep), branch-protection creation
|
||||
/// below will fail loudly rather than silently no-op — see
|
||||
/// [`Client::create_repo`]'s doc comment.
|
||||
/// this client protects. Provisioned by this daemon's own periodic pass
|
||||
/// ([`objects`]), and checked again by [`Client::create_repo`] before it
|
||||
/// applies a rule naming the team.
|
||||
const OPERATORS_TEAM: &str = "operators";
|
||||
|
||||
/// The org owning per-agent config repos — where this daemon creates them,
|
||||
|
|
@ -105,7 +103,7 @@ const OPERATORS_TEAM: &str = "operators";
|
|||
/// config repo created in `agents` is invisible to every one of those paths,
|
||||
/// and nothing errors, because both orgs exist and both accept a repo.
|
||||
///
|
||||
/// The merge gate survives the distinction: `hive-c0re` provisions the
|
||||
/// The merge gate survives the distinction: [`objects`] provisions the
|
||||
/// `operators` team in **both** orgs precisely so branch protection can be
|
||||
/// applied in either.
|
||||
const CONFIG_ORG: &str = "agent-configs";
|
||||
|
|
@ -288,12 +286,9 @@ impl Client {
|
|||
/// collaborator, not on this whitelist) still cannot push directly,
|
||||
/// so this doesn't loosen the "can't merge your own PR" guarantee.
|
||||
///
|
||||
/// Assumes [`OPERATORS_TEAM`] already exists in [`CONFIG_ORG`]
|
||||
/// (provisioned by `hive-c0re::forge::repos::ensure_operators_team`
|
||||
/// on its own startup sweep, not re-provisioned here) — if it
|
||||
/// doesn't yet, this fails loudly rather than silently leaving the
|
||||
/// repo unprotected, which is the correct failure mode for a
|
||||
/// prerequisite that's supposed to already be there.
|
||||
/// Needs [`OPERATORS_TEAM`] to exist in [`CONFIG_ORG`], which
|
||||
/// [`Client::create_repo`] ensures first — if it still doesn't, this
|
||||
/// fails loudly rather than silently leaving the repo unprotected.
|
||||
async fn apply_operator_branch_protection(&self, repo: &str) -> Result<()> {
|
||||
let rule = CreateBranchProtectionOption {
|
||||
apply_to_admins: None,
|
||||
|
|
@ -357,7 +352,12 @@ impl Client {
|
|||
/// doesn't exist yet"), unlike collaborator-add and config-seed, which
|
||||
/// are genuinely separate operations against an already-existing repo.
|
||||
/// Idempotent — safe to call again.
|
||||
///
|
||||
/// Ensures the org and its [`OPERATORS_TEAM`] first: the periodic
|
||||
/// [`objects`] pass does too, but an agent can be created before that
|
||||
/// pass has succeeded once, and the rule below names the team.
|
||||
pub async fn create_repo(&self, repo: &str) -> Result<String> {
|
||||
self.ensure_merge_gate_prerequisites().await?;
|
||||
self.ensure_org_repo(repo).await?;
|
||||
self.apply_operator_branch_protection(repo).await?;
|
||||
tracing::info!(%repo, "swarm forge: created repo in {CONFIG_ORG} with operator merge gate");
|
||||
|
|
|
|||
1234
swarm-controller/src/forge/objects.rs
Normal file
1234
swarm-controller/src/forge/objects.rs
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -1910,6 +1910,25 @@ fn spawn_matrix_account_backfill(
|
|||
});
|
||||
}
|
||||
|
||||
/// Start the forge's periodic passes — the swarm-wide objects and agents'
|
||||
/// tokens — when a forge is configured. Lifted out of `main` for
|
||||
/// `clippy::too_many_lines`.
|
||||
fn spawn_forge_workers(
|
||||
jobq: &Arc<Mutex<hive_jobq::scheduler::Scheduler<SwarmNodeKind, SwarmResourceKind>>>,
|
||||
forge_client: Option<Arc<forge::Client>>,
|
||||
) {
|
||||
let Some(client) = forge_client else {
|
||||
return;
|
||||
};
|
||||
forge::objects::spawn(Arc::clone(&client));
|
||||
let sched = Arc::clone(jobq);
|
||||
forge::agent_token::spawn(client, move |agents| {
|
||||
if let Err(e) = queue_forge_token_mints(&sched, agents) {
|
||||
tracing::warn!(error = %format!("{e:#}"), "agent forge tokens: queueing failed");
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
/// Check an agent's forge token now, and mint one if it is missing or stale.
|
||||
///
|
||||
/// The periodic pass (`forge::agent_token::spawn`) does the same every five
|
||||
|
|
@ -2342,14 +2361,7 @@ async fn main() -> Result<()> {
|
|||
);
|
||||
}
|
||||
|
||||
if let Some(client) = forge_client.clone() {
|
||||
let sched = Arc::clone(&jobq);
|
||||
forge::agent_token::spawn(client, move |agents| {
|
||||
if let Err(e) = queue_forge_token_mints(&sched, agents) {
|
||||
tracing::warn!(error = %format!("{e:#}"), "agent forge tokens: queueing failed");
|
||||
}
|
||||
});
|
||||
}
|
||||
spawn_forge_workers(&jobq, forge_client.clone());
|
||||
let config_prs = forge_client.clone().map(config_pr::spawn);
|
||||
let state_forge = keep_forge_for_state(forge_client, webhook_secret.clone());
|
||||
|
||||
|
|
|
|||
|
|
@ -69,11 +69,23 @@ fn secret_path() -> std::path::PathBuf {
|
|||
/// tests setting it race — which is not hypothetical here, it is how the
|
||||
/// first version of this module's tests failed.
|
||||
fn secret_path_from(raw: Option<&str>) -> std::path::PathBuf {
|
||||
state_dir_from(raw).join(SECRET_FILE)
|
||||
}
|
||||
|
||||
/// This daemon's state directory, read the same way [`secret_path`] reads
|
||||
/// it. Other state files (the forge avatar marker in
|
||||
/// `crate::forge::objects`) live beside the secret, so both share this one
|
||||
/// reading of `STATE_DIRECTORY`.
|
||||
pub fn state_dir() -> std::path::PathBuf {
|
||||
state_dir_from(std::env::var("STATE_DIRECTORY").ok().as_deref())
|
||||
}
|
||||
|
||||
fn state_dir_from(raw: Option<&str>) -> std::path::PathBuf {
|
||||
let dir = raw
|
||||
.and_then(|raw| raw.split(':').next())
|
||||
.filter(|first| !first.is_empty())
|
||||
.unwrap_or(DEFAULT_STATE_DIR);
|
||||
std::path::PathBuf::from(dir).join(SECRET_FILE)
|
||||
std::path::PathBuf::from(dir)
|
||||
}
|
||||
|
||||
/// Load the swarm's webhook HMAC secret, generating and persisting it if the
|
||||
|
|
|
|||
Loading…
Reference in a new issue