Watch
0
0
Fork
You've already forked hyperhive
0

swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them

The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
This commit is contained in:
atlas 2026-09-24 15:00:10 +02:00
commit 20419ccd41
18 changed files with 1456 additions and 822 deletions

View file

@ -1,7 +1,7 @@
//! Optional Forgejo wiring — per-agent account alignment,
//! config-repo mirroring, meta read-access grants. Also seeds
//! `internal/docs` — a private repo every agent gets read-only
//! collaborator access to for operator-curated shared content.
//! config-repo mirroring, meta read-access grants, and each agent's
//! read-only grant on `internal/docs` (the swarm-controller creates that
//! repo, with the other swarm-wide orgs, teams and repos).
//! No-op when `hive-forge` isn't running. Full design: `docs/integrations/forge.md`.
mod ci_runner;
@ -17,9 +17,9 @@ pub use pr_merge::{
};
pub use reconcile::{reconcile_config_apply, reconcile_config_status};
pub use repos::{
clone_config_into_proposed, create_agent_repo, ensure_config_repo, ensure_knowledge_repo,
ensure_meta_remote, ensure_repo, ensure_shared_docs_repo, fast_forward_applied_main,
fetch_config_main_into_applied, meta_read_access, push_config, push_meta, shared_docs_access,
clone_config_into_proposed, create_agent_repo, ensure_config_repo, ensure_meta_remote,
ensure_repo, fast_forward_applied_main, fetch_config_main_into_applied, meta_read_access,
push_config, push_meta, shared_docs_access,
};
pub use users::core_token;
@ -30,10 +30,9 @@ use anyhow::{Context, Result};
use forgejo_api::{Auth, Forgejo};
use url::Url;
use repos::{ensure_mirrors, ensure_operators_team, ensure_org};
use users::{
ensure_config_org_avatar, ensure_core_avatar, ensure_core_user_and_token,
ensure_repo_creation_disabled, ensure_user_email,
ensure_core_avatar, ensure_core_user_and_token, ensure_repo_creation_disabled,
ensure_user_email,
};
const FORGE_CONTAINER: &str = "hive-forge";
@ -107,36 +106,30 @@ pub(crate) const CONFIG_ORG: &str = "agent-configs";
/// that every agent gets read-only access to. Agents use it as a
/// common reference without the operator having to bake content into
/// the system prompt or rely on `/shared`. Only the manager + operator
/// (i.e. `core` user) can push.
/// (i.e. `core` user) can push. The swarm-controller creates the org
/// and the repo; this hive only grants its agents read access.
const SHARED_ORG: &str = "internal";
/// The shared docs repo inside `SHARED_ORG`. Cloneable by every agent
/// at `{forge_http_base()}/internal/docs.git`.
const SHARED_DOCS_REPO: &str = "docs";
/// The hive-wide knowledge repo inside `SHARED_ORG`. Public — agents
/// can fork it and open PRs without explicit collaborator grants.
/// Bind-mounted read-only into every container at `/knowledge`.
/// See `hive-c0re/src/workers/knowledge.rs`.
const KNOWLEDGE_REPO: &str = crate::knowledge::REPO;
/// Forgejo org that owns agent-created repos. Agents can't create
/// repos with their own token (`max_repo_creation = 0`); instead hive-c0re
/// creates them here and adds the requesting agent as a **write** member
/// (not owner/admin). Because the org — not the agent — owns the repo,
/// perms stay c0re-managed and branch protection (referencing
/// [`OPERATORS_TEAM`]) can block the author from merging their own PR. This
/// is the "agents namespace" repos land in by default.
/// is the "agents namespace" repos land in by default. The swarm-controller
/// ensures the org itself.
const AGENTS_ORG: &str = "agents";
/// Operator merge-gate team inside [`AGENTS_ORG`]. Provisioned **empty** by
/// hive-c0re (so perms can be set before anyone joins); the operator adds
/// herself via the forge UI / hivectl. Branch protection on agents-org repos
/// references this team by name for the merge/approval whitelist, so the
/// rule never hardcodes a specific reviewer agent (which may not exist).
/// the swarm-controller (so perms can be set before anyone joins); the
/// operator adds herself via the forge UI. Branch protection on agents-org
/// repos references this team by name for the merge/approval whitelist, so
/// the rule never hardcodes a specific reviewer agent (which may not exist).
const OPERATORS_TEAM: &str = "operators";
/// Forgejo orgs hive-c0re ensures on startup. The meta repo lives at
/// `core/meta` (the `core` user's own namespace — no org needed).
const SEEDED_ORGS: &[&str] = &[CONFIG_ORG, SHARED_ORG, AGENTS_ORG];
/// Leak `s` to get a `&'static str` warning `kind` for the small, bounded
/// set of per-org boot warnings in [`ensure_all`] (one per seeded org, at
/// set of boot warnings in [`ensure_all`] keyed by a runtime name (at
/// most a handful per process). [`crate::warnings::set_boot_warning`]
/// requires a `'static` kind so distinct orgs/repos don't clobber each
/// other's banner entry; leaking a few short strings once per boot is
@ -277,44 +270,15 @@ pub async fn sync_agent(name: &str, core_token: Option<&str>) -> bool {
ok
}
/// The `core_token.is_some()` half of [`ensure_all`]: orgs, teams, the meta
/// repo, shared/knowledge repos, avatars, and CI runner registration — every
/// step that needs an authenticated forge client. Split out purely to keep
/// The `core_token.is_some()` half of [`ensure_all`]: the meta repo, the
/// local knowledge clone, the core avatar, and CI runner registration —
/// every step that needs an authenticated forge client. The swarm-wide
/// objects (orgs, the `operators` team, mirrors, `internal/docs`,
/// `internal/knowledge`, the `agent-configs` avatar) are the
/// swarm-controller's, not this hive's. Split out purely to keep
/// `ensure_all` under clippy's function-length limit; not meant to be called
/// from anywhere else.
async fn ensure_all_orgs_and_repos(token: &str) {
for org in SEEDED_ORGS {
if let Err(e) = ensure_org(org, token).await {
tracing::warn!(%org, error = ?e, "forge: ensure_org failed");
crate::warnings::set_boot_warning(
static_kind(format!("forge_ensure_org_{org}")),
"crit",
format!("forge: org {org} provisioning failed: {e}"),
);
}
}
// Seed the operator-declared pull-mirrors (nix `forge.mirrors` +
// the CI-auto `actions/checkout`, forwarded via the
// `HYPERHIVE_FORGE_MIRRORS` env). Each ensures its own dest org, so
// this is independent of the SEEDED_ORGS loop above.
ensure_mirrors(token).await;
// Provision the operator merge-gate team (empty) inside BOTH the
// agents org and the agent-configs org so branch protection in each
// can reference it before anyone joins. Gitea teams are org-scoped —
// missing the agent-configs copy 422'd every config-repo protection
// apply, leaving those repos unprotected and letting operator-merged
// config PRs bypass the deploy pipeline. The operator adds herself as
// a member out-of-band.
for org in [AGENTS_ORG, CONFIG_ORG] {
if let Err(e) = ensure_operators_team(org, token).await {
tracing::warn!(%org, error = ?e, "forge: ensure_operators_team failed");
crate::warnings::set_boot_warning(
static_kind(format!("forge_ensure_operators_team_{org}")),
"crit",
format!("forge: operators team in {org} provisioning failed: {e}"),
);
}
}
// Meta repo lives at core/meta — pushed from git_commit in
// meta.rs on every deploy/lock-update. Make sure it exists
// before the first push hits a 404.
@ -326,26 +290,8 @@ async fn ensure_all_orgs_and_repos(token: &str) {
format!("forge: core/meta repo provisioning failed: {e}"),
);
}
// Seed the shared docs repo. internal is already in
// SEEDED_ORGS above so the org exists; ensure the repo itself.
if let Err(e) = ensure_shared_docs_repo(token).await {
tracing::warn!(error = ?e, "forge: ensure_shared_docs_repo failed");
crate::warnings::set_boot_warning(
"forge_ensure_shared_docs_repo",
"warn",
format!("forge: shared docs repo provisioning failed: {e}"),
);
}
// Seed the hive-wide knowledge repo.
if let Err(e) = ensure_knowledge_repo(token).await {
tracing::warn!(error = ?e, "forge: ensure_knowledge_repo failed");
crate::warnings::set_boot_warning(
"forge_ensure_knowledge_repo",
"crit",
format!("forge: knowledge repo provisioning failed: {e}"),
);
}
// Clone knowledge repo locally so it can be bind-mounted into agents.
// The swarm-controller creates and seeds the repo itself.
if let Err(e) = crate::knowledge::ensure_local_clone(token).await {
tracing::warn!(error = ?e, "knowledge: ensure_local_clone failed");
crate::warnings::set_boot_warning(
@ -362,14 +308,6 @@ async fn ensure_all_orgs_and_repos(token: &str) {
format!("forge: core avatar upload failed: {e}"),
);
}
if let Err(e) = ensure_config_org_avatar(token).await {
tracing::warn!(error = ?e, "forge: ensure_config_org_avatar failed");
crate::warnings::set_boot_warning(
"forge_ensure_config_org_avatar",
"warn",
format!("forge: agent-configs org avatar upload failed: {e}"),
);
}
// Register the hive-ci Actions runner (off the container's boot path;
// no-op when CI is disabled or the runner already holds valid creds).
ci_runner::ensure_ci_runner_registered(token).await;
@ -419,7 +357,8 @@ async fn wait_until_ready() -> bool {
/// each has a forgejo user + token, plus an `agent-configs/<name>`
/// repo mirroring its applied config. Also seeds the `core` admin
/// user (hive-c0re's own identity for pushing the meta repo + driving
/// the API), the `agent-configs` org, and the `core/meta` repo.
/// the API) and the `core/meta` repo. The `agent-configs` org itself is
/// the swarm-controller's to ensure.
/// Called once at hive-c0re startup. Per-step failures are logged
/// but don't abort the sweep.
pub async fn ensure_all() {
@ -495,9 +434,9 @@ pub async fn ensure_all() {
/// An org-level hook covers every repo in `agent-configs` automatically,
/// so no per-repo setup is needed as new agents are provisioned.
///
/// Called at startup beside `knowledge::remove_webhook`, its opposite: that
/// repo's one hook is the controller's now, this one has not moved yet. No-op
/// when the core token is absent (forge not yet provisioned).
/// Called at startup. The knowledge repo's one hook is the controller's
/// now; this one has not moved yet. No-op when the core token is absent
/// (forge not yet provisioned).
///
/// # Errors
///

View file

@ -10,9 +10,7 @@ use std::sync::Mutex;
use anyhow::{Context, Result};
use forgejo_api::structs::{
AddCollaboratorOption, AddCollaboratorOptionPermission, CreateBranchProtectionOption,
CreateOrgOption, CreateRepoOption, CreateTeamOption, CreateTeamOptionPermission,
EditBranchProtectionOption, EditRepoOption, EditTeamOption, EditTeamOptionPermission,
MigrateRepoOptions, MigrateRepoOptionsService, Repository,
CreateRepoOption, EditBranchProtectionOption, Repository,
};
use forgejo_api::{ApiErrorKind, ForgejoError};
use reqwest::StatusCode;
@ -20,8 +18,8 @@ use reqwest::StatusCode;
use crate::coordinator::Coordinator;
use super::{
AGENTS_ORG, CONFIG_ORG, KNOWLEDGE_REPO, OPERATORS_TEAM, SHARED_DOCS_REPO, SHARED_ORG, api,
core_auth_header, core_token, forge_git_url, forge_http_base, is_present,
AGENTS_ORG, CONFIG_ORG, OPERATORS_TEAM, SHARED_DOCS_REPO, SHARED_ORG, api, core_auth_header,
core_token, forge_git_url, forge_http_base, is_present,
};
/// Creation options for an empty repo defaulting to `main`.
@ -42,48 +40,6 @@ fn repo_option(name: &str, private: bool) -> CreateRepoOption {
}
}
/// `EditRepoOption` with every field unset — repo edits only ever
/// change the one field the caller sets on top (Forgejo leaves `None`
/// fields untouched).
fn sparse_edit_repo_option() -> EditRepoOption {
EditRepoOption {
allow_fast_forward_only_merge: None,
allow_manual_merge: None,
allow_merge_commits: None,
allow_rebase: None,
allow_rebase_explicit: None,
allow_rebase_update: None,
allow_squash_merge: None,
archived: None,
autodetect_manual_merge: None,
default_allow_maintainer_edit: None,
default_branch: None,
default_delete_branch_after_merge: None,
default_merge_style: None,
default_update_style: None,
description: None,
enable_prune: None,
external_tracker: None,
external_wiki: None,
globally_editable_wiki: None,
has_actions: None,
has_issues: None,
has_packages: None,
has_projects: None,
has_pull_requests: None,
has_releases: None,
has_wiki: None,
ignore_whitespace_conflicts: None,
internal_tracker: None,
mirror_interval: None,
name: None,
private: None,
template: None,
website: None,
wiki_branch: None,
}
}
/// Whether a create-style call failed because the object already
/// exists. Forgejo signals this as HTTP 409 (conflict) or 422
/// (validation). The typed client surfaces those as
@ -104,59 +60,6 @@ fn is_already_exists(e: &ForgejoError) -> bool {
}
}
/// Whether an error is specifically an HTTP 409 conflict (and NOT a
/// 422): the migrate endpoint's 422 is a validation error (bad
/// `clone_addr` / service) and must surface, so it can't share
/// [`is_already_exists`]'s 422 tolerance.
fn is_conflict(e: &ForgejoError) -> bool {
match e {
ForgejoError::ApiError(api) => {
matches!(api.error_kind(), ApiErrorKind::Other(s) if *s == StatusCode::CONFLICT)
}
ForgejoError::UnexpectedStatusCode(s) => *s == StatusCode::CONFLICT,
_ => false,
}
}
/// Whether a rendered error message names the "already exists" case. The
/// discriminator that tells a *benign* already-exists 422 apart from a
/// *real* validation 422 (invalid units, etc.). Case-insensitive.
fn message_says_already_exists(rendered: &str) -> bool {
rendered.to_lowercase().contains("already exists")
}
/// Whether `e` is Forgejo saying the resource already exists — matching
/// the 409 conflict shape *and* the 422 shape this Forgejo build actually
/// returns for a duplicate team: `validation failed: team already exists`.
/// A 409-only [`is_conflict`] check misses that 422, so the caller fires
/// a spurious "provisioning failed" warning every boot and skips its
/// settings reconcile. This still surfaces *other* 422s (bad request
/// body) as real failures — only a 422 whose message names the
/// already-exists case is folded in.
fn is_already_exists_lenient(e: &ForgejoError) -> bool {
if is_conflict(e) {
return true;
}
let is_unprocessable = match e {
ForgejoError::ApiError(api) => matches!(api.error_kind(), ApiErrorKind::ValidationFailed),
ForgejoError::UnexpectedStatusCode(s) => *s == StatusCode::UNPROCESSABLE_ENTITY,
_ => false,
};
is_unprocessable && message_says_already_exists(&e.to_string())
}
/// Whether an error is Forgejo saying 404 — the resource is absent,
/// as opposed to a transport / auth / server failure.
fn is_not_found(e: &ForgejoError) -> bool {
match e {
ForgejoError::ApiError(api) => {
matches!(api.error_kind(), ApiErrorKind::NotFound { .. })
}
ForgejoError::UnexpectedStatusCode(s) => *s == StatusCode::NOT_FOUND,
_ => false,
}
}
/// Fold a repo-creation result's "already exists" (409 / 422) into
/// success. `label` is `<owner>/<name>` — purely for log + error
/// context.
@ -174,28 +77,6 @@ fn created_or_exists(res: Result<Repository, ForgejoError>, label: &str) -> Resu
}
}
/// Set an existing repo to public visibility. No-op if the repo is
/// already public. Used for `internal/knowledge` which may have been
/// created as private on an older deployment.
async fn set_repo_public(owner: &str, repo: &str, token: &str) -> Result<()> {
let mut edit = sparse_edit_repo_option();
edit.private = Some(false);
api(token)?
.repo_edit(owner, repo, edit)
.await
.with_context(|| format!("edit {owner}/{repo} (set public)"))?;
tracing::debug!(%owner, %repo, "forge: repo set to public");
Ok(())
}
/// Create `name` inside org `org` as a public repo. Idempotent.
async fn ensure_org_repo_public(org: &str, name: &str, token: &str) -> Result<()> {
let res = api(token)?
.create_org_repo(org, repo_option(name, false))
.await;
created_or_exists(res, &format!("{org}/{name}"))
}
/// Create a repo in the token-owner's own namespace. `token` belongs
/// to the user we want the repo owned by (we use `core`'s token for
/// `core/meta`). Idempotent.
@ -356,13 +237,6 @@ fn record_branch_protection_result(name: &str, ok: bool) {
}
}
/// Ensure the `internal/docs` repo exists. Called once at startup
/// after `ensure_org(SHARED_ORG)`. Idempotent — `ensure_org_repo`
/// treats 409 as success.
pub async fn ensure_shared_docs_repo(core_token: &str) -> Result<()> {
ensure_org_repo(SHARED_ORG, SHARED_DOCS_REPO, core_token).await
}
/// Grant agent `name` read-only collaborator access to `internal/docs`.
/// Idempotent: re-adding an existing collaborator succeeds (Forgejo
/// answers 204 either way). Mirrors `meta_read_access` so agents can
@ -380,19 +254,6 @@ pub async fn shared_docs_access(name: &str, core_token: &str) -> Result<()> {
Ok(())
}
/// Ensure the `internal/knowledge` repo exists and is public.
/// Called once at startup after `ensure_org(SHARED_ORG)`. Idempotent.
///
/// The repo is created as public so any agent with a forge account can
/// fork it and open PRs to contribute. Existing deployments that ended
/// up with a private repo are patched to public on the next hive-c0re
/// startup via `set_repo_public`.
pub async fn ensure_knowledge_repo(core_token: &str) -> Result<()> {
ensure_org_repo_public(SHARED_ORG, KNOWLEDGE_REPO, core_token).await?;
// Ensure public even if the repo already existed as private (older deployment).
set_repo_public(SHARED_ORG, KNOWLEDGE_REPO, core_token).await
}
/// Grant agent `name` read-only collaborator access to `core/meta` on
/// the forge so the agent can clone/fetch the meta flake. Idempotent:
/// re-adding an existing collaborator succeeds (Forgejo answers 204
@ -732,261 +593,6 @@ async fn run_config_push(
.context("invoke git push agent-configs")
}
/// Create an org named `name` (`org_create`). Idempotent: HTTP 422
/// ("user already exists") / 409 is treated as success.
pub(super) async fn ensure_org(name: &str, admin_token: &str) -> Result<()> {
let org = CreateOrgOption {
description: None,
email: None,
full_name: None,
location: None,
repo_admin_change_team_access: None,
username: name.to_owned(),
visibility: None,
website: None,
};
match api(admin_token)?.org_create(org).await {
Ok(_) => {
tracing::info!(%name, "forge: created org");
Ok(())
}
Err(e) if is_already_exists(&e) => {
tracing::debug!(%name, "forge: org already exists");
Ok(())
}
Err(e) => Err(e).with_context(|| format!("create org {name}")),
}
}
/// One operator-declared pull-mirror, forwarded from the nix
/// `services.hyperhive.forge.mirrors` option as JSON in
/// `HYPERHIVE_FORGE_MIRRORS`.
#[derive(serde::Deserialize)]
struct Mirror {
/// Upstream clone URL to mirror from (e.g. `https://github.com/actions/checkout`).
upstream: String,
/// Local `<owner>/<repo>` the mirror is created at.
dest: String,
}
/// Ensure each `HYPERHIVE_FORGE_MIRRORS` entry exists as a real Forgejo
/// pull-mirror. The env carries the JSON-encoded nix `forge.mirrors` list
/// (plus the CI-auto `actions/checkout` entry). Absent/empty env = no-op.
/// Per-mirror failures warn and continue — never abort the startup sweep.
pub(super) async fn ensure_mirrors(admin_token: &str) {
let raw = match std::env::var("HYPERHIVE_FORGE_MIRRORS") {
Ok(s) if !s.trim().is_empty() => s,
_ => return,
};
let mirrors: Vec<Mirror> = match serde_json::from_str(&raw) {
Ok(m) => m,
Err(e) => {
tracing::warn!(error = ?e, "forge: HYPERHIVE_FORGE_MIRRORS is not valid JSON; skipping mirror seed");
return;
}
};
for m in mirrors {
let Some((owner, repo)) = m.dest.split_once('/') else {
tracing::warn!(dest = %m.dest, "forge: mirror dest is not <owner>/<repo>; skipping");
continue;
};
// Create the dest org first (idempotent); the mirror can't land
// without its owner existing.
if let Err(e) = ensure_org(owner, admin_token).await {
tracing::warn!(%owner, error = ?e, "forge: ensure_org for mirror failed");
continue;
}
if let Err(e) = ensure_mirror_repo(&m.upstream, owner, repo, admin_token).await {
tracing::warn!(dest = %m.dest, error = ?e, "forge: ensure_mirror_repo failed");
}
}
}
/// Periodic sync interval for pull-mirrors. Forgejo syncs mirrors
/// on-access by default, which re-introduces external DNS latency on
/// every `git clone` (the hive-ci runner shares the host netns and is
/// therefore affected by host resolver blips). A fixed periodic interval
/// isolates CI from transient DNS failures — a stale mirror is
/// acceptable; a broken clone because of a momentary DNS blip is not.
const MIRROR_INTERVAL: &str = "8h0m0s";
/// Create `owner/repo` as a pull-mirror of `upstream` via the migrate API.
/// Idempotent: if the repo already exists this function patches its
/// `mirror_interval` to ensure it matches (covers mirrors that were
/// created before the interval was introduced). A 409 on the migrate
/// call (a race between the existence check and the migrate) is also
/// success.
async fn ensure_mirror_repo(
upstream: &str,
owner: &str,
repo: &str,
admin_token: &str,
) -> Result<()> {
let client = api(admin_token)?;
match client.repo_get(owner, repo).await {
Ok(_) => {
// Mirror already present. Patch interval so mirrors seeded before
// this field was introduced (or with a different value) converge.
let mut edit = sparse_edit_repo_option();
edit.mirror_interval = Some(MIRROR_INTERVAL.to_owned());
match client.repo_edit(owner, repo, edit).await {
Ok(_) => {
tracing::debug!(%owner, %repo, interval = MIRROR_INTERVAL, "forge: pull-mirror interval updated");
}
Err(e) => {
tracing::warn!(
%owner, %repo, error = %e,
"forge: failed to set mirror_interval on existing pull-mirror"
);
}
}
return Ok(());
}
// Absent — fall through to migrate.
Err(e) if is_not_found(&e) => {}
// Anything else (transport, auth, 5xx) leaves the repo's existence
// unknown: migrating anyway would fold a 409 into success and skip
// the interval patch this pass. Surface it instead.
Err(e) => {
return Err(e).with_context(|| format!("get pull-mirror {owner}/{repo}"));
}
}
let opts = MigrateRepoOptions {
auth_password: None,
auth_token: None,
auth_username: None,
clone_addr: upstream.to_owned(),
description: None,
issues: None,
labels: None,
lfs: None,
lfs_endpoint: None,
milestones: None,
mirror: Some(true),
// Periodic refresh instead of on-access sync — keeps CI isolated
// from external DNS failures at clone time.
mirror_interval: Some(MIRROR_INTERVAL.to_owned()),
private: Some(false),
pull_requests: None,
releases: None,
repo_name: repo.to_owned(),
repo_owner: Some(owner.to_owned()),
service: Some(MigrateRepoOptionsService::Git),
uid: None,
wiki: None,
};
match client.repo_migrate(opts).await {
Ok(_) => {
tracing::info!(%owner, %repo, %upstream, interval = MIRROR_INTERVAL, "forge: created pull-mirror");
Ok(())
}
// 409 = a race created it between our existence check and here (the
// check is the real idempotency guard). NOT 422: for the migrate
// endpoint 422 is a validation error (bad clone_addr / service), so
// it must surface via the error arm, not be swallowed as "already
// exists" — hence `is_conflict`, not `is_already_exists`.
Err(e) if is_conflict(&e) => {
tracing::debug!(%owner, %repo, "forge: pull-mirror already exists (race)");
Ok(())
}
Err(e) => Err(e).with_context(|| format!("migrate pull-mirror {owner}/{repo}")),
}
}
/// Provision the [`OPERATORS_TEAM`] inside `org` as an **empty** team.
/// Branch protection on that org's repos references it as the
/// merge/approval whitelist; the operator adds herself as a member via the
/// forge UI / hivectl. `includes_all_repositories` so the gate applies to
/// every repo in the org; `write` is enough to approve + merge. hive-c0re
/// never manages membership. Idempotent (409 = already exists).
///
/// Must run for BOTH [`AGENTS_ORG`] and [`CONFIG_ORG`]: Gitea teams are
/// org-scoped, so a config-repo branch-protection rule referencing
/// `operators` needs the team to exist in `agent-configs` too. Missing it
/// there 422'd every `apply_config_repo_branch_protection`, leaving config
/// repos unprotected — operator-merged config PRs then bypassed the deploy
/// pipeline and silently didn't apply.
///
/// Uses `is_conflict` (409 only) — NOT `is_already_exists` (which also
/// folds 422 into "already exists"). A 422 from `org_create_team` is a
/// real validation error (bad request shape, missing units, etc.) that
/// must surface so it can be fixed; the previous 422-swallowing hid the
/// true cause and left the team silently uncreated every boot.
/// Repo-unit access flags for the `operators` team.
/// Explicit list so Forgejo doesn't reject a null/absent `units` field;
/// a `write`-permission team needs at least `repo.code` + `repo.pulls`
/// to review and merge PRs.
const OPERATORS_TEAM_UNITS: &[&str] = &[
"repo.code",
"repo.issues",
"repo.pulls",
"repo.releases",
"repo.wiki",
"repo.projects",
"repo.packages",
];
pub(super) async fn ensure_operators_team(org: &str, token: &str) -> Result<()> {
let units_vec: Vec<String> = OPERATORS_TEAM_UNITS.iter().map(|&s| s.to_owned()).collect();
let team = CreateTeamOption {
can_create_org_repo: Some(false),
description: Some("hyperhive operators — merge gate for agent repos".to_owned()),
includes_all_repositories: Some(true),
name: OPERATORS_TEAM.to_owned(),
permission: Some(CreateTeamOptionPermission::Write),
units: Some(units_vec.clone()),
units_map: None,
};
let client = api(token)?;
match client.org_create_team(org, team).await {
Ok(_) => {
tracing::info!(%org, "forge: created {OPERATORS_TEAM} team");
Ok(())
}
// Team already exists — reconcile settings to desired state so a
// team created with an older/wrong shape self-heals on next boot.
// Forgejo signals the duplicate as a 409 conflict OR (this build) a
// 422 `validation failed: team already exists`; both mean the same
// thing, so fold both in via `is_already_exists_lenient` — a 409-only
// check missed the 422 and warned every boot. List teams to find the
// id (required by org_edit_team), then unconditionally PATCH to the
// desired settings. Members are a separate endpoint; untouched here.
Err(e) if is_already_exists_lenient(&e) => {
let (_headers, teams) = client
.org_list_teams(org)
.await
.with_context(|| format!("list teams for {org}"))?;
let team_id = teams
.into_iter()
.find(|t| t.name.as_deref() == Some(OPERATORS_TEAM))
.and_then(|t| t.id)
.with_context(|| {
format!("{OPERATORS_TEAM} team not found in {org} after 409 conflict")
})?;
let edit = EditTeamOption {
can_create_org_repo: Some(false),
description: Some("hyperhive operators — merge gate for agent repos".to_owned()),
includes_all_repositories: Some(true),
name: OPERATORS_TEAM.to_owned(),
permission: Some(EditTeamOptionPermission::Write),
units: Some(units_vec),
units_map: None,
};
client
.org_edit_team(team_id, edit)
.await
.with_context(|| format!("reconcile {org}/{OPERATORS_TEAM} team settings"))?;
tracing::debug!(%org, "forge: reconciled {OPERATORS_TEAM} team settings");
Ok(())
}
// Any OTHER error (a 422 that is NOT already-exists, or a transport
// / auth failure): surface it — don't mask a real failure. A
// non-already-exists 422 means the request body is invalid (e.g.
// Forgejo rejected the units list); it repeats every boot until fixed.
Err(e) => Err(e).with_context(|| format!("create team {org}/{OPERATORS_TEAM}")),
}
}
/// Add `user` as a collaborator on `owner/repo` at `permission`.
/// Idempotent: adding an existing collaborator just updates its
/// permission (Forgejo answers 204 either way; a 201 from older
@ -1224,27 +830,3 @@ pub async fn create_agent_repo(agent: &str, repo: &str, core_token: &str) -> Res
tracing::info!(%agent, %repo, "forge: created agent repo in {AGENTS_ORG} with operator merge gate");
Ok(format!("{AGENTS_ORG}/{repo}"))
}
#[cfg(test)]
mod tests {
use super::message_says_already_exists;
#[test]
fn already_exists_message_is_recognised() {
// The exact 422 body this Forgejo build returns for a duplicate team.
assert!(message_says_already_exists(
"validation failed: team already exists [org_id: 10, name: operators]"
));
// Case-insensitive.
assert!(message_says_already_exists("Repository Already Exists"));
}
#[test]
fn real_validation_error_is_not_treated_as_already_exists() {
// A genuine bad-request 422 must still surface, not be folded in.
assert!(!message_says_already_exists(
"validation failed: units must not be empty"
));
assert!(!message_says_already_exists("not found"));
}
}

View file

@ -12,7 +12,7 @@ use forgejo_api::structs::{EditUserOption, UpdateUserAvatarOption};
use forgejo_api::{ApiErrorKind, ForgejoError};
use reqwest::StatusCode;
use super::{CONFIG_ORG, api, forge_admin};
use super::{api, forge_admin};
const TOKEN_NAME_PREFIX: &str = "hyperhive";
// Where the host-side `core` admin token lives. Used by hive-c0re itself
@ -20,21 +20,18 @@ const TOKEN_NAME_PREFIX: &str = "hyperhive";
// in `crate::paths`; aliased here under the long-standing name.
use crate::paths::FORGE_CORE_TOKEN as CORE_TOKEN_PATH;
// Forge provisioning markers (`forge/core-avatar-set`,
// `forge/agent-configs-avatar-set`, `forge/email-aligned-<name>`) live
// in `crate::paths` — one-shot guards: the upload/align runs once, the
// marker is written, subsequent startups skip. Delete one to force its
// step to re-run.
// Avatar PNG paths are resolved by `core_avatar_png_path` /
// `config_org_avatar_png_path` below — only-used-here, so they live
// in this module rather than the shared `hive_sh4re::assets` helpers
// (which now hold only `prompt_template`, the one asset path every
// crate needs — the agent icon, the last other one, moved off the
// `forge/email-aligned-<name>`) live in `crate::paths` — one-shot guards:
// the upload/align runs once, the marker is written, subsequent startups
// skip. Delete one to force its step to re-run.
// The avatar PNG path is resolved by `core_avatar_png_path` below —
// only-used-here, so it lives in this module rather than the shared
// `hive_sh4re::assets` helpers (which now hold only `prompt_template`, the
// one asset path every crate needs — the agent icon, the last other one, moved off the
// shared-asset model entirely; see `hive-agent::web_ui::screen`).
/// `$HIVE_ASSETS_DIR/branding/hyperhive.png` — the core mark,
/// rasterised. Not independently configurable (unlike the org avatar
/// below) — override the whole `services.hyperhive.c0re.assets`
/// package to change it.
/// rasterised. Not independently configurable — override the whole
/// `services.hyperhive.c0re.assets` package to change it.
fn core_avatar_png_path() -> std::path::PathBuf {
let dir = std::env::var("HIVE_ASSETS_DIR").expect(
"HIVE_ASSETS_DIR is unset — the hyperhive NixOS module sets it from \
@ -44,21 +41,6 @@ fn core_avatar_png_path() -> std::path::PathBuf {
std::path::PathBuf::from(dir).join("branding/hyperhive.png")
}
/// Path to the `agent-configs` org's avatar PNG. Independently
/// configurable via `services.hyperhive.c0re.orgAvatarPng` — the nix
/// module resolves `HIVE_ORG_AVATAR_PNG` to that option's value, or
/// the bundled `agent-configs.png` from `assets` when unset, so an
/// operator can override just this image without replacing the whole
/// `assets` package.
fn config_org_avatar_png_path() -> std::path::PathBuf {
std::path::PathBuf::from(std::env::var("HIVE_ORG_AVATAR_PNG").expect(
"HIVE_ORG_AVATAR_PNG is unset — the hyperhive NixOS module sets it \
unconditionally on the hive-c0re unit (from \
services.hyperhive.c0re.orgAvatarPng or the bundled default), so \
this process was started outside that unit",
))
}
/// Bootstrap `core` token scopes — adds `read:admin,write:admin` on
/// top of the agent scopes (swarm-controller's `AGENT_TOKEN_SCOPES`)
/// so the host daemon can drive `/api/v1/admin/*`. Site-admin
@ -369,39 +351,6 @@ pub(super) async fn ensure_core_avatar(token: &str) -> Result<()> {
Ok(())
}
/// Set the `agent-configs` org's Forgejo avatar to the
/// configs-stack glyph once. Sibling to `ensure_core_avatar`:
/// one-shot, marker-guarded, best-effort. Uses `org_update_avatar`
/// (`POST /api/v1/orgs/{org}/avatar`, base64-PNG payload).
pub(super) async fn ensure_config_org_avatar(token: &str) -> Result<()> {
let marker = crate::paths::forge_config_org_avatar_marker();
if marker.exists() {
return Ok(());
}
let png_path = config_org_avatar_png_path();
let png_bytes = tokio::fs::read(&png_path)
.await
.with_context(|| format!("read {CONFIG_ORG} avatar PNG from {}", png_path.display()))?;
api(token)?
.org_update_avatar(
CONFIG_ORG,
UpdateUserAvatarOption {
image: Some(base64::engine::general_purpose::STANDARD.encode(&png_bytes)),
},
)
.await
.with_context(|| format!("set {CONFIG_ORG} avatar"))?;
if let Some(parent) = marker.parent() {
std::fs::create_dir_all(parent).ok();
}
std::fs::write(marker, "").ok();
tracing::info!(
org = CONFIG_ORG,
"forge: set org avatar to configs-stack logo"
);
Ok(())
}
/// Outcome of probing whether the persisted core token still works
/// against the *current* forge. Existence on disk is not validity: a
/// token minted before a forge rebuild / re-provision is unknown to the

View file

@ -184,13 +184,8 @@ async fn run_matrix_sweep() -> Result<()> {
/// `tokio::spawn` block it replaced used: no-op (not an error) when the
/// HMAC secret, core token, or hive domain aren't available yet.
///
/// The node now does one of each: it still registers the config-PR hook,
/// and it *removes* the knowledge one. A knowledge push is delivered to
/// the swarm controller, which addresses an event to each hive over the
/// queue — so a hive holding its own registration is holding a shared
/// resource only one party can own. The removal runs every boot rather
/// than behind a marker because it is already idempotent: it is a no-op
/// the moment the hook is gone.
/// Registers the config-PR hook only. The knowledge hook is the swarm
/// controller's, which addresses an event to each hive over the queue.
async fn run_webhook_register() -> Result<()> {
let Ok(webhook_secret) = crate::webhook_secret::load_or_generate() else {
tracing::debug!("webhook secret unavailable; skipping hook registration");
@ -206,9 +201,6 @@ async fn run_webhook_register() -> Result<()> {
tracing::debug!("HYPERHIVE_HIVE_DOMAIN unset; skipping webhook registration");
return Ok(());
};
if let Err(e) = crate::workers::knowledge::remove_webhook(&token, &domain).await {
tracing::warn!(error = ?e, "knowledge: remove_webhook failed");
}
if let Err(e) = crate::forge::ensure_config_pr_webhook(&token, &domain, &webhook_secret).await {
tracing::warn!(error = ?e, "forge: ensure_config_pr_webhook failed");
}

View file

@ -80,12 +80,6 @@ pub fn forge_core_avatar_marker() -> PathBuf {
forge_dir().join("core-avatar-set")
}
/// `forge/agent-configs-avatar-set` — marker: agent-configs org avatar set.
#[must_use]
pub fn forge_config_org_avatar_marker() -> PathBuf {
forge_dir().join("agent-configs-avatar-set")
}
/// `forge/email-aligned-<name>` — marker: `<name>`'s forge email aligned.
#[must_use]
pub fn forge_email_aligned_marker(name: &str) -> PathBuf {
@ -319,14 +313,10 @@ pub fn agent_runtime_dir(name: &str) -> PathBuf {
/// within the same filesystem is atomic.
pub fn relocate_legacy_state() {
let root = state_root();
let moves: [(&str, PathBuf); 8] = [
let moves: [(&str, PathBuf); 7] = [
("broker.sqlite", db_dir().join("broker.sqlite")),
("build_logs.sqlite", db_dir().join("build_logs.sqlite")),
("forge-core-avatar-set", forge_core_avatar_marker()),
(
"forge-agent-configs-avatar-set",
forge_config_org_avatar_marker(),
),
("matrix-sender-token", matrix_sender_token()),
("matrix-space-room-id", matrix_space_room_id()),
("matrix-creds", matrix_creds_dir()),

View file

@ -15,7 +15,9 @@
//! `/webhook/knowledge`. A webhook has exactly one target URL, so with
//! more than one hive that was last-writer-wins rather than idempotent —
//! every hive but the most recent silently stopped receiving deliveries.
//! [`remove_webhook`] is the migration off it.
//!
//! The repo itself (public, README-seeded) is the swarm controller's too;
//! this module only clones and pulls it.
use anyhow::{Context, Result};
@ -35,38 +37,12 @@ pub use crate::paths::KNOWLEDGE_DIR as LOCAL_DIR;
/// read-only from [`LOCAL_DIR`] into every agent container.
pub const CONTAINER_MOUNT: &str = "/knowledge";
/// Default README pushed to a freshly created `internal/knowledge` repo.
/// Short explanation + empty table-of-contents with an HTML comment instructing contributors
/// to add entries when they create new files.
const README_CONTENT: &str = "\
# knowledge
Hive-wide reference documents: conventions, runbooks, and anything that \
every agent should know.
## How to contribute
1. Fork this repo into your own namespace on the forge.
2. Create a branch, add or update a document.
3. Open a pull request — the operator reviews and merges.
4. Every agent container updates automatically on merge.
Do **not** push directly to `main` — agents have read-only access.
## Contents
<!-- Add an entry here each time you create a new document:
- [Title](path/to/file.md) - one-line description
-->
";
/// Clone `internal/knowledge` to [`LOCAL_DIR`] if it is not already a git
/// repository. `core_token` authenticates the HTTPS clone so private repos
/// work. Idempotent — skips if `LOCAL_DIR/.git` exists.
///
/// When the upstream repo is empty (freshly created), seeds it with a
/// README.md before returning so the local clone is always non-empty and
/// agents see a useful starting document.
/// The swarm controller seeds a README into a freshly created repo. A clone
/// taken before that lands is empty; the next [`pull`] fills it.
///
/// Called once at hive-c0re startup after `forge::ensure_all`.
pub async fn ensure_local_clone(core_token: &str) -> Result<()> {
@ -87,129 +63,6 @@ pub async fn ensure_local_clone(core_token: &str) -> Result<()> {
anyhow::bail!("git clone {ORG}/{REPO} failed: {stderr}");
}
tracing::info!("knowledge: cloned {ORG}/{REPO} to {LOCAL_DIR}");
// If the repo is brand new (no commits), seed it with a README.
let head_out = tokio::process::Command::new("git")
.args(["-C", LOCAL_DIR, "rev-parse", "HEAD"])
.output()
.await
.context("git rev-parse HEAD")?;
if !head_out.status.success() {
seed_readme(core_token).await?;
}
Ok(())
}
/// Write the initial README.md, commit, and push to `internal/knowledge`.
/// Called only when the upstream repo is empty.
async fn seed_readme(core_token: &str) -> Result<()> {
let readme = std::path::Path::new(LOCAL_DIR).join("README.md");
std::fs::write(&readme, README_CONTENT).context("write README.md")?;
// Set a minimal git identity for the seed commit.
for (k, v) in [("user.email", "core@hive"), ("user.name", "hive-c0re")] {
let out = tokio::process::Command::new("git")
.args(["-C", LOCAL_DIR, "config", k, v])
.output()
.await
.with_context(|| format!("git config {k}"))?;
if !out.status.success() {
let stderr = String::from_utf8_lossy(&out.stderr).trim().to_owned();
anyhow::bail!("git config {k} failed: {stderr}");
}
}
for args in [
vec!["add", "README.md"],
vec!["commit", "-m", "init: seed README"],
] {
let out = tokio::process::Command::new("git")
.args(["-C", LOCAL_DIR].iter().chain(args.iter()))
.output()
.await
.with_context(|| format!("git {args:?}"))?;
if !out.status.success() {
let stderr = String::from_utf8_lossy(&out.stderr).trim().to_owned();
anyhow::bail!("git {args:?} failed: {stderr}");
}
}
let url = forge_git_url(&format!("{ORG}/{REPO}"));
let out = crate::lifecycle::git_command_authed(&core_auth_header(core_token))
.args(["-C", LOCAL_DIR, "push", &url, "HEAD:main"])
.output()
.await
.context("git push knowledge README")?;
if out.status.success() {
tracing::info!("knowledge: seeded README.md and pushed to {ORG}/{REPO}");
Ok(())
} else {
let stderr = String::from_utf8_lossy(&out.stderr).trim().to_owned();
anyhow::bail!("git push {ORG}/{REPO} failed: {stderr}")
}
}
/// Delete this hive's own `internal/knowledge` push webhook if it is
/// still registered, so the swarm controller is the only party holding
/// one.
///
/// # Why this is a migration and not just a deletion
///
/// Not registering any more fixes nothing on a hive that has already
/// run: the hook it created persists on the forge, so the contention
/// this removes would survive on exactly the deployments that have it
/// while fresh installs looked fixed. The hive that created a hook is
/// the one that removes it.
///
/// # It removes only its OWN hook, never a neighbour's
///
/// The match is the full URL, not the `/webhook/knowledge` suffix. A
/// hook with that suffix and a different base belongs to *another hive* —
/// one that may not have been upgraded yet — and deleting it would break
/// its knowledge sync until it was. Reaping a neighbour's registration is
/// the very behaviour this issue is about; doing it in the name of fixing
/// it would just invert the direction.
///
/// (The predecessor did reap by suffix, to clear loopback hooks left by
/// an older single-hive layout. That was safe when a hive was alone on
/// its forge and is not safe now.)
///
/// A listing failure is an error rather than a silent skip: there is no
/// create attempt left to fall through to, so swallowing it would leave
/// the hook in place with nothing said. The caller logs and continues —
/// boot does not depend on this.
pub async fn remove_webhook(core_token: &str, hive_domain: &str) -> Result<()> {
// The typed client carries no per-request timeout, so each call is
// wrapped in one: this runs as a detached startup task, and a forge
// that accepts connections but never answers would otherwise hang it
// forever.
const HTTP_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(10);
let own_url = format!("https://{hive_domain}/webhook/knowledge");
let client = crate::forge::api(core_token)?;
let hooks = tokio::time::timeout(HTTP_TIMEOUT, client.repo_list_hooks(ORG, REPO).all())
.await
.map_err(anyhow::Error::from)
.and_then(|r| r.map_err(anyhow::Error::from))
.with_context(|| format!("list webhooks for {ORG}/{REPO}"))?;
for h in &hooks {
let hook_url = h
.config
.as_ref()
.and_then(|c| c.get("url"))
.map_or("", String::as_str);
if hook_url == own_url
&& let Some(id) = h.id
{
tokio::time::timeout(HTTP_TIMEOUT, client.repo_delete_hook(ORG, REPO, id).send())
.await
.map_err(anyhow::Error::from)
.and_then(|r| r.map_err(anyhow::Error::from))
.with_context(|| format!("delete webhook {id} for {ORG}/{REPO}"))?;
tracing::info!(
%own_url,
"knowledge: removed this hive's push webhook — the swarm controller owns it now"
);
}
}
Ok(())
}
@ -256,8 +109,7 @@ pub async fn pull(coord: &Coordinator) -> Result<()> {
.await;
// Discard any local drift before pulling — both tracked (`reset --hard`)
// and untracked (`clean -fd`). Nothing in this codebase writes to
// `LOCAL_DIR` after the initial clone (`seed_readme` runs once, at
// creation, before any pull) — this working tree exists to mirror
// `LOCAL_DIR` after the initial clone — this working tree exists to mirror
// `origin/main`, not to be edited in place. A tracked file left dirty
// by any other means (a stray manual edit on the host, an interrupted
// prior operation, or — the actual root cause here — two unsynchronized
@ -280,7 +132,7 @@ pub async fn pull(coord: &Coordinator) -> Result<()> {
.await;
let before = head_sha().await;
// The repo is public (`ensure_knowledge_repo` makes it so), so an
// The repo is public (the swarm controller makes it so), so an
// unauthenticated pull is enough — but authenticate when a token is around,
// which keeps this working if the repo is ever made private again.
let auth = crate::forge::core_token().map(|t| core_auth_header(&t));