mara ruled on #4849 (c88934): "remove create_repo tool". The tool ran in
hive-c0re with the hive's core token, so it only ever worked for agents
on the hive that runs the forge.
Removed:
- the create_repo MCP tool and CreateRepoArgs (hive-agent-mcp)
- wire variants Request::CreateRepo and Response::RepoCreated
(hive-core-agent-sock)
- hive-c0re's handle_create_repo, its valid_repo_name check and the
dispatch arm
- forge::create_agent_repo and apply_operator_branch_protection, which
had no other caller, plus AGENTS_ORG and OPERATORS_TEAM, whose only
users they were
- the tool's docs (docs/tools/forge.md repo management, docs/turn-loop/
mcp.md, the conventions tool-group table) and the doc comments that
named it (hive-sock-client's response timeout, ensure_repo_creation_
disabled, the security doc's merge-gate bullet)
ToolGroup::Forge is kept with no tools, the same way b88a5b24 kept
Lifecycle, so existing meta/capabilities.json grants still parse.
Forge state is untouched: existing agents/* repos keep their collaborators
and operators-team branch protection. The swarm-controller's own
create_repo (config-org repos) is a different path and is unchanged.
Closes#4849
Before b5d07d4d, hive-c0re minted a new `hyperhive-<unix-seconds>` token
for an agent on every spawn and rebuild and never revoked one, so every
live agent's forge user carries a pile of write-scoped tokens nothing
holds. Nothing in the tree lists or deletes them.
On each start, swarm-controller now walks the store's hive-agent-*
roster (the one the swarm-agent mint pass walks), and for every agent
whose swarm-agent token that pass would keep, deletes each token named
exactly `hyperhive-<digits>`. It logs the count per agent and a total.
- `core` is refused by name in both the roster filter and the per-user
delete: hive-c0re still names core's live admin token
`hyperhive-<unix-seconds>`.
- An agent whose swarm-agent token is not current is skipped, because
consumers fall back to `<state>/forge-token`, the last hyperhive-*
token, until the swarm token is fetched.
- A failed list or delete is logged and skipped; the sweep does not
retry and never blocks startup. A second start deletes nothing.
Tokens on forge users of agents no longer on the store roster (already
destroyed) are not reached.
Closes#4644
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.
swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.
hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.
nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.
Refs #3782
POST /api/forge/users/{name}/admin reads the account and, when it is not
already a site admin, sets `admin` with admin_edit_user. It never
creates one: a human's account is made by their first authelia login,
so a missing one answers 404, saying the user has not logged in via SSO
yet. An existing admin is a success with nothing sent.
The edit carries `admin` alone. repo_creation_lockdown's login_name +
source_id = 0 would turn an SSO-made account into a local one: in
Forgejo 16 a source_id sets the login type.
An agent's name is refused, and so is any name when the roster can't be
read: a site admin ignores max_repo_creation, the lockdown that keeps an
agent's token from creating a repo and self-merging in it.
Refs #3782
The stub-client helper in forge.rs built Client { api } after #4703
added a url field, breaking compilation of swarm-controller's test
target on main. Clone the URL the helper already builds for
Forgejo::new so it can also populate Client.url.
Closes#4709
A MintAgentForgeToken node mints a fixed-name swarm-agent token with the
admin API, keeps it when the stored value's last eight and the normalised
scopes match the forge's list, and otherwise deletes and re-creates it.
The token is stored at swarm/agents/<agent>/forge-token. Agent creation
inserts the node, and a pass at start and every five minutes inserts it
for every agent holding a store identity whose token is missing or stale.
Refs #3782
disable_repo_creation's ownership guard now matches either the aligned
{agent}@hyperhive.local email or the legacy {agent}@hive.local one
hive-c0re::forge::users::ensure_user_exists used before its own
ensure_user_email alignment pass existed. That pass only runs on the
forge-host hive, and only once it has a core token and has ticked, so a
real pre-existing agent can still carry the old email when this node reads
it. Without this, the guard would bail "not this agent's" on a genuine
agent during that rollout window (argus, round 2).
Verified separately (not a code change): swarm-controller's forge account
is a site admin (created with --admin), and Forgejo's
convert.toUser/ToUser only hides an account's email when the caller isn't
the admin and isn't the account itself (services/convert/user.go), so
user_get already returns the real email regardless of hide_email on the
target account. No endpoint change needed for that half of the review.
disable_repo_creation now reads the account once before PATCHing
max_repo_creation/source_id, and fails the node if the email isn't the
{agent}@hyperhive.local marker create_agent_user itself sets. The 409/422
create-fold (#4681) only proves some account with that name exists, not
that this node created it, so a pre-existing non-agent account sharing an
agent's chosen name could otherwise get locked onto local auth with repo
creation disabled. Leaves the fold untouched (Option B, per atlas/argus on
#4693); the read moves into disable_repo_creation instead.
Also fixes the nix-sandboxed cargo-test check: forgejo_api::Forgejo::new
builds a reqwest client that eagerly resolves TLS roots via
rustls-native-certs even for the tests' plain-http loopback stub server,
which panics with "No CA certificates were loaded from the system" in the
CA-less build sandbox. Gives that check's nativeBuildInputs pkgs.cacert and
sets SSL_CERT_FILE, same pattern this repo's runtime deployment already
uses for the same reqwest/rustls resolution.
`ensure_agent_user` (the `CreateForgeUser` node) created the agent's
forge account but never set `max_repo_creation = 0`, relying on
hive-c0re's per-hive `ensure_repo_creation_disabled` pass to lock it
down later. That pass is being removed (#3507, #4669), and it is the
only guard against an agent token creating, owning and self-merging in
its own repo.
The node now PATCHes `max_repo_creation = 0` via `admin_edit_user`
after the create, on both the created and the already-exists path, and
a refused PATCH fails the node with Forgejo's message. The body mirrors
hive-c0re's `sparse_edit_user_option` (`login_name` + `source_id = 0`,
everything else unset).
The PATCH is unconditional: Forgejo's API never returns
`max_repo_creation` (it is in `EditUserOption` only, not `User`), so
there is no current value to verify against first.
Closes#4689
list_repos_with_open_issues only read one repo_search page (forgejo's
30-row default), so a repo past position 30 in the default alpha sort
silently dropped out of the issue-report/repo-dropdown data source, with
no truncation signal. The comment claiming RepoSearchQuery has no page
field was wrong -- Request::page()/page_size() are generic builder
methods independent of the query struct.
Adds page_search_results, a small paging loop over a fetch closure for
search-shaped (data: Option<Vec<T>>, no header) responses that can't use
the existing .all() helper (that needs a (Headers, Vec<T>) response
shape). Pages until a short page or a 40-page bound, erroring on the
bound rather than truncating again.
Closes#4675
get_issue_report mapped every error from Client::issue_report to 503
(StatusUnavailable) unconditionally. issue_report calls
issue_list_issues(org, repo, ...) first, and a Forgejo 404 there — a
typo'd or deleted repo — was flattened into the same 503 a genuine
forge outage produces, which tells a client to retry a request that
will never succeed.
Downcast the anyhow error back to forgejo_api::ForgejoError (same
pattern main.rs's wanted_error_status uses for
swarm_queue_client::Error) and check it structurally against the
ApiErrorKind::NotFound / UnexpectedStatusCode(404) shapes forgejo's
generated client produces for a 404, rather than string-matching the
rendered message. Only that case answers 404, naming the org/repo;
every other forge failure still answers 503. Updates the route's
OpenAPI response list to document the 404.
Closes#4701
Forgejo answers 422 for six different causes on admin user create
(ErrUserAlreadyExist, ErrEmailAlreadyUsed, ErrNameReserved,
ErrNameCharsNotAllowed, ErrEmailInvalid, ErrNamePatternNotAllowed) and
several on repo create, but is_already_exists() treated every one of
them as a conflict. A reserved or otherwise-refused name silently
folded to Done, so CreateForgeUser reported success with no user
created, and the graph's real failure only surfaced one node later as
a misleading AddRepoMember error.
ensure_agent_user and ensure_org_repo now trust a 409 unconditionally
(folds_into_success) but confirm a 422 with a follow-up user_get /
repo_get before folding it to success; an unconfirmed 422 fails with
forgejo's own message at error level. Webhook registration still uses
the old is_already_exists — it has no comparable follow-up read, so it
is out of scope here.
Closes#4678
Per mara's review call on this PR: "the view should be filled by a single
backend call." AgentsPage.tsx was doing three fetches (/api/agents,
/api/config-prs, /api/agents/status) and joining them client-side by name.
Moves the config-PR join server-side instead: AgentStatusRow gains a
config_pr field, populated by get_agents_status's handler from
AppState::config_prs after agent_status::AgentStatusReader::view() returns
- not inside that module, which has no forge client and stays that way (see
the field's doc comment for why the handler is the right layer for this
merge, not the reader).
AgentsPage.tsx now does exactly one fetch and no client-side joining at all
- the wire row is the table row. Dropped the separate AgentStatusRow TS
interface (folded into AgentRow, which now mirrors the backend type
field-for-field) and the /api/agents + /api/config-prs fetches entirely;
neither is needed once /api/agents/status already returns every roster
agent with its config PR attached.
ConfigPrStatus gained Deserialize (previously Serialize-only) since
AgentStatusRow derives both and a struct's derive requires every field to
support it.
`POST /admin/hooks` reads `is_system_webhook` out of the config map and
defaults it to false, which creates a forgejo *default* webhook — a
template copied into repos created later — instead of a live
instance-wide one. `GET /admin/hooks` returns only hooks with the flag
set, so `list_hook_urls` could never see what the create had just made:
every controller start listed zero hooks and created another default
webhook (12 in 26h), while the instance-wide push observation the scope
exists for never fired at all.
Send the key on the `Instance` arm only, via a per-scope
`extra_create_config()` so repo and org scopes stay unchanged.
Closes#3807
The swarm writes agent-configs/<agent> when it creates an agent, before
any hive is told to deploy it. setup_proposed authored a second copy of
those same bytes locally, so an agent's initial config had two sources
of truth, each unaware of the other and free to disagree. It now clones
that repo and falls back to the template only when there is nothing
there to take.
Preferred-source rather than a new-path-only variant because
provision_container is the Provision node for the swarm deploy and the
approval flow both, and cannot tell them apart. The approval flow
creates agent-configs/<agent> only after the first spawn
(forge_after_first_spawn), so it finds nothing and lands on the
template: the fallback becomes unreachable when hive-level create is
removed, rather than becoming something someone has to find and delete.
clone, not the neighbouring init+fetch. A failed fetch leaves an empty
.git behind, and that .git is exactly the byte setup_proposed reads to
decide whether seeding is still needed, so the fallback would have seen
a seeded repo. git removes a directory it created when a clone fails.
--branch main also makes an empty repo fail cleanly instead of cloning
to an unborn HEAD that would look seeded.
`ensure_hook` reaches its create call from two different places — the
list step failed, or it succeeded and matched nothing — and both were
logged at `debug!`. At the level the journal keeps, that made a repeated
registration indistinguishable from a first, correct one, and it is a
repeated registration that is being observed: the instance-scoped hook
logs the create arm on every process start while the repo-scoped one
correctly goes quiet.
The fold that keeps this harmless (`is_already_exists` swallowing a
duplicate create) is an assumption about the forge rather than a
guarantee, so the failed-list arm becomes a `warn!`, and the
matched-nothing arm logs the number of hooks it did see: `listed=0` is a
permission or scope problem, a non-zero count with no match means the
recorded url is not the one being compared.
No behaviour change — this makes the existing behaviour legible.
WIP — compiles per an earlier build, but the verifying build/test run was
cut short by a graceful stop. Re-run gate.sh before pushing.
The controller created every agent config repo, its collaborator entry,
its branch protection and its seeded agent.nix/flake.nix in the agents
org. Config repos live in agent-configs, which is where hive-c0re
reconciles, merges and mirrors them — so a repo created in agents is
invisible to all of those, and nothing errors, because both orgs exist
and both accept a repo.
Root cause was a doc comment asserting something false: AGENTS_ORG
claimed to be the same org hive-c0re uses for its config-repo path. It is
not — hive-c0re's agents org is the namespace repos an AGENT ASKS FOR
land in, and its config repos use agent-configs. The whole flow inherited
the wrong premise from that sentence.
The merge gate survives the move: hive-c0re provisions the operators team
in both orgs, with a comment recording that missing the agent-configs
copy once left every config repo unprotected.
mara, PR #3438 review: 'remove the identity function. agent names
are unique and all repos go into agent-configs namespace anyway'. Right --
repo==agent isn't a convention worth a name once every call site can just
say so; call create_repo/add_repo_member/seed_agent_config with &agent
directly.
Every repo in this graph is `agents/<agent>` -- the node payloads were
carrying the same string under two names, and `create_agent` opened with
a `let repo = agent.clone()` that said so out loud.
All four node kinds now carry `agent` alone, and `forge::agent_repo` is
the single home for the naming convention. The identity it returns is the
point: a caller holding an agent name never writes a repo name itself, so
changing the convention later is one edit rather than a search.
`forge::Client`'s methods keep taking a repo, because they are a general
forge client and `add_repo_member(repo, user)` is a real signature -- the
derivation belongs at the call site that knows the two are the same here,
not baked into an API that has no reason to assume it.
`data()`'s four arms are now identical and merged into one or-pattern.
Left as an explicit list rather than a catch-all so a fifth variant fails
to compile here instead of silently rendering as an agent name.
`POST /api/agents` now requires `hive` alongside `name`. It is parsed as
an `Ident` like `name` already was, and then checked against the roster
loaded from `SWARM_CONTROLLER_HIVES` -- a hive that is not in this swarm
is a 400 naming the ones that are, rather than a typo accepted and
forgotten. The roster check is what makes the field worth having; without
it nothing notices until a deploy message is addressed to a hive that
does not exist.
`hive` is an address, not an attribute of the agent: it is where a deploy
message goes over the queue, so nothing writes it into the agent's config
repo. A config naming its own hive would be a second statement of where
the agent lives, free to drift from the queue that actually delivers to
it.
It rides on the `InitAgentConfigRepo` node payload because the graph is
the only thing carrying the operator's choice forward from the API
boundary; seeding does not consume it. The node that routes on it is the
deploy node in #3124.
The refusal is asserted by effect -- the test checks that *nothing was
queued*, not just the status code, since a version that queued the graph
and then complained would satisfy a status-only assertion while still
creating the agent.
This is a breaking change for every existing caller: the swarm-UI create
page posts `{name}` only and needs its hive dropdown to land alongside.
The endpoint landed inert: nothing pointed at it, so the only way to see
it work was to mint an HMAC by hand. Register the two swarm-wide hooks
at startup so a real forge event produces a journal line.
Registered ALONGSIDE the per-hive hooks, not instead of them. Every hive
keeps receiving and acting on its own deliveries; the controller gets a
copy and logs it. Moving the registration is a later step and has to be:
fan-out swarm->hive does not exist yet, so a hook moved now would point
at a receiver that forwards nowhere, silently on both sides.
Deliberately no stale-hook deletion arm, unlike the two per-hive
registrars this otherwise mirrors: theirs delete hooks matching their own
path with a foreign base, and the hives' hooks are not stale.
The route prefix is what keeps this safe. Both hive-side registrars
delete any hook ending in /webhook/knowledge or /webhook/config-pr with a
different base, so a swarm hook under those paths would be deleted by
every hive on every boot. Serving them under /webhook/forge/ avoids it,
and a test pins it -- there is nothing else that can.
SWARM_CONTROLLER_PUBLIC_URL is set only where the swarm vhost is served,
because a hook whose target_url nothing answers is worse than no hook.