Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/swarm-controller
Author SHA1 Message Date
atlas
b90be9e65e swarm UI: delete linked accounts — R2 fixes
Addresses argus review comment 90297 on PR #4899:

- swarm-controller/README.md: list the three DELETE routes (including
  matrix's ?revoke=true) beside the PUT/GET ones already documented.
- LinkedAccounts.tsx: a delete answering 404 means the account is
  already gone, so treat it as the delete's end state — re-fetch and
  close the dialog instead of showing an error.
- matrix_account.rs: matrix_logout treats a 401 M_UNKNOWN_TOKEN as the
  token already being revoked and proceeds with the delete; every
  other logout failure still keeps the account. Adds unit tests and
  updates docs/swarm/ui.md to match.
2026-10-03 00:56:39 +02:00
atlas
bbf931207f swarm UI: delete linked accounts
Each row of an agent's linked accounts, except its own `main` matrix
account, gets a delete action. swarm-controller serves DELETE beside each
PUT (matrix-accounts/{account}, forge-accounts/{label}, github-account),
answers 404 for an account the store does not hold, refuses `main`, and
removes every version through `delete_all_versions`.

The matrix confirmation has a revoke checkbox, off by default: the
controller logs the stored token out at its homeserver first, and keeps
the account when that fails or no homeserver is stored.

The controller's policy gains `delete` on each agent's
`metadata/.../matrix/+`, `forge/+` and `github-token`, pinned in
bao-grants.nix.

Refs #4855
2026-10-03 00:56:39 +02:00
atlas
0cee0382e9 config repos: drop the hive's branch-protection edit and core from the merge gate
Folds in #4853: hive-c0re no longer writes branch protection on
`agent-configs` repos (apply_config_repo_branch_protection,
config_repo_protection_edit, record_branch_protection_result and the
now-unused main_branch_protection_option go). swarm-controller's
CreateRepo rule and its forge-objects convergence own `main`'s gate.

Nothing merges into config `main` as `core` any more, so the converged
rule's merge user list is empty. `push_config` still pushes `main` as
core, through push rights, not the merge whitelist.

Refs #4850
Refs #4853
2026-10-02 23:13:03 +02:00
atlas
efbfec6d01 config PRs: remove the hive's config-PR webhook, poll and core merge
An operator's merge on the forge deploys a config PR through
swarm-controller's DeployRequest{rev}. The hive-side path that queued a
MergeConfigPr approval and merged the PR as `core` goes:

- the `/webhook/config-pr` receiver, its HMAC secret, the WebhookRegister
  boot node and the org-hook registration; the hive vhost's `/webhook/`
  location
- the 5-minute config-PR poll
- ApprovalKind::MergeConfigPr, its dashboard card, and the deploy DAG it
  drove (DeployWindow, MergeVerify, DeployApply, FinalizeDeploy,
  DeployTail), with verify_commit, the two-phase meta deploy, rollback
  refs, the PR-failure comment and forge/pr_merge.rs
- `fetched_sha`, `sha_short`/`pr_number` on approval events, and
  `sha`/`tag` on HelperEvent::ApprovalResolved: only the merge path set
  them

`config_repo`, `merged_pr_for_commit` and `post_pr_comment` move to
forge/pr_comment.rs for the merged-rev deploy's refusal comment.
Approvals v5 drops stored `merge_config_pr` rows; a test reopens a v4
database holding them.

Closes #4850
2026-10-02 23:13:03 +02:00
atlas
a5ea015bc6 config PRs: document the operator merge, deploy only merges into main
docs: the config-repo `main` merge gate (merge = `core` + team
`operators`, approvals = `operators`), the operator merge in the forge
UI and what it deploys, the hand-added `operators` membership, and that
with no eval-verify a failed rebuild leaves `applied/main` at the merged
commit. The swarm README lists the converged gate and the merge deploy.

`merged()` also requires `pull_request.base.ref == "main"`: the hive
deploys its config repo's `main`, so a merge into another branch would
only cost a forge fetch and a refusal comment (argus, #4894).

Refs #4850
2026-10-02 23:13:03 +02:00
atlas
f1c695c212 config PRs: an operator's Forgejo merge deploys the merged rev
A config PR merged in the Forgejo UI changed nothing on the hive: the
hive's webhook ignores `closed`, its poll then cancels the dashboard
card, and `applied/main` stays where it was.

swarm-controller reads `merged`/`merge_commit_sha` off the
`pull_request` delivery it already receives for `agent-configs`, finds
the hive placing the agent by scanning every hive's wanted state (the
scan `declarations_elsewhere` already ran, factored out), and queues a
`TriggerDeploy` carrying the rev. Zero or several claimants deploy
nothing and log the claimants.

`DeployRequest` gains `rev: Option<String>` with `serde(default)`, so
rev-less payloads from either side keep decoding.

hive-c0re, given a rev for an agent it runs: a no-op when
`applied/main` already is the rev (a dashboard merge deploys its own
PR); otherwise it fetches the forge `main` with the core token,
requires the rev to descend from `applied/main` (the ancestry gate,
factored out of `run_deploy_merge_verify`), fast-forwards by CAS and
queues the usual relocking rebuild. No eval-verify on this path, per
mara (#4850 c90075). A refusal is commented on the PR that merged the
rev, found by commit.

swarm-controller's forge-objects pass converges every config repo's
`main` rule to merge whitelist `operators` + `core` and approval
whitelist `operators`. The hive's boot PATCH stops forcing
`enable_approvals_whitelist` off, so the two do not fight.

Refs #4850
2026-10-02 23:13:03 +02:00
atlas
9224c0bd15 swarm UI: fetch linked accounts only for the opened agent
The detail panel makes one request for the agent it shows,
GET /api/hives/{hive}/agents/{agent}/linked-accounts, which returns every
matrix, forge and github account of that agent as names and hosts. The
all-agents route and the table's matrix-column rows are removed, so the
table makes no linked-accounts request. The panel stays keyed by
hive/agent. The bao grant is unchanged.

Refs #4855
2026-10-02 22:36:11 +02:00
atlas
7ccde4647b swarm UI: one request for every agent's linked accounts
GET /api/agents/linked-accounts returns one entry per agent that
/api/agents/status has a row for, as {hive, agent, accounts}, from one
store login. The agents page fetches it once (and again when a link dialog
closes) and hands each table row and the detail panel its agent's slice,
so the page makes no per-agent request. The per-agent route had no caller
left and is removed. The bao grant is unchanged: the same list on each
agent's matrix and forge metadata directories.

Refs #4855
2026-10-02 21:20:40 +02:00
atlas
3380c1915f swarm UI: show the accounts linked to each agent
GET /api/hives/{hive}/agents/{agent}/linked-accounts returns one row per
account linked to the agent, as kind, name and host: each matrix account
under swarm/agents/<agent>/matrix (with its homeserver, and the agent's own
`main` marked reserved), each forge label under swarm/agents/<agent>/forge
(with its url), and github when swarm/agents/<agent>/github-token exists
(host github.com, which is not stored). No credential field is in the
response type.

Listing those two directories needs a new controller grant: `list` on
secret/metadata/swarm/agents/+/matrix and .../+/forge only, pinned in
bao-grants.nix as the only metadata stanzas under agents/ beside the queue
revocation. Checked against a dev OpenBao 2.6.3: the grant lists those two
directories and is refused on agents/, agents/<agent>/, and a leaf.

The swarm UI agent detail panel shows all rows under "accounts"; the table
view's matrix column shows the matrix rows. The link badges stay.

Refs #4855
2026-10-02 21:20:40 +02:00
atlas
8e23feb01b github: PATs live in swarm bao; the agent fetches them itself
An operator links an agent's GitHub personal access token in the swarm UI
(LinkGithubAccountForm, "link github account" on /agents). swarm-controller's
PUT /api/hives/{hive}/agents/{agent}/github-account stores it at
swarm/agents/<agent>/github-token (swarm_secret_client::github), a flat leaf
under the agent's prefix that the agent's existing read grant already covers:
no policy change, and no list grant, since there is one token per agent.

In the agent, hive-agent-github-token (oneshot + 2-minute timer, as the agent
user, under its own store certificate, ordered before hive-github-notify)
reads that path and writes <state>/github-token, 0600 and agent-owned, the
file the gh wrapper, git credential helper and hive-github-notify already
read. It replaces the file by rename only when the bytes changed and never
deletes it: a hive-written github-token stays until a token is linked in the
swarm UI. It is installed only with a store address and
services.hyperhive.agent.github.enable.

Removed: the dashboard's CR3D3NTIALS page (credentials.html/js/css, its
build entries and H0M3 tile; GITHUB was its only tab), hive-c0re's
dashboard/matrix_accounts.rs with GET/POST /api/github-account,
priv_client::write_agent_github_token, the host socket's
SetAgentGithubToken and `hivectl github set-token`, and hive-priv's
WriteAgentGithubToken with write_agent_state_file, its only caller gone.

Docs: integrations/github.md and swarm/ui.md describe the swarm path,
swarm/credentials.md gains the store-path row, and the hive UI docs,
hivectl docs and security.md's hive-priv table drop the removed pieces.

Closes #4347
2026-10-02 17:48:27 +02:00
atlas
77533be234 docs(swarm): address review
swarm-controller/README.md "What it does" was still missing two route
groups argus caught: linked external matrix/forge accounts
(PUT .../matrix-accounts/{account}, .../forge-accounts/{label}) and
config-PR status (GET /api/config-prs, /api/agents/{name}/config-pr).
Verified against the router at main.rs:2874-2899.
2026-10-02 12:50:34 +02:00
atlas
270430a4b4 docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
2026-10-02 12:50:34 +02:00
atlas
2c7e586f47 forge: external forge accounts live in swarm bao; the agent fetches them itself
An operator now links an agent's external forge account (label, base URL,
token) in the swarm UI. swarm-controller stores it at
swarm/agents/<agent>/forge/<label>. There is no index: the store's
listing of the agent's forge/ directory is the set of accounts.

In the agent, hive-agent-forge-accounts (oneshot + 2-minute timer, as
the agent user, under its own store certificate) lists
swarm/agents/<agent>/forge/ with the `list` #4866 grants an agent on its
own metadata subtree, reads each account, and writes
<state>/forge-<label>-token and forge-<label>.json in the names and shape
hive-forge -f already reads. An empty listing (a 404, which `bao kv list
-format=json` answers with `{}` and an empty stderr) is zero accounts; a
denial or an unreachable store fails the unit. It never deletes: files
for labels not listed, including ones the hive wrote, stay as they are.

Removed: the dashboard FORGES tab (credentials.js/html section and its
CSS), hive-c0re's extra_forges.rs and its routes, priv_client's
extra-forge calls, and hive-priv's WriteAgentExtraForgeAccount /
DeleteAgentExtraForgeAccount with their helpers. The GITHUB tab and
WriteAgentGithubToken stay.

Also: persistence.md's matrix avatar note names the exit-75 restart on a
changed account listing, not the dashboard, as what brings a linked
account up.

Refs #4348
2026-10-01 18:05:33 +02:00
atlas
97fb76ce99 matrix: the agent's daemon pulls its linked accounts from bao itself
hive-matrix-daemon now learns which external matrix accounts it has from
the swarm secret store, under the agent's own certificate, and the hive
push chain for matrix is gone.

The daemon lists swarm/agents/<agent>/matrix/ (the `list` its policy
grants on its own metadata subtree), reads each account's homeserver
from its credential, and brings the accounts up with their tokens from
the store. Every two minutes it lists again and exits with 75 when the
set of linked accounts changed; the unit restarts on 75 without counting
a failure. A listed name whose credential reads as absent is skipped and
logged once. At start it removes the matrix-token-<a> /
matrix-account-<a>.json pairs a hive delivered (a sidecar marks a pair
as delivered; a declared tokenFile keeps its token).

Removed: CredentialNotice and the $SWARM.credential.* subject and NATS
grant, the controller's publish and its queue precondition on the PUT
route, hive-c0re's credential subscription arm and workers/credential.rs,
priv_client::write_agent_matrix_token, hive-priv's WriteAgentMatrixToken
and its helpers, and the daemon's state-dir account discovery.

Kept: WriteAgentGithubToken and the external-forge path
(WriteAgentExtraForgeAccount, extra_forges.rs) are untouched, and a
declared matrixAccounts tokenFile is still read when the store has no
token for that account.

Refs #4348
2026-10-01 17:43:28 +02:00
atlas
e04616eb70 swarm-secret-client: agents may list their own subtree; controller rewrites agent policies
render_agent gains a second stanza: list on
secret/metadata/swarm/agents/<agent>/*, next to the existing read on
secret/data/swarm/agents/<agent>/*. An agent can now learn which
credentials it holds by listing its own subtree. Metadata read, writes
and every other principal's paths stay refused.

An agent's policy was only written when it was minted, so existing agents
would never get the new stanza. swarm-controller now rewrites every
agent's policy at start (read_policy::ensure_agent_policies), with the
same 30s / 24h retry as ensure_hive_access. The roster is the store's
hive-agent-* cert-auth roles, listed with the controller's existing
`list` on auth/cert/certs; the writes use its existing grant on
sys/policies/acl/hive-*. Only the policy is written: mint_and_verify
also reissues the certificate, so the pass does not call it.

Refs #4348
2026-10-01 17:43:28 +02:00
atlas
7d217f8267 remove the create_repo agent tool
mara ruled on #4849 (c88934): "remove create_repo tool". The tool ran in
hive-c0re with the hive's core token, so it only ever worked for agents
on the hive that runs the forge.

Removed:

- the create_repo MCP tool and CreateRepoArgs (hive-agent-mcp)
- wire variants Request::CreateRepo and Response::RepoCreated
  (hive-core-agent-sock)
- hive-c0re's handle_create_repo, its valid_repo_name check and the
  dispatch arm
- forge::create_agent_repo and apply_operator_branch_protection, which
  had no other caller, plus AGENTS_ORG and OPERATORS_TEAM, whose only
  users they were
- the tool's docs (docs/tools/forge.md repo management, docs/turn-loop/
  mcp.md, the conventions tool-group table) and the doc comments that
  named it (hive-sock-client's response timeout, ensure_repo_creation_
  disabled, the security doc's merge-gate bullet)

ToolGroup::Forge is kept with no tools, the same way b88a5b24 kept
Lifecycle, so existing meta/capabilities.json grants still parse.

Forge state is untouched: existing agents/* repos keep their collaborators
and operators-team branch protection. The swarm-controller's own
create_repo (config-org repos) is a different path and is unchanged.

Closes #4849
2026-10-01 09:04:17 +02:00
atlas
9e06e191a3 swarm-controller: first credential-renewal pass waits for the queue connection
The renewal pass ran the moment swarm-controller started, before its
queue client had connected, so `WantedWriter::view` refused with
"client state Pending" and the pass logged
`agent credential renewal: pass failed; retrying next tick`. That fired
twice in 24h on muede-lpt2, both under a second after start, and would
trip a Grafana rule on that WARN on ordinary restarts.

The first pass now waits until the queue client is connected, polling
`swarm_queue_client::ensure_connected` every 5s the way
`AgentIconReader::create_when_connected` does. The wait is bounded by
one RECONCILE_INTERVAL (5 min): the old code's retry after a false start
also came one interval later, so a queue that never connects gets its
first pass no later than before. Hitting the bound logs one WARN and runs
the pass anyway. Only startup waits; a later disconnect still fails a
pass and logs the WARN.

Refs #4717
2026-09-30 15:46:56 +02:00
atlas
ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00
atlas
78d8d69c7f swarm-controller: mint each hive's matrix sender token
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.

swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.

swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.

hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.

The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
9a5a947f8d swarm-controller: fix broken rustdoc intra-doc link to AGENT_TOKEN_NAME
legacy_tokens.rs referenced [`AGENT_TOKEN_NAME`] unqualified, but the
const lives in the sibling agent_token module and isn't in scope here;
rustdoc's broken-intra-doc-links lint (denied) failed the docs build.
2026-09-29 22:13:32 +02:00
atlas
e9206505d4 swarm-controller: retry the legacy forge token sweep from the mint pass
The sweep ran once at startup and was never retried: a boot where
store::connect or the roster read failed (e.g. the controller up
before bao) left the tokens live until the next restart. It now runs
off forge::agent_token::spawn's five-minute mint pass, reusing that
pass's roster observation instead of a second store/roster read, so a
failed first tick retries on the next one. Idempotent, so a re-run
after a partial sweep deletes nothing extra. Drops the one-shot
startup spawn; one call path.

Also updates the forge.md and agent_token.rs docs that still said the
legacy hyperhive-<seconds> tokens stay until manually removed.

Refs #4644
2026-09-29 22:13:32 +02:00
atlas
23e0c313b8 swarm-controller: sweep agents' legacy hyperhive-* forge tokens at start
Before b5d07d4d, hive-c0re minted a new `hyperhive-<unix-seconds>` token
for an agent on every spawn and rebuild and never revoked one, so every
live agent's forge user carries a pile of write-scoped tokens nothing
holds. Nothing in the tree lists or deletes them.

On each start, swarm-controller now walks the store's hive-agent-*
roster (the one the swarm-agent mint pass walks), and for every agent
whose swarm-agent token that pass would keep, deletes each token named
exactly `hyperhive-<digits>`. It logs the count per agent and a total.

- `core` is refused by name in both the roster filter and the per-user
  delete: hive-c0re still names core's live admin token
  `hyperhive-<unix-seconds>`.
- An agent whose swarm-agent token is not current is skipped, because
  consumers fall back to `<state>/forge-token`, the last hyperhive-*
  token, until the swarm token is fetched.
- A failed list or delete is logged and skipped; the sweep does not
  retry and never blocks startup. A second start deletes nothing.

Tokens on forge users of agents no longer on the store roster (already
destroyed) are not reached.

Closes #4644
2026-09-29 22:13:32 +02:00
atlas
544a8dd228 swarm-controller: check every first declaration, and fail closed on the roster
Exempting offline and paused from the name rules let a reserved name be
placed anyway: a first `offline` declared it, and the `up` after it
passed as an agent already declared on the hive. `paused` alone sufficed,
since the hive deploys an absent agent declared paused. A first
declaration in any placing state now runs the name rules and the
placed-elsewhere check; an agent already declared on the hive skips both.

A rule-breaking name whose roster read fails, including when no identity
bridge is configured, was accepted with a warning nobody sees, as in
create_agent. No later step on this route refuses the name, so it now
refuses with 503.

The gate is taken only for a placing declaration, so a destroy and its
credential revocation no longer wait on creations.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
d3755603f4 swarm-controller: check only a first declaration, and never a stop's name
The state route ran the name rules and the placed-elsewhere check on
every up/offline/paused declaration. An agent already declared on the
hive whose name breaks a rule, and which is not in the roster, could then
only be destroyed from the swarm.

Now an agent this hive already declares in a placing state passes both
checks. For a first declaration the placed-elsewhere check still applies
to up, offline and paused alike, and the name rules to up only: offline
and paused are how an operator stops an agent. The route reads this hive's
declaration to tell, and refuses with 503/500 when it cannot.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
8b892f508b swarm-controller: refuse a state declaration create_agent would refuse
PUT /api/hives/{hive}/agents/{agent}/state wrote any identifier into any
hive's wanted state, and the hive first-deploys a declared agent it has
no container for. That skipped create_agent's checks on the name.

A declaration that places the agent (up, paused, offline) is now refused
with 400 for a new name that breaks a naming rule, and with 409 for a
name the swarm has placed on another hive. Both reuse create_agent's
helpers (broken_name_rules/name_verdict, placements_elsewhere), under the
same gate, and an unreadable wanted state on another hive refuses with
503/500 as creation does. `destroyed` places nothing and is not checked,
so the revocation path is unchanged.

An agent with no declaration at all is still accepted: swarm-ui declares
state for agents that predate swarm-level creation, which is how the swarm
adopts them. Whether to refuse such names instead is left open on #4804.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
efed43b672 swarm-controller: read queued placements before published ones
create_agent read the other hives' wanted state first and snapshotted the
queued SetAgentWanted nodes second. A node finishing between the two was
in neither — not yet declared at the first read, already terminal at the
second — so a second hive could get the same name.

The queue is now snapshotted first: a SetAgentWanted node only turns
terminal after its declaration is written, so anything terminal by then is
visible to the read that follows. The order lives in
`placements_elsewhere`, and a test that finishes a node between the two
reads fails with them swapped.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
5785c0024c Make agent creation swarm-only and refuse a name placed on another hive
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.

Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.

policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
f6ac007dc0 Drop the last references to the removed hive matrix login form
`PutMatrixAccountRequest`'s no-`Debug` comment named the deleted
`MatrixLoginForm` as its precedent; it now states the reason directly.
The daemon's missing-sidecar warning told operators to re-login via the
dashboard, which no longer has that form; it now points at re-linking
the account from the swarm UI, whose stored credential carries the
homeserver the sidecar is written from. Behaviour is unchanged.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
6a1d85c24f hive-dashboard: remove the MATRIX credentials tab and its login route
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.

Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.

The GITHUB tab is the credentials page's default tab now.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
215a8aedc4 swarm-controller: link AgentState by path in the revocation doc comment
rustdoc runs with -D rustdoc::broken-intra-doc-links and the type is
imported inside the function body, not at module scope, so the bare
link resolved to nothing and failed the workspace-doc derivation.
2026-09-28 23:46:17 +02:00
atlas
c21ec7719d swarm: tests and docs for the queue-credential revocation
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.

docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.
2026-09-28 23:46:17 +02:00
atlas
8caf688ee4 swarm: revoke an agent's queue credential when it is declared destroyed
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.

Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.

Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.

The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
2026-09-28 23:46:17 +02:00
atlas
7bc4b25f16 swarm-controller: backfill a missing queue-secret mint time instead of re-minting it
A stored queue secret with no `minted_at` counted as due, and no secret
minted before the renewal pass has one, so the first pass after deploy
would re-mint every agent's secret. Every reconnect before that agent's
next restart would then be refused.

Such a secret is now stamped instead: `minted_at = now` is written beside
the unchanged `value`, and its 45-day clock starts there. Only a secret
whose recorded mint time is at least 45 days old gets a new value.

The decision is `secret_step` (Keep / Backfill / Remint), and the pass
reports an unstamped secret as `Observed::Unstamped`. Backfill and
re-mint log different lines.
2026-09-28 22:24:07 +02:00
atlas
2115ec2bb3 swarm-controller: re-issue agent certificates and re-mint queue secrets at half-life
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:

- `MintAgentIdentity` (the node agent creation uses) when the stored
  certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
  read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
  swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
  re-decides, writes a fresh value with `minted_at`, reads it back, and logs
  the agent and the old age.

When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.

`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.

Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.

Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.

docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
2026-09-28 21:35:01 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
e974194e3a swarm: let an agent publish its own icon
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.

hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.

swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.

Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
2026-09-28 13:47:37 +02:00
atlas
9513058a71 swarm: serve an agent's icon at swarm scope
An agent is not fixed to a hive, so its icon cannot be resolved as
hive -> agent. This adds the swarm-level half: an `agent-icons` KV
bucket keyed by the agent name alone — no hive token, so an agent that
moves hives keeps its icon and one that is stopped still has one — and
`GET /api/agents/<name>/icon` on swarm-controller serving it
same-origin, like every other `/api/*` route swarm-ui calls.

404 is the "this agent has no icon" answer, the same contract the
per-agent harness's own `GET /icon` has for an unconfigured agent.
Until the agent-side publisher lands, that is every agent's answer:
the publisher runs inside the container and an agent's NATS grants are
hive-scoped, which cannot authorise a write to a single-token agent
key. The read side needs no grant change — the controller already
holds `$KV.*.>` and `$JS.API.DIRECT.GET.*.>`.

The response carries `Content-Security-Policy: sandbox` and `nosniff`:
the body is an operator-authored SVG served from this daemon's own
origin, and an SVG can carry script.

Hive-side icon serving is untouched.

Refs #4502
2026-09-28 13:47:37 +02:00
atlas
0c1fb44a4f swarm-controller: relay an agent's hive-free subjects beside the old ones
The terminal and turn-state relays now subscribe to `$SWARM.term.<agent>`
and `$SWARM.agent-state.<agent>`, the subjects a verified agent token is
granted, as well as the hive-scoped `<prefix>.<hive>.<agent>` an agent
still on its hive's shared credential publishes to. Both at once, so a
swarm-ui terminal keeps working whichever credential an agent connected
with and in whatever order hosts deploy.
2026-09-28 08:24:52 +02:00
atlas
fc97c237dc swarm-queue-client: one agent-token spelling, and no hive in AgentCredential
`swarm_queue_client::agent_token::format_agent_token` / `parse_agent_token`
are the spelling an agent presents its own queue secret in,
`swarm-agent.<agent>.<secret>`, and the one the auth-callout responder
reads back. The prefix is what separates it from an OIDC access token,
which may itself contain `.`. Parsing distinguishes "not an agent token"
(no prefix) from "a malformed one"; the error names the problem and never
the value. The module is store-free, so the agent formats its token
without linking the secret-store client.

`swarm_secret_client::queue::AgentCredential` loses `hive`: an agent's
identity is not tied to a hive, and nothing reads the field. Objects
already in the store carry it and still decode, since unknown fields are
ignored; a test parses one. The controller stops writing it.

With the credential no longer naming a hive, and the agent's policy
naming none since #4762, nothing in the mint consumes one. `hive` goes
from `mint_and_verify`, from the `MintAgentIdentity` node, and from
`POST /api/agents/{name}/identity`, which now takes no body and no longer
checks a hive against the roster; a caller that still sends one is not
refused, the body is ignored. `swarmctl agent mint-identity` loses
`--hive`, so passing it is now a usage error.
2026-09-28 08:24:52 +02:00
atlas
6170e74a31 swarm-bao: agent certificates issued by a store-generated agent CA
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.

- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
  tunes the mount, generates the root once (guarded on an empty issuer
  list, no replace branch), upserts the `swarm-agent` role (client
  certificates named `hive-agent-*` only, 90 days), caches the CA at
  /var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
  (client-ca.pem, then the agent CA). Host cert-auth roles still pin
  client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
  composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
  changed after openbao started, swarm-bao-agent-pki restarts
  openbao.service in the container; under `seal = "shamir"` it prints
  the step instead. Once swarm-bao-certs has a cached CA, later boots
  start openbao with it and do not restart.
- The controller policy gains exactly `update` on
  pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
  the leaf (the store generates the key), writes the agent's cert-auth
  role pinning the issuing CA bao returned, and writes the agent's
  policy as render_agent alone: the hive-shared queue credential stanza
  is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
  other pki role names); swarm-controller gets
  SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.

Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.

Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.

Closes #4756
2026-09-27 22:59:27 +02:00
atlas
0f58cdbde2 swarm-queue-client: install aws-lc-rs as the process rustls provider
rustls is built with both `ring` (async-nats's `ring` feature) and
`aws-lc-rs` (reqwest's `rustls` feature), so it cannot pick a
process-level default by itself. Since the queue started requiring TLS
(1d261b3f), async-nats builds its config with `ClientConfig::builder()`,
which panics without an installed default. The panic kills the async-nats
connector task, and every queue client (swarm-controller, hive-c0re, all
hive-agents) has sat in `Pending` since the 2026-09-25 23:04Z deploy.

Add `swarm_queue_client::install_crypto_provider()`, which installs
aws-lc-rs and ignores the "already installed" error. It is called first in
`main` of every binary that links async-nats: hive-agent, hive-c0re,
swarm-controller, swarm-nats-auth. `connect()` also calls it, so a new
binary that dials through this crate is covered without remembering to.

aws-lc-rs because reqwest already falls back to it when no default is
installed, so HTTPS in these processes keeps its current provider. The
other rustls users in the tree reach it only through reqwest, which never
panics here.

Closes #4738
2026-09-27 04:17:24 +02:00
atlas
7b1fe5f9d3 hive-sock-client, web proxy, HTTP clients: bound connect and response waits
hive-sock-client: each attempt now bounds connect (5s), write (10s) and
the wait for the response (60s by default). The response bound is per
call through the new `request_within`, which hive-agent's serve-loop
`Recv` uses with its 180s long-poll plus 30s headroom. A response
timeout is terminal rather than retried: the server holds the request,
so a retry re-sends something it may still act on and multiplies the
wait by the backoff schedule.

Outbound HTTP: the matrix login/whoami clients in swarm-controller and
hive-c0re's dashboard (5s connect, 30s request), the authelia-bridge
client (5s/30s; ensuring an identity runs an argon2 hash first) and the
ci-runner forge calls (5s/15s, config_pr_poll's forge budget) get a
connect_timeout and a request timeout. Timeout errors name the bound
that fired.

hive-agent's unix-socket extra web proxy bounds the connect (5s) and
the wait for the response head (30s, the http sibling's budget); the
body read stays unbounded.

Refs #4723
2026-09-27 03:46:53 +02:00
atlas
3202cde704 nix: gate hive-c0re on deploy.hive-controller.enable, drop hyperhive.enable
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.

Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:

- Dropped: each swarm service and its glue keeps only its own deploy
  toggle (authelia, bao and its PKI glue, grafana, victorialogs,
  victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
  the controller/nats/matrix-ctl/publisher/services-issuer identities),
  the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
  its assertion, which feed hive-c0re and write under its state dir, plus
  their policy-order entry; the network identity assertions; hive-tls's
  two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
  (`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
  `deploy.matrix.enable` plus their client-identity checks. They read
  deploy.matrix's token file and registration script; their deploy.bao
  inputs are the client-half options a hive sets to read a store it does
  not run, so gating on deploy.bao.enable would drop the tested
  remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
  and fires wherever the hive, the store or the homeserver runs: each
  turns the name into an identifier with no fallback.

On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.

Refs #4500
2026-09-26 01:19:49 +02:00
atlas
20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00
atlas
2776e121e5 swarm-controller: mint each agent's matrix account with the swarm's token
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.

The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.

`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
2026-09-25 08:31:01 +02:00
atlas
f1f59ea165 swarm-controller: make an existing forge user a site admin
POST /api/forge/users/{name}/admin reads the account and, when it is not
already a site admin, sets `admin` with admin_edit_user. It never
creates one: a human's account is made by their first authelia login,
so a missing one answers 404, saying the user has not logged in via SSO
yet. An existing admin is a success with nothing sent.

The edit carries `admin` alone. repo_creation_lockdown's login_name +
source_id = 0 would turn an SSO-made account into a local one: in
Forgejo 16 a source_id sets the login type.

An agent's name is refused, and so is any name when the roster can't be
read: a site admin ignores max_repo_creation, the lockdown that keeps an
agent's token from creating a repo and self-merging in it.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
ef494af188 hivectl: drop forge create-user; SSO makes a human's forge account
The forge now creates a human's account on their first authelia login,
so the verb has no job left. Deletes it, HostRequest::ForgeCreateUser,
its handler, provision_user_token, change_user_password and the hive's
TOKEN_SCOPES. change_user_password also passed the password as an
argument to `forgejo admin user change-password`, so it showed in the
container's process list.

ensure_user_exists and mint_token stay for the `core` bootstrap, their
one caller now. ensure_user_exists loses its password parameter: only the
deleted path set one.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
21c17772b8 swarm-controller: fix test helper's forge::Client initializer
The stub-client helper in forge.rs built Client { api } after #4703
added a url field, breaking compilation of swarm-controller's test
target on main. Clone the URL the helper already builds for
Forgejo::new so it can also populate Client.url.

Closes #4709
2026-09-24 20:29:42 +02:00
atlas
22f0acfd6d swarm-controller: the forge-token backfill creates a missing forge user
An agent with a hive-agent-* store identity but no forge user was
observed as NoForgeUser and dropped by plan(), so it never got a token.
plan() now keeps it, and queue_forge_token_mints inserts CreateForgeUser
ahead of MintAgentForgeToken with after_ok, the edge declare_agent_job
already uses. ensure_agent_user folds an existing user into success, so
the extra node is a no-op for agents that have one.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
52c8c0b0de swarm-controller: mint each agent's forge token and store it in bao
A MintAgentForgeToken node mints a fixed-name swarm-agent token with the
admin API, keeps it when the stored value's last eight and the normalised
scopes match the forge's list, and otherwise deletes and re-creates it.
The token is stored at swarm/agents/<agent>/forge-token. Agent creation
inserts the node, and a pass at start and every five minutes inserts it
for every agent holding a store identity whose token is missing or stale.

Refs #3782
2026-09-24 17:48:53 +02:00