Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/swarm-controller
Author SHA1 Message Date
atlas
0f58cdbde2 swarm-queue-client: install aws-lc-rs as the process rustls provider
rustls is built with both `ring` (async-nats's `ring` feature) and
`aws-lc-rs` (reqwest's `rustls` feature), so it cannot pick a
process-level default by itself. Since the queue started requiring TLS
(1d261b3f), async-nats builds its config with `ClientConfig::builder()`,
which panics without an installed default. The panic kills the async-nats
connector task, and every queue client (swarm-controller, hive-c0re, all
hive-agents) has sat in `Pending` since the 2026-09-25 23:04Z deploy.

Add `swarm_queue_client::install_crypto_provider()`, which installs
aws-lc-rs and ignores the "already installed" error. It is called first in
`main` of every binary that links async-nats: hive-agent, hive-c0re,
swarm-controller, swarm-nats-auth. `connect()` also calls it, so a new
binary that dials through this crate is covered without remembering to.

aws-lc-rs because reqwest already falls back to it when no default is
installed, so HTTPS in these processes keeps its current provider. The
other rustls users in the tree reach it only through reqwest, which never
panics here.

Closes #4738
2026-09-27 04:17:24 +02:00
atlas
7b1fe5f9d3 hive-sock-client, web proxy, HTTP clients: bound connect and response waits
hive-sock-client: each attempt now bounds connect (5s), write (10s) and
the wait for the response (60s by default). The response bound is per
call through the new `request_within`, which hive-agent's serve-loop
`Recv` uses with its 180s long-poll plus 30s headroom. A response
timeout is terminal rather than retried: the server holds the request,
so a retry re-sends something it may still act on and multiplies the
wait by the backoff schedule.

Outbound HTTP: the matrix login/whoami clients in swarm-controller and
hive-c0re's dashboard (5s connect, 30s request), the authelia-bridge
client (5s/30s; ensuring an identity runs an argon2 hash first) and the
ci-runner forge calls (5s/15s, config_pr_poll's forge budget) get a
connect_timeout and a request timeout. Timeout errors name the bound
that fired.

hive-agent's unix-socket extra web proxy bounds the connect (5s) and
the wait for the response head (30s, the http sibling's budget); the
body read stays unbounded.

Refs #4723
2026-09-27 03:46:53 +02:00
atlas
3202cde704 nix: gate hive-c0re on deploy.hive-controller.enable, drop hyperhive.enable
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.

Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:

- Dropped: each swarm service and its glue keeps only its own deploy
  toggle (authelia, bao and its PKI glue, grafana, victorialogs,
  victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
  the controller/nats/matrix-ctl/publisher/services-issuer identities),
  the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
  its assertion, which feed hive-c0re and write under its state dir, plus
  their policy-order entry; the network identity assertions; hive-tls's
  two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
  (`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
  `deploy.matrix.enable` plus their client-identity checks. They read
  deploy.matrix's token file and registration script; their deploy.bao
  inputs are the client-half options a hive sets to read a store it does
  not run, so gating on deploy.bao.enable would drop the tested
  remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
  and fires wherever the hive, the store or the homeserver runs: each
  turns the name into an identifier with no fallback.

On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.

Refs #4500
2026-09-26 01:19:49 +02:00
atlas
20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00
atlas
2776e121e5 swarm-controller: mint each agent's matrix account with the swarm's token
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.

The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.

`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
2026-09-25 08:31:01 +02:00
atlas
f1f59ea165 swarm-controller: make an existing forge user a site admin
POST /api/forge/users/{name}/admin reads the account and, when it is not
already a site admin, sets `admin` with admin_edit_user. It never
creates one: a human's account is made by their first authelia login,
so a missing one answers 404, saying the user has not logged in via SSO
yet. An existing admin is a success with nothing sent.

The edit carries `admin` alone. repo_creation_lockdown's login_name +
source_id = 0 would turn an SSO-made account into a local one: in
Forgejo 16 a source_id sets the login type.

An agent's name is refused, and so is any name when the roster can't be
read: a site admin ignores max_repo_creation, the lockdown that keeps an
agent's token from creating a repo and self-merging in it.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
ef494af188 hivectl: drop forge create-user; SSO makes a human's forge account
The forge now creates a human's account on their first authelia login,
so the verb has no job left. Deletes it, HostRequest::ForgeCreateUser,
its handler, provision_user_token, change_user_password and the hive's
TOKEN_SCOPES. change_user_password also passed the password as an
argument to `forgejo admin user change-password`, so it showed in the
container's process list.

ensure_user_exists and mint_token stay for the `core` bootstrap, their
one caller now. ensure_user_exists loses its password parameter: only the
deleted path set one.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
21c17772b8 swarm-controller: fix test helper's forge::Client initializer
The stub-client helper in forge.rs built Client { api } after #4703
added a url field, breaking compilation of swarm-controller's test
target on main. Clone the URL the helper already builds for
Forgejo::new so it can also populate Client.url.

Closes #4709
2026-09-24 20:29:42 +02:00
atlas
22f0acfd6d swarm-controller: the forge-token backfill creates a missing forge user
An agent with a hive-agent-* store identity but no forge user was
observed as NoForgeUser and dropped by plan(), so it never got a token.
plan() now keeps it, and queue_forge_token_mints inserts CreateForgeUser
ahead of MintAgentForgeToken with after_ok, the edge declare_agent_job
already uses. ensure_agent_user folds an existing user into success, so
the extra node is a no-op for agents that have one.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
52c8c0b0de swarm-controller: mint each agent's forge token and store it in bao
A MintAgentForgeToken node mints a fixed-name swarm-agent token with the
admin API, keeps it when the stored value's last eight and the normalised
scopes match the forge's list, and otherwise deletes and re-creates it.
The token is stored at swarm/agents/<agent>/forge-token. Agent creation
inserts the node, and a pass at start and every five minutes inserts it
for every agent holding a store identity whose token is missing or stale.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
2fda529ca8 swarm-controller: accept the legacy pre-alignment email in the lockdown guard
disable_repo_creation's ownership guard now matches either the aligned
{agent}@hyperhive.local email or the legacy {agent}@hive.local one
hive-c0re::forge::users::ensure_user_exists used before its own
ensure_user_email alignment pass existed. That pass only runs on the
forge-host hive, and only once it has a core token and has ticked, so a
real pre-existing agent can still carry the old email when this node reads
it. Without this, the guard would bail "not this agent's" on a genuine
agent during that rollout window (argus, round 2).

Verified separately (not a code change): swarm-controller's forge account
is a site admin (created with --admin), and Forgejo's
convert.toUser/ToUser only hides an account's email when the caller isn't
the admin and isn't the account itself (services/convert/user.go), so
user_get already returns the real email regardless of hide_email on the
target account. No endpoint change needed for that half of the review.
2026-09-24 17:30:17 +02:00
atlas
58f50ed506 swarm-controller: guard the lockdown PATCH against a forge name collision
disable_repo_creation now reads the account once before PATCHing
max_repo_creation/source_id, and fails the node if the email isn't the
{agent}@hyperhive.local marker create_agent_user itself sets. The 409/422
create-fold (#4681) only proves some account with that name exists, not
that this node created it, so a pre-existing non-agent account sharing an
agent's chosen name could otherwise get locked onto local auth with repo
creation disabled. Leaves the fold untouched (Option B, per atlas/argus on
#4693); the read moves into disable_repo_creation instead.

Also fixes the nix-sandboxed cargo-test check: forgejo_api::Forgejo::new
builds a reqwest client that eagerly resolves TLS roots via
rustls-native-certs even for the tests' plain-http loopback stub server,
which panics with "No CA certificates were loaded from the system" in the
CA-less build sandbox. Gives that check's nativeBuildInputs pkgs.cacert and
sets SSL_CERT_FILE, same pattern this repo's runtime deployment already
uses for the same reqwest/rustls resolution.
2026-09-24 17:30:17 +02:00
atlas
04d5bf5f3e swarm-controller: disable repo creation on the agent forge users it creates
`ensure_agent_user` (the `CreateForgeUser` node) created the agent's
forge account but never set `max_repo_creation = 0`, relying on
hive-c0re's per-hive `ensure_repo_creation_disabled` pass to lock it
down later. That pass is being removed (#3507, #4669), and it is the
only guard against an agent token creating, owning and self-merging in
its own repo.

The node now PATCHes `max_repo_creation = 0` via `admin_edit_user`
after the create, on both the created and the already-exists path, and
a refused PATCH fails the node with Forgejo's message. The body mirrors
hive-c0re's `sparse_edit_user_option` (`login_name` + `source_id = 0`,
everything else unset).

The PATCH is unconditional: Forgejo's API never returns
`max_repo_creation` (it is in `EditUserOption` only, not `User`), so
there is no current value to verify against first.

Closes #4689
2026-09-24 17:30:17 +02:00
atlas
065b11a99d swarm-controller: page repo_search past forgejo's first page
list_repos_with_open_issues only read one repo_search page (forgejo's
30-row default), so a repo past position 30 in the default alpha sort
silently dropped out of the issue-report/repo-dropdown data source, with
no truncation signal. The comment claiming RepoSearchQuery has no page
field was wrong -- Request::page()/page_size() are generic builder
methods independent of the query struct.

Adds page_search_results, a small paging loop over a fetch closure for
search-shaped (data: Option<Vec<T>>, no header) responses that can't use
the existing .all() helper (that needs a (Headers, Vec<T>) response
shape). Pages until a short page or a 40-page bound, erroring on the
bound rather than truncating again.

Closes #4675
2026-09-24 17:27:12 +02:00
atlas
244db367c4 swarm-controller: answer 404, not 503, for an issue-report on a nonexistent repo
get_issue_report mapped every error from Client::issue_report to 503
(StatusUnavailable) unconditionally. issue_report calls
issue_list_issues(org, repo, ...) first, and a Forgejo 404 there — a
typo'd or deleted repo — was flattened into the same 503 a genuine
forge outage produces, which tells a client to retry a request that
will never succeed.

Downcast the anyhow error back to forgejo_api::ForgejoError (same
pattern main.rs's wanted_error_status uses for
swarm_queue_client::Error) and check it structurally against the
ApiErrorKind::NotFound / UnexpectedStatusCode(404) shapes forgejo's
generated client produces for a 404, rather than string-matching the
rendered message. Only that case answers 404, naming the org/repo;
every other forge failure still answers 503. Updates the route's
OpenAPI response list to document the 404.

Closes #4701
2026-09-24 16:38:47 +02:00
atlas
88e571a463 swarm-controller: refuse a new agent name the forge would reject
Every agent becomes a Forgejo user of the same name, and nothing upstream
of `CreateForgeUser` knew what Forgejo refuses: `admin`, `api`, `foo-` or a
41-character name passed name validation and failed one node into
provisioning with Forgejo's 422.

`nix/reserved-names.nix` gains the 22 reserved usernames of Forgejo
v16.0.5 (`models/user/user.go:639-680`) that `[a-z0-9-]` can spell, and
the bare `-` (`models/repo/repo.go:67`). The dot and underscore entries are
left out, since our charset cannot produce them. The header's admission
rule grows a third class — a username the forge refuses — because that is
a failure behind the refusal.

The shape rules are not literals, so they live in
`hive_types::forge_username_violation`: no leading `-`, no `--`, no
trailing `-`, at most 40 characters. Beside `is_reserved_name`, not in
`Ident::parse`: an `Ident` is also a hive, label, account and subagent
name, and parsing runs on every read of an existing name.

`create_agent` used to WARN on a reserved name, deliberately: an operator
with agents already created under a colliding name would otherwise be
unable to re-run creation. That reason is kept, and narrowed to what it
protects. A name breaking either rule is now refused with a 400 naming the
rule when the name is NOT in the swarm roster, and still only warned
about when it is, so re-creating an existing agent keeps working. The
roster is read only for a rule-breaking name; when it cannot be read, a new
name and an existing one look alike, and this warns as before. The
hive-collision warning is unchanged.
2026-09-24 15:16:32 +02:00
atlas
6a87b25488 swarm-controller: answer 503 on a disconnected queue, not 500 or a hang
put_matrix_account writes the credential to bao before checking that
the queue it must notify is actually connected — only that a queue is
configured, via state.status.as_ref(). While the queue is Pending or
Disconnected, client.flush().await hangs (async-nats does not process
commands during the initial connect retry), so the request hangs until
nginx times out and the credential is already stored. Add the
ensure_connected check every sibling queue route already makes
(wanted.rs, term_stream.rs, agent_state_stream.rs), placed immediately
before store::connect() so a request that cannot be delivered never
reaches the store.

Same shape, lower impact, in webhook::announce_knowledge_change: it
awaits publish/flush inline in the webhook handler, so a disconnected
queue can run past forgejo's short delivery timeout. Add the same
ensure_connected guard, warn and return.

set_agent_state and get_hive_wanted map every writer error to 500,
including swarm_queue_client::Error::NotConnected surfaced through
WantedWriter::view/set. Add wanted_error_status, which downcasts the
anyhow::Error back to the concrete type and maps NotConnected to 503
(retryable) while leaving every other failure at 500. Update both
routes' OpenAPI descriptions to say so.

Closes #4688
2026-09-24 15:15:59 +02:00
atlas
295925abb6 swarm-controller: don't fold every 422 into "already exists" on user/repo create
Forgejo answers 422 for six different causes on admin user create
(ErrUserAlreadyExist, ErrEmailAlreadyUsed, ErrNameReserved,
ErrNameCharsNotAllowed, ErrEmailInvalid, ErrNamePatternNotAllowed) and
several on repo create, but is_already_exists() treated every one of
them as a conflict. A reserved or otherwise-refused name silently
folded to Done, so CreateForgeUser reported success with no user
created, and the graph's real failure only surfaced one node later as
a misleading AddRepoMember error.

ensure_agent_user and ensure_org_repo now trust a 409 unconditionally
(folds_into_success) but confirm a 422 with a follow-up user_get /
repo_get before folding it to success; an unconfirmed 422 fails with
forgejo's own message at error level. Webhook registration still uses
the old is_already_exists — it has no comparable follow-up read, so it
is out of scope here.

Closes #4678
2026-09-24 13:45:51 +02:00
atlas
8b6dc72526 swarm-controller: pin what the backfill route queues
The mint route shipped without tests while the create route beside it
has three, so the two properties that make it a backfill rather than a
second creation route were unasserted: that it refuses before queuing,
and that it queues the mint node and nothing else.

Queuing a config-repo scaffold or a deploy against an agent that already
exists is the failure the second of those catches, and it is invisible
from the status code -- a version that queued the whole creation graph
would answer 200 with a node id just the same.

The agent name gets its own arm. create_agent's hive is checked against
the swarm roster, which incidentally rejects a name that is not an
identifier; an agent name has no roster to check against, so the parse
is the only thing between a traversal and a store path built out of it.

MintAgentIdentityResponse derives Clone + Debug to match
CreateAgentResponse -- expect_err on the refusal arms needs Debug on the
success type.
2026-09-21 20:38:55 +02:00
atlas
1442168715 swarmctl: re-mint an existing agent's store identity
Agent creation at swarm level is event-driven and nothing sweeps for
agents missing a credential, so an agent created before a credential
joined the mint never receives one -- nothing comes back around to it.
Without a way to re-run the mint by hand, the only route to giving an
existing agent its queue credential would be to delete and recreate the
agent.

POST /api/agents/{name}/identity enqueues the same MintAgentIdentity
node POST /api/agents declares, rather than writing inline: a second
code path that mints an identity is a second place for the four strings
that have to agree to disagree. swarmctl agent mint-identity is the
operator end, the same POST-and-print-the-node-id shape agent create
already has.

--hive is required on both ends. Neither the CLI nor the controller
keeps a roster of which agent runs where, and the credentials this mints
name a hive, so a default would be a guess that hands an agent subjects
on a hive it does not run on.

Documents the backfill as a runbook step, and fills in the renewal cell
the credential matrix requires for the new row.
2026-09-21 20:38:55 +02:00
atlas
ffd5018b18 swarm: mint a per-agent queue credential beside the agent's store identity
Every agent on a hive authenticates to the swarm queue with the same
hive-scoped secret today, so at the auth callout one agent is
indistinguishable from its co-hived neighbours and no subject can be
scoped to one of them.

Mint a secret per agent instead, at swarm level, into
secret/swarm/agents/<agent>/queue -- inside the stanza every agent's ACL
document already grants, so no policy changes and no existing agent's
document is rewritten. It is written by the same node that already mints
the agent's certificate, and read back under the agent's own token
before that node reports success.

The secret is not derived from the agent's mTLS identity: the two
credentials answer different questions and coupling their lifetimes
would mean renewing either implied renewing the other. Nothing here
rotates a queue secret -- a re-run keeps the existing value and only
corrects the principal it names, because this function is re-run
deliberately against agents that are already connected. Revoking one
means deleting the path.

Nothing reads the new credential yet; this is the minting half.
2026-09-21 20:38:55 +02:00
atlas
8cc7f90c98 log: send records natively to journald, keep stdout off-unit
A record written to stdout carries no priority, so journald files the
whole stream at one level and the swarm log store shows `info` whatever
level `tracing` gave it. Under a systemd unit the process's stdout
already *is* the journal, so the fix is to speak the journal protocol
directly and let each record carry its own severity.

New `hive-log` crate holds the one sink chooser, called by `hive-c0re`,
`hive-agent` and `swarm-controller`. It builds the same `EnvFilter`
those binaries always built, then installs exactly one layer — never
both, since a journald layer stacked on the `fmt` layer under a unit
stores every record twice.

The choice is an fstat compare, not a presence test: a child inherits
`$JOURNAL_STREAM` even when its own stdout was redirected elsewhere, so
the variable existing proves nothing. The crate parses `dev:inode` out
of it and compares both numbers against an fstat of stdout, the
descriptor the `fmt` layer writes to by default. No match, unset, or
unparseable takes the `fmt` branch. A journald layer that fails to
construct despite a match falls back to `fmt` and warns through it —
a process must never fail to start because of its logger.
2026-09-21 15:52:57 +02:00
müde
78ebd7db3f fix doc-comment pointers broken by the module-eval split
Both referenced the now-deleted nix/module-eval.nix. The nix side of
each pairing is spread across multiple files post-split, so drop the
cross-reference rather than chase it across files.
2026-09-20 04:31:08 +02:00
iris
8aafe4eaee swarm-controller: trim oversized comments, drop a trivial wrapper fn
mara: "pls dont make comments longer than functions or fns that just
call a single other fn". Trimmed wanted.rs's module doc, apply()'s
doc, and declare_new_agent's doc down to what's non-obvious; deleted
WantedWriter::store (a one-line call to open_or_create with a
10-line doc comment above it) and inlined its body into its two
callers.
2026-09-18 22:59:29 +02:00
iris
c908a71f3e swarm-controller: drop the Intent split, one unconditional wanted-state write
mara: "i dont want any logic differene between the two cases" and
"do not refuse to recreate an agent". wanted.rs goes back to a
single write path (set), with no terminal-state refusal at all,
used identically by the pause/resume endpoint and the agent-creation
job node.

The agent-creation node's own idempotency requirement (re-running
create against a name that already has a declaration must not
silently pause it) now lives entirely in declare_new_agent: it reads
the current declaration first and only writes Paused when the agent
has no entry, or its entry is Destroyed (recreating a previously-
destroyed name is the fresh deploy that state's own doc comment
names as the way back).
2026-09-18 22:59:29 +02:00
iris
75af27abe0 swarm-controller: make agent-creation's wanted-state write idempotent
Intent::Create now no-ops when the agent already has a non-Destroyed
declaration, instead of unconditionally overwriting it to the
create path's state. Without this, re-running create against a name
that already has a wanted-state entry (an operator migrating a
pre-existing agent into this bookkeeping, or a retried request)
would silently pause an agent already running under some other
state. write() skips the network round-trip entirely when apply
returns the declaration unchanged.

mara: this does not match the expectation that agent creation is
idempotent so pre existing agents can be migrated
2026-09-18 22:59:29 +02:00
iris
48aa7a1e79 swarm-controller: let creating an agent reuse a destroyed name
Review caught a path the new declaration node breaks: an operator may
already declare an agent `Destroyed` over the per-agent state endpoint,
and `apply` then refuses any transition off that state. Before the
wanted-state node existed the refusal was inert at creation time, but
now creating an agent under a previously-destroyed name builds its
identity, repo and config, fails the declaration, and silently cancels
the deploy — while the caller sees a 200 and a job id.

`AgentState::Destroyed`'s own doc comment already sanctions this case
("no state that brings a destroyed agent back short of a fresh deploy");
nothing implemented it. Give `apply` an `Intent`, keep the refusal for
redeclares, and add `WantedWriter::create` for the one caller that is a
fresh deploy. A separate method rather than a parameter on `set`, so no
other caller can reach the override by passing an argument wrong.
2026-09-18 22:59:29 +02:00
iris
bfd019fe8d swarm-controller: declare a new agent paused at creation
Creating an agent queued its identity, forge repo and deploy, but never
wrote a wanted-state declaration for it — so the agent showed up in
swarm-ui as "no declaration", and the hive brought it up with nothing
saying whether it should be driving turns.

Add a `SetAgentWanted` job node that declares the agent `Paused` in its
hive's wanted-state bucket, using the same `WantedWriter::set` primitive
the per-agent state HTTP handler already uses. A fresh agent therefore
sits paused until the operator explicitly flips it to `Up`.

The node is a root — it needs only the hive and agent names known at
request time — but the deploy trigger now waits on it, so the pause is
in the store before the hive brings the container up rather than landing
some time after a freshly deployed agent has already started taking
turns.
2026-09-18 22:59:29 +02:00
atlas
676c45bc93 swarm: mint, publish and login-verify an agent's store identity at create
`swarm/agents/<agent>/bao-mtls` did not exist, and neither did any
per-agent identity at the secret store: `policy::agent_object_name`,
`render_agent` and `render_agent_with_queue` had been written and never
called outside their own tests. An agent's only "per-agent" secret today
is read under the HIVE's certificate, through a wide grant on
`swarm/agents/*` — so "per-agent" was presentational.

The swarm now mints the certificate, so no hive ever needs the capability
to mint one. `swarm-controller` is the service that does it: it already
logs in to the store, and its existing grant already covers exactly the
three objects written here (`create/update` on
`secret/data/swarm/agents/*`, `sys/policies/acl/hive-*` and
`auth/cert/certs/hive-*`). No new bao grant, and nothing co-located — a
cert-auth role pins its authority by value, per role, so the controller
issues from its own CA on its own host and pins that CA in the role it
writes. No existing role changes.

The mint node does not report success on a write. After publishing it
connects again, with the leaf it just issued and under the role it just
wrote, and reads the path back — so the policy, the role, the common name
and the leaf are exercised in production on every agent creation. A
certificate this code mints that the role this code writes will not accept
turns the job node red at creation time instead of surfacing later as an
agent container that cannot start.

`TriggerDeploy` gains an `after_any` edge on the mint, not `after_ok`: a
hive cannot pass down a certificate the swarm has not published, but a
host with no authority configured must still create agents exactly as it
does today.

The private key is generated in memory and never written to disk on the
controller — `SecretStore::connect_with_identity` takes the PEM the minter
is already holding, so nothing is written out purely to be logged in with.

Refs #4137
2026-09-18 15:05:24 +02:00
atlas
c74249f371 matrix: make the hive-internal main account an ordinary matrixAccounts entry
`matrixAccounts` is meant to be the agent's full account list, but the
hive-internal `main` account was outside it: the nix module emitted only
the extras and `hive-matrix-daemon` prepended a `main` it synthesized
from the per-agent paths, with the option schema forbidding the name
outright.

nix/agent-modules/matrix.nix now declares `main` itself, as an ordinary
entry under `matrix.enable`, from the state-dir paths the module already
used for its token path-watcher (now a shared `stateDir` binding) plus
`matrix.url`. The whole set, `main` included, is serialized to
HIVE_MATRIX_ACCOUNTS.

accounts::configured therefore synthesizes `main` only when the parsed
list carries none, and otherwise takes the declared one verbatim —
hoisting it to index 0, since the daemon reads index 0 as the primary
and nix serializes an attrset, so `main` sorts wherever its key falls.
Declared xor synthesized: an agent whose harness predates this entry
keeps working, a current one gets its own, and there is no arrangement
where `main` is duplicated or missing.

The reserved-name assertion is replaced rather than dropped: the name
must now be legal (the module uses it), but `main`'s tokenFile stays
pinned to `<state>/matrix-token`, since hive-c0re provisions the
hive-internal token there and nowhere else — a retarget would evaluate
fine and then never restore. The other two fields are mkDefault and free
to override.

Refs #4475
2026-09-18 09:34:44 +02:00
atlas
6ea81aae65 hive-c0re: render the new agent option paths into generated agent flakes
meta.rs writes each agent's flake, and it still named the pre-move
`hyperhive.*` paths — so every agent rebuild would print a rename
deprecation warning about a line no human wrote and no operator could fix.
A warning nobody can act on trains everyone to ignore the ones that matter,
which is the whole value of the alias shims.

Repoints the FORWARDED_VAR_OPTIONS table and every other emitted option
assignment (otel.*, docs.source, claudeCodePath, github.enable, user.name,
claudeMemoryMaxBytes) to `services.hyperhive.agent.*`, with the test
expectations that pin the rendered text. The flake input named `hyperhive`
(`hyperhive.url`, `hyperhive.inputs.nixpkgs.follows`,
`hyperhive.nixosConfigurations.*`), hive-tier `services.hyperhive.*` paths,
and the `@hyperhive.local` git identity share the word and are untouched.

Also repoints the same option paths where they appear in comments, rustdoc
and runtime message strings across the other crates — a refusal message
naming `hyperhive.allowedRecipients` sends an operator to a path that will
stop existing. Prose under docs/ is deliberately not in this commit.

Refs #4473
2026-09-17 20:19:30 +02:00
atlas
d8f6d99bf9 swarm-queue-client: request the bearer-authz scope when minting an agent token
`swarm-logs query` got a bare nginx 401 from the swarm log store on every
query. The agent OIDC client is registered for `authelia.bearer.authz`
(`swarm-authelia.nix`'s `agentClients` sets `bearerAuthz`), but registration
is not issuance: the token request asked for no scope, so the token came back
carrying none, and authelia's `/api/authz/auth-request` refuses that exactly
as it refuses an unauthenticated caller.

The same failure is already recorded in `swarm-otel.nix` against the
collector's client, on the same scope string — prometheus asks for no scopes
unless told to, and every scrape was refused at introspection. This is that
bug one layer down, so it gets the same shape of fix.

`scope` becomes an opt-in parameter alongside `audience`, not a hardcoded
value or a config field: the two travel together (registered ≠ requested
applies to both) and only the destination decides whether either is needed.
`None` keeps every other caller byte-identical — the NATS connect callback,
`auth.rs`'s bridge client and the OTLP push client all pass it.

Refs #4464
2026-09-17 09:51:44 +02:00
atlas
72bff26336 fix tracker refs flagged by check-issue-refs.sh
Reworded the two hits scripts/check-issue-refs.sh found — a bare hash-
number tag in .forgejo/workflows/ci.yml's push-trigger comment and a
hyperhive#4345 tag in matrix_account.rs's doc comment — into prose that
stands on its own, per the lint's own rule. Neither carried semantic
weight beyond what the prose already says once reworded.

Refs #4345
2026-09-16 14:53:44 +02:00
atlas
87af0f38d3 swarm-controller: make the homeserver default fn pure, fixing test race
homeserver_or_configured_default read DEFAULT_HOMESERVER_ENV internally,
so its three unit tests raced each other by set_var/remove_var-ing the
same process env var with no synchronization under cargo test's default
parallelism (argus, PR #4443 review).

Take the default as a plain parameter instead of reading the env var
inside the function. The one env read moves to a new
configured_default_homeserver() helper, called once at the edge
(main.rs's startup diagnostic); homeserver_or_configured_default itself
is now pure and its tests need no env mutation at all.

Refs #4345
2026-09-16 14:51:30 +02:00
atlas
c894192e4f swarm-controller: add configured default matrix homeserver URL
Adds services.hyperhive.deploy.swarm-controller.matrixHomeserverUrl,
threaded to the daemon as SWARM_CONTROLLER_MATRIX_HOMESERVER_URL, and a
Rust helper (homeserver_or_configured_default) that lets a caller-supplied
homeserver keep overriding it. Config plumbing only: put_matrix_account
does not call the helper yet, so this is a no-op for every current caller.

Refs #4345
2026-09-16 14:43:13 +02:00
atlas
1ea3d87d7a swarm: publish each agent's turn-state header on its own subject
The swarm can already tell whether an agent is alive — the `agent-status`
KV bucket republishes once a minute — but not what it is doing right now.
A header bar wants the second thing, and a minute-old answer to "is this
agent thinking" is the wrong answer most of the time it is read.

`hive-agent` now publishes a turn-state header to
`$SWARM.agent-state.<hive>.<agent>`, a core subject beside the terminal
rows it already sends. It goes out **on transition, not on a timer**: the
publisher watches the event bus, rebuilds the header, and sends only when
the serialised result differs from the last one it sent — so a second
periodic writer, which is the problem this exists to fix, is not what
replaces the bucket.

The payload is the published contract a swarm-level renderer is written
against, so the test asserts on the serialised JSON keys rather than on
Rust field names. Two fields deliberately depart from the per-agent web
UI's `StateSnapshot`: `turn_state_since` is an ISO 8601 UTC string rather
than unix seconds, matching the sibling `$SWARM.term` subject's stamp, and
`agent_state` carries the swarm's own `AgentState` vocabulary rather than
a `paused` boolean, so a reader can compare actual against wanted without
translating. `turn_state` and `agent_state` stay two separate fields:
neither vocabulary contains the other's values.

Swarm-side, `GET /api/agents/{name}/state/stream` relays the subject as
SSE, resolving the agent's hive at request time exactly as the terminal
stream does and passing the bytes through without parsing them.

The broker grant is a second `--agent-publish-subject` rather than a
widening of the existing one, so the terminal family and the header family
stay independently revocable, and a `module-eval` arm pins the rendered
flag and its argument together — the doubled dollar included, since a
single one expands to nothing in `ExecStart` and yields a grant that
matches nothing.

Refs #3802
2026-09-14 15:12:23 +02:00
iris
c43a457752 term_stream: drop hive from the URL, resolve it from agent_status
mara's review point on #4351: an agent isn't pinned to a hive forever
(it can move), so a URL naming one would go stale the moment it did.
Resolve the hive at request time from the agent-status bucket instead
-- the same source AgentStatusRow.hive already comes from -- rather
than trusting a caller-supplied value. Route is now
GET /api/agents/{name}/term/stream; a never-reported agent now answers
404 (no hive on record) instead of silently guessing.
2026-09-13 18:08:33 +02:00
iris
1d902a0992 swarm-controller: add GET /api/agents/{hive}/{name}/term/stream
hive-agent already publishes classified TermMsg rows to the core NATS
subject $SWARM.term.{hive}.{agent} -- live only, no retention, by
design. This endpoint subscribes that subject per request and relays
each row over SSE, opaque to this daemon (no TermMsg dependency, same
pass-through shape crate::status already uses for hive snapshots).

No replay/history: the publish side never grew JetStream retention, and
this route's own job (a live tail) never needed it.
2026-09-13 18:08:33 +02:00
atlas
45e73f8636 swarm: say the read policy names the hive it is written for
`policy::render()` became `render(hive)` when a hive gained read on its own
entry, so two places now describe a document that no longer exists: this
module's header said it "is the same for every hive and depends on nothing",
and the security doc said the grant reaches the agent-credential prefix and
nothing else.

The module header is the load-bearing one. It sits above `write_policy_for`
and says, to anyone about to touch that function, that the render is
hive-independent — which is an invitation to hoist it to a shared constant
and hand every hive the stanza naming one of them.
2026-09-12 11:41:01 +02:00
atlas
395ecbdf41 swarm-secret-client: name the hive queue credential, and grant a hive its own kind
The agreement half of delivering the agent queue principal's client secret
through the store. No producer yet, so nothing writes this path — the unit
that does lands in the same PR, with the write grant it needs.

queue.rs is the sibling matrix.rs prescribes for a second kind of secret
rather than another field on a shared struct. Keyed per HIVE, not per agent:
the queue identity is minted once per hive at deploy time and says which hive
an agent belongs to, never which agent.

The client id rides with the secret for matrix.rs's stated reason — a
credential has to be reconstructable from the store alone, and deriving
`hive-<name>-agent` on the reading side is the split spelling the authelia
module warns denies every agent as a timeout.

policy.rs's render() takes the hive name now and emits a second, narrow
stanza for that hive's own path. The agent stanza is untouched: an agent's
path does not name its hive, so narrowing it still needs the enumeration
docs/trust-boundary/security.md rejects. A hive path does name its principal,
so scoping it costs nothing and drifts nowhere.

every_hive_gets_a_byte_identical_document is replaced rather than deleted.
Its surviving half is that the text is a function of the deploy-time name
alone, so a re-emission cannot drift; the new arms are that one hive's
document cannot reach another's path, and that a name which could close the
stanza is refused — live again now that a name reaches the document text.

Refs #3853
2026-09-12 10:56:50 +02:00
damocles
7773f10646 remove the two remaining as_str-style wrappers (DeliveryKind::parse, Wanted::parse) 2026-09-12 00:06:31 +02:00
damocles
46456f75ce remove as_str() legacy wrappers, callers use .into() directly 2026-09-12 00:06:31 +02:00
damocles
299add158f convert hand-written enum as_str matches to strum derives workspace-wide 2026-09-12 00:06:31 +02:00
atlas
b8c5840299 swarm-controller: retry hive provisioning until the store is up
The read policy and cert-auth role for each hive were written once, at
startup. On the deploy that surfaced this, the store was still coming
up, the pass logged its warning and moved on, and no hive could log in
until someone restarted the daemon — while cert auth answered "no chain
matching all constraints", which reads like a certificate problem
rather than a role that was never created.

The bootstrap unit in swarm-bao.nix lost the same race and won on its
retry 30s later. A daemon that boots alongside its store loses that race
routinely; on a normal boot it is the ordinary case.

The two passes fold into one `provision()` that logs in once instead of
twice for two loops over the same list, keeping policy before role since
the role names the policy. `ensure_hive_access` still awaits the first
pass, so a store that is already up leaves nothing deferred, and only a
pass that could not reach the store at all spawns the retry.

The retry is `config_pr::spawn`'s idiom from this same crate: an
interval task whose first tick is immediate. Its cadence and bound match
the bootstrap unit's — 30s, ~a day — because the two halves of one race
should not disagree about how long a wait is worth.

`Error::MissingEnv` is what keeps it from spinning forever: no `BAO_*`
set means a deployment that runs no store, where asking again changes
nothing, so it returns Ok. Everything else is retryable, including an
authority file that is not placed yet — the unit that writes it starts
alongside this one. Both cases previously landed in the same "not
managed here" line, so a store that was late looked exactly like one
that was never configured.

Per-hive failures keep their old behaviour: logged, skipped, Ok. A store
that refuses one hive's write refuses it again, so the next start really
is the right retry for those, and the module doc still says so.

Closes #4176.
2026-09-11 09:03:30 +02:00
iris
513554fe9a swarm: add a declared "paused" agent wanted state
mara (#4170): swarm-ui's wanted-state dropdown could only ever declare
up/offline/destroy, with no way to swarm-declare the existing hive-local
turn-loop pause (`hivectl agent pause|resume`).

`AgentState::Paused` is not a fifth peer of Up/Offline/Destroyed on the
power axis this enum otherwise answers — it's Up plus an orthogonal
turn-loop pause. `hive-c0re`'s `workers::wanted` reconcile loop now
decides the two axes independently (`decide` for power, the new
`decide_pause` for the marker), so a stopped agent declared Paused
converges with both a Start and a Pause in the same pass.

Known, deliberate limitation: a Paused declaration on an agent this
hive has never deployed only reaches Deploy this pass — writing the
pause marker into a harness dir that may not exist yet was judged not
worth the risk, so it converges on the next pass once the agent is
present instead.

swarm-ui's WantedMenu gains a fourth "paused" option (warning-tone
badge). No separate "resume" entry — selecting "up" from a paused row
already clears the marker via the same decide_pause path.

Pause/resume marker writes go through one shared
Coordinator::set_paused_by_name helper, used by both the interactive
dashboard pause/resume handlers and this reconcile loop, instead of
each duplicating the parse-name/write-marker/track-rescan shape.
swarm-ui's "offline" and "paused" confirm dialogs share one
confirmTarget state and one ConfirmDialog instead of two near-identical
copies.

Closes #4170
2026-09-11 01:35:40 +02:00
atlas
4007fc965d swarm-controller: create each hive's cert-auth role at startup
A hive holds an mTLS pair and a policy naming what it may read, and still
cannot log in: nothing creates the role that maps its certificate to that
policy. The one pre-shared credential in the system therefore buys no
access.

Minting happens here rather than in nix, which was the first plan. Nix
mints from the store's own container, and that path is gated on the
bootstrap token -- so onboarding a hive later would mean placing the one
genuinely pre-shared secret again. Doing it from the controller costs a
public certificate authority as an input and makes the bootstrap token
one-time.

A startup pass, not a hook: the hive list is loaded once and a config
change means a redeploy, so the roles are as static as the list. Only the
policy is derived from something that moves.

The subject is the hive's name because glue-bao-tls.nix mints a hive's
client leaf with its name as the CN, and cert auth matches on that.

Per-hive failures are logged and skipped, matching the queue, bridge and
forge connects above it: a controller whose store is unreachable still
serves everything else, and the next start retries.

Not covered by a test: ensure_hive_roles is IO from end to end, and the
seam that would make it assertable is the one the read-grant sink already
has. Said here rather than implied by a green suite.
2026-09-10 00:25:07 +02:00
damocles
a4f72365c7 check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00
damocles
e1e913015d check-issue-refs: blanket-ban tracker tags in markdown too, no exceptions 2026-09-09 21:02:48 +02:00
atlas
3752482524 swarm: let every hive read every agent's credential, and say so
A hive reads its agents' credentials with its own certificate, and nothing
said which paths that certificate may read, so the read half of a delivery
answered 403.

The grant is wide on purpose. An agent's path does not name the hive
hosting it -- agents move -- so a per-hive grant has to be an enumeration
the controller re-emits whenever the roster changes, and an enumeration
that can drift or land out of order advertises a boundary it does not
hold. A wide grant that says what it is beats a narrow one that only looks
narrow. mara's call, on the PR: rather a too-lax scope than one that
pretends to be strict.

What that buys, beyond honesty: the document is identical for every hive
and depends on nothing, so it is written once at startup beside the rest of
a hive's provisioning instead of on every declaration. No derived state, no
re-emission, and the ordering hazard that came with one stops existing.

What still holds is read-only. A hive cannot write an agent's credential,
so it cannot hand itself an agent's identity, and the grant reaches nothing
in the store outside the agent-credential prefix.

The fact is documented where someone meets the boundary rather than only in
this message, and the two ways to narrow it later -- scope per hive, or
give agents their own store identity -- are tracked.
2026-09-09 18:40:41 +02:00
atlas
e638db262e swarm-controller: keep a hive's read grant in step with its declaration
The controller writes an agent's credential; the hive fetches it back with
its own certificate. Nothing said which paths that certificate may read, so
the read half of a delivery answers 403 with no way to tell why.

The grant is derived from the declaration, so it is re-rendered at the one
place the declaration changes -- WantedWriter::set -- rather than at its
caller, which would work today and break on the second caller.

Emitted before the KV write: a grant that lands late is a 403 on an agent's
first fetch, while one that shrinks early only affects an agent already
being torn down. A failed write then leaves a superset the next declaration
re-renders.

Destroyed agents are filtered out. The declared set is a hive's whole
history -- a destroyed entry stays so that redeclaring it Up is refused as
the terminal transition it is -- so granting every declared agent would
leave a torn-down agent's credentials readable forever.

The sink is a trait because a missed emission is that same untraceable 403:
the double pins which agents were published, and the no-sink and refusing
arms pin the two deployments that are not a happy path. Not covered: the
call site inside set(), which needs a live queue.

The cert role moves to its own module on the way past. It is the
controller's identity at the store, not something the matrix route owns,
and the policy writer needs the same login.
2026-09-09 18:40:41 +02:00