render_agent gains a second stanza: list on
secret/metadata/swarm/agents/<agent>/*, next to the existing read on
secret/data/swarm/agents/<agent>/*. An agent can now learn which
credentials it holds by listing its own subtree. Metadata read, writes
and every other principal's paths stay refused.
An agent's policy was only written when it was minted, so existing agents
would never get the new stanza. swarm-controller now rewrites every
agent's policy at start (read_policy::ensure_agent_policies), with the
same 30s / 24h retry as ensure_hive_access. The roster is the store's
hive-agent-* cert-auth roles, listed with the controller's existing
`list` on auth/cert/certs; the writes use its existing grant on
sys/policies/acl/hive-*. Only the policy is written: mint_and_verify
also reissues the certificate, so the pass does not call it.
Refs #4348
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
The sweep ran once at startup and was never retried: a boot where
store::connect or the roster read failed (e.g. the controller up
before bao) left the tokens live until the next restart. It now runs
off forge::agent_token::spawn's five-minute mint pass, reusing that
pass's roster observation instead of a second store/roster read, so a
failed first tick retries on the next one. Idempotent, so a re-run
after a partial sweep deletes nothing extra. Drops the one-shot
startup spawn; one call path.
Also updates the forge.md and agent_token.rs docs that still said the
legacy hyperhive-<seconds> tokens stay until manually removed.
Refs #4644
Before b5d07d4d, hive-c0re minted a new `hyperhive-<unix-seconds>` token
for an agent on every spawn and rebuild and never revoked one, so every
live agent's forge user carries a pile of write-scoped tokens nothing
holds. Nothing in the tree lists or deletes them.
On each start, swarm-controller now walks the store's hive-agent-*
roster (the one the swarm-agent mint pass walks), and for every agent
whose swarm-agent token that pass would keep, deletes each token named
exactly `hyperhive-<digits>`. It logs the count per agent and a total.
- `core` is refused by name in both the roster filter and the per-user
delete: hive-c0re still names core's live admin token
`hyperhive-<unix-seconds>`.
- An agent whose swarm-agent token is not current is skipped, because
consumers fall back to `<state>/forge-token`, the last hyperhive-*
token, until the swarm token is fetched.
- A failed list or delete is logged and skipped; the sweep does not
retry and never blocks startup. A second start deletes nothing.
Tokens on forge users of agents no longer on the store roster (already
destroyed) are not reached.
Closes#4644
Exempting offline and paused from the name rules let a reserved name be
placed anyway: a first `offline` declared it, and the `up` after it
passed as an agent already declared on the hive. `paused` alone sufficed,
since the hive deploys an absent agent declared paused. A first
declaration in any placing state now runs the name rules and the
placed-elsewhere check; an agent already declared on the hive skips both.
A rule-breaking name whose roster read fails, including when no identity
bridge is configured, was accepted with a warning nobody sees, as in
create_agent. No later step on this route refuses the name, so it now
refuses with 503.
The gate is taken only for a placing declaration, so a destroy and its
credential revocation no longer wait on creations.
Refs #4804
The state route ran the name rules and the placed-elsewhere check on
every up/offline/paused declaration. An agent already declared on the
hive whose name breaks a rule, and which is not in the roster, could then
only be destroyed from the swarm.
Now an agent this hive already declares in a placing state passes both
checks. For a first declaration the placed-elsewhere check still applies
to up, offline and paused alike, and the name rules to up only: offline
and paused are how an operator stops an agent. The route reads this hive's
declaration to tell, and refuses with 503/500 when it cannot.
Refs #4804
PUT /api/hives/{hive}/agents/{agent}/state wrote any identifier into any
hive's wanted state, and the hive first-deploys a declared agent it has
no container for. That skipped create_agent's checks on the name.
A declaration that places the agent (up, paused, offline) is now refused
with 400 for a new name that breaks a naming rule, and with 409 for a
name the swarm has placed on another hive. Both reuse create_agent's
helpers (broken_name_rules/name_verdict, placements_elsewhere), under the
same gate, and an unreadable wanted state on another hive refuses with
503/500 as creation does. `destroyed` places nothing and is not checked,
so the revocation path is unchanged.
An agent with no declaration at all is still accepted: swarm-ui declares
state for agents that predate swarm-level creation, which is how the swarm
adopts them. Whether to refuse such names instead is left open on #4804.
Refs #4804
create_agent read the other hives' wanted state first and snapshotted the
queued SetAgentWanted nodes second. A node finishing between the two was
in neither — not yet declared at the first read, already terminal at the
second — so a second hive could get the same name.
The queue is now snapshotted first: a SetAgentWanted node only turns
terminal after its declaration is written, so anything terminal by then is
visible to the read that follows. The order lives in
`placements_elsewhere`, and a test that finishes a node between the two
reads fails with them swapped.
Refs #4396
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.
Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.
policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.
Refs #4396
rustdoc runs with -D rustdoc::broken-intra-doc-links and the type is
imported inside the function body, not at module scope, so the bare
link resolved to nothing and failed the workspace-doc derivation.
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.
docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.
Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.
Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.
The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:
- `MintAgentIdentity` (the node agent creation uses) when the stored
certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
re-decides, writes a fresh value with `minted_at`, reads it back, and logs
the agent and the old age.
When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.
`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.
Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.
Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.
docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.
Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.
Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.
Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.
hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.
swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.
Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
An agent is not fixed to a hive, so its icon cannot be resolved as
hive -> agent. This adds the swarm-level half: an `agent-icons` KV
bucket keyed by the agent name alone — no hive token, so an agent that
moves hives keeps its icon and one that is stopped still has one — and
`GET /api/agents/<name>/icon` on swarm-controller serving it
same-origin, like every other `/api/*` route swarm-ui calls.
404 is the "this agent has no icon" answer, the same contract the
per-agent harness's own `GET /icon` has for an unconfigured agent.
Until the agent-side publisher lands, that is every agent's answer:
the publisher runs inside the container and an agent's NATS grants are
hive-scoped, which cannot authorise a write to a single-token agent
key. The read side needs no grant change — the controller already
holds `$KV.*.>` and `$JS.API.DIRECT.GET.*.>`.
The response carries `Content-Security-Policy: sandbox` and `nosniff`:
the body is an operator-authored SVG served from this daemon's own
origin, and an SVG can carry script.
Hive-side icon serving is untouched.
Refs #4502
`swarm_queue_client::agent_token::format_agent_token` / `parse_agent_token`
are the spelling an agent presents its own queue secret in,
`swarm-agent.<agent>.<secret>`, and the one the auth-callout responder
reads back. The prefix is what separates it from an OIDC access token,
which may itself contain `.`. Parsing distinguishes "not an agent token"
(no prefix) from "a malformed one"; the error names the problem and never
the value. The module is store-free, so the agent formats its token
without linking the secret-store client.
`swarm_secret_client::queue::AgentCredential` loses `hive`: an agent's
identity is not tied to a hive, and nothing reads the field. Objects
already in the store carry it and still decode, since unknown fields are
ignored; a test parses one. The controller stops writing it.
With the credential no longer naming a hive, and the agent's policy
naming none since #4762, nothing in the mint consumes one. `hive` goes
from `mint_and_verify`, from the `MintAgentIdentity` node, and from
`POST /api/agents/{name}/identity`, which now takes no body and no longer
checks a hive against the roster; a caller that still sends one is not
refused, the body is ignored. `swarmctl agent mint-identity` loses
`--hive`, so passing it is now a usage error.
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.
- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
tunes the mount, generates the root once (guarded on an empty issuer
list, no replace branch), upserts the `swarm-agent` role (client
certificates named `hive-agent-*` only, 90 days), caches the CA at
/var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
(client-ca.pem, then the agent CA). Host cert-auth roles still pin
client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
changed after openbao started, swarm-bao-agent-pki restarts
openbao.service in the container; under `seal = "shamir"` it prints
the step instead. Once swarm-bao-certs has a cached CA, later boots
start openbao with it and do not restart.
- The controller policy gains exactly `update` on
pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
the leaf (the store generates the key), writes the agent's cert-auth
role pinning the issuing CA bao returned, and writes the agent's
policy as render_agent alone: the hive-shared queue credential stanza
is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
other pki role names); swarm-controller gets
SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.
Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.
Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.
Closes#4756
rustls is built with both `ring` (async-nats's `ring` feature) and
`aws-lc-rs` (reqwest's `rustls` feature), so it cannot pick a
process-level default by itself. Since the queue started requiring TLS
(1d261b3f), async-nats builds its config with `ClientConfig::builder()`,
which panics without an installed default. The panic kills the async-nats
connector task, and every queue client (swarm-controller, hive-c0re, all
hive-agents) has sat in `Pending` since the 2026-09-25 23:04Z deploy.
Add `swarm_queue_client::install_crypto_provider()`, which installs
aws-lc-rs and ignores the "already installed" error. It is called first in
`main` of every binary that links async-nats: hive-agent, hive-c0re,
swarm-controller, swarm-nats-auth. `connect()` also calls it, so a new
binary that dials through this crate is covered without remembering to.
aws-lc-rs because reqwest already falls back to it when no default is
installed, so HTTPS in these processes keeps its current provider. The
other rustls users in the tree reach it only through reqwest, which never
panics here.
Closes#4738
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.
swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.
hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.
nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.
Refs #3782
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.
The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.
`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
POST /api/forge/users/{name}/admin reads the account and, when it is not
already a site admin, sets `admin` with admin_edit_user. It never
creates one: a human's account is made by their first authelia login,
so a missing one answers 404, saying the user has not logged in via SSO
yet. An existing admin is a success with nothing sent.
The edit carries `admin` alone. repo_creation_lockdown's login_name +
source_id = 0 would turn an SSO-made account into a local one: in
Forgejo 16 a source_id sets the login type.
An agent's name is refused, and so is any name when the roster can't be
read: a site admin ignores max_repo_creation, the lockdown that keeps an
agent's token from creating a repo and self-merging in it.
Refs #3782
An agent with a hive-agent-* store identity but no forge user was
observed as NoForgeUser and dropped by plan(), so it never got a token.
plan() now keeps it, and queue_forge_token_mints inserts CreateForgeUser
ahead of MintAgentForgeToken with after_ok, the edge declare_agent_job
already uses. ensure_agent_user folds an existing user into success, so
the extra node is a no-op for agents that have one.
Refs #3782
A MintAgentForgeToken node mints a fixed-name swarm-agent token with the
admin API, keeps it when the stored value's last eight and the normalised
scopes match the forge's list, and otherwise deletes and re-creates it.
The token is stored at swarm/agents/<agent>/forge-token. Agent creation
inserts the node, and a pass at start and every five minutes inserts it
for every agent holding a store identity whose token is missing or stale.
Refs #3782
Every agent becomes a Forgejo user of the same name, and nothing upstream
of `CreateForgeUser` knew what Forgejo refuses: `admin`, `api`, `foo-` or a
41-character name passed name validation and failed one node into
provisioning with Forgejo's 422.
`nix/reserved-names.nix` gains the 22 reserved usernames of Forgejo
v16.0.5 (`models/user/user.go:639-680`) that `[a-z0-9-]` can spell, and
the bare `-` (`models/repo/repo.go:67`). The dot and underscore entries are
left out, since our charset cannot produce them. The header's admission
rule grows a third class — a username the forge refuses — because that is
a failure behind the refusal.
The shape rules are not literals, so they live in
`hive_types::forge_username_violation`: no leading `-`, no `--`, no
trailing `-`, at most 40 characters. Beside `is_reserved_name`, not in
`Ident::parse`: an `Ident` is also a hive, label, account and subagent
name, and parsing runs on every read of an existing name.
`create_agent` used to WARN on a reserved name, deliberately: an operator
with agents already created under a colliding name would otherwise be
unable to re-run creation. That reason is kept, and narrowed to what it
protects. A name breaking either rule is now refused with a 400 naming the
rule when the name is NOT in the swarm roster, and still only warned
about when it is, so re-creating an existing agent keeps working. The
roster is read only for a rule-breaking name; when it cannot be read, a new
name and an existing one look alike, and this warns as before. The
hive-collision warning is unchanged.
put_matrix_account writes the credential to bao before checking that
the queue it must notify is actually connected — only that a queue is
configured, via state.status.as_ref(). While the queue is Pending or
Disconnected, client.flush().await hangs (async-nats does not process
commands during the initial connect retry), so the request hangs until
nginx times out and the credential is already stored. Add the
ensure_connected check every sibling queue route already makes
(wanted.rs, term_stream.rs, agent_state_stream.rs), placed immediately
before store::connect() so a request that cannot be delivered never
reaches the store.
Same shape, lower impact, in webhook::announce_knowledge_change: it
awaits publish/flush inline in the webhook handler, so a disconnected
queue can run past forgejo's short delivery timeout. Add the same
ensure_connected guard, warn and return.
set_agent_state and get_hive_wanted map every writer error to 500,
including swarm_queue_client::Error::NotConnected surfaced through
WantedWriter::view/set. Add wanted_error_status, which downcasts the
anyhow::Error back to the concrete type and maps NotConnected to 503
(retryable) while leaving every other failure at 500. Update both
routes' OpenAPI descriptions to say so.
Closes#4688
The mint route shipped without tests while the create route beside it
has three, so the two properties that make it a backfill rather than a
second creation route were unasserted: that it refuses before queuing,
and that it queues the mint node and nothing else.
Queuing a config-repo scaffold or a deploy against an agent that already
exists is the failure the second of those catches, and it is invisible
from the status code -- a version that queued the whole creation graph
would answer 200 with a node id just the same.
The agent name gets its own arm. create_agent's hive is checked against
the swarm roster, which incidentally rejects a name that is not an
identifier; an agent name has no roster to check against, so the parse
is the only thing between a traversal and a store path built out of it.
MintAgentIdentityResponse derives Clone + Debug to match
CreateAgentResponse -- expect_err on the refusal arms needs Debug on the
success type.
Agent creation at swarm level is event-driven and nothing sweeps for
agents missing a credential, so an agent created before a credential
joined the mint never receives one -- nothing comes back around to it.
Without a way to re-run the mint by hand, the only route to giving an
existing agent its queue credential would be to delete and recreate the
agent.
POST /api/agents/{name}/identity enqueues the same MintAgentIdentity
node POST /api/agents declares, rather than writing inline: a second
code path that mints an identity is a second place for the four strings
that have to agree to disagree. swarmctl agent mint-identity is the
operator end, the same POST-and-print-the-node-id shape agent create
already has.
--hive is required on both ends. Neither the CLI nor the controller
keeps a roster of which agent runs where, and the credentials this mints
name a hive, so a default would be a guess that hands an agent subjects
on a hive it does not run on.
Documents the backfill as a runbook step, and fills in the renewal cell
the credential matrix requires for the new row.
A record written to stdout carries no priority, so journald files the
whole stream at one level and the swarm log store shows `info` whatever
level `tracing` gave it. Under a systemd unit the process's stdout
already *is* the journal, so the fix is to speak the journal protocol
directly and let each record carry its own severity.
New `hive-log` crate holds the one sink chooser, called by `hive-c0re`,
`hive-agent` and `swarm-controller`. It builds the same `EnvFilter`
those binaries always built, then installs exactly one layer — never
both, since a journald layer stacked on the `fmt` layer under a unit
stores every record twice.
The choice is an fstat compare, not a presence test: a child inherits
`$JOURNAL_STREAM` even when its own stdout was redirected elsewhere, so
the variable existing proves nothing. The crate parses `dev:inode` out
of it and compares both numbers against an fstat of stdout, the
descriptor the `fmt` layer writes to by default. No match, unset, or
unparseable takes the `fmt` branch. A journald layer that fails to
construct despite a match falls back to `fmt` and warns through it —
a process must never fail to start because of its logger.
mara: "pls dont make comments longer than functions or fns that just
call a single other fn". Trimmed wanted.rs's module doc, apply()'s
doc, and declare_new_agent's doc down to what's non-obvious; deleted
WantedWriter::store (a one-line call to open_or_create with a
10-line doc comment above it) and inlined its body into its two
callers.
mara: "i dont want any logic differene between the two cases" and
"do not refuse to recreate an agent". wanted.rs goes back to a
single write path (set), with no terminal-state refusal at all,
used identically by the pause/resume endpoint and the agent-creation
job node.
The agent-creation node's own idempotency requirement (re-running
create against a name that already has a declaration must not
silently pause it) now lives entirely in declare_new_agent: it reads
the current declaration first and only writes Paused when the agent
has no entry, or its entry is Destroyed (recreating a previously-
destroyed name is the fresh deploy that state's own doc comment
names as the way back).
Review caught a path the new declaration node breaks: an operator may
already declare an agent `Destroyed` over the per-agent state endpoint,
and `apply` then refuses any transition off that state. Before the
wanted-state node existed the refusal was inert at creation time, but
now creating an agent under a previously-destroyed name builds its
identity, repo and config, fails the declaration, and silently cancels
the deploy — while the caller sees a 200 and a job id.
`AgentState::Destroyed`'s own doc comment already sanctions this case
("no state that brings a destroyed agent back short of a fresh deploy");
nothing implemented it. Give `apply` an `Intent`, keep the refusal for
redeclares, and add `WantedWriter::create` for the one caller that is a
fresh deploy. A separate method rather than a parameter on `set`, so no
other caller can reach the override by passing an argument wrong.
Creating an agent queued its identity, forge repo and deploy, but never
wrote a wanted-state declaration for it — so the agent showed up in
swarm-ui as "no declaration", and the hive brought it up with nothing
saying whether it should be driving turns.
Add a `SetAgentWanted` job node that declares the agent `Paused` in its
hive's wanted-state bucket, using the same `WantedWriter::set` primitive
the per-agent state HTTP handler already uses. A fresh agent therefore
sits paused until the operator explicitly flips it to `Up`.
The node is a root — it needs only the hive and agent names known at
request time — but the deploy trigger now waits on it, so the pause is
in the store before the hive brings the container up rather than landing
some time after a freshly deployed agent has already started taking
turns.
`swarm/agents/<agent>/bao-mtls` did not exist, and neither did any
per-agent identity at the secret store: `policy::agent_object_name`,
`render_agent` and `render_agent_with_queue` had been written and never
called outside their own tests. An agent's only "per-agent" secret today
is read under the HIVE's certificate, through a wide grant on
`swarm/agents/*` — so "per-agent" was presentational.
The swarm now mints the certificate, so no hive ever needs the capability
to mint one. `swarm-controller` is the service that does it: it already
logs in to the store, and its existing grant already covers exactly the
three objects written here (`create/update` on
`secret/data/swarm/agents/*`, `sys/policies/acl/hive-*` and
`auth/cert/certs/hive-*`). No new bao grant, and nothing co-located — a
cert-auth role pins its authority by value, per role, so the controller
issues from its own CA on its own host and pins that CA in the role it
writes. No existing role changes.
The mint node does not report success on a write. After publishing it
connects again, with the leaf it just issued and under the role it just
wrote, and reads the path back — so the policy, the role, the common name
and the leaf are exercised in production on every agent creation. A
certificate this code mints that the role this code writes will not accept
turns the job node red at creation time instead of surfacing later as an
agent container that cannot start.
`TriggerDeploy` gains an `after_any` edge on the mint, not `after_ok`: a
hive cannot pass down a certificate the swarm has not published, but a
host with no authority configured must still create agents exactly as it
does today.
The private key is generated in memory and never written to disk on the
controller — `SecretStore::connect_with_identity` takes the PEM the minter
is already holding, so nothing is written out purely to be logged in with.
Refs #4137
homeserver_or_configured_default read DEFAULT_HOMESERVER_ENV internally,
so its three unit tests raced each other by set_var/remove_var-ing the
same process env var with no synchronization under cargo test's default
parallelism (argus, PR #4443 review).
Take the default as a plain parameter instead of reading the env var
inside the function. The one env read moves to a new
configured_default_homeserver() helper, called once at the edge
(main.rs's startup diagnostic); homeserver_or_configured_default itself
is now pure and its tests need no env mutation at all.
Refs #4345
Adds services.hyperhive.deploy.swarm-controller.matrixHomeserverUrl,
threaded to the daemon as SWARM_CONTROLLER_MATRIX_HOMESERVER_URL, and a
Rust helper (homeserver_or_configured_default) that lets a caller-supplied
homeserver keep overriding it. Config plumbing only: put_matrix_account
does not call the helper yet, so this is a no-op for every current caller.
Refs #4345
The swarm can already tell whether an agent is alive — the `agent-status`
KV bucket republishes once a minute — but not what it is doing right now.
A header bar wants the second thing, and a minute-old answer to "is this
agent thinking" is the wrong answer most of the time it is read.
`hive-agent` now publishes a turn-state header to
`$SWARM.agent-state.<hive>.<agent>`, a core subject beside the terminal
rows it already sends. It goes out **on transition, not on a timer**: the
publisher watches the event bus, rebuilds the header, and sends only when
the serialised result differs from the last one it sent — so a second
periodic writer, which is the problem this exists to fix, is not what
replaces the bucket.
The payload is the published contract a swarm-level renderer is written
against, so the test asserts on the serialised JSON keys rather than on
Rust field names. Two fields deliberately depart from the per-agent web
UI's `StateSnapshot`: `turn_state_since` is an ISO 8601 UTC string rather
than unix seconds, matching the sibling `$SWARM.term` subject's stamp, and
`agent_state` carries the swarm's own `AgentState` vocabulary rather than
a `paused` boolean, so a reader can compare actual against wanted without
translating. `turn_state` and `agent_state` stay two separate fields:
neither vocabulary contains the other's values.
Swarm-side, `GET /api/agents/{name}/state/stream` relays the subject as
SSE, resolving the agent's hive at request time exactly as the terminal
stream does and passing the bytes through without parsing them.
The broker grant is a second `--agent-publish-subject` rather than a
widening of the existing one, so the terminal family and the header family
stay independently revocable, and a `module-eval` arm pins the rendered
flag and its argument together — the doubled dollar included, since a
single one expands to nothing in `ExecStart` and yields a grant that
matches nothing.
Refs #3802
hive-agent already publishes classified TermMsg rows to the core NATS
subject $SWARM.term.{hive}.{agent} -- live only, no retention, by
design. This endpoint subscribes that subject per request and relays
each row over SSE, opaque to this daemon (no TermMsg dependency, same
pass-through shape crate::status already uses for hive snapshots).
No replay/history: the publish side never grew JetStream retention, and
this route's own job (a live tail) never needed it.
The read policy and cert-auth role for each hive were written once, at
startup. On the deploy that surfaced this, the store was still coming
up, the pass logged its warning and moved on, and no hive could log in
until someone restarted the daemon — while cert auth answered "no chain
matching all constraints", which reads like a certificate problem
rather than a role that was never created.
The bootstrap unit in swarm-bao.nix lost the same race and won on its
retry 30s later. A daemon that boots alongside its store loses that race
routinely; on a normal boot it is the ordinary case.
The two passes fold into one `provision()` that logs in once instead of
twice for two loops over the same list, keeping policy before role since
the role names the policy. `ensure_hive_access` still awaits the first
pass, so a store that is already up leaves nothing deferred, and only a
pass that could not reach the store at all spawns the retry.
The retry is `config_pr::spawn`'s idiom from this same crate: an
interval task whose first tick is immediate. Its cadence and bound match
the bootstrap unit's — 30s, ~a day — because the two halves of one race
should not disagree about how long a wait is worth.
`Error::MissingEnv` is what keeps it from spinning forever: no `BAO_*`
set means a deployment that runs no store, where asking again changes
nothing, so it returns Ok. Everything else is retryable, including an
authority file that is not placed yet — the unit that writes it starts
alongside this one. Both cases previously landed in the same "not
managed here" line, so a store that was late looked exactly like one
that was never configured.
Per-hive failures keep their old behaviour: logged, skipped, Ok. A store
that refuses one hive's write refuses it again, so the next start really
is the right retry for those, and the module doc still says so.
Closes#4176.
A hive holds an mTLS pair and a policy naming what it may read, and still
cannot log in: nothing creates the role that maps its certificate to that
policy. The one pre-shared credential in the system therefore buys no
access.
Minting happens here rather than in nix, which was the first plan. Nix
mints from the store's own container, and that path is gated on the
bootstrap token -- so onboarding a hive later would mean placing the one
genuinely pre-shared secret again. Doing it from the controller costs a
public certificate authority as an input and makes the bootstrap token
one-time.
A startup pass, not a hook: the hive list is loaded once and a config
change means a redeploy, so the roles are as static as the list. Only the
policy is derived from something that moves.
The subject is the hive's name because glue-bao-tls.nix mints a hive's
client leaf with its name as the CN, and cert auth matches on that.
Per-hive failures are logged and skipped, matching the queue, bridge and
forge connects above it: a controller whose store is unreachable still
serves everything else, and the next start retries.
Not covered by a test: ensure_hive_roles is IO from end to end, and the
seam that would make it assertable is the one the read-grant sink already
has. Said here rather than implied by a green suite.
A hive reads its agents' credentials with its own certificate, and nothing
said which paths that certificate may read, so the read half of a delivery
answered 403.
The grant is wide on purpose. An agent's path does not name the hive
hosting it -- agents move -- so a per-hive grant has to be an enumeration
the controller re-emits whenever the roster changes, and an enumeration
that can drift or land out of order advertises a boundary it does not
hold. A wide grant that says what it is beats a narrow one that only looks
narrow. mara's call, on the PR: rather a too-lax scope than one that
pretends to be strict.
What that buys, beyond honesty: the document is identical for every hive
and depends on nothing, so it is written once at startup beside the rest of
a hive's provisioning instead of on every declaration. No derived state, no
re-emission, and the ordering hazard that came with one stops existing.
What still holds is read-only. A hive cannot write an agent's credential,
so it cannot hand itself an agent's identity, and the grant reaches nothing
in the store outside the agent-credential prefix.
The fact is documented where someone meets the boundary rather than only in
this message, and the two ways to narrow it later -- scope per hive, or
give agents their own store identity -- are tracked.
The controller writes an agent's credential; the hive fetches it back with
its own certificate. Nothing said which paths that certificate may read, so
the read half of a delivery answers 403 with no way to tell why.
The grant is derived from the declaration, so it is re-rendered at the one
place the declaration changes -- WantedWriter::set -- rather than at its
caller, which would work today and break on the second caller.
Emitted before the KV write: a grant that lands late is a 403 on an agent's
first fetch, while one that shrinks early only affects an agent already
being torn down. A failed write then leaves a superset the next declaration
re-renders.
Destroyed agents are filtered out. The declared set is a hive's whole
history -- a destroyed entry stays so that redeclaring it Up is refused as
the terminal transition it is -- so granting every declared agent would
leave a torn-down agent's credentials readable forever.
The sink is a trait because a missed emission is that same untraceable 403:
the double pins which agents were published, and the no-sink and refusing
arms pin the two deployments that are not a happy path. Not covered: the
call site inside set(), which needs a live queue.
The cert role moves to its own module on the way past. It is the
controller's identity at the store, not something the matrix route owns,
and the policy writer needs the same login.
The swarm UI had nowhere to POST an external matrix account to: this daemon
had no matrix-account code at all and no `swarm-secret-client` dependency, so
the last leg of #3726 — a credential reaching an agent — had no entry point.
`PUT /api/hives/{hive}/agents/{agent}/matrix-accounts/{account}` writes the
credential to the store under the agent's own path and publishes a
`CredentialNotice` on that hive's credential subject. All three path names are
load-bearing: agent + account locate the secret, hive routes the notice. The
account is a path segment rather than a body field so that splitting the 1:1
account-to-agent mapping later is a new route, not a changed payload.
Store first, notify second, and the order cannot be swapped: a notice that
overtakes its own write reaches a hive that reads nothing, and the hive
deliberately does not retry. The publish is followed by a flush for the reason
`publish_deploy` flushes — `publish` hands the message to the connection's
write buffer and returns, so the response could otherwise outrun the notice it
reports as sent.
The store client is built per request rather than held in `AppState`, matching
what the hive side does inside `deliver`: a login that expires is not worth
caching for a route this cold.
`swarm_hive` is `declaration_target`'s two name checks, extracted so this
handler makes them identically rather than in a second copy free to drift.
`declaration_target` still tests the writer first, so a deployment with no
queue answers 503 whatever the caller spelled.
## The nix half
#4081 minted the controller's leaf and gave it `baoClientCertFile` /
`baoClientKeyFile`, deliberately stopping there — the leaf is minted whether or
not a controller runs on that host. Nothing consumed those options, so the
identity never reached the process. Measured before writing: `git grep
baoClientCertFile` returned 5 sites and zero consumers, against a control
(`tokenEndpoint`, 4 hits in the same file) proving the search can see
consumption where it exists.
The unit now gets `BAO_ADDR` / `BAO_CLIENT_CERT` / `BAO_CLIENT_KEY` /
`BAO_CACERT` and the matching `LoadCredential` entries, following
`hive-c0re/environment.nix`'s `%d` credential shape.
The gate is `deploy.swarm-controller.baoClientCertFile`, NOT
`deploy.bao.clientCertFile`. The latter is the hive reader's identity and its
policy scopes a hive's own secrets; wiring it here would evaluate, deploy, and
fail only when the daemon tried to write an agent's credential.
Two `module-eval` arms cover exactly that. The presence arm asserts the
`LoadCredential` *source path* (`…:/var/lib/swarm-bao-pki/controller.pem`) and
not just the `%d` name, because a `%d`-only assertion passes while the daemon
holds the wrong policy. The absence arm (`controllerNoStore`) is what makes the
presence arm mean anything.
`RestrictAddressFamilies` already covers the store client; its own comment asks
for the family to be added with the client, and AF_INET/AF_INET6 are present.
Contributes to #3726
hyperhive#3896. The backend for start/stop (Up/Offline wanted-state
declarations) already existed and was merged (#3905's writer, the
PUT /api/hives/{hive}/agents/{agent}/state route) — nothing here was
waiting on Paused/Destroyed, which I'd mistakenly conflated with this
issue in an earlier comment (that's #3803, a different feature).
swarm-controller: merges each row's declared wanted state into
GET /api/agents/status, same shape as the config_pr merge (one read
per distinct hive, not per agent, since a declaration is a hive's
whole agent map).
swarm-ui: AgentsPage gets a "wanted" column — clicking the current-
state badge toggles it (Badge's own chip-plus-control shape, same as
its own header comment's pause/resume example), backed by the PUT
route above. A row with no declaration yet reads its implied current
state off the agent's own last-reported running flag. Stop asks for
confirmation (native window.confirm — no confirm-dialog component
exists in swarm-ui yet); start doesn't.
Per mara's review call on this PR: "the view should be filled by a single
backend call." AgentsPage.tsx was doing three fetches (/api/agents,
/api/config-prs, /api/agents/status) and joining them client-side by name.
Moves the config-PR join server-side instead: AgentStatusRow gains a
config_pr field, populated by get_agents_status's handler from
AppState::config_prs after agent_status::AgentStatusReader::view() returns
- not inside that module, which has no forge client and stays that way (see
the field's doc comment for why the handler is the right layer for this
merge, not the reader).
AgentsPage.tsx now does exactly one fetch and no client-side joining at all
- the wire row is the table row. Dropped the separate AgentStatusRow TS
interface (folded into AgentRow, which now mirrors the backend type
field-for-field) and the /api/agents + /api/config-prs fetches entirely;
neither is needed once /api/agents/status already returns every roster
agent with its config PR attached.
ConfigPrStatus gained Deserialize (previously Serialize-only) since
AgentStatusRow derives both and a struct's derive requires every field to
support it.
The hive-side loop landed without anything to converge to: nothing wrote
`$KV.hive-wanted.<hive>`, so in production only the "no key" branch ran.
This is the writer.
`WantedWriter` mirrors `StatusReader` — that module reads what hives report,
this one writes what they are told, so it holds a client rather than a bucket
handle and resolves the store on first use. It shares the status reader's
connection: the controller has exactly one by design, and a second connect
would double the auth-callout traffic and give the two paths independent
reconnect state.
The value under a hive's key is the map of every agent on that hive, so a
plain `put` of a single-agent change would drop a concurrent change to a
different agent, with only one revision of history to not recover from.
Writes are read-modify-write against the entry revision, and only
`WrongLastRevision` / `AlreadyExists` count as a lost race — every other
error returns immediately rather than spinning the retry loop and then
blaming a concurrent writer that never existed.
`apply` is split out and tested because it holds the invariant: declaring
one agent preserves the rest, and a current value that will not decode is an
error rather than a fresh start. Overwriting a document nobody can read
discards every other agent's declaration.
Two routes, no swarmctl verb and no jobq node: `create_agent` needs a graph
because it is multi-step, and one CAS'd write is not.
`build_app` is extracted from `main` in the same change because `main` sat at
exactly the `too_many_lines` limit, so adding an endpoint tripped a lint
about the startup sequence. The route list is the part that grows.