Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph

4,957 commits

Author SHA1 Message Date
atlas
b0e26e7e44 hive-runtime: shared runtime crate with claude and acp backends
A `Runtime` trait (run / compact / archive) with two backends:

- claude: a pass-through to hive_claude's InfiniteSession and
  SessionStore, so a claude turn is the same spawn, session handling
  and errors as before.
- acp: a generic Agent Client Protocol client. It spawns the command,
  args and env from RuntimeSpec (HIVE_RUNTIME / HIVE_ACP_COMMAND /
  HIVE_ACP_ARGS / HIVE_ACP_ENV), refuses an agent whose
  mcpCapabilities.http is not true, passes the claude --mcp-config
  servers as ACP mcpServers, keeps one session id in a file
  (session/load after a restart, session/new otherwise), and maps
  session/update into claude stream-json events plus usage_update into
  Telemetry. Permission requests are answered by a caller-supplied
  policy on the ACP tool kind. compact returns Unsupported for now.

The crate depends on no hyperhive binary crate, so the subagent
daemon can move onto it without pulling in hive-agent.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
d98d407bf1 docs: fix upgrade-block bullets for the store-first sweep
`ensure_hive_user` no longer short-circuits on the token file it
already has, and a hive's per-hive store token now wins over the file
whenever it's present. Both upgrade-guide bullets described the old
file-first order and needed rewording to match.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
78d8d69c7f swarm-controller: mint each hive's matrix sender token
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.

swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.

swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.

hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.

The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
9a5a947f8d swarm-controller: fix broken rustdoc intra-doc link to AGENT_TOKEN_NAME
legacy_tokens.rs referenced [`AGENT_TOKEN_NAME`] unqualified, but the
const lives in the sibling agent_token module and isn't in scope here;
rustdoc's broken-intra-doc-links lint (denied) failed the docs build.
2026-09-29 22:13:32 +02:00
atlas
e9206505d4 swarm-controller: retry the legacy forge token sweep from the mint pass
The sweep ran once at startup and was never retried: a boot where
store::connect or the roster read failed (e.g. the controller up
before bao) left the tokens live until the next restart. It now runs
off forge::agent_token::spawn's five-minute mint pass, reusing that
pass's roster observation instead of a second store/roster read, so a
failed first tick retries on the next one. Idempotent, so a re-run
after a partial sweep deletes nothing extra. Drops the one-shot
startup spawn; one call path.

Also updates the forge.md and agent_token.rs docs that still said the
legacy hyperhive-<seconds> tokens stay until manually removed.

Refs #4644
2026-09-29 22:13:32 +02:00
atlas
23e0c313b8 swarm-controller: sweep agents' legacy hyperhive-* forge tokens at start
Before b5d07d4d, hive-c0re minted a new `hyperhive-<unix-seconds>` token
for an agent on every spawn and rebuild and never revoked one, so every
live agent's forge user carries a pile of write-scoped tokens nothing
holds. Nothing in the tree lists or deletes them.

On each start, swarm-controller now walks the store's hive-agent-*
roster (the one the swarm-agent mint pass walks), and for every agent
whose swarm-agent token that pass would keep, deletes each token named
exactly `hyperhive-<digits>`. It logs the count per agent and a total.

- `core` is refused by name in both the roster filter and the per-user
  delete: hive-c0re still names core's live admin token
  `hyperhive-<unix-seconds>`.
- An agent whose swarm-agent token is not current is skipped, because
  consumers fall back to `<state>/forge-token`, the last hyperhive-*
  token, until the swarm token is fetched.
- A failed list or delete is logged and skipped; the sweep does not
  retry and never blocks startup. A second start deletes nothing.

Tokens on forge users of agents no longer on the store roster (already
destroyed) are not reached.

Closes #4644
2026-09-29 22:13:32 +02:00
atlas
41d66e2e05 swarm-grafana: provision alert rules for forge reconcile + swarm terminal publish failures
Two Grafana-managed rules in a "hyperhive" folder, one group evaluated
every 1m, each a LogsQL stats count over the VictoriaLogs datasource:

- forge reconcile failing: `swarm forge objects: write failed; retrying
  next pass` over 10m, for 15m. The pass runs every 5m, so one failed
  pass stays under `for` and a pass failing on every tick fires.
- swarm terminal publish skipped: `swarm terminal: publish skipped` over
  15m, for 5m, one instance per hive/agent.

No contact point or notification policy. Grafana 13 routes to its
built-in `empty` receiver when none is provisioned, so firing rules
are visible under Alerting and sent nowhere, with no send errors.

Refs #4717
2026-09-29 22:12:28 +02:00
atlas
10d579ecaf nix: move the remaining service containers onto the swarm-container module
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.

hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.

Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.

module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.

Refs #3773
2026-09-29 20:17:28 +02:00
atlas
53d9ebdede nix: list nix/container-modules/ among the module trees
Refs #3773
2026-09-29 19:51:46 +02:00
atlas
cce41c1d79 nix: move the in-container modules to nix/container-modules/
Refs #3773
2026-09-29 19:51:46 +02:00
atlas
024067f3f8 nix: share the service-container settings through one in-container module
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.

`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).

swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.

Refs #3773
2026-09-29 19:51:46 +02:00
atlas
544a8dd228 swarm-controller: check every first declaration, and fail closed on the roster
Exempting offline and paused from the name rules let a reserved name be
placed anyway: a first `offline` declared it, and the `up` after it
passed as an agent already declared on the hive. `paused` alone sufficed,
since the hive deploys an absent agent declared paused. A first
declaration in any placing state now runs the name rules and the
placed-elsewhere check; an agent already declared on the hive skips both.

A rule-breaking name whose roster read fails, including when no identity
bridge is configured, was accepted with a warning nobody sees, as in
create_agent. No later step on this route refuses the name, so it now
refuses with 503.

The gate is taken only for a placing declaration, so a destroy and its
credential revocation no longer wait on creations.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
d3755603f4 swarm-controller: check only a first declaration, and never a stop's name
The state route ran the name rules and the placed-elsewhere check on
every up/offline/paused declaration. An agent already declared on the
hive whose name breaks a rule, and which is not in the roster, could then
only be destroyed from the swarm.

Now an agent this hive already declares in a placing state passes both
checks. For a first declaration the placed-elsewhere check still applies
to up, offline and paused alike, and the name rules to up only: offline
and paused are how an operator stops an agent. The route reads this hive's
declaration to tell, and refuses with 503/500 when it cannot.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
8b892f508b swarm-controller: refuse a state declaration create_agent would refuse
PUT /api/hives/{hive}/agents/{agent}/state wrote any identifier into any
hive's wanted state, and the hive first-deploys a declared agent it has
no container for. That skipped create_agent's checks on the name.

A declaration that places the agent (up, paused, offline) is now refused
with 400 for a new name that breaks a naming rule, and with 409 for a
name the swarm has placed on another hive. Both reuse create_agent's
helpers (broken_name_rules/name_verdict, placements_elsewhere), under the
same gate, and an unreadable wanted state on another hive refuses with
503/500 as creation does. `destroyed` places nothing and is not checked,
so the revocation path is unchanged.

An agent with no declaration at all is still accepted: swarm-ui declares
state for agents that predate swarm-level creation, which is how the swarm
adopts them. Whether to refuse such names instead is left open on #4804.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
8fb7751da9 swarm-bao: don't let the viewer restart enqueue fail the granter
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
2026-09-29 18:09:53 +02:00
atlas
22c96282b9 swarm-bao: re-run the operator viewer unit when the granter step succeeds
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.

The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.

OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.

The viewer's log message no longer tells the operator to restart it.

Refs #4772
2026-09-29 17:45:58 +02:00
atlas
efed43b672 swarm-controller: read queued placements before published ones
create_agent read the other hives' wanted state first and snapshotted the
queued SetAgentWanted nodes second. A node finishing between the two was
in neither — not yet declared at the first read, already terminal at the
second — so a second hive could get the same name.

The queue is now snapshotted first: a SetAgentWanted node only turns
terminal after its declaration is written, so anything terminal by then is
visible to the read that follows. The order lives in
`placements_elsewhere`, and a test that finishes a node between the two
reads fails with them swapped.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
5785c0024c Make agent creation swarm-only and refuse a name placed on another hive
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.

Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.

policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
1d8ec00ddc hive-priv: remove the RestartMatrixDaemon command
Its only client was hive-c0re's matrix-account-login handler, removed
earlier on this branch, so nothing sends `restart_matrix_daemon` any
more. Drops the `PrivRequest` variant, the hive-priv handler and its
`systemctl --machine=h-<agent> restart hive-matrix-daemon.service`
helper, and the row in the hive-priv op table in security.md.

`PrivRequest` is internally tagged by `op` name, so no other variant's
encoding changes. A new token file still re-fires the daemon through
its `matrix-token*` path unit.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
f6ac007dc0 Drop the last references to the removed hive matrix login form
`PutMatrixAccountRequest`'s no-`Debug` comment named the deleted
`MatrixLoginForm` as its precedent; it now states the reason directly.
The daemon's missing-sidecar warning told operators to re-login via the
dashboard, which no longer has that form; it now points at re-linking
the account from the swarm UI, whose stored credential carries the
homeserver the sidecar is written from. Behaviour is unchanged.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
6a1d85c24f hive-dashboard: remove the MATRIX credentials tab and its login route
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.

Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.

The GITHUB tab is the credentials page's default tab now.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
93bbec015f hivectl, hive-c0re: remove dead matrix create-user/promote-user/reset-password
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.

Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.

Refs #4585
2026-09-29 13:13:59 +02:00
atlas
69ae23f801 swarmctl: add user reset-password
authelia's file backend has no self-service reset (no SMTP notifier),
so the only way a human account got a new password after the old one
was forgotten was hand-editing users.yml as root. `user add` already
hashes a password into the file; this verb does the same for an
existing user instead of refusing on the name.

Mirrors `user add`'s UX exactly: no password flag, authelia generates
and hashes it (never crosses argv), and it's printed once and never
stored. Refuses on an unknown user before ever invoking authelia. Same
publish path as add/update, so the same atomic write and no-restart
(authelia watches the file) behaviour apply.

Split the digest-replacement into users::reset_password so it's
testable without a command line or a running authelia, same pattern
as apply_update.
2026-09-29 12:31:21 +02:00
atlas
ccb5bd3b38 hive-agent: read the per-agent queue secret from bao in process
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.

hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.

A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).

Closes #4783
2026-09-29 10:18:07 +02:00
atlas
f7437a4773 dashboard: log a warning when tombstone purge fails to fail pending approvals
post_purge_tombstone discarded fail_pending_for_agent's error with
let _ =. Mirrors #4740's fix for the identical discard in
job_queue/exec.rs's run_destroy_bookkeeping: warn and continue, since
the purge itself has already succeeded by this point.

Refs #4747
2026-09-29 09:37:21 +02:00
atlas
c04cabdf80 hive-c0re: warn on a du failure that isn't just a missing path
du_bytes returned None uniformly for every du failure, and both
du_bytes and measure_agent_disk had no logging at all, so a container
rootfs du couldn't read (permission, I/O error, unparseable output)
silently reported as 0 disk with no signal in the journal.

Distinguish the expected case — the path plain doesn't exist, e.g. a
destroyed-but-kept agent's rootfs or a not-yet-created state dir — from
an actual read failure, and warn only on the latter, with the path plus
whichever of the failure (spawn error, exit status, stderr, unparseable
stdout) applies.
2026-09-29 09:37:16 +02:00
atlas
535073011b hive-forge: bound the web-router client's connect and request waits
The raw reqwest client hive-forge uses for Forgejo's web-router-only
routes (attachment downloads, Actions artifact/log routes) had no
timeout, so a hung Forgejo response blocked the calling CLI invocation
indefinitely.

Adds a 5s connect_timeout (matching #4737's outbound-HTTP sites) plus
a per-call request timeout: 15s (config_pr_poll's forge budget) for
the small JSON calls (get_api_json, post_json_web), and 10 minutes for
get_bytes_named/get_bytes_raw, which download attachments, Actions
artifact zips and persisted job logs that can be large.
reqwest::blocking has no separate read/idle timeout, so a single
whole-request budget has to cover those downloads; the smaller JSON
budget would cut them off partway through.

Refs #4746, #4737
2026-09-29 09:17:47 +02:00
atlas
202f7f7c83 checks: run the agent bao-fetch unit scripts against a stub bao
`script-test-agent-bao-fetch` executes the rendered `ExecStart` of
`hive-agent-forge-token` and `hive-agent-queue-credential` under
`umask 0377`, with a stub `bao` first on the unit's own PATH. It covers
every error branch, the happy path, forge rotation/unchanged, and 0400
files already in place: the redirect failure #4736 fixed, whose live
symptom was `bao.err: Permission denied` reported as a refused
certificate. The TLS-alert fixture is the `unknown certificate
authority` error h-atlas's identity check got from the store.

#4736 claimed this test but never committed it; the two module-eval
suites for these units only evaluate the config.

Closes #4748
2026-09-29 09:15:31 +02:00
atlas
916c441e82 swarm-otel: bound the collector's restart backoff
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.

Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
2026-09-29 09:15:12 +02:00
flake-bot
375123f5d2 nix flake update 2026-09-29 05:01:16 +02:00
atlas
5409f8160b hive-c0re: bound prebuild_toplevel's nix build with a timeout
prebuild_toplevel awaited nix build with a bare child.wait().await, so
a wedged nix-daemon (unreachable remote builder, stuck build slot)
hung the rebuild job forever with no way for the job queue to recover
short of restarting hive-c0re.

Mirrors #4741's hive-priv fix: the child now leads its own process
group, and after PREBUILD_TIMEOUT (1h, four times CI's observed
cold-cache flake check) the whole group is SIGKILLed and the call
fails with a named timeout error instead of hanging.

Refs #4723, #4741
2026-09-29 02:04:53 +02:00
atlas
642be57678 swarm: narrow the revoke grant to the queue leaf, not the whole agent prefix
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.

Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
2026-09-28 23:46:17 +02:00
atlas
215a8aedc4 swarm-controller: link AgentState by path in the revocation doc comment
rustdoc runs with -D rustdoc::broken-intra-doc-links and the type is
imported inside the function body, not at module scope, so the bare
link resolved to nothing and failed the workspace-doc derivation.
2026-09-28 23:46:17 +02:00
atlas
c21ec7719d swarm: tests and docs for the queue-credential revocation
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.

docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.
2026-09-28 23:46:17 +02:00
atlas
8caf688ee4 swarm: revoke an agent's queue credential when it is declared destroyed
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.

Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.

Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.

The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
2026-09-28 23:46:17 +02:00
iris
88c386c96c swarm-ui: move dynamic agent-term tabs into the top nav bar
mara, PR review: "the dynamic tab should be in the top bar, not a new
one below". Moves the .shell-tabs group from its own sticky row under
the header into .shell-nav itself, right after the nav indicator.

Also adds a third re-measure effect for the sliding nav indicator,
keyed on tabs.length: with tabs inline in the same flex row the
indicator measures, closing a background tab (no navigation) can
shrink the row without the hop effect's own re-measure ever firing.
Same reflow-not-navigation reasoning as the existing resize-listener
effect.
2026-09-28 23:17:44 +02:00
iris
7bf4cfbc4c swarm-ui: fix specificity claim in AgentTermPreview.css comment
argus's review on PR #4784 caught a wrong technical claim: the
comment said the .ui-agent-term-preview-full override rules relied
on source order because they had equal specificity to the
un-modified rules above. They don't — each override selector adds
one more class (the .ui-agent-term-preview-full prefix) than what it
overrides, so they're strictly more specific and win regardless of
file order. Corrected the comment to say so.
2026-09-28 23:17:44 +02:00
iris
c460e91bd1 swarm-ui: full-tab agent terminal (#4506)
Adds a full, non-capped agent terminal reachable from a new expand
trigger on the embedded AgentTermPreview (the detail-panel preview on
AgentsPage stays as-is, just gains the trigger). Opens
/agents/:name/terminal in a new dynamic tab in Shell's header, next to
the static nav row — tabs persist across a reload via useDynamicTabs,
a small localStorage-backed hook built on @hive/shared's existing
settings-storage primitive.

AgentTermPreview gains two new props to support both mounts from one
component: fullHeight (drops the 12em preview cap, fills its page)
and showHeaderBadges (default true — lets a future caller that
already shows turn_state/model/ctx/cost elsewhere suppress this
cluster; AgentsPage doesn't use it, see below).

Deviation from the originally posted plan (issue comment 80596): that
plan proposed AgentsPage's embedded preview pass showHeaderBadges as
false, reasoning the detail panel already duplicates that info.
Checked the actual code before implementing — it doesn't; AgentRow/
AgentTypes.ts carry none of turn_state/model/ctx/cost, and
AgentTermPreview's own floating badges are the only place swarm-ui
shows them. Left the badges visible there instead of shipping a
regression the plan's own stated justification didn't hold up to.

Also fixed a same-tab pub/sub race found by actually rendering a cold
load of /agents/:name/terminal (headless chromium, not just reasoning
about the code): useLocalSetting subscribes inside a useEffect, and
mount effects fire children-before-parents, so a descendant's
mount-time write (AgentTerminalPage registering its own tab) can beat
an ancestor's (Shell's) subscription into existence, leaving Shell's
tab row silently empty on a direct/reload load. Fixed by having
useDynamicTabs re-sync from storage on every location change, not
just on notify() — the fix lives in the new hook itself, not in the
shared settings-storage primitive theme/motion overrides also use.
2026-09-28 23:17:44 +02:00
atlas
433429ebfd bao: disable the unused approle auth method, declaratively
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.

Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.

nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
2026-09-28 22:58:36 +02:00
atlas
7bc4b25f16 swarm-controller: backfill a missing queue-secret mint time instead of re-minting it
A stored queue secret with no `minted_at` counted as due, and no secret
minted before the renewal pass has one, so the first pass after deploy
would re-mint every agent's secret. Every reconnect before that agent's
next restart would then be refused.

Such a secret is now stamped instead: `minted_at = now` is written beside
the unchanged `value`, and its 45-day clock starts there. Only a secret
whose recorded mint time is at least 45 days old gets a new value.

The decision is `secret_step` (Keep / Backfill / Remint), and the pass
reports an unstamped secret as `Observed::Unstamped`. Backfill and
re-mint log different lines.
2026-09-28 22:24:07 +02:00
atlas
2115ec2bb3 swarm-controller: re-issue agent certificates and re-mint queue secrets at half-life
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:

- `MintAgentIdentity` (the node agent creation uses) when the stored
  certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
  read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
  swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
  re-decides, writes a fresh value with `minted_at`, reads it back, and logs
  the agent and the old age.

When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.

`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.

Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.

Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.

docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
2026-09-28 21:35:01 +02:00
atlas
b14ff2796c bao: OIDC login to the browser UI via authelia, as a metadata-only viewer
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.

- authelia registers an interactive client `swarm-bao-ui`
  (glue-bao-ui-oidc-client.nix) with redirect
  `https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
  publisher carries its secret to
  `secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
  mount with listing visibility `unauth`, asked before attempted like
  cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
  and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
  policy, and once the granter may configure `auth/oidc/config` (checked
  through `sys/capabilities-self`), writes the mount's config from the
  published secret and the role binding `groups=admins` to the viewer.
  Before the bootstrap step re-runs it writes the policy, logs the step
  and exits 0.

Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.

module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
2026-09-28 19:56:38 +02:00
atlas
3b27a2e5a1 docs: fix vale prose-lint errors in bao UI docs
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Passive voice and sentence-initial 'So' hits from PR #4776's vale error job.
2026-09-28 19:31:02 +02:00
atlas
3898ca33c7 bao: serve the browser UI to admins via a loopback-only listener
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.

The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.

Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.

Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.

On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
2026-09-28 19:31:02 +02:00
atlas
a86379eac5 ci: internal jobs skip on the public forge instead of waiting for a hive-ci runner
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
A job-level if: is evaluated by whichever runner picks the job up. The
public forge's copy has only a nixos runner, so a job guarded by
if: vars.PUBLIC_FORGE != 'true' and pinned to runs-on: [hive-ci] never
gets picked up there to evaluate the guard at all - it just sits queued.

runs-on now carries the same PUBLIC_FORGE variable: hive-ci internally,
nixos on the public copy. There the nixos runner picks the job up,
evaluates the existing if:, and skips it immediately. Internally
PUBLIC_FORGE is unset, so runs-on still resolves to hive-ci and nothing
changes.
2026-09-28 19:28:23 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
97c4771514 ci: run nothing but a main-only bin-cache push on the public forge
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Guards internal-only jobs (ci.yml, coverage.yml, flake-update.yml) with
`vars.PUBLIC_CACHE != 'true'` and adds public-cache.yml, guarded to the
inverse, to build the deployed closures and push them to the public
attic cache on a push to main.

`PUBLIC_CACHE` is a repo Actions variable, opt-in only on the public
copy: unset here, it leaves internal CI's `!=` guards true so internal
jobs always run. Job-level `if:` cannot see the `github`/`forgejo`
context at all on this runner (confirmed empirically — a
`github.server_url` comparison always evaluates false at job level,
though the identical comparison resolves correctly inside a step),
so `vars.*`, which is available at job level, is the only usable
opt-in signal here.
2026-09-28 17:24:06 +02:00
iris
9a444c27fa agent ui: add a clickable new-session menu entry
Same gap #4611 fixed for logout: the Preact rewrite's StatusChips menu
never picked up new-session as a click path, only /new-session typed
twice into the terminal. postNewSession already existed in
termActions.ts. Wire it into the status menu with the same arm-then-
confirm click pattern logout/cancel-turn already use.

Closes #4612
2026-09-28 13:53:58 +02:00
atlas
3b8df27dcf swarm-ui: show each agent's icon on its card
Each agent card now leads with the agent's icon, loaded as an `<img>`
from `GET /api/agents/<name>/icon`: the same 5em square, background and
fallback as the hive dashboard's container row. An agent with no icon
(the route's 404), or any other failed load, shows the dimmed hyperhive
mark (`/favicon.svg`) instead of a broken image.

Only ever an `<img>`, never inline markup: the body is an agent-authored
SVG, and an image load does not run its script.
2026-09-28 13:47:37 +02:00
atlas
e974194e3a swarm: let an agent publish its own icon
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.

hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.

swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.

Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
2026-09-28 13:47:37 +02:00