Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph

4,948 commits

Author SHA1 Message Date
atlas
cce41c1d79 nix: move the in-container modules to nix/container-modules/
Refs #3773
2026-09-29 19:51:46 +02:00
atlas
024067f3f8 nix: share the service-container settings through one in-container module
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.

`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).

swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.

Refs #3773
2026-09-29 19:51:46 +02:00
atlas
544a8dd228 swarm-controller: check every first declaration, and fail closed on the roster
Exempting offline and paused from the name rules let a reserved name be
placed anyway: a first `offline` declared it, and the `up` after it
passed as an agent already declared on the hive. `paused` alone sufficed,
since the hive deploys an absent agent declared paused. A first
declaration in any placing state now runs the name rules and the
placed-elsewhere check; an agent already declared on the hive skips both.

A rule-breaking name whose roster read fails, including when no identity
bridge is configured, was accepted with a warning nobody sees, as in
create_agent. No later step on this route refuses the name, so it now
refuses with 503.

The gate is taken only for a placing declaration, so a destroy and its
credential revocation no longer wait on creations.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
d3755603f4 swarm-controller: check only a first declaration, and never a stop's name
The state route ran the name rules and the placed-elsewhere check on
every up/offline/paused declaration. An agent already declared on the
hive whose name breaks a rule, and which is not in the roster, could then
only be destroyed from the swarm.

Now an agent this hive already declares in a placing state passes both
checks. For a first declaration the placed-elsewhere check still applies
to up, offline and paused alike, and the name rules to up only: offline
and paused are how an operator stops an agent. The route reads this hive's
declaration to tell, and refuses with 503/500 when it cannot.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
8b892f508b swarm-controller: refuse a state declaration create_agent would refuse
PUT /api/hives/{hive}/agents/{agent}/state wrote any identifier into any
hive's wanted state, and the hive first-deploys a declared agent it has
no container for. That skipped create_agent's checks on the name.

A declaration that places the agent (up, paused, offline) is now refused
with 400 for a new name that breaks a naming rule, and with 409 for a
name the swarm has placed on another hive. Both reuse create_agent's
helpers (broken_name_rules/name_verdict, placements_elsewhere), under the
same gate, and an unreadable wanted state on another hive refuses with
503/500 as creation does. `destroyed` places nothing and is not checked,
so the revocation path is unchanged.

An agent with no declaration at all is still accepted: swarm-ui declares
state for agents that predate swarm-level creation, which is how the swarm
adopts them. Whether to refuse such names instead is left open on #4804.

Refs #4804
2026-09-29 19:16:13 +02:00
atlas
8fb7751da9 swarm-bao: don't let the viewer restart enqueue fail the granter
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
2026-09-29 18:09:53 +02:00
atlas
22c96282b9 swarm-bao: re-run the operator viewer unit when the granter step succeeds
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.

The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.

OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.

The viewer's log message no longer tells the operator to restart it.

Refs #4772
2026-09-29 17:45:58 +02:00
atlas
efed43b672 swarm-controller: read queued placements before published ones
create_agent read the other hives' wanted state first and snapshotted the
queued SetAgentWanted nodes second. A node finishing between the two was
in neither — not yet declared at the first read, already terminal at the
second — so a second hive could get the same name.

The queue is now snapshotted first: a SetAgentWanted node only turns
terminal after its declaration is written, so anything terminal by then is
visible to the read that follows. The order lives in
`placements_elsewhere`, and a test that finishes a node between the two
reads fails with them swapped.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
5785c0024c Make agent creation swarm-only and refuse a name placed on another hive
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.

Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.

policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
1d8ec00ddc hive-priv: remove the RestartMatrixDaemon command
Its only client was hive-c0re's matrix-account-login handler, removed
earlier on this branch, so nothing sends `restart_matrix_daemon` any
more. Drops the `PrivRequest` variant, the hive-priv handler and its
`systemctl --machine=h-<agent> restart hive-matrix-daemon.service`
helper, and the row in the hive-priv op table in security.md.

`PrivRequest` is internally tagged by `op` name, so no other variant's
encoding changes. A new token file still re-fires the daemon through
its `matrix-token*` path unit.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
f6ac007dc0 Drop the last references to the removed hive matrix login form
`PutMatrixAccountRequest`'s no-`Debug` comment named the deleted
`MatrixLoginForm` as its precedent; it now states the reason directly.
The daemon's missing-sidecar warning told operators to re-login via the
dashboard, which no longer has that form; it now points at re-linking
the account from the swarm UI, whose stored credential carries the
homeserver the sidecar is written from. Behaviour is unchanged.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
6a1d85c24f hive-dashboard: remove the MATRIX credentials tab and its login route
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.

Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.

The GITHUB tab is the credentials page's default tab now.

Refs #4348
2026-09-29 13:54:18 +02:00
atlas
93bbec015f hivectl, hive-c0re: remove dead matrix create-user/promote-user/reset-password
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.

Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.

Refs #4585
2026-09-29 13:13:59 +02:00
atlas
69ae23f801 swarmctl: add user reset-password
authelia's file backend has no self-service reset (no SMTP notifier),
so the only way a human account got a new password after the old one
was forgotten was hand-editing users.yml as root. `user add` already
hashes a password into the file; this verb does the same for an
existing user instead of refusing on the name.

Mirrors `user add`'s UX exactly: no password flag, authelia generates
and hashes it (never crosses argv), and it's printed once and never
stored. Refuses on an unknown user before ever invoking authelia. Same
publish path as add/update, so the same atomic write and no-restart
(authelia watches the file) behaviour apply.

Split the digest-replacement into users::reset_password so it's
testable without a command line or a running authelia, same pattern
as apply_update.
2026-09-29 12:31:21 +02:00
atlas
ccb5bd3b38 hive-agent: read the per-agent queue secret from bao in process
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.

hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.

A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).

Closes #4783
2026-09-29 10:18:07 +02:00
atlas
f7437a4773 dashboard: log a warning when tombstone purge fails to fail pending approvals
post_purge_tombstone discarded fail_pending_for_agent's error with
let _ =. Mirrors #4740's fix for the identical discard in
job_queue/exec.rs's run_destroy_bookkeeping: warn and continue, since
the purge itself has already succeeded by this point.

Refs #4747
2026-09-29 09:37:21 +02:00
atlas
c04cabdf80 hive-c0re: warn on a du failure that isn't just a missing path
du_bytes returned None uniformly for every du failure, and both
du_bytes and measure_agent_disk had no logging at all, so a container
rootfs du couldn't read (permission, I/O error, unparseable output)
silently reported as 0 disk with no signal in the journal.

Distinguish the expected case — the path plain doesn't exist, e.g. a
destroyed-but-kept agent's rootfs or a not-yet-created state dir — from
an actual read failure, and warn only on the latter, with the path plus
whichever of the failure (spawn error, exit status, stderr, unparseable
stdout) applies.
2026-09-29 09:37:16 +02:00
atlas
535073011b hive-forge: bound the web-router client's connect and request waits
The raw reqwest client hive-forge uses for Forgejo's web-router-only
routes (attachment downloads, Actions artifact/log routes) had no
timeout, so a hung Forgejo response blocked the calling CLI invocation
indefinitely.

Adds a 5s connect_timeout (matching #4737's outbound-HTTP sites) plus
a per-call request timeout: 15s (config_pr_poll's forge budget) for
the small JSON calls (get_api_json, post_json_web), and 10 minutes for
get_bytes_named/get_bytes_raw, which download attachments, Actions
artifact zips and persisted job logs that can be large.
reqwest::blocking has no separate read/idle timeout, so a single
whole-request budget has to cover those downloads; the smaller JSON
budget would cut them off partway through.

Refs #4746, #4737
2026-09-29 09:17:47 +02:00
atlas
202f7f7c83 checks: run the agent bao-fetch unit scripts against a stub bao
`script-test-agent-bao-fetch` executes the rendered `ExecStart` of
`hive-agent-forge-token` and `hive-agent-queue-credential` under
`umask 0377`, with a stub `bao` first on the unit's own PATH. It covers
every error branch, the happy path, forge rotation/unchanged, and 0400
files already in place: the redirect failure #4736 fixed, whose live
symptom was `bao.err: Permission denied` reported as a refused
certificate. The TLS-alert fixture is the `unknown certificate
authority` error h-atlas's identity check got from the store.

#4736 claimed this test but never committed it; the two module-eval
suites for these units only evaluate the config.

Closes #4748
2026-09-29 09:15:31 +02:00
atlas
916c441e82 swarm-otel: bound the collector's restart backoff
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.

Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
2026-09-29 09:15:12 +02:00
flake-bot
375123f5d2 nix flake update 2026-09-29 05:01:16 +02:00
atlas
5409f8160b hive-c0re: bound prebuild_toplevel's nix build with a timeout
prebuild_toplevel awaited nix build with a bare child.wait().await, so
a wedged nix-daemon (unreachable remote builder, stuck build slot)
hung the rebuild job forever with no way for the job queue to recover
short of restarting hive-c0re.

Mirrors #4741's hive-priv fix: the child now leads its own process
group, and after PREBUILD_TIMEOUT (1h, four times CI's observed
cold-cache flake check) the whole group is SIGKILLed and the call
fails with a named timeout error instead of hanging.

Refs #4723, #4741
2026-09-29 02:04:53 +02:00
atlas
642be57678 swarm: narrow the revoke grant to the queue leaf, not the whole agent prefix
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.

Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
2026-09-28 23:46:17 +02:00
atlas
215a8aedc4 swarm-controller: link AgentState by path in the revocation doc comment
rustdoc runs with -D rustdoc::broken-intra-doc-links and the type is
imported inside the function body, not at module scope, so the bare
link resolved to nothing and failed the workspace-doc derivation.
2026-09-28 23:46:17 +02:00
atlas
c21ec7719d swarm: tests and docs for the queue-credential revocation
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.

docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.
2026-09-28 23:46:17 +02:00
atlas
8caf688ee4 swarm: revoke an agent's queue credential when it is declared destroyed
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.

Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.

Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.

The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
2026-09-28 23:46:17 +02:00
iris
88c386c96c swarm-ui: move dynamic agent-term tabs into the top nav bar
mara, PR review: "the dynamic tab should be in the top bar, not a new
one below". Moves the .shell-tabs group from its own sticky row under
the header into .shell-nav itself, right after the nav indicator.

Also adds a third re-measure effect for the sliding nav indicator,
keyed on tabs.length: with tabs inline in the same flex row the
indicator measures, closing a background tab (no navigation) can
shrink the row without the hop effect's own re-measure ever firing.
Same reflow-not-navigation reasoning as the existing resize-listener
effect.
2026-09-28 23:17:44 +02:00
iris
7bf4cfbc4c swarm-ui: fix specificity claim in AgentTermPreview.css comment
argus's review on PR #4784 caught a wrong technical claim: the
comment said the .ui-agent-term-preview-full override rules relied
on source order because they had equal specificity to the
un-modified rules above. They don't — each override selector adds
one more class (the .ui-agent-term-preview-full prefix) than what it
overrides, so they're strictly more specific and win regardless of
file order. Corrected the comment to say so.
2026-09-28 23:17:44 +02:00
iris
c460e91bd1 swarm-ui: full-tab agent terminal (#4506)
Adds a full, non-capped agent terminal reachable from a new expand
trigger on the embedded AgentTermPreview (the detail-panel preview on
AgentsPage stays as-is, just gains the trigger). Opens
/agents/:name/terminal in a new dynamic tab in Shell's header, next to
the static nav row — tabs persist across a reload via useDynamicTabs,
a small localStorage-backed hook built on @hive/shared's existing
settings-storage primitive.

AgentTermPreview gains two new props to support both mounts from one
component: fullHeight (drops the 12em preview cap, fills its page)
and showHeaderBadges (default true — lets a future caller that
already shows turn_state/model/ctx/cost elsewhere suppress this
cluster; AgentsPage doesn't use it, see below).

Deviation from the originally posted plan (issue comment 80596): that
plan proposed AgentsPage's embedded preview pass showHeaderBadges as
false, reasoning the detail panel already duplicates that info.
Checked the actual code before implementing — it doesn't; AgentRow/
AgentTypes.ts carry none of turn_state/model/ctx/cost, and
AgentTermPreview's own floating badges are the only place swarm-ui
shows them. Left the badges visible there instead of shipping a
regression the plan's own stated justification didn't hold up to.

Also fixed a same-tab pub/sub race found by actually rendering a cold
load of /agents/:name/terminal (headless chromium, not just reasoning
about the code): useLocalSetting subscribes inside a useEffect, and
mount effects fire children-before-parents, so a descendant's
mount-time write (AgentTerminalPage registering its own tab) can beat
an ancestor's (Shell's) subscription into existence, leaving Shell's
tab row silently empty on a direct/reload load. Fixed by having
useDynamicTabs re-sync from storage on every location change, not
just on notify() — the fix lives in the new hook itself, not in the
shared settings-storage primitive theme/motion overrides also use.
2026-09-28 23:17:44 +02:00
atlas
433429ebfd bao: disable the unused approle auth method, declaratively
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.

Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.

nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
2026-09-28 22:58:36 +02:00
atlas
7bc4b25f16 swarm-controller: backfill a missing queue-secret mint time instead of re-minting it
A stored queue secret with no `minted_at` counted as due, and no secret
minted before the renewal pass has one, so the first pass after deploy
would re-mint every agent's secret. Every reconnect before that agent's
next restart would then be refused.

Such a secret is now stamped instead: `minted_at = now` is written beside
the unchanged `value`, and its 45-day clock starts there. Only a secret
whose recorded mint time is at least 45 days old gets a new value.

The decision is `secret_step` (Keep / Backfill / Remint), and the pass
reports an unstamped secret as `Observed::Unstamped`. Backfill and
re-mint log different lines.
2026-09-28 22:24:07 +02:00
atlas
2115ec2bb3 swarm-controller: re-issue agent certificates and re-mint queue secrets at half-life
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:

- `MintAgentIdentity` (the node agent creation uses) when the stored
  certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
  read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
  swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
  re-decides, writes a fresh value with `minted_at`, reads it back, and logs
  the agent and the old age.

When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.

`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.

Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.

Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.

docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
2026-09-28 21:35:01 +02:00
atlas
b14ff2796c bao: OIDC login to the browser UI via authelia, as a metadata-only viewer
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.

- authelia registers an interactive client `swarm-bao-ui`
  (glue-bao-ui-oidc-client.nix) with redirect
  `https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
  publisher carries its secret to
  `secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
  mount with listing visibility `unauth`, asked before attempted like
  cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
  and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
  policy, and once the granter may configure `auth/oidc/config` (checked
  through `sys/capabilities-self`), writes the mount's config from the
  published secret and the role binding `groups=admins` to the viewer.
  Before the bootstrap step re-runs it writes the policy, logs the step
  and exits 0.

Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.

module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
2026-09-28 19:56:38 +02:00
atlas
3b27a2e5a1 docs: fix vale prose-lint errors in bao UI docs
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Passive voice and sentence-initial 'So' hits from PR #4776's vale error job.
2026-09-28 19:31:02 +02:00
atlas
3898ca33c7 bao: serve the browser UI to admins via a loopback-only listener
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.

The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.

Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.

Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.

On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
2026-09-28 19:31:02 +02:00
atlas
a86379eac5 ci: internal jobs skip on the public forge instead of waiting for a hive-ci runner
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
A job-level if: is evaluated by whichever runner picks the job up. The
public forge's copy has only a nixos runner, so a job guarded by
if: vars.PUBLIC_FORGE != 'true' and pinned to runs-on: [hive-ci] never
gets picked up there to evaluate the guard at all - it just sits queued.

runs-on now carries the same PUBLIC_FORGE variable: hive-ci internally,
nixos on the public copy. There the nixos runner picks the job up,
evaluates the existing if:, and skips it immediately. Internally
PUBLIC_FORGE is unset, so runs-on still resolves to hive-ci and nothing
changes.
2026-09-28 19:28:23 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
97c4771514 ci: run nothing but a main-only bin-cache push on the public forge
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Guards internal-only jobs (ci.yml, coverage.yml, flake-update.yml) with
`vars.PUBLIC_CACHE != 'true'` and adds public-cache.yml, guarded to the
inverse, to build the deployed closures and push them to the public
attic cache on a push to main.

`PUBLIC_CACHE` is a repo Actions variable, opt-in only on the public
copy: unset here, it leaves internal CI's `!=` guards true so internal
jobs always run. Job-level `if:` cannot see the `github`/`forgejo`
context at all on this runner (confirmed empirically — a
`github.server_url` comparison always evaluates false at job level,
though the identical comparison resolves correctly inside a step),
so `vars.*`, which is available at job level, is the only usable
opt-in signal here.
2026-09-28 17:24:06 +02:00
iris
9a444c27fa agent ui: add a clickable new-session menu entry
Same gap #4611 fixed for logout: the Preact rewrite's StatusChips menu
never picked up new-session as a click path, only /new-session typed
twice into the terminal. postNewSession already existed in
termActions.ts. Wire it into the status menu with the same arm-then-
confirm click pattern logout/cancel-turn already use.

Closes #4612
2026-09-28 13:53:58 +02:00
atlas
3b8df27dcf swarm-ui: show each agent's icon on its card
Each agent card now leads with the agent's icon, loaded as an `<img>`
from `GET /api/agents/<name>/icon`: the same 5em square, background and
fallback as the hive dashboard's container row. An agent with no icon
(the route's 404), or any other failed load, shows the dimmed hyperhive
mark (`/favicon.svg`) instead of a broken image.

Only ever an `<img>`, never inline markup: the body is an agent-authored
SVG, and an image load does not run its script.
2026-09-28 13:47:37 +02:00
atlas
e974194e3a swarm: let an agent publish its own icon
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.

hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.

swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.

Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
2026-09-28 13:47:37 +02:00
atlas
4dbab024da swarm: drop the tracker tag from agent_icon's module doc
`scripts/check-issue-refs.sh` blocks a `#N`-shaped reference in source.
The quoted ruling reads identically without it, so the tag goes rather
than a lint suppression.
2026-09-28 13:47:37 +02:00
atlas
9513058a71 swarm: serve an agent's icon at swarm scope
An agent is not fixed to a hive, so its icon cannot be resolved as
hive -> agent. This adds the swarm-level half: an `agent-icons` KV
bucket keyed by the agent name alone — no hive token, so an agent that
moves hives keeps its icon and one that is stopped still has one — and
`GET /api/agents/<name>/icon` on swarm-controller serving it
same-origin, like every other `/api/*` route swarm-ui calls.

404 is the "this agent has no icon" answer, the same contract the
per-agent harness's own `GET /icon` has for an unconfigured agent.
Until the agent-side publisher lands, that is every agent's answer:
the publisher runs inside the container and an agent's NATS grants are
hive-scoped, which cannot authorise a write to a single-token agent
key. The read side needs no grant change — the controller already
holds `$KV.*.>` and `$JS.API.DIRECT.GET.*.>`.

The response carries `Content-Security-Policy: sandbox` and `nosniff`:
the body is an operator-authored SVG served from this daemon's own
origin, and an SVG can carry script.

Hive-side icon serving is untouched.

Refs #4502
2026-09-28 13:47:37 +02:00
iris
2e67656de8 docs: drop leftover 'nothing nests' framing from C0NTAINERS blurb
Per mara's review on #4769: don't extend stale info to say it's stale,
just remove it. The sentence only described the absence of the removed
topology tree, giving the operator nothing actionable.
2026-09-28 13:18:12 +02:00
iris
d9beb965a5 docs/web-ui: drop the redundant Topology tree section
C0NTAINERS already states the flat/alphabetical/no-nesting fact;
the section repeated it with zero new operator info. Fix the two
dangling references (README.md reading-path index, swarm.js
comment) rather than leaving them pointing at a removed heading.
2026-09-28 13:18:12 +02:00
iris
ef4e49805c docs/web-ui/dashboard.md: drop stale topology-tree history per mara review
state current fact only, not the old-vs-new narrative
2026-09-28 13:18:12 +02:00
iris
d447974536 docs/web-ui/dashboard.md: fix vale hits in topology-tree rewrite
- avoid 'backend' (Microsoft.Avoid)
- reword 'was removed' passive voice (write-good.Passive)
2026-09-28 13:18:12 +02:00
iris
6b5c8cc51a dashboard: flatten the agent list, drop dead topology-tree machinery (#4638)
The backend dropped the agent hierarchy's parent field, so every
container is a root and buildAgentTree/treePrefixDom could only ever
produce a single-level flat list — the .tree-prefix CSS lane rules
already matched nothing. Replaced with sortedContainerRows, a plain
alphabetical sort, and dropped the now-dead .tree-prefix/.tree-lane/
data-depth CSS and the depth/isLast/ancestorIsLast fields from the row
fingerprint and buildContainerLi's signature. No visible behavior
change - the rendered list was already flat, just via dead machinery.

Docs updated to describe the simpler implementation directly instead
of narrating the removal (kept the heading name since swarm.js still
points a comment at it).
2026-09-28 13:18:12 +02:00
iris
c9e12ffcc0 swarm-grafana: revert busiest-agents table, restore bargauges (#4658)
c9ba9062 replaced 5 bargauge panels with one table panel using a
joinByField+organize+sortBy transformations chain and a
vizConfig.group="table" viz spec. That construct had zero precedent
anywhere else in this repo's dashboards and its live render was
explicitly flagged as unverified at merge time (no in-container way
to render Grafana to confirm).

#4658 reports a dashboard crash: "'' not found in: reduce,
filterFieldsByName, ... transpose" - a Grafana transform-registry
lookup failing on an empty transformer id. That registry's full id
list matches the error text exactly, and the table panel is the only
panel in the whole swarm-grafana/dashboards tree using a non-empty
transformations array, so it's the leading suspect: some part of how
the dashboard-v2beta1 schema expects a QueryGroup's transformations
to be shaped likely differs from the classic {id, options} shape used
here, and the file-based provisioner may not be normalizing the
mismatch the same way a schema-v2-aware loader does.

Reverting to the prior, previously-uneventful bargauge panels
(byte-for-byte c9ba9062^'s version of this section) to stop the
active crash while the correct v2beta1 transformation shape gets
verified against a live Grafana render instead of guessed at again.
2026-09-28 11:52:33 +02:00
atlas
a2b4acfb6a swarm-bao: say "run the bootstrap step" when the bootstrap token is dead, not "sealed"
swarm-bao-granter-role used its token for `bao auth list` with no check,
so an expired, revoked or policy-less token in bootstrapTokenFile exited
2 with a raw 403 and no hint. It now checks whether the store is up when
that call fails: if it is, the token is at fault, and the unit prints the
one-time bootstrap step and exits 4. A missing token file is still a
ConditionPathExists skip, so the two read differently in the journal.

The twelve granterLogin units printed "the store is sealed or
unreachable" on a healthy store, because `bao status` exited 1 there:
the CLI resolves a token helper under $HOME before asking, systemd sets
no HOME for a unit without User=, and the fallback shells out to
`getent`/`sh`, neither of which is on the unit's PATH ("failed to get
token helper: error expanding config path "": exec: "sh": executable
file not found in $PATH"). The check now runs with HOME=/var/empty and
keeps its stderr, so a genuinely unreachable store says why. When the
store is up and the login is refused, the units now name
swarm-bao-granter-role as the unit that writes the missing role.

setup.md's post-step restart used 'swarm-bao-*-policy.service', which
misses swarm-bao-agent-pki. It now names that unit too, and a
module-eval case fails when the restart misses any unit that logs in as
the granter.

Refs #4704
2026-09-28 11:03:08 +02:00