AgentSession becomes hive_runtime::AgentRuntime, picked at startup from
the environment. Unset HIVE_RUNTIME keeps the claude backend, built
from the same title, store and PercentPolicy as before; drive_turn,
the 401 retry and the error mapping see the claude errors unchanged.
On acp the harness keeps the session id in the harness dir, answers
permission requests like the claude built-in allow-list (no built-in
shell, web fetch only with web_tools), maps runtime errors to Failed,
and reports an operator /compact as skipped instead of done.
Refs #4391
A `Runtime` trait (run / compact / archive) with two backends:
- claude: a pass-through to hive_claude's InfiniteSession and
SessionStore, so a claude turn is the same spawn, session handling
and errors as before.
- acp: a generic Agent Client Protocol client. It spawns the command,
args and env from RuntimeSpec (HIVE_RUNTIME / HIVE_ACP_COMMAND /
HIVE_ACP_ARGS / HIVE_ACP_ENV), refuses an agent whose
mcpCapabilities.http is not true, passes the claude --mcp-config
servers as ACP mcpServers, keeps one session id in a file
(session/load after a restart, session/new otherwise), and maps
session/update into claude stream-json events plus usage_update into
Telemetry. Permission requests are answered by a caller-supplied
policy on the ACP tool kind. compact returns Unsupported for now.
The crate depends on no hyperhive binary crate, so the subagent
daemon can move onto it without pulling in hive-agent.
Refs #4391
`ensure_hive_user` no longer short-circuits on the token file it
already has, and a hive's per-hive store token now wins over the file
whenever it's present. Both upgrade-guide bullets described the old
file-first order and needed rewording to match.
Refs #4427
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
legacy_tokens.rs referenced [`AGENT_TOKEN_NAME`] unqualified, but the
const lives in the sibling agent_token module and isn't in scope here;
rustdoc's broken-intra-doc-links lint (denied) failed the docs build.
The sweep ran once at startup and was never retried: a boot where
store::connect or the roster read failed (e.g. the controller up
before bao) left the tokens live until the next restart. It now runs
off forge::agent_token::spawn's five-minute mint pass, reusing that
pass's roster observation instead of a second store/roster read, so a
failed first tick retries on the next one. Idempotent, so a re-run
after a partial sweep deletes nothing extra. Drops the one-shot
startup spawn; one call path.
Also updates the forge.md and agent_token.rs docs that still said the
legacy hyperhive-<seconds> tokens stay until manually removed.
Refs #4644
Before b5d07d4d, hive-c0re minted a new `hyperhive-<unix-seconds>` token
for an agent on every spawn and rebuild and never revoked one, so every
live agent's forge user carries a pile of write-scoped tokens nothing
holds. Nothing in the tree lists or deletes them.
On each start, swarm-controller now walks the store's hive-agent-*
roster (the one the swarm-agent mint pass walks), and for every agent
whose swarm-agent token that pass would keep, deletes each token named
exactly `hyperhive-<digits>`. It logs the count per agent and a total.
- `core` is refused by name in both the roster filter and the per-user
delete: hive-c0re still names core's live admin token
`hyperhive-<unix-seconds>`.
- An agent whose swarm-agent token is not current is skipped, because
consumers fall back to `<state>/forge-token`, the last hyperhive-*
token, until the swarm token is fetched.
- A failed list or delete is logged and skipped; the sweep does not
retry and never blocks startup. A second start deletes nothing.
Tokens on forge users of agents no longer on the store roster (already
destroyed) are not reached.
Closes#4644
Two Grafana-managed rules in a "hyperhive" folder, one group evaluated
every 1m, each a LogsQL stats count over the VictoriaLogs datasource:
- forge reconcile failing: `swarm forge objects: write failed; retrying
next pass` over 10m, for 15m. The pass runs every 5m, so one failed
pass stays under `for` and a pass failing on every tick fires.
- swarm terminal publish skipped: `swarm terminal: publish skipped` over
15m, for 5m, one instance per hive/agent.
No contact point or notification policy. Grafana 13 routes to its
built-in `empty` receiver when none is provisioned, so firing rules
are visible under Alerting and sent nowhere, with no send errors.
Refs #4717
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.
hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.
Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.
module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.
Refs #3773
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.
`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).
swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.
Refs #3773
Exempting offline and paused from the name rules let a reserved name be
placed anyway: a first `offline` declared it, and the `up` after it
passed as an agent already declared on the hive. `paused` alone sufficed,
since the hive deploys an absent agent declared paused. A first
declaration in any placing state now runs the name rules and the
placed-elsewhere check; an agent already declared on the hive skips both.
A rule-breaking name whose roster read fails, including when no identity
bridge is configured, was accepted with a warning nobody sees, as in
create_agent. No later step on this route refuses the name, so it now
refuses with 503.
The gate is taken only for a placing declaration, so a destroy and its
credential revocation no longer wait on creations.
Refs #4804
The state route ran the name rules and the placed-elsewhere check on
every up/offline/paused declaration. An agent already declared on the
hive whose name breaks a rule, and which is not in the roster, could then
only be destroyed from the swarm.
Now an agent this hive already declares in a placing state passes both
checks. For a first declaration the placed-elsewhere check still applies
to up, offline and paused alike, and the name rules to up only: offline
and paused are how an operator stops an agent. The route reads this hive's
declaration to tell, and refuses with 503/500 when it cannot.
Refs #4804
PUT /api/hives/{hive}/agents/{agent}/state wrote any identifier into any
hive's wanted state, and the hive first-deploys a declared agent it has
no container for. That skipped create_agent's checks on the name.
A declaration that places the agent (up, paused, offline) is now refused
with 400 for a new name that breaks a naming rule, and with 409 for a
name the swarm has placed on another hive. Both reuse create_agent's
helpers (broken_name_rules/name_verdict, placements_elsewhere), under the
same gate, and an unreadable wanted state on another hive refuses with
503/500 as creation does. `destroyed` places nothing and is not checked,
so the revocation path is unchanged.
An agent with no declaration at all is still accepted: swarm-ui declares
state for agents that predate swarm-level creation, which is how the swarm
adopts them. Whether to refuse such names instead is left open on #4804.
Refs #4804
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.
The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.
OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.
The viewer's log message no longer tells the operator to restart it.
Refs #4772
create_agent read the other hives' wanted state first and snapshotted the
queued SetAgentWanted nodes second. A node finishing between the two was
in neither — not yet declared at the first read, already terminal at the
second — so a second hive could get the same name.
The queue is now snapshotted first: a SetAgentWanted node only turns
terminal after its declaration is written, so anything terminal by then is
visible to the read that follows. The order lives in
`placements_elsewhere`, and a test that finishes a node between the two
reads fails with them swapped.
Refs #4396
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.
Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.
policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.
Refs #4396
Its only client was hive-c0re's matrix-account-login handler, removed
earlier on this branch, so nothing sends `restart_matrix_daemon` any
more. Drops the `PrivRequest` variant, the hive-priv handler and its
`systemctl --machine=h-<agent> restart hive-matrix-daemon.service`
helper, and the row in the hive-priv op table in security.md.
`PrivRequest` is internally tagged by `op` name, so no other variant's
encoding changes. A new token file still re-fires the daemon through
its `matrix-token*` path unit.
Refs #4348
`PutMatrixAccountRequest`'s no-`Debug` comment named the deleted
`MatrixLoginForm` as its precedent; it now states the reason directly.
The daemon's missing-sidecar warning told operators to re-login via the
dashboard, which no longer has that form; it now points at re-linking
the account from the swarm UI, whose stored credential carries the
homeserver the sidecar is written from. Behaviour is unchanged.
Refs #4348
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.
Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.
The GITHUB tab is the credentials page's default tab now.
Refs #4348
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.
Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.
Refs #4585
authelia's file backend has no self-service reset (no SMTP notifier),
so the only way a human account got a new password after the old one
was forgotten was hand-editing users.yml as root. `user add` already
hashes a password into the file; this verb does the same for an
existing user instead of refusing on the name.
Mirrors `user add`'s UX exactly: no password flag, authelia generates
and hashes it (never crosses argv), and it's printed once and never
stored. Refuses on an unknown user before ever invoking authelia. Same
publish path as add/update, so the same atomic write and no-restart
(authelia watches the file) behaviour apply.
Split the digest-replacement into users::reset_password so it's
testable without a command line or a running authelia, same pattern
as apply_update.
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.
hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.
A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).
Closes#4783
post_purge_tombstone discarded fail_pending_for_agent's error with
let _ =. Mirrors #4740's fix for the identical discard in
job_queue/exec.rs's run_destroy_bookkeeping: warn and continue, since
the purge itself has already succeeded by this point.
Refs #4747
du_bytes returned None uniformly for every du failure, and both
du_bytes and measure_agent_disk had no logging at all, so a container
rootfs du couldn't read (permission, I/O error, unparseable output)
silently reported as 0 disk with no signal in the journal.
Distinguish the expected case — the path plain doesn't exist, e.g. a
destroyed-but-kept agent's rootfs or a not-yet-created state dir — from
an actual read failure, and warn only on the latter, with the path plus
whichever of the failure (spawn error, exit status, stderr, unparseable
stdout) applies.
The raw reqwest client hive-forge uses for Forgejo's web-router-only
routes (attachment downloads, Actions artifact/log routes) had no
timeout, so a hung Forgejo response blocked the calling CLI invocation
indefinitely.
Adds a 5s connect_timeout (matching #4737's outbound-HTTP sites) plus
a per-call request timeout: 15s (config_pr_poll's forge budget) for
the small JSON calls (get_api_json, post_json_web), and 10 minutes for
get_bytes_named/get_bytes_raw, which download attachments, Actions
artifact zips and persisted job logs that can be large.
reqwest::blocking has no separate read/idle timeout, so a single
whole-request budget has to cover those downloads; the smaller JSON
budget would cut them off partway through.
Refs #4746, #4737
`script-test-agent-bao-fetch` executes the rendered `ExecStart` of
`hive-agent-forge-token` and `hive-agent-queue-credential` under
`umask 0377`, with a stub `bao` first on the unit's own PATH. It covers
every error branch, the happy path, forge rotation/unchanged, and 0400
files already in place: the redirect failure #4736 fixed, whose live
symptom was `bao.err: Permission denied` reported as a refused
certificate. The TLS-alert fixture is the `unknown certificate
authority` error h-atlas's identity check got from the store.
#4736 claimed this test but never committed it; the two module-eval
suites for these units only evaluate the config.
Closes#4748
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.
Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
prebuild_toplevel awaited nix build with a bare child.wait().await, so
a wedged nix-daemon (unreachable remote builder, stuck build slot)
hung the rebuild job forever with no way for the job queue to recover
short of restarting hive-c0re.
Mirrors #4741's hive-priv fix: the child now leads its own process
group, and after PREBUILD_TIMEOUT (1h, four times CI's observed
cold-cache flake check) the whole group is SIGKILLed and the call
fails with a named timeout error instead of hanging.
Refs #4723, #4741
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.
Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
rustdoc runs with -D rustdoc::broken-intra-doc-links and the type is
imported inside the function body, not at module scope, so the bare
link resolved to nothing and failed the workspace-doc derivation.
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.
docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.
Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.
Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.
The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
mara, PR review: "the dynamic tab should be in the top bar, not a new
one below". Moves the .shell-tabs group from its own sticky row under
the header into .shell-nav itself, right after the nav indicator.
Also adds a third re-measure effect for the sliding nav indicator,
keyed on tabs.length: with tabs inline in the same flex row the
indicator measures, closing a background tab (no navigation) can
shrink the row without the hop effect's own re-measure ever firing.
Same reflow-not-navigation reasoning as the existing resize-listener
effect.
argus's review on PR #4784 caught a wrong technical claim: the
comment said the .ui-agent-term-preview-full override rules relied
on source order because they had equal specificity to the
un-modified rules above. They don't — each override selector adds
one more class (the .ui-agent-term-preview-full prefix) than what it
overrides, so they're strictly more specific and win regardless of
file order. Corrected the comment to say so.
Adds a full, non-capped agent terminal reachable from a new expand
trigger on the embedded AgentTermPreview (the detail-panel preview on
AgentsPage stays as-is, just gains the trigger). Opens
/agents/:name/terminal in a new dynamic tab in Shell's header, next to
the static nav row — tabs persist across a reload via useDynamicTabs,
a small localStorage-backed hook built on @hive/shared's existing
settings-storage primitive.
AgentTermPreview gains two new props to support both mounts from one
component: fullHeight (drops the 12em preview cap, fills its page)
and showHeaderBadges (default true — lets a future caller that
already shows turn_state/model/ctx/cost elsewhere suppress this
cluster; AgentsPage doesn't use it, see below).
Deviation from the originally posted plan (issue comment 80596): that
plan proposed AgentsPage's embedded preview pass showHeaderBadges as
false, reasoning the detail panel already duplicates that info.
Checked the actual code before implementing — it doesn't; AgentRow/
AgentTypes.ts carry none of turn_state/model/ctx/cost, and
AgentTermPreview's own floating badges are the only place swarm-ui
shows them. Left the badges visible there instead of shipping a
regression the plan's own stated justification didn't hold up to.
Also fixed a same-tab pub/sub race found by actually rendering a cold
load of /agents/:name/terminal (headless chromium, not just reasoning
about the code): useLocalSetting subscribes inside a useEffect, and
mount effects fire children-before-parents, so a descendant's
mount-time write (AgentTerminalPage registering its own tab) can beat
an ancestor's (Shell's) subscription into existence, leaving Shell's
tab row silently empty on a direct/reload load. Fixed by having
useDynamicTabs re-sync from storage on every location change, not
just on notify() — the fix lives in the new hook itself, not in the
shared settings-storage primitive theme/motion overrides also use.
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.
Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.
nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
A stored queue secret with no `minted_at` counted as due, and no secret
minted before the renewal pass has one, so the first pass after deploy
would re-mint every agent's secret. Every reconnect before that agent's
next restart would then be refused.
Such a secret is now stamped instead: `minted_at = now` is written beside
the unchanged `value`, and its 45-day clock starts there. Only a secret
whose recorded mint time is at least 45 days old gets a new value.
The decision is `secret_step` (Keep / Backfill / Remint), and the pass
reports an unstamped secret as `Observed::Unstamped`. Backfill and
re-mint log different lines.
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:
- `MintAgentIdentity` (the node agent creation uses) when the stored
certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
re-decides, writes a fresh value with `minted_at`, reads it back, and logs
the agent and the old age.
When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.
`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.
Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.
Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.
docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.
- authelia registers an interactive client `swarm-bao-ui`
(glue-bao-ui-oidc-client.nix) with redirect
`https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
publisher carries its secret to
`secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
mount with listing visibility `unauth`, asked before attempted like
cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
policy, and once the granter may configure `auth/oidc/config` (checked
through `sys/capabilities-self`), writes the mount's config from the
published secret and the role binding `groups=admins` to the viewer.
Before the bootstrap step re-runs it writes the policy, logs the step
and exits 0.
Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.
module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.
The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.
Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.
Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.
On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
A job-level if: is evaluated by whichever runner picks the job up. The
public forge's copy has only a nixos runner, so a job guarded by
if: vars.PUBLIC_FORGE != 'true' and pinned to runs-on: [hive-ci] never
gets picked up there to evaluate the guard at all - it just sits queued.
runs-on now carries the same PUBLIC_FORGE variable: hive-ci internally,
nixos on the public copy. There the nixos runner picks the job up,
evaluates the existing if:, and skips it immediately. Internally
PUBLIC_FORGE is unset, so runs-on still resolves to hive-ci and nothing
changes.
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.
Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.
Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.
Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
Guards internal-only jobs (ci.yml, coverage.yml, flake-update.yml) with
`vars.PUBLIC_CACHE != 'true'` and adds public-cache.yml, guarded to the
inverse, to build the deployed closures and push them to the public
attic cache on a push to main.
`PUBLIC_CACHE` is a repo Actions variable, opt-in only on the public
copy: unset here, it leaves internal CI's `!=` guards true so internal
jobs always run. Job-level `if:` cannot see the `github`/`forgejo`
context at all on this runner (confirmed empirically — a
`github.server_url` comparison always evaluates false at job level,
though the identical comparison resolves correctly inside a step),
so `vars.*`, which is available at job level, is the only usable
opt-in signal here.
Same gap #4611 fixed for logout: the Preact rewrite's StatusChips menu
never picked up new-session as a click path, only /new-session typed
twice into the terminal. postNewSession already existed in
termActions.ts. Wire it into the status menu with the same arm-then-
confirm click pattern logout/cancel-turn already use.
Closes#4612
Each agent card now leads with the agent's icon, loaded as an `<img>`
from `GET /api/agents/<name>/icon`: the same 5em square, background and
fallback as the hive dashboard's container row. An agent with no icon
(the route's 404), or any other failed load, shows the dimmed hyperhive
mark (`/favicon.svg`) instead of a broken image.
Only ever an `<img>`, never inline markup: the body is an agent-authored
SVG, and an image load does not run its script.