A plain removal breaks any out-of-tree host config that still sets the
option: eval fails with "option does not exist" and no pointer to what
replaced it. mkRemovedOptionModule gives a clear evaluation error instead.
swarm-controller is the only minter of swarm/hives/<hive>/matrix/sender-token
since #4820, so the swarm-matrix-ctl policy's first stanza granted a write no
code performs. The policy keeps its one used stanza, the swarm appservice token
that `swarm-matrix-ctl appservice publish` writes. deploy.bao.matrixCtlHiveName
only named the hive in the dropped stanza and goes with it.
hive-matrix.nix no longer calls the per-hive appservice's as_token hive-c0re's
authority: tuwunel loads the registration and creates the sender account, and
no client presents that token.
A `compact` command that errors, or sends nothing, used to leave the
session as full as before: `compacted` stayed false, so every following
turn ran the checkpoint and another `/compact` again, each one waiting
out the 600s turn idle window when the command was silent.
- The `/compact` turn gets its own idle bound, COMPACT_IDLE (3 min, or
the turn's idle window if shorter), through the turn's existing
watchdog, so a silent command is cancelled (or killed) like any
stalled turn.
- When the command fails or hits that bound, compaction falls back to
the no-command path: the session is archived and the next turn starts
a new one. The checkpoint turn runs there only if it has not already
run in this compaction.
- Compaction returns early when attaching produced a new session (a
failed `session/load` or a never-answered first prompt): there is
nothing in it to compact.
Refs #4391
An ACP agent's session is now compacted like a claude one: proactively
once a turn crosses the percent-of-window watermark, and on the
operator's /compact or the agent's compact tool. Before, the ACP
backend's compact returned Unsupported and no watermark applied to it.
- AcpRuntime takes the same CompactionPolicy as ClaudeRuntime;
make_session builds one PercentPolicy (with CHECKPOINT_PROMPT) and hands
it to whichever backend runs.
- The runtime keeps the commands each session advertises in
available_commands_update. If `compact` is among them, compaction sends
the prompt `/compact` on the same session, which is how the ACP spec
runs an advertised command. A proactive compaction runs the checkpoint
turn first, as InfiniteSession does.
- With no compact command, the checkpoint turn runs, the session is
archived, and the next turn starts a new one with the system prompt.
- Error::Unsupported had no producer left, so it and drive_turn's
"/compact skipped" arm are gone.
Refs #4391
`Canceller::cancel()` does nothing and returns `false` when no ACP turn
is in flight: from `start` until the run's task begins its turn, and
from a turn's end until `after_turn` releases the name. `interrupt`
ignored that return, recorded `Cancelled` and replied that the run was
cancelled, so a non-goal run went on to do its whole first turn while
`status` read as cancelled.
On ACP, `interrupt` now records the stop, cancels, and if nothing was
in flight withdraws the stop under the same `stops` lock, restores the
tracking entry, and refuses with the claude path's "still starting, try
again shortly". The claude arm is unchanged.
Review finding on #4824 (argus).
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.
The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.
Refs #4391
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:
- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
hive-matrix container, Command::Mint and src/mint.rs. The binary, its
appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
role stay. bao-matrix-reader's checks on the deleted unit are removed;
the leaf-identity and no-token-in-env checks now look at
swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
(register/appservice-login/password-login with the local as_token), with
read_appservice_token, paths::matrix_appservice_token and the helpers
only it used. ensure_hive_user now takes the store's token, keeps the
file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
unchanged apart from no longer reading the local as_token.
This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.
Closes#4813Closes#4814
`/api/cancel` stopped a turn only by SIGINTing a child process whose argv0 is
`claude`, so on an ACP agent it reported "no claude process to interrupt"
and the turn ran on. The serve loop now stores the session's canceller on
the `Bus` when the runtime has one, and `/api/cancel` uses it: the agent is
sent `session/cancel`, and the next wake prompt carries the interrupted hint,
as for a signalled claude turn. Without a canceller (claude), the SIGINT path
is untouched.
An ACP turn stopped by the idle watchdog (`HIVE_TURN_IDLE_SECS`) becomes
`TurnError::AgentStall(note)`, handled like `ApiStall`: park for
`HIVE_STALL_SLEEP_SECS`, requeue, and record `api_stall`. The TurnEnd note
is the runtime's message (how long the agent was silent, whether it had to
be killed, and that a silently retried provider error such as HTTP 429
looks like this), not the claude-specific `ApiStall` text. A cancel the
agent ignores is a `Failed` turn.
Refs #4391
`Runtime` gets a fourth operation, `canceller()`: a handle that stops the
turn in flight from outside `run`. The ACP backend returns one; claude
returns `None`, because the harness stops a claude turn by signalling the
`claude` process, and that path is unchanged.
Both stops send the agent `session/cancel`:
- `Canceller::cancel()`, when asked from outside. The turn then ends
normally, reported with stop reason `cancelled` whatever reason the agent
gives. opencode 1.15.10, for one, answers a cancelled prompt with
`end_turn` (`acp/agent.ts` `prompt()` always returns `end_turn`).
- The idle watchdog, once no `session/update` has arrived for
`Config::idle_timeout`, the same field claude's watchdog reads. The turn
fails with `AcpError::IdleTimeout`.
An agent that has not answered the prompt 10s after `session/cancel` is
killed (`IdleKilled` / `CancelIgnored`), and the next turn respawns it.
The watchdog is also what ends a turn stuck on a provider HTTP 429.
opencode 1.15.10 retries a retryable provider error with no attempt limit
(`session/retry.ts` `policy`, `session/processor.ts` `Effect.retry`) and
forwards neither `session.status` nor `session.error` over ACP (its
`handleEvent` only handles `permission.asked` and `message.part.*`). So the
ACP client sees nothing at all until the provider recovers. The
`IdleTimeout` message says a silently retried provider error looks like
this.
Refs #4391
A new session is only recorded once its first prompt is answered, so a
failed first prompt made the next turn run `session/new` again and left the
first session behind in the agent: in the same process, and after a restart.
The new session's id is now kept in `<session file>.pending` until that
prompt is answered. The next turn attaches it (in-process, or through
`session/load` after a restart) and still prepends the system prompt, since
the session has not had an answered prompt yet. Recording the session
removes the pending file, and `archive` drops it.
Raised by argus in the #4812 review.
Refs #4391
A permission request counted as an MCP tool call whenever its title
looked like `<server>_<tool>`, whatever its kind, and acp_permits
allowed MCP calls before looking at the kind. So an `execute` request
titled e.g. `hyperhive_x`, or a `fetch` without web_tools, was allowed.
MCP tool calls come with kind `other` (opencode's toToolKind maps every
tool it doesn't name, MCP tools included, to "other"). The runtime now
sets PermissionAsk::mcp_server only for kind `other`, and acp_permits
allows an MCP server's tool only under the `other` arm.
Refs #4391
acp_permits allows tools of the session's MCP servers, the read, edit
and search kinds, and fetch with web_tools; every other kind, including
execute, other and kinds it doesn't know, is refused. Before, anything
but execute (and fetch without web_tools) was allowed, which was only
safe while the preset's own config denied the risky built-ins.
Refs #4391
The session id was written to the session file right after session/new,
but the system prompt rides on the first prompt only. If that prompt
failed, the next turn (or the next harness start) resumed the recorded
session as not-new and the system prompt never reached it.
The id is now written after the first session/prompt gets its reply.
A failed first prompt leaves nothing recorded, so the next turn starts
a new session and sends the system prompt again. Chosen over a separate
"system prompt delivered" marker: one file, and "recorded" already means
"usable".
Tests drive AcpRuntime against a scripted sh agent that fails the first
prompt: in-process and across a restart, the retry is a new session
carrying the system prompt.
Also: PermissionPolicy now sees a PermissionAsk (kind plus the MCP
server the tool belongs to, matched by name against the servers handed
to the session), so a caller can tell MCP tool calls from other `other`
requests.
Refs #4391
services.hyperhive.agent.runtime ("claude" default | "acp") and
acp.{command,args,env}, rendered into HIVE_RUNTIME / HIVE_ACP_* only
for acp, so a claude agent's unit is unchanged. acp implies useApiKey.
acp.presets.opencode runs `opencode acp` from nixpkgs against an
OpenAI-compatible provider from acp.opencode.{provider,model,
contextWindow,outputLimit}: the config is rendered to the store with
the API key as an {env:VAR} reference, so the key is read at runtime
from backendEnvironmentFile. OPENCODE_PERMISSION denies opencode's
built-in bash, task, todowrite and websearch, and makes webfetch ask.
Refs #4391
AgentSession becomes hive_runtime::AgentRuntime, picked at startup from
the environment. Unset HIVE_RUNTIME keeps the claude backend, built
from the same title, store and PercentPolicy as before; drive_turn,
the 401 retry and the error mapping see the claude errors unchanged.
On acp the harness keeps the session id in the harness dir, answers
permission requests like the claude built-in allow-list (no built-in
shell, web fetch only with web_tools), maps runtime errors to Failed,
and reports an operator /compact as skipped instead of done.
Refs #4391
A `Runtime` trait (run / compact / archive) with two backends:
- claude: a pass-through to hive_claude's InfiniteSession and
SessionStore, so a claude turn is the same spawn, session handling
and errors as before.
- acp: a generic Agent Client Protocol client. It spawns the command,
args and env from RuntimeSpec (HIVE_RUNTIME / HIVE_ACP_COMMAND /
HIVE_ACP_ARGS / HIVE_ACP_ENV), refuses an agent whose
mcpCapabilities.http is not true, passes the claude --mcp-config
servers as ACP mcpServers, keeps one session id in a file
(session/load after a restart, session/new otherwise), and maps
session/update into claude stream-json events plus usage_update into
Telemetry. Permission requests are answered by a caller-supplied
policy on the ACP tool kind. compact returns Unsupported for now.
The crate depends on no hyperhive binary crate, so the subagent
daemon can move onto it without pulling in hive-agent.
Refs #4391
`ensure_hive_user` no longer short-circuits on the token file it
already has, and a hive's per-hive store token now wins over the file
whenever it's present. Both upgrade-guide bullets described the old
file-first order and needed rewording to match.
Refs #4427
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
legacy_tokens.rs referenced [`AGENT_TOKEN_NAME`] unqualified, but the
const lives in the sibling agent_token module and isn't in scope here;
rustdoc's broken-intra-doc-links lint (denied) failed the docs build.
The sweep ran once at startup and was never retried: a boot where
store::connect or the roster read failed (e.g. the controller up
before bao) left the tokens live until the next restart. It now runs
off forge::agent_token::spawn's five-minute mint pass, reusing that
pass's roster observation instead of a second store/roster read, so a
failed first tick retries on the next one. Idempotent, so a re-run
after a partial sweep deletes nothing extra. Drops the one-shot
startup spawn; one call path.
Also updates the forge.md and agent_token.rs docs that still said the
legacy hyperhive-<seconds> tokens stay until manually removed.
Refs #4644
Before b5d07d4d, hive-c0re minted a new `hyperhive-<unix-seconds>` token
for an agent on every spawn and rebuild and never revoked one, so every
live agent's forge user carries a pile of write-scoped tokens nothing
holds. Nothing in the tree lists or deletes them.
On each start, swarm-controller now walks the store's hive-agent-*
roster (the one the swarm-agent mint pass walks), and for every agent
whose swarm-agent token that pass would keep, deletes each token named
exactly `hyperhive-<digits>`. It logs the count per agent and a total.
- `core` is refused by name in both the roster filter and the per-user
delete: hive-c0re still names core's live admin token
`hyperhive-<unix-seconds>`.
- An agent whose swarm-agent token is not current is skipped, because
consumers fall back to `<state>/forge-token`, the last hyperhive-*
token, until the swarm token is fetched.
- A failed list or delete is logged and skipped; the sweep does not
retry and never blocks startup. A second start deletes nothing.
Tokens on forge users of agents no longer on the store roster (already
destroyed) are not reached.
Closes#4644
Two Grafana-managed rules in a "hyperhive" folder, one group evaluated
every 1m, each a LogsQL stats count over the VictoriaLogs datasource:
- forge reconcile failing: `swarm forge objects: write failed; retrying
next pass` over 10m, for 15m. The pass runs every 5m, so one failed
pass stays under `for` and a pass failing on every tick fires.
- swarm terminal publish skipped: `swarm terminal: publish skipped` over
15m, for 5m, one instance per hive/agent.
No contact point or notification policy. Grafana 13 routes to its
built-in `empty` receiver when none is provisioned, so firing rules
are visible under Alerting and sent nowhere, with no send errors.
Refs #4717
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.
hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.
Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.
module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.
Refs #3773
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.
`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).
swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.
Refs #3773
Exempting offline and paused from the name rules let a reserved name be
placed anyway: a first `offline` declared it, and the `up` after it
passed as an agent already declared on the hive. `paused` alone sufficed,
since the hive deploys an absent agent declared paused. A first
declaration in any placing state now runs the name rules and the
placed-elsewhere check; an agent already declared on the hive skips both.
A rule-breaking name whose roster read fails, including when no identity
bridge is configured, was accepted with a warning nobody sees, as in
create_agent. No later step on this route refuses the name, so it now
refuses with 503.
The gate is taken only for a placing declaration, so a destroy and its
credential revocation no longer wait on creations.
Refs #4804
The state route ran the name rules and the placed-elsewhere check on
every up/offline/paused declaration. An agent already declared on the
hive whose name breaks a rule, and which is not in the roster, could then
only be destroyed from the swarm.
Now an agent this hive already declares in a placing state passes both
checks. For a first declaration the placed-elsewhere check still applies
to up, offline and paused alike, and the name rules to up only: offline
and paused are how an operator stops an agent. The route reads this hive's
declaration to tell, and refuses with 503/500 when it cannot.
Refs #4804
PUT /api/hives/{hive}/agents/{agent}/state wrote any identifier into any
hive's wanted state, and the hive first-deploys a declared agent it has
no container for. That skipped create_agent's checks on the name.
A declaration that places the agent (up, paused, offline) is now refused
with 400 for a new name that breaks a naming rule, and with 409 for a
name the swarm has placed on another hive. Both reuse create_agent's
helpers (broken_name_rules/name_verdict, placements_elsewhere), under the
same gate, and an unreadable wanted state on another hive refuses with
503/500 as creation does. `destroyed` places nothing and is not checked,
so the revocation path is unchanged.
An agent with no declaration at all is still accepted: swarm-ui declares
state for agents that predate swarm-level creation, which is how the swarm
adopts them. Whether to refuse such names instead is left open on #4804.
Refs #4804
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.
The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.
OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.
The viewer's log message no longer tells the operator to restart it.
Refs #4772
create_agent read the other hives' wanted state first and snapshotted the
queued SetAgentWanted nodes second. A node finishing between the two was
in neither — not yet declared at the first read, already terminal at the
second — so a second hive could get the same name.
The queue is now snapshotted first: a SetAgentWanted node only turns
terminal after its declaration is written, so anything terminal by then is
visible to the read that follows. The order lives in
`placements_elsewhere`, and a test that finishes a node between the two
reads fails with them swapped.
Refs #4396
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.
Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.
policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.
Refs #4396
Its only client was hive-c0re's matrix-account-login handler, removed
earlier on this branch, so nothing sends `restart_matrix_daemon` any
more. Drops the `PrivRequest` variant, the hive-priv handler and its
`systemctl --machine=h-<agent> restart hive-matrix-daemon.service`
helper, and the row in the hive-priv op table in security.md.
`PrivRequest` is internally tagged by `op` name, so no other variant's
encoding changes. A new token file still re-fires the daemon through
its `matrix-token*` path unit.
Refs #4348
`PutMatrixAccountRequest`'s no-`Debug` comment named the deleted
`MatrixLoginForm` as its precedent; it now states the reason directly.
The daemon's missing-sidecar warning told operators to re-login via the
dashboard, which no longer has that form; it now points at re-linking
the account from the swarm UI, whose stored credential carries the
homeserver the sidecar is written from. Behaviour is unchanged.
Refs #4348
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.
Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.
The GITHUB tab is the credentials page's default tab now.
Refs #4348
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.
Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.
Refs #4585
authelia's file backend has no self-service reset (no SMTP notifier),
so the only way a human account got a new password after the old one
was forgotten was hand-editing users.yml as root. `user add` already
hashes a password into the file; this verb does the same for an
existing user instead of refusing on the name.
Mirrors `user add`'s UX exactly: no password flag, authelia generates
and hashes it (never crosses argv), and it's printed once and never
stored. Refuses on an unknown user before ever invoking authelia. Same
publish path as add/update, so the same atomic write and no-restart
(authelia watches the file) behaviour apply.
Split the digest-replacement into users::reset_password so it's
testable without a command line or a running authelia, same pattern
as apply_update.
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.
hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.
A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).
Closes#4783
post_purge_tombstone discarded fail_pending_for_agent's error with
let _ =. Mirrors #4740's fix for the identical discard in
job_queue/exec.rs's run_destroy_bookkeeping: warn and continue, since
the purge itself has already succeeded by this point.
Refs #4747
du_bytes returned None uniformly for every du failure, and both
du_bytes and measure_agent_disk had no logging at all, so a container
rootfs du couldn't read (permission, I/O error, unparseable output)
silently reported as 0 disk with no signal in the journal.
Distinguish the expected case — the path plain doesn't exist, e.g. a
destroyed-but-kept agent's rootfs or a not-yet-created state dir — from
an actual read failure, and warn only on the latter, with the path plus
whichever of the failure (spawn error, exit status, stderr, unparseable
stdout) applies.
The raw reqwest client hive-forge uses for Forgejo's web-router-only
routes (attachment downloads, Actions artifact/log routes) had no
timeout, so a hung Forgejo response blocked the calling CLI invocation
indefinitely.
Adds a 5s connect_timeout (matching #4737's outbound-HTTP sites) plus
a per-call request timeout: 15s (config_pr_poll's forge budget) for
the small JSON calls (get_api_json, post_json_web), and 10 minutes for
get_bytes_named/get_bytes_raw, which download attachments, Actions
artifact zips and persisted job logs that can be large.
reqwest::blocking has no separate read/idle timeout, so a single
whole-request budget has to cover those downloads; the smaller JSON
budget would cut them off partway through.
Refs #4746, #4737
`script-test-agent-bao-fetch` executes the rendered `ExecStart` of
`hive-agent-forge-token` and `hive-agent-queue-credential` under
`umask 0377`, with a stub `bao` first on the unit's own PATH. It covers
every error branch, the happy path, forge rotation/unchanged, and 0400
files already in place: the redirect failure #4736 fixed, whose live
symptom was `bao.err: Permission denied` reported as a refused
certificate. The TLS-alert fixture is the `unknown certificate
authority` error h-atlas's identity check got from the store.
#4736 claimed this test but never committed it; the two module-eval
suites for these units only evaluate the config.
Closes#4748
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.
Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
prebuild_toplevel awaited nix build with a bare child.wait().await, so
a wedged nix-daemon (unreachable remote builder, stuck build slot)
hung the rebuild job forever with no way for the job queue to recover
short of restarting hive-c0re.
Mirrors #4741's hive-priv fix: the child now leads its own process
group, and after PREBUILD_TIMEOUT (1h, four times CI's observed
cold-cache flake check) the whole group is SIGKILLed and the call
fails with a named timeout error instead of hanging.
Refs #4723, #4741
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.
Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
rustdoc runs with -D rustdoc::broken-intra-doc-links and the type is
imported inside the function body, not at module scope, so the bare
link resolved to nothing and failed the workspace-doc derivation.
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.
docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.