Review fixes for #4879 (argus):
- mcp.md: the 'Waking the agent' cross-ref pointed at Core tools, which
never mentions UpsertTodo/HIVE_AGENT_SOCKET. Point it at
docs/tools/bash.md's 'Completion as a todo (loose-ends v2)' section,
which documents the actual upsert/signal/clear mechanism.
- conventions.md 'Wake injection': still framed AgentRequest::Wake (a
type that no longer exists) as the live wake surface with matrix/forge
as callers. hive_core_agent_sock::Request::Wake has exactly one
non-test reference on origin/main (the handler at
socket_server/mod.rs:307) and no client; matrix/bash/forge all moved
to the in-agent todo socket. Rewritten to match, linking mcp.md's
'Waking the agent' section instead of duplicating it.
docs/tools/matrix.md:136-137 has the same stale AgentRequest::Wake claim
(and contradicts its own :159-165) but is out of scope (#4136) — noted
as a follow-up in the PR body instead of edited.
mcp.md:
- matrix and subagent extra MCP servers are http (hive-matrix-daemon,
hive-subagent-daemon), not stdio; screen is the one entry that still
uses the stdio default (nix/agent-modules/matrix.nix:290-297,
screen.nix:22-25)
- set_status is always-on, not meta-group-gated; mark_todos_done (also
always-on) was undocumented (hive-sh4re/src/permissions.rs:151,
hive-agent-mcp/src/mcp/mod.rs:425)
- System messages: HelperEvent has 3 variants, not the 8 previously
listed; ApprovalResolved/ContainerCrash routing and the swarm-wide
NATS notices stream (swarm_notices.rs) replace the old per-agent
todo-wake description for rebuilt/killed/destroyed/logged_in/needs_login
- get_loose_ends's approval rows are manager-only; PendingMessages and
UnreadMatrix were missing from the description
(hive-sh4re/src/inbox.rs:127-184)
- subagent spawning runs on hive-runtime (claude or ACP), not
claude-only (hive-subagent-mcp/src/session.rs:77)
- Waking section: matrix/bash/forge all moved to the in-agent todo
socket; the host Wake request has no built-in caller left today
agent-hierarchy.md:
- distinguished the swarm-wide agent roster (swarm-controller's
identity store, authoritative) from the hive-local topology.json
(a derived, reconciled cache scoping ManageRootAgent's bind-mounts),
linking README's framing
- noted services.hyperhive.ruthless (a hive can run with no manager at
all)
- Wire-protocol bullet: the only privileged Request variants left are
the scheduling ops; Kill/Start/Restart/Update/GetLogs don't exist on
this socket
- Prompt/tools: prompt::render hardcodes the agent role for every
container today (role:manager blocks are dead code); the tool
allow-list has no Flavor switch, it's HIVE_TOOL_GROUPS same as any
agent
Not touched: agent-hierarchy.md:140-200 (Harness systemd unit shape,
kept in place — see PR follow-ups) and docs/agent-lifecycle/approvals.md
(blocked on #4853).
authelia 4.39.20 exits at startup on `users: {}` ("users: non zero value
required"), and the first-boot unit seeded exactly that, so a swarm with
no users crash-looped authelia and answered 502 until `swarmctl user add`
ran.
The first-boot unit now writes one subject, `swarm.placeholder`, when
the users database is absent, empty, or exactly `users: {}`:
- `disabled: true` — authelia returns "user not found" for a disabled
user before any password check (file_user_provider.go,
CheckUserPassword).
- password: an argon2id digest with an all-zero key. It decodes (authelia
rejects a non-digest at startup) and no known password hashes to it.
- the `.` keeps it out of agent names (`[a-z0-9-]`), and `swarmctl user
add` refuses it as already existing. Neither writer removes users, and
both round-trip `disabled`.
A file with any user in it is never touched.
The docs that described the crash-loop (sso.md, gateway.md, setup.md,
the sso-unavailable error page) now describe the placeholder; the
writers' load_store docs and the seed fixtures follow. module-eval
nats-authelia asserts the seed branch.
swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.
Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
"not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
account, and declares the agent paused (main.rs:1822-1920, 247-248)
Refs #3902
- frontend/README.md: list the missing packages/swarm-ui package, fix
the vanilla-JS claim (all packages depend on preact), and the
deprecated hyperhive.frontend.extraFiles spelling
- hive-c0re/README.md: hive-c0re no longer provisions per-agent
forge/matrix accounts (swarm-controller does); it wires gateway
vhosts and reconciles forge/matrix config
- hive-metric/README.md: OTEL_EXPORTER_OTLP_HEADERS is never set by
the harness and has no way to be set
- hive-agent-sock/README.md: fix the deprecated
hyperhive.extraMcpServers spelling
- docs/process/gotchas.md: fix a dangling hive-ag3nt/ path, the real
directory is hive-agent/
Also fixes pre-existing vale error-level alerts (passive voice,
Microsoft.Auto, Microsoft.Contractions) in the same files so
prose-lint-errors passes clean.
The old FORGES tab wrote forge-<label>-token/forge-<label>.json directly;
the swarm-fetch unit that replaced it never deletes a pair for a label it
doesn't list, so those files keep being read by hive-forge -f <label>
until the operator links an account under the same label in the swarm UI,
which overwrites both files (nix/agent-modules/forge-accounts.nix:11-13,159-176).
hive-matrix-daemon removes undeclared matrix-token-<name> files on
every start (hive-matrix-mcp/src/main.rs:100), so there is no window
where an account linked through the old per-agent MATRIX tab keeps
working — it must be declared via matrixAccounts at swarm/agent
level.
Frame the turn loop runtime-neutrally: every turn runs through
hive-runtime, on claude (default) or an ACP agent. The loop steps,
harness binary shape and hive-agent README now say so; claude-only
failure detection gets its own heading; claude-invocation.md opens with
its scope and links the ACP side to hive-runtime/README.md.
Fact fixes: on-boot file paths (/run/hive-config, not /run/hive),
hive-claude is a crates.io dependency with no README here, hive-agent
has no client.rs or forge_notify.rs, the claude launch-config layer is
hive-agent's mcp_config.rs, agent forge/matrix accounts are
swarm-controller's, hive-c0re's dashboard is dashboard/, and the
deprecated hyperhive.gui.enable / hyperhive.extraMcpServers spellings.
Refs #3902
- hivectl/README.md: split matrix.rs/github.rs into accurate per-module
lines (matrix.rs invites a matrix user to the hive Space/room,
cli.rs:289-296; it is not per-agent). Fixed the summary line's
'provisioning verbs' to name the actual verb groups left after the
facts pass.
- web-ui/README.md: the swarm/ui.md pointer no longer promises roster/
creating-agents/linking-accounts content that page doesn't have.
- Reverted the unrelated doesn't/does not drive-by at hivectl/README.md:5.
docs/web-ui/README.md now frames itself as the per-hive dashboard (host
approvals, container state, rebuild queue) and points to swarm/ui.md for
swarm-wide day to day, matching the README's swarm-first reframe.
Credentials page fixes: it has only a GitHub PAT tab now (2c7e586f
removed the FORGES tab and moved external forge accounts to the swarm
UI; credentials.html never had a matrix tab).
hivectl/README.md: agents.rs has no create verb (hivectl-cli.md has no
'create' entry; agents.rs's create is swarm-level, swarmctl agent
create). forge.rs reconciles config, it doesn't provision an account;
matrix.rs/github.rs do agent-scoped invites/token writes; gateway.rs
manages the gateway's own htpasswd users — split out of the former
single 'per-integration account/token provisioning' line.
Fix the 4 pre-existing vale error-level hits on main (docs/README.md:83,
docs/getting-started/setup.md:11,127, docs/swarm/bao.md:4) that fail CI's
prose-lint-errors job for every docs PR regardless of its own diff.
An operator now links an agent's external forge account (label, base URL,
token) in the swarm UI. swarm-controller stores it at
swarm/agents/<agent>/forge/<label>. There is no index: the store's
listing of the agent's forge/ directory is the set of accounts.
In the agent, hive-agent-forge-accounts (oneshot + 2-minute timer, as
the agent user, under its own store certificate) lists
swarm/agents/<agent>/forge/ with the `list` #4866 grants an agent on its
own metadata subtree, reads each account, and writes
<state>/forge-<label>-token and forge-<label>.json in the names and shape
hive-forge -f already reads. An empty listing (a 404, which `bao kv list
-format=json` answers with `{}` and an empty stderr) is zero accounts; a
denial or an unreachable store fails the unit. It never deletes: files
for labels not listed, including ones the hive wrote, stay as they are.
Removed: the dashboard FORGES tab (credentials.js/html section and its
CSS), hive-c0re's extra_forges.rs and its routes, priv_client's
extra-forge calls, and hive-priv's WriteAgentExtraForgeAccount /
DeleteAgentExtraForgeAccount with their helpers. The GITHUB tab and
WriteAgentGithubToken stay.
Also: persistence.md's matrix avatar note names the exit-75 restart on a
changed account listing, not the dashboard, as what brings a linked
account up.
Refs #4348
hive-matrix-daemon now learns which external matrix accounts it has from
the swarm secret store, under the agent's own certificate, and the hive
push chain for matrix is gone.
The daemon lists swarm/agents/<agent>/matrix/ (the `list` its policy
grants on its own metadata subtree), reads each account's homeserver
from its credential, and brings the accounts up with their tokens from
the store. Every two minutes it lists again and exits with 75 when the
set of linked accounts changed; the unit restarts on 75 without counting
a failure. A listed name whose credential reads as absent is skipped and
logged once. At start it removes the matrix-token-<a> /
matrix-account-<a>.json pairs a hive delivered (a sidecar marks a pair
as delivered; a declared tokenFile keeps its token).
Removed: CredentialNotice and the $SWARM.credential.* subject and NATS
grant, the controller's publish and its queue precondition on the PUT
route, hive-c0re's credential subscription arm and workers/credential.rs,
priv_client::write_agent_matrix_token, hive-priv's WriteAgentMatrixToken
and its helpers, and the daemon's state-dir account discovery.
Kept: WriteAgentGithubToken and the external-forge path
(WriteAgentExtraForgeAccount, extra_forges.rs) are untouched, and a
declared matrixAccounts tokenFile is still read when the store has no
token for that account.
Refs #4348
render_agent gains a second stanza: list on
secret/metadata/swarm/agents/<agent>/*, next to the existing read on
secret/data/swarm/agents/<agent>/*. An agent can now learn which
credentials it holds by listing its own subtree. Metadata read, writes
and every other principal's paths stay refused.
An agent's policy was only written when it was minted, so existing agents
would never get the new stanza. swarm-controller now rewrites every
agent's policy at start (read_policy::ensure_agent_policies), with the
same 30s / 24h retry as ensure_hive_access. The roster is the store's
hive-agent-* cert-auth roles, listed with the controller's existing
`list` on auth/cert/certs; the writes use its existing grant on
sys/policies/acl/hive-*. Only the policy is written: mint_and_verify
also reissues the certificate, so the pass does not call it.
Refs #4348
mara ruled on #4849 (c88934): "remove create_repo tool". The tool ran in
hive-c0re with the hive's core token, so it only ever worked for agents
on the hive that runs the forge.
Removed:
- the create_repo MCP tool and CreateRepoArgs (hive-agent-mcp)
- wire variants Request::CreateRepo and Response::RepoCreated
(hive-core-agent-sock)
- hive-c0re's handle_create_repo, its valid_repo_name check and the
dispatch arm
- forge::create_agent_repo and apply_operator_branch_protection, which
had no other caller, plus AGENTS_ORG and OPERATORS_TEAM, whose only
users they were
- the tool's docs (docs/tools/forge.md repo management, docs/turn-loop/
mcp.md, the conventions tool-group table) and the doc comments that
named it (hive-sock-client's response timeout, ensure_repo_creation_
disabled, the security doc's merge-gate bullet)
ToolGroup::Forge is kept with no tools, the same way b88a5b24 kept
Lifecycle, so existing meta/capabilities.json grants still parse.
Forge state is untouched: existing agents/* repos keep their collaborators
and operators-team branch protection. The swarm-controller's own
create_repo (config-org repos) is a different path and is unchanged.
Closes#4849
An opencode ACP agent got its provider API key only from the hand-placed
backendEnvironmentFile. It now also reads it from the swarm secret store
at swarm/agents/<agent>/acp-provider, field api_key, under its own
certificate, and sets it in the spawned ACP agent's environment only.
Nothing is written to disk.
Precedence: a value already in the process environment (the env file)
wins and the store is not asked. Otherwise the stored key is used when
present. With no store, nothing stored, or a failed read, the agent is
spawned without the key as before, and one line is logged without the
value.
The variable name comes from the existing per-agent option
acp.opencode.provider.apiKeyEnv, exported as HIVE_ACP_API_KEY_ENV on the
harness only for the opencode preset. Other ACP commands are unchanged.
The read lives in hive-runtime, where the ACP child is spawned, so both
hive-agent and hive-subagent-daemon use it. The subagent daemon unit
gets the key name and, when the agent has a store, the agent's store
identity (the same credentials queue-identity.nix gives the harness).
No new option or setting. Closes#4841.
An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.
hive-agent's existing turn-metrics meter records three new instruments:
- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
every backend: the two numbers the web UI's ctx% divides.
The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.
Refs #4845
addresses argus's review on #4842 (c88740): fetch's success path was
silent when the response parsed but windows() found nothing to record.
should_warn_empty tracks the same once-per-streak shape as the
existing token-skip logging, reset the moment a response has windows
again, and never logs the body.
Also notes in the observability doc that a past resets_at means the
paired percent gauge is stale.
A detached task polls GET https://api.anthropic.com/api/oauth/usage every
5 minutes with the OAuth access token from ~/.claude/.credentials.json and
records, per window the response names (five_hour, seven_day,
seven_day_sonnet, ...):
- hyperhive.agent.claude_usage.percent (%, 0-100)
- hyperhive.agent.claude_usage.resets_at (s, unix seconds)
both labelled window=<name>. Endpoint, the anthropic-beta:
oauth-2025-04-20 header and the {utilization, resets_at} window shape are
taken from the claude-code 2.1.283 bundle's own /usage fetch.
The token is only read, never refreshed: claude owns refresh-token
rotation and a second refresher can log the agent out. An expired token
is skipped until claude's next turn refreshes it. API-key agents
(HIVE_USE_API_KEY, the ACP default) and agents with no credentials file
skip quietly, and the task does nothing when OTEL is not configured.
Request failures and non-2xx statuses warn with the status or transport
error only, never the body or the token.
An ACP agent's model and effort pickers now list what its session offers
(its `model` and `thought_level` config options) instead of the claude
model list and EFFORT_LEVELS. A pick goes through the same Bus::set_model /
Bus::set_effort -> Config.model / Config.effort path as on claude; before
each prompt the ACP runtime sets it with `session/set_config_option`, model
first, and only when the session offers that value. Options are re-read
from the set response and from `config_option_update`, so the effort picker
disappears when the chosen model offers no effort levels, and is hidden
while a newly picked model waits for the next turn.
The session's options reach the web UI through a `Choices` handle from a
new `Runtime::choices`, registered on the bus the way `canceller` is.
/api/model and /api/effort accept only offered values on ACP. The claude
path is unchanged.
Refs #4391
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.
Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:
- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao
A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.
module-eval-bao-grants gains one case per consumer.
Refs #4662
An ACP agent's session is now compacted like a claude one: proactively
once a turn crosses the percent-of-window watermark, and on the
operator's /compact or the agent's compact tool. Before, the ACP
backend's compact returned Unsupported and no watermark applied to it.
- AcpRuntime takes the same CompactionPolicy as ClaudeRuntime;
make_session builds one PercentPolicy (with CHECKPOINT_PROMPT) and hands
it to whichever backend runs.
- The runtime keeps the commands each session advertises in
available_commands_update. If `compact` is among them, compaction sends
the prompt `/compact` on the same session, which is how the ACP spec
runs an advertised command. A proactive compaction runs the checkpoint
turn first, as InfiniteSession does.
- With no compact command, the checkpoint turn runs, the session is
archived, and the next turn starts a new one with the system prompt.
- Error::Unsupported had no producer left, so it and drive_turn's
"/compact skipped" arm are gone.
Refs #4391
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.
The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.
Refs #4391
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:
- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
hive-matrix container, Command::Mint and src/mint.rs. The binary, its
appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
role stay. bao-matrix-reader's checks on the deleted unit are removed;
the leaf-identity and no-token-in-env checks now look at
swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
(register/appservice-login/password-login with the local as_token), with
read_appservice_token, paths::matrix_appservice_token and the helpers
only it used. ensure_hive_user now takes the store's token, keeps the
file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
unchanged apart from no longer reading the local as_token.
This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.
Closes#4813Closes#4814
`/api/cancel` stopped a turn only by SIGINTing a child process whose argv0 is
`claude`, so on an ACP agent it reported "no claude process to interrupt"
and the turn ran on. The serve loop now stores the session's canceller on
the `Bus` when the runtime has one, and `/api/cancel` uses it: the agent is
sent `session/cancel`, and the next wake prompt carries the interrupted hint,
as for a signalled claude turn. Without a canceller (claude), the SIGINT path
is untouched.
An ACP turn stopped by the idle watchdog (`HIVE_TURN_IDLE_SECS`) becomes
`TurnError::AgentStall(note)`, handled like `ApiStall`: park for
`HIVE_STALL_SLEEP_SECS`, requeue, and record `api_stall`. The TurnEnd note
is the runtime's message (how long the agent was silent, whether it had to
be killed, and that a silently retried provider error such as HTTP 429
looks like this), not the claude-specific `ApiStall` text. A cancel the
agent ignores is a `Failed` turn.
Refs #4391
`ensure_hive_user` no longer short-circuits on the token file it
already has, and a hive's per-hive store token now wins over the file
whenever it's present. Both upgrade-guide bullets described the old
file-first order and needed rewording to match.
Refs #4427
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
The sweep ran once at startup and was never retried: a boot where
store::connect or the roster read failed (e.g. the controller up
before bao) left the tokens live until the next restart. It now runs
off forge::agent_token::spawn's five-minute mint pass, reusing that
pass's roster observation instead of a second store/roster read, so a
failed first tick retries on the next one. Idempotent, so a re-run
after a partial sweep deletes nothing extra. Drops the one-shot
startup spawn; one call path.
Also updates the forge.md and agent_token.rs docs that still said the
legacy hyperhive-<seconds> tokens stay until manually removed.
Refs #4644
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.
The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.
OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.
The viewer's log message no longer tells the operator to restart it.
Refs #4772
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.
Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.
policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.
Refs #4396
Its only client was hive-c0re's matrix-account-login handler, removed
earlier on this branch, so nothing sends `restart_matrix_daemon` any
more. Drops the `PrivRequest` variant, the hive-priv handler and its
`systemctl --machine=h-<agent> restart hive-matrix-daemon.service`
helper, and the row in the hive-priv op table in security.md.
`PrivRequest` is internally tagged by `op` name, so no other variant's
encoding changes. A new token file still re-fires the daemon through
its `matrix-token*` path unit.
Refs #4348
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.
Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.
The GITHUB tab is the credentials page's default tab now.
Refs #4348
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.
Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.
Refs #4585
authelia's file backend has no self-service reset (no SMTP notifier),
so the only way a human account got a new password after the old one
was forgotten was hand-editing users.yml as root. `user add` already
hashes a password into the file; this verb does the same for an
existing user instead of refusing on the name.
Mirrors `user add`'s UX exactly: no password flag, authelia generates
and hashes it (never crosses argv), and it's printed once and never
stored. Refuses on an unknown user before ever invoking authelia. Same
publish path as add/update, so the same atomic write and no-restart
(authelia watches the file) behaviour apply.
Split the digest-replacement into users::reset_password so it's
testable without a command line or a running authelia, same pattern
as apply_update.