Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/docs/swarm
Author SHA1 Message Date
atlas
8ad2af735e swarm-controller: refuse linking over an existing account
The matrix, forge and github link routes wrote their credential
unconditionally, so linking a name that was already linked replaced the
working account. For matrix that lost the device the agent's crypto store
belongs to (#4838).

Each route now reads the account's store path first and answers 409,
naming the existing account, when something is stored there. Nothing is
written. Replacing an account takes the delete from #4899, then a link.

The matrix route checks before password mode's login, so a refused link
mints no new device at the homeserver.

The check is a read then a write, not an atomic step; two concurrent
links to one name can still both pass it.

Closes #4856
2026-10-03 13:49:03 +02:00
atlas
e21546d08d swarm-controller: cap a subagent terminal stream's disk use
The controller-created term-sub-<agent> stream had max_age only, so a
publishing agent could grow it without bound for 24h. Add a 64 MiB
max_bytes cap with discard: Old (oldest rows drop first, publish never
fails on the cap), and size max_message_size off the queue's live
max_payload rather than a hardcoded guess.
2026-10-03 01:34:01 +02:00
atlas
318f67cda9 swarm-controller: create every agent's subagent stream
swarm-controller now creates `term-sub-<agent>` for every agent a hive is
declared to run, at start and every minute after, with the config
`swarm_queue_client::subagent_term::open_or_create` spells (subjects
`$SWARM.term.<agent>.sub.>`, max_age 24h). An existing stream is opened as
it is, as the controller does for its other streams and buckets, under the
`$JS.API.STREAM.CREATE.*` grant it already holds.

The agent no longer creates the stream: its token is granted publish on
`$SWARM.term.<agent>.sub.>` and no `$JS.API.STREAM.CREATE|INFO` subject,
and the subagent daemon only publishes. A `CREATE` carries the stream's
config in its payload, which no subject grant narrows, so the agent could
otherwise pick the stream's subjects and limits.
2026-10-03 01:34:01 +02:00
atlas
d6f94e5247 swarm: show subagent terminals in the swarm UI
An agent's subagent daemon publishes each subagent's output as terminal
rows on `$SWARM.term.<agent>.sub.<subagent>`, as the agent, into a
per-agent stream it creates itself; swarm-controller lists an agent's
subagents from that stream's subjects and relays one subagent's rows as
SSE; the swarm UI lists them under the agent's terminal preview and
reuses AgentTermPreview, full-screen tab included, with no input.

- swarm-nats.nix: the agent token may also publish
  `$SWARM.term.{agent}.sub.>` and `$JS.API.STREAM.CREATE|INFO` on
  `term-sub-{agent}`, and nothing else of JetStream. A module-eval arm
  pins the agent-token grant as an exact list.
- mcp.nix: hive-subagent-daemon loads the agent's store identity
  (`hive-agent-bao-cert/-key/-server-ca`, the ones hive-agent loads)
  whenever the agent has a store, not only on the opencode preset. The
  agent's own queue secret lives in the store, so this is the credential
  the harness connects with.
- hive-subagent-mcp: `swarm_term` reads the agent's queue secret under
  that identity, connects with the agent token, opens or creates
  `term-sub-<agent>` (max_age 24h), and publishes classified rows from
  the sink every subagent line already passes through. The sink only
  queues (bounded, drop-and-count); a missing store, refused credential,
  failed stream create or failed publish is a log line.
- The stream-json classifier (`stream_enrich`) and the `TermMsg` row
  types plus `fit` move from the hive-agent binary into hive-sh4re, so
  the subagent daemon publishes the rows AgentTermPreview already
  renders. hive-agent keeps its LiveEvent classifier on top.
- swarm-controller: `GET /api/agents/{name}/subagents` and
  `GET /api/agents/{name}/subagents/{subagent}/term/stream`.
- docs/swarm: what the UI shows and what the queue carries.

Closes #4827
2026-10-03 01:34:01 +02:00
atlas
b90be9e65e swarm UI: delete linked accounts — R2 fixes
Addresses argus review comment 90297 on PR #4899:

- swarm-controller/README.md: list the three DELETE routes (including
  matrix's ?revoke=true) beside the PUT/GET ones already documented.
- LinkedAccounts.tsx: a delete answering 404 means the account is
  already gone, so treat it as the delete's end state — re-fetch and
  close the dialog instead of showing an error.
- matrix_account.rs: matrix_logout treats a 401 M_UNKNOWN_TOKEN as the
  token already being revoked and proceeds with the delete; every
  other logout failure still keeps the account. Adds unit tests and
  updates docs/swarm/ui.md to match.
2026-10-03 00:56:39 +02:00
atlas
bbf931207f swarm UI: delete linked accounts
Each row of an agent's linked accounts, except its own `main` matrix
account, gets a delete action. swarm-controller serves DELETE beside each
PUT (matrix-accounts/{account}, forge-accounts/{label}, github-account),
answers 404 for an account the store does not hold, refuses `main`, and
removes every version through `delete_all_versions`.

The matrix confirmation has a revoke checkbox, off by default: the
controller logs the stored token out at its homeserver first, and keeps
the account when that fails or no homeserver is stored.

The controller's policy gains `delete` on each agent's
`metadata/.../matrix/+`, `forge/+` and `github-token`, pinned in
bao-grants.nix.

Refs #4855
2026-10-03 00:56:39 +02:00
atlas
ed9c0f53ed docs: fix argus R4 review nits on config-PR merge docs
- hive-c0re/README.md: drop deleted webhook_secret.rs from the module
  list
- docs/swarm/README.md, docs/agent-lifecycle/approvals.md: rewrite
  temporal wording (legacy/older-release phrasing) as current
  behaviour
- docs/swarm/README.md: note that a lost closed delivery deploys
  nothing and how the operator recovers
2026-10-02 23:36:57 +02:00
atlas
56f592d525 config PRs: each hive removes its own stale config-PR org hook on boot
The hook a hive registered on the agent-configs org points at its
/webhook/config-pr route, which no hive serves any more, so every
config-repo event fails delivery to it. The forge sweep now deletes it.

Only a hook whose URL equals this hive's own exactly is removed. The same
path on another base belongs to another hive and is left alone. A forge
error is logged and boot continues; once the hook is gone the step is a
no-op.

Closes #4850
2026-10-02 23:13:04 +02:00
atlas
a88ed9f24e docs: config changes are operator merges on the forge
Rewrites the config-change flow around the forge merge and the
DeployRequest{rev} deploy, drops the MergeConfigPr approval, its deploy
DAG, the hive's `/webhook/` route and the `core` merge allowlist from
the docs, and states that operators join the `operators` team by hand.

Refs #4850
2026-10-02 23:13:04 +02:00
atlas
a5ea015bc6 config PRs: document the operator merge, deploy only merges into main
docs: the config-repo `main` merge gate (merge = `core` + team
`operators`, approvals = `operators`), the operator merge in the forge
UI and what it deploys, the hand-added `operators` membership, and that
with no eval-verify a failed rebuild leaves `applied/main` at the merged
commit. The swarm README lists the converged gate and the merge deploy.

`merged()` also requires `pull_request.base.ref == "main"`: the hive
deploys its config repo's `main`, so a merge into another branch would
only cost a forge fetch and a refusal comment (argus, #4894).

Refs #4850
2026-10-02 23:13:03 +02:00
atlas
9224c0bd15 swarm UI: fetch linked accounts only for the opened agent
The detail panel makes one request for the agent it shows,
GET /api/hives/{hive}/agents/{agent}/linked-accounts, which returns every
matrix, forge and github account of that agent as names and hosts. The
all-agents route and the table's matrix-column rows are removed, so the
table makes no linked-accounts request. The panel stays keyed by
hive/agent. The bao grant is unchanged.

Refs #4855
2026-10-02 22:36:11 +02:00
atlas
7ccde4647b swarm UI: one request for every agent's linked accounts
GET /api/agents/linked-accounts returns one entry per agent that
/api/agents/status has a row for, as {hive, agent, accounts}, from one
store login. The agents page fetches it once (and again when a link dialog
closes) and hands each table row and the detail panel its agent's slice,
so the page makes no per-agent request. The per-agent route had no caller
left and is removed. The bao grant is unchanged: the same list on each
agent's matrix and forge metadata directories.

Refs #4855
2026-10-02 21:20:40 +02:00
atlas
3380c1915f swarm UI: show the accounts linked to each agent
GET /api/hives/{hive}/agents/{agent}/linked-accounts returns one row per
account linked to the agent, as kind, name and host: each matrix account
under swarm/agents/<agent>/matrix (with its homeserver, and the agent's own
`main` marked reserved), each forge label under swarm/agents/<agent>/forge
(with its url), and github when swarm/agents/<agent>/github-token exists
(host github.com, which is not stored). No credential field is in the
response type.

Listing those two directories needs a new controller grant: `list` on
secret/metadata/swarm/agents/+/matrix and .../+/forge only, pinned in
bao-grants.nix as the only metadata stanzas under agents/ beside the queue
revocation. Checked against a dev OpenBao 2.6.3: the grant lists those two
directories and is refused on agents/, agents/<agent>/, and a leaf.

The swarm UI agent detail panel shows all rows under "accounts"; the table
view's matrix column shows the matrix rows. The link badges stay.

Refs #4855
2026-10-02 21:20:40 +02:00
atlas
6fd91b0e69 docs(swarm-ui): state that every quick-link's option defaults to a value
argus (PR #4895): the link table didn't say an entry can exist without
the service running anywhere in the swarm. Every option in the table
defaults to a value, so a reader shouldn't take an entry's presence as
proof the service is up.
2026-10-02 20:50:15 +02:00
atlas
72e9e1bb4a swarm-ui: build the forge link from swarm.forge.domain
The swarm-controller builds the Forge quick link from
services.hyperhive.swarm.forge.domain, replacing hive-forge/default.nix's
per-host entry, so all seven swarm-service links come from swarm-level
options.

Also drops the remaining references to the removed matrix GUI switch:
the HiveUrls / Urls / hive_urls docs, the hivectl.md `open` note and
the gateway.md vhost-map rows, which name `gatewayHost` instead. The
grafana, victoriametrics and victorialogs modules' comments no longer
mention a quick-link they do not define.

Refs #4885
2026-10-02 20:02:02 +02:00
atlas
6786d54e5a swarm-ui: link the secret store's web UI from swarm.bao.ui.domain
The swarm-controller adds a Bao quick link built from
services.hyperhive.swarm.bao.ui.domain, so the popover links the
store's browser UI whichever host runs bao.

Refs #4885
2026-10-02 20:02:02 +02:00
atlas
f60f8af33b matrix: serve the web client unconditionally; link it from the swarm domain
Removes services.hyperhive.deploy.matrix.gui.enable and its
swarm.matrix.gui.enable alias; both are mkRemovedOptionModule stubs. A
host running the homeserver serves fluffychat at gatewayHost's vhost,
and the hive's /matrix/ redirect follows the same condition.

The swarm-controller builds the Matrix quick link from
swarm.matrix.gatewayHost, replacing hive-matrix.nix's per-host entry.
HIVE_MATRIX_PUBLIC_URL is set on every hive with a gatewayHost, so
`hivectl open matrix` resolves off the homeserver's host too.

Drops HIVE_MATRIX_GUI_ENABLED and the dashboard's matrix_gui_enabled
field; nothing in the frontend reads it.

Refs #4885
2026-10-02 20:02:02 +02:00
atlas
8628e0ecdd swarm-ui: build authelia/grafana/metrics/logs links from swarm domains
The swarm-controller module builds the Authelia, Grafana, Metrics and
Logs quick links from services.hyperhive.swarm.<service>.domain, on the
controller's host, instead of each service module adding its entry only
on the host that runs it. A controller whose swarm runs those services
on other hosts lists them in its /api/links popover.

Forge, matrix and bao links are not moved yet: forge waits on #4891,
matrix and bao on whether their GUI gate becomes swarm-level.

Refs #4885
2026-10-02 19:45:02 +02:00
atlas
f63ab954c9 Merge remote-tracking branch 'forge/main' into docs/networking-scheduler-pass
# Conflicts:
#	docs/networking/gateway.md
2026-10-02 19:00:42 +02:00
atlas
7488c337cc docs(swarm): drop stale github-token exception in credentials.md
The sentence claiming a per-agent github token sits outside this page
and never passes through the store contradicted the
swarm/agents/<agent>/github-token row already in the table: the token
is minted by swarm-controller and read by hive-agent-github-token
through bao like every other row.

Refs #4347
2026-10-02 17:56:06 +02:00
atlas
8e23feb01b github: PATs live in swarm bao; the agent fetches them itself
An operator links an agent's GitHub personal access token in the swarm UI
(LinkGithubAccountForm, "link github account" on /agents). swarm-controller's
PUT /api/hives/{hive}/agents/{agent}/github-account stores it at
swarm/agents/<agent>/github-token (swarm_secret_client::github), a flat leaf
under the agent's prefix that the agent's existing read grant already covers:
no policy change, and no list grant, since there is one token per agent.

In the agent, hive-agent-github-token (oneshot + 2-minute timer, as the agent
user, under its own store certificate, ordered before hive-github-notify)
reads that path and writes <state>/github-token, 0600 and agent-owned, the
file the gh wrapper, git credential helper and hive-github-notify already
read. It replaces the file by rename only when the bytes changed and never
deletes it: a hive-written github-token stays until a token is linked in the
swarm UI. It is installed only with a store address and
services.hyperhive.agent.github.enable.

Removed: the dashboard's CR3D3NTIALS page (credentials.html/js/css, its
build entries and H0M3 tile; GITHUB was its only tab), hive-c0re's
dashboard/matrix_accounts.rs with GET/POST /api/github-account,
priv_client::write_agent_github_token, the host socket's
SetAgentGithubToken and `hivectl github set-token`, and hive-priv's
WriteAgentGithubToken with write_agent_state_file, its only caller gone.

Docs: integrations/github.md and swarm/ui.md describe the swarm path,
swarm/credentials.md gains the store-path row, and the hive UI docs,
hivectl docs and security.md's hive-priv table drop the removed pieces.

Closes #4347
2026-10-02 17:48:27 +02:00
atlas
5ec658e0fa docs(networking): drop remaining stale authelia-empty-user claims
error-pages.nix paragraph (gateway.md:433) blamed a dead authelia
upstream on an empty user set; the real reason the route earns a
custom page is that a bare 502 there blames the proxy while the
gateway itself is fine. gateway.md:38 dropped 'yet' from the
placeholder-while-empty phrasing. services.md:109 corrected
'seeds an empty users database' to the disabled placeholder subject
swarm-authelia.nix actually seeds (swarm-authelia.nix:873).

Refs #3902
2026-10-02 17:25:37 +02:00
atlas
cabde572e9 docs(web-ui): scope dashboard.md to what the hive UI renders
Per mara's review: the hive UI doc covers only what hive-c0re's pages
render. Removed the M4TR1X page section (the hive gateway redirects
/matrix/ to the swarm matrix client, which the swarm UI's quick links
open), the swarm-UI forge/matrix account-linking lines from the
CR3D3NTIALS section, the infra-services hivectl paragraph, and the H0M3
Matrix/Forge absence line. Added a single pointer to docs/swarm/ui.md,
and stated the account-linking and Matrix quick-link facts there.

Refs #3902
2026-10-02 14:44:34 +02:00
atlas
fb2fff0668 fix(nix): require swarm domain only when a hyperhive service is enabled
The swarm.domain assertion in hive-network.nix fired on every host that
imported the module, so a host that enables nothing failed eval. It now
fires only when one of the hyperhive service switches is on (every
deploy.*.enable that runs something, gateway, gateway.dns, network,
otel, snapshotStore). The requirement itself is unchanged: any host that
runs a hyperhive service still needs swarm.domain.

The core-toggle module-eval suite gains a case: a missing swarm.domain is
refused on a hive and on a swarm-service-only host, and a host enabling
nothing passes every assertion.

Closes #4887
2026-10-02 13:09:45 +02:00
atlas
99905f50b0 nix(authelia): start with a disabled placeholder user when the user set is empty
authelia 4.39.20 exits at startup on `users: {}` ("users: non zero value
required"), and the first-boot unit seeded exactly that, so a swarm with
no users crash-looped authelia and answered 502 until `swarmctl user add`
ran.

The first-boot unit now writes one subject, `swarm.placeholder`, when
the users database is absent, empty, or exactly `users: {}`:

- `disabled: true` — authelia returns "user not found" for a disabled
  user before any password check (file_user_provider.go,
  CheckUserPassword).
- password: an argon2id digest with an all-zero key. It decodes (authelia
  rejects a non-digest at startup) and no known password hashes to it.
- the `.` keeps it out of agent names (`[a-z0-9-]`), and `swarmctl user
  add` refuses it as already existing. Neither writer removes users, and
  both round-trip `disabled`.

A file with any user in it is never touched.

The docs that described the crash-loop (sso.md, gateway.md, setup.md,
the sso-unavailable error page) now describe the placeholder; the
writers' load_store docs and the seed fixtures follow. module-eval
nats-authelia asserts the seed branch.
2026-10-02 12:50:47 +02:00
atlas
7b226f21dd docs(swarm): state behaviour instead of denying absent options 2026-10-02 12:50:34 +02:00
atlas
9c52809439 docs(swarm): drop statements about absent fields 2026-10-02 12:50:34 +02:00
atlas
ff2aa1e27b docs(swarm): state requirements, drop upgrade narration 2026-10-02 12:50:34 +02:00
atlas
4a7a1c341b docs(swarm): drop mentions of nonexistent fields 2026-10-02 12:50:34 +02:00
atlas
6ed61da8b8 docs(swarm): state swarm domain as required; drop changelog wording
Refs #3902
2026-10-02 12:50:34 +02:00
atlas
270430a4b4 docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
2026-10-02 12:50:34 +02:00
atlas
9545151b8d docs: state current behaviour, drop change-log wording
Refs #3902
2026-10-02 11:43:10 +02:00
atlas
9d804ae094 docs: fix vale errors on main
Fix the 4 pre-existing vale error-level hits on main (docs/README.md:83,
docs/getting-started/setup.md:11,127, docs/swarm/bao.md:4) that fail CI's
prose-lint-errors job for every docs PR regardless of its own diff.
2026-10-01 23:34:47 +02:00
müde
a7991c9242 docs: setup.md as a short all-local checklist; bao internals move to swarm/bao.md 2026-10-01 20:41:47 +02:00
atlas
2c7e586f47 forge: external forge accounts live in swarm bao; the agent fetches them itself
An operator now links an agent's external forge account (label, base URL,
token) in the swarm UI. swarm-controller stores it at
swarm/agents/<agent>/forge/<label>. There is no index: the store's
listing of the agent's forge/ directory is the set of accounts.

In the agent, hive-agent-forge-accounts (oneshot + 2-minute timer, as
the agent user, under its own store certificate) lists
swarm/agents/<agent>/forge/ with the `list` #4866 grants an agent on its
own metadata subtree, reads each account, and writes
<state>/forge-<label>-token and forge-<label>.json in the names and shape
hive-forge -f already reads. An empty listing (a 404, which `bao kv list
-format=json` answers with `{}` and an empty stderr) is zero accounts; a
denial or an unreachable store fails the unit. It never deletes: files
for labels not listed, including ones the hive wrote, stay as they are.

Removed: the dashboard FORGES tab (credentials.js/html section and its
CSS), hive-c0re's extra_forges.rs and its routes, priv_client's
extra-forge calls, and hive-priv's WriteAgentExtraForgeAccount /
DeleteAgentExtraForgeAccount with their helpers. The GITHUB tab and
WriteAgentGithubToken stay.

Also: persistence.md's matrix avatar note names the exit-75 restart on a
changed account listing, not the dashboard, as what brings a linked
account up.

Refs #4348
2026-10-01 18:05:33 +02:00
atlas
e04616eb70 swarm-secret-client: agents may list their own subtree; controller rewrites agent policies
render_agent gains a second stanza: list on
secret/metadata/swarm/agents/<agent>/*, next to the existing read on
secret/data/swarm/agents/<agent>/*. An agent can now learn which
credentials it holds by listing its own subtree. Metadata read, writes
and every other principal's paths stay refused.

An agent's policy was only written when it was minted, so existing agents
would never get the new stanza. swarm-controller now rewrites every
agent's policy at start (read_policy::ensure_agent_policies), with the
same 30s / 24h retry as ensure_hive_access. The roster is the store's
hive-agent-* cert-auth roles, listed with the controller's existing
`list` on auth/cert/certs; the writes use its existing grant on
sys/policies/acl/hive-*. Only the policy is written: mint_and_verify
also reissues the certificate, so the pass does not call it.

Refs #4348
2026-10-01 17:43:28 +02:00
atlas
7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00
atlas
78d8d69c7f swarm-controller: mint each hive's matrix sender token
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.

swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.

swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.

hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.

The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
22c96282b9 swarm-bao: re-run the operator viewer unit when the granter step succeeds
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.

The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.

OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.

The viewer's log message no longer tells the operator to restart it.

Refs #4772
2026-09-29 17:45:58 +02:00
atlas
69ae23f801 swarmctl: add user reset-password
authelia's file backend has no self-service reset (no SMTP notifier),
so the only way a human account got a new password after the old one
was forgotten was hand-editing users.yml as root. `user add` already
hashes a password into the file; this verb does the same for an
existing user instead of refusing on the name.

Mirrors `user add`'s UX exactly: no password flag, authelia generates
and hashes it (never crosses argv), and it's printed once and never
stored. Refuses on an unknown user before ever invoking authelia. Same
publish path as add/update, so the same atomic write and no-restart
(authelia watches the file) behaviour apply.

Split the digest-replacement into users::reset_password so it's
testable without a command line or a running authelia, same pattern
as apply_update.
2026-09-29 12:31:21 +02:00
atlas
ccb5bd3b38 hive-agent: read the per-agent queue secret from bao in process
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.

hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.

A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).

Closes #4783
2026-09-29 10:18:07 +02:00
atlas
c21ec7719d swarm: tests and docs for the queue-credential revocation
Pins the three properties the revocation rests on and cannot check
against a store: that only `Destroyed` revokes (a revocation on
`Offline` or `Paused` would give an agent that stops and never
restarts), that a 404 is absence while a 403 stays a failure, and that
the delete addresses `secret/metadata/` -- the path that takes every
version, which is the string the grant has to match.

docs/swarm/credentials.md gains the revocation section and its table
cell stops describing the deletion as something an operator does by
hand.
2026-09-28 23:46:17 +02:00
atlas
7bc4b25f16 swarm-controller: backfill a missing queue-secret mint time instead of re-minting it
A stored queue secret with no `minted_at` counted as due, and no secret
minted before the renewal pass has one, so the first pass after deploy
would re-mint every agent's secret. Every reconnect before that agent's
next restart would then be refused.

Such a secret is now stamped instead: `minted_at = now` is written beside
the unchanged `value`, and its 45-day clock starts there. Only a secret
whose recorded mint time is at least 45 days old gets a new value.

The decision is `secret_step` (Keep / Backfill / Remint), and the pass
reports an unstamped secret as `Observed::Unstamped`. Backfill and
re-mint log different lines.
2026-09-28 22:24:07 +02:00
atlas
2115ec2bb3 swarm-controller: re-issue agent certificates and re-mint queue secrets at half-life
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:

- `MintAgentIdentity` (the node agent creation uses) when the stored
  certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
  read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
  swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
  re-decides, writes a fresh value with `minted_at`, reads it back, and logs
  the agent and the old age.

When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.

`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.

Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.

Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.

docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
2026-09-28 21:35:01 +02:00
atlas
b14ff2796c bao: OIDC login to the browser UI via authelia, as a metadata-only viewer
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.

- authelia registers an interactive client `swarm-bao-ui`
  (glue-bao-ui-oidc-client.nix) with redirect
  `https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
  publisher carries its secret to
  `secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
  mount with listing visibility `unauth`, asked before attempted like
  cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
  and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
  policy, and once the granter may configure `auth/oidc/config` (checked
  through `sys/capabilities-self`), writes the mount's config from the
  published secret and the role binding `groups=admins` to the viewer.
  Before the bootstrap step re-runs it writes the policy, logs the step
  and exits 0.

Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.

module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
2026-09-28 19:56:38 +02:00
atlas
3b27a2e5a1 docs: fix vale prose-lint errors in bao UI docs
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Passive voice and sentence-initial 'So' hits from PR #4776's vale error job.
2026-09-28 19:31:02 +02:00
atlas
3898ca33c7 bao: serve the browser UI to admins via a loopback-only listener
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.

The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.

Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.

Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.

On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
2026-09-28 19:31:02 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
3b8df27dcf swarm-ui: show each agent's icon on its card
Each agent card now leads with the agent's icon, loaded as an `<img>`
from `GET /api/agents/<name>/icon`: the same 5em square, background and
fallback as the hive dashboard's container row. An agent with no icon
(the route's 404), or any other failed load, shows the dimmed hyperhive
mark (`/favicon.svg`) instead of a broken image.

Only ever an `<img>`, never inline markup: the body is an agent-authored
SVG, and an image load does not run its script.
2026-09-28 13:47:37 +02:00