glue-swarm-bao-otel-oidc-client.nix:34 still pointed clientId's
declaration at ./swarm-bao.nix after the split moved it to
./swarm-bao-service.nix. Line 17, which points the config block's
deploy.bao.enable gate at ./swarm-bao.nix, is unchanged -- that part
stayed.
`swarm.bao` (what the secret store is to every hive: container name,
domain, UI domain and OIDC client, port, collector client id and
telemetry port) moves to nix/host-modules/swarm-bao-service.nix, together
with `domainBase`, the only helper it reads besides `cfg`. Everything
else -- the `deploy.bao` options, the removed-option import, the whole
`config` block including `containers.swarm-bao`, and the helpers only
they read -- stays in nix/host-modules/swarm-bao.nix, which default.nix
now imports alongside the new file.
Both halves read `cfg` (`swarm.bao.ui.oidc.redirectUri` defaults from
`cfg.ui.domain`; the config block reads `cfg` throughout). It is an
option read, so each file binds it from `config.services.hyperhive.swarm.bao`.
The service file has no `hyperhiveCfg`, so its `swarmDomain` reads
`config.services.hyperhive.swarm.domain` directly, as
swarm-nats-service.nix does.
A pure move: option paths, option definitions and config are unchanged
apart from the comment above `deploy.bao`, which now names the file
`swarm.bao` lives in, and the pointer in swarm-nats-service.nix to the
`domainBase` rationale, which moved with it.
Refs #3742
`swarm.grafana` (what the metrics UI is to every hive: container name,
domain, metrics port, OIDC client) moves to
nix/host-modules/swarm-grafana-service.nix, together with `domainBase`,
the only helper it reads besides `cfg`. Everything else -- the
`deploy.grafana` options, the whole `config` block including
`containers.swarm-grafana`, and the helpers only they read -- stays in
nix/host-modules/swarm-grafana.nix, which default.nix now imports
alongside the new file.
Both halves read `cfg` (`swarm.grafana.oidc.redirectUri` defaults from
`cfg.domain`; the config block reads `cfg` throughout). It is an option
read, so each file binds it from `config.services.hyperhive.swarm.grafana`.
The service file has no `hyperhiveCfg`, so its `swarmDomain` reads
`config.services.hyperhive.swarm.domain` directly, as
swarm-nats-service.nix does.
A pure move: option paths, option definitions and config are unchanged
apart from the two comments on either side of the cut, which now name the
file the other half lives in.
Refs #3742
`swarm.nats` (what the queue is to every hive: domain, ports, client id)
moves to nix/host-modules/swarm-nats-service.nix, together with the only
two helpers it reads, `swarmDomain` and `domainBase`. Everything else --
the `deploy.nats` options, the whole `config` block including
`containers.swarm-nats`, and the helpers only they read -- stays in
nix/host-modules/swarm-nats.nix, which default.nix now imports alongside
the new file.
A pure move: option paths, option definitions and config are unchanged
apart from the transition comment above `deploy.nats`, which now names the
file `swarm.nats` lives in. The nats fixtures evaluate to the same host and
container toplevel derivations before and after.
Refs #3742
The swarm collector's journald receiver read only the units listed in
`services.hyperhive.swarm.otel.journaldUnits`. A unit nobody listed
never reached the store, and a misspelt entry shipped nothing without
an error. The list existed to keep an operator's desktop session out of
a store every swarm operator can read, but the receiver can only match
positively, so the only way to express "not user sessions" was to name
every service instead.
The receiver now reads the whole host journal, and a new
`filter/exclude-user-sessions` processor in the `logs/<swarm>` pipeline
drops records whose `_SYSTEMD_SLICE` is `user-<uid>.slice` (session
scopes and `user@<uid>.service`). The per-hive `logs/<hive>` pipelines
carry agent-container journals only and get no filter.
`journaldUnits` is removed with `mkRemovedOptionModule`, together with
its non-empty assertion and the entry each host module added. The four
module-eval membership checks go with it, replaced by one structural
case in swarm-otel-core.
Closes#3646
An opencode ACP agent got its provider API key only from the hand-placed
backendEnvironmentFile. It now also reads it from the swarm secret store
at swarm/agents/<agent>/acp-provider, field api_key, under its own
certificate, and sets it in the spawned ACP agent's environment only.
Nothing is written to disk.
Precedence: a value already in the process environment (the env file)
wins and the store is not asked. Otherwise the stored key is used when
present. With no store, nothing stored, or a failed read, the agent is
spawned without the key as before, and one line is logged without the
value.
The variable name comes from the existing per-agent option
acp.opencode.provider.apiKeyEnv, exported as HIVE_ACP_API_KEY_ENV on the
harness only for the opencode preset. Other ACP commands are unchanged.
The read lives in hive-runtime, where the ACP child is spawned, so both
hive-agent and hive-subagent-daemon use it. The subagent daemon unit
gets the key name and, when the agent has a store, the agent's store
identity (the same credentials queue-identity.nix gives the harness).
No new option or setting. Closes#4841.
An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.
hive-agent's existing turn-metrics meter records three new instruments:
- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
every backend: the two numbers the web UI's ctx% divides.
The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.
Refs #4845
forge-avatar-sync passed the base64-encoded icon as a jq --arg and then
as a curl -d command-line argument. Linux caps a single exec argument at
MAX_ARG_STRLEN (128 KiB), so any icon whose rasterized PNG base64-encodes
past that (red's does, at 133300 bytes) makes jq fail with E2BIG before
curl is ever reached, and the avatar upload silently never happens. The
base64 and the JSON payload now go through temp files instead of argv.
Closes#4839
Adds rules to the `hyperhive` alerting folder next to the two from #4811:
- swarm queue connect failed (per hive/agent, 1h window, fires at once:
the line is logged once per agent process and the connection is never
retried, so the window is how long the rule stays firing)
- swarm agent state publish skipped (per hive/agent, 15m, for 5m, as the
swarm terminal rule)
- swarm queue credential missing (per hive/agent, 15m, for 5m)
- agent credential renewal failing (10m, for 15m, as the forge reconcile
rule: same 5m pass cadence, one line per failed pass)
- no log records from <hive>, one per `swarm.hives` entry: fires when
the hive's ungrouped count over 10m is below 1, and on no data. A new
`logAbsenceRule` helper wraps `logCountRule` with the inverted
threshold and noDataState = Alerting.
No contact point and no notification policy: the rules show under
Alerting -> Alert rules and are delivered nowhere.
Refs #4717
Refs #3900
Filter the ACP model picker's list by services.hyperhive.agent.availableModels: an
unconfigured agent (env absent) shows every model the session offers, a
configured list narrows the picker to whatever it names that the session
also offers (in the session's own order), and a configured list matching
none of the session's models (the claude names on an ACP agent that never
touched the option) falls back to showing everything, with one warning
naming the mismatch.
The nix option's default, the model-vs-availableModels build assertion and
hive-subagent-mcp's check_model rail are unchanged — this only touches the
web UI's picker.
The restart decision was a shell variable set by comparing the fetched
value with the file just before overwriting it. A run that wrote the
file and then failed before the restart (the matrix unit's registration
render, or `systemctl --machine` finding no bus yet) left a retry that
saw an unchanged file and never restarted the consumer.
The file is now written only when the value differs, so its mtime marks
the last real change, and `refresh_consumer <machine> <unit> <path>`
compares that mtime with the consumer's ActiveEnterTimestamp on every
run, the shape the openbao client-CA refresh in swarm-bao.nix already
uses. A consumer that started after the last change is left alone; a
running one is try-restarted, a failed one reset and started, all with
--no-block, and nothing happens while the container is down.
The helper's comment block also exceeded the 30-line limit
(`comment-block lint` failed on d871467d); its per-function notes now
sit beside the functions.
module-eval-bao-grants asserts the gated write, the path the refresh is
keyed on, and the mtime-vs-start comparison for each consumer.
Refs #4662
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.
Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:
- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao
A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.
module-eval-bao-grants gains one case per consumer.
Refs #4662
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.
- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
hive-agent-bao-identity and hive-agent-forge-token use it.
queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
forge-token.nix is a fetch unit with the same short budget that was
added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
(--no-block) a loaded nginx that is not active; an active nginx keeps
the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
fetch unit, swarm-services-cert included.
Refs #4662
A plain removal breaks any out-of-tree host config that still sets the
option: eval fails with "option does not exist" and no pointer to what
replaced it. mkRemovedOptionModule gives a clear evaluation error instead.
swarm-controller is the only minter of swarm/hives/<hive>/matrix/sender-token
since #4820, so the swarm-matrix-ctl policy's first stanza granted a write no
code performs. The policy keeps its one used stanza, the swarm appservice token
that `swarm-matrix-ctl appservice publish` writes. deploy.bao.matrixCtlHiveName
only named the hive in the dropped stanza and goes with it.
hive-matrix.nix no longer calls the per-hive appservice's as_token hive-c0re's
authority: tuwunel loads the registration and creates the sender account, and
no client presents that token.
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.
The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.
Refs #4391
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:
- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
hive-matrix container, Command::Mint and src/mint.rs. The binary, its
appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
role stay. bao-matrix-reader's checks on the deleted unit are removed;
the leaf-identity and no-token-in-env checks now look at
swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
(register/appservice-login/password-login with the local as_token), with
read_appservice_token, paths::matrix_appservice_token and the helpers
only it used. ensure_hive_user now takes the store's token, keeps the
file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
unchanged apart from no longer reading the local as_token.
This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.
Closes#4813Closes#4814
services.hyperhive.agent.runtime ("claude" default | "acp") and
acp.{command,args,env}, rendered into HIVE_RUNTIME / HIVE_ACP_* only
for acp, so a claude agent's unit is unchanged. acp implies useApiKey.
acp.presets.opencode runs `opencode acp` from nixpkgs against an
OpenAI-compatible provider from acp.opencode.{provider,model,
contextWindow,outputLimit}: the config is rendered to the store with
the API key as an {env:VAR} reference, so the key is read at runtime
from backendEnvironmentFile. OPENCODE_PERMISSION denies opencode's
built-in bash, task, todowrite and websearch, and makes webfetch ask.
Refs #4391
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
Two Grafana-managed rules in a "hyperhive" folder, one group evaluated
every 1m, each a LogsQL stats count over the VictoriaLogs datasource:
- forge reconcile failing: `swarm forge objects: write failed; retrying
next pass` over 10m, for 15m. The pass runs every 5m, so one failed
pass stays under `for` and a pass failing on every tick fires.
- swarm terminal publish skipped: `swarm terminal: publish skipped` over
15m, for 5m, one instance per hive/agent.
No contact point or notification policy. Grafana 13 routes to its
built-in `empty` receiver when none is provisioned, so firing rules
are visible under Alerting and sent nowhere, with no send errors.
Refs #4717
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.
hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.
Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.
module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.
Refs #3773
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.
`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).
swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.
Refs #3773
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.
The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.
OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.
The viewer's log message no longer tells the operator to restart it.
Refs #4772
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.
Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.
Refs #4585
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.
hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.
A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).
Closes#4783
`script-test-agent-bao-fetch` executes the rendered `ExecStart` of
`hive-agent-forge-token` and `hive-agent-queue-credential` under
`umask 0377`, with a stub `bao` first on the unit's own PATH. It covers
every error branch, the happy path, forge rotation/unchanged, and 0400
files already in place: the redirect failure #4736 fixed, whose live
symptom was `bao.err: Permission denied` reported as a refused
certificate. The TLS-alert fixture is the `unknown certificate
authority` error h-atlas's identity check got from the store.
#4736 claimed this test but never committed it; the two module-eval
suites for these units only evaluate the config.
Closes#4748
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.
Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.
Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.
Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.
Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.
The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.
Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.
nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:
- `MintAgentIdentity` (the node agent creation uses) when the stored
certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
re-decides, writes a fresh value with `minted_at`, reads it back, and logs
the agent and the old age.
When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.
`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.
Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.
Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.
docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.
- authelia registers an interactive client `swarm-bao-ui`
(glue-bao-ui-oidc-client.nix) with redirect
`https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
publisher carries its secret to
`secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
mount with listing visibility `unauth`, asked before attempted like
cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
policy, and once the granter may configure `auth/oidc/config` (checked
through `sys/capabilities-self`), writes the mount's config from the
published secret and the role binding `groups=admins` to the viewer.
Before the bootstrap step re-runs it writes the policy, logs the step
and exits 0.
Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.
module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.
The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.
Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.
Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.
On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.
Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.
Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.
Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.
hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.
swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.
Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
c9ba9062 replaced 5 bargauge panels with one table panel using a
joinByField+organize+sortBy transformations chain and a
vizConfig.group="table" viz spec. That construct had zero precedent
anywhere else in this repo's dashboards and its live render was
explicitly flagged as unverified at merge time (no in-container way
to render Grafana to confirm).
#4658 reports a dashboard crash: "'' not found in: reduce,
filterFieldsByName, ... transpose" - a Grafana transform-registry
lookup failing on an empty transformer id. That registry's full id
list matches the error text exactly, and the table panel is the only
panel in the whole swarm-grafana/dashboards tree using a non-empty
transformations array, so it's the leading suspect: some part of how
the dashboard-v2beta1 schema expects a QueryGroup's transformations
to be shaped likely differs from the classic {id, options} shape used
here, and the file-based provisioner may not be normalizing the
mismatch the same way a schema-v2-aware loader does.
Reverting to the prior, previously-uneventful bargauge panels
(byte-for-byte c9ba9062^'s version of this section) to stop the
active crash while the correct v2beta1 transformation shape gets
verified against a live Grafana render instead of guessed at again.
swarm-bao-granter-role used its token for `bao auth list` with no check,
so an expired, revoked or policy-less token in bootstrapTokenFile exited
2 with a raw 403 and no hint. It now checks whether the store is up when
that call fails: if it is, the token is at fault, and the unit prints the
one-time bootstrap step and exits 4. A missing token file is still a
ConditionPathExists skip, so the two read differently in the journal.
The twelve granterLogin units printed "the store is sealed or
unreachable" on a healthy store, because `bao status` exited 1 there:
the CLI resolves a token helper under $HOME before asking, systemd sets
no HOME for a unit without User=, and the fallback shells out to
`getent`/`sh`, neither of which is on the unit's PATH ("failed to get
token helper: error expanding config path "": exec: "sh": executable
file not found in $PATH"). The check now runs with HOME=/var/empty and
keeps its stderr, so a genuinely unreachable store says why. When the
store is up and the login is refused, the units now name
swarm-bao-granter-role as the unit that writes the missing role.
setup.md's post-step restart used 'swarm-bao-*-policy.service', which
misses swarm-bao-agent-pki. It now names that unit too, and a
module-eval case fails when the restart misses any unit that logs in as
the granter.
Refs #4704
matrixHomeserverUrl re-derived gatewayHost's own default
(chat.${swarmDomain}) instead of reading gatewayHost itself, so a
deployment that pins gatewayHost away from that default (the exact
case hive-matrix.nix's own option doc describes) left the controller
calling a name nothing serves. Read matrix.gatewayHost directly,
mirroring hive-matrix.nix's ctlHomeserverUrl. Add module-eval cases
that pin gatewayHost and that null it out, alongside the existing
unpinned default case.
Closes#4757
An `auth_token` spelled `swarm-agent.<agent>.<secret>` is no longer sent
to introspection. The responder reads `swarm/agents/<agent>/queue` with
an identity of its own, checks that the stored object names the same
agent, compares the secret in constant time, and grants the subjects
`--agent-token-publish-subject` lists with `{agent}` expanded. Every
other outcome denies: a malformed token, no store identity, nothing
stored, a failed or slow lookup, a different secret. A token without
the prefix takes the OIDC path unchanged.
The journal's `auth request` line names such a caller `agent:<agent>`;
the hive-shared credential keeps `hive-<h>-agent`.
The new principal: a `swarm-nats-auth` cert-auth role and policy with
read on `secret/data/swarm/agents/+/queue` alone, a leaf signed by the
store's PKI glue, and `glue-nats-auth-bao-identity.nix` pairing the two.
The copy unit delivers the identity into the queue's container, and an
absent leaf is delivered empty so the responder still starts and only
agent tokens are refused.
The policy and role are written by `swarm-bao-nats-auth-policy`, logged in
as the bao granter: both names fall under its `swarm-*` globs, so the
deploy writes them with no operator step. module-eval counts it among the
granting units, so every generic granting-unit case covers it.
The secret compare uses `subtle`, already in the lock file through the
TLS stack; no workspace crate offered one directly.
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.
- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
tunes the mount, generates the root once (guarded on an empty issuer
list, no replace branch), upserts the `swarm-agent` role (client
certificates named `hive-agent-*` only, 90 days), caches the CA at
/var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
(client-ca.pem, then the agent CA). Host cert-auth roles still pin
client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
changed after openbao started, swarm-bao-agent-pki restarts
openbao.service in the container; under `seal = "shamir"` it prints
the step instead. Once swarm-bao-certs has a cached CA, later boots
start openbao with it and do not restart.
- The controller policy gains exactly `update` on
pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
the leaf (the store generates the key), writes the agent's cert-auth
role pinning the issuing CA bao returned, and writes the agent's
policy as render_agent alone: the hive-shared queue credential stanza
is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
other pki role names); swarm-controller gets
SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.
Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.
Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.
Closes#4756
#4756 moves agent client certificates onto a PKI mount of their own,
`pki-agents`, whose root bao generates internally. The unit that sets
that mount up runs as the bao granter, and the granter's policy is only
written while #4754's one-time bootstrap token is in place. Adding these
grants after an operator has done that step would cost a second token
placement, so they go into the granter's policy here, before it.
Six stanzas: enable and tune the mount, list its issuers, read its CA,
generate its root internally, and write `swarm-*` roles on it. No root
delete or sudo: agent cert-auth roles pin that root by value, so
replacing it must not be something a deploy can do.
Adds deploy.bao.agentPkiMountPath (default `pki-agents`), which the
stanzas are rendered from. The module-eval case pinning the granter's
policy now lists all seventeen stanzas, and the "grants nothing outside"
case also refuses the agent mount's root, issue, sign and a roles/*
glob.
Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).
A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.
swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.
The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.
module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.
Refs #4704
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).
- hive-priv gains `EnsureAgentSocketDir { name }`, called from
`set_nspawn_flags` in every start path. It creates
`/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
directory is left alone. hive-c0re's own create_dir_all went: its /run
is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
the agent user and sets 0751, the same way it already handles state/ and
harness/. It refuses a symlink or non-directory there, since `test -d`
and chmod follow links. No host-side passwd parse, and no window where
the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
(`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
`converge_start_preamble` + `start_with_fallback`. It was a bare start,
so after a reboot the manager's bind sources existed only because of the
tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
`priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
`SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
does the same unlink and returns Ok, for an older hive-c0re.
Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.
Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.
Closes#4742
hive-agent-forge-token and hive-agent-queue-credential both run with
UMask=0377. Their scripts captured bao's stderr in `err="$(mktemp)"`,
which under that umask is created 0400; the very next `2>"$err"` on the
`bao login` line cannot reopen it for writing, so bash fails the
redirect with "Permission denied" before bao ever runs. The `if !`
around the login then took the only error branch it had and printed
"this agent's certificate was refused by the swarm secret store" — the
store was never contacted. No agent has fetched either credential.
The stderr file now lives in each unit's own 0700 RuntimeDirectory and
is removed before every redirect into it, so the redirect creates it —
the idiom forge-token.nix already used for its staging file.
The login's error branch now says which of these happened, then quotes
bao's output:
- `$err` could not be created, so bao never ran;
- the store answered with HTTP 4xx (refusal) or another status;
- the store sent a TLS alert rejecting the certificate;
- no answer at all (network, DNS, or local TLS).
Unreadable cert/key credentials are reported before bao runs.
bao.nix has the same fetch shape but no UMask=, so its mktemp file is
0600 and writable; it is untouched.
Closes#4735
The header comment grew to 39 lines across two audit-driven rounds,
over the 30-line comment-block-lint max. Trimmed to 30: merged the
value-as-argument rationale with its /proc/cmdline justification into
one paragraph, cut the usage example to one call instead of two, and
condensed the args paragraph — no content dropped, just restatement.
persistence.md's first-boot-migration marker paragraph used "it is"
and "there is not" — vale's Microsoft.Contractions rule (this repo's
config) wants the contracted forms, and "is not" also collides with
"is nothing" as a literal substring, which is what actually tripped
the error. Reworded to "it's" / "there's nothing", no meaning change.
Lint-only: no script logic changed, gates re-run below are all lint
checks (no module-eval, no cargo).
Refs #4723
The pipe contract had a gap: if a producer piped into
atomic_write_secret exited non-zero after writing partial output,
cat still saw a clean EOF and wrote that partial content through to
the live target via mv — pipefail only reported the failure
afterward, once the bad write was already committed. The helper now
takes the value as its 4th argument and writes it itself with printf
(a shell builtin, so the value never touches an external process's
own argv/environ, same as a function argument never does), so there
is no pipe left to fail silently.
Callers that compute the value with a command now capture it into a
variable first (`value=$(cmd)`), which fails under `set -e` before
atomic_write_secret is ever called — swarm-bao.nix's pin.env site is
the one that needed this (`pin_env_value="BAO_HSM_PIN=$(cat ...)"`).
All seven call sites converted; output is byte-identical (same
printf '%s\n' framing, now applied inside the helper instead of by
each caller).
Refs #4723
swarm-bao-token's pin.env write (the BAO_HSM_PIN EnvironmentFile for
openbao's pkcs11 seal) had the same write-then-chmod-on-live-path
shape as the sites already converted: printf > path directly on the
live file, chmod after. Same fix, same helper. Content
("BAO_HSM_PIN=<user-pin>\n") and final mode (0400, root-owned — no
chown, same as before) are unchanged; pin.env stays at the same path,
so openbao's EnvironmentFile= reference needs no change.
Refs #4723