An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.
hive-agent's existing turn-metrics meter records three new instruments:
- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
every backend: the two numbers the web UI's ctx% divides.
The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.
Refs #4845
Adds rules to the `hyperhive` alerting folder next to the two from #4811:
- swarm queue connect failed (per hive/agent, 1h window, fires at once:
the line is logged once per agent process and the connection is never
retried, so the window is how long the rule stays firing)
- swarm agent state publish skipped (per hive/agent, 15m, for 5m, as the
swarm terminal rule)
- swarm queue credential missing (per hive/agent, 15m, for 5m)
- agent credential renewal failing (10m, for 15m, as the forge reconcile
rule: same 5m pass cadence, one line per failed pass)
- no log records from <hive>, one per `swarm.hives` entry: fires when
the hive's ungrouped count over 10m is below 1, and on no data. A new
`logAbsenceRule` helper wraps `logCountRule` with the inverted
threshold and noDataState = Alerting.
No contact point and no notification policy: the rules show under
Alerting -> Alert rules and are delivered nowhere.
Refs #4717
Refs #3900
The restart decision was a shell variable set by comparing the fetched
value with the file just before overwriting it. A run that wrote the
file and then failed before the restart (the matrix unit's registration
render, or `systemctl --machine` finding no bus yet) left a retry that
saw an unchanged file and never restarted the consumer.
The file is now written only when the value differs, so its mtime marks
the last real change, and `refresh_consumer <machine> <unit> <path>`
compares that mtime with the consumer's ActiveEnterTimestamp on every
run, the shape the openbao client-CA refresh in swarm-bao.nix already
uses. A consumer that started after the last change is left alone; a
running one is try-restarted, a failed one reset and started, all with
--no-block, and nothing happens while the container is down.
The helper's comment block also exceeded the 30-line limit
(`comment-block lint` failed on d871467d); its per-function notes now
sit beside the functions.
module-eval-bao-grants asserts the gated write, the path the refresh is
keyed on, and the mtime-vs-start comparison for each consumer.
Refs #4662
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.
Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:
- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao
A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.
module-eval-bao-grants gains one case per consumer.
Refs #4662
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.
- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
hive-agent-bao-identity and hive-agent-forge-token use it.
queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
forge-token.nix is a fetch unit with the same short budget that was
added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
(--no-block) a loaded nginx that is not active; an active nginx keeps
the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
fetch unit, swarm-services-cert included.
Refs #4662
A plain removal breaks any out-of-tree host config that still sets the
option: eval fails with "option does not exist" and no pointer to what
replaced it. mkRemovedOptionModule gives a clear evaluation error instead.
swarm-controller is the only minter of swarm/hives/<hive>/matrix/sender-token
since #4820, so the swarm-matrix-ctl policy's first stanza granted a write no
code performs. The policy keeps its one used stanza, the swarm appservice token
that `swarm-matrix-ctl appservice publish` writes. deploy.bao.matrixCtlHiveName
only named the hive in the dropped stanza and goes with it.
hive-matrix.nix no longer calls the per-hive appservice's as_token hive-c0re's
authority: tuwunel loads the registration and creates the sender account, and
no client presents that token.
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:
- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
hive-matrix container, Command::Mint and src/mint.rs. The binary, its
appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
role stay. bao-matrix-reader's checks on the deleted unit are removed;
the leaf-identity and no-token-in-env checks now look at
swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
(register/appservice-login/password-login with the local as_token), with
read_appservice_token, paths::matrix_appservice_token and the helpers
only it used. ensure_hive_user now takes the store's token, keeps the
file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
unchanged apart from no longer reading the local as_token.
This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.
Closes#4813Closes#4814
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
Two Grafana-managed rules in a "hyperhive" folder, one group evaluated
every 1m, each a LogsQL stats count over the VictoriaLogs datasource:
- forge reconcile failing: `swarm forge objects: write failed; retrying
next pass` over 10m, for 15m. The pass runs every 5m, so one failed
pass stays under `for` and a pass failing on every tick fires.
- swarm terminal publish skipped: `swarm terminal: publish skipped` over
15m, for 5m, one instance per hive/agent.
No contact point or notification policy. Grafana 13 routes to its
built-in `empty` receiver when none is provisioned, so firing rules
are visible under Alerting and sent nowhere, with no send errors.
Refs #4717
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.
hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.
Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.
module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.
Refs #3773
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.
`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).
swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.
Refs #3773
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.
The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.
OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.
The viewer's log message no longer tells the operator to restart it.
Refs #4772
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.
Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.
Refs #4585
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.
Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.
Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.
Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.
Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.
The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.
Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.
nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:
- `MintAgentIdentity` (the node agent creation uses) when the stored
certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
re-decides, writes a fresh value with `minted_at`, reads it back, and logs
the agent and the old age.
When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.
`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.
Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.
Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.
docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.
- authelia registers an interactive client `swarm-bao-ui`
(glue-bao-ui-oidc-client.nix) with redirect
`https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
publisher carries its secret to
`secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
mount with listing visibility `unauth`, asked before attempted like
cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
policy, and once the granter may configure `auth/oidc/config` (checked
through `sys/capabilities-self`), writes the mount's config from the
published secret and the role binding `groups=admins` to the viewer.
Before the bootstrap step re-runs it writes the policy, logs the step
and exits 0.
Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.
module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.
The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.
Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.
Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.
On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.
Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.
Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.
Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.
hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.
swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.
Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
c9ba9062 replaced 5 bargauge panels with one table panel using a
joinByField+organize+sortBy transformations chain and a
vizConfig.group="table" viz spec. That construct had zero precedent
anywhere else in this repo's dashboards and its live render was
explicitly flagged as unverified at merge time (no in-container way
to render Grafana to confirm).
#4658 reports a dashboard crash: "'' not found in: reduce,
filterFieldsByName, ... transpose" - a Grafana transform-registry
lookup failing on an empty transformer id. That registry's full id
list matches the error text exactly, and the table panel is the only
panel in the whole swarm-grafana/dashboards tree using a non-empty
transformations array, so it's the leading suspect: some part of how
the dashboard-v2beta1 schema expects a QueryGroup's transformations
to be shaped likely differs from the classic {id, options} shape used
here, and the file-based provisioner may not be normalizing the
mismatch the same way a schema-v2-aware loader does.
Reverting to the prior, previously-uneventful bargauge panels
(byte-for-byte c9ba9062^'s version of this section) to stop the
active crash while the correct v2beta1 transformation shape gets
verified against a live Grafana render instead of guessed at again.
swarm-bao-granter-role used its token for `bao auth list` with no check,
so an expired, revoked or policy-less token in bootstrapTokenFile exited
2 with a raw 403 and no hint. It now checks whether the store is up when
that call fails: if it is, the token is at fault, and the unit prints the
one-time bootstrap step and exits 4. A missing token file is still a
ConditionPathExists skip, so the two read differently in the journal.
The twelve granterLogin units printed "the store is sealed or
unreachable" on a healthy store, because `bao status` exited 1 there:
the CLI resolves a token helper under $HOME before asking, systemd sets
no HOME for a unit without User=, and the fallback shells out to
`getent`/`sh`, neither of which is on the unit's PATH ("failed to get
token helper: error expanding config path "": exec: "sh": executable
file not found in $PATH"). The check now runs with HOME=/var/empty and
keeps its stderr, so a genuinely unreachable store says why. When the
store is up and the login is refused, the units now name
swarm-bao-granter-role as the unit that writes the missing role.
setup.md's post-step restart used 'swarm-bao-*-policy.service', which
misses swarm-bao-agent-pki. It now names that unit too, and a
module-eval case fails when the restart misses any unit that logs in as
the granter.
Refs #4704
matrixHomeserverUrl re-derived gatewayHost's own default
(chat.${swarmDomain}) instead of reading gatewayHost itself, so a
deployment that pins gatewayHost away from that default (the exact
case hive-matrix.nix's own option doc describes) left the controller
calling a name nothing serves. Read matrix.gatewayHost directly,
mirroring hive-matrix.nix's ctlHomeserverUrl. Add module-eval cases
that pin gatewayHost and that null it out, alongside the existing
unpinned default case.
Closes#4757
An `auth_token` spelled `swarm-agent.<agent>.<secret>` is no longer sent
to introspection. The responder reads `swarm/agents/<agent>/queue` with
an identity of its own, checks that the stored object names the same
agent, compares the secret in constant time, and grants the subjects
`--agent-token-publish-subject` lists with `{agent}` expanded. Every
other outcome denies: a malformed token, no store identity, nothing
stored, a failed or slow lookup, a different secret. A token without
the prefix takes the OIDC path unchanged.
The journal's `auth request` line names such a caller `agent:<agent>`;
the hive-shared credential keeps `hive-<h>-agent`.
The new principal: a `swarm-nats-auth` cert-auth role and policy with
read on `secret/data/swarm/agents/+/queue` alone, a leaf signed by the
store's PKI glue, and `glue-nats-auth-bao-identity.nix` pairing the two.
The copy unit delivers the identity into the queue's container, and an
absent leaf is delivered empty so the responder still starts and only
agent tokens are refused.
The policy and role are written by `swarm-bao-nats-auth-policy`, logged in
as the bao granter: both names fall under its `swarm-*` globs, so the
deploy writes them with no operator step. module-eval counts it among the
granting units, so every generic granting-unit case covers it.
The secret compare uses `subtle`, already in the lock file through the
TLS stack; no workspace crate offered one directly.
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.
- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
tunes the mount, generates the root once (guarded on an empty issuer
list, no replace branch), upserts the `swarm-agent` role (client
certificates named `hive-agent-*` only, 90 days), caches the CA at
/var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
(client-ca.pem, then the agent CA). Host cert-auth roles still pin
client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
changed after openbao started, swarm-bao-agent-pki restarts
openbao.service in the container; under `seal = "shamir"` it prints
the step instead. Once swarm-bao-certs has a cached CA, later boots
start openbao with it and do not restart.
- The controller policy gains exactly `update` on
pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
the leaf (the store generates the key), writes the agent's cert-auth
role pinning the issuing CA bao returned, and writes the agent's
policy as render_agent alone: the hive-shared queue credential stanza
is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
other pki role names); swarm-controller gets
SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.
Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.
Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.
Closes#4756
#4756 moves agent client certificates onto a PKI mount of their own,
`pki-agents`, whose root bao generates internally. The unit that sets
that mount up runs as the bao granter, and the granter's policy is only
written while #4754's one-time bootstrap token is in place. Adding these
grants after an operator has done that step would cost a second token
placement, so they go into the granter's policy here, before it.
Six stanzas: enable and tune the mount, list its issuers, read its CA,
generate its root internally, and write `swarm-*` roles on it. No root
delete or sudo: agent cert-auth roles pin that root by value, so
replacing it must not be something a deploy can do.
Adds deploy.bao.agentPkiMountPath (default `pki-agents`), which the
stanzas are rendered from. The module-eval case pinning the granter's
policy now lists all seventeen stanzas, and the "grants nothing outside"
case also refuses the agent mount's root, issue, sign and a roles/*
glob.
Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).
A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.
swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.
The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.
module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.
Refs #4704
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).
- hive-priv gains `EnsureAgentSocketDir { name }`, called from
`set_nspawn_flags` in every start path. It creates
`/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
directory is left alone. hive-c0re's own create_dir_all went: its /run
is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
the agent user and sets 0751, the same way it already handles state/ and
harness/. It refuses a symlink or non-directory there, since `test -d`
and chmod follow links. No host-side passwd parse, and no window where
the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
(`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
`converge_start_preamble` + `start_with_fallback`. It was a bare start,
so after a reboot the manager's bind sources existed only because of the
tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
`priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
`SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
does the same unlink and returns Ok, for an older hive-c0re.
Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.
Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.
Closes#4742
The header comment grew to 39 lines across two audit-driven rounds,
over the 30-line comment-block-lint max. Trimmed to 30: merged the
value-as-argument rationale with its /proc/cmdline justification into
one paragraph, cut the usage example to one call instead of two, and
condensed the args paragraph — no content dropped, just restatement.
persistence.md's first-boot-migration marker paragraph used "it is"
and "there is not" — vale's Microsoft.Contractions rule (this repo's
config) wants the contracted forms, and "is not" also collides with
"is nothing" as a literal substring, which is what actually tripped
the error. Reworded to "it's" / "there's nothing", no meaning change.
Lint-only: no script logic changed, gates re-run below are all lint
checks (no module-eval, no cargo).
Refs #4723
The pipe contract had a gap: if a producer piped into
atomic_write_secret exited non-zero after writing partial output,
cat still saw a clean EOF and wrote that partial content through to
the live target via mv — pipefail only reported the failure
afterward, once the bad write was already committed. The helper now
takes the value as its 4th argument and writes it itself with printf
(a shell builtin, so the value never touches an external process's
own argv/environ, same as a function argument never does), so there
is no pipe left to fail silently.
Callers that compute the value with a command now capture it into a
variable first (`value=$(cmd)`), which fails under `set -e` before
atomic_write_secret is ever called — swarm-bao.nix's pin.env site is
the one that needed this (`pin_env_value="BAO_HSM_PIN=$(cat ...)"`).
All seven call sites converted; output is byte-identical (same
printf '%s\n' framing, now applied inside the helper instead of by
each caller).
Refs #4723
swarm-bao-token's pin.env write (the BAO_HSM_PIN EnvironmentFile for
openbao's pkcs11 seal) had the same write-then-chmod-on-live-path
shape as the sites already converted: printf > path directly on the
live file, chmod after. Same fix, same helper. Content
("BAO_HSM_PIN=<user-pin>\n") and final mode (0400, root-owned — no
chown, same as before) are unchanged; pin.env stays at the same path,
so openbao's EnvironmentFile= reference needs no change.
Refs #4723
atomic_write_secret's cleanup trap used RETURN, which never fires when
set -e aborts the function mid-body (a failing cat/chmod/chown), so a
secret-bearing temp file was left behind instead of being removed.
The write now runs in a subshell with its own EXIT trap, invoked via a
named handler (so `local rc=$?` is a normal, shellcheck-visible
assignment) that only removes the temp file when the subshell's exit
status is nonzero — the subshell's trap table is private, so a calling
unit's own EXIT trap is untouched. Reproduced the leftover-tmp bug
against the prior commit, confirmed it's gone, and confirmed both the
success path and a caller's own EXIT trap still work as before.
swarm-bao.nix's swarm-bao-forwarder-oidc unit had the identical
write-then-chmod-on-live-path defect as the four sites already fixed
here (fetches an OIDC client secret from swarm-bao, printfs it to the
live host path, chowns/chmods after) and was missed by the original
sweep. Converted it to atomic_write_secret; content and final
owner/mode (root:root, 0400) are unchanged.
Refs #4723
Four host glue units fetched a secret from swarm-bao and rendered it
with `> path; chmod`: a reader racing the write could see a truncated
file, and briefly one at the wrong mode before the chmod landed.
glue-matrix-bao-token.nix, glue-queue-agent-credential.nix (both
files), swarm-grafana.nix and swarm-otel.nix now write to a same-
directory temp file, set its final mode/owner, then `mv -f` it over
the target — a shared `atomic_write_secret` helper
(nix/host-modules/lib/atomic-write-secret.nix) so the five call sites
share one implementation.
The first-boot `/root/.claude` migration in nix/agent-modules/user.nix
wrote its done-marker unconditionally, so a failed `cp` (disk full,
permission error) left the marker behind and no boot ever retried the
copy. The marker is now written only when there was nothing to
migrate or the copy succeeded; `cp -an`'s no-clobber semantics already
make a retry after a partial copy safe.
Refs #4723
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.
Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:
- Dropped: each swarm service and its glue keeps only its own deploy
toggle (authelia, bao and its PKI glue, grafana, victorialogs,
victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
the controller/nats/matrix-ctl/publisher/services-issuer identities),
the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
its assertion, which feed hive-c0re and write under its state dir, plus
their policy-order entry; the network identity assertions; hive-tls's
two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
(`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
`deploy.matrix.enable` plus their client-identity checks. They read
deploy.matrix's token file and registration script; their deploy.bao
inputs are the client-half options a hive sets to read a store it does
not run, so gating on deploy.bao.enable would drop the tested
remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
and fires wherever the hive, the store or the homeserver runs: each
turns the name into an identifier with no fallback.
On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.
Refs #4500
Every gateway asked the store's `pki/issue/swarm-services` for the whole
swarm's service set, so a private key on any gateway host could serve a
valid certificate for services that host does not front and never has.
`swarm.localServiceDomains` derives the per-host subset by filtering
`swarm.serviceDomains` against the vhosts this host actually renders —
the deploy flags those vhosts are already guarded on, read once rather
than copied into a second filter. The leaf request and the coverage
guard that decides whether to re-issue both read it, so they cannot
disagree about which names the leaf owes.
The sub-CA's name constraint and the role's `allowed_domains` stay the
swarm-wide set: every host's subset is inside it, and narrowing the
constraint per host would turn one signing into N.
The store's `swarm-services` role issues the services leaf for 720h, and
`swarm-services-cert` only ever ran at boot or rebuild: it is a
`RemainAfterExit` oneshot wanted by `multi-user.target` and no timer
targeted it. A hive not rebuilt within 30 days served an expired leaf.
`swarm-services-cert-renew` runs the same script from a daily timer. It
is a unit of its own because a timer starting the `RemainAfterExit` unit
is a no-op, and restarting that unit instead would propagate through
`hive-gateway-self-signed-cert`'s `Requires=` to nginx, so a sealed store
would take the gateway down over a still-valid leaf. Nothing requires or
orders against the new unit; it has no `Restart=`, so a failure stays in
`systemctl --failed` until the next tick, and the script only moves files
into place after the store has answered.
The re-issue threshold was `checkend 2592000`, the whole 30-day
lifetime, so every run re-issued. It is now half the role's lifetime,
read from a new internal option `deploy.bao.servicesPkiLeafTtlHours`
that the role's `ttl`/`max_ttl` also read. Boot and timer share the
script and so the threshold. The services-root re-check reads the
same option, at the store's own replacement threshold (hours × 3600),
so the hive asks for a new leaf when the store replaces its root. A
`flock` keeps the two runs from interleaving one issuance's key with another's leaf.
`checks.module-eval-hive-tls` pins the timer, that the unit it starts
re-runs the issuance without `RemainAfterExit`, that nothing depends on
it, and that both the leaf and root thresholds move with the option.
Closes#4587
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.
swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.
hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.
nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.
Refs #3782
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.
The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.
`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
The matrix container renders the swarm registration before tuwunel starts
(`requiredBy` it, no network), tuwunel loads it as a second `.yaml`
credential, and a publish unit hands its token to the store. tuwunel 1.9.1
refuses only a duplicate id or as_token, not overlapping non-exclusive
namespaces, so it sits beside the hive's `hyperhive` registration.
`admin_execute` promotes exactly `@swarm` at boot. The module-eval pin
narrows from "no account" to that one list, compared whole so a second
entry fails; `admin_execute_errors_ignore` is pinned too. matrix-ctl may
write the swarm token's leaf, and the controller may only read it.
Everything is gated on matrix-ctl's store identity: with nobody to publish
the token, the registration would be an admin credential nobody reads.
bao's own in-container collector labelled every log line and metric it
forwards `service.name=swarm-bao` (`processors.resource.attributes`,
keyed off `swarm.bao.machine`), while every panel in the shipped
Grafana bao dashboard queries the literal `service.name="bao"` — no
panel has matched since the scrape moved into that collector.
Rename it in the collector instead of templating the dashboard: a new
`collectorServiceName` binding in swarm-bao.nix, deliberately not
`cfg.machine` (that value names the container/receiver, not bao's
display identity), stamps `service.name="bao"` directly. bao.json is
back to its origin/main shape, unchanged.
Adds a module-eval case to checks.module-eval-grafana that reads the
collector's own evaluated config and asserts its service.name matches
every selector the shipped dashboard uses; verified invert-proof by
setting the value back to cfg.machine and confirming that specific
case (and only it) fails.
[oauth2_client] turns on auto-registration through the authelia login
source, with the account named after authelia's preferred_username.
DISABLE_REGISTRATION stays true: forgejo 16's auto-registration checks
only ALLOW_ONLY_INTERNAL_REGISTRATION, so local sign-up stays off.
ACCOUNT_LINKING is `login`, forgejo's default, set explicitly. With
`auto`, an SSO login whose name matches an existing local account would
be handed that account, and agents, `core` and `swarm-controller` all
have one. `login` asks for that account's own password instead.
Refs #3782
argus review on #4713: the comment said every container block sets
--link-journal=host, and warned that dropping it silently blinds this
receiver. #4713 does exactly that for swarm-bao (its journal stays
inside the container for its own collector instead), so read on its
own the comment told a debugger to revert the fix. State the
exception.
--link-journal=host bind-mounts a host directory journald never
writes into for this container (empty, root:nogroup, confirmed on
the live host — #4527). Dropping it falls back to nixpkgs' default
--link-journal=try-guest, the same shape every agent container
already uses, so the in-container collector's journald receiver
(directory = /var/log/journal) now reads a journal that is actually
written.
merge stays true and services.hyperhive.swarm.otel.journaldUnits is
untouched so this deploy changes exactly one thing; comments that
described the old host-linked shape are rewritten to match.
Adds a module-eval assertion (bao-otel-collector.nix) that the
container's extraFlags never re-add --link-journal=host.
Refs #4527, #4499.
`swarm-bao-forwarder-oidc` fetches the store container's collector secret,
one path, and was the last reader still logging in with
`deploy.bao.clientCertFile`: the hive's own leaf, whose policy reads every
agent's credentials, the hive's tree and every service's OIDC secret. The
four-way split gave grafana's and the swarm collector's readers leaves of
their own and left this one behind.
It now holds `forwarder-oidc.pem`, minted by `swarm-bao-pki`, and logs in
under the `swarm-forwarder-oidc` cert-auth role, whose policy reads
`secret/data/swarm/services/<store forwarder client id>/oidc/client` and
nothing else. The role is written by `swarm-bao-forwarder-oidc-policy`
from the bootstrap token, which gains the two grants that unit calls, and
the reader is ordered after it. The subject is reserved as a hive name. A
store host whose pair is null is refused at eval rather than falling back
to the hive's leaf.
The hive's own role and `client.pem` are untouched; nothing is revoked.
Every hive with hyperhive enabled ran its own hive-forge container, and
its gateway answered forge.<swarm> with its own bridge IP, so on a
multi-host swarm each hive talked to its own forge.
deploy.forgejo.enable defaults to false and allSwarmServices sets it with
mkDefault, like authelia and bao; singleHostSwarm gets it through that.
The forge's OIDC client moves to a glue module gated on authelia, so a
split authelia/forge swarm still registers it. CI now requires the forge
on the same host, and the controller's forgeTokenFile defaults to null
where the forge is not.
Closes#4705
Refs #3782