Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/nix/module-eval
Author SHA1 Message Date
atlas
b68fd7306e refresh-consumer: key the restart on the file's mtime, not a pre-write compare
The restart decision was a shell variable set by comparing the fetched
value with the file just before overwriting it. A run that wrote the
file and then failed before the restart (the matrix unit's registration
render, or `systemctl --machine` finding no bus yet) left a retry that
saw an unchanged file and never restarted the consumer.

The file is now written only when the value differs, so its mtime marks
the last real change, and `refresh_consumer <machine> <unit> <path>`
compares that mtime with the consumer's ActiveEnterTimestamp on every
run, the shape the openbao client-CA refresh in swarm-bao.nix already
uses. A consumer that started after the last change is left alone; a
running one is try-restarted, a failed one reset and started, all with
--no-block, and nothing happens while the container is down.

The helper's comment block also exceeded the 30-line limit
(`comment-block lint` failed on d871467d); its per-function notes now
sit beside the functions.

module-eval-bao-grants asserts the gated write, the path the refresh is
keyed on, and the mtime-vs-start comparison for each consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
b3b42d3279 credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
b58a0d8ba9 bao: drop matrix-ctl's per-hive sender-token grant
swarm-controller is the only minter of swarm/hives/<hive>/matrix/sender-token
since #4820, so the swarm-matrix-ctl policy's first stanza granted a write no
code performs. The policy keeps its one used stanza, the swarm appservice token
that `swarm-matrix-ctl appservice publish` writes. deploy.bao.matrixCtlHiveName
only named the hive in the dropped stanza and goes with it.

hive-matrix.nix no longer calls the per-hive appservice's as_token hive-c0re's
authority: tuwunel loads the registration and creates the sender account, and
no client presents that token.
2026-09-30 07:42:22 +02:00
atlas
84d4808d54 hive-subagent-mcp: run an agent's subagents on its runtime
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.

The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.

Refs #4391
2026-09-30 07:41:12 +02:00
atlas
ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00
atlas
78d8d69c7f swarm-controller: mint each hive's matrix sender token
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.

swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.

swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.

hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.

The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
10d579ecaf nix: move the remaining service containers onto the swarm-container module
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.

hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.

Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.

module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.

Refs #3773
2026-09-29 20:17:28 +02:00
atlas
cce41c1d79 nix: move the in-container modules to nix/container-modules/
Refs #3773
2026-09-29 19:51:46 +02:00
atlas
024067f3f8 nix: share the service-container settings through one in-container module
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.

`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).

swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.

Refs #3773
2026-09-29 19:51:46 +02:00
atlas
8fb7751da9 swarm-bao: don't let the viewer restart enqueue fail the granter
ExecStartPost (not postStart, which can't take a prefix) with a
leading -: a failed enqueue -- an already-running viewer unit,
or systemctl itself failing -- must not mark the granter failed
or trigger its own Restart=on-failure.
2026-09-29 18:09:53 +02:00
atlas
22c96282b9 swarm-bao: re-run the operator viewer unit when the granter step succeeds
swarm-bao-operator-viewer-policy exits 0 while the granter may not
configure auth/oidc, so it never retries on its own. On 2026-09-29 the
operator fixed the granter (swarm-bao-granter-role succeeded at 15:29Z),
but the viewer unit had last run on 2026-09-28 19:11Z on that exit-0
branch. auth/oidc/config and the viewer role stayed unwritten and OIDC
login failed until a manual restart.

The granter unit now restarts the viewer unit from ExecStartPost, which
runs only after its script exits 0. Restart rather than start, because
the viewer unit is RemainAfterExit and a start would be a no-op.
--no-block, because the viewer unit is ordered after the granter and a
blocking restart would deadlock. The link is one-way, so the viewer's
own Restart=on-failure never re-runs the granter.

OnSuccess= would not fire (the granter stays active under
RemainAfterExit), and Wants=/PartOf= either no-op on an active unit or
also propagate a failed restart and every stop.

The viewer's log message no longer tells the operator to restart it.

Refs #4772
2026-09-29 17:45:58 +02:00
atlas
ccb5bd3b38 hive-agent: read the per-agent queue secret from bao in process
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.

hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.

A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).

Closes #4783
2026-09-29 10:18:07 +02:00
atlas
916c441e82 swarm-otel: bound the collector's restart backoff
The collector's oidc/* authenticators call out to authelia at startup, so a
restart that races authelia's own (a redeploy that touches both, a store
outage) can fail immediately. nixpkgs' upstream opentelemetry-collector
module sets Restart=always with no RestartSec, so systemd's defaults
(100ms RestartSec, 5-in-10s start limit) burn the whole allowance in well
under a second and leave the unit in start-limit-hit, dead until someone
resets it by hand.

Sets RestartSec=5 plus an explicit startLimitBurst/startLimitIntervalSec
window (12/120s) sized so the burst can never trip while authelia comes
back — same values host-modules/otel.nix already uses for the sibling
host-tier collector, which depends on authelia the same way. Pins the
[Unit]-vs-[Service] placement and the window relation in
module-eval-swarm-otel-core, mirroring module-eval-hive-otel's existing
case for the host tier.
2026-09-29 09:15:12 +02:00
atlas
642be57678 swarm: narrow the revoke grant to the queue leaf, not the whole agent prefix
revoke_queue_credential only ever deletes swarm/agents/<agent>/queue
(agent_queue_path + a literal "queue" suffix), never anything else
under an agent's prefix. secret/metadata/swarm/agents/+/queue matches
that exactly — `+` is bao's single-segment glob, the same form
swarm-nats-auth's read grant already uses for the data-side path.

Also rewords the module-eval test's stale note about a read/list grant
handing a "write-only principal" the version history: the controller
has held read on secret/data/swarm/agents/* since the mint-and-verify
read-before-write change, so it was never write-only on that path.
2026-09-28 23:46:17 +02:00
atlas
8caf688ee4 swarm: revoke an agent's queue credential when it is declared destroyed
A per-agent queue credential is minted at agent creation and nothing has
ever removed it. An agent declared destroyed loses its container and
keeps its credential: a bearer secret recovered from a snapshot or a
stale capture still authenticates as that agent, so the set of usable
credentials only grows.

Delete the path the mint published, on the one transition that ends an
agent's life. It mirrors step 3 of `mint_and_verify` and no other step:
the leaf, the ACL document and the cert role are what a hive uses to
collect an agent's secrets and are re-minted on every run of the mint.

Every version, not the newest. The mint rewrites the path when the
principal it names needs correcting, so KV v2's plain delete would leave
the identical secret readable at ?version=N. That is a separately-ACL'd
path, hence the second stanza in the controller's grant -- `delete` on
metadata discloses nothing, and `update` on the data path already lets
this principal destroy any agent credential's usability.

The destroy is not blocked by a failed revocation: the declaration is
already published and refusing the call would leave an operator with an
agent they cannot tear down. The failure is logged at error instead,
naming the agent, since a silent orphan is the fault being removed.
2026-09-28 23:46:17 +02:00
atlas
433429ebfd bao: disable the unused approle auth method, declaratively
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.

Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.

nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
2026-09-28 22:58:36 +02:00
atlas
b14ff2796c bao: OIDC login to the browser UI via authelia, as a metadata-only viewer
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.

- authelia registers an interactive client `swarm-bao-ui`
  (glue-bao-ui-oidc-client.nix) with redirect
  `https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
  publisher carries its secret to
  `secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
  mount with listing visibility `unauth`, asked before attempted like
  cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
  and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
  policy, and once the granter may configure `auth/oidc/config` (checked
  through `sys/capabilities-self`), writes the mount's config from the
  published secret and the role binding `groups=admins` to the viewer.
  Before the bootstrap step re-runs it writes the policy, logs the step
  and exits 0.

Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.

module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
2026-09-28 19:56:38 +02:00
atlas
3898ca33c7 bao: serve the browser UI to admins via a loopback-only listener
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.

The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.

Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.

Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.

On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
2026-09-28 19:31:02 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
e974194e3a swarm: let an agent publish its own icon
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.

hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.

swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.

Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
2026-09-28 13:47:37 +02:00
atlas
a2b4acfb6a swarm-bao: say "run the bootstrap step" when the bootstrap token is dead, not "sealed"
swarm-bao-granter-role used its token for `bao auth list` with no check,
so an expired, revoked or policy-less token in bootstrapTokenFile exited
2 with a raw 403 and no hint. It now checks whether the store is up when
that call fails: if it is, the token is at fault, and the unit prints the
one-time bootstrap step and exits 4. A missing token file is still a
ConditionPathExists skip, so the two read differently in the journal.

The twelve granterLogin units printed "the store is sealed or
unreachable" on a healthy store, because `bao status` exited 1 there:
the CLI resolves a token helper under $HOME before asking, systemd sets
no HOME for a unit without User=, and the fallback shells out to
`getent`/`sh`, neither of which is on the unit's PATH ("failed to get
token helper: error expanding config path "": exec: "sh": executable
file not found in $PATH"). The check now runs with HOME=/var/empty and
keeps its stderr, so a genuinely unreachable store says why. When the
store is up and the login is refused, the units now name
swarm-bao-granter-role as the unit that writes the missing role.

setup.md's post-step restart used 'swarm-bao-*-policy.service', which
misses swarm-bao-agent-pki. It now names that unit too, and a
module-eval case fails when the restart misses any unit that logs in as
the granter.

Refs #4704
2026-09-28 11:03:08 +02:00
atlas
3fc7a785fa swarm-controller: default matrixHomeserverUrl from swarm.matrix.gatewayHost
matrixHomeserverUrl re-derived gatewayHost's own default
(chat.${swarmDomain}) instead of reading gatewayHost itself, so a
deployment that pins gatewayHost away from that default (the exact
case hive-matrix.nix's own option doc describes) left the controller
calling a name nothing serves. Read matrix.gatewayHost directly,
mirroring hive-matrix.nix's ctlHomeserverUrl. Add module-eval cases
that pin gatewayHost and that null it out, alongside the existing
unpinned default case.

Closes #4757
2026-09-28 10:16:50 +02:00
atlas
9bad58d86d swarm-nats-auth: verify an agent's own token against the store
An `auth_token` spelled `swarm-agent.<agent>.<secret>` is no longer sent
to introspection. The responder reads `swarm/agents/<agent>/queue` with
an identity of its own, checks that the stored object names the same
agent, compares the secret in constant time, and grants the subjects
`--agent-token-publish-subject` lists with `{agent}` expanded. Every
other outcome denies: a malformed token, no store identity, nothing
stored, a failed or slow lookup, a different secret. A token without
the prefix takes the OIDC path unchanged.

The journal's `auth request` line names such a caller `agent:<agent>`;
the hive-shared credential keeps `hive-<h>-agent`.

The new principal: a `swarm-nats-auth` cert-auth role and policy with
read on `secret/data/swarm/agents/+/queue` alone, a leaf signed by the
store's PKI glue, and `glue-nats-auth-bao-identity.nix` pairing the two.
The copy unit delivers the identity into the queue's container, and an
absent leaf is delivered empty so the responder still starts and only
agent tokens are refused.

The policy and role are written by `swarm-bao-nats-auth-policy`, logged in
as the bao granter: both names fall under its `swarm-*` globs, so the
deploy writes them with no operator step. module-eval counts it among the
granting units, so every generic granting-unit case covers it.

The secret compare uses `subtle`, already in the lock file through the
TLS stack; no workspace crate offered one directly.
2026-09-28 08:24:52 +02:00
atlas
6170e74a31 swarm-bao: agent certificates issued by a store-generated agent CA
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.

- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
  tunes the mount, generates the root once (guarded on an empty issuer
  list, no replace branch), upserts the `swarm-agent` role (client
  certificates named `hive-agent-*` only, 90 days), caches the CA at
  /var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
  (client-ca.pem, then the agent CA). Host cert-auth roles still pin
  client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
  composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
  changed after openbao started, swarm-bao-agent-pki restarts
  openbao.service in the container; under `seal = "shamir"` it prints
  the step instead. Once swarm-bao-certs has a cached CA, later boots
  start openbao with it and do not restart.
- The controller policy gains exactly `update` on
  pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
  the leaf (the store generates the key), writes the agent's cert-auth
  role pinning the issuing CA bao returned, and writes the agent's
  policy as render_agent alone: the hive-shared queue credential stanza
  is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
  other pki role names); swarm-controller gets
  SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.

Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.

Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.

Closes #4756
2026-09-27 22:59:27 +02:00
atlas
5cd7f866f4 swarm-bao: grant the agent PKI mount (for #4756)
#4756 moves agent client certificates onto a PKI mount of their own,
`pki-agents`, whose root bao generates internally. The unit that sets
that mount up runs as the bao granter, and the granter's policy is only
written while #4754's one-time bootstrap token is in place. Adding these
grants after an operator has done that step would cost a second token
placement, so they go into the granter's policy here, before it.

Six stanzas: enable and tune the mount, list its issuers, read its CA,
generate its root internally, and write `swarm-*` roles on it. No root
delete or sudo: agent cert-auth roles pin that root by value, so
replacing it must not be something a deploy can do.

Adds deploy.bao.agentPkiMountPath (default `pki-agents`), which the
stanzas are rendered from. The module-eval case pinning the granter's
policy now lists all seventeen stanzas, and the "grants nothing outside"
case also refuses the agent mount's root, issue, sign and a roles/*
glob.
2026-09-27 22:57:46 +02:00
atlas
e9cec0da21 swarm-bao: write every swarm-* grant as a bao granter, not with a 24h token
Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).

A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.

swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.

The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.

module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.

Refs #4704
2026-09-27 22:57:46 +02:00
atlas
2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00
atlas
1948e7ad23 module-eval: read the services-leaf narrowing off a forge host
`checks.module-eval-core-toggle` is red on main: "a gateway's services
leaf asks for only the swarm names that host fronts" read its leaf
request off `bare`, on the premise that `bare` fronts forge. That was
true where the property was written, before 978164dc (#4705) defaulted
`deploy.forgejo.enable` to false. The two merged in sequence, and on main
`bare` renders only the `_` and `h1.t.local` vhosts, so its
`localServiceDomains` is [] and the script carries `alt_names=''`,
never `alt_names=forge.t.local`.

The narrowing itself is right: a host that runs no forge fronts no
swarm name and asks for no services leaf. Only the fixture was stale.
The property now reads a hive with `deploy.forgejo.enable = true`,
whose request is `common_name=forge.t.local alt_names=forge.t.local`
while `auth.t.local` stays in the swarm-wide set.

The renewal timer from 0649673e is not involved: the check's derivation
is identical at 0649673e and at its parent, and ed2ec52f replayed onto
978164dc^ holds while replayed onto 978164dc it fails.

Refs #4587
2026-09-26 01:39:22 +02:00
atlas
3202cde704 nix: gate hive-c0re on deploy.hive-controller.enable, drop hyperhive.enable
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.

Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:

- Dropped: each swarm service and its glue keeps only its own deploy
  toggle (authelia, bao and its PKI glue, grafana, victorialogs,
  victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
  the controller/nats/matrix-ctl/publisher/services-issuer identities),
  the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
  its assertion, which feed hive-c0re and write under its state dir, plus
  their policy-order entry; the network identity assertions; hive-tls's
  two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
  (`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
  `deploy.matrix.enable` plus their client-identity checks. They read
  deploy.matrix's token file and registration script; their deploy.bao
  inputs are the client-half options a hive sets to read a store it does
  not run, so gating on deploy.bao.enable would drop the tested
  remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
  and fires wherever the hive, the store or the homeserver runs: each
  turns the name into an identifier with no fallback.

On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.

Refs #4500
2026-09-26 01:19:49 +02:00
atlas
ed2ec52fe5 swarm-tls: narrow each gateway's services leaf to the names it fronts
Every gateway asked the store's `pki/issue/swarm-services` for the whole
swarm's service set, so a private key on any gateway host could serve a
valid certificate for services that host does not front and never has.

`swarm.localServiceDomains` derives the per-host subset by filtering
`swarm.serviceDomains` against the vhosts this host actually renders —
the deploy flags those vhosts are already guarded on, read once rather
than copied into a second filter. The leaf request and the coverage
guard that decides whether to re-issue both read it, so they cannot
disagree about which names the leaf owes.

The sub-CA's name constraint and the role's `allowed_domains` stay the
swarm-wide set: every host's subset is inside it, and narrowing the
constraint per host would turn one signing into N.
2026-09-25 23:41:10 +02:00
atlas
0649673ebf hive-tls: renew the swarm-services leaf on a daily timer
The store's `swarm-services` role issues the services leaf for 720h, and
`swarm-services-cert` only ever ran at boot or rebuild: it is a
`RemainAfterExit` oneshot wanted by `multi-user.target` and no timer
targeted it. A hive not rebuilt within 30 days served an expired leaf.

`swarm-services-cert-renew` runs the same script from a daily timer. It
is a unit of its own because a timer starting the `RemainAfterExit` unit
is a no-op, and restarting that unit instead would propagate through
`hive-gateway-self-signed-cert`'s `Requires=` to nginx, so a sealed store
would take the gateway down over a still-valid leaf. Nothing requires or
orders against the new unit; it has no `Restart=`, so a failure stays in
`systemctl --failed` until the next tick, and the script only moves files
into place after the store has answered.

The re-issue threshold was `checkend 2592000`, the whole 30-day
lifetime, so every run re-issued. It is now half the role's lifetime,
read from a new internal option `deploy.bao.servicesPkiLeafTtlHours`
that the role's `ttl`/`max_ttl` also read. Boot and timer share the
script and so the threshold. The services-root re-check reads the
same option, at the store's own replacement threshold (hours × 3600),
so the hive asks for a new leaf when the store replaces its root. A
`flock` keeps the two runs from interleaving one issuance's key with another's leaf.

`checks.module-eval-hive-tls` pins the timer, that the unit it starts
re-runs the issuance without `RemainAfterExit`, that nothing depends on
it, and that both the leaf and root thresholds move with the option.

Closes #4587
2026-09-25 23:38:36 +02:00
atlas
afde9f380f module-eval: give the forgejo-mirror fixtures deploy.forgejo.enable
#4708 defaulted deploy.forgejo.enable to false, and the hive-forge module's
whole config block (including the mirror-declared-without-a-controller
warning) now hangs off it. controllerWithMirrors and mirrorsNoController
both relied on the old always-on default to reach that warning path.
2026-09-25 08:36:05 +02:00
atlas
20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00
atlas
ab153bda2f hive-matrix-mcp: read the main account's token from the store too
The daemon reads each account's token from `swarm/agents/<agent>/matrix/`
as the agent itself, inside its own container, and falls back to the file
only when the store has none. This is #4519's read, without its `main`
carve-out: the swarm now mints `main` there and no hive writes the file.

The daemon unit gets the agent's store identity, spelled the way
forge-token.nix spells it. A timer re-starts it while it is down: a token
the swarm mints or replaces in the store changes no file, so the path
watcher never fires for it, and a daemon that exited on a replaced token
would otherwise stay down until the container restarts.
2026-09-25 08:31:01 +02:00
atlas
2776e121e5 swarm-controller: mint each agent's matrix account with the swarm's token
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.

The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.

`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
2026-09-25 08:31:01 +02:00
atlas
9308752a09 hive-matrix: load the swarm's appservice and promote its sender
The matrix container renders the swarm registration before tuwunel starts
(`requiredBy` it, no network), tuwunel loads it as a second `.yaml`
credential, and a publish unit hands its token to the store. tuwunel 1.9.1
refuses only a duplicate id or as_token, not overlapping non-exclusive
namespaces, so it sits beside the hive's `hyperhive` registration.

`admin_execute` promotes exactly `@swarm` at boot. The module-eval pin
narrows from "no account" to that one list, compared whole so a second
entry fails; `admin_execute_errors_ignore` is pinned too. matrix-ctl may
write the swarm token's leaf, and the controller may only read it.

Everything is gated on matrix-ctl's store identity: with nobody to publish
the token, the registration would be an admin credential nobody reads.
2026-09-25 08:31:01 +02:00
atlas
cb176c7be7 swarm-bao: stamp collector's service.name as "bao"
bao's own in-container collector labelled every log line and metric it
forwards `service.name=swarm-bao` (`processors.resource.attributes`,
keyed off `swarm.bao.machine`), while every panel in the shipped
Grafana bao dashboard queries the literal `service.name="bao"` — no
panel has matched since the scrape moved into that collector.

Rename it in the collector instead of templating the dashboard: a new
`collectorServiceName` binding in swarm-bao.nix, deliberately not
`cfg.machine` (that value names the container/receiver, not bao's
display identity), stamps `service.name="bao"` directly. bao.json is
back to its origin/main shape, unchanged.

Adds a module-eval case to checks.module-eval-grafana that reads the
collector's own evaluated config and asserts its service.name matches
every selector the shipped dashboard uses; verified invert-proof by
setting the value back to cfg.machine and confirming that specific
case (and only it) fails.
2026-09-25 08:30:25 +02:00
atlas
113f3fe6e2 hive-forge: a first authelia login creates the forge account
[oauth2_client] turns on auto-registration through the authelia login
source, with the account named after authelia's preferred_username.
DISABLE_REGISTRATION stays true: forgejo 16's auto-registration checks
only ALLOW_ONLY_INTERNAL_REGISTRATION, so local sign-up stays off.

ACCOUNT_LINKING is `login`, forgejo's default, set explicitly. With
`auto`, an SSO login whose name matches an existing local account would
be handed that account, and agents, `core` and `swarm-controller` all
have one. `login` asks for that account's own password instead.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
44009dc51d swarm-bao: stop linking the container's journal onto the host
--link-journal=host bind-mounts a host directory journald never
writes into for this container (empty, root:nogroup, confirmed on
the live host — #4527). Dropping it falls back to nixpkgs' default
--link-journal=try-guest, the same shape every agent container
already uses, so the in-container collector's journald receiver
(directory = /var/log/journal) now reads a journal that is actually
written.

merge stays true and services.hyperhive.swarm.otel.journaldUnits is
untouched so this deploy changes exactly one thing; comments that
described the old host-linked shape are rewritten to match.

Adds a module-eval assertion (bao-otel-collector.nix) that the
container's extraFlags never re-add --link-journal=host.

Refs #4527, #4499.
2026-09-25 04:02:08 +02:00
atlas
92e1909caf swarm-bao: give the store forwarder's OIDC reader its own bao identity
`swarm-bao-forwarder-oidc` fetches the store container's collector secret,
one path, and was the last reader still logging in with
`deploy.bao.clientCertFile`: the hive's own leaf, whose policy reads every
agent's credentials, the hive's tree and every service's OIDC secret. The
four-way split gave grafana's and the swarm collector's readers leaves of
their own and left this one behind.

It now holds `forwarder-oidc.pem`, minted by `swarm-bao-pki`, and logs in
under the `swarm-forwarder-oidc` cert-auth role, whose policy reads
`secret/data/swarm/services/<store forwarder client id>/oidc/client` and
nothing else. The role is written by `swarm-bao-forwarder-oidc-policy`
from the bootstrap token, which gains the two grants that unit calls, and
the reader is ordered after it. The subject is reserved as a hive name. A
store host whose pair is null is refused at eval rather than falling back
to the hive's leaf.

The hive's own role and `client.pem` are untouched; nothing is revoked.
2026-09-25 00:37:31 +02:00
atlas
978164dc53 nix: run the forge on one host per swarm (deploy.forgejo.enable)
Every hive with hyperhive enabled ran its own hive-forge container, and
its gateway answered forge.<swarm> with its own bridge IP, so on a
multi-host swarm each hive talked to its own forge.

deploy.forgejo.enable defaults to false and allSwarmServices sets it with
mkDefault, like authelia and bao; singleHostSwarm gets it through that.
The forge's OIDC client moves to a glue module gated on authelia, so a
split authelia/forge swarm still registers it. CI now requires the forge
on the same host, and the controller's forgeTokenFile defaults to null
where the forge is not.

Closes #4705
Refs #3782
2026-09-24 23:56:07 +02:00
atlas
6d0c30ade2 module-eval: pin the agent forge-token fetch and tea-login's removal
Refs #3782
2026-09-24 17:48:53 +02:00
atlas
a5259146dc swarm: default every queue URL to the queue's name on every hive
A remote hive dialled nothing until an operator copied the queue's URL
into it, though the URL is the same string everywhere. statusPublish.natsUrl,
queue.agentNatsUrl and controller.queue.natsUrl now default to
tls://<swarm.nats.domain>:<port> unconditionally.

The statusPublish assertion treated a URL without a secret as a half
config. With the URL a default on every hive, only the secret claims
publishing: the assertion now refuses a secret without a URL or token
endpoint, and hive-c0re's status environment is gated on the secret too,
so a hive without one publishes nothing instead of reading a missing
credential.
2026-09-24 17:26:31 +02:00
atlas
0081d75c86 swarm-bao: grant the bootstrap token the queue's pki role, policy and login role
swarm-bao-nats-tls-policy acts with the bootstrap token, and main's
module-eval-bao-grants now fails any such unit whose calls the policy
file does not grant. Adds its three paths and counts it among the units
the check must see.
2026-09-24 17:26:31 +02:00
atlas
1d261b3fed swarm-nats: give the queue a name, a bao-issued leaf, and require TLS
The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.

- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
  `swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
  every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
  IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
  may only `update` `pki/issue/swarm-nats`, written by
  `swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
  paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
  as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
  the container, restarts nats when it rotates, and re-runs daily.
  It joins glue-bao-readers-policy-order, so it is ordered after its policy
  unit (`after` and `wants`, never `requires`) where the store is on the
  same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
  `allow_non_tls`. `validateConfig` is now off in every mode, because the
  build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
  host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
  `singleHostSwarm`, and the auth responder all dial
  `tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
  to the NATS connection too, so hive-c0re and the controller trust the
  leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
  swarm needs.

module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.

Closes #4626
2026-09-24 17:26:31 +02:00
atlas
e3c595857e swarm-bao: keep the bootstrap policy in one file, checked against the units using it (#4698)
setup.md's copy of the swarm-bootstrap policy still granted only the
controller's first six paths, while eight units now act with the
bootstrap token. The policy moves to
nix/host-modules/swarm-bao-bootstrap-policy.hcl, now covering every path
those units call. module-eval-bao-grants reads that file and fails
when a unit whose script uses the token calls a path the file does not
grant.
2026-09-24 15:55:38 +02:00
atlas
0bfe354b6d nix: move the reader-after-policy edges into their own colocation glue
The four readers' ordering after their policy units only applies where
the store and that reader share a host, so it is colocation glue and
does not belong in the reader modules (two of which are main modules).
glue-bao-readers-policy-order.nix now sets the after+wants edges, gated
on deploy.bao.enable AND the reader's own gate, so a store host without
a reader gains no stub unit.
2026-09-24 15:15:15 +02:00
atlas
cde3fec956 module-eval: pin each swarm-bao reader's ordering after its policy unit 2026-09-24 15:15:15 +02:00
atlas
aa719da571 nix: ship the journals of the units an apply can leave failed
The units on the path a deploy takes to TLS, the store's grants and the
swarm collector itself were not on the host collector's journald
allowlist, so an ingest outage one of them caused showed in the store
only as every source going quiet at once.

Each module names its own units, per the option's rule:
- hive-tls.nix: hive-tls-ca, swarm-services-cert
- hive-gateway: hive-gateway-self-signed-cert (self-signed mode only)
- swarm-bao.nix: the seven grant units beside
  swarm-bao-services-issuer-policy
- swarm-otel.nix: container@<machine>, and
  nixos-rebuild-switch-to-configuration, the transient unit nixos-rebuild
  runs the activation in and whose syslog lines carry its status

The module-eval arm pins each unit as both listed and defined, since a
listed name that matches nothing is silent.
2026-09-24 15:14:44 +02:00