Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/credentials.md
atlas 7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00

24 KiB

Credentials: the target shape

The swarm's credential store is bao. This page describes the shape every credential is meant to have — who mints it, who reads it, and how it renews — not what's on disk today. secrets.md remains the map of the files that exist right now; this page replaces it, and secrets.md gets deleted, once the swarm's credential path matches what's described below.

Public material is a value. The store hands a certificate or a public nkey to every client that connects, so it's a fine place for that material. Nothing below is about those.

No secret the store holds is ever written to disk. That's the invariant, and everything else in this page follows from it. A value pulled from the store — bao — lives in the memory of the process that asked for it and nowhere else: not in a state directory, not in a bind-mounted file, not in a systemd credential, not in a rendered config, not for a moment before a unit deletes it. The target isn't a shorter list of secret files. It's the store, plus one file per identity.

Those files are mTLS client certificates, one per identity, and they're the only credential on disk. Each has to be a file, and the reason is the whole asymmetry: the certificate is what authenticates a principal to the store, so it's the one credential nothing can fetch from the store. Something has to exist on disk before the first request, or there's nothing to make the request with. Every identity — an agent, a hive, a swarm-level service — needs one; a host running several holds several, and its only power is to ask the store for the rest. swarm-bao.nix:529-533 states the rule for the nix option that carries it: this is "the credential an operator places by hand", and "a path, never a value." A literal in a nix expression lands in the nix store — world-readable and permanent — so that option takes a path to the certificate on disk, never the certificate's bytes.

The hive hands an agent an identity, never a secret. Its hive passes an agent container an mTLS certificate, and from then on the agent authenticates to the store under its own name, pulling what it needs when it needs it. No process reads a secret on another principal's behalf: the principal that needs a value is the principal that authenticates for it.

One per-agent credential file — the github token — sits outside this page: it's operator-supplied and never passes through the store, so the table below doesn't govern it.

Per secret, the target specifies minter, reader, and renewal strategy. Those three are the contract, and the reader is a process pulling a store path at runtime — not a path on disk, and not a unit whose job is to turn a store value into a file. A renewal cell may never read NONE: state the strategy for every credential, including the mTLS leaf.

Renewal is two columns. Automatic re-mint is something replacing the stored value without an operator. Automatic re-pull is the reader picking up a replaced value without a restart. ✅ or ❌ says which exist today, and the rest of the cell says how.

store path minter reader — pulls at runtime, holds in memory automatic re-mint automatic re-pull
swarm/agents/<agent>/matrix/main swarm-controller, with the swarm's appservice token, at agent creation and in a five-minute pass the agent container itself, under the certificate its hive passed in ✅ the pass re-mints when the stored token is missing, unknown to the homeserver, or someone else's ✅ hive-matrix-daemon exits when the homeserver rejects its token, and a five-minute timer restarts it, which reads the store again
swarm/agents/<agent>/matrix/<account> swarm-controller the agent container itself, under the certificate its hive passed in must be stated must be stated
swarm/controller/swarm-controller/matrix/appservice-token swarm-matrix-ctl, inside the hive-matrix container, once swarm-controller, under its own certificate ❌ swarm-matrix-ctl mints it once; the container keeps its copy and republishes it when the store's differs ✅ the controller reads it on every five-minute matrix pass
swarm/controller/swarm-controller/oidc/client authelia, at its first boot, where the controller registers its client; swarm-secret-publish copies it in swarm-controller, under its own certificate, once at start ❌ authelia mints it once. A re-mint is republished by swarm-secret-publish's path unit ❌ read once at start; the controller holds the old value until it restarts
swarm/agents/<agent>/bao-mtls the store's agent PKI mount (deploy.bao.agentPkiMountPath), which generates the key, at swarm-controller's request at agent creation hive-c0re, under the hive's own certificate, when it writes the agent's container config ✅ swarm-controller's five-minute pass re-issues a live agent's leaf once it's past half its validity (45 of 90 days, read from the certificate itself) ❌ hive-c0re reads it when it writes the container config, so the agent presents a new leaf from its next start; the old leaf stays valid until it expires
swarm/agents/<agent>/queue swarm-controller, at agent creation hive-agent in the agent container, under the agent's own certificate, held in memory — the identity it presents to the swarm queue, naming that one agent rather than its hive ✅ swarm-controller's five-minute pass re-mints a live agent's secret once it's 45 days old by minted_at on the stored object; a secret with no minted_at gets one stamped, value unchanged. The pass skips agents declared Destroyed — declaring an agent destroyed deletes every version of the path instead, the undo of the mint rather than another one ✅ hive-agent reads the path before its first connect and again on every reconnect attempt, so a reconnect after a re-mint presents the new secret. An open connection keeps the secret it connected with; after a revocation the agent keeps retrying under the queue client's backoff
swarm/agents/<agent>/forge-token swarm-controller, at agent creation and in a pass every 5 minutes over every agent with a store identity the agent container itself, under its own certificate, fetched to /run/hive-agent-forge-token/token ✅ the controller re-mints when the stored token is missing or no longer matches the forge (last eight characters and scopes) ✅ the agent re-fetches on a 10-minute timer
swarm/hives/<hive>/matrix/appservice-token one minter, on the authelia host the hive process that presents the token to its homeserver, under the hive's own certificate must be stated must be stated
swarm/hives/<hive>/matrix/sender-token swarm-controller, with the swarm's appservice token, for every hive in its directory in a five-minute pass swarm-controller under its own certificate, before it decides whether to mint, and hive-c0re's stored_sender_token(), under the hive's own certificate ✅ the controller's pass re-mints when the stored token is missing, unknown to the homeserver, or someone else's ✅ hive-c0re's matrix sweep reads the store every run and overwrites its token file when the store's token differs
swarm/hives/<hive>/queue/agent authelia swarm-bao-queue-agent on the hive's host, under its own per-hive certificate; no agent's policy reaches it must be stated must be stated
swarm/services/<clientId>/oidc/client authelia the service process that presents the client secret, under the certificate of the host it runs on must be stated must be stated
(not in the store) a hive's mTLS leaf the store's own PKI, or an operator placing it by hand its own client, off disk — the exception above, because it's what makes every other row's pull possible must be stated must be stated

Two rows share the matrix/ prefix and have different minters, on purpose. An agent's main account is on the swarm's own homeserver, and swarm-controller creates it with the swarm's appservice token. Every other account under that prefix is somewhere else entirely, and an operator hands the controller a credential for it. The operator-facing route refuses the name main for exactly this reason: two writers, one name, and that refusal is what keeps them apart.

The swarm's appservice token sits under controller/, the one kind no hive's policy reads. Its sender is the homeserver's admin. Under agents/, hives/ or services/ every hive could read it. matrix-ctl may write and read that one leaf; the controller may only read it.

An agent's mTLS leaf is in the store; a hive's isn't, and the difference isn't an inconsistency. The rule the exception protects is that nothing can fetch from the store the credential it would need in order to fetch. A hive's leaf is that credential, so it can only come off disk. The hive, which already holds one, reads an agent's — so publishing it costs nothing and buys the property this page asks for: the swarm mints it, the hive only carries it, and no hive ever needs the capability to mint an identity. swarm-controller proves the leaf it publishes before the creation job reports success, by logging in with it and reading the row back.

Backfilling an agent that predates a credential. Agent creation at swarm level is purely event-driven — swarm-controller mints an agent's store identity on the job graph POST /api/agents inserts, and nothing sweeps for agents that already exist. An agent created before a credential joined that mint therefore never receives one, and nothing will ever come back around to it. Re-run the mint for one agent with:

swarmctl agent mint-identity <agent>

The queue secret half is idempotent — an agent that already has one keeps exactly the value it holds, so running this against an already-migrated agent doesn't drop its queue connection. The certificate half isn't: the agent gets a fresh leaf and picks it up on its next boot.

The renewal pass in the table doesn't replace this step: it only re-issues a certificate or re-mints a queue secret that already exists, and only for an agent some hive's wanted state declares as anything but destroyed.

The forge token needs no such step: swarm-controller checks every agent that has a store identity at start and every five minutes. For any whose stored token is missing or stale it creates the forge user if there isn't one, then mints the token. To check one agent now:

swarmctl agent mint-forge-token <agent>

An agent without a store identity gets no swarm token and keeps using the forge-token file in its state dir, if it has one; run mint-identity for it first.

⚠️ Run this for every existing agent before deploying a hive-side change that makes a container require a credential it may not have. A container whose credential is absent doesn't start — that's deliberate, and it's what makes the backfill a step rather than a suggestion.

Revoking an agent's queue credential. Declaring an agent destroyed — PUT /api/hives/{hive}/agents/{agent}/state — deletes swarm/agents/<agent>/queue and every version it ever held, right after the declaration lands. The credential is a bearer secret, so a soft delete would leave the same value readable at an older version number; the controller removes the path's metadata, which takes the versions with it.

An agent recreated under the same name draws a fresh secret, because the mint keeps an existing value only when it finds one at that path (agent_identity::mint_and_verify, step 3) and the revocation left nothing to find.

Two things this deliberately doesn't do. It doesn't touch the agent's mTLS leaf, its ACL document or its cert-auth role — the mint rewrites all three on every run, so a re-created agent gets new ones regardless. And it doesn't fail the destroy: the declaration is already published by the time the revocation runs, so a store that refuses the delete gets a swarm-controller log line at error naming the agent, and the teardown continues. Declare the agent destroyed again to re-run the revocation.

Who reads that row, and what happens to it. hive-c0re reads it every time it writes an agent's container configuration (lifecycle::agent_identity), stages the certificate and its key 0600 outside every bind-mounted tree, and passes both to the container as systemd credentials — the same mechanism, and for the same mode reason, as the per-hive queue secret. A bind mount would hand the agent's unprivileged user a file it lacks the rights to open; the container manager reads a credential as root and re-exposes it under the consuming unit's own user.

Inside the container, hive-agent-bao-identity.service logs in with that certificate and reads this row back before reporting success, so an agent locked out of its own identity says so at boot rather than at whichever pull needed the store first. The unit exists whenever services.hyperhive.agent.bao.addr has a value, which the hive's meta flake sets from its own store address — the same all-or-nothing gate the per-hive queue credential beside it uses, and the reason the delivery above never lands in a container with nothing to read it. It fails loudly where the hive-side readers treat a missing value as a normal state, which is deliberate: a missing queue secret means a swarm whose publisher has yet to run, while a refused certificate means an agent that believes it reaches the store and never does.

Both kinds of unit retry a failed login every 30 seconds for a day (nix/host-modules/lib/store-retry.nix). A unit ordered after one waits for a single attempt, not for the retries: an attempt that fails ends that start, so the dependent starts with whatever is already on disk. When a later attempt lands a changed value, the homeserver, Grafana and collector readers restart the service that loaded it (nix/host-modules/lib/refresh-consumer.nix).

Progressive enhancement

New functionality has to match this shape immediately — no PR introducing a credential gets a pass on any of the rules below. A PR can move existing functionality step by step, as long as each individual step moves toward the target shape; a step that doesn't isn't allowed just because it's existing.

A pull request that touches a credential can't:

  • add a minter outside the swarm's existing mint path
  • persist a store-provided secret to disk — a state directory, a bind mount, a rendered config
  • add a credential whose renewal strategy is NONE — state the strategy, even if it's "operator reissues and restarts the reader"
  • read a secret on another principal's behalf and hand it over — the principal that needs the value authenticates for it
  • give a host or container an out-of-band credential that isn't the store mTLS leaf — one out-of-band credential per principal is the whole point of the store

While the swarm's credential path is still moving to this shape, a PR that moves a secret into bao may leave its renewal strategy unresolved, provided it opens a follow-up issue to settle renewal. That's a migration-era allowance, not a standing exception to the renewal-strategy rule above.