A five-minute pass over every agent some hive's wanted state declares as anything but destroyed queues, per agent: - `MintAgentIdentity` (the node agent creation uses) when the stored certificate at swarm/agents/<agent>/bao-mtls is past half its validity, read from its own notBefore/notAfter: day 45 of the role's 90; - the new `RenewAgentQueueCredential` node when the queue secret at swarm/agents/<agent>/queue is 45 days old or has no mint time. The node re-decides, writes a fresh value with `minted_at`, reads it back, and logs the agent and the old age. When both are due the secret node runs after_any the certificate node, because mint_and_verify compares the queue secret it read with the one it reads back. A credential that is not stored is never created here. `queue::AgentCredential` gains an optional `minted_at` (unix seconds); agent creation now sets it. Stored objects without it decode unchanged and count as due, so every existing queue secret is re-minted on the first pass. Both replacements reach the agent at its next start. The old certificate stays valid until it expires; the old queue secret does not, so a queue reconnect before that restart is denied. Adds x509-cert 0.2 (with der_derive and flagset) to read the validity. docs/swarm/credentials.md: the renewal column splits into automatic re-mint and automatic re-pull, filled from the code as it stands.
184 lines
18 KiB
Markdown
184 lines
18 KiB
Markdown
# Credentials: the target shape
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
The swarm's credential store is bao. This page describes the shape every
|
|
credential is meant to have — who mints it, who reads it, and how it renews
|
|
— not what's on disk today. [`secrets.md`](secrets.md) remains the map of
|
|
the files that exist right now; this page replaces it, and `secrets.md` gets
|
|
deleted, once the swarm's credential path matches what's described below.
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
**Public material is a value.** The store hands a certificate or a public
|
|
nkey to every client that connects, so it's a fine place for that material.
|
|
Nothing below is about those.
|
|
|
|
**No secret the store holds is ever written to disk.** That's the invariant,
|
|
and everything else in this page follows from it. A value pulled from the
|
|
store — bao — lives in the memory of the process that asked for it and
|
|
nowhere else: not in a state directory, not in a bind-mounted file, not in a
|
|
systemd credential, not in a rendered config, not for a moment before a unit
|
|
deletes it. The target isn't a shorter list of secret files. It's the store,
|
|
plus one file per identity.
|
|
|
|
**Those files are mTLS client certificates, one per identity, and they're
|
|
the only credential on disk.** Each has to be a file, and the reason is the
|
|
whole asymmetry: the certificate is what authenticates a principal to the
|
|
store, so it's the one credential nothing can fetch from the store.
|
|
Something has to exist on disk before the first request, or there's nothing
|
|
to make the request with. Every identity — an agent, a hive, a swarm-level
|
|
service — needs one; a host running several holds several, and its only
|
|
power is to ask the store for the rest.
|
|
`swarm-bao.nix:529-533` states the rule for the nix option that carries it:
|
|
this is _"the credential an operator places by hand"_, and _"a path, never a
|
|
value."_ A literal in a nix expression lands in the nix store —
|
|
world-readable and permanent — so that option takes a path to the
|
|
certificate on disk, never the certificate's bytes.
|
|
|
|
**The hive hands an agent an identity, never a secret.** Its hive passes an
|
|
agent container an mTLS certificate, and from then on the agent authenticates
|
|
to the store under its own name, pulling what it needs when it needs it. No
|
|
process reads a secret on another principal's behalf: the principal that
|
|
needs a value is the principal that authenticates for it.
|
|
|
|
One per-agent credential file — the github token — sits outside this page:
|
|
it's operator-supplied and never passes through the store, so the table
|
|
below doesn't govern it.
|
|
|
|
**Per secret, the target specifies minter, reader, and renewal strategy.**
|
|
Those three are the contract, and the reader is a process pulling a store
|
|
path at runtime — not a path on disk, and not a unit whose job is to turn a
|
|
store value into a file. A renewal cell may never read `NONE`: state the
|
|
strategy for every credential, including the mTLS leaf.
|
|
|
|
Renewal is two columns. **Automatic re-mint** is something replacing the
|
|
stored value without an operator. **Automatic re-pull** is the reader picking up
|
|
a replaced value without a restart. ✅ or ❌ says which exist today, and the rest
|
|
of the cell says how.
|
|
|
|
<!-- vale write-good.Passive = NO -->
|
|
|
|
| store path | minter | reader — pulls at runtime, holds in memory | automatic re-mint | automatic re-pull |
|
|
| ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
| `swarm/agents/<agent>/matrix/main` | `swarm-controller`, with the swarm's appservice token, at agent creation and in a five-minute pass | the agent container itself, under the certificate its hive passed in | ✅ the pass re-mints when the stored token is missing, unknown to the homeserver, or someone else's | ✅ `hive-matrix-daemon` exits when the homeserver rejects its token, and a five-minute timer restarts it, which reads the store again |
|
|
| `swarm/agents/<agent>/matrix/<account>` | `swarm-controller` | the agent container itself, under the certificate its hive passed in | must be stated | must be stated |
|
|
| `swarm/controller/swarm-controller/matrix/appservice-token` | `swarm-matrix-ctl`, inside the `hive-matrix` container, once | `swarm-controller`, under its own certificate | ❌ `swarm-matrix-ctl` mints it once; the container keeps its copy and republishes it when the store's differs | ✅ the controller reads it on every five-minute matrix pass |
|
|
| `swarm/controller/swarm-controller/oidc/client` | authelia, at its first boot, where the controller registers its client; `swarm-secret-publish` copies it in | `swarm-controller`, under its own certificate, once at start | ❌ authelia mints it once. A re-mint is republished by `swarm-secret-publish`'s path unit | ❌ read once at start; the controller holds the old value until it restarts |
|
|
| `swarm/agents/<agent>/bao-mtls` | the store's agent PKI mount (`deploy.bao.agentPkiMountPath`), which generates the key, at `swarm-controller`'s request at agent creation | `hive-c0re`, under the hive's own certificate, when it writes the agent's container config | ✅ `swarm-controller`'s five-minute pass re-issues a live agent's leaf once it's past half its validity (45 of 90 days, read from the certificate itself) | ❌ `hive-c0re` reads it when it writes the container config, so the agent presents a new leaf from its next start; the old leaf stays valid until it expires |
|
|
| `swarm/agents/<agent>/queue` | `swarm-controller`, at agent creation | the agent container itself, under its own certificate — the identity it presents to the swarm queue, naming that one agent rather than its hive | ✅ `swarm-controller`'s five-minute pass re-mints a live agent's secret once it's 45 days old or has no recorded mint time (`minted_at` on the stored object) | ❌ fetched when the container starts. The queue checks the secret only at connect, so an open connection survives a re-mint, but a reconnect before the next restart is denied |
|
|
| `swarm/agents/<agent>/forge-token` | `swarm-controller`, at agent creation and in a pass every 5 minutes over every agent with a store identity | the agent container itself, under its own certificate, fetched to `/run/hive-agent-forge-token/token` | ✅ the controller re-mints when the stored token is missing or no longer matches the forge (last eight characters and scopes) | ✅ the agent re-fetches on a 10-minute timer |
|
|
| `swarm/hives/<hive>/matrix/appservice-token` | one minter, on the authelia host | the hive process that presents the token to its homeserver, under the hive's own certificate | must be stated | must be stated |
|
|
| `swarm/hives/<hive>/matrix/sender-token` | `swarm-matrix-ctl`, in the `hive-matrix` container | `swarm-matrix-ctl` itself, under its own certificate, before it decides whether to mint, and hive-c0re's `stored_sender_token()`, under the hive's own certificate | must be stated | must be stated |
|
|
| `swarm/hives/<hive>/queue/agent` | authelia | `swarm-bao-queue-agent` on the hive's host, under its own per-hive certificate; no agent's policy reaches it | must be stated | must be stated |
|
|
| `swarm/services/<clientId>/oidc/client` | authelia | the service process that presents the client secret, under the certificate of the host it runs on | must be stated | must be stated |
|
|
| _(not in the store)_ a hive's mTLS leaf | the store's own PKI, or an operator placing it by hand | its own client, off disk — the exception above, because it's what makes every other row's pull possible | must be stated | must be stated |
|
|
|
|
<!-- vale write-good.Passive = YES -->
|
|
|
|
**Two rows share the `matrix/` prefix and have different minters, on
|
|
purpose.** An agent's `main` account is on the swarm's own homeserver, and
|
|
`swarm-controller` creates it with the swarm's appservice token. Every other
|
|
account under that prefix is somewhere else entirely, and an operator hands
|
|
the controller a credential for it. The operator-facing route refuses the
|
|
name `main` for exactly this reason: two writers, one name, and that refusal
|
|
is what keeps them apart.
|
|
|
|
**The swarm's appservice token sits under `controller/`, the one kind no
|
|
hive's policy reads.** Its sender is the homeserver's admin. Under
|
|
`agents/`, `hives/` or `services/` every hive could read it. matrix-ctl may
|
|
write and read that one leaf; the controller may only read it.
|
|
|
|
**An agent's mTLS leaf is in the store; a hive's isn't, and the difference
|
|
isn't an inconsistency.** The rule the exception protects is that nothing
|
|
can fetch from the store the credential it would need in order to fetch. A
|
|
hive's leaf is that credential, so it can only come off disk. The _hive_,
|
|
which already holds one, reads an agent's — so publishing it costs nothing
|
|
and buys the property this page asks for: the swarm mints it, the hive only
|
|
carries it, and no hive ever needs the capability to mint an identity.
|
|
`swarm-controller` proves the leaf it publishes before the creation job
|
|
reports success, by logging in with it and reading the row back.
|
|
|
|
**Backfilling an agent that predates a credential.** Agent creation at swarm
|
|
level is purely event-driven — `swarm-controller` mints an agent's store
|
|
identity on the job graph `POST /api/agents` inserts, and nothing sweeps for
|
|
agents that already exist. An agent created before a credential joined that
|
|
mint therefore never receives one, and nothing will ever come back around to
|
|
it. Re-run the mint for one agent with:
|
|
|
|
```sh
|
|
swarmctl agent mint-identity <agent>
|
|
```
|
|
|
|
The queue secret half is idempotent — an agent that already has one keeps exactly the
|
|
value it holds, so running this against an already-migrated agent doesn't drop
|
|
its queue connection. The certificate half isn't: the agent gets a fresh leaf
|
|
and picks it up on its next boot.
|
|
|
|
The renewal pass in the table doesn't replace this step: it only re-issues a
|
|
certificate or re-mints a queue secret that already exists, and only for an
|
|
agent some hive's wanted state declares as anything but `destroyed`.
|
|
|
|
The forge token needs no such step: `swarm-controller` checks every agent
|
|
that has a store identity at start and every five minutes. For any whose
|
|
stored token is missing or stale it creates the forge user if there isn't one,
|
|
then mints the token. To check one agent now:
|
|
|
|
```sh
|
|
swarmctl agent mint-forge-token <agent>
|
|
```
|
|
|
|
An agent without a store identity gets no swarm token and keeps using the
|
|
`forge-token` file in its state dir, if it has one; run `mint-identity` for it
|
|
first.
|
|
|
|
⚠️ **Run this for every existing agent before deploying a hive-side change
|
|
that makes a container require a credential it may not have.** A container
|
|
whose credential is absent doesn't start — that's deliberate, and it's what
|
|
makes the backfill a step rather than a suggestion.
|
|
|
|
**Who reads that row, and what happens to it.** `hive-c0re` reads it every
|
|
time it writes an agent's container configuration
|
|
(`lifecycle::agent_identity`), stages the certificate and its key `0600`
|
|
outside every bind-mounted tree, and passes both to the container as systemd
|
|
credentials — the same mechanism, and for the same mode reason, as the
|
|
per-hive queue secret. A bind mount would hand the agent's unprivileged user
|
|
a file it lacks the rights to open; the container manager reads a credential
|
|
as root and re-exposes it under the consuming unit's own user.
|
|
|
|
Inside the container, `hive-agent-bao-identity.service` logs in with that
|
|
certificate and reads this row back before reporting success, so an agent
|
|
locked out of its own identity says so at boot rather than at whichever pull
|
|
needed the store first. The unit exists whenever
|
|
`services.hyperhive.agent.bao.addr` has a value, which the hive's meta flake
|
|
sets from its own store address — the same all-or-nothing gate the per-hive
|
|
queue credential beside it uses, and the reason the delivery above never lands
|
|
in a container with nothing to read it. It fails loudly where the hive-side
|
|
readers degrade quietly, which is deliberate: a missing queue secret means a
|
|
swarm whose publisher has yet to run, while a refused certificate means an
|
|
agent that believes it reaches the store and never does.
|
|
|
|
## Progressive enhancement
|
|
|
|
New functionality has to match this shape immediately — no PR introducing a
|
|
credential gets a pass on any of the rules below. A PR can move existing
|
|
functionality step by step, as long as each individual step moves toward the
|
|
target shape; a step that doesn't isn't allowed just because it's existing.
|
|
|
|
A pull request that touches a credential can't:
|
|
|
|
- add a minter outside the swarm's existing mint path
|
|
- persist a store-provided secret to disk — a state directory, a bind mount,
|
|
a rendered config
|
|
- add a credential whose renewal strategy is `NONE` — state the strategy,
|
|
even if it's "operator reissues and restarts the reader"
|
|
- read a secret on another principal's behalf and hand it over — the
|
|
principal that needs the value authenticates for it
|
|
- give a host or container an out-of-band credential that isn't the store
|
|
mTLS leaf — one out-of-band credential per principal is the whole point of
|
|
the store
|
|
|
|
While the swarm's credential path is still moving to this shape, a PR that
|
|
moves a secret into bao may leave its renewal strategy unresolved, provided
|
|
it opens a follow-up issue to settle renewal. That's a migration-era
|
|
allowance, not a standing exception to the renewal-strategy rule above.
|