| Filename | Latest commit message | Latest commit date |
|---|---|---|
`swarm/agents/<agent>/bao-mtls` did not exist, and neither did any per-agent identity at the secret store: `policy::agent_object_name`, `render_agent` and `render_agent_with_queue` had been written and never called outside their own tests. An agent's only "per-agent" secret today is read under the HIVE's certificate, through a wide grant on `swarm/agents/*` — so "per-agent" was presentational. The swarm now mints the certificate, so no hive ever needs the capability to mint one. `swarm-controller` is the service that does it: it already logs in to the store, and its existing grant already covers exactly the three objects written here (`create/update` on `secret/data/swarm/agents/*`, `sys/policies/acl/hive-*` and `auth/cert/certs/hive-*`). No new bao grant, and nothing co-located — a cert-auth role pins its authority by value, per role, so the controller issues from its own CA on its own host and pins that CA in the role it writes. No existing role changes. The mint node does not report success on a write. After publishing it connects again, with the leaf it just issued and under the role it just wrote, and reads the path back — so the policy, the role, the common name and the leaf are exercised in production on every agent creation. A certificate this code mints that the role this code writes will not accept turns the job node red at creation time instead of surfacing later as an agent container that cannot start. `TriggerDeploy` gains an `after_any` edge on the mint, not `after_ok`: a hive cannot pass down a certificate the swarm has not published, but a host with no authority configured must still create agents exactly as it does today. The private key is generated in memory and never written to disk on the controller — `SecretStore::connect_with_identity` takes the PEM the minter is already holding, so nothing is written out purely to be logged in with. Refs #4137 |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-secret-client
Reading and writing a swarm credential in the secret store, over vaultrs.
The HTTP is that crate's job. What this one owns is the agreements both
ends of the store have to state identically: where a credential lives, which
field its bytes are in, and how this deployment's environment becomes a
logged-in client.
The surrounding picture — which secret is minted where, and why delivery is a
copy rather than a bind mount — is docs/swarm/secrets.md.
Why a crate and not a module per binary
The store has two Rust ends and they are peers: the swarm controller writes a credential, a hive reads it. Neither is senior to the other, so a path formatted at each call site is an agreement with no owner — it holds right up until one side is edited alone, and then it fails as a missing key rather than as a mismatch.
The third end is what settles it. nix/host-modules/glue-matrix-bao-token.nix
reads the store with bao kv get -field=value. That reader is a shell line in
a nix module: it cannot be renamed by the same refactor as a Rust struct, and
no Rust test reaches it. So the field name is pinned by a test against the
serialised literal rather than left to the struct definition.
The identity is a certificate, and the role is the hive's name
Authentication is the store's cert auth method.
nix/host-modules/glue-bao-tls.nix mints the client certificate with its CN
set to the hive's name, because a cert-auth role matches on the CN. So
cert_role is not a free choice for the caller: a hive passes its own name,
and the policy attached to that role is what scopes what it may read.
The listener's tls_require_and_verify_client_cert is a different thing and
not a substitute. It decides who may open a connection; it says nothing about
who the connection belongs to, and a store with the option set and no cert
mount configured refuses every login made here. That refusal is what
Error::Vault out of connect means.
Configuration
Settings::from_env reads BAO_ADDR, BAO_CLIENT_CERT, BAO_CLIENT_KEY,
and the optional BAO_CACERT.
The BAO_ spellings are read explicitly rather than left to vaultrs. Its
own defaults look for VAULT_ADDR / VAULT_CLIENT_CERT / VAULT_CLIENT_KEY,
which no unit in this tree sets. Falling through to them builds a client with
no identity at all, and that surfaces as a TLS handshake failure — a place
that names neither the variable nor the reason.
Empty is as absent as unset. systemd renders an unset nix option as
Environment=BAO_CACERT=, so empty is the shape a missing value arrives in.
The certificate and key are paths, not values, and are read at connect time — same rule as every other credential in this tree, for the same reason: a value in a nix expression is rendered into the world-readable store.
Reading the environment is separate from connecting (Settings::from_lookup)
because every one of those failures is a misconfiguration an operator has to
read an error about, and none of them needs a reachable store to happen.
Names that arrive from elsewhere
matrix::account_path is fallible, which for a string formatter needs saying:
its segments are an agent name from the topology and an account name from that
agent's own config. A / turns one agent's segment into another agent's
directory and .. walks out of the prefix entirely, so the charset it accepts
is deliberately narrower than what the store would.
One module per kind of secret
client moves whatever type a caller names; it decodes nothing itself. What a
stored object holds is stated in the module that also builds its path —
matrix today, and a second kind of swarm secret gets a module beside it.
The split is deliberate. A single shared struct that grows one field per consumer ends up carrying, on every path, a field only one path's reader has ever heard of; and the two things a kind of secret must pin — where it lives and what is in it — are one agreement that reads worse split across modules.
What this crate does not do
It has no opinion on what a caller may read. That is the policy attached to the cert role, and it lives in the store.
It holds the token minted at login and renews nothing. A handle is built per credential, so the login is the cheap part of a rare operation — a caller that wanted to keep one alive across a token's lifetime would need more than this.