hyperhive/swarm-queue-client
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 7a03ce096a hive-c0re: deliver an agent's credential from the store to its state dir
mara on #4015: "not merging code without callers", and on the same PR
"see issue, we decided what the first thing should be". #3726 decided it:
the controller writes a token to the store and tells the hive; the hive
reads it back and writes /agents/<agent>/state/matrix-token-<account> at
0600, where matrix.nix's existing systemd.paths glob re-fires the daemon.
So this is the hive half of that, and the library's first caller.

The notice names a credential and never carries one, and deploy_subject's
own doc is why: the auth-callout responder scopes publish and leaves sub
unrestricted, so a hive that wanted another's messages could subscribe to
them. A secret in that payload would be readable swarm-wide. The value is
read from the store under the reading hive's own certificate, where the
store's policy is what actually scopes it.

Two boundaries guard the two addresses, and they are not the same check.
`path::matrix_account` guards the address in the store. `Ident` guards the
address on disk -- `agent_state_dir` takes one, so an unvalidated name off
the queue cannot reach a directory. I had written the first and assumed it
covered both; the compiler refused the `&str` and was right. `token_path`
now takes the newtype so a call site cannot forget.

The write is atomic because the path-watcher fires on the file appearing:
written in place it would be visible while partial, and the daemon would
read a truncated credential exactly once, which is the hardest possible
failure to reproduce. The temp name is dot-prefixed so it cannot match the
`matrix-token*` glob on its way past.

The publish grant is here because without it the failure is invisible.
policy.rs already says why for its siblings: a refused publish reaches the
client as a timeout, so the symptom is a hive that never receives a
credential with nothing in either log naming a permission. Two tests: the
controller may publish, a hive may not -- its own subject included. A
forged notice leaks nothing, but it would make a hive fetch and overwrite
a token file for a name the forger chose.

Refs #3726
2026-09-03 00:29:52 +02:00
..
src hive-c0re: deliver an agent's credential from the store to its state dir 2026-09-03 00:29:52 +02:00
Cargo.toml move otel_http_client from swarm-queue-client into swarm-controller 2026-08-29 11:17:24 +02:00
README.md treefmt: apply prettier 2026-09-02 15:25:07 +02:00

swarm-queue-client

Connecting to the swarm message queue as an authenticated client. Shared by every process that participates: the swarm controller reads hive status out of the queue, a hive publishes its own status into it.

Why a crate and not a module per binary

The connect is identical for every participant — mint an authelia token, present it at CONNECT for the auth_callout responder to introspect, let async-nats re-run the callback on each connection attempt. Only the use differs.

Two copies of that would be two copies of credential handling, and a token-refresh fix would have to be found twice. The same reasoning already put hive-sock-client in its own crate rather than in each daemon that speaks to a unix socket.

The two properties that constrain the code

A token expires. Authelia issues client_credentials access tokens with expires_in: 3599. Authentication happens at CONNECT, so a long-lived connection is fine — but a reconnect an hour later needs a token minted an hour later.

The refresh therefore lives in the auth callback, not in a timer. async-nats invokes it per connection attempt, so there is no window in which the client holds a token it minted for a previous connection. The alternative — mint once, own the reconnect loop — fails in the way this subsystem exists to prevent: the process keeps serving while its data quietly stops moving, and nothing says so until someone reads a dashboard.

Configuration

QueueConfig::from_env(prefix) reads <prefix>_NATS_URL, <prefix>_OIDC_TOKEN_ENDPOINT, <prefix>_OIDC_CLIENT_ID and <prefix>_OIDC_CLIENT_SECRET_FILE.

The prefix is a parameter because the variables belong to the consuming unit — a NixOS module sets them alongside its other options. What is shared is the rule, not the spelling: all four together or none at all. A half-set environment is a hard error, because the failure it would otherwise produce is the expensive kind — the process comes up "fine", never connects, and the data it was supposed to move silently stops.

The client secret is a path, not a value: putting it in the environment would publish it to anything that can read /proc/<pid>/environ. It is read per token request rather than cached, so a rotation the operator believes took effect actually did.

What this crate does not do

It ends at a connected client. jetstream/kv are off by default — what a consumer does with the connection is its own business, and its Cargo.toml is where that requirement should be visible. The auth-callout responder speaks the connect and nothing else, and pays for nothing else.

The one exception: the kv feature

kv adds status, which holds the name and the creation config of the hive-status bucket — nothing more.

It is here because that bucket has two ends in two crates: a hive writes its own key, the controller reads every key. The name being a repeated literal is the mild half of the problem; the sharp half is that either end may arrive first on a fresh swarm, so both create the bucket if it is missing. Two Configs that drift means whichever end created it wins and the other opens a bucket it did not ask for — no error, no log, just a retention policy nobody chose.

An agreement between two crates has to live in one of them, and neither end of this bucket is senior to the other. Behind a default-off feature, the consumer that needs none of it still pays nothing.