hyperhive/swarm-queue-client/README.md
atlas 22659234c4 refactor(swarm-queue-client): share the hive-status bucket's name and shape
The bucket has two ends in two crates: a hive writes its own key, the
controller reads every key. `swarm-controller` declared the name as a
private const with a doc comment arguing that "reader and writer must
name the same bucket" — an argument the writer, in another crate, could
not obey.

The name is the mild half. Both ends do get-or-create, because either may
come up first on a fresh swarm and neither can assume the other has run.
Two `Config`s that drift means whichever end created the bucket wins and
the other's `get_key_value` succeeds against a bucket it did not ask for:
no error, no log, just a retention policy nobody chose. Sharing the
constructor gives that race one outcome.

Behind a default-off `kv` feature, so the crate's other consumer — the
auth-callout responder, which speaks the connect and nothing else — still
pulls neither `jetstream` nor `kv`. That was the actual reason the
feature was excluded when this crate was extracted; the flag preserves
it. The surface is deliberately narrow: one bucket's name and creation
config, not a general KV facade.
2026-08-16 13:12:59 +02:00

72 lines
3.4 KiB
Markdown

# swarm-queue-client
Connecting to the swarm message queue as an authenticated client. Shared by
every process that participates: the swarm controller reads hive status out of
the queue, a hive publishes its own status into it.
## Why a crate and not a module per binary
The *connect* is identical for every participant — mint an authelia token,
present it at CONNECT for the `auth_callout` responder to introspect, let
`async-nats` re-run the callback on each connection attempt. Only the **use**
differs.
Two copies of that would be two copies of credential handling, and a
token-refresh fix would have to be found twice. The same reasoning already put
`hive-sock-client` in its own crate rather than in each daemon that speaks to a
unix socket.
## The two properties that constrain the code
**A token expires.** Authelia issues `client_credentials` access tokens with
`expires_in: 3599`. Authentication happens at CONNECT, so a long-lived
connection is fine — but a *reconnect* an hour later needs a token minted an
hour later.
**The refresh therefore lives in the auth callback, not in a timer.**
`async-nats` invokes it per connection attempt, so there is no window in which
the client holds a token it minted for a previous connection. The alternative —
mint once, own the reconnect loop — fails in the way this subsystem exists to
prevent: the process keeps serving while its data quietly stops moving, and
nothing says so until someone reads a dashboard.
## Configuration
`QueueConfig::from_env(prefix)` reads `<prefix>_NATS_URL`,
`<prefix>_OIDC_TOKEN_ENDPOINT`, `<prefix>_OIDC_CLIENT_ID` and
`<prefix>_OIDC_CLIENT_SECRET_FILE`.
The prefix is a parameter because the variables belong to the consuming unit —
a NixOS module sets them alongside its other options. What is shared is the
rule, not the spelling: **all four together or none at all.** A half-set
environment is a hard error, because the failure it would otherwise produce is
the expensive kind — the process comes up "fine", never connects, and the data
it was supposed to move silently stops.
The client secret is a **path, not a value**: putting it in the environment
would publish it to anything that can read `/proc/<pid>/environ`. It is read
per token request rather than cached, so a rotation the operator believes took
effect actually did.
## What this crate does not do
It ends at a connected client. `jetstream`/`kv` are **off by default** — what a
consumer does with the connection is its own business, and its `Cargo.toml` is
where that requirement should be visible. The auth-callout responder speaks the
connect and nothing else, and pays for nothing else.
## The one exception: the `kv` feature
`kv` adds `status`, which holds the name and the creation config of the
`hive-status` bucket — nothing more.
It is here because that bucket has **two ends in two crates**: a hive writes its
own key, the controller reads every key. The name being a repeated literal is
the mild half of the problem; the sharp half is that either end may arrive first
on a fresh swarm, so both create the bucket if it is missing. Two `Config`s that
drift means whichever end created it wins and the other opens a bucket it did
not ask for — no error, no log, just a retention policy nobody chose.
An agreement between two crates has to live in one of them, and neither end of
this bucket is senior to the other. Behind a default-off feature, the consumer
that needs none of it still pays nothing.