| Filename | Latest commit message | Latest commit date |
|---|---|---|
The controller could not emit an event at all: a reader's grant is `reader_subjects()`, which is `$JS.API.*` only, so a publish to any event subject would be refused — and a NATS refusal reaches the client as a timeout, so the visible symptom would have been a hive that never hears about a change, with nothing in any log naming a permission. Adds `swarm_queue_client::events`, following `status::BUCKET`: three crates must agree on these strings (the controller publishes, a hive subscribes, the callout responder decides whether the publish is permitted), and a literal repeated across crates is an agreement nothing checks. The responder speaks neither jetstream nor kv, so the module is unconditional and carries no NATS types, exactly as the bucket name is. The grant takes the wildcard form from the same function the publisher calls, so the two cannot drift; a separate wildcard constant would have re-created the disagreement this module exists to prevent. Tests pin that a reader gets the subject and that a hive does NOT — a hive able to publish here could tell a neighbour the knowledge repo changed when it had not, which is an unauthenticated write into someone else's control path. That one asserts on the subject root rather than a rendered subject, so a future event leaf fails it too instead of passing because the test only knew about `knowledge`. Both assertions mutation-tested: removing the grant fails the reader test, granting a hive the subject fails the denial test, each on its own assertion line, and the unmutated tree is green. |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-queue-client
Connecting to the swarm message queue as an authenticated client. Shared by every process that participates: the swarm controller reads hive status out of the queue, a hive publishes its own status into it.
Why a crate and not a module per binary
The connect is identical for every participant — mint an authelia token,
present it at CONNECT for the auth_callout responder to introspect, let
async-nats re-run the callback on each connection attempt. Only the use
differs.
Two copies of that would be two copies of credential handling, and a
token-refresh fix would have to be found twice. The same reasoning already put
hive-sock-client in its own crate rather than in each daemon that speaks to a
unix socket.
The two properties that constrain the code
A token expires. Authelia issues client_credentials access tokens with
expires_in: 3599. Authentication happens at CONNECT, so a long-lived
connection is fine — but a reconnect an hour later needs a token minted an
hour later.
The refresh therefore lives in the auth callback, not in a timer.
async-nats invokes it per connection attempt, so there is no window in which
the client holds a token it minted for a previous connection. The alternative —
mint once, own the reconnect loop — fails in the way this subsystem exists to
prevent: the process keeps serving while its data quietly stops moving, and
nothing says so until someone reads a dashboard.
Configuration
QueueConfig::from_env(prefix) reads <prefix>_NATS_URL,
<prefix>_OIDC_TOKEN_ENDPOINT, <prefix>_OIDC_CLIENT_ID and
<prefix>_OIDC_CLIENT_SECRET_FILE.
The prefix is a parameter because the variables belong to the consuming unit — a NixOS module sets them alongside its other options. What is shared is the rule, not the spelling: all four together or none at all. A half-set environment is a hard error, because the failure it would otherwise produce is the expensive kind — the process comes up "fine", never connects, and the data it was supposed to move silently stops.
The client secret is a path, not a value: putting it in the environment
would publish it to anything that can read /proc/<pid>/environ. It is read
per token request rather than cached, so a rotation the operator believes took
effect actually did.
What this crate does not do
It ends at a connected client. jetstream/kv are off by default — what a
consumer does with the connection is its own business, and its Cargo.toml is
where that requirement should be visible. The auth-callout responder speaks the
connect and nothing else, and pays for nothing else.
The one exception: the kv feature
kv adds status, which holds the name and the creation config of the
hive-status bucket — nothing more.
It is here because that bucket has two ends in two crates: a hive writes its
own key, the controller reads every key. The name being a repeated literal is
the mild half of the problem; the sharp half is that either end may arrive first
on a fresh swarm, so both create the bucket if it is missing. Two Configs that
drift means whichever end created it wins and the other opens a bucket it did
not ask for — no error, no log, just a retention policy nobody chose.
An agreement between two crates has to live in one of them, and neither end of this bucket is senior to the other. Behind a default-off feature, the consumer that needs none of it still pays nothing.