| Filename | Latest commit message | Latest commit date |
|---|---|---|
Review call: the event was addressed per hive — `$SWARM.events.<hive>.knowledge`, published in a loop over the roster, granted through a wildcard. It does not need to be. The payload is empty and the event means the same thing to every hive, so one publish to one subject delivers exactly what N publishes to N subjects did, and core NATS already fans out to whoever is subscribed. A hive that was down misses it either way and reconciles on its next periodic pull. That deletes rather than reshuffles: the roster loop, the wildcard, and the shared subject-building function whose entire purpose was keeping the grant and the publish from drifting apart. With one literal there is nothing to disagree about. The per-hive shape was justified by the callout policy's rule that an extra subject must contain the hive name. That rule governs `extra_hive_subjects` — what a HIVE may publish. This subject lives in the controller's reader grant, which the rule does not constrain, so a real rule was carried across into a decision it had no authority over. Knowledge becomes its own category rather than a leaf under a general event namespace, since a namespace shaped for events that do not exist yet is a decision made before there is anything to decide from. The empty config-PR match arm goes with it: an arm with no body claims this is where the deploy path is handled, and it is not. The deny test stays and matters more, not less: with one shared subject a forged event would reach the whole swarm where a per-hive one reached a single hive. |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-queue-client
Connecting to the swarm message queue as an authenticated client. Shared by every process that participates: the swarm controller reads hive status out of the queue, a hive publishes its own status into it.
Why a crate and not a module per binary
The connect is identical for every participant — mint an authelia token,
present it at CONNECT for the auth_callout responder to introspect, let
async-nats re-run the callback on each connection attempt. Only the use
differs.
Two copies of that would be two copies of credential handling, and a
token-refresh fix would have to be found twice. The same reasoning already put
hive-sock-client in its own crate rather than in each daemon that speaks to a
unix socket.
The two properties that constrain the code
A token expires. Authelia issues client_credentials access tokens with
expires_in: 3599. Authentication happens at CONNECT, so a long-lived
connection is fine — but a reconnect an hour later needs a token minted an
hour later.
The refresh therefore lives in the auth callback, not in a timer.
async-nats invokes it per connection attempt, so there is no window in which
the client holds a token it minted for a previous connection. The alternative —
mint once, own the reconnect loop — fails in the way this subsystem exists to
prevent: the process keeps serving while its data quietly stops moving, and
nothing says so until someone reads a dashboard.
Configuration
QueueConfig::from_env(prefix) reads <prefix>_NATS_URL,
<prefix>_OIDC_TOKEN_ENDPOINT, <prefix>_OIDC_CLIENT_ID and
<prefix>_OIDC_CLIENT_SECRET_FILE.
The prefix is a parameter because the variables belong to the consuming unit — a NixOS module sets them alongside its other options. What is shared is the rule, not the spelling: all four together or none at all. A half-set environment is a hard error, because the failure it would otherwise produce is the expensive kind — the process comes up "fine", never connects, and the data it was supposed to move silently stops.
The client secret is a path, not a value: putting it in the environment
would publish it to anything that can read /proc/<pid>/environ. It is read
per token request rather than cached, so a rotation the operator believes took
effect actually did.
What this crate does not do
It ends at a connected client. jetstream/kv are off by default — what a
consumer does with the connection is its own business, and its Cargo.toml is
where that requirement should be visible. The auth-callout responder speaks the
connect and nothing else, and pays for nothing else.
The one exception: the kv feature
kv adds status, which holds the name and the creation config of the
hive-status bucket — nothing more.
It is here because that bucket has two ends in two crates: a hive writes its
own key, the controller reads every key. The name being a repeated literal is
the mild half of the problem; the sharp half is that either end may arrive first
on a fresh swarm, so both create the bucket if it is missing. Two Configs that
drift means whichever end created it wins and the other opens a bucket it did
not ask for — no error, no log, just a retention policy nobody chose.
An agreement between two crates has to live in one of them, and neither end of this bucket is senior to the other. Behind a default-off feature, the consumer that needs none of it still pays nothing.