| Filename | Latest commit message | Latest commit date |
|---|---|---|
The deploy event is a nudge with no second path: core NATS is at-most-once, so a hive that was down when the controller published simply never learns that an agent is meant to exist here. This adds the repair path — one boot-time DAG node that reads this hive's own key in the `hive-wanted` bucket and converges the agents it names. Two semantics settled on the issue thread, and both are places where a plausible implementation is the wrong one: - **Absence is not a deletion order.** No bucket, no key, or an agent the value does not name all mean the controller has said nothing. Swarm-side lifecycle does not yet cover agents that predate it, so "converge to exactly this set" would tear down every agent the swarm has not adopted. `plan` only ever inspects the agents a declaration names. - **An unrecognised state is inert.** `AgentState` is an open enum: a value this build cannot read deserialises into `Unrecognised` and is left alone. A closed enum would force "not `Up`" onto a state like `paused`, so a controller that learned a new value would take agents down on every hive not yet updated. Divergence is measured against the hive's **stored power intent**, not the container's observed running state — an agent that is down while its intent says `Up` is already the boot reconcile's work, and a loop reading `is_running` would insert a start DAG behind that reconcile's back on every boot. A hive that already agrees with its declaration queues nothing at all. `queue_first_deploy` is extracted from the deploy-event path rather than open-coded here, for the power-intent seed: without it `first_deploy`'s tail `Reconcile` seeds `Wanted` from a container that exists but has not started yet, which locks the agent to `Offline` on its first reconcile. The read is authorised as-is: `store.get` takes async-nats' direct-get arm (the KV bucket is created with `allow_direct`), which is exactly the `$JS.API.DIRECT.GET.KV_hive-wanted.$KV.hive-wanted.<hive>` subject `swarm-nats-auth` grants a hive. The fallback subject is not granted, and a refused NATS request surfaces as a timeout rather than an error. Nothing writes the bucket yet — the controller-side writer is the other half of #3124, so this does not close it. |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
swarm-queue-client
Connecting to the swarm message queue as an authenticated client. Shared by every process that participates: the swarm controller reads hive status out of the queue, a hive publishes its own status into it.
Why a crate and not a module per binary
The connect is identical for every participant — mint an authelia token,
present it at CONNECT for the auth_callout responder to introspect, let
async-nats re-run the callback on each connection attempt. Only the use
differs.
Two copies of that would be two copies of credential handling, and a
token-refresh fix would have to be found twice. The same reasoning already put
hive-sock-client in its own crate rather than in each daemon that speaks to a
unix socket.
The two properties that constrain the code
A token expires. Authelia issues client_credentials access tokens with
expires_in: 3599. Authentication happens at CONNECT, so a long-lived
connection is fine — but a reconnect an hour later needs a token minted an
hour later.
The refresh therefore lives in the auth callback, not in a timer.
async-nats invokes it per connection attempt, so there is no window in which
the client holds a token it minted for a previous connection. The alternative —
mint once, own the reconnect loop — fails in the way this subsystem exists to
prevent: the process keeps serving while its data quietly stops moving, and
nothing says so until someone reads a dashboard.
Configuration
QueueConfig::from_env(prefix) reads <prefix>_NATS_URL,
<prefix>_OIDC_TOKEN_ENDPOINT, <prefix>_OIDC_CLIENT_ID and
<prefix>_OIDC_CLIENT_SECRET_FILE.
The prefix is a parameter because the variables belong to the consuming unit — a NixOS module sets them alongside its other options. What is shared is the rule, not the spelling: all four together or none at all. A half-set environment is a hard error, because the failure it would otherwise produce is the expensive kind — the process comes up "fine", never connects, and the data it was supposed to move silently stops.
The client secret is a path, not a value: putting it in the environment
would publish it to anything that can read /proc/<pid>/environ. It is read
per token request rather than cached, so a rotation the operator believes took
effect actually did.
What this crate does not do
It ends at a connected client. jetstream/kv are off by default — what a
consumer does with the connection is its own business, and its Cargo.toml is
where that requirement should be visible. The auth-callout responder speaks the
connect and nothing else, and pays for nothing else.
The one exception: the kv feature
kv adds status, which holds the name and the creation config of the
hive-status bucket — nothing more.
It is here because that bucket has two ends in two crates: a hive writes its
own key, the controller reads every key. The name being a repeated literal is
the mild half of the problem; the sharp half is that either end may arrive first
on a fresh swarm, so both create the bucket if it is missing. Two Configs that
drift means whichever end created it wins and the other opens a bucket it did
not ask for — no error, no log, just a retention policy nobody chose.
An agreement between two crates has to live in one of them, and neither end of this bucket is senior to the other. Behind a default-off feature, the consumer that needs none of it still pays nothing.