The swarm UI had nowhere to POST an external matrix account to: this daemon had no matrix-account code at all and no `swarm-secret-client` dependency, so the last leg of #3726 — a credential reaching an agent — had no entry point. `PUT /api/hives/{hive}/agents/{agent}/matrix-accounts/{account}` writes the credential to the store under the agent's own path and publishes a `CredentialNotice` on that hive's credential subject. All three path names are load-bearing: agent + account locate the secret, hive routes the notice. The account is a path segment rather than a body field so that splitting the 1:1 account-to-agent mapping later is a new route, not a changed payload. Store first, notify second, and the order cannot be swapped: a notice that overtakes its own write reaches a hive that reads nothing, and the hive deliberately does not retry. The publish is followed by a flush for the reason `publish_deploy` flushes — `publish` hands the message to the connection's write buffer and returns, so the response could otherwise outrun the notice it reports as sent. The store client is built per request rather than held in `AppState`, matching what the hive side does inside `deliver`: a login that expires is not worth caching for a route this cold. `swarm_hive` is `declaration_target`'s two name checks, extracted so this handler makes them identically rather than in a second copy free to drift. `declaration_target` still tests the writer first, so a deployment with no queue answers 503 whatever the caller spelled. ## The nix half #4081 minted the controller's leaf and gave it `baoClientCertFile` / `baoClientKeyFile`, deliberately stopping there — the leaf is minted whether or not a controller runs on that host. Nothing consumed those options, so the identity never reached the process. Measured before writing: `git grep baoClientCertFile` returned 5 sites and zero consumers, against a control (`tokenEndpoint`, 4 hits in the same file) proving the search can see consumption where it exists. The unit now gets `BAO_ADDR` / `BAO_CLIENT_CERT` / `BAO_CLIENT_KEY` / `BAO_CACERT` and the matching `LoadCredential` entries, following `hive-c0re/environment.nix`'s `%d` credential shape. The gate is `deploy.swarm-controller.baoClientCertFile`, NOT `deploy.bao.clientCertFile`. The latter is the hive reader's identity and its policy scopes a hive's own secrets; wiring it here would evaluate, deploy, and fail only when the daemon tried to write an agent's credential. Two `module-eval` arms cover exactly that. The presence arm asserts the `LoadCredential` *source path* (`…:/var/lib/swarm-bao-pki/controller.pem`) and not just the `%d` name, because a `%d`-only assertion passes while the daemon holds the wrong policy. The absence arm (`controllerNoStore`) is what makes the presence arm mean anything. `RestrictAddressFamilies` already covers the store client; its own comment asks for the family to be added with the client, and AF_INET/AF_INET6 are present. Contributes to #3726
114 lines
5.5 KiB
TOML
114 lines
5.5 KiB
TOML
[package]
|
|
name = "swarm-controller"
|
|
version.workspace = true
|
|
readme = "README.md"
|
|
edition.workspace = true
|
|
|
|
[[bin]]
|
|
name = "swarm-controller"
|
|
path = "src/main.rs"
|
|
|
|
[dependencies]
|
|
anyhow.workspace = true
|
|
# `kv` (which pulls `jetstream`) on top of the workspace's feature set: the
|
|
# queue is this daemon's *store*, not just its transport - a hive's last
|
|
# status snapshot is read out of a JetStream KV bucket. Declared here rather
|
|
# than in the workspace entry so the auth-callout responder, which speaks
|
|
# neither, does not claim to need them.
|
|
async-nats = { workspace = true, features = ["kv"] }
|
|
axum.workspace = true
|
|
# swarm-controller's own forge client (`forge.rs`) — self-contained,
|
|
# deliberately not sharing code with `hive-c0re::forge` across the crate
|
|
# boundary (see #3306's design discussion: forcing that split now, over a
|
|
# few idempotent CRUD-ish calls, is premature plumbing).
|
|
forgejo-api.workspace = true
|
|
# Only for base64-encoding file content for `forge.rs`'s
|
|
# `repo_change_files` calls — forgejo's content API takes base64, never
|
|
# raw bytes.
|
|
base64.workspace = true
|
|
futures-util.workspace = true
|
|
# RFC 9457 `application/problem+json` error bodies. Same version + `axum`
|
|
# feature as hive-c0re: the two daemons answer the same operator UIs, so a
|
|
# reader that handles one's failures has to handle the other's.
|
|
problem_details = { version = "0.9.0", features = ["axum"] }
|
|
# The graph itself, held directly rather than behind a c0re-style wrapper
|
|
# module — that layering (`hive-c0re::job_queue`) is partially legacy (predates
|
|
# `hive-jobq`'s extraction into its own crate) and this daemon does not need it
|
|
# repeated. Driven by `hive_jobq::scheduler::Scheduler` (`spawn_jobq_worker`),
|
|
# same shape `hive-c0re/src/job_queue/scheduler.rs` uses over its own graph.
|
|
hive-jobq.workspace = true
|
|
hive-jobq-wire.workspace = true
|
|
# The jobq-rollup OTEL exporter, wired up in `main` via
|
|
# `hive_jobq_metrics::spawn_exporter` — moved to its own crate (rather than
|
|
# living here as `jobq_metrics.rs`) specifically so a future second caller
|
|
# (e.g. hive-c0re, for its own per-hive job graph) doesn't have to depend on
|
|
# this whole binary to reuse it.
|
|
hive-jobq-metrics.workspace = true
|
|
# Direct OTEL SDK use in `vcs_metrics.rs` — sync counters recorded off
|
|
# webhook deliveries, a different shape from `hive-jobq-metrics`'s
|
|
# observable-gauge rollup, so it isn't a fit for that crate's API and lives
|
|
# here instead. Same three crates that pairing already pulls in transitively,
|
|
# named directly since this module builds its own `SdkMeterProvider`.
|
|
opentelemetry.workspace = true
|
|
opentelemetry_sdk.workspace = true
|
|
opentelemetry-otlp.workspace = true
|
|
# The forge webhook HMAC (`webhook.rs`). Kept in this crate rather than
|
|
# shared with hive-c0re's equivalent: c0re's copy is scheduled to be deleted
|
|
# with its webhook routes once registration moves here, so the second holder
|
|
# is departing, not arriving — see that module's docs.
|
|
hmac.workspace = true
|
|
sha2.workspace = true
|
|
# Validates `POST /api/agents`' `name` before it becomes `agent`/`repo`
|
|
# everywhere downstream — see `create_agent`'s doc comment for why this is
|
|
# defense-in-depth, not the only gate (per an argus review finding).
|
|
hive-types.workspace = true
|
|
# `auth`'s bridge client — same crate the bridge itself uses to define the
|
|
# request/response shape, so the two ends cannot drift. `forge.rs` also
|
|
# uses this directly for `StatusCode` in its error-classification helpers.
|
|
#
|
|
# `blocking` on top of the workspace default: `vcs_metrics`'s OTLP exporter
|
|
# runs on a thread with no tokio reactor (see that module's doc), so the
|
|
# `reqwest::blocking::Client` it hands to `AuthenticatedHttpClient` has to
|
|
# come from the blocking half of this crate, not the async one every other
|
|
# consumer here uses.
|
|
reqwest = { workspace = true, features = ["blocking"] }
|
|
serde.workspace = true
|
|
serde_json.workspace = true
|
|
swarm-authelia-bridge-sock.workspace = true
|
|
# The queue connect (token mint + auth callback + reconnect) is shared with
|
|
# every other participant - a hive publishing its own status runs the same
|
|
# code with a different client id. Two copies of credential handling is one
|
|
# token-refresh fix that has to be found twice.
|
|
#
|
|
# `kv` for the same reason one level in: the status bucket's name and
|
|
# creation config are shared with the hive that writes it, so this end does
|
|
# not get to declare them privately.
|
|
#
|
|
swarm-queue-client = { workspace = true, features = ["kv"] }
|
|
# `matrix_account.rs` writes the credential this daemon's route accepts. Same
|
|
# crate the hive reads it back with, which is the point: the path, the field
|
|
# name and the object's shape are agreements between the two ends, and a
|
|
# second spelling here would be a store this hive could not read.
|
|
swarm-secret-client.workspace = true
|
|
# `otel_http_client.rs`'s `AuthenticatedHttpClient` — an
|
|
# `opentelemetry_http::HttpClient` impl authenticated with this crate's own
|
|
# `swarm-queue-client` identity. Lives in this crate rather than
|
|
# `swarm-queue-client` (mara's own call during review) since this daemon is
|
|
# its only caller; see that module's doc for the full rationale. `opentelemetry-http`
|
|
# is also needed directly (not just transitively) so `main.rs` can name
|
|
# `Box<dyn opentelemetry_http::HttpClient>` when handing
|
|
# `vcs_metrics::authenticated_http_client()`'s output to
|
|
# `hive_jobq_metrics::spawn_exporter`.
|
|
opentelemetry-http.workspace = true
|
|
async-trait.workspace = true
|
|
bytes.workspace = true
|
|
http.workspace = true
|
|
tokio.workspace = true
|
|
tracing.workspace = true
|
|
tracing-subscriber.workspace = true
|
|
url.workspace = true
|
|
utoipa.workspace = true
|
|
utoipa-axum.workspace = true
|
|
|
|
[lints]
|
|
workspace = true
|