`swarm/agents/<agent>/bao-mtls` did not exist, and neither did any per-agent identity at the secret store: `policy::agent_object_name`, `render_agent` and `render_agent_with_queue` had been written and never called outside their own tests. An agent's only "per-agent" secret today is read under the HIVE's certificate, through a wide grant on `swarm/agents/*` — so "per-agent" was presentational. The swarm now mints the certificate, so no hive ever needs the capability to mint one. `swarm-controller` is the service that does it: it already logs in to the store, and its existing grant already covers exactly the three objects written here (`create/update` on `secret/data/swarm/agents/*`, `sys/policies/acl/hive-*` and `auth/cert/certs/hive-*`). No new bao grant, and nothing co-located — a cert-auth role pins its authority by value, per role, so the controller issues from its own CA on its own host and pins that CA in the role it writes. No existing role changes. The mint node does not report success on a write. After publishing it connects again, with the leaf it just issued and under the role it just wrote, and reads the path back — so the policy, the role, the common name and the leaf are exercised in production on every agent creation. A certificate this code mints that the role this code writes will not accept turns the job node red at creation time instead of surfacing later as an agent container that cannot start. `TriggerDeploy` gains an `after_any` edge on the mint, not `after_ok`: a hive cannot pass down a certificate the swarm has not published, but a host with no authority configured must still create agents exactly as it does today. The private key is generated in memory and never written to disk on the controller — `SecretStore::connect_with_identity` takes the PEM the minter is already holding, so nothing is written out purely to be logged in with. Refs #4137
124 lines
6.1 KiB
TOML
124 lines
6.1 KiB
TOML
[package]
|
|
name = "swarm-controller"
|
|
version.workspace = true
|
|
readme = "README.md"
|
|
edition.workspace = true
|
|
|
|
[[bin]]
|
|
name = "swarm-controller"
|
|
path = "src/main.rs"
|
|
|
|
[dependencies]
|
|
anyhow.workspace = true
|
|
# `kv` (which pulls `jetstream`) on top of the workspace's feature set: the
|
|
# queue is this daemon's *store*, not just its transport - a hive's last
|
|
# status snapshot is read out of a JetStream KV bucket. Declared here rather
|
|
# than in the workspace entry so the auth-callout responder, which speaks
|
|
# neither, does not claim to need them.
|
|
async-nats = { workspace = true, features = ["kv"] }
|
|
axum.workspace = true
|
|
# swarm-controller's own forge client (`forge.rs`) — self-contained,
|
|
# deliberately not sharing code with `hive-c0re::forge` across the crate
|
|
# boundary (see #3306's design discussion: forcing that split now, over a
|
|
# few idempotent CRUD-ish calls, is premature plumbing).
|
|
forgejo-api.workspace = true
|
|
# Only for base64-encoding file content for `forge.rs`'s
|
|
# `repo_change_files` calls — forgejo's content API takes base64, never
|
|
# raw bytes.
|
|
base64.workspace = true
|
|
futures-util.workspace = true
|
|
# RFC 9457 `application/problem+json` error bodies. Same version + `axum`
|
|
# feature as hive-c0re: the two daemons answer the same operator UIs, so a
|
|
# reader that handles one's failures has to handle the other's.
|
|
problem_details = { version = "0.9.0", features = ["axum"] }
|
|
# The graph itself, held directly rather than behind a c0re-style wrapper
|
|
# module — that layering (`hive-c0re::job_queue`) is partially legacy (predates
|
|
# `hive-jobq`'s extraction into its own crate) and this daemon does not need it
|
|
# repeated. Driven by `hive_jobq::scheduler::Scheduler` (`spawn_jobq_worker`),
|
|
# same shape `hive-c0re/src/job_queue/scheduler.rs` uses over its own graph.
|
|
hive-jobq.workspace = true
|
|
hive-jobq-wire.workspace = true
|
|
# The jobq-rollup OTEL exporter, wired up in `main` via
|
|
# `hive_jobq_metrics::spawn_exporter` — moved to its own crate (rather than
|
|
# living here as `jobq_metrics.rs`) specifically so a future second caller
|
|
# (e.g. hive-c0re, for its own per-hive job graph) doesn't have to depend on
|
|
# this whole binary to reuse it.
|
|
hive-jobq-metrics.workspace = true
|
|
# Direct OTEL SDK use in `vcs_metrics.rs` — sync counters recorded off
|
|
# webhook deliveries, a different shape from `hive-jobq-metrics`'s
|
|
# observable-gauge rollup, so it isn't a fit for that crate's API and lives
|
|
# here instead. Same three crates that pairing already pulls in transitively,
|
|
# named directly since this module builds its own `SdkMeterProvider`.
|
|
opentelemetry.workspace = true
|
|
opentelemetry_sdk.workspace = true
|
|
opentelemetry-otlp.workspace = true
|
|
# The forge webhook HMAC (`webhook.rs`). Kept in this crate rather than
|
|
# shared with hive-c0re's equivalent: c0re's copy is scheduled to be deleted
|
|
# with its webhook routes once registration moves here, so the second holder
|
|
# is departing, not arriving — see that module's docs.
|
|
hmac.workspace = true
|
|
sha2.workspace = true
|
|
# Validates `POST /api/agents`' `name` before it becomes `agent`/`repo`
|
|
# everywhere downstream — see `create_agent`'s doc comment for why this is
|
|
# defense-in-depth, not the only gate (per an argus review finding).
|
|
hive-types.workspace = true
|
|
# `auth`'s bridge client — same crate the bridge itself uses to define the
|
|
# request/response shape, so the two ends cannot drift. `forge.rs` also
|
|
# uses this directly for `StatusCode` in its error-classification helpers.
|
|
#
|
|
# `blocking` on top of the workspace default: `vcs_metrics`'s OTLP exporter
|
|
# runs on a thread with no tokio reactor (see that module's doc), so the
|
|
# `reqwest::blocking::Client` it hands to `AuthenticatedHttpClient` has to
|
|
# come from the blocking half of this crate, not the async one every other
|
|
# consumer here uses.
|
|
reqwest = { workspace = true, features = ["blocking"] }
|
|
serde.workspace = true
|
|
serde_json.workspace = true
|
|
strum.workspace = true
|
|
swarm-authelia-bridge-sock.workspace = true
|
|
# The queue connect (token mint + auth callback + reconnect) is shared with
|
|
# every other participant - a hive publishing its own status runs the same
|
|
# code with a different client id. Two copies of credential handling is one
|
|
# token-refresh fix that has to be found twice.
|
|
#
|
|
# `kv` for the same reason one level in: the status bucket's name and
|
|
# creation config are shared with the hive that writes it, so this end does
|
|
# not get to declare them privately.
|
|
#
|
|
swarm-queue-client = { workspace = true, features = ["kv"] }
|
|
# `matrix_account.rs` writes the credential this daemon's route accepts. Same
|
|
# crate the hive reads it back with, which is the point: the path, the field
|
|
# name and the object's shape are agreements between the two ends, and a
|
|
# second spelling here would be a store this hive could not read.
|
|
swarm-secret-client.workspace = true
|
|
# `agent_identity.rs` signs the per-agent client leaf the store authenticates
|
|
# an agent container by. The one runtime signer in this tree — every other CA
|
|
# here is a deploy-time `openssl` oneshot — because this one issues per agent
|
|
# creation and must keep the key it generates in memory, never on disk.
|
|
rcgen.workspace = true
|
|
# `rcgen::CertificateParams`'s validity window is a `time::OffsetDateTime`,
|
|
# which `agent_identity::validity` builds. Same workspace pin every other
|
|
# holder of a timestamp here uses.
|
|
time.workspace = true
|
|
# `otel_http_client.rs`'s `AuthenticatedHttpClient` — an
|
|
# `opentelemetry_http::HttpClient` impl authenticated with this crate's own
|
|
# `swarm-queue-client` identity. Lives in this crate rather than
|
|
# `swarm-queue-client` (mara's own call during review) since this daemon is
|
|
# its only caller; see that module's doc for the full rationale. `opentelemetry-http`
|
|
# is also needed directly (not just transitively) so `main.rs` can name
|
|
# `Box<dyn opentelemetry_http::HttpClient>` when handing
|
|
# `vcs_metrics::authenticated_http_client()`'s output to
|
|
# `hive_jobq_metrics::spawn_exporter`.
|
|
opentelemetry-http.workspace = true
|
|
async-trait.workspace = true
|
|
bytes.workspace = true
|
|
http.workspace = true
|
|
tokio.workspace = true
|
|
tracing.workspace = true
|
|
tracing-subscriber.workspace = true
|
|
url.workspace = true
|
|
utoipa.workspace = true
|
|
utoipa-axum.workspace = true
|
|
|
|
[lints]
|
|
workspace = true
|