An agent's store identity was signed in swarm-controller's memory by a CA a controller-host unit generated on disk, and the listener never trusted that CA. Agent leaves now come from the store itself: a `pki-agents` PKI mount whose root openbao generates internally, so the agent CA's key never exists outside the store. - swarm-bao-agent-pki (new, store host, as the bao granter): enables and tunes the mount, generates the root once (guarded on an empty issuer list, no replace branch), upserts the `swarm-agent` role (client certificates named `hive-agent-*` only, 90 days), caches the CA at /var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle. - The listener's tls_client_ca_file is a new listener-client-ca.pem (client-ca.pem, then the agent CA). Host cert-auth roles still pin client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs composes the same bundle before openbao starts. - openbao reads tls_client_ca_file only at start, so when the bundle changed after openbao started, swarm-bao-agent-pki restarts openbao.service in the container; under `seal = "shamir"` it prints the step instead. Once swarm-bao-certs has a cached CA, later boots start openbao with it and do not restart. - The controller policy gains exactly `update` on pki-agents/issue/swarm-agent. mint_and_verify now asks that role for the leaf (the store generates the key), writes the agent's cert-auth role pinning the issuing CA bao returned, and writes the agent's policy as render_agent alone: the hive-shared queue credential stanza is gone. - deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the other pki role names); swarm-controller gets SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options. Deleted: swarm-controller-agent-ca and its options (agentCaFile, agentCaKeyFile), env, LoadCredential entries and assertion; agent_identity's Authority, rcgen signing and validity window; the rcgen and time dependencies of swarm-controller (rcgen leaves the workspace); policy::render_agent_with_queue and its tests. The CN-prefix assertion policy.rs said was owed is not: agent and host roles pin different CAs. Migration is re-creating each agent after deploy; that overwrites the stale role and policy. Closes #4756
122 lines
6 KiB
TOML
122 lines
6 KiB
TOML
[package]
|
|
name = "swarm-controller"
|
|
version.workspace = true
|
|
readme = "README.md"
|
|
edition.workspace = true
|
|
|
|
[[bin]]
|
|
name = "swarm-controller"
|
|
path = "src/main.rs"
|
|
|
|
[dependencies]
|
|
anyhow.workspace = true
|
|
# `kv` (which pulls `jetstream`) on top of the workspace's feature set: the
|
|
# queue is this daemon's *store*, not just its transport - a hive's last
|
|
# status snapshot is read out of a JetStream KV bucket. Declared here rather
|
|
# than in the workspace entry so the auth-callout responder, which speaks
|
|
# neither, does not claim to need them.
|
|
async-nats = { workspace = true, features = ["kv"] }
|
|
axum.workspace = true
|
|
# swarm-controller's own forge client (`forge.rs`) — self-contained,
|
|
# deliberately not sharing code with `hive-c0re::forge` across the crate
|
|
# boundary (see #3306's design discussion: forcing that split now, over a
|
|
# few idempotent CRUD-ish calls, is premature plumbing).
|
|
forgejo-api.workspace = true
|
|
# Only for base64-encoding file content for `forge.rs`'s
|
|
# `repo_change_files` calls — forgejo's content API takes base64, never
|
|
# raw bytes.
|
|
base64.workspace = true
|
|
# `agent_identity.rs` mints the per-agent queue secret, the one value in this
|
|
# tree this daemon invents rather than receives. Straight from the kernel's
|
|
# CSPRNG — see the workspace entry for why this and not `rand`.
|
|
getrandom.workspace = true
|
|
futures-util.workspace = true
|
|
# RFC 9457 `application/problem+json` error bodies. Same version + `axum`
|
|
# feature as hive-c0re: the two daemons answer the same operator UIs, so a
|
|
# reader that handles one's failures has to handle the other's.
|
|
problem_details = { version = "0.9.0", features = ["axum"] }
|
|
# The graph itself, held directly rather than behind a c0re-style wrapper
|
|
# module — that layering (`hive-c0re::job_queue`) is partially legacy (predates
|
|
# `hive-jobq`'s extraction into its own crate) and this daemon does not need it
|
|
# repeated. Driven by `hive_jobq::scheduler::Scheduler` (`spawn_jobq_worker`),
|
|
# same shape `hive-c0re/src/job_queue/scheduler.rs` uses over its own graph.
|
|
hive-jobq.workspace = true
|
|
hive-jobq-wire.workspace = true
|
|
hive-log.workspace = true
|
|
# The jobq-rollup OTEL exporter, wired up in `main` via
|
|
# `hive_jobq_metrics::spawn_exporter` — moved to its own crate (rather than
|
|
# living here as `jobq_metrics.rs`) specifically so a future second caller
|
|
# (e.g. hive-c0re, for its own per-hive job graph) doesn't have to depend on
|
|
# this whole binary to reuse it.
|
|
hive-jobq-metrics.workspace = true
|
|
# Direct OTEL SDK use in `vcs_metrics.rs` — sync counters recorded off
|
|
# webhook deliveries, a different shape from `hive-jobq-metrics`'s
|
|
# observable-gauge rollup, so it isn't a fit for that crate's API and lives
|
|
# here instead. Same three crates that pairing already pulls in transitively,
|
|
# named directly since this module builds its own `SdkMeterProvider`.
|
|
opentelemetry.workspace = true
|
|
opentelemetry_sdk.workspace = true
|
|
opentelemetry-otlp.workspace = true
|
|
# The forge webhook HMAC (`webhook.rs`). Kept in this crate rather than
|
|
# shared with hive-c0re's equivalent: c0re's copy is scheduled to be deleted
|
|
# with its webhook routes once registration moves here, so the second holder
|
|
# is departing, not arriving — see that module's docs.
|
|
hmac.workspace = true
|
|
sha2.workspace = true
|
|
# Validates `POST /api/agents`' `name` before it becomes `agent`/`repo`
|
|
# everywhere downstream — see `create_agent`'s doc comment for why this is
|
|
# defense-in-depth, not the only gate (per an argus review finding).
|
|
hive-types.workspace = true
|
|
# `auth`'s bridge client — same crate the bridge itself uses to define the
|
|
# request/response shape, so the two ends cannot drift. `forge.rs` also
|
|
# uses this directly for `StatusCode` in its error-classification helpers.
|
|
#
|
|
# `blocking` on top of the workspace default: `vcs_metrics`'s OTLP exporter
|
|
# runs on a thread with no tokio reactor (see that module's doc), so the
|
|
# `reqwest::blocking::Client` it hands to `AuthenticatedHttpClient` has to
|
|
# come from the blocking half of this crate, not the async one every other
|
|
# consumer here uses.
|
|
reqwest = { workspace = true, features = ["blocking"] }
|
|
serde.workspace = true
|
|
serde_json.workspace = true
|
|
strum.workspace = true
|
|
swarm-authelia-bridge-sock.workspace = true
|
|
# The queue connect (token mint + auth callback + reconnect) is shared with
|
|
# every other participant - a hive publishing its own status runs the same
|
|
# code with a different client id. Two copies of credential handling is one
|
|
# token-refresh fix that has to be found twice.
|
|
#
|
|
# `kv` for the same reason one level in: the status bucket's name and
|
|
# creation config are shared with the hive that writes it, so this end does
|
|
# not get to declare them privately.
|
|
#
|
|
swarm-queue-client = { workspace = true, features = ["kv"] }
|
|
# `matrix_account.rs` writes the credential this daemon's route accepts. Same
|
|
# crate the hive reads it back with, which is the point: the path, the field
|
|
# name and the object's shape are agreements between the two ends, and a
|
|
# second spelling here would be a store this hive could not read.
|
|
swarm-secret-client.workspace = true
|
|
# The appservice calls `matrix_account::agent_token` mints agents' accounts
|
|
# with — shared with `swarm-matrix-ctl`, which pins the same device id.
|
|
swarm-matrix-client.workspace = true
|
|
# `otel_http_client.rs`'s `AuthenticatedHttpClient` — an
|
|
# `opentelemetry_http::HttpClient` impl authenticated with this crate's own
|
|
# `swarm-queue-client` identity. Lives in this crate rather than
|
|
# `swarm-queue-client` (mara's own call during review) since this daemon is
|
|
# its only caller; see that module's doc for the full rationale. `opentelemetry-http`
|
|
# is also needed directly (not just transitively) so `main.rs` can name
|
|
# `Box<dyn opentelemetry_http::HttpClient>` when handing
|
|
# `vcs_metrics::authenticated_http_client()`'s output to
|
|
# `hive_jobq_metrics::spawn_exporter`.
|
|
opentelemetry-http.workspace = true
|
|
async-trait.workspace = true
|
|
bytes.workspace = true
|
|
http.workspace = true
|
|
tokio.workspace = true
|
|
tracing.workspace = true
|
|
url.workspace = true
|
|
utoipa.workspace = true
|
|
utoipa-axum.workspace = true
|
|
|
|
[lints]
|
|
workspace = true
|