swarm: revoke an agent's queue credential when it is declared destroyed
A per-agent queue credential is minted at agent creation and nothing has ever removed it. An agent declared destroyed loses its container and keeps its credential: a bearer secret recovered from a snapshot or a stale capture still authenticates as that agent, so the set of usable credentials only grows. Delete the path the mint published, on the one transition that ends an agent's life. It mirrors step 3 of `mint_and_verify` and no other step: the leaf, the ACL document and the cert role are what a hive uses to collect an agent's secrets and are re-minted on every run of the mint. Every version, not the newest. The mint rewrites the path when the principal it names needs correcting, so KV v2's plain delete would leave the identical secret readable at ?version=N. That is a separately-ACL'd path, hence the second stanza in the controller's grant -- `delete` on metadata discloses nothing, and `update` on the data path already lets this principal destroy any agent credential's usability. The destroy is not blocked by a failed revocation: the declaration is already published and refusing the call would leave an operator with an agent they cannot tear down. The failure is logged at error instead, naming the agent, since a silent orphan is the fault being removed.
This commit is contained in:
parent
88c386c96c
commit
8caf688ee4
5 changed files with 229 additions and 12 deletions
|
|
@ -220,6 +220,54 @@ pub async fn mint_and_verify(agent: &str) -> Result<()> {
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Revoke `agent`'s queue credential: delete the path
|
||||
/// [`mint_and_verify`]'s step 3 published, and everything ever written at it.
|
||||
///
|
||||
/// The undo of that one step and of no other. The leaf, the ACL document and
|
||||
/// the cert-auth role that make up the rest of an agent's identity stay where
|
||||
/// they are — they are what a *hive* uses to collect an agent's secrets, they
|
||||
/// are minted afresh on every run of the mint, and tearing them down is not
|
||||
/// what the queue credential outliving its holder is about.
|
||||
///
|
||||
/// **Deletes every version, not the newest one.** The mint rewrites this path
|
||||
/// whenever the principal it names has to be corrected, so a soft delete would
|
||||
/// leave the identical secret sitting in version history, readable at
|
||||
/// `?version=N` by anything that can read the path at all — a value still
|
||||
/// recoverable has not been revoked. See
|
||||
/// [`SecretStore::delete_all_versions`][swarm_secret_client::SecretStore::delete_all_versions].
|
||||
///
|
||||
/// **Idempotent**: revoking an agent that never had a credential, or one
|
||||
/// already revoked, succeeds and says so. A teardown that runs twice is
|
||||
/// ordinary, and a second run that failed would be a worse fault than the one
|
||||
/// this exists to fix.
|
||||
///
|
||||
/// # Errors
|
||||
/// When the store cannot be reached or refuses the delete. The caller decides
|
||||
/// what that costs — `set_agent_state` logs it and lets the destroy proceed,
|
||||
/// since a credential that is still live is a smaller harm than an agent that
|
||||
/// cannot be torn down.
|
||||
pub async fn revoke_queue_credential(agent: &str) -> Result<()> {
|
||||
let queue_path = queue::agent_queue_path(agent)?;
|
||||
let store = crate::store::connect()
|
||||
.await
|
||||
.context("logging in to the swarm secret store")?;
|
||||
store
|
||||
.delete_all_versions(&queue_path)
|
||||
.await
|
||||
.with_context(|| format!("revoking the agent queue credential at {queue_path}"))?;
|
||||
// At `info` and unconditional: an unlogged revocation is indistinguishable
|
||||
// from a leak, and this line is the only record an operator has that the
|
||||
// credential stopped being usable. It cannot say whether one was there —
|
||||
// the controller's grant on these paths is write-only by design, so it
|
||||
// deletes blind.
|
||||
tracing::info!(
|
||||
agent,
|
||||
%queue_path,
|
||||
"agent queue credential revoked: every version of the path deleted"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// The consumer of everything [`mint_and_verify`] wrote: log in **as the
|
||||
/// agent**, with the leaf just issued, and read back both paths just
|
||||
/// published.
|
||||
|
|
|
|||
|
|
@ -1050,9 +1050,55 @@ async fn set_agent_state(
|
|||
tracing::warn!(hive = %hive, agent = %agent, error = %format!("{e:#}"), "declaring agent state failed");
|
||||
error_problem(wanted_error_status(&e), &format!("{e:#}"))
|
||||
})?;
|
||||
|
||||
// After the declaration and never before it: the write above is the
|
||||
// destroy order, and until it lands the agent is not being torn down, so
|
||||
// its credential has to keep working. Sequenced rather than spawned so
|
||||
// the revocation attempt is finished — and its outcome logged — by the
|
||||
// time the operator's call returns.
|
||||
if revokes_queue_credential(req.state) {
|
||||
revoke_or_complain(&hive, &agent).await;
|
||||
}
|
||||
Ok(Json(render(&declaration)))
|
||||
}
|
||||
|
||||
/// Whether declaring an agent into `state` ends its queue credential's life.
|
||||
///
|
||||
/// Exhaustive on purpose, like every other match on [`AgentState`] in this
|
||||
/// tree: a state added later has to say whether it revokes rather than
|
||||
/// inheriting "no" from a catch-all. `Offline` and `Paused` deliberately do
|
||||
/// not — both are states an agent comes back from, and revoking on either
|
||||
/// would mean an agent that could be stopped but never restarted.
|
||||
fn revokes_queue_credential(state: swarm_queue_client::wanted::AgentState) -> bool {
|
||||
use swarm_queue_client::wanted::AgentState;
|
||||
match state {
|
||||
AgentState::Destroyed => true,
|
||||
AgentState::Up | AgentState::Offline | AgentState::Paused => false,
|
||||
}
|
||||
}
|
||||
|
||||
/// Revoke the credential, and make a failure impossible to miss without
|
||||
/// letting it stop the teardown.
|
||||
///
|
||||
/// **The destroy is not blocked by this.** The declaration is already
|
||||
/// published, the hive will converge on it, and refusing the operator's call
|
||||
/// at this point would leave them with an agent they cannot destroy — the
|
||||
/// larger of the two harms. What is left is that the failure must not be
|
||||
/// quiet, or the outcome is the silent orphan this whole path exists to
|
||||
/// remove: hence `error`, naming the agent, with the cause chain attached.
|
||||
async fn revoke_or_complain(hive: &str, agent: &str) {
|
||||
if let Err(e) = agent_identity::revoke_queue_credential(agent).await {
|
||||
tracing::error!(
|
||||
hive,
|
||||
agent,
|
||||
error = %format!("{e:#}"),
|
||||
"revoking this agent's queue credential FAILED; it was destroyed anyway, so \
|
||||
the credential outlives it and may still authenticate. Declaring the agent \
|
||||
destroyed again re-runs this revocation."
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// What this swarm currently declares for a hive.
|
||||
///
|
||||
/// Read back from the bucket rather than from a second copy kept here: the
|
||||
|
|
|
|||
Loading…
Reference in a new issue