Watch
0
0
Fork
You've already forked hyperhive
0

swarm-controller: answer 503 on a disconnected queue, not 500 or a hang

put_matrix_account writes the credential to bao before checking that
the queue it must notify is actually connected — only that a queue is
configured, via state.status.as_ref(). While the queue is Pending or
Disconnected, client.flush().await hangs (async-nats does not process
commands during the initial connect retry), so the request hangs until
nginx times out and the credential is already stored. Add the
ensure_connected check every sibling queue route already makes
(wanted.rs, term_stream.rs, agent_state_stream.rs), placed immediately
before store::connect() so a request that cannot be delivered never
reaches the store.

Same shape, lower impact, in webhook::announce_knowledge_change: it
awaits publish/flush inline in the webhook handler, so a disconnected
queue can run past forgejo's short delivery timeout. Add the same
ensure_connected guard, warn and return.

set_agent_state and get_hive_wanted map every writer error to 500,
including swarm_queue_client::Error::NotConnected surfaced through
WantedWriter::view/set. Add wanted_error_status, which downcasts the
anyhow::Error back to the concrete type and maps NotConnected to 503
(retryable) while leaving every other failure at 500. Update both
routes' OpenAPI descriptions to say so.

Closes #4688
This commit is contained in:
atlas 2026-09-24 14:05:18 +02:00 • committed by mara
commit 6a87b25488
3 changed files with 327 additions and 14 deletions

View file

@ -419,6 +419,16 @@ async fn announce_knowledge_change(state: &AppState) {
let client = status.queue_client();
let subject = swarm_queue_client::KNOWLEDGE_SUBJECT;
// A queue that is configured but not (yet, or no longer) connected does
// not fail `publish`/`flush` below outright — it *hangs* them, per
// `ensure_connected`'s own doc — and this fn is awaited inline from the
// webhook handler, so that hang would sit inside forgejo's own delivery
// timeout instead of failing soft the way every other branch here does.
if let Err(e) = swarm_queue_client::ensure_connected(&client) {
tracing::warn!(%subject, error = %e, "webhook: swarm queue is not connected; not publishing");
return;
}
// One publish, not one per hive: every subscriber gets the same empty
// event, so the roster is not consulted at all. The controller does not
// need to know who the hives are in order to say the repository moved.