swarm-controller: answer 503 on a disconnected queue, not 500 or a hang
put_matrix_account writes the credential to bao before checking that the queue it must notify is actually connected — only that a queue is configured, via state.status.as_ref(). While the queue is Pending or Disconnected, client.flush().await hangs (async-nats does not process commands during the initial connect retry), so the request hangs until nginx times out and the credential is already stored. Add the ensure_connected check every sibling queue route already makes (wanted.rs, term_stream.rs, agent_state_stream.rs), placed immediately before store::connect() so a request that cannot be delivered never reaches the store. Same shape, lower impact, in webhook::announce_knowledge_change: it awaits publish/flush inline in the webhook handler, so a disconnected queue can run past forgejo's short delivery timeout. Add the same ensure_connected guard, warn and return. set_agent_state and get_hive_wanted map every writer error to 500, including swarm_queue_client::Error::NotConnected surfaced through WantedWriter::view/set. Add wanted_error_status, which downcasts the anyhow::Error back to the concrete type and maps NotConnected to 503 (retryable) while leaving every other failure at 500. Update both routes' OpenAPI descriptions to say so. Closes #4688
This commit is contained in:
parent
0bfe354b6d
commit
6a87b25488
3 changed files with 327 additions and 14 deletions
|
|
@ -419,6 +419,16 @@ async fn announce_knowledge_change(state: &AppState) {
|
|||
let client = status.queue_client();
|
||||
let subject = swarm_queue_client::KNOWLEDGE_SUBJECT;
|
||||
|
||||
// A queue that is configured but not (yet, or no longer) connected does
|
||||
// not fail `publish`/`flush` below outright — it *hangs* them, per
|
||||
// `ensure_connected`'s own doc — and this fn is awaited inline from the
|
||||
// webhook handler, so that hang would sit inside forgejo's own delivery
|
||||
// timeout instead of failing soft the way every other branch here does.
|
||||
if let Err(e) = swarm_queue_client::ensure_connected(&client) {
|
||||
tracing::warn!(%subject, error = %e, "webhook: swarm queue is not connected; not publishing");
|
||||
return;
|
||||
}
|
||||
|
||||
// One publish, not one per hive: every subscriber gets the same empty
|
||||
// event, so the roster is not consulted at all. The controller does not
|
||||
// need to know who the hives are in order to say the repository moved.
|
||||
|
|
|
|||
Loading…
Reference in a new issue