hive-c0re: fail a nixos-container destroy that leaves the container in place

lifecycle::destroy logged a failed `nixos-container destroy` and returned
Ok, so the DestroyContainer node went green, the agent was unregistered,
and with purge the after_ok PurgeState deleted its state while the
container config and root still existed.

Propagate the error unless the container list, read after the failure,
no longer names the container. An unreadable list fails too.
This commit is contained in:
atlas 2026-09-24 11:00:51 +02:00 • committed by mara
commit 18bd8dd2c7
3 changed files with 85 additions and 3 deletions

View file

@ -267,7 +267,9 @@ async fn sync_meta_after_lifecycle(coord: &Coordinator) -> Result<()> {
///
/// The only fallible step is the destroy itself: once the container is gone the
/// un-registration cannot meaningfully fail, and returning early would strand
/// the roster claiming an agent that no longer exists.
/// the roster claiming an agent that no longer exists. A destroy that leaves the
/// container in place fails this node before un-registration, and its failure
/// is what keeps the `after_ok` purge and bookkeeping from running.
async fn run_destroy_container(coord: &Arc<Coordinator>, agent: &str) -> Result<()> {
crate::lifecycle::destroy(agent).await?;
coord.unregister_agent(agent);