crash_watch's 10s poll and auto_update's ensure_root_agent both read
lifecycle::list().await.unwrap_or_default(), which turned a failed read
into 'zero containers'. In crash_watch that made every previously-running
agent look like it crashed simultaneously (prev.difference(current) over
an empty current), and left prev empty for the next cycle too, so a
second wave of false 'agent logged in' / 'agent needs login' events fired
against the next successful read. In ensure_root_agent it read as
'manager container missing' and called lifecycle::spawn on a manager
that might already exist.
Both sites now treat a list error as its own outcome: log it at warn and
skip the cycle's decision entirely. crash_watch leaves prev exactly as
the last good read produced it. ensure_root_agent attempts no spawn.
Factors each site's decision into a pure helper (plan_cycle /
plan_root_agent) matching the check_not_live / confirm_gone_after_failed_destroy
pattern, with unit tests for the error case, a control for the readable
case, and (for crash_watch) an invert-proof run locally against the old
unwrap_or_default logic before reverting.