Both accept loops returned on the first accept() error. The AgentSocket
stayed in Coordinator.agents, so mcp_sockets::sync_on_start skipped the
agent, and only a Start node or a hive-c0re restart bound it again. The
unit sets no LimitNOFILE, so a single EMFILE at the 1024-fd soft limit
cut the agent off from send/recv/ack with one warn line as the only
trace.
Both loops now share accept_until_fatal. Connection errors
(ECONNABORTED/ECONNRESET/ECONNREFUSED) retry at once, any other error
retries after 1 s, following axum's serve loop that the dashboard
already runs. Errors that leave the listening fd unusable (EBADF,
EFAULT, EINVAL, ENOTSOCK, EOPNOTSUPP) log at error and exit the
process: nothing re-binds a listener while hive-c0re runs, the manager
listener has no re-bind path at all, and the unit's Restart=on-failure
restart runs start_manager and sync_on_start, which re-bind every
socket.
mcp_sockets.rs now states that invariant instead of asserting that a
listener can only disappear on restart.
LimitNOFILE is left unset: per-agent fd use is a listener plus a
Recv long-poll plus short-lived requests, and a higher limit would raise
the memory ceiling of the per-line bound (bound x open connections).
Closes#4721