feat(#2290): maintain /etc/tmpfiles.d/hyperhive-agents.conf for boot safety
Root cause of the boot outage: container@h-* units try to start before
hive-c0re reaches ensure_runtime, so bind-mount source dirs are missing.
Fix: hive-c0re (via hive-priv, which runs as root) writes
/etc/tmpfiles.d/hyperhive-agents.conf whenever the agent set changes.
systemd-tmpfiles-setup.service (sysinit.target) reads it at every boot
BEFORE any container units start, pre-creating:
/run/hyperhive/agents/<name> — MCP socket dir (bind -> /run/hive)
/run/hive-agent/<name> — web socket dir (bind -> /run/hive-agent)
This alone removes the outage class: even if hive-c0re is slow to start,
the bind-mount sources exist and container units can activate.
Added:
- PrivRequest::SyncAgentTmpfiles { agents } in hive-sh4re
- sync_agent_tmpfiles() in hive-priv: generates content, writes atomically,
calls systemd-tmpfiles --create to apply immediately
- priv_client::sync_agent_tmpfiles() wrapper
- lifecycle::sync_tmpfiles() best-effort helper (list + priv call)
- Call sites: hive-c0re startup, handle_spawn success, destroy success
This commit is contained in:
parent
9cd408de8b
commit
9d1f5ebe76
7 changed files with 128 additions and 0 deletions
|
|
@ -722,6 +722,33 @@ pub async fn list() -> Result<Vec<String>> {
|
|||
.collect())
|
||||
}
|
||||
|
||||
/// Sync `/etc/tmpfiles.d/hyperhive-agents.conf` with the currently-known
|
||||
/// agent set (from `nixos-container list`). Strips the `h-` prefix to get
|
||||
/// logical names. Best-effort: errors are logged but never propagated — a
|
||||
/// failed tmpfiles write shouldn't block a spawn or destroy.
|
||||
///
|
||||
/// Called at hive-c0re startup and after each spawn / destroy so the file
|
||||
/// always reflects the live agent set. `systemd-tmpfiles-setup.service`
|
||||
/// reads the file at boot (before any container units start), pre-creating
|
||||
/// bind-mount source dirs so container@h-* units don't race hive-c0re.
|
||||
pub async fn sync_tmpfiles() {
|
||||
let agents = match list().await {
|
||||
Ok(containers) => containers
|
||||
.into_iter()
|
||||
.filter_map(|c| c.strip_prefix(AGENT_PREFIX).map(str::to_owned))
|
||||
.collect::<Vec<_>>(),
|
||||
Err(e) => {
|
||||
tracing::warn!(error = ?e, "sync_tmpfiles: list failed; skipping");
|
||||
return;
|
||||
}
|
||||
};
|
||||
if let Err(e) = crate::priv_client::sync_agent_tmpfiles(&agents).await {
|
||||
tracing::warn!(error = ?e, "sync_tmpfiles: priv call failed");
|
||||
} else {
|
||||
tracing::debug!(count = agents.len(), "sync_tmpfiles: ok");
|
||||
}
|
||||
}
|
||||
|
||||
/// Build the per-line callback for `create_container_streaming` /
|
||||
/// `update_container_streaming`. Both ops share identical dispatch logic
|
||||
/// (stdout → info + `append_stdout`, stderr → warn + `append_stderr`); this
|
||||
|
|
|
|||
Loading…
Reference in a new issue