hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290) that pre-created every agent's bind sources. The start preamble already creates them for every c0re-driven start, and on this host only hive-c0re starts agent containers. The file was also the reason the socket dir's owner had to be declared there, which is how it spent its life at `0777 root root` whenever the uid could not be resolved (#4742). - hive-priv gains `EnsureAgentSocketDir { name }`, called from `set_nspawn_flags` in every start path. It creates `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a directory is left alone. hive-c0re's own create_dir_all went: its /run is read-only under ProtectSystem=strict. - The container's `hive-agent-user-migrate` activation chowns that dir to the agent user and sets 0751, the same way it already handles state/ and harness/. It refuses a symlink or non-directory there, since `test -d` and chmod follow links. No host-side passwd parse, and no window where the dir is world-writable. - `/run/hyperhive/agents/<name>` stays created by hive-c0re itself (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re binds as hive-core, so it must not become root- or agent-owned. - The `/run/hive-agent` parent is declared in hive-priv.nix, `0755 root:root`, instead of hive-gateway's hive-core rule. hive-priv is its only writer now, and hive-priv's ReadWritePaths needs it to exist. - The manager start in `ensure_root_agent` now goes through `converge_start_preamble` + `start_with_fallback`. It was a bare start, so after a reboot the manager's bind sources existed only because of the tmpfiles file, and its limits drop-in did not exist at all. - Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`, `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles body builder and their tests, plus the three call sites. - Legacy: hive-priv unlinks the file at every start, ignoring ENOENT. `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that does the same unlink and returns Ok, for an older hive-c0re. Salvaged from #4752: the boundary.md correction that nginx only dials, because ProtectSystem=strict makes its /run read-only. Behaviour change: a manual `nixos-container start h-<name>` right after a reboot, before hive-c0re has started that agent, now fails on a missing bind source instead of starting. Closes #4742
This commit is contained in:
parent
e7456a49ff
commit
2252c55df8
21 changed files with 369 additions and 421 deletions
|
|
@ -347,9 +347,6 @@ async fn run_destroy_bookkeeping(coord: &Arc<Coordinator>, agent: &str, purge: b
|
|||
// roster, so any schedule that still targets the just-destroyed agent now
|
||||
// drops that ghost column live (no page reload needed).
|
||||
coord.emit_schedules_snapshot();
|
||||
// Update tmpfiles.d to remove the destroyed agent's dirs from the boot-time
|
||||
// pre-creation list. Best-effort: failure is logged only.
|
||||
tokio::spawn(crate::lifecycle::sync_tmpfiles());
|
||||
}
|
||||
|
||||
/// Emit this agent's rebuild-complete todo. `ok` is not computed — it is which
|
||||
|
|
|
|||
|
|
@ -445,14 +445,11 @@ async fn set_nspawn_flags(
|
|||
// gateway sees it. Subdir bind (not socket file) keeps the inode
|
||||
// visible after the harness unlinks a stale socket on rebind.
|
||||
// Applies to manager and sub-agents alike.
|
||||
// hive-priv creates it (`/run` is read-only to hive-c0re) as `0751 root`
|
||||
// and leaves an existing one alone; the container's activation hands it
|
||||
// to the agent user. Don't chown or chmod it here.
|
||||
let socket_dir = crate::agent_sockets::agent_dir_for(agent_name);
|
||||
std::fs::create_dir_all(&socket_dir)
|
||||
.with_context(|| format!("create {}", socket_dir.display()))?;
|
||||
// Ownership is NOT repaired here. The dir's owner + mode are declared by
|
||||
// the tmpfiles.d entry (`SyncAgentTmpfiles`), which is the mechanism that
|
||||
// re-applies on every boot and every spawn — so a chown made here was
|
||||
// silently reverted the next time any agent was spawned or destroyed.
|
||||
// This `create_dir_all` only covers the window before that sync lands.
|
||||
crate::priv_client::ensure_agent_socket_dir(agent_name).await?;
|
||||
binds.push(BindMount {
|
||||
host_path: socket_dir.to_string_lossy().into_owned(),
|
||||
container_path: socket_dir.to_string_lossy().into_owned(),
|
||||
|
|
|
|||
|
|
@ -178,79 +178,6 @@ pub fn network_isolation_from_env() -> Result<hive_priv_sock::NetworkIsolation>
|
|||
network_isolation_from_vars(bridge.as_deref(), subnet.as_deref())
|
||||
}
|
||||
|
||||
/// Read the agent user's `(uid, gid)` from the container's nixos-managed
|
||||
/// `/etc/passwd`. Returns `None` when the passwd file cannot be read (not
|
||||
/// built yet, or this process cannot reach the path), or when it is read
|
||||
/// but holds no usable entry for the agent (missing user, e.g. a legacy
|
||||
/// container that still runs as root).
|
||||
///
|
||||
/// Used by `forge` + `matrix` after writing per-agent state files so
|
||||
/// the bind-mounted host file ends up readable by the agent user
|
||||
/// without waiting for the next container activation to run the chown
|
||||
/// fixup.
|
||||
///
|
||||
/// Notes:
|
||||
/// - Reads the *container-local* passwd at
|
||||
/// `/var/lib/nixos-containers/<container>/etc/passwd`, not the host's.
|
||||
/// The container's user-namespace shares uids with the host (no
|
||||
/// `PrivateUsers`), so the uid is directly usable in host-side
|
||||
/// `chown(2)`.
|
||||
/// - Best-effort: caller treats `None` as "skip the chown".
|
||||
/// - ⚠️ Every `None` is logged with which of those cases produced it: the
|
||||
/// tmpfiles caller answers `None` by declaring the agent's socket dir
|
||||
/// `0777`, and a fallback that widens a directory has to say why it fired.
|
||||
#[must_use]
|
||||
pub fn agent_uid_gid(agent_name: &str) -> Option<(u32, u32)> {
|
||||
let container = container_name(agent_name);
|
||||
let passwd_path = format!("/var/lib/nixos-containers/{container}/etc/passwd");
|
||||
let content = match std::fs::read_to_string(&passwd_path) {
|
||||
Ok(content) => content,
|
||||
Err(e) => {
|
||||
tracing::warn!(
|
||||
agent = %agent_name,
|
||||
path = %passwd_path,
|
||||
error = ?e,
|
||||
"agent_uid_gid: cannot read the container's passwd"
|
||||
);
|
||||
return None;
|
||||
}
|
||||
};
|
||||
let ids = parse_passwd_uid_gid(&content, agent_name);
|
||||
if ids.is_none() {
|
||||
tracing::warn!(
|
||||
agent = %agent_name,
|
||||
path = %passwd_path,
|
||||
lines = content.lines().count(),
|
||||
"agent_uid_gid: passwd read but no usable entry for this agent"
|
||||
);
|
||||
}
|
||||
ids
|
||||
}
|
||||
|
||||
/// Find `user`'s `(uid, gid)` in a `passwd(5)` body.
|
||||
///
|
||||
/// Separate from the read so it can be tested at all — the caller's half is a
|
||||
/// host path no test can stand up.
|
||||
fn parse_passwd_uid_gid(content: &str, user: &str) -> Option<(u32, u32)> {
|
||||
for line in content.lines() {
|
||||
let mut parts = line.split(':');
|
||||
// Skip, never abort: an unusable row for this same user must not hide
|
||||
// a usable one below it.
|
||||
if parts.next() != Some(user) {
|
||||
continue;
|
||||
}
|
||||
let _password = parts.next();
|
||||
let (Some(uid), Some(gid)) = (parts.next(), parts.next()) else {
|
||||
continue;
|
||||
};
|
||||
let (Ok(uid), Ok(gid)) = (uid.parse(), gid.parse()) else {
|
||||
continue;
|
||||
};
|
||||
return Some((uid, gid));
|
||||
}
|
||||
None
|
||||
}
|
||||
|
||||
fn validate(name: &str) -> Result<()> {
|
||||
if name.is_empty() {
|
||||
bail!("agent name must not be empty");
|
||||
|
|
@ -509,7 +436,7 @@ pub async fn kill(name: &str) -> Result<()> {
|
|||
/// exit polls the unit state for [`START_SETTLE_TIMEOUT`] before concluding
|
||||
/// the start actually failed. Every caller (dashboard restart, reconcile
|
||||
/// start, the cold-start fallback) gets this truth for free.
|
||||
pub async fn start(name: &str) -> Result<()> {
|
||||
async fn start(name: &str) -> Result<()> {
|
||||
validate(name)?;
|
||||
if priv_run("start", name).await.is_ok() {
|
||||
return Ok(());
|
||||
|
|
@ -900,48 +827,9 @@ pub fn agent_names(listed: Result<Vec<String>>) -> Result<Vec<String>> {
|
|||
.collect())
|
||||
}
|
||||
|
||||
/// Sync `/etc/tmpfiles.d/hyperhive-agents.conf` with the currently-known
|
||||
/// agent set (from `nixos-container list`). Strips the `h-` prefix to get
|
||||
/// logical names. Best-effort: errors are logged but never propagated — a
|
||||
/// failed tmpfiles write shouldn't block a spawn or destroy.
|
||||
///
|
||||
/// Called at hive-c0re startup and after each spawn / destroy so the file
|
||||
/// always reflects the live agent set. `systemd-tmpfiles-setup.service`
|
||||
/// reads the file at boot (before any container units start), pre-creating
|
||||
/// bind-mount source dirs so container@h-* units don't race hive-c0re.
|
||||
pub async fn sync_tmpfiles() {
|
||||
let agents = match list().await {
|
||||
Ok(containers) => containers
|
||||
.into_iter()
|
||||
.filter_map(|c| c.strip_prefix(AGENT_PREFIX).map(str::to_owned))
|
||||
.map(|name| {
|
||||
// Resolved here, not in hive-priv: the mapping lives in the
|
||||
// container's /etc/passwd, which is c0re's to read. `None`
|
||||
// until the container's first boot renders it.
|
||||
let (uid, gid) = match agent_uid_gid(&name) {
|
||||
Some((uid, gid)) => (Some(uid), Some(gid)),
|
||||
None => (None, None),
|
||||
};
|
||||
hive_priv_sock::AgentTmpfilesEntry { name, uid, gid }
|
||||
})
|
||||
.collect::<Vec<_>>(),
|
||||
Err(e) => {
|
||||
tracing::warn!(error = ?e, "sync_tmpfiles: list failed; skipping");
|
||||
return;
|
||||
}
|
||||
};
|
||||
if let Err(e) = crate::priv_client::sync_agent_tmpfiles(&agents).await {
|
||||
tracing::warn!(error = ?e, "sync_tmpfiles: priv call failed");
|
||||
} else {
|
||||
tracing::debug!(count = agents.len(), "sync_tmpfiles: ok");
|
||||
}
|
||||
}
|
||||
|
||||
/// Ensure the per-agent runtime directory `/run/hyperhive/agents/<name>`
|
||||
/// exists. The directory is also written by `SyncAgentTmpfiles` (run at
|
||||
/// boot + spawn/destroy), but explicit creation in start/spawn paths guards
|
||||
/// against races where hive-c0re starts a container before tmpfiles.d has
|
||||
/// applied the new entry.
|
||||
/// exists. It is a bind source and `/run` is a tmpfs, so every start path
|
||||
/// creates it before `nixos-container start`.
|
||||
///
|
||||
/// Pure filesystem op — no `Coordinator` dependency — so callers that only
|
||||
/// need the dir do not have to hold an `Arc<Coordinator>`.
|
||||
|
|
|
|||
|
|
@ -283,50 +283,6 @@ fn an_unknown_state_is_never_reported_as_failed() {
|
|||
}
|
||||
}
|
||||
|
||||
/// A malformed row for the wanted user must not end the scan.
|
||||
///
|
||||
/// The old implementation used `?` on the field reads, which are only reached
|
||||
/// once the name matches — so a short or unparseable row *for this agent*
|
||||
/// returned `None` from the whole function and hid a usable entry below it.
|
||||
/// The caller answers that `None` by declaring the socket dir `0777`.
|
||||
#[test]
|
||||
fn passwd_parse_keeps_scanning_past_a_malformed_row_for_the_same_user() {
|
||||
// The two `atlas` rows above the real one are what matters: a malformed
|
||||
// row for the SAME user is what used to end the scan, so the usable entry
|
||||
// below it was never reached. Rows for other users were always skipped
|
||||
// fine. Both unusable shapes are here because they leave the parser by
|
||||
// different arms — missing fields, and fields that will not parse — and a
|
||||
// mutation run showed the second arm untested when only the first was.
|
||||
let body = "\
|
||||
root:x:0:0:System administrator:/root:/bin/sh
|
||||
# a comment line, not a passwd entry
|
||||
atlas:x
|
||||
atlas:x:notanumber:994::/:/bin/sh
|
||||
atlas:x:1000:994:atlas:/home/atlas:/bin/sh
|
||||
";
|
||||
assert_eq!(parse_passwd_uid_gid(body, "atlas"), Some((1000, 994)));
|
||||
// Control: the same body must NOT answer for a user it does not carry,
|
||||
// or the assertion above is satisfied by any parse at all.
|
||||
assert_eq!(parse_passwd_uid_gid(body, "nobody-here"), None);
|
||||
}
|
||||
|
||||
/// A matching name with unusable ids is not a match. Without this the parser
|
||||
/// could return a partly-parsed row and the caller would chown to it.
|
||||
#[test]
|
||||
fn passwd_parse_rejects_unusable_ids() {
|
||||
assert_eq!(
|
||||
parse_passwd_uid_gid("atlas:x:notanumber:994::/:/bin/sh", "atlas"),
|
||||
None
|
||||
);
|
||||
assert_eq!(parse_passwd_uid_gid("atlas:x:1000", "atlas"), None);
|
||||
assert_eq!(parse_passwd_uid_gid("", "atlas"), None);
|
||||
// Presence control for the three absences above.
|
||||
assert_eq!(
|
||||
parse_passwd_uid_gid("atlas:x:1000:994::/:/bin/sh", "atlas"),
|
||||
Some((1000, 994))
|
||||
);
|
||||
}
|
||||
|
||||
/// A failed `nixos-container destroy` whose container is still listed fails,
|
||||
/// and keeps hive-priv's error in the chain. The destroy DAG gates the state
|
||||
/// purge on this call, so this is the case that must not read as success.
|
||||
|
|
|
|||
|
|
@ -287,10 +287,6 @@ async fn cmd_serve(
|
|||
if let Err(e) = auto_update::ensure_root_agent(&coord).await {
|
||||
tracing::warn!(error = ?e, "auto-spawn root agent failed");
|
||||
}
|
||||
// Sync /etc/tmpfiles.d/hyperhive-agents.conf so agent runtime dirs are
|
||||
// pre-declared for the next boot. Best-effort background task — a failure
|
||||
// here must not block hive-c0re startup. See lifecycle::sync_tmpfiles.
|
||||
tokio::spawn(crate::lifecycle::sync_tmpfiles());
|
||||
// Auto-update in the background — don't block service start.
|
||||
// Sub-agent rebuilds can take tens of seconds; we want the admin
|
||||
// socket up immediately.
|
||||
|
|
|
|||
|
|
@ -8,8 +8,8 @@
|
|||
|
||||
use anyhow::{Context as _, Result, bail};
|
||||
use hive_priv_sock::{
|
||||
AgentTmpfilesEntry, BindMount, CredentialMount, InfraAction, InfraContainer, JournalQuery,
|
||||
NetworkIsolation, PRIV_SOCK, PrivEvent, PrivRequest, PrivResponse, PrivStream,
|
||||
BindMount, CredentialMount, InfraAction, InfraContainer, JournalQuery, NetworkIsolation,
|
||||
PRIV_SOCK, PrivEvent, PrivRequest, PrivResponse, PrivStream,
|
||||
};
|
||||
use std::os::fd::{AsRawFd as _, OwnedFd, RawFd};
|
||||
|
||||
|
|
@ -306,6 +306,14 @@ pub async fn write_nspawn_flags(
|
|||
.await?)
|
||||
}
|
||||
|
||||
/// See [`PrivRequest::EnsureAgentSocketDir`].
|
||||
pub async fn ensure_agent_socket_dir(name: &str) -> Result<()> {
|
||||
ok(call(&PrivRequest::EnsureAgentSocketDir {
|
||||
name: name.to_owned(),
|
||||
})
|
||||
.await?)
|
||||
}
|
||||
|
||||
pub async fn write_resource_limits(
|
||||
container: &str,
|
||||
memory_max: &str,
|
||||
|
|
@ -616,22 +624,6 @@ pub async fn send_agent_snapshot_to_file(
|
|||
Ok(stdout)
|
||||
}
|
||||
|
||||
/// Write `/etc/tmpfiles.d/hyperhive-agents.conf` for `agents` and immediately
|
||||
/// apply it with `systemd-tmpfiles --create`. Each entry carries the agent's
|
||||
/// container uid/gid so the socket dir's ownership is *declared* here rather
|
||||
/// than corrected afterwards. See [`PrivRequest::SyncAgentTmpfiles`].
|
||||
///
|
||||
/// # Errors
|
||||
///
|
||||
/// Returns an error if the priv socket call fails, if any agent name is
|
||||
/// invalid, or if `systemd-tmpfiles --create` exits non-zero.
|
||||
pub async fn sync_agent_tmpfiles(agents: &[AgentTmpfilesEntry]) -> Result<()> {
|
||||
ok(call(&PrivRequest::SyncAgentTmpfiles {
|
||||
agents: agents.to_vec(),
|
||||
})
|
||||
.await?)
|
||||
}
|
||||
|
||||
/// Parse `(referenced, exclusive)` bytes from `btrfs qgroup show -f --raw`
|
||||
/// output (a qgroup row is `<id-with-slash> <rfer> <excl> …`).
|
||||
///
|
||||
|
|
|
|||
|
|
@ -313,8 +313,6 @@ async fn handle_spawn(coord: &Arc<Coordinator>, name: &str) -> Result<HostRespon
|
|||
None,
|
||||
)
|
||||
.await;
|
||||
// Update tmpfiles.d so the new agent's dirs survive a reboot.
|
||||
tokio::spawn(lifecycle::sync_tmpfiles());
|
||||
}
|
||||
Err(e) => {
|
||||
// Spawn failed: register_agent was never called, so there is
|
||||
|
|
|
|||
|
|
@ -179,7 +179,19 @@ pub async fn ensure_root_agent(coord: &Arc<Coordinator>) -> Result<()> {
|
|||
if let Err(e) = coord.power.set(MANAGER_NAME, crate::power::Wanted::Up) {
|
||||
tracing::warn!(error = ?e, "agent_power: set manager wanted=up failed");
|
||||
}
|
||||
if let Err(e) = lifecycle::start(MANAGER_NAME).await {
|
||||
// Through the preamble: after a reboot its bind sources and limits
|
||||
// drop-in in `/run` are gone, and nothing else recreates them.
|
||||
let hive = coord.hive_env();
|
||||
let started = async {
|
||||
let paths = Coordinator::agent_paths(
|
||||
MANAGER_NAME,
|
||||
crate::paths::agent_runtime_dir(MANAGER_NAME),
|
||||
)?;
|
||||
let token = lifecycle::converge_start_preamble(MANAGER_NAME, &hive, &paths).await?;
|
||||
lifecycle::start_with_fallback(token).await
|
||||
}
|
||||
.await;
|
||||
if let Err(e) = started {
|
||||
tracing::warn!(error = ?e, "manager start failed");
|
||||
}
|
||||
}
|
||||
|
|
|
|||
Loading…
Reference in a new issue