Watch
0
0
Fork
You've already forked hyperhive
0

hive-c0re: fail on an unparseable resource-limits or topology file, write both atomically

resource-limits.json and topology.json were read with parse errors
folded into an empty map, and written in place with std::fs::write. One
truncated resource-limits.json followed by a single set_limits call
rewrote the file with only that agent's entry, erasing every other
agent's CPU and memory overrides without a log line. topology.json had
the same shape: reconcile rebuilt it from the live set, losing pending
(provisioned, never spawned) names.

- agent_config::read_map / write_map are generic over the stored type.
  tool-groups and capabilities behave as before.
- resource_limits::read / effective return an error for an existing but
  unreadable file; a missing file is still the empty map. set_limits
  fails without writing on such a file, and writes atomically.
- topology: reconcile fails without writing on an unreadable file and
  writes atomically. all_agents logs the error and returns no agents,
  so a ManageRootAgent holder starts without cross-agent mounts.

Read-path behaviour on an unreadable resource-limits.json, per caller:
- write_dropins (every spawn / swap / WriteDropin): logs the error and
  keeps the limits drop-in already under /run; the agent still starts.
  With no drop-in yet (first start since boot) it writes the hive
  defaults, because no drop-in means an uncapped container.
- render_flake: propagates, so sync_agents (and spawn/rebuild/destroy
  jobs) fail. An empty map would give tighter-capped agents the hive
  memoryMaxBytes.
- container_view::build_all: logs the error each scan and renders the
  rows at the hive defaults (no ContainerView wire change).
- set_resource_limits reply: propagates.

Closes #4731
This commit is contained in:
atlas 2026-09-26 15:16:31 +02:00 • committed by mara
commit 6fac00dcc5
8 changed files with 332 additions and 119 deletions

View file

@ -506,10 +506,11 @@ async fn handle_set_resource_limits(
crate::lifecycle::write_dropins(name.as_str(), &hive, &paths).await?;
let (cpu, mem) = crate::resource_limits::effective(
&crate::resource_limits::resource_limits_path(),
name.as_str(),
&hive.agent_cpu_quota,
&hive.agent_memory_max,
);
)?;
Ok(HostResponse::messages(vec![format!(
"{name}: CPUQuota={cpu} MemoryMax={mem} — the cgroup cap itself is live now (restart \
the container if it's running and needs the new cap immediately), but the derived \