resource-limits.json and topology.json were read with parse errors
folded into an empty map, and written in place with std::fs::write. One
truncated resource-limits.json followed by a single set_limits call
rewrote the file with only that agent's entry, erasing every other
agent's CPU and memory overrides without a log line. topology.json had
the same shape: reconcile rebuilt it from the live set, losing pending
(provisioned, never spawned) names.
- agent_config::read_map / write_map are generic over the stored type.
tool-groups and capabilities behave as before.
- resource_limits::read / effective return an error for an existing but
unreadable file; a missing file is still the empty map. set_limits
fails without writing on such a file, and writes atomically.
- topology: reconcile fails without writing on an unreadable file and
writes atomically. all_agents logs the error and returns no agents,
so a ManageRootAgent holder starts without cross-agent mounts.
Read-path behaviour on an unreadable resource-limits.json, per caller:
- write_dropins (every spawn / swap / WriteDropin): logs the error and
keeps the limits drop-in already under /run; the agent still starts.
With no drop-in yet (first start since boot) it writes the hive
defaults, because no drop-in means an uncapped container.
- render_flake: propagates, so sync_agents (and spawn/rebuild/destroy
jobs) fail. An empty map would give tighter-capped agents the hive
memoryMaxBytes.
- container_view::build_all: logs the error each scan and renders the
rows at the hive defaults (no ContainerView wire change).
- set_resource_limits reply: propagates.
Closes#4731
set_nspawn_flags propagated has_cap's error, so one corrupt
capabilities.json failed every agent's Start, spawn and Swap. It now
goes through holds_manage_root_agent, which logs the error (agent and
file) and treats the capability as absent: the agent starts without the
cross-agent, /applied and /meta mounts. caps_for/has_cap take the file
path so that seam is testable against a tempfile.
- meta.rs: a comment at the render_flake reads records why they
propagate (an empty tool-groups map renders toolGroups = null, i.e.
AGENT_DEFAULT, which fails open for narrower explicit entries).
- capabilities::read doc: states when set_caps/remove_agent rewrite
the file instead of saying remove_agent repairs it.
- set/remove corrupt-file tests assert ErrorKind::InvalidData.
tool_groups::read and capabilities::read returned an empty map when
their file existed but didn't parse. Every set_*/remove_agent is a
read-modify-write, and write() rewrote the file in place, so a crash or
ENOSPC mid-write left a truncated file, and the next write (e.g. the
manager-spawn seed of ruth's tool groups) replaced it with a map holding
only one agent. The scheduling and approval gates then denied every
other agent, recoverable only from meta git history.
- Both registries now read through agent_config::read_map: a missing
file is still the empty map, any other read failure or a parse
failure is an io::Error. set_groups / set_caps / remove_agent fail
without writing.
- Writes go through agent_config::write_map: temp file in the same
directory, fsync, rename, fsync the directory. hive-c0re had no
shared atomic-write helper (the existing tmp+rename sites are inline
and don't fsync).
- Callers of read / groups_for / has_cap now handle the error:
* dashboard GET /api/tool-groups, /api/capabilities,
/api/permissions/stale return 500 instead of an empty table;
* the SSE permission snapshots are skipped with a warn;
* render_flake returns Result, so sync_agents fails instead of
rendering every agent without its tool groups / capabilities;
* set_nspawn_flags propagates has_cap's error;
* the socket tool-group gates deny with the read error as message;
* seed_manager_tool_groups logs and does not seed.
- capabilities::write had no callers left once set_caps writes through
write_map, and is removed.
Closes#4719