Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/hive-c0re/src/agent_config/resource_limits.rs
atlas 6fac00dcc5 hive-c0re: fail on an unparseable resource-limits or topology file, write both atomically
resource-limits.json and topology.json were read with parse errors
folded into an empty map, and written in place with std::fs::write. One
truncated resource-limits.json followed by a single set_limits call
rewrote the file with only that agent's entry, erasing every other
agent's CPU and memory overrides without a log line. topology.json had
the same shape: reconcile rebuilt it from the live set, losing pending
(provisioned, never spawned) names.

- agent_config::read_map / write_map are generic over the stored type.
  tool-groups and capabilities behave as before.
- resource_limits::read / effective return an error for an existing but
  unreadable file; a missing file is still the empty map. set_limits
  fails without writing on such a file, and writes atomically.
- topology: reconcile fails without writing on an unreadable file and
  writes atomically. all_agents logs the error and returns no agents,
  so a ManageRootAgent holder starts without cross-agent mounts.

Read-path behaviour on an unreadable resource-limits.json, per caller:
- write_dropins (every spawn / swap / WriteDropin): logs the error and
  keeps the limits drop-in already under /run; the agent still starts.
  With no drop-in yet (first start since boot) it writes the hive
  defaults, because no drop-in means an uncapped container.
- render_flake: propagates, so sync_agents (and spawn/rebuild/destroy
  jobs) fail. An empty map would give tighter-capped agents the hive
  memoryMaxBytes.
- container_view::build_all: logs the error each scan and renders the
  rows at the hive defaults (no ContainerView wire change).
- set_resource_limits reply: propagates.

Closes #4731
2026-09-27 02:41:38 +02:00

577 lines
22 KiB
Rust

//! Per-agent CPU/memory limit overrides. Stored at
//! `/var/lib/hyperhive/meta/resource-limits.json` alongside
//! `topology.json`, `tool-groups.json` and `capabilities.json`.
//!
//! Format: a JSON object mapping agent name to an object with optional
//! `cpu_quota` / `memory_max` strings, passed verbatim to systemd's
//! `CPUQuota=` / `MemoryMax=` in the per-container drop-in:
//!
//! ```json
//! {
//! "sock": { "cpu_quota": "400%", "memory_max": "8G" }
//! }
//! ```
//!
//! Fallback is **per field**: an absent file, an absent agent, or an
//! absent field all fall back to the hive-wide
//! `services.hyperhive.agentCpuQuota` / `agentMemoryMax`. So an agent
//! can raise only its memory cap and keep tracking the hive default for
//! CPU — see [`effective`].
//!
//! Read path: `lifecycle::host_config::write_dropins`, on every spawn
//! and every rebuild.
//!
//! Why host-side JSON and not an option in the agent's own `agent.nix`:
//! the drop-in lands on the *host's* `container@h-<name>.service`, so
//! c0re would have to `nix eval` the agent's whole nixosConfiguration
//! just to read two strings. It is also the wrong trust boundary —
//! a resource *cap* should not be sourced from the capped party.
use std::collections::BTreeMap;
use std::path::{Path, PathBuf};
const RESOURCE_LIMITS_FILE: &str = "resource-limits.json";
/// One agent's overrides. Both fields optional and independent; `None`
/// means "use the hive-wide default for this field".
#[derive(Debug, Clone, Default, PartialEq, Eq, serde::Serialize, serde::Deserialize)]
pub struct AgentLimits {
/// systemd `CPUQuota=` value, e.g. `"400%"`.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub cpu_quota: Option<String>,
/// systemd `MemoryMax=` value, e.g. `"8G"`.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub memory_max: Option<String>,
}
impl AgentLimits {
/// True when neither field is set — such an entry is dropped rather
/// than persisted as an empty object.
#[must_use]
pub fn is_empty(&self) -> bool {
self.cpu_quota.is_none() && self.memory_max.is_none()
}
}
#[must_use]
pub fn resource_limits_path() -> PathBuf {
crate::paths::meta_root().join(RESOURCE_LIMITS_FILE)
}
/// Read the per-agent limit map. An absent file is the empty map —
/// callers treat a missing entry as "hive-wide defaults". A file that
/// exists but can't be read or parsed is an error: an override can be
/// tighter than the hive default, so reading it as empty is not a safe
/// fallback.
pub fn read() -> std::io::Result<BTreeMap<String, AgentLimits>> {
super::read_map(&resource_limits_path())
}
/// Resolve the effective values for an agent from the override file at
/// `path` (normally [`resource_limits_path`]), filling each unset field
/// from the hive-wide default. This is the single place the fallback
/// rule lives; `write_dropins` calls it and passes the result straight
/// to systemd. Errors as [`read`] does.
///
/// Use [`effective_from`] when resolving more than one agent in a row.
pub fn effective(
path: &Path,
name: &str,
hive_cpu_quota: &str,
hive_memory_max: &str,
) -> std::io::Result<(String, String)> {
let limits = super::read_map(path)?;
Ok(effective_from(
&limits,
name,
hive_cpu_quota,
hive_memory_max,
))
}
/// [`effective`] against an already-loaded map — the multi-agent form.
///
/// `container_view::build_all` renders every agent on each SSE scan, so
/// it loads the map once and calls this per agent rather than re-reading
/// the same small file N times per scan.
#[must_use]
pub fn effective_from(
limits: &BTreeMap<String, AgentLimits>,
name: &str,
hive_cpu_quota: &str,
hive_memory_max: &str,
) -> (String, String) {
match limits.get(name) {
Some(l) => resolve(l, hive_cpu_quota, hive_memory_max),
None => (hive_cpu_quota.to_owned(), hive_memory_max.to_owned()),
}
}
/// Pure core of [`effective`], split out so the fallback matrix is
/// testable without touching the filesystem.
#[must_use]
fn resolve(limits: &AgentLimits, hive_cpu_quota: &str, hive_memory_max: &str) -> (String, String) {
let cpu = limits
.cpu_quota
.clone()
.unwrap_or_else(|| hive_cpu_quota.to_owned());
let mem = limits
.memory_max
.clone()
.unwrap_or_else(|| hive_memory_max.to_owned());
(cpu, mem)
}
/// Effective `MemoryMax=` string for `name` against an already-loaded
/// map — the memory-only half of [`effective_from`], split out because
/// [`effective_memory_bytes_from`] doesn't have a CPU quota to pass.
/// No single-agent convenience wrapper (unlike [`effective`] /
/// [`effective_from`]): every current caller already has a loaded map
/// in hand, so one would just be dead code.
#[must_use]
fn effective_memory_max_from(
limits: &BTreeMap<String, AgentLimits>,
name: &str,
hive_memory_max: &str,
) -> String {
limits
.get(name)
.and_then(|l| l.memory_max.clone())
.unwrap_or_else(|| hive_memory_max.to_owned())
}
/// Effective `MemoryMax=` for `name`, as a raw byte count, against an
/// already-loaded map — the form `render_flake_with_lookup` uses so it
/// doesn't re-read `resource-limits.json` once per agent (same
/// reasoning as [`effective_from`] / `container_view::build_all`). No
/// single-agent convenience wrapper, for the same reason as
/// [`effective_memory_max_from`] — every current caller already has a
/// loaded map in hand.
///
/// Logs a `warn!` naming `name` when the effective value is a RAM
/// percentage: unlike `"infinity"`, a percentage IS a real, resolvable
/// cap — resolving one just needs host `MemTotal`, which this module
/// doesn't track — so today it degrades to `None` (no derived heap
/// ceiling) rather than silently guessing wrong against `MemTotal`.
#[must_use]
pub fn effective_memory_bytes_from(
limits: &BTreeMap<String, AgentLimits>,
name: &str,
hive_memory_max: &str,
) -> Option<u64> {
let mem = effective_memory_max_from(limits, name, hive_memory_max);
if mem.ends_with('%') {
tracing::warn!(
agent = name,
memory_max = %mem,
"effective MemoryMax= is a RAM percentage; can't derive a JSC heap ceiling from it \
without host MemTotal — leaving BUN_JSC_forceRAMSize unset for this agent"
);
}
parse_bytes(&mem)
}
/// Parse a systemd `MemoryMax=`-style byte-size value (`"4G"`, `"512M"`,
/// a bare byte count, optionally with a decimal like `"1.5G"`) into a
/// raw byte count. systemd's `K`/`M`/`G`/`T` suffixes are IEC binary
/// (1024-based), not decimal — this matches. Pure integer arithmetic
/// throughout (via `u128` headroom) rather than `f64`: a `MemoryMax=`
/// value is always a small non-negative decimal (enforced by
/// [`is_plain_number`] upstream in [`validate_memory_max`]), so floats
/// would only add rounding / sign-loss risk for no benefit. Returns
/// `None` for `"infinity"` and percentages; see
/// [`effective_memory_bytes_from`] for why those can't be turned into a
/// byte count here.
#[must_use]
pub fn parse_bytes(value: &str) -> Option<u64> {
if value == "infinity" || value.ends_with('%') {
return None;
}
let (mantissa, exponent) = [('K', 1u32), ('M', 2), ('G', 3), ('T', 4)]
.into_iter()
.find_map(|(suffix, exp)| {
value
.strip_suffix([suffix, suffix.to_ascii_lowercase()])
.map(|m| (m, exp))
})
.unwrap_or((value, 0));
let scale = u128::from(1024u64.checked_pow(exponent)?);
let (int_part, frac_part) = mantissa.split_once('.').unwrap_or((mantissa, ""));
let whole: u128 = int_part.parse().ok()?;
let mut bytes = whole.checked_mul(scale)?;
if !frac_part.is_empty() {
let frac_num: u128 = frac_part.parse().ok()?;
let frac_denom = 10u128.checked_pow(u32::try_from(frac_part.len()).ok()?)?;
bytes = bytes.checked_add(frac_num.checked_mul(scale)? / frac_denom)?;
}
u64::try_from(bytes).ok()
}
/// Set one agent's overrides and persist the map atomically. An entry
/// with both fields unset is removed rather than stored, so "reset to
/// hive defaults" and "never configured" are the same state on disk.
///
/// # Errors
///
/// Fails without writing when the existing file can't be read or parsed
/// (`InvalidData` for a parse failure), and otherwise when the meta dir
/// or the file can't be written.
pub fn set_limits(name: &str, limits: &AgentLimits) -> std::io::Result<()> {
set_limits_at(&resource_limits_path(), name, limits)
}
fn set_limits_at(path: &Path, name: &str, limits: &AgentLimits) -> std::io::Result<()> {
let mut current: BTreeMap<String, AgentLimits> = super::read_map(path)?;
if limits.is_empty() {
current.remove(name);
} else {
current.insert(name.to_owned(), limits.clone());
}
super::write_map(path, &current)
}
/// Validate a systemd `CPUQuota=` value. Percentages only, and values
/// above 100% are legitimate (one full core is 100%).
///
/// Validated because the value is written verbatim into a drop-in: a
/// typo makes systemd reject the unit, which means the container stops
/// starting at all. Better to refuse at the CLI than to brick a spawn.
///
/// # Errors
///
/// Returns a human-readable message naming the offending value when it
/// isn't a percentage. The string is surfaced straight to the operator,
/// so it names the expected shape rather than just saying "invalid".
pub fn validate_cpu_quota(value: &str) -> Result<(), String> {
if is_percentage(value) {
return Ok(());
}
Err(format!(
"invalid CPUQuota {value:?}: expected a percentage such as \"200%\" \
(100% = one full core)"
))
}
/// Validate a systemd `MemoryMax=` value — a byte count with an optional
/// `K`/`M`/`G`/`T` suffix, a percentage of physical memory, or the literal
/// `infinity` — and return the exact string to store and pass to systemd.
///
/// That's not always `value` itself: a size is accepted case-insensitively
/// and with an optional redundant trailing `B` (`8gb`, `8Gb`, `8GB`, `8g`,
/// `8G` are all the same size to a human), but only `8G` is what systemd's
/// own parser actually takes — confirmed directly against a running
/// systemd 260: `MemoryMax=8GB` and `MemoryMax=8g` both fail unit
/// activation with "Invalid argument", while `MemoryMax=8G` starts fine.
/// Accepting the friendly spellings and normalizing them here means every
/// caller downstream only ever sees the one form systemd accepts, instead
/// of everyone re-deriving that systemd is this particular about case and
/// the redundant `B`.
///
/// # Errors
///
/// Returns a human-readable message naming the offending value and the
/// three accepted shapes. Same operator-facing contract as
/// [`validate_cpu_quota`].
pub fn validate_memory_max(value: &str) -> Result<String, String> {
if value == "infinity" || is_percentage(value) {
return Ok(value.to_owned());
}
if let Some(normalized) = normalized_byte_size(value) {
return Ok(normalized);
}
Err(format!(
"invalid MemoryMax {value:?}: expected a size such as \"8G\", a percentage \
such as \"50%\", or \"infinity\""
))
}
/// A decimal number followed by `%`.
fn is_percentage(value: &str) -> bool {
value.strip_suffix('%').is_some_and(is_plain_number)
}
/// A decimal number with an optional `K`/`M`/`G`/`T` suffix — accepted
/// case-insensitively and with an optional trailing `B` (`gb`/`Gb`/`GB`/`g`
/// are all read as `G`) — returned in the single uppercase-letter-no-
/// redundant-`B` form systemd itself accepts. See [`validate_memory_max`]
/// for why. A trailing `B` with no multiplier before it (`1048576B`) is a
/// bare byte count and kept as-is — systemd accepts `B` as a unit on its
/// own, unlike the redundant `KB`/`MB`/`GB`/`TB` combinations.
fn normalized_byte_size(value: &str) -> Option<String> {
if let Some(before_b) = value.strip_suffix(['B', 'b']) {
return match before_b.strip_suffix(['K', 'M', 'G', 'T', 'k', 'm', 'g', 't']) {
Some(mantissa) if is_plain_number(mantissa) => {
let unit = before_b[mantissa.len()..].to_ascii_uppercase();
Some(format!("{mantissa}{unit}"))
}
Some(_) => None,
None => is_plain_number(before_b).then(|| format!("{before_b}B")),
};
}
match value.strip_suffix(['K', 'M', 'G', 'T', 'k', 'm', 'g', 't']) {
Some(mantissa) if is_plain_number(mantissa) => {
let unit = value[mantissa.len()..].to_ascii_uppercase();
Some(format!("{mantissa}{unit}"))
}
Some(_) => None,
None => is_plain_number(value).then(|| value.to_owned()),
}
}
/// Digits, optionally followed by a single `.` and more digits. Hand
/// rolled rather than pulling in a regex dependency for two patterns;
/// deliberately rejects the exponent/sign forms `f64::from_str` accepts,
/// since systemd wouldn't take them either.
fn is_plain_number(value: &str) -> bool {
let mut parts = value.splitn(2, '.');
let int = parts.next().unwrap_or_default();
if int.is_empty() || !int.bytes().all(|b| b.is_ascii_digit()) {
return false;
}
match parts.next() {
None => true,
Some(frac) => !frac.is_empty() && frac.bytes().all(|b| b.is_ascii_digit()),
}
}
#[cfg(test)]
mod tests {
use super::*;
const HIVE_CPU: &str = "200%";
const HIVE_MEM: &str = "4G";
fn limits(cpu: Option<&str>, mem: Option<&str>) -> AgentLimits {
AgentLimits {
cpu_quota: cpu.map(ToOwned::to_owned),
memory_max: mem.map(ToOwned::to_owned),
}
}
#[test]
fn unset_agent_gets_hive_defaults() {
let (cpu, mem) = resolve(&AgentLimits::default(), HIVE_CPU, HIVE_MEM);
assert_eq!(cpu, "200%");
assert_eq!(mem, "4G");
}
#[test]
fn both_fields_override() {
let (cpu, mem) = resolve(&limits(Some("400%"), Some("8G")), HIVE_CPU, HIVE_MEM);
assert_eq!(cpu, "400%");
assert_eq!(mem, "8G");
}
/// The point of per-field fallback: overriding memory must not drag
/// CPU along with it.
#[test]
fn partial_override_keeps_other_field_on_hive_default() {
let (cpu, mem) = resolve(&limits(None, Some("8G")), HIVE_CPU, HIVE_MEM);
assert_eq!(cpu, "200%", "cpu should still track the hive default");
assert_eq!(mem, "8G");
let (cpu, mem) = resolve(&limits(Some("400%"), None), HIVE_CPU, HIVE_MEM);
assert_eq!(cpu, "400%");
assert_eq!(mem, "4G", "memory should still track the hive default");
}
#[test]
fn empty_entry_is_reported_empty() {
assert!(AgentLimits::default().is_empty());
assert!(!limits(None, Some("8G")).is_empty());
assert!(!limits(Some("400%"), None).is_empty());
}
const TRUNCATED: &str =
"{\n \"sock\": { \"cpu_quota\": \"50%\", \"memory_max\": \"1G\" },\n \"iris\": { \"mem";
fn corrupt_file() -> (tempfile::TempDir, PathBuf) {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join(RESOURCE_LIMITS_FILE);
std::fs::write(&path, TRUNCATED).expect("seed");
(dir, path)
}
#[test]
fn effective_of_a_corrupt_file_is_an_error() {
let (_dir, path) = corrupt_file();
let err = effective(&path, "sock", HIVE_CPU, HIVE_MEM)
.expect_err("a corrupt file must not resolve to hive defaults");
assert_eq!(err.kind(), std::io::ErrorKind::InvalidData);
}
#[test]
fn set_limits_leaves_a_corrupt_file_untouched() {
let (_dir, path) = corrupt_file();
let err = set_limits_at(&path, "ruth", &limits(Some("400%"), None))
.expect_err("a corrupt file must not be overwritten");
assert_eq!(err.kind(), std::io::ErrorKind::InvalidData);
assert_eq!(std::fs::read(&path).expect("read"), TRUNCATED.as_bytes());
}
#[test]
fn missing_file_resolves_to_hive_defaults() {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join(RESOURCE_LIMITS_FILE);
let (cpu, mem) = effective(&path, "sock", HIVE_CPU, HIVE_MEM).expect("missing file");
assert_eq!((cpu.as_str(), mem.as_str()), (HIVE_CPU, HIVE_MEM));
}
#[test]
fn set_limits_keeps_other_agents_and_leaves_no_temp_file() {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join(RESOURCE_LIMITS_FILE);
set_limits_at(&path, "sock", &limits(None, Some("8G"))).expect("set on missing file");
set_limits_at(&path, "iris", &limits(Some("50%"), None)).expect("set");
let (cpu, mem) = effective(&path, "sock", HIVE_CPU, HIVE_MEM).expect("read");
assert_eq!((cpu.as_str(), mem.as_str()), (HIVE_CPU, "8G"));
let (cpu, _) = effective(&path, "iris", HIVE_CPU, HIVE_MEM).expect("read");
assert_eq!(cpu, "50%");
let entries: Vec<_> = std::fs::read_dir(dir.path())
.expect("read_dir")
.map(|e| e.expect("entry").file_name())
.collect();
assert_eq!(
entries,
vec![std::ffi::OsString::from(RESOURCE_LIMITS_FILE)]
);
}
/// Absent fields must deserialize to `None`, not fail — an entry
/// written by an older version with only one field must still load.
#[test]
fn partial_entry_deserializes() {
let map: BTreeMap<String, AgentLimits> =
serde_json::from_str(r#"{"sock":{"memory_max":"8G"}}"#).expect("parses");
assert_eq!(map["sock"], limits(None, Some("8G")));
}
#[test]
fn empty_fields_are_not_serialized() {
let map = BTreeMap::from([("sock".to_owned(), limits(None, Some("8G")))]);
let text = serde_json::to_string(&map).expect("serializes");
assert_eq!(text, r#"{"sock":{"memory_max":"8G"}}"#);
}
#[test]
fn accepts_valid_cpu_quotas() {
for v in ["100%", "200%", "400%", "50%", "12.5%"] {
assert!(validate_cpu_quota(v).is_ok(), "{v} should be valid");
}
}
#[test]
fn rejects_invalid_cpu_quotas() {
for v in ["", "200", "%", "abc", "200%%", "-50%", "2e2%", "200 %"] {
assert!(validate_cpu_quota(v).is_err(), "{v} should be rejected");
}
}
#[test]
fn accepts_valid_memory_maxes() {
for v in [
"8G", "512M", "1024", "2T", "4096K", "50%", "infinity", "1.5G",
] {
assert!(validate_memory_max(v).is_ok(), "{v} should be valid");
}
}
/// mara hit this directly: "16GB" reads as an obviously
/// valid size to a human, but systemd's own parser takes only the
/// single-letter form — confirmed against a running systemd 260,
/// `MemoryMax=16GB` and `MemoryMax=16g` both fail unit activation,
/// `MemoryMax=16G` doesn't. Accept the friendly spellings and normalize
/// them to what systemd actually wants, rather than rejecting input a
/// human reasonably expects to work.
#[test]
fn friendly_size_spellings_normalize_to_systemds_own_form() {
for (input, want) in [
("8GB", "8G"),
("8Gb", "8G"),
("8gb", "8G"),
("8g", "8G"),
("512mb", "512M"),
("1048576b", "1048576B"),
("1.5gb", "1.5G"),
] {
assert_eq!(validate_memory_max(input).as_deref(), Ok(want));
}
}
#[test]
fn rejects_invalid_memory_maxes() {
for v in ["", "G", "abc", "-8G", "8 G", "Infinity", "8Gi", "8KG"] {
assert!(validate_memory_max(v).is_err(), "{v} should be rejected");
}
}
#[test]
fn parse_bytes_handles_plain_and_suffixed_values() {
assert_eq!(parse_bytes("1024"), Some(1024));
assert_eq!(parse_bytes("4G"), Some(4 * 1024 * 1024 * 1024));
assert_eq!(parse_bytes("512M"), Some(512 * 1024 * 1024));
assert_eq!(parse_bytes("4096K"), Some(4096 * 1024));
assert_eq!(parse_bytes("2T"), Some(2 * 1024 * 1024 * 1024 * 1024));
assert_eq!(
parse_bytes("1.5G"),
Some(1024 * 1024 * 1024 + 512 * 1024 * 1024)
);
// Lowercase suffixes accepted, matching `normalized_byte_size`.
assert_eq!(parse_bytes("4g"), Some(4 * 1024 * 1024 * 1024));
}
#[test]
fn parse_bytes_rejects_infinity_and_percentages() {
assert_eq!(parse_bytes("infinity"), None);
assert_eq!(parse_bytes("50%"), None);
}
#[test]
fn parse_bytes_rejects_garbage() {
for v in ["", "abc", "-8G", "8Gi"] {
assert_eq!(parse_bytes(v), None, "{v} should not parse");
}
}
#[test]
fn effective_memory_bytes_falls_back_to_hive_default() {
// Empty map, i.e. no per-agent override — exercises the
// "use the hive-wide default" arm end to end.
let empty = BTreeMap::new();
assert_eq!(
effective_memory_bytes_from(&empty, "nobody-configured-this-agent", "4G"),
Some(4 * 1024 * 1024 * 1024)
);
assert_eq!(
effective_memory_bytes_from(&empty, "nobody-configured-this-agent", "infinity"),
None
);
}
#[test]
fn effective_memory_bytes_from_prefers_per_agent_override() {
let map = BTreeMap::from([("sock".to_owned(), limits(None, Some("8G")))]);
assert_eq!(
effective_memory_bytes_from(&map, "sock", HIVE_MEM),
Some(8 * 1024 * 1024 * 1024)
);
// A different agent not in the map still falls back to the
// hive-wide default from the same loaded map (no re-read).
assert_eq!(
effective_memory_bytes_from(&map, "iris", HIVE_MEM),
Some(4 * 1024 * 1024 * 1024)
);
}
/// A percentage cap is real and resolvable in principle, but this
/// module has no host `MemTotal` to resolve it against — must
/// degrade to `None` (no derived heap ceiling) rather than silently
/// treating it as unbounded or guessing a number. `warn!` firing is
/// exercised for coverage but not asserted on (no tracing test
/// subscriber wired up here) — the `None` return is the contract.
#[test]
fn effective_memory_bytes_from_returns_none_for_percentage() {
let map = BTreeMap::from([("sock".to_owned(), limits(None, Some("50%")))]);
assert_eq!(effective_memory_bytes_from(&map, "sock", HIVE_MEM), None);
}
}