hyperhive/hive-c0re/src/workers/auto_update.rs
atlas 2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00

707 lines
30 KiB
Rust

//! Boot reconcile: on `hive-c0re serve` boot, (a) run the config path
//! for agents whose per-agent rev marker is stale — one `Boot` DAG
//! (meta hyperhive lock bump) that grows a `Rebuild` subgraph for each
//! stale agent whose `wanted` power intent is `Up` — and (b)
//! converge every other drifted agent to its persisted `wanted` via
//! `Reconcile` DAGs. Two rules keep boot-time nix work minimal:
//!
//! 1. **Stale but wanted-offline agents** get no rebuild — their
//! rebuild happens the first time they're started (the start
//! submit path upgrades a stale start to rebuild+start). The sweep
//! parent still runs whenever *any* marker is stale so the meta
//! hyperhive lock is bumped for those later start-upgrades.
//! 2. **Agents whose rev marker matches** the current hyperhive flake
//! path are skipped — nothing changed, no nix work to do.
//!
//! Booting with no config change performs no meta commit — only
//! reconciles. See `docs/scheduler/coordinator.md::Boot reconcile`.
use std::path::Path;
use std::sync::Arc;
use anyhow::Result;
use crate::coordinator::Coordinator;
use crate::lifecycle::{self, AGENT_PREFIX, MANAGER_NAME};
use crate::tool_groups;
/// Resolve the current rev of `hyperhive_flake`. For a path on disk we
/// canonicalize (following symlinks) so a /etc/hyperhive → /nix/store/...
/// update yields a different string. For anything else we return None.
#[must_use]
pub fn current_flake_rev(hyperhive_flake: &str) -> Option<String> {
let path = Path::new(hyperhive_flake);
if !path.exists() {
return None;
}
std::fs::canonicalize(path)
.ok()
.map(|p| p.display().to_string())
}
/// Returns true when the applied repo has commits that have not yet been
/// deployed (i.e. the applied HEAD differs from the sha currently locked in
/// meta's flake.lock). This is the semantic the dashboard `needs_update` chip
/// conveys: "there is a config change ready to apply via rebuild."
///
/// Async on purpose: this runs per agent inside `container_view::build_all`,
/// which fires on the ~10s dashboard sweep, every `AgentStatus` request, and
/// every `rescan_containers_and_emit` after a lifecycle step. A synchronous
/// `git` fork here blocks a tokio worker for the whole exec — under
/// nix-build disk saturation that's long enough that concurrent sweeps
/// starved the runtime and stalled the per-agent sockets.
pub async fn agent_config_pending(name: &str, deployed_sha: Option<&str>) -> bool {
let applied = crate::paths::applied_dir(name);
let applied_head = tokio::process::Command::new("git")
.args(["-C", &applied.to_string_lossy(), "rev-parse", "HEAD"])
.output()
.await
.ok()
.filter(|o| o.status.success())
.and_then(|o| String::from_utf8(o.stdout).ok())
.map(|s| s.trim().to_owned());
match (applied_head.as_deref(), deployed_sha) {
(Some(head), Some(sha)) => !head.starts_with(sha) && !sha.starts_with(head),
_ => false,
}
}
/// Whether this hive is "ruthless" — running with no root/manager agent at
/// all (no ruth). When true, hive-c0re skips the root-agent create/start
/// sweep entirely. Controlled by the host option
/// `services.hyperhive.ruthless`, threaded in via the `HYPERHIVE_RUTHLESS`
/// env var. Defaults to `false` when the var is unset (back-compat: the
/// root agent was always auto-managed before this opt-out existed); only
/// an explicit `true` / `1` / `yes` enables ruthless mode.
fn ruthless() -> bool {
match std::env::var("HYPERHIVE_RUTHLESS") {
Ok(v) => matches!(v.trim().to_ascii_lowercase().as_str(), "true" | "1" | "yes"),
Err(_) => false,
}
}
/// What `ensure_root_agent` does with a `lifecycle::list()` result, before
/// it even gets to check whether the manager is present.
#[derive(Debug, PartialEq, Eq)]
enum RootAgentPlan {
/// The manager is in this readable list.
Present,
/// The manager is absent from this readable list: spawn it.
Absent,
/// The list could not be read. That is not "manager absent" — it
/// proves nothing either way, so the auto-spawn decision is skipped
/// for this attempt rather than risking a spawn over a manager that's
/// actually there.
Unreadable,
}
/// Turn a `lifecycle::list()` result into `ensure_root_agent`'s plan.
fn plan_root_agent(list_result: anyhow::Result<Vec<String>>) -> RootAgentPlan {
match list_result {
Ok(names)
if names
.iter()
.any(|c| c.strip_prefix(AGENT_PREFIX) == Some(MANAGER_NAME)) =>
{
RootAgentPlan::Present
}
Ok(_) => RootAgentPlan::Absent,
Err(_) => RootAgentPlan::Unreadable,
}
}
/// Auto-create the manager container on startup if it isn't already there.
/// hive-c0re manages the manager end-to-end: operators no longer declare
/// `containers.h-ruth` in their host NixOS config. Bypasses the approval
/// queue — the root/manager is auto-managed by default. Operators who
/// don't want a root agent at all set `services.hyperhive.ruthless = true`,
/// which short-circuits this whole function. Idempotent.
pub async fn ensure_root_agent(coord: &Arc<Coordinator>) -> Result<()> {
if ruthless() {
tracing::info!(
"ruthless mode (services.hyperhive.ruthless = true) - skipping root agent create/start"
);
return Ok(());
}
// Before the create/start branch, not inside it: on a hive whose root
// container already exists this is the only run that can still seed the
// grant, and it has to land before her next rebuild bakes the binds.
seed_manager_capabilities();
let list_result = lifecycle::list().await;
if let Err(e) = &list_result {
tracing::warn!(
error = ?e,
"manager container list unreadable — skipping auto-spawn check for this attempt"
);
}
let plan = plan_root_agent(list_result);
let current_rev = current_flake_rev(&coord.hyperhive_flake);
if plan == RootAgentPlan::Unreadable {
// An unreadable list is not "manager absent" — spawning on it would
// both waste `provision_container`'s work and hit `nixos-container
// create`'s "already exists" failure if the manager is actually
// there. Skip the decision this attempt; there is no periodic
// retry for this boot-time call, so the next chance is the next
// `hive-c0re` restart.
return Ok(());
}
if plan == RootAgentPlan::Present {
// Container exists already. If it predates the unified lifecycle
// (no applied flake on disk) we must rebuild — otherwise it's
// running whatever the host-declarative config was at create
// time, with a wrong systemd unit and port.
let applied_flake = crate::paths::applied_dir(MANAGER_NAME).join("flake.nix");
if !applied_flake.exists() && current_rev.is_some() {
tracing::warn!(
"manager container exists but no applied flake — forcing rebuild to migrate"
);
if let Err(e) = coord.job_queue.insert_job(|b| {
crate::job_queue::templates::rebuild(b, MANAGER_NAME, true);
Vec::new()
}) {
tracing::warn!(error = ?e, "manager migration rebuild insert failed");
}
} else {
tracing::debug!("manager container already present");
}
// hive-c0re auto-manages the root/manager by default, so a
// present-but-stopped root (e.g. a first-start failure on a fresh
// install) is brought back up here: the startup sweep's rebuild only
// restarts a container that was already running, so without this it
// stays down until a manual `nixos-container start`. The sub-agent
// `was_running` guard is intentionally left untouched. (Operators
// opt out of this whole auto-management with
// `services.hyperhive.ruthless = true`, gated at the top of
// this function.)
if !lifecycle::is_running(MANAGER_NAME).await {
tracing::info!("manager container present but not running — starting");
if let Err(e) = coord.power.set(MANAGER_NAME, crate::power::Wanted::Up) {
tracing::warn!(error = ?e, "agent_power: set manager wanted=up failed");
}
// Through the preamble: after a reboot its bind sources and limits
// drop-in in `/run` are gone, and nothing else recreates them.
let hive = coord.hive_env();
let started = async {
let paths = Coordinator::agent_paths(
MANAGER_NAME,
crate::paths::agent_runtime_dir(MANAGER_NAME),
)?;
let token = lifecycle::converge_start_preamble(MANAGER_NAME, &hive, &paths).await?;
lifecycle::start_with_fallback(token).await
}
.await;
if let Err(e) = started {
tracing::warn!(error = ?e, "manager start failed");
}
}
return Ok(());
}
tracing::info!("manager container missing — spawning");
// lifecycle::spawn creates the runtime dir internally; no manual
// ensure_agent_runtime_dir needed here.
let runtime = crate::paths::agent_runtime_dir(MANAGER_NAME);
let hive = coord.hive_env();
let paths = Coordinator::agent_paths(MANAGER_NAME, runtime)?;
lifecycle::spawn(MANAGER_NAME, &hive, &paths).await?;
seed_manager_tool_groups();
if let Err(e) = coord.power.set(MANAGER_NAME, crate::power::Wanted::Up) {
tracing::warn!(error = ?e, "agent_power: set manager wanted=up failed");
}
if let Some(rev) = current_rev {
let _ = std::fs::write(crate::paths::applied_rev_marker(MANAGER_NAME), &rev);
}
Ok(())
}
/// Give ruth her privileged tool groups on the one path that creates her.
///
/// `effective_tool_groups()` has no manager-flavour fallback, so an agent
/// with no entry in `tool-groups.json` is an agent with no privileged
/// tools. Ruth needs hers from her first turn, and this is the only place
/// she is brought into existence — so it is written once, here, rather
/// than re-checked on every hive-c0re boot.
///
/// Skips a name that already has an entry: a destroy+recreate under the
/// same name must not silently reset an operator's chosen group set back
/// to the default. An unreadable file is logged and left alone: ruth has
/// already been spawned, and seeding stays best-effort like the write
/// below.
fn seed_manager_tool_groups() {
match tool_groups::groups_for(MANAGER_NAME) {
Ok(groups) if groups.is_empty() => {}
Ok(_) => {
tracing::debug!("manager tool groups already set — leaving as-is");
return;
}
Err(e) => {
tracing::warn!(
error = ?e,
"tool-groups file unreadable — not seeding ruth's tool groups"
);
return;
}
}
let all_groups: Vec<String> = hive_sh4re::permissions::ToolGroup::MANAGER_DEFAULT
.iter()
.map(|g| <&str>::from(*g).to_owned())
.collect();
match tool_groups::set_groups(MANAGER_NAME, &all_groups) {
Ok(()) => tracing::info!("seeded ruth's tool groups to MANAGER_DEFAULT (all groups)"),
Err(e) => tracing::warn!(
error = ?e,
"failed to seed ruth's tool groups — she will start without privileged tools"
),
}
}
/// The capability the root agent is seeded with, as the `snake_case` string
/// `capabilities.json` stores and `capabilities::has_cap` compares against.
/// Taken from the enum rather than written out, so the seed cannot drift into
/// a name `capabilities::prune_unknown` would silently drop.
fn manager_seed_caps() -> Vec<String> {
vec![<&str>::from(hive_sh4re::permissions::Capability::ManageRootAgent).to_owned()]
}
/// Whether the root agent's default capability grant should be written.
///
/// `store_written` is "`capabilities.json` exists" — see
/// [`seed_manager_capabilities`] for why that, and not the absence of an
/// entry for the manager, is the condition.
fn should_seed_manager_caps(store_written: bool) -> bool {
!store_written
}
/// Give ruth the `ManageRootAgent` capability on the path that deploys her.
///
/// This is the capability-store half of the grant `roles.json` used to make:
/// `topology::reconcile_roles` seeded `can_manage_top_level_agents` onto
/// `MANAGER_NAME` on every meta sync, and that role is what put the other
/// agents' state/config dirs, `/applied` and `/meta` into her nspawn binds.
/// Collapsing the role into the capability means that without a seed here she
/// loses those recovery mounts at her next container rebuild — silently, and
/// only then, because nspawn bakes bind flags at container start.
///
/// **Seeded once, not re-ensured on every boot.** The role kept an empty-list
/// tombstone so an explicit revoke stuck; the capability store *deletes* an
/// entry that has been emptied rather than tombstoning it, so "the manager
/// has no entry" cannot tell a fresh hive apart from a deliberate revoke.
/// File existence can: every grant and revoke goes through
/// `meta::commit_capabilities` → `capabilities::set_caps`, which
/// writes the file even when the result is an empty `{}`. So while the file
/// is absent nobody has ever had a say, and once it exists this is inert
/// forever — including on the destroy+recreate path, matching
/// [`seed_manager_tool_groups`]'s refusal to reset an operator's choice.
///
/// Written with plain `set_caps` rather than `meta::commit_capabilities`: the
/// role's own seed wrote the file and left it for `stage_generated_meta_files`
/// to `git add` into the next deploy commit, and taking `META_LOCK` on the
/// boot path would be a new ordering constraint for no gain.
fn seed_manager_capabilities() {
if !should_seed_manager_caps(crate::capabilities::capabilities_path().exists()) {
tracing::debug!(
"capabilities.json already written — leaving the manager's capabilities as-is"
);
return;
}
let caps = manager_seed_caps();
match crate::capabilities::set_caps(MANAGER_NAME, &caps) {
Ok(()) => tracing::info!(
caps = ?caps,
"seeded ruth's default capabilities (recovery bind mounts for every agent)"
),
Err(e) => tracing::warn!(
error = ?e,
"failed to seed ruth's capabilities — she will come up without the recovery mounts"
),
}
}
/// Boot reconcile (see the module doc): classify every agent by rev
/// freshness + persisted `wanted` intent, submit one `Boot` DAG
/// (hyperhive lock bump growing an in-DAG rebuild subgraph per stale
/// wanted-up agent) when anything is stale, and `Reconcile` DAGs for
/// agents whose observed power state drifted from `wanted`. Returns Ok even
/// if some submissions failed.
pub async fn run(coord: Arc<Coordinator>) -> Result<()> {
let containers = match lifecycle::list().await {
Ok(c) => c,
Err(e) => {
tracing::warn!(error = ?e, "boot reconcile: nixos-container list failed");
return Ok(());
}
};
let current_rev = current_flake_rev(&coord.hyperhive_flake);
// Resolve container names to logical agent names, then sort. The
// parent field this used to depth-sort by is gone; with no
// hierarchy left to respect, alphabetical is the whole order — and it
// is exactly what the depth sort already produced once every agent
// was a root.
let mut logical_names: Vec<String> = containers
.iter()
.filter_map(|c| c.strip_prefix(AGENT_PREFIX).map(str::to_owned))
.collect();
logical_names.sort();
// Classify. `get_or_seed` doubles as the one-time migration: an
// agent without an `agent_power` row is seeded from its observed
// state (running ⇒ Up), after which the DB is authoritative.
let mut any_stale = false;
// stale ∧ wanted=Up → sweep rebuild. The bool is the agent's observed
// running state, used to order running agents first before submit.
let mut fanout: Vec<(String, bool)> = Vec::new();
let mut drifted: Vec<String> = Vec::new(); // fresh ∧ wanted≠observed → reconcile
let mut n_deferred = 0usize;
let mut n_skipped = 0usize;
for name in &logical_names {
let running = lifecycle::is_running(name).await;
let wanted = match coord.power.get_or_seed(name, running) {
Ok(w) => Some(w),
Err(e) => {
tracing::warn!(%name, error = ?e, "agent_power read failed — reconciling so it surfaces");
None
}
};
let fresh = current_rev.as_ref().is_some_and(|rev| {
std::fs::read_to_string(crate::paths::applied_rev_marker(name))
.is_ok_and(|stored| stored == rev.as_str())
});
match boot_action(wanted, fresh, running) {
BootAction::Rebuild => {
any_stale = true;
fanout.push((name.clone(), running));
}
BootAction::Defer { reconcile } => {
any_stale = true;
n_deferred += 1;
tracing::debug!(%name, "boot reconcile: stale but offline — deferring rebuild to on-start");
if reconcile {
drifted.push(name.clone());
}
}
BootAction::Skip { reconcile } => {
n_skipped += 1;
if reconcile {
drifted.push(name.clone());
}
}
BootAction::Unreadable => drifted.push(name.clone()),
}
}
tracing::info!(
total = containers.len(),
rebuilds = fanout.len(),
reconciles = drifted.len(),
deferred = n_deferred,
up_to_date = n_skipped,
"boot reconcile"
);
// Rebuild running agents first. All fanout entries are wanted=Up;
// among them, warm the live/serving agents onto the fresh config before the
// stopped-but-wanted-up ones so the scarce build slots hit uptime-critical
// agents first. Stable sort keeps the alphabetical order within each
// running/stopped group. The `drifted` reconciles aren't sorted
// — they hold no build slot and run concurrently, so their order is moot.
fanout.sort_by_key(|(_, running)| !running);
let fanout: Vec<String> = fanout.into_iter().map(|(name, _)| name).collect();
submit_boot_tree(&coord, any_stale, fanout, drifted, n_deferred, n_skipped);
submit_startup_sweep_nodes(&coord);
Ok(())
}
/// What the boot sweep does with one agent.
#[derive(Debug, PartialEq, Eq)]
enum BootAction {
/// Stale ∧ wanted up: rebuild against the post-bump lock. The DAG's
/// tail `Reconcile` brings the agent (back) up, covering both the
/// running-stale and stopped-but-wanted-up cases.
Rebuild,
/// Stale ∧ wanted offline: no boot-time nix work — the start submit
/// path upgrades a stale start to a rebuild.
Defer { reconcile: bool },
/// Already on the current rev.
Skip { reconcile: bool },
/// The intent could not be read.
Unreadable,
}
/// `wanted` is `None` when the store could not answer for this agent.
///
/// That case has no neutral stand-in, which is why it is a variant rather
/// than a fallback value: `Wanted::from_running(running)` is precisely the
/// value for which `reconcile_action` returns `Noop`, so guessing it here
/// puts the agent beyond every check below — a corrupt row then leaves one
/// `warn!` per boot and no other trace. The queued reconcile's own
/// `get_or_seed` fails as a per-agent node instead.
fn boot_action(wanted: Option<crate::power::Wanted>, fresh: bool, running: bool) -> BootAction {
let Some(wanted) = wanted else {
return BootAction::Unreadable;
};
let reconcile =
crate::power::reconcile_action(wanted, running) != crate::power::ReconcileAction::Noop;
if fresh {
BootAction::Skip { reconcile }
} else if wanted == crate::power::Wanted::Up {
BootAction::Rebuild
} else {
BootAction::Defer { reconcile }
}
}
/// Submit the boot-time forge/matrix/webhook/knowledge/wanted-state sweeps as
/// DAG nodes — `ForgeSweep`, `MatrixSweep`, `WebhookRegister`,
/// `KnowledgePull`, `WantedPull`. Unlike
/// [`submit_boot_tree`] this runs on **every** boot, quiet or not: these
/// aren't config-drift work, they're startup housekeeping that always needs
/// to happen, and the point of moving them here is exactly so they show up
/// as real work on the dashboard instead of an invisible `tokio::spawn` that
/// only surfaces on failure. Five independent, build-slot- and lease-exempt
/// roots — no dependency edges between them, matching the existing
/// `Reconcile`-root pattern in [`boot_nodes`].
fn submit_startup_sweep_nodes(coord: &Arc<Coordinator>) {
use crate::job_queue::NodeKind;
if let Err(e) = coord.job_queue.insert_job(|b| {
let _ = b.node(NodeKind::ForgeSweep);
let _ = b.node(NodeKind::MatrixSweep);
let _ = b.node(NodeKind::WebhookRegister);
let _ = b.node(NodeKind::KnowledgePull);
let _ = b.node(NodeKind::WantedPull);
Vec::new()
}) {
tracing::warn!(error = ?e, "boot: startup sweep DAG insert failed");
}
}
/// The boot DAG's node declarations, split out of [`submit_boot_tree`] so they
/// can be exercised without a live [`Coordinator`].
///
/// That split is not cosmetic: this path constructs nodes outside
/// `job_queue/`, so it is the one place a resource declaration can be forgotten
/// without any in-module test noticing. It has happened once already — the
/// sweep `MetaLock` and the boot `Reconcile`s silently declared nothing when
/// kind-derived resources were removed, which drops the agent lease a boot
/// reconcile needs to not race another DAG's container ops.
pub(crate) fn boot_nodes(
builder: &crate::job_queue::JobBuilder,
any_stale: bool,
fanout: Vec<String>,
drifted: Vec<String>,
) {
use crate::job_queue::NodeKind;
use crate::job_queue::resource::Resource;
// Sweep whenever ANY marker is stale — even when every stale agent is
// wanted-offline: the hyperhive lock bump must land now so their later
// start-upgrade rebuilds build against it. No stale agents ⇒ no MetaLock
// ⇒ no meta commit on a no-change boot. The `fanout` list rides the
// MetaLock into `run_meta_lock`, which appends the rebuild subgraphs.
if any_stale {
let _ = builder
.node(NodeKind::MetaLock {
sweep: true,
fanout: Some(fanout),
// A sweep bumps `hyperhive` alone (`lock_update_hyperhive`),
// so it names no inputs.
inputs: Vec::new(),
})
.needs(Resource::BuildSlot)
.needs(Resource::MetaWindow);
}
// One boot Reconcile per drifted agent — independent roots.
for name in drifted {
// Name the lease before the agent string moves into the kind.
let lease = Resource::Agent(name.clone());
let _ = builder
.node(NodeKind::Reconcile { agent: name })
.needs(lease);
}
}
/// Submit this boot's work as **one DAG** (no anchor node, no per-agent
/// child DAGs). Node 0 is the sweep `MetaLock` (only when
/// something is stale) — its executor bumps the hyperhive lock, then grows
/// one rebuild subgraph per stale agent into *this same* DAG (rooted on the
/// `MetaLock`, so they build against the post-bump lock; see
/// `exec::run_meta_lock`). Every drifted agent gets a boot `Reconcile` as an
/// independent root — a boot reconcile needs no lock bump, so it converges
/// concurrently with the sweep. No-op when there's nothing to do.
fn submit_boot_tree(
coord: &Arc<Coordinator>,
any_stale: bool,
fanout: Vec<String>,
drifted: Vec<String>,
n_deferred: usize,
n_skipped: usize,
) {
// Fully-quiet boot (nothing stale, nothing drifted) inserts nothing.
if !any_stale && drifted.is_empty() {
return;
}
// The summary the sweep used to hand the container as its `reason` is a log
// line now: it was only ever stored on a node nobody read, and the counts
// are worth having where they can actually be seen.
tracing::info!(
rebuilds = fanout.len(),
reconciles = drifted.len(),
deferred = n_deferred,
up_to_date = n_skipped,
"boot: sweep"
);
// The sweep's own rebuild subgraphs emit their `Rebuilt` events as they
// land; the boot DAG as a whole has no terminal side effect, so no tail.
// The subgraphs also carry their own per-agent crash-watch suppression
// during their `Swap` (applied at claim time); a reconcile-only boot needs
// no transient.
if let Err(e) = coord.job_queue.insert_job(|b| {
boot_nodes(b, any_stale, fanout, drifted);
Vec::new()
}) {
tracing::warn!(error = ?e, "boot: sweep DAG insert failed");
}
coord.emit_rebuild_queue_snapshot();
}
#[cfg(test)]
mod tests {
use super::{
BootAction, MANAGER_NAME, RootAgentPlan, boot_action, manager_seed_caps, plan_root_agent,
should_seed_manager_caps,
};
use crate::lifecycle::AGENT_PREFIX;
use crate::power::Wanted;
// -----------------------------------------------------------------------
// `ensure_root_agent`'s spawn-or-not decision. An unreadable list must
// never be read as "manager absent" — see `plan_root_agent`'s doc.
/// The regression this fix exists for: a read failure must not spawn.
#[test]
fn unreadable_list_skips_without_spawning() {
let result: anyhow::Result<Vec<String>> =
Err(anyhow::anyhow!("connect to hive-priv socket"));
assert_eq!(plan_root_agent(result), RootAgentPlan::Unreadable);
}
/// Control: a genuinely empty, readable list still plans a spawn.
#[test]
fn readable_list_missing_manager_spawns() {
let result: anyhow::Result<Vec<String>> = Ok(Vec::new());
assert_eq!(plan_root_agent(result), RootAgentPlan::Absent);
}
#[test]
fn readable_list_with_manager_present_is_a_noop() {
let result: anyhow::Result<Vec<String>> = Ok(vec![format!("{AGENT_PREFIX}{MANAGER_NAME}")]);
assert_eq!(plan_root_agent(result), RootAgentPlan::Present);
}
// -----------------------------------------------------------------------
// Root-agent capability seed. `capabilities_path()` resolves under
// `paths::meta_root()`, which is hardcoded to `/var/lib/hyperhive` with no
// test override, so — same as `capabilities.rs`'s own tests — these pin
// the pure decision (`should_seed_manager_caps`, `manager_seed_caps`)
// rather than round-tripping the real file. Together they are the
// replacement for the deleted `reconcile_roles_in_seeds_root_when_absent`
// and `reconcile_roles_in_does_not_reseed_after_explicit_revoke`.
#[test]
fn the_manager_is_seeded_while_the_capability_store_has_never_been_written() {
assert!(should_seed_manager_caps(false));
assert_eq!(manager_seed_caps(), vec!["manage_root_agent".to_owned()]);
}
/// The revoke-sticks half, and the reason the condition is the file rather
/// than the manager's entry: the store deletes an emptied entry instead of
/// keeping a tombstone, but a revoke still writes the file — so once it
/// exists the seed must never fire again.
#[test]
fn a_written_capability_store_is_never_reseeded() {
assert!(!should_seed_manager_caps(true));
}
/// A seed the store would drop as unrecognised grants nothing while
/// logging success, so pin that what we write is a name `prune_unknown`
/// keeps — i.e. one of `Capability::ALL`'s `snake_case` spellings.
#[test]
fn the_seeded_name_is_a_capability_the_store_recognises() {
for cap in manager_seed_caps() {
assert!(
hive_sh4re::permissions::Capability::ALL
.iter()
.any(|known| <&str>::from(*known) == cap),
"{cap} is not a known capability"
);
}
}
// The regression the `Unreadable` variant exists for: the arm used to
// substitute `from_running(running)`, which is `Noop` against BOTH
// observations — so no combination of inputs could reach a reconcile.
#[test]
fn an_unreadable_intent_is_never_settled_whatever_the_agent_is_doing() {
for fresh in [true, false] {
for running in [true, false] {
assert_eq!(
boot_action(None, fresh, running),
BootAction::Unreadable,
"fresh={fresh} running={running}"
);
}
}
}
#[test]
fn a_fresh_agent_doing_what_it_should_is_skipped_without_a_reconcile() {
assert_eq!(
boot_action(Some(Wanted::Up), true, true),
BootAction::Skip { reconcile: false }
);
assert_eq!(
boot_action(Some(Wanted::Offline), true, false),
BootAction::Skip { reconcile: false }
);
}
#[test]
fn a_fresh_agent_that_drifted_from_its_intent_is_reconciled() {
assert_eq!(
boot_action(Some(Wanted::Up), true, false),
BootAction::Skip { reconcile: true }
);
assert_eq!(
boot_action(Some(Wanted::Offline), true, true),
BootAction::Skip { reconcile: true }
);
}
#[test]
fn a_stale_agent_wanted_up_is_rebuilt_running_or_not() {
assert_eq!(
boot_action(Some(Wanted::Up), false, true),
BootAction::Rebuild
);
assert_eq!(
boot_action(Some(Wanted::Up), false, false),
BootAction::Rebuild
);
}
#[test]
fn a_stale_agent_wanted_offline_defers_and_still_reports_drift() {
assert_eq!(
boot_action(Some(Wanted::Offline), false, false),
BootAction::Defer { reconcile: false }
);
assert_eq!(
boot_action(Some(Wanted::Offline), false, true),
BootAction::Defer { reconcile: true }
);
}
}