Watch
0
0
Fork
You've already forked hyperhive
0

hive-screen-mcp: bound grim/wtype and VNC calls; hive-c0re: make messages match the code

hive-screen-mcp ran `grim` / `wtype` through an unbounded
`Command::output()` and spoke RFB to neatvnc with no deadline, so a
wedged compositor or a VNC server that accepts and never speaks held the
agent's turn forever. Each subprocess now has a 30s limit and is killed
when it hits it; each RFB exchange has a 10s limit. Both come back to the
agent as the tool's text result, like every other failure in this crate.

hive-c0re operator-facing text that described behaviour the code lacks:
- `hivectl matrix reset-password` printed a "next: hivectl matrix
  create-user" hint that fails for every target (agents are refused, a
  non-agent hits M_USER_IN_USE). The line is gone.
- a failed `nixos-container update` appended the container's journal
  tail, read with `journalctl -M`. Since the job DAG, `update` only runs
  from the `Swap` node on a stopped container, so the read always came
  back empty. The helper is removed; the error still points at the
  build log.
- the matrix sweep comment in main.rs said it re-provisions agent token
  files; `ensure_all` creates no agent accounts or tokens.
- the knowledge-pull comments named a webhook caller that no longer
  exists and claimed a race was "fixed at its source".
- `handle_spawn`'s doc and the `mcp_sockets` module doc named callers of
  `register_agent` / a rollback that do not exist.

Refs #4723
This commit is contained in:
atlas 2026-09-26 23:28:09 +02:00
commit 19cc1b12e2
6 changed files with 150 additions and 80 deletions

View file

@ -935,48 +935,9 @@ async fn priv_run_inner(kind: &str, name: &str, node_id: Option<u64>) -> Result<
match result {
Ok(()) => Ok(()),
Err(e) => {
let journal = if kind == "update" {
container_journal_tail(&container).await
} else {
String::new()
};
match log_id {
Some(id) => bail!("{e:#}; see build log #{id}{journal}"),
None => bail!("{e:#}{journal}"),
}
}
}
}
/// On a failed `nixos-container update`, the stderr nixos-container
/// itself prints is often terse ("failed to reload container") — the
/// real reason (which unit failed `switch-to-configuration` during
/// the reload phase) lands in the *container's* own journal, not on
/// the host. Fetch the tail of it so a failed rebuild self-documents
/// the failing unit in the error string, no second round-trip.
///
/// Scoped to `update`: that's the reload-phase case, and the
/// container is still up (running the old generation) so
/// `journalctl -M` works. Best-effort — returns "" for other verbs
/// or when the journal can't be read (machine gone, journalctl
/// missing); it never produces an error of its own.
async fn container_journal_tail(container: &str) -> String {
// `-M` enters the container namespace and needs root, so the read
// is delegated to hive-priv (hive-c0re itself runs unprivileged).
let res = crate::priv_client::read_container_journal(
container,
hive_priv_sock::JournalQuery {
lines: 40,
..Default::default()
Err(e) => match log_id {
Some(id) => bail!("{e:#}; see build log #{id}"),
None => bail!("{e:#}"),
},
)
.await;
match res {
Ok((stdout, _)) if !stdout.is_empty() => format!(
"\n--- last 40 journal lines from container '{container}' ---\n{}",
stdout.trim_end()
),
_ => String::new(),
}
}