Watch
0
0
Fork
You've already forked hyperhive
0

hive-screen-mcp: bound grim/wtype and VNC calls; hive-c0re: make messages match the code

hive-screen-mcp ran `grim` / `wtype` through an unbounded
`Command::output()` and spoke RFB to neatvnc with no deadline, so a
wedged compositor or a VNC server that accepts and never speaks held the
agent's turn forever. Each subprocess now has a 30s limit and is killed
when it hits it; each RFB exchange has a 10s limit. Both come back to the
agent as the tool's text result, like every other failure in this crate.

hive-c0re operator-facing text that described behaviour the code lacks:
- `hivectl matrix reset-password` printed a "next: hivectl matrix
  create-user" hint that fails for every target (agents are refused, a
  non-agent hits M_USER_IN_USE). The line is gone.
- a failed `nixos-container update` appended the container's journal
  tail, read with `journalctl -M`. Since the job DAG, `update` only runs
  from the `Swap` node on a stopped container, so the read always came
  back empty. The helper is removed; the error still points at the
  build log.
- the matrix sweep comment in main.rs said it re-provisions agent token
  files; `ensure_all` creates no agent accounts or tokens.
- the knowledge-pull comments named a webhook caller that no longer
  exists and claimed a race was "fixed at its source".
- `handle_spawn`'s doc and the `mcp_sockets` module doc named callers of
  `register_agent` / a rollback that do not exist.

Refs #4723
This commit is contained in:
atlas 2026-09-26 23:28:09 +02:00
commit 19cc1b12e2
6 changed files with 150 additions and 80 deletions

View file

@ -344,17 +344,10 @@ async fn cmd_serve(
}
}
});
// Knowledge periodic pull: hourly fallback in case the webhook is
// missed (e.g. hive-c0re was down during a push). Deliberately does
// NOT also fire an immediate pull at startup the way this task used
// to: `auto_update::run`'s `NodeKind::KnowledgePull` DAG node (spawned
// separately, a few lines up) already does that unconditionally on
// every boot. The two used to run concurrently with no lock between
// them, both `git pull --ff-only`-ing the same working tree — a real
// race, and the likely root cause of the "local changes would be
// overwritten" wedge this file's `pull()` now defends against
// (`reset --hard` before every pull). Removing the redundant caller
// fixes the race at its source instead of just self-healing after it.
// Knowledge periodic pull: hourly fallback in case a knowledge event from
// the swarm is missed (e.g. hive-c0re was down when it was sent). Sleeps
// before its first pull: the boot `NodeKind::KnowledgePull` DAG node
// already pulls once at startup.
let mut knowledge_shutdown = coord.shutdown_rx();
let knowledge_coord = coord.clone();
tokio::spawn(async move {
@ -394,15 +387,11 @@ async fn cmd_serve(
}
}
});
// Matrix user sweep: same shape — ensure every container has
// an account on the local matrix-tuwunel homeserver with an
// access_token persisted to `<state>/matrix-token`. No-op when
// the hive-matrix container isn't running.
//
// Re-submitted every 30 minutes so that token files deleted by
// `hive-matrix-daemon` (stale-token recovery — `M_UNKNOWN_TOKEN`)
// get re-provisioned without requiring a hive-c0re restart. The
// startup pass is the boot `MatrixSweep` node.
// Matrix sweep (`matrix::ensure_all`): the hive's own account, the hive
// Space and chat room, and every agent container's invite to both. It
// creates no agent accounts or tokens — swarm-controller mints those.
// No-op when the hive-matrix container isn't running. Re-submitted every
// 30 minutes; the startup pass is the boot `MatrixSweep` node.
let mut matrix_shutdown = coord.shutdown_rx();
let matrix_coord = coord.clone();
tokio::spawn(async move {