hive-screen-mcp: bound grim/wtype and VNC calls; hive-c0re: make messages match the code
hive-screen-mcp ran `grim` / `wtype` through an unbounded `Command::output()` and spoke RFB to neatvnc with no deadline, so a wedged compositor or a VNC server that accepts and never speaks held the agent's turn forever. Each subprocess now has a 30s limit and is killed when it hits it; each RFB exchange has a 10s limit. Both come back to the agent as the tool's text result, like every other failure in this crate. hive-c0re operator-facing text that described behaviour the code lacks: - `hivectl matrix reset-password` printed a "next: hivectl matrix create-user" hint that fails for every target (agents are refused, a non-agent hits M_USER_IN_USE). The line is gone. - a failed `nixos-container update` appended the container's journal tail, read with `journalctl -M`. Since the job DAG, `update` only runs from the `Swap` node on a stopped container, so the read always came back empty. The helper is removed; the error still points at the build log. - the matrix sweep comment in main.rs said it re-provisions agent token files; `ensure_all` creates no agent accounts or tokens. - the knowledge-pull comments named a webhook caller that no longer exists and claimed a race was "fixed at its source". - `handle_spawn`'s doc and the `mcp_sockets` module doc named callers of `register_agent` / a rollback that do not exist. Refs #4723
This commit is contained in:
parent
3385aaf026
commit
19cc1b12e2
6 changed files with 150 additions and 80 deletions
|
|
@ -935,48 +935,9 @@ async fn priv_run_inner(kind: &str, name: &str, node_id: Option<u64>) -> Result<
|
|||
|
||||
match result {
|
||||
Ok(()) => Ok(()),
|
||||
Err(e) => {
|
||||
let journal = if kind == "update" {
|
||||
container_journal_tail(&container).await
|
||||
} else {
|
||||
String::new()
|
||||
};
|
||||
match log_id {
|
||||
Some(id) => bail!("{e:#}; see build log #{id}{journal}"),
|
||||
None => bail!("{e:#}{journal}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// On a failed `nixos-container update`, the stderr nixos-container
|
||||
/// itself prints is often terse ("failed to reload container") — the
|
||||
/// real reason (which unit failed `switch-to-configuration` during
|
||||
/// the reload phase) lands in the *container's* own journal, not on
|
||||
/// the host. Fetch the tail of it so a failed rebuild self-documents
|
||||
/// the failing unit in the error string, no second round-trip.
|
||||
///
|
||||
/// Scoped to `update`: that's the reload-phase case, and the
|
||||
/// container is still up (running the old generation) so
|
||||
/// `journalctl -M` works. Best-effort — returns "" for other verbs
|
||||
/// or when the journal can't be read (machine gone, journalctl
|
||||
/// missing); it never produces an error of its own.
|
||||
async fn container_journal_tail(container: &str) -> String {
|
||||
// `-M` enters the container namespace and needs root, so the read
|
||||
// is delegated to hive-priv (hive-c0re itself runs unprivileged).
|
||||
let res = crate::priv_client::read_container_journal(
|
||||
container,
|
||||
hive_priv_sock::JournalQuery {
|
||||
lines: 40,
|
||||
..Default::default()
|
||||
Err(e) => match log_id {
|
||||
Some(id) => bail!("{e:#}; see build log #{id}"),
|
||||
None => bail!("{e:#}"),
|
||||
},
|
||||
)
|
||||
.await;
|
||||
match res {
|
||||
Ok((stdout, _)) if !stdout.is_empty() => format!(
|
||||
"\n--- last 40 journal lines from container '{container}' ---\n{}",
|
||||
stdout.trim_end()
|
||||
),
|
||||
_ => String::new(),
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -344,17 +344,10 @@ async fn cmd_serve(
|
|||
}
|
||||
}
|
||||
});
|
||||
// Knowledge periodic pull: hourly fallback in case the webhook is
|
||||
// missed (e.g. hive-c0re was down during a push). Deliberately does
|
||||
// NOT also fire an immediate pull at startup the way this task used
|
||||
// to: `auto_update::run`'s `NodeKind::KnowledgePull` DAG node (spawned
|
||||
// separately, a few lines up) already does that unconditionally on
|
||||
// every boot. The two used to run concurrently with no lock between
|
||||
// them, both `git pull --ff-only`-ing the same working tree — a real
|
||||
// race, and the likely root cause of the "local changes would be
|
||||
// overwritten" wedge this file's `pull()` now defends against
|
||||
// (`reset --hard` before every pull). Removing the redundant caller
|
||||
// fixes the race at its source instead of just self-healing after it.
|
||||
// Knowledge periodic pull: hourly fallback in case a knowledge event from
|
||||
// the swarm is missed (e.g. hive-c0re was down when it was sent). Sleeps
|
||||
// before its first pull: the boot `NodeKind::KnowledgePull` DAG node
|
||||
// already pulls once at startup.
|
||||
let mut knowledge_shutdown = coord.shutdown_rx();
|
||||
let knowledge_coord = coord.clone();
|
||||
tokio::spawn(async move {
|
||||
|
|
@ -394,15 +387,11 @@ async fn cmd_serve(
|
|||
}
|
||||
}
|
||||
});
|
||||
// Matrix user sweep: same shape — ensure every container has
|
||||
// an account on the local matrix-tuwunel homeserver with an
|
||||
// access_token persisted to `<state>/matrix-token`. No-op when
|
||||
// the hive-matrix container isn't running.
|
||||
//
|
||||
// Re-submitted every 30 minutes so that token files deleted by
|
||||
// `hive-matrix-daemon` (stale-token recovery — `M_UNKNOWN_TOKEN`)
|
||||
// get re-provisioned without requiring a hive-c0re restart. The
|
||||
// startup pass is the boot `MatrixSweep` node.
|
||||
// Matrix sweep (`matrix::ensure_all`): the hive's own account, the hive
|
||||
// Space and chat room, and every agent container's invite to both. It
|
||||
// creates no agent accounts or tokens — swarm-controller mints those.
|
||||
// No-op when the hive-matrix container isn't running. Re-submitted every
|
||||
// 30 minutes; the startup pass is the boot `MatrixSweep` node.
|
||||
let mut matrix_shutdown = coord.shutdown_rx();
|
||||
let matrix_coord = coord.clone();
|
||||
tokio::spawn(async move {
|
||||
|
|
|
|||
|
|
@ -287,8 +287,8 @@ async fn dispatch(req: &HostRequest, coord: Arc<Coordinator>) -> HostResponse {
|
|||
}
|
||||
}
|
||||
|
||||
/// Create + start the container for `name`, rolling back socket
|
||||
/// registration and notifying the manager on failure.
|
||||
/// Create + start the container for `name` and bind its MCP listener. On a
|
||||
/// failed spawn nothing was registered: post a swarm notice and return the error.
|
||||
async fn handle_spawn(coord: &Arc<Coordinator>, name: &str) -> Result<HostResponse> {
|
||||
tracing::info!(%name, "spawn");
|
||||
let agent_dir = crate::paths::agent_runtime_dir(name);
|
||||
|
|
@ -762,7 +762,6 @@ async fn handle_matrix_reset_password(name: &str) -> Result<HostResponse> {
|
|||
Ok(HostResponse::messages(vec![
|
||||
format!("matrix: password for @{name}:{server_name} reset"),
|
||||
format!("password persisted at: {}", pw_path.display()),
|
||||
format!("next: hivectl matrix create-user {name} # mints a fresh access token"),
|
||||
]))
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -111,9 +111,8 @@ pub async fn pull(coord: &Coordinator) -> Result<()> {
|
|||
// `LOCAL_DIR` after the initial clone — this working tree exists to mirror
|
||||
// `origin/main`, not to be edited in place. A tracked file left dirty
|
||||
// by any other means (a stray manual edit on the host, an interrupted
|
||||
// prior operation, or — the actual root cause here — two unsynchronized
|
||||
// boot-time pull callers racing on this same working tree, since fixed
|
||||
// in `main.rs`) would otherwise abort the `--ff-only` merge below with
|
||||
// prior operation, or two pulls running on this working tree at once)
|
||||
// would otherwise abort the `--ff-only` merge below with
|
||||
// "local changes would be overwritten"; an untracked file left behind
|
||||
// the same ways aborts it with "untracked working tree files would be
|
||||
// overwritten" instead — same wedge, just `clean`'s failure message
|
||||
|
|
|
|||
|
|
@ -6,8 +6,8 @@
|
|||
//! re-register all running agents.
|
||||
//!
|
||||
//! After startup, listeners are managed event-driven:
|
||||
//! - `run_create` calls `register_agent` eagerly on first-spawn.
|
||||
//! - `run_reconcile` calls `register_agent` immediately after `start_with_fallback`.
|
||||
//! - `run_start` (the job queue's `Start` node) and `server::handle_spawn`
|
||||
//! call `register_agent` once the container is started.
|
||||
//! - `kill`/`destroy` paths call `unregister_agent`.
|
||||
//!
|
||||
//! No recurring poll is needed because c0re owns the listener lifecycle. An
|
||||
|
|
|
|||
Loading…
Reference in a new issue