diff --git a/docs/conventions.md b/docs/conventions.md index 1a0d0283..9bd03501 100644 --- a/docs/conventions.md +++ b/docs/conventions.md @@ -121,7 +121,7 @@ in `hive-c0re::dashboard_events::DashboardEvent`. agent. Always returns a list (`Messages { messages }`) — empty when nothing's pending, single-pop when `max = None` (default 1, the single-message behaviour), batched up to `max` when caller asks for -more (server-side cap is 5; values above clamp silently). +more (server-side cap is 32; values above clamp silently). `wait_seconds` long-polls for the first message; once one arrives — or one is already pending — the call drains up to `max` in total before returning, so a single `Recv` call coalesces a burst. diff --git a/docs/turn-loop.md b/docs/turn-loop.md index e3a7db7f..7818e4cd 100644 --- a/docs/turn-loop.md +++ b/docs/turn-loop.md @@ -520,7 +520,7 @@ ttl_seconds?, to?)`, `answer(id, answer)`, `ack_until(up_to)`. - `recv` — drain inbox. Without `wait_seconds` (or `0`) returns immediately. Positive value parks the turn up to that many seconds (cap 180) — incoming messages wake instantly. `max` (default 1, cap - 5) drains up to N rows; `wait_seconds` applies to the first, then + 32) drains up to N rows; `wait_seconds` applies to the first, then drains up to `max` total. Each returned row is prefixed with `[msg #]` (broker row id; note the highest id seen, then pass it to `ack_until` to bulk-triage the batch). **Graceful shutdown**: when the harness diff --git a/hive-ag3nt/prompts/system.md b/hive-ag3nt/prompts/system.md index df06b634..7f0a19db 100644 --- a/hive-ag3nt/prompts/system.md +++ b/hive-ag3nt/prompts/system.md @@ -2,7 +2,7 @@ You are hyperhive agent `{label}` (qualified: `{qualified_label}`){hive_identity Tools (hyperhive surface): -- `mcp__hyperhive__recv(wait_seconds?, max?)` — drain inbox messages (returns `(empty)` if nothing pending). Without `wait_seconds` (or with `0`) it returns immediately — a cheap "anything pending?" peek you can sprinkle between tool calls. To **wait** for work when you have nothing else useful to do this turn, call with a long wait (e.g. `wait_seconds: 180`, the max) — incoming messages wake you instantly, otherwise the call returns empty at the timeout. That's strictly better than a fixed `sleep` shell command: lower latency on new work, no busy-loop. `max` (default 1, cap 5) drains several queued messages in one call — the wake prompt tells you the pending count. +- `mcp__hyperhive__recv(wait_seconds?, max?)` — drain inbox messages (returns `(empty)` if nothing pending). Without `wait_seconds` (or with `0`) it returns immediately — a cheap "anything pending?" peek you can sprinkle between tool calls. To **wait** for work when you have nothing else useful to do this turn, call with a long wait (e.g. `wait_seconds: 180`, the max) — incoming messages wake you instantly, otherwise the call returns empty at the timeout. That's strictly better than a fixed `sleep` shell command: lower latency on new work, no busy-loop. `max` (default 1, cap 32) drains several queued messages in one call — the wake prompt tells you the pending count. - `mcp__hyperhive__ack_until(up_to)` — bulk-mark inbox messages handled: every message with broker id `<= up_to` (ids show as `[msg #]` in wake prompts and recv output) is acked in one call, pending and delivered alike. Use it to clear a backlog you've already triaged — e.g. a redelivered flood after a container restart — instead of draining it one recv at a time: note the highest `[msg #N]` you've seen, then `ack_until(up_to: N)`. Acked messages never redeliver; messages newer than `up_to` stay queued. Only affects your own inbox. - `mcp__hyperhive__send(to, body, in_reply_to?)` — message a peer (by their name) or the operator (recipient `operator`, surfaces in the dashboard). Use `to: "*"` to broadcast to all agents (they receive a hint that it's a broadcast and may not need action). Use `to: ""` to address your structural parent without hardcoding their name — hive-c0re rewrites it at delivery time per `topology.json`, falling back to `operator` if you're a root agent. Use `to: ""` to fan-out to every direct child of yours per `topology.json` (no-op for leaf agents). Both sentinels let the operator reparent at runtime with zero change on your side. Optional `in_reply_to: ` threads this message under a prior one — the dashboard and per-agent inbox render it with a `↳ reply` link. Some agents have a per-agent allow-list (`hyperhive.allowedRecipients` in their `agent.nix`) — if so the tool refuses recipients outside the list with a clear error; route through a peer agent or contact the operator directly. - (some agents only) **extra MCP tools** surfaced as `mcp____` — these are agent-specific (matrix client, scraper, db connector, etc.) declared in your `agent.nix` under `hyperhive.extraMcpServers`. Treat them as first-class tools alongside the hyperhive surface; the operator already auto-approved them at deploy time. diff --git a/hive-ag3nt/src/mcp.rs b/hive-ag3nt/src/mcp.rs index fc78cf90..8061e906 100644 --- a/hive-ag3nt/src/mcp.rs +++ b/hive-ag3nt/src/mcp.rs @@ -613,7 +613,7 @@ pub struct RecvArgs { /// Maximum number of messages to pop in this round-trip. Default /// (None) is 1 (single-message behaviour — exactly what you want /// when you're called to drive a turn off the first wake). Pass - /// a higher value (capped at 5 server-side) when you've been + /// a higher value (capped at 32 server-side) when you've been /// told the inbox has more queued (the wake prompt mentions /// pending count) and want to drain everything in one tool call. /// Once the long-poll wakes up, the call drains up to `max` in @@ -806,7 +806,7 @@ impl AgentServer { `wait_seconds` (capped at 180) to park the turn waiting for new work — incoming \ messages wake you instantly, otherwise the call returns empty at the timeout. \ That's strictly better than a fixed shell `sleep`. \n\n\ - **Batch drain**: pass `max: N` (capped at 5) to drain up to N messages in one \ + **Batch drain**: pass `max: N` (capped at 32) to drain up to N messages in one \ round-trip. Use this when the wake prompt told you the inbox has more queued, or \ any time you expect a burst — one tool call beats N consecutive single recvs. \ `wait_seconds` still applies to the FIRST message; once one arrives the call drains \ diff --git a/hive-ag3nt/src/serve_common.rs b/hive-ag3nt/src/serve_common.rs index bdff6cf7..abc460f9 100644 --- a/hive-ag3nt/src/serve_common.rs +++ b/hive-ag3nt/src/serve_common.rs @@ -31,13 +31,9 @@ pub fn format_wake_prompt( let pending = if unread == 0 { String::new() } else { - // Suggested batch size is clamped to the server-side recv cap - // so the hint never asks for more than one round-trip can - // deliver. - let batch = unread.min(u64::from(hive_sh4re::RECV_BATCH_MAX)); format!( "\n\n({unread} more message(s) pending in your inbox — call `mcp__hyperhive__recv` \ - with `max: {batch}` to drain the next batch before acting. If the \ + with `max: {unread}` to drain them all in one round-trip before acting. If the \ backlog is stale/already handled, `ack_until(up_to: )` \ clears everything up to that id in one call instead.)" ) diff --git a/hive-c0re/src/socket_server.rs b/hive-c0re/src/socket_server.rs index b21f60e3..df590644 100644 --- a/hive-c0re/src/socket_server.rs +++ b/hive-c0re/src/socket_server.rs @@ -149,10 +149,14 @@ async fn serve(stream: UnixStream, agent: String, coord: Arc) -> Re /// positive `wait_seconds`. pub(crate) const RECV_LONG_POLL_MAX: std::time::Duration = std::time::Duration::from_mins(3); -/// Server-side hard cap on `Recv.max` — canonical value lives in -/// `hive_sh4re::RECV_BATCH_MAX` so the harness's wake-prompt hint and -/// this enforcement site can't drift apart. -pub(crate) const RECV_BATCH_MAX: u32 = hive_sh4re::RECV_BATCH_MAX; +/// Server-side hard cap on `Recv.max`. Bounds the size of a single +/// round-trip so a confused caller can't drain the entire inbox in +/// one go and blow past wire-buffer sizes; everything above the cap +/// silently clamps. 32 is comfortably above the burst sizes we've +/// seen in practice (post-rebuild rescue, multi-agent reply storms) +/// and well under the per-message `MESSAGE_MAX_BYTES` * N envelope +/// budget. +pub(crate) const RECV_BATCH_MAX: u32 = 32; pub(crate) fn recv_timeout(wait_seconds: Option) -> std::time::Duration { match wait_seconds { diff --git a/hive-sh4re/src/lib.rs b/hive-sh4re/src/lib.rs index 09931b0d..3913e120 100644 --- a/hive-sh4re/src/lib.rs +++ b/hive-sh4re/src/lib.rs @@ -7,16 +7,6 @@ pub mod paths; pub mod priv_proto; pub mod wire_time; -/// Server-side hard cap on `Recv.max` (see `AgentRequest::Recv`). Bounds -/// the size of a single round-trip so a confused caller can't drain the -/// entire inbox in one go and blow past wire-buffer sizes; everything -/// above the cap silently clamps. 5 keeps individual turns small — a big -/// backlog is drained over several recv calls instead of one giant pop. -/// Lives here so both the enforcing side (hive-c0re's `socket_server`) and -/// the hinting side (hive-ag3nt's wake prompt + tool docs) reference one -/// constant instead of a scattered magic value. -pub const RECV_BATCH_MAX: u32 = 5; - // ----------------------------------------------------------------------------- // Host admin socket — /run/hyperhive/host.sock // -----------------------------------------------------------------------------