| Filename | Latest commit message | Latest commit date |
|---|---|---|
hive-sock-client: each attempt now bounds connect (5s), write (10s) and the wait for the response (60s by default). The response bound is per call through the new `request_within`, which hive-agent's serve-loop `Recv` uses with its 180s long-poll plus 30s headroom. A response timeout is terminal rather than retried: the server holds the request, so a retry re-sends something it may still act on and multiplies the wait by the backoff schedule. Outbound HTTP: the matrix login/whoami clients in swarm-controller and hive-c0re's dashboard (5s connect, 30s request), the authelia-bridge client (5s/30s; ensuring an identity runs an argon2 hash first) and the ci-runner forge calls (5s/15s, config_pr_poll's forge budget) get a connect_timeout and a request timeout. Timeout errors name the bound that fired. hive-agent's unix-socket extra web proxy bounds the connect (5s) and the wait for the response head (30s, the http sibling's budget); the body read stays unbounded. Refs #4723 |
||
| .. | ||
| src | ||
| Cargo.toml | ||
| README.md | ||
hive-sock-client
One JSON-line-over-unix-socket client, shared by every daemon that speaks to a hyperhive socket.
The wire protocol is the same everywhere: connect, write one line of JSON, read one line of JSON back. Before this crate existed, five daemons each carried their own copy of that — two of them byte-identical — and the retry, response-handling and error-context behaviour drifted between them.
The crate is generic over the request and response types, so it is protocol-agnostic: the host-served control socket and the harness's in-agent socket both use it, with their own wire-type crates.
Shapes
request/request_retried— write, then decode the response line.request_retriedadditionally reports how many retries it took, for callers (MCP tool handlers) that want to tell the model a socket flake happened so it doesn't retry at the LLM level.notify— write, half-close, drain the response line and discard it. For fire-and-forget ops where the reply carries nothing the caller acts on. The drain is not optional: without it the server's write-back lands on a closed socket.
Retry is a policy value, not a fork
Retry::RideOutRestart backs off 2/4/8/16/30s (60s total), sized to ride
out a service restart. It is for callers with no natural retry of their
own — a stall there costs a surfaced tool error and the tokens to handle
it.
Retry::None fails fast. It is for callers already inside a poll loop,
where the poll interval is the retry and a second backoff would only
stack sleeps and delay the rest of the batch.
Two named policies, not a configurable schedule: nobody needs a third yet, and naming them keeps the reason for the choice at the call site.
Error contract
Errors always name the socket path. That detail is load-bearing: a permission problem on a socket that reads as "is the daemon running?" sends the operator to fix the wrong thing.
A refused or missing socket additionally gets a "may be restarting" hint, and an exhausted retry schedule records how long it tried, so a surfaced error reads as the likely transient it usually is.
Serialisation and deserialisation failures are never retried — retrying identical bytes reproduces the same failure. They are raised outside the retry loop, so only connect/IO/short-read failures ever reach it.