Six places in the tree hand-rolled the same connect / write one JSON line / read one JSON line back. Two of them — the harness serve loop's client and the MCP server's — were byte-identical apart from a six-line wrapper, ~145 lines of literal copy-paste. The other four each reimplemented a subset, and the subsets had drifted: some named the socket path in their errors and some did not, one classified transient against fatal failures and the rest retried nothing at all, two drained the response and two decoded it. That duplication was defended when the daemons were split out, on the grounds that a daemon's socket etiquette should stay visible in the crate that depends on it. The etiquette genuinely does differ. The code does not, and five copies is where "each daemon documents its own etiquette" stops paying for itself. `hive-sock-client` now owns the transport once, generic over the request and response types so it is protocol-agnostic: the host-served control socket and the harness's in-agent socket both use it with their own wire-type crates. The two real differences become values instead of forks. Retry is `Retry::RideOutRestart` (2/4/8/16/30s, sized to ride out a service restart) for callers with no natural retry of their own, or `Retry::None` for callers already inside a poll loop where the poll interval is the retry — and the reason each caller picked one is a comment at the call site rather than a reimplementation. The response is either decoded (`request`) or half-closed and drained (`notify`, where the drain exists so the server's write-back doesn't land on a closed socket). Whether a failure propagates or is logged and swallowed stays at the call site, because that is the caller's choice and not a property of the transport. Errors always name the socket path now, everywhere. That detail is load-bearing: a permission problem on a socket that reads as "is the daemon running?" sends the operator to fix the wrong thing. The transient-against-fatal enum is gone rather than moved. Serialising happens before the retry loop and deserialising after it, so only connect, I/O and short-read failures can reach the loop at all — a deterministic failure is now unretryable by construction instead of by classification. It is deliberately a new crate and not part of `hive-agent-sock`. The `*-sock` crates are pure wire types by convention — `hive-agent-sock` depends on serde and nothing else — and the two largest copies talk to the host socket, whose types live in a different crate entirely. A transport in either wire-type crate would drag tokio into it and point the wrong way besides. No wire-format change: same JSON line in, same line out.
2.3 KiB
hive-sock-client
One JSON-line-over-unix-socket client, shared by every daemon that speaks to a hyperhive socket.
The wire protocol is the same everywhere: connect, write one line of JSON, read one line of JSON back. Before this crate existed, five daemons each carried their own copy of that — two of them byte-identical — and the retry, response-handling and error-context behaviour drifted between them.
The crate is generic over the request and response types, so it is protocol-agnostic: the host-served control socket and the harness's in-agent socket both use it, with their own wire-type crates.
Shapes
request/request_retried— write, then decode the response line.request_retriedadditionally reports how many retries it took, for callers (MCP tool handlers) that want to tell the model a socket flake happened so it doesn't retry at the LLM level.notify— write, half-close, drain the response line and discard it. For fire-and-forget ops where the reply carries nothing the caller acts on. The drain is not optional: without it the server's write-back lands on a closed socket.
Retry is a policy value, not a fork
Retry::RideOutRestart backs off 2/4/8/16/30s (60s total), sized to ride
out a service restart. It is for callers with no natural retry of their
own — a stall there costs a surfaced tool error and the tokens to handle
it.
Retry::None fails fast. It is for callers already inside a poll loop,
where the poll interval is the retry and a second backoff would only
stack sleeps and delay the rest of the batch.
Two named policies, not a configurable schedule: nobody needs a third yet, and naming them keeps the reason for the choice at the call site.
Error contract
Errors always name the socket path. That detail is load-bearing: a permission problem on a socket that reads as "is the daemon running?" sends the operator to fix the wrong thing.
A refused or missing socket additionally gets a "may be restarting" hint, and an exhausted retry schedule records how long it tried, so a surfaced error reads as the likely transient it usually is.
Serialisation and deserialisation failures are never retried — retrying identical bytes reproduces the same failure. They are raised outside the retry loop, so only connect/IO/short-read failures ever reach it.