hyperhive/hive-sock-client/README.md
atlas 39b95c2ede treefmt: apply prettier
Pure `nix fmt` output from the commit before this one — no hand edits.
203 files: 52 md, 42 tsx, 32 js, 32 css, 21 ts, 13 html, 8 json, 3 mjs.

Reproduce with `nix develop -c nix fmt` on the parent commit; the result
should be byte-identical to this tree.

None of the 13 `.prettierignore` entries appears here — verified by
intersecting the changed-file list against the ignore file, with a
control proving the intersection finds a match when one exists.
2026-09-02 15:25:07 +02:00

2.3 KiB

hive-sock-client

One JSON-line-over-unix-socket client, shared by every daemon that speaks to a hyperhive socket.

The wire protocol is the same everywhere: connect, write one line of JSON, read one line of JSON back. Before this crate existed, five daemons each carried their own copy of that — two of them byte-identical — and the retry, response-handling and error-context behaviour drifted between them.

The crate is generic over the request and response types, so it is protocol-agnostic: the host-served control socket and the harness's in-agent socket both use it, with their own wire-type crates.

Shapes

  • request / request_retried — write, then decode the response line. request_retried additionally reports how many retries it took, for callers (MCP tool handlers) that want to tell the model a socket flake happened so it doesn't retry at the LLM level.
  • notify — write, half-close, drain the response line and discard it. For fire-and-forget ops where the reply carries nothing the caller acts on. The drain is not optional: without it the server's write-back lands on a closed socket.

Retry is a policy value, not a fork

Retry::RideOutRestart backs off 2/4/8/16/30s (60s total), sized to ride out a service restart. It is for callers with no natural retry of their own — a stall there costs a surfaced tool error and the tokens to handle it.

Retry::None fails fast. It is for callers already inside a poll loop, where the poll interval is the retry and a second backoff would only stack sleeps and delay the rest of the batch.

Two named policies, not a configurable schedule: nobody needs a third yet, and naming them keeps the reason for the choice at the call site.

Error contract

Errors always name the socket path. That detail is load-bearing: a permission problem on a socket that reads as "is the daemon running?" sends the operator to fix the wrong thing.

A refused or missing socket additionally gets a "may be restarting" hint, and an exhausted retry schedule records how long it tried, so a surfaced error reads as the likely transient it usually is.

Serialisation and deserialisation failures are never retried — retrying identical bytes reproduces the same failure. They are raised outside the retry loop, so only connect/IO/short-read failures ever reach it.