hyperhive/hive-sock-client/README.md
atlas 39b95c2ede treefmt: apply prettier
Pure `nix fmt` output from the commit before this one — no hand edits.
203 files: 52 md, 42 tsx, 32 js, 32 css, 21 ts, 13 html, 8 json, 3 mjs.

Reproduce with `nix develop -c nix fmt` on the parent commit; the result
should be byte-identical to this tree.

None of the 13 `.prettierignore` entries appears here — verified by
intersecting the changed-file list against the ignore file, with a
control proving the intersection finds a match when one exists.
2026-09-02 15:25:07 +02:00

52 lines
2.3 KiB
Markdown

# hive-sock-client
One JSON-line-over-unix-socket client, shared by every daemon that speaks
to a hyperhive socket.
The wire protocol is the same everywhere: connect, write one line of JSON,
read one line of JSON back. Before this crate existed, five daemons each
carried their own copy of that — two of them byte-identical — and the
retry, response-handling and error-context behaviour drifted between them.
The crate is generic over the request and response types, so it is
protocol-agnostic: the host-served control socket and the harness's
in-agent socket both use it, with their own wire-type crates.
## Shapes
- `request` / `request_retried` — write, then decode the response line.
`request_retried` additionally reports how many retries it took, for
callers (MCP tool handlers) that want to tell the model a socket flake
happened so it doesn't retry at the LLM level.
- `notify` — write, half-close, drain the response line and discard it.
For fire-and-forget ops where the reply carries nothing the caller acts
on. The drain is not optional: without it the server's write-back lands
on a closed socket.
## Retry is a policy value, not a fork
`Retry::RideOutRestart` backs off 2/4/8/16/30s (60s total), sized to ride
out a service restart. It is for callers with no natural retry of their
own — a stall there costs a surfaced tool error and the tokens to handle
it.
`Retry::None` fails fast. It is for callers already inside a poll loop,
where the poll interval _is_ the retry and a second backoff would only
stack sleeps and delay the rest of the batch.
Two named policies, not a configurable schedule: nobody needs a third yet,
and naming them keeps the _reason_ for the choice at the call site.
## Error contract
Errors always name the socket path. That detail is load-bearing: a
permission problem on a socket that reads as "is the daemon running?"
sends the operator to fix the wrong thing.
A refused or missing socket additionally gets a "may be restarting" hint,
and an exhausted retry schedule records how long it tried, so a surfaced
error reads as the likely transient it usually is.
Serialisation and deserialisation failures are never retried — retrying
identical bytes reproduces the same failure. They are raised outside the
retry loop, so only connect/IO/short-read failures ever reach it.