hyperhive/hive-sock-client
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas be3411e180 feat(#3245): gate rustdoc in nix flake check, and clear the workspace
Nothing in the gate read doc-comments: clippy doesn't check intra-doc
links, cargo test doesn't, and no check built docs. So a [`Foo`] pointing
at a renamed, moved or deleted item rendered as plain text and had no
discoverer but a human happening to read the comment.

That matters here more than in most repos, because the convention is to
put a thing's authoritative description in one doc-comment and point at
it from everywhere else -- the design leans on the pointers being real,
and a dangling link is worse than no link since it names something and
sends the reader looking.

Adds `docs-rustdoc` to nix/checks.nix: craneLib.cargoDoc over
--workspace --no-deps --document-private-items, denying six rustdoc
lints. Listed explicitly rather than -D warnings so a new lint appearing
upstream cannot red the build on a class nobody has triaged.

--document-private-items is load-bearing rather than thoroughness for
its own sake: most of this workspace's doc-comments live on private
items and //! module headers, so without it rustdoc checks a small
fraction of the links and the gate sits green while the rot continues.

Then fixes every error it reports, 40 to 0 across nine crates. The
classes differ and so do the fixes:

- public item, wrong scope -> qualify. Node and Node::parent are both
  public; the link failed only because scheduler.rs does not import
  Node. Six sites become [`crate::Node::parent`].
- private item -> downgrade to backticks. Nothing was made public to
  satisfy a lint; changing API surface to appease a doc check would be
  the tail wagging the dog.
- genuinely dead -> [`JobBuilder::insert_into`] names a method that does
  not exist. Insertion is Scheduler::insert_job.
- prose that looks like markup -> argv[0] parsed as a link, and
  <args>/<hex>/<name> parsed as HTML tags.

Note for future fixes: pub(crate) resolves in an intra-doc link, a plain
private fn in a binary crate does not (wait_for_nodes resolved,
connect_hint did not, same crate, same shape).

The check does not ride the clippy/test artifact cache. It takes
cargoArtifacts, but rustdoc needs its own flavour of dependency
metadata, which cargo build does not produce, so a --no-deps docs build
still compiles dependencies it never documents. Measured at 6m47s cold;
that reasoning is recorded in the check's own comment so the next reader
does not re-derive it.

Verified by running the check's exact command against the pre-cleanup
tree first: 40 errors, build failed. A gate that cannot fail is not
evidence, and building it before the cleanup makes that proof free.
2026-08-14 02:30:55 +02:00
..
src feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
Cargo.toml refactor(sock): one socket client, retry as a policy value 2026-07-26 22:44:48 +02:00
README.md refactor(sock): one socket client, retry as a policy value 2026-07-26 22:44:48 +02:00

hive-sock-client

One JSON-line-over-unix-socket client, shared by every daemon that speaks to a hyperhive socket.

The wire protocol is the same everywhere: connect, write one line of JSON, read one line of JSON back. Before this crate existed, five daemons each carried their own copy of that — two of them byte-identical — and the retry, response-handling and error-context behaviour drifted between them.

The crate is generic over the request and response types, so it is protocol-agnostic: the host-served control socket and the harness's in-agent socket both use it, with their own wire-type crates.

Shapes

  • request / request_retried — write, then decode the response line. request_retried additionally reports how many retries it took, for callers (MCP tool handlers) that want to tell the model a socket flake happened so it doesn't retry at the LLM level.
  • notify — write, half-close, drain the response line and discard it. For fire-and-forget ops where the reply carries nothing the caller acts on. The drain is not optional: without it the server's write-back lands on a closed socket.

Retry is a policy value, not a fork

Retry::RideOutRestart backs off 2/4/8/16/30s (60s total), sized to ride out a service restart. It is for callers with no natural retry of their own — a stall there costs a surfaced tool error and the tokens to handle it.

Retry::None fails fast. It is for callers already inside a poll loop, where the poll interval is the retry and a second backoff would only stack sleeps and delay the rest of the batch.

Two named policies, not a configurable schedule: nobody needs a third yet, and naming them keeps the reason for the choice at the call site.

Error contract

Errors always name the socket path. That detail is load-bearing: a permission problem on a socket that reads as "is the daemon running?" sends the operator to fix the wrong thing.

A refused or missing socket additionally gets a "may be restarting" hint, and an exhausted retry schedule records how long it tried, so a surfaced error reads as the likely transient it usually is.

Serialisation and deserialisation failures are never retried — retrying identical bytes reproduces the same failure. They are raised outside the retry loop, so only connect/IO/short-read failures ever reach it.