hyperhive/hive-agent
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 7a826f9ee2 refactor(sock): one socket client, retry as a policy value
Six places in the tree hand-rolled the same connect / write one JSON
line / read one JSON line back. Two of them — the harness serve loop's
client and the MCP server's — were byte-identical apart from a six-line
wrapper, ~145 lines of literal copy-paste. The other four each
reimplemented a subset, and the subsets had drifted: some named the
socket path in their errors and some did not, one classified transient
against fatal failures and the rest retried nothing at all, two drained
the response and two decoded it.

That duplication was defended when the daemons were split out, on the
grounds that a daemon's socket etiquette should stay visible in the
crate that depends on it. The etiquette genuinely does differ. The code
does not, and five copies is where "each daemon documents its own
etiquette" stops paying for itself.

`hive-sock-client` now owns the transport once, generic over the
request and response types so it is protocol-agnostic: the host-served
control socket and the harness's in-agent socket both use it with their
own wire-type crates. The two real differences become values instead of
forks. Retry is `Retry::RideOutRestart` (2/4/8/16/30s, sized to ride out
a service restart) for callers with no natural retry of their own, or
`Retry::None` for callers already inside a poll loop where the poll
interval is the retry — and the reason each caller picked one is a
comment at the call site rather than a reimplementation. The response is
either decoded (`request`) or half-closed and drained (`notify`, where
the drain exists so the server's write-back doesn't land on a closed
socket). Whether a failure propagates or is logged and swallowed stays
at the call site, because that is the caller's choice and not a property
of the transport.

Errors always name the socket path now, everywhere. That detail is
load-bearing: a permission problem on a socket that reads as "is the
daemon running?" sends the operator to fix the wrong thing.

The transient-against-fatal enum is gone rather than moved. Serialising
happens before the retry loop and deserialising after it, so only
connect, I/O and short-read failures can reach the loop at all — a
deterministic failure is now unretryable by construction instead of by
classification.

It is deliberately a new crate and not part of `hive-agent-sock`. The
`*-sock` crates are pure wire types by convention — `hive-agent-sock`
depends on serde and nothing else — and the two largest copies talk to
the host socket, whose types live in a different crate entirely. A
transport in either wire-type crate would drag tokio into it and point
the wrong way besides.

No wire-format change: same JSON line in, same line out.
2026-07-26 22:44:48 +02:00
..
prompts prompts: fix rebase conflict + qualify inbox tool names consistently 2026-07-26 17:33:41 +02:00
src refactor(sock): one socket client, retry as a policy value 2026-07-26 22:44:48 +02:00
Cargo.toml refactor(sock): one socket client, retry as a policy value 2026-07-26 22:44:48 +02:00
README.md remove hive-agent-wake — no shipped consumer 2026-07-25 20:05:32 +02:00

hive-agent

The in-container harness serve-loop binary — one instance per agent. Long-polls the broker inbox and drives one claude --print turn per inbox message, over the hive-claude driver. There is one role here (agent); the Surface trait + AgentSurface zero-sized type tag keep the turn loop generic and testable for future roles without a parallel copy of the loop.

When to use it

You don't call into this crate from elsewhere — it's the top-level binary systemd starts per agent container. Look here when you need to understand or change: what happens between "a message lands in the inbox" and "claude produces a reply", how login/auth is bootstrapped, how the per-agent web UI is served, or how turn/event stats get recorded. Architecture detail lives in docs/turn-loop.md; this README is just the map of the module tree.

Shape

  • turn.rs — the turn-loop policy layer: renders the system prompt + MCP config, invokes hive-claude, classifies the outcome, and feeds the event/turn-stats sinks.
  • client.rs — broker client (inbox poll, ack, send) speaking the hive-sh4re wire protocol.
  • login.rs / login_session.rs — first-run and session-resume auth flow for the claude CLI.
  • mcp_config.rs — renders the per-turn --mcp-config / --allowedTools blob from tool groups + capabilities.
  • todos.rs / reminders.rs / todo_server.rs — the harness-local loose-ends v2 stores (sqlite-backed) and the in-agent socket server extra MCP daemons + hive-agent-mcp dial into for todo/reminder ops.
  • vacuum.rs — periodic sqlite vacuum sweep for the harness-local stores.
  • events.rs / turn_stats.rs / stats.rs — append-only event sink and per-turn telemetry recording (context usage, cost, tool favorites) under harness/.
  • forge_notify.rs — subscribes to forge notifications and wakes the harness on new activity.
  • prompt.rs — system-prompt renderer (persona + tool docs + environment facts).
  • web_ui/ — the per-agent dashboard (terminal pane, status, schedules) served over the built-in hive-agent web port.
  • paths.rs — canonical path resolution for state/harness dirs and the harness-local sqlite files.

Sibling: hive-agent-mcp (the MCP server this loop points claude at every turn). Both are described together in docs/turn-loop.md::Harness binary shape.