| Filename | Latest commit message | Latest commit date |
|---|---|---|
The controller connects to the swarm queue as its own client and serves what each hive last said about itself at GET /api/hives/status. THE QUEUE IS THE STORE. A hive publishes into the `hive-status` JetStream KV bucket (history 1) and the controller reads it per request, keeping no copy. A cache here would be a second answer to the same question, free to disagree with the first, and the disagreement surfaces as a hive reading healthy on a dashboard while the bucket says otherwise. Whichever side arrives first creates the bucket; both want the same shape. Rows come from the roster rather than from the bucket, so an empty bucket renders as a swarm nobody has heard from instead of a healthy one, and `never_reported` stays distinct from `stale` - went quiet is a fault, never spoke is usually a deployment that has not happened. Freshness is derived at read time and never stored as a flag, because a stored `healthy` boolean goes stale silently the moment nothing arrives, which is the failure this endpoint is designed against. The timestamp is the NATS server's, applied when the value landed, so a publisher cannot make itself look fresher than it is. Authentication is per connection attempt, not per process. Authelia issues `client_credentials` tokens that expire in 3599s, and auth happens at CONNECT, so a long-lived connection is fine but a reconnect an hour later needs a token minted an hour later. `with_auth_callback` is re-run by async-nats for each attempt, which handles expiry by construction rather than by a timer - the alternative fails in the way this subsystem exists to prevent, with the controller still serving while its data quietly stops updating. Three failure shapes are deliberate: - A half-set environment is fatal; an absent one is not. Silently behaving like an unconfigured host is how every hive ends up reading `never_reported` with nothing to point at. - The endpoint answers 503 rather than an empty list when the store cannot be read. "I cannot reach the store" and "every hive is silent" are different answers, and rendering the second turns a local fault into an apparent swarm-wide outage. - `retry_on_initial_connect` makes the daemon and the queue bootable in either order, and the status handler refuses when the client is not Connected rather than issuing a request into it - a request made in that window does not fail, it waits, so every poll would hang and learn nothing. `Pending` is the state a never-connected client is in, which is why the test is `!= Connected` and not `== Disconnected`. The rendering rules are a pure function over a map, so the semantics are tested against a table rather than against a running server. The KV read, the credential rotation and the 503 paths are covered behaviourally instead: a real NATS server with a rotating token endpoint, asserting that the controller recovers only when the credential rotates, and mutation-tested by holding the credential wrong for the same window. |
||
| .. | ||
| crates | ||
| swarm | ||
| tools | ||
| turn-loop | ||
| web-ui | ||
| agent-hierarchy.md | ||
| approvals.md | ||
| boundary.md | ||
| ci.md | ||
| conventions.md | ||
| coordinator.md | ||
| forge.md | ||
| gateway.md | ||
| github.md | ||
| gotchas.md | ||
| knowledge.md | ||
| matrix.md | ||
| network.md | ||
| observability.md | ||
| persistence.md | ||
| pr-review-gate.md | ||
| README.md | ||
| security.md | ||
| setup.md | ||
| snapshot-store.md | ||
| terminal-rendering.md | ||
| web-ui.md | ||
hyperhive docs
Depth reference for hyperhive — the substrate, not the pitch (that's the
top-level README / website).
Every page here stands alone; pick the one matching your task rather than
reading top to bottom. For the auto-generated NixOS options reference
(every services.hyperhive.* / hyperhive.* option, host and agent), see
the options site instead —
this tree is prose, that one's generated straight from the module
declarations.
Getting started
- Bringing a fresh hive online? →
setup.md(first-runhivectlbootstrap). - What does the dashboard look like, and how do I use it? →
web-ui/— the operator-facing starting point; its own sub-pages (shape,dashboard,agent,css-vars) go deeper into implementation. - What tools does an agent (or the operator) have available? →
tools/—hivectl(yours) plus every agent's MCP tool surface (bash, forge, lifecycle, matrix, scheduling).
Dashboard & agent UI internals
- How does the per-agent terminal classify + colour events? →
terminal-rendering.md.
Turn loop, config, approvals
- How does claude get its prompt, and what tools does it have? →
turn-loop/— the loop, binary shape, turn outcomes; sub-pages:claude-invocation,config,mcp. - How do config changes flow from manager to operator to container? →
approvals.md(two-step spawn, approval state machine,flake.lockvalidation). - What state survives destroy / purge / restart? →
persistence.md.
Trust boundary & security
- What's the operator/agent trust boundary? What's a capability? →
boundary.md. - Agent trust model, prompt-injection threat model, credential
isolation? →
security.md. - Who can do what to whom — agent hierarchy and privilege? →
agent-hierarchy.md.
Accounts & integrations
- How do per-agent forge accounts work? What does
forge_notifypoll, and how does it format wake messages? →forge.md(the hive's own Forgejo);tools/forge.mdfor thehive-forgeCLI verbs agents actually call. - How does the matrix-tuwunel container work? Multiple accounts per
agent? →
matrix.md(the homeserver);tools/matrix.mdfor the MCP tool surface andhyperhive.matrixAccounts. - How do I give an agent a GitHub account (
gh+git push)? How is the PAT injected? →github.md. - What does
hivectldo? Provisioning, gateway users, container shells? →tools/hivectl.md(the curated guide);tools/hivectl-cli.mdfor the exhaustive, auto-generated flag reference.
Networking & swarms
- What nginx vhosts does the gateway serve? How does matrix
discovery work? →
gateway.md. - How does DNS resolution work in agent containers? What's the
bridge network for? →
network.md. - How do I connect two hives into a swarm? →
swarm/(peer hives, TLS trust). - Where do agent snapshots go? How does the swarm's
btrfs receiveendpoint authenticate a pushing hive? →snapshot-store.md.
Scheduler, CI, observability
- How does the rebuild queue work? What are queue kinds and
sources? →
coordinator.md. - How does the CI runner work? What's the auto-registration flow? →
ci.md. - How do I export Claude Code metrics (tokens, cost, tool calls) to
Prometheus/Grafana? →
observability.md.
Crate reference
- What does a specific Rust crate do, on its own terms? →
crates/— every workspace crate's ownREADME.md, one level up from source (hyperhive#3051); the crate itself is still the source of truth, this is just a walkable mirror.
Process & conventions
- Naming, commit style, wire protocol, the
data-asyncpattern? →conventions.md. - Why does the nspawn flag look like that? →
gotchas.md(bind mounts, conf flags, other NixOS/nspawn quirks). - What is
/knowledge? How does the hive-wide knowledge repo sync, and how do I contribute a document? →knowledge.md.