hyperhive/docs
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas bb0afcd256 nix: the store's own collector scrapes its metrics listener
bao's metrics were scraped by the SWARM collector over loopback, via a
`swarm.otel.scrapeTargets.bao` entry gated on `deploy.swarm-otel.enable`
— "does the swarm's collector run on THIS host". It had to be: loopback
only reaches a reader that landed on the same host.

What that rendered everywhere else was nothing at all. Off that host the
metrics listener was not emitted, so the store's metrics reached the
store nowhere, and a host with no entry is indistinguishable from a host
nobody asked to scrape.

Moves the scrape into the collector this container already runs, per
mara on #4537: "move the existing scraper to the local collector". The
container shares the host netns (privateNetwork = false), so the scrape
still dials 127.0.0.1 — the listener keeps its address, its
`metrics_only` narrowing and its loopback-only bind, and the API
listener's `tls_require_and_verify_client_cert` is untouched.

The listener and its `prometheus_retention_time` lose their gate: the
reader ships with the store now, so there is no host where the endpoint
has none. The metrics pipeline reuses the logs pipeline's `resource`
processor and `otlphttp` exporter, so both signals carry the same
`service.name` and leave by the one hop.

Logs are unaffected: `journaldUnits` and --link-journal=host stay until
every sibling swarm container has a collector of its own.

The module-eval absence arm "a store with no collector beside it serves
no metrics" is inverted rather than dropped — the condition it asserted
is the bug. Three cases join it: the job is in swarm-bao AND gone from
swarm-otel (a move, not a copy), the scrape target and listener are both
pinned to loopback, and the metrics pipeline shares its exporter with
the logs one.
2026-09-21 17:19:52 +02:00
..
agent-lifecycle remove the list_containers and request_update_meta_inputs MCP tools 2026-09-20 22:47:46 +02:00
crates check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00
getting-started swarm-matrix-ctl: one control binary for the matrix container, not one per job 2026-09-20 22:07:16 +02:00
integrations matrix: one sender account and one sender token per hive 2026-09-20 22:07:16 +02:00
networking docs: suppress reviewed write-good.Passive false positives 2026-09-20 16:24:11 +02:00
process remove the list_containers and request_update_meta_inputs MCP tools 2026-09-20 22:47:46 +02:00
scheduler docs/scheduler/observability.md: clear write-good.Passive hits from the merge 2026-09-20 19:05:13 +02:00
swarm nix: the store's own collector scrapes its metrics listener 2026-09-21 17:19:52 +02:00
tools remove the list_containers and request_update_meta_inputs MCP tools 2026-09-20 22:47:46 +02:00
trust-boundary matrix: one sender account and one sender token per hive 2026-09-20 22:07:16 +02:00
turn-loop remove the list_containers and request_update_meta_inputs MCP tools 2026-09-20 22:47:46 +02:00
web-ui docs: suppress reviewed write-good.Passive false positives 2026-09-20 16:24:11 +02:00
README.md docs: repoint agent-tier option paths to services.hyperhive.agent.* 2026-09-18 03:05:43 +02:00

hyperhive docs

Depth reference for hyperhive — the substrate, not the pitch (that's the top-level README / website). Every page here stands alone; pick the one matching your task rather than reading top to bottom. For the autogenerated NixOS options reference (every services.hyperhive.* / hyperhive.* option, host and agent), see the options site instead — this tree is prose, that one's generated straight from the module declarations.

Getting started

  • Bringing a fresh hive online?getting-started/setup.md (first-run hivectl bootstrap).
  • What does the dashboard look like, and how do I use it?web-ui/ — the operator-facing starting point; its own sub-pages (shape, dashboard, agent, css-vars, terminal-rendering) go deeper into implementation.
  • What tools does an agent (or the operator) have available?tools/hivectl (yours) plus every agent's MCP tool surface (bash, forge, lifecycle, matrix, scheduling).

Agent lifecycle

Trust boundary & security

Accounts & integrations

  • How do per-agent forge accounts work? What does forge_notify poll, and how does it format wake messages?integrations/forge.md (the hive's own Forgejo); tools/forge.md for the hive-forge CLI verbs agents actually call.
  • How does the matrix-tuwunel container work? Multiple accounts per agent?integrations/matrix.md (the homeserver); tools/matrix.md for the MCP tool surface and services.hyperhive.agent.matrixAccounts.
  • How do I give an agent a GitHub account (gh + git push)? how's the PAT injected?integrations/github.md (operator content up top; the gh/git-push + notification-poller mechanics are in a collapsed "Implementation" section at the bottom).
  • What's /knowledge? How does the hive-wide knowledge repo sync, and how do I contribute a document?integrations/knowledge.md.
  • What does hivectl do? Provisioning, gateway users, container shells?tools/hivectl.md (the curated guide); tools/hivectl-cli.md for the exhaustive, autogenerated flag reference.

Networking & swarms

  • What nginx vhosts does the gateway serve? How does matrix discovery work?networking/gateway.md.
  • How does DNS resolution work in agent containers? What's the bridge network for?networking/network.md.
  • How do I connect two hives into a swarm?swarm/ (peer hives, TLS trust).
  • Where do agent snapshots go? How does the swarm's btrfs receive endpoint authenticate a pushing hive?networking/snapshot-store.md.
  • Who mints each credential, who reads it, and how does it rotate — and where's that shape headed?swarm/credentials.md (current state, target state, and the progressive-enhancement rule); swarm/secrets.md for where each file lives today.

Scheduler, CI, observability

  • what's the job queue, as a general idea (not hive-c0re specifics)?scheduler/jobq.md — operator-facing, no implementation detail.
  • How does the rebuild queue work? What are the concrete step kinds, queue sources, scheduler internals?scheduler/coordinator.md.
  • How does the CI runner work? What's the autoregistration flow?scheduler/ci.md.
  • How do I export Claude Code metrics (tokens, cost, tool calls) to Prometheus/Grafana?scheduler/observability.md.

Crate reference

  • What does a specific Rust crate do, on its own terms?crates/ — every workspace crate's own README.md, one level up from source; the crate itself is still the source of truth, this is just a walkable mirror.

Process & conventions