Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
..
agent-lifecycle Make agent creation swarm-only and refuse a name placed on another hive 2026-09-29 15:47:40 +02:00
crates check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00
getting-started matrix: swarm-controller is the only minter 2026-09-30 00:46:46 +02:00
integrations matrix: swarm-controller is the only minter 2026-09-30 00:46:46 +02:00
networking hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf 2026-09-27 18:55:33 +02:00
process agents: pull the forge token from bao; drop tea-login 2026-09-24 17:48:53 +02:00
scheduler Make agent creation swarm-only and refuse a name placed on another hive 2026-09-29 15:47:40 +02:00
swarm credential units: restart consumers on a changed credential; fix the ordering claim 2026-09-30 07:45:47 +02:00
tools hive-subagent-mcp: run an agent's subagents on its runtime 2026-09-30 07:41:12 +02:00
trust-boundary hive-priv: remove the RestartMatrixDaemon command 2026-09-29 13:54:18 +02:00
turn-loop hive-runtime: compact ACP sessions through the agent's compact command 2026-09-30 07:41:47 +02:00
web-ui hive-agent: dashboard Cancel and idle stalls for ACP agents 2026-09-29 23:25:48 +02:00
README.md docs: retire the agent hierarchy from every page that described it 2026-09-21 22:08:47 +02:00

hyperhive docs

Depth reference for hyperhive — the substrate, not the pitch (that's the top-level README / website). Every page here stands alone; pick the one matching your task rather than reading top to bottom. For the autogenerated NixOS options reference (every services.hyperhive.* / hyperhive.* option, host and agent), see the options site instead — this tree is prose, that one's generated straight from the module declarations.

Getting started

  • Bringing a fresh hive online? → getting-started/setup.md (first-run hivectl bootstrap).
  • What does the dashboard look like, and how do I use it? → web-ui/ — the operator-facing starting point; its own sub-pages (shape, dashboard, agent, css-vars, terminal-rendering) go deeper into implementation.
  • What tools does an agent (or the operator) have available? → tools/ — hivectl (yours) plus every agent's MCP tool surface (bash, forge, lifecycle, matrix, scheduling).

Agent lifecycle

Trust boundary & security

Accounts & integrations

  • How do per-agent forge accounts work? What does forge_notify poll, and how does it format wake messages? → integrations/forge.md (the hive's own Forgejo); tools/forge.md for the hive-forge CLI verbs agents actually call.
  • How does the matrix-tuwunel container work? Multiple accounts per agent? → integrations/matrix.md (the homeserver); tools/matrix.md for the MCP tool surface and services.hyperhive.agent.matrixAccounts.
  • How do I give an agent a GitHub account (gh + git push)? how's the PAT injected? → integrations/github.md (operator content up top; the gh/git-push + notification-poller mechanics are in a collapsed "Implementation" section at the bottom).
  • What's /knowledge? How does the hive-wide knowledge repo sync, and how do I contribute a document? → integrations/knowledge.md.
  • What does hivectl do? Provisioning, gateway users, container shells? → tools/hivectl.md (the curated guide); tools/hivectl-cli.md for the exhaustive, autogenerated flag reference.

Networking & swarms

  • What nginx vhosts does the gateway serve? How does matrix discovery work? → networking/gateway.md.
  • How does DNS resolution work in agent containers? What's the bridge network for? → networking/network.md.
  • How do I connect two hives into a swarm? → swarm/ (peer hives, TLS trust).
  • Where do agent snapshots go? How does the swarm's btrfs receive endpoint authenticate a pushing hive? → networking/snapshot-store.md.
  • Who mints each credential, who reads it, and how does it rotate — and where's that shape headed? → swarm/credentials.md (current state, target state, and the progressive-enhancement rule); swarm/secrets.md for where each file lives today.

Scheduler, CI, observability

  • what's the job queue, as a general idea (not hive-c0re specifics)? → scheduler/jobq.md — operator-facing, no implementation detail.
  • How does the rebuild queue work? What are the concrete step kinds, queue sources, scheduler internals? → scheduler/coordinator.md.
  • How does the CI runner work? What's the autoregistration flow? → scheduler/ci.md.
  • How do I export Claude Code metrics (tokens, cost, tool calls) to Prometheus/Grafana? → scheduler/observability.md.

Crate reference

  • What does a specific Rust crate do, on its own terms? → crates/ — every workspace crate's own README.md, one level up from source; the crate itself is still the source of truth, this is just a walkable mirror.

Process & conventions