a swarm o agents, each in its own nspawn cage, gossiping over unix sockets. config changes flow as git commits, the operator approves them in a browser, every deploy is a tag. cyberpunk-themed dashboard included. 💜
  • Rust 66.9%
  • Nix 17.6%
  • JavaScript 6.1%
  • TypeScript 4.3%
  • CSS 3.5%
  • Other 1.6%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas b8c5840299 swarm-controller: retry hive provisioning until the store is up
The read policy and cert-auth role for each hive were written once, at
startup. On the deploy that surfaced this, the store was still coming
up, the pass logged its warning and moved on, and no hive could log in
until someone restarted the daemon — while cert auth answered "no chain
matching all constraints", which reads like a certificate problem
rather than a role that was never created.

The bootstrap unit in swarm-bao.nix lost the same race and won on its
retry 30s later. A daemon that boots alongside its store loses that race
routinely; on a normal boot it is the ordinary case.

The two passes fold into one `provision()` that logs in once instead of
twice for two loops over the same list, keeping policy before role since
the role names the policy. `ensure_hive_access` still awaits the first
pass, so a store that is already up leaves nothing deferred, and only a
pass that could not reach the store at all spawns the retry.

The retry is `config_pr::spawn`'s idiom from this same crate: an
interval task whose first tick is immediate. Its cadence and bound match
the bootstrap unit's — 30s, ~a day — because the two halves of one race
should not disagree about how long a wait is worth.

`Error::MissingEnv` is what keeps it from spinning forever: no `BAO_*`
set means a deployment that runs no store, where asking again changes
nothing, so it returns Ok. Everything else is retryable, including an
authority file that is not placed yet — the unit that writes it starts
alongside this one. Both cases previously landed in the same "not
managed here" line, so a store that was late looked exactly like one
that was never configured.

Per-hive failures keep their old behaviour: logged, skipped, Ok. A store
that refuses one hive's write refuses it again, so the next start really
is the right retry for those, and the module doc still says so.

Closes #4176.
2026-09-11 09:03:30 +02:00
.forgejo/workflows ci: drop bare issue tags from ci.yml comments (mara, #4146) 2026-09-09 22:55:28 +02:00
branding docs(#1182): remove component-diagram.svg; trim README; link to website + options 2026-06-03 19:06:06 +02:00
claude-plugins docs, prompts, hive-forge: stop handing readers the renamed verbs 2026-09-10 17:22:57 +02:00
docs docs/setup: contract "cannot" in the KV-mount troubleshooting block 2026-09-11 00:55:03 +02:00
frontend swarm: add a declared "paused" agent wanted state 2026-09-11 01:35:40 +02:00
hive-agent hive-agent: add speech-bubble icon to assistant text output rows 2026-09-10 19:00:21 +02:00
hive-agent-mcp docs: three comments point at a nix directory that does not exist 2026-09-02 14:19:03 +02:00
hive-agent-sock treefmt: apply prettier 2026-09-02 15:25:07 +02:00
hive-bash-mcp raise mcp streamable-http session keepalive from 5m to 24h 2026-08-31 12:53:07 +02:00
hive-c0re swarm: add a declared "paused" agent wanted state 2026-09-11 01:35:40 +02:00
hive-core-agent-sock treefmt: apply prettier 2026-09-02 15:25:07 +02:00
hive-forge docs, prompts, hive-forge: stop handing readers the renamed verbs 2026-09-10 17:22:57 +02:00
hive-forge-notify check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00
hive-host-sock host.sock: push a live agent-status stream instead of poll-only 2026-09-07 18:43:47 +02:00
hive-jobq treefmt: apply prettier 2026-09-02 15:25:07 +02:00
hive-jobq-metrics move otel_http_client from swarm-queue-client into swarm-controller 2026-08-29 11:17:24 +02:00
hive-jobq-wire address review: move parse_states/filter_nodes_by_state to hive-jobq-wire, rename placeholder enums, trim core-mirroring framing 2026-08-16 16:59:54 +02:00
hive-matrix-mcp raise mcp streamable-http session keepalive from 5m to 24h 2026-08-31 12:53:07 +02:00
hive-metric docs: restructure into topic subdirectories, collapse duplicated index 2026-09-02 01:55:37 +02:00
hive-priv check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00
hive-priv-sock docs: restructure into topic subdirectories, collapse duplicated index 2026-09-02 01:55:37 +02:00
hive-screen-mcp treefmt: apply prettier 2026-09-02 15:25:07 +02:00
hive-sh4re add per-agent url to host.sock agent status rows 2026-09-07 20:30:38 +02:00
hive-sock-client treefmt: apply prettier 2026-09-02 15:25:07 +02:00
hive-subagent-mcp subagent: fix a real test race on the process-wide OTEL_RESOURCE_ATTRIBUTES env var 2026-09-09 23:47:21 +02:00
hive-types docs: restructure into topic subdirectories, collapse duplicated index 2026-09-02 01:55:37 +02:00
hivectl rewrite generated CLI docs' passive voice to active 2026-09-08 14:56:26 +02:00
nix swarm-bao: create the KV mount the controller writes credentials through 2026-09-11 00:16:46 +02:00
scripts check-issue-refs: scan .yml/.yaml too, closes #4148 2026-09-09 23:25:54 +02:00
swagger-ui-theme treefmt: apply prettier 2026-09-02 15:25:07 +02:00
swarm-authelia-bridge check-issue-refs: catch full forge issue URLs too, drop internal links from docs entirely 2026-09-09 21:15:28 +02:00
swarm-authelia-bridge-sock feat(swarm-authelia-bridge): report a heal as its own outcome 2026-08-23 19:00:41 +02:00
swarm-controller swarm-controller: retry hive provisioning until the store is up 2026-09-11 09:03:30 +02:00
swarm-nats-auth deploy: move the queue's callout identity out of swarm.nats 2026-09-07 14:24:52 +02:00
swarm-queue-client swarm: add a declared "paused" agent wanted state 2026-09-11 01:35:40 +02:00
swarm-secret-client swarm-secret-client: write hive policies to the modern ACL path 2026-09-11 01:34:31 +02:00
swarmctl docs: clear the remaining error-level vale lints 2026-09-09 22:55:28 +02:00
.gitignore docs: address review — redundancy proof for Passive, wave-2 split, re-enable Contractions 2026-09-07 11:56:31 +02:00
.mailmap chore(#2165): add damocles@pr1ma + lexis@pr1ma mailmap entries 2026-07-04 13:50:16 +02:00
.prettierignore hive-forge: add markdown-docs generator and CI freshness check 2026-09-02 19:38:34 +02:00
.prettierrc temp: add prettier configs 2026-07-02 23:33:11 +02:00
.vale.ini Disable Microsoft.HeadingColons and Microsoft.Percentages, fix Plurals hits 2026-09-07 14:19:24 +02:00
Cargo.lock swarm-secret-client: write hive policies to the modern ACL path 2026-09-11 01:34:31 +02:00
Cargo.toml swarm-secret-client: write hive policies to the modern ACL path 2026-09-11 01:34:31 +02:00
CLAUDE.md subagent: add status tool, cut docs down to operator-facing + no cli flags 2026-09-09 23:45:12 +02:00
clippy.toml hivectl: wireguard mesh setup verbs (#1756) 2026-06-19 14:37:50 +02:00
flake.lock flake: bump nixpkgs 569d5785 -> 5dfba623 2026-09-02 14:16:57 +02:00
flake.nix nix: wire the independent hive-subagent-daemon systemd unit and MCP server 2026-09-09 23:45:12 +02:00
README.md docs/hive-c0re: fix ask/answer removal doc gaps argus caught on #3741 2026-08-30 03:02:31 +02:00

hyperhive

a swarm of claude-code agents, each in its own nspawn cage, gossiping over unix sockets. config changes flow as git commits, the operator approves them in a browser, every deploy is a tag. cyberpunk-themed dashboard included. 💜

Claude code is great in one window, exponentielle across many — but only if you can keep the agents from stepping on each other, give them durable identity, and stop them from eating production. hyperhive is the substrate.

  • identity = unix socket
  • communication = sqlite-backed broker (send / recv / remind)
  • config = git (manager proposes, operator approves, deploys land as tagged commits)
  • blast radius = container
every hive (NixOS host, runs hive-c0re.service)
│
├── operator
│   ├── browser → :80 (hive-gateway)    dashboard + per-agent UIs
│   │                                   /agent/<name>/ → per-agent unix socket
│   └── CLI     → /run/hyperhive/host.sock   admin protocol
│
├── hive-c0re  (Rust daemon: lifecycle / broker / approvals /
│               auto-update / dashboard / sockets)
│
├── hive-gateway (optional)   nginx — proxies :80 → c0re dashboard + per-agent sockets
│
└── agent containers
    ├── h-ruth     manager (privileged MCP surface, approval gating)
    └── h-<name>   sub-agent (claude + MCP tools + per-agent web UI + unix socket)

one host per swarm (optional — connects hives; can be any hive, including
one that's also running the tree above)
│
├── hive-forge             Forgejo — swarm-wide singleton, per-agent accounts + config mirror
├── hive-matrix            tuwunel — swarm-wide singleton, Matrix homeserver + per-agent accounts
├── swarm-controller       cross-hive state: hive directory, agent roster, jobs
├── swarm-ui               swarm-wide SPA, served straight off the gateway (no own container)
├── swarm-authelia         SSO — one login gates swarm-ui + Grafana + more
├── swarm-nats             message queue (JetStream KV: hive-status, …)
├── swarm-otel             telemetry collector, sole holder of the upstream credential
├── swarm-victoriametrics  metrics store
├── swarm-victorialogs     log store
└── swarm-grafana          dashboards over the metrics/log stores, own OIDC login

→ website · → docs · → options reference

Depth lives in docs/ (rendered at hyperhive.darkest.space/docs/) — start at docs/README.md and pick the page matching your task rather than reading front to back.

Quick start

Minimal flake.nix for a host that runs hive-c0re:

{
  inputs = {
    nixpkgs.url = "github:NixOS/nixpkgs/nixos-26.05";
    hyperhive.url = "git+https://forge.darkest.space/hyperhive/hyperhive";
    # Pin hyperhive to your own nixpkgs instead of the one it ships with
    # (see "Overriding nixpkgs" below) — recommended for most hosts:
    hyperhive.inputs.nixpkgs.follows = "nixpkgs";
  };

  outputs = { nixpkgs, hyperhive, ... }: {
    nixosConfigurations.my-host = nixpkgs.lib.nixosSystem {
      system = "x86_64-linux";
      modules = [
        hyperhive.nixosModules.default  # hive-c0re + hive-forge + hive-gateway in one import
        ({ ... }: {
          services.hyperhive.enable = true;
          # services.hyperhive.c0re.operatorPronouns = "they/them";  # default: "she/her"

          # ... rest of your host config
          system.stateVersion = "25.11";
        })
      ];
    };
  };
}

hive-c0re opens its admin socket + dashboard, auto-creates the manager container, and auto-rebuilds any container whose hyperhive rev goes stale. claude-code is unfree — hyperhive scopes the whitelist to itself, nothing for the operator to set.

Overriding nixpkgs

hyperhive pins its own nixpkgs so it builds standalone in CI. Add hyperhive.inputs.nixpkgs.follows = "nixpkgs" (as in the quick-start above) to build it against your host's nixpkgs instead — one less nixpkgs evaluation, no version drift from the rest of your system. Standard flake follows pattern; works as long as your channel is reasonably close to the nixos-26.05 hyperhive develops against. Drop it again if a much older/newer channel hits breakage hyperhive's CI doesn't catch.

For the full list of host and agent NixOS options see the options reference.

Operator CLI

hivectl is the operator-facing host CLI for ad-hoc administration that doesn't go through the broker (built alongside hive-c0re when the host module is enabled):

sudo hivectl forge create-user mara                       # provisions a forge user
sudo hivectl forge create-user mara --password 'hunter2'  # … with a fixed password
sudo hivectl matrix create-user mara                      # provisions a matrix user
sudo hivectl matrix create-user mara --password-stdin     # … reading one line from stdin

For a name that's a managed agent, hivectl persists the resulting token to that agent's state dir, the same as the boot sweep does. For a non-agent name (e.g. the operator's own forge/matrix account), it prints the token to stdout and writes nothing.

Build / deploy

nix develop -c cargo check
nix flake check        # rust + nix + toml fmt + clippy

# deploy from a host config that imports hyperhive.nixosModules.default
nix flake update --update-input hyperhive
sudo nixos-rebuild switch --flake .#<host>