a swarm o agents, each in its own nspawn cage, gossiping over unix sockets. config changes flow as git commits, the operator approves them in a browser, every deploy is a tag. cyberpunk-themed dashboard included. 💜
  • Rust 68.7%
  • Nix 15.7%
  • JavaScript 8.4%
  • CSS 3.7%
  • TypeScript 1.9%
  • Other 1.6%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
atlas 8891b46943 feat(swarm-controller): aggregate per-hive status from the swarm queue
The controller connects to the swarm queue as its own client and serves
what each hive last said about itself at GET /api/hives/status.

THE QUEUE IS THE STORE. A hive publishes into the `hive-status` JetStream
KV bucket (history 1) and the controller reads it per request, keeping no
copy. A cache here would be a second answer to the same question, free to
disagree with the first, and the disagreement surfaces as a hive reading
healthy on a dashboard while the bucket says otherwise. Whichever side
arrives first creates the bucket; both want the same shape.

Rows come from the roster rather than from the bucket, so an empty bucket
renders as a swarm nobody has heard from instead of a healthy one, and
`never_reported` stays distinct from `stale` - went quiet is a fault,
never spoke is usually a deployment that has not happened. Freshness is
derived at read time and never stored as a flag, because a stored
`healthy` boolean goes stale silently the moment nothing arrives, which
is the failure this endpoint is designed against. The timestamp is the
NATS server's, applied when the value landed, so a publisher cannot make
itself look fresher than it is.

Authentication is per connection attempt, not per process. Authelia
issues `client_credentials` tokens that expire in 3599s, and auth happens
at CONNECT, so a long-lived connection is fine but a reconnect an hour
later needs a token minted an hour later. `with_auth_callback` is re-run
by async-nats for each attempt, which handles expiry by construction
rather than by a timer - the alternative fails in the way this subsystem
exists to prevent, with the controller still serving while its data
quietly stops updating.

Three failure shapes are deliberate:

- A half-set environment is fatal; an absent one is not. Silently
  behaving like an unconfigured host is how every hive ends up reading
  `never_reported` with nothing to point at.
- The endpoint answers 503 rather than an empty list when the store
  cannot be read. "I cannot reach the store" and "every hive is silent"
  are different answers, and rendering the second turns a local fault
  into an apparent swarm-wide outage.
- `retry_on_initial_connect` makes the daemon and the queue bootable in
  either order, and the status handler refuses when the client is not
  Connected rather than issuing a request into it - a request made in
  that window does not fail, it waits, so every poll would hang and
  learn nothing. `Pending` is the state a never-connected client is in,
  which is why the test is `!= Connected` and not `== Disconnected`.

The rendering rules are a pure function over a map, so the semantics are
tested against a table rather than against a running server. The KV read,
the credential rotation and the 503 paths are covered behaviourally
instead: a real NATS server with a rotating token endpoint, asserting
that the controller recovers only when the credential rotates, and
mutation-tested by holding the credential wrong for the same window.
2026-08-15 18:37:23 +02:00
.forgejo/workflows docs: stop claiming tracker-tag/comment-block lint are non-required 2026-07-23 00:10:59 +02:00
branding docs(#1182): remove component-diagram.svg; trim README; link to website + options 2026-06-03 19:06:06 +02:00
claude-plugins skills(headless-screenshot): document two false-positive traps 2026-08-15 09:59:55 +02:00
docs feat(swarm-controller): aggregate per-hive status from the swarm queue 2026-08-15 18:37:23 +02:00
frontend swarm-ui: header links menu for swarm-wide services (hyperhive#3289) 2026-08-15 14:24:52 +02:00
hive-agent hive-agent: remove dead bash-task- wake_from special-case 2026-08-14 12:09:02 +02:00
hive-agent-mcp hive-agent: guarantee a wake after a self-requested /compact 2026-08-13 23:17:10 +02:00
hive-agent-sock hive-agent: guarantee a wake after a self-requested /compact 2026-08-13 23:17:10 +02:00
hive-bash-mcp feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
hive-c0re cut the comments back to what the code cannot say 2026-08-14 23:16:59 +02:00
hive-core-agent-sock hive-c0re/hivectl/hive-agent: pause as a job-queue DAG node (closes #3056) 2026-08-11 23:47:09 +02:00
hive-forge hive-forge: point to per-line review comments instead of inlining them, fix pr reviews line numbers 2026-08-15 12:34:23 +02:00
hive-forge-notify feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
hive-host-sock fix(3179): the gateway's config files get their own state dir 2026-08-12 10:29:27 +02:00
hive-jobq feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
hive-jobq-wire jobq: a generic per-state roll-up, served beside the graph 2026-08-03 20:37:06 +02:00
hive-matrix-mcp feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
hive-metric cut the comments back to what the code cannot say 2026-08-14 23:16:59 +02:00
hive-priv hive-priv: replace json! with typed structs for account sidecar files 2026-08-13 23:16:17 +02:00
hive-priv-sock feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
hive-screen-mcp docs(#2627): add README for hive-screen-mcp 2026-07-23 14:17:47 +02:00
hive-sh4re feat(3088): move the gateway's nginx + dnsmasq onto the host 2026-08-11 18:01:03 +02:00
hive-sock-client feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
hive-types docs(#2627): add READMEs for the remaining infra crates 2026-07-23 13:16:29 +02:00
hivectl feat(#3245): gate rustdoc in nix flake check, and clear the workspace 2026-08-14 02:30:55 +02:00
nix feat(swarm-controller): aggregate per-hive status from the swarm queue 2026-08-15 18:37:23 +02:00
scripts scripts: cover .tsx in the tracker-tag and comment-block lints 2026-08-11 21:01:26 +02:00
swagger-ui-theme move swagger-ui-theme/ out of hive-c0re/ 2026-08-02 21:24:57 +02:00
swarm-controller feat(swarm-controller): aggregate per-hive status from the swarm queue 2026-08-15 18:37:23 +02:00
swarm-nats-auth fix(swarm): the doc link pointed at a cfg(test) item 2026-08-15 09:34:33 +02:00
swarmctl feat(3216): swarmctl shell completions 2026-08-12 21:09:29 +02:00
.gitignore fix(review): drop libnull.rlib artifact + add Errors doc to ensure_config_pr_webhook 2026-07-11 12:19:52 +02:00
.mailmap chore(#2165): add damocles@pr1ma + lexis@pr1ma mailmap entries 2026-07-04 13:50:16 +02:00
.prettierignore docs: give turn-loop/ a README.md landing page 2026-08-03 12:55:18 +02:00
.prettierrc temp: add prettier configs 2026-07-02 23:33:11 +02:00
Cargo.lock feat(swarm-controller): aggregate per-hive status from the swarm queue 2026-08-15 18:37:23 +02:00
Cargo.toml chore(swarm): move swarm-nats-auth's deps to the workspace 2026-08-15 09:34:33 +02:00
CLAUDE.md swarmctl: add CLI reference docs, same pattern as hivectl 2026-08-11 21:55:56 +02:00
clippy.toml hivectl: wireguard mesh setup verbs (#1756) 2026-06-19 14:37:50 +02:00
flake.lock nix flake update 2026-07-13 13:58:53 +02:00
flake.nix wip: nix unit + secret delivery for the callout responder 2026-08-15 09:34:33 +02:00
README.md docs(readme): drop Rust function-name citation for plain-language behavior 2026-08-15 12:45:22 +02:00

hyperhive

a swarm of claude-code agents, each in its own nspawn cage, gossiping over unix sockets. config changes flow as git commits, the operator approves them in a browser, every deploy is a tag. cyberpunk-themed dashboard included. 💜

Claude code is great in one window, exponentielle across many — but only if you can keep the agents from stepping on each other, give them durable identity, and stop them from eating production. hyperhive is the substrate.

  • identity = unix socket
  • communication = sqlite-backed broker (send / recv / ask / answer / remind)
  • config = git (manager proposes, operator approves, deploys land as tagged commits)
  • blast radius = container
host (NixOS, runs hive-c0re.service)
│
├── operator
│   ├── browser → :80 (hive-gateway)    dashboard + per-agent UIs
│   │                                   /agent/<name>/ → per-agent unix socket
│   └── CLI     → /run/hyperhive/host.sock   admin protocol
│
├── hive-c0re  (Rust daemon: lifecycle / broker / approvals /
│               auto-update / dashboard / sockets)
│
├── optional containers
│   ├── hive-gateway   nginx — proxies :80 → c0re dashboard + per-agent sockets
│   ├── hive-forge     Forgejo — per-agent accounts, config mirror (agent-configs/)
│   └── hive-matrix    tuwunel — Matrix homeserver + per-agent accounts
│
└── agent containers
    ├── h-ruth     manager (privileged MCP surface, approval gating)
    └── h-<name>   sub-agent (claude + MCP tools + per-agent web UI + unix socket)

→ website · → docs · → options reference

Depth lives in docs/ (rendered at hyperhive.darkest.space/docs/) — start at docs/README.md and pick the page matching your task rather than reading front to back.

Quick start

Minimal flake.nix for a host that runs hive-c0re:

{
  inputs = {
    nixpkgs.url = "github:NixOS/nixpkgs/nixos-26.05";
    hyperhive.url = "git+https://forge.darkest.space/hyperhive/hyperhive";
    # Pin hyperhive to your own nixpkgs instead of the one it ships with
    # (see "Overriding nixpkgs" below) — recommended for most hosts:
    hyperhive.inputs.nixpkgs.follows = "nixpkgs";
  };

  outputs = { nixpkgs, hyperhive, ... }: {
    nixosConfigurations.my-host = nixpkgs.lib.nixosSystem {
      system = "x86_64-linux";
      modules = [
        hyperhive.nixosModules.default  # hive-c0re + hive-forge + hive-gateway in one import
        ({ ... }: {
          services.hyperhive.enable = true;
          # services.hyperhive.c0re.operatorPronouns = "they/them";  # default: "she/her"

          # ... rest of your host config
          system.stateVersion = "25.11";
        })
      ];
    };
  };
}

hive-c0re opens its admin socket + dashboard, auto-creates the manager container, and auto-rebuilds any container whose hyperhive rev goes stale. claude-code is unfree — hyperhive scopes the whitelist to itself, nothing for the operator to set.

Overriding nixpkgs

hyperhive pins its own nixpkgs so it builds standalone in CI. Add hyperhive.inputs.nixpkgs.follows = "nixpkgs" (as in the quick-start above) to build it against your host's nixpkgs instead — one less nixpkgs evaluation, no version drift from the rest of your system. Standard flake follows pattern; works as long as your channel is reasonably close to the nixos-26.05 hyperhive develops against. Drop it again if a much older/newer channel hits breakage hyperhive's CI doesn't catch.

For the full list of host and agent NixOS options see the options reference.

Operator CLI

hivectl is the operator-facing host CLI for ad-hoc administration that doesn't go through the broker (built alongside hive-c0re when the host module is enabled):

sudo hivectl forge create-user mara                       # provisions a forge user
sudo hivectl forge create-user mara --password 'hunter2'  # … with a fixed password
sudo hivectl matrix create-user mara                      # provisions a matrix user
sudo hivectl matrix create-user mara --password-stdin     # … reading one line from stdin

For a name that's a managed agent, hivectl persists the resulting token to that agent's state dir, the same as the boot sweep does. For a non-agent name (e.g. the operator's own forge/matrix account), it prints the token to stdout and writes nothing.

Build / deploy

nix develop -c cargo check
nix flake check        # rust + nix + toml fmt + clippy

# deploy from a host config that imports hyperhive.nixosModules.default
nix flake update --update-input hyperhive
sudo nixos-rebuild switch --flake .#<host>