Watch
0
0
Fork
You've already forked hyperhive
0
hyperhive/docs/swarm/README.md
atlas 318f67cda9 swarm-controller: create every agent's subagent stream
swarm-controller now creates `term-sub-<agent>` for every agent a hive is
declared to run, at start and every minute after, with the config
`swarm_queue_client::subagent_term::open_or_create` spells (subjects
`$SWARM.term.<agent>.sub.>`, max_age 24h). An existing stream is opened as
it is, as the controller does for its other streams and buckets, under the
`$JS.API.STREAM.CREATE.*` grant it already holds.

The agent no longer creates the stream: its token is granted publish on
`$SWARM.term.<agent>.sub.>` and no `$JS.API.STREAM.CREATE|INFO` subject,
and the subagent daemon only publishes. A `CREATE` carries the stream's
config in its payload, which no subject grant narrows, so the agent could
otherwise pick the stream's subjects and limits.
2026-10-03 01:34:01 +02:00

25 KiB

The swarm

The swarm is where things live: agent identities and accounts, secrets, the job graph that creates and places agents, telemetry, and the UI you drive it all from. Hives are the substrate — NixOS hosts that run agent containers on the swarm's behalf. Every hive belongs to a swarm; a single host is a swarm of one.

This page is for the operator. It covers the control plane, how a hive joins the directory, and how hives and agents report upward. The steps for a fresh swarm are in setup.md; the all-local config is the README quick start.

Option reference: services.hyperhive.swarm.* (swarm-wide facts, identical on every host) and services.hyperhive.deploy.* (whether this host runs grafana, the controller, authelia, …) → options reference, or nix build .#docs-swarm / .#docs-deploy.

Where each piece lives

  • control plane — swarm-controller: hive directory, agent roster, job graph, agent creation. → below
  • swarm UI — the operator's day-to-day surface, on the swarm apex, admins only. → ui.md
  • shared services — one forge, homeserver, SSO, queue and metrics/logs stack, each on whichever host you put it. → services.md
  • SSO — the secrets authelia generates and how each reaches its reader. → sso.md
  • secrets — every credential the swarm holds, who mints it and where it lives → secrets.md · the per-secret minter/reader/renewal contract → credentials.md · how the store comes up, who writes its grants and how it unseals → bao.md
  • swarm CA — the root every hive's internal TLS chains to, and what to hand a peer (trust-bundle.pem, never ca.pem). → ca.md
  • snapshot store — the swarm's one btrfs receive endpoint, swarm.snapshotStore.{address,port}. → snapshot-store.md

Swarm controller

swarm-controller is the swarm's control plane: it holds the hive directory, the agent roster and their wanted state, and the job graph. Creating an agent — from the swarm UI or swarmctl agent create --hive <h> — queues its SSO identity, forge user, config repo, store identity and matrix account, then sends the hive a deploy message. A new agent starts paused.

services.hyperhive.deploy.swarm-controller.enable runs it on this host; singleHostSwarm turns it on. Otherwise off: set it on the one host that runs the controller. A swarm has one controller, so enabling it states a fact about swarm topology, not about whether this host runs a hive.

What it serves, why it's a unix socket rather than a port, and the socket-directory constraint that governs where socketPath may point: swarm-controller/README.md.

Per-hive status (GET /api/hives/status)

One row per hive in swarm.hives, saying when it last reported and what it said. Hives publish upward through the swarm queue; the controller never reaches down to collect, so a hive that can't reach the swarm still knows its own state — you just can't see it from here.

A hive publishes only once it holds all three status-publish coordinates below. A hive without them reads never_reported — it's not broken, it just has nothing to say upward.

freshness what to do about it
fresh nothing — reported within staleAfterSeconds
stale the hive stopped reporting. Its last payload is still shown, so check age_seconds and the payload for what it managed to say
never_reported this hive has never reported at all — normally a deployment that hasn't happened, not an outage
unknown something is publishing under a name that's not in swarm.hives — a typo in the roster, or a hive removed from the roster but still running

Every row also carries last_seen_unix and age_seconds if you want to apply your own threshold. The timestamp is the one the queue recorded on arrival, not one the hive put in its own payload.

Set services.hyperhive.swarm.controller.staleAfterSeconds (default 120) above the rate hives publish at, or everything reads stale between reports. Hives publish once a minute, so the default tolerates one missed report and flags two. It takes effect on the next request; nothing has to re-publish.

Making a hive report

Three options on the hive. They sit in two namespaces, because two of them are facts about this machine and one is the swarm's single address:

option what to set it to
deploy.hive-controller.statusPublish.natsUrl tls://<swarm.nats.domain>:<swarm.nats.port>
swarm.statusPublish.tokenEndpoint the swarm IdP's /api/oidc/token
deploy.hive-controller.statusPublish.clientSecretFile path to this hive's client secret, plaintext

The queue URL and the token endpoint default to the swarm's own addresses on every hive, so there is nothing to set for them. The secret is what turns publishing on: a hive without it doesn't publish. A secret without the other two is an eval error rather than a hive that quietly never reports.

On a host that runs the queue and the IdP, the secret defaults to the local one. Any other hive needs the secret to physically be there, because the swarm doesn't distribute it. Copy hive-<hiveName>.secret out of the swarm host's deploy.authelia.hostClientSecretDir with whatever secret management the deployment already uses.

The queue URL is the same string on every hive. The queue accepts TLS only, with a certificate for swarm.nats.domain (default nats.<swarm.domain>) and no other name or address, so a URL with an IP address or nats:// fails. The queue's host resolves the name itself. A multi-host swarm needs one upstream DNS record, nats.<swarm.domain> pointing at the queue host's mesh address, the same contract as bao.<swarm.domain>. The queue's port is open on wg-hive when that host is on the mesh.

The identity isn't a choice — a hive authenticates as hive-<hiveName> and publishes under hiveName, the same name that keys swarm.hives.

If a hive stops reporting, its own dashboard is the place to look: a failure to publish raises a warning banner there after three consecutive misses. It stays warn rather than crit on purpose — a hive that can't reach the queue isn't itself unhealthy, so it doesn't start calling itself degraded for being unable to say it's fine.

The endpoint answers 503 when this host has no swarm queue configured, or has one and can't read it — deliberately not an empty list, which would look like a silent swarm rather than a controller that can't see. The body says which. Status survives a controller restart: it's stored in the queue, not in the daemon.

Giving the agents the queue too

An agent authenticates as its own client, not as its hive, so it needs its own coordinates. Two of them are options on the hive; the other two arrive with the credential itself and aren't configurable.

option what to set it to
deploy.hive-controller.queue.agentNatsUrl where the queue listens, as an agent container reaches it
swarm.statusPublish.tokenEndpoint the swarm IdP's /api/oidc/token — the same one the hive uses

On every hive, agentNatsUrl defaults to the same tls://<swarm.nats.domain>:<swarm.nats.port> as the hive's own. On the queue's host the name resolves inside a container to the bridge address, where the firewall opens the port; on any other hive it resolves through the host's DNS, like the store's name. ⚠️ Never a loopback address here: an agent has its own network namespace, so 127.0.0.1 reaches the agent.

The harness sees four variables, and treats them as all-or-none: HIVE_AGENT_NATS_URL and HIVE_AGENT_OIDC_TOKEN_ENDPOINT from the two options above, plus HIVE_AGENT_OIDC_CLIENT_SECRET_FILE and HIVE_AGENT_OIDC_CLIENT_ID_FILE, which point into the unit's own credentials directory. The last two come from the delivered credential rather than from config — see secrets.md for how it gets there. A hive lacking the queue's address for its agents sets none of the four and each agent logs that it has none; a half-set environment logs an error and the harness keeps serving.

What an agent publishes over the queue

What an agent does with that connection is publish its terminal. Every row its own web UI renders also goes to $SWARM.term.<agent>, one subject per agent, so a swarm-level terminal can follow one agent without subscribing to the swarm's whole traffic. An agent connected with its own queue credential gets that subject. An agent without one, or whose own credential the queue refused, connects with its hive's shared client and publishes to $SWARM.term.<hive>.<agent> instead, the <hive> being the one that client id names. The swarm controller relays both. Publishing only: an agent talks about itself here and reads nothing. Rows aren't retained — a subscriber that wasn't listening missed them, the same as on the agent's own live stream.

An agent's subagents publish their terminals too, from the agent's subagent daemon: each subagent's rows go to $SWARM.term.<agent>.sub.<subagent>, classified the same way. The queue grants that family to the agent's own queue credential alone, so the daemon reads it from the store under the agent's store identity, exactly as the harness does, and publishes nothing without it. The queue keeps these rows in the stream term-sub-<agent> for 24 hours, and the swarm lists an agent's subagents from that stream's subjects. The swarm controller creates that stream for every agent a hive's wanted state names, within a minute of the agent appearing there. The agent's grant is publish on its own $SWARM.term.<agent>.sub.> and no JetStream subject, so the stream's subjects and limits are the controller's and never the agent's. Subagents publish output only and read nothing.

The queue would refuse a row too large for its max_payload outright and take the connection down with it, so the harness drops such a row's body before sending and leaves a marker in its place; the summary, level and icon still arrive. The harness logs and skips a row that's too large even without its body.

The second thing an agent publishes is its turn-state header, on $SWARM.agent-state.<agent> (or $SWARM.agent-state.<hive>.<agent>) — same shape of subject, same grant mechanics, same lack of retention. It carries what a header bar wants: what the turn loop is doing (turn_state, plus turn_state_since as an ISO 8601 UTC stamp), which model (model and the resolved id the last turn actually ran on), the context budget and the last turn's context and cost token blocks, and agent_state.

agent_state reuses the swarm's own wanted-state vocabulary (up/offline/paused/destroyed) so a reader can compare what an agent is against what the swarm declared it should be without translating between two spellings. ⚠️ From inside the container only two of those four are sayable: the harness reports up, or paused when the pause marker is present. offline and destroyed are hive-c0re's observations — a stopped agent publishes nothing and a destroyed one doesn't exist — so a view that needs the full four-state picture takes them from the agent-status bucket and uses this subject to sharpen the rest.

Headers go out on transition, not on a timer: the harness rebuilds the header whenever its event bus moves and publishes only when the result differs from what it last sent. That's the whole point of the subject — the agent-status bucket already republishes once a minute, which is far too slow for "is this agent thinking right now." The cost of a core subject is that a subscriber attaching mid-idle sees nothing until the next change, so a renderer opens with the bucket's snapshot and lets this stream refine it.

Swarm-side, GET /api/agents/<name>/state/stream relays that subject as SSE, resolving the agent's hive at request time exactly as the terminal stream does. The payload passes through opaquely — the controller never parses a header.

The third thing an agent publishes is its icon, the same SVG its own GET /icon serves. It goes into the agent-icons KV bucket under the key <agent>, with no hive in it, so the swarm can show the icon of an agent that's stopped or has moved hives. Only an agent connected with its own queue credential publishes it: the queue grants that credential $KV.agent-icons.<agent> and no other key, and grants the hive's shared client none of the bucket. The harness writes once per start, because a config change reaches an agent by restarting its container. An agent with no icon deletes its key. Swarm-side, GET /api/agents/<name>/icon serves the stored bytes, and 404 means the agent has no icon. swarm-ui's agent cards load it as an <img> and show the dimmed hyperhive mark for an agent without one, as the hive dashboard does.

Swarm-wide forge objects

The controller also keeps the forge objects that are one per swarm, not one per hive. It ensures them at start and every five minutes after (swarm-controller/src/forge/objects.rs):

  • the orgs agent-configs, internal and agents, plus each mirror's owner org;
  • the empty operators merge-gate team in agents and agent-configs;
  • the main merge gate on every agent-configs repo: merge and approval whitelists = the operators team, and no user. The controller leaves a repo with no main rule alone;
  • the pull-mirrors from deploy.forgejo.mirrors on the controller's host (with the actions/checkout one deploy.forgejo.ci.enable adds);
  • internal/docs (private) and internal/knowledge (public, with a seed README.md while empty);
  • the agent-configs org avatar (deploy.swarm-controller.configOrgAvatarPng).

A pass that can't finish logs a warn line per object plus swarm forge objects: pass incomplete in journalctl -u swarm-controller, and retries on the next tick. While the controller is down the objects stay as they are.

Swarm-wide forge webhooks

At startup the controller registers two Forgejo hooks pointing at itself — a push hook on internal/knowledge and a pull_request hook on the agent-configs org, both under https://<swarm.domain>/webhook/forge/.

The controller interprets a delivery and sends hives a specific message — the knowledge repo changed, deploy agent foo at rev abc123 — rather than forwarding forge payloads for each hive to re-derive. Approval happens once, at the swarm level: a hive receives a decision, not an event to adjudicate.

A config-pr delivery reporting a PR merged into main queues a deploy of its merge_commit_sha on the one hive whose wanted state places the agent; with no such hive, or several, the controller deploys nothing. See approvals.md § Config changes. A merge whose closed delivery never arrives deploys nothing either; the operator recovers by redelivering that delivery from the forge hook page, which re-enters the same handler and re-queues the deploy.

internal/knowledge is on that path. The controller's is the only hook on it (knowledge.md covers clearing a leftover). A webhook has exactly one target URL, so a second registration would take delivery away from the first rather than add a recipient.

The agent-configs org is on it too. The controller's hook is the only one that acts on config PRs. Each hive deletes its own /webhook/config-pr org hook on startup, since a route no hive serves would otherwise sit on the org collecting failed deliveries. Removing the controller's hook stops forge-UI merges from deploying until its next start recreates it.

Nothing to configure. The controller registers the hooks only when this host also serves the swarm UI vhost — that's what publishes the endpoint, and a hook the forge can't reach would collect failed deliveries while looking healthy. The controller generates the HMAC secret on first start and keeps it (see docs/agent-lifecycle/persistence.md).

To check it's working, push to internal/knowledge and look for webhook: verified delivery in journalctl -u swarm-controller. A refused delivery logs webhook: refused delivery with the reason.

Hives: the substrate

  • hive — one host running hive-c0re and its agent containers (deploy.hive-controller.enable). Addressed as <hiveName>.<swarm.domain>.
  • swarm — every hive in services.hyperhive.swarm.hives, plus the shared services and controller. You can qualify an agent as agent@hive-domain.
  • peer hive — any hive in the directory other than this one. Peers are derived, not declared: the directory lists every hive including yourself, and hiveName says which one you are.

Hive identity config

services.hyperhive = {
  swarm.domain = "example.com";     # required — the swarm's DNS domain
  hiveName = "pr1ma";               # required — this hive's label in it
  swarm.name = "constellat1on";     # shared swarm display name (optional)

  # required — the directory, identical on every host in the swarm.
  # Names only: each entry's `domain` defaults to <name>.<swarm.domain>.
  swarm.hives = {
    pr1ma = { };                                     # this host, per hiveName
    lab   = { };                                     # a second hive
    edge  = { domain = "edge.elsewhere.example"; };  # addressed off-convention
  };
};

Swarm options are identical on every host in the swarm, so swarm.domain is required on every host that runs any hyperhive service; a host that imports the module and enables nothing evaluates without it. hiveName is required on a host that runs a hive, the secret store or the homeserver. Eval fails with a hint naming each. Neither defaults, because a guessed value here is a wrong hostname that evaluates cleanly and deploys.

The directory is one attrset, identical on every host. It describes every hive in the swarm, including this one, keyed by hiveName; what differs between hosts is only hiveName. It must contain an entry for this host's hiveName, and an empty directory fails that check too — eval names the missing hive. Peers are every entry but this host's, so a directory without this host would make every hive a peer, itself included.

# hive A                          # hive B
hiveName = "pr1ma";               hiveName = "edge";
swarm.hives = { … };              swarm.hives = { … };   # byte-identical

One entry per hive means two hosts can't hold different facts about the same third hive, such as a stale endpoint.

domain defaults to <name>.<swarm.domain>, so a conventional directory is names only. Set it only for a hive addressed by something else. This hive's own services.hyperhive.domain comes from its entry; it drives HYPERHIVE_HIVE_DOMAIN in every container so agents can form qualified labels (iris@pr1ma.example.com). Setting services.hyperhive.domain directly overrides the directory entry, with a deprecation warning: the option is local to this host while every host shares the directory, so a value written only here is invisible to the rest of the swarm.

swarm.name is display only — the dashboard chrome header and per-agent system prompts — and federated hives at different domains can share one. hiveName surfaces in the same places but is also the leftmost label of the hive's domain. swarm.name names the group; hiveName names this hive.

The env-var chain and qualify() / qualified_label() semantics: conventions.md § Hive identity.

Trust inside a swarm comes from the swarm root (ca.md): every hive chains to it. Trusting a hive whose root this swarm doesn't own — another swarm's — has no mechanism.

What the directory feeds

  1. Swarm-wide hive roster — swarm-controller reads this same directory and serves it at GET /api/hives; the swarm UI's front page renders it (ui.md).

  2. Matrix federation — when this host runs the homeserver (deploy.matrix.enable), tuwunel federates with the peer's matrix server (discovered via the peer's .well-known/matrix/server delegation, which the gateway serves). Federation validates the peer's TLS certificate against the matrix container's trust bundle, independent of this directory.

    ⚠️ On a hive with self-signed gateway certificates, the container also trusts this hive's trust-bundle.pem, bound in at runtime and added to the public CAs; it ends at the swarm root, so a peer whose certificate chains to the same root validates. A peer outside this swarm's root needs a CA-issued certificate (ACME). Federation firewall and TLS requirements: integrations/matrix.md.

  3. WireGuard mesh (optional) — deploy.wireguard.enable reads each entry's wireguardPublicKey/wireguardEndpoint/wireguardAddress to configure wg-hive. See below.

WireGuard inter-hive mesh (optional)

The peer config above uses public HTTPS for all inter-hive traffic. For private deployments — or to reduce latency and TLS overhead on intra-swarm traffic — hyperhive can configure a host-to-host WireGuard mesh.

Generating keys

On each hive host:

wg genkey | install -m 0400 /dev/stdin /etc/wireguard/hive.key
wg pubkey < /etc/wireguard/hive.key   # → share this with peer operators

Config example (two hives)

# hive A (pr1ma.example.com, mesh IP 10.100.0.1)
services.hyperhive = {
  deploy.wireguard = {
    enable        = true;
    privateKeyFile = "/etc/wireguard/hive.key";
    address        = "10.100.0.1/24";
    listenPort     = 51820;          # optional, default 51820
  };

  # The same `hives` attrset both hosts hold — mesh fields included,
  # since "where this hive can be dialled" is a fact about that hive.
  swarm.hives = {
    pr1ma = {
      domain             = "pr1ma.example.com";
      wireguardPublicKey = "base64keyA=";
      wireguardEndpoint  = "198.51.100.1:51820";
      wireguardAddress   = "10.100.0.1/32";
    };
    edge = {
      domain             = "edge.corp";
      wireguardPublicKey = "base64keyB=";
      wireguardEndpoint  = "203.0.113.42:51820";
      wireguardAddress   = "10.100.0.2/32";
    };
  };
};

# hive B (edge.corp, mesh IP 10.100.0.2)
services.hyperhive = {
  deploy.wireguard = {
    enable        = true;
    privateKeyFile = "/etc/wireguard/hive.key";
    address        = "10.100.0.2/24";
  };

  swarm.hives = { /* … identical to hive A's … */ };
};

What the mesh does

  • hyperhive configures networking.wireguard.interfaces.wg-hive on the host (not inside agent containers; containers reach peers via the host's routing table).
  • It opens UDP port 51820 (or listenPort) on the host firewall.
  • swarm-wireguard.nix reads each entry's wireguardAddress directly from services.hyperhive.swarm.peerHives to build wg-hive's allowedIPs, so intra-swarm traffic can route over the mesh address rather than the public domain.
  • It sets persistentKeepalive = 25 by default; override or null to disable (not needed when both sides have public IPs and no NAT).

NAT / one-sided endpoints

If one host is behind NAT and can't accept incoming connections, only that host needs a null wireguardEndpoint on the peer config — the other side initiates. With keepalive on, the NAT hole stays open.

If both hosts are behind NAT, you need a STUN relay or a third host (exit node); hyperhive sets up neither.

Cross-references