hyperhive/docs/snapshot-store.md
atlas 433b294099 refactor(nix): swarm.peers becomes swarm.hives, a directory of every hive
One attrset describing every hive in the swarm including this one,
identical on every host, with hiveName selecting which entry is us.
"My peers" is derived (swarm.peerHives) rather than declared.

Every field in the old per-host peer list was intrinsic to the hive it
described, never to the pair -- so the list was a directory each host
kept its own copy of. Beyond the deduplication it removes a bug class:
two hosts could hold different endpoints for the same third hive with
nothing to detect the disagreement.

Drops the per-hive caCert. Trust inside a swarm derives from the swarm
root, which every hive chains to. What that genuinely removes is
trusting a hive whose root this swarm does not own -- a cross-swarm
problem that wants a mechanism of its own, not a field that happened to
work.

The matrix container's certificateFiles block goes with it and could
NOT be migrated: that list is read at build time and the swarm root is
a runtime file (its key must never enter the store), so there is no
build-time name to put there. caCert being a nix path was precisely
what made it the build-time distribution channel. Agents are unaffected
-- hive-tls folds the root into the hive trust bundle and the meta
renderer embeds that one file. Tracked separately.

Migration is an assertion plus warnings, not a rename: hives is peers
union {self}, and the set gains a member no existing config has written
down. A rename migrates a name and a default can re-root a meaning;
neither can conjure a new member. The warning explains, the self-entry
assertion stops the build.
2026-08-05 20:44:16 +02:00

8.8 KiB

Snapshot store

The swarm's btrfs receive endpoint. Hives push agent snapshots to it over the WireGuard mesh; a destination hive later pulls one back to complete a migration.

Two things it is not, both worth stating because both are easy to assume:

  • It is not the swarm controller, and does not depend on one. It is a NixOS host role: a btrfs subvolume tree, a socket-activated receiver, and the wg-hive interface the swarm module already brings up. That is why it can be deployed before any controller exists.
  • It is not a backup product. It happens to hold the data a backup would hold, and it should be operated accordingly (see Operating it) --- but nothing in it does scheduling, verification, or restore orchestration.

Enabling it

services.hyperhive.snapshotStore = {
  enable = true;
  path = "/var/lib/hyperhive-snapshots";  # must be on btrfs
  port = 51821;
};

# The mesh is a hard requirement, and is asserted:
services.hyperhive.swarm.wireguard = {
  enable = true;
  address = "10.100.0.9/24";
  privateKeyFile = "/etc/wireguard/hive.key";
};

The store host is a swarm member like any other: it gets an entry in services.hyperhive.swarm.hives, the same directory every host holds. See swarm/ for the mesh itself.

Note that the mesh is gated on swarm.wireguard.enable, not on c0re.enable --- a store host runs no hive and would otherwise get no wg-hive interface at all.

Pointing a hive at it

The block above configures the host that receives. Every hive that pushes separately needs to be told where the store is:

services.hyperhive.swarm.snapshotStore = {
  address = "10.100.0.9";  # the store's mesh address, no prefix
  port = 51821;            # optional; must match the receiver's port
};

Two deliberate asymmetries in that pair, both easy to misread as inconsistency:

  • address has no default. It is a deployment fact a pushing hive cannot derive, and a wrong guess means streaming an agent's state at whatever happens to answer. Unset, a push fails naming this option.
  • port does default (51821), because it is a convention both ends read from the same option docs --- a default there is coordination, not a guess.

Note the option lives under swarm.* while the receiving host's lives under services.hyperhive.snapshotStore. That is the distinction the two namespaces carry throughout: swarm.* describes the swarm as seen from this host, and a bare services.hyperhive.<service> describes a role this host performs. A store host sets both --- one to run the receiver, one only if it also runs a hive that pushes.

With it set, hivectl agent <name> subvol snapshot push <label> [--parent <label>] streams a snapshot straight into the store. There is no destination argument, because a swarm has exactly one store (see One subvolume per agent, not per hive), and no credential argument, because the mesh is the authentication.

The mesh is the authentication

There are no certificates here, and no key material of its own. That is deliberate rather than an omission.

WireGuard's cryptokey routing already binds a peer's source address to its public key: the swarm module configures each peer with allowedIPs = [ peer.wireguardAddress ], so a packet arriving from that address provably came from the holder of that private key. A packet that reaches the receiver has therefore already been authenticated by the kernel.

Layering TLS client certs on top would authenticate the same fact a second time, and add a credential with an expiry --- a migration that fails because a renewal quietly didn't happen, discovered on the day you need to move an agent.

One subvolume per agent, not per hive

The destination is keyed by agent.

This is not cosmetic. After a migration, an agent's next incremental send arrives from a different hive than the previous one. Keying by hive would split that agent's snapshot chain across two directories, and btrfs send -p would fail to find its parent --- breaking exactly the case the store exists to serve.

What the sender can and cannot choose

A btrfs send stream carries no notion of which agent it belongs to, and the subvolume name inside it is chosen by the sender. So the protocol is one agent <name> header line, then the raw stream.

The rule that matters:

The receiver owns the destination root. The sender-supplied name is validated, never used as a path.

Validation is a whitelist --- [A-Za-z0-9_-]+ and nothing else. No slash and no dot means neither directory traversal nor an absolute path can survive it. It is deliberately a whitelist and not a list of forbidden characters: a blocklist only ever excludes the attacks somebody already thought of.

Reachability

The receiver is socket-activated, and the socket binds this host's mesh address, never a wildcard. Both the mesh being enabled and the address being set are assertions, not documentation --- bound to 0.0.0.0 this socket is an unauthenticated remote write into agent state.

Binding is not sufficient on its own. NixOS's firewall is default-deny and filters in netfilter, before a packet reaches a bound socket, so the port is opened explicitly --- and scoped to the mesh interface:

networking.firewall.interfaces.wg-hive.allowedTCPPorts = [ cfg.port ];

A host-wide allowedTCPPorts would open the port on every interface including a public NIC, leaving only the socket's bind address between the internet and a root btrfs receive.

Operating it

Confinement is the deployment's job

btrfs receive needs CAP_SYS_ADMIN, so the receiver runs as root. The unit sets ProtectSystem=strict, ProtectHome, PrivateTmp and a narrow ReadWritePaths --- but those are defence in depth, not a boundary: a process holding CAP_SYS_ADMIN can call mount(2) and undo the namespace they set up.

The boundary is the machine. The intended deployments are:

  • a swarm: the store is its own small VM. The machine is the boundary, which is stronger than anything the unit could assert about itself.
  • all-in-one / local: the store runs as a container on the c0re host.

The second is worth keeping deliberately, and not only for convenience: it means the confined path is exercised by every local deployment. The usual failure mode for an isolated variant is that nobody runs it day to day, so it rots and is discovered broken in production.

⚠️ The assumption to keep true over time: the store host runs nothing else. That is true on day one and quietly false the day someone notices the box has spare disk. Nothing in the config objects when it stops being true.

It holds every agent's state from every hive

Which makes it the highest-value target in the swarm by a wide margin, and means it should get the treatment a backup host gets --- restricted access, and a decision (rather than an omission) on encryption at rest.

The trap is the label: this box holds backup-grade data while not being called a backup, so it can end up with backup-grade exposure and non-backup-grade controls. Nobody puts a migration staging area on the access-review list.

What a snapshot contains

The snapshot covers an agent's state subvolume, which is the parent of state/, claude/ and harness/. Consequences:

  • The Claude session (claude/) travels, so a restored agent keeps its live --continue session rather than needing to log in again.
  • harness/ travels too, including harness/bash-tasks/. Task output is part of an agent's working continuity, so this is wanted --- but it means anything that has ever leaked into a task's captured output is in the retained snapshots as well.

It does not cover the agent's applied config (/applied/<name>/) or its topology entry, both of which live outside the subvolume. A restore therefore yields an agent's memory without its definition; closing that gap is tracked separately.

Retention

Retention lives on the sending side (last-N by count, swept periodically), not here. Count rather than age is deliberate: a count is bounded by construction, whereas an age policy silently scales disk usage with how hot a hive runs.

Per-agent or per-hive btrfs qgroup quotas are not configured yet. Without them one runaway hive can fill the store and take out every other hive's snapshots.

Not built yet

The pull side. Push is safe with minimal authorisation because a hive can only ever write to a chain it owns. Pull is the direction that needs a policy: unrestricted, any compromised hive could read every agent's state from every other hive. It needs a notion of which hive currently owns which agent, and that ownership record lands with the swarm controller work.

With a single hive the question is trivial --- the only peer owns everything it sends --- which is why the receive half ships first.