The swarm collector's OIDC client secret only existed where authelia did: `swarm-otel-oidc-secret.service` copied the minted plaintext out of authelia's container tree, reachable only because the two share a host's network namespace. A swarm that placed authelia elsewhere delivered nothing, and the option's own description said so — "a deployment that places authelia elsewhere points this at a file it delivers itself." Same gap as #3853 and #4234, and this is the swarm-otel twin of #4234's fix for Grafana. Mirrors PR #4361 (Grafana) almost exactly: - `swarm-bao-otel-oidc.service` reads `swarm/services/<client-id>/oidc/client` out of the store, in every deployment, replacing the co-located copy unit outright — one delivery route, not two, per the ruling that landed under #4234. - Client registration moved out of `swarm-otel.nix`'s own `config` block (gated on this host running the collector) into `glue-swarm-otel-oidc-client.nix` (gated on this host running authelia), the same split `glue-grafana-oidc-client.nix` made. It was broken the same way: a split deployment registered the client nowhere at all, so authelia never minted a secret for the publisher to send on. - The publisher's `services` prefix (write grant in `swarm-bao.nix`, hive read grant in `policy::render`) already covers any service's path — nothing to add there. `swarm-secret-publisher.nix` only grew `serviceClientIds` by one entry. One judgement call, stated rather than buried: the store-reading unit renders only where this host holds a client identity (`deploy.bao.clientCertFile`/`clientKeyFile`), rather than asserting it the way `swarm-grafana.nix` does. Grafana's local login form is disabled unconditionally, so a Grafana with no OIDC secret has no way in at all — that earns a hard refusal. This collector without a credential still receives every hive's telemetry; only its own pushes to the stores go out unauthenticated and get refused there, an already-supported degrade the module's own `haveCollectorSecret` flag named before this change. So the reading unit follows the shape `glue-matrix-bao-token.nix` and `glue-queue-agent-credential.nix` use for their own optional readers: no unit when the identity is absent, not a build refusal. Fixtures mirror #4361's: `otelBaoWithAuthelia`/`otelBaoRemoteAuthelia` are the positive pair (co-located and split, both reading through the store), `otelNoIdentity` is the negative — no reading unit, no assertion firing, `clientSecretFile` left null. Refs #4258
27 KiB
Security model
Agent trust model
The sections below document specific mechanisms (the state-file endpoint, nixbld isolation, privilege separation). This section frames the model they serve: what hyperhive defends, what it deliberately doesn't, and where the operator is accepting risk. It's the reference for "is it safe to give an agent capability X?."
The trust boundary is the container, not credential storage
An agent is trusted code running inside its own nspawn container. The
boundary that matters is the container: a sub-agent can't see the host
netns, another agent's container, or another agent's state dir. Within its
own container the agent is privileged — it has passwordless sudo by
default. Isolating credentials from the agent itself is therefore not a
goal: an agent can read its own tokens, its own /home/<name>/.claude, and
run arbitrary commands as root inside its container. (The narrow exception is
cross-tenant leakage — for example the unsandboxed-nix-build 0600 token policy
below stops a build's nixbld user reading the agent's own forge token, and the
state-file endpoint stops one agent proxying another's files. Those harden the
boundary; they don't sandbox the agent from itself.)
The corollary: don't reason about security as "can the agent be stopped from touching its credentials." Reason about it as "what's the blast radius if this agent does the worst possible thing with everything it can reach."
Scoped tokens bound the blast radius
Each agent gets its own scoped credentials, never shared:
- forge token → that agent's Forgejo account only (its own repos +
collaborator grants; can't act as another agent or as
core). - matrix token → that agent's matrix account only.
Its own account's scope bounds a compromised/confused agent's reach on the forge or matrix, not the swarm's. This is the main thing standing between "one agent does something dumb" and "the whole hive is affected."
Identity vs. secret (matrix). The scoping is on the secret, not the
identity: an agent's matrix token is private to its own account, but its
matrix identities — the public handles (name, user_id @user:server,
homeserver) — are intentionally readable by any agent via GetAgentMeta, so
peers can find and address one another on a shared matrix instance. Only the
public handle crosses that boundary; the token never does.
The swarm secret store isn't a boundary between hives
A hive reads its agents' credentials out of the swarm's secret store with its own certificate. Every hive's policy grants read on every agent's credentials, not only on the agents it hosts — so a compromised hive can read the matrix token of an agent running on a different hive.
That's deliberate and it's the interim state, not the intent. An agent's credential path doesn't name the hive hosting it (agents move between hives), so a per-hive grant has to be an enumeration the controller re-emits whenever the roster changes — and an enumeration that can drift or land out of order advertises a boundary it doesn't actually hold. A wide grant that says what it is beats a narrow one that only looks narrow.
What still holds: the grant is read-only (a hive can't write an agent's credential, so it can't hand itself an agent's identity), and it reaches three prefixes and nothing else in the store — every agent's credentials, the reader's own entry under the hive namespace, which names the hive asking and so widens nothing between them, and every swarm service's OIDC client secret.
That third prefix has the same shape of reason as the first, and the same honest
cost. A swarm service (Grafana and the swarm collector are the two there today)
registers one client for the whole swarm, so its credential's path names
the service and never the host — and which hive runs a given service is a
deploy.* fact, per-host by definition, so nothing swarm-wide exists to scope
the grant to. The host running such a service has no store identity of its own
either; it reads with the certificate of the hive it is. Any hive can
therefore read any swarm service's client secret, which is worth what it
buys: every swarm service gets its credential the same way from anywhere,
instead of only where the identity provider happens to sit. Giving such a
service its own store identity is what would remove this rather than re-scope
it.
A tracked follow-up narrows this, with the two candidate directions: scope the grant per hive (and pay for the re-emission), or give each agent container its own store identity so credentials never pass through a hive at all.
Threat model: prompt injection → confused deputy
The realistic adversary never needs to breach the container. They supply untrusted input the agent reads and acts on: a poisoned issue or PR comment, a cloned repo's README/CI, a scraped webpage, a crafted matrix message. The agent is the trusted, capable party; the input is the untrusted part. A successful injection turns the agent into a confused deputy — it uses its legitimate capabilities (push, comment, deploy, run shell) on the attacker's behalf.
Mitigations are therefore about bounding capability and inserting human checkpoints, not about sandboxing the agent from its own tools:
- Operator merges, not the agent — an agent may push branches, but a
human (the operator) merges the PR, keeping a person in the loop on the
highest-value action. On the internal forge this is technically enforced,
not just convention: agents can't create repos (
max_repo_creation = 0), so every repo iscore-created with branch protection on by default — merges restricted to the operators team + a required operators-team approval (apply_operator_branch_protection/ the config-repo equivalent) — and an agent (a write collaborator, not a repo admin) can neither change those settings nor merge its own PR. It's not set up for external VCS (GitHub etc.), though — there, operator-merge is process + accepted risk, not a technical control. - Approvals — config changes, schedule additions, and other
blast-radius-y operations route through the operator approval queue
(see
approvals.md).
Capability = accepted risk
Every capability granted to an agent is a risk the operator is explicitly accepting. The rule of thumb:
Don't give an agent access to something you can't afford to lose.
If an agent can deploy to prod, you are accepting the risk of a dropped production database (via injection or plain error). If that's unacceptable, the answer is don't grant the capability — not "grant it and hope the sandbox holds," because there is no sandbox between an agent and the tools you handed it.
No autosandboxing of external tokens
hyperhive provisions and scopes its own per-agent forge + matrix tokens. It does not automatically sandbox or scope external credentials (GitHub PATs, cloud keys, third-party API tokens). The scope of an external token is operator-accepted risk: if you drop a broadly scoped GitHub token into an agent's config, that agent has exactly that reach, with no hyperhive layer narrowing it. Scope external tokens tightly at the source (the external provider) before handing them over.
State-file endpoint security model
GET /api/state-file?path=<p> serves files from agent state dirs and
the shared space to authenticated dashboard users (browser, operator).
It accepts two allow-listed root prefixes and rejects all other paths
before touching the filesystem:
/var/lib/hyperhive/agents/<n>/state/— per-agent durable notes (canonical host form or the in-container view/agents/<n>/state/)/var/lib/hyperhive/shared/— shared docs (/shared/in-container)
/state/... without an agent prefix is explicitly not accepted — it's
ambiguous from the host's perspective.
Defense-in-depth layers (in order):
- Allow-list prefix check — rejects without touching the filesystem if the path doesn't match either root.
- No symlinks below the matched root — each path component is
checked with
symlink_metadatabefore canonicalize. A sub-agent that plantsln -s /other/secret /agents/me/state/peekcan't proxy another agent's file through this endpoint (canonicalize would happily resolve the symlink to a still-within-allow-list path). - Canonicalize as belt-and-braces — resolves
../.traversal and rejects if the result escapes the roots. state/subdir constraint — underAGENTS_ROOT, the second path component must bestate/. Applied, proposed git repos and config dirs are off-limits.- World-readable check — file must have
mode & 0o004set. A0600file insidestate/would otherwise be accessible to any operator with dashboard access.
scan_validated_paths (broker-message ingest, linkifier) uses the same
resolve_state_path helper so security rules stay in sync — the
dashboard renders anchors only for tokens that passed the same checks the
read endpoint enforces.
The same invariant holds wherever an agent-supplied name reaches a filesystem
path: the agent socket's GetAgentMeta takes name as a serde-validated
hive_types::Ident (or falls back to Ident::parse for the "self" case)
before building agent_notes_dir(name), so a .. component can't traverse.
Nix builds and credential isolation
Background
Agent containers bind-mount the host's nix-daemon socket. The host daemon may
have sandbox-fallback = false (strict NixOS defaults), which causes nix build
inside nspawn containers to fail — containers lack kernel user namespaces, so nix
can't set up its build sandbox. the agent modules set sandbox-fallback = true
so that builds fall back to unsandboxed execution rather than failing outright.
Threat model
Unsandboxed nix builds run as nixbld users (non-root, typically UIDs 30001-30010).
Without sandbox isolation, a build derivation's builder script has read access to
any file in the container that the nixbld user can read.
The blast radius also has a network dimension. hive-ci runs its unsandboxed
builds of untrusted PR code in its own private netns behind the hive bridge: a
build reaches the forge only through the gateway and can't reach host-loopback
services — including the core dashboard at 127.0.0.1:<dashboard_port>, which
has no application-layer auth of its own (see docs/scheduler/ci.md). The 0600
token policy bounds file reads; network isolation bounds network reach.
What's NOT exposed:
/home/<name>/.claude/—.credentials.json,history.jsonlandsettings.jsonare0600;projects/andsessions/are0700. All owned by the per-agent user<name>, so nixbld users can't read any of them. ⚠️ The directory itself is0755, on purpose:hive-coreis a different user and needs read+execute to list it soclaude_has_sessioncan detect a valid session (ensure_claude_dir,hive-c0re/src/lifecycle/setup.rs). The per-file mode is therefore the whole protection here — anything added to this directory at a default mode is world-readable, which isn't hypothetical:plugins/and.last-cleanupalready are.$HYPERHIVE_STATE_DIR/forge-token(=/agents/<name>/state/forge-token) — written at mode0600and chowned to the per-agent uid:gid (seehive-c0re/src/forge/mod.rs's module doc for exactly where). nixbld users can't read it.
Policy: all credential files written to agent state directories MUST be mode
0600 or stricter. Don't create world-readable secret files in agent state dirs.
Long-term fix
The proper fix is to enable user namespaces inside nspawn containers
(--private-users=inherit in EXTRA_NSPAWN_FLAGS) so nix can set up its real
sandbox and sandbox-fallback becomes a true last resort. This requires verifying
bind-mount compatibility with user namespace UID mapping and is tracked as a TODO.
hive-c0re privilege separation
Background
hive-c0re runs as the unprivileged system user hive-core
(/var/lib/hyperhive owned by hive-core:hive-core). It can't
directly invoke nixos-container, journalctl -M, or act on a system
unit (systemctl reload nginx) — those require root. hive-priv fills
this gap.
⚠️ ReloadGatewayNginx acts on a host unit, so nothing implicitly
scopes it. Its containment is the unit name hard-coded in hive-priv:
a caller can't name the unit, so the verb can't be steered at another
service. A privileged verb needs something bounding what it can act
on; when that isn't a namespace, it has to be a constant the caller
can't supply.
hive-priv
hive-priv is a minimal privileged helper that runs as root, socket-activated
at /run/hive/priv.sock (mode 0660, group hive-core — only the
hive-core user can connect). hive-c0re calls it via priv_client
for every operation that genuinely requires root.
Narrow interface — PrivRequest variants map 1:1 to specific
known operations; there is no arbitrary command pass-through:
| Operation | What it runs |
|---|---|
StartContainer / StopContainer |
nixos-container start/stop <name> |
KillContainer |
machinectl kill <machine> --signal=SIGKILL (nixos-container has no kill verb) |
CreateContainer / UpdateContainer |
nixos-container create/update <name> --flake <ref> |
DestroyContainer |
nixos-container destroy <name> |
ListContainers |
nixos-container list |
ReadContainerJournal |
journalctl -M <container> -n <n> [filters...] |
ReloadGatewayNginx |
systemctl reload/start/reset-failed nginx (host unit; the unit name is hard-coded, not a parameter) |
WriteNspawnFlags |
write /etc/nixos-containers/<container>.conf (bind-mount list + network isolation vars) |
WriteResourceLimits |
write CPUQuota=/MemoryMax=/CPUWeight=/IOWeight= systemd drop-in for agent container |
RemoveServiceDropin |
remove container@<name>.service.d/ drop-in on destroy |
DaemonReload |
systemctl daemon-reload |
RunForgeAdmin |
nixos-container run hive-forge -- runuser -u forgejo -- forgejo admin <args> |
WriteAgentForgeToken / WriteAgentMatrixToken |
write 0600 credential file into agent state dir |
RestartMatrixDaemon |
systemctl --machine=h-<name> restart hive-matrix-daemon.service |
ControlInfraContainer |
systemctl <action> container@<container>.service — the InfraContainer enum is the allowlist, and serde rejects unknown names at the wire boundary (hive-c0re has no variant, so no request can name it) |
SyncAgentTmpfiles |
write /etc/tmpfiles.d/hyperhive-agents.conf for the agent set, then systemd-tmpfiles --create |
SetAgentPaused |
create / remove the <state>/<name>/harness/paused marker that parks an agent's turn loop |
WriteAgentGithubToken |
write 0600 github-token into agent state dir (same semantics as the forge/matrix token writes) |
WriteAgentExtraForgeAccount / DeleteAgentExtraForgeAccount |
write / remove forge-<label>-token + a forge-<label>.json base-URL sidecar, both 0600. hive-priv validates label as a plain identifier before it reaches the filename — an unchecked one traverses out of the state dir |
RegisterCiRunner |
write /run/hive-ci/runner-token (host path, root-owned) then systemctl --machine=hive-ci restart gitea-runner-hive.service. Only the registration token crosses; the forge admin token never enters the container |
EnsureAgentSubvolume |
btrfs subvolume create <state>/<name> for a new agent — no-op when the path exists or the filesystem isn't btrfs |
UpgradeAgentSubvolume |
migrate an existing plain state dir into a subvolume: create, cp -a --reflink=auto, atomic swap. Operator opt-in, and the caller stops the agent first |
DeleteAgentSubvolume |
btrfs subvolume delete <state>/<name> — purge path only, never a plain destroy |
EnsureBtrfsQuota |
btrfs quota enable <AGENT_STATE_ROOT>. Operator opt-in — enabling forces a full rescan |
ReadSubvolumeUsage |
btrfs qgroup show -f --raw <state>/<name> |
SetSubvolumeQuota |
btrfs qgroup limit <bytes|none> <state>/<name> |
SnapshotAgentSubvolume |
btrfs subvolume snapshot -r <agent_root> <snapshot_path> — the frozen point-in-time copy a migration streams from |
DeleteAgentSnapshot |
btrfs subvolume delete <snapshot_path> |
SendAgentSnapshotToFile |
btrfs send [-p <parent>] <snapshot> into a bare filename under MIGRATE_STAGING_ROOT; refuses to overwrite an existing export |
SendAgentSnapshotToFd |
btrfs send [-p <parent>] <snapshot> into a file descriptor passed with the request (SCM_RIGHTS). hive-c0re connects to the peer and hands over the connected socket, so hive-priv never learns an address or protocol. hive-priv requires exactly one descriptor here, and refuses one that arrives alongside any other operation |
Container allowlist — hive-priv validates every request against
an allowlist before any operation: the allowlist accepts only names
matching the agent-name convention (char-validated) or the known
sibling service containers (hive-forge, hive-matrix, hive-ci),
and rejects arbitrary container names. hive-gateway is a host unit,
not a container, so it's not in this list — see ReloadGatewayNginx
above for how its access is scoped instead.
Socket-activated — systemd starts hive-priv on the first
incoming connection (LISTEN_FDS=1); it's not running between calls.
The ProtectSystem=strict + ReadWritePaths sandbox limits filesystem
writes to only the paths hive-priv legitimately needs.
Privilege boundary summary
| Component | Runs as | Privilege needed for |
|---|---|---|
hive-c0re |
hive-core |
broker, HTTP dashboard, scheduling, approvals |
hive-priv |
root |
container lifecycle, journal reads, bind mounts, cred writes |
hive-ag3nt (per-container) |
per-agent user | turn execution, MCP serving |