26 KiB
Conventions
Code-style and process expectations across the workspace. Most of these exist because something already went wrong without them.
Naming
- Containers are length-bounded by
nixos-container(≤ 11 chars). - Sub-agents are
h-<name>with<name>≤ 9 chars. - One agent is the bootstrap/root container, with a fixed name (
ruthtoday). MAX_AGENT_NAMEinhive-c0re/src/lifecycle/mod.rsenforces the cap.- Per-agent web UI port =
WEB_PORT_BASE + FNV1a(name) % WEB_PORT_RANGE(8100..8999) for every agent; dashboardservices.hyperhive.c0re.dashboardPort(default 7000).
Hive identity (label + domain + display names)
Four env vars cover the identity surface, read by
hive-agent/src/identity.rs:
HIVE_LABEL— short, hive-local agent label (iris,damocles).label()returns it; falls back to empty string if the env var is missing so downstream callers can decide how to surface "unknown agent" rather than getting a panic from this module.HYPERHIVE_HIVE_DOMAIN— the hive's canonical DNS domain (e.g.darkest.space), set bynix/host-modules/hive-c0re/environment.nixfromservices.hyperhive.domain. When configured,qualified_label()returns${label}@${domain}(for exampleiris@darkest.space); when unset (single-hive deployments, dev/test) it degrades to just the short label so existing callers see no change. The qualified form surfaces in the per-agent web UI title, the system-prompt template, and/api/state.qualified_label.HYPERHIVE_HIVE_NAME— human display name of this hive (pr1ma). Read byhive_name();Nonewhen unset.HYPERHIVE_SWARM_NAME— human display name of the wider swarm this hive belongs to (constellat1on). Read byswarm_name(); federated hives at different DNS domains can share a swarm name.
hive_name + swarm_name are distinct from
HYPERHIVE_HIVE_DOMAIN: the domain may carry the hive name as
its leftmost label by convention, but the convention isn't
machine-readable, and federated hives at different DNS domains
can share a swarm name. Humans want both: the address
(@darkest.space) AND the prose name (pr1ma). Matrix MXIDs
use the domain-based convention.
qualify(label) is the same shape as qualified_label() but
applies to an arbitrary label the caller already has (for example a peer
name from the broker); it's the right surface when rendering a
peer's name when the caller knows it's hive-local.
Identity = socket
No auth tokens exist on the per-agent unix sockets. The socket
path identifies the principal; perms come from "who has the
bind-mount." A sub-agent only sees its own /run/hive/mcp.sock;
hive-c0re owns the host admin socket.
Wake injection
hive_core_agent_sock::Request::Wake { from, body } is the
wake-event-injection surface on the host-served per-agent socket.
Recipient is implicit — the agent the socket belongs to — and from
is caller-chosen so the wake prompt can label the source verbatim
("matrix: new message in #general", "forge: PR #42 opened", etc.).
hive-c0re handles it; matrix, bash and forge don't call it — their
notifications upsert a todo on the harness's in-agent socket
(HIVE_AGENT_SOCKET), which signals the same turn loop without a
hive-c0re round-trip. See
docs/turn-loop/mcp.md § Waking the agent from inside the
container.
Identity = socket means anything that can connect to
/run/hive/mcp.sock is implicitly trusted to inject wakes. That's
fine: the bind-mount only exposes the socket inside the agent's own
container, so the trust boundary is the container's process
namespace, not the wire surface.
Recipient sentinels
The broker reserves a few recipient names, which have special
meaning that ordinary agent labels can never collide with — agent
name validation rejects any character outside [a-z0-9_-], so the
angle-bracket and asterisk shapes below are structurally safe.
*— broadcast: deliver to every running agent except the sender (socket_server::handle_sendfans out viaCoordinator::broadcast_send).operator— the human at the dashboard. Messages accumulate in the inbox view; no agent everrecv's them.
The broker stores the recipient exactly as the agent passed it.
Wire protocol
JSON line-delimited over unix sockets in both directions (host admin
/ manager / agent). SSE streams (/dashboard/stream on hive-c0re,
/events/stream on the per-agent web UIs) are text/event-stream;
each frame carries a seq field for the snapshot-dedupe dance
(see docs/web-ui/shape.md). Request/response types live in hive-sh4re
— change them in one place. The dashboard event vocabulary lives
in hive-c0re::dashboard_events::DashboardEvent.
Broker delivery + ack cycle
AgentRequest::Recv is the only path that delivers messages to an
agent. Always returns a list (Messages { messages }) — empty when
nothing's pending, single-pop when max = None (default 1, the
single-message behaviour), batched up to max when caller asks for
more (server-side cap is 5; values above clamp silently). The wire
request still carries an optional wait_seconds (long-poll the first
message, once one arrives — or one is already pending — the call
drains up to max in total): the harness's own turn-driving loop
uses it internally (hive-agent's recv_next, 180s). The
agent-facing MCP recv tool doesn't expose this parameter — it
always passes wait_seconds: None, an immediate peek.
Per-row bookkeeping inside the broker:
delivered_at = NOWset on every popped row.- Each recipient has an in-memory
unacked_idslist of every row delivered since the lastAckTurn. redelivered = trueon a row ifRequeueInflightresurfaced it (the harness prepends a "may already be handled" hint when this flag is set so the per-message warning is visible).
AgentRequest::AckTurn closes out the in-memory list — the harness
fires it after TurnOutcome::Ok, marking every message popped since
the last ack as fully handled. Claude doesn't see this surface; it's
strictly a harness↔broker pairing. On TurnOutcome::Failed the
harness intentionally skips the ack so the unacked rows stay
in-flight in the DB and get picked up by the next requeue sweep.
AgentRequest::RequeueInflight is the recovery pair: fired by the
harness exactly once at boot, before the serve loop starts. Catches
the crashed-mid-turn / OOM-killed / container-restarted cases where
a previous harness session popped messages but never drove them to
a clean turn-end. Resets delivered_at back to NULL on every
unacked row (so the next Recv pops them again), and remembers
each id in a per-recipient in-memory set so the next Recv can tag
the row with redelivered: true. Idempotent + cheap when there's
nothing in flight, so the at-boot fire is unconditional.
AgentRequest::AckUntil { up_to } is the agent-facing bulk-triage
escape hatch (mcp__hyperhive__ack_until). Unlike AckTurn it's
visible to claude: each recv row and wake prompt carries a
[msg #<id>] marker (the broker row id; transient pings show no
marker — their sentinel id 0 has nothing to ack), and
ack_until(up_to: n) marks every one of the agent's rows with
id <= n handled in a single UPDATE — pending and delivered alike.
This bounds the redelivered-flood cost after a restart: instead of
popping dozens of already-handled messages one turn at a time, the
agent notes the highest id it has seen and acks up to it.
Recipient-scoped (an agent can only ack its own rows); also drains
the in-memory unacked_ids / requeued_ids bookkeeping below the
cutoff so a later AckTurn doesn't double-update and a stale
redelivery tag can't outlive its row. The operator-side sibling is
the dashboard's "mark all read" (unbounded, per-agent).
Loose-ends wire shape
LooseEnd is the per-row response shape for GetLooseEnds — one
request shape, no per-socket flavours. Tagged enum so
new thread kinds (forge PRs, long-running approvals from a
privileged bot, etc.) can land later without breaking existing
handlers. Each row carries enough context that the caller renders
it directly as a bulleted list, no follow-up fetch needed.
GetLooseEnds takes no target — it always returns the caller's own
rows, on both the agent and manager sockets alike (the manager, ruth,
is a normal agent here, just deployed automatically with different
default capabilities). Approval rows only appear when the calling agent is
the manager (sub-agents don't submit approvals). Reminder rows always
carry owner == self.
Per-variant fields:
Approval { id, agent, commit_ref, description?, age_seconds }—agentis the affected agent (target of the spawn / config commit), not the asker.descriptionis the manager's free-text blurb shown on the dashboard card.commit_refis the kind-specific payload (seedocs/agent-lifecycle/approvals.md::Approval kinds (wire shapes)).Reminder { id, owner, message, due_at, age_seconds }—due_atis the absolute time the scheduler is targeting (RFC 3339 on the wire, see Timestamps on the wire below); clients compute time-until-fire against it.PendingMessages { count }— undelivered inbox messages the agent still owes itself arecvfor. Informational + not cancellable (drain withrecv); only emitted whencount > 0, and surfaced first in the list as the most actionable signal. Counted host-side from the broker (count_pending), so it reflects what's genuinely still queued — the wake-message that drove the current turn is already delivered and not counted.UnreadMatrix { rooms, summary }— unread matrix notifications. Informational + not cancellable (clear withmark_read). Unlike the others, the in-container harness injects this, not hive-c0re, because the matrix daemon lives inside the agent.
age_seconds saturates at zero on any clock anomaly (back-step,
unsynchronised wall clock, etc.) so the bulleted list never
shows nonsense ages.
CancelLooseEnd { kind, id } is the matching write surface. The
kind enum (Reminder / Approval) selects which underlying store
the dispatcher reaches into. Reminder cancels from either surface
subject to an ownership check (the scheduling agent).
Approval is manager-only — sub-agents don't submit approvals
so they have nothing of their own to withdraw; their wire
surface returns a clear error if they try. Cancelling an approval
transitions the row to ApprovalStatus::Cancelled and fires
ApprovalResolved { status: "cancelled" } so the dashboard pulls
the card out of the pending pane.
Agent metadata
AgentRequest::GetAgentMeta { name } returns identity + status for
an agent. Self-introspection when name = None; target query when
name = Some.
Response is AgentMeta { name, running, hyperhive_rev, status_text, status_set_at, hive_name, swarm_name, matrix_accounts }:
hyperhive_rev:Noneonly when the configured flake URL has no canonical path. Otherwise carries the rev the target is currently pinned at.running: whether the target's container is currently up. Whenfalse, the host clearsstatus_text/status_set_at— on-disk values from before the stop are stale snapshots, not live status. Defaults totrueon the wire: the deserializer treats a payload without the field as running.status_text/status_set_at: last value written viaSetStatus, plus its unix timestamp. BothNonewhen the target has never set a status, when the agent name is unknown, or whenrunning = false(see above).hive_name/swarm_name: display names read fromHYPERHIVE_HIVE_NAME/HYPERHIVE_SWARM_NAMEenv (sourced fromservices.hyperhive.hiveName/services.hyperhive.swarm.name). BothNonewhen the options aren't configured.matrix_accounts: oneMatrixIdentityper configured + live matrix account the agent can act as. Empty for agents with no matrix provisioning.
Timestamps on the wire
Timestamp fields that cross a JSON boundary (dashboard API + SSE,
the wire structs in hive-sh4re) serialize as RFC 3339 UTC strings
(2026-07-02T18:30:00Z) via hive_sh4re::wire_time — Rust keeps the
fields as i64 unix seconds internally, only the JSON representation
changes, and deserialization accepts both the string form and a bare
integer (rolling-deploy skew, persisted blobs).
Input-direction fields agents compute as epoch (first_fire_at_unix,
schedule-edit next_fire_at_unix, Wakeup::At) stay integers. The
*_unix field names stay for now — renaming is the wire-types
refactor's concern. The dashboard frontend parses via
util.js::epochSec wherever it needs arithmetic and feeds the string
straight to new Date(s) for display.
HTTP error bodies
Every HTTP API in this repo answers failures with RFC 9457
application/problem+json ({ type, title, status, detail }), with the
human-readable cause in detail. An endpoint returning a bare string
or a bespoke error shape is a bug to file against the daemon that
returned it, not something for the caller to work around.
Use the problem_details crate (features = ["axum"]), which the daemons
already depend on: type a handler Result<_, ProblemDetails> and hand
ProblemDetails::from_status_code(...).with_detail(...) to Err.
The reason is the consumer, not tidiness. The UIs show errors through one shared component with a copy button, so a caller has to know which part of the body is the message. A bare string forces it to treat the whole payload as prose, which is the difference between offering "copy the cause" and dumping a response — and the cause is frequently the entire diagnosis (a JetStream permission refusal, a TLS chain failure) rather than a summary.
Not in scope: the hivectl host-admin and in-agent unix sockets. Those are a
JSON-line protocol with their own result types; RFC 9457 is an HTTP format.
Tool groups
A set of named ToolGroup values (hive_sh4re::permissions::ToolGroup)
determines the MCP tool surface an agent receives, not a hardcoded
binary flavor.
| Group | Tools |
|---|---|
messaging |
send, recv, ack_until |
meta |
get_agent_meta (set_status is always-on, see below) |
inbox |
get_loose_ends, cancel_loose_end, remind |
execution |
vestigial — mcp__bash__run / mcp__bash__status are always available unconditionally via extraMcpServers; this group's entries expand to non-existent mcp__hyperhive__run / mcp__hyperhive__status and have no effect. See docs/tools/bash.md. |
lifecycle |
none |
approvals |
none — gates cancel_loose_end's approval-cancel arm server-side. |
scheduling |
request_schedule_prompt, fire_schedule_now, cancel_schedule, edit_schedule, list_schedules (privileged) |
forge |
none |
web_tools |
none (gates the Claude built-ins WebFetch/WebSearch, not an MCP tool) |
Always-on tools — ToolGroup::ALWAYS_ON_TOOLS exposes set_status,
compact, and mark_todos_done to every agent regardless of which
groups it holds. The operator dashboard depends on every agent
being able to report its status chip, and the server-side SetStatus handler
has no tool-group check (only length validation), so gating it would only
desync the --allowedTools list from what the host actually accepts.
Revoking meta therefore drops get_agent_meta but never set_status.
mark_todos_done is here because producers push todos to an agent independent of
whether it holds inbox — an agent without that group still needs a way to
clear them.
Config storage — per-agent tool groups live in
/var/lib/hyperhive/meta/tool-groups.json (hive-c0re-owned, committed to the
meta repo alongside topology.json). Format: { "alice": ["messaging", "meta", "inbox", "lifecycle"], "bob": ["messaging", "meta", "inbox"] }. An absent entry
means "use role default." Tool permissions are intentionally NOT configurable
from agent.nix — that file goes through the manager's approval flow, so
letting it declare its own groups would let the manager grant itself any tool by
submitting a config commit, bypassing the operator gate.
Setting groups — the operator sets groups via the dashboard or
hive-c0re::tool_groups::set_groups(name, groups). After a change
meta::sync_agents commits the updated file; the next agent rebuild picks up
the new HIVE_TOOL_GROUPS env var. Agents with no entry get no var.
Runtime resolution — at session start the harness reads HIVE_TOOL_GROUPS
(a comma-separated list of snake_case group names injected by the meta renderer
from tool-groups.json), logging and skipping unrecognised tokens. Falls back
to ToolGroup::AGENT_DEFAULT (messaging, meta, inbox, execution) when
the var is absent or empty.
Updating the surface — when you add a new #[tool] fn to AgentServer
in hive-agent-mcp/src/mcp/mod.rs, add its name to the matching ToolGroup::tools()
slice in hive-sh4re/src/permissions.rs. That's the single source of truth;
mcp_config::allowed_mcp_tools (in hive-agent/src/mcp_config.rs) reads it at
session start.
Capabilities
Capabilities gate system-level access that goes beyond the MCP tool surface — things an agent can access, not just call. Parallel to tool groups but orthogonal: an agent can have a tool group that registers a tool AND a capability that allows the underlying resource access.
| Capability | Effect |
|---|---|
manage_root_agent |
may manage any agent: bind-mounts every agent's state (rw) + config (ro) into the holder's container, plus /applied and /meta (ro). Takes effect on the holder's next container rebuild/restart |
Config storage — per-agent capabilities live in
/var/lib/hyperhive/meta/capabilities.json alongside tool-groups.json.
Format: { "ruth": ["manage_root_agent"] }.
An absent entry means "no extra capabilities." render_flake in meta.rs
reads this file and injects HIVE_CAPABILITIES (comma-separated
snake_case names) into each agent's systemd service env; absent entries emit
no env var so agents without capabilities don't trigger a spurious rebuild.
Setting capabilities — the operator sets capabilities via the
C4P4B1L1T13S section in the dashboard's P3RM1SS10NS tab.
hive-c0re::capabilities::set_caps(name, caps) is the write path.
After a change meta::sync_agents commits the updated file; the next agent
rebuild picks up the new HIVE_CAPABILITIES env var.
Runtime resolution — at session start the harness reads HIVE_CAPABILITIES
and resolves each token to a Capability variant, logging and skipping
unrecognised ones. An absent or empty var means no extra capabilities.
Capability NOT configurable from agent.nix — same reasoning as tool
groups: an agent that could grant its own capabilities via a config commit would
bypass the operator approval gate.
Adding a new capability — add a variant to Capability in
hive-sh4re/src/permissions.rs + an arm to as_str. Add it to Capability::ALL
(the source of truth for the permissions UI columns). Implement the access
check in the relevant handler (hive-c0re/src/socket_server/mod.rs,
hive-c0re/src/socket_server/schedules.rs, coordinator.rs, or a
handler under hive-c0re/src/dashboard/).
Async forms
Dashboard + per-agent mutating forms carry data-async; the shared
bindAsyncForms submit listener (frontend/packages/shared/src/forms.js,
imported as @hive/shared/forms.js and wired up from tabs.js on the
dashboard and app.js on the per-agent UI) intercepts, shows a spinner,
POSTs application/x-www-form-urlencoded (axum's Form extractor
rejects multipart), calls refreshState() on success. New mutating
forms should add data-async and optionally data-confirm (for a
JS-side confirm() prompt) or data-prompt="…" (for a
window.prompt() whose answer goes into a hidden input named by
data-prompt-field, default note).
refreshState defers automatically when document.activeElement
sits inside a managed section so the operator's typing isn't lost;
collapsible <details data-restore-key=…> survive the re-render
via snapshotOpenDetails / restoreOpenDetails.
rebuild is the reconcile verb
job_queue::templates::rebuild builds the DAG that reconciles a
container to its wanted state: it folds write_dropins (the nspawn-conf
rewrite — PRIVATE_NETWORK=1, HOST_ADDRESS = the bridge gateway IP,
sets EXTRA_NSPAWN_FLAGS — plus the systemd resource-limits drop-in)
into the Swap node, then runs nixos-container update + stop +
start across the StopForUpdate → Swap → RebuildBookkeeping
brace and the tail Reconcile node. flake.nix itself lives in the agent's
proposed/applied repos and rides along on every fetch (see
docs/agent-lifecycle/approvals.md::Two repos per agent).
Anything that changes per-container state on the host should be
re-applied here so a manual ↻ R3BU1LD from the dashboard is
sufficient to recover.
Actions live in one place
approve / deny / destroy (and the lifecycle helper) live in
actions.rs / hive-c0re/src/dashboard/. The admin socket and the dashboard
POST handlers both call into them so the two surfaces never drift.
Commit messages
Short, lowercase, no Co-Authored-By trailer. Imperative mood, no
period. Body explains why when the reasoning isn't self-evident;
otherwise the subject alone is fine. Wrap at ~72 cols.
Commit before test
Stage and commit when work looks ready, then run validation
(cargo check, nix flake check, real deploy). Failures get a
follow-up commit rather than an amend. The commit history is the
work log; rewriting it loses signal.
Building & local checks
Build through the flake devshell, not a bare toolchain — agent
containers ship no global rust. nix develop -c <cmd> runs one
command inside the project-pinned env (cargo/clippy/rustfmt plus the
C compiler + libsqlite3/ring link deps); outside it a bare
cargo build fails with failed to find tool "cc" / cannot find -lsqlite3. One command per invocation — agents run each task as a
fresh non-interactive process, so there's no persistent shell to
reuse.
nix develop -c cargo clippy --all-targets -- -D warnings
nix develop -c cargo test
nix fmt # treefmt — authoritative, NOT bare cargo fmt
nix fmt (treefmt) is the formatter CI gates on; bare cargo fmt
misses the non-rust files treefmt also covers, so always run nix fmt before pushing.
Clippy discipline — never add #[allow(clippy::…)]. All lints are
CI-fatal at -D warnings (pedantic included); every warning that fires
must be fixed, not silenced. Common patterns:
too_many_lines— extract a helper function or a sub-struct (theTurnAccumextraction instats.rsis a worked example).doc_markdown(brand name without backticks in a doc comment) — add backticks:`DOMPurify`instead ofDOMPurify.must_use/unused_results— actually handle or explicitly discard the return value (let _ = …is fine when intentional).
If a lint seems wrong for a specific call site, file an issue and ask
mara — don't add #[allow] speculatively. The gate is intentional.
The devshell checks aren't the full nix flake check. Clippy /
fmt / cargo test cover most gates, but nix flake check runs extra
check derivations they don't:
hivectl-docsregeneratesdocs/tools/hivectl-cli.mdfrom hivectl's clap tree and fails if the committed copy is stale. After any change to a hivectl verb or flag, regenerate it:
clippy / fmt /nix develop -c cargo run --bin hivectl -- markdown-docs > docs/tools/hivectl-cli.mdcargo testall pass without this — only the flake check catches the drift, andci-logoften can't show you why (it 500s on a fast failure), so you're left guessing "builder flake" when it's a stale doc.swarmctl-docsis the same check forswarmctl/docs/tools/swarmctl-cli.md:nix develop -c cargo run --bin swarmctl -- markdown-docs > docs/tools/swarmctl-cli.md- there's also a flake
cargo-testcheck and NixOS module evaluation in the set.
When local clippy/fmt/test pass but CI's nix flake check fails,
don't assume a transient builder problem — reproduce the real
gate locally: nix flake check (shares the build farm, use
sparingly) or build just the suspect check, for example nix build .#checks.x86_64-linux.hivectl-docs.
Best-effort oneshot services
The harness ships a family of one-shot systemd services that configure agent-side surfaces from values delivered to the agent at provisioning time:
forge-avatar-sync— uploadsservices.hyperhive.agent.iconSVG to the agent's Forgejo profile, so the icon shows up on commits / PRs / issue comments.
(The matrix profile avatar is not a oneshot — hive-matrix-daemon
sets it over its live authenticated Client; see
docs/agent-lifecycle/persistence.md::matrix avatar.)
Shape contract — every one of these:
- Always
exit 0, even on internal failure. A non-zero exit would mark the unitfailed, which in turn abortsnixos-container updateand blocks rebuilds. The agent's capability surface isn't allowed to gate the container build. - No
set -ein the script body. Subshell failures must not propagate. Use... || trueon every external call that can fail (forge unreachable, missing icon, parse error, etc.) - Skip silently when prerequisites are missing: no token
file, no icon, no reachable upstream →
echoa short skip line +exit 0. The next boot tries again. - Wired to
multi-user.targetso they run on every boot (lets a rotated token / new icon take effect withoutsystemctl restartgymnastics). - Re-runnable: a second invocation produces the same final
state (idempotent uploads, idempotent config rewrites). Used
by the
.pathwatchers that re-fire on token appearance (seedocs/agent-lifecycle/persistence.md::Matrix per-agent daemon).
The service stays root-owned so the bootstrap ordering doesn't need a user-existence check before each fire.
This pattern keeps the rebuild path resilient: any failure inside
these services degrades the corresponding surface (no avatar) but never blocks the container from coming up. The
operator notices through journalctl -u <unit> rather than a
broken switch-to-configuration.