A config PR merged in the Forgejo UI changed nothing on the hive: the
hive's webhook ignores `closed`, its poll then cancels the dashboard
card, and `applied/main` stays where it was.
swarm-controller reads `merged`/`merge_commit_sha` off the
`pull_request` delivery it already receives for `agent-configs`, finds
the hive placing the agent by scanning every hive's wanted state (the
scan `declarations_elsewhere` already ran, factored out), and queues a
`TriggerDeploy` carrying the rev. Zero or several claimants deploy
nothing and log the claimants.
`DeployRequest` gains `rev: Option<String>` with `serde(default)`, so
rev-less payloads from either side keep decoding.
hive-c0re, given a rev for an agent it runs: a no-op when
`applied/main` already is the rev (a dashboard merge deploys its own
PR); otherwise it fetches the forge `main` with the core token,
requires the rev to descend from `applied/main` (the ancestry gate,
factored out of `run_deploy_merge_verify`), fast-forwards by CAS and
queues the usual relocking rebuild. No eval-verify on this path, per
mara (#4850 c90075). A refusal is commented on the PR that merged the
rev, found by commit.
swarm-controller's forge-objects pass converges every config repo's
`main` rule to merge whitelist `operators` + `core` and approval
whitelist `operators`. The hive's boot PATCH stops forcing
`enable_approvals_whitelist` off, so the two do not fight.
Refs #4850
The swarm-controller builds the Forge quick link from
services.hyperhive.swarm.forge.domain, replacing hive-forge/default.nix's
per-host entry, so all seven swarm-service links come from swarm-level
options.
Also drops the remaining references to the removed matrix GUI switch:
the HiveUrls / Urls / hive_urls docs, the hivectl.md `open` note and
the gateway.md vhost-map rows, which name `gatewayHost` instead. The
grafana, victoriametrics and victorialogs modules' comments no longer
mention a quick-link they do not define.
Refs #4885
Removes services.hyperhive.deploy.matrix.gui.enable and its
swarm.matrix.gui.enable alias; both are mkRemovedOptionModule stubs. A
host running the homeserver serves fluffychat at gatewayHost's vhost,
and the hive's /matrix/ redirect follows the same condition.
The swarm-controller builds the Matrix quick link from
swarm.matrix.gatewayHost, replacing hive-matrix.nix's per-host entry.
HIVE_MATRIX_PUBLIC_URL is set on every hive with a gatewayHost, so
`hivectl open matrix` resolves off the homeserver's host too.
Drops HIVE_MATRIX_GUI_ENABLED and the dashboard's matrix_gui_enabled
field; nothing in the frontend reads it.
Refs #4885
hive-c0re/server.rs and hive-host-sock/lib.rs still described the
deleted behindGateway option as if a forge URL could be None for that
reason. Reword both to match the HiveUrls doc's publicUrl wording.
Refs #4885
The forge always sits behind the gateway, so `deploy.forgejo.behindGateway`
(and its `swarm.forge.behindGateway` rename alias) is removed and its
true-branch behaviour is now unconditional within `deploy.forgejo.enable`:
https ROOT_URL on the gateway's httpsPort, the forge vhost and local DNS
name, the swarm-ui quick link, the published metrics scrape target, forgejo
metrics, the authelia `/metrics` rule, and `publicUrl` defaulting to
`https://<forge.domain>`.
Removed with it: the direct-port `http://<domain>:<httpPort>/` ROOT_URL
branch, the hive-ci assertion that the option is true, the core-toggle
cases that only exercised the false branch (the services-leaf case reads
`bare`, which never enabled the forge either). `hivectl open forge` now
points at `swarm.forge.publicUrl`, which can still be set to null.
Refs #4885
`verify_hmac` is the only authentication gate on the public
`/webhook/config-pr` endpoint, and neither of its guard clauses nor the
handler's "unavailable" → 503 / otherwise → 401 mapping had a test.
The tests build a real `AppState` over the tempdir `Coordinator` that
the socket_server schedule tests already use. That helper is widened
from `pub(in crate::socket_server)` to `pub(crate)` and re-exported from
`socket_server` under `#[cfg(test)]`, so non-test code is unchanged.
Removing the missing-secret guard (replacing it with
`unwrap_or_default()`) fails
verify_hmac_rejects_every_delivery_when_no_secret_is_loaded and
config_pr_answers_503_when_no_secret_is_loaded. Deleting the
missing-header guard fails
verify_hmac_rejects_a_delivery_with_no_usable_signature_header. Breaking
the "unavailable" match fails the 503 test.
Without the missing-header guard, an unsigned delivery is still refused
by `verify_signature` (no `sha256=` prefix). The test therefore pins the
guard's own message rather than the bare rejection.
Closes#4654
An operator links an agent's GitHub personal access token in the swarm UI
(LinkGithubAccountForm, "link github account" on /agents). swarm-controller's
PUT /api/hives/{hive}/agents/{agent}/github-account stores it at
swarm/agents/<agent>/github-token (swarm_secret_client::github), a flat leaf
under the agent's prefix that the agent's existing read grant already covers:
no policy change, and no list grant, since there is one token per agent.
In the agent, hive-agent-github-token (oneshot + 2-minute timer, as the agent
user, under its own store certificate, ordered before hive-github-notify)
reads that path and writes <state>/github-token, 0600 and agent-owned, the
file the gh wrapper, git credential helper and hive-github-notify already
read. It replaces the file by rename only when the bytes changed and never
deletes it: a hive-written github-token stays until a token is linked in the
swarm UI. It is installed only with a store address and
services.hyperhive.agent.github.enable.
Removed: the dashboard's CR3D3NTIALS page (credentials.html/js/css, its
build entries and H0M3 tile; GITHUB was its only tab), hive-c0re's
dashboard/matrix_accounts.rs with GET/POST /api/github-account,
priv_client::write_agent_github_token, the host socket's
SetAgentGithubToken and `hivectl github set-token`, and hive-priv's
WriteAgentGithubToken with write_agent_state_file, its only caller gone.
Docs: integrations/github.md and swarm/ui.md describe the swarm path,
swarm/credentials.md gains the store-path row, and the hive UI docs,
hivectl docs and security.md's hive-priv table drop the removed pieces.
Closes#4347
The previous commit removed dashboard.md's M4TR1X section. These
comments and the `deploy.matrix.gui.enable` option description still
described a hive-dashboard M4TR1X tab or cited that section. They now
state what the option does: it serves the client on the gateway vhost
and adds the swarm UI's Matrix quick link (docs/swarm/ui.md::Quick links).
Refs #3902
- frontend/README.md: list the missing packages/swarm-ui package, fix
the vanilla-JS claim (all packages depend on preact), and the
deprecated hyperhive.frontend.extraFiles spelling
- hive-c0re/README.md: hive-c0re no longer provisions per-agent
forge/matrix accounts (swarm-controller does); it wires gateway
vhosts and reconciles forge/matrix config
- hive-metric/README.md: OTEL_EXPORTER_OTLP_HEADERS is never set by
the harness and has no way to be set
- hive-agent-sock/README.md: fix the deprecated
hyperhive.extraMcpServers spelling
- docs/process/gotchas.md: fix a dangling hive-ag3nt/ path, the real
directory is hive-agent/
Also fixes pre-existing vale error-level alerts (passive voice,
Microsoft.Auto, Microsoft.Contractions) in the same files so
prose-lint-errors passes clean.
An operator now links an agent's external forge account (label, base URL,
token) in the swarm UI. swarm-controller stores it at
swarm/agents/<agent>/forge/<label>. There is no index: the store's
listing of the agent's forge/ directory is the set of accounts.
In the agent, hive-agent-forge-accounts (oneshot + 2-minute timer, as
the agent user, under its own store certificate) lists
swarm/agents/<agent>/forge/ with the `list` #4866 grants an agent on its
own metadata subtree, reads each account, and writes
<state>/forge-<label>-token and forge-<label>.json in the names and shape
hive-forge -f already reads. An empty listing (a 404, which `bao kv list
-format=json` answers with `{}` and an empty stderr) is zero accounts; a
denial or an unreachable store fails the unit. It never deletes: files
for labels not listed, including ones the hive wrote, stay as they are.
Removed: the dashboard FORGES tab (credentials.js/html section and its
CSS), hive-c0re's extra_forges.rs and its routes, priv_client's
extra-forge calls, and hive-priv's WriteAgentExtraForgeAccount /
DeleteAgentExtraForgeAccount with their helpers. The GITHUB tab and
WriteAgentGithubToken stay.
Also: persistence.md's matrix avatar note names the exit-75 restart on a
changed account listing, not the dashboard, as what brings a linked
account up.
Refs #4348
hive-matrix-daemon now learns which external matrix accounts it has from
the swarm secret store, under the agent's own certificate, and the hive
push chain for matrix is gone.
The daemon lists swarm/agents/<agent>/matrix/ (the `list` its policy
grants on its own metadata subtree), reads each account's homeserver
from its credential, and brings the accounts up with their tokens from
the store. Every two minutes it lists again and exits with 75 when the
set of linked accounts changed; the unit restarts on 75 without counting
a failure. A listed name whose credential reads as absent is skipped and
logged once. At start it removes the matrix-token-<a> /
matrix-account-<a>.json pairs a hive delivered (a sidecar marks a pair
as delivered; a declared tokenFile keeps its token).
Removed: CredentialNotice and the $SWARM.credential.* subject and NATS
grant, the controller's publish and its queue precondition on the PUT
route, hive-c0re's credential subscription arm and workers/credential.rs,
priv_client::write_agent_matrix_token, hive-priv's WriteAgentMatrixToken
and its helpers, and the daemon's state-dir account discovery.
Kept: WriteAgentGithubToken and the external-forge path
(WriteAgentExtraForgeAccount, extra_forges.rs) are untouched, and a
declared matrixAccounts tokenFile is still read when the store has no
token for that account.
Refs #4348
mara ruled on #4849 (c88934): "remove create_repo tool". The tool ran in
hive-c0re with the hive's core token, so it only ever worked for agents
on the hive that runs the forge.
Removed:
- the create_repo MCP tool and CreateRepoArgs (hive-agent-mcp)
- wire variants Request::CreateRepo and Response::RepoCreated
(hive-core-agent-sock)
- hive-c0re's handle_create_repo, its valid_repo_name check and the
dispatch arm
- forge::create_agent_repo and apply_operator_branch_protection, which
had no other caller, plus AGENTS_ORG and OPERATORS_TEAM, whose only
users they were
- the tool's docs (docs/tools/forge.md repo management, docs/turn-loop/
mcp.md, the conventions tool-group table) and the doc comments that
named it (hive-sock-client's response timeout, ensure_repo_creation_
disabled, the security doc's merge-gate bullet)
ToolGroup::Forge is kept with no tools, the same way b88a5b24 kept
Lifecycle, so existing meta/capabilities.json grants still parse.
Forge state is untouched: existing agents/* repos keep their collaborators
and operators-team branch protection. The swarm-controller's own
create_repo (config-org repos) is a different path and is unchanged.
Closes#4849
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:
- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
hive-matrix container, Command::Mint and src/mint.rs. The binary, its
appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
role stay. bao-matrix-reader's checks on the deleted unit are removed;
the leaf-identity and no-token-in-env checks now look at
swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
(register/appservice-login/password-login with the local as_token), with
read_appservice_token, paths::matrix_appservice_token and the helpers
only it used. ensure_hive_user now takes the store's token, keeps the
file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
unchanged apart from no longer reading the local as_token.
This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.
Closes#4813Closes#4814
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.
swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.
swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.
hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.
The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.
Refs #4427
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.
Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.
policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.
Refs #4396
The CR3D3NTIALS page's MATRIX tab was the only caller of
`POST /api/matrix-account-login` (provision/log in an external matrix
account through the hive) and `GET /api/matrix-accounts` (its account
list). External matrix accounts are linked from the swarm UI now
(`LinkMatrixAccountForm` -> swarm-controller), so the hive-side UI and
both routes go. `priv_client::restart_matrix_daemon` had no other caller
and goes with them.
Already-provisioned credentials keep working: the `matrix-token-<name>`
files and `matrix-account-<name>.json` sidecars the old route wrote are
still discovered by hive-matrix-mcp (`accounts::configured` ->
`discover_token_accounts`), the `matrix-token*` path unit still re-fires
the daemon, and `WriteAgentMatrixToken` stays for the swarm credential
worker. Removing that usage waits on moving the existing creds to
swarm level.
The GITHUB tab is the credentials page's default tab now.
Refs #4348
Human matrix accounts come from SSO, not hivectl. Matrix homeserver
admin will come from authelia's admins group (sync tracked in #4585);
password reset moves to swarm level (#4798). promote-user and
reset-password were already broken from the hive: the hive's sender
account has no admin sender to call the admin room with, only the
swarm's does.
Removes the three hivectl matrix verbs, their HostRequest variants,
their hive-c0re handlers, and the admin-room helpers (discover room id,
send-and-poll, event-id extraction, password/success parsing) that
only they used. sync-admin and invite are unchanged.
Refs #4585
post_purge_tombstone discarded fail_pending_for_agent's error with
let _ =. Mirrors #4740's fix for the identical discard in
job_queue/exec.rs's run_destroy_bookkeeping: warn and continue, since
the purge itself has already succeeded by this point.
Refs #4747
du_bytes returned None uniformly for every du failure, and both
du_bytes and measure_agent_disk had no logging at all, so a container
rootfs du couldn't read (permission, I/O error, unparseable output)
silently reported as 0 disk with no signal in the journal.
Distinguish the expected case — the path plain doesn't exist, e.g. a
destroyed-but-kept agent's rootfs or a not-yet-created state dir — from
an actual read failure, and warn only on the latter, with the path plus
whichever of the failure (spawn error, exit status, stderr, unparseable
stdout) applies.
prebuild_toplevel awaited nix build with a bare child.wait().await, so
a wedged nix-daemon (unreachable remote builder, stuck build slot)
hung the rebuild job forever with no way for the job queue to recover
short of restarting hive-c0re.
Mirrors #4741's hive-priv fix: the child now leads its own process
group, and after PREBUILD_TIMEOUT (1h, four times CI's observed
cold-cache flake check) the whole group is SIGKILLed and the call
fails with a named timeout error instead of hanging.
Refs #4723, #4741
hive-screen-mcp ran `grim` / `wtype` through an unbounded
`Command::output()` and spoke RFB to neatvnc with no deadline, so a
wedged compositor or a VNC server that accepts and never speaks held the
agent's turn forever. Each subprocess now has a 30s limit and is killed
when it hits it; each RFB exchange has a 10s limit. Both come back to the
agent as the tool's text result, like every other failure in this crate.
hive-c0re operator-facing text that described behaviour the code lacks:
- `hivectl matrix reset-password` printed a "next: hivectl matrix
create-user" hint that fails for every target (agents are refused, a
non-agent hits M_USER_IN_USE). The line is gone.
- a failed `nixos-container update` appended the container's journal
tail, read with `journalctl -M`. Since the job DAG, `update` only runs
from the `Swap` node on a stopped container, so the read always came
back empty. The helper is removed; the error still points at the
build log.
- the matrix sweep comment in main.rs said it re-provisions agent token
files; `ensure_all` creates no agent accounts or tokens.
- the knowledge-pull comments named a webhook caller that no longer
exists and claimed a race was "fixed at its source".
- `handle_spawn`'s doc and the `mcp_sockets` module doc named callers of
`register_agent` / a rollback that do not exist.
Refs #4723
mara chose to accept extra queued sweep passes over adding a new
hive-jobq primitive (or a hive-c0re one-off) for "don't queue another
of this kind". Every sweep caller now plain-inserts its node; the
capacity-1 Dep::Resource per sweep kind (MatrixSweep/KnowledgeTree)
still keeps two passes of the same kind from running concurrently, it
just no longer collapses a tick that lands while one is live or queued
into the existing one.
The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.
Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.
- templates::matrix_sweep / templates::knowledge_pull build the node
with its resource; boot, the periodic loops and the swarm event all
use them.
- JobQueue::insert_unless_live folds a submission into a live node of
the same kind instead of queueing another. Periodic ticks fold into a
queued or running pass. The swarm knowledge event folds into a queued
pull only, and queues one behind a running pull, which may have
fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
boot MatrixSweep node is the startup pass, as KnowledgePull already
was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
a hung pass would otherwise hold its resource against every later one,
and own the sweep-health banners, so every pass reports to them.
Replaces the SweepLock version of this branch, per review.
Refs #4723
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).
- hive-priv gains `EnsureAgentSocketDir { name }`, called from
`set_nspawn_flags` in every start path. It creates
`/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
directory is left alone. hive-c0re's own create_dir_all went: its /run
is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
the agent user and sets 0751, the same way it already handles state/ and
harness/. It refuses a symlink or non-directory there, since `test -d`
and chmod follow links. No host-side passwd parse, and no window where
the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
(`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
`converge_start_preamble` + `start_with_fallback`. It was a bare start,
so after a reboot the manager's bind sources existed only because of the
tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
`priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
`SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
does the same unlink and returns Ok, for an older hive-c0re.
Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.
Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.
Closes#4742
An agent name that is not a valid Ident made `Coordinator::agent_paths`
panic. Job payloads carry names as plain strings (the swarm's published
wanted state is one source), and a panic inside a job-queue node never
reaches `complete_growing`, so the node's resources (the deploy
window included) were held until hive-c0re restarted. `agent_paths` now
returns an error; the job-queue nodes, the admin-socket spawn and
set-limits paths, the root-agent spawn and the dashboard set-limits
handler propagate it.
`lifecycle::list().await.unwrap_or_default()` turned a failed container
list into "no agents":
- meta-update cascade: the lock bump committed and zero rebuilds fanned
out, reported as success. The cascade is now resolved before the lock
bump and a list failure fails the node.
- dashboard update-all: queued nothing and returned 200 "ok". Now 500
with the error.
- container rescan: every row was emitted as removed and the cache
emptied. Now the last snapshot stands; `hivectl status` gets an error.
- dashboard journal: answered 404 "no managed container". Now 500.
- spawn/rebuild port-collision check: silently skipped. Now fails.
- startup migration: the per-agent phases ran over nothing, and phase 3
handed an empty agent list to `meta::sync_agents`, which renders the
meta flake with exactly the agents it is given. Both now log the list
failure and skip.
The hive-jobq scheduler still leaks a node's resources on any executor
panic; that root is not addressed here.
Refs #4723
invite_user_id mapped every 403 M_FORBIDDEN to Ok(()). The membership
pre-check already skips invited/joined users, so the 403s that reach the
POST are mostly real refusals (banned target, sender without power),
including `hivectl matrix invite`. A 403 is now success only when a
membership re-read shows the user invited or joined; otherwise it is an
error carrying the status and body.
admin_room_send_and_poll read the send response's event_id with
unwrap_or_default() and, when it was missing, walked every recent event
unanchored, so an older bot reply (an earlier reset password) could be
returned as this command's result. A send response without an event_id
is now an error.
run_destroy_bookkeeping discarded fail_pending_for_agent's error; it now
warns like its neighbouring steps.
ensure_config_pr_webhook returned as soon as a hook with the target URL
existed, so a regenerated webhook-secret never reached Forgejo and every
config-PR delivery failed HMAC until the 5-minute poll caught up.
Forgejo's edit-hook API ignores `secret` and never returns it, so the
SHA-256 of the secret last registered is recorded at
forge/config-pr-webhook-secret-sha256; when it doesn't match, the
same-URL hook is deleted and recreated. The paths.rs doc claiming
re-registration on change now describes this.
Refs #4723
rustls is built with both `ring` (async-nats's `ring` feature) and
`aws-lc-rs` (reqwest's `rustls` feature), so it cannot pick a
process-level default by itself. Since the queue started requiring TLS
(1d261b3f), async-nats builds its config with `ClientConfig::builder()`,
which panics without an installed default. The panic kills the async-nats
connector task, and every queue client (swarm-controller, hive-c0re, all
hive-agents) has sat in `Pending` since the 2026-09-25 23:04Z deploy.
Add `swarm_queue_client::install_crypto_provider()`, which installs
aws-lc-rs and ignores the "already installed" error. It is called first in
`main` of every binary that links async-nats: hive-agent, hive-c0re,
swarm-controller, swarm-nats-auth. `connect()` also calls it, so a new
binary that dials through this crate is covered without remembering to.
aws-lc-rs because reqwest already falls back to it when no default is
installed, so HTTPS in these processes keeps its current provider. The
other rustls users in the tree reach it only through reqwest, which never
panics here.
Closes#4738
Both accept loops returned on the first accept() error. The AgentSocket
stayed in Coordinator.agents, so mcp_sockets::sync_on_start skipped the
agent, and only a Start node or a hive-c0re restart bound it again. The
unit sets no LimitNOFILE, so a single EMFILE at the 1024-fd soft limit
cut the agent off from send/recv/ack with one warn line as the only
trace.
Both loops now share accept_until_fatal. Connection errors
(ECONNABORTED/ECONNRESET/ECONNREFUSED) retry at once, any other error
retries after 1 s, following axum's serve loop that the dashboard
already runs. Errors that leave the listening fd unusable (EBADF,
EFAULT, EINVAL, ENOTSOCK, EOPNOTSUPP) log at error and exit the
process: nothing re-binds a listener while hive-c0re runs, the manager
listener has no re-bind path at all, and the unit's Restart=on-failure
restart runs start_manager and sync_on_start, which re-bind every
socket.
mcp_sockets.rs now states that invariant instead of asserting that a
listener can only disappear on restart.
LimitNOFILE is left unset: per-agent fd use is a listener plus a
Recv long-poll plus short-lived requests, and a higher limit would raise
the memory ceiling of the per-line bound (bound x open connections).
Closes#4721
serve() read each request with read_line into an unbounded String, so
size checks such as the 4 KiB Send body limit ran only after the whole
line was buffered. Any process in an agent container can write to
/run/hive/mcp.sock, and one stream with no newline grew hive-c0re's heap
until the host OOM-killed it.
Read at most 16 MiB per line. The largest request a client sends is an
OperatorMsg from the agent web UI's /send, whose body axum's default
limit caps at 2 MiB (the gateway's nginx admits 10 MiB); JSON escaping
can roughly double that, and the cap is 4x the result. A longer line
gets a Response::Err, the same shape as the parse-error path, and the
connection is closed because the rest of the line is still unread.
The line is read as bytes and parsed with serde_json::from_slice, so a
line with invalid UTF-8 now gets a parse-error response instead of the
connection being dropped.
Closes#4718
hive-sock-client: each attempt now bounds connect (5s), write (10s) and
the wait for the response (60s by default). The response bound is per
call through the new `request_within`, which hive-agent's serve-loop
`Recv` uses with its 180s long-poll plus 30s headroom. A response
timeout is terminal rather than retried: the server holds the request,
so a retry re-sends something it may still act on and multiplies the
wait by the backoff schedule.
Outbound HTTP: the matrix login/whoami clients in swarm-controller and
hive-c0re's dashboard (5s connect, 30s request), the authelia-bridge
client (5s/30s; ensuring an identity runs an argon2 hash first) and the
ci-runner forge calls (5s/15s, config_pr_poll's forge budget) get a
connect_timeout and a request timeout. Timeout errors name the bound
that fired.
hive-agent's unix-socket extra web proxy bounds the connect (5s) and
the wait for the response head (30s, the http sibling's budget); the
body read stays unbounded.
Refs #4723
resource-limits.json and topology.json were read with parse errors
folded into an empty map, and written in place with std::fs::write. One
truncated resource-limits.json followed by a single set_limits call
rewrote the file with only that agent's entry, erasing every other
agent's CPU and memory overrides without a log line. topology.json had
the same shape: reconcile rebuilt it from the live set, losing pending
(provisioned, never spawned) names.
- agent_config::read_map / write_map are generic over the stored type.
tool-groups and capabilities behave as before.
- resource_limits::read / effective return an error for an existing but
unreadable file; a missing file is still the empty map. set_limits
fails without writing on such a file, and writes atomically.
- topology: reconcile fails without writing on an unreadable file and
writes atomically. all_agents logs the error and returns no agents,
so a ManageRootAgent holder starts without cross-agent mounts.
Read-path behaviour on an unreadable resource-limits.json, per caller:
- write_dropins (every spawn / swap / WriteDropin): logs the error and
keeps the limits drop-in already under /run; the agent still starts.
With no drop-in yet (first start since boot) it writes the hive
defaults, because no drop-in means an uncapped container.
- render_flake: propagates, so sync_agents (and spawn/rebuild/destroy
jobs) fail. An empty map would give tighter-capped agents the hive
memoryMaxBytes.
- container_view::build_all: logs the error each scan and renders the
rows at the hive defaults (no ContainerView wire change).
- set_resource_limits reply: propagates.
Closes#4731
set_nspawn_flags propagated has_cap's error, so one corrupt
capabilities.json failed every agent's Start, spawn and Swap. It now
goes through holds_manage_root_agent, which logs the error (agent and
file) and treats the capability as absent: the agent starts without the
cross-agent, /applied and /meta mounts. caps_for/has_cap take the file
path so that seam is testable against a tempfile.
- meta.rs: a comment at the render_flake reads records why they
propagate (an empty tool-groups map renders toolGroups = null, i.e.
AGENT_DEFAULT, which fails open for narrower explicit entries).
- capabilities::read doc: states when set_caps/remove_agent rewrite
the file instead of saying remove_agent repairs it.
- set/remove corrupt-file tests assert ErrorKind::InvalidData.
tool_groups::read and capabilities::read returned an empty map when
their file existed but didn't parse. Every set_*/remove_agent is a
read-modify-write, and write() rewrote the file in place, so a crash or
ENOSPC mid-write left a truncated file, and the next write (e.g. the
manager-spawn seed of ruth's tool groups) replaced it with a map holding
only one agent. The scheduling and approval gates then denied every
other agent, recoverable only from meta git history.
- Both registries now read through agent_config::read_map: a missing
file is still the empty map, any other read failure or a parse
failure is an io::Error. set_groups / set_caps / remove_agent fail
without writing.
- Writes go through agent_config::write_map: temp file in the same
directory, fsync, rename, fsync the directory. hive-c0re had no
shared atomic-write helper (the existing tmp+rename sites are inline
and don't fsync).
- Callers of read / groups_for / has_cap now handle the error:
* dashboard GET /api/tool-groups, /api/capabilities,
/api/permissions/stale return 500 instead of an empty table;
* the SSE permission snapshots are skipped with a warn;
* render_flake returns Result, so sync_agents fails instead of
rendering every agent without its tool groups / capabilities;
* set_nspawn_flags propagates has_cap's error;
* the socket tool-group gates deny with the read error as message;
* seed_manager_tool_groups logs and does not seed.
- capabilities::write had no callers left once set_caps writes through
write_map, and is removed.
Closes#4719
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.
swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.
hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.
nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.
Refs #3782
The swarm mints each agent's `main` account now, so the hive's own mint
goes: `ensure_user_for`, `finish_user_provisioning`, `sync_agent`,
`sync_agent_standalone`, `token_path`, `legacy_password_path`,
`auto_reset_password` and `token_file_present`, and the calls from the startup sweep and the
rebuild bookkeeping. Both mints pinned the device `hyperhive-<agent>`, so
leaving this one would have each re-login kill the other's token.
`hivectl matrix create-user` refuses an agent's name and says where its
account comes from. Everything that still uses the hive's appservice token
stays: the hive's own account, the Space and chat room, and operator
accounts.
The forge now creates a human's account on their first authelia login,
so the verb has no job left. Deletes it, HostRequest::ForgeCreateUser,
its handler, provision_user_token, change_user_password and the hive's
TOKEN_SCOPES. change_user_password also passed the password as an
argument to `forgejo admin user change-password`, so it showed in the
container's process list.
ensure_user_exists and mint_token stay for the `core` bootstrap, their
one caller now. ensure_user_exists loses its password parameter: only the
deleted path set one.
Refs #3782
hive-priv writes an agent's matrix-token 0600 and owned by the agent,
so hive-c0re, running as hive-core, cannot read it. The read-based
token-present guard in ensure_user_for therefore never fired, and every
matrix sweep (boot, every 30 minutes, every rebuild) re-minted each
agent's token through the appservice login and restarted its
hive-matrix-daemon.
Decide presence with a stat instead: a non-empty regular file counts as
present. hive-core can stat the file through the 0755 state dir.
Closes#4665
crash_watch's 10s poll and auto_update's ensure_root_agent both read
lifecycle::list().await.unwrap_or_default(), which turned a failed read
into 'zero containers'. In crash_watch that made every previously-running
agent look like it crashed simultaneously (prev.difference(current) over
an empty current), and left prev empty for the next cycle too, so a
second wave of false 'agent logged in' / 'agent needs login' events fired
against the next successful read. In ensure_root_agent it read as
'manager container missing' and called lifecycle::spawn on a manager
that might already exist.
Both sites now treat a list error as its own outcome: log it at warn and
skip the cycle's decision entirely. crash_watch leaves prev exactly as
the last good read produced it. ensure_root_agent attempts no spawn.
Factors each site's decision into a pure helper (plan_cycle /
plan_root_agent) matching the check_not_live / confirm_gone_after_failed_destroy
pattern, with unit tests for the error case, a control for the readable
case, and (for crash_watch) an invert-proof run locally against the old
unwrap_or_default logic before reverting.
Delete ensure_user_for and mint_and_persist_agent_token, the user step
of sync_agent (the per-rebuild re-mint, #4644) and of
forge_after_first_spawn, and the hive-priv WriteAgentForgeToken request
that wrote the token into the agent's state dir. hivectl forge
create-user now refuses an agent and points at swarmctl agent
mint-forge-token. mint_token, ensure_user_exists and TOKEN_SCOPES stay:
provision_user_token and the core bootstrap still call them.
Refs #3782
forge-token.nix fetches swarm/agents/<agent>/forge-token under the
agent's own store identity into /run/hive-agent-forge-token/token, and
re-fetches on a timer so a rotation lands. hive-forge, the git
credential helper, hive-forge-notify, forge-avatar-sync and the web UI
read that file first and fall back to <state>/forge-token.
tea-login is deleted: it copied the token into ~/.config/tea, which
docs/swarm/credentials.md forbids for a store secret. hive-forge covers
the same verbs. swarmctl gains agent mint-forge-token.
Refs #3782
The PR retiring loose-ends' manager-only visibility replaced the false
claim with prose narrating its retirement, which still adds lines for
a removal. Delete the dead claim outright instead of documenting that
it used to be true: drop the three added sentences in
agent-hierarchy.md, and shorten socket_server/mod.rs's authority
comment to state only what's true now (tool-group membership for the
orchestration verbs) rather than adding an explanatory clause for the
removed capability check.
Follow-up to 729c5b4f42, which dropped the
`agent` parameter from get_loose_ends and deleted Capability::QueryAgentState
with it. Three prose sites still describe the interface that commit removed:
- socket_server/mod.rs's module doc claimed authority on this socket derives
from "capabilities for the hive-wide queries and tool-group membership for
the orchestration verbs". There are no capability-gated queries left —
`git grep -n has_cap -- hive-c0re/src/socket_server/` returns nothing, and
require_group("scheduling") is the only gate in dispatch_orchestration.
- agent-hierarchy.md listed loose-ends visibility ("manager sees hive-wide,
sub-agents only their own") as one of the manager-only overrides that
"exist across hive-c0re today". It doesn't, and the sentence pointed a
reader at loose_ends.rs for owner-check logic that is no longer there.
- conventions.md's Loose-ends wire shape still spoke of "the agent-flavour
and manager-flavour requests" and "the agent-flavour list". There is one
request shape.
No behaviour change: prose only.
Refs #4480
post_purge_tombstone's live-container guard used lifecycle::list().await
.unwrap_or_default(), so a failed list() (hive-priv socket unreachable,
restarting, ...) read as 'no live containers' and let the purge proceed —
deleting agent_state_dir/applied_dir for an agent whose container may
still be running. Factor the decision into check_not_live(), a pure
helper that refuses on both 'still listed' and 'list unreadable', with
unit tests covering both plus the allowed case.
Closes#4671
lifecycle::destroy logged a failed `nixos-container destroy` and returned
Ok, so the DestroyContainer node went green, the agent was unregistered,
and with purge the after_ok PurgeState deleted its state while the
container config and root still existed.
Propagate the error unless the container list, read after the failure,
no longer names the container. An unreadable list fails too.
`is_forbidden` and the webhook-secret load/regenerate path were the last two
entries on the shortlist in hyperhive/hyperhive#3950; the other three landed in
hyperhive/hyperhive#4650. Both are classification logic whose failure mode is
silence, which is why they are worth a test rather than a coverage line.
`is_forbidden` gates the arms that tell an operator a Forgejo admin PATCH was
refused for want of a scope, and which credential to delete and re-mint to fix
it. One of those PATCHes is `ensure_repo_creation_disabled` — the lockdown that
stops an agent creating a repo it owns and self-merging in it. Two tests: a 403
is recognised in both shapes the typed client produces (the spec-listed
`Forbidden` kind and the bare `UnexpectedStatusCode`), and nothing else is —
not a 401, whose remedy is the automatic re-mint one function down, and not a
transport error that never reached the forge at all.
`load_or_generate` grows the path-taking half `load_or_generate_at`, the same
seam `swarm-controller`'s `webhook::load_or_generate_at` already has and for
the same stated reason. Three tests over it: a valid stored secret is returned
verbatim and never rotated (the newline this module writes itself makes the
trim load-bearing, not defensive); a malformed one is replaced by a secret that
reaches *disk*, not just the caller, and is then stable; and each near miss —
empty, whitespace, 63 chars, 65 chars, right length with a non-hex char — is
refused. That last one is the security case: `Hmac::new_from_slice` accepts a
key of any length, empty included, so a relaxed check fails nowhere and just
keys every signature off a guessable value.
Every test was confirmed able to fail: six mutations of the code under test,
each watched red, then reverted. The two halves of the validity check and the
two arms of `is_forbidden` were broken separately, so neither test passes on
one arm alone.
- hive-c0re::webhook_secret::verify_signature — the HMAC comparison
verify_hmac (the only gate on the public webhook endpoint) delegates
to. Correct-signature and mismatched-signature (wrong secret, tampered
body) cases.
- hive-forge credential-helper get (host= check) — subprocess
integration tests since the check is inlined in run(), which reads
real stdin/env and prints real stdout. Host mismatch (error, token
withheld), host match (credentials printed), and no host= line
(backwards compat) cases.
- hive-priv::{validate_credential_name, validate_snapshot_name,
ensure_plain_filename} — the only gate on the root-privileged socket.
Empty/charset/dot/slash rejection, hive- prefix requirement, and
./../slash rejection respectively.
The prose added by this branch named the tracker item in seventeen
places, which check-issue-refs.sh rejects: a `#N` tag is dead weight for
anyone reading the public mirror, where no issue data exists. Each one
now states the fact it was pointing at — the parent field is gone — so
the sentence stands on its own.
Two of those lines also carried a rustdoc break: `[`write`]` in
topology.rs is ambiguous between the module's own `write` fn and the
`write!` macro, which `-D rustdoc::broken-intra-doc-links` fails. Spelled
`[`write()`]`, per rustdoc's own suggestion.
The host_config.rs rewrite is two lines rather than three so the doc
block stays under check-comment-blocks.sh's 30-line ceiling.
The topology doc keeps its filename and its second half (manager
special-casing, harness unit shape) — both are cross-referenced from
other pages and neither is about the parent field. Its first half is
rewritten: what topology.json is now, and a table of what the removal
took with it, so a reader who finds `<parent>` or `set-parent` in an old
issue thread learns it went away rather than moved.
The dashboard's tree-rendering section is marked dormant rather than
deleted: the walk is still in swarm.js and retiring it is the frontend
owner's call.
`topology.json` was a map of `name -> parent | null`, and that value fed
the whole agent hierarchy: `<parent>` / `<children>` recipient sentinels,
the reparenting API (CLI verb, wire verb, dashboard endpoints, DAG node),
the dashboard tree, the rebuild depth sort, and an unconditional
bind-mount grant giving every agent RW on its direct children's state.
Per the operator's ruling the field goes, and with it all of the above.
The file survives as what remains once the value is gone: the roster of
agent names, which is the set `ManageRootAgent` grants mounts over. It is
now a JSON array; `read` still accepts the old map shape and keeps its
keys, so a hive that upgrades across this does not blank its roster (and
so no capability holder loses its mounts for the length of that window).
Two sites kept their behaviour under a different recipient rather than
losing it. Both addressed `<parent>`, which the broker already resolved to
`operator` for a root agent, and every agent is now what that fallback
called a root:
- the harness's turn-failure / plugin-failure notification
(`Surface::send_to_parent` -> `send_to_operator`), and
- the send allow-list's always-permitted escape hatch, so an agent with a
restrictive allow-list still has a way to say it is stuck.
What is NOT preserved, deliberately: an agent with no capability no longer
sees any other agent's dirs. `ManageRootAgent`'s own grant is unchanged --
still every agent in the roster, still state RW + config RO, still no
`harness`.
The dashboard's reparenting control (the M0V3 picker) is deleted with its
CSS. The tree rendering that reads `ContainerView.parent` is left for the
frontend owner -- it degrades to a flat list with the field gone.
swarm-logs covers every host-tier unit this tool could reach, so the
second, capability-gated path into host journald earns nothing and is
removed outright rather than disabled behind a flag.
Removed end to end: the MCP tool definition + handler, the
GetHostJournal/HostJournal wire variants, hive-c0re's
dispatch_host_journal handler, the ReadHostJournal capability, and the
harness-side capability->--allowedTools gate. get_host_journal was the
only capability that mapped to an MCP tool, so allowed_capability_tools
could only ever return an empty vec; it goes too rather than linger as a
function that provably does nothing.
capabilities::has_cap/caps_for stay: #4624 gave ManageRootAgent's
bind-mount enforcement (hive-c0re/src/lifecycle/host_config.rs) a
second caller of has_cap, so they're no longer callerless once this
lands on top of it.
hive-sh4re's journal module (JournalPriority) had no consumer outside
this tool and is deleted.
An existing capabilities.json still naming read_host_journal does not
error: capabilities::prune_unknown drops unrecognised names with a
warn!, and an agent left with no capabilities has its entry removed. No
migration step is needed.
Untouched: hive-c0re/src/dashboard/journal.rs's
read_host_journal_response, which matches the name but is the private
helper behind the operator-only GET /api/journal-host dashboard route
and carries no capability check.