Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph

4,875 commits

Author SHA1 Message Date
atlas
244db367c4 swarm-controller: answer 404, not 503, for an issue-report on a nonexistent repo
get_issue_report mapped every error from Client::issue_report to 503
(StatusUnavailable) unconditionally. issue_report calls
issue_list_issues(org, repo, ...) first, and a Forgejo 404 there — a
typo'd or deleted repo — was flattened into the same 503 a genuine
forge outage produces, which tells a client to retry a request that
will never succeed.

Downcast the anyhow error back to forgejo_api::ForgejoError (same
pattern main.rs's wanted_error_status uses for
swarm_queue_client::Error) and check it structurally against the
ApiErrorKind::NotFound / UnexpectedStatusCode(404) shapes forgejo's
generated client produces for a 404, rather than string-matching the
rendered message. Only that case answers 404, naming the org/repo;
every other forge failure still answers 503. Updates the route's
OpenAPI response list to document the 404.

Closes #4701
2026-09-24 16:38:47 +02:00
atlas
d048fee698 docs/setup: point at the bootstrap policy file instead of inlining a copy (#4698) 2026-09-24 15:55:38 +02:00
atlas
e3c595857e swarm-bao: keep the bootstrap policy in one file, checked against the units using it (#4698)
setup.md's copy of the swarm-bootstrap policy still granted only the
controller's first six paths, while eight units now act with the
bootstrap token. The policy moves to
nix/host-modules/swarm-bao-bootstrap-policy.hcl, now covering every path
those units call. module-eval-bao-grants reads that file and fails
when a unit whose script uses the token calls a path the file does not
grant.
2026-09-24 15:55:38 +02:00
atlas
88e571a463 swarm-controller: refuse a new agent name the forge would reject
Every agent becomes a Forgejo user of the same name, and nothing upstream
of `CreateForgeUser` knew what Forgejo refuses: `admin`, `api`, `foo-` or a
41-character name passed name validation and failed one node into
provisioning with Forgejo's 422.

`nix/reserved-names.nix` gains the 22 reserved usernames of Forgejo
v16.0.5 (`models/user/user.go:639-680`) that `[a-z0-9-]` can spell, and
the bare `-` (`models/repo/repo.go:67`). The dot and underscore entries are
left out, since our charset cannot produce them. The header's admission
rule grows a third class — a username the forge refuses — because that is
a failure behind the refusal.

The shape rules are not literals, so they live in
`hive_types::forge_username_violation`: no leading `-`, no `--`, no
trailing `-`, at most 40 characters. Beside `is_reserved_name`, not in
`Ident::parse`: an `Ident` is also a hive, label, account and subagent
name, and parsing runs on every read of an existing name.

`create_agent` used to WARN on a reserved name, deliberately: an operator
with agents already created under a colliding name would otherwise be
unable to re-run creation. That reason is kept, and narrowed to what it
protects. A name breaking either rule is now refused with a 400 naming the
rule when the name is NOT in the swarm roster, and still only warned
about when it is, so re-creating an existing agent keeps working. The
roster is read only for a rule-breaking name; when it cannot be read, a new
name and an existing one look alike, and this warns as before. The
hive-collision warning is unchanged.
2026-09-24 15:16:32 +02:00
atlas
6a87b25488 swarm-controller: answer 503 on a disconnected queue, not 500 or a hang
put_matrix_account writes the credential to bao before checking that
the queue it must notify is actually connected — only that a queue is
configured, via state.status.as_ref(). While the queue is Pending or
Disconnected, client.flush().await hangs (async-nats does not process
commands during the initial connect retry), so the request hangs until
nginx times out and the credential is already stored. Add the
ensure_connected check every sibling queue route already makes
(wanted.rs, term_stream.rs, agent_state_stream.rs), placed immediately
before store::connect() so a request that cannot be delivered never
reaches the store.

Same shape, lower impact, in webhook::announce_knowledge_change: it
awaits publish/flush inline in the webhook handler, so a disconnected
queue can run past forgejo's short delivery timeout. Add the same
ensure_connected guard, warn and return.

set_agent_state and get_hive_wanted map every writer error to 500,
including swarm_queue_client::Error::NotConnected surfaced through
WantedWriter::view/set. Add wanted_error_status, which downcasts the
anyhow::Error back to the concrete type and maps NotConnected to 503
(retryable) while leaving every other failure at 500. Update both
routes' OpenAPI descriptions to say so.

Closes #4688
2026-09-24 15:15:59 +02:00
atlas
0bfe354b6d nix: move the reader-after-policy edges into their own colocation glue
The four readers' ordering after their policy units only applies where
the store and that reader share a host, so it is colocation glue and
does not belong in the reader modules (two of which are main modules).
glue-bao-readers-policy-order.nix now sets the after+wants edges, gated
on deploy.bao.enable AND the reader's own gate, so a store host without
a reader gains no stub unit.
2026-09-24 15:15:15 +02:00
atlas
cde3fec956 module-eval: pin each swarm-bao reader's ordering after its policy unit 2026-09-24 15:15:15 +02:00
atlas
c92bb0dce7 nix: order each swarm-bao secret reader after the policy unit writing its role
The four readers (matrix-token, queue-agent, grafana-oidc, otel-oidc) log
in against a cert-auth role that their own swarm-bao-*-policy unit writes.
They were ordered after swarm-bao-pki and the store's container but not
after that unit, so on an apply a reader could log in before its role
existed and be refused by `allowed_common_names` until a retry landed
after the role did.

After= plus Wants= on the policy unit, never Requires=: the policy unit
skips by ConditionPathExists once the bootstrap token is gone, and a
skipped unit counts as done for ordering.
2026-09-24 15:15:15 +02:00
atlas
aa719da571 nix: ship the journals of the units an apply can leave failed
The units on the path a deploy takes to TLS, the store's grants and the
swarm collector itself were not on the host collector's journald
allowlist, so an ingest outage one of them caused showed in the store
only as every source going quiet at once.

Each module names its own units, per the option's rule:
- hive-tls.nix: hive-tls-ca, swarm-services-cert
- hive-gateway: hive-gateway-self-signed-cert (self-signed mode only)
- swarm-bao.nix: the seven grant units beside
  swarm-bao-services-issuer-policy
- swarm-otel.nix: container@<machine>, and
  nixos-rebuild-switch-to-configuration, the transient unit nixos-rebuild
  runs the activation in and whose syslog lines carry its status

The module-eval arm pins each unit as both listed and defined, since a
listed name that matches nothing is silent.
2026-09-24 15:14:44 +02:00
atlas
fdb847cd87 hive-priv: stop run_forge_admin's bail! from reproducing --password
The comment above the bail! named the gap itself: stderr went through
redact_secret_line, but args.join(" ") did not. hive-c0re passes a
live --password value as an argument on user create and
change-password, so any non-zero forgejo admin exit put the plaintext
password in the error string, and from there into hive-c0re's warn!
log (journal + VictoriaLogs) and hivectl's returned error.

Add describe_forge_admin, a pure function that names the invocation by
its leading verb path and stops at the first flag, same approach as
hive-c0re's own describe_forge_admin (forge/mod.rs, from #2936) and
for the same reason: the verbs are a closed set this crate chooses,
argument values never are, so an allowlist over shape excludes any
future secret-bearing flag by construction instead of by someone
remembering to redact its value.

Unit tests cover: a --password value dropped, the --password=value
form also dropped (it starts with '-', so the take_while excludes the
whole argument), and a control that the verb path still appears.
Closes #4670.
2026-09-24 14:45:51 +02:00
atlas
9986da4c7f hive-forge: attach-issue/attach-comment fail on a missing download URL
print_url only printed when browser_download_url was Some, so a 2xx
upload response with no URL printed nothing and still exited 0 — a
caller doing url=$(hive-forge attach-issue …) got an empty string with
no signal anything went wrong.

require_url now errors, naming the attachment's id/name, when the
forge returns no download URL.
2026-09-24 13:50:24 +02:00
atlas
6ac402ed2d hive-forge: labels remove refuses unknown names and verifies removal landed
labels <n> remove used to silently skip an unknown label name (lookup_id
returning None just did nothing) and discard each remove call's Result,
so a typo or a forge-rejected removal (403, insufficient permission)
both exited 0 with the label untouched.

Resolve names through the same resolve_ids labels add already uses, so
an unknown name is refused before anything is removed. After removing,
re-fetch and diff against what was requested — following assign.rs's
same before/after check — and bail! naming any label still present.
2026-09-24 13:50:24 +02:00
atlas
6d10c1367e hive-priv: don't leave a partial snapshot export at dest on failure
send_agent_snapshot_to_file created `dest` with create_new(true) before
checking whether the parent snapshot exists, and left dest behind on any
spawn()/wait_with_output() failure too. Cleanup only ran on the one
remaining path: btrfs send spawning fine and exiting non-zero.

Every other failure listed in the wire doc left a zero-byte export, and
because dest already existed, a retry against the same --dest always hit
the no-overwrite guard — the only way out was an operator manually
removing the file.

Fixes #4687.

- check_send_snapshot_preconditions() validates the snapshot and its
  optional parent before dest is touched at all.
- open_export_dest() combines that check with the create_new open, so the
  ordering can't drift apart again, and is unit-testable without btrfs.
- PartialExportGuard removes dest on drop unless disarmed, covering every
  failure path after dest is created (including the spawn/wait `?`s that
  previously leaked it), not just the btrfs-exit-nonzero case.
2026-09-24 13:49:02 +02:00
atlas
295925abb6 swarm-controller: don't fold every 422 into "already exists" on user/repo create
Forgejo answers 422 for six different causes on admin user create
(ErrUserAlreadyExist, ErrEmailAlreadyUsed, ErrNameReserved,
ErrNameCharsNotAllowed, ErrEmailInvalid, ErrNamePatternNotAllowed) and
several on repo create, but is_already_exists() treated every one of
them as a conflict. A reserved or otherwise-refused name silently
folded to Done, so CreateForgeUser reported success with no user
created, and the graph's real failure only surfaced one node later as
a misleading AddRepoMember error.

ensure_agent_user and ensure_org_repo now trust a 409 unconditionally
(folds_into_success) but confirm a 422 with a follow-up user_get /
repo_get before folding it to success; an unconfirmed 422 fails with
forgejo's own message at error level. Webhook registration still uses
the old is_already_exists — it has no comparable follow-up read, so it
is out of scope here.

Closes #4678
2026-09-24 13:45:51 +02:00
atlas
e3fefb8c5f docs: fix markdown indent treefmt wants after removing the ownership-checks bullet 2026-09-24 12:24:32 +02:00
atlas
a65dbee982 docs: delete two more stale manager-override claims
destroy has no manager-name check (actions.rs:837 says the root
container is 'destroyable like any other'), and crash-watch has no
name check at all (polls every managed container uniformly). The
pointer to broker.rs/actions.rs/crash_watch.rs for 'owner-check logic'
is also stale — none of the three hold any.
2026-09-24 12:24:32 +02:00
atlas
d2c4c8d4a3 docs: drop dead-claim narration for the retired loose-ends override
The PR retiring loose-ends' manager-only visibility replaced the false
claim with prose narrating its retirement, which still adds lines for
a removal. Delete the dead claim outright instead of documenting that
it used to be true: drop the three added sentences in
agent-hierarchy.md, and shorten socket_server/mod.rs's authority
comment to state only what's true now (tool-group membership for the
orchestration verbs) rather than adding an explanatory clause for the
removed capability check.
2026-09-24 12:24:32 +02:00
atlas
f53508cb2b docs, hive-c0re: retire prose describing the removed get_loose_ends target
Follow-up to 729c5b4f42, which dropped the
`agent` parameter from get_loose_ends and deleted Capability::QueryAgentState
with it. Three prose sites still describe the interface that commit removed:

- socket_server/mod.rs's module doc claimed authority on this socket derives
  from "capabilities for the hive-wide queries and tool-group membership for
  the orchestration verbs". There are no capability-gated queries left —
  `git grep -n has_cap -- hive-c0re/src/socket_server/` returns nothing, and
  require_group("scheduling") is the only gate in dispatch_orchestration.
- agent-hierarchy.md listed loose-ends visibility ("manager sees hive-wide,
  sub-agents only their own") as one of the manager-only overrides that
  "exist across hive-c0re today". It doesn't, and the sentence pointed a
  reader at loose_ends.rs for owner-check logic that is no longer there.
- conventions.md's Loose-ends wire shape still spoke of "the agent-flavour
  and manager-flavour requests" and "the agent-flavour list". There is one
  request shape.

No behaviour change: prose only.

Refs #4480
2026-09-24 12:24:32 +02:00
atlas
7dd52207de dashboard: fail purge-tombstone closed when the container list is unreadable
post_purge_tombstone's live-container guard used lifecycle::list().await
.unwrap_or_default(), so a failed list() (hive-priv socket unreachable,
restarting, ...) read as 'no live containers' and let the purge proceed —
deleting agent_state_dir/applied_dir for an agent whose container may
still be running. Factor the decision into check_not_live(), a pure
helper that refuses on both 'still listed' and 'list unreadable', with
unit tests covering both plus the allowed case.

Closes #4671
2026-09-24 12:20:08 +02:00
atlas
11cd2038c9 hive-priv: refuse symlinks when writing the bridge-DNS marker
The marker lives in the container's own /etc, which root inside the
container controls, and hive-priv writes it as host root. The write
followed a symlink at the leaf and at etc, so an absolute symlink
planted in the container redirected a host-root truncate-and-write to a
host path.

Open etc with O_DIRECTORY|O_NOFOLLOW and the leaf relative to that fd
with O_NOFOLLOW, and require a regular file (opened O_NONBLOCK so a
FIFO cannot hang the helper). The marker keeps mode 0644.

Closes #4667
2026-09-24 12:19:10 +02:00
atlas
18bd8dd2c7 hive-c0re: fail a nixos-container destroy that leaves the container in place
lifecycle::destroy logged a failed `nixos-container destroy` and returned
Ok, so the DestroyContainer node went green, the agent was unregistered,
and with purge the after_ok PurgeState deleted its state while the
container config and root still existed.

Propagate the error unless the container list, read after the failure,
no longer names the container. An unreadable list fails too.
2026-09-24 12:18:34 +02:00
atlas
5b54b6ddc8 swarm-bao: fail the unit when the services root cannot be read, instead of replacing it
A failed `bao read` of the services root made the checkend pipeline
non-zero, so the unit deleted a working root and minted a new trust
anchor. A failed `bao list` of the issuers likewise looked like an empty
mount. Both now fail the unit, which retries on its own restart budget;
the root is replaced only when openssl parsed the returned certificate
and -checkend said it expires inside a leaf's window.

Closes #4663
2026-09-24 11:00:20 +02:00
atlas
c0c031a5e4 swarm-bao: let the journald receiver read the host-linked journal
The container's collector has never shipped a line. `journalctl --follow`
— which the journald receiver passes unconditionally — scopes itself to
the current boot unless `--merge` is given too, and `--link-journal=host`
makes /var/log/journal the HOST's journal tree, where this container's
current boot has no entry. journalctl exited 1 with "No journal boot
entry found for the specified boot (+0)" and the receiver respawned it
every ~2s, so nothing was ever read and nothing was ever exported.

`merge = true` is the receiver's key for `--merge` (buildArgs() in
pkg/stanza/operator/input/journald/config_linux.go at tag
receiver/journaldreceiver/v0.151.0, the deployed collector's version),
and --merge is what clears the implicit boot scope in journalctl.c
(systemd v260.4, the version on the host).

The module-eval arm pins both halves: the boot filter is gone AND the
directory is still the host-linked one — either alone is satisfiable by
the broken config.
2026-09-23 23:19:31 +02:00
atlas
33da51382e subagent: make interrupt cancel a goal run, not just its turn
`interrupt` killed the child and trusted that to end the run. It didn't:
the default signal is SIGINT, and `claude --print` handles SIGINT by
writing its terminal `result` event and exiting **zero**. A zero exit is
`TurnEnd::Complete`, so `spawn_and_track`'s killed-turn early return —
the code that was supposed to stop the run — never fired, and the loop
went straight on to spawn the next goal turn. An interrupted run carried
on to completion; the interrupt changed nothing about the outcome.

Measured, not inferred: `claude --print --verbose --output-format
stream-json`, SIGINT'd mid-turn, exits 0 on every run.

So record the reason instead of inferring it. `interrupt` writes
`StopReason::Cancelled` — the same path `goal_reached`/`need_help`
already use — while it still holds the `running` lock, so there is no
instant in which the name has lost its cancel handle but not yet gained
its stop reason. `plan_after_turn` checks stop reasons ahead of the goal,
so the run stops whichever way the child ends up exiting. `force`'s
SIGKILL still produces a `TurnEnd::Killed`, and `status` still prefers
the kill record it already had.

`continue`'s budget reset is untouched: it clears the stop reason and
hands back a fresh allowance, which is what makes a cancel a pause an
operator can undo.

The test spawns the real continuation loop against a fake claude that
exits 0 on SIGINT, interrupts it, and asserts turn two never starts —
it reports `left: 2` without the fix.
2026-09-23 23:14:13 +02:00
müde
11ec050ca8 hive-tls: give swarm-services-cert the cmp it tests the root with
`cmp` lives in diffutils, not coreutils, so the root-changed test exited
127 with "command not found". Inside `if ! cmp -s`, a 127 reads as
"differs" and errexit never sees it, so `rootchanged` was 1 on every run
and the trust-bundle rebuild it guards bounced `hive-tls-ca` after each
issuance — the exact "only a changed one, or every boot would bounce a
unit with nothing to do" the comment there rules out.

Found in the journal of a hive that had just issued a leaf successfully:
the unit logged "the services root changed" on a run where the store had
left the issuer alone.
2026-09-23 22:18:32 +02:00
müde
5704c0c583 nix: unbreak the swarm-services leaf the gateway waits on
Three defects in the store-issued path, each of which alone kept nginx
from starting at all. The gateway's cert import `Requires=` this leaf, so
a leaf that is never issued is not a name mismatch — it is an empty
listener, and the swarm's own forge stopped answering on :443.

`swarm-services-cert` declared `Before=hive-tls-ca` for the trust
bundle's sake while also being `After=` the store's container, which is
itself `After=hive-tls-ca`. systemd resolved the cycle the only way it
can, by deleting the job, so the leaf went unissued on every activation.
The edge is gone; the bundle converges the other way round, through the
restart this unit already performed when the root it wrote was new.

The `pki` mount was enabled without `-max-lease-ttl`, so bao clamped the
30-year root to the 768h default and then refused every issue call,
because a leaf of the mount's own default length would outlive the CA
signing it. The mount is tuned on every run, the role pins a 720h leaf,
and a root that can no longer cover one is replaced rather than left to
refuse forever. A hive tests its own copy of that certificate against the
same threshold, so both ends reach a fresh leaf without signalling.

`swarm-bao-pki` mints the services-issuer leaf that opens the mount, but
only `swarm-bao-certs` required it. `RemainAfterExit` plus an
already-active unit means an activation that ADDS a leaf mints nothing —
which is how a host whose config named `services-issuer.pem` came to have
no such file. A target wants it now, like every sibling granting unit.
2026-09-23 21:36:46 +02:00
atlas
22a87f7268 docs/swarm/{ca,secrets}.md: reword the 8 vale write-good.Passive / Microsoft.Contractions hits
Active-voice / contraction rewrites only, no meaning changes; the
allowed_domains/SANs sentence and the published-cert table cell were
checked against the parked services-issuer/role-type/vhost-scope
questions on #4622 and don't touch any of them.
2026-09-23 21:00:02 +02:00
atlas
f4df4fc4a9 nix: issue the swarm-services leaf from bao's pki mount
The `pki` mount had no issuer and no principal could log in to it, so the
swarm's service certificates were still minted by two openssl hops from a
root key on disk. Close both halves and retire the openssl path with them.

The mount now generates its own root, once. The granting unit asks bao
whether an issuer already exists (`bao list pki/issuers`) before calling
`pki/root/generate/internal`, so a rebuild or a reboot re-asserts the role
and the grant without touching the anchor — a root that changed per boot
would invalidate every certificate issued under it and every browser
taught to trust it. The guard asks the store rather than looking for a
marker file on this host's disk: a file is a claim about a mount that may
have been restored from a snapshot or disabled and re-enabled underneath
it.

`swarm-services-issuer` stops being an inert policy. A fourth cert-auth
role attaches it, following the shape the controller, the publisher and
matrix-ctl already use, and glue-bao-tls.nix signs the leaf carrying its
CN — that credential is what opens the mount, so it cannot come out of it.

`swarm-services-cert.service` logs in with that leaf, calls
`pki/issue/swarm-services`, and writes the result to the path
hive-tls.nix already wrote and the gateway already copies from. The
sub-CA layer does not move; it stops existing. The role's
`allowed_domains`, read from the same `swarm.serviceDomains` the SANs
come from, enforces at issue time what the sub-CA encoded in x509
`nameConstraints`, and with the root inside the mount there is nothing
left for an intermediate to be an intermediate of.

Not a flag day: the issuing root is published beside the leaf as
`swarm-services-root.pem` (0644) and joins `trust-bundle.pem`, where the
swarm root still sits. A leaf chaining to the old sub-CA and one issued
by the store both verify against the same bundle, so hives can be
rebuilt in any order. The same file is what an operator hands a browser
— readable without a store login, which matters because every listener
demands a client certificate.

The eval-time warning about uncovered service names is gone rather than
reworded. It fired on "this host does not hold the swarm root key", which
was the reason a hive could end up serving its own leaf on a
swarm-service name. Every hive now asks the store with its own identity,
so that stopped being the thing that decides.

Closes #4586
2026-09-23 21:00:02 +02:00
atlas
ebcc5bde89 swarm-bao: shrink the read-grant comment under the comment-block max 2026-09-23 18:55:02 +02:00
atlas
87174863ca swarm-bao: grant the controller read on the agent credential prefix
mint_and_verify reads the queue credential back before writing it, so a
re-run keeps the value a live agent already authenticates with. that read
is read_optional, which maps only a 404 to absence — so with create/update
alone every mint aborted on a 403 at its first store read.

read on the same paths the stanza already grants create and update, and
nothing else: no list, no delete, no patch.
2026-09-23 18:55:02 +02:00
atlas
e2ff4f5281 swarm-bao: give the store's collector an explicit self-telemetry port
8890, so it stops claiming the hive collector's 8888 in the shared netns.
Extends the module-eval port case to all three tiers.
2026-09-23 18:04:33 +02:00
atlas
5cec69bd21 otel.nix: trim the StartLimit comment block to the load-bearing points
Cut ~27 lines of blackout-measurement and cross-reference narrative
(already in the PR body / issue) down to the three things a reader
actually needs at this call site: the [Unit]-vs-[Service] trap, why
Restart is absent, and the window-vs-burst constraint.
2026-09-23 17:22:51 +02:00
atlas
596c0c17bc otel: back the hive collector off a failed bind instead of burning its start limit
Every deploy on a hive host, the replacement opentelemetry-collector
reaches bind() while the outgoing process still holds
127.0.0.1:8888 (its self-scrape endpoint). nixpkgs sets
Restart = "always" with no RestartSec, so the unit spends its five
default attempts in under two seconds, hits start-limit-hit and stops
retrying — ~27s of telemetry blackout per deploy.

RestartSec = 5 with a 12-attempt burst over a 120s window rides the
race out instead: the blackout ends within one interval of the port
coming free, and 55s of it being held is survivable where 2s was not.

StartLimitBurst/StartLimitIntervalSec go at the systemd.services attr
level, which NixOS renders into [Unit]; under serviceConfig systemd
ignores them silently. module-eval-hive-otel asserts the placement.
2026-09-23 17:22:51 +02:00
atlas
2c066cc871 hive-c0re: pin the two security boundaries that had no test
`is_forbidden` and the webhook-secret load/regenerate path were the last two
entries on the shortlist in hyperhive/hyperhive#3950; the other three landed in
hyperhive/hyperhive#4650. Both are classification logic whose failure mode is
silence, which is why they are worth a test rather than a coverage line.

`is_forbidden` gates the arms that tell an operator a Forgejo admin PATCH was
refused for want of a scope, and which credential to delete and re-mint to fix
it. One of those PATCHes is `ensure_repo_creation_disabled` — the lockdown that
stops an agent creating a repo it owns and self-merging in it. Two tests: a 403
is recognised in both shapes the typed client produces (the spec-listed
`Forbidden` kind and the bare `UnexpectedStatusCode`), and nothing else is —
not a 401, whose remedy is the automatic re-mint one function down, and not a
transport error that never reached the forge at all.

`load_or_generate` grows the path-taking half `load_or_generate_at`, the same
seam `swarm-controller`'s `webhook::load_or_generate_at` already has and for
the same stated reason. Three tests over it: a valid stored secret is returned
verbatim and never rotated (the newline this module writes itself makes the
trim load-bearing, not defensive); a malformed one is replaced by a secret that
reaches *disk*, not just the caller, and is then stable; and each near miss —
empty, whitespace, 63 chars, 65 chars, right length with a non-hex char — is
refused. That last one is the security case: `Hmac::new_from_slice` accepts a
key of any length, empty included, so a relaxed check fails nowhere and just
keys every signature off a guessable value.

Every test was confirmed able to fail: six mutations of the code under test,
each watched red, then reverted. The two halves of the validity check and the
two arms of `is_forbidden` were broken separately, so neither test passes on
one arm alone.
2026-09-23 17:00:44 +02:00
atlas
e0673b6192 tests: cover the three security boundaries that had none (#3950)
- hive-c0re::webhook_secret::verify_signature — the HMAC comparison
  verify_hmac (the only gate on the public webhook endpoint) delegates
  to. Correct-signature and mismatched-signature (wrong secret, tampered
  body) cases.
- hive-forge credential-helper get (host= check) — subprocess
  integration tests since the check is inlined in run(), which reads
  real stdin/env and prints real stdout. Host mismatch (error, token
  withheld), host match (credentials printed), and no host= line
  (backwards compat) cases.
- hive-priv::{validate_credential_name, validate_snapshot_name,
  ensure_plain_filename} — the only gate on the root-privileged socket.
  Empty/charset/dot/slash rejection, hive- prefix requirement, and
  ./../slash rejection respectively.
2026-09-23 14:37:54 +02:00
atlas
549156e55f swarm-otel: persist the journald cursor across collector restarts
The swarm-tier collector's journald receiver had no storage extension, so
it started each run with no cursor: journalctl --follow --lines=0 ships
only what arrives after the receiver starts. Every collector restart
therefore dropped whatever was written to the journal while it was down,
silently — no error and no replay.

Wire the receiver to a file_storage extension, matching the agent-tier
collector in nix/agent-modules/otel.nix, so a restart resumes from the
persisted cursor instead.

Refs #4527.
2026-09-23 12:52:32 +02:00
atlas
9d339b1b58 module-eval: name the grafana helper the refusal readers are shaped after
Comment-only. The helper it points at is `grafanaRefusedFor`, not an
unnamed one.
2026-09-23 10:11:42 +02:00
atlas
d3e4951cc8 swarm-bao: refuse a remote reader that named seven of the eight leaves
The four-way client-cert split gives each store reader its own leaf, and
three of the four readers render only where their own leaf exists. On a
host that mints its own PKI glue-bao-tls.nix defaults all eight, so there
is nothing to do; on a hand-configured remote-store hive, omitting one
pair used to mean that unit silently did not render — a privilege-
narrowing unit absent from a green build, with the missing unit as the
only evidence.

Each of the three now asserts its own pair, shaped after
swarm-grafana.nix's haveClientIdentity assertion and named to the pair it
needs. What differs from Grafana's is the gate: these fire only where the
host demonstrably reads the store (it holds deploy.bao.clientCertFile and
clientKeyFile) and the consumer is on. A host with no store identity is
the supported no-store deployment and still evaluates; the collector's
no-secret degrade is untouched, because that host holds no clientCertFile
either.

Also rewords three passive-voice sentences in docs/swarm/secrets.md that
vale flagged, and documents what the refusal costs and where it stays
silent.
2026-09-23 10:11:42 +02:00
atlas
f1445b4c8b swarm-bao: give each hive-cert consumer its own bao identity
Four units read one path each out of the store, and all four logged in
holding `deploy.bao.clientCertFile` — the hive's own leaf. Bao identifies
a principal by the subject of the certificate it presents, so four
readers behind one certificate were ONE principal, and the only grant
expressible was the union of what the four need: read on
`swarm/agents/*`, `swarm/hives/<hive>/*` and `swarm/services/*`. The unit
fetching Grafana's OIDC client secret could fetch every agent credential
in the swarm; the one fetching this hive's matrix token could fetch
Grafana's. Least privilege was not misconfigured here, it was
unrepresentable.

Each now holds a leaf, a cert-auth role and a policy of its own, and each
policy is the single `secret/data/…` path that unit's own script names —
spelled to the leaf, not to a prefix, the way matrix-ctl's already is.
Following the four exemplars in-tree rather than building a mechanism:
`signLeaf` mints the leaves, `swarm-bao.nix` writes the roles from the
bootstrap token, the consumers name their own pair.

Two of the four are written PER HIVE and two are not, which is the shape
of the paths rather than a preference. A matrix appservice token and a
queue credential live under `swarm/hives/<name>/` and every hive runs a
reader for its own, so one role for all of them would have to be granted
`hives/*` — letting one hive read another's, a reach no hive has today.
An OIDC client secret lives under `swarm/services/<client-id>/` and a
swarm registers each exactly once, so one role each is enough. The
per-hive subjects are `<prefix>-<hive>` and swarm.nix reserves every
composed spelling as a hive name, so a hive cannot be named into another
hive's role.

The shared leaf stays: hive-c0re still passes it into its container, the
`bao` CLI wrapper still defaults to it, and the three
`glue-*-bao-identity.nix` files derive the PKI directory from it.

module-eval-bao-grants gains a negative arm per principal — each pins the
three stanzas the hive's leaf carried and the two wildcards a later
widening would reach for, so a policy that grows fails here rather than
in a store. Plus the consuming side: repointing a unit back at the hive's
leaf would evaluate, deploy and log in, and silently restore the union.

A hive that reads a store on another machine now places one leaf per
principal instead of one shared by four. That cost is the point, and
docs/swarm/secrets.md lists the pairs.
2026-09-23 10:11:42 +02:00
flake-bot
ded5b08258 nix flake update 2026-09-22 23:15:29 +02:00
atlas
d547ae58e6 docs: describe the flat container list, not a dormant tree
mara's ruling on the review: documentation describes functionality as
is. The dashboard doc carried a "Topology tree" section marked dormant,
still spelling out the indent lanes, joints and continuation bars the
renderer paints — for a renderer that, with no parent field to walk,
puts every container at depth 0 and emits no prefix column at all. A
section labelled dormant is still a section describing a feature the
code does not have.

Each one now states what the page renders today: SW4RM's C0NTAINERS is
a flat alphabetical list, one row per container, no indent and no glyph;
the tree section says nothing nests and names the code that decides so;
the selection bar lists the bulk actions it has, without a note about
the M0V3 picker it doesn't (agent-hierarchy.md's removal table is where
that record belongs).

Two more the -U15 context sweep turned up outside that section, neither
naming a removed identifier so neither reachable by grep: approvals.md
told an agent to clone "the child's" config repo, and hivectl.md sold
`agent restart` as a way round "the agent hierarchy".
2026-09-21 22:56:56 +02:00
atlas
336ed5a010 docs+comments: say what changed instead of tagging the tracker item
The prose added by this branch named the tracker item in seventeen
places, which check-issue-refs.sh rejects: a `#N` tag is dead weight for
anyone reading the public mirror, where no issue data exists. Each one
now states the fact it was pointing at — the parent field is gone — so
the sentence stands on its own.

Two of those lines also carried a rustdoc break: `[`write`]` in
topology.rs is ambiguous between the module's own `write` fn and the
`write!` macro, which `-D rustdoc::broken-intra-doc-links` fails. Spelled
`[`write()`]`, per rustdoc's own suggestion.

The host_config.rs rewrite is two lines rather than three so the doc
block stays under check-comment-blocks.sh's 30-line ceiling.
2026-09-21 22:43:16 +02:00
atlas
179f873722 docs: retire the agent hierarchy from every page that described it
The topology doc keeps its filename and its second half (manager
special-casing, harness unit shape) — both are cross-referenced from
other pages and neither is about the parent field. Its first half is
rewritten: what topology.json is now, and a table of what the removal
took with it, so a reader who finds `<parent>` or `set-parent` in an old
issue thread learns it went away rather than moved.

The dashboard's tree-rendering section is marked dormant rather than
deleted: the walk is still in swarm.js and retiring it is the frontend
owner's call.
2026-09-21 22:08:47 +02:00
atlas
d94bc2188d topology: drop the parent field and the hierarchy it fed
`topology.json` was a map of `name -> parent | null`, and that value fed
the whole agent hierarchy: `<parent>` / `<children>` recipient sentinels,
the reparenting API (CLI verb, wire verb, dashboard endpoints, DAG node),
the dashboard tree, the rebuild depth sort, and an unconditional
bind-mount grant giving every agent RW on its direct children's state.

Per the operator's ruling the field goes, and with it all of the above.
The file survives as what remains once the value is gone: the roster of
agent names, which is the set `ManageRootAgent` grants mounts over. It is
now a JSON array; `read` still accepts the old map shape and keeps its
keys, so a hive that upgrades across this does not blank its roster (and
so no capability holder loses its mounts for the length of that window).

Two sites kept their behaviour under a different recipient rather than
losing it. Both addressed `<parent>`, which the broker already resolved to
`operator` for a root agent, and every agent is now what that fallback
called a root:

- the harness's turn-failure / plugin-failure notification
  (`Surface::send_to_parent` -> `send_to_operator`), and
- the send allow-list's always-permitted escape hatch, so an agent with a
  restrictive allow-list still has a way to say it is stuck.

What is NOT preserved, deliberately: an agent with no capability no longer
sees any other agent's dirs. `ManageRootAgent`'s own grant is unchanged --
still every agent in the roster, still state RW + config RO, still no
`harness`.

The dashboard's reparenting control (the M0V3 picker) is deleted with its
CSS. The tree rendering that reads `ContainerView.parent` is left for the
frontend owner -- it degrades to a flat list with the field gone.
2026-09-21 22:08:47 +02:00
atlas
392f16cbc0 hive-forge: ci-rerun --run refuses on a PR-triggered run too
--run <id> against a run that was itself PR-triggered has the identical
defect #4632 fixed for --pr: a workflow_dispatch run writes no commit
status, so it can't clear the red (pull_request) check on that PR's sha
no matter how the dispatched run turns out.

resolve_run already distinguishes this case -- it recognizes a run's
prettyref as a PR pseudo-ref (#<n>) via pr_number_from_run_ref, then
used to call branch_for_pr to keep going. It now bails with the same
pr_refusal_message --pr uses instead, before ever building a dispatch
request. branch_for_pr has no other caller (--pr refuses before
touching it too, since #4632), so it's removed rather than left dead.

Push (non-PR-triggered) and --branch are untouched.

Regenerated docs/tools/forge-cli.md from clap help; corrected
docs/tools/forge.md's claim that --run always works.
2026-09-21 21:37:22 +02:00
atlas
8de85729cc hive-forge: make ci-rerun --pr help text match the refusal it prints
The doc comment still described the old (never-shipped) behavior --
exactly the lie this whole PR exists to remove, just in --help instead
of the module doc or forge.md. Say what the flag actually does: refuse,
because a workflow_dispatch run writes no commit status and can't clear
a red (pull_request) check.

Regenerated docs/tools/forge-cli.md (derived from clap help via
'hive-forge markdown-docs').
2026-09-21 20:55:34 +02:00
atlas
92100ac1f8 hive-forge: ci-rerun --pr refuses instead of lying about a fix it can't achieve
A workflow_dispatch run writes no commit status, so ci-rerun --pr could
never clear the red (pull_request) check it claimed to be re-running for
- it dispatched a fresh run and printed a success message regardless,
even though the check stays red no matter how that run turns out.

--pr now refuses up front, before dispatching, naming the mechanism and
the working alternative (re-run from the web UI). --run and --branch are
unchanged: --run's own PR-pseudo-ref resolution and --branch's direct
dispatch are both untouched.

Fixes the exit-code/honesty defect from #4613; the workflow_dispatch vs.
pull_request event-type question (whether to close+reopen the PR to fire
a real pull_request event) is a separate, parked decision.
2026-09-21 20:55:34 +02:00
atlas
afdfce67ec agent: fetch this agent's own swarm-queue credential from the store
Every agent on a hive authenticates to the swarm queue with the same
hive-scoped OIDC client, so at the auth callout one agent is
indistinguishable from its co-hived neighbours. The commit before this
one mints a secret per agent at swarm level into
secret/swarm/agents/<agent>/queue; nothing read it.

Read it here, and read it from the container itself. A hive courier in
the path would be the hive vouching for which agent this is, which is
the property a per-agent credential exists to remove -- so the agent
logs in to the store with the certificate hive-agent-bao-identity
already proves it can log in with, and reads its own path. The store
certificate is for reaching the store and nothing else: what the new
unit writes to /run is the secret it read back, and nothing hands a
BAO_CLIENT_* path to anything queue-shaped.

The read needs no policy change. render_agent grants read on
secret/data/swarm/agents/<agent>/*, which covers this path and the
bao-mtls one beside it alike -- which is also why this unit degrades
where the identity check fails. A refusal this unit sees and that check
did not cannot be a policy that drifted; it is an object not yet minted,
the ordinary state of every agent created before its swarm knew to mint
one.

The harness resolves the path and reports which credential this agent
can present. It does not yet present it: the auth-callout responder
still verifies only the hive-scoped token, and an agent offering a
credential nothing on the other end reads back would simply be refused.
Teaching swarm-nats-auth to read the same path is the next slice.
2026-09-21 20:44:52 +02:00
atlas
10427467b3 docs: clear the vale errors the queue-credential prose introduced
`vale --minAlertLevel=error docs` is a CI gate and origin/main passes it
with zero findings, so the twelve this branch added were a red build, not
a backlog: eight Microsoft.Contractions, one write-good.So, two
write-good.Passive, across the new credential-matrix row, the backfill
runbook and the mint-identity help text.

Contractions and the sentence that started with "So" are mechanical.
The two passive hits are rewrites: "`--hive` is required" becomes
"`--hive` has no default", which is the actual claim -- the flag has no
value to fall back on -- and the help text's "the agent's queue secret is
left exactly as it is" becomes "it leaves an existing queue secret
exactly as it stands", which also names who does the leaving.

docs/tools/swarmctl-cli.md is regenerated, not hand-edited; the wording
lives in swarmctl's clap doc comments.
2026-09-21 20:38:55 +02:00
atlas
8b6dc72526 swarm-controller: pin what the backfill route queues
The mint route shipped without tests while the create route beside it
has three, so the two properties that make it a backfill rather than a
second creation route were unasserted: that it refuses before queuing,
and that it queues the mint node and nothing else.

Queuing a config-repo scaffold or a deploy against an agent that already
exists is the failure the second of those catches, and it is invisible
from the status code -- a version that queued the whole creation graph
would answer 200 with a node id just the same.

The agent name gets its own arm. create_agent's hive is checked against
the swarm roster, which incidentally rejects a name that is not an
identifier; an agent name has no roster to check against, so the parse
is the only thing between a traversal and a store path built out of it.

MintAgentIdentityResponse derives Clone + Debug to match
CreateAgentResponse -- expect_err on the refusal arms needs Debug on the
success type.
2026-09-21 20:38:55 +02:00