Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph

4,922 commits

Author SHA1 Message Date
iris
88c386c96c swarm-ui: move dynamic agent-term tabs into the top nav bar
mara, PR review: "the dynamic tab should be in the top bar, not a new
one below". Moves the .shell-tabs group from its own sticky row under
the header into .shell-nav itself, right after the nav indicator.

Also adds a third re-measure effect for the sliding nav indicator,
keyed on tabs.length: with tabs inline in the same flex row the
indicator measures, closing a background tab (no navigation) can
shrink the row without the hop effect's own re-measure ever firing.
Same reflow-not-navigation reasoning as the existing resize-listener
effect.
2026-09-28 23:17:44 +02:00
iris
7bf4cfbc4c swarm-ui: fix specificity claim in AgentTermPreview.css comment
argus's review on PR #4784 caught a wrong technical claim: the
comment said the .ui-agent-term-preview-full override rules relied
on source order because they had equal specificity to the
un-modified rules above. They don't — each override selector adds
one more class (the .ui-agent-term-preview-full prefix) than what it
overrides, so they're strictly more specific and win regardless of
file order. Corrected the comment to say so.
2026-09-28 23:17:44 +02:00
iris
c460e91bd1 swarm-ui: full-tab agent terminal (#4506)
Adds a full, non-capped agent terminal reachable from a new expand
trigger on the embedded AgentTermPreview (the detail-panel preview on
AgentsPage stays as-is, just gains the trigger). Opens
/agents/:name/terminal in a new dynamic tab in Shell's header, next to
the static nav row — tabs persist across a reload via useDynamicTabs,
a small localStorage-backed hook built on @hive/shared's existing
settings-storage primitive.

AgentTermPreview gains two new props to support both mounts from one
component: fullHeight (drops the 12em preview cap, fills its page)
and showHeaderBadges (default true — lets a future caller that
already shows turn_state/model/ctx/cost elsewhere suppress this
cluster; AgentsPage doesn't use it, see below).

Deviation from the originally posted plan (issue comment 80596): that
plan proposed AgentsPage's embedded preview pass showHeaderBadges as
false, reasoning the detail panel already duplicates that info.
Checked the actual code before implementing — it doesn't; AgentRow/
AgentTypes.ts carry none of turn_state/model/ctx/cost, and
AgentTermPreview's own floating badges are the only place swarm-ui
shows them. Left the badges visible there instead of shipping a
regression the plan's own stated justification didn't hold up to.

Also fixed a same-tab pub/sub race found by actually rendering a cold
load of /agents/:name/terminal (headless chromium, not just reasoning
about the code): useLocalSetting subscribes inside a useEffect, and
mount effects fire children-before-parents, so a descendant's
mount-time write (AgentTerminalPage registering its own tab) can beat
an ancestor's (Shell's) subscription into existence, leaving Shell's
tab row silently empty on a direct/reload load. Fixed by having
useDynamicTabs re-sync from storage on every location change, not
just on notify() — the fix lives in the new hook itself, not in the
shared settings-storage primitive theme/motion overrides also use.
2026-09-28 23:17:44 +02:00
atlas
433429ebfd bao: disable the unused approle auth method, declaratively
No Rust ever minted a secret_id; approle was dead attack surface. The
bootstrap step's check-then-enable case becomes check-then-disable: if
approle is mounted, `bao auth disable approle`; otherwise a no-op.

Disabling costs `delete`+`sudo` on `sys/auth/approle`, not
`create`/`update` — verified against `bao auth disable -output-policy`
on a live dev store. The bootstrap policy grant is narrowed to match.

nix/module-eval/bao-grants.nix pins the new shape: the bootstrap policy
may disable approle, and the granter's role unit never enables it.
2026-09-28 22:58:36 +02:00
atlas
7bc4b25f16 swarm-controller: backfill a missing queue-secret mint time instead of re-minting it
A stored queue secret with no `minted_at` counted as due, and no secret
minted before the renewal pass has one, so the first pass after deploy
would re-mint every agent's secret. Every reconnect before that agent's
next restart would then be refused.

Such a secret is now stamped instead: `minted_at = now` is written beside
the unchanged `value`, and its 45-day clock starts there. Only a secret
whose recorded mint time is at least 45 days old gets a new value.

The decision is `secret_step` (Keep / Backfill / Remint), and the pass
reports an unstamped secret as `Observed::Unstamped`. Backfill and
re-mint log different lines.
2026-09-28 22:24:07 +02:00
atlas
2115ec2bb3 swarm-controller: re-issue agent certificates and re-mint queue secrets at half-life
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:

- `MintAgentIdentity` (the node agent creation uses) when the stored
  certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
  read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
  swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
  re-decides, writes a fresh value with `minted_at`, reads it back, and logs
  the agent and the old age.

When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.

`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.

Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.

Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.

docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
2026-09-28 21:35:01 +02:00
atlas
b14ff2796c bao: OIDC login to the browser UI via authelia, as a metadata-only viewer
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.

- authelia registers an interactive client `swarm-bao-ui`
  (glue-bao-ui-oidc-client.nix) with redirect
  `https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
  publisher carries its secret to
  `secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
  mount with listing visibility `unauth`, asked before attempted like
  cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
  and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
  policy, and once the granter may configure `auth/oidc/config` (checked
  through `sys/capabilities-self`), writes the mount's config from the
  published secret and the role binding `groups=admins` to the viewer.
  Before the bootstrap step re-runs it writes the policy, logs the step
  and exits 0.

Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.

module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
2026-09-28 19:56:38 +02:00
atlas
3b27a2e5a1 docs: fix vale prose-lint errors in bao UI docs
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Passive voice and sentence-initial 'So' hits from PR #4776's vale error job.
2026-09-28 19:31:02 +02:00
atlas
3898ca33c7 bao: serve the browser UI to admins via a loopback-only listener
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.

The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.

Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.

Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.

On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
2026-09-28 19:31:02 +02:00
atlas
a86379eac5 ci: internal jobs skip on the public forge instead of waiting for a hive-ci runner
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
A job-level if: is evaluated by whichever runner picks the job up. The
public forge's copy has only a nixos runner, so a job guarded by
if: vars.PUBLIC_FORGE != 'true' and pinned to runs-on: [hive-ci] never
gets picked up there to evaluate the guard at all - it just sits queued.

runs-on now carries the same PUBLIC_FORGE variable: hive-ci internally,
nixos on the public copy. There the nixos runner picks the job up,
evaluates the existing if:, and skips it immediately. Internally
PUBLIC_FORGE is unset, so runs-on still resolves to hive-ci and nothing
changes.
2026-09-28 19:28:23 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
97c4771514 ci: run nothing but a main-only bin-cache push on the public forge
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Guards internal-only jobs (ci.yml, coverage.yml, flake-update.yml) with
`vars.PUBLIC_CACHE != 'true'` and adds public-cache.yml, guarded to the
inverse, to build the deployed closures and push them to the public
attic cache on a push to main.

`PUBLIC_CACHE` is a repo Actions variable, opt-in only on the public
copy: unset here, it leaves internal CI's `!=` guards true so internal
jobs always run. Job-level `if:` cannot see the `github`/`forgejo`
context at all on this runner (confirmed empirically — a
`github.server_url` comparison always evaluates false at job level,
though the identical comparison resolves correctly inside a step),
so `vars.*`, which is available at job level, is the only usable
opt-in signal here.
2026-09-28 17:24:06 +02:00
iris
9a444c27fa agent ui: add a clickable new-session menu entry
Same gap #4611 fixed for logout: the Preact rewrite's StatusChips menu
never picked up new-session as a click path, only /new-session typed
twice into the terminal. postNewSession already existed in
termActions.ts. Wire it into the status menu with the same arm-then-
confirm click pattern logout/cancel-turn already use.

Closes #4612
2026-09-28 13:53:58 +02:00
atlas
3b8df27dcf swarm-ui: show each agent's icon on its card
Each agent card now leads with the agent's icon, loaded as an `<img>`
from `GET /api/agents/<name>/icon`: the same 5em square, background and
fallback as the hive dashboard's container row. An agent with no icon
(the route's 404), or any other failed load, shows the dimmed hyperhive
mark (`/favicon.svg`) instead of a broken image.

Only ever an `<img>`, never inline markup: the body is an agent-authored
SVG, and an image load does not run its script.
2026-09-28 13:47:37 +02:00
atlas
e974194e3a swarm: let an agent publish its own icon
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.

hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.

swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.

Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
2026-09-28 13:47:37 +02:00
atlas
4dbab024da swarm: drop the tracker tag from agent_icon's module doc
`scripts/check-issue-refs.sh` blocks a `#N`-shaped reference in source.
The quoted ruling reads identically without it, so the tag goes rather
than a lint suppression.
2026-09-28 13:47:37 +02:00
atlas
9513058a71 swarm: serve an agent's icon at swarm scope
An agent is not fixed to a hive, so its icon cannot be resolved as
hive -> agent. This adds the swarm-level half: an `agent-icons` KV
bucket keyed by the agent name alone — no hive token, so an agent that
moves hives keeps its icon and one that is stopped still has one — and
`GET /api/agents/<name>/icon` on swarm-controller serving it
same-origin, like every other `/api/*` route swarm-ui calls.

404 is the "this agent has no icon" answer, the same contract the
per-agent harness's own `GET /icon` has for an unconfigured agent.
Until the agent-side publisher lands, that is every agent's answer:
the publisher runs inside the container and an agent's NATS grants are
hive-scoped, which cannot authorise a write to a single-token agent
key. The read side needs no grant change — the controller already
holds `$KV.*.>` and `$JS.API.DIRECT.GET.*.>`.

The response carries `Content-Security-Policy: sandbox` and `nosniff`:
the body is an operator-authored SVG served from this daemon's own
origin, and an SVG can carry script.

Hive-side icon serving is untouched.

Refs #4502
2026-09-28 13:47:37 +02:00
iris
2e67656de8 docs: drop leftover 'nothing nests' framing from C0NTAINERS blurb
Per mara's review on #4769: don't extend stale info to say it's stale,
just remove it. The sentence only described the absence of the removed
topology tree, giving the operator nothing actionable.
2026-09-28 13:18:12 +02:00
iris
d9beb965a5 docs/web-ui: drop the redundant Topology tree section
C0NTAINERS already states the flat/alphabetical/no-nesting fact;
the section repeated it with zero new operator info. Fix the two
dangling references (README.md reading-path index, swarm.js
comment) rather than leaving them pointing at a removed heading.
2026-09-28 13:18:12 +02:00
iris
ef4e49805c docs/web-ui/dashboard.md: drop stale topology-tree history per mara review
state current fact only, not the old-vs-new narrative
2026-09-28 13:18:12 +02:00
iris
d447974536 docs/web-ui/dashboard.md: fix vale hits in topology-tree rewrite
- avoid 'backend' (Microsoft.Avoid)
- reword 'was removed' passive voice (write-good.Passive)
2026-09-28 13:18:12 +02:00
iris
6b5c8cc51a dashboard: flatten the agent list, drop dead topology-tree machinery (#4638)
The backend dropped the agent hierarchy's parent field, so every
container is a root and buildAgentTree/treePrefixDom could only ever
produce a single-level flat list — the .tree-prefix CSS lane rules
already matched nothing. Replaced with sortedContainerRows, a plain
alphabetical sort, and dropped the now-dead .tree-prefix/.tree-lane/
data-depth CSS and the depth/isLast/ancestorIsLast fields from the row
fingerprint and buildContainerLi's signature. No visible behavior
change - the rendered list was already flat, just via dead machinery.

Docs updated to describe the simpler implementation directly instead
of narrating the removal (kept the heading name since swarm.js still
points a comment at it).
2026-09-28 13:18:12 +02:00
iris
c9e12ffcc0 swarm-grafana: revert busiest-agents table, restore bargauges (#4658)
c9ba9062 replaced 5 bargauge panels with one table panel using a
joinByField+organize+sortBy transformations chain and a
vizConfig.group="table" viz spec. That construct had zero precedent
anywhere else in this repo's dashboards and its live render was
explicitly flagged as unverified at merge time (no in-container way
to render Grafana to confirm).

#4658 reports a dashboard crash: "'' not found in: reduce,
filterFieldsByName, ... transpose" - a Grafana transform-registry
lookup failing on an empty transformer id. That registry's full id
list matches the error text exactly, and the table panel is the only
panel in the whole swarm-grafana/dashboards tree using a non-empty
transformations array, so it's the leading suspect: some part of how
the dashboard-v2beta1 schema expects a QueryGroup's transformations
to be shaped likely differs from the classic {id, options} shape used
here, and the file-based provisioner may not be normalizing the
mismatch the same way a schema-v2-aware loader does.

Reverting to the prior, previously-uneventful bargauge panels
(byte-for-byte c9ba9062^'s version of this section) to stop the
active crash while the correct v2beta1 transformation shape gets
verified against a live Grafana render instead of guessed at again.
2026-09-28 11:52:33 +02:00
atlas
a2b4acfb6a swarm-bao: say "run the bootstrap step" when the bootstrap token is dead, not "sealed"
swarm-bao-granter-role used its token for `bao auth list` with no check,
so an expired, revoked or policy-less token in bootstrapTokenFile exited
2 with a raw 403 and no hint. It now checks whether the store is up when
that call fails: if it is, the token is at fault, and the unit prints the
one-time bootstrap step and exits 4. A missing token file is still a
ConditionPathExists skip, so the two read differently in the journal.

The twelve granterLogin units printed "the store is sealed or
unreachable" on a healthy store, because `bao status` exited 1 there:
the CLI resolves a token helper under $HOME before asking, systemd sets
no HOME for a unit without User=, and the fallback shells out to
`getent`/`sh`, neither of which is on the unit's PATH ("failed to get
token helper: error expanding config path "": exec: "sh": executable
file not found in $PATH"). The check now runs with HOME=/var/empty and
keeps its stderr, so a genuinely unreachable store says why. When the
store is up and the login is refused, the units now name
swarm-bao-granter-role as the unit that writes the missing role.

setup.md's post-step restart used 'swarm-bao-*-policy.service', which
misses swarm-bao-agent-pki. It now names that unit too, and a
module-eval case fails when the restart misses any unit that logs in as
the granter.

Refs #4704
2026-09-28 11:03:08 +02:00
atlas
3fc7a785fa swarm-controller: default matrixHomeserverUrl from swarm.matrix.gatewayHost
matrixHomeserverUrl re-derived gatewayHost's own default
(chat.${swarmDomain}) instead of reading gatewayHost itself, so a
deployment that pins gatewayHost away from that default (the exact
case hive-matrix.nix's own option doc describes) left the controller
calling a name nothing serves. Read matrix.gatewayHost directly,
mirroring hive-matrix.nix's ctlHomeserverUrl. Add module-eval cases
that pin gatewayHost and that null it out, alongside the existing
unpinned default case.

Closes #4757
2026-09-28 10:16:50 +02:00
atlas
9cbee27e6d docs: reword vale-flagged prose in docs/swarm/README.md 2026-09-28 08:24:52 +02:00
atlas
727065c960 hive-agent: present this agent's own queue credential, then fall back
When the per-agent secret `queue-identity.nix` fetched is present, the
harness connects with `swarm-agent.<agent>.<secret>` as a static token
and publishes on `$SWARM.term.<agent>` and `$SWARM.agent-state.<agent>`.
When it is absent, or that first connect fails for any reason, a refusal
from a responder that does not verify agent tokens included, it connects
with the hive's shared OIDC client and publishes on the hive-scoped
subjects as before. Which one it took is logged once per connect.

`swarm_queue_client::connect_with_token` is the static-token connect: no
retry on the initial attempt, so the caller sees the refusal and can
fall back. Reconnects share the existing backoff, now a named function.

Closes #4630
2026-09-28 08:24:52 +02:00
atlas
0c1fb44a4f swarm-controller: relay an agent's hive-free subjects beside the old ones
The terminal and turn-state relays now subscribe to `$SWARM.term.<agent>`
and `$SWARM.agent-state.<agent>`, the subjects a verified agent token is
granted, as well as the hive-scoped `<prefix>.<hive>.<agent>` an agent
still on its hive's shared credential publishes to. Both at once, so a
swarm-ui terminal keeps working whichever credential an agent connected
with and in whatever order hosts deploy.
2026-09-28 08:24:52 +02:00
atlas
9bad58d86d swarm-nats-auth: verify an agent's own token against the store
An `auth_token` spelled `swarm-agent.<agent>.<secret>` is no longer sent
to introspection. The responder reads `swarm/agents/<agent>/queue` with
an identity of its own, checks that the stored object names the same
agent, compares the secret in constant time, and grants the subjects
`--agent-token-publish-subject` lists with `{agent}` expanded. Every
other outcome denies: a malformed token, no store identity, nothing
stored, a failed or slow lookup, a different secret. A token without
the prefix takes the OIDC path unchanged.

The journal's `auth request` line names such a caller `agent:<agent>`;
the hive-shared credential keeps `hive-<h>-agent`.

The new principal: a `swarm-nats-auth` cert-auth role and policy with
read on `secret/data/swarm/agents/+/queue` alone, a leaf signed by the
store's PKI glue, and `glue-nats-auth-bao-identity.nix` pairing the two.
The copy unit delivers the identity into the queue's container, and an
absent leaf is delivered empty so the responder still starts and only
agent tokens are refused.

The policy and role are written by `swarm-bao-nats-auth-policy`, logged in
as the bao granter: both names fall under its `swarm-*` globs, so the
deploy writes them with no operator step. module-eval counts it among the
granting units, so every generic granting-unit case covers it.

The secret compare uses `subtle`, already in the lock file through the
TLS stack; no workspace crate offered one directly.
2026-09-28 08:24:52 +02:00
atlas
fc97c237dc swarm-queue-client: one agent-token spelling, and no hive in AgentCredential
`swarm_queue_client::agent_token::format_agent_token` / `parse_agent_token`
are the spelling an agent presents its own queue secret in,
`swarm-agent.<agent>.<secret>`, and the one the auth-callout responder
reads back. The prefix is what separates it from an OIDC access token,
which may itself contain `.`. Parsing distinguishes "not an agent token"
(no prefix) from "a malformed one"; the error names the problem and never
the value. The module is store-free, so the agent formats its token
without linking the secret-store client.

`swarm_secret_client::queue::AgentCredential` loses `hive`: an agent's
identity is not tied to a hive, and nothing reads the field. Objects
already in the store carry it and still decode, since unknown fields are
ignored; a test parses one. The controller stops writing it.

With the credential no longer naming a hive, and the agent's policy
naming none since #4762, nothing in the mint consumes one. `hive` goes
from `mint_and_verify`, from the `MintAgentIdentity` node, and from
`POST /api/agents/{name}/identity`, which now takes no body and no longer
checks a hive against the roster; a caller that still sends one is not
refused, the body is ignored. `swarmctl agent mint-identity` loses
`--hive`, so passing it is now a usage error.
2026-09-28 08:24:52 +02:00
atlas
6170e74a31 swarm-bao: agent certificates issued by a store-generated agent CA
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.

- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
  tunes the mount, generates the root once (guarded on an empty issuer
  list, no replace branch), upserts the `swarm-agent` role (client
  certificates named `hive-agent-*` only, 90 days), caches the CA at
  /var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
  (client-ca.pem, then the agent CA). Host cert-auth roles still pin
  client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
  composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
  changed after openbao started, swarm-bao-agent-pki restarts
  openbao.service in the container; under `seal = "shamir"` it prints
  the step instead. Once swarm-bao-certs has a cached CA, later boots
  start openbao with it and do not restart.
- The controller policy gains exactly `update` on
  pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
  the leaf (the store generates the key), writes the agent's cert-auth
  role pinning the issuing CA bao returned, and writes the agent's
  policy as render_agent alone: the hive-shared queue credential stanza
  is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
  other pki role names); swarm-controller gets
  SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.

Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.

Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.

Closes #4756
2026-09-27 22:59:27 +02:00
atlas
5cd7f866f4 swarm-bao: grant the agent PKI mount (for #4756)
#4756 moves agent client certificates onto a PKI mount of their own,
`pki-agents`, whose root bao generates internally. The unit that sets
that mount up runs as the bao granter, and the granter's policy is only
written while #4754's one-time bootstrap token is in place. Adding these
grants after an operator has done that step would cost a second token
placement, so they go into the granter's policy here, before it.

Six stanzas: enable and tune the mount, list its issuers, read its CA,
generate its root internally, and write `swarm-*` roles on it. No root
delete or sudo: agent cert-auth roles pin that root by value, so
replacing it must not be something a deploy can do.

Adds deploy.bao.agentPkiMountPath (default `pki-agents`), which the
stanzas are rendered from. The module-eval case pinning the granter's
policy now lists all seventeen stanzas, and the "grants nothing outside"
case also refuses the agent mount's root, issue, sign and a roles/*
glob.
2026-09-27 22:57:46 +02:00
atlas
e9cec0da21 swarm-bao: write every swarm-* grant as a bao granter, not with a 24h token
Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).

A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.

swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.

The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.

module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.

Refs #4704
2026-09-27 22:57:46 +02:00
atlas
19cc1b12e2 hive-screen-mcp: bound grim/wtype and VNC calls; hive-c0re: make messages match the code
hive-screen-mcp ran `grim` / `wtype` through an unbounded
`Command::output()` and spoke RFB to neatvnc with no deadline, so a
wedged compositor or a VNC server that accepts and never speaks held the
agent's turn forever. Each subprocess now has a 30s limit and is killed
when it hits it; each RFB exchange has a 10s limit. Both come back to the
agent as the tool's text result, like every other failure in this crate.

hive-c0re operator-facing text that described behaviour the code lacks:
- `hivectl matrix reset-password` printed a "next: hivectl matrix
  create-user" hint that fails for every target (agents are refused, a
  non-agent hits M_USER_IN_USE). The line is gone.
- a failed `nixos-container update` appended the container's journal
  tail, read with `journalctl -M`. Since the job DAG, `update` only runs
  from the `Swap` node on a stopped container, so the read always came
  back empty. The helper is removed; the error still points at the
  build log.
- the matrix sweep comment in main.rs said it re-provisions agent token
  files; `ensure_all` creates no agent accounts or tokens.
- the knowledge-pull comments named a webhook caller that no longer
  exists and claimed a race was "fixed at its source".
- `handle_spawn`'s doc and the `mcp_sockets` module doc named callers of
  `register_agent` / a rollback that do not exist.

Refs #4723
2026-09-27 20:47:06 +02:00
atlas
3385aaf026 docs: drop "simply" from the knowledge sweep page 2026-09-27 20:15:55 +02:00
atlas
313d582c01 job_queue: drop insert_unless_live, accept extra queue passes
mara chose to accept extra queued sweep passes over adding a new
hive-jobq primitive (or a hive-c0re one-off) for "don't queue another
of this kind". Every sweep caller now plain-inserts its node; the
capacity-1 Dep::Resource per sweep kind (MatrixSweep/KnowledgeTree)
still keeps two passes of the same kind from running concurrently, it
just no longer collapses a tick that lands while one is live or queued
into the existing one.
2026-09-27 20:15:55 +02:00
atlas
1d4c77d2c8 hive-c0re: serialise the matrix and knowledge sweeps through the job queue
The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
2026-09-27 20:15:55 +02:00
atlas
2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00
flake-bot
e7456a49ff nix flake update 2026-09-27 14:33:24 +02:00
atlas
7ac6819652 hive-c0re: fail on a malformed agent name and on an unreadable container list
An agent name that is not a valid Ident made `Coordinator::agent_paths`
panic. Job payloads carry names as plain strings (the swarm's published
wanted state is one source), and a panic inside a job-queue node never
reaches `complete_growing`, so the node's resources (the deploy
window included) were held until hive-c0re restarted. `agent_paths` now
returns an error; the job-queue nodes, the admin-socket spawn and
set-limits paths, the root-agent spawn and the dashboard set-limits
handler propagate it.

`lifecycle::list().await.unwrap_or_default()` turned a failed container
list into "no agents":
- meta-update cascade: the lock bump committed and zero rebuilds fanned
  out, reported as success. The cascade is now resolved before the lock
  bump and a list failure fails the node.
- dashboard update-all: queued nothing and returned 200 "ok". Now 500
  with the error.
- container rescan: every row was emitted as removed and the cache
  emptied. Now the last snapshot stands; `hivectl status` gets an error.
- dashboard journal: answered 404 "no managed container". Now 500.
- spawn/rebuild port-collision check: silently skipped. Now fails.
- startup migration: the per-agent phases ran over nothing, and phase 3
  handed an empty agent list to `meta::sync_agents`, which renders the
  meta flake with exactly the agents it is given. Both now log the list
  failure and skip.

The hive-jobq scheduler still leaks a node's resources on any executor
panic; that root is not addressed here.

Refs #4723
2026-09-27 05:13:21 +02:00
atlas
ee25b7de20 hive-forge-notify: a failed assigned-count poll leaves the rollup todo unchanged
update_assigned_rollup treated a failed count_assigned as zero and issued
ClearTodo, deleting an acked rollup todo on every transient forge error
(and, on a one-sided failure, corrupting the breakdown for one poll).
count_assigned now returns Result<u64, String> instead of Option<u64> so
the caller can log the cause, and decide_rollup only acts once both
counts are known — either failing leaves the existing todo state as-is
and logs a warn! naming which poll failed.

Closes #4720
2026-09-27 05:12:39 +02:00
atlas
fc85d0ee7e hive-priv: publish config files by rename, bound the toplevel build
Every file hive-priv writes for another reader now goes through one
StagedFile: a temp with a unique dot-prefixed `.partial` name in the
destination directory, created O_EXCL|O_NOFOLLOW at its final mode
(and owner, where one is set), fsynced, renamed onto the final name
relative to a directory fd, then the directory fsynced. Dropping it
unpublished unlinks the temp.

- write_resource_limits wrote its systemd drop-in in place, so a
  concurrent daemon-reload could load a truncated file.
- sync_agent_tmpfiles staged through a fixed `.tmp` name, so two
  overlapping syncs shared one inode and one could publish the other's
  bytes, or a mix.
- write_nspawn_flags rewrote /etc/nixos-containers/<c>.conf in place;
  it now keeps the file's existing owner and mode.
- write_agent_dir_file (agent tokens, sidecars, pause marker) already
  renamed, but through a fixed temp opened O_TRUNC, with no fsync.
- write_bridge_dns_marker_in truncated and wrote the marker in place;
  a leaf the container planted (symlink, FIFO) is now replaced by the
  rename instead of refused.
- register_ci_runner must keep writing in place (nspawn pins the
  bind-mounted inode); it now creates the file 0600 when the tmpfiles
  seed is missing, instead of at the umask's mode until the chmod.

nix_build_toplevel now runs nix as its own process-group leader and,
after one hour, SIGKILLs the group and fails with "timed out after
3600s". One hour is four times CI's observed cold-cache flake check.

Refs #4723
2026-09-27 05:03:28 +02:00
atlas
bfd8189900 hive-c0re: stop reporting refused invites as success; re-register the config-PR hook when its secret changes
invite_user_id mapped every 403 M_FORBIDDEN to Ok(()). The membership
pre-check already skips invited/joined users, so the 403s that reach the
POST are mostly real refusals (banned target, sender without power),
including `hivectl matrix invite`. A 403 is now success only when a
membership re-read shows the user invited or joined; otherwise it is an
error carrying the status and body.

admin_room_send_and_poll read the send response's event_id with
unwrap_or_default() and, when it was missing, walked every recent event
unanchored, so an older bot reply (an earlier reset password) could be
returned as this command's result. A send response without an event_id
is now an error.

run_destroy_bookkeeping discarded fail_pending_for_agent's error; it now
warns like its neighbouring steps.

ensure_config_pr_webhook returned as soon as a hook with the target URL
existed, so a regenerated webhook-secret never reached Forgejo and every
config-PR delivery failed HMAC until the 5-minute poll caught up.
Forgejo's edit-hook API ignores `secret` and never returns it, so the
SHA-256 of the secret last registered is recorded at
forge/config-pr-webhook-secret-sha256; when it doesn't match, the
same-URL hook is deleted and recreated. The paths.rs doc claiming
re-registration on change now describes this.

Refs #4723
2026-09-27 04:41:31 +02:00
atlas
0f58cdbde2 swarm-queue-client: install aws-lc-rs as the process rustls provider
rustls is built with both `ring` (async-nats's `ring` feature) and
`aws-lc-rs` (reqwest's `rustls` feature), so it cannot pick a
process-level default by itself. Since the queue started requiring TLS
(1d261b3f), async-nats builds its config with `ClientConfig::builder()`,
which panics without an installed default. The panic kills the async-nats
connector task, and every queue client (swarm-controller, hive-c0re, all
hive-agents) has sat in `Pending` since the 2026-09-25 23:04Z deploy.

Add `swarm_queue_client::install_crypto_provider()`, which installs
aws-lc-rs and ignores the "already installed" error. It is called first in
`main` of every binary that links async-nats: hive-agent, hive-c0re,
swarm-controller, swarm-nats-auth. `connect()` also calls it, so a new
binary that dials through this crate is covered without remembering to.

aws-lc-rs because reqwest already falls back to it when no default is
installed, so HTTPS in these processes keeps its current provider. The
other rustls users in the tree reach it only through reqwest, which never
panics here.

Closes #4738
2026-09-27 04:17:24 +02:00
atlas
386741d38f hive-c0re: keep agent and manager listeners alive across accept errors
Both accept loops returned on the first accept() error. The AgentSocket
stayed in Coordinator.agents, so mcp_sockets::sync_on_start skipped the
agent, and only a Start node or a hive-c0re restart bound it again. The
unit sets no LimitNOFILE, so a single EMFILE at the 1024-fd soft limit
cut the agent off from send/recv/ack with one warn line as the only
trace.

Both loops now share accept_until_fatal. Connection errors
(ECONNABORTED/ECONNRESET/ECONNREFUSED) retry at once, any other error
retries after 1 s, following axum's serve loop that the dashboard
already runs. Errors that leave the listening fd unusable (EBADF,
EFAULT, EINVAL, ENOTSOCK, EOPNOTSUPP) log at error and exit the
process: nothing re-binds a listener while hive-c0re runs, the manager
listener has no re-bind path at all, and the unit's Restart=on-failure
restart runs start_manager and sync_on_start, which re-bind every
socket.

mcp_sockets.rs now states that invariant instead of asserting that a
listener can only disappear on restart.

LimitNOFILE is left unset: per-agent fd use is a listener plus a
Recv long-poll plus short-lived requests, and a higher limit would raise
the memory ceiling of the per-line bound (bound x open connections).

Closes #4721
2026-09-27 03:48:56 +02:00
atlas
65a1f8d902 hive-c0re: bound request lines on the agent and manager sockets
serve() read each request with read_line into an unbounded String, so
size checks such as the 4 KiB Send body limit ran only after the whole
line was buffered. Any process in an agent container can write to
/run/hive/mcp.sock, and one stream with no newline grew hive-c0re's heap
until the host OOM-killed it.

Read at most 16 MiB per line. The largest request a client sends is an
OperatorMsg from the agent web UI's /send, whose body axum's default
limit caps at 2 MiB (the gateway's nginx admits 10 MiB); JSON escaping
can roughly double that, and the cap is 4x the result. A longer line
gets a Response::Err, the same shape as the parse-error path, and the
connection is closed because the rest of the line is still unread.

The line is read as bytes and parsed with serde_json::from_slice, so a
line with invalid UTF-8 now gets a parse-error response instead of the
connection being dropped.

Closes #4718
2026-09-27 03:48:56 +02:00
atlas
7b1fe5f9d3 hive-sock-client, web proxy, HTTP clients: bound connect and response waits
hive-sock-client: each attempt now bounds connect (5s), write (10s) and
the wait for the response (60s by default). The response bound is per
call through the new `request_within`, which hive-agent's serve-loop
`Recv` uses with its 180s long-poll plus 30s headroom. A response
timeout is terminal rather than retried: the server holds the request,
so a retry re-sends something it may still act on and multiplies the
wait by the backoff schedule.

Outbound HTTP: the matrix login/whoami clients in swarm-controller and
hive-c0re's dashboard (5s connect, 30s request), the authelia-bridge
client (5s/30s; ensuring an identity runs an argon2 hash first) and the
ci-runner forge calls (5s/15s, config_pr_poll's forge budget) get a
connect_timeout and a request timeout. Timeout errors name the bound
that fired.

hive-agent's unix-socket extra web proxy bounds the connect (5s) and
the wait for the response head (30s, the http sibling's budget); the
body read stays unbounded.

Refs #4723
2026-09-27 03:46:53 +02:00
atlas
6fac00dcc5 hive-c0re: fail on an unparseable resource-limits or topology file, write both atomically
resource-limits.json and topology.json were read with parse errors
folded into an empty map, and written in place with std::fs::write. One
truncated resource-limits.json followed by a single set_limits call
rewrote the file with only that agent's entry, erasing every other
agent's CPU and memory overrides without a log line. topology.json had
the same shape: reconcile rebuilt it from the live set, losing pending
(provisioned, never spawned) names.

- agent_config::read_map / write_map are generic over the stored type.
  tool-groups and capabilities behave as before.
- resource_limits::read / effective return an error for an existing but
  unreadable file; a missing file is still the empty map. set_limits
  fails without writing on such a file, and writes atomically.
- topology: reconcile fails without writing on an unreadable file and
  writes atomically. all_agents logs the error and returns no agents,
  so a ManageRootAgent holder starts without cross-agent mounts.

Read-path behaviour on an unreadable resource-limits.json, per caller:
- write_dropins (every spawn / swap / WriteDropin): logs the error and
  keeps the limits drop-in already under /run; the agent still starts.
  With no drop-in yet (first start since boot) it writes the hive
  defaults, because no drop-in means an uncapped container.
- render_flake: propagates, so sync_agents (and spawn/rebuild/destroy
  jobs) fail. An empty map would give tighter-capped agents the hive
  memoryMaxBytes.
- container_view::build_all: logs the error each scan and renders the
  rows at the hive defaults (no ContainerView wire change).
- set_resource_limits reply: propagates.

Closes #4731
2026-09-27 02:41:38 +02:00
atlas
97cf8a1b2b agent-modules: write bao's stderr where UMask=0377 lets it, name what failed
hive-agent-forge-token and hive-agent-queue-credential both run with
UMask=0377. Their scripts captured bao's stderr in `err="$(mktemp)"`,
which under that umask is created 0400; the very next `2>"$err"` on the
`bao login` line cannot reopen it for writing, so bash fails the
redirect with "Permission denied" before bao ever runs. The `if !`
around the login then took the only error branch it had and printed
"this agent's certificate was refused by the swarm secret store" — the
store was never contacted. No agent has fetched either credential.

The stderr file now lives in each unit's own 0700 RuntimeDirectory and
is removed before every redirect into it, so the redirect creates it —
the idiom forge-token.nix already used for its staging file.

The login's error branch now says which of these happened, then quotes
bao's output:
- `$err` could not be created, so bao never ran;
- the store answered with HTTP 4xx (refusal) or another status;
- the store sent a TLS alert rejecting the certificate;
- no answer at all (network, DNS, or local TLS).
Unreadable cert/key credentials are reported before bao runs.

bao.nix has the same fetch shape but no UMask=, so its mktemp file is
0600 and writable; it is untouched.

Closes #4735
2026-09-26 21:50:39 +02:00
atlas
c6a778f666 lint: tighten atomic-write-secret.nix's header comment; fix vale contractions in persistence.md
The header comment grew to 39 lines across two audit-driven rounds,
over the 30-line comment-block-lint max. Trimmed to 30: merged the
value-as-argument rationale with its /proc/cmdline justification into
one paragraph, cut the usage example to one call instead of two, and
condensed the args paragraph — no content dropped, just restatement.

persistence.md's first-boot-migration marker paragraph used "it is"
and "there is not" — vale's Microsoft.Contractions rule (this repo's
config) wants the contracted forms, and "is not" also collides with
"is nothing" as a literal substring, which is what actually tripped
the error. Reworded to "it's" / "there's nothing", no meaning change.

Lint-only: no script logic changed, gates re-run below are all lint
checks (no module-eval, no cargo).

Refs #4723
2026-09-26 21:50:03 +02:00