Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/docs
Author SHA1 Message Date
atlas
2115ec2bb3 swarm-controller: re-issue agent certificates and re-mint queue secrets at half-life
A five-minute pass over every agent some hive's wanted state declares as
anything but destroyed queues, per agent:

- `MintAgentIdentity` (the node agent creation uses) when the stored
  certificate at swarm/agents/<agent>/bao-mtls is past half its validity,
  read from its own notBefore/notAfter: day 45 of the role's 90;
- the new `RenewAgentQueueCredential` node when the queue secret at
  swarm/agents/<agent>/queue is 45 days old or has no mint time. The node
  re-decides, writes a fresh value with `minted_at`, reads it back, and logs
  the agent and the old age.

When both are due the secret node runs after_any the certificate node,
because mint_and_verify compares the queue secret it read with the one it
reads back. A credential that is not stored is never created here.

`queue::AgentCredential` gains an optional `minted_at` (unix seconds);
agent creation now sets it. Stored objects without it decode unchanged and
count as due, so every existing queue secret is re-minted on the first pass.

Both replacements reach the agent at its next start. The old certificate
stays valid until it expires; the old queue secret does not, so a queue
reconnect before that restart is denied.

Adds x509-cert 0.2 (with der_derive and flagset) to read the validity.

docs/swarm/credentials.md: the renewal column splits into automatic re-mint
and automatic re-pull, filled from the code as it stands.
2026-09-28 21:35:01 +02:00
atlas
b14ff2796c bao: OIDC login to the browser UI via authelia, as a metadata-only viewer
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The bao UI at bao-ui.<swarm> took a raw store token and nothing else.
It now offers an OIDC tab: authelia's `admins` group logs in and lands
on `swarm-operator-viewer`, which is list+read on `secret/metadata/*`
and nothing under `secret/data/` or `sys/`.

- authelia registers an interactive client `swarm-bao-ui`
  (glue-bao-ui-oidc-client.nix) with redirect
  `https://bao-ui.<swarm>/ui/vault/auth/oidc/oidc/callback`; the secret
  publisher carries its secret to
  `secret/swarm/services/swarm-bao-ui/oidc/client`.
- `swarm-bao-granter-role` (bootstrap token) enables the `oidc` auth
  mount with listing visibility `unauth`, asked before attempted like
  cert/approle; `bao-bootstrap-policy.hcl` gains `sys/auth/oidc`.
- The granter's policy gains `auth/oidc/config`, `auth/oidc/role/swarm-*`
  and read on that one secret leaf. It still holds no `sys/auth`.
- New granting unit `swarm-bao-operator-viewer-policy` writes the viewer
  policy, and once the granter may configure `auth/oidc/config` (checked
  through `sys/capabilities-self`), writes the mount's config from the
  published secret and the role binding `groups=admins` to the viewer.
  Before the bootstrap step re-runs it writes the policy, logs the step
  and exits 0.

Route (a) per mara on #4775: enabling the auth method stays a
bootstrap-token step, re-run once on the live store.

module-eval pins the viewer policy's single metadata stanza, that the
granter's policy has no sys/auth path, the oidc enable in the bootstrap
unit, the exit-0 path, the config/role contents, and the client
registration + publish.
2026-09-28 19:56:38 +02:00
atlas
3b27a2e5a1 docs: fix vale prose-lint errors in bao UI docs
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
Passive voice and sentence-initial 'So' hits from PR #4776's vale error job.
2026-09-28 19:31:02 +02:00
atlas
3898ca33c7 bao: serve the browser UI to admins via a loopback-only listener
openbao gains a second listener, `ui`, on 127.0.0.1:<deploy.bao.uiPort>
(default 8204) with TLS off and no client-certificate requirement, and
`ui = true`. The existing listeners are unchanged. An nginx inside the
store's container, on 127.0.0.1:<deploy.bao.uiProxyPort> (default 8206),
forwards only /ui/ and /v1/ to it, redirects / to /ui/, answers 403 on
sys/unseal, sys/seal, sys/step-down, sys/rekey* and sys/generate-root*,
and 404 on everything else.

The gateway on the store's host serves `swarm.bao.ui.domain` (default
bao-ui.<swarm>) behind the authelia auth_request subrequest, proxying to
that nginx; the name joins serviceDomains and localNames like every
other gateway-published swarm service. authelia gets an access_control
rule restricting that name to group:admins, rendered wherever authelia
runs, since the default policy admits any session.

Trade-off, ruled by the operator on the parent issue: the UI listener
asks for no client certificate, so on that door a bao token alone is the
credential.

Three comments and a doc line claimed every API listener demands a
client certificate; they now except the loopback UI listener. The
module-eval case counting declared listeners excludes `ui` by name, as
it already did `metrics`.

On a self-signed gateway, the UI's name is a swarm service name, so its
host requests the services leaf from the store. `swarm-services-cert`
sits Before= and RequiredBy= the gateway's cert import, which nginx
Requires=. On a host whose only swarm name is the UI, that would hold
nginx, and with it the stream passthrough every reader dials, on a login
to a store that may be sealed. hive-tls drops those two edges exactly
when the UI is the only local swarm name: nginx starts on the existing
hive-leaf fallback, and the script's existing re-import reloads nginx
once the leaf issues. Every other host keeps both edges.
2026-09-28 19:31:02 +02:00
atlas
e94406cdb9 swarm-controller: read the queue client secret from the store, drop the file
Some checks were skipped
public bin cache / build + push to preem:grid (push) Has been skipped
The controller's OIDC client secret (client `swarm-controller`, used for
the queue connection, the auth-bridge bearer and the OTLP push) came from
an operator-placed file, `deploy.swarm-controller.queue.clientSecretFile`,
handed in by `LoadCredential=`.

Now `swarm-secret-publish`, which already copies authelia's minted OIDC
secrets into the store, also publishes this one, to
`swarm/controller/swarm-controller/oidc/client`. That path sits under
`controller/`, which no hive's policy reads. The controller reads it once
at start with its existing store certificate and holds it in memory, as
`swarm_queue_client::ClientSecret::Value`. If the store is down, it
retries for about a minute and then fails the start, so `Restart=` tries
again.

Policy delta: the controller gets `read` on that leaf, and the publisher
gets `create`/`update` on that leaf.

Removed: the `queue.clientSecretFile` option (both spellings, now removed
options with a message), its singleHostSwarm default, the credential and
placeholder, and the path watcher plus its restart oneshot. A controller
without a store identity is now an eval error, because it has no other
way to get the secret.
2026-09-28 19:01:05 +02:00
atlas
3b8df27dcf swarm-ui: show each agent's icon on its card
Each agent card now leads with the agent's icon, loaded as an `<img>`
from `GET /api/agents/<name>/icon`: the same 5em square, background and
fallback as the hive dashboard's container row. An agent with no icon
(the route's 404), or any other failed load, shows the dimmed hyperhive
mark (`/favicon.svg`) instead of a broken image.

Only ever an `<img>`, never inline markup: the body is an agent-authored
SVG, and an image load does not run its script.
2026-09-28 13:47:37 +02:00
atlas
e974194e3a swarm: let an agent publish its own icon
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.

hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.

swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.

Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
2026-09-28 13:47:37 +02:00
iris
2e67656de8 docs: drop leftover 'nothing nests' framing from C0NTAINERS blurb
Per mara's review on #4769: don't extend stale info to say it's stale,
just remove it. The sentence only described the absence of the removed
topology tree, giving the operator nothing actionable.
2026-09-28 13:18:12 +02:00
iris
d9beb965a5 docs/web-ui: drop the redundant Topology tree section
C0NTAINERS already states the flat/alphabetical/no-nesting fact;
the section repeated it with zero new operator info. Fix the two
dangling references (README.md reading-path index, swarm.js
comment) rather than leaving them pointing at a removed heading.
2026-09-28 13:18:12 +02:00
iris
ef4e49805c docs/web-ui/dashboard.md: drop stale topology-tree history per mara review
state current fact only, not the old-vs-new narrative
2026-09-28 13:18:12 +02:00
iris
d447974536 docs/web-ui/dashboard.md: fix vale hits in topology-tree rewrite
- avoid 'backend' (Microsoft.Avoid)
- reword 'was removed' passive voice (write-good.Passive)
2026-09-28 13:18:12 +02:00
iris
6b5c8cc51a dashboard: flatten the agent list, drop dead topology-tree machinery (#4638)
The backend dropped the agent hierarchy's parent field, so every
container is a root and buildAgentTree/treePrefixDom could only ever
produce a single-level flat list — the .tree-prefix CSS lane rules
already matched nothing. Replaced with sortedContainerRows, a plain
alphabetical sort, and dropped the now-dead .tree-prefix/.tree-lane/
data-depth CSS and the depth/isLast/ancestorIsLast fields from the row
fingerprint and buildContainerLi's signature. No visible behavior
change - the rendered list was already flat, just via dead machinery.

Docs updated to describe the simpler implementation directly instead
of narrating the removal (kept the heading name since swarm.js still
points a comment at it).
2026-09-28 13:18:12 +02:00
atlas
a2b4acfb6a swarm-bao: say "run the bootstrap step" when the bootstrap token is dead, not "sealed"
swarm-bao-granter-role used its token for `bao auth list` with no check,
so an expired, revoked or policy-less token in bootstrapTokenFile exited
2 with a raw 403 and no hint. It now checks whether the store is up when
that call fails: if it is, the token is at fault, and the unit prints the
one-time bootstrap step and exits 4. A missing token file is still a
ConditionPathExists skip, so the two read differently in the journal.

The twelve granterLogin units printed "the store is sealed or
unreachable" on a healthy store, because `bao status` exited 1 there:
the CLI resolves a token helper under $HOME before asking, systemd sets
no HOME for a unit without User=, and the fallback shells out to
`getent`/`sh`, neither of which is on the unit's PATH ("failed to get
token helper: error expanding config path "": exec: "sh": executable
file not found in $PATH"). The check now runs with HOME=/var/empty and
keeps its stderr, so a genuinely unreachable store says why. When the
store is up and the login is refused, the units now name
swarm-bao-granter-role as the unit that writes the missing role.

setup.md's post-step restart used 'swarm-bao-*-policy.service', which
misses swarm-bao-agent-pki. It now names that unit too, and a
module-eval case fails when the restart misses any unit that logs in as
the granter.

Refs #4704
2026-09-28 11:03:08 +02:00
atlas
9cbee27e6d docs: reword vale-flagged prose in docs/swarm/README.md 2026-09-28 08:24:52 +02:00
atlas
727065c960 hive-agent: present this agent's own queue credential, then fall back
When the per-agent secret `queue-identity.nix` fetched is present, the
harness connects with `swarm-agent.<agent>.<secret>` as a static token
and publishes on `$SWARM.term.<agent>` and `$SWARM.agent-state.<agent>`.
When it is absent, or that first connect fails for any reason, a refusal
from a responder that does not verify agent tokens included, it connects
with the hive's shared OIDC client and publishes on the hive-scoped
subjects as before. Which one it took is logged once per connect.

`swarm_queue_client::connect_with_token` is the static-token connect: no
retry on the initial attempt, so the caller sees the refusal and can
fall back. Reconnects share the existing backoff, now a named function.

Closes #4630
2026-09-28 08:24:52 +02:00
atlas
fc97c237dc swarm-queue-client: one agent-token spelling, and no hive in AgentCredential
`swarm_queue_client::agent_token::format_agent_token` / `parse_agent_token`
are the spelling an agent presents its own queue secret in,
`swarm-agent.<agent>.<secret>`, and the one the auth-callout responder
reads back. The prefix is what separates it from an OIDC access token,
which may itself contain `.`. Parsing distinguishes "not an agent token"
(no prefix) from "a malformed one"; the error names the problem and never
the value. The module is store-free, so the agent formats its token
without linking the secret-store client.

`swarm_secret_client::queue::AgentCredential` loses `hive`: an agent's
identity is not tied to a hive, and nothing reads the field. Objects
already in the store carry it and still decode, since unknown fields are
ignored; a test parses one. The controller stops writing it.

With the credential no longer naming a hive, and the agent's policy
naming none since #4762, nothing in the mint consumes one. `hive` goes
from `mint_and_verify`, from the `MintAgentIdentity` node, and from
`POST /api/agents/{name}/identity`, which now takes no body and no longer
checks a hive against the roster; a caller that still sends one is not
refused, the body is ignored. `swarmctl agent mint-identity` loses
`--hive`, so passing it is now a usage error.
2026-09-28 08:24:52 +02:00
atlas
6170e74a31 swarm-bao: agent certificates issued by a store-generated agent CA
An agent's store identity was signed in swarm-controller's memory by a CA
a controller-host unit generated on disk, and the listener never trusted
that CA. Agent leaves now come from the store itself: a `pki-agents` PKI
mount whose root openbao generates internally, so the agent CA's key
never exists outside the store.

- swarm-bao-agent-pki (new, store host, as the bao granter): enables and
  tunes the mount, generates the root once (guarded on an empty issuer
  list, no replace branch), upserts the `swarm-agent` role (client
  certificates named `hive-agent-*` only, 90 days), caches the CA at
  /var/lib/swarm-bao-tls/agent-ca.pem and composes the listener bundle.
- The listener's tls_client_ca_file is a new listener-client-ca.pem
  (client-ca.pem, then the agent CA). Host cert-auth roles still pin
  client-ca.pem, so an agent leaf satisfies no host role. swarm-bao-certs
  composes the same bundle before openbao starts.
- openbao reads tls_client_ca_file only at start, so when the bundle
  changed after openbao started, swarm-bao-agent-pki restarts
  openbao.service in the container; under `seal = "shamir"` it prints
  the step instead. Once swarm-bao-certs has a cached CA, later boots
  start openbao with it and do not restart.
- The controller policy gains exactly `update` on
  pki-agents/issue/swarm-agent. mint_and_verify now asks that role for
  the leaf (the store generates the key), writes the agent's cert-auth
  role pinning the issuing CA bao returned, and writes the agent's
  policy as render_agent alone: the hive-shared queue credential stanza
  is gone.
- deploy.bao.agentPkiRoleName (must start `swarm-`, asserted with the
  other pki role names); swarm-controller gets
  SWARM_CONTROLLER_AGENT_PKI_MOUNT/_ROLE from the deploy.bao options.

Deleted: swarm-controller-agent-ca and its options (agentCaFile,
agentCaKeyFile), env, LoadCredential entries and assertion;
agent_identity's Authority, rcgen signing and validity window; the
rcgen and time dependencies of swarm-controller (rcgen leaves the
workspace); policy::render_agent_with_queue and its tests. The CN-prefix
assertion policy.rs said was owed is not: agent and host roles pin
different CAs.

Migration is re-creating each agent after deploy; that overwrites the
stale role and policy.

Closes #4756
2026-09-27 22:59:27 +02:00
atlas
e9cec0da21 swarm-bao: write every swarm-* grant as a bao granter, not with a 24h token
Every unit that writes a bao policy or cert-auth role ran only while the
operator-placed bootstrap token existed, and skipped silently otherwise.
The token lives 24h, so on any real swarm a PR adding or changing a grant
deployed with its unit skipped, and each one needed a manual token refresh
(plus a root `bao policy write` when it added a path).

A `bao-granter` principal now writes them. Its leaf is minted by
swarm-bao-pki on the store host (0600 root, never copied off it), and its
policy covers `swarm-*` policies, `swarm-*` cert-auth roles and
`pki/roles/swarm-*` by glob, plus the mount and services-root paths the
controller's unit already used. All ten granting units
(controller, secret-publisher, matrix-ctl, matrix-token, queue-agent,
grafana-oidc, otel-oidc, forwarder-oidc, services-issuer, nats-tls) log in
with it instead of reading the token. They keep the 2880 x 30s retry, now
require swarm-bao-pki, and when the store refuses the granter they fail
and print the one-time step instead of skipping.

swarm-bao-granter-role is the one unit left on the token. It enables the
auth mounts (moved out of the controller's unit) and writes the granter's
own policy and role. The bootstrap policy is renamed `bao-bootstrap` and
shrinks to those five stanzas; it is shipped at
/etc/hyperhive/bao-bootstrap-policy.hcl. The old name `swarm-bootstrap`
matched the granter's own `swarm-*` glob.

The granter's CN joins certAuthCns, so no hive can be named into its role.
An assertion keeps both pki role names under `swarm-`. With no client CA
the granting units no longer render, and a warning says so.

module-eval pins the granter's policy stanza by stanza, what it cannot
reach, that every call a granting unit makes is granted, and that only
swarm-bao-granter-role reads the token.

Refs #4704
2026-09-27 22:57:46 +02:00
atlas
3385aaf026 docs: drop "simply" from the knowledge sweep page 2026-09-27 20:15:55 +02:00
atlas
313d582c01 job_queue: drop insert_unless_live, accept extra queue passes
mara chose to accept extra queued sweep passes over adding a new
hive-jobq primitive (or a hive-c0re one-off) for "don't queue another
of this kind". Every sweep caller now plain-inserts its node; the
capacity-1 Dep::Resource per sweep kind (MatrixSweep/KnowledgeTree)
still keeps two passes of the same kind from running concurrently, it
just no longer collapses a tick that lands while one is live or queued
into the existing one.
2026-09-27 20:15:55 +02:00
atlas
1d4c77d2c8 hive-c0re: serialise the matrix and knowledge sweeps through the job queue
The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
2026-09-27 20:15:55 +02:00
atlas
2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00
atlas
c6a778f666 lint: tighten atomic-write-secret.nix's header comment; fix vale contractions in persistence.md
The header comment grew to 39 lines across two audit-driven rounds,
over the 30-line comment-block-lint max. Trimmed to 30: merged the
value-as-argument rationale with its /proc/cmdline justification into
one paragraph, cut the usage example to one call instead of two, and
condensed the args paragraph — no content dropped, just restatement.

persistence.md's first-boot-migration marker paragraph used "it is"
and "there is not" — vale's Microsoft.Contractions rule (this repo's
config) wants the contracted forms, and "is not" also collides with
"is nothing" as a literal substring, which is what actually tripped
the error. Reworded to "it's" / "there's nothing", no meaning change.

Lint-only: no script logic changed, gates re-run below are all lint
checks (no module-eval, no cargo).

Refs #4723
2026-09-26 21:50:03 +02:00
atlas
ed53e9abcc host-modules: write credential files atomically; agent-modules: retry a failed .claude migration
Four host glue units fetched a secret from swarm-bao and rendered it
with `> path; chmod`: a reader racing the write could see a truncated
file, and briefly one at the wrong mode before the chmod landed.
glue-matrix-bao-token.nix, glue-queue-agent-credential.nix (both
files), swarm-grafana.nix and swarm-otel.nix now write to a same-
directory temp file, set its final mode/owner, then `mv -f` it over
the target — a shared `atomic_write_secret` helper
(nix/host-modules/lib/atomic-write-secret.nix) so the five call sites
share one implementation.

The first-boot `/root/.claude` migration in nix/agent-modules/user.nix
wrote its done-marker unconditionally, so a failed `cp` (disk full,
permission error) left the marker behind and no boot ever retried the
copy. The marker is now written only when there was nothing to
migrate or the copy succeeded; `cp -an`'s no-clobber semantics already
make a retry after a partial copy safe.

Refs #4723
2026-09-26 21:50:03 +02:00
atlas
3202cde704 nix: gate hive-c0re on deploy.hive-controller.enable, drop hyperhive.enable
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.

Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:

- Dropped: each swarm service and its glue keeps only its own deploy
  toggle (authelia, bao and its PKI glue, grafana, victorialogs,
  victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
  the controller/nats/matrix-ctl/publisher/services-issuer identities),
  the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
  its assertion, which feed hive-c0re and write under its state dir, plus
  their policy-order entry; the network identity assertions; hive-tls's
  two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
  (`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
  `deploy.matrix.enable` plus their client-identity checks. They read
  deploy.matrix's token file and registration script; their deploy.bao
  inputs are the client-half options a hive sets to read a store it does
  not run, so gating on deploy.bao.enable would drop the tested
  remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
  and fires wherever the hive, the store or the homeserver runs: each
  turns the name into an identifier with no fallback.

On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.

Refs #4500
2026-09-26 01:19:49 +02:00
atlas
ed2ec52fe5 swarm-tls: narrow each gateway's services leaf to the names it fronts
Every gateway asked the store's `pki/issue/swarm-services` for the whole
swarm's service set, so a private key on any gateway host could serve a
valid certificate for services that host does not front and never has.

`swarm.localServiceDomains` derives the per-host subset by filtering
`swarm.serviceDomains` against the vhosts this host actually renders —
the deploy flags those vhosts are already guarded on, read once rather
than copied into a second filter. The leaf request and the coverage
guard that decides whether to re-issue both read it, so they cannot
disagree about which names the leaf owes.

The sub-CA's name constraint and the role's `allowed_domains` stay the
swarm-wide set: every host's subset is inside it, and narrowing the
constraint per host would turn one signing into N.
2026-09-25 23:41:10 +02:00
atlas
0649673ebf hive-tls: renew the swarm-services leaf on a daily timer
The store's `swarm-services` role issues the services leaf for 720h, and
`swarm-services-cert` only ever ran at boot or rebuild: it is a
`RemainAfterExit` oneshot wanted by `multi-user.target` and no timer
targeted it. A hive not rebuilt within 30 days served an expired leaf.

`swarm-services-cert-renew` runs the same script from a daily timer. It
is a unit of its own because a timer starting the `RemainAfterExit` unit
is a no-op, and restarting that unit instead would propagate through
`hive-gateway-self-signed-cert`'s `Requires=` to nginx, so a sealed store
would take the gateway down over a still-valid leaf. Nothing requires or
orders against the new unit; it has no `Restart=`, so a failure stays in
`systemctl --failed` until the next tick, and the script only moves files
into place after the store has answered.

The re-issue threshold was `checkend 2592000`, the whole 30-day
lifetime, so every run re-issued. It is now half the role's lifetime,
read from a new internal option `deploy.bao.servicesPkiLeafTtlHours`
that the role's `ttl`/`max_ttl` also read. Boot and timer share the
script and so the threshold. The services-root re-check reads the
same option, at the store's own replacement threshold (hours × 3600),
so the hive asks for a new leaf when the store replaces its root. A
`flock` keeps the two runs from interleaving one issuance's key with another's leaf.

`checks.module-eval-hive-tls` pins the timer, that the unit it starts
re-runs the issuance without `RemainAfterExit`, that nothing depends on
it, and that both the leaf and root thresholds move with the option.

Closes #4587
2026-09-25 23:38:36 +02:00
atlas
20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00
atlas
85ba45b2de docs: agents' matrix accounts come from the swarm
The integration, tool, persistence and setup pages described hive-c0re
minting each agent's token into its state dir. They now describe the
swarm's appservice, its admin sender, the controller's mint and pass, the
daemon's store read and re-start timer, and the two new credential rows.
ruth's matrix account comes from the same pass once she holds the store
identity the setup page already has the operator mint.
2026-09-25 08:31:01 +02:00
atlas
89a5dd752c hive-c0re: stop minting agents' matrix accounts
The swarm mints each agent's `main` account now, so the hive's own mint
goes: `ensure_user_for`, `finish_user_provisioning`, `sync_agent`,
`sync_agent_standalone`, `token_path`, `legacy_password_path`,
`auto_reset_password` and `token_file_present`, and the calls from the startup sweep and the
rebuild bookkeeping. Both mints pinned the device `hyperhive-<agent>`, so
leaving this one would have each re-login kill the other's token.

`hivectl matrix create-user` refuses an agent's name and says where its
account comes from. Everything that still uses the hive's appservice token
stays: the hive's own account, the Space and chat room, and operator
accounts.
2026-09-25 08:31:01 +02:00
atlas
0cbb7db2c0 docs: a first SSO login makes a human's forge account
setup.md said Swarm SSO creates the operator's forge account, which was
not true until the previous commits. It now says how: sign in to the
forge once through authelia, then `swarmctl forge make-admin <you>`.
sso.md says what that first login does and why ACCOUNT_LINKING is
`login`. README, hivectl.md and forge.md drop `hivectl forge
create-user`, and the swarmctl README gains `forge make-admin`.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
d56d8f2b36 swarmctl: forge make-admin
Calls POST /api/forge/users/{name}/admin and prints what it found. It
fails with the controller's message when the user has not logged in via
SSO yet, and when the name is an agent's.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
ef494af188 hivectl: drop forge create-user; SSO makes a human's forge account
The forge now creates a human's account on their first authelia login,
so the verb has no job left. Deletes it, HostRequest::ForgeCreateUser,
its handler, provision_user_token, change_user_password and the hive's
TOKEN_SCOPES. change_user_password also passed the password as an
argument to `forgejo admin user change-password`, so it showed in the
container's process list.

ensure_user_exists and mint_token stay for the `core` bootstrap, their
one caller now. ensure_user_exists loses its password parameter: only the
deleted path set one.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
92e1909caf swarm-bao: give the store forwarder's OIDC reader its own bao identity
`swarm-bao-forwarder-oidc` fetches the store container's collector secret,
one path, and was the last reader still logging in with
`deploy.bao.clientCertFile`: the hive's own leaf, whose policy reads every
agent's credentials, the hive's tree and every service's OIDC secret. The
four-way split gave grafana's and the swarm collector's readers leaves of
their own and left this one behind.

It now holds `forwarder-oidc.pem`, minted by `swarm-bao-pki`, and logs in
under the `swarm-forwarder-oidc` cert-auth role, whose policy reads
`secret/data/swarm/services/<store forwarder client id>/oidc/client` and
nothing else. The role is written by `swarm-bao-forwarder-oidc-policy`
from the bootstrap token, which gains the two grants that unit calls, and
the reader is ordered after it. The subject is reserved as a hive name. A
store host whose pair is null is refused at eval rather than falling back
to the hive's leaf.

The hive's own role and `client.pem` are untouched; nothing is revoked.
2026-09-25 00:37:31 +02:00
atlas
978164dc53 nix: run the forge on one host per swarm (deploy.forgejo.enable)
Every hive with hyperhive enabled ran its own hive-forge container, and
its gateway answered forge.<swarm> with its own bridge IP, so on a
multi-host swarm each hive talked to its own forge.

deploy.forgejo.enable defaults to false and allSwarmServices sets it with
mkDefault, like authelia and bao; singleHostSwarm gets it through that.
The forge's OIDC client moves to a glue module gated on authelia, so a
split authelia/forge swarm still registers it. CI now requires the forge
on the same host, and the controller's forgeTokenFile defaults to null
where the forge is not.

Closes #4705
Refs #3782
2026-09-24 23:56:07 +02:00
atlas
3ee5a960b4 docs/setup: ruth needs a store identity and a rebuild, not a hand-made forge user
The rewritten ruth step skipped her store identity, without which
neither the backfill nor her container's fetch can reach her token. With
the backfill now creating a missing forge user, two commands cover her:
swarmctl agent mint-identity ruth --hive <hive>, then hivectl agent ruth
rebuild so hive-c0re hands the identity to her container.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
abc942cff3 docs: the swarm mints agent forge tokens; hive-c0re and tea-login no longer do
credentials.md gains the forge-token row and drops the claim that the
forge token never passes through the store. setup.md says plainly that
an agent spawned on the hive alone, ruth's bootstrap included, now gets
no forge user from anything. CLI references regenerated.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
dd32a395f7 agents: pull the forge token from bao; drop tea-login
forge-token.nix fetches swarm/agents/<agent>/forge-token under the
agent's own store identity into /run/hive-agent-forge-token/token, and
re-fetches on a timer so a rotation lands. hive-forge, the git
credential helper, hive-forge-notify, forge-avatar-sync and the web UI
read that file first and fall back to <state>/forge-token.

tea-login is deleted: it copied the token into ~/.config/tea, which
docs/swarm/credentials.md forbids for a store secret. hive-forge covers
the same verbs. swarmctl gains agent mint-forge-token.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
a5259146dc swarm: default every queue URL to the queue's name on every hive
A remote hive dialled nothing until an operator copied the queue's URL
into it, though the URL is the same string everywhere. statusPublish.natsUrl,
queue.agentNatsUrl and controller.queue.natsUrl now default to
tls://<swarm.nats.domain>:<port> unconditionally.

The statusPublish assertion treated a URL without a secret as a half
config. With the URL a default on every hive, only the secret claims
publishing: the assertion now refuses a secret without a URL or token
endpoint, and hive-c0re's status environment is gated on the secret too,
so a hive without one publishes nothing instead of reading a missing
credential.
2026-09-24 17:26:31 +02:00
atlas
1d261b3fed swarm-nats: give the queue a name, a bao-issued leaf, and require TLS
The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.

- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
  `swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
  every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
  IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
  may only `update` `pki/issue/swarm-nats`, written by
  `swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
  paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
  as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
  the container, restarts nats when it rotates, and re-runs daily.
  It joins glue-bao-readers-policy-order, so it is ordered after its policy
  unit (`after` and `wants`, never `requires`) where the store is on the
  same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
  `allow_non_tls`. `validateConfig` is now off in every mode, because the
  build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
  host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
  `singleHostSwarm`, and the auth responder all dial
  `tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
  to the NATS connection too, so hive-c0re and the controller trust the
  leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
  swarm needs.

module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.

Closes #4626
2026-09-24 17:26:31 +02:00
atlas
d048fee698 docs/setup: point at the bootstrap policy file instead of inlining a copy (#4698) 2026-09-24 15:55:38 +02:00
atlas
e3fefb8c5f docs: fix markdown indent treefmt wants after removing the ownership-checks bullet 2026-09-24 12:24:32 +02:00
atlas
a65dbee982 docs: delete two more stale manager-override claims
destroy has no manager-name check (actions.rs:837 says the root
container is 'destroyable like any other'), and crash-watch has no
name check at all (polls every managed container uniformly). The
pointer to broker.rs/actions.rs/crash_watch.rs for 'owner-check logic'
is also stale — none of the three hold any.
2026-09-24 12:24:32 +02:00
atlas
d2c4c8d4a3 docs: drop dead-claim narration for the retired loose-ends override
The PR retiring loose-ends' manager-only visibility replaced the false
claim with prose narrating its retirement, which still adds lines for
a removal. Delete the dead claim outright instead of documenting that
it used to be true: drop the three added sentences in
agent-hierarchy.md, and shorten socket_server/mod.rs's authority
comment to state only what's true now (tool-group membership for the
orchestration verbs) rather than adding an explanatory clause for the
removed capability check.
2026-09-24 12:24:32 +02:00
atlas
f53508cb2b docs, hive-c0re: retire prose describing the removed get_loose_ends target
Follow-up to 729c5b4f42, which dropped the
`agent` parameter from get_loose_ends and deleted Capability::QueryAgentState
with it. Three prose sites still describe the interface that commit removed:

- socket_server/mod.rs's module doc claimed authority on this socket derives
  from "capabilities for the hive-wide queries and tool-group membership for
  the orchestration verbs". There are no capability-gated queries left —
  `git grep -n has_cap -- hive-c0re/src/socket_server/` returns nothing, and
  require_group("scheduling") is the only gate in dispatch_orchestration.
- agent-hierarchy.md listed loose-ends visibility ("manager sees hive-wide,
  sub-agents only their own") as one of the manager-only overrides that
  "exist across hive-c0re today". It doesn't, and the sentence pointed a
  reader at loose_ends.rs for owner-check logic that is no longer there.
- conventions.md's Loose-ends wire shape still spoke of "the agent-flavour
  and manager-flavour requests" and "the agent-flavour list". There is one
  request shape.

No behaviour change: prose only.

Refs #4480
2026-09-24 12:24:32 +02:00
atlas
33da51382e subagent: make interrupt cancel a goal run, not just its turn
`interrupt` killed the child and trusted that to end the run. It didn't:
the default signal is SIGINT, and `claude --print` handles SIGINT by
writing its terminal `result` event and exiting **zero**. A zero exit is
`TurnEnd::Complete`, so `spawn_and_track`'s killed-turn early return —
the code that was supposed to stop the run — never fired, and the loop
went straight on to spawn the next goal turn. An interrupted run carried
on to completion; the interrupt changed nothing about the outcome.

Measured, not inferred: `claude --print --verbose --output-format
stream-json`, SIGINT'd mid-turn, exits 0 on every run.

So record the reason instead of inferring it. `interrupt` writes
`StopReason::Cancelled` — the same path `goal_reached`/`need_help`
already use — while it still holds the `running` lock, so there is no
instant in which the name has lost its cancel handle but not yet gained
its stop reason. `plan_after_turn` checks stop reasons ahead of the goal,
so the run stops whichever way the child ends up exiting. `force`'s
SIGKILL still produces a `TurnEnd::Killed`, and `status` still prefers
the kill record it already had.

`continue`'s budget reset is untouched: it clears the stop reason and
hands back a fresh allowance, which is what makes a cancel a pause an
operator can undo.

The test spawns the real continuation loop against a fake claude that
exits 0 on SIGINT, interrupts it, and asserts turn two never starts —
it reports `left: 2` without the fix.
2026-09-23 23:14:13 +02:00
atlas
22a87f7268 docs/swarm/{ca,secrets}.md: reword the 8 vale write-good.Passive / Microsoft.Contractions hits
Active-voice / contraction rewrites only, no meaning changes; the
allowed_domains/SANs sentence and the published-cert table cell were
checked against the parked services-issuer/role-type/vhost-scope
questions on #4622 and don't touch any of them.
2026-09-23 21:00:02 +02:00
atlas
f4df4fc4a9 nix: issue the swarm-services leaf from bao's pki mount
The `pki` mount had no issuer and no principal could log in to it, so the
swarm's service certificates were still minted by two openssl hops from a
root key on disk. Close both halves and retire the openssl path with them.

The mount now generates its own root, once. The granting unit asks bao
whether an issuer already exists (`bao list pki/issuers`) before calling
`pki/root/generate/internal`, so a rebuild or a reboot re-asserts the role
and the grant without touching the anchor — a root that changed per boot
would invalidate every certificate issued under it and every browser
taught to trust it. The guard asks the store rather than looking for a
marker file on this host's disk: a file is a claim about a mount that may
have been restored from a snapshot or disabled and re-enabled underneath
it.

`swarm-services-issuer` stops being an inert policy. A fourth cert-auth
role attaches it, following the shape the controller, the publisher and
matrix-ctl already use, and glue-bao-tls.nix signs the leaf carrying its
CN — that credential is what opens the mount, so it cannot come out of it.

`swarm-services-cert.service` logs in with that leaf, calls
`pki/issue/swarm-services`, and writes the result to the path
hive-tls.nix already wrote and the gateway already copies from. The
sub-CA layer does not move; it stops existing. The role's
`allowed_domains`, read from the same `swarm.serviceDomains` the SANs
come from, enforces at issue time what the sub-CA encoded in x509
`nameConstraints`, and with the root inside the mount there is nothing
left for an intermediate to be an intermediate of.

Not a flag day: the issuing root is published beside the leaf as
`swarm-services-root.pem` (0644) and joins `trust-bundle.pem`, where the
swarm root still sits. A leaf chaining to the old sub-CA and one issued
by the store both verify against the same bundle, so hives can be
rebuilt in any order. The same file is what an operator hands a browser
— readable without a store login, which matters because every listener
demands a client certificate.

The eval-time warning about uncovered service names is gone rather than
reworded. It fired on "this host does not hold the swarm root key", which
was the reason a hive could end up serving its own leaf on a
swarm-service name. Every hive now asks the store with its own identity,
so that stopped being the thing that decides.

Closes #4586
2026-09-23 21:00:02 +02:00
atlas
d3e4951cc8 swarm-bao: refuse a remote reader that named seven of the eight leaves
The four-way client-cert split gives each store reader its own leaf, and
three of the four readers render only where their own leaf exists. On a
host that mints its own PKI glue-bao-tls.nix defaults all eight, so there
is nothing to do; on a hand-configured remote-store hive, omitting one
pair used to mean that unit silently did not render — a privilege-
narrowing unit absent from a green build, with the missing unit as the
only evidence.

Each of the three now asserts its own pair, shaped after
swarm-grafana.nix's haveClientIdentity assertion and named to the pair it
needs. What differs from Grafana's is the gate: these fire only where the
host demonstrably reads the store (it holds deploy.bao.clientCertFile and
clientKeyFile) and the consumer is on. A host with no store identity is
the supported no-store deployment and still evaluates; the collector's
no-secret degrade is untouched, because that host holds no clientCertFile
either.

Also rewords three passive-voice sentences in docs/swarm/secrets.md that
vale flagged, and documents what the refusal costs and where it stays
silent.
2026-09-23 10:11:42 +02:00
atlas
f1445b4c8b swarm-bao: give each hive-cert consumer its own bao identity
Four units read one path each out of the store, and all four logged in
holding `deploy.bao.clientCertFile` — the hive's own leaf. Bao identifies
a principal by the subject of the certificate it presents, so four
readers behind one certificate were ONE principal, and the only grant
expressible was the union of what the four need: read on
`swarm/agents/*`, `swarm/hives/<hive>/*` and `swarm/services/*`. The unit
fetching Grafana's OIDC client secret could fetch every agent credential
in the swarm; the one fetching this hive's matrix token could fetch
Grafana's. Least privilege was not misconfigured here, it was
unrepresentable.

Each now holds a leaf, a cert-auth role and a policy of its own, and each
policy is the single `secret/data/…` path that unit's own script names —
spelled to the leaf, not to a prefix, the way matrix-ctl's already is.
Following the four exemplars in-tree rather than building a mechanism:
`signLeaf` mints the leaves, `swarm-bao.nix` writes the roles from the
bootstrap token, the consumers name their own pair.

Two of the four are written PER HIVE and two are not, which is the shape
of the paths rather than a preference. A matrix appservice token and a
queue credential live under `swarm/hives/<name>/` and every hive runs a
reader for its own, so one role for all of them would have to be granted
`hives/*` — letting one hive read another's, a reach no hive has today.
An OIDC client secret lives under `swarm/services/<client-id>/` and a
swarm registers each exactly once, so one role each is enough. The
per-hive subjects are `<prefix>-<hive>` and swarm.nix reserves every
composed spelling as a hive name, so a hive cannot be named into another
hive's role.

The shared leaf stays: hive-c0re still passes it into its container, the
`bao` CLI wrapper still defaults to it, and the three
`glue-*-bao-identity.nix` files derive the PKI directory from it.

module-eval-bao-grants gains a negative arm per principal — each pins the
three stanzas the hive's leaf carried and the two wildcards a later
widening would reach for, so a policy that grows fails here rather than
in a store. Plus the consuming side: repointing a unit back at the hive's
leaf would evaluate, deploy and log in, and silently restore the union.

A hive that reads a store on another machine now places one leaf per
principal instead of one shared by four. That cost is the point, and
docs/swarm/secrets.md lists the pairs.
2026-09-23 10:11:42 +02:00