An agent's subagent daemon publishes each subagent's output as terminal
rows on `$SWARM.term.<agent>.sub.<subagent>`, as the agent, into a
per-agent stream it creates itself; swarm-controller lists an agent's
subagents from that stream's subjects and relays one subagent's rows as
SSE; the swarm UI lists them under the agent's terminal preview and
reuses AgentTermPreview, full-screen tab included, with no input.
- swarm-nats.nix: the agent token may also publish
`$SWARM.term.{agent}.sub.>` and `$JS.API.STREAM.CREATE|INFO` on
`term-sub-{agent}`, and nothing else of JetStream. A module-eval arm
pins the agent-token grant as an exact list.
- mcp.nix: hive-subagent-daemon loads the agent's store identity
(`hive-agent-bao-cert/-key/-server-ca`, the ones hive-agent loads)
whenever the agent has a store, not only on the opencode preset. The
agent's own queue secret lives in the store, so this is the credential
the harness connects with.
- hive-subagent-mcp: `swarm_term` reads the agent's queue secret under
that identity, connects with the agent token, opens or creates
`term-sub-<agent>` (max_age 24h), and publishes classified rows from
the sink every subagent line already passes through. The sink only
queues (bounded, drop-and-count); a missing store, refused credential,
failed stream create or failed publish is a log line.
- The stream-json classifier (`stream_enrich`) and the `TermMsg` row
types plus `fit` move from the hive-agent binary into hive-sh4re, so
the subagent daemon publishes the rows AgentTermPreview already
renders. hive-agent keeps its LiveEvent classifier on top.
- swarm-controller: `GET /api/agents/{name}/subagents` and
`GET /api/agents/{name}/subagents/{subagent}/term/stream`.
- docs/swarm: what the UI shows and what the queue carries.
Closes#4827
authelia 4.39.20 exits at startup on `users: {}` ("users: non zero value
required"), and the first-boot unit seeded exactly that, so a swarm with
no users crash-looped authelia and answered 502 until `swarmctl user add`
ran.
The first-boot unit now writes one subject, `swarm.placeholder`, when
the users database is absent, empty, or exactly `users: {}`:
- `disabled: true` — authelia returns "user not found" for a disabled
user before any password check (file_user_provider.go,
CheckUserPassword).
- password: an argon2id digest with an all-zero key. It decodes (authelia
rejects a non-digest at startup) and no known password hashes to it.
- the `.` keeps it out of agent names (`[a-z0-9-]`), and `swarmctl user
add` refuses it as already existing. Neither writer removes users, and
both round-trip `disabled`.
A file with any user in it is never touched.
The docs that described the crash-loop (sso.md, gateway.md, setup.md,
the sso-unavailable error page) now describe the placeholder; the
writers' load_store docs and the seed fixtures follow. module-eval
nats-authelia asserts the seed branch.
The auth callout grants an agent that presents its own queue credential
one more subject, `$KV.agent-icons.<agent>`: its own key in the
agent-icons bucket and no other. The hive's shared agent client is
granted none of the bucket, since every agent on a hive presents it.
hive-agent writes `/etc/hyperhive/icon.svg`, the file its `GET /icon`
serves, to that key once per start, as a JetStream publish straight to
the subject (what `kv::Store::put` sends, minus the bucket lookup), so
the one subject is the whole grant. No icon deletes the key. A failed
write, including one that arrives before the bucket exists, is retried
with backoff until acked. An agent connected with the hive's shared
client publishes nothing.
swarm-controller creates the bucket as soon as its queue connection is
up, instead of on the first icon read, so an agent's write does not
wait for someone to look.
Measured against a local nats-server with a user allowed publish on
`$KV.agent-icons.atlas` only: the write to its own key is stored and
readable, a write to `$KV.agent-icons.argus` is refused (the ack times
out), the DEL marker makes the key read as absent, and a write before
the bucket exists fails with "no responders".
An `auth_token` spelled `swarm-agent.<agent>.<secret>` is no longer sent
to introspection. The responder reads `swarm/agents/<agent>/queue` with
an identity of its own, checks that the stored object names the same
agent, compares the secret in constant time, and grants the subjects
`--agent-token-publish-subject` lists with `{agent}` expanded. Every
other outcome denies: a malformed token, no store identity, nothing
stored, a failed or slow lookup, a different secret. A token without
the prefix takes the OIDC path unchanged.
The journal's `auth request` line names such a caller `agent:<agent>`;
the hive-shared credential keeps `hive-<h>-agent`.
The new principal: a `swarm-nats-auth` cert-auth role and policy with
read on `secret/data/swarm/agents/+/queue` alone, a leaf signed by the
store's PKI glue, and `glue-nats-auth-bao-identity.nix` pairing the two.
The copy unit delivers the identity into the queue's container, and an
absent leaf is delivered empty so the responder still starts and only
agent tokens are refused.
The policy and role are written by `swarm-bao-nats-auth-policy`, logged in
as the bao granter: both names fall under its `swarm-*` globs, so the
deploy writes them with no operator step. module-eval counts it among the
granting units, so every generic granting-unit case covers it.
The secret compare uses `subtle`, already in the lock file through the
TLS stack; no workspace crate offered one directly.
`swarm.authelia.url` defaulted to `https://<domain>` only when this host
ran the container, and to `null` otherwise — so the address a client is
given was a statement about co-location rather than about the swarm. A
swarm has one SSO provider; every hive addresses the same name and
resolution decides which address that reaches, exactly as
`swarm.otel.domain` already works.
The option stays nullable: "this swarm has no IdP" is still expressible,
it is just now something an operator states rather than something not
running the container produces. The Grafana fixture that exercised the
no-IdP refusal says it explicitly.
Closes#4536
The single module-eval derivation forced ~62 full nixosSystem
fixtures live at once to compute its cases list: 10.6GB peak RSS /
5m25s to evaluate, by far the dominant cost in nix flake check.
Splits it into 21 independent checks.module-eval-* derivations
(1-7 fixtures each) sharing builders/helpers via module-eval/lib.nix,
so no single derivation needs more than a handful of fixtures live
at once. A few cases spanning two clusters carry a small duplicated
fixture rather than threading shared state through lib.nix.