Commit graph hyperhive/nix
Author SHA1 Message Date
atlas
ed2ec52fe5 swarm-tls: narrow each gateway's services leaf to the names it fronts
Every gateway asked the store's `pki/issue/swarm-services` for the whole
swarm's service set, so a private key on any gateway host could serve a
valid certificate for services that host does not front and never has.

`swarm.localServiceDomains` derives the per-host subset by filtering
`swarm.serviceDomains` against the vhosts this host actually renders —
the deploy flags those vhosts are already guarded on, read once rather
than copied into a second filter. The leaf request and the coverage
guard that decides whether to re-issue both read it, so they cannot
disagree about which names the leaf owes.

The sub-CA's name constraint and the role's `allowed_domains` stay the
swarm-wide set: every host's subset is inside it, and narrowing the
constraint per host would turn one signing into N.
2026-09-25 23:41:10 +02:00
atlas
0649673ebf hive-tls: renew the swarm-services leaf on a daily timer
The store's `swarm-services` role issues the services leaf for 720h, and
`swarm-services-cert` only ever ran at boot or rebuild: it is a
`RemainAfterExit` oneshot wanted by `multi-user.target` and no timer
targeted it. A hive not rebuilt within 30 days served an expired leaf.

`swarm-services-cert-renew` runs the same script from a daily timer. It
is a unit of its own because a timer starting the `RemainAfterExit` unit
is a no-op, and restarting that unit instead would propagate through
`hive-gateway-self-signed-cert`'s `Requires=` to nginx, so a sealed store
would take the gateway down over a still-valid leaf. Nothing requires or
orders against the new unit; it has no `Restart=`, so a failure stays in
`systemctl --failed` until the next tick, and the script only moves files
into place after the store has answered.

The re-issue threshold was `checkend 2592000`, the whole 30-day
lifetime, so every run re-issued. It is now half the role's lifetime,
read from a new internal option `deploy.bao.servicesPkiLeafTtlHours`
that the role's `ttl`/`max_ttl` also read. Boot and timer share the
script and so the threshold. The services-root re-check reads the
same option, at the store's own replacement threshold (hours × 3600),
so the hive asks for a new leaf when the store replaces its root. A
`flock` keeps the two runs from interleaving one issuance's key with another's leaf.

`checks.module-eval-hive-tls` pins the timer, that the unit it starts
re-runs the issuance without `RemainAfterExit`, that nothing depends on
it, and that both the leaf and root thresholds move with the option.

Closes #4587
2026-09-25 23:38:36 +02:00
atlas
afde9f380f module-eval: give the forgejo-mirror fixtures deploy.forgejo.enable
#4708 defaulted deploy.forgejo.enable to false, and the hive-forge module's
whole config block (including the mirror-declared-without-a-controller
warning) now hangs off it. controllerWithMirrors and mirrorsNoController
both relied on the old always-on default to reach that warning path.
2026-09-25 08:36:05 +02:00
atlas
20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00
atlas
ab153bda2f hive-matrix-mcp: read the main account's token from the store too
The daemon reads each account's token from `swarm/agents/<agent>/matrix/`
as the agent itself, inside its own container, and falls back to the file
only when the store has none. This is #4519's read, without its `main`
carve-out: the swarm now mints `main` there and no hive writes the file.

The daemon unit gets the agent's store identity, spelled the way
forge-token.nix spells it. A timer re-starts it while it is down: a token
the swarm mints or replaces in the store changes no file, so the path
watcher never fires for it, and a daemon that exited on a replaced token
would otherwise stay down until the container restarts.
2026-09-25 08:31:01 +02:00
atlas
2776e121e5 swarm-controller: mint each agent's matrix account with the swarm's token
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.

The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.

`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
2026-09-25 08:31:01 +02:00
atlas
9308752a09 hive-matrix: load the swarm's appservice and promote its sender
The matrix container renders the swarm registration before tuwunel starts
(`requiredBy` it, no network), tuwunel loads it as a second `.yaml`
credential, and a publish unit hands its token to the store. tuwunel 1.9.1
refuses only a duplicate id or as_token, not overlapping non-exclusive
namespaces, so it sits beside the hive's `hyperhive` registration.

`admin_execute` promotes exactly `@swarm` at boot. The module-eval pin
narrows from "no account" to that one list, compared whole so a second
entry fails; `admin_execute_errors_ignore` is pinned too. matrix-ctl may
write the swarm token's leaf, and the controller may only read it.

Everything is gated on matrix-ctl's store identity: with nobody to publish
the token, the registration would be an admin credential nobody reads.
2026-09-25 08:31:01 +02:00
atlas
cb176c7be7 swarm-bao: stamp collector's service.name as "bao"
bao's own in-container collector labelled every log line and metric it
forwards `service.name=swarm-bao` (`processors.resource.attributes`,
keyed off `swarm.bao.machine`), while every panel in the shipped
Grafana bao dashboard queries the literal `service.name="bao"` — no
panel has matched since the scrape moved into that collector.

Rename it in the collector instead of templating the dashboard: a new
`collectorServiceName` binding in swarm-bao.nix, deliberately not
`cfg.machine` (that value names the container/receiver, not bao's
display identity), stamps `service.name="bao"` directly. bao.json is
back to its origin/main shape, unchanged.

Adds a module-eval case to checks.module-eval-grafana that reads the
collector's own evaluated config and asserts its service.name matches
every selector the shipped dashboard uses; verified invert-proof by
setting the value back to cfg.machine and confirming that specific
case (and only it) fails.
2026-09-25 08:30:25 +02:00
atlas
113f3fe6e2 hive-forge: a first authelia login creates the forge account
[oauth2_client] turns on auto-registration through the authelia login
source, with the account named after authelia's preferred_username.
DISABLE_REGISTRATION stays true: forgejo 16's auto-registration checks
only ALLOW_ONLY_INTERNAL_REGISTRATION, so local sign-up stays off.

ACCOUNT_LINKING is `login`, forgejo's default, set explicitly. With
`auto`, an SSO login whose name matches an existing local account would
be handed that account, and agents, `core` and `swarm-controller` all
have one. `login` asks for that account's own password instead.

Refs #3782
2026-09-25 08:29:56 +02:00
atlas
c4edb78aa2 swarm-otel: correct the hostJournalDir comment for swarm-bao's exception
argus review on #4713: the comment said every container block sets
--link-journal=host, and warned that dropping it silently blinds this
receiver. #4713 does exactly that for swarm-bao (its journal stays
inside the container for its own collector instead), so read on its
own the comment told a debugger to revert the fix. State the
exception.
2026-09-25 04:02:08 +02:00
atlas
44009dc51d swarm-bao: stop linking the container's journal onto the host
--link-journal=host bind-mounts a host directory journald never
writes into for this container (empty, root:nogroup, confirmed on
the live host — #4527). Dropping it falls back to nixpkgs' default
--link-journal=try-guest, the same shape every agent container
already uses, so the in-container collector's journald receiver
(directory = /var/log/journal) now reads a journal that is actually
written.

merge stays true and services.hyperhive.swarm.otel.journaldUnits is
untouched so this deploy changes exactly one thing; comments that
described the old host-linked shape are rewritten to match.

Adds a module-eval assertion (bao-otel-collector.nix) that the
container's extraFlags never re-add --link-journal=host.

Refs #4527, #4499.
2026-09-25 04:02:08 +02:00
atlas
92e1909caf swarm-bao: give the store forwarder's OIDC reader its own bao identity
`swarm-bao-forwarder-oidc` fetches the store container's collector secret,
one path, and was the last reader still logging in with
`deploy.bao.clientCertFile`: the hive's own leaf, whose policy reads every
agent's credentials, the hive's tree and every service's OIDC secret. The
four-way split gave grafana's and the swarm collector's readers leaves of
their own and left this one behind.

It now holds `forwarder-oidc.pem`, minted by `swarm-bao-pki`, and logs in
under the `swarm-forwarder-oidc` cert-auth role, whose policy reads
`secret/data/swarm/services/<store forwarder client id>/oidc/client` and
nothing else. The role is written by `swarm-bao-forwarder-oidc-policy`
from the bootstrap token, which gains the two grants that unit calls, and
the reader is ordered after it. The subject is reserved as a hive name. A
store host whose pair is null is refused at eval rather than falling back
to the hive's leaf.

The hive's own role and `client.pem` are untouched; nothing is revoked.
2026-09-25 00:37:31 +02:00
atlas
978164dc53 nix: run the forge on one host per swarm (deploy.forgejo.enable)
Every hive with hyperhive enabled ran its own hive-forge container, and
its gateway answered forge.<swarm> with its own bridge IP, so on a
multi-host swarm each hive talked to its own forge.

deploy.forgejo.enable defaults to false and allSwarmServices sets it with
mkDefault, like authelia and bao; singleHostSwarm gets it through that.
The forge's OIDC client moves to a glue module gated on authelia, so a
split authelia/forge swarm still registers it. CI now requires the forge
on the same host, and the controller's forgeTokenFile defaults to null
where the forge is not.

Closes #4705
Refs #3782
2026-09-24 23:56:07 +02:00
atlas
6d0c30ade2 module-eval: pin the agent forge-token fetch and tea-login's removal
Refs #3782
2026-09-24 17:48:53 +02:00
atlas
dd32a395f7 agents: pull the forge token from bao; drop tea-login
forge-token.nix fetches swarm/agents/<agent>/forge-token under the
agent's own store identity into /run/hive-agent-forge-token/token, and
re-fetches on a timer so a rotation lands. hive-forge, the git
credential helper, hive-forge-notify, forge-avatar-sync and the web UI
read that file first and fall back to <state>/forge-token.

tea-login is deleted: it copied the token into ~/.config/tea, which
docs/swarm/credentials.md forbids for a store secret. hive-forge covers
the same verbs. swarmctl gains agent mint-forge-token.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
58f50ed506 swarm-controller: guard the lockdown PATCH against a forge name collision
disable_repo_creation now reads the account once before PATCHing
max_repo_creation/source_id, and fails the node if the email isn't the
{agent}@hyperhive.local marker create_agent_user itself sets. The 409/422
create-fold (#4681) only proves some account with that name exists, not
that this node created it, so a pre-existing non-agent account sharing an
agent's chosen name could otherwise get locked onto local auth with repo
creation disabled. Leaves the fold untouched (Option B, per atlas/argus on
#4693); the read moves into disable_repo_creation instead.

Also fixes the nix-sandboxed cargo-test check: forgejo_api::Forgejo::new
builds a reqwest client that eagerly resolves TLS roots via
rustls-native-certs even for the tests' plain-http loopback stub server,
which panics with "No CA certificates were loaded from the system" in the
CA-less build sandbox. Gives that check's nativeBuildInputs pkgs.cacert and
sets SSL_CERT_FILE, same pattern this repo's runtime deployment already
uses for the same reqwest/rustls resolution.
2026-09-24 17:30:17 +02:00
atlas
a5259146dc swarm: default every queue URL to the queue's name on every hive
A remote hive dialled nothing until an operator copied the queue's URL
into it, though the URL is the same string everywhere. statusPublish.natsUrl,
queue.agentNatsUrl and controller.queue.natsUrl now default to
tls://<swarm.nats.domain>:<port> unconditionally.

The statusPublish assertion treated a URL without a secret as a half
config. With the URL a default on every hive, only the secret claims
publishing: the assertion now refuses a secret without a URL or token
endpoint, and hive-c0re's status environment is gated on the secret too,
so a hive without one publishes nothing instead of reading a missing
credential.
2026-09-24 17:26:31 +02:00
atlas
0081d75c86 swarm-bao: grant the bootstrap token the queue's pki role, policy and login role
swarm-bao-nats-tls-policy acts with the bootstrap token, and main's
module-eval-bao-grants now fails any such unit whose calls the policy
file does not grant. Adds its three paths and counts it among the units
the check must see.
2026-09-24 17:26:31 +02:00
atlas
1d261b3fed swarm-nats: give the queue a name, a bao-issued leaf, and require TLS
The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.

- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
  `swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
  every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
  IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
  may only `update` `pki/issue/swarm-nats`, written by
  `swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
  paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
  as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
  the container, restarts nats when it rotates, and re-runs daily.
  It joins glue-bao-readers-policy-order, so it is ordered after its policy
  unit (`after` and `wants`, never `requires`) where the store is on the
  same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
  `allow_non_tls`. `validateConfig` is now off in every mode, because the
  build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
  host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
  `singleHostSwarm`, and the auth responder all dial
  `tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
  to the NATS connection too, so hive-c0re and the controller trust the
  leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
  swarm needs.

module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.

Closes #4626
2026-09-24 17:26:31 +02:00
atlas
e3c595857e swarm-bao: keep the bootstrap policy in one file, checked against the units using it (#4698)
setup.md's copy of the swarm-bootstrap policy still granted only the
controller's first six paths, while eight units now act with the
bootstrap token. The policy moves to
nix/host-modules/swarm-bao-bootstrap-policy.hcl, now covering every path
those units call. module-eval-bao-grants reads that file and fails
when a unit whose script uses the token calls a path the file does not
grant.
2026-09-24 15:55:38 +02:00
atlas
88e571a463 swarm-controller: refuse a new agent name the forge would reject
Every agent becomes a Forgejo user of the same name, and nothing upstream
of `CreateForgeUser` knew what Forgejo refuses: `admin`, `api`, `foo-` or a
41-character name passed name validation and failed one node into
provisioning with Forgejo's 422.

`nix/reserved-names.nix` gains the 22 reserved usernames of Forgejo
v16.0.5 (`models/user/user.go:639-680`) that `[a-z0-9-]` can spell, and
the bare `-` (`models/repo/repo.go:67`). The dot and underscore entries are
left out, since our charset cannot produce them. The header's admission
rule grows a third class — a username the forge refuses — because that is
a failure behind the refusal.

The shape rules are not literals, so they live in
`hive_types::forge_username_violation`: no leading `-`, no `--`, no
trailing `-`, at most 40 characters. Beside `is_reserved_name`, not in
`Ident::parse`: an `Ident` is also a hive, label, account and subagent
name, and parsing runs on every read of an existing name.

`create_agent` used to WARN on a reserved name, deliberately: an operator
with agents already created under a colliding name would otherwise be
unable to re-run creation. That reason is kept, and narrowed to what it
protects. A name breaking either rule is now refused with a 400 naming the
rule when the name is NOT in the swarm roster, and still only warned
about when it is, so re-creating an existing agent keeps working. The
roster is read only for a rule-breaking name; when it cannot be read, a new
name and an existing one look alike, and this warns as before. The
hive-collision warning is unchanged.
2026-09-24 15:16:32 +02:00
atlas
0bfe354b6d nix: move the reader-after-policy edges into their own colocation glue
The four readers' ordering after their policy units only applies where
the store and that reader share a host, so it is colocation glue and
does not belong in the reader modules (two of which are main modules).
glue-bao-readers-policy-order.nix now sets the after+wants edges, gated
on deploy.bao.enable AND the reader's own gate, so a store host without
a reader gains no stub unit.
2026-09-24 15:15:15 +02:00
atlas
cde3fec956 module-eval: pin each swarm-bao reader's ordering after its policy unit 2026-09-24 15:15:15 +02:00
atlas
c92bb0dce7 nix: order each swarm-bao secret reader after the policy unit writing its role
The four readers (matrix-token, queue-agent, grafana-oidc, otel-oidc) log
in against a cert-auth role that their own swarm-bao-*-policy unit writes.
They were ordered after swarm-bao-pki and the store's container but not
after that unit, so on an apply a reader could log in before its role
existed and be refused by `allowed_common_names` until a retry landed
after the role did.

After= plus Wants= on the policy unit, never Requires=: the policy unit
skips by ConditionPathExists once the bootstrap token is gone, and a
skipped unit counts as done for ordering.
2026-09-24 15:15:15 +02:00
atlas
aa719da571 nix: ship the journals of the units an apply can leave failed
The units on the path a deploy takes to TLS, the store's grants and the
swarm collector itself were not on the host collector's journald
allowlist, so an ingest outage one of them caused showed in the store
only as every source going quiet at once.

Each module names its own units, per the option's rule:
- hive-tls.nix: hive-tls-ca, swarm-services-cert
- hive-gateway: hive-gateway-self-signed-cert (self-signed mode only)
- swarm-bao.nix: the seven grant units beside
  swarm-bao-services-issuer-policy
- swarm-otel.nix: container@<machine>, and
  nixos-rebuild-switch-to-configuration, the transient unit nixos-rebuild
  runs the activation in and whose syslog lines carry its status

The module-eval arm pins each unit as both listed and defined, since a
listed name that matches nothing is silent.
2026-09-24 15:14:44 +02:00
atlas
5b54b6ddc8 swarm-bao: fail the unit when the services root cannot be read, instead of replacing it
A failed `bao read` of the services root made the checkend pipeline
non-zero, so the unit deleted a working root and minted a new trust
anchor. A failed `bao list` of the issuers likewise looked like an empty
mount. Both now fail the unit, which retries on its own restart budget;
the root is replaced only when openssl parsed the returned certificate
and -checkend said it expires inside a leaf's window.

Closes #4663
2026-09-24 11:00:20 +02:00
atlas
c0c031a5e4 swarm-bao: let the journald receiver read the host-linked journal
The container's collector has never shipped a line. `journalctl --follow`
— which the journald receiver passes unconditionally — scopes itself to
the current boot unless `--merge` is given too, and `--link-journal=host`
makes /var/log/journal the HOST's journal tree, where this container's
current boot has no entry. journalctl exited 1 with "No journal boot
entry found for the specified boot (+0)" and the receiver respawned it
every ~2s, so nothing was ever read and nothing was ever exported.

`merge = true` is the receiver's key for `--merge` (buildArgs() in
pkg/stanza/operator/input/journald/config_linux.go at tag
receiver/journaldreceiver/v0.151.0, the deployed collector's version),
and --merge is what clears the implicit boot scope in journalctl.c
(systemd v260.4, the version on the host).

The module-eval arm pins both halves: the boot filter is gone AND the
directory is still the host-linked one — either alone is satisfiable by
the broken config.
2026-09-23 23:19:31 +02:00
müde
11ec050ca8 hive-tls: give swarm-services-cert the cmp it tests the root with
`cmp` lives in diffutils, not coreutils, so the root-changed test exited
127 with "command not found". Inside `if ! cmp -s`, a 127 reads as
"differs" and errexit never sees it, so `rootchanged` was 1 on every run
and the trust-bundle rebuild it guards bounced `hive-tls-ca` after each
issuance — the exact "only a changed one, or every boot would bounce a
unit with nothing to do" the comment there rules out.

Found in the journal of a hive that had just issued a leaf successfully:
the unit logged "the services root changed" on a run where the store had
left the issuer alone.
2026-09-23 22:18:32 +02:00
müde
5704c0c583 nix: unbreak the swarm-services leaf the gateway waits on
Three defects in the store-issued path, each of which alone kept nginx
from starting at all. The gateway's cert import `Requires=` this leaf, so
a leaf that is never issued is not a name mismatch — it is an empty
listener, and the swarm's own forge stopped answering on :443.

`swarm-services-cert` declared `Before=hive-tls-ca` for the trust
bundle's sake while also being `After=` the store's container, which is
itself `After=hive-tls-ca`. systemd resolved the cycle the only way it
can, by deleting the job, so the leaf went unissued on every activation.
The edge is gone; the bundle converges the other way round, through the
restart this unit already performed when the root it wrote was new.

The `pki` mount was enabled without `-max-lease-ttl`, so bao clamped the
30-year root to the 768h default and then refused every issue call,
because a leaf of the mount's own default length would outlive the CA
signing it. The mount is tuned on every run, the role pins a 720h leaf,
and a root that can no longer cover one is replaced rather than left to
refuse forever. A hive tests its own copy of that certificate against the
same threshold, so both ends reach a fresh leaf without signalling.

`swarm-bao-pki` mints the services-issuer leaf that opens the mount, but
only `swarm-bao-certs` required it. `RemainAfterExit` plus an
already-active unit means an activation that ADDS a leaf mints nothing —
which is how a host whose config named `services-issuer.pem` came to have
no such file. A target wants it now, like every sibling granting unit.
2026-09-23 21:36:46 +02:00
atlas
f4df4fc4a9 nix: issue the swarm-services leaf from bao's pki mount
The `pki` mount had no issuer and no principal could log in to it, so the
swarm's service certificates were still minted by two openssl hops from a
root key on disk. Close both halves and retire the openssl path with them.

The mount now generates its own root, once. The granting unit asks bao
whether an issuer already exists (`bao list pki/issuers`) before calling
`pki/root/generate/internal`, so a rebuild or a reboot re-asserts the role
and the grant without touching the anchor — a root that changed per boot
would invalidate every certificate issued under it and every browser
taught to trust it. The guard asks the store rather than looking for a
marker file on this host's disk: a file is a claim about a mount that may
have been restored from a snapshot or disabled and re-enabled underneath
it.

`swarm-services-issuer` stops being an inert policy. A fourth cert-auth
role attaches it, following the shape the controller, the publisher and
matrix-ctl already use, and glue-bao-tls.nix signs the leaf carrying its
CN — that credential is what opens the mount, so it cannot come out of it.

`swarm-services-cert.service` logs in with that leaf, calls
`pki/issue/swarm-services`, and writes the result to the path
hive-tls.nix already wrote and the gateway already copies from. The
sub-CA layer does not move; it stops existing. The role's
`allowed_domains`, read from the same `swarm.serviceDomains` the SANs
come from, enforces at issue time what the sub-CA encoded in x509
`nameConstraints`, and with the root inside the mount there is nothing
left for an intermediate to be an intermediate of.

Not a flag day: the issuing root is published beside the leaf as
`swarm-services-root.pem` (0644) and joins `trust-bundle.pem`, where the
swarm root still sits. A leaf chaining to the old sub-CA and one issued
by the store both verify against the same bundle, so hives can be
rebuilt in any order. The same file is what an operator hands a browser
— readable without a store login, which matters because every listener
demands a client certificate.

The eval-time warning about uncovered service names is gone rather than
reworded. It fired on "this host does not hold the swarm root key", which
was the reason a hive could end up serving its own leaf on a
swarm-service name. Every hive now asks the store with its own identity,
so that stopped being the thing that decides.

Closes #4586
2026-09-23 21:00:02 +02:00
atlas
ebcc5bde89 swarm-bao: shrink the read-grant comment under the comment-block max 2026-09-23 18:55:02 +02:00
atlas
87174863ca swarm-bao: grant the controller read on the agent credential prefix
mint_and_verify reads the queue credential back before writing it, so a
re-run keeps the value a live agent already authenticates with. that read
is read_optional, which maps only a 404 to absence — so with create/update
alone every mint aborted on a 403 at its first store read.

read on the same paths the stanza already grants create and update, and
nothing else: no list, no delete, no patch.
2026-09-23 18:55:02 +02:00
atlas
e2ff4f5281 swarm-bao: give the store's collector an explicit self-telemetry port
8890, so it stops claiming the hive collector's 8888 in the shared netns.
Extends the module-eval port case to all three tiers.
2026-09-23 18:04:33 +02:00
atlas
5cec69bd21 otel.nix: trim the StartLimit comment block to the load-bearing points
Cut ~27 lines of blackout-measurement and cross-reference narrative
(already in the PR body / issue) down to the three things a reader
actually needs at this call site: the [Unit]-vs-[Service] trap, why
Restart is absent, and the window-vs-burst constraint.
2026-09-23 17:22:51 +02:00
atlas
596c0c17bc otel: back the hive collector off a failed bind instead of burning its start limit
Every deploy on a hive host, the replacement opentelemetry-collector
reaches bind() while the outgoing process still holds
127.0.0.1:8888 (its self-scrape endpoint). nixpkgs sets
Restart = "always" with no RestartSec, so the unit spends its five
default attempts in under two seconds, hits start-limit-hit and stops
retrying — ~27s of telemetry blackout per deploy.

RestartSec = 5 with a 12-attempt burst over a 120s window rides the
race out instead: the blackout ends within one interval of the port
coming free, and 55s of it being held is survivable where 2s was not.

StartLimitBurst/StartLimitIntervalSec go at the systemd.services attr
level, which NixOS renders into [Unit]; under serviceConfig systemd
ignores them silently. module-eval-hive-otel asserts the placement.
2026-09-23 17:22:51 +02:00
atlas
549156e55f swarm-otel: persist the journald cursor across collector restarts
The swarm-tier collector's journald receiver had no storage extension, so
it started each run with no cursor: journalctl --follow --lines=0 ships
only what arrives after the receiver starts. Every collector restart
therefore dropped whatever was written to the journal while it was down,
silently — no error and no replay.

Wire the receiver to a file_storage extension, matching the agent-tier
collector in nix/agent-modules/otel.nix, so a restart resumes from the
persisted cursor instead.

Refs #4527.
2026-09-23 12:52:32 +02:00
atlas
9d339b1b58 module-eval: name the grafana helper the refusal readers are shaped after
Comment-only. The helper it points at is `grafanaRefusedFor`, not an
unnamed one.
2026-09-23 10:11:42 +02:00
atlas
d3e4951cc8 swarm-bao: refuse a remote reader that named seven of the eight leaves
The four-way client-cert split gives each store reader its own leaf, and
three of the four readers render only where their own leaf exists. On a
host that mints its own PKI glue-bao-tls.nix defaults all eight, so there
is nothing to do; on a hand-configured remote-store hive, omitting one
pair used to mean that unit silently did not render — a privilege-
narrowing unit absent from a green build, with the missing unit as the
only evidence.

Each of the three now asserts its own pair, shaped after
swarm-grafana.nix's haveClientIdentity assertion and named to the pair it
needs. What differs from Grafana's is the gate: these fire only where the
host demonstrably reads the store (it holds deploy.bao.clientCertFile and
clientKeyFile) and the consumer is on. A host with no store identity is
the supported no-store deployment and still evaluates; the collector's
no-secret degrade is untouched, because that host holds no clientCertFile
either.

Also rewords three passive-voice sentences in docs/swarm/secrets.md that
vale flagged, and documents what the refusal costs and where it stays
silent.
2026-09-23 10:11:42 +02:00
atlas
f1445b4c8b swarm-bao: give each hive-cert consumer its own bao identity
Four units read one path each out of the store, and all four logged in
holding `deploy.bao.clientCertFile` — the hive's own leaf. Bao identifies
a principal by the subject of the certificate it presents, so four
readers behind one certificate were ONE principal, and the only grant
expressible was the union of what the four need: read on
`swarm/agents/*`, `swarm/hives/<hive>/*` and `swarm/services/*`. The unit
fetching Grafana's OIDC client secret could fetch every agent credential
in the swarm; the one fetching this hive's matrix token could fetch
Grafana's. Least privilege was not misconfigured here, it was
unrepresentable.

Each now holds a leaf, a cert-auth role and a policy of its own, and each
policy is the single `secret/data/…` path that unit's own script names —
spelled to the leaf, not to a prefix, the way matrix-ctl's already is.
Following the four exemplars in-tree rather than building a mechanism:
`signLeaf` mints the leaves, `swarm-bao.nix` writes the roles from the
bootstrap token, the consumers name their own pair.

Two of the four are written PER HIVE and two are not, which is the shape
of the paths rather than a preference. A matrix appservice token and a
queue credential live under `swarm/hives/<name>/` and every hive runs a
reader for its own, so one role for all of them would have to be granted
`hives/*` — letting one hive read another's, a reach no hive has today.
An OIDC client secret lives under `swarm/services/<client-id>/` and a
swarm registers each exactly once, so one role each is enough. The
per-hive subjects are `<prefix>-<hive>` and swarm.nix reserves every
composed spelling as a hive name, so a hive cannot be named into another
hive's role.

The shared leaf stays: hive-c0re still passes it into its container, the
`bao` CLI wrapper still defaults to it, and the three
`glue-*-bao-identity.nix` files derive the PKI directory from it.

module-eval-bao-grants gains a negative arm per principal — each pins the
three stanzas the hive's leaf carried and the two wildcards a later
widening would reach for, so a policy that grows fails here rather than
in a store. Plus the consuming side: repointing a unit back at the hive's
leaf would evaluate, deploy and log in, and silently restore the union.

A hive that reads a store on another machine now places one leaf per
principal instead of one shared by four. That cost is the point, and
docs/swarm/secrets.md lists the pairs.
2026-09-23 10:11:42 +02:00
atlas
179f873722 docs: retire the agent hierarchy from every page that described it
The topology doc keeps its filename and its second half (manager
special-casing, harness unit shape) — both are cross-referenced from
other pages and neither is about the parent field. Its first half is
rewritten: what topology.json is now, and a table of what the removal
took with it, so a reader who finds `<parent>` or `set-parent` in an old
issue thread learns it went away rather than moved.

The dashboard's tree-rendering section is marked dormant rather than
deleted: the walk is still in swarm.js and retiring it is the frontend
owner's call.
2026-09-21 22:08:47 +02:00
atlas
afdfce67ec agent: fetch this agent's own swarm-queue credential from the store
Every agent on a hive authenticates to the swarm queue with the same
hive-scoped OIDC client, so at the auth callout one agent is
indistinguishable from its co-hived neighbours. The commit before this
one mints a secret per agent at swarm level into
secret/swarm/agents/<agent>/queue; nothing read it.

Read it here, and read it from the container itself. A hive courier in
the path would be the hive vouching for which agent this is, which is
the property a per-agent credential exists to remove -- so the agent
logs in to the store with the certificate hive-agent-bao-identity
already proves it can log in with, and reads its own path. The store
certificate is for reaching the store and nothing else: what the new
unit writes to /run is the secret it read back, and nothing hands a
BAO_CLIENT_* path to anything queue-shaped.

The read needs no policy change. render_agent grants read on
secret/data/swarm/agents/<agent>/*, which covers this path and the
bao-mtls one beside it alike -- which is also why this unit degrades
where the identity check fails. A refusal this unit sees and that check
did not cannot be a policy that drifted; it is an object not yet minted,
the ordinary state of every agent created before its swarm knew to mint
one.

The harness resolves the path and reports which credential this agent
can present. It does not yet present it: the auth-callout responder
still verifies only the hive-scoped token, and an agent offering a
credential nothing on the other end reads back would simply be refused.
Teaching swarm-nats-auth to read the same path is the next slice.
2026-09-21 20:44:52 +02:00
atlas
5ec0ce90fd nix: make swarm.authelia.url non-nullable, trim its docs
Review response on #4620: not having SSO is not a supported
deployment, so the type should not permit it, and the docs paragraph
explaining why SSO is always present is redundant once the type says
so.

- swarm.authelia.url drops types.nullOr.
- Every consumer's null-arm is gone: two option defaults
  (swarm-controller's and swarm's own statusPublish.tokenEndpoint)
  that produced an empty/null placeholder when the URL was null now
  unconditionally compute the real derived URL. Five now-dead
  "assertion = ... != null" guards (swarm-authelia's bridge,
  swarm-grafana, swarm-otel, swarm-nats, hive-forge, hive-matrix) are
  removed as unreachable — in every case the same URL was already
  interpolated unconditionally a few lines below the guard.
- grafanaNoSso, the module-eval fixture whose sole purpose was
  exercising the now-unsupported no-IdP refusal, is removed along
  with its dedicated test case; swarm.authelia.url = null is a type
  error now, not a value that reaches that assertion.
- docs/swarm/services.md: cut the clause about setting the option to
  null and the sentence explaining why the URL is co-location-
  independent — both redundant now that the type enforces it.
2026-09-21 18:14:28 +02:00
atlas
4b6214305f nix: address the swarm IdP by its domain, not by who runs it
`swarm.authelia.url` defaulted to `https://<domain>` only when this host
ran the container, and to `null` otherwise — so the address a client is
given was a statement about co-location rather than about the swarm. A
swarm has one SSO provider; every hive addresses the same name and
resolution decides which address that reaches, exactly as
`swarm.otel.domain` already works.

The option stays nullable: "this swarm has no IdP" is still expressible,
it is just now something an operator states rather than something not
running the container produces. The Grafana fixture that exercised the
no-IdP refusal says it explicitly.

Closes #4536
2026-09-21 18:14:28 +02:00
atlas
10fb79efc9 nix: let nixpkgs own the store collector's Restart
`services.opentelemetry-collector` already defines
`serviceConfig.Restart = "always"` at the normal priority. The forwarder
added in #4537 defined `"on-failure"` beside it, and two definitions at
one priority are a conflict the module system refuses to resolve — so
`containers.swarm-bao` stopped evaluating at all (#4615).

Drop our definition rather than force a value over it, which is what the
sibling collector in swarm-otel.nix already does. The resolved value is
`"always"`, which is the one we want here: a forwarder that exits for any
reason, clean or not, has stopped shipping the store's journal.
`RestartSec` stays — it conflicts with nothing and is what keeps the
restarts during first-boot secret delivery under the default start-rate
limit.
2026-09-21 17:29:46 +02:00
atlas
aaedff20a9 agent-modules/mcp: give subagent daemon a longer default bash timeout
A subagent's backgrounded bash child is reaped along with the rest of
its process tree at turn end, which silently orphans anything still
running past the CLI's default 2-minute timeout — the run reports
normally but the log file is empty or truncated. Raise the default on
hive-subagent-daemon's own unit rather than in managed-settings, since
managed-settings is also read by the main agent's session and the
operator's ruling is explicit that the main agent's environment stays
as is.
2026-09-21 17:20:59 +02:00
atlas
bb0afcd256 nix: the store's own collector scrapes its metrics listener
bao's metrics were scraped by the SWARM collector over loopback, via a
`swarm.otel.scrapeTargets.bao` entry gated on `deploy.swarm-otel.enable`
— "does the swarm's collector run on THIS host". It had to be: loopback
only reaches a reader that landed on the same host.

What that rendered everywhere else was nothing at all. Off that host the
metrics listener was not emitted, so the store's metrics reached the
store nowhere, and a host with no entry is indistinguishable from a host
nobody asked to scrape.

Moves the scrape into the collector this container already runs, per
mara on #4537: "move the existing scraper to the local collector". The
container shares the host netns (privateNetwork = false), so the scrape
still dials 127.0.0.1 — the listener keeps its address, its
`metrics_only` narrowing and its loopback-only bind, and the API
listener's `tls_require_and_verify_client_cert` is untouched.

The listener and its `prometheus_retention_time` lose their gate: the
reader ships with the store now, so there is no host where the endpoint
has none. The metrics pipeline reuses the logs pipeline's `resource`
processor and `otlphttp` exporter, so both signals carry the same
`service.name` and leave by the one hop.

Logs are unaffected: `journaldUnits` and --link-journal=host stay until
every sibling swarm container has a collector of its own.

The module-eval absence arm "a store with no collector beside it serves
no metrics" is inverted rather than dropped — the condition it asserted
is the bug. Three cases join it: the job is in swarm-bao AND gone from
swarm-otel (a move, not a copy), the scrape target and listener are both
pinned to loopback, and the metrics pipeline shares its exporter with
the logs one.
2026-09-21 17:19:52 +02:00
atlas
d4313fc34d nix: the store's journal forwarder has no gate to have
Both earlier versions asked the wrong host. `hyperhive.otel.enable` asked
whether this host runs a HIVE collector; `deploy.swarm-otel.enable` asked
whether this host runs the SWARM one. Neither answers the question the
forwarder actually has — "is there a collector to forward to" — and that
question cannot be false: a swarm always runs at least one instance of
every swarm-level service. So the forwarder renders under the condition
already enclosing it, that the store is deployed here, and nothing else.

`scrapeHere` deliberately keeps its `deploy.*` gate one line up. It is a
loopback metrics listener, which genuinely only works where the scraper
is — the two are different tiers, and the name says so.

The module-eval case that pins it is the split topology: the swarm
collector on another host, nothing local naming it, and the forwarder
still enabled and still addressed at `swarm.otel.domain`'s route. Both
removed gates render nothing in that fixture, which the co-located ones
they shipped with could not show.
2026-09-21 17:19:52 +02:00
atlas
ca8fc4ca64 nix: address swarm-bao's journal forwarder by swarm name
The forwarder pointed at the hive bridge address and was gated on the
hive's `otel.enable`, so it existed only where a hive collector stood
beside it. It now exports to `swarm.otel.domain` — the gateway-served
name that resolves locally when co-located and over the network
otherwise — on the swarm tier's own producer route, and is gated on
`deploy.swarm-otel.enable` like its sibling `scrapeHere`.

Refs #4526
2026-09-21 17:19:52 +02:00
atlas
e5224a6725 nix: give swarm-bao its own otel collector
Every container is supposed to run a collector that passes its logs and
metrics to the next hop. swarm-bao did not: its journal reached the store
only because `--link-journal=host` puts it in the host tree, where the
swarm collector — a different container — reads it through a unit
allowlist. That is the topology being retired, and in this deployment it
delivers nothing: no `_SYSTEMD_UNIT` value in the seven-day store mentions
openbao at all.

So the store's container now runs its own journal forwarder, copied from
an agent container's (nix/agent-modules/otel.nix): the whole journal, no
unit allowlist, pushed to the same first hop every agent on the host
already exports to. A local collector reads the local journal, so there is
nothing for a list of unit names to disagree with.

The `swarm.otel.journaldUnits` entry and `--link-journal=host` both stay.
Every sibling swarm container still rides the shared collector, and they
come out once each of them has a forwarder of its own.

Closes #4526
2026-09-21 17:19:52 +02:00
iris
c9ba906284 swarm-grafana: replace busiest-agents bargauges with an actual table
mara: 'i asked for a table. ask when doing something different.' Right
call — the PR body flagged the bargauge substitution as an open question,
not a decision, and I should have waited for an answer instead of
treating the silent absence of an objection as one.

Single 'Busiest agents' table panel (replaces the 5 per-metric bargauges):
5 table-format instant queries (turns, input, output, cache-read, cost)
joined on the agent label (joinByField), renamed to the screenshot's own
column names via organize, sorted by turns descending via sortBy — same
column set and sort order as the attached /stats screenshot.

This is a genuinely novel schema shape for this repo: no table panel,
transformation, or field-override config exists anywhere else in
nix/host-modules/swarm-grafana/dashboards/*.json to verify the join/
organize/sortBy option shapes against. Structurally verified (valid
JSON, jq empty, unique panel ids, every panel referenced exactly once,
nix fmt clean, dashboard-description lint clean) but the join/rename
field-name mechanics (whether Grafana names the joined columns exactly
'Value #A'/'Value #B'/etc.) are built from general Grafana schema
knowledge, not a working local precedent — flagging that explicitly so
the actual render gets checked before merge.
2026-09-20 23:40:21 +02:00