Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/nix
Author SHA1 Message Date
atlas
72e9e1bb4a swarm-ui: build the forge link from swarm.forge.domain
The swarm-controller builds the Forge quick link from
services.hyperhive.swarm.forge.domain, replacing hive-forge/default.nix's
per-host entry, so all seven swarm-service links come from swarm-level
options.

Also drops the remaining references to the removed matrix GUI switch:
the HiveUrls / Urls / hive_urls docs, the hivectl.md `open` note and
the gateway.md vhost-map rows, which name `gatewayHost` instead. The
grafana, victoriametrics and victorialogs modules' comments no longer
mention a quick-link they do not define.

Refs #4885
2026-10-02 20:02:02 +02:00
atlas
6786d54e5a swarm-ui: link the secret store's web UI from swarm.bao.ui.domain
The swarm-controller adds a Bao quick link built from
services.hyperhive.swarm.bao.ui.domain, so the popover links the
store's browser UI whichever host runs bao.

Refs #4885
2026-10-02 20:02:02 +02:00
atlas
f60f8af33b matrix: serve the web client unconditionally; link it from the swarm domain
Removes services.hyperhive.deploy.matrix.gui.enable and its
swarm.matrix.gui.enable alias; both are mkRemovedOptionModule stubs. A
host running the homeserver serves fluffychat at gatewayHost's vhost,
and the hive's /matrix/ redirect follows the same condition.

The swarm-controller builds the Matrix quick link from
swarm.matrix.gatewayHost, replacing hive-matrix.nix's per-host entry.
HIVE_MATRIX_PUBLIC_URL is set on every hive with a gatewayHost, so
`hivectl open matrix` resolves off the homeserver's host too.

Drops HIVE_MATRIX_GUI_ENABLED and the dashboard's matrix_gui_enabled
field; nothing in the frontend reads it.

Refs #4885
2026-10-02 20:02:02 +02:00
atlas
8628e0ecdd swarm-ui: build authelia/grafana/metrics/logs links from swarm domains
The swarm-controller module builds the Authelia, Grafana, Metrics and
Logs quick links from services.hyperhive.swarm.<service>.domain, on the
controller's host, instead of each service module adding its entry only
on the host that runs it. A controller whose swarm runs those services
on other hosts lists them in its /api/links popover.

Forge, matrix and bao links are not moved yet: forge waits on #4891,
matrix and bao on whether their GUI gate becomes swarm-level.

Refs #4885
2026-10-02 19:45:02 +02:00
atlas
ac592a5d23 forge: always behind the gateway; drop behindGateway
The forge always sits behind the gateway, so `deploy.forgejo.behindGateway`
(and its `swarm.forge.behindGateway` rename alias) is removed and its
true-branch behaviour is now unconditional within `deploy.forgejo.enable`:
https ROOT_URL on the gateway's httpsPort, the forge vhost and local DNS
name, the swarm-ui quick link, the published metrics scrape target, forgejo
metrics, the authelia `/metrics` rule, and `publicUrl` defaulting to
`https://<forge.domain>`.

Removed with it: the direct-port `http://<domain>:<httpPort>/` ROOT_URL
branch, the hive-ci assertion that the option is true, the core-toggle
cases that only exercised the false branch (the services-leaf case reads
`bare`, which never enabled the forge either). `hivectl open forge` now
points at `swarm.forge.publicUrl`, which can still be set to null.

Refs #4885
2026-10-02 18:50:42 +02:00
atlas
f8a8acae93 github swarm bao: address argus review on #4892
- docs/web-ui/README.md: drop the removed Credentials tile from the
  H0M3 hub list.
- api-error.ts, hive-warn.js: rewrite comments pointing at
  dashboard/src/credentials.js and credentials.html, now deleted, to
  state what the code does instead.
- swarm-secret-client/src/github.rs: correct the Credential.value doc
  to the actual read command (bao kv get -format=json | jq
  .data.data.value), keeping the load-bearing-field-name point.
- github-token.nix, agent-github-bao.nix, LinkGithubAccountForm.tsx:
  restate added comments as current behaviour instead of changelog
  wording ("has always had", "holds the token now").

Refs #4347
2026-10-02 17:54:45 +02:00
atlas
8e23feb01b github: PATs live in swarm bao; the agent fetches them itself
An operator links an agent's GitHub personal access token in the swarm UI
(LinkGithubAccountForm, "link github account" on /agents). swarm-controller's
PUT /api/hives/{hive}/agents/{agent}/github-account stores it at
swarm/agents/<agent>/github-token (swarm_secret_client::github), a flat leaf
under the agent's prefix that the agent's existing read grant already covers:
no policy change, and no list grant, since there is one token per agent.

In the agent, hive-agent-github-token (oneshot + 2-minute timer, as the agent
user, under its own store certificate, ordered before hive-github-notify)
reads that path and writes <state>/github-token, 0600 and agent-owned, the
file the gh wrapper, git credential helper and hive-github-notify already
read. It replaces the file by rename only when the bytes changed and never
deletes it: a hive-written github-token stays until a token is linked in the
swarm UI. It is installed only with a store address and
services.hyperhive.agent.github.enable.

Removed: the dashboard's CR3D3NTIALS page (credentials.html/js/css, its
build entries and H0M3 tile; GITHUB was its only tab), hive-c0re's
dashboard/matrix_accounts.rs with GET/POST /api/github-account,
priv_client::write_agent_github_token, the host socket's
SetAgentGithubToken and `hivectl github set-token`, and hive-priv's
WriteAgentGithubToken with write_agent_state_file, its only caller gone.

Docs: integrations/github.md and swarm/ui.md describe the swarm path,
swarm/credentials.md gains the store-path row, and the hive UI docs,
hivectl docs and security.md's hive-priv table drop the removed pieces.

Closes #4347
2026-10-02 17:48:27 +02:00
atlas
2b2608a491 docs(turn-loop): move harness systemd unit shape out of agent-roster.md
Moves the "Harness systemd unit shape" section from
docs/agent-lifecycle/agent-roster.md into docs/turn-loop/README.md: it
describes the per-agent harness systemd unit (env vars, PATH wiring,
serviceConfig), which is turn-loop material, not roster material.

Fixes two facts while moving: the ExecStart package is `hive-agent`,
not `hyperhive` (no package by that name exists); and `ruth.nix`
doesn't set any forge subscription default — it only defaults
`services.hyperhive.agent.docs.enable`.

Updates the inbound pointers in docs/turn-loop/config.md and the
module comment at nix/agent-modules/agent-service.nix.

Refs #3902
2026-10-02 14:47:31 +02:00
atlas
46f1f3cbdb docs: point matrix GUI comments at the swarm UI, drop the M4TR1X tab
The previous commit removed dashboard.md's M4TR1X section. These
comments and the `deploy.matrix.gui.enable` option description still
described a hive-dashboard M4TR1X tab or cited that section. They now
state what the option does: it serves the client on the gateway vhost
and adds the swarm UI's Matrix quick link (docs/swarm/ui.md::Quick links).

Refs #3902
2026-10-02 14:44:34 +02:00
atlas
fb2fff0668 fix(nix): require swarm domain only when a hyperhive service is enabled
The swarm.domain assertion in hive-network.nix fired on every host that
imported the module, so a host that enables nothing failed eval. It now
fires only when one of the hyperhive service switches is on (every
deploy.*.enable that runs something, gateway, gateway.dns, network,
otel, snapshotStore). The requirement itself is unchanged: any host that
runs a hyperhive service still needs swarm.domain.

The core-toggle module-eval suite gains a case: a missing swarm.domain is
refused on a hive and on a swarm-service-only host, and a host enabling
nothing passes every assertion.

Closes #4887
2026-10-02 13:09:45 +02:00
atlas
eded8f2e4e docs(agents): rename agent-hierarchy.md to agent-roster.md
Per mara's #4879 review (89815): the system has no agent hierarchy,
just a flat set scoped by the capability store, so the filename no
longer matched. The file already read "Agent roster & privileges"
after the earlier facts pass; rename it to match, and update the
four inbound references (docs/README.md, coordinator.md, config.md,
agent-service.nix).
2026-10-02 12:52:38 +02:00
atlas
99905f50b0 nix(authelia): start with a disabled placeholder user when the user set is empty
authelia 4.39.20 exits at startup on `users: {}` ("users: non zero value
required"), and the first-boot unit seeded exactly that, so a swarm with
no users crash-looped authelia and answered 502 until `swarmctl user add`
ran.

The first-boot unit now writes one subject, `swarm.placeholder`, when
the users database is absent, empty, or exactly `users: {}`:

- `disabled: true` — authelia returns "user not found" for a disabled
  user before any password check (file_user_provider.go,
  CheckUserPassword).
- password: an argon2id digest with an all-zero key. It decodes (authelia
  rejects a non-digest at startup) and no known password hashes to it.
- the `.` keeps it out of agent names (`[a-z0-9-]`), and `swarmctl user
  add` refuses it as already existing. Neither writer removes users, and
  both round-trip `disabled`.

A file with any user in it is never touched.

The docs that described the crash-loop (sso.md, gateway.md, setup.md,
the sso-unavailable error page) now describe the placeholder; the
writers' load_store docs and the seed fixtures follow. module-eval
nats-authelia asserts the seed branch.
2026-10-02 12:50:47 +02:00
atlas
bf81241744 nix: drop upgrade narration from swarm option docs 2026-10-02 11:41:56 +02:00
atlas
c8ff9a1943 nix: state the swarm requirement without a no-fallback clause 2026-10-02 11:41:56 +02:00
atlas
82463388c5 nix: require swarm.domain on every host
Swarm options are the same on all hosts, so the swarm.domain assertion is no longer gated on deploy.hive-controller.enable.

Refs #4872
2026-10-02 11:41:56 +02:00
atlas
2c7e586f47 forge: external forge accounts live in swarm bao; the agent fetches them itself
An operator now links an agent's external forge account (label, base URL,
token) in the swarm UI. swarm-controller stores it at
swarm/agents/<agent>/forge/<label>. There is no index: the store's
listing of the agent's forge/ directory is the set of accounts.

In the agent, hive-agent-forge-accounts (oneshot + 2-minute timer, as
the agent user, under its own store certificate) lists
swarm/agents/<agent>/forge/ with the `list` #4866 grants an agent on its
own metadata subtree, reads each account, and writes
<state>/forge-<label>-token and forge-<label>.json in the names and shape
hive-forge -f already reads. An empty listing (a 404, which `bao kv list
-format=json` answers with `{}` and an empty stderr) is zero accounts; a
denial or an unreachable store fails the unit. It never deletes: files
for labels not listed, including ones the hive wrote, stay as they are.

Removed: the dashboard FORGES tab (credentials.js/html section and its
CSS), hive-c0re's extra_forges.rs and its routes, priv_client's
extra-forge calls, and hive-priv's WriteAgentExtraForgeAccount /
DeleteAgentExtraForgeAccount with their helpers. The GITHUB tab and
WriteAgentGithubToken stay.

Also: persistence.md's matrix avatar note names the exit-75 restart on a
changed account listing, not the dashboard, as what brings a linked
account up.

Refs #4348
2026-10-01 18:05:33 +02:00
atlas
97fb76ce99 matrix: the agent's daemon pulls its linked accounts from bao itself
hive-matrix-daemon now learns which external matrix accounts it has from
the swarm secret store, under the agent's own certificate, and the hive
push chain for matrix is gone.

The daemon lists swarm/agents/<agent>/matrix/ (the `list` its policy
grants on its own metadata subtree), reads each account's homeserver
from its credential, and brings the accounts up with their tokens from
the store. Every two minutes it lists again and exits with 75 when the
set of linked accounts changed; the unit restarts on 75 without counting
a failure. A listed name whose credential reads as absent is skipped and
logged once. At start it removes the matrix-token-<a> /
matrix-account-<a>.json pairs a hive delivered (a sidecar marks a pair
as delivered; a declared tokenFile keeps its token).

Removed: CredentialNotice and the $SWARM.credential.* subject and NATS
grant, the controller's publish and its queue precondition on the PUT
route, hive-c0re's credential subscription arm and workers/credential.rs,
priv_client::write_agent_matrix_token, hive-priv's WriteAgentMatrixToken
and its helpers, and the daemon's state-dir account discovery.

Kept: WriteAgentGithubToken and the external-forge path
(WriteAgentExtraForgeAccount, extra_forges.rs) are untouched, and a
declared matrixAccounts tokenFile is still read when the store has no
token for that account.

Refs #4348
2026-10-01 17:43:28 +02:00
atlas
4e8225c058 nix: split hive-forge into service and deploy-mode files
`swarm.forge` (what the forge is to every hive: ports, domain, public and
root URLs, OIDC client id and callback) moves to
nix/host-modules/hive-forge/service.nix, together with the rename of the
old `services.hyperhive.forge` tree and the removed `swarm.forge.sso.enable`,
both of which name `swarm.forge` paths. Everything else -- the
`deploy.forgejo` options, the whole `config` block including
`containers.hive-forge`, and the helpers only they read -- stays in
nix/host-modules/hive-forge/default.nix, which now imports ./service.nix.
Importing it from the directory's own default.nix, as hive-c0re/ and
hive-gateway/ do with their option files, keeps the flake's standalone
`nixosModules.hive-forge` export whole.

Both halves read `cfg`, `gatewayCfg`, `swarmDomain` and `deployCfg`. They
are option reads, so each file binds them from `config`. `ssoSourceName`,
`defaultRootUrl` and `effectiveRootUrl` are not options and both halves
need them (the service half builds `sso.redirectUri` from them, the deploy
half registers the login source and sets ROOT_URL), so they are duplicated,
with a note at each copy. `ssoRedirectUri` is read only by the service
half and moves.

`swarm.forge.publicUrl` and `swarm.forge.sso.redirectUri` default from
`deploy.forgejo.behindGateway`; both move as they are.

A pure move: option paths, option definitions and config are unchanged
apart from comments: the two on either side of the cut, the
`ssoRedirectUri` comment and the duplication notes, and the rename
precedent in ./deploy.nix, which now names ./hive-forge/service.nix.

Refs #3742
2026-10-01 13:00:52 +02:00
atlas
a5eb3c15c4 ops: update option pointers after otel split
Four comments pointed at ./swarm-otel.nix for something the split moved
to ./swarm-otel-service.nix: `domain` (otel.nix), `domainBase`
(swarm-ui.nix), the `clientId`/`audience` options
(glue-swarm-otel-oidc-client.nix), and `producerName` (swarm-otel.nix's
own "Read-only option below"). Every other pointer to ./swarm-otel.nix
names its `config` block, units, exporters, authenticators or
assertions, which stayed.

Refs #3742
2026-10-01 13:00:40 +02:00
atlas
3a0a7346b1 nix: split swarm-otel into service and deploy-mode files
`swarm.otel` (what the swarm collector is to every hive: its domain,
receiver ports, producer names, client id and the audiences that client
may present) moves to nix/host-modules/swarm-otel-service.nix, together
with the `journaldUnits` removal module and the helpers its defaults
read. Everything else -- the `deploy.swarm-otel` options, the whole
`config` block including `containers.swarm-otel`, and the helpers only
they read -- stays in nix/host-modules/swarm-otel.nix, which default.nix
now imports after the new file.

The service file needs `domainBase` (for `domain`), `baoCfg` (for
`storeProducerName`) and `cfg` plus `pushAudiences` (for `audience`).
`swarmDomain` and `domainBase` move. `cfg`, `hyperhiveCfg`, `baoCfg`,
`vmCfg`, `vlCfg`, `metricsPushUrl`, `logsPushUrl` and `pushAudiences`
are read on both sides and none is an option, so each is bound in both
files; the comments on the push URLs say the two copies must agree.
`swarmCfg` was bound and never read, so neither file carries it.

A pure move: option paths, option definitions and config are unchanged.
Comments changed: the transition comment above `deploy.swarm-otel` now
names the file `swarm.otel` lives in, and the push-URL and
`pushAudiences` comments now describe the per-file readers and the
duplicate. Three otel fixtures evaluate to the same host and container
toplevel derivations, `swarm.otel.*` and `deploy.swarm-otel.*` values
before and after.

Refs #3742
2026-10-01 13:00:40 +02:00
atlas
fc8b8fa2f5 ops: update option pointer after hive-matrix split
nix/module-eval/bao-controller.nix:49 quoted `gatewayHost`'s description
as `hive-matrix.nix`'s own doc. The option now lives in
hive-matrix-service.nix, so the pointer names that file.

Refs #3742
2026-10-01 13:00:25 +02:00
atlas
a539dceab1 nix: split hive-matrix into service and deploy-mode files
`swarm.matrix` (what the homeserver is to every hive: server name, ports,
API URL, gateway host, encryption policy, OIDC client id) moves to
nix/host-modules/hive-matrix-service.nix. Everything else -- the
`imports` block with its two renames and the removed `sso.enable`, the
`deploy.matrix` options, the whole `config` block including
`containers.hive-matrix`, and every other let binding -- stays in
nix/host-modules/hive-matrix.nix, which default.nix now imports alongside
the new file.

The `swarm.matrix` block reads three let bindings, and the `config` block
reads all three too: `cfg` (`apiUrl` defaults from `cfg.httpPort`),
`swarmDomain` (`gatewayHost`'s default) and `deployCfg` (`apiUrl` reads
`deployCfg.matrix.enable`). All three are option reads, so each file binds
them from `config.services.hyperhive.*`. Nothing is duplicated.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above `deploy.matrix`, which now names the file
`swarm.matrix` lives in.

Refs #3742
2026-10-01 13:00:25 +02:00
atlas
4181d33cad ops: cross-reference authelia unit literals in both split files 2026-10-01 10:28:57 +02:00
atlas
ba56bfe32e nix: split swarm-authelia into service and deploy-mode files
`swarm.authelia` (what the SSO provider is to every hive: ports, domain,
`url`, the OIDC client register, the published names and the bridge's
address) moves to nix/host-modules/swarm-authelia-service.nix. Everything
else -- the `deploy.authelia` options, the whole `config` block including
`containers.swarm-authelia`, and the helpers only they read -- stays in
nix/host-modules/swarm-authelia.nix, which default.nix now imports
alongside the new file.

Both halves read four `let` bindings. `cfg`, `swarmDomain` and `deployCfg`
are option reads, so each file binds them from `config`; the service file
has no `hyperhiveCfg`, so it spells the paths out, as
swarm-nats-service.nix does. `instance` and `unitName` are literals, not
options, so the service file carries its own copy of the two (`unit`'s
default reads `unitName`). `deployCfg` is in the service file only for
`bridgeUrl`'s default, which is moved as it is.

`hyperhiveDomain` had no reader and is dropped rather than carried into
either file.

A pure move: option paths, option definitions and config are unchanged
apart from three comments that pointed "above"/"below" across the new
file boundary and now name the file. Authelia's container toplevel, the
host toplevel (with `c0re.hyperhiveFlake` pinned, since the flake source
path lands in /etc/hyperhive/serve.json), the `swarm.authelia` and
`deploy.authelia` values and option set, and the eleven
module-eval checks that enable authelia evaluate to the same derivations
before and after.

Refs #3742
2026-10-01 10:25:19 +02:00
atlas
eef7b70c0e nix: split swarm-victorialogs into service and deploy-mode files
`swarm.victorialogs` (what the log store is to every hive: container name,
domain, port) moves to nix/host-modules/swarm-victorialogs-service.nix,
together with the only two helpers it reads, `swarmDomain` and
`domainBase`. Everything else -- the `deploy.victorialogs` options, the
whole `config` block including `containers.swarm-victorialogs`, the file
header and the helpers only they read (`swarmAuthRequest` among them) --
stays in nix/host-modules/swarm-victorialogs.nix, which default.nix now
imports alongside the new file.

`hyperhiveCfg` (an alias for `config.services.hyperhive`, not an option) is
read by both halves, so it is duplicated into the service file rather than
shared.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above the `deploy.victorialogs` options, which now
names the file `swarm.victorialogs` lives in. Fixtures enabling the store
evaluate to the same host and container toplevel derivations before and
after.

Refs #3742
2026-10-01 10:03:55 +02:00
atlas
f18d8099f5 nix: split swarm-victoriametrics into service and deploy-mode files
`swarm.victoriametrics` (what the metrics store is to every hive: container
name, domain, port) moves to nix/host-modules/swarm-victoriametrics-service.nix,
together with the only two helpers it reads, `swarmDomain` and `domainBase`.
Everything else -- the `deploy.victoriametrics` options, the whole `config`
block including `containers.swarm-victoriametrics`, the file header and the
helpers only they read -- stays in nix/host-modules/swarm-victoriametrics.nix,
which default.nix now imports alongside the new file.

`hyperhiveCfg` (an alias for `config.services.hyperhive`, not an option) is
read by both halves, so it is duplicated into the service file rather than
shared.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above the `deploy.victoriametrics` options, which now
names the file `swarm.victoriametrics` lives in. Fixtures enabling the store
evaluate to the same host and container toplevel derivations before and
after.

Refs #3742
2026-10-01 10:03:55 +02:00
atlas
9a815e9658 ops: update option pointers after grafana/bao split
glue-swarm-bao-otel-oidc-client.nix:34 still pointed clientId's
declaration at ./swarm-bao.nix after the split moved it to
./swarm-bao-service.nix. Line 17, which points the config block's
deploy.bao.enable gate at ./swarm-bao.nix, is unchanged -- that part
stayed.
2026-10-01 09:53:04 +02:00
atlas
3bfba1925c nix: split swarm-bao into service and deploy-mode files
`swarm.bao` (what the secret store is to every hive: container name,
domain, UI domain and OIDC client, port, collector client id and
telemetry port) moves to nix/host-modules/swarm-bao-service.nix, together
with `domainBase`, the only helper it reads besides `cfg`. Everything
else -- the `deploy.bao` options, the removed-option import, the whole
`config` block including `containers.swarm-bao`, and the helpers only
they read -- stays in nix/host-modules/swarm-bao.nix, which default.nix
now imports alongside the new file.

Both halves read `cfg` (`swarm.bao.ui.oidc.redirectUri` defaults from
`cfg.ui.domain`; the config block reads `cfg` throughout). It is an
option read, so each file binds it from `config.services.hyperhive.swarm.bao`.
The service file has no `hyperhiveCfg`, so its `swarmDomain` reads
`config.services.hyperhive.swarm.domain` directly, as
swarm-nats-service.nix does.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above `deploy.bao`, which now names the file
`swarm.bao` lives in, and the pointer in swarm-nats-service.nix to the
`domainBase` rationale, which moved with it.

Refs #3742
2026-10-01 09:37:36 +02:00
atlas
eed53a2b59 nix: split swarm-grafana into service and deploy-mode files
`swarm.grafana` (what the metrics UI is to every hive: container name,
domain, metrics port, OIDC client) moves to
nix/host-modules/swarm-grafana-service.nix, together with `domainBase`,
the only helper it reads besides `cfg`. Everything else -- the
`deploy.grafana` options, the whole `config` block including
`containers.swarm-grafana`, and the helpers only they read -- stays in
nix/host-modules/swarm-grafana.nix, which default.nix now imports
alongside the new file.

Both halves read `cfg` (`swarm.grafana.oidc.redirectUri` defaults from
`cfg.domain`; the config block reads `cfg` throughout). It is an option
read, so each file binds it from `config.services.hyperhive.swarm.grafana`.
The service file has no `hyperhiveCfg`, so its `swarmDomain` reads
`config.services.hyperhive.swarm.domain` directly, as
swarm-nats-service.nix does.

A pure move: option paths, option definitions and config are unchanged
apart from the two comments on either side of the cut, which now name the
file the other half lives in.

Refs #3742
2026-10-01 09:30:21 +02:00
atlas
07ca06dcca nix: split swarm-nats into service and deploy-mode files
`swarm.nats` (what the queue is to every hive: domain, ports, client id)
moves to nix/host-modules/swarm-nats-service.nix, together with the only
two helpers it reads, `swarmDomain` and `domainBase`. Everything else --
the `deploy.nats` options, the whole `config` block including
`containers.swarm-nats`, and the helpers only they read -- stays in
nix/host-modules/swarm-nats.nix, which default.nix now imports alongside
the new file.

A pure move: option paths, option definitions and config are unchanged
apart from the transition comment above `deploy.nats`, which now names the
file `swarm.nats` lives in. The nats fixtures evaluate to the same host and
container toplevel derivations before and after.

Refs #3742
2026-10-01 09:04:35 +02:00
atlas
6b1e825c0a swarm-otel: ship the whole host journal, drop user sessions after it
The swarm collector's journald receiver read only the units listed in
`services.hyperhive.swarm.otel.journaldUnits`. A unit nobody listed
never reached the store, and a misspelt entry shipped nothing without
an error. The list existed to keep an operator's desktop session out of
a store every swarm operator can read, but the receiver can only match
positively, so the only way to express "not user sessions" was to name
every service instead.

The receiver now reads the whole host journal, and a new
`filter/exclude-user-sessions` processor in the `logs/<swarm>` pipeline
drops records whose `_SYSTEMD_SLICE` is `user-<uid>.slice` (session
scopes and `user@<uid>.service`). The per-hive `logs/<hive>` pipelines
carry agent-container journals only and get no filter.

`journaldUnits` is removed with `mkRemovedOptionModule`, together with
its non-empty assertion and the entry each host module added. The four
module-eval membership checks go with it, replaced by one structural
case in swarm-otel-core.

Closes #3646
2026-09-30 23:01:49 +02:00
atlas
c5b21403a6 hive-runtime: read the ACP provider key from bao
An opencode ACP agent got its provider API key only from the hand-placed
backendEnvironmentFile. It now also reads it from the swarm secret store
at swarm/agents/<agent>/acp-provider, field api_key, under its own
certificate, and sets it in the spawned ACP agent's environment only.
Nothing is written to disk.

Precedence: a value already in the process environment (the env file)
wins and the store is not asked. Otherwise the stored key is used when
present. With no store, nothing stored, or a failed read, the agent is
spawned without the key as before, and one line is logged without the
value.

The variable name comes from the existing per-agent option
acp.opencode.provider.apiKeyEnv, exported as HIVE_ACP_API_KEY_ENV on the
harness only for the opencode preset. Other ACP commands are unchanged.

The read lives in hive-runtime, where the ACP child is spawned, so both
hive-agent and hive-subagent-daemon use it. The subagent daemon unit
gets the key name and, when the agent has a store, the agent's store
identity (the same credentials queue-identity.nix gives the harness).

No new option or setting. Closes #4841.
2026-09-30 22:55:03 +02:00
atlas
c2bdf30e05 hive-agent: export ACP-reported cost and context fill over OTLP
An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.

hive-agent's existing turn-metrics meter records three new instruments:

- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
  every backend: the two numbers the web UI's ctx% divides.

The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.

Refs #4845
2026-09-30 22:24:06 +02:00
atlas
b0d92ccbd8 fix(forge): pass avatar image via files, not argv (E2BIG over 128 KiB)
forge-avatar-sync passed the base64-encoded icon as a jq --arg and then
as a curl -d command-line argument. Linux caps a single exec argument at
MAX_ARG_STRLEN (128 KiB), so any icon whose rasterized PNG base64-encodes
past that (red's does, at 133300 bytes) makes jq fail with E2BIG before
curl is ever reached, and the avatar upload silently never happens. The
base64 and the JSON payload now go through temp files instead of argv.

Closes #4839
2026-09-30 19:17:59 +02:00
atlas
b9667d5975 swarm-grafana: alert rules for queue, renewal and missing-log WARNs
Adds rules to the `hyperhive` alerting folder next to the two from #4811:

- swarm queue connect failed (per hive/agent, 1h window, fires at once:
  the line is logged once per agent process and the connection is never
  retried, so the window is how long the rule stays firing)
- swarm agent state publish skipped (per hive/agent, 15m, for 5m, as the
  swarm terminal rule)
- swarm queue credential missing (per hive/agent, 15m, for 5m)
- agent credential renewal failing (10m, for 15m, as the forge reconcile
  rule: same 5m pass cadence, one line per failed pass)
- no log records from <hive>, one per `swarm.hives` entry: fires when
  the hive's ungrouped count over 10m is below 1, and on no data. A new
  `logAbsenceRule` helper wraps `logCountRule` with the inverted
  threshold and noDataState = Alerting.

No contact point and no notification policy: the rules show under
Alerting -> Alert rules and are delivered nowhere.

Refs #4717
Refs #3900
2026-09-30 15:00:28 +02:00
iris
d74f567c59 nix: runtime option description: ACP cancel, compact and model/effort now work 2026-09-30 13:12:49 +02:00
atlas
8a8da5ec8a hive-agent: ACP model picker honours availableModels (#4391)
Filter the ACP model picker's list by services.hyperhive.agent.availableModels: an
unconfigured agent (env absent) shows every model the session offers, a
configured list narrows the picker to whatever it names that the session
also offers (in the session's own order), and a configured list matching
none of the session's models (the claude names on an ACP agent that never
touched the option) falls back to showing everything, with one warning
naming the mismatch.

The nix option's default, the model-vs-availableModels build assertion and
hive-subagent-mcp's check_model rail are unchanged — this only touches the
web UI's picker.
2026-09-30 10:53:53 +02:00
atlas
b68fd7306e refresh-consumer: key the restart on the file's mtime, not a pre-write compare
The restart decision was a shell variable set by comparing the fetched
value with the file just before overwriting it. A run that wrote the
file and then failed before the restart (the matrix unit's registration
render, or `systemctl --machine` finding no bus yet) left a retry that
saw an unchanged file and never restarted the consumer.

The file is now written only when the value differs, so its mtime marks
the last real change, and `refresh_consumer <machine> <unit> <path>`
compares that mtime with the consumer's ActiveEnterTimestamp on every
run, the shape the openbao client-CA refresh in swarm-bao.nix already
uses. A consumer that started after the last change is left alone; a
running one is try-restarted, a failed one reset and started, all with
--no-block, and nothing happens while the container is down.

The helper's comment block also exceeded the 30-line limit
(`comment-block lint` failed on d871467d); its per-function notes now
sit beside the functions.

module-eval-bao-grants asserts the gated write, the path the refresh is
keyed on, and the mtime-vs-start comparison for each consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
b3b42d3279 credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
3db1233da0 bao: mkRemovedOptionModule for matrixCtlHiveName
A plain removal breaks any out-of-tree host config that still sets the
option: eval fails with "option does not exist" and no pointer to what
replaced it. mkRemovedOptionModule gives a clear evaluation error instead.
2026-09-30 07:42:22 +02:00
atlas
b58a0d8ba9 bao: drop matrix-ctl's per-hive sender-token grant
swarm-controller is the only minter of swarm/hives/<hive>/matrix/sender-token
since #4820, so the swarm-matrix-ctl policy's first stanza granted a write no
code performs. The policy keeps its one used stanza, the swarm appservice token
that `swarm-matrix-ctl appservice publish` writes. deploy.bao.matrixCtlHiveName
only named the hive in the dropped stanza and goes with it.

hive-matrix.nix no longer calls the per-hive appservice's as_token hive-c0re's
authority: tuwunel loads the registration and creates the sender account, and
no client presents that token.
2026-09-30 07:42:22 +02:00
atlas
84d4808d54 hive-subagent-mcp: run an agent's subagents on its runtime
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.

The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.

Refs #4391
2026-09-30 07:41:12 +02:00
atlas
ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00
atlas
1b24edf4b4 nix: runtime option, acp.* and an opencode preset
services.hyperhive.agent.runtime ("claude" default | "acp") and
acp.{command,args,env}, rendered into HIVE_RUNTIME / HIVE_ACP_* only
for acp, so a claude agent's unit is unchanged. acp implies useApiKey.

acp.presets.opencode runs `opencode acp` from nixpkgs against an
OpenAI-compatible provider from acp.opencode.{provider,model,
contextWindow,outputLimit}: the config is rendered to the store with
the API key as an {env:VAR} reference, so the key is read at runtime
from backendEnvironmentFile. OPENCODE_PERMISSION denies opencode's
built-in bash, task, todowrite and websearch, and makes webfetch ask.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
78d8d69c7f swarm-controller: mint each hive's matrix sender token
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.

swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.

swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.

hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.

The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
41d66e2e05 swarm-grafana: provision alert rules for forge reconcile + swarm terminal publish failures
Two Grafana-managed rules in a "hyperhive" folder, one group evaluated
every 1m, each a LogsQL stats count over the VictoriaLogs datasource:

- forge reconcile failing: `swarm forge objects: write failed; retrying
  next pass` over 10m, for 15m. The pass runs every 5m, so one failed
  pass stays under `for` and a pass failing on every tick fires.
- swarm terminal publish skipped: `swarm terminal: publish skipped` over
  15m, for 5m, one instance per hive/agent.

No contact point or notification policy. Grafana 13 routes to its
built-in `empty` receiver when none is provisioned, so firing rules
are visible under Alerting and sent nowhere, with no send errors.

Refs #4717
2026-09-29 22:12:28 +02:00
atlas
10d579ecaf nix: move the remaining service containers onto the swarm-container module
hive-ci, hive-forge, hive-matrix, swarm-authelia, swarm-bao, swarm-grafana,
swarm-nats, swarm-otel and swarm-victorialogs now import
./swarm-container.nix and drop their own copies of the stateVersion,
firewall and resolvconf lines. Each binds privateNetwork once in its
top-level let and passes it to both the host attr and the in-container
option, as swarm-victoriametrics already does.

hive-forge (25.11) and swarm-otel (the host's value) keep their own
stateVersion over the module's mkDefault. hive-ci sets privateNetwork =
true and writesOwnResolvConf = false, which leaves its firewall and
resolvconf on, as before. hive-matrix keeps its useHostResolvConf
override and static resolv.conf; its resolvconf mkForce now comes from
the module default.

Every container's system.build.toplevel drvPath and host-side attrs
evaluate identical to the parent commit.

module-eval-swarm-services-switch gains a fixture with all ten service
containers and checks that each one's in-container privateNetwork equals
its host-side value, that the nine on the host netns run no firewall or
resolvconf, that hive-ci keeps both, and that hive-forge keeps its pinned
stateVersion.

Refs #3773
2026-09-29 20:17:28 +02:00
atlas
cce41c1d79 nix: move the in-container modules to nix/container-modules/
Refs #3773
2026-09-29 19:51:46 +02:00
atlas
024067f3f8 nix: share the service-container settings through one in-container module
The ten hand-rolled `containers.<name>` blocks each repeat the same
in-container lines: `system.stateVersion`, a firewall turned off because
the container shares the host netns, and resolvconf forced off because
something in the container writes /etc/resolv.conf itself.

`nix/host-modules/swarm-container.nix` now owns those lines. It is
imported inside the container's own config and exposes
`services.hyperhive.swarmContainer.{privateNetwork,writesOwnResolvConf}`
for the host module to set. `stateVersion` is a `mkDefault`, so the two
containers on another value can keep theirs. `--link-journal=host` stays
per module, and so do the host-side attrs (autoStart, ephemeral,
privateNetwork, bindMounts).

swarm-victoriametrics is converted as the first user. Its container
toplevel drvPath is unchanged. A module-eval case now forces that
container's config, which nothing in the suite read before.

Refs #3773
2026-09-29 19:51:46 +02:00