Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph

5,003 commits

Author SHA1 Message Date
atlas
eef7b70c0e nix: split swarm-victorialogs into service and deploy-mode files
`swarm.victorialogs` (what the log store is to every hive: container name,
domain, port) moves to nix/host-modules/swarm-victorialogs-service.nix,
together with the only two helpers it reads, `swarmDomain` and
`domainBase`. Everything else -- the `deploy.victorialogs` options, the
whole `config` block including `containers.swarm-victorialogs`, the file
header and the helpers only they read (`swarmAuthRequest` among them) --
stays in nix/host-modules/swarm-victorialogs.nix, which default.nix now
imports alongside the new file.

`hyperhiveCfg` (an alias for `config.services.hyperhive`, not an option) is
read by both halves, so it is duplicated into the service file rather than
shared.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above the `deploy.victorialogs` options, which now
names the file `swarm.victorialogs` lives in. Fixtures enabling the store
evaluate to the same host and container toplevel derivations before and
after.

Refs #3742
2026-10-01 10:03:55 +02:00
atlas
f18d8099f5 nix: split swarm-victoriametrics into service and deploy-mode files
`swarm.victoriametrics` (what the metrics store is to every hive: container
name, domain, port) moves to nix/host-modules/swarm-victoriametrics-service.nix,
together with the only two helpers it reads, `swarmDomain` and `domainBase`.
Everything else -- the `deploy.victoriametrics` options, the whole `config`
block including `containers.swarm-victoriametrics`, the file header and the
helpers only they read -- stays in nix/host-modules/swarm-victoriametrics.nix,
which default.nix now imports alongside the new file.

`hyperhiveCfg` (an alias for `config.services.hyperhive`, not an option) is
read by both halves, so it is duplicated into the service file rather than
shared.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above the `deploy.victoriametrics` options, which now
names the file `swarm.victoriametrics` lives in. Fixtures enabling the store
evaluate to the same host and container toplevel derivations before and
after.

Refs #3742
2026-10-01 10:03:55 +02:00
atlas
9a815e9658 ops: update option pointers after grafana/bao split
glue-swarm-bao-otel-oidc-client.nix:34 still pointed clientId's
declaration at ./swarm-bao.nix after the split moved it to
./swarm-bao-service.nix. Line 17, which points the config block's
deploy.bao.enable gate at ./swarm-bao.nix, is unchanged -- that part
stayed.
2026-10-01 09:53:04 +02:00
atlas
3bfba1925c nix: split swarm-bao into service and deploy-mode files
`swarm.bao` (what the secret store is to every hive: container name,
domain, UI domain and OIDC client, port, collector client id and
telemetry port) moves to nix/host-modules/swarm-bao-service.nix, together
with `domainBase`, the only helper it reads besides `cfg`. Everything
else -- the `deploy.bao` options, the removed-option import, the whole
`config` block including `containers.swarm-bao`, and the helpers only
they read -- stays in nix/host-modules/swarm-bao.nix, which default.nix
now imports alongside the new file.

Both halves read `cfg` (`swarm.bao.ui.oidc.redirectUri` defaults from
`cfg.ui.domain`; the config block reads `cfg` throughout). It is an
option read, so each file binds it from `config.services.hyperhive.swarm.bao`.
The service file has no `hyperhiveCfg`, so its `swarmDomain` reads
`config.services.hyperhive.swarm.domain` directly, as
swarm-nats-service.nix does.

A pure move: option paths, option definitions and config are unchanged
apart from the comment above `deploy.bao`, which now names the file
`swarm.bao` lives in, and the pointer in swarm-nats-service.nix to the
`domainBase` rationale, which moved with it.

Refs #3742
2026-10-01 09:37:36 +02:00
atlas
eed53a2b59 nix: split swarm-grafana into service and deploy-mode files
`swarm.grafana` (what the metrics UI is to every hive: container name,
domain, metrics port, OIDC client) moves to
nix/host-modules/swarm-grafana-service.nix, together with `domainBase`,
the only helper it reads besides `cfg`. Everything else -- the
`deploy.grafana` options, the whole `config` block including
`containers.swarm-grafana`, and the helpers only they read -- stays in
nix/host-modules/swarm-grafana.nix, which default.nix now imports
alongside the new file.

Both halves read `cfg` (`swarm.grafana.oidc.redirectUri` defaults from
`cfg.domain`; the config block reads `cfg` throughout). It is an option
read, so each file binds it from `config.services.hyperhive.swarm.grafana`.
The service file has no `hyperhiveCfg`, so its `swarmDomain` reads
`config.services.hyperhive.swarm.domain` directly, as
swarm-nats-service.nix does.

A pure move: option paths, option definitions and config are unchanged
apart from the two comments on either side of the cut, which now name the
file the other half lives in.

Refs #3742
2026-10-01 09:30:21 +02:00
atlas
07ca06dcca nix: split swarm-nats into service and deploy-mode files
`swarm.nats` (what the queue is to every hive: domain, ports, client id)
moves to nix/host-modules/swarm-nats-service.nix, together with the only
two helpers it reads, `swarmDomain` and `domainBase`. Everything else --
the `deploy.nats` options, the whole `config` block including
`containers.swarm-nats`, and the helpers only they read -- stays in
nix/host-modules/swarm-nats.nix, which default.nix now imports alongside
the new file.

A pure move: option paths, option definitions and config are unchanged
apart from the transition comment above `deploy.nats`, which now names the
file `swarm.nats` lives in. The nats fixtures evaluate to the same host and
container toplevel derivations before and after.

Refs #3742
2026-10-01 09:04:35 +02:00
atlas
7d217f8267 remove the create_repo agent tool
mara ruled on #4849 (c88934): "remove create_repo tool". The tool ran in
hive-c0re with the hive's core token, so it only ever worked for agents
on the hive that runs the forge.

Removed:

- the create_repo MCP tool and CreateRepoArgs (hive-agent-mcp)
- wire variants Request::CreateRepo and Response::RepoCreated
  (hive-core-agent-sock)
- hive-c0re's handle_create_repo, its valid_repo_name check and the
  dispatch arm
- forge::create_agent_repo and apply_operator_branch_protection, which
  had no other caller, plus AGENTS_ORG and OPERATORS_TEAM, whose only
  users they were
- the tool's docs (docs/tools/forge.md repo management, docs/turn-loop/
  mcp.md, the conventions tool-group table) and the doc comments that
  named it (hive-sock-client's response timeout, ensure_repo_creation_
  disabled, the security doc's merge-gate bullet)

ToolGroup::Forge is kept with no tools, the same way b88a5b24 kept
Lifecycle, so existing meta/capabilities.json grants still parse.

Forge state is untouched: existing agents/* repos keep their collaborators
and operators-team branch protection. The swarm-controller's own
create_repo (config-org repos) is a different path and is unchanged.

Closes #4849
2026-10-01 09:04:17 +02:00
flake-bot
bab15e2ba6 nix flake update 2026-10-01 05:01:18 +02:00
atlas
6b1e825c0a swarm-otel: ship the whole host journal, drop user sessions after it
The swarm collector's journald receiver read only the units listed in
`services.hyperhive.swarm.otel.journaldUnits`. A unit nobody listed
never reached the store, and a misspelt entry shipped nothing without
an error. The list existed to keep an operator's desktop session out of
a store every swarm operator can read, but the receiver can only match
positively, so the only way to express "not user sessions" was to name
every service instead.

The receiver now reads the whole host journal, and a new
`filter/exclude-user-sessions` processor in the `logs/<swarm>` pipeline
drops records whose `_SYSTEMD_SLICE` is `user-<uid>.slice` (session
scopes and `user@<uid>.service`). The per-hive `logs/<hive>` pipelines
carry agent-container journals only and get no filter.

`journaldUnits` is removed with `mkRemovedOptionModule`, together with
its non-empty assertion and the entry each host module added. The four
module-eval membership checks go with it, replaced by one structural
case in swarm-otel-core.

Closes #3646
2026-09-30 23:01:49 +02:00
atlas
710f5b7b8c hive-runtime: re-export ACP_API_KEY_ENV_ENV
docs-rustdoc failed: the doc link on AcpCommand::api_key_env pointed at
ACP_API_KEY_ENV_ENV, which wasn't re-exported from lib.rs like its
sibling *_ENV consts, so the link resolved to a private item.
2026-09-30 22:55:03 +02:00
atlas
c5b21403a6 hive-runtime: read the ACP provider key from bao
An opencode ACP agent got its provider API key only from the hand-placed
backendEnvironmentFile. It now also reads it from the swarm secret store
at swarm/agents/<agent>/acp-provider, field api_key, under its own
certificate, and sets it in the spawned ACP agent's environment only.
Nothing is written to disk.

Precedence: a value already in the process environment (the env file)
wins and the store is not asked. Otherwise the stored key is used when
present. With no store, nothing stored, or a failed read, the agent is
spawned without the key as before, and one line is logged without the
value.

The variable name comes from the existing per-agent option
acp.opencode.provider.apiKeyEnv, exported as HIVE_ACP_API_KEY_ENV on the
harness only for the opencode preset. Other ACP commands are unchanged.

The read lives in hive-runtime, where the ACP child is spawned, so both
hive-agent and hive-subagent-daemon use it. The subagent daemon unit
gets the key name and, when the agent has a store, the agent's store
identity (the same credentials queue-identity.nix gives the harness).

No new option or setting. Closes #4841.
2026-09-30 22:55:03 +02:00
atlas
c2bdf30e05 hive-agent: export ACP-reported cost and context fill over OTLP
An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.

hive-agent's existing turn-metrics meter records three new instruments:

- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
  every backend: the two numbers the web UI's ctx% divides.

The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.

Refs #4845
2026-09-30 22:24:06 +02:00
atlas
5d7042e655 hive-agent: warn on usage responses with no recognisable windows
addresses argus's review on #4842 (c88740): fetch's success path was
silent when the response parsed but windows() found nothing to record.
should_warn_empty tracks the same once-per-streak shape as the
existing token-skip logging, reset the moment a response has windows
again, and never logs the body.

Also notes in the observability doc that a past resets_at means the
paired percent gauge is stale.
2026-09-30 20:53:03 +02:00
atlas
8a735ddcb3 hive-agent: publish Claude subscription usage (5h/7d %) as metrics
A detached task polls GET https://api.anthropic.com/api/oauth/usage every
5 minutes with the OAuth access token from ~/.claude/.credentials.json and
records, per window the response names (five_hour, seven_day,
seven_day_sonnet, ...):

- hyperhive.agent.claude_usage.percent   (%, 0-100)
- hyperhive.agent.claude_usage.resets_at (s, unix seconds)

both labelled window=<name>. Endpoint, the anthropic-beta:
oauth-2025-04-20 header and the {utilization, resets_at} window shape are
taken from the claude-code 2.1.283 bundle's own /usage fetch.

The token is only read, never refreshed: claude owns refresh-token
rotation and a second refresher can log the agent out. An expired token
is skipped until claude's next turn refreshes it. API-key agents
(HIVE_USE_API_KEY, the ACP default) and agents with no credentials file
skip quietly, and the task does nothing when OTEL is not configured.
Request failures and non-2xx statuses warn with the status or transport
error only, never the body or the token.
2026-09-30 20:53:03 +02:00
atlas
b0d92ccbd8 fix(forge): pass avatar image via files, not argv (E2BIG over 128 KiB)
forge-avatar-sync passed the base64-encoded icon as a jq --arg and then
as a curl -d command-line argument. Linux caps a single exec argument at
MAX_ARG_STRLEN (128 KiB), so any icon whose rasterized PNG base64-encodes
past that (red's does, at 133300 bytes) makes jq fail with E2BIG before
curl is ever reached, and the avatar upload silently never happens. The
base64 and the JSON payload now go through temp files instead of argv.

Closes #4839
2026-09-30 19:17:59 +02:00
atlas
9e06e191a3 swarm-controller: first credential-renewal pass waits for the queue connection
The renewal pass ran the moment swarm-controller started, before its
queue client had connected, so `WantedWriter::view` refused with
"client state Pending" and the pass logged
`agent credential renewal: pass failed; retrying next tick`. That fired
twice in 24h on muede-lpt2, both under a second after start, and would
trip a Grafana rule on that WARN on ordinary restarts.

The first pass now waits until the queue client is connected, polling
`swarm_queue_client::ensure_connected` every 5s the way
`AgentIconReader::create_when_connected` does. The wait is bounded by
one RECONCILE_INTERVAL (5 min): the old code's retry after a false start
also came one interval later, so a queue that never connects gets its
first pass no later than before. Hitting the bound logs one WARN and runs
the pass anyway. Only startup waits; a later disconnect still fails a
pass and logs the WARN.

Refs #4717
2026-09-30 15:46:56 +02:00
atlas
b9667d5975 swarm-grafana: alert rules for queue, renewal and missing-log WARNs
Adds rules to the `hyperhive` alerting folder next to the two from #4811:

- swarm queue connect failed (per hive/agent, 1h window, fires at once:
  the line is logged once per agent process and the connection is never
  retried, so the window is how long the rule stays firing)
- swarm agent state publish skipped (per hive/agent, 15m, for 5m, as the
  swarm terminal rule)
- swarm queue credential missing (per hive/agent, 15m, for 5m)
- agent credential renewal failing (10m, for 15m, as the forge reconcile
  rule: same 5m pass cadence, one line per failed pass)
- no log records from <hive>, one per `swarm.hives` entry: fires when
  the hive's ungrouped count over 10m is below 1, and on no data. A new
  `logAbsenceRule` helper wraps `logCountRule` with the inverted
  threshold and noDataState = Alerting.

No contact point and no notification policy: the rules show under
Alerting -> Alert rules and are delivered nowhere.

Refs #4717
Refs #3900
2026-09-30 15:00:28 +02:00
atlas
7ccbeaa161 hive-runtime: drop an ACP refusal on respawn and when the held value is wanted
Choice::refused was cleared only when the agent accepted a
session/set_config_option, so it outlived what it described in two cases:
a respawned agent process kept the dead one's refusal masking the wanted
value in the picker, and a refusal survived a turn whose wanted value was
the one the session already held (or none), so re-picking the refused
value snapped the picker back to the session's value.

start() now resets the refusals along with the offered options, and
choose() clears the category's refusal when the wanted value is none or
the one the session holds.

Refs #4832
2026-09-30 14:54:39 +02:00
atlas
25a09dd3ec hive-runtime, hive-agent: clear the ACP picker's pending model on a refusal
When an ACP agent answered session/set_config_option with an error,
choose() only logged it, so the session's current value never moved and
offered_pickers() kept treating the refused model as pending: the picker
showed the refused model as current and hid the effort picker.

hive-runtime now keeps the refused value per category on Choices, exposed
as Choice::refused, set on the RPC error and cleared when the agent next
accepts a value for that category. offered_pickers() treats a refused
value as not settable, so the picker shows the session's actual model and
its effort levels.

The fake ACP agent gains REFUSE, which errors the first
session/set_config_option to that value.

Refs #4832 (ACP half only)
2026-09-30 14:54:39 +02:00
iris
d74f567c59 nix: runtime option description: ACP cancel, compact and model/effort now work 2026-09-30 13:12:49 +02:00
atlas
9de5ba4279 hive-agent: warn once on a model-list mismatch; test the ACP set-model/effort branch
filter_offered_models logged its "no match" warning on every pickers() call;
pickers() runs on every /api/state poll (every 4s per open dashboard tab)
and on every post_set_model/post_set_effort call. Guard it with a
process-lifetime AtomicBool, threaded in explicitly so tests can use their
own instead of racing each other on a shared static.

Extract post_set_model/post_set_effort's ACP accept/refuse decision into a
pure accept_or_refuse() helper and add tests for it: an offered value is
accepted, an unoffered one is refused with 400, and an empty offered list
(before the first turn, or effort while a model switch is pending) refuses
with 400 too.

Not fixed here: a session refusing session/set_config_option leaves the
picker showing the refused value as pending indefinitely. Fixing it needs
hive-runtime to expose a real refused/pending distinction that Choice and
SessionChoices don't carry today, threaded through to hive-agent's
model_pending computation, plus a fake-agent test fixture that can
simulate an RPC refusal — a cross-crate change, not a small one.
2026-09-30 10:53:53 +02:00
atlas
8a8da5ec8a hive-agent: ACP model picker honours availableModels (#4391)
Filter the ACP model picker's list by services.hyperhive.agent.availableModels: an
unconfigured agent (env absent) shows every model the session offers, a
configured list narrows the picker to whatever it names that the session
also offers (in the session's own order), and a configured list matching
none of the session's models (the claude names on an ACP agent that never
touched the option) falls back to showing everything, with one warning
naming the mismatch.

The nix option's default, the model-vs-availableModels build assertion and
hive-subagent-mcp's check_model rail are unchanged — this only touches the
web UI's picker.
2026-09-30 10:53:53 +02:00
atlas
f189724a4c hive-runtime, hive-agent: model/effort picker for ACP agents from configOptions
An ACP agent's model and effort pickers now list what its session offers
(its `model` and `thought_level` config options) instead of the claude
model list and EFFORT_LEVELS. A pick goes through the same Bus::set_model /
Bus::set_effort -> Config.model / Config.effort path as on claude; before
each prompt the ACP runtime sets it with `session/set_config_option`, model
first, and only when the session offers that value. Options are re-read
from the set response and from `config_option_update`, so the effort picker
disappears when the chosen model offers no effort levels, and is hidden
while a newly picked model waits for the next turn.

The session's options reach the web UI through a `Choices` handle from a
new `Runtime::choices`, registered on the bus the way `canceller` is.
/api/model and /api/effort accept only offered values on ACP. The claude
path is unchanged.

Refs #4391
2026-09-30 10:53:53 +02:00
atlas
abd547f600 hive-subagent-mcp: give the ACP test agent's replies a token usage
The daemon's own mock ACP agent answered with a bare `end_turn`, no
usage, same as hive-runtime's test agent did before EmptyEndTurn: two
subagent tests failed because the mock's turns now looked like the
provider-error case the runtime backend reports.
2026-09-30 10:13:54 +02:00
atlas
af8e0fe681 hive-runtime: exempt ACP compaction turns from the empty end_turn check
A checkpoint or `compact` command turn can end with a bare `end_turn`
as a normal answer, since the agent does the work on its side. Treating
it as `EmptyEndTurn` made `compact_session` count every such compaction
as failed and archive the session. `prompt()` and `turn()` now take a
`TurnKind`; only `run()`'s ordinary turns get the check.

The test agent's `BLANK` env answers one prompt text with a bare
`end_turn`, for the compact and checkpoint tests.
2026-09-30 09:27:40 +02:00
atlas
f9c6a56ab9 hive-runtime: report an empty ACP end_turn as a turn error (#4819)
opencode answers `end_turn` even when its provider rejected the request
with a non-retryable error (401, 400), and forwards nothing over ACP, so
the turn looked like an empty success. A turn that ends with `end_turn`,
no event, no `usage_update` and no `usage` in the prompt response now
fails with `AcpError::EmptyEndTurn`.

A turn that really produced nothing and reported no usage is reported
the same way; that false positive is accepted.

The test agent's plain `end_turn` replies now carry a response `usage`,
so its ordinary turns stay successes; a response `usage` feeds only cost
telemetry, not the compaction watermark. A new `blank` mode keeps the
empty reply for the error case.
2026-09-30 09:11:36 +02:00
atlas
c01ddd8238 hive-subagent-mcp: fix build against AcpRuntime's new policy parameter
#4822 added a CompactionPolicy generic parameter and a policy argument
to AcpRuntime::new. #4822 and #4824 (which added acp_runtime's call)
were reviewed and CI'd independently and merged textually clean, so
neither build caught that the two conflict: forge/main fails to
compile with 'missing generics for struct AcpRuntime' and 'this
function takes 4 arguments but 3 were supplied'.

Parameterize AcpRuntime<hive_claude::NeverCompact>, matching this
crate's documented no-mid-turn-compaction design (session.rs's module
doc: subagents are bounded, single-batch work, not sessions long-lived
enough to need in-place compaction).
2026-09-30 08:11:51 +02:00
atlas
b68fd7306e refresh-consumer: key the restart on the file's mtime, not a pre-write compare
The restart decision was a shell variable set by comparing the fetched
value with the file just before overwriting it. A run that wrote the
file and then failed before the restart (the matrix unit's registration
render, or `systemctl --machine` finding no bus yet) left a retry that
saw an unchanged file and never restarted the consumer.

The file is now written only when the value differs, so its mtime marks
the last real change, and `refresh_consumer <machine> <unit> <path>`
compares that mtime with the consumer's ActiveEnterTimestamp on every
run, the shape the openbao client-CA refresh in swarm-bao.nix already
uses. A consumer that started after the last change is left alone; a
running one is try-restarted, a failed one reset and started, all with
--no-block, and nothing happens while the container is down.

The helper's comment block also exceeded the 30-line limit
(`comment-block lint` failed on d871467d); its per-function notes now
sit beside the functions.

module-eval-bao-grants asserts the gated write, the path the refresh is
keyed on, and the mtime-vs-start comparison for each consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
b3b42d3279 credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
3db1233da0 bao: mkRemovedOptionModule for matrixCtlHiveName
A plain removal breaks any out-of-tree host config that still sets the
option: eval fails with "option does not exist" and no pointer to what
replaced it. mkRemovedOptionModule gives a clear evaluation error instead.
2026-09-30 07:42:22 +02:00
atlas
b58a0d8ba9 bao: drop matrix-ctl's per-hive sender-token grant
swarm-controller is the only minter of swarm/hives/<hive>/matrix/sender-token
since #4820, so the swarm-matrix-ctl policy's first stanza granted a write no
code performs. The policy keeps its one used stanza, the swarm appservice token
that `swarm-matrix-ctl appservice publish` writes. deploy.bao.matrixCtlHiveName
only named the hive in the dropped stanza and goes with it.

hive-matrix.nix no longer calls the per-hive appservice's as_token hive-c0re's
authority: tuwunel loads the registration and creates the sender account, and
no client presents that token.
2026-09-30 07:42:22 +02:00
atlas
b8ec0f4f85 hive-runtime: replace the session when an ACP compact command fails
A `compact` command that errors, or sends nothing, used to leave the
session as full as before: `compacted` stayed false, so every following
turn ran the checkpoint and another `/compact` again, each one waiting
out the 600s turn idle window when the command was silent.

- The `/compact` turn gets its own idle bound, COMPACT_IDLE (3 min, or
  the turn's idle window if shorter), through the turn's existing
  watchdog, so a silent command is cancelled (or killed) like any
  stalled turn.
- When the command fails or hits that bound, compaction falls back to
  the no-command path: the session is archived and the next turn starts
  a new one. The checkpoint turn runs there only if it has not already
  run in this compaction.
- Compaction returns early when attaching produced a new session (a
  failed `session/load` or a never-answered first prompt): there is
  nothing in it to compact.

Refs #4391
2026-09-30 07:41:47 +02:00
atlas
cec35bfbb1 hive-runtime: compact ACP sessions through the agent's compact command
An ACP agent's session is now compacted like a claude one: proactively
once a turn crosses the percent-of-window watermark, and on the
operator's /compact or the agent's compact tool. Before, the ACP
backend's compact returned Unsupported and no watermark applied to it.

- AcpRuntime takes the same CompactionPolicy as ClaudeRuntime;
  make_session builds one PercentPolicy (with CHECKPOINT_PROMPT) and hands
  it to whichever backend runs.
- The runtime keeps the commands each session advertises in
  available_commands_update. If `compact` is among them, compaction sends
  the prompt `/compact` on the same session, which is how the ACP spec
  runs an advertised command. A proactive compaction runs the checkpoint
  turn first, as InfiniteSession does.
- With no compact command, the checkpoint turn runs, the session is
  archived, and the next turn starts a new one with the system prompt.
- Error::Unsupported had no producer left, so it and drive_turn's
  "/compact skipped" arm are gone.

Refs #4391
2026-09-30 07:41:47 +02:00
atlas
3bef1dfab6 hive-subagent-mcp: refuse an ACP interrupt with no turn in flight
`Canceller::cancel()` does nothing and returns `false` when no ACP turn
is in flight: from `start` until the run's task begins its turn, and
from a turn's end until `after_turn` releases the name. `interrupt`
ignored that return, recorded `Cancelled` and replied that the run was
cancelled, so a non-goal run went on to do its whole first turn while
`status` read as cancelled.

On ACP, `interrupt` now records the stop, cancels, and if nothing was
in flight withdraws the stop under the same `stops` lock, restores the
tracking entry, and refuses with the claude path's "still starting, try
again shortly". The claude arm is unchanged.

Review finding on #4824 (argus).
2026-09-30 07:41:12 +02:00
atlas
84d4808d54 hive-subagent-mcp: run an agent's subagents on its runtime
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.

The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.

Refs #4391
2026-09-30 07:41:12 +02:00
flake-bot
a483b23ccf nix flake update 2026-09-30 05:01:14 +02:00
atlas
ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00
atlas
91e47732a6 hive-agent: dashboard Cancel and idle stalls for ACP agents
`/api/cancel` stopped a turn only by SIGINTing a child process whose argv0 is
`claude`, so on an ACP agent it reported "no claude process to interrupt"
and the turn ran on. The serve loop now stores the session's canceller on
the `Bus` when the runtime has one, and `/api/cancel` uses it: the agent is
sent `session/cancel`, and the next wake prompt carries the interrupted hint,
as for a signalled claude turn. Without a canceller (claude), the SIGINT path
is untouched.

An ACP turn stopped by the idle watchdog (`HIVE_TURN_IDLE_SECS`) becomes
`TurnError::AgentStall(note)`, handled like `ApiStall`: park for
`HIVE_STALL_SLEEP_SECS`, requeue, and record `api_stall`. The TurnEnd note
is the runtime's message (how long the agent was silent, whether it had to
be killed, and that a silently retried provider error such as HTTP 429
looks like this), not the claude-specific `ApiStall` text. A cancel the
agent ignores is a `Failed` turn.

Refs #4391
2026-09-29 23:25:48 +02:00
atlas
bcbeb8ac9c hive-runtime: cancel and idle watchdog for ACP turns
`Runtime` gets a fourth operation, `canceller()`: a handle that stops the
turn in flight from outside `run`. The ACP backend returns one; claude
returns `None`, because the harness stops a claude turn by signalling the
`claude` process, and that path is unchanged.

Both stops send the agent `session/cancel`:

- `Canceller::cancel()`, when asked from outside. The turn then ends
  normally, reported with stop reason `cancelled` whatever reason the agent
  gives. opencode 1.15.10, for one, answers a cancelled prompt with
  `end_turn` (`acp/agent.ts` `prompt()` always returns `end_turn`).
- The idle watchdog, once no `session/update` has arrived for
  `Config::idle_timeout`, the same field claude's watchdog reads. The turn
  fails with `AcpError::IdleTimeout`.

An agent that has not answered the prompt 10s after `session/cancel` is
killed (`IdleKilled` / `CancelIgnored`), and the next turn respawns it.

The watchdog is also what ends a turn stuck on a provider HTTP 429.
opencode 1.15.10 retries a retryable provider error with no attempt limit
(`session/retry.ts` `policy`, `session/processor.ts` `Effect.retry`) and
forwards neither `session.status` nor `session.error` over ACP (its
`handleEvent` only handles `permission.asked` and `message.part.*`). So the
ACP client sees nothing at all until the provider recovers. The
`IdleTimeout` message says a silently retried provider error looks like
this.

Refs #4391
2026-09-29 23:22:50 +02:00
atlas
8ec51ca774 hive-runtime: retry an ACP session whose first prompt failed on that session
A new session is only recorded once its first prompt is answered, so a
failed first prompt made the next turn run `session/new` again and left the
first session behind in the agent: in the same process, and after a restart.

The new session's id is now kept in `<session file>.pending` until that
prompt is answered. The next turn attaches it (in-process, or through
`session/load` after a restart) and still prepends the system prompt, since
the session has not had an answered prompt yet. Recording the session
removes the pending file, and `archive` drops it.

Raised by argus in the #4812 review.

Refs #4391
2026-09-29 23:20:24 +02:00
atlas
bed72ce280 hive-runtime, hive-agent: MCP permission only for kind other
A permission request counted as an MCP tool call whenever its title
looked like `<server>_<tool>`, whatever its kind, and acp_permits
allowed MCP calls before looking at the kind. So an `execute` request
titled e.g. `hyperhive_x`, or a `fetch` without web_tools, was allowed.

MCP tool calls come with kind `other` (opencode's toToolKind maps every
tool it doesn't name, MCP tools included, to "other"). The runtime now
sets PermissionAsk::mcp_server only for kind `other`, and acp_permits
allows an MCP server's tool only under the `other` arm.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
868fc789d1 hive-agent: default-deny ACP permission requests
acp_permits allows tools of the session's MCP servers, the read, edit
and search kinds, and fetch with web_tools; every other kind, including
execute, other and kinds it doesn't know, is refused. Before, anything
but execute (and fetch without web_tools) was allowed, which was only
safe while the preset's own config denied the risky built-ins.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
a2ab40cc69 hive-runtime: record a new ACP session only once its first prompt is answered
The session id was written to the session file right after session/new,
but the system prompt rides on the first prompt only. If that prompt
failed, the next turn (or the next harness start) resumed the recorded
session as not-new and the system prompt never reached it.

The id is now written after the first session/prompt gets its reply.
A failed first prompt leaves nothing recorded, so the next turn starts
a new session and sends the system prompt again. Chosen over a separate
"system prompt delivered" marker: one file, and "recorded" already means
"usable".

Tests drive AcpRuntime against a scripted sh agent that fails the first
prompt: in-process and across a restart, the retry is a new session
carrying the system prompt.

Also: PermissionPolicy now sees a PermissionAsk (kind plus the MCP
server the tool belongs to, matched by name against the servers handed
to the session), so a caller can tell MCP tool calls from other `other`
requests.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
1b24edf4b4 nix: runtime option, acp.* and an opencode preset
services.hyperhive.agent.runtime ("claude" default | "acp") and
acp.{command,args,env}, rendered into HIVE_RUNTIME / HIVE_ACP_* only
for acp, so a claude agent's unit is unchanged. acp implies useApiKey.

acp.presets.opencode runs `opencode acp` from nixpkgs against an
OpenAI-compatible provider from acp.opencode.{provider,model,
contextWindow,outputLimit}: the config is rendered to the store with
the API key as an {env:VAR} reference, so the key is read at runtime
from backendEnvironmentFile. OPENCODE_PERMISSION denies opencode's
built-in bash, task, todowrite and websearch, and makes webfetch ask.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
f0110be76c hive-agent: drive turns through hive-runtime
AgentSession becomes hive_runtime::AgentRuntime, picked at startup from
the environment. Unset HIVE_RUNTIME keeps the claude backend, built
from the same title, store and PercentPolicy as before; drive_turn,
the 401 retry and the error mapping see the claude errors unchanged.

On acp the harness keeps the session id in the harness dir, answers
permission requests like the claude built-in allow-list (no built-in
shell, web fetch only with web_tools), maps runtime errors to Failed,
and reports an operator /compact as skipped instead of done.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
b0e26e7e44 hive-runtime: shared runtime crate with claude and acp backends
A `Runtime` trait (run / compact / archive) with two backends:

- claude: a pass-through to hive_claude's InfiniteSession and
  SessionStore, so a claude turn is the same spawn, session handling
  and errors as before.
- acp: a generic Agent Client Protocol client. It spawns the command,
  args and env from RuntimeSpec (HIVE_RUNTIME / HIVE_ACP_COMMAND /
  HIVE_ACP_ARGS / HIVE_ACP_ENV), refuses an agent whose
  mcpCapabilities.http is not true, passes the claude --mcp-config
  servers as ACP mcpServers, keeps one session id in a file
  (session/load after a restart, session/new otherwise), and maps
  session/update into claude stream-json events plus usage_update into
  Telemetry. Permission requests are answered by a caller-supplied
  policy on the ACP tool kind. compact returns Unsupported for now.

The crate depends on no hyperhive binary crate, so the subagent
daemon can move onto it without pulling in hive-agent.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
d98d407bf1 docs: fix upgrade-block bullets for the store-first sweep
`ensure_hive_user` no longer short-circuits on the token file it
already has, and a hive's per-hive store token now wins over the file
whenever it's present. Both upgrade-guide bullets described the old
file-first order and needed rewording to match.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
78d8d69c7f swarm-controller: mint each hive's matrix sender token
A hive whose homeserver runs on another host has no local
matrix-appservice-token, so hive-c0re's matrix sweep returned before
reaching the store read in ensure_hive_user: no @hive-<name>: token, no
Space, no chat room, no invites, and a sweep-health banner.

swarm-controller now mints @hive-<name>: with the swarm appservice
token for every hive in its directory, as a MintHiveSenderToken job
node queued by a five-minute pass, and stores it at
swarm/hives/<name>/matrix/sender-token, the same matrix::Credential
swarm-matrix-ctl writes there. It is keep-if-live, reusing agent_token's
classify/plan: a stored token whoami confirms as @hive-<name>: is left
alone, so only an absent or dead one is minted. agent_token's probe and
mint steps are lifted into probe_at/mint_at so both passes share them.

swarm-matrix-ctl mint still writes the path for its own hive when it is
empty. If both mint an empty path at once, one token is invalidated
(same pinned device); the next pass classifies it Revoked and re-mints.

hive-c0re's ensure_all no longer returns when there is no local
as_token. ensure_hive_user reads the store first on every sweep and
overwrites its token file when the store's token differs, keeps the
file when the store has none, mints with the local as_token only when
neither holds one, and fails with one error when there is nothing at
all. The decision is sender_source, unit-tested.

The controller's bao policy gains create/read/update on
swarm/hives/+/matrix/sender-token (`+`, since `*` is a glob only at the
end of a path), pinned in module-eval.

Refs #4427
2026-09-29 22:14:40 +02:00
atlas
9a5a947f8d swarm-controller: fix broken rustdoc intra-doc link to AGENT_TOKEN_NAME
legacy_tokens.rs referenced [`AGENT_TOKEN_NAME`] unqualified, but the
const lives in the sibling agent_token module and isn't in scope here;
rustdoc's broken-intra-doc-links lint (denied) failed the docs build.
2026-09-29 22:13:32 +02:00