Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/docs
Author SHA1 Message Date
atlas
067f4e5699 docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix
gateway.md: split the opener into what/audience/enable; vhost map in two
tables (swarm-service vhosts declared by their own modules, then the hive
vhost) matching vhosts.nix and the service modules; gateway.enable exists
and is set with mkDefault by the modules that need it; Basic auth scope,
dashboard /health/ prefix, error-page rendering, matrix body limit and
forge link source corrected; nginx internals grouped under one Internals
section with their headings unchanged.

network.md: gateway and dnsmasq run on the host, not in a container;
network.enable is set by the modules that need it; shared-netns firewall
rule covers every swarm service container; hive-priv writes the nspawn
conf; domain sentence rewritten; removed options moved into <details>.

jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS
show different graphs drawn by the same component.

observability.md: swarm tier first; history narration cut; network access
deduplicated into a link to network.md; options link made absolute.

matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and
SSO options under deploy.matrix; .well-known is served on the hive domain;
roadmap sentence deleted; stale hive-c0re provisioning claims fixed;
serverName upgrade note moved into <details>.

Refs #3902
2026-10-02 17:25:36 +02:00
atlas
2b2608a491 docs(turn-loop): move harness systemd unit shape out of agent-roster.md
Moves the "Harness systemd unit shape" section from
docs/agent-lifecycle/agent-roster.md into docs/turn-loop/README.md: it
describes the per-agent harness systemd unit (env vars, PATH wiring,
serviceConfig), which is turn-loop material, not roster material.

Fixes two facts while moving: the ExecStart package is `hive-agent`,
not `hyperhive` (no package by that name exists); and `ruth.nix`
doesn't set any forge subscription default — it only defaults
`services.hyperhive.agent.docs.enable`.

Updates the inbound pointers in docs/turn-loop/config.md and the
module comment at nix/agent-modules/agent-service.nix.

Refs #3902
2026-10-02 14:47:31 +02:00
atlas
cabde572e9 docs(web-ui): scope dashboard.md to what the hive UI renders
Per mara's review: the hive UI doc covers only what hive-c0re's pages
render. Removed the M4TR1X page section (the hive gateway redirects
/matrix/ to the swarm matrix client, which the swarm UI's quick links
open), the swarm-UI forge/matrix account-linking lines from the
CR3D3NTIALS section, the infra-services hivectl paragraph, and the H0M3
Matrix/Forge absence line. Added a single pointer to docs/swarm/ui.md,
and stated the account-linking and Matrix quick-link facts there.

Refs #3902
2026-10-02 14:44:34 +02:00
atlas
6a60312da7 docs(web-ui): fix argus's request-changes items + 5 more absence/changelog lines
Removes the remaining #system/Settings stale facts and absent-thing
mentions argus's review flagged, plus 5 more lines mara's rule-3 audit
found in the same file (named a nonexistent field/state instead of
stating current behaviour).

Refs #3902
2026-10-02 14:44:34 +02:00
atlas
843c9ad016 docs(web-ui): state current dashboard behaviour, drop changelog wording
Rewrites 12 sentences in docs/web-ui/dashboard.md that described
current state as a change from something earlier (moved/no longer/
gone/was) or stated what a field/page doesn't exist without saying
what replaced it. Verified each against the current code at forge/main
before rewriting.

Refs #3902
2026-10-02 14:44:34 +02:00
atlas
fb2fff0668 fix(nix): require swarm domain only when a hyperhive service is enabled
The swarm.domain assertion in hive-network.nix fired on every host that
imported the module, so a host that enables nothing failed eval. It now
fires only when one of the hyperhive service switches is on (every
deploy.*.enable that runs something, gateway, gateway.dns, network,
otel, snapshotStore). The requirement itself is unchanged: any host that
runs a hyperhive service still needs swarm.domain.

The core-toggle module-eval suite gains a case: a missing swarm.domain is
refused on a hive and on a swarm-service-only host, and a host enabling
nothing passes every assertion.

Closes #4887
2026-10-02 13:09:45 +02:00
atlas
eded8f2e4e docs(agents): rename agent-hierarchy.md to agent-roster.md
Per mara's #4879 review (89815): the system has no agent hierarchy,
just a flat set scoped by the capability store, so the filename no
longer matched. The file already read "Agent roster & privileges"
after the earlier facts pass; rename it to match, and update the
four inbound references (docs/README.md, coordinator.md, config.md,
agent-service.nix).
2026-10-02 12:52:38 +02:00
atlas
967bd34622 docs(agents): drop change-log section and temporal wording 2026-10-02 12:52:38 +02:00
atlas
78aacf13ce docs(agents): drop nonexistent-thing mentions, state current behaviour 2026-10-02 12:52:38 +02:00
atlas
578ee90096 docs: state current behaviour, drop change-log wording
Refs #3902
2026-10-02 12:52:38 +02:00
atlas
ac35f0f2cf docs(agents): fix dangling xref and sweep conventions.md Wake injection
Review fixes for #4879 (argus):
- mcp.md: the 'Waking the agent' cross-ref pointed at Core tools, which
  never mentions UpsertTodo/HIVE_AGENT_SOCKET. Point it at
  docs/tools/bash.md's 'Completion as a todo (loose-ends v2)' section,
  which documents the actual upsert/signal/clear mechanism.
- conventions.md 'Wake injection': still framed AgentRequest::Wake (a
  type that no longer exists) as the live wake surface with matrix/forge
  as callers. hive_core_agent_sock::Request::Wake has exactly one
  non-test reference on origin/main (the handler at
  socket_server/mod.rs:307) and no client; matrix/bash/forge all moved
  to the in-agent todo socket. Rewritten to match, linking mcp.md's
  'Waking the agent' section instead of duplicating it.

docs/tools/matrix.md:136-137 has the same stale AgentRequest::Wake claim
(and contradicts its own :159-165) but is out of scope (#4136) — noted
as a follow-up in the PR body instead of edited.
2026-10-02 12:52:38 +02:00
atlas
4e31550dad docs(agents): facts pass on agent-hierarchy.md and mcp.md
mcp.md:
- matrix and subagent extra MCP servers are http (hive-matrix-daemon,
  hive-subagent-daemon), not stdio; screen is the one entry that still
  uses the stdio default (nix/agent-modules/matrix.nix:290-297,
  screen.nix:22-25)
- set_status is always-on, not meta-group-gated; mark_todos_done (also
  always-on) was undocumented (hive-sh4re/src/permissions.rs:151,
  hive-agent-mcp/src/mcp/mod.rs:425)
- System messages: HelperEvent has 3 variants, not the 8 previously
  listed; ApprovalResolved/ContainerCrash routing and the swarm-wide
  NATS notices stream (swarm_notices.rs) replace the old per-agent
  todo-wake description for rebuilt/killed/destroyed/logged_in/needs_login
- get_loose_ends's approval rows are manager-only; PendingMessages and
  UnreadMatrix were missing from the description
  (hive-sh4re/src/inbox.rs:127-184)
- subagent spawning runs on hive-runtime (claude or ACP), not
  claude-only (hive-subagent-mcp/src/session.rs:77)
- Waking section: matrix/bash/forge all moved to the in-agent todo
  socket; the host Wake request has no built-in caller left today

agent-hierarchy.md:
- distinguished the swarm-wide agent roster (swarm-controller's
  identity store, authoritative) from the hive-local topology.json
  (a derived, reconciled cache scoping ManageRootAgent's bind-mounts),
  linking README's framing
- noted services.hyperhive.ruthless (a hive can run with no manager at
  all)
- Wire-protocol bullet: the only privileged Request variants left are
  the scheduling ops; Kill/Start/Restart/Update/GetLogs don't exist on
  this socket
- Prompt/tools: prompt::render hardcodes the agent role for every
  container today (role:manager blocks are dead code); the tool
  allow-list has no Flavor switch, it's HIVE_TOOL_GROUPS same as any
  agent

Not touched: agent-hierarchy.md:140-200 (Harness systemd unit shape,
kept in place — see PR follow-ups) and docs/agent-lifecycle/approvals.md
(blocked on #4853).
2026-10-02 12:52:38 +02:00
atlas
99905f50b0 nix(authelia): start with a disabled placeholder user when the user set is empty
authelia 4.39.20 exits at startup on `users: {}` ("users: non zero value
required"), and the first-boot unit seeded exactly that, so a swarm with
no users crash-looped authelia and answered 502 until `swarmctl user add`
ran.

The first-boot unit now writes one subject, `swarm.placeholder`, when
the users database is absent, empty, or exactly `users: {}`:

- `disabled: true` — authelia returns "user not found" for a disabled
  user before any password check (file_user_provider.go,
  CheckUserPassword).
- password: an argon2id digest with an all-zero key. It decodes (authelia
  rejects a non-digest at startup) and no known password hashes to it.
- the `.` keeps it out of agent names (`[a-z0-9-]`), and `swarmctl user
  add` refuses it as already existing. Neither writer removes users, and
  both round-trip `disabled`.

A file with any user in it is never touched.

The docs that described the crash-loop (sso.md, gateway.md, setup.md,
the sso-unavailable error page) now describe the placeholder; the
writers' load_store docs and the seed fixtures follow. module-eval
nats-authelia asserts the seed branch.
2026-10-02 12:50:47 +02:00
atlas
7b226f21dd docs(swarm): state behaviour instead of denying absent options 2026-10-02 12:50:34 +02:00
atlas
9c52809439 docs(swarm): drop statements about absent fields 2026-10-02 12:50:34 +02:00
atlas
ff2aa1e27b docs(swarm): state requirements, drop upgrade narration 2026-10-02 12:50:34 +02:00
atlas
4a7a1c341b docs(swarm): drop mentions of nonexistent fields 2026-10-02 12:50:34 +02:00
atlas
6ed61da8b8 docs(swarm): state swarm domain as required; drop changelog wording
Refs #3902
2026-10-02 12:50:34 +02:00
atlas
270430a4b4 docs(swarm): facts + structure pass
swarm/README.md opens with the swarm and its control plane; hive identity
and the directory follow as the substrate. Upgrade notes move into a
<details> block, the per-agent queue publishing detail into another, and
the one-paragraph pointer sections collapse into a link list.

Fact fixes, checked against origin/main:
- an empty swarm.hives fails eval (swarm.nix:341-354); it does not mean
  "not in a swarm"
- swarm.domain is required with a hive (hive-network.nix:156,188), hiveName
  with a hive, store or homeserver (hyperhive.nix:161-166)
- the matrix container trusts the hive's trust-bundle.pem at runtime under
  self-signed certs (hive-matrix.nix:1046-1052, lib/hive-ca-trust.nix:76-85)
- singleHostSwarm also defaults the controller, localHostsEntry, the nats
  callout keys and the bao bootstrap token path (local-defaults.nix:72-129)
- swarm-controller serves far more than /health: roster, wanted state, job
  graph, agent creation and credential mints (main.rs:2874-2899)
- swarmctl user add needs --email for the forge account and refuses an
  existing user (setup.md:67-71, swarmctl/src/main.rs:425-430); document
  agent mint-identity and mint-forge-token
- agent creation also mints store identity, forge token and matrix
  account, and declares the agent paused (main.rs:1822-1920, 247-248)

Refs #3902
2026-10-02 12:50:34 +02:00
atlas
f688cfcdf0 docs(matrix): state where the account prefix appears 2026-10-02 11:43:10 +02:00
atlas
79ada8aace docs: state current behaviour (wake path, subagent todo) 2026-10-02 11:43:10 +02:00
atlas
7d352486cd docs: state current behaviour without change-log wording 2026-10-02 11:43:10 +02:00
atlas
788542ba75 docs: drop statements about absent things, state current behaviour 2026-10-02 11:43:10 +02:00
atlas
abda781eef docs: re-pad tables after wording edits 2026-10-02 11:43:10 +02:00
atlas
37b8ca20d1 docs: state current behaviour, drop remaining change-log wording
Refs #3902
2026-10-02 11:43:10 +02:00
atlas
9545151b8d docs: state current behaviour, drop change-log wording
Refs #3902
2026-10-02 11:43:10 +02:00
atlas
62553cf3be docs: state current behaviour, drop change-log wording
Refs #3902
2026-10-02 11:41:34 +02:00
atlas
acfa2a7a56 docs(readmes): fix stale facts in remaining crate READMEs + gotchas
- frontend/README.md: list the missing packages/swarm-ui package, fix
  the vanilla-JS claim (all packages depend on preact), and the
  deprecated hyperhive.frontend.extraFiles spelling
- hive-c0re/README.md: hive-c0re no longer provisions per-agent
  forge/matrix accounts (swarm-controller does); it wires gateway
  vhosts and reconciles forge/matrix config
- hive-metric/README.md: OTEL_EXPORTER_OTLP_HEADERS is never set by
  the harness and has no way to be set
- hive-agent-sock/README.md: fix the deprecated
  hyperhive.extraMcpServers spelling
- docs/process/gotchas.md: fix a dangling hive-ag3nt/ path, the real
  directory is hive-agent/

Also fixes pre-existing vale error-level alerts (passive voice,
Microsoft.Auto, Microsoft.Contractions) in the same files so
prose-lint-errors passes clean.
2026-10-02 11:40:07 +02:00
atlas
dc41dea64d docs: state current behaviour, drop change-log wording
Refs #3902
2026-10-02 11:39:53 +02:00
atlas
49b6f0f0c3 docs(web-ui): say precisely what keeps old FORGES-tab accounts working
The old FORGES tab wrote forge-<label>-token/forge-<label>.json directly;
the swarm-fetch unit that replaced it never deletes a pair for a label it
doesn't list, so those files keep being read by hive-forge -f <label>
until the operator links an account under the same label in the swarm UI,
which overwrites both files (nix/agent-modules/forge-accounts.nix:11-13,159-176).
2026-10-02 11:39:53 +02:00
atlas
c952572376 docs(web-ui): drop the stale MATRIX-tab compat claim
hive-matrix-daemon removes undeclared matrix-token-<name> files on
every start (hive-matrix-mcp/src/main.rs:100), so there is no window
where an account linked through the old per-agent MATRIX tab keeps
working — it must be declared via matrixAccounts at swarm/agent
level.
2026-10-02 11:39:53 +02:00
atlas
fd74cbd495 docs(turn-loop): facts + structure pass
Frame the turn loop runtime-neutrally: every turn runs through
hive-runtime, on claude (default) or an ACP agent. The loop steps,
harness binary shape and hive-agent README now say so; claude-only
failure detection gets its own heading; claude-invocation.md opens with
its scope and links the ACP side to hive-runtime/README.md.

Fact fixes: on-boot file paths (/run/hive-config, not /run/hive),
hive-claude is a crates.io dependency with no README here, hive-agent
has no client.rs or forge_notify.rs, the claude launch-config layer is
hive-agent's mcp_config.rs, agent forge/matrix accounts are
swarm-controller's, hive-c0re's dashboard is dashboard/, and the
deprecated hyperhive.gui.enable / hyperhive.extraMcpServers spellings.

Refs #3902
2026-10-02 07:52:13 +02:00
atlas
816d0d5d3c docs(web-ui): address review
- hivectl/README.md: split matrix.rs/github.rs into accurate per-module
  lines (matrix.rs invites a matrix user to the hive Space/room,
  cli.rs:289-296; it is not per-agent). Fixed the summary line's
  'provisioning verbs' to name the actual verb groups left after the
  facts pass.
- web-ui/README.md: the swarm/ui.md pointer no longer promises roster/
  creating-agents/linking-accounts content that page doesn't have.
- Reverted the unrelated doesn't/does not drive-by at hivectl/README.md:5.
2026-10-02 07:50:46 +02:00
atlas
8e90c79413 docs(web-ui): facts + swarm-UI framing for operator day-to-day surfaces
docs/web-ui/README.md now frames itself as the per-hive dashboard (host
approvals, container state, rebuild queue) and points to swarm/ui.md for
swarm-wide day to day, matching the README's swarm-first reframe.

Credentials page fixes: it has only a GitHub PAT tab now (2c7e586f
removed the FORGES tab and moved external forge accounts to the swarm
UI; credentials.html never had a matrix tab).

hivectl/README.md: agents.rs has no create verb (hivectl-cli.md has no
'create' entry; agents.rs's create is swarm-level, swarmctl agent
create). forge.rs reconciles config, it doesn't provision an account;
matrix.rs/github.rs do agent-scoped invites/token writes; gateway.rs
manages the gateway's own htpasswd users — split out of the former
single 'per-integration account/token provisioning' line.
2026-10-02 07:50:46 +02:00
atlas
9d804ae094 docs: fix vale errors on main
Fix the 4 pre-existing vale error-level hits on main (docs/README.md:83,
docs/getting-started/setup.md:11,127, docs/swarm/bao.md:4) that fail CI's
prose-lint-errors job for every docs PR regardless of its own diff.
2026-10-01 23:34:47 +02:00
müde
a7991c9242 docs: setup.md as a short all-local checklist; bao internals move to swarm/bao.md 2026-10-01 20:41:47 +02:00
atlas
2c7e586f47 forge: external forge accounts live in swarm bao; the agent fetches them itself
An operator now links an agent's external forge account (label, base URL,
token) in the swarm UI. swarm-controller stores it at
swarm/agents/<agent>/forge/<label>. There is no index: the store's
listing of the agent's forge/ directory is the set of accounts.

In the agent, hive-agent-forge-accounts (oneshot + 2-minute timer, as
the agent user, under its own store certificate) lists
swarm/agents/<agent>/forge/ with the `list` #4866 grants an agent on its
own metadata subtree, reads each account, and writes
<state>/forge-<label>-token and forge-<label>.json in the names and shape
hive-forge -f already reads. An empty listing (a 404, which `bao kv list
-format=json` answers with `{}` and an empty stderr) is zero accounts; a
denial or an unreachable store fails the unit. It never deletes: files
for labels not listed, including ones the hive wrote, stay as they are.

Removed: the dashboard FORGES tab (credentials.js/html section and its
CSS), hive-c0re's extra_forges.rs and its routes, priv_client's
extra-forge calls, and hive-priv's WriteAgentExtraForgeAccount /
DeleteAgentExtraForgeAccount with their helpers. The GITHUB tab and
WriteAgentGithubToken stay.

Also: persistence.md's matrix avatar note names the exit-75 restart on a
changed account listing, not the dashboard, as what brings a linked
account up.

Refs #4348
2026-10-01 18:05:33 +02:00
atlas
97fb76ce99 matrix: the agent's daemon pulls its linked accounts from bao itself
hive-matrix-daemon now learns which external matrix accounts it has from
the swarm secret store, under the agent's own certificate, and the hive
push chain for matrix is gone.

The daemon lists swarm/agents/<agent>/matrix/ (the `list` its policy
grants on its own metadata subtree), reads each account's homeserver
from its credential, and brings the accounts up with their tokens from
the store. Every two minutes it lists again and exits with 75 when the
set of linked accounts changed; the unit restarts on 75 without counting
a failure. A listed name whose credential reads as absent is skipped and
logged once. At start it removes the matrix-token-<a> /
matrix-account-<a>.json pairs a hive delivered (a sidecar marks a pair
as delivered; a declared tokenFile keeps its token).

Removed: CredentialNotice and the $SWARM.credential.* subject and NATS
grant, the controller's publish and its queue precondition on the PUT
route, hive-c0re's credential subscription arm and workers/credential.rs,
priv_client::write_agent_matrix_token, hive-priv's WriteAgentMatrixToken
and its helpers, and the daemon's state-dir account discovery.

Kept: WriteAgentGithubToken and the external-forge path
(WriteAgentExtraForgeAccount, extra_forges.rs) are untouched, and a
declared matrixAccounts tokenFile is still read when the store has no
token for that account.

Refs #4348
2026-10-01 17:43:28 +02:00
atlas
e04616eb70 swarm-secret-client: agents may list their own subtree; controller rewrites agent policies
render_agent gains a second stanza: list on
secret/metadata/swarm/agents/<agent>/*, next to the existing read on
secret/data/swarm/agents/<agent>/*. An agent can now learn which
credentials it holds by listing its own subtree. Metadata read, writes
and every other principal's paths stay refused.

An agent's policy was only written when it was minted, so existing agents
would never get the new stanza. swarm-controller now rewrites every
agent's policy at start (read_policy::ensure_agent_policies), with the
same 30s / 24h retry as ensure_hive_access. The roster is the store's
hive-agent-* cert-auth roles, listed with the controller's existing
`list` on auth/cert/certs; the writes use its existing grant on
sys/policies/acl/hive-*. Only the policy is written: mint_and_verify
also reissues the certificate, so the pass does not call it.

Refs #4348
2026-10-01 17:43:28 +02:00
atlas
7d217f8267 remove the create_repo agent tool
mara ruled on #4849 (c88934): "remove create_repo tool". The tool ran in
hive-c0re with the hive's core token, so it only ever worked for agents
on the hive that runs the forge.

Removed:

- the create_repo MCP tool and CreateRepoArgs (hive-agent-mcp)
- wire variants Request::CreateRepo and Response::RepoCreated
  (hive-core-agent-sock)
- hive-c0re's handle_create_repo, its valid_repo_name check and the
  dispatch arm
- forge::create_agent_repo and apply_operator_branch_protection, which
  had no other caller, plus AGENTS_ORG and OPERATORS_TEAM, whose only
  users they were
- the tool's docs (docs/tools/forge.md repo management, docs/turn-loop/
  mcp.md, the conventions tool-group table) and the doc comments that
  named it (hive-sock-client's response timeout, ensure_repo_creation_
  disabled, the security doc's merge-gate bullet)

ToolGroup::Forge is kept with no tools, the same way b88a5b24 kept
Lifecycle, so existing meta/capabilities.json grants still parse.

Forge state is untouched: existing agents/* repos keep their collaborators
and operators-team branch protection. The swarm-controller's own
create_repo (config-org repos) is a different path and is unchanged.

Closes #4849
2026-10-01 09:04:17 +02:00
atlas
c5b21403a6 hive-runtime: read the ACP provider key from bao
An opencode ACP agent got its provider API key only from the hand-placed
backendEnvironmentFile. It now also reads it from the swarm secret store
at swarm/agents/<agent>/acp-provider, field api_key, under its own
certificate, and sets it in the spawned ACP agent's environment only.
Nothing is written to disk.

Precedence: a value already in the process environment (the env file)
wins and the store is not asked. Otherwise the stored key is used when
present. With no store, nothing stored, or a failed read, the agent is
spawned without the key as before, and one line is logged without the
value.

The variable name comes from the existing per-agent option
acp.opencode.provider.apiKeyEnv, exported as HIVE_ACP_API_KEY_ENV on the
harness only for the opencode preset. Other ACP commands are unchanged.

The read lives in hive-runtime, where the ACP child is spawned, so both
hive-agent and hive-subagent-daemon use it. The subagent daemon unit
gets the key name and, when the agent has a store, the agent's store
identity (the same credentials queue-identity.nix gives the harness).

No new option or setting. Closes #4841.
2026-09-30 22:55:03 +02:00
atlas
c2bdf30e05 hive-agent: export ACP-reported cost and context fill over OTLP
An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.

hive-agent's existing turn-metrics meter records three new instruments:

- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
  every backend: the two numbers the web UI's ctx% divides.

The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.

Refs #4845
2026-09-30 22:24:06 +02:00
atlas
5d7042e655 hive-agent: warn on usage responses with no recognisable windows
addresses argus's review on #4842 (c88740): fetch's success path was
silent when the response parsed but windows() found nothing to record.
should_warn_empty tracks the same once-per-streak shape as the
existing token-skip logging, reset the moment a response has windows
again, and never logs the body.

Also notes in the observability doc that a past resets_at means the
paired percent gauge is stale.
2026-09-30 20:53:03 +02:00
atlas
8a735ddcb3 hive-agent: publish Claude subscription usage (5h/7d %) as metrics
A detached task polls GET https://api.anthropic.com/api/oauth/usage every
5 minutes with the OAuth access token from ~/.claude/.credentials.json and
records, per window the response names (five_hour, seven_day,
seven_day_sonnet, ...):

- hyperhive.agent.claude_usage.percent   (%, 0-100)
- hyperhive.agent.claude_usage.resets_at (s, unix seconds)

both labelled window=<name>. Endpoint, the anthropic-beta:
oauth-2025-04-20 header and the {utilization, resets_at} window shape are
taken from the claude-code 2.1.283 bundle's own /usage fetch.

The token is only read, never refreshed: claude owns refresh-token
rotation and a second refresher can log the agent out. An expired token
is skipped until claude's next turn refreshes it. API-key agents
(HIVE_USE_API_KEY, the ACP default) and agents with no credentials file
skip quietly, and the task does nothing when OTEL is not configured.
Request failures and non-2xx statuses warn with the status or transport
error only, never the body or the token.
2026-09-30 20:53:03 +02:00
atlas
f189724a4c hive-runtime, hive-agent: model/effort picker for ACP agents from configOptions
An ACP agent's model and effort pickers now list what its session offers
(its `model` and `thought_level` config options) instead of the claude
model list and EFFORT_LEVELS. A pick goes through the same Bus::set_model /
Bus::set_effort -> Config.model / Config.effort path as on claude; before
each prompt the ACP runtime sets it with `session/set_config_option`, model
first, and only when the session offers that value. Options are re-read
from the set response and from `config_option_update`, so the effort picker
disappears when the chosen model offers no effort levels, and is hidden
while a newly picked model waits for the next turn.

The session's options reach the web UI through a `Choices` handle from a
new `Runtime::choices`, registered on the bus the way `canceller` is.
/api/model and /api/effort accept only offered values on ACP. The claude
path is unchanged.

Refs #4391
2026-09-30 10:53:53 +02:00
atlas
7eb966fe2b credential units: restart consumers on a changed credential; fix the ordering claim
The previous commit's comments said a unit in auto-restart keeps its
start job, so anything ordered after it waits for the whole 24h retry
window. That is wrong under the default RestartMode=normal: each failed
attempt passes through `failed`, which ends that start job. `After=`
dependents proceed after one attempt, `Requires=` dependents fail with
`dependency`, and the retries continue as fresh start jobs. The
2026-09-24 journal shows it with the already-2880 swarm-services-cert:
nginx got "Dependency failed" 1ms after the first failure, and
switch-to-configuration exited before the first restart was scheduled.
The comments in lib/store-retry.nix, glue-matrix-bao-token.nix,
glue-queue-agent-credential.nix, swarm-otel.nix and swarm-grafana.nix
now say that, and so does docs/swarm/credentials.md.

Because dependents start after one attempt, a consumer that loads its
credential at start never sees a value a later attempt lands, or a
rotated one. nix/host-modules/lib/refresh-consumer.nix adds
`secret_differs` and `refresh_consumer`, and the four fetch units whose
consumers take a start-time copy call them after the write, only when
the value changed:

- swarm-bao-matrix-token -> tuwunel.service in hive-matrix
- swarm-bao-otel-oidc -> opentelemetry-collector.service in swarm-otel
- swarm-bao-grafana-oidc -> grafana.service in the grafana container
- swarm-bao-forwarder-oidc -> opentelemetry-collector.service in swarm-bao

A running consumer is try-restarted, a failed one is reset and started,
all with --no-block. Inline in the fetch script rather than a
PathChanged path unit because the fetch script is the only writer and
already knows whether the value changed, and it is the same shape as
this PR's nginx hook and swarm-bao-nats-tls's restart of nats.

module-eval-bao-grants gains one case per consumer.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
cec35bfbb1 hive-runtime: compact ACP sessions through the agent's compact command
An ACP agent's session is now compacted like a claude one: proactively
once a turn crosses the percent-of-window watermark, and on the
operator's /compact or the agent's compact tool. Before, the ACP
backend's compact returned Unsupported and no watermark applied to it.

- AcpRuntime takes the same CompactionPolicy as ClaudeRuntime;
  make_session builds one PercentPolicy (with CHECKPOINT_PROMPT) and hands
  it to whichever backend runs.
- The runtime keeps the commands each session advertises in
  available_commands_update. If `compact` is among them, compaction sends
  the prompt `/compact` on the same session, which is how the ACP spec
  runs an advertised command. A proactive compaction runs the checkpoint
  turn first, as InfiniteSession does.
- With no compact command, the checkpoint turn runs, the session is
  archived, and the next turn starts a new one with the system prompt.
- Error::Unsupported had no producer left, so it and drive_turn's
  "/compact skipped" arm are gone.

Refs #4391
2026-09-30 07:41:47 +02:00
atlas
84d4808d54 hive-subagent-mcp: run an agent's subagents on its runtime
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.

The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.

Refs #4391
2026-09-30 07:41:12 +02:00
atlas
ddb7d7196d matrix: swarm-controller is the only minter
Every hive is in a swarm and every swarm runs matrix, so every swarm has a
swarm-controller, and since #4810 its hive_sender pass mints each hive's
@hive-<hive>: sender token into the store every five minutes. The two
other minters of that token go:

- swarm-matrix-ctl mint: the systemd.services.swarm-matrix-ctl unit in the
  hive-matrix container, Command::Mint and src/mint.rs. The binary, its
  appservice render/publish verbs, ctlPackage, ctlActive and the ctl cert
  role stay. bao-matrix-reader's checks on the deleted unit are removed;
  the leaf-identity and no-token-in-env checks now look at
  swarm-matrix-appservice-publish, which runs under the same identity.
- the hive-side mint ladder in hive-c0re's ensure_hive_user
  (register/appservice-login/password-login with the local as_token), with
  read_appservice_token, paths::matrix_appservice_token and the helpers
  only it used. ensure_hive_user now takes the store's token, keeps the
  file when the store has none or can't be reached, and fails otherwise.
- hivectl matrix sync-admin: the verb, HostRequest::MatrixSyncAdmin and
  handle_matrix_sync_admin. The periodic MatrixSweep (ensure_all) is
  unchanged apart from no longer reading the local as_token.

This removes the double-mint race #4810's review flagged: two minters
logging in on one pinned device could leave a dead token in the store
until the next pass.

Closes #4813
Closes #4814
2026-09-30 00:46:46 +02:00
atlas
91e47732a6 hive-agent: dashboard Cancel and idle stalls for ACP agents
`/api/cancel` stopped a turn only by SIGINTing a child process whose argv0 is
`claude`, so on an ACP agent it reported "no claude process to interrupt"
and the turn ran on. The serve loop now stores the session's canceller on
the `Bus` when the runtime has one, and `/api/cancel` uses it: the agent is
sent `session/cancel`, and the next wake prompt carries the interrupted hint,
as for a signalled claude turn. Without a canceller (claude), the SIGINT path
is untouched.

An ACP turn stopped by the idle watchdog (`HIVE_TURN_IDLE_SECS`) becomes
`TurnError::AgentStall(note)`, handled like `ApiStall`: park for
`HIVE_STALL_SLEEP_SECS`, requeue, and record `api_stall`. The TurnEnd note
is the runtime's message (how long the agent was silent, whether it had to
be killed, and that a silently retried provider error such as HTTP 429
looks like this), not the claude-specific `ApiStall` text. A cancel the
agent ignores is a `Failed` turn.

Refs #4391
2026-09-29 23:25:48 +02:00