hive-agent: export ACP-reported cost and context fill over OTLP
An ACP agent's `usage_update` carries `cost.{amount,currency}`, the
session's running total (opencode sums every assistant message in the
session). hive-runtime now reads it and turns the running total into
what each report added: a new session counts from zero, a session loaded
into a freshly started agent only baselines on its first report, and a
falling total adds nothing. The spend is held on the runtime until
`Runtime::take_reported_cost` drains it; claude's runtime reports none,
since the claude binary already exports `claude_code.cost.usage`.
hive-agent's existing turn-metrics meter records three new instruments:
- `hyperhive.agent.cost.usage` (counter, `model` + `currency`), ACP only;
- `hyperhive.agent.context.used` / `.size` (gauges, no attributes), for
every backend: the two numbers the web UI's ctx% divides.
The `hyperhive · agents` dashboard gets ACP cost panels on its cost tab
and a context-fill panel on its health tab.
Refs #4845
This commit is contained in:
parent
5d7042e655
commit
c2bdf30e05
8 changed files with 605 additions and 26 deletions
|
|
@ -355,22 +355,27 @@ When OTEL is enabled, the harness itself (`hive-agent`) exports one small set
|
|||
of metrics per claude turn, recorded the moment the turn ends (not polled).
|
||||
These are deliberately the fields Claude Code's own built-in export (see
|
||||
above) can't know about — the harness's own wall-clock timing, what woke the
|
||||
turn, its own outcome classification, the loose-ends backlog, and session
|
||||
boundaries. Token usage, cost, and tool-call counts are **not** duplicated
|
||||
here; that's already covered by Claude's own export.
|
||||
turn, its own outcome classification, the loose-ends backlog, session
|
||||
boundaries, and how full the context window is. Token usage and tool-call
|
||||
counts are **not** duplicated here; that's already covered by Claude's own
|
||||
export. Cost is here only for agents on an ACP runtime, which have no export
|
||||
of their own — see [ACP-reported cost](#acp-reported-cost).
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
|
||||
| Metric | Unit | Kind | Attributes |
|
||||
| ---------------------------------------- | ---- | --------- | ------------------------------------------------------------------------------------------------------------------ |
|
||||
| `hyperhive.agent.turn.duration` | `ms` | histogram | `wake_from`, `result_kind`, `model` |
|
||||
| `hyperhive.agent.turn.count` | — | counter | `wake_from`, `result_kind`, `model` |
|
||||
| `hyperhive.agent.session.count` | — | counter | `model` (incremented once per fresh, non-`--continue`'d session) |
|
||||
| `hyperhive.agent.loose_ends.threads` | — | gauge | none |
|
||||
| `hyperhive.agent.loose_ends.reminders` | — | gauge | none |
|
||||
| `hyperhive.agent.claude_md.lines` | — | gauge | none — recorded from the `CLAUDE.md`-size watch's own ~15-minute tick, **not** per turn like the rows above |
|
||||
| `hyperhive.agent.claude_usage.percent` | `%` | gauge | `window` (`five_hour`, `seven_day`, … as the usage endpoint names them) — polled every 5 minutes, **not** per turn |
|
||||
| `hyperhive.agent.claude_usage.resets_at` | `s` | gauge | `window` — unix seconds at which that window resets; same 5-minute poll |
|
||||
| Metric | Unit | Kind | Attributes |
|
||||
| ---------------------------------------- | --------- | --------- | ------------------------------------------------------------------------------------------------------------------ |
|
||||
| `hyperhive.agent.turn.duration` | `ms` | histogram | `wake_from`, `result_kind`, `model` |
|
||||
| `hyperhive.agent.turn.count` | — | counter | `wake_from`, `result_kind`, `model` |
|
||||
| `hyperhive.agent.session.count` | — | counter | `model` (incremented once per fresh, non-`--continue`'d session) |
|
||||
| `hyperhive.agent.loose_ends.threads` | — | gauge | none |
|
||||
| `hyperhive.agent.loose_ends.reminders` | — | gauge | none |
|
||||
| `hyperhive.agent.claude_md.lines` | — | gauge | none — recorded from the `CLAUDE.md`-size watch's own ~15-minute tick, **not** per turn like the rows above |
|
||||
| `hyperhive.agent.claude_usage.percent` | `%` | gauge | `window` (`five_hour`, `seven_day`, … as the usage endpoint names them) — polled every 5 minutes, **not** per turn |
|
||||
| `hyperhive.agent.claude_usage.resets_at` | `s` | gauge | `window` — unix seconds at which that window resets; same 5-minute poll |
|
||||
| `hyperhive.agent.context.used` | `{token}` | gauge | none — tokens in the context window at turn end, the numerator of the web UI's ctx% |
|
||||
| `hyperhive.agent.context.size` | `{token}` | gauge | none — the context window those tokens fill, ctx%'s denominator |
|
||||
| `hyperhive.agent.cost.usage` | — | counter | `model`, `currency` — ACP agents only, see below |
|
||||
|
||||
For the two `claude_usage` gauges the harness polls the Claude subscription
|
||||
usage endpoint (`GET /api/oauth/usage`, the one behind claude's own `/usage`)
|
||||
|
|
@ -381,6 +386,12 @@ session (API-key backends, not yet logged in) or an expired token skips the
|
|||
poll, so a gauge keeps its last value until the next successful poll — a
|
||||
`resets_at` in the past means the paired `percent` is stale.
|
||||
|
||||
The two `context` gauges hold the two numbers the web UI divides for its ctx%:
|
||||
the last turn's context tokens (input, cache read and cache creation), and the
|
||||
API-reported window, else the model's default. A turn that parsed no usage
|
||||
leaves both where the previous turn put them. The dashboard's **Context
|
||||
window used by agent** panel (health tab) divides one by the other.
|
||||
|
||||
Resource attributes (`service.name`, `agent`, `hive`, `swarm`) come from the
|
||||
same container-wide `OTEL_RESOURCE_ATTRIBUTES` as everything else in this
|
||||
section — nothing extra to configure. Cadence follows
|
||||
|
|
@ -389,6 +400,32 @@ section — nothing extra to configure. Cadence follows
|
|||
the harness flushes the batched points to the collector, not how often it
|
||||
records them (every turn, always).
|
||||
|
||||
### ACP-reported cost
|
||||
|
||||
An ACP agent reports what its session has cost so far in the `cost` field of
|
||||
its `usage_update` notifications — a running total, not a per-turn figure
|
||||
(opencode sums every assistant message in the session). The harness counts
|
||||
what each report adds to the last one and exports that as
|
||||
`hyperhive.agent.cost.usage`, so `increase()` over it gives the money spent
|
||||
in a range. It follows a few rules:
|
||||
|
||||
- A new session's first report counts in full, since the session started at
|
||||
zero.
|
||||
- A session loaded into a freshly started agent (after a harness restart or
|
||||
an agent crash) has an unknown total until it reports one, so that first
|
||||
report only sets the baseline — the turn it covers goes uncounted.
|
||||
- A total lower than the previous one means the agent dropped part of the
|
||||
session's history; it adds nothing and becomes the new baseline.
|
||||
- Cost from a compaction the harness runs between turns counts toward the
|
||||
next turn.
|
||||
|
||||
The value is in whatever `currency` the agent names; opencode always sends
|
||||
`USD`, and the `hyperhive · agents` dashboard's ACP cost panels (cost tab)
|
||||
filter on it. A claude agent never records this metric: claude's own export
|
||||
already carries its cost as `claude_code.cost.usage`, and keeping the two
|
||||
apart means no turn is ever counted twice. Add both names together for a
|
||||
whole-swarm spend figure.
|
||||
|
||||
## Hive-scoped metrics (hive-c0re)
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
|
|
|||
Loading…
Reference in a new issue