hyperhive/docs/scheduler/observability.md
iris 09e4e2f5e9 docs: fix genuine passive-voice hits in docs/scheduler
Seventh batch of the ongoing write-good.Passive pass (hyperhive#4042):
read all 65 hits across jobq.md/ci.md/observability.md/coordinator.md
in context and rewrote 41 with a clearly nameable actor -- mostly
hive-c0re, nix/the nix module, the harness, or a specific fn/type
named right there or nearby (coordinator.md's node-inventory table
and DAG-shape descriptions name concrete Rust items constantly, so
the actor is almost always sitting in the same sentence).

Left 24 alone: predicate-adjective-copula state descriptions ("is
stuck", "is gone", "is unaffected", "is done", etc. -- the largest
recurring bucket this batch, especially in observability.md's
scope/status descriptions), negative-capability idioms ("no X is
needed/left", "X can't be written down"), the established "is
tracked as a follow-up" idiom, a firewall-shorthand notation
("bridge->127.0.0.0/8 is dropped") where rewriting would break the
compact rule-like format, a CLI-flag "(repeatable)" annotation ("May
be repeated"), a Rust type-signature fact ("`moves` is typed ..."),
a hypothetical/counterfactual maintenance-burden clause, a
readiness-condition list ("a node is ready when ... every dep is
satisfied"), and one deliberately-parallel idiom pair
("When OTEL is enabled" used identically twice as a section-opening
convention -- fixing one would break the parallelism, not the
opposite).

One self-caught regression: an early attempt to fix "used by every
`Reconcile` node's start action" (a reduced participial clause, not
flagged) into "is used by every `Reconcile` node's start action"
introduced a brand-new flagged passive. Caught by the post-edit vale
count (expected 65->24, got 65->25) not matching, same discipline as
the docs/turn-loop batch's tail-truncation catch -- re-ran with
active voice instead ("Every `Reconcile` node's start action uses
this fallback").

Verified via vale before/after: 65 -> 24 write-good.Passive hits,
exactly the 24 left alone above; error count and other warning
categories unchanged. Re-read every changed line in full surrounding
context after editing before running the final vale check.
2026-09-08 20:23:54 +02:00

387 lines
22 KiB
Markdown

# Observability (OpenTelemetry)
hyperhive has built-in support for exporting per-agent Claude Code statistics —
token usage, cost, tool call counts — to any OTLP-compatible collector via
Claude Code's built-in OpenTelemetry integration.
This is a **hive-wide** setting: one switch in the host NixOS config enables it
for every agent container simultaneously. No per-agent opt-in or opt-out exists.
## Enabling export
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel";
};
```
`enable` is the single gate. `endpoint` is where telemetry ends up after it
leaves the swarm — optional, because the swarm's own metrics store
(`deploy.victoriametrics`) is a destination in its own right. With both,
telemetry goes to both. See
[`swarm/services.md`](../swarm/services.md#metrics-victoriametrics--grafana).
**Telemetry leaves a hive exactly one way: through the collector that
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
they export unauthenticated to a bridge address only their own containers can
reach. That collector forwards to the swarm's
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), which is
the single process holding the upstream credential and the only writer to the
swarm's store. No agent holds a copy, and neither does this hive.
The hive collector reaches the swarm collector by its gateway name
(`swarm.otel.domain`, default `otel.<swarm domain>`) — the same DNS-and-CA-trust
shape every hive-to-swarm-service hop uses, not a URL an operator has to point
anywhere. A hive that doesn't run the swarm's services still resolves that
name through the gateway; nothing here needs setting for the split-host case.
⚠️ **The collector is therefore in the path of all telemetry.** It runs on the
same host as the agents and restarts on failure, and telemetry isn't the
control plane — degraded telemetry isn't degraded operation — but the export
no longer survives independently of anything host-side.
### what the agent→collector hop is and isn't
**It has no application-level auth.** The receiver takes any OTLP that reaches
it; what bounds who can reach it's the firewall — `exposeHostPorts` opens the
port on the bridge interface only — "unauthenticated to a bridge address"
means _reachable from an agent container_, not _presents a credential_.
The consequence, stated because it's a choice rather than an oversight: **any
agent can push arbitrary OTLP, and it's forwarded on under the operator's
credential.** Neither tier can tell a container's genuine Claude Code stats
from anything else shaped like OTLP arriving on that port — including data
smuggled out in resource attributes on an otherwise-legitimate export.
That's a **different risk from the one the collector fixes**, and strictly
smaller than what preceded it: before, every agent held the upstream credential
itself, so it could do all of the above _and_ use the token anywhere else. The
collector removes the token and keeps the pipe. Agents are inside the trust
boundary (`docs/trust-boundary/security.md`: capability = accepted risk), so an agent being
able to _send_ is an accepted extension of that boundary — but it's not
closed by this design, and nothing here should be read as closing it.
**The `agent` label is self-reported, and no planned authentication changes
that.** Treat it as a convenience for grouping dashboards, never as evidence of
which container produced a sample: any agent that can reach this hive's
collector can label its data as any other agent.
Worth spelling out, because two different hops are in play and only one of them
is getting a credential:
- **agent→collector** (this section's hop) stays open on the bridge. Nothing
downstream can tell one agent's export from another's.
- **hive→swarm** is where the planned ingest auth goes. The swarm tier stamps
`hive=` from the connection it authenticated, so _that_ label becomes
unforgeable.
A verified `hive` is reachable and a verified `agent` isn't — and that falls
out of the topology rather than being a gap someone forgot to close. The swarm
runs one collector, and the mechanism gives it no finer grain: a bearer-token
check never reveals _which_ token matched, and a receiver reads request metadata
rather than the claims it authenticated with.
If you need per-agent numbers you can act on, take them from the agent's own
turn-stats rather than from a metric label.
## Options reference
The nix module (`nix/host-modules/otel.nix`) generates every
`services.hyperhive.otel.*` option's full type/default/description/
example straight into [`/options/`](/options/) (host options — `nix build
.#docs-host` for a local render). The build keeps that page honest in a
way a hand-copied version here can't be, so it's the reference, not this
doc. What follows is what a flat per-option listing can't express: the
two-tier architecture, the security model, and how the options interact.
## The two collectors
Telemetry crosses two collectors, and which one you configure depends on what
the host is:
| | runs where | receives from | does |
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
| **hive tier**`otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
| **swarm tier**`deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's store and exports upstream |
An all-local host runs both, and needs nothing said about the hop between them.
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel"; # the upstream
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
};
```
**Why the hive tier isn't optional.** Exporting straight to `endpoint` means
every agent needs the credential to authenticate — and the harness delivers
that token into the agent's own `~/.claude/settings.json`, a file the agent can
read. `0600` protects it from other containers, not from the agent itself. As
long as the direct path stays _selectable_, that hole stays selectable; an
option that can reintroduce it's a hole with extra steps.
**Why the tiers stay separate on one box.** They're not collapsed when
co-located: an all-local hive is a statement about _where_ processes run, not
about the shape of the deployment. A boundary that disappears locally is one
the local deployment stops testing.
**`endpoint` keeps meaning "where telemetry goes upstream."** Neither tier
redefines it — the agent-facing value is _derived_
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
keeps working unchanged. The otel module contributes the bridge port to `exposeHostPorts`
automatically; there is nothing to open by hand.
### Authenticated ingest
The swarm tier gives **each hive its own receiver**, and stamps the `hive` label
from whichever receiver accepted a sample. A hive therefore can't report
metrics as another hive, and can't relabel its own by editing what it sends —
the label isn't taken from the payload at all.
**On an all-local swarm there is nothing to set.** Each hive already has an
identity, and its collector reads the secret that host's own authelia minted.
**On a hive that doesn't host the swarm's services**, the secret has to arrive
somehow — copy it across and name it:
```nix
services.hyperhive.otel.clientSecretFile = "/run/secrets/hive-telemetry.secret";
```
**No unauthenticated mode exists.** A hive always presents an identity, so a
missing credential is a build error rather than a quieter fallback — the
collector has no anonymous route to accept samples on, and every path it serves
belongs to exactly one hive.
Getting the secret wrong shows up as the hive's collector logging 401s from the
swarm tier and no metrics appearing for that hive.
⚠️ **`endpoint` must be valid for `protocol`.** The upstream exporter follows
`otel.protocol` (`grpc` → the gRPC exporter, otherwise OTLP/HTTP), and the gRPC
exporter takes an _address_: `https://host/path` is a legal
`OTEL_EXPORTER_OTLP_ENDPOINT` for HTTP but fails as gRPC with _"missing port in
address"_. nix validates the collector's config at build time, so a mismatch is
a build error naming the reason rather than telemetry silently going nowhere.
## Network access
Agent containers can only reach the host on ports 80 and 443 by default. To let
them reach some other host-local service you run yourself — a database, a
scratch HTTP endpoint — open its port on the bridge:
```nix
services.hyperhive.network.exposeHostPorts = [ 5432 ];
```
and point whatever consumes it at `10.42.0.1:5432` rather than loopback: inside
a container, loopback is the _container_. The bridge IP is the host's address on
the `hive-br0` bridge. The service must also bind an address the bridge can
reach — a `127.0.0.1`-only listener stays unreachable no matter what the
firewall allows. See `docs/networking/network.md::Reaching host services` for details.
⚠️ **None of this is needed for hyperhive's own telemetry**`otel.enable`
contributes the collector's port and derives the agent-facing endpoint itself.
## Built-in resource labels
The harness sets the OTLP variables (`OTEL_EXPORTER_OTLP_ENDPOINT`, `_PROTOCOL`,
`OTEL_RESOURCE_ATTRIBUTES`, the temporality preference) **container
wide** — in systemd's `DefaultEnvironment` and in `/etc/profile` — so every
process in an agent container exports to the hive's collector without any
per-tool wiring. That covers Claude Code, `hive-metric`, and anything you run
yourself from a tool call or `hivectl shell`.
Every agent's export therefore includes these resource attributes
automatically:
| Attribute | Value |
| -------------- | ------------------------------------------------------------ |
| `service.name` | `hyperhive-agent` (constant) |
| `agent` | agent logical name (for example `iris`) |
| `hive` | hive display name (`services.hyperhive.hiveName`) |
| `swarm` | swarm display name (`services.hyperhive.swarm.name`, if set) |
Append additional labels via `extraResourceAttributes` (see option
reference above); pass custom per-data-point labels with
`hive-metric --labels` (see below).
## Host-emitted container-resource metrics (hive-c0re)
When OTEL is enabled, **hive-c0re itself** also exports each agent
container's resource load — the same cgroup gauges shown on the dashboard
LOAD tab — to this hive's own collector, exactly like an agent does and with
no separate toggle. These come from the host, not the in-container Claude SDK,
so they cover containers even when their agent is idle.
Emitted via the OpenTelemetry Rust SDK, using the
[semconv `container.*`](https://opentelemetry.io/docs/specs/semconv/system/container-metrics/)
metric names + the standard `container.name` attribute where a spec metric
exists, so off-the-shelf OTEL/Grafana container dashboards work. Resource
`service.name = hyperhive-c0re`; hive-c0re tags each data point `container.name`
(= the `h-<agent>` machine) and the hive `agent` label:
| Metric | Unit | Kind | Source |
| ----------------------------------- | ---- | ------- | ----------------------------------------------------------------------------------------------------------- |
| `container.cpu.time` | `s` | counter | cumulative `cpu.stat` `usage_usec` → seconds |
| `container.memory.usage` | `By` | gauge | `memory.current` |
| `hyperhive.container.memory.limit` | `By` | gauge | `memory.max` (custom — semconv has no `.limit` metric; omitted when unlimited) |
| `hyperhive.container.memory.peak` | `By` | gauge | `memory.peak` (custom — no semconv metric; omitted if unavailable) |
| `hyperhive.container.storage.usage` | `By` | gauge | state dir + writable rootfs (custom — semconv only has `disk.io`; omitted until the slow disk sampler runs) |
| `hyperhive.container.cpu.percent` | `%` | gauge | host-normalised percent (custom — the value the dashboard LOAD tab shows, no `rate()` needed) |
The `hyperhive.`-prefixed metrics have no semconv equivalent (memory
limit + peak, on-disk footprint, and an instantaneous cpu percent kept
alongside the spec `container.cpu.time` counter for convenience). Hive
labels (`hive`, `swarm`, …) ride on the resource via
`extraResourceAttributes`.
Cadence follows `metricIntervalMs` (default 60s). Transport is OTLP/HTTP
(JSON) to the hive collector's bridge address, with no auth header — that
first hop is unauthenticated for every producer on this host, and the upstream
credential stays on the swarm tier.
## Agent-emitted per-turn metrics (`hive-agent`)
When OTEL is enabled, the harness itself (`hive-agent`) exports one small set
of metrics per claude turn, recorded the moment the turn ends (not polled).
These are deliberately the fields Claude Code's own built-in export (see
above) can't know about — the harness's own wall-clock timing, what woke the
turn, its own outcome classification, the loose-ends backlog, and session
boundaries. Token usage, cost, and tool-call counts are **not** duplicated
here; that's already covered by Claude's own export.
| Metric | Unit | Kind | Attributes |
| -------------------------------------- | ---- | --------- | ----------------------------------------------------------------------------------------------------------- |
| `hyperhive.agent.turn.duration` | `ms` | histogram | `wake_from`, `result_kind`, `model` |
| `hyperhive.agent.turn.count` | — | counter | `wake_from`, `result_kind`, `model` |
| `hyperhive.agent.session.count` | — | counter | `model` (incremented once per fresh, non-`--continue`'d session) |
| `hyperhive.agent.loose_ends.threads` | — | gauge | none |
| `hyperhive.agent.loose_ends.reminders` | — | gauge | none |
| `hyperhive.agent.claude_md.lines` | — | gauge | none — recorded from the `CLAUDE.md`-size watch's own ~15-minute tick, **not** per turn like the rows above |
Resource attributes (`service.name`, `agent`, `hive`, `swarm`) come from the
same container-wide `OTEL_RESOURCE_ATTRIBUTES` as everything else in this
section — nothing extra to configure. Cadence follows
`HYPERHIVE_OTEL_METRIC_INTERVAL_MS` (default 60s, same variable + default as
`hive-c0re`'s container-resource export above) — that only controls how often
the harness flushes the batched points to the collector, not how often it
records them (every turn, always).
## Hive-scoped metrics (hive-c0re)
Everything above is measured **per agent**, tagged with the hive it runs in.
These three are measured per **hive**, and carry no `agent` label — so a hive
that hosts no agents still reports, and "this hive is quiet" is
distinguishable from "this hive is gone." Select them with
`{hive!="",agent=""}`.
| Metric | Unit | Kind | Meaning |
| ------------------------- | ---- | ----- | -------------------------------------------------------------------------------------------------------------------- |
| `process.uptime` | `s` | gauge | seconds since this hive's `hive-c0re` started exporting; a restart reads as a drop to ~0 |
| `hyperhive.hive.degraded` | `1` | gauge | `1` while the hive reports itself unhealthy — the same verdict `/health/ready` gives and the swarm status view shows |
| `hyperhive.hive.warnings` | `1` | gauge | how many warnings are currently raised, split by a `level` attribute (`warn`, `crit`) |
hive-c0re reports both levels every cycle, `0` included, so a healthy hive is
visible as zeros rather than as missing series.
`hyperhive.hive.degraded` is what a dashboard should alert on: it's
`hive-c0re`'s own readiness verdict, so it stays in step with `/health/ready`
and with what the swarm controller sees. `hyperhive.hive.warnings` is the
detail behind it — `warn`-level entries mean "an operator should look" and do
**not** set `degraded`.
Same cadence, transport and resource labels as the container metrics above.
## VCS activity metrics (`swarm-controller`)
`swarm-controller` registers a single instance-wide Forgejo webhook (a
"global/system" hook, not scoped to any one org or repo) and counts commit
and push activity as deliveries arrive — occurrence-driven, not polled.
Forgejo's own native `/metrics` endpoint has no equivalent: it exposes
counts of durable rows (issues, comments, repos), and Forgejo doesn't
store either a commit or a push anywhere as a row to count.
| Metric | Unit | Kind | Attributes |
| ---------------------------- | ---- | ------- | ------------------- |
| `hyperhive.vcs.commit.count` | — | counter | `repo` (`org/repo`) |
| `hyperhive.vcs.push.count` | — | counter | `repo` (`org/repo`) |
A push with zero commits (a branch delete, or a force-push that doesn't add
new commits) still increments `push.count`; `commit.count` only advances
when the delivery actually carries commits. Same enable signal (`OTEL_EXPORTER_OTLP_ENDPOINT`), cadence variable
(`HYPERHIVE_OTEL_METRIC_INTERVAL_MS`) and `HYPERHIVE_OTEL_EXTRA_RESOURCE_ATTRIBUTES`
resource-attribute channel as `swarm-controller`'s other OTEL exporter (its
`hive-jobq-metrics`-backed job-graph rollup, undocumented here — see that
crate's own doc comment) — `swarm-controller` sets `service.name =
swarm-controller` directly rather than reading it from the container
environment, since it's a standalone daemon, not a per-agent harness process.
## Agent-emitted custom metrics (`hive-metric`)
Agents can push arbitrary labeled metrics to the same OTEL collector via the
`hive-metric` CLI tool, available in every agent container when
`services.hyperhive.otel.enable = true`.
### Usage
```text
hive-metric <name> <value> [--type counter|gauge] [--temporality delta|cumulative] [--labels key=value...]
```
- `<name>` — metric name (for example `tasks_completed`, `latency_ms`).
- `<value>` — numeric value (f64; integers and floats both accepted).
- `--type counter|gauge` — metric kind: `counter` (increasing sum, default) or
`gauge` (instantaneous point-in-time value).
- `--temporality delta|cumulative` — counter reporting mode (`counter` only,
ignored for `gauge`): `delta` (this call's own contribution, default — send
`1` each time and the collector accumulates) or `cumulative` (this call
reports the running total, which a stateless one-shot CLI can't track
itself).
- `--labels key=value` — extra per-data-point labels. May be repeated.
`hive-metric` inherits the resource labels (agent, hive, swarm, service.name)
automatically from `OTEL_RESOURCE_ATTRIBUTES` — don't re-specify them.
### Examples
```text
# Counter: one more task finished (delta is the default — no flag needed)
hive-metric tasks_completed 1 --labels phase=scan
# Gauge: current queue depth (absolute value — must use --type gauge)
hive-metric queue_depth 17 --type gauge
# Float gauge with multiple labels (instantaneous measurement)
hive-metric api_latency_ms 142.5 --type gauge --labels model=sonnet --labels tier=api
```
### Error when OTEL isn't configured
When `services.hyperhive.otel.enable = false` (the default), the
`OTEL_EXPORTER_OTLP_ENDPOINT` env var isn't set and `hive-metric` exits
with an informative error message. No silently dropped metrics.
### Wire format
`hive-metric` always uses **OTLP HTTP/JSON** (`application/json` POST to
`$OTEL_EXPORTER_OTLP_ENDPOINT/v1/metrics`), regardless of the
`OTEL_EXPORTER_OTLP_PROTOCOL` setting. `hive-metric` forwards auth headers from
`OTEL_EXPORTER_OTLP_HEADERS` verbatim.
## Metrics temporality
The harness configures OTEL export with **cumulative** temporality by default
(`OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative`),
overriding Claude Code's default of DELTA. This avoids silent metric drops in
Prometheus-family backends (including Grafana LGTM / Mimir) that don't ship a
delta-to-cumulative processor.
**`hive-metric` counters are the one exception**, reporting delta by default
(see above) — programmatically set on the exporter, which overrides this
container-wide env var for that tool specifically. `--type gauge` is
unaffected either way; gauges have no temporality. The hive-tier collector
runs a `deltatocumulative` processor ahead of export, so a delta
`hive-metric` counter still lands in VictoriaMetrics as a cumulative
series — the standard `rate()`/`increase()` idioms work on it exactly like
any other counter in this system, no special query needed.