hyperhive/docs/observability.md
atlas 16d578e692 docs(#3265): observability.md still said the endpoint was required
Review catch: this PR relaxed the `otel.endpoint` assertion and staled the
canonical OTEL reference in the same stroke — `docs/observability.md` is
what CLAUDE.md points readers at for "what OTEL options are available",
and it still said required-full-stop while the new swarm/services.md
section said a local store satisfies it.

Also corrects the option's own description in otel.nix, which said the
same thing and renders into the generated options doc. Grepping the
reviewer's phrasing did not find that one; grepping the claim did.

Records the second destination where the "endpoint is where telemetry
ultimately goes" paragraph makes its claim, rather than only in the new
section a reader may not reach.
2026-08-16 22:27:05 +02:00

307 lines
14 KiB
Markdown

# Observability (OpenTelemetry)
hyperhive has built-in support for exporting per-agent Claude Code statistics —
token usage, cost, tool call counts — to any OTLP-compatible collector via
Claude Code's built-in OpenTelemetry integration.
This is a **hive-wide** setting: one switch in the host NixOS config enables it
for every agent container simultaneously. There is no per-agent opt-in or opt-out.
## Enabling export
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel";
};
```
`enable` is the single gate. `endpoint` is where telemetry goes upstream —
required when enabled *unless* this host runs the swarm's own metrics store
(`swarm.victoriametrics.enable`), which is a destination in its own right. With
both, telemetry goes to both. See
[`swarm/services.md`](swarm/services.md#metrics-victoriametrics--grafana).
**There is exactly one way telemetry leaves a hive: through the collector that
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
they export unauthenticated to a bridge address only their own containers can
reach, and the collector forwards upstream with the auth header. So the
upstream credential exists in one place, on the host, and no agent ever holds
a copy.
⚠️ **The collector is therefore in the path of all telemetry.** It runs on the
same host as the agents and restarts on failure, and telemetry is not the
control plane — degraded telemetry is not degraded operation — but the export
no longer survives independently of anything host-side.
### what the agent→collector hop is and isn't
**It has no application-level auth.** The receiver takes any OTLP that reaches
it; what bounds who can reach it is the firewall — `exposeHostPorts` opens the
port on the bridge interface only. So "unauthenticated to a bridge address"
means *reachable from an agent container*, not *presents a credential*.
The consequence, stated because it is a choice rather than an oversight: **any
agent can push arbitrary OTLP, and the collector forwards it upstream under the
operator's credential.** It cannot tell a container's genuine Claude Code stats
from anything else shaped like OTLP arriving on that port — including data
smuggled out in resource attributes on an otherwise-legitimate export.
That is a **different risk from the one the collector fixes**, and strictly
smaller than what preceded it: before, every agent held the upstream credential
itself, so it could do all of the above *and* use the token anywhere else. The
collector removes the token and keeps the pipe. Agents are inside the trust
boundary (`docs/security.md`: capability = accepted risk), so an agent being
able to *send* is an accepted extension of that boundary — but it is not
closed by this design, and nothing here should be read as closing it.
## Options reference
### `services.hyperhive.otel.enable` — bool, default `false`
Master switch. When true, all other options below take effect.
### `services.hyperhive.otel.endpoint` — string, required when enabled unless the swarm store runs here
Upstream OTLP endpoint URL. Set as `OTEL_EXPORTER_OTLP_ENDPOINT` for every
agent. Example: `"https://collector.example.com/otel"`.
Leave it empty **only** on a host running `swarm.victoriametrics.enable` — the
local store is then the destination and the collector writes there instead.
With neither, `enable` is refused at eval: telemetry with nowhere to go is a
misconfiguration, not a quiet no-op.
### `services.hyperhive.otel.protocol` — enum, default `"http/protobuf"`
OTLP wire protocol, passed as `OTEL_EXPORTER_OTLP_PROTOCOL`. Accepted values:
- `"http/protobuf"` (default)
- `"http/json"`
- `"grpc"`
### `services.hyperhive.otel.headersCredential` — string or null, default `null`
Absolute path to a secret file on the host holding the upstream auth header as
`NAME=value` (e.g. `Authorization=Bearer <token>`).
**Only the host collector reads it.** It arrives as an `EnvironmentFile` on the
collector's unit, so the value is never read by nix, never copied into the
store or the generated config, never passed in argv — and **never forwarded
into an agent container**. An agent cannot read the hive's upstream credential
because it is never given one.
Leave `null` if the upstream needs no auth header; the collector then sends
none rather than an empty one.
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel";
headersCredential = "/run/secrets/otel-headers";
};
```
### `services.hyperhive.otel.extraResourceAttributes` — string, default `""`
Extra comma-separated entries appended to `OTEL_RESOURCE_ATTRIBUTES` after the
built-in labels (`service.name`, `agent`, `hive`, `swarm`). Example:
```nix
extraResourceAttributes = "deployment.environment=prod,team=platform";
```
### `services.hyperhive.otel.debug` — bool, default `false`
When `true`, sets `CLAUDE_CODE_OTEL_DIAG_STDERR=1` in every agent container,
causing the OTEL SDK to emit diagnostic messages to stderr. Useful when
troubleshooting collector connectivity or endpoint config errors. Leave `false`
in normal operation — SDK errors from a misconfigured endpoint would otherwise
appear in every agent's journal unconditionally.
Only meaningful when `enable` is true.
### `services.hyperhive.otel.metricIntervalMs` — positive int or null, default `null`
Metric export interval in milliseconds, set as `OTEL_METRIC_EXPORT_INTERVAL`
for every agent. Claude Code's default is 60000 (60 s). Leave `null` to keep
that default.
Each agent runs claude as a short-lived per-turn process; claude force-flushes
metrics on process exit, so interval tuning is not required for metrics to be
exported. A lower value gives more frequent intermediate flushes within
long-running turns — cosmetic, not a correctness knob.
## The host collector
`enable` starts an OpenTelemetry collector on the host. It is not optional and
there is no second path — that is the whole point:
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel"; # the upstream
headersCredential = "/run/secrets/otel-headers"; # only the host reads it
};
```
**Why it isn't a knob.** Exporting straight to `endpoint` means every agent
needs the credential to authenticate — and the harness delivers that token into
the agent's own `~/.claude/settings.json`, a file the agent can read. `0600`
protects it from other containers, not from the agent itself. As long as the
direct path stays *selectable*, that hole stays selectable; an option that can
reintroduce it is a hole with extra steps.
**`endpoint` keeps meaning "where telemetry goes upstream."** The collector
does not redefine it — the agent-facing value is *derived*
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
keeps working unchanged. The bridge port is contributed to `exposeHostPorts`
automatically; there is nothing to open by hand.
What the collector *added* is a second destination: on a host running the
swarm's metrics store it writes there too, so `endpoint` is no longer the only
place telemetry can land — and no longer the only way to have one.
### `services.hyperhive.otel.collector.port` — port, default `4318`
The OTLP/HTTP port the collector listens on, bound to the bridge IP only.
### `services.hyperhive.otel.collector.upstreamHeaderName` — string, default `"Authorization"`
Name of the header the collector sends upstream. The **value** comes from the
credential file at runtime (`EnvironmentFile``${env:<name>}`), never from
nix — so header names are config and header values are secrets, which is the
only split the collector's static header map can express.
⚠️ **`endpoint` must be valid for `protocol`.** The upstream exporter follows
`otel.protocol` (`grpc` → the gRPC exporter, otherwise OTLP/HTTP), and the gRPC
exporter takes an *address*: `https://host/path` is a legal
`OTEL_EXPORTER_OTLP_ENDPOINT` for HTTP but fails as gRPC with *"missing port in
address"*. The collector's config is validated at build time, so a mismatch is
a build error naming the reason rather than telemetry silently going nowhere.
## Network access
Agent containers can only reach the host on ports 80 and 443 by default. If
your OTLP collector runs on a non-standard port on the same host (e.g. a local
dev collector on `:4318`), open that port via:
```nix
services.hyperhive.network.exposeHostPorts = [ 4318 ];
```
Then point the endpoint at the bridge IP rather than loopback:
```nix
services.hyperhive.otel.endpoint = "http://10.42.0.1:4318";
```
The bridge IP is the host's address on the `hvbr0` bridge, typically
`10.42.0.1`. See `docs/network.md::Reaching host services` for details.
⚠️ **You do not need either line for hyperhive's own telemetry**`otel.enable`
contributes the collector's port and derives the endpoint itself. The above is
for pointing something *else* at a host-local service.
## Built-in resource labels
Every agent's export includes these resource attributes automatically:
| Attribute | Value |
|-----------|-------|
| `service.name` | `hyperhive-agent` (constant) |
| `agent` | agent logical name (e.g. `iris`) |
| `hive` | hive display name (`services.hyperhive.hiveName`) |
| `swarm` | swarm display name (`services.hyperhive.swarm.name`, if set) |
Additional labels can be appended via `extraResourceAttributes` (see option
reference above); custom per-data-point labels can be passed with
`hive-metric --labels` (see below).
## Host-emitted container-resource metrics (hive-c0re)
When OTEL is enabled, **hive-c0re itself** also exports each agent
container's resource load — the same cgroup gauges shown on the dashboard
LOAD tab — to the configured `endpoint`, reusing the same
`services.hyperhive.otel` config (no separate toggle). These come from the
host, not the in-container Claude SDK, so they cover containers even when
their agent is idle.
Emitted via the OpenTelemetry Rust SDK, using the
[semconv `container.*`](https://opentelemetry.io/docs/specs/semconv/system/container-metrics/)
metric names + the standard `container.name` attribute where a spec metric
exists, so off-the-shelf OTEL/Grafana container dashboards work. Resource
`service.name = hyperhive-c0re`; each data point is tagged `container.name`
(= the `h-<agent>` machine) and the hive `agent` label:
| Metric | Unit | Kind | Source |
|--------|------|------|--------|
| `container.cpu.time` | `s` | counter | cumulative `cpu.stat` `usage_usec` → seconds |
| `container.memory.usage` | `By` | gauge | `memory.current` |
| `hyperhive.container.memory.limit` | `By` | gauge | `memory.max` (custom — semconv has no `.limit` metric; omitted when unlimited) |
| `hyperhive.container.memory.peak` | `By` | gauge | `memory.peak` (custom — no semconv metric; omitted if unavailable) |
| `hyperhive.container.storage.usage` | `By` | gauge | state dir + writable rootfs (custom — semconv only has `disk.io`; omitted until the slow disk sampler runs) |
| `hyperhive.container.cpu.percent` | `%` | gauge | host-normalised percent (custom — the value the dashboard LOAD tab shows, no `rate()` needed) |
The `hyperhive.`-prefixed metrics have no semconv equivalent (memory
limit + peak, on-disk footprint, and an instantaneous cpu percent kept
alongside the spec `container.cpu.time` counter for convenience). Hive
labels (`hive`, `swarm`, …) ride on the resource via
`extraResourceAttributes`.
Cadence follows `metricIntervalMs` (default 60s). Transport is OTLP/HTTP
(JSON) to `<endpoint>`; the auth header is loaded onto hive-c0re's own unit
via systemd `LoadCredential` (from the same `headersCredential` file) and
sent as `Authorization`.
## Agent-emitted custom metrics (`hive-metric`)
Agents can push arbitrary labeled metrics to the same OTEL collector via the
`hive-metric` CLI tool, available in every agent container when
`services.hyperhive.otel.enable = true`.
### Usage
```text
hive-metric <name> <value> [--type counter|gauge] [--labels key=value...]
```
- `<name>` — metric name (e.g. `tasks_completed`, `latency_ms`).
- `<value>` — numeric value (f64; integers and floats both accepted).
- `--type counter|gauge` — metric kind: `counter` (cumulative sum, default) or
`gauge` (instantaneous point-in-time value).
- `--labels key=value` — extra per-data-point labels. May be repeated.
The resource labels (agent, hive, swarm, service.name) are inherited
automatically from `OTEL_RESOURCE_ATTRIBUTES` — do not re-specify them.
### Examples
```text
# Counter: cumulative tasks finished (default type — no --type flag needed)
hive-metric tasks_completed 1 --labels phase=scan
# Gauge: current queue depth (absolute value — must use --type gauge)
hive-metric queue_depth 17 --type gauge
# Float gauge with multiple labels (instantaneous measurement)
hive-metric api_latency_ms 142.5 --type gauge --labels model=sonnet --labels tier=api
```
### Error when OTEL is not configured
When `services.hyperhive.otel.enable = false` (the default), the
`OTEL_EXPORTER_OTLP_ENDPOINT` env var is not set and `hive-metric` exits
with an informative error message. No silently-dropped metrics.
### Wire format
`hive-metric` always uses **OTLP HTTP/JSON** (`application/json` POST to
`$OTEL_EXPORTER_OTLP_ENDPOINT/v1/metrics`), regardless of the
`OTEL_EXPORTER_OTLP_PROTOCOL` setting. Auth headers from
`OTEL_EXPORTER_OTLP_HEADERS` are forwarded verbatim.
## Metrics temporality
OTEL export is always configured with **cumulative** temporality
(`OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative`),
overriding Claude Code's default of DELTA. This avoids silent metric drops in
Prometheus-family backends (including Grafana LGTM / Mimir) that don't ship a
delta-to-cumulative processor.