hyperhive/docs/observability.md
atlas 865a1cc9ac docs(otel): say what the agent->collector hop is and is not
argus on #3280: the receiver has no auth extension - the nixpkgs module
passes settings straight through and nothing wires one on - so
'unauthenticated to a bridge address' means reachable from an agent
container, not presents a credential. The firewall is the whole access
control.

Consequence, stated because it is a choice rather than an oversight:
any agent can push arbitrary OTLP and the collector forwards it under
the operator's credential, including data smuggled out in resource
attributes. That is a different risk from the one the collector fixes,
and strictly smaller than what preceded it - before, every agent held
the credential itself and could do all of that plus use the token
anywhere else. The collector removes the token and keeps the pipe.

Same principle this PR already applies to the availability trade: state
it where the reader is, rather than let it be discovered.
2026-08-15 11:46:24 +02:00

13 KiB

Observability (OpenTelemetry)

hyperhive has built-in support for exporting per-agent Claude Code statistics — token usage, cost, tool call counts — to any OTLP-compatible collector via Claude Code's built-in OpenTelemetry integration.

This is a hive-wide setting: one switch in the host NixOS config enables it for every agent container simultaneously. There is no per-agent opt-in or opt-out.

Enabling export

services.hyperhive.otel = {
  enable   = true;
  endpoint = "https://collector.example.com/otel";
};

enable is the single gate. endpoint (required when enabled) is where telemetry ultimately goes.

There is exactly one way telemetry leaves a hive: through the collector that enable starts on the host. Agents never talk to endpoint themselves — they export unauthenticated to a bridge address only their own containers can reach, and the collector forwards upstream with the auth header. So the upstream credential exists in one place, on the host, and no agent ever holds a copy.

⚠️ The collector is therefore in the path of all telemetry. It runs on the same host as the agents and restarts on failure, and telemetry is not the control plane — degraded telemetry is not degraded operation — but the export no longer survives independently of anything host-side.

what the agent→collector hop is and isn't

It has no application-level auth. The receiver takes any OTLP that reaches it; what bounds who can reach it is the firewall — exposeHostPorts opens the port on the bridge interface only. So "unauthenticated to a bridge address" means reachable from an agent container, not presents a credential.

The consequence, stated because it is a choice rather than an oversight: any agent can push arbitrary OTLP, and the collector forwards it upstream under the operator's credential. It cannot tell a container's genuine Claude Code stats from anything else shaped like OTLP arriving on that port — including data smuggled out in resource attributes on an otherwise-legitimate export.

That is a different risk from the one the collector fixes, and strictly smaller than what preceded it: before, every agent held the upstream credential itself, so it could do all of the above and use the token anywhere else. The collector removes the token and keeps the pipe. Agents are inside the trust boundary (docs/security.md: capability = accepted risk), so an agent being able to send is an accepted extension of that boundary — but it is not closed by this design, and nothing here should be read as closing it.

Options reference

services.hyperhive.otel.enable — bool, default false

Master switch. When true, all other options below take effect.

services.hyperhive.otel.endpoint — string, required when enabled

OTLP collector endpoint URL. Set as OTEL_EXPORTER_OTLP_ENDPOINT for every agent. Example: "https://collector.example.com/otel".

services.hyperhive.otel.protocol — enum, default "http/protobuf"

OTLP wire protocol, passed as OTEL_EXPORTER_OTLP_PROTOCOL. Accepted values:

  • "http/protobuf" (default)
  • "http/json"
  • "grpc"

services.hyperhive.otel.headersCredential — string or null, default null

Absolute path to a secret file on the host holding the upstream auth header as NAME=value (e.g. Authorization=Bearer <token>).

Only the host collector reads it. It arrives as an EnvironmentFile on the collector's unit, so the value is never read by nix, never copied into the store or the generated config, never passed in argv — and never forwarded into an agent container. An agent cannot read the hive's upstream credential because it is never given one.

Leave null if the upstream needs no auth header; the collector then sends none rather than an empty one.

services.hyperhive.otel = {
  enable           = true;
  endpoint         = "https://collector.example.com/otel";
  headersCredential = "/run/secrets/otel-headers";
};

services.hyperhive.otel.extraResourceAttributes — string, default ""

Extra comma-separated entries appended to OTEL_RESOURCE_ATTRIBUTES after the built-in labels (service.name, agent, hive, swarm). Example:

extraResourceAttributes = "deployment.environment=prod,team=platform";

services.hyperhive.otel.debug — bool, default false

When true, sets CLAUDE_CODE_OTEL_DIAG_STDERR=1 in every agent container, causing the OTEL SDK to emit diagnostic messages to stderr. Useful when troubleshooting collector connectivity or endpoint config errors. Leave false in normal operation — SDK errors from a misconfigured endpoint would otherwise appear in every agent's journal unconditionally.

Only meaningful when enable is true.

services.hyperhive.otel.metricIntervalMs — positive int or null, default null

Metric export interval in milliseconds, set as OTEL_METRIC_EXPORT_INTERVAL for every agent. Claude Code's default is 60000 (60 s). Leave null to keep that default.

Each agent runs claude as a short-lived per-turn process; claude force-flushes metrics on process exit, so interval tuning is not required for metrics to be exported. A lower value gives more frequent intermediate flushes within long-running turns — cosmetic, not a correctness knob.

The host collector

enable starts an OpenTelemetry collector on the host. It is not optional and there is no second path — that is the whole point:

services.hyperhive.otel = {
  enable            = true;
  endpoint          = "https://collector.example.com/otel";  # the upstream
  headersCredential = "/run/secrets/otel-headers";           # only the host reads it
};

Why it isn't a knob. Exporting straight to endpoint means every agent needs the credential to authenticate — and the harness delivers that token into the agent's own ~/.claude/settings.json, a file the agent can read. 0600 protects it from other containers, not from the agent itself. As long as the direct path stays selectable, that hole stays selectable; an option that can reintroduce it is a hole with extra steps.

endpoint keeps meaning "where telemetry ultimately goes." The collector does not redefine it — the agent-facing value is derived (http://<bridgeIp>:<collector.port>), so an existing deployment's endpoint keeps working unchanged. The bridge port is contributed to exposeHostPorts automatically; there is nothing to open by hand.

services.hyperhive.otel.collector.port — port, default 4318

The OTLP/HTTP port the collector listens on, bound to the bridge IP only.

services.hyperhive.otel.collector.upstreamHeaderName — string, default "Authorization"

Name of the header the collector sends upstream. The value comes from the credential file at runtime (EnvironmentFile${env:<name>}), never from nix — so header names are config and header values are secrets, which is the only split the collector's static header map can express.

⚠️ endpoint must be valid for protocol. The upstream exporter follows otel.protocol (grpc → the gRPC exporter, otherwise OTLP/HTTP), and the gRPC exporter takes an address: https://host/path is a legal OTEL_EXPORTER_OTLP_ENDPOINT for HTTP but fails as gRPC with "missing port in address". The collector's config is validated at build time, so a mismatch is a build error naming the reason rather than telemetry silently going nowhere.

Network access

Agent containers can only reach the host on ports 80 and 443 by default. If your OTLP collector runs on a non-standard port on the same host (e.g. a local dev collector on :4318), open that port via:

services.hyperhive.network.exposeHostPorts = [ 4318 ];

Then point the endpoint at the bridge IP rather than loopback:

services.hyperhive.otel.endpoint = "http://10.42.0.1:4318";

The bridge IP is the host's address on the hvbr0 bridge, typically 10.42.0.1. See docs/network.md::Reaching host services for details.

⚠️ You do not need either line for hyperhive's own telemetryotel.enable contributes the collector's port and derives the endpoint itself. The above is for pointing something else at a host-local service.

Built-in resource labels

Every agent's export includes these resource attributes automatically:

Attribute Value
service.name hyperhive-agent (constant)
agent agent logical name (e.g. iris)
hive hive display name (services.hyperhive.hiveName)
swarm swarm display name (services.hyperhive.swarm.name, if set)

Additional labels can be appended via extraResourceAttributes (see option reference above); custom per-data-point labels can be passed with hive-metric --labels (see below).

Host-emitted container-resource metrics (hive-c0re)

When OTEL is enabled, hive-c0re itself also exports each agent container's resource load — the same cgroup gauges shown on the dashboard LOAD tab — to the configured endpoint, reusing the same services.hyperhive.otel config (no separate toggle). These come from the host, not the in-container Claude SDK, so they cover containers even when their agent is idle.

Emitted via the OpenTelemetry Rust SDK, using the semconv container.* metric names + the standard container.name attribute where a spec metric exists, so off-the-shelf OTEL/Grafana container dashboards work. Resource service.name = hyperhive-c0re; each data point is tagged container.name (= the h-<agent> machine) and the hive agent label:

Metric Unit Kind Source
container.cpu.time s counter cumulative cpu.stat usage_usec → seconds
container.memory.usage By gauge memory.current
hyperhive.container.memory.limit By gauge memory.max (custom — semconv has no .limit metric; omitted when unlimited)
hyperhive.container.memory.peak By gauge memory.peak (custom — no semconv metric; omitted if unavailable)
hyperhive.container.storage.usage By gauge state dir + writable rootfs (custom — semconv only has disk.io; omitted until the slow disk sampler runs)
hyperhive.container.cpu.percent % gauge host-normalised percent (custom — the value the dashboard LOAD tab shows, no rate() needed)

The hyperhive.-prefixed metrics have no semconv equivalent (memory limit + peak, on-disk footprint, and an instantaneous cpu percent kept alongside the spec container.cpu.time counter for convenience). Hive labels (hive, swarm, …) ride on the resource via extraResourceAttributes.

Cadence follows metricIntervalMs (default 60s). Transport is OTLP/HTTP (JSON) to <endpoint>; the auth header is loaded onto hive-c0re's own unit via systemd LoadCredential (from the same headersCredential file) and sent as Authorization.

Agent-emitted custom metrics (hive-metric)

Agents can push arbitrary labeled metrics to the same OTEL collector via the hive-metric CLI tool, available in every agent container when services.hyperhive.otel.enable = true.

Usage

hive-metric <name> <value> [--type counter|gauge] [--labels key=value...]
  • <name> — metric name (e.g. tasks_completed, latency_ms).
  • <value> — numeric value (f64; integers and floats both accepted).
  • --type counter|gauge — metric kind: counter (cumulative sum, default) or gauge (instantaneous point-in-time value).
  • --labels key=value — extra per-data-point labels. May be repeated. The resource labels (agent, hive, swarm, service.name) are inherited automatically from OTEL_RESOURCE_ATTRIBUTES — do not re-specify them.

Examples

# Counter: cumulative tasks finished (default type — no --type flag needed)
hive-metric tasks_completed 1 --labels phase=scan

# Gauge: current queue depth (absolute value — must use --type gauge)
hive-metric queue_depth 17 --type gauge

# Float gauge with multiple labels (instantaneous measurement)
hive-metric api_latency_ms 142.5 --type gauge --labels model=sonnet --labels tier=api

Error when OTEL is not configured

When services.hyperhive.otel.enable = false (the default), the OTEL_EXPORTER_OTLP_ENDPOINT env var is not set and hive-metric exits with an informative error message. No silently-dropped metrics.

Wire format

hive-metric always uses OTLP HTTP/JSON (application/json POST to $OTEL_EXPORTER_OTLP_ENDPOINT/v1/metrics), regardless of the OTEL_EXPORTER_OTLP_PROTOCOL setting. Auth headers from OTEL_EXPORTER_OTLP_HEADERS are forwarded verbatim.

Metrics temporality

OTEL export is always configured with cumulative temporality (OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative), overriding Claude Code's default of DELTA. This avoids silent metric drops in Prometheus-family backends (including Grafana LGTM / Mimir) that don't ship a delta-to-cumulative processor.