6.5 KiB
Observability (OpenTelemetry)
hyperhive has built-in support for exporting per-agent Claude Code statistics — token usage, cost, tool call counts — to any OTLP-compatible collector via Claude Code's built-in OpenTelemetry integration.
This is a hive-wide setting: one switch in the host NixOS config enables it for every agent container simultaneously. There is no per-agent opt-in or opt-out.
Enabling export
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel";
};
enable is the single gate. endpoint (required when enabled) is the OTLP
HTTP endpoint; hive-c0re injects it as OTEL_EXPORTER_OTLP_ENDPOINT into
every agent's systemd service via the generated meta flake.
Each agent's harness (hive-ag3nt) exports directly to the collector — the pipeline keeps working even when hive-c0re is down.
Options reference
services.hyperhive.otel.enable — bool, default false
Master switch. When true, all other options below take effect.
services.hyperhive.otel.endpoint — string, required when enabled
OTLP collector endpoint URL. Set as OTEL_EXPORTER_OTLP_ENDPOINT for every
agent. Example: "https://collector.example.com/otel".
services.hyperhive.otel.protocol — enum, default "http/protobuf"
OTLP wire protocol, passed as OTEL_EXPORTER_OTLP_PROTOCOL. Accepted values:
"http/protobuf"(default)"http/json""grpc"
services.hyperhive.otel.headersCredential — string or null, default null
Absolute path to a secret file on the host whose contents become
OTEL_EXPORTER_OTLP_HEADERS (e.g. Authorization=Bearer <token>).
hive-c0re forwards this host file into each agent container via
systemd-nspawn --load-credential=otel-headers:<path>; the inner harness unit
inherits it by name. The token is never copied into the nix store, generated
config, a bind mount, or argv.
Leave null if the endpoint needs no auth header. A configured-but-missing
file is skipped with a log warning — OTEL still exports, just without the auth
header.
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel";
headersCredential = "/run/secrets/otel-headers";
};
services.hyperhive.otel.extraResourceAttributes — string, default ""
Extra comma-separated entries appended to OTEL_RESOURCE_ATTRIBUTES after the
built-in labels (service.name, agent, hive, swarm). Example:
extraResourceAttributes = "deployment.environment=prod,team=platform";
services.hyperhive.otel.debug — bool, default false
When true, sets CLAUDE_CODE_OTEL_DIAG_STDERR=1 in every agent container,
causing the OTEL SDK to emit diagnostic messages to stderr. Useful when
troubleshooting collector connectivity or endpoint config errors. Leave false
in normal operation — SDK errors from a misconfigured endpoint would otherwise
appear in every agent's journal unconditionally.
Only meaningful when enable is true.
services.hyperhive.otel.metricIntervalMs — positive int or null, default null
Metric export interval in milliseconds, set as OTEL_METRIC_EXPORT_INTERVAL
for every agent. Claude Code's default is 60000 (60 s). Leave null to keep
that default.
Each agent runs claude as a short-lived per-turn process; claude force-flushes metrics on process exit, so interval tuning is not required for metrics to be exported. A lower value gives more frequent intermediate flushes within long-running turns — cosmetic, not a correctness knob.
Network access
Agent containers can only reach the host on ports 80 and 443 by default. If
your OTLP collector runs on a non-standard port on the same host (e.g. a local
dev collector on :4318), open that port via:
services.hyperhive.network.exposeHostPorts = [ 4318 ];
Then point the endpoint at the bridge IP rather than loopback:
services.hyperhive.otel.endpoint = "http://10.42.0.1:4318";
The bridge IP is the host's address on the hvbr0 bridge, typically
10.42.0.1. See docs/network.md::Reaching host services for details.
Built-in resource labels
Every agent's export includes these resource attributes automatically:
| Attribute | Value |
|---|---|
service.name |
hyperhive-agent (constant) |
agent |
agent logical name (e.g. iris) |
hive |
hive display name (services.hyperhive.hiveName) |
swarm |
swarm display name (services.hyperhive.swarmName, if set) |
Additional labels can be appended via extraResourceAttributes (see option
reference above); custom per-data-point labels can be passed with
hive-metric --labels (see below).
Agent-emitted custom metrics (hive-metric)
Agents can push arbitrary labeled metrics to the same OTEL collector via the
hive-metric CLI tool, available in every agent container when
services.hyperhive.otel.enable = true.
Usage
hive-metric <name> <value> [--type gauge|counter] [--labels key=value...]
<name>— metric name (e.g.tasks_completed,latency_ms).<value>— numeric value (f64; integers and floats both accepted).--type gauge|counter— metric kind:gauge(instantaneous, default) orcounter(monotonically increasing cumulative sum).--labels key=value— extra per-data-point labels. May be repeated. The resource labels (agent, hive, swarm, service.name) are inherited automatically fromOTEL_RESOURCE_ATTRIBUTES— do not re-specify them.
Examples
# Gauge: current queue depth
hive-metric queue_depth 17
# Counter: cumulative tasks finished, with a custom label
hive-metric tasks_completed 1 --type counter --labels phase=scan
# Float gauge with multiple labels
hive-metric api_latency_ms 142.5 --labels model=sonnet --labels tier=api
Error when OTEL is not configured
When services.hyperhive.otel.enable = false (the default), the
OTEL_EXPORTER_OTLP_ENDPOINT env var is not set and hive-metric exits
with an informative error message. No silently-dropped metrics.
Wire format
hive-metric always uses OTLP HTTP/JSON (application/json POST to
$OTEL_EXPORTER_OTLP_ENDPOINT/v1/metrics), regardless of the
OTEL_EXPORTER_OTLP_PROTOCOL setting. Auth headers from
OTEL_EXPORTER_OTLP_HEADERS are forwarded verbatim.
Metrics temporality
OTEL export is always configured with cumulative temporality
(OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative),
overriding Claude Code's default of DELTA. This avoids silent metric drops in
Prometheus-family backends (including Grafana LGTM / Mimir) that don't ship a
delta-to-cumulative processor.