350 lines
17 KiB
Markdown
350 lines
17 KiB
Markdown
# Observability (OpenTelemetry)
|
|
|
|
hyperhive has built-in support for exporting per-agent Claude Code statistics —
|
|
token usage, cost, tool call counts — to any OTLP-compatible collector via
|
|
Claude Code's built-in OpenTelemetry integration.
|
|
|
|
This is a **hive-wide** setting: one switch in the host NixOS config enables it
|
|
for every agent container simultaneously. There is no per-agent opt-in or opt-out.
|
|
|
|
## Enabling export
|
|
|
|
```nix
|
|
services.hyperhive.otel = {
|
|
enable = true;
|
|
endpoint = "https://collector.example.com/otel";
|
|
};
|
|
```
|
|
|
|
`enable` is the single gate. `endpoint` is where telemetry ends up after it
|
|
leaves the swarm — optional, because the swarm's own metrics store
|
|
(`swarm.victoriametrics.enable`) is a destination in its own right. With both,
|
|
telemetry goes to both. See
|
|
[`swarm/services.md`](swarm/services.md#metrics-victoriametrics--grafana).
|
|
|
|
**There is exactly one way telemetry leaves a hive: through the collector that
|
|
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
|
|
they export unauthenticated to a bridge address only their own containers can
|
|
reach. That collector forwards to the swarm's
|
|
([`swarm/services.md`](swarm/services.md#telemetry-collector-otel)), which is
|
|
the single process holding the upstream credential and the only writer to the
|
|
swarm's store. No agent holds a copy, and neither does this hive.
|
|
|
|
The hive collector reaches the swarm collector by its gateway name
|
|
(`swarm.otel.domain`, default `otel.<swarm domain>`) — the same DNS-and-CA-trust
|
|
shape every hive-to-swarm-service hop uses, not a URL an operator has to point
|
|
anywhere. A hive that does not run the swarm's services still resolves that
|
|
name through the gateway; nothing here needs setting for the split-host case.
|
|
|
|
⚠️ **The collector is therefore in the path of all telemetry.** It runs on the
|
|
same host as the agents and restarts on failure, and telemetry is not the
|
|
control plane — degraded telemetry is not degraded operation — but the export
|
|
no longer survives independently of anything host-side.
|
|
|
|
### what the agent→collector hop is and isn't
|
|
|
|
**It has no application-level auth.** The receiver takes any OTLP that reaches
|
|
it; what bounds who can reach it is the firewall — `exposeHostPorts` opens the
|
|
port on the bridge interface only. So "unauthenticated to a bridge address"
|
|
means *reachable from an agent container*, not *presents a credential*.
|
|
|
|
The consequence, stated because it is a choice rather than an oversight: **any
|
|
agent can push arbitrary OTLP, and it is forwarded on under the operator's
|
|
credential.** Neither tier can tell a container's genuine Claude Code stats
|
|
from anything else shaped like OTLP arriving on that port — including data
|
|
smuggled out in resource attributes on an otherwise-legitimate export.
|
|
|
|
That is a **different risk from the one the collector fixes**, and strictly
|
|
smaller than what preceded it: before, every agent held the upstream credential
|
|
itself, so it could do all of the above *and* use the token anywhere else. The
|
|
collector removes the token and keeps the pipe. Agents are inside the trust
|
|
boundary (`docs/security.md`: capability = accepted risk), so an agent being
|
|
able to *send* is an accepted extension of that boundary — but it is not
|
|
closed by this design, and nothing here should be read as closing it.
|
|
|
|
## Options reference
|
|
|
|
Every `services.hyperhive.otel.*` option's full type/default/description/
|
|
example is generated straight from the nix module (`nix/host-modules/
|
|
otel.nix`) into [`/options/`](/options/) (host options — `nix build
|
|
.#docs-host` for a local render). That page is kept honest by the build in a
|
|
way a hand-copied version here cannot be, so it is the reference, not this
|
|
doc. What follows is what a flat per-option listing can't express: the
|
|
two-tier architecture, the security model, and how the options interact.
|
|
|
|
## The two collectors
|
|
|
|
Telemetry crosses two collectors, and which one you configure depends on what
|
|
the host is:
|
|
|
|
| | runs where | receives from | does |
|
|
|---|---|---|---|
|
|
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
|
|
| **swarm tier** — `swarm.otel.enable` | once per swarm | every hive's collector | writes the swarm's store and exports upstream |
|
|
|
|
An all-local host runs both, and needs nothing said about the hop between them.
|
|
|
|
```nix
|
|
services.hyperhive.otel = {
|
|
enable = true;
|
|
endpoint = "https://collector.example.com/otel"; # the upstream
|
|
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
|
|
};
|
|
```
|
|
|
|
**Why the hive tier isn't optional.** Exporting straight to `endpoint` means
|
|
every agent needs the credential to authenticate — and the harness delivers
|
|
that token into the agent's own `~/.claude/settings.json`, a file the agent can
|
|
read. `0600` protects it from other containers, not from the agent itself. As
|
|
long as the direct path stays *selectable*, that hole stays selectable; an
|
|
option that can reintroduce it is a hole with extra steps.
|
|
|
|
**Why the tiers stay separate on one box.** They are not collapsed when
|
|
co-located: an all-local hive is a statement about *where* processes run, not
|
|
about the shape of the deployment. A boundary that disappears locally is one
|
|
the local deployment stops testing.
|
|
|
|
**`endpoint` keeps meaning "where telemetry goes upstream."** Neither tier
|
|
redefines it — the agent-facing value is *derived*
|
|
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
|
|
keeps working unchanged. The bridge port is contributed to `exposeHostPorts`
|
|
automatically; there is nothing to open by hand.
|
|
|
|
### Authenticated ingest
|
|
|
|
The swarm tier gives **each hive its own receiver**, and stamps the `hive` label
|
|
from whichever receiver accepted a sample. A hive therefore cannot report
|
|
metrics as another hive, and cannot relabel its own by editing what it sends —
|
|
the label is not taken from the payload at all.
|
|
|
|
**On an all-local swarm there is nothing to set.** Each hive already has an
|
|
identity, and its collector reads the secret that host's own authelia minted.
|
|
|
|
**On a hive that does not host the swarm's services**, the secret has to arrive
|
|
somehow — copy it across and name it:
|
|
|
|
```nix
|
|
services.hyperhive.otel.clientSecretFile = "/run/secrets/hive-telemetry.secret";
|
|
```
|
|
|
|
**There is no unauthenticated mode.** A hive always presents an identity, so a
|
|
missing credential is a build error rather than a quieter fallback — the
|
|
collector has no anonymous route to accept samples on, and every path it serves
|
|
belongs to exactly one hive.
|
|
|
|
Getting the secret wrong shows up as the hive's collector logging 401s from the
|
|
swarm tier and no metrics appearing for that hive.
|
|
|
|
⚠️ **`endpoint` must be valid for `protocol`.** The upstream exporter follows
|
|
`otel.protocol` (`grpc` → the gRPC exporter, otherwise OTLP/HTTP), and the gRPC
|
|
exporter takes an *address*: `https://host/path` is a legal
|
|
`OTEL_EXPORTER_OTLP_ENDPOINT` for HTTP but fails as gRPC with *"missing port in
|
|
address"*. The collector's config is validated at build time, so a mismatch is
|
|
a build error naming the reason rather than telemetry silently going nowhere.
|
|
|
|
## Network access
|
|
|
|
Agent containers can only reach the host on ports 80 and 443 by default. To let
|
|
them reach some other host-local service you run yourself — a database, a
|
|
scratch HTTP endpoint — open its port on the bridge:
|
|
|
|
```nix
|
|
services.hyperhive.network.exposeHostPorts = [ 5432 ];
|
|
```
|
|
|
|
and point whatever consumes it at `10.42.0.1:5432` rather than loopback: inside
|
|
a container, loopback is the *container*. The bridge IP is the host's address on
|
|
the `hive-br0` bridge. The service must also bind an address the bridge can
|
|
reach — a `127.0.0.1`-only listener stays unreachable no matter what the
|
|
firewall allows. See `docs/network.md::Reaching host services` for details.
|
|
|
|
⚠️ **None of this is needed for hyperhive's own telemetry** — `otel.enable`
|
|
contributes the collector's port and derives the agent-facing endpoint itself.
|
|
|
|
## Built-in resource labels
|
|
|
|
The OTLP variables (`OTEL_EXPORTER_OTLP_ENDPOINT`, `_PROTOCOL`,
|
|
`OTEL_RESOURCE_ATTRIBUTES`, the temporality preference) are set **container
|
|
wide** — in systemd's `DefaultEnvironment` and in `/etc/profile` — so every
|
|
process in an agent container exports to the hive's collector without any
|
|
per-tool wiring. That covers Claude Code, `hive-metric`, and anything you run
|
|
yourself from a tool call or `hivectl shell`.
|
|
|
|
Every agent's export therefore includes these resource attributes
|
|
automatically:
|
|
|
|
| Attribute | Value |
|
|
|-----------|-------|
|
|
| `service.name` | `hyperhive-agent` (constant) |
|
|
| `agent` | agent logical name (e.g. `iris`) |
|
|
| `hive` | hive display name (`services.hyperhive.hiveName`) |
|
|
| `swarm` | swarm display name (`services.hyperhive.swarm.name`, if set) |
|
|
|
|
Additional labels can be appended via `extraResourceAttributes` (see option
|
|
reference above); custom per-data-point labels can be passed with
|
|
`hive-metric --labels` (see below).
|
|
|
|
## Host-emitted container-resource metrics (hive-c0re)
|
|
|
|
When OTEL is enabled, **hive-c0re itself** also exports each agent
|
|
container's resource load — the same cgroup gauges shown on the dashboard
|
|
LOAD tab — to this hive's own collector, exactly like an agent does and with
|
|
no separate toggle. These come from the host, not the in-container Claude SDK,
|
|
so they cover containers even when their agent is idle.
|
|
|
|
Emitted via the OpenTelemetry Rust SDK, using the
|
|
[semconv `container.*`](https://opentelemetry.io/docs/specs/semconv/system/container-metrics/)
|
|
metric names + the standard `container.name` attribute where a spec metric
|
|
exists, so off-the-shelf OTEL/Grafana container dashboards work. Resource
|
|
`service.name = hyperhive-c0re`; each data point is tagged `container.name`
|
|
(= the `h-<agent>` machine) and the hive `agent` label:
|
|
|
|
| Metric | Unit | Kind | Source |
|
|
|--------|------|------|--------|
|
|
| `container.cpu.time` | `s` | counter | cumulative `cpu.stat` `usage_usec` → seconds |
|
|
| `container.memory.usage` | `By` | gauge | `memory.current` |
|
|
| `hyperhive.container.memory.limit` | `By` | gauge | `memory.max` (custom — semconv has no `.limit` metric; omitted when unlimited) |
|
|
| `hyperhive.container.memory.peak` | `By` | gauge | `memory.peak` (custom — no semconv metric; omitted if unavailable) |
|
|
| `hyperhive.container.storage.usage` | `By` | gauge | state dir + writable rootfs (custom — semconv only has `disk.io`; omitted until the slow disk sampler runs) |
|
|
| `hyperhive.container.cpu.percent` | `%` | gauge | host-normalised percent (custom — the value the dashboard LOAD tab shows, no `rate()` needed) |
|
|
|
|
The `hyperhive.`-prefixed metrics have no semconv equivalent (memory
|
|
limit + peak, on-disk footprint, and an instantaneous cpu percent kept
|
|
alongside the spec `container.cpu.time` counter for convenience). Hive
|
|
labels (`hive`, `swarm`, …) ride on the resource via
|
|
`extraResourceAttributes`.
|
|
|
|
Cadence follows `metricIntervalMs` (default 60s). Transport is OTLP/HTTP
|
|
(JSON) to the hive collector's bridge address, with no auth header — that
|
|
first hop is unauthenticated for every producer on this host, and the upstream
|
|
credential stays on the swarm tier.
|
|
|
|
## Agent-emitted per-turn metrics (`hive-agent`)
|
|
|
|
When OTEL is enabled, the harness itself (`hive-agent`) exports one small set
|
|
of metrics per claude turn, recorded the moment the turn ends (not polled).
|
|
These are deliberately the fields Claude Code's own built-in export (see
|
|
above) can't know about — the harness's own wall-clock timing, what woke the
|
|
turn, its own outcome classification, the loose-ends backlog, and session
|
|
boundaries. Token usage, cost, and tool-call counts are **not** duplicated
|
|
here; that's already covered by Claude's own export.
|
|
|
|
| Metric | Unit | Kind | Attributes |
|
|
|--------|------|------|------------|
|
|
| `hyperhive.agent.turn.duration` | `ms` | histogram | `wake_from`, `result_kind`, `model` |
|
|
| `hyperhive.agent.turn.count` | — | counter | `wake_from`, `result_kind`, `model` |
|
|
| `hyperhive.agent.session.count` | — | counter | `model` (incremented once per fresh, non-`--continue`'d session) |
|
|
| `hyperhive.agent.loose_ends.threads` | — | gauge | none |
|
|
| `hyperhive.agent.loose_ends.reminders` | — | gauge | none |
|
|
| `hyperhive.agent.claude_md.lines` | — | gauge | none — recorded from the `CLAUDE.md`-size watch's own ~15-minute tick, **not** per turn like the rows above |
|
|
|
|
Resource attributes (`service.name`, `agent`, `hive`, `swarm`) come from the
|
|
same container-wide `OTEL_RESOURCE_ATTRIBUTES` as everything else in this
|
|
section — nothing extra to configure. Cadence follows
|
|
`HYPERHIVE_OTEL_METRIC_INTERVAL_MS` (default 60s, same variable + default as
|
|
`hive-c0re`'s container-resource export above) — that only controls how often
|
|
the batched points are flushed to the collector, not how often they're
|
|
recorded (every turn, always).
|
|
|
|
## Hive-scoped metrics (hive-c0re)
|
|
|
|
Everything above is measured **per agent**, tagged with the hive it runs in.
|
|
These three are measured per **hive**, and carry no `agent` label — so a hive
|
|
that hosts no agents still reports, and "this hive is quiet" is
|
|
distinguishable from "this hive is gone". Select them with
|
|
`{hive!="",agent=""}`.
|
|
|
|
| Metric | Unit | Kind | Meaning |
|
|
|--------|------|------|---------|
|
|
| `process.uptime` | `s` | gauge | seconds since this hive's `hive-c0re` started exporting; a restart reads as a drop to ~0 |
|
|
| `hyperhive.hive.degraded` | `1` | gauge | `1` while the hive reports itself unhealthy — the same verdict `/health/ready` gives and the swarm status view shows |
|
|
| `hyperhive.hive.warnings` | `1` | gauge | how many warnings are currently raised, split by a `level` attribute (`warn`, `crit`) |
|
|
|
|
Both levels are reported every cycle, `0` included, so a healthy hive is
|
|
visible as zeros rather than as missing series.
|
|
|
|
`hyperhive.hive.degraded` is what a dashboard should alert on: it is
|
|
`hive-c0re`'s own readiness verdict, so it stays in step with `/health/ready`
|
|
and with what the swarm controller sees. `hyperhive.hive.warnings` is the
|
|
detail behind it — `warn`-level entries mean "an operator should look" and do
|
|
**not** set `degraded`.
|
|
|
|
Same cadence, transport and resource labels as the container metrics above.
|
|
|
|
## VCS activity metrics (`swarm-controller`)
|
|
|
|
`swarm-controller` registers a single instance-wide Forgejo webhook (a
|
|
"global/system" hook, not scoped to any one org or repo) and counts commit
|
|
and push activity as deliveries arrive — occurrence-driven, not polled.
|
|
Forgejo's own native `/metrics` endpoint has no equivalent: it exposes
|
|
counts of durable rows (issues, comments, repos), and neither a commit nor a
|
|
push is stored anywhere as a row to count.
|
|
|
|
| Metric | Unit | Kind | Attributes |
|
|
|--------|------|------|------------|
|
|
| `hyperhive.vcs.commit.count` | — | counter | `repo` (`org/repo`) |
|
|
| `hyperhive.vcs.push.count` | — | counter | `repo` (`org/repo`) |
|
|
|
|
A push with zero commits (a branch delete, or a force-push that doesn't add
|
|
new commits) still increments `push.count`; `commit.count` only advances
|
|
when the delivery actually carries commits. Same enable signal (`OTEL_EXPORTER_OTLP_ENDPOINT`), cadence variable
|
|
(`HYPERHIVE_OTEL_METRIC_INTERVAL_MS`) and `HYPERHIVE_OTEL_EXTRA_RESOURCE_ATTRIBUTES`
|
|
resource-attribute channel as `swarm-controller`'s other OTEL exporter (its
|
|
`hive-jobq-metrics`-backed job-graph rollup, undocumented here — see that
|
|
crate's own doc comment) — `service.name = swarm-controller` is set
|
|
directly rather than read from the container environment, since
|
|
`swarm-controller` is a standalone daemon, not a per-agent harness process.
|
|
|
|
## Agent-emitted custom metrics (`hive-metric`)
|
|
|
|
Agents can push arbitrary labeled metrics to the same OTEL collector via the
|
|
`hive-metric` CLI tool, available in every agent container when
|
|
`services.hyperhive.otel.enable = true`.
|
|
|
|
### Usage
|
|
|
|
```text
|
|
hive-metric <name> <value> [--type counter|gauge] [--labels key=value...]
|
|
```
|
|
|
|
- `<name>` — metric name (e.g. `tasks_completed`, `latency_ms`).
|
|
- `<value>` — numeric value (f64; integers and floats both accepted).
|
|
- `--type counter|gauge` — metric kind: `counter` (cumulative sum, default) or
|
|
`gauge` (instantaneous point-in-time value).
|
|
- `--labels key=value` — extra per-data-point labels. May be repeated.
|
|
The resource labels (agent, hive, swarm, service.name) are inherited
|
|
automatically from `OTEL_RESOURCE_ATTRIBUTES` — do not re-specify them.
|
|
|
|
### Examples
|
|
|
|
```text
|
|
# Counter: cumulative tasks finished (default type — no --type flag needed)
|
|
hive-metric tasks_completed 1 --labels phase=scan
|
|
|
|
# Gauge: current queue depth (absolute value — must use --type gauge)
|
|
hive-metric queue_depth 17 --type gauge
|
|
|
|
# Float gauge with multiple labels (instantaneous measurement)
|
|
hive-metric api_latency_ms 142.5 --type gauge --labels model=sonnet --labels tier=api
|
|
```
|
|
|
|
### Error when OTEL is not configured
|
|
|
|
When `services.hyperhive.otel.enable = false` (the default), the
|
|
`OTEL_EXPORTER_OTLP_ENDPOINT` env var is not set and `hive-metric` exits
|
|
with an informative error message. No silently-dropped metrics.
|
|
|
|
### Wire format
|
|
|
|
`hive-metric` always uses **OTLP HTTP/JSON** (`application/json` POST to
|
|
`$OTEL_EXPORTER_OTLP_ENDPOINT/v1/metrics`), regardless of the
|
|
`OTEL_EXPORTER_OTLP_PROTOCOL` setting. Auth headers from
|
|
`OTEL_EXPORTER_OTLP_HEADERS` are forwarded verbatim.
|
|
|
|
## Metrics temporality
|
|
|
|
OTEL export is always configured with **cumulative** temporality
|
|
(`OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative`),
|
|
overriding Claude Code's default of DELTA. This avoids silent metric drops in
|
|
Prometheus-family backends (including Grafana LGTM / Mimir) that don't ship a
|
|
delta-to-cumulative processor.
|