feat(#2007): export per-agent container cpu/mem/disk via otel
hive-c0re already samples each agent container's cgroup load for the dashboard (stats/container_stats.rs); this rides those gauges out to the configured OTLP endpoint, reusing the existing services.hyperhive.otel config (endpoint + auth header) — no new toggle. - New stats/otel_metrics.rs: exports via the OpenTelemetry Rust SDK (same crates as hive-metric) with the semconv container.* metric names + container.name attribute so off-the-shelf OTel/Grafana dashboards work, plus the hive agent label. container.cpu.time (counter, s, from cumulative cpu.stat usage_usec), container.memory.usage, container.memory.usage.limit; memory peak / on-disk storage / instantaneous cpu percent stay hyperhive.* custom (no semconv equivalent). Observable instruments read a shared snapshot an async task refreshes (gather() is async; SDK callbacks sync). - container_stats: expose cpu_time_usec (cumulative) on ContainerResource. - The OTLP auth header is loaded onto hive-c0re's own unit via systemd LoadCredential and read from $CREDENTIALS_DIRECTORY/otel-headers. - docs/observability.md documents the host-emitted semconv metrics. Host-side export, so it covers containers even when their agent is idle.
This commit is contained in:
parent
c238ffe1ff
commit
419c9659a3
9 changed files with 396 additions and 2 deletions
3
Cargo.lock
generated
3
Cargo.lock
generated
|
|
@ -1556,6 +1556,9 @@ dependencies = [
|
|||
"indicatif",
|
||||
"libc",
|
||||
"listenfd",
|
||||
"opentelemetry",
|
||||
"opentelemetry-otlp",
|
||||
"opentelemetry_sdk",
|
||||
"petgraph",
|
||||
"problem_details",
|
||||
"reqwest 0.12.28",
|
||||
|
|
|
|||
Loading…
Reference in a new issue