feat(hive-c0re): export metrics whose subject is the hive, not an agent

Every hive-labelled series in the store also carries an agent label, so a
hive is only ever visible as the sum of its agents — and a hive whose c0re
has stopped is indistinguishable from one that simply hosts none.

Adds three instruments to the exporter hive-c0re already runs, each a
projection of a value the process computes anyway: process.uptime (the
semconv name — the spec defines it as a double gauge in seconds, which is
exactly this instrument), hyperhive.hive.degraded, and
hyperhive.hive.warnings split by level. None carries an agent attribute;
that absence is what makes them selectable as hive-scoped.

The health pair reads warnings::readiness() rather than deriving its own
verdict, and degraded ships as a series instead of being left for a
dashboard query to compute from warnings{level="crit"} — either would put
the "what counts as unhealthy" rule in a second place that disagrees
silently the first time a degrading condition is added.
This commit is contained in:
atlas 2026-08-24 19:13:09 +02:00 committed by mara
commit 5eca0cc516
3 changed files with 222 additions and 2 deletions

View file

@ -245,6 +245,31 @@ section — nothing extra to configure. Cadence follows
the batched points are flushed to the collector, not how often they're the batched points are flushed to the collector, not how often they're
recorded (every turn, always). recorded (every turn, always).
## Hive-scoped metrics (hive-c0re)
Everything above is measured **per agent**, tagged with the hive it runs in.
These three are measured per **hive**, and carry no `agent` label — so a hive
that hosts no agents still reports, and "this hive is quiet" is
distinguishable from "this hive is gone". Select them with
`{hive!="",agent=""}`.
| Metric | Unit | Kind | Meaning |
|--------|------|------|---------|
| `process.uptime` | `s` | gauge | seconds since this hive's `hive-c0re` started exporting; a restart reads as a drop to ~0 |
| `hyperhive.hive.degraded` | `1` | gauge | `1` while the hive reports itself unhealthy — the same verdict `/health/ready` gives and the swarm status view shows |
| `hyperhive.hive.warnings` | `1` | gauge | how many warnings are currently raised, split by a `level` attribute (`warn`, `crit`) |
Both levels are reported every cycle, `0` included, so a healthy hive is
visible as zeros rather than as missing series.
`hyperhive.hive.degraded` is what a dashboard should alert on: it is
`hive-c0re`'s own readiness verdict, so it stays in step with `/health/ready`
and with what the swarm controller sees. `hyperhive.hive.warnings` is the
detail behind it — `warn`-level entries mean "an operator should look" and do
**not** set `degraded`.
Same cadence, transport and resource labels as the container metrics above.
## Agent-emitted custom metrics (`hive-metric`) ## Agent-emitted custom metrics (`hive-metric`)
Agents can push arbitrary labeled metrics to the same OTEL collector via the Agents can push arbitrary labeled metrics to the same OTEL collector via the

View file

@ -21,9 +21,14 @@
//! Bridging async→sync: [`container_stats::gather`] is async, but the SDK's //! Bridging async→sync: [`container_stats::gather`] is async, but the SDK's
//! observable-instrument callbacks are sync. So an async task refreshes a //! observable-instrument callbacks are sync. So an async task refreshes a
//! shared snapshot on an interval, and the (sync) callbacks read it. //! shared snapshot on an interval, and the (sync) callbacks read it.
//!
//! Also emits instruments whose subject is the hive itself rather than a
//! container in it — see [`register_hive_instruments`]. They carry **no
//! `agent` attribute**, and that absence is the point: it is what makes them
//! selectable as `{hive!="", agent=""}`.
use std::sync::{Arc, Mutex, OnceLock}; use std::sync::{Arc, Mutex, OnceLock};
use std::time::Duration; use std::time::{Duration, Instant};
use anyhow::{Context, Result}; use anyhow::{Context, Result};
use opentelemetry::KeyValue; use opentelemetry::KeyValue;
@ -33,6 +38,7 @@ use opentelemetry_sdk::Resource;
use opentelemetry_sdk::metrics::{PeriodicReader, SdkMeterProvider}; use opentelemetry_sdk::metrics::{PeriodicReader, SdkMeterProvider};
use super::container_stats::{self, ContainerResource}; use super::container_stats::{self, ContainerResource};
use super::warnings;
/// Default export cadence when `HYPERHIVE_OTEL_METRIC_INTERVAL_MS` is unset. /// Default export cadence when `HYPERHIVE_OTEL_METRIC_INTERVAL_MS` is unset.
const DEFAULT_INTERVAL: Duration = Duration::from_mins(1); const DEFAULT_INTERVAL: Duration = Duration::from_mins(1);
@ -45,6 +51,12 @@ static SNAPSHOT: OnceLock<Arc<Mutex<Vec<ContainerResource>>>> = OnceLock::new();
/// exports only while the provider lives. /// exports only while the provider lives.
static PROVIDER: OnceLock<SdkMeterProvider> = OnceLock::new(); static PROVIDER: OnceLock<SdkMeterProvider> = OnceLock::new();
/// When this exporter started, i.e. the origin `hyperhive.hive.uptime` counts
/// from. Set once in [`spawn_exporter`], which runs during hive-c0re startup —
/// so it is the process's age to within the startup sequence, and the metric's
/// doc says exporter-start rather than claiming a precision it does not have.
static STARTED: OnceLock<Instant> = OnceLock::new();
/// Spawn the container-resource OTEL exporter if OTEL is configured /// Spawn the container-resource OTEL exporter if OTEL is configured
/// (`OTEL_EXPORTER_OTLP_ENDPOINT` non-empty). No-op otherwise. Call once at /// (`OTEL_EXPORTER_OTLP_ENDPOINT` non-empty). No-op otherwise. Call once at
/// startup. /// startup.
@ -53,6 +65,7 @@ pub fn spawn_exporter(hyperhive_flake: &str) {
tracing::debug!("otel container-metrics: no endpoint configured, exporter disabled"); tracing::debug!("otel container-metrics: no endpoint configured, exporter disabled");
return; return;
}; };
let _ = STARTED.set(Instant::now());
let snapshot = SNAPSHOT let snapshot = SNAPSHOT
.get_or_init(|| Arc::new(Mutex::new(Vec::new()))) .get_or_init(|| Arc::new(Mutex::new(Vec::new())))
.clone(); .clone();
@ -225,6 +238,103 @@ fn register_instruments(provider: &SdkMeterProvider, snapshot: Arc<Mutex<Vec<Con
} }
}) })
.build(); .build();
register_hive_instruments(&meter);
}
/// Register the instruments whose subject is **this hive**, not a container in
/// it. Each observes with an empty attribute set (`hive` / `swarm` ride the
/// resource), so no data point here carries `agent` — see the module doc for
/// why that absence is the point.
///
/// Each is a projection of a value the process already computes; none is a new
/// source of truth. In particular the health pair comes from
/// [`warnings::readiness`], which is the one place that decides what counts as
/// unhealthy — deriving `degraded` from `warnings{level="crit"}` in a
/// dashboard query instead would put that rule in two places, and they would
/// disagree silently the first time a second degrading condition is added.
///
/// The two health instruments each read the registry separately, because the
/// multi-instrument batch observer (one callback feeding several instruments)
/// does not exist in this SDK: the API crate has no `register_callback`, and
/// the SDK's is `pub(crate)`, reachable only via these per-instrument
/// builders. Two consecutive reads inside one collection pass can therefore
/// straddle a registry change. Both readings are individually valid and the
/// next export agrees again; the shared-snapshot trick the container
/// instruments use would move that window rather than close it, and exists
/// only because `gather()` is async — which `readiness()` is not.
fn register_hive_instruments(meter: &opentelemetry::metrics::Meter) {
// Seconds since the exporter started. A hive that is up reports this every
// interval whether or not it hosts a single agent, which is what makes
// "no agents" distinguishable from "not reporting"; a restart reads as a
// drop back to ~0.
//
// The semconv name, not a `hyperhive.` one: the spec defines
// `process.uptime` as a double gauge in seconds, which is exactly this
// instrument, and inventing a name for a quantity the spec already names
// is the same mistake as claiming a spec name that does not exist (see
// `host_arch` below). Same split the container family already uses — spec
// names where a spec metric exists, `hyperhive.` for the rest.
meter
.f64_observable_gauge("process.uptime")
.with_unit("s")
.with_callback(|obs| {
if let Some(started) = STARTED.get() {
obs.observe(started.elapsed().as_secs_f64(), HIVE_ATTRS);
}
})
.build();
// The hive's own readiness verdict, 1 when degraded. Emitted as a series
// rather than left for a query to compute — see the note above.
meter
.u64_observable_gauge("hyperhive.hive.degraded")
.with_unit("1")
.with_callback(|obs| {
obs.observe(u64::from(warnings::readiness().is_degraded()), HIVE_ATTRS);
})
.build();
// How many warnings are currently raised, split by level. Every level is
// reported every cycle, zero included: an instrument that goes silent when
// there is nothing to report makes "healthy" and "not reporting" the same
// observation, which is the confusion this whole family exists to remove.
meter
.u64_observable_gauge("hyperhive.hive.warnings")
.with_unit("1")
.with_callback(|obs| {
for (level, n) in warning_counts(&warnings::snapshot()) {
obs.observe(n, &[KeyValue::new("level", level)]);
}
})
.build();
}
/// Attributes for the hive-scoped family: **none**.
///
/// Named rather than written as a bare `&[]` at each call site because the
/// emptiness is the load-bearing part — it is what makes these series
/// selectable as `{hive!="", agent=""}`. A bare empty slice reads like nothing
/// was decided; this one has a test asserting no `agent` key ever appears in
/// it, so "helpfully" labelling the hive family per-agent goes red.
const HIVE_ATTRS: &[KeyValue] = &[];
/// One count per warning level, **including the levels at zero**.
///
/// The zeros are the point. An instrument that reports only the levels
/// currently raised goes silent on a healthy hive, and a silent series is
/// indistinguishable from a hive that has stopped reporting — the exact
/// confusion the `hyperhive.hive.*` family exists to remove.
fn warning_counts(raised: &[crate::host_stats::ServerWarning]) -> Vec<(&'static str, u64)> {
warnings::LEVELS
.into_iter()
.map(|level| {
// Summed rather than `count()`ed to stay cast-free: `count()`
// yields a `usize` and the instrument takes a `u64`.
let n: u64 = raised.iter().filter(|w| w.level == level).map(|_| 1).sum();
(level, n)
})
.collect()
} }
/// Per-data-point attributes: the spec `container.name` plus the hive `agent` /// Per-data-point attributes: the spec `container.name` plus the hive `agent`
@ -332,6 +442,77 @@ fn parse_kv(s: &str) -> Vec<(String, String)> {
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
use crate::host_stats::ServerWarning;
fn warning(level: &'static str) -> ServerWarning {
ServerWarning {
kind: "t",
level,
message: String::new(),
}
}
/// The hive family must not be labelled per-agent: `{hive!="", agent=""}`
/// is the whole selector that distinguishes "a measurement of the hive"
/// from "a measurement of an agent that happens to live in one", and a
/// single `agent` attribute here collapses the two.
///
/// Paired with a control on [`attrs`], because this assertion is an
/// *absence*: without proof that the same check finds an `agent` key when
/// one is genuinely present, an empty result would equally mean the test
/// is looking for the wrong thing.
#[test]
fn hive_attrs_carry_no_agent_but_container_attrs_do() {
let has_agent = |kvs: &[KeyValue]| kvs.iter().any(|kv| kv.key.as_str() == "agent");
assert!(
!has_agent(HIVE_ATTRS),
"a hive-scoped series must not carry an `agent` attribute"
);
// Control: the same predicate, on the per-container attributes, which
// are supposed to carry exactly that key.
let sample = ContainerResource {
name: "atlas".to_owned(),
cpu_pct: 0.0,
cpu_time_usec: None,
mem_current_bytes: 0,
mem_peak_bytes: None,
mem_max_bytes: None,
disk_bytes: None,
};
assert!(
has_agent(&attrs(&sample)),
"control failed: the predicate cannot find an `agent` key that IS \
present, so the assertion above proves nothing"
);
}
/// A healthy hive still reports every level, at zero. Reporting only the
/// levels currently raised would make a healthy hive and a hive that has
/// stopped exporting produce the same thing — no series.
#[test]
fn warning_counts_reports_every_level_even_when_none_are_raised() {
let counts = warning_counts(&[]);
assert_eq!(counts.len(), warnings::LEVELS.len());
for (level, n) in counts {
assert_eq!(n, 0, "{level} should be reported as zero, not omitted");
}
}
#[test]
fn warning_counts_splits_by_level() {
let raised = [
warning(warnings::LEVEL_WARN),
warning(warnings::LEVEL_CRIT),
warning(warnings::LEVEL_WARN),
];
let counts = warning_counts(&raised);
assert_eq!(
counts,
vec![(warnings::LEVEL_WARN, 2), (warnings::LEVEL_CRIT, 1)]
);
}
#[test] #[test]
fn parse_kv_keeps_bearer_value() { fn parse_kv_keeps_bearer_value() {

View file

@ -97,6 +97,20 @@ pub const STATUS_OK: &str = "ok";
/// `status` value for a hive with at least one `crit`-level warning set. /// `status` value for a hive with at least one `crit`-level warning set.
pub const STATUS_DEGRADED: &str = "degraded"; pub const STATUS_DEGRADED: &str = "degraded";
/// Warning level meaning "an operator should look" — does not degrade the
/// hive's readiness.
pub const LEVEL_WARN: &str = "warn";
/// Warning level meaning "this hive is not healthy" — degrades readiness.
pub const LEVEL_CRIT: &str = "crit";
/// Every level a warning can carry, in increasing severity.
///
/// Named because a consumer that *reports on* levels has to enumerate them,
/// and enumerating them as literals in another module is the agreement
/// nothing checks: the day a third level is added, that consumer keeps
/// compiling and silently stops covering it.
pub const LEVELS: [&str; 2] = [LEVEL_WARN, LEVEL_CRIT];
/// What this hive currently says about its own health. /// What this hive currently says about its own health.
/// ///
/// One type with one producer ([`readiness`]) because there is more than /// One type with one producer ([`readiness`]) because there is more than
@ -133,7 +147,7 @@ impl Readiness {
#[must_use] #[must_use]
pub fn readiness() -> Readiness { pub fn readiness() -> Readiness {
let warnings = snapshot(); let warnings = snapshot();
let status = if warnings.iter().any(|w| w.level == "crit") { let status = if warnings.iter().any(|w| w.level == LEVEL_CRIT) {
STATUS_DEGRADED STATUS_DEGRADED
} else { } else {
STATUS_OK STATUS_OK