turn duration (ms) — p50 / p95 / avg
context tokens (last inference per turn) — avg / max
token cost per bucket (sum across inferences)
turns by model per bucket — model drives token cost
result trend per bucket — errors / rate-limits / compactions over time