hyperhive/nix/host-modules/swarm-grafana/dashboards/logstore.json
atlas fc870c459a grafana: count log lines that arrive with no severity
The mapping above needs something that says whether it is still working
after whoever wrote it has gone. A timeseries rather than a stat, so a
regression is a line lifting off zero rather than a number nobody reads.

Two series, and the split is the point: rows that carried a PRIORITY and
arrived with no severity anyway (a broken mapping — this one must reach
zero and stay there), against rows that never had a priority to map. The
latter is Claude Code's own OTLP telemetry, which emits log records with no
severity set at the source; mapping cannot reach it, so it is named rather
than folded into one number that never goes to zero.

Needs no provisioning change — logstore.json is already in the shipped
dashboard list.
2026-09-20 14:23:56 +02:00

589 lines
18 KiB
JSON

{
"title": "hyperhive · logs (victorialogs)",
"uid": "hyperhive-swarm-logstore",
"description": "The swarm's log pipeline: whether the store is healthy (top) and which sources are actually feeding it (bottom). One of the per-service boards split out of the combined swarm-services page. The two halves answer different questions and are on one board on purpose — aggregate ingestion stays normal on host-tier units while a whole tier ships nothing, so a healthy top half is not evidence about the bottom one.",
"editable": true,
"refresh": "1m",
"schemaVersion": 39,
"tags": ["hyperhive", "swarm", "victorialogs", "logs"],
"time": {
"from": "now-24h",
"to": "now"
},
"timezone": "utc",
"version": 1,
"panels": [
{
"id": 1,
"type": "timeseries",
"title": "Log rows ingested (range)",
"description": "Rows accepted by the log store over the range. ⚠️ This counts what ARRIVED, never what became findable — the pipeline shipped for days with every record carrying a placeholder in place of its message, while this counter climbed normally. A healthy line here is necessary and not sufficient; the check that closes it is a query returning readable lines.",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"gridPos": {
"h": 8,
"w": 24,
"x": 0,
"y": 4
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"expr": "sum(increase(vl_rows_ingested_total[$__range])) or vector(0)",
"legendFormat": "rows"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "palette-classic"
},
"custom": {
"drawStyle": "line",
"lineWidth": 1,
"fillOpacity": 10,
"showPoints": "never",
"spanNulls": false,
"axisSoftMin": 0
}
},
"overrides": []
},
"options": {
"legend": {
"displayMode": "list",
"placement": "bottom",
"showLegend": true,
"calcs": []
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
}
},
{
"id": 2,
"type": "timeseries",
"title": "Log store size on disk",
"description": "Compressed on-disk size, split by the store's own `type` label (storage vs indexdb). Retention is 30 days, so this is expected to climb to a plateau rather than forever; a straight line past that is the signal.",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"gridPos": {
"h": 8,
"w": 24,
"x": 0,
"y": 12
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"expr": "vl_data_size_bytes",
"legendFormat": "{{type}}"
}
],
"fieldConfig": {
"defaults": {
"unit": "bytes",
"decimals": 1,
"color": {
"mode": "palette-classic"
},
"custom": {
"drawStyle": "line",
"lineWidth": 1,
"fillOpacity": 10,
"showPoints": "never",
"spanNulls": false,
"axisSoftMin": 0
}
},
"overrides": []
},
"options": {
"legend": {
"displayMode": "list",
"placement": "bottom",
"showLegend": true,
"calcs": []
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
}
},
{
"id": 3,
"type": "stat",
"title": "Log store free disk",
"description": "Free space on the volume holding the log store. It shares a filesystem with the metrics store, so the two panels move together — a drop here that the metrics one does not show means something outside this swarm is filling the disk.",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"gridPos": {
"h": 4,
"w": 12,
"x": 0,
"y": 0
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"expr": "vl_free_disk_space_bytes",
"instant": true,
"legendFormat": "free"
}
],
"fieldConfig": {
"defaults": {
"unit": "bytes",
"decimals": 1,
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 5000000000
}
]
}
},
"overrides": []
},
"options": {
"graphMode": "none",
"colorMode": "value",
"textMode": "value",
"justifyMode": "auto",
"reduceOptions": {
"calcs": ["lastNotNull"],
"fields": "",
"values": false
}
}
},
{
"id": 4,
"type": "stat",
"title": "Log store errors",
"description": "Internal errors plus rejected HTTP requests. Healthy value is zero. ⚠️ A wrong OTLP path answers 400 and lands here, and so does a right path with a bad payload — the status code alone does not separate them, only the store's own log does.",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"gridPos": {
"h": 4,
"w": 12,
"x": 12,
"y": 0
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"expr": "sum(vl_errors_total) or vector(0)",
"instant": true,
"legendFormat": "internal"
},
{
"refId": "B",
"datasource": {
"type": "prometheus",
"uid": "@datasourceUid@"
},
"expr": "sum(vl_http_errors_total) or vector(0)",
"instant": true,
"legendFormat": "http"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 1
}
]
}
},
"overrides": []
},
"options": {
"graphMode": "none",
"colorMode": "value",
"textMode": "value",
"justifyMode": "auto",
"reduceOptions": {
"calcs": ["lastNotNull"],
"fields": "",
"values": false
}
},
"x-zero-is-healthy": true
},
{
"id": 5,
"type": "stat",
"title": "Rows in range · all sources (control)",
"description": "Deliberately ungrouped, and it is what makes the two tables below readable. An empty table renders identically whether the query is malformed or the source genuinely never shipped: nonzero here with an empty table means the QUERY is broken; zero here means the store really is empty for this range. Read it against `Log rows ingested` above too — that counter is the same quantity measured from the store's own metrics rather than by querying, so the two disagreeing means rows arrived that a query cannot reach.",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"gridPos": {
"h": 4,
"w": 24,
"x": 0,
"y": 20
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"queryType": "stats",
"expr": "* | stats count() as rows"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
},
"overrides": []
},
"options": {
"graphMode": "none",
"colorMode": "value",
"textMode": "value",
"justifyMode": "auto",
"reduceOptions": {
"calcs": ["lastNotNull"],
"fields": "",
"values": false
}
}
},
{
"id": 6,
"type": "bargauge",
"title": "Rows by unit",
"description": "Every `_SYSTEMD_UNIT` present in the range, discovered rather than enumerated — nothing here names a unit, so a source that starts shipping appears on its own. ⚠️ Absence means NOT COLLECTED, which is not the same as not running: the collector's journald receiver only reads the units listed in `services.hyperhive.swarm.otel.journaldUnits`, so a unit missing from that list can be running and logging and still never reach this board.",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"gridPos": {
"h": 12,
"w": 12,
"x": 0,
"y": 33
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"queryType": "stats",
"expr": "* | stats by (_SYSTEMD_UNIT) count() as rows",
"legendFormat": "{{_SYSTEMD_UNIT}}"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "continuous-BlPu"
},
"mappings": []
},
"overrides": []
},
"options": {
"displayMode": "gradient",
"orientation": "horizontal",
"showUnfilled": true,
"valueMode": "text",
"minVizWidth": 8,
"minVizHeight": 12,
"reduceOptions": {
"calcs": ["lastNotNull"],
"fields": "",
"values": false
}
}
},
{
"id": 7,
"type": "bargauge",
"title": "Rows by host",
"description": "The same count grouped by `_HOSTNAME`. Read together with the unit table this separates two very different failures: one host contributing nothing is a collector or journal-linkage problem on that host, while a unit missing across every host is a configuration problem in what gets collected.",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"gridPos": {
"h": 12,
"w": 12,
"x": 12,
"y": 33
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"queryType": "stats",
"expr": "* | stats by (_HOSTNAME) count() as rows",
"legendFormat": "{{_HOSTNAME}}"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "continuous-BlPu"
},
"mappings": []
},
"overrides": []
},
"options": {
"displayMode": "gradient",
"orientation": "horizontal",
"showUnfilled": true,
"valueMode": "text",
"minVizWidth": 8,
"minVizHeight": 12,
"reduceOptions": {
"calcs": ["lastNotNull"],
"fields": "",
"values": false
}
}
},
{
"id": 8,
"type": "timeseries",
"title": "Rows over time, by unit",
"description": "The tables say who shipped over the whole range; this says when. A line that stops is a source that died, and it is the only view here that distinguishes that from one which never existed — a source absent for the entire range looks the same as an unknown one in a table, but shows up here as a line that ends.",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"gridPos": {
"h": 9,
"w": 24,
"x": 0,
"y": 45
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"queryType": "statsRange",
"expr": "* | stats by (_SYSTEMD_UNIT) count() as rows",
"legendFormat": "{{_SYSTEMD_UNIT}}"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "palette-classic"
},
"custom": {
"drawStyle": "line",
"lineWidth": 1,
"fillOpacity": 10,
"showPoints": "never",
"spanNulls": false,
"axisSoftMin": 0
}
},
"overrides": []
},
"options": {
"legend": {
"displayMode": "list",
"placement": "bottom",
"showLegend": true,
"calcs": []
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
}
},
{
"id": 9,
"type": "timeseries",
"title": "Log rows with no severity (must be zero)",
"description": "This panel is a check rather than a view: the collector maps journald's PRIORITY onto an OpenTelemetry severity at every tier, and this is what says out loud whether a line got one. Both series are rows whose `severity_text` is absent or the literal `Unspecified`, which is what VictoriaLogs stores when nothing set a severity — split by whether the row had a PRIORITY to map in the first place.\n\n**`journald_priority_unmapped`** is a row that carries PRIORITY and still arrived without a severity. It is the regression line: nonzero means a journald receiver somewhere lost its severity_parser operator, so read it against `nix/journald-severity.nix` and the two receivers that import it.\n\n**`no_priority_field`** is a row that never had a priority — today that is Claude Code's own OTLP telemetry (`scope.name: com.anthropic.claude_code.events`), which emits log records with no severity set at the source. Mapping cannot reach it; it needs a fix where it is produced, so this series is expected nonzero until that lands and is NOT evidence the mapping broke.\n\nRange-scoped over time on purpose. An instantaneous zero is also what a broken query returns — a line you can watch go to zero and stay there is the part that survives whoever wrote it. To sanity-check the query itself, read it against `Rows in range · all sources (control)` above: that panel nonzero while both of these are flat zero means these queries stopped matching, not that every line grew a severity.",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"gridPos": {
"h": 9,
"w": 24,
"x": 0,
"y": 24
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"queryType": "statsRange",
"expr": "(severity_text:\"\" OR severity_text:\"Unspecified\") PRIORITY:* | stats count() as journald_priority_unmapped"
},
{
"refId": "B",
"datasource": {
"type": "victoriametrics-logs-datasource",
"uid": "@logsDatasourceUid@"
},
"queryType": "statsRange",
"expr": "(severity_text:\"\" OR severity_text:\"Unspecified\") PRIORITY:\"\" | stats count() as no_priority_field"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"decimals": 0,
"color": {
"mode": "palette-classic"
},
"custom": {
"drawStyle": "line",
"lineWidth": 2,
"fillOpacity": 10,
"showPoints": "never",
"spanNulls": false,
"axisSoftMin": 0
}
},
"overrides": [
{
"matcher": {
"id": "byName",
"options": "journald_priority_unmapped"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "red"
}
}
]
},
{
"matcher": {
"id": "byName",
"options": "no_priority_field"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "orange"
}
}
]
}
]
},
"options": {
"legend": {
"displayMode": "table",
"placement": "bottom",
"showLegend": true,
"calcs": ["max", "lastNotNull"]
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
}
}
]
}