Watch
0
0
Fork
You've already forked hyperhive
0

docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix

gateway.md: split the opener into what/audience/enable; vhost map in two
tables (swarm-service vhosts declared by their own modules, then the hive
vhost) matching vhosts.nix and the service modules; gateway.enable exists
and is set with mkDefault by the modules that need it; Basic auth scope,
dashboard /health/ prefix, error-page rendering, matrix body limit and
forge link source corrected; nginx internals grouped under one Internals
section with their headings unchanged.

network.md: gateway and dnsmasq run on the host, not in a container;
network.enable is set by the modules that need it; shared-netns firewall
rule covers every swarm service container; hive-priv writes the nspawn
conf; domain sentence rewritten; removed options moved into <details>.

jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS
show different graphs drawn by the same component.

observability.md: swarm tier first; history narration cut; network access
deduplicated into a link to network.md; options link made absolute.

matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and
SSO options under deploy.matrix; .well-known is served on the hive domain;
roadmap sentence deleted; stale hive-c0re provisioning claims fixed;
serverName upgrade note moved into <details>.

Refs #3902
This commit is contained in:
atlas 2026-10-01 23:31:46 +02:00
commit 067f4e5699
5 changed files with 433 additions and 633 deletions

View file

@ -1,11 +1,13 @@
# The job queue, for operators
Every container operation — rebuild, first-spawn, a config-PR deploy,
power changes — runs through one shared job queue. This page explains
what the job queue _is_, as a general idea, independent of what any one
subsystem uses it for. For the hive-c0re-specific step catalogue and the
engineering internals (scheduler, leases, resource windows) see
[`coordinator.md`](coordinator.md) instead.
Long-running work runs through a job graph. The swarm controller keeps one
for swarm-level work — creating an agent's identity, forge user and config
repo. Each hive's hive-c0re keeps its own for container operations —
rebuild, first-spawn, a config-PR deploy, power changes. This page explains
what the job queue _is_, as a general idea, independent of what either uses
it for. For the hive-c0re step catalogue and the engineering internals
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)
instead.
## What the job queue is, in the abstract
@ -24,18 +26,18 @@ Two ideas are all there is to it:
The engine's whole job is: whenever a step's ordering and resource needs
are both satisfied, run it. It has no opinion on what the steps _do_ —
that's supplied by whoever builds the graph. hive-c0re is the one thing
building graphs on it today, but nothing about the engine is specific to
containers or rebuilds; there's nothing stopping another subsystem from
using the same engine for its own unrelated queue.
that's supplied by whoever builds the graph. The swarm controller and
hive-c0re each build their own graph on it, and nothing about the engine
is specific to either.
## Watching it happen
Each **row** you see in a queue view (the **BU1LDS** page's R3BU1LD QU3U3
— see [`web-ui/dashboard.md`](../web-ui/dashboard.md) — and swarm-ui's
`/jobs` page both render the same underlying graph) is one job; the rows
nested under it are that job's steps, in order (occasionally a couple run
side by side). A step shows one of:
Two views, one per graph, drawn by the same component: the swarm UI's
`/jobs` page shows the swarm controller's graph, and the hive dashboard's
**BU1LDS** page (R3BU1LD QU3U3 — see
[`web-ui/dashboard.md`](../web-ui/dashboard.md)) shows that hive's. Each
**row** is one job; the rows nested under it are that job's steps, in
order (occasionally a couple run side by side). A step shows one of:
<!-- vale write-good.Passive = NO -->

View file

@ -1,48 +1,74 @@
# Observability (OpenTelemetry)
hyperhive has built-in support for exporting each agent container's telemetry to
any OTLP-compatible collector: Claude Code statistics — token usage, cost, tool
call counts — via Claude Code's built-in OpenTelemetry integration, and the
container's journal, forwarded by a collector running inside it.
The swarm has one telemetry pipeline. The swarm's collector writes the
swarm's metrics and log stores and, if you name one, exports upstream; each
hive runs a collector of its own that forwards its agents' telemetry there.
What flows through it: each agent's Claude Code statistics (token usage,
cost, tool call counts) via Claude Code's built-in OpenTelemetry
integration, each agent container's journal, and the hyperhive metrics
catalogued below.
This is a **hive-wide** setting: one switch in the host NixOS config enables it
for every agent container simultaneously. No per-agent opt-in or opt-out exists.
## The two collectors
## Enabling export
Telemetry crosses two collectors, and which one you configure depends on what
the host is:
| | runs where | receives from | does |
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's stores and exports upstream |
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
An all-local host runs both, and needs nothing said about the hop between them.
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel";
enable = true;
endpoint = "https://collector.example.com/otel"; # the upstream
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
};
```
`enable` is the single gate. `endpoint` is where telemetry ends up after it
leaves the swarm — optional, because the swarm's own metrics store
(`deploy.victoriametrics`) is a destination in its own right. With both,
telemetry goes to both. See
## Enabling export
`otel.enable` is the single gate on a hive: one switch in the host config
covers every agent container on it, with no per-agent opt-in or opt-out.
`endpoint` is where telemetry ends up after it leaves the swarm — optional,
because the swarm's own metrics store (`deploy.victoriametrics`) is a
destination in its own right. With both, telemetry goes to both. See
[`swarm/services.md`](../swarm/services.md#metrics-victoriametrics--grafana).
**Telemetry leaves a hive exactly one way: through the collector that
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
they export unauthenticated to a bridge address only their own containers can
reach. That collector forwards to the swarm's
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), which is
the single process holding the upstream credential and the only writer to the
swarm's store. No agent holds a copy, and neither does this hive.
reach (`http://<bridgeIp>:<collector port>`; the otel module opens that port
on the bridge itself). That collector forwards to the swarm's
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), the
single process holding the upstream credential and the only writer to the
swarm's stores. No agent holds a copy, and neither does the hive.
The hive collector reaches the swarm collector by its gateway name
(`swarm.otel.domain`, default `otel.<swarm domain>`) — the same DNS-and-CA-trust
shape every hive-to-swarm-service hop uses, not a URL an operator has to point
anywhere. A hive that doesn't run the swarm's services still resolves that
name through the gateway; nothing here needs setting for the split-host case.
shape every hive-to-swarm-service hop uses. On the host running the swarm
collector the hive's dnsmasq answers that name; elsewhere it resolves through
ordinary DNS. Nothing here needs setting for the split-host case.
⚠️ **The collector is therefore in the path of all telemetry.** It runs on the
same host as the agents and restarts on failure, and telemetry isn't the
control plane — degraded telemetry isn't degraded operation — but the export
no longer survives independently of anything host-side.
⚠️ **The hive collector carries every bit of that hive's telemetry.** It
runs on the same host as the agents and restarts on failure. Telemetry isn't
the control plane, so degraded telemetry isn't degraded operation.
### what the agent→collector hop is and isn't
### Why two tiers
**The hive tier isn't optional.** Exporting straight to `endpoint` would mean
every agent needs the credential — and the harness delivers that token into
the agent's own `~/.claude/settings.json`, a file the agent can read. `0600`
protects it from other containers, not from the agent itself. An option that
could select the direct path would reopen that hole.
**The tiers stay separate on one box.** An all-local hive is a statement about
_where_ processes run, not about the shape of the deployment. A boundary that
disappears locally is one the local deployment stops testing.
### What the agent→collector hop is and isn't
**It has no application-level auth.** The receiver takes any OTLP that reaches
it; what bounds who can reach it's the firewall — `exposeHostPorts` opens the
@ -55,10 +81,9 @@ credential.** Neither tier can tell a container's genuine Claude Code stats
from anything else shaped like OTLP arriving on that port — including data
smuggled out in resource attributes on an otherwise-legitimate export.
That's a **different risk from the one the collector fixes**, and strictly
smaller than what preceded it: before, every agent held the upstream credential
itself, so it could do all of the above _and_ use the token anywhere else. The
collector removes the token and keeps the pipe. Agents are inside the trust
That's a **different risk from the one the collector fixes**: an agent that
held the upstream credential could do everything above _and_ use the token
anywhere else. The collector removes the token and keeps the pipe. Agents are inside the trust
boundary (`docs/trust-boundary/security.md`: capability = accepted risk), so an agent being
able to _send_ is an accepted extension of that boundary — but it's not
closed by this design, and don't read anything here as closing it.
@ -74,7 +99,7 @@ hive's collector can label its data as any other agent.
agent container forwards its own journal through this port — every unit in it at
`info` and above, not an allowlist. That's the harness, the MCP daemons and
whatever a tool call spawned, so command lines and error text now leave the
container where before only counts did. The trust boundary is unchanged (same
container, not just counts. The trust boundary is unchanged (same
destination, same credential, and an agent could already send arbitrary OTLP);
what changes is how much detail leaves by default.
@ -98,54 +123,6 @@ rather than the claims it authenticated with.
If you need per-agent numbers you can act on, take them from the agent's own
turn-stats rather than from a metric label.
## Options reference
The nix module (`nix/host-modules/otel.nix`) generates every
`services.hyperhive.otel.*` option's full type/default/description/
example straight into [`/options/`](/options/) (host options — `nix build
.#docs-host` for a local render). The build keeps that page honest in a
way a hand-copied version here can't be, so it's the reference, not this
doc. What follows is what a flat per-option listing can't express: the
two-tier architecture, the security model, and how the options interact.
## The two collectors
Telemetry crosses two collectors, and which one you configure depends on what
the host is:
| | runs where | receives from | does |
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's store and exports upstream |
An all-local host runs both, and needs nothing said about the hop between them.
```nix
services.hyperhive.otel = {
enable = true;
endpoint = "https://collector.example.com/otel"; # the upstream
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
};
```
**Why the hive tier isn't optional.** Exporting straight to `endpoint` means
every agent needs the credential to authenticate — and the harness delivers
that token into the agent's own `~/.claude/settings.json`, a file the agent can
read. `0600` protects it from other containers, not from the agent itself. As
long as the direct path stays _selectable_, that hole stays selectable; an
option that can reintroduce it's a hole with extra steps.
**Why the tiers stay separate on one box.** They're not collapsed when
co-located: an all-local hive is a statement about _where_ processes run, not
about the shape of the deployment. A boundary that disappears locally is one
the local deployment stops testing.
**`endpoint` keeps meaning "where telemetry goes upstream."** Neither tier
redefines it — the agent-facing value is _derived_
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
keeps working unchanged. The otel module contributes the bridge port to `exposeHostPorts`
automatically; there is nothing to open by hand.
### Authenticated ingest
The swarm tier gives **each hive its own receiver**, and stamps the `hive` label
@ -180,26 +157,19 @@ a build error naming the reason rather than telemetry silently going nowhere.
## Network access
Agent containers can only reach the host on ports 80 and 443 by default. To let
them reach some other host-local service you run yourself — a database, a
scratch HTTP endpoint — open its port on the bridge:
Agent containers reach the host only on ports 80 and 443 (plus DNS). The
collector port needs nothing from you: `otel.enable` opens it on the bridge.
To let agents reach some other host-local service of your own →
[`networking/network.md`](../networking/network.md#reaching-host-services-exposehostports).
```nix
services.hyperhive.network.exposeHostPorts = [ 5432 ];
```
## Options reference
and point whatever consumes it at `10.42.0.1:5432` rather than loopback: inside
a container, loopback is the _container_. The bridge IP is the host's address on
the `hive-br0` bridge. The service must also bind an address the bridge can
reach — a `127.0.0.1`-only listener stays unreachable no matter what the
firewall allows. See `docs/networking/network.md::Reaching host services` for details.
<!-- vale write-good.Passive = NO -->
⚠️ **None of this is needed for hyperhive's own telemetry** — `otel.enable`
contributes the collector's port and derives the agent-facing endpoint itself.
<!-- vale write-good.Passive = YES -->
The nix module (`nix/host-modules/otel.nix`) generates every
`services.hyperhive.otel.*` option's full type, default, description and
example into the [options reference](https://hyperhive.darkest.space/options/)
(host options — `nix build .#docs-host` for a local render). That's the
reference; this page covers what a per-option listing can't: the two-tier
architecture, the security model, and how the options interact.
## Built-in resource labels