docs(networking): facts + structure pass on gateway, network, jobq, observability, matrix
gateway.md: split the opener into what/audience/enable; vhost map in two tables (swarm-service vhosts declared by their own modules, then the hive vhost) matching vhosts.nix and the service modules; gateway.enable exists and is set with mkDefault by the modules that need it; Basic auth scope, dashboard /health/ prefix, error-page rendering, matrix body limit and forge link source corrected; nginx internals grouped under one Internals section with their headings unchanged. network.md: gateway and dnsmasq run on the host, not in a container; network.enable is set by the modules that need it; shared-netns firewall rule covers every swarm service container; hive-priv writes the nspawn conf; domain sentence rewritten; removed options moved into <details>. jobq.md: swarm-controller runs its own graph; swarm UI /jobs and BU1LDS show different graphs drawn by the same component. observability.md: swarm tier first; history narration cut; network access deduplicated into a link to network.md; options link made absolute. matrix.md: swarm.matrix vs deploy.matrix namespaces; tuning, firewall and SSO options under deploy.matrix; .well-known is served on the hive domain; roadmap sentence deleted; stale hive-c0re provisioning claims fixed; serverName upgrade note moved into <details>. Refs #3902
This commit is contained in:
parent
2b2608a491
commit
067f4e5699
5 changed files with 433 additions and 633 deletions
|
|
@ -1,11 +1,13 @@
|
|||
# The job queue, for operators
|
||||
|
||||
Every container operation — rebuild, first-spawn, a config-PR deploy,
|
||||
power changes — runs through one shared job queue. This page explains
|
||||
what the job queue _is_, as a general idea, independent of what any one
|
||||
subsystem uses it for. For the hive-c0re-specific step catalogue and the
|
||||
engineering internals (scheduler, leases, resource windows) see
|
||||
[`coordinator.md`](coordinator.md) instead.
|
||||
Long-running work runs through a job graph. The swarm controller keeps one
|
||||
for swarm-level work — creating an agent's identity, forge user and config
|
||||
repo. Each hive's hive-c0re keeps its own for container operations —
|
||||
rebuild, first-spawn, a config-PR deploy, power changes. This page explains
|
||||
what the job queue _is_, as a general idea, independent of what either uses
|
||||
it for. For the hive-c0re step catalogue and the engineering internals
|
||||
(scheduler, leases, resource windows) see [`coordinator.md`](coordinator.md)
|
||||
instead.
|
||||
|
||||
## What the job queue is, in the abstract
|
||||
|
||||
|
|
@ -24,18 +26,18 @@ Two ideas are all there is to it:
|
|||
|
||||
The engine's whole job is: whenever a step's ordering and resource needs
|
||||
are both satisfied, run it. It has no opinion on what the steps _do_ —
|
||||
that's supplied by whoever builds the graph. hive-c0re is the one thing
|
||||
building graphs on it today, but nothing about the engine is specific to
|
||||
containers or rebuilds; there's nothing stopping another subsystem from
|
||||
using the same engine for its own unrelated queue.
|
||||
that's supplied by whoever builds the graph. The swarm controller and
|
||||
hive-c0re each build their own graph on it, and nothing about the engine
|
||||
is specific to either.
|
||||
|
||||
## Watching it happen
|
||||
|
||||
Each **row** you see in a queue view (the **BU1LDS** page's R3BU1LD QU3U3
|
||||
— see [`web-ui/dashboard.md`](../web-ui/dashboard.md) — and swarm-ui's
|
||||
`/jobs` page both render the same underlying graph) is one job; the rows
|
||||
nested under it are that job's steps, in order (occasionally a couple run
|
||||
side by side). A step shows one of:
|
||||
Two views, one per graph, drawn by the same component: the swarm UI's
|
||||
`/jobs` page shows the swarm controller's graph, and the hive dashboard's
|
||||
**BU1LDS** page (R3BU1LD QU3U3 — see
|
||||
[`web-ui/dashboard.md`](../web-ui/dashboard.md)) shows that hive's. Each
|
||||
**row** is one job; the rows nested under it are that job's steps, in
|
||||
order (occasionally a couple run side by side). A step shows one of:
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
|
|
|
|||
|
|
@ -1,48 +1,74 @@
|
|||
# Observability (OpenTelemetry)
|
||||
|
||||
hyperhive has built-in support for exporting each agent container's telemetry to
|
||||
any OTLP-compatible collector: Claude Code statistics — token usage, cost, tool
|
||||
call counts — via Claude Code's built-in OpenTelemetry integration, and the
|
||||
container's journal, forwarded by a collector running inside it.
|
||||
The swarm has one telemetry pipeline. The swarm's collector writes the
|
||||
swarm's metrics and log stores and, if you name one, exports upstream; each
|
||||
hive runs a collector of its own that forwards its agents' telemetry there.
|
||||
What flows through it: each agent's Claude Code statistics (token usage,
|
||||
cost, tool call counts) via Claude Code's built-in OpenTelemetry
|
||||
integration, each agent container's journal, and the hyperhive metrics
|
||||
catalogued below.
|
||||
|
||||
This is a **hive-wide** setting: one switch in the host NixOS config enables it
|
||||
for every agent container simultaneously. No per-agent opt-in or opt-out exists.
|
||||
## The two collectors
|
||||
|
||||
## Enabling export
|
||||
Telemetry crosses two collectors, and which one you configure depends on what
|
||||
the host is:
|
||||
|
||||
| | runs where | receives from | does |
|
||||
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
|
||||
| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's stores and exports upstream |
|
||||
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
|
||||
|
||||
An all-local host runs both, and needs nothing said about the hop between them.
|
||||
|
||||
```nix
|
||||
services.hyperhive.otel = {
|
||||
enable = true;
|
||||
endpoint = "https://collector.example.com/otel";
|
||||
enable = true;
|
||||
endpoint = "https://collector.example.com/otel"; # the upstream
|
||||
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
|
||||
};
|
||||
```
|
||||
|
||||
`enable` is the single gate. `endpoint` is where telemetry ends up after it
|
||||
leaves the swarm — optional, because the swarm's own metrics store
|
||||
(`deploy.victoriametrics`) is a destination in its own right. With both,
|
||||
telemetry goes to both. See
|
||||
## Enabling export
|
||||
|
||||
`otel.enable` is the single gate on a hive: one switch in the host config
|
||||
covers every agent container on it, with no per-agent opt-in or opt-out.
|
||||
`endpoint` is where telemetry ends up after it leaves the swarm — optional,
|
||||
because the swarm's own metrics store (`deploy.victoriametrics`) is a
|
||||
destination in its own right. With both, telemetry goes to both. See
|
||||
[`swarm/services.md`](../swarm/services.md#metrics-victoriametrics--grafana).
|
||||
|
||||
**Telemetry leaves a hive exactly one way: through the collector that
|
||||
`enable` starts on the host.** Agents never talk to `endpoint` themselves —
|
||||
they export unauthenticated to a bridge address only their own containers can
|
||||
reach. That collector forwards to the swarm's
|
||||
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), which is
|
||||
the single process holding the upstream credential and the only writer to the
|
||||
swarm's store. No agent holds a copy, and neither does this hive.
|
||||
reach (`http://<bridgeIp>:<collector port>`; the otel module opens that port
|
||||
on the bridge itself). That collector forwards to the swarm's
|
||||
([`swarm/services.md`](../swarm/services.md#telemetry-collector-otel)), the
|
||||
single process holding the upstream credential and the only writer to the
|
||||
swarm's stores. No agent holds a copy, and neither does the hive.
|
||||
|
||||
The hive collector reaches the swarm collector by its gateway name
|
||||
(`swarm.otel.domain`, default `otel.<swarm domain>`) — the same DNS-and-CA-trust
|
||||
shape every hive-to-swarm-service hop uses, not a URL an operator has to point
|
||||
anywhere. A hive that doesn't run the swarm's services still resolves that
|
||||
name through the gateway; nothing here needs setting for the split-host case.
|
||||
shape every hive-to-swarm-service hop uses. On the host running the swarm
|
||||
collector the hive's dnsmasq answers that name; elsewhere it resolves through
|
||||
ordinary DNS. Nothing here needs setting for the split-host case.
|
||||
|
||||
⚠️ **The collector is therefore in the path of all telemetry.** It runs on the
|
||||
same host as the agents and restarts on failure, and telemetry isn't the
|
||||
control plane — degraded telemetry isn't degraded operation — but the export
|
||||
no longer survives independently of anything host-side.
|
||||
⚠️ **The hive collector carries every bit of that hive's telemetry.** It
|
||||
runs on the same host as the agents and restarts on failure. Telemetry isn't
|
||||
the control plane, so degraded telemetry isn't degraded operation.
|
||||
|
||||
### what the agent→collector hop is and isn't
|
||||
### Why two tiers
|
||||
|
||||
**The hive tier isn't optional.** Exporting straight to `endpoint` would mean
|
||||
every agent needs the credential — and the harness delivers that token into
|
||||
the agent's own `~/.claude/settings.json`, a file the agent can read. `0600`
|
||||
protects it from other containers, not from the agent itself. An option that
|
||||
could select the direct path would reopen that hole.
|
||||
|
||||
**The tiers stay separate on one box.** An all-local hive is a statement about
|
||||
_where_ processes run, not about the shape of the deployment. A boundary that
|
||||
disappears locally is one the local deployment stops testing.
|
||||
|
||||
### What the agent→collector hop is and isn't
|
||||
|
||||
**It has no application-level auth.** The receiver takes any OTLP that reaches
|
||||
it; what bounds who can reach it's the firewall — `exposeHostPorts` opens the
|
||||
|
|
@ -55,10 +81,9 @@ credential.** Neither tier can tell a container's genuine Claude Code stats
|
|||
from anything else shaped like OTLP arriving on that port — including data
|
||||
smuggled out in resource attributes on an otherwise-legitimate export.
|
||||
|
||||
That's a **different risk from the one the collector fixes**, and strictly
|
||||
smaller than what preceded it: before, every agent held the upstream credential
|
||||
itself, so it could do all of the above _and_ use the token anywhere else. The
|
||||
collector removes the token and keeps the pipe. Agents are inside the trust
|
||||
That's a **different risk from the one the collector fixes**: an agent that
|
||||
held the upstream credential could do everything above _and_ use the token
|
||||
anywhere else. The collector removes the token and keeps the pipe. Agents are inside the trust
|
||||
boundary (`docs/trust-boundary/security.md`: capability = accepted risk), so an agent being
|
||||
able to _send_ is an accepted extension of that boundary — but it's not
|
||||
closed by this design, and don't read anything here as closing it.
|
||||
|
|
@ -74,7 +99,7 @@ hive's collector can label its data as any other agent.
|
|||
agent container forwards its own journal through this port — every unit in it at
|
||||
`info` and above, not an allowlist. That's the harness, the MCP daemons and
|
||||
whatever a tool call spawned, so command lines and error text now leave the
|
||||
container where before only counts did. The trust boundary is unchanged (same
|
||||
container, not just counts. The trust boundary is unchanged (same
|
||||
destination, same credential, and an agent could already send arbitrary OTLP);
|
||||
what changes is how much detail leaves by default.
|
||||
|
||||
|
|
@ -98,54 +123,6 @@ rather than the claims it authenticated with.
|
|||
If you need per-agent numbers you can act on, take them from the agent's own
|
||||
turn-stats rather than from a metric label.
|
||||
|
||||
## Options reference
|
||||
|
||||
The nix module (`nix/host-modules/otel.nix`) generates every
|
||||
`services.hyperhive.otel.*` option's full type/default/description/
|
||||
example straight into [`/options/`](/options/) (host options — `nix build
|
||||
.#docs-host` for a local render). The build keeps that page honest in a
|
||||
way a hand-copied version here can't be, so it's the reference, not this
|
||||
doc. What follows is what a flat per-option listing can't express: the
|
||||
two-tier architecture, the security model, and how the options interact.
|
||||
|
||||
## The two collectors
|
||||
|
||||
Telemetry crosses two collectors, and which one you configure depends on what
|
||||
the host is:
|
||||
|
||||
| | runs where | receives from | does |
|
||||
| ------------------------------------ | ---------------------- | --------------------------------- | --------------------------------------------------------------------- |
|
||||
| **hive tier** — `otel.enable` | every hive with agents | that hive's agents, on the bridge | forwards to the swarm tier. Holds no credential, picks no destination |
|
||||
| **swarm tier** — `deploy.swarm-otel` | once per swarm | every hive's collector | writes the swarm's store and exports upstream |
|
||||
|
||||
An all-local host runs both, and needs nothing said about the hop between them.
|
||||
|
||||
```nix
|
||||
services.hyperhive.otel = {
|
||||
enable = true;
|
||||
endpoint = "https://collector.example.com/otel"; # the upstream
|
||||
headersCredential = "/run/secrets/otel-headers"; # only the swarm tier reads it
|
||||
};
|
||||
```
|
||||
|
||||
**Why the hive tier isn't optional.** Exporting straight to `endpoint` means
|
||||
every agent needs the credential to authenticate — and the harness delivers
|
||||
that token into the agent's own `~/.claude/settings.json`, a file the agent can
|
||||
read. `0600` protects it from other containers, not from the agent itself. As
|
||||
long as the direct path stays _selectable_, that hole stays selectable; an
|
||||
option that can reintroduce it's a hole with extra steps.
|
||||
|
||||
**Why the tiers stay separate on one box.** They're not collapsed when
|
||||
co-located: an all-local hive is a statement about _where_ processes run, not
|
||||
about the shape of the deployment. A boundary that disappears locally is one
|
||||
the local deployment stops testing.
|
||||
|
||||
**`endpoint` keeps meaning "where telemetry goes upstream."** Neither tier
|
||||
redefines it — the agent-facing value is _derived_
|
||||
(`http://<bridgeIp>:<collector.port>`), so an existing deployment's `endpoint`
|
||||
keeps working unchanged. The otel module contributes the bridge port to `exposeHostPorts`
|
||||
automatically; there is nothing to open by hand.
|
||||
|
||||
### Authenticated ingest
|
||||
|
||||
The swarm tier gives **each hive its own receiver**, and stamps the `hive` label
|
||||
|
|
@ -180,26 +157,19 @@ a build error naming the reason rather than telemetry silently going nowhere.
|
|||
|
||||
## Network access
|
||||
|
||||
Agent containers can only reach the host on ports 80 and 443 by default. To let
|
||||
them reach some other host-local service you run yourself — a database, a
|
||||
scratch HTTP endpoint — open its port on the bridge:
|
||||
Agent containers reach the host only on ports 80 and 443 (plus DNS). The
|
||||
collector port needs nothing from you: `otel.enable` opens it on the bridge.
|
||||
To let agents reach some other host-local service of your own →
|
||||
[`networking/network.md`](../networking/network.md#reaching-host-services-exposehostports).
|
||||
|
||||
```nix
|
||||
services.hyperhive.network.exposeHostPorts = [ 5432 ];
|
||||
```
|
||||
## Options reference
|
||||
|
||||
and point whatever consumes it at `10.42.0.1:5432` rather than loopback: inside
|
||||
a container, loopback is the _container_. The bridge IP is the host's address on
|
||||
the `hive-br0` bridge. The service must also bind an address the bridge can
|
||||
reach — a `127.0.0.1`-only listener stays unreachable no matter what the
|
||||
firewall allows. See `docs/networking/network.md::Reaching host services` for details.
|
||||
|
||||
<!-- vale write-good.Passive = NO -->
|
||||
|
||||
⚠️ **None of this is needed for hyperhive's own telemetry** — `otel.enable`
|
||||
contributes the collector's port and derives the agent-facing endpoint itself.
|
||||
|
||||
<!-- vale write-good.Passive = YES -->
|
||||
The nix module (`nix/host-modules/otel.nix`) generates every
|
||||
`services.hyperhive.otel.*` option's full type, default, description and
|
||||
example into the [options reference](https://hyperhive.darkest.space/options/)
|
||||
(host options — `nix build .#docs-host` for a local render). That's the
|
||||
reference; this page covers what a per-option listing can't: the two-tier
|
||||
architecture, the security model, and how the options interact.
|
||||
|
||||
## Built-in resource labels
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue