Compare commits

...
19 changed files with 309 additions and 423 deletions

View file

@ -142,6 +142,51 @@ One-shot rows fire once (if past due, on the next worker pass) and are deleted b
`targets` is its own table (`scheduled_prompt_targets`) so partial cancellation flips a single row and the dashboard can show last-fired / last-result per recipient. Cancelling every target reaps the parent row on the next worker pass. `targets` is its own table (`scheduled_prompt_targets`) so partial cancellation flips a single row and the dashboard can show last-fired / last-result per recipient. Cancelling every target reaps the parent row on the next worker pass.
### Missing-target failure
When a target name doesn't resolve to a known agent (container
destroyed, operator typo, etc.) the worker:
1. Records `last_result = "no such agent: <name>"` on the
per-target row.
2. Sends a single advisory `Message` from `system` to `operator`
naming the schedule, target, and reason.
3. Continues fanning out to the other live targets.
Transient broker errors (sqlite lock contention, etc.) get the same
`last_result` annotation plus a `tracing::warn`, and then:
- **Recurring rows** re-arm to the next interval slot — the retry
self-heals on the next worker pass.
- **One-shot rows** are deleted unconditionally after their single
fan-out pass; a broker error on a one-shot is not retried (the
operator advisory and `last_result` are the only audit trail).
### Reminder delivery: file-path semantics
A reminder may carry a `file_path` (the agent-visible path inside its
container, e.g. `/agents/<name>/state/foo.md`). On delivery hive-c0re:
1. **Translates** the container path to the host path
(`/var/lib/hyperhive/agents/<name>/state/foo.md`) so c0re can write
from outside the container.
2. **Validates** the path: rejects anything outside the agent's own state
subtree, containing `..` (path traversal), or with an empty relative
tail. On rejection the write is skipped and the original message is
delivered inline with a warning — the reminder still fires.
3. **Defends against symlink escape**: after `create_dir_all`, the parent
dir is canonicalized and re-verified to live under the agent's host
state root. The final file is opened with
`O_NOFOLLOW | O_CREAT | O_TRUNC` so an existing symlink at the
basename cannot redirect the write to an arbitrary host path.
4. **Writes the body to disk** and delivers a short pointer message in its
place, keeping the agent's inbox / wake-prompt small while the bulky
payload is read out of band.
Atomicity of the inbox INSERT + `reminders.sent_at` UPDATE is handled
inside `Broker::deliver_reminders_batch`; the scheduler only computes the
body strings before calling it.
### Destroy semantics ### Destroy semantics
`HostRequest::Destroy { name, purge }` is the lifecycle tear-down, `HostRequest::Destroy { name, purge }` is the lifecycle tear-down,

View file

@ -113,6 +113,67 @@ Key operations:
--- ---
## Container lifecycle (`lifecycle.rs`)
Every container operation ultimately calls into `lifecycle.rs`. Two paths exist:
**rebuild** (existing container) and **spawn** (first-time creation).
### Rebuild path (existing container)
Goal: apply the new system profile and any `EXTRA_NSPAWN_FLAGS` / drop-in changes
in a single start, with minimum downtime.
`nixos-container update` only runs `systemctl reload container@<c>` when the
container is already up (per `isContainerRunning` in `nixos-container.pl`). Stopping
first turns `update` into a boot-style operation: it builds + `nix-env --set`s the
new profile and skips the in-container `switch-to-configuration`. The subsequent
`start` then applies both the new profile and any `EXTRA_NSPAWN_FLAGS` changes in
one go, rather than the double-bounce a live `update` would trigger.
Sequence for a running container:
1. `prebuild_toplevel` — build the new `system.build.toplevel` **before** stopping.
The container keeps serving the previous generation while eval + fetch + build
happen out-of-band. `nixos-container update` then finds the result cached and
skips straight to the profile-swap. Build failures surface here, before the
running container is touched.
2. `nixos-container stop` — bring the container down.
3. `nixos-container update --flake meta#<name>` — profile-swap (near-instant after
the prebuild).
4. `nixos-container start` — boot into the new generation; the in-container
activation script transitions old → new.
If the container is already stopped, step 1 is skipped (no downtime to shave — no
point evaluating the flake twice).
### Cold-start fallback
`start` after `update` can exit non-zero when packages are **removed** between
generations: the old-generation activation script references units that no longer
exist in the new closure, causing systemd to exit non-zero. The container may be
half-started at that point.
Fallback: `stop` (graceful SIGTERM drain) → `kill` (SIGKILL any lingering processes)
`start` (clean cold-start, no generation transition, new activation runs cleanly).
Both errors are preserved and surfaced if the cold-start also fails.
### Spawn path (new container)
For a first-time `create`, `nixos-container create` is atomic: if the build fails,
no container record is left to clean up. A separate prebuild would just duplicate
the eval, so it's skipped. Sequence: `create --flake meta#<name>` → write nspawn
flags → `systemctl daemon-reload``start`.
### Prebuild attr path
`nix build` does not auto-resolve `meta#<name>` against `nixosConfigurations` the
way `nixos-container` does internally. The explicit attr path
`<flake-root>#nixosConfigurations.<name>.config.system.build.toplevel` is required;
using the bare `meta#<name>` ref would make nix look in `packages`, `legacyPackages`,
or the flake root directly — none of which exist in the rendered meta flake.
---
## See also ## See also
- `docs/approvals.md` — approval flow + scheduled prompts - `docs/approvals.md` — approval flow + scheduled prompts

View file

@ -8,6 +8,37 @@ handling live in [`docs/gateway.md`](gateway.md); this file owns the
per-agent integration story and the notification pump that wakes per-agent integration story and the notification pump that wakes
each agent on relevant activity. each agent on relevant activity.
## Token scopes
Two scope sets live in `hive-c0re::forge`:
**`TOKEN_SCOPES`** (per-agent tokens):
| Scope | Why |
|-------|-----|
| `write:repository` | Create, clone, push, delete repos; merge PRs. |
| `write:issue` | Open / comment / review issues **and** pull requests (Forgejo namespaces PR conversation under issues). |
| `write:user` | Edit own profile, create repos under own user. |
| `write:organization` | Create + manage orgs (lets agents share a forge namespace). |
| `read:user` | Token-owner endpoint used for self-identification at harness startup. |
| `write:misc` | Hooks, attachments, the rest of the long tail. |
| `read:notification` | Poll `GET /notifications` for unread events. |
| `write:notification` | Mark notifications read via `PATCH /notifications/threads/{id}`. |
**`CORE_TOKEN_SCOPES`** (hive-c0re's own `core` user): everything in
`TOKEN_SCOPES` plus `read:admin` and `write:admin`. Site-admin
membership alone isn't sufficient — Forgejo's token scope gate runs
before the user-permission check, so `/api/v1/admin/*` returns
`403 Forbidden` for any token without the admin scope bits, even when
the bearer is a site admin.
**Migration note**: if `PATCH /api/v1/admin/users/{name}` returns 403
on an existing deploy, the core token predates the admin-scope
addition. Delete `/var/lib/hyperhive/forge-core-token` and restart
hive-c0re to re-mint with the new scopes.
---
## Per-agent forge accounts ## Per-agent forge accounts
Each agent gets its own Forgejo user + access token, provisioned at Each agent gets its own Forgejo user + access token, provisioned at

View file

@ -114,6 +114,35 @@ TCP loopback upstream in `agents.conf` (deterministic port from
`agent_web_port(name)`). A future cleanup will drop the TCP fallback `agent_web_port(name)`). A future cleanup will drop the TCP fallback
once every agent's flipped. once every agent's flipped.
## Agent port map (`agent-ports.json`)
`/var/lib/hyperhive/agent-ports.json` is a flat JSON object keyed by
logical agent name → TCP web port:
```json
{
"iris": 8178,
"atlas": 8304,
"argus": 8267,
"damocles": 8549
}
```
Written alongside `agents.conf` on every topology change. Ports come from
`lifecycle::agent_web_port(name)` — a pure FNV-1a hash of the name,
reproducible from the name alone. The manager is excluded: the gateway
routes `/` directly to c0re's dashboard upstream, not through a
per-agent `/agent/<name>/` prefix.
The file doubles as a human-readable audit artifact — `cat agent-ports.json`
shows every registered sub-agent and its deterministic port assignment. TCP
loopback upstreams in `agents.conf` reference these ports for agents that
haven't opted into unix-socket mode yet.
Both `agent-ports.json` and `agents.conf` use atomic `<path>.tmp` +
`rename()` writes so a crashing c0re process never leaves a partial or
unparseable file behind.
## Dashboard link shape (gateway vs direct) ## Dashboard link shape (gateway vs direct)
When the gateway is in front, the SW4RM tab builds per-agent links When the gateway is in front, the SW4RM tab builds per-agent links

View file

@ -127,6 +127,28 @@ sentinel files (`hyperhive-rate-limited`, `hyperhive-needs-login`) if the
JSON is absent, so existing containers keep working through the transition JSON is absent, so existing containers keep working through the transition
window before their next rebuild. window before their next rebuild.
### `/var/lib/hyperhive/build_logs.sqlite` (host)
Full stdout + stderr capture for every `nixos-container` / `nix
build` invocation the lifecycle layer fires. One row per invocation;
the row accumulates lines as the child runs.
Replaces the legacy 32-line stderr ring buffer that `lifecycle::run`
kept. The ring tail routinely truncated real eval errors ("tried
alternatives" blocks alone are often 30+ lines), so failures bailed
with an arbitrary tail whose full stream only lived in the host
journal. With this table the dashboard can surface the entire log.
Two indices:
- `(agent, started_at)` — backs the per-agent latest-N lookup used
by the agent card chip.
- `(status, finished_at)` — backs the retention sweep that runs
as part of the existing hourly vacuum.
Writes are best-effort: `append_stdout` / `append_stderr` / `finish`
log a warning on sqlite error and let the build continue. A failed
log row never blocks a rebuild.
### `/state/hyperhive-model` (per agent) ### `/state/hyperhive-model` (per agent)
Single-line text file holding the claude model name currently Single-line text file holding the claude model name currently

View file

@ -1,5 +1,42 @@
# Security model # Security model
## State-file endpoint security model
`GET /api/state-file?path=<p>` serves files from agent state dirs and
the shared space to authenticated dashboard users (browser, operator).
Two allow-listed root prefixes are accepted; all other paths are rejected
before touching the filesystem:
- `/var/lib/hyperhive/agents/<n>/state/` — per-agent durable notes
(canonical host form or the in-container view `/agents/<n>/state/`)
- `/var/lib/hyperhive/shared/` — shared docs (`/shared/` in-container)
`/state/...` without an agent prefix is explicitly *not* accepted — it is
ambiguous from the host's perspective.
Defense-in-depth layers (in order):
1. **Allow-list prefix check** — rejects without touching the filesystem
if the path doesn't match either root.
2. **No symlinks below the matched root** — each path component is
checked with `symlink_metadata` before canonicalize. A sub-agent
that plants `ln -s /other/secret /agents/me/state/peek` can't proxy
another agent's file through this endpoint (canonicalize would
happily resolve the symlink to a still-within-allow-list path).
3. **Canonicalize as belt-and-braces** — resolves `..`/`.` traversal
and rejects if the result escapes the roots.
4. **`state/` subdir constraint** — under `AGENTS_ROOT`, the second
path component must be `state/`. Applied, proposed git repos and
config dirs are off-limits.
5. **World-readable check** — file must have `mode & 0o004` set.
A `0600` file inside `state/` would otherwise be accessible to any
operator with dashboard access.
`scan_validated_paths` (broker-message ingest, linkifier) uses the same
`resolve_state_path` helper so security rules stay in sync — the
dashboard renders anchors only for tokens that passed the same checks the
read endpoint enforces.
## Nix builds and credential isolation ## Nix builds and credential isolation
### Background ### Background

View file

@ -46,6 +46,14 @@
~200 broker messages wrapped in `{ seq, events }`) on the ~200 broker messages wrapped in `{ seq, events }`) on the
dashboard and `GET /events/history` (last 2000 `LiveEvent`s dashboard and `GET /events/history` (last 2000 `LiveEvent`s
also wrapped in `{ seq, events }`) on the agent. also wrapped in `{ seq, events }`) on the agent.
**One unified channel**: browsers cap concurrent SSE
connections per origin (~6 in Chrome). Using one channel per
domain would exhaust this budget on a live hive; dispatching
by `kind` on the client is a one-liner. Per-domain splits are
reserved for high-volume sub-streams most consumers skip (none
exist yet). The broker's intra-process channel stays separate
from the dashboard channel to avoid coupling `recv_blocking_batch`
(hot path inside the harness turn loop) to presentation concerns.
**SSE multiplexing**: the dashboard uses a **SSE multiplexing**: the dashboard uses a
`SharedWorker` (`stream-worker.js`) to hold one upstream `SharedWorker` (`stream-worker.js`) to hold one upstream
`EventSource` per URL. All same-origin tabs share this worker `EventSource` per URL. All same-origin tabs share this worker

View file

@ -1,33 +1,7 @@
//! `/var/lib/hyperhive/agent-ports.json` writer. Port map for //! `/var/lib/hyperhive/agent-ports.json` writer — flat map of
//! per-agent `/agent/<name>/` TCP routing. Written alongside //! agent name → TCP web port. Written alongside `agents.conf` on
//! `agents.conf` (see `gateway_nginx.rs`) on every topology change; //! every topology change. JSON shape, port derivation (FNV-1a hash),
//! `gateway_nginx::render` reads it indirectly via //! atomicity, and manager exclusion: `docs/gateway.md::Agent port map`.
//! `lifecycle::agent_web_port` to populate TCP upstreams for agents
//! that haven't opted in to unix-socket mode yet. Also kept as a
//! human-readable audit file — `cat agent-ports.json` shows every
//! registered sub-agent and its deterministic port assignment.
//!
//! Shape (flat object keyed by logical agent name → web port):
//!
//! ```json
//! {
//! "iris": 8178,
//! "atlas": 8304,
//! "argus": 8267,
//! "damocles": 8549
//! }
//! ```
//!
//! Ports come from [`crate::lifecycle::agent_web_port`] — pure
//! FNV-1a(name) hash so the value is reproducible from the name
//! alone. Manager is excluded from the map: the gateway routes `/`
//! straight to it via the c0re dashboard upstream rather than a
//! per-agent `/agent/<name>/` prefix.
//!
//! Atomicity: write to a sibling `.tmp` file + rename so a partial
//! write never leaves an unparseable file in place. The gateway's
//! `nginx` worker can read mid-write and Just Work because `rename()`
//! is atomic on the same filesystem.
use std::collections::BTreeMap; use std::collections::BTreeMap;
use std::path::PathBuf; use std::path::PathBuf;

View file

@ -1,23 +1,7 @@
//! Sqlite-backed full build-log capture. One row per `nixos-container` //! Sqlite-backed full build-log capture — stdout + stderr per
//! / `nix build` invocation that the host-side lifecycle layer fires; //! `nixos-container` / `nix build` invocation, accumulated live.
//! the row accumulates stdout + stderr line-by-line as the child runs. //! Schema, indices, retention, and the rationale for replacing
//! //! the old ring buffer: `docs/persistence.md::/var/lib/hyperhive/build_logs.sqlite`.
//! Replaces the legacy 32-line stderr ring buffer in
//! `lifecycle::run` / `lifecycle::prebuild_toplevel`. The ring tail
//! routinely truncated the actual eval error (a "tried alternatives"
//! block alone is often 30+ lines), so failures bailed with an
//! arbitrary tail and the full stream only lived in the host journal.
//! With this table the dashboard can surface the entire log.
//!
//! Storage lives next to the broker / approvals dbs (one file at
//! `<db_path>/build_logs.sqlite`). Two indices:
//! `(agent, started_at)` for the per-agent latest-N lookup that backs
//! the agent card chip; `(status, finished_at)` for the retention
//! sweep that runs as part of the existing hourly vacuum.
//!
//! Writes are best-effort: every `append_*` / `finish` call logs a
//! warning on sqlite error and lets the build continue. A failed log
//! row never breaks a rebuild.
use std::path::Path; use std::path::Path;
use std::sync::{Arc, Mutex, OnceLock}; use std::sync::{Arc, Mutex, OnceLock};

View file

@ -1,19 +1,7 @@
//! Per-container state watcher. Polls every managed container on a //! Per-container crash and login-state watcher. Polls every managed
//! fixed interval, tracks two orthogonal state-sets across ticks, //! container on a 10s interval. Fires `ContainerCrash`, `LoggedIn`,
//! and emits a `HelperEvent` to the manager on each transition: //! and `NeedsLogin` helper events. Event semantics and the
//! //! `RECENT_TRANSIENT_GRACE` window: `docs/approvals.md::Helper events`.
//! - **running**: container is up. running → stopped without an
//! operator-initiated transient (`Stopping` / `Restarting` /
//! `Destroying` / `Rebuilding`) → `ContainerCrash`.
//! - **logged-in**: claude session dir is populated. ! → ✓ →
//! `LoggedIn`; ✓ → ! → `NeedsLogin` (rare — usually only fires
//! on a fresh spawn / purge).
//!
//! `NeedsUpdate` events are now fired from the apply-commit path
//! directly rather than via rev-marker polling.
//!
//! D-Bus subscription would be lower-latency for the first axis,
//! but polling is simpler and a 10s detection delay is fine.
use std::collections::HashSet; use std::collections::HashSet;
use std::sync::Arc; use std::sync::Arc;

View file

@ -94,13 +94,8 @@ pub async fn serve(port: u16, coord: Arc<Coordinator>) -> Result<()> {
// /static/dashboard.css → dist/static/dashboard.css, etc.). // /static/dashboard.css → dist/static/dashboard.css, etc.).
.fallback_service(ServeDir::new(&static_dir)) .fallback_service(ServeDir::new(&static_dir))
.with_state(AppState { coord }); .with_state(AppState { coord });
// Bind loopback-only. External access funnels through // Binds loopback-only; external access via gateway.
// hive-gateway (in-host-netns nginx container), which proxies // Rationale: docs/gateway.md::Firewall posture.
// `/` → `127.0.0.1:<dashboardPort>` upstream. Operators who opt
// out of the gateway lose remote dashboard access — that's by
// design; the c0re HTTP surface is privileged (approve / deny /
// destroy, etc.) and any external exposure needs to pass through
// a real reverse proxy with auth.
let addr = SocketAddr::from(([127, 0, 0, 1], port)); let addr = SocketAddr::from(([127, 0, 0, 1], port));
let listener = bind_with_retry(addr).await?; let listener = bind_with_retry(addr).await?;
tracing::info!(%addr, "dashboard listening"); tracing::info!(%addr, "dashboard listening");
@ -108,27 +103,10 @@ pub async fn serve(port: u16, coord: Arc<Coordinator>) -> Result<()> {
Ok(()) Ok(())
} }
// --------------------------------------------------------------------------- // SPA shape + SSE channels: docs/web-ui/shape.md.
// The dashboard is an SPA. Its HTML shell + bundled JS / CSS / favicon
// live in the directory pointed at by `HIVE_STATIC_DIR` (set by the
// hive-c0re NixOS module to `${frontend}/dashboard`), served by the
// `tower_http::ServeDir` fallback declared in `serve()`. The dynamic
// surface — `/api/state` and the action endpoints — is owned here.
// The JS app fetches state on load, re-fetches after every async-form
// submit, and listens on `/dashboard/stream` for the unified live event
// channel.
// ---------------------------------------------------------------------------
/// `SO_REUSEADDR` bind with retry. Mirrors the per-agent variant in /// `SO_REUSEADDR` bind with retry. Retry mechanics, attempt-cap
/// `hive-ag3nt::web_ui::bind_with_retry`: hive-c0re restarts also /// rationale, and log-level cadence: `docs/web-ui/shape.md::Listener bind`.
/// race the previous process's socket release, and the retry has no
/// attempt cap — capping was the proximate cause of a silent
/// give-up on a long stale socket. Genuine port collisions
/// don't reach this layer (dashboard is bound to a fixed configured
/// port, no per-agent hashing), so any persistent `AddrInUse` always
/// reflects a recoverable stale socket. WARN for the first dozen
/// attempts; INFO after that to avoid spamming the journal during a
/// long hold; INFO on eventual success when we did have to retry.
async fn bind_with_retry(addr: SocketAddr) -> Result<tokio::net::TcpListener> { async fn bind_with_retry(addr: SocketAddr) -> Result<tokio::net::TcpListener> {
let mut delay_ms = 250u64; let mut delay_ms = 250u64;
let mut attempts = 0u32; let mut attempts = 0u32;
@ -1211,54 +1189,10 @@ struct StateFileQuery {
path: String, path: String,
} }
/// Bounded-size read of a file under one of two allow-listed /// Resolve a caller-supplied path against the allow-listed roots
/// roots: `/var/lib/hyperhive/agents/<n>/state/` (per-agent durable /// (`agents/<n>/state/` and `shared/`). Applies defense-in-depth
/// notes — the only writable path agents have outside their /// symlink + traversal checks before serving. Security model and
/// container) and `/var/lib/hyperhive/shared/` (shared docs). Both /// all five layers: `docs/security.md::State-file endpoint`.
/// path forms are accepted:
/// - canonical host: `/var/lib/hyperhive/agents/alice/state/foo.md`
/// - container view: `/agents/alice/state/foo.md`
/// - shared: `/shared/foo.md`
///
/// `/state/...` on its own is *not* accepted — the in-container
/// mount is ambiguous from the host's perspective (we don't know
/// which agent's `/state` it refers to) and using it would silently
/// resolve to the wrong file.
///
/// Path is canonicalised before the allow-list check so `..`
/// traversal and symlink games can't escape the roots. Files larger
/// than `MAX_BYTES` are truncated with a banner so a runaway log
/// can't OOM the browser.
/// Resolve a caller-supplied path string to a canonical host path
/// that has been verified against the allow-list. Returns `Err`
/// with a human-readable reason for every failure mode (path
/// outside roots, canonicalize failure, escape via symlink,
/// per-agent subdir not `state`, symlink anywhere below the root,
/// file not world-readable). Shared by `get_state_file` (read) and
/// `scan_validated_paths` (linkify candidates in message bodies)
/// so both apply identical security rules and the linkifier
/// doesn't render a path the reader will refuse to serve.
///
/// Defense-in-depth layers (in order):
/// 1. Caller-supplied prefix has to match the allow-list (agents/
/// or shared/), else reject without touching the fs.
/// 2. No symlinks below the matched root. Walked pre-canonicalize
/// via `symlink_metadata` on each component so a sub-agent that
/// plants `ln -s /var/lib/hyperhive/agents/other/state/secret
/// /agents/me/state/peek` can't proxy a different agent's file
/// through this endpoint (canonicalize would happily resolve
/// the symlink to a path inside the allow-list).
/// 3. Canonicalize is run anyway as a belt-and-braces check —
/// resolves `..`/`.` traversal and rejects if the result
/// escapes the roots.
/// 4. Under `AGENTS_ROOT`, the second path component must be
/// `state/` — agents' applied/proposed git repos and config dirs
/// are off-limits.
/// 5. The target's metadata is fetched once and returned to the
/// caller so they don't restat. If the target is a regular
/// file it must be world-readable (mode & 0o004); a 0600 file
/// inside `state/` could leak through this endpoint to anyone
/// holding the dashboard URL otherwise.
fn resolve_state_path( fn resolve_state_path(
raw: &str, raw: &str,
) -> std::result::Result<(std::path::PathBuf, std::fs::Metadata), String> { ) -> std::result::Result<(std::path::PathBuf, std::fs::Metadata), String> {
@ -1571,21 +1505,10 @@ pub(crate) fn emit_meta_inputs_snapshot(coord: &Coordinator) {
}); });
} }
/// Scan `body` for path-shaped tokens, validate each against the /// Scan `body` for path-shaped tokens and return those that pass the
/// allow-list, return the unique set of tokens that resolve to a /// allow-list + `is_file` check via `resolve_state_path`. Called at
/// regular file. Called at broker-message ingest time so the /// broker-message ingest so the dashboard event already carries the
/// dashboard event already carries the verified set — no client- /// verified set; security rules stay in sync with the read endpoint.
/// side probe endpoint required, and historical messages get the
/// same treatment on `/dashboard/history` backfill.
///
/// Tokenisation: split on whitespace + a handful of trailing
/// punctuation chars (`,;:)]}`) that commonly follow paths in
/// natural-language text but aren't part of the path itself. Any
/// token starting with `/agents/`, `/shared/`, or
/// `/var/lib/hyperhive/{agents,shared}/` is a candidate. The
/// allow-list + `is_file` check happens via the same
/// `resolve_state_path` helper the read endpoint uses, so the
/// security rules can't drift.
pub fn scan_validated_paths(body: &str) -> Vec<String> { pub fn scan_validated_paths(body: &str) -> Vec<String> {
const PREFIXES: [&str; 4] = [ const PREFIXES: [&str; 4] = [
"/agents/", "/agents/",

View file

@ -1,27 +1,8 @@
//! Unified dashboard event channel. //! Unified dashboard event channel — all near-real-time browser events
//! //! flow through `Coordinator.dashboard_events`. Each event carries a
//! Anything the browser wants to react to in near-real-time flows through //! monotonic `seq` for client-side dedupe against `/api/state` snapshots.
//! `Coordinator.dashboard_events`. Each event is stamped with a monotonic //! Design rationale (single channel, broker forwarder isolation):
//! per-process `seq` so the client can dedupe its buffered live traffic //! `docs/web-ui/shape.md::One unified channel`.
//! against snapshot/history responses (drop frames with
//! `seq <= snapshot.seq`).
//!
//! Why one channel instead of one-per-domain: browsers cap concurrent
//! SSE connections per origin (~6 in chrome) and dispatch-by-kind on the
//! client is a one-liner. Splits get reserved for high-volume sub-streams
//! that most consumers don't care about (none yet).
//!
//! Message-broker traffic (`Sent` / `Delivered`) lives on this channel
//! too. A background forwarder task in `main.rs` subscribes to the broker
//! and re-emits each `MessageEvent` as a `DashboardEvent::Sent` /
//! `DashboardEvent::Delivered` with a freshly-stamped seq. Keeping the
//! broker's intra-process channel separate avoids coupling the broker
//! (used by `recv_blocking_batch` inside the harness loop) to dashboard
//! presentation concerns.
//!
//! New mutation kinds (approval added/resolved, question added/answered,
//! transient changed, etc.) land here as additional variants. The client
//! dispatches by `kind` and updates the relevant section.
use serde::Serialize; use serde::Serialize;

View file

@ -1,21 +1,6 @@
//! Optional Forgejo wiring. When the `hive-forge` nixos-container is //! Optional Forgejo wiring — per-agent user + token provisioning,
//! present and running, hive-c0re ensures every agent (and the //! config-repo mirroring, meta read-access grants. No-op when
//! manager) has a corresponding forgejo user with an API token //! `hive-forge` isn't running. Full design: `docs/forge.md`.
//! written to `<agent-state>/forge-token` — visible inside the
//! container as `/state/forge-token`. Idempotent: skips creation
//! when the user already exists, skips token issuance when the file
//! is already there.
//!
//! It also mirrors each agent's hive-c0re-owned *applied* config repo
//! into the private `agent-configs` org (`push_config`), so every
//! deploy / approval tag core plants is visible on the forge. Each
//! agent is a read-only collaborator on `core/meta` (the meta flake)
//! so they can fetch their deployment context; the `agent-configs`
//! repos remain core-only.
//!
//! No-op when `hive-forge` isn't enabled (detected via
//! `nixos-container list`), so operators who don't run the bundled
//! forge pay nothing.
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
@ -56,32 +41,15 @@ const CONFIG_ORG: &str = "agent-configs";
/// Forgejo orgs hive-c0re ensures on startup. The meta repo lives at /// Forgejo orgs hive-c0re ensures on startup. The meta repo lives at
/// `core/meta` (the `core` user's own namespace — no org needed). /// `core/meta` (the `core` user's own namespace — no org needed).
const SEEDED_ORGS: &[&str] = &[CONFIG_ORG]; const SEEDED_ORGS: &[&str] = &[CONFIG_ORG];
/// Forgejo scopes the agent's token gets. Broad-but-not-admin: every /// Per-agent token scopes (broad-but-not-admin). See
/// repo / PR / issue thing an agent needs day-to-day, no admin /// `docs/forge.md::Token scopes` for the per-scope rationale.
/// surface.
/// - `write:repository` — create, clone, push, delete repos in the
/// user's own namespace; merge PRs.
/// - `write:issue` — open / comment / review issues *and* pull
/// requests (forgejo namespaces PR conversation under issues).
/// - `write:user` — edit own profile, create repos under own user.
/// - `write:organization` — create + manage orgs (lets agents share
/// a forge namespace).
/// - `read:user` — token-owner endpoint clients call to introspect.
/// - `write:misc` — hooks, attachments, the rest of the long tail.
/// - `read:notification` — required by `forge_notify` to poll
/// `GET /notifications` for unread PR/review events.
/// - `write:notification` — required by `forge_notify` to mark
/// notifications as read via `PATCH /notifications/threads/{id}`.
const TOKEN_SCOPES: &str = "read:user,write:user,read:notification,write:notification,write:repository,write:issue,write:organization,write:misc"; const TOKEN_SCOPES: &str = "read:user,write:user,read:notification,write:notification,write:repository,write:issue,write:organization,write:misc";
/// Scopes for the bootstrap `core` token used by hive-c0re itself. /// Bootstrap `core` token scopes — adds `read:admin,write:admin` on
/// Adds `read:admin,write:admin` on top of `TOKEN_SCOPES` so the host /// top of `TOKEN_SCOPES` so the host daemon can drive
/// daemon can drive `/api/v1/admin/*` endpoints (PATCH user email on /// `/api/v1/admin/*`. Site-admin membership alone isn't enough: the
/// agent provision, future webhook + org admin work). Site-admin /// token's own scope gate runs before the user-permission check.
/// membership alone isn't enough — the token's own scope gate runs /// See `docs/forge.md::Token scopes`.
/// before the user-permission check, so `403 Forbidden` comes back
/// for any `/admin/users/*` call from a non-admin-scoped token
/// even if the bearer is an admin user.
const CORE_TOKEN_SCOPES: &str = "read:admin,write:admin,read:user,write:user,read:notification,write:notification,write:repository,write:issue,write:organization,write:misc"; const CORE_TOKEN_SCOPES: &str = "read:admin,write:admin,read:user,write:user,read:notification,write:notification,write:repository,write:issue,write:organization,write:misc";
/// Token file inside the agent's bind-mounted state dir (visible as /// Token file inside the agent's bind-mounted state dir (visible as
@ -271,10 +239,8 @@ async fn ensure_user_email(name: &str) {
tracing::debug!(%name, %email, "forge: user email aligned"); tracing::debug!(%name, %email, "forge: user email aligned");
} }
Ok(status) if status == reqwest::StatusCode::FORBIDDEN => { Ok(status) if status == reqwest::StatusCode::FORBIDDEN => {
// Almost certainly an existing-deployment migration case: // Core token missing admin scope — see
// an older core token may have been minted without admin // `docs/forge.md::Token scopes` migration note.
// scope, so /admin/users/* now returns 403 even though
// `core` is a site admin. Tell the operator how to fix it.
tracing::warn!( tracing::warn!(
%name, %email, %status, %name, %email, %status,
"forge: PATCH user email forbidden — core token likely missing admin scope. \ "forge: PATCH user email forbidden — core token likely missing admin scope. \

View file

@ -1,30 +1,9 @@
//! Runtime nginx include-file generator for the gateway's per-agent //! Runtime nginx include-file generator for the gateway's per-agent
//! `/agent/<name>/` location blocks. //! `/agent/<name>/` location blocks. Writes
//! //! `/var/lib/hyperhive/gateway/agents.conf` on every topology change.
//! Writes `/var/lib/hyperhive/gateway/agents.conf` on every topology change. //! UDS vs TCP upstream selection, reload trigger (`systemd-run
//! The gateway container bind-mounts `/var/lib/hyperhive/gateway/` (NOT the //! --machine=hive-gateway`), and idempotency:
//! whole parent dir) at `/run/hive-state/`; nginx includes //! `docs/gateway.md::Per-agent unix-socket upstream`.
//! `/run/hive-state/agents.conf`. After each write, c0re triggers
//! `nginx -s reload` inside the gateway container via
//! `systemd-run --machine=hive-gateway` from the host — no `nixos-rebuild
//! switch` needed when agents start, stop, or flip `useUnixSocket`.
//! (A path unit inside the container was tried first but `IN_MOVED_TO`
//! from the atomic rename does not cross the nspawn mount-namespace
//! boundary — see `docs/gateway.md` for the failure analysis.)
//!
//! Upstream selection mirrors `agent_sockets::build_map`: an agent
//! gets a UDS upstream when its `.bound` marker exists (harness has
//! bound the unix socket); otherwise falls back to the deterministic
//! TCP port from `lifecycle::agent_web_port`. Proxy headers are
//! emitted in full so the generated file is self-contained nginx
//! config — no dependency on which `recommendedProxySettings` knobs
//! the host config has on.
//!
//! `write()` is idempotent: if the rendered body equals what's already
//! on disk, the rename is skipped and the path unit doesn't fire.
//! Same atomic `<path>.tmp` + `rename()` shape as `agent_ports` /
//! `agent_sockets` — a crashing c0re process never leaves a partial
//! file the gateway's nginx would fail to parse.
use anyhow::{Context, Result}; use anyhow::{Context, Result};
use std::fmt::Write as _; use std::fmt::Write as _;

View file

@ -408,38 +408,13 @@ pub async fn rebuild_no_meta(
let container = container_name(name); let container = container_name(name);
let flake_ref = format!("{}#{name}", crate::meta::meta_dir().display()); let flake_ref = format!("{}#{name}", crate::meta::meta_dir().display());
if container_exists(name).await { if container_exists(name).await {
// Existing container: preserve the prior running state across // Rebuild strategy: stop-before-update + pre-build.
// rebuild, and apply both the new system profile // See `docs/coordinator.md::Container lifecycle`.
// AND any `/etc/nixos-containers/<c>.conf` / drop-in changes
// in a single start rather than `update`'s reload-then-outer-
// restart double-bounce.
//
// `nixos-container update` only runs `systemctl reload
// container@<c>` when the container is up (per the
// `isContainerRunning` check in nixos-container.pl), so
// stopping first makes `update` boot-style: build + nix-env
// --set the new profile, skip the in-container
// switch-to-configuration, let the next `start` apply both
// the new profile and the new EXTRA_NSPAWN_FLAGS in one go.
// If the container was already stopped, `update` builds + sets
// the profile and we leave it stopped.
let was_running = is_running(name).await; let was_running = is_running(name).await;
set_nspawn_flags(&container, agent_dir, claude_dir, notes_dir)?; set_nspawn_flags(&container, agent_dir, claude_dir, notes_dir)?;
set_resource_limits(&container)?; set_resource_limits(&container)?;
systemd_daemon_reload().await?; systemd_daemon_reload().await?;
if was_running { if was_running {
// Pre-build the system toplevel **before** stopping the
// running container so the agent keeps serving its
// previous generation while the eval + fetch + build
// happens out-of-band. `nixos-container update` then
// finds the toplevel cached and skips straight to the
// profile-swap + restart — downtime collapses to that
// window only. Build failures surface here, before we
// touch the container.
//
// When the container is already stopped there's no
// downtime to shave — let `update` do the build inline
// rather than evaluating the flake twice for nothing.
on_step("nix build"); on_step("nix build");
prebuild_toplevel(name, &flake_ref).await?; prebuild_toplevel(name, &flake_ref).await?;
on_step("nixos-container stop"); on_step("nixos-container stop");
@ -448,18 +423,8 @@ pub async fn rebuild_no_meta(
on_step("nixos-container update"); on_step("nixos-container update");
run(&["update", &container, "--flake", &flake_ref]).await?; run(&["update", &container, "--flake", &flake_ref]).await?;
if was_running { if was_running {
// Normal path: start into the new generation. The activation // Cold-start fallback on activation errors.
// script runs inside the container to transition old → new. // See `docs/coordinator.md::Cold-start fallback`.
// This can fail when packages are removed between generations —
// the old-generation activation references units that no longer
// exist in the new closure, causing systemd to exit non-zero.
//
// Fallback: stop + kill + start (cold-start). The activation
// script can fail when packages are removed between generations —
// `start` exits non-zero but the container may be half-started.
// `stop` requests a graceful SIGTERM drain; `kill` then SIGKILLs
// any lingering processes so the next `start` enters a clean state
// without a generation transition, letting the activation succeed.
on_step("nixos-container start"); on_step("nixos-container start");
if let Err(start_err) = run(&["start", &container]).await { if let Err(start_err) = run(&["start", &container]).await {
tracing::warn!( tracing::warn!(
@ -494,11 +459,8 @@ pub async fn rebuild_no_meta(
Ok(()) Ok(())
} }
} else { } else {
// First spawn: no running container, no downtime to shave. // Spawn path: create is atomic, no prebuild needed.
// `nixos-container create` builds + creates atomically — if // See `docs/coordinator.md::Spawn path`.
// the build fails, no container record is left around to
// clean up — so a pre-build adds nothing but a duplicate
// eval.
on_step("nixos-container create"); on_step("nixos-container create");
run(&["create", &container, "--flake", &flake_ref]).await?; run(&["create", &container, "--flake", &flake_ref]).await?;
set_nspawn_flags(&container, agent_dir, claude_dir, notes_dir)?; set_nspawn_flags(&container, agent_dir, claude_dir, notes_dir)?;
@ -509,29 +471,12 @@ pub async fn rebuild_no_meta(
} }
} }
/// Pre-build the agent's `system.build.toplevel` derivation against /// Pre-build `system.build.toplevel` against `meta#<name>` so the
/// `meta#<name>` so the subsequent `nixos-container update` / /// subsequent `nixos-container update` finds the result cached and
/// `create` finds the result already in the store. The container /// skips straight to the profile-swap. Store-warming only — container
/// itself is untouched — this is purely a store-warming pass. /// is untouched. See `docs/coordinator.md::Rebuild path` for why
/// /// the prebuild happens before stop, and `docs/coordinator.md::Prebuild
/// Streams nix's stdout to INFO and stderr to WARN like the /// attr path` for why the explicit nixosConfigurations attr is required.
/// `nixos-container` shellouts so progress shows up in journald as
/// it happens. `--no-link` keeps us from littering the working
/// directory with `result` symlinks. Per-derivation cost: pure
/// cache hit when nothing changed (handful of seconds for the
/// eval), expensive only on the rebuild that actually has work.
///
/// Attr path is `<flake-root>#nixosConfigurations.<name>.config.
/// system.build.toplevel` — `nix build` won't auto-resolve the bare
/// `<name>` against `nixosConfigurations` like `nixos-container` does
/// internally, so we have to spell the path out explicitly. Falling
/// back to `meta#<name>` (the shape `nixos-container update --flake
/// meta#<name>` uses) makes nix look for `packages.<system>.<name>`,
/// `legacyPackages.<system>.<name>`, or `<name>` at the flake root —
/// none of which exist in the rendered meta flake.
///
/// Returns the same error shape as the other nixos-container
/// helpers so callers can use `?` without translation.
async fn prebuild_toplevel(name: &str, flake_ref: &str) -> Result<()> { async fn prebuild_toplevel(name: &str, flake_ref: &str) -> Result<()> {
use tokio::io::{AsyncBufReadExt, BufReader}; use tokio::io::{AsyncBufReadExt, BufReader};
// Split `<root>#<name>` so we can re-emit with the explicit // Split `<root>#<name>` so we can re-emit with the explicit

View file

@ -1,20 +1,8 @@
//! Single hive-c0re-owned flake at `/var/lib/hyperhive/meta/` that //! Single hive-c0re-owned flake at `/var/lib/hyperhive/meta/` that
//! consumes every agent's applied repo as a flake input and exports one //! exports one `nixosConfiguration` per agent and drives the system-wide
//! `nixosConfiguration` per agent. Containers run against //! deploy audit trail. Flow (`sync_agents`, two-phase `prepare_deploy` /
//! `--flake /var/lib/hyperhive/meta#<name>`; lifecycle ops here drive the //! `finalize_deploy` / `abort_deploy`, `lock_update_hyperhive`):
//! lock file so meta's git log is the system-wide deploy audit trail. //! `docs/approvals.md::Meta flake`.
//!
//! Flow:
//! - `sync_agents` (idempotent) — render `flake.nix` for the current
//! agent set, init the repo on first call, relock if the rendered
//! contents changed, commit. Used by spawn / destroy / startup
//! migration.
//! - `prepare_deploy` + `finalize_deploy` / `abort_deploy` — two-phase
//! for the `request_apply_commit` path so a failed
//! `nixos-container update` leaves no orphan commit in meta. Prepare
//! writes the new lock without committing; finalize commits with the
//! deploy message; abort `git restore`s the lock back.
//! - `lock_update_hyperhive` — one-shot for the auto-update path.
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};

View file

@ -1,23 +1,7 @@
//! Startup auto-migration from the pre-meta layout. Runs before //! Startup auto-migration from the pre-meta layout. Four idempotent
//! `auto_update::run` and consists of four phases, each idempotent: //! phases: applied repo, proposed repo, meta repo, container repoint.
//! //! Kill-switch: `HIVE_SKIP_META_MIGRATION=1`. Full migration sequence
//! 1. Per-agent applied repo: rewrite `flake.nix` to the module-only //! and phase details: `docs/approvals.md::Migration from the pre-tag`.
//! boilerplate if it isn't already, commit, relocate `deployed/0`
//! to HEAD so `setup_applied`'s existence check passes.
//! 2. Per-agent proposed repo: ensure the `applied` git remote
//! points at `/applied/<n>/.git` (re-runs `setup_proposed`'s
//! `ensure_applied_remote` indirectly via a host-side git call).
//! 3. Meta repo: `meta::sync_agents` over the current agent list —
//! init the repo on first call, rerender + relock if anything
//! drifted.
//! 4. Container repoint: for every existing container, run
//! `nixos-container update <c> --flake meta#<name>` so it
//! activates against the meta flake. Guarded by a marker file
//! so the (expensive) phase 4 only runs once across hive-c0re
//! restarts.
//!
//! Env kill-switch: `HIVE_SKIP_META_MIGRATION=1` skips the whole
//! migration. Use when smoke-testing one agent at a time by hand.
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
use std::sync::Arc; use std::sync::Arc;

View file

@ -1,34 +1,7 @@
//! Background loop that drains due reminders out of the broker and //! Background loop that drains due reminders from the broker and
//! delivers them as inbox messages. Mirrors the `events_vacuum` / //! delivers them as inbox messages. 5s poll cadence, shutdown-aware.
//! `crash_watch` shape — a single `spawn(coord)` entry point started //! File-path semantics (path translation, traversal + symlink defense,
//! from `main.rs`. //! pointer delivery): `docs/approvals.md::Reminder delivery`.
//!
//! File-path semantics: a reminder may carry a `file_path` (the
//! agent-visible path inside its container). On delivery we:
//!
//! - Translate the container path (`/agents/<agent>/state/foo.md`) to
//! the host path (`/var/lib/hyperhive/agents/<agent>/state/foo.md`)
//! so hive-c0re can write to it from outside the container.
//! - Reject anything that isn't under the agent's own state subtree,
//! contains `..` (path traversal), or has an empty relative tail.
//! Falling outside the allowed prefix means the file write is
//! skipped and the original message is delivered inline (with a
//! noted warning) — the reminder still fires, just without the
//! payload split.
//! - Defend against symlink escape: after `create_dir_all`, the
//! parent dir is canonicalized and re-verified to live under the
//! agent's host state root. Then we open the final file with
//! `O_NOFOLLOW | O_CREAT | O_TRUNC` so an existing-symlink basename
//! can't redirect the write either. Without this an agent could
//! `ln -s /etc /agents/foo/state/escape` and bounce a write to an
//! arbitrary host path.
//! - Write the reminder body to disk and deliver a short pointer
//! message in its place, so the agent's inbox/wake-prompt stays
//! small and the bulky payload can be read out of band.
//!
//! Atomicity of the inbox INSERT + `reminders.sent_at` UPDATE is handled
//! inside `Broker::deliver_reminders_batch`; this module only computes the
//! body strings before calling it.
use std::io::Write; use std::io::Write;
use std::os::unix::fs::OpenOptionsExt; use std::os::unix::fs::OpenOptionsExt;

View file

@ -1,39 +1,7 @@
//! Background loop that drains due `scheduled_prompts` rows //! Background loop that drains due `scheduled_prompts` rows and fans
//! and fans the body out as inbox `Message`s to each active //! the body to each active target. 5s poll cadence, shutdown-aware.
//! target. Mirrors `reminder_scheduler::spawn` shape: //! Catch-up clamp, missing-target handling, and broker-error retry
//! single `spawn(coord)` entry, 5s poll cadence, shutdown-aware. //! semantics: `docs/approvals.md::Scheduled prompt worker`.
//!
//! ## Catch-up semantics
//!
//! When hive-c0re comes back from being down, a recurring row
//! whose `next_fire_at` is well in the past would otherwise fire
//! N delayed pulses in a row. Instead we fire ONCE and let
//! `ScheduledPrompts::rearm` bump `next_fire_at` to the next
//! interval slot ≥ `now`, recording the skipped-cycle count in
//! the per-target `last_result` so operators see how many
//! firings were caught up rather than losing the signal.
//!
//! ## Missing-target failure
//!
//! When a target name doesn't resolve to a known agent (the
//! container has been destroyed, the operator typo'd a name,
//! etc.) the worker:
//! 1. records `last_result = "no such agent: <name>"` against
//! the per-target row,
//! 2. sends a single advisory `Message` from `system` to
//! `operator` describing the schedule + target + reason,
//! 3. continues fanning out to the other (live) targets.
//!
//! Transient broker errors (sqlite lock contention, etc.) get
//! the per-target `last_result` annotated AND a `tracing::warn`,
//! but the post-fire bookkeeping treats the row the same way it
//! does on a clean fire:
//! - **recurring** rows re-arm — the next interval slot tries
//! the broker send again, so transient errors self-heal.
//! - **one-shots** delete unconditionally after their single
//! fan-out pass; a broker failure on a one-shot is NOT
//! retried (the operator advisory + `last_result` are the only
//! audit trail).
use std::sync::Arc; use std::sync::Arc;
use std::time::Duration; use std::time::Duration;