diff --git a/docs/approvals.md b/docs/approvals.md index 44d8a122..75ad859b 100644 --- a/docs/approvals.md +++ b/docs/approvals.md @@ -142,51 +142,6 @@ One-shot rows fire once (if past due, on the next worker pass) and are deleted b `targets` is its own table (`scheduled_prompt_targets`) so partial cancellation flips a single row and the dashboard can show last-fired / last-result per recipient. Cancelling every target reaps the parent row on the next worker pass. -### Missing-target failure - -When a target name doesn't resolve to a known agent (container -destroyed, operator typo, etc.) the worker: - -1. Records `last_result = "no such agent: "` on the - per-target row. -2. Sends a single advisory `Message` from `system` to `operator` - naming the schedule, target, and reason. -3. Continues fanning out to the other live targets. - -Transient broker errors (sqlite lock contention, etc.) get the same -`last_result` annotation plus a `tracing::warn`, and then: - -- **Recurring rows** re-arm to the next interval slot — the retry - self-heals on the next worker pass. -- **One-shot rows** are deleted unconditionally after their single - fan-out pass; a broker error on a one-shot is not retried (the - operator advisory and `last_result` are the only audit trail). - -### Reminder delivery: file-path semantics - -A reminder may carry a `file_path` (the agent-visible path inside its -container, e.g. `/agents//state/foo.md`). On delivery hive-c0re: - -1. **Translates** the container path to the host path - (`/var/lib/hyperhive/agents//state/foo.md`) so c0re can write - from outside the container. -2. **Validates** the path: rejects anything outside the agent's own state - subtree, containing `..` (path traversal), or with an empty relative - tail. On rejection the write is skipped and the original message is - delivered inline with a warning — the reminder still fires. -3. **Defends against symlink escape**: after `create_dir_all`, the parent - dir is canonicalized and re-verified to live under the agent's host - state root. The final file is opened with - `O_NOFOLLOW | O_CREAT | O_TRUNC` so an existing symlink at the - basename cannot redirect the write to an arbitrary host path. -4. **Writes the body to disk** and delivers a short pointer message in its - place, keeping the agent's inbox / wake-prompt small while the bulky - payload is read out of band. - -Atomicity of the inbox INSERT + `reminders.sent_at` UPDATE is handled -inside `Broker::deliver_reminders_batch`; the scheduler only computes the -body strings before calling it. - ### Destroy semantics `HostRequest::Destroy { name, purge }` is the lifecycle tear-down, diff --git a/docs/coordinator.md b/docs/coordinator.md index 421e0312..5063356e 100644 --- a/docs/coordinator.md +++ b/docs/coordinator.md @@ -113,67 +113,6 @@ Key operations: --- -## Container lifecycle (`lifecycle.rs`) - -Every container operation ultimately calls into `lifecycle.rs`. Two paths exist: -**rebuild** (existing container) and **spawn** (first-time creation). - -### Rebuild path (existing container) - -Goal: apply the new system profile and any `EXTRA_NSPAWN_FLAGS` / drop-in changes -in a single start, with minimum downtime. - -`nixos-container update` only runs `systemctl reload container@` when the -container is already up (per `isContainerRunning` in `nixos-container.pl`). Stopping -first turns `update` into a boot-style operation: it builds + `nix-env --set`s the -new profile and skips the in-container `switch-to-configuration`. The subsequent -`start` then applies both the new profile and any `EXTRA_NSPAWN_FLAGS` changes in -one go, rather than the double-bounce a live `update` would trigger. - -Sequence for a running container: - -1. `prebuild_toplevel` — build the new `system.build.toplevel` **before** stopping. - The container keeps serving the previous generation while eval + fetch + build - happen out-of-band. `nixos-container update` then finds the result cached and - skips straight to the profile-swap. Build failures surface here, before the - running container is touched. -2. `nixos-container stop` — bring the container down. -3. `nixos-container update --flake meta#` — profile-swap (near-instant after - the prebuild). -4. `nixos-container start` — boot into the new generation; the in-container - activation script transitions old → new. - -If the container is already stopped, step 1 is skipped (no downtime to shave — no -point evaluating the flake twice). - -### Cold-start fallback - -`start` after `update` can exit non-zero when packages are **removed** between -generations: the old-generation activation script references units that no longer -exist in the new closure, causing systemd to exit non-zero. The container may be -half-started at that point. - -Fallback: `stop` (graceful SIGTERM drain) → `kill` (SIGKILL any lingering processes) -→ `start` (clean cold-start, no generation transition, new activation runs cleanly). -Both errors are preserved and surfaced if the cold-start also fails. - -### Spawn path (new container) - -For a first-time `create`, `nixos-container create` is atomic: if the build fails, -no container record is left to clean up. A separate prebuild would just duplicate -the eval, so it's skipped. Sequence: `create --flake meta#` → write nspawn -flags → `systemctl daemon-reload` → `start`. - -### Prebuild attr path - -`nix build` does not auto-resolve `meta#` against `nixosConfigurations` the -way `nixos-container` does internally. The explicit attr path -`#nixosConfigurations..config.system.build.toplevel` is required; -using the bare `meta#` ref would make nix look in `packages`, `legacyPackages`, -or the flake root directly — none of which exist in the rendered meta flake. - ---- - ## See also - `docs/approvals.md` — approval flow + scheduled prompts diff --git a/docs/forge.md b/docs/forge.md index c086fc05..ee15c5b1 100644 --- a/docs/forge.md +++ b/docs/forge.md @@ -8,37 +8,6 @@ handling live in [`docs/gateway.md`](gateway.md); this file owns the per-agent integration story and the notification pump that wakes each agent on relevant activity. -## Token scopes - -Two scope sets live in `hive-c0re::forge`: - -**`TOKEN_SCOPES`** (per-agent tokens): - -| Scope | Why | -|-------|-----| -| `write:repository` | Create, clone, push, delete repos; merge PRs. | -| `write:issue` | Open / comment / review issues **and** pull requests (Forgejo namespaces PR conversation under issues). | -| `write:user` | Edit own profile, create repos under own user. | -| `write:organization` | Create + manage orgs (lets agents share a forge namespace). | -| `read:user` | Token-owner endpoint used for self-identification at harness startup. | -| `write:misc` | Hooks, attachments, the rest of the long tail. | -| `read:notification` | Poll `GET /notifications` for unread events. | -| `write:notification` | Mark notifications read via `PATCH /notifications/threads/{id}`. | - -**`CORE_TOKEN_SCOPES`** (hive-c0re's own `core` user): everything in -`TOKEN_SCOPES` plus `read:admin` and `write:admin`. Site-admin -membership alone isn't sufficient — Forgejo's token scope gate runs -before the user-permission check, so `/api/v1/admin/*` returns -`403 Forbidden` for any token without the admin scope bits, even when -the bearer is a site admin. - -**Migration note**: if `PATCH /api/v1/admin/users/{name}` returns 403 -on an existing deploy, the core token predates the admin-scope -addition. Delete `/var/lib/hyperhive/forge-core-token` and restart -hive-c0re to re-mint with the new scopes. - ---- - ## Per-agent forge accounts Each agent gets its own Forgejo user + access token, provisioned at diff --git a/docs/gateway.md b/docs/gateway.md index e974ec2b..b1971bff 100644 --- a/docs/gateway.md +++ b/docs/gateway.md @@ -114,35 +114,6 @@ TCP loopback upstream in `agents.conf` (deterministic port from `agent_web_port(name)`). A future cleanup will drop the TCP fallback once every agent's flipped. -## Agent port map (`agent-ports.json`) - -`/var/lib/hyperhive/agent-ports.json` is a flat JSON object keyed by -logical agent name → TCP web port: - -```json -{ - "iris": 8178, - "atlas": 8304, - "argus": 8267, - "damocles": 8549 -} -``` - -Written alongside `agents.conf` on every topology change. Ports come from -`lifecycle::agent_web_port(name)` — a pure FNV-1a hash of the name, -reproducible from the name alone. The manager is excluded: the gateway -routes `/` directly to c0re's dashboard upstream, not through a -per-agent `/agent//` prefix. - -The file doubles as a human-readable audit artifact — `cat agent-ports.json` -shows every registered sub-agent and its deterministic port assignment. TCP -loopback upstreams in `agents.conf` reference these ports for agents that -haven't opted into unix-socket mode yet. - -Both `agent-ports.json` and `agents.conf` use atomic `.tmp` + -`rename()` writes so a crashing c0re process never leaves a partial or -unparseable file behind. - ## Dashboard link shape (gateway vs direct) When the gateway is in front, the SW4RM tab builds per-agent links diff --git a/docs/persistence.md b/docs/persistence.md index 386e9e0c..f5de1554 100644 --- a/docs/persistence.md +++ b/docs/persistence.md @@ -127,28 +127,6 @@ sentinel files (`hyperhive-rate-limited`, `hyperhive-needs-login`) if the JSON is absent, so existing containers keep working through the transition window before their next rebuild. -### `/var/lib/hyperhive/build_logs.sqlite` (host) - -Full stdout + stderr capture for every `nixos-container` / `nix -build` invocation the lifecycle layer fires. One row per invocation; -the row accumulates lines as the child runs. - -Replaces the legacy 32-line stderr ring buffer that `lifecycle::run` -kept. The ring tail routinely truncated real eval errors ("tried -alternatives" blocks alone are often 30+ lines), so failures bailed -with an arbitrary tail whose full stream only lived in the host -journal. With this table the dashboard can surface the entire log. - -Two indices: -- `(agent, started_at)` — backs the per-agent latest-N lookup used - by the agent card chip. -- `(status, finished_at)` — backs the retention sweep that runs - as part of the existing hourly vacuum. - -Writes are best-effort: `append_stdout` / `append_stderr` / `finish` -log a warning on sqlite error and let the build continue. A failed -log row never blocks a rebuild. - ### `/state/hyperhive-model` (per agent) Single-line text file holding the claude model name currently diff --git a/docs/security.md b/docs/security.md index d740f066..831bed24 100644 --- a/docs/security.md +++ b/docs/security.md @@ -1,42 +1,5 @@ # Security model -## State-file endpoint security model - -`GET /api/state-file?path=

` serves files from agent state dirs and -the shared space to authenticated dashboard users (browser, operator). -Two allow-listed root prefixes are accepted; all other paths are rejected -before touching the filesystem: - -- `/var/lib/hyperhive/agents//state/` — per-agent durable notes - (canonical host form or the in-container view `/agents//state/`) -- `/var/lib/hyperhive/shared/` — shared docs (`/shared/` in-container) - -`/state/...` without an agent prefix is explicitly *not* accepted — it is -ambiguous from the host's perspective. - -Defense-in-depth layers (in order): - -1. **Allow-list prefix check** — rejects without touching the filesystem - if the path doesn't match either root. -2. **No symlinks below the matched root** — each path component is - checked with `symlink_metadata` before canonicalize. A sub-agent - that plants `ln -s /other/secret /agents/me/state/peek` can't proxy - another agent's file through this endpoint (canonicalize would - happily resolve the symlink to a still-within-allow-list path). -3. **Canonicalize as belt-and-braces** — resolves `..`/`.` traversal - and rejects if the result escapes the roots. -4. **`state/` subdir constraint** — under `AGENTS_ROOT`, the second - path component must be `state/`. Applied, proposed git repos and - config dirs are off-limits. -5. **World-readable check** — file must have `mode & 0o004` set. - A `0600` file inside `state/` would otherwise be accessible to any - operator with dashboard access. - -`scan_validated_paths` (broker-message ingest, linkifier) uses the same -`resolve_state_path` helper so security rules stay in sync — the -dashboard renders anchors only for tokens that passed the same checks the -read endpoint enforces. - ## Nix builds and credential isolation ### Background diff --git a/docs/web-ui/shape.md b/docs/web-ui/shape.md index 29dce188..15817abd 100644 --- a/docs/web-ui/shape.md +++ b/docs/web-ui/shape.md @@ -46,14 +46,6 @@ ~200 broker messages wrapped in `{ seq, events }`) on the dashboard and `GET /events/history` (last 2000 `LiveEvent`s also wrapped in `{ seq, events }`) on the agent. - **One unified channel**: browsers cap concurrent SSE - connections per origin (~6 in Chrome). Using one channel per - domain would exhaust this budget on a live hive; dispatching - by `kind` on the client is a one-liner. Per-domain splits are - reserved for high-volume sub-streams most consumers skip (none - exist yet). The broker's intra-process channel stays separate - from the dashboard channel to avoid coupling `recv_blocking_batch` - (hot path inside the harness turn loop) to presentation concerns. **SSE multiplexing**: the dashboard uses a `SharedWorker` (`stream-worker.js`) to hold one upstream `EventSource` per URL. All same-origin tabs share this worker diff --git a/hive-c0re/src/agent_ports.rs b/hive-c0re/src/agent_ports.rs index c856a128..88f1486b 100644 --- a/hive-c0re/src/agent_ports.rs +++ b/hive-c0re/src/agent_ports.rs @@ -1,7 +1,33 @@ -//! `/var/lib/hyperhive/agent-ports.json` writer — flat map of -//! agent name → TCP web port. Written alongside `agents.conf` on -//! every topology change. JSON shape, port derivation (FNV-1a hash), -//! atomicity, and manager exclusion: `docs/gateway.md::Agent port map`. +//! `/var/lib/hyperhive/agent-ports.json` writer. Port map for +//! per-agent `/agent//` TCP routing. Written alongside +//! `agents.conf` (see `gateway_nginx.rs`) on every topology change; +//! `gateway_nginx::render` reads it indirectly via +//! `lifecycle::agent_web_port` to populate TCP upstreams for agents +//! that haven't opted in to unix-socket mode yet. Also kept as a +//! human-readable audit file — `cat agent-ports.json` shows every +//! registered sub-agent and its deterministic port assignment. +//! +//! Shape (flat object keyed by logical agent name → web port): +//! +//! ```json +//! { +//! "iris": 8178, +//! "atlas": 8304, +//! "argus": 8267, +//! "damocles": 8549 +//! } +//! ``` +//! +//! Ports come from [`crate::lifecycle::agent_web_port`] — pure +//! FNV-1a(name) hash so the value is reproducible from the name +//! alone. Manager is excluded from the map: the gateway routes `/` +//! straight to it via the c0re dashboard upstream rather than a +//! per-agent `/agent//` prefix. +//! +//! Atomicity: write to a sibling `.tmp` file + rename so a partial +//! write never leaves an unparseable file in place. The gateway's +//! `nginx` worker can read mid-write and Just Work because `rename()` +//! is atomic on the same filesystem. use std::collections::BTreeMap; use std::path::PathBuf; diff --git a/hive-c0re/src/build_logs.rs b/hive-c0re/src/build_logs.rs index f629cd93..e54d7f01 100644 --- a/hive-c0re/src/build_logs.rs +++ b/hive-c0re/src/build_logs.rs @@ -1,7 +1,23 @@ -//! Sqlite-backed full build-log capture — stdout + stderr per -//! `nixos-container` / `nix build` invocation, accumulated live. -//! Schema, indices, retention, and the rationale for replacing -//! the old ring buffer: `docs/persistence.md::/var/lib/hyperhive/build_logs.sqlite`. +//! Sqlite-backed full build-log capture. One row per `nixos-container` +//! / `nix build` invocation that the host-side lifecycle layer fires; +//! the row accumulates stdout + stderr line-by-line as the child runs. +//! +//! Replaces the legacy 32-line stderr ring buffer in +//! `lifecycle::run` / `lifecycle::prebuild_toplevel`. The ring tail +//! routinely truncated the actual eval error (a "tried alternatives" +//! block alone is often 30+ lines), so failures bailed with an +//! arbitrary tail and the full stream only lived in the host journal. +//! With this table the dashboard can surface the entire log. +//! +//! Storage lives next to the broker / approvals dbs (one file at +//! `/build_logs.sqlite`). Two indices: +//! `(agent, started_at)` for the per-agent latest-N lookup that backs +//! the agent card chip; `(status, finished_at)` for the retention +//! sweep that runs as part of the existing hourly vacuum. +//! +//! Writes are best-effort: every `append_*` / `finish` call logs a +//! warning on sqlite error and lets the build continue. A failed log +//! row never breaks a rebuild. use std::path::Path; use std::sync::{Arc, Mutex, OnceLock}; diff --git a/hive-c0re/src/crash_watch.rs b/hive-c0re/src/crash_watch.rs index ae3ce18e..85feb1e6 100644 --- a/hive-c0re/src/crash_watch.rs +++ b/hive-c0re/src/crash_watch.rs @@ -1,7 +1,19 @@ -//! Per-container crash and login-state watcher. Polls every managed -//! container on a 10s interval. Fires `ContainerCrash`, `LoggedIn`, -//! and `NeedsLogin` helper events. Event semantics and the -//! `RECENT_TRANSIENT_GRACE` window: `docs/approvals.md::Helper events`. +//! Per-container state watcher. Polls every managed container on a +//! fixed interval, tracks two orthogonal state-sets across ticks, +//! and emits a `HelperEvent` to the manager on each transition: +//! +//! - **running**: container is up. running → stopped without an +//! operator-initiated transient (`Stopping` / `Restarting` / +//! `Destroying` / `Rebuilding`) → `ContainerCrash`. +//! - **logged-in**: claude session dir is populated. ! → ✓ → +//! `LoggedIn`; ✓ → ! → `NeedsLogin` (rare — usually only fires +//! on a fresh spawn / purge). +//! +//! `NeedsUpdate` events are now fired from the apply-commit path +//! directly rather than via rev-marker polling. +//! +//! D-Bus subscription would be lower-latency for the first axis, +//! but polling is simpler and a 10s detection delay is fine. use std::collections::HashSet; use std::sync::Arc; diff --git a/hive-c0re/src/dashboard.rs b/hive-c0re/src/dashboard.rs index de780ea1..15794ef6 100644 --- a/hive-c0re/src/dashboard.rs +++ b/hive-c0re/src/dashboard.rs @@ -94,8 +94,13 @@ pub async fn serve(port: u16, coord: Arc) -> Result<()> { // /static/dashboard.css → dist/static/dashboard.css, etc.). .fallback_service(ServeDir::new(&static_dir)) .with_state(AppState { coord }); - // Binds loopback-only; external access via gateway. - // Rationale: docs/gateway.md::Firewall posture. + // Bind loopback-only. External access funnels through + // hive-gateway (in-host-netns nginx container), which proxies + // `/` → `127.0.0.1:` upstream. Operators who opt + // out of the gateway lose remote dashboard access — that's by + // design; the c0re HTTP surface is privileged (approve / deny / + // destroy, etc.) and any external exposure needs to pass through + // a real reverse proxy with auth. let addr = SocketAddr::from(([127, 0, 0, 1], port)); let listener = bind_with_retry(addr).await?; tracing::info!(%addr, "dashboard listening"); @@ -103,10 +108,27 @@ pub async fn serve(port: u16, coord: Arc) -> Result<()> { Ok(()) } -// SPA shape + SSE channels: docs/web-ui/shape.md. +// --------------------------------------------------------------------------- +// The dashboard is an SPA. Its HTML shell + bundled JS / CSS / favicon +// live in the directory pointed at by `HIVE_STATIC_DIR` (set by the +// hive-c0re NixOS module to `${frontend}/dashboard`), served by the +// `tower_http::ServeDir` fallback declared in `serve()`. The dynamic +// surface — `/api/state` and the action endpoints — is owned here. +// The JS app fetches state on load, re-fetches after every async-form +// submit, and listens on `/dashboard/stream` for the unified live event +// channel. +// --------------------------------------------------------------------------- -/// `SO_REUSEADDR` bind with retry. Retry mechanics, attempt-cap -/// rationale, and log-level cadence: `docs/web-ui/shape.md::Listener bind`. +/// `SO_REUSEADDR` bind with retry. Mirrors the per-agent variant in +/// `hive-ag3nt::web_ui::bind_with_retry`: hive-c0re restarts also +/// race the previous process's socket release, and the retry has no +/// attempt cap — capping was the proximate cause of a silent +/// give-up on a long stale socket. Genuine port collisions +/// don't reach this layer (dashboard is bound to a fixed configured +/// port, no per-agent hashing), so any persistent `AddrInUse` always +/// reflects a recoverable stale socket. WARN for the first dozen +/// attempts; INFO after that to avoid spamming the journal during a +/// long hold; INFO on eventual success when we did have to retry. async fn bind_with_retry(addr: SocketAddr) -> Result { let mut delay_ms = 250u64; let mut attempts = 0u32; @@ -1189,10 +1211,54 @@ struct StateFileQuery { path: String, } -/// Resolve a caller-supplied path against the allow-listed roots -/// (`agents//state/` and `shared/`). Applies defense-in-depth -/// symlink + traversal checks before serving. Security model and -/// all five layers: `docs/security.md::State-file endpoint`. +/// Bounded-size read of a file under one of two allow-listed +/// roots: `/var/lib/hyperhive/agents//state/` (per-agent durable +/// notes — the only writable path agents have outside their +/// container) and `/var/lib/hyperhive/shared/` (shared docs). Both +/// path forms are accepted: +/// - canonical host: `/var/lib/hyperhive/agents/alice/state/foo.md` +/// - container view: `/agents/alice/state/foo.md` +/// - shared: `/shared/foo.md` +/// +/// `/state/...` on its own is *not* accepted — the in-container +/// mount is ambiguous from the host's perspective (we don't know +/// which agent's `/state` it refers to) and using it would silently +/// resolve to the wrong file. +/// +/// Path is canonicalised before the allow-list check so `..` +/// traversal and symlink games can't escape the roots. Files larger +/// than `MAX_BYTES` are truncated with a banner so a runaway log +/// can't OOM the browser. +/// Resolve a caller-supplied path string to a canonical host path +/// that has been verified against the allow-list. Returns `Err` +/// with a human-readable reason for every failure mode (path +/// outside roots, canonicalize failure, escape via symlink, +/// per-agent subdir not `state`, symlink anywhere below the root, +/// file not world-readable). Shared by `get_state_file` (read) and +/// `scan_validated_paths` (linkify candidates in message bodies) +/// so both apply identical security rules and the linkifier +/// doesn't render a path the reader will refuse to serve. +/// +/// Defense-in-depth layers (in order): +/// 1. Caller-supplied prefix has to match the allow-list (agents/ +/// or shared/), else reject without touching the fs. +/// 2. No symlinks below the matched root. Walked pre-canonicalize +/// via `symlink_metadata` on each component so a sub-agent that +/// plants `ln -s /var/lib/hyperhive/agents/other/state/secret +/// /agents/me/state/peek` can't proxy a different agent's file +/// through this endpoint (canonicalize would happily resolve +/// the symlink to a path inside the allow-list). +/// 3. Canonicalize is run anyway as a belt-and-braces check — +/// resolves `..`/`.` traversal and rejects if the result +/// escapes the roots. +/// 4. Under `AGENTS_ROOT`, the second path component must be +/// `state/` — agents' applied/proposed git repos and config dirs +/// are off-limits. +/// 5. The target's metadata is fetched once and returned to the +/// caller so they don't restat. If the target is a regular +/// file it must be world-readable (mode & 0o004); a 0600 file +/// inside `state/` could leak through this endpoint to anyone +/// holding the dashboard URL otherwise. fn resolve_state_path( raw: &str, ) -> std::result::Result<(std::path::PathBuf, std::fs::Metadata), String> { @@ -1505,10 +1571,21 @@ pub(crate) fn emit_meta_inputs_snapshot(coord: &Coordinator) { }); } -/// Scan `body` for path-shaped tokens and return those that pass the -/// allow-list + `is_file` check via `resolve_state_path`. Called at -/// broker-message ingest so the dashboard event already carries the -/// verified set; security rules stay in sync with the read endpoint. +/// Scan `body` for path-shaped tokens, validate each against the +/// allow-list, return the unique set of tokens that resolve to a +/// regular file. Called at broker-message ingest time so the +/// dashboard event already carries the verified set — no client- +/// side probe endpoint required, and historical messages get the +/// same treatment on `/dashboard/history` backfill. +/// +/// Tokenisation: split on whitespace + a handful of trailing +/// punctuation chars (`,;:)]}`) that commonly follow paths in +/// natural-language text but aren't part of the path itself. Any +/// token starting with `/agents/`, `/shared/`, or +/// `/var/lib/hyperhive/{agents,shared}/` is a candidate. The +/// allow-list + `is_file` check happens via the same +/// `resolve_state_path` helper the read endpoint uses, so the +/// security rules can't drift. pub fn scan_validated_paths(body: &str) -> Vec { const PREFIXES: [&str; 4] = [ "/agents/", diff --git a/hive-c0re/src/dashboard_events.rs b/hive-c0re/src/dashboard_events.rs index fd05b19d..83a3fcac 100644 --- a/hive-c0re/src/dashboard_events.rs +++ b/hive-c0re/src/dashboard_events.rs @@ -1,8 +1,27 @@ -//! Unified dashboard event channel — all near-real-time browser events -//! flow through `Coordinator.dashboard_events`. Each event carries a -//! monotonic `seq` for client-side dedupe against `/api/state` snapshots. -//! Design rationale (single channel, broker forwarder isolation): -//! `docs/web-ui/shape.md::One unified channel`. +//! Unified dashboard event channel. +//! +//! Anything the browser wants to react to in near-real-time flows through +//! `Coordinator.dashboard_events`. Each event is stamped with a monotonic +//! per-process `seq` so the client can dedupe its buffered live traffic +//! against snapshot/history responses (drop frames with +//! `seq <= snapshot.seq`). +//! +//! Why one channel instead of one-per-domain: browsers cap concurrent +//! SSE connections per origin (~6 in chrome) and dispatch-by-kind on the +//! client is a one-liner. Splits get reserved for high-volume sub-streams +//! that most consumers don't care about (none yet). +//! +//! Message-broker traffic (`Sent` / `Delivered`) lives on this channel +//! too. A background forwarder task in `main.rs` subscribes to the broker +//! and re-emits each `MessageEvent` as a `DashboardEvent::Sent` / +//! `DashboardEvent::Delivered` with a freshly-stamped seq. Keeping the +//! broker's intra-process channel separate avoids coupling the broker +//! (used by `recv_blocking_batch` inside the harness loop) to dashboard +//! presentation concerns. +//! +//! New mutation kinds (approval added/resolved, question added/answered, +//! transient changed, etc.) land here as additional variants. The client +//! dispatches by `kind` and updates the relevant section. use serde::Serialize; diff --git a/hive-c0re/src/forge.rs b/hive-c0re/src/forge.rs index fd8c69ad..01cf713d 100644 --- a/hive-c0re/src/forge.rs +++ b/hive-c0re/src/forge.rs @@ -1,6 +1,21 @@ -//! Optional Forgejo wiring — per-agent user + token provisioning, -//! config-repo mirroring, meta read-access grants. No-op when -//! `hive-forge` isn't running. Full design: `docs/forge.md`. +//! Optional Forgejo wiring. When the `hive-forge` nixos-container is +//! present and running, hive-c0re ensures every agent (and the +//! manager) has a corresponding forgejo user with an API token +//! written to `/forge-token` — visible inside the +//! container as `/state/forge-token`. Idempotent: skips creation +//! when the user already exists, skips token issuance when the file +//! is already there. +//! +//! It also mirrors each agent's hive-c0re-owned *applied* config repo +//! into the private `agent-configs` org (`push_config`), so every +//! deploy / approval tag core plants is visible on the forge. Each +//! agent is a read-only collaborator on `core/meta` (the meta flake) +//! so they can fetch their deployment context; the `agent-configs` +//! repos remain core-only. +//! +//! No-op when `hive-forge` isn't enabled (detected via +//! `nixos-container list`), so operators who don't run the bundled +//! forge pay nothing. use std::path::{Path, PathBuf}; @@ -41,15 +56,32 @@ const CONFIG_ORG: &str = "agent-configs"; /// Forgejo orgs hive-c0re ensures on startup. The meta repo lives at /// `core/meta` (the `core` user's own namespace — no org needed). const SEEDED_ORGS: &[&str] = &[CONFIG_ORG]; -/// Per-agent token scopes (broad-but-not-admin). See -/// `docs/forge.md::Token scopes` for the per-scope rationale. +/// Forgejo scopes the agent's token gets. Broad-but-not-admin: every +/// repo / PR / issue thing an agent needs day-to-day, no admin +/// surface. +/// - `write:repository` — create, clone, push, delete repos in the +/// user's own namespace; merge PRs. +/// - `write:issue` — open / comment / review issues *and* pull +/// requests (forgejo namespaces PR conversation under issues). +/// - `write:user` — edit own profile, create repos under own user. +/// - `write:organization` — create + manage orgs (lets agents share +/// a forge namespace). +/// - `read:user` — token-owner endpoint clients call to introspect. +/// - `write:misc` — hooks, attachments, the rest of the long tail. +/// - `read:notification` — required by `forge_notify` to poll +/// `GET /notifications` for unread PR/review events. +/// - `write:notification` — required by `forge_notify` to mark +/// notifications as read via `PATCH /notifications/threads/{id}`. const TOKEN_SCOPES: &str = "read:user,write:user,read:notification,write:notification,write:repository,write:issue,write:organization,write:misc"; -/// Bootstrap `core` token scopes — adds `read:admin,write:admin` on -/// top of `TOKEN_SCOPES` so the host daemon can drive -/// `/api/v1/admin/*`. Site-admin membership alone isn't enough: the -/// token's own scope gate runs before the user-permission check. -/// See `docs/forge.md::Token scopes`. +/// Scopes for the bootstrap `core` token used by hive-c0re itself. +/// Adds `read:admin,write:admin` on top of `TOKEN_SCOPES` so the host +/// daemon can drive `/api/v1/admin/*` endpoints (PATCH user email on +/// agent provision, future webhook + org admin work). Site-admin +/// membership alone isn't enough — the token's own scope gate runs +/// before the user-permission check, so `403 Forbidden` comes back +/// for any `/admin/users/*` call from a non-admin-scoped token +/// even if the bearer is an admin user. const CORE_TOKEN_SCOPES: &str = "read:admin,write:admin,read:user,write:user,read:notification,write:notification,write:repository,write:issue,write:organization,write:misc"; /// Token file inside the agent's bind-mounted state dir (visible as @@ -239,8 +271,10 @@ async fn ensure_user_email(name: &str) { tracing::debug!(%name, %email, "forge: user email aligned"); } Ok(status) if status == reqwest::StatusCode::FORBIDDEN => { - // Core token missing admin scope — see - // `docs/forge.md::Token scopes` migration note. + // Almost certainly an existing-deployment migration case: + // an older core token may have been minted without admin + // scope, so /admin/users/* now returns 403 even though + // `core` is a site admin. Tell the operator how to fix it. tracing::warn!( %name, %email, %status, "forge: PATCH user email forbidden — core token likely missing admin scope. \ diff --git a/hive-c0re/src/gateway_nginx.rs b/hive-c0re/src/gateway_nginx.rs index 764545a8..1087119e 100644 --- a/hive-c0re/src/gateway_nginx.rs +++ b/hive-c0re/src/gateway_nginx.rs @@ -1,9 +1,30 @@ //! Runtime nginx include-file generator for the gateway's per-agent -//! `/agent//` location blocks. Writes -//! `/var/lib/hyperhive/gateway/agents.conf` on every topology change. -//! UDS vs TCP upstream selection, reload trigger (`systemd-run -//! --machine=hive-gateway`), and idempotency: -//! `docs/gateway.md::Per-agent unix-socket upstream`. +//! `/agent//` location blocks. +//! +//! Writes `/var/lib/hyperhive/gateway/agents.conf` on every topology change. +//! The gateway container bind-mounts `/var/lib/hyperhive/gateway/` (NOT the +//! whole parent dir) at `/run/hive-state/`; nginx includes +//! `/run/hive-state/agents.conf`. After each write, c0re triggers +//! `nginx -s reload` inside the gateway container via +//! `systemd-run --machine=hive-gateway` from the host — no `nixos-rebuild +//! switch` needed when agents start, stop, or flip `useUnixSocket`. +//! (A path unit inside the container was tried first but `IN_MOVED_TO` +//! from the atomic rename does not cross the nspawn mount-namespace +//! boundary — see `docs/gateway.md` for the failure analysis.) +//! +//! Upstream selection mirrors `agent_sockets::build_map`: an agent +//! gets a UDS upstream when its `.bound` marker exists (harness has +//! bound the unix socket); otherwise falls back to the deterministic +//! TCP port from `lifecycle::agent_web_port`. Proxy headers are +//! emitted in full so the generated file is self-contained nginx +//! config — no dependency on which `recommendedProxySettings` knobs +//! the host config has on. +//! +//! `write()` is idempotent: if the rendered body equals what's already +//! on disk, the rename is skipped and the path unit doesn't fire. +//! Same atomic `.tmp` + `rename()` shape as `agent_ports` / +//! `agent_sockets` — a crashing c0re process never leaves a partial +//! file the gateway's nginx would fail to parse. use anyhow::{Context, Result}; use std::fmt::Write as _; diff --git a/hive-c0re/src/lifecycle.rs b/hive-c0re/src/lifecycle.rs index ea8f3eb7..9dc9d080 100644 --- a/hive-c0re/src/lifecycle.rs +++ b/hive-c0re/src/lifecycle.rs @@ -408,13 +408,38 @@ pub async fn rebuild_no_meta( let container = container_name(name); let flake_ref = format!("{}#{name}", crate::meta::meta_dir().display()); if container_exists(name).await { - // Rebuild strategy: stop-before-update + pre-build. - // See `docs/coordinator.md::Container lifecycle`. + // Existing container: preserve the prior running state across + // rebuild, and apply both the new system profile + // AND any `/etc/nixos-containers/.conf` / drop-in changes + // in a single start rather than `update`'s reload-then-outer- + // restart double-bounce. + // + // `nixos-container update` only runs `systemctl reload + // container@` when the container is up (per the + // `isContainerRunning` check in nixos-container.pl), so + // stopping first makes `update` boot-style: build + nix-env + // --set the new profile, skip the in-container + // switch-to-configuration, let the next `start` apply both + // the new profile and the new EXTRA_NSPAWN_FLAGS in one go. + // If the container was already stopped, `update` builds + sets + // the profile and we leave it stopped. let was_running = is_running(name).await; set_nspawn_flags(&container, agent_dir, claude_dir, notes_dir)?; set_resource_limits(&container)?; systemd_daemon_reload().await?; if was_running { + // Pre-build the system toplevel **before** stopping the + // running container so the agent keeps serving its + // previous generation while the eval + fetch + build + // happens out-of-band. `nixos-container update` then + // finds the toplevel cached and skips straight to the + // profile-swap + restart — downtime collapses to that + // window only. Build failures surface here, before we + // touch the container. + // + // When the container is already stopped there's no + // downtime to shave — let `update` do the build inline + // rather than evaluating the flake twice for nothing. on_step("nix build"); prebuild_toplevel(name, &flake_ref).await?; on_step("nixos-container stop"); @@ -423,8 +448,18 @@ pub async fn rebuild_no_meta( on_step("nixos-container update"); run(&["update", &container, "--flake", &flake_ref]).await?; if was_running { - // Cold-start fallback on activation errors. - // See `docs/coordinator.md::Cold-start fallback`. + // Normal path: start into the new generation. The activation + // script runs inside the container to transition old → new. + // This can fail when packages are removed between generations — + // the old-generation activation references units that no longer + // exist in the new closure, causing systemd to exit non-zero. + // + // Fallback: stop + kill + start (cold-start). The activation + // script can fail when packages are removed between generations — + // `start` exits non-zero but the container may be half-started. + // `stop` requests a graceful SIGTERM drain; `kill` then SIGKILLs + // any lingering processes so the next `start` enters a clean state + // without a generation transition, letting the activation succeed. on_step("nixos-container start"); if let Err(start_err) = run(&["start", &container]).await { tracing::warn!( @@ -459,8 +494,11 @@ pub async fn rebuild_no_meta( Ok(()) } } else { - // Spawn path: create is atomic, no prebuild needed. - // See `docs/coordinator.md::Spawn path`. + // First spawn: no running container, no downtime to shave. + // `nixos-container create` builds + creates atomically — if + // the build fails, no container record is left around to + // clean up — so a pre-build adds nothing but a duplicate + // eval. on_step("nixos-container create"); run(&["create", &container, "--flake", &flake_ref]).await?; set_nspawn_flags(&container, agent_dir, claude_dir, notes_dir)?; @@ -471,12 +509,29 @@ pub async fn rebuild_no_meta( } } -/// Pre-build `system.build.toplevel` against `meta#` so the -/// subsequent `nixos-container update` finds the result cached and -/// skips straight to the profile-swap. Store-warming only — container -/// is untouched. See `docs/coordinator.md::Rebuild path` for why -/// the prebuild happens before stop, and `docs/coordinator.md::Prebuild -/// attr path` for why the explicit nixosConfigurations attr is required. +/// Pre-build the agent's `system.build.toplevel` derivation against +/// `meta#` so the subsequent `nixos-container update` / +/// `create` finds the result already in the store. The container +/// itself is untouched — this is purely a store-warming pass. +/// +/// Streams nix's stdout to INFO and stderr to WARN like the +/// `nixos-container` shellouts so progress shows up in journald as +/// it happens. `--no-link` keeps us from littering the working +/// directory with `result` symlinks. Per-derivation cost: pure +/// cache hit when nothing changed (handful of seconds for the +/// eval), expensive only on the rebuild that actually has work. +/// +/// Attr path is `#nixosConfigurations..config. +/// system.build.toplevel` — `nix build` won't auto-resolve the bare +/// `` against `nixosConfigurations` like `nixos-container` does +/// internally, so we have to spell the path out explicitly. Falling +/// back to `meta#` (the shape `nixos-container update --flake +/// meta#` uses) makes nix look for `packages..`, +/// `legacyPackages..`, or `` at the flake root — +/// none of which exist in the rendered meta flake. +/// +/// Returns the same error shape as the other nixos-container +/// helpers so callers can use `?` without translation. async fn prebuild_toplevel(name: &str, flake_ref: &str) -> Result<()> { use tokio::io::{AsyncBufReadExt, BufReader}; // Split `#` so we can re-emit with the explicit diff --git a/hive-c0re/src/meta.rs b/hive-c0re/src/meta.rs index 9b602407..a9d013b7 100644 --- a/hive-c0re/src/meta.rs +++ b/hive-c0re/src/meta.rs @@ -1,8 +1,20 @@ //! Single hive-c0re-owned flake at `/var/lib/hyperhive/meta/` that -//! exports one `nixosConfiguration` per agent and drives the system-wide -//! deploy audit trail. Flow (`sync_agents`, two-phase `prepare_deploy` / -//! `finalize_deploy` / `abort_deploy`, `lock_update_hyperhive`): -//! `docs/approvals.md::Meta flake`. +//! consumes every agent's applied repo as a flake input and exports one +//! `nixosConfiguration` per agent. Containers run against +//! `--flake /var/lib/hyperhive/meta#`; lifecycle ops here drive the +//! lock file so meta's git log is the system-wide deploy audit trail. +//! +//! Flow: +//! - `sync_agents` (idempotent) — render `flake.nix` for the current +//! agent set, init the repo on first call, relock if the rendered +//! contents changed, commit. Used by spawn / destroy / startup +//! migration. +//! - `prepare_deploy` + `finalize_deploy` / `abort_deploy` — two-phase +//! for the `request_apply_commit` path so a failed +//! `nixos-container update` leaves no orphan commit in meta. Prepare +//! writes the new lock without committing; finalize commits with the +//! deploy message; abort `git restore`s the lock back. +//! - `lock_update_hyperhive` — one-shot for the auto-update path. use std::path::{Path, PathBuf}; diff --git a/hive-c0re/src/migrate.rs b/hive-c0re/src/migrate.rs index 7bebc8cb..5e1ae421 100644 --- a/hive-c0re/src/migrate.rs +++ b/hive-c0re/src/migrate.rs @@ -1,7 +1,23 @@ -//! Startup auto-migration from the pre-meta layout. Four idempotent -//! phases: applied repo, proposed repo, meta repo, container repoint. -//! Kill-switch: `HIVE_SKIP_META_MIGRATION=1`. Full migration sequence -//! and phase details: `docs/approvals.md::Migration from the pre-tag`. +//! Startup auto-migration from the pre-meta layout. Runs before +//! `auto_update::run` and consists of four phases, each idempotent: +//! +//! 1. Per-agent applied repo: rewrite `flake.nix` to the module-only +//! boilerplate if it isn't already, commit, relocate `deployed/0` +//! to HEAD so `setup_applied`'s existence check passes. +//! 2. Per-agent proposed repo: ensure the `applied` git remote +//! points at `/applied//.git` (re-runs `setup_proposed`'s +//! `ensure_applied_remote` indirectly via a host-side git call). +//! 3. Meta repo: `meta::sync_agents` over the current agent list — +//! init the repo on first call, rerender + relock if anything +//! drifted. +//! 4. Container repoint: for every existing container, run +//! `nixos-container update --flake meta#` so it +//! activates against the meta flake. Guarded by a marker file +//! so the (expensive) phase 4 only runs once across hive-c0re +//! restarts. +//! +//! Env kill-switch: `HIVE_SKIP_META_MIGRATION=1` skips the whole +//! migration. Use when smoke-testing one agent at a time by hand. use std::path::{Path, PathBuf}; use std::sync::Arc; diff --git a/hive-c0re/src/reminder_scheduler.rs b/hive-c0re/src/reminder_scheduler.rs index fd731dfd..de370e4f 100644 --- a/hive-c0re/src/reminder_scheduler.rs +++ b/hive-c0re/src/reminder_scheduler.rs @@ -1,7 +1,34 @@ -//! Background loop that drains due reminders from the broker and -//! delivers them as inbox messages. 5s poll cadence, shutdown-aware. -//! File-path semantics (path translation, traversal + symlink defense, -//! pointer delivery): `docs/approvals.md::Reminder delivery`. +//! Background loop that drains due reminders out of the broker and +//! delivers them as inbox messages. Mirrors the `events_vacuum` / +//! `crash_watch` shape — a single `spawn(coord)` entry point started +//! from `main.rs`. +//! +//! File-path semantics: a reminder may carry a `file_path` (the +//! agent-visible path inside its container). On delivery we: +//! +//! - Translate the container path (`/agents//state/foo.md`) to +//! the host path (`/var/lib/hyperhive/agents//state/foo.md`) +//! so hive-c0re can write to it from outside the container. +//! - Reject anything that isn't under the agent's own state subtree, +//! contains `..` (path traversal), or has an empty relative tail. +//! Falling outside the allowed prefix means the file write is +//! skipped and the original message is delivered inline (with a +//! noted warning) — the reminder still fires, just without the +//! payload split. +//! - Defend against symlink escape: after `create_dir_all`, the +//! parent dir is canonicalized and re-verified to live under the +//! agent's host state root. Then we open the final file with +//! `O_NOFOLLOW | O_CREAT | O_TRUNC` so an existing-symlink basename +//! can't redirect the write either. Without this an agent could +//! `ln -s /etc /agents/foo/state/escape` and bounce a write to an +//! arbitrary host path. +//! - Write the reminder body to disk and deliver a short pointer +//! message in its place, so the agent's inbox/wake-prompt stays +//! small and the bulky payload can be read out of band. +//! +//! Atomicity of the inbox INSERT + `reminders.sent_at` UPDATE is handled +//! inside `Broker::deliver_reminders_batch`; this module only computes the +//! body strings before calling it. use std::io::Write; use std::os::unix::fs::OpenOptionsExt; diff --git a/hive-c0re/src/scheduled_prompts_worker.rs b/hive-c0re/src/scheduled_prompts_worker.rs index 611a57d4..b241f164 100644 --- a/hive-c0re/src/scheduled_prompts_worker.rs +++ b/hive-c0re/src/scheduled_prompts_worker.rs @@ -1,7 +1,39 @@ -//! Background loop that drains due `scheduled_prompts` rows and fans -//! the body to each active target. 5s poll cadence, shutdown-aware. -//! Catch-up clamp, missing-target handling, and broker-error retry -//! semantics: `docs/approvals.md::Scheduled prompt worker`. +//! Background loop that drains due `scheduled_prompts` rows +//! and fans the body out as inbox `Message`s to each active +//! target. Mirrors `reminder_scheduler::spawn` shape: +//! single `spawn(coord)` entry, 5s poll cadence, shutdown-aware. +//! +//! ## Catch-up semantics +//! +//! When hive-c0re comes back from being down, a recurring row +//! whose `next_fire_at` is well in the past would otherwise fire +//! N delayed pulses in a row. Instead we fire ONCE and let +//! `ScheduledPrompts::rearm` bump `next_fire_at` to the next +//! interval slot ≥ `now`, recording the skipped-cycle count in +//! the per-target `last_result` so operators see how many +//! firings were caught up rather than losing the signal. +//! +//! ## Missing-target failure +//! +//! When a target name doesn't resolve to a known agent (the +//! container has been destroyed, the operator typo'd a name, +//! etc.) the worker: +//! 1. records `last_result = "no such agent: "` against +//! the per-target row, +//! 2. sends a single advisory `Message` from `system` to +//! `operator` describing the schedule + target + reason, +//! 3. continues fanning out to the other (live) targets. +//! +//! Transient broker errors (sqlite lock contention, etc.) get +//! the per-target `last_result` annotated AND a `tracing::warn`, +//! but the post-fire bookkeeping treats the row the same way it +//! does on a clean fire: +//! - **recurring** rows re-arm — the next interval slot tries +//! the broker send again, so transient errors self-heal. +//! - **one-shots** delete unconditionally after their single +//! fan-out pass; a broker failure on a one-shot is NOT +//! retried (the operator advisory + `last_result` are the only +//! audit trail). use std::sync::Arc; use std::time::Duration;