hyperhive/docs/ci.md
atlas 9d816431dc fix(#981): validate runner credentials on every boot, purge stale .runner
hive-ci-register.service now runs unconditionally on every boot (not
just when .runner is absent). Before fetching a registration token it
validates existing .runner credentials via the forge admin API:
- 200: runner still registered, write dummy token and exit
- 404: runner deleted from forge, purge .runner and re-register
- 000: forge unreachable, keep credentials (runner surfaces the error)
- other non-200 or malformed .runner: purge and re-register

Removes ConditionPathExists so stale credentials from a wiped forge
no longer block the runner indefinitely. Updates docs/ci.md to match.
2026-06-02 00:27:47 +02:00

3.2 KiB

hive-ci: Forgejo Actions Runner

The hive-ci module runs a Forgejo Actions runner in a hive-ci nixos-container, executing CI jobs from .forgejo/workflows/ci.yml (e.g., nix flake check on every PR).

Operator bootstrap

Set services.hyperhive.ci.enable = true in the host NixOS config. That's it — no manual token provisioning.

Requirements:

  • services.hyperhive.forge.enable = true must also be set (the runner registers against hive-forge).
  • Optional: tune services.hyperhive.ci.name (runner name in forge admin panel), concurrency (parallel job capacity), labels (workflow targeting).

Container design

  • Shared host netns: container reaches hive-forge at http://127.0.0.1:<httpPort> (same as hive-gateway).
  • Non-ephemeral: runner credentials persist across restarts (written to container's stateDir on first registration, reused thereafter).
  • Sandbox fallback: nspawn containers can't create user-namespaces, so nix's sandboxing would always fail. Module sets nix.settings.sandbox-fallback = true in the container — nix builds run unsandboxed (safe because the container is already isolated). See docs/gotchas.md.

Auto-registration flow

hive-ci-register.service is a oneshot that runs on every boot before gitea-runner-hive.service. It handles both first-run registration and stale-credential detection.

Every boot

  1. Read the core admin token from /run/hive-ci/core-token (bind-mounted from /var/lib/hyperhive/forge-core-token).
  2. If .runner exists: validate the runner ID against GET /api/v1/admin/runners/{id} using the core token:
    • 200: runner still registered — write dummy TOKEN=placeholder and exit; runner reuses .runner credentials.
    • 404: runner was deleted from forge (e.g. after a wipe) — delete .runner, proceed to re-registration below.
    • 000 (forge unreachable): keep existing .runner; the runner itself will surface the connectivity error.
    • other non-200: treat as stale, delete .runner and re-register.
    • malformed .runner (no id field): delete and re-register.
  3. If .runner is absent (first boot or purged above): fetch a fresh registration token from GET /api/v1/admin/runners/registration-token. Retries for 30s in case forge is still starting. Writes TOKEN=<real> to /run/hive-ci/runner-token.
  4. gitea-actions-runner reads the token, registers itself, and persists credentials to .runner. On subsequent boots step 2 validates these credentials and fast-paths past registration.

CI workflow

The single CI job is defined in .forgejo/workflows/ci.yml:

name: CI
on:
  pull_request:
    branches: ["**"]
jobs:
  check:
    name: nix flake check
    runs-on: [hive-ci]
    steps:
      - uses: actions/checkout@v3
      - name: check
        run: nix flake check

This runs on every PR, executing all flake checks (treefmt, rustfmt, cargo test, cargo clippy, module evaluation). No --no-build: the checks' derivations are the canonical source of truth.

References

  • nix/modules/hive-ci.nix: runner configuration, auto-registration script, container setup.
  • .forgejo/workflows/ci.yml: workflow definition.
  • docs/gotchas.md: nix sandboxing limitations in containers.