hive-forge-notify grows a second binary, hive-github-notify. The two share the notification half of the job — tolerant parse, classification, formatting, dedupe, todo delivery — and nothing else: each binary owns its host's protocol outright. Two binaries rather than one multi-source daemon, and rather than a cargo feature. A feature would unify across the workspace and cost every crate its build cache. Two binaries keep the decision in nix: forge.nix installs the forge unit, github.nix installs the github one under hyperhive.github.enable, so a hive built without that module has no github poller in its closure at all — GitHub access is separable (a tier, a policy boundary), not merely switched off. Both binaries ship from the existing derivation, so packages.nix is untouched. The split is real at the code level too, not just at the unit level. source.rs is a trait; the impls live in the binaries that use them, so neither binary links the other's protocol code and the library names no host at all. The forge-only assigned-issue rollup moves into the forge binary for the same reason: it asks the forge what is assigned to this agent, which is not a notification-protocol concern. At runtime the github unit needs a PAT at <state>/github-token, the same dashboard-provisioned token the gh wrapper and the git credential helper already use. No PAT: it logs why and exits 0, which is why the unit is Restart=on-failure and not always. Forgejo's notifications API is modelled on GitHub's, so one tolerant parse serves both — the differences (string thread ids, PullRequest vs Pull) are absorbed by lenient deserializers rather than a second parse path. Thread ids normalise to String at the parse boundary; they are only ever opaque keys. Todo keys gain a per-source prefix so the two hosts cannot collide, and the forge's is deliberately empty to keep existing forge todo keys stable across the deploy that lands this. The github loop honours the server's X-Poll-Interval, re-arming only when the server asks for a slower cadence than ours; the hint is read before the status check, because it arrives on error and empty pages too and that is exactly when it matters. Reading the notification stream needs the notifications scope on the PAT, which a token minted for push access typically lacks; the failure mode is silence, so docs/github.md says so explicitly.
324 lines
17 KiB
Markdown
324 lines
17 KiB
Markdown
# hive-forge
|
|
|
|
Private Forgejo instance running in a nixos-container, used as the
|
|
swarm's persistent code-collaboration surface (issues, PRs, reviews,
|
|
attachments). Configured via `services.hyperhive.forge.*`. Container
|
|
shape, ROOT_URL / sub-domain routing, and operator-vs-in-cluster URL
|
|
handling live in [`docs/gateway.md`](gateway.md); this file owns the
|
|
per-agent integration story and the notification pump that wakes
|
|
each agent on relevant activity.
|
|
|
|
## Token scopes
|
|
|
|
Two scope sets live in `hive-c0re::forge`:
|
|
|
|
**`TOKEN_SCOPES`** (per-agent tokens):
|
|
|
|
| Scope | Why |
|
|
| -------------------- | ------------------------------------------------------------------------------------------------------- |
|
|
| `write:repository` | Create, clone, push, delete repos; merge PRs. |
|
|
| `write:issue` | Open / comment / review issues **and** pull requests (Forgejo namespaces PR conversation under issues). |
|
|
| `write:user` | Edit own profile, create repos under own user. |
|
|
| `write:organization` | Create + manage orgs (lets agents share a forge namespace). |
|
|
| `read:user` | Token-owner endpoint used for self-identification at harness startup. |
|
|
| `write:misc` | Hooks, attachments, the rest of the long tail. |
|
|
| `read:notification` | Poll `GET /notifications` for unread events. |
|
|
| `write:notification` | Mark notifications read via `PATCH /notifications/threads/{id}`. |
|
|
|
|
**`CORE_TOKEN_SCOPES`** (hive-c0re's own `core` user): everything in
|
|
`TOKEN_SCOPES` plus `read:admin` and `write:admin`. Site-admin
|
|
membership alone isn't sufficient — Forgejo's token scope gate runs
|
|
before the user-permission check, so `/api/v1/admin/*` returns
|
|
`403 Forbidden` for any token without the admin scope bits, even when
|
|
the bearer is a site admin.
|
|
|
|
**Migration note**: if `PATCH /api/v1/admin/users/{name}` returns 403
|
|
on an existing deploy, the core token predates the admin-scope
|
|
addition. Delete `/var/lib/hyperhive/forge-core-token` and restart
|
|
hive-c0re to re-mint with the new scopes.
|
|
|
|
---
|
|
|
|
## Per-agent forge accounts
|
|
|
|
Each agent gets its own Forgejo user + access token, provisioned at
|
|
boot by `hive-c0re::forge`. The provisioning flow is idempotent:
|
|
existing accounts + tokens are reused, so container destroy/recreate
|
|
doesn't lose forge identity. The token is written to
|
|
`<state>/forge-token` (one line, no trailing newline) inside the
|
|
agent container so `hive-forge` CLI + `forge_notify` poller can
|
|
read it without touching c0re's host-side credential store.
|
|
|
|
Two things live in the `agent-configs` Forgejo organization:
|
|
|
|
- A config repo per agent (`agent-configs/<name>`). As of #1787 the
|
|
agent is a **write collaborator on its own** repo — it can push
|
|
config-change branches and (once #1838 P2 lands) open config PRs — but
|
|
`main` is branch-protected core-only: only hive-c0re's verify-and-ff-push
|
|
merge handler lands on `main`, an operator-team approval is required, and
|
|
the agent can neither push `main` directly nor self-merge. `main` is
|
|
fast-forward-only — hive-c0re never force-pushes (the merge handler's ff
|
|
push lands fine; the `push_config` mirror pushes `main` + the add-only
|
|
status tags without force, and treats a non-fast-forward rejection of
|
|
`main` after a rolled-back deploy as expected — the forge keeps the
|
|
approved history, the `failed/<id>` tag records the divergence).
|
|
Repos stay private, so an agent can't read another
|
|
agent's config. (Agents remain read-only collaborators on `core/meta`.)
|
|
hive-c0re also references this repo as the agent's **persistent meta
|
|
flake input** (`agent-<n>.url = git+http://<forge>/agent-configs/<n>.git`;
|
|
see [approvals.md § Meta flake](approvals.md)), fetching it as the `core`
|
|
user via a git credential helper that reads the live forge-core token —
|
|
so the config lives on the forge, not a hand-synced local checkout.
|
|
- The dashboard links each container's "config" anchor to this
|
|
config repo, so operators can click straight from the SW4RM tab into
|
|
the rendered repo without an extra `git` step.
|
|
|
|
The `hive-forge` CLI (separate workspace crate, see
|
|
[`README.md`](../README.md) file map) wraps the Forgejo REST API
|
|
with the per-agent token; agents call it for issue / PR / comment
|
|
ops as if it were a peer. All REST calls across the workspace
|
|
(`hive-forge` verbs, hive-c0re provisioning, this poller) go through
|
|
the typed `forgejo-api` crate; only non-`/api/v1` web-router routes
|
|
(attachment / artifact downloads, log streaming) and the poller's
|
|
enrichment fetches of server-provided subject URLs stay on raw
|
|
reqwest.
|
|
|
|
## Notification poller (`hive-forge-notify/src/notify.rs`)
|
|
|
|
Its own long-running per-agent daemon (`hive-forge-notify`, a sibling of
|
|
`hive-bash-daemon` / `hive-matrix-daemon`) — it used
|
|
to be a background task inside the `hive-agent` serve loop. Polls
|
|
`GET /api/v1/notifications?all=false` every 30 seconds (Forgejo's
|
|
unread-only filter), formats each notification as a broker
|
|
`Wake { from: "forge" }` message, and delivers it to the agent's own
|
|
inbox so claude's normal turn loop picks it up.
|
|
|
|
The crate builds a second, independent binary for a different host — see
|
|
[github.md](github.md#notifications).
|
|
|
|
The host-specific calls — list unread, mark read, resolve own login —
|
|
live behind `Source` in `hive-forge-notify/src/source.rs`;
|
|
classification, formatting, dedupe and todo delivery are shared. Todo
|
|
keys here are the bare thread ids, and must stay that way: renaming them
|
|
would orphan every in-flight forge todo on the first restart after a
|
|
deploy.
|
|
|
|
### Mark-read on delivery
|
|
|
|
On a **successful** broker delivery, `forge_notify` marks the thread
|
|
read on forge straight away (`PATCH /notifications/threads/{id}`). The
|
|
broker inbox is the durable work queue now — each delivered wake is a
|
|
sqlite row with its own ack lifecycle — so the forge unread flag no
|
|
longer needs to track whether the agent has *processed* a
|
|
notification. Clearing it on delivery keeps forge's unread set **tiny
|
|
by construction**: at rest it holds only threads that failed to
|
|
deliver plus whatever arrived since the last 30s poll.
|
|
|
|
That size property is the whole point. A container rebuild starts the
|
|
poller with no memory of what it delivered, re-scans `?all=false`, and
|
|
finds nothing stale — the delivered threads are already read on forge.
|
|
Forge's own read-state is thus the durable, cross-rebuild record of
|
|
what's been delivered; there is **no persisted cursor**. (This
|
|
replaced an earlier design that left threads unread and leaned on a
|
|
persisted dedup cursor: a rebuild that lost the cursor re-delivered the
|
|
entire still-unread backlog as fresh wakes — the notification flood of
|
|
#2593 / #2106.)
|
|
|
|
**Read-before-comment coupling, dropped on purpose.** The old design
|
|
left threads unread so the hive-forge read-before-comment guard (which
|
|
keys off forge unread-state) would force the agent to view a thread
|
|
before commenting. That coupling is gone: the broker wake already
|
|
carries the notification body, so *delivery is the read*. An agent that
|
|
wants the full thread still runs `hive-forge comments` / `view`; the
|
|
guard no longer blocks a first comment on a freshly-delivered thread.
|
|
|
|
**In-process dedupe (tiny, ephemeral).** A single-process map (thread
|
|
id → last-delivered `updated_at`) guards the narrow window where a
|
|
mark-read call *transiently fails* and the thread reappears unread in
|
|
the next poll before its `updated_at` bumps — so a flaky PATCH doesn't
|
|
re-fire the wake. It is **not persisted** and resets on restart (forge
|
|
read-state covers the durable case). Each poll prunes it to the ids in
|
|
the single `limit=UNREAD_FETCH_LIMIT` (50) fetch page, so it can never
|
|
exceed that many entries (a debug assertion pins the invariant; the
|
|
fetch limit and the bound are the same constant). A failed *delivery*
|
|
is left unread and out of the map, so it resurfaces next tick.
|
|
|
|
Self-echo notifications (the agent's own writes, see below) are marked
|
|
read directly without a delivery — same `mark_read` call, no wake.
|
|
|
|
### Activation gates (graceful no-ops)
|
|
|
|
The poller starts disabled and stays that way for any of:
|
|
|
|
- `HIVE_FORGE_URL` not set (no forge configured for this hive), or
|
|
not parseable as a URL.
|
|
- `<state>/forge-token` missing or empty (agent has no forge
|
|
account — pre-provisioning or destroy-without-purge race).
|
|
- Initial client construction fails (the typed `forgejo-api` client
|
|
for the API calls, or the plain reqwest client kept for the
|
|
best-effort enrichment fetches of server-provided subject URLs;
|
|
both extremely unlikely; treated as fatal-to-the-task only).
|
|
|
|
Disabled = the spawned task returns immediately. All other failure
|
|
modes (HTTP errors, parse errors, mark-read failures) are
|
|
best-effort: logged at debug/warn and retried next tick.
|
|
|
|
### Self-notification filtering
|
|
|
|
Forgejo fires notifications for the agent's own actions (it opened a
|
|
PR, posted a comment, submitted a review). Surfacing those would
|
|
loop claude on its own writes. The comment/review case is dropped
|
|
silently (mark-read without delivery):
|
|
|
|
- **Self-authored comments / reviews** — comment payload's
|
|
`user.login` matches `own_login`.
|
|
- **Self-authored creations** (an agent opening its own PR/issue) — the
|
|
already-fetched subject payload's poster `user.login` matches
|
|
`own_login`. Only _creations_ are dropped; a later state change on the
|
|
agent's own subject is driven by someone else and still surfaces.
|
|
|
|
`own_login` is fetched at startup via `GET /api/v1/user`. On fetch
|
|
failure the filter degrades open (no filtering) rather than crashing
|
|
the task — a noisy inbox beats a silently-stuck poller — but the fetch
|
|
is **re-attempted on each poll tick** until it succeeds, so a boot-time
|
|
failure (the forge not yet reachable) self-heals instead of leaving
|
|
self-echo filtering off for the whole process lifetime.
|
|
|
|
### Body excerpt + truncation + heading escape
|
|
|
|
The wake message embeds the comment / review / new-item body so the
|
|
agent sees actual content without a follow-up fetch. Three pipeline
|
|
steps in order:
|
|
|
|
1. **Truncate** to `BODY_TRUNCATE = 500` chars at a char-boundary;
|
|
appends `…` when cut. Truncation happens BEFORE escape so the
|
|
mention-overflow diff (next step) compares like-for-like against
|
|
the raw body.
|
|
2. **Mention overflow extraction** — when truncation actually
|
|
trimmed content, walk the full body line-by-line and surface any
|
|
`@username` lines that fell outside the embed window. Rendered as
|
|
a trailing `mentions (truncated from body):\n > <line>` block.
|
|
Mention detection requires the `@` to be at line start or
|
|
following a non-username byte, so email-style `foo@bar.com` does
|
|
NOT count.
|
|
3. **ATX heading escape** — for each line that's a strict
|
|
`CommonMark` ATX heading (1-6 leading `#`s followed by a space,
|
|
tab, or end-of-line), prepend `\` so the embedded body doesn't
|
|
blow into a top-level h1/h2 inside the wrapper message when the
|
|
dashboard renders it. Lines like `#tag`, `#123`, `#!/bin/bash`
|
|
are NOT headings — no escape, no cosmetic noise. Indented
|
|
"headings" inside lists / nested quotes keep their leading
|
|
whitespace.
|
|
|
|
The strict ATX rule is deliberate: `\#tag` and `#tag` render
|
|
identically, so an over-eager escape just adds visual clutter
|
|
without changing behavior. Setext-style headings (`title\n====`)
|
|
are not handled — rarer in practice, would need multi-line
|
|
lookahead.
|
|
|
|
### Wrapper format
|
|
|
|
Five shapes, distinguished by the notification's classification:
|
|
|
|
| Trigger | Wrapper |
|
|
| ----------------------------------- | --------------------------------------------------------------------------------- |
|
|
| Comment on issue / PR | `[comment on PR #N owner/repo] title\nurl: ...\n\nauthor: body\nassignee: ...` |
|
|
| Review submission | `[PR approved #N owner/repo] title\nurl: ...\n\nauthor: body\nassignee: ...` |
|
|
| New issue / PR | `[new PR #N owner/repo] title\nurl: ...\n\n<body excerpt>\nassignee: ...` |
|
|
| Later activity (open, not creation) | `[activity on PR #N owner/repo] title\nurl: ...\n\n<body excerpt>\nassignee: ...` |
|
|
| State change | `[PR merged #N owner/repo] title\nurl: ...\nassignee: ...` |
|
|
|
|
Review labels come from the Forgejo `state` field: `APPROVED` →
|
|
`approved`, `REQUEST_CHANGES` → `changes requested`, `COMMENT` →
|
|
`review comment`. `PENDING` is dropped (review saved but not
|
|
submitted yet — no peer-visible event). Unknown states fall back to
|
|
the generic comment wrapper.
|
|
|
|
A review submitted with **no body** renders `reviewed by: <author>` in
|
|
place of the `<author>: <body>` line — deliberately worded to not collide
|
|
with the meta-suffix `reviewer:` line (requested reviewers, below).
|
|
|
|
### Merge/close vs a later comment
|
|
|
|
A notification carrying a `latest_comment_url` normally takes the comment
|
|
path. But a merged/closed subject **keeps** its `latest_comment_url` set,
|
|
so a just-merged PR that had any prior discussion would route to the
|
|
comment path and render `[comment on PR]` (with a stale pre-merge comment
|
|
body) instead of `[PR merged]` — the agent never learns its PR merged
|
|
(#2495). So when the notification IS the merge/close transition — its
|
|
event time (`updated_at`) is within `NEW_ITEM_TOLERANCE_SECS` of the
|
|
subject's `closed_at` (set for both `merged` and `closed`) — the
|
|
state-change path wins even with a comment url present
|
|
(`state_change_is_current`). A genuine **later** comment on an
|
|
already-closed subject bumps `updated_at` well past `closed_at`, so it
|
|
stays on the comment path and keeps its comment body. Missing/unparseable
|
|
timestamps default to the state-change path, so a merge is never silently
|
|
hidden behind a stale comment.
|
|
|
|
#### Merge racing a comment
|
|
|
|
The one gap the timestamp cut leaves: a genuine comment posted **within
|
|
`NEW_ITEM_TOLERANCE_SECS` of the merge** bumps `updated_at` close enough
|
|
to `closed_at` that `state_change_is_current` returns `true` — so it takes
|
|
the state-change path and its body would be dropped. Best of both worlds:
|
|
on the merge/close path we fetch the `latest_comment_url` comment and, when
|
|
its `created_at` is strictly **after** the subject's `closed_at`
|
|
(`comment_is_after_close`) — i.e. it raced the merge rather than being the
|
|
pre-merge last comment the subject keeps — append it as a
|
|
`comment by <author>: <excerpt>` block before the meta suffix
|
|
(`fresh_post_close_comment_tail`). So the wake carries **both** `[PR merged]`
|
|
and the racing comment. The kept pre-merge comment (created before
|
|
`closed_at`) is left off, a self-authored racing comment is dropped (don't
|
|
echo the agent's own write), and a missing/unparseable `created_at`/
|
|
`closed_at` appends nothing (conservative — only surface a comment we can
|
|
positively place after the close). Cost: one extra comment fetch on
|
|
merge/close notifications, acceptable given how rare they are.
|
|
|
|
### "new" vs "activity on"
|
|
|
|
A review submitted with **no body** carries no `latest_comment_url`,
|
|
so it misses the comment path and lands on the state-change path with
|
|
`state == "open"` — exactly like a freshly opened PR. Labeling that
|
|
`new PR` is misleading: agents dismiss it as a duplicate of the
|
|
original open notification and miss the review (#1637). So the `open`
|
|
state only earns the `new <kind>` label when the notification's event
|
|
time (`updated_at`) is within `NEW_ITEM_TOLERANCE_SECS` (120s) of the
|
|
subject's `created_at`. Anything later is labeled `activity on <kind>`
|
|
— neutral and non-misleading, since we can't cheaply say _what_ the
|
|
activity was without an extra reviews fetch. Missing/unparseable
|
|
timestamps default to `new` (preserve prior behavior rather than mask a
|
|
genuine new item). Timestamps are parsed by a small dependency-free
|
|
RFC 3339 → epoch-seconds helper (`parse_rfc3339_secs`).
|
|
|
|
Number is extracted from `subject.html_url`'s last path segment
|
|
(strips `#anchor` first); repo slug from `repository.full_name`.
|
|
Both degrade gracefully when absent (number → blank, repo → blank)
|
|
so unexpected Forgejo shapes don't crash the formatter.
|
|
|
|
### Meta suffix
|
|
|
|
Every wrapper ends with one or more of:
|
|
|
|
- `assignee: <list>` — always present; `unassigned` when empty so
|
|
the line shape is stable.
|
|
- `reviewer: <list>` — PR notifications only, present only when
|
|
`requested_reviewers` is non-empty.
|
|
|
|
### Review-request override
|
|
|
|
For new PRs, the kind label flips to `[review requested #N
|
|
owner/repo]` when `own_login` appears in `requested_reviewers`,
|
|
regardless of the Forgejo `reason` field. Forgejo doesn't reliably
|
|
set `reason == "review_requested"` (often null instead), so the
|
|
fallback checks the subject payload directly. Detection is gated on
|
|
`is_new` so the label only fires once on PR creation, not on every
|
|
subsequent comment.
|
|
|
|
### Subscription management
|
|
|
|
The poller does **not** auto-unsubscribe from repo watches — it
|
|
delivers every unread notification it's handed. Bounding the
|
|
firehose (dropping broad repo watches an agent doesn't need) is done
|
|
explicitly via a hive-forge CLI subscription verb, not by the poller
|
|
guessing which watches to drop. See the `subscription` verb in
|
|
[`docs/tools/forge.md`](tools/forge.md).
|