Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/nix/agent-modules
Author SHA1 Message Date
atlas
b0d92ccbd8 fix(forge): pass avatar image via files, not argv (E2BIG over 128 KiB)
forge-avatar-sync passed the base64-encoded icon as a jq --arg and then
as a curl -d command-line argument. Linux caps a single exec argument at
MAX_ARG_STRLEN (128 KiB), so any icon whose rasterized PNG base64-encodes
past that (red's does, at 133300 bytes) makes jq fail with E2BIG before
curl is ever reached, and the avatar upload silently never happens. The
base64 and the JSON payload now go through temp files instead of argv.

Closes #4839
2026-09-30 19:17:59 +02:00
iris
d74f567c59 nix: runtime option description: ACP cancel, compact and model/effort now work 2026-09-30 13:12:49 +02:00
atlas
8a8da5ec8a hive-agent: ACP model picker honours availableModels (#4391)
Filter the ACP model picker's list by services.hyperhive.agent.availableModels: an
unconfigured agent (env absent) shows every model the session offers, a
configured list narrows the picker to whatever it names that the session
also offers (in the session's own order), and a configured list matching
none of the session's models (the claude names on an ACP agent that never
touched the option) falls back to showing everything, with one warning
naming the mismatch.

The nix option's default, the model-vs-availableModels build assertion and
hive-subagent-mcp's check_model rail are unchanged — this only touches the
web UI's picker.
2026-09-30 10:53:53 +02:00
atlas
b3b42d3279 credential units: 24h retry shape; start a failed nginx when the cert lands
Six credential-fetch units retried 4 times at 15s, so an apply during
which the store or gateway was down for more than about a minute left
them in start-limit-hit, and nothing started them again once the store
came back. The swarm-services leaf could also land after nginx had
already given up on it, and the hook that propagates a new leaf only
reloaded a running nginx, so a stopped one stayed down until a second
apply.

- nix/host-modules/lib/store-retry.nix: the 2880 x 30s / 25h window
  shape swarm-services-cert already had, as one attrset.
- swarm-services-cert, swarm-bao-otel-oidc, swarm-bao-forwarder-oidc,
  swarm-bao-matrix-token, swarm-bao-queue-agent, swarm-bao-grafana-oidc,
  hive-agent-bao-identity and hive-agent-forge-token use it.
  queue-identity.nix no longer has a fetch unit (ccb5bd3b), and
  forge-token.nix is a fetch unit with the same short budget that was
  added after the census in #4662.
- The swarm-services-cert propagation hook now reset-fails and starts
  (--no-block) a loaded nginx that is not active; an active nginx keeps
  the re-import + reload.
- module-eval-bao-grants: one case pinning the shape on every host-side
  fetch unit, swarm-services-cert included.

Refs #4662
2026-09-30 07:45:47 +02:00
atlas
84d4808d54 hive-subagent-mcp: run an agent's subagents on its runtime
The subagent daemon now reads the parent agent's runtime at startup
(`hive_runtime::RuntimeSpec`, from the harness's `HIVE_RUNTIME` /
`HIVE_ACP_*`, which `mcp.nix` forwards onto its unit). On claude
nothing changes. On ACP, each run drives an `AcpRuntime` whose session
id is kept per name under the harness dir: `start` archives the old one,
`continue` loads it (and fails when none is recorded), `interrupt` sends
`session/cancel`, a role goes in front of the first prompt, and
permission requests get the answers a claude subagent's tool list
gives. The unit loads `backendEnvironmentFile` on ACP only, so the
agent can authenticate.

The end-of-turn handling moves out of the claude loop into `after_turn`
unchanged, so both loops share it.

Refs #4391
2026-09-30 07:41:12 +02:00
atlas
1b24edf4b4 nix: runtime option, acp.* and an opencode preset
services.hyperhive.agent.runtime ("claude" default | "acp") and
acp.{command,args,env}, rendered into HIVE_RUNTIME / HIVE_ACP_* only
for acp, so a claude agent's unit is unchanged. acp implies useApiKey.

acp.presets.opencode runs `opencode acp` from nixpkgs against an
OpenAI-compatible provider from acp.opencode.{provider,model,
contextWindow,outputLimit}: the config is rendered to the store with
the API key as an {env:VAR} reference, so the key is read at runtime
from backendEnvironmentFile. OPENCODE_PERMISSION denies opencode's
built-in bash, task, todowrite and websearch, and makes webfetch ask.

Refs #4391
2026-09-29 22:29:36 +02:00
atlas
ccb5bd3b38 hive-agent: read the per-agent queue secret from bao in process
The harness now reads swarm/agents/<agent>/queue from the store itself,
under the agent's own store certificate, and holds it in memory only.
It reads once before the first connect and again on every reconnect
attempt (async-nats `ConnectOptions::with_auth_callback`), so an agent
whose secret was re-minted reconnects with the new value instead of
being refused until the container restarts.

hive-agent-queue-credential.service, the /run file it wrote, and
HIVE_AGENT_QUEUE_AGENT_SECRET_FILE are gone; queue-identity.nix now
hands hive-agent.service the store address, its certificate paths and
the agent name.

A failed or empty read before the first connect still falls back to
the hive's shared client. Each read is bounded by a 10s timeout, and
retries wait out the existing reconnect backoff (500ms doubling, capped
at 60s).

Closes #4783
2026-09-29 10:18:07 +02:00
atlas
2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00
atlas
97cf8a1b2b agent-modules: write bao's stderr where UMask=0377 lets it, name what failed
hive-agent-forge-token and hive-agent-queue-credential both run with
UMask=0377. Their scripts captured bao's stderr in `err="$(mktemp)"`,
which under that umask is created 0400; the very next `2>"$err"` on the
`bao login` line cannot reopen it for writing, so bash fails the
redirect with "Permission denied" before bao ever runs. The `if !`
around the login then took the only error branch it had and printed
"this agent's certificate was refused by the swarm secret store" — the
store was never contacted. No agent has fetched either credential.

The stderr file now lives in each unit's own 0700 RuntimeDirectory and
is removed before every redirect into it, so the redirect creates it —
the idiom forge-token.nix already used for its staging file.

The login's error branch now says which of these happened, then quotes
bao's output:
- `$err` could not be created, so bao never ran;
- the store answered with HTTP 4xx (refusal) or another status;
- the store sent a TLS alert rejecting the certificate;
- no answer at all (network, DNS, or local TLS).
Unreadable cert/key credentials are reported before bao runs.

bao.nix has the same fetch shape but no UMask=, so its mktemp file is
0600 and writable; it is untouched.

Closes #4735
2026-09-26 21:50:39 +02:00
atlas
ed53e9abcc host-modules: write credential files atomically; agent-modules: retry a failed .claude migration
Four host glue units fetched a secret from swarm-bao and rendered it
with `> path; chmod`: a reader racing the write could see a truncated
file, and briefly one at the wrong mode before the chmod landed.
glue-matrix-bao-token.nix, glue-queue-agent-credential.nix (both
files), swarm-grafana.nix and swarm-otel.nix now write to a same-
directory temp file, set its final mode/owner, then `mv -f` it over
the target — a shared `atomic_write_secret` helper
(nix/host-modules/lib/atomic-write-secret.nix) so the five call sites
share one implementation.

The first-boot `/root/.claude` migration in nix/agent-modules/user.nix
wrote its done-marker unconditionally, so a failed `cp` (disk full,
permission error) left the marker behind and no boot ever retried the
copy. The marker is now written only when there was nothing to
migrate or the copy succeeded; `cp -an`'s no-clobber semantics already
make a retry after a partial copy safe.

Refs #4723
2026-09-26 21:50:03 +02:00
atlas
ab153bda2f hive-matrix-mcp: read the main account's token from the store too
The daemon reads each account's token from `swarm/agents/<agent>/matrix/`
as the agent itself, inside its own container, and falls back to the file
only when the store has none. This is #4519's read, without its `main`
carve-out: the swarm now mints `main` there and no hive writes the file.

The daemon unit gets the agent's store identity, spelled the way
forge-token.nix spells it. A timer re-starts it while it is down: a token
the swarm mints or replaces in the store changes no file, so the path
watcher never fires for it, and a daemon that exited on a replaced token
would otherwise stay down until the container restarts.
2026-09-25 08:31:01 +02:00
atlas
dd32a395f7 agents: pull the forge token from bao; drop tea-login
forge-token.nix fetches swarm/agents/<agent>/forge-token under the
agent's own store identity into /run/hive-agent-forge-token/token, and
re-fetches on a timer so a rotation lands. hive-forge, the git
credential helper, hive-forge-notify, forge-avatar-sync and the web UI
read that file first and fall back to <state>/forge-token.

tea-login is deleted: it copied the token into ~/.config/tea, which
docs/swarm/credentials.md forbids for a store secret. hive-forge covers
the same verbs. swarmctl gains agent mint-forge-token.

Refs #3782
2026-09-24 17:48:53 +02:00
atlas
1d261b3fed swarm-nats: give the queue a name, a bao-issued leaf, and require TLS
The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.

- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
  `swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
  every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
  IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
  may only `update` `pki/issue/swarm-nats`, written by
  `swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
  paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
  as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
  the container, restarts nats when it rotates, and re-runs daily.
  It joins glue-bao-readers-policy-order, so it is ordered after its policy
  unit (`after` and `wants`, never `requires`) where the store is on the
  same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
  `allow_non_tls`. `validateConfig` is now off in every mode, because the
  build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
  host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
  `singleHostSwarm`, and the auth responder all dial
  `tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
  to the NATS connection too, so hive-c0re and the controller trust the
  leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
  swarm needs.

module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.

Closes #4626
2026-09-24 17:26:31 +02:00
atlas
179f873722 docs: retire the agent hierarchy from every page that described it
The topology doc keeps its filename and its second half (manager
special-casing, harness unit shape) — both are cross-referenced from
other pages and neither is about the parent field. Its first half is
rewritten: what topology.json is now, and a table of what the removal
took with it, so a reader who finds `<parent>` or `set-parent` in an old
issue thread learns it went away rather than moved.

The dashboard's tree-rendering section is marked dormant rather than
deleted: the walk is still in swarm.js and retiring it is the frontend
owner's call.
2026-09-21 22:08:47 +02:00
atlas
afdfce67ec agent: fetch this agent's own swarm-queue credential from the store
Every agent on a hive authenticates to the swarm queue with the same
hive-scoped OIDC client, so at the auth callout one agent is
indistinguishable from its co-hived neighbours. The commit before this
one mints a secret per agent at swarm level into
secret/swarm/agents/<agent>/queue; nothing read it.

Read it here, and read it from the container itself. A hive courier in
the path would be the hive vouching for which agent this is, which is
the property a per-agent credential exists to remove -- so the agent
logs in to the store with the certificate hive-agent-bao-identity
already proves it can log in with, and reads its own path. The store
certificate is for reaching the store and nothing else: what the new
unit writes to /run is the secret it read back, and nothing hands a
BAO_CLIENT_* path to anything queue-shaped.

The read needs no policy change. render_agent grants read on
secret/data/swarm/agents/<agent>/*, which covers this path and the
bao-mtls one beside it alike -- which is also why this unit degrades
where the identity check fails. A refusal this unit sees and that check
did not cannot be a policy that drifted; it is an object not yet minted,
the ordinary state of every agent created before its swarm knew to mint
one.

The harness resolves the path and reports which credential this agent
can present. It does not yet present it: the auth-callout responder
still verifies only the hive-scoped token, and an agent offering a
credential nothing on the other end reads back would simply be refused.
Teaching swarm-nats-auth to read the same path is the next slice.
2026-09-21 20:44:52 +02:00
atlas
aaedff20a9 agent-modules/mcp: give subagent daemon a longer default bash timeout
A subagent's backgrounded bash child is reaped along with the rest of
its process tree at turn end, which silently orphans anything still
running past the CLI's default 2-minute timeout — the run reports
normally but the log file is empty or truncated. Raise the default on
hive-subagent-daemon's own unit rather than in managed-settings, since
managed-settings is also read by the main agent's session and the
operator's ruling is explicit that the main agent's environment stays
as is.
2026-09-21 17:20:59 +02:00
atlas
72a375b0db otel: map journald PRIORITY onto a severity at every journald receiver
Records reached VictoriaLogs carrying the journal's raw PRIORITY and
severity_text "Unspecified" — every line in the store, at every tier, with
no level a query or a dashboard could read. VictoriaLogs has no ingest
parameter naming a level field; it auto-detects one by field name, so the
mapping has to happen in the collector.

A stanza severity_parser on each journald receiver, from one shared file
rather than a copy per tier: the two receivers are unrelated config (a
fixed stanza in the agent container, a parameterised block inside the
swarm-otel container) and a drifted copy fails silently — every line still
arrives, labelled as the wrong thing.

Two details that are easy to get wrong and quiet when wrong. PRIORITY
counts down in urgency where the OTEL severity counts up, so the table is
written as a table. And overwrite_text is required: without it the parser
sets the severity number and leaves the text as the raw digit, so
severity_text arrives as the literal "6" — populated, and not a level
anything renders.
2026-09-20 14:23:56 +02:00
atlas
ed1fce3329 nix: gate the avatar-sync path unit on the same condition as its service
`systemd.paths.forge-avatar-sync` was gated on `agent.icon != null` alone,
while the `systemd.services.forge-avatar-sync` it triggers is gated on
`agent.icon != null && agent.forge.url != null`. An agent with an icon and no
forge URL therefore rendered a `.path` unit, pulled into multi-user.target,
watching for a forge-token whose arrival would activate a unit that does not
exist.

The module already documents the fixed behaviour: `forge.url`'s own option
description says the tea-login and avatar-sync units are "not generated at all"
when it is null --- an absent integration, never a misdirected one. That
sentence was true of the oneshot and false of its watcher.

Latent, not live: hive-c0re renders `forge.url` into every agent's config from
the host's `HIVE_FORGE_URL`, so on a real hive it is always set and the
asymmetric arm is unreachable. It is reachable wherever the agent modules are
evaluated outside a hive.

A module-eval case pins both halves absent for an agent with an icon and no
forge; it fails on the parent commit, where the path unit renders.
2026-09-19 13:26:17 +02:00
atlas
b6e180dfbf agent: make claudePlugins additive instead of replacing
The base set (skill-creator + base@hyperhive) was declared via the
option's `default`, so a per-agent definition of claudePlugins
replaced it wholesale. Move the base set to a plain `config`
definition instead: a listOf option merges multiple plain definitions
by concatenation, so an agent's own list now adds to the base set
rather than replacing it, while lib.mkForce / lib.mkOverride on the
agent side still replace the whole merged list deliberately (mkDefault
was ruled out explicitly).

Also de-dup at the JSON-render site with lib.unique, so an agent that
names a base-set entry itself doesn't get it installed twice, and
reword the option doc, which still claimed the old REPLACES semantics.

Four module-eval cases cover the unset / agent-adds / mkForce-replaces
/ duplicate-entry shapes.

Refs #4467
2026-09-19 10:48:29 +02:00
atlas
837e658d4a swarm: courier an agent's store identity into its container, and log in with it
`swarm-controller` mints an agent's mTLS leaf at creation and publishes it
at `swarm/agents/<agent>/bao-mtls`. Nothing read it back. This adds the
hop that carries it the rest of the way, and the in-container consumer
that proves the hop works.

Host side, `lifecycle::agent_identity` reads the row under *this hive's*
own certificate — the hive is a principal the store already knows — and
stages the leaf and its key `0600` under a new `agent-identity/<name>`
state dir, deliberately outside every bind-mounted tree. Both files go in
as systemd credentials rather than binds, the same answer and the same
mode reason as the queue secret beside it: the staged key is unreadable
to the unprivileged agent user, and the container manager reads a
`--load-credential` source as root before re-exposing it under the
consuming unit's own `User=`. The agent is never asked to authenticate in
order to obtain the thing it authenticates with.

Container side, `hive-agent-bao-identity.service` logs in with that
certificate and reads the agent's own path back, failing the unit when
either step does not succeed. It fails loudly where the hive-side readers
degrade quietly, because a refused certificate means an agent that
believes it reaches the store and never does — a cause only the login
itself can name.

The address is the whole switch, no separate `enable`, matching how
`queue.nix` and `logs.nix` already gate themselves. A hive with a store
forwards `HIVE_AGENT_BAO_ADDR` and every agent on it gets the check; a
hive without one forwards nothing and no agent does. That is what keeps
the delivery from landing in a container with nothing to read it.

The hive can now reach an agent's identity, so hive privilege covers
agent privilege. Accepted, not mitigated: the alternative is an agent
fetching its own credential with a credential it does not yet have.

Refs #4137
2026-09-19 01:55:31 +02:00
atlas
99b141f5f2 matrix: drop the per-agent matrix.enable; accounts are the enable signal
`services.hyperhive.agent.matrix.enable` was a second source of truth for
a fact the account set already carried: after ①-③ the hive-internal
`main` account is an ordinary `matrixAccounts` entry, so "does this agent
have matrix" and "does this agent have an account" were the same question
asked twice, with the boolean able to disagree.

The option is gone and a non-empty `matrixAccounts` now gates the daemon
unit, its token path-watcher and the injected `extraMcpServers.matrix`
entry.

That is only a real condition because `matrixAccounts.main` is itself
gated: it is declared when `matrix.url != null`, never unconditionally. A
`main` with no homeserver is an account the daemon can never log in as,
so declaring one always would have made the signal trivially true and
turned matrix on for every agent in every hive. With the URL gate, the
empty set is reachable exactly for an agent the hive gave no homeserver
and whose operator declared no account of its own — the state the old
`enable = false` expressed.

Assertions: "extras require enable" is deleted, having become the
definition of the thing it checked (an external-only account with its own
homeserver is now rendered rather than rejected). `main.tokenFile` stays
pinned, re-guarded on `? main` instead of on the flag, since `main` is
absent whenever the URL is null and an unguarded index would throw there.

Both spellings of the option get `mkRemovedOptionModule`, following
../host-modules/deploy.nix's registrationTokenFile pair rather than a
silent delete: the definition whose meaning changes is `false`, and left
undeclared it would be ignored and hand the agent the tools its operator
turned off. Failing the eval with the replacement spelling is the only
outcome that cannot.

module-eval gains the three arms — URL, nothing, external-only — with the
middle one carrying why it exists: it is the only thing in the suite that
would notice `main` becoming unconditional again.

Refs #4475
2026-09-18 10:35:16 +02:00
atlas
c74249f371 matrix: make the hive-internal main account an ordinary matrixAccounts entry
`matrixAccounts` is meant to be the agent's full account list, but the
hive-internal `main` account was outside it: the nix module emitted only
the extras and `hive-matrix-daemon` prepended a `main` it synthesized
from the per-agent paths, with the option schema forbidding the name
outright.

nix/agent-modules/matrix.nix now declares `main` itself, as an ordinary
entry under `matrix.enable`, from the state-dir paths the module already
used for its token path-watcher (now a shared `stateDir` binding) plus
`matrix.url`. The whole set, `main` included, is serialized to
HIVE_MATRIX_ACCOUNTS.

accounts::configured therefore synthesizes `main` only when the parsed
list carries none, and otherwise takes the declared one verbatim —
hoisting it to index 0, since the daemon reads index 0 as the primary
and nix serializes an attrset, so `main` sorts wherever its key falls.
Declared xor synthesized: an agent whose harness predates this entry
keeps working, a current one gets its own, and there is no arrangement
where `main` is duplicated or missing.

The reserved-name assertion is replaced rather than dropped: the name
must now be legal (the module uses it), but `main`'s tokenFile stays
pinned to `<state>/matrix-token`, since hive-c0re provisions the
hive-internal token there and nowhere else — a retarget would evaluate
fine and then never restore. The other two fields are mkDefault and free
to override.

Refs #4475
2026-09-18 09:34:44 +02:00
atlas
3662eda440 nix: move the agent option namespace under services.hyperhive.agent
Every per-agent harness option lived at the top-level `hyperhive.*` while
the host tier has always been `services.hyperhive.*`. Move all 52 agent-tier
option leaves (33 top-level names across 16 modules) to
`services.hyperhive.agent.*`, repoint every read, and keep existing agent
configs evaluating through one `mkRenamedOptionModule` per old leaf path in
the new nix/agent-modules/renamed-options.nix.

The shims are per leaf rather than per namespace: `user`, `mcp`, `otel`,
`queue`, `docs`, `forge`, `frontend`, `github`, `gui`, `logs`, `matrix` and
`cargo` are plain attrsets of declarations, not submodule-typed options, so
a parent-path rename would not reach their children. Three read-only
options (`frontend.mergedDist`, `queue.clientIdFile`,
`queue.clientSecretFile`) deliberately get no shim — a rename contributes a
definition, which a read-only option refuses; the exclusions are commented
in place.

Refs #4473
2026-09-17 20:19:30 +02:00
atlas
a39399f037 swarm-logs: an agent's CLI for the swarm log store
An agent can reach VictoriaLogs only through the gateway, and since the
machine query route landed the way to read it has been to hand-roll a
client_credentials token request and a curl, per query. This is the CLI
that closes that: `swarm-logs query '<LogsQL>'`, matched log lines on
stdout, so the answer pipes into grep like any other command's.

Built to the plan posted on the tracker thread: own crate, own
docs/tools reference generated off the clap tree, `query` as the one
verb, and the JSON error body surfaced on a non-200 rather than
swallowed. No `tail`: streaming is a different endpoint with a different
response shape, and folding it in here would be a fatter scope than the
ask.

Minting the token is NOT implemented here — swarm-queue-client already
owns the client_credentials request, its error type and its CA handling,
and a token-endpoint fix has to be findable in one place. What this crate
adds is the agent-shaped half: the client id arrives as a *file* beside
the secret, so nothing outside nix/agent-modules/queue.nix spells
`hive-<name>-agent` twice. That is the same problem hive-agent's
swarm_queue module solves, and swarm-logs/src/auth.rs is its `decide`
restated over this binary's inputs.

⚠️ The plan named one thing to verify empirically before calling the auth
settled: whether authelia's bearer policy for the logs vhost accepts the
agent client's audience. Measured from inside a container: it does not.
The client minted a token fine but with `aud: []` and `scp: []`, asking
for the logs URL as an audience answered `invalid_target`, and presenting
the audience-less token to the gateway answered a bare 401. So
swarm-authelia.nix's agentClients gains `authelia.bearer.authz` and the
query URL as a second audience — authelia authorises a bearer token by
the URL being requested, and that URL is now one binding read by three
places rather than three spellings of one address.

The URL reaches an agent the same way its queue coordinates do: computed
on the host (a container cannot derive a gateway address), forwarded by
hive_c0re::meta into the container's option set, and consumed by a new
agent module that installs the binary *wrapped* with its coordinates —
the shape swarm-controller.nix installs swarmctl in. Gated on the queue
credential as well as on the URL: a binary that can only answer 401 is
worse than no binary, because an agent reads a 401 as "no logs", which is
the exact confusion the store's machine route was added to end.
2026-09-17 01:02:14 +02:00
atlas
c3f5479ca8 agent-modules/network: accept unicast DHCP renewal replies unconditionally
A unicast DHCP renewal reply currently reaches dhcpcd only by matching
the firewall's ESTABLISHED,RELATED conntrack rule against the outbound
request. When that conntrack entry has already expired the reply is
dropped silently, with no log line anywhere. The client's broadcast
paths (DISCOVER, rebind) bypass netfilter entirely via a raw BPF
socket and never depend on this state — only the unicast renewal path
does.

This removes that dependency by accepting DHCP client traffic
unconditionally, gated on the firewall being enabled at all. It does
not identify or claim to fix the cause of any particular observed
renewal failure.

Refs #3389
2026-09-16 19:54:41 +02:00
atlas
46a6317f13 subagent: gate the daemon's model against availableModels
The subagent daemon put `model` straight onto claude's argv with no
validation, so an agent could spawn nested sessions on any model the
operator had deliberately kept off its harness. Forward the existing
`hyperhive.availableModels` onto the daemon unit as
HIVE_AVAILABLE_MODELS (same rail `HIVE_TOOL_GROUPS` uses) and check
`start`/`continue` against it before building the config.

Default open: an absent var restricts nothing, so an agent deployed
before this keeps working. An omitted `model` is always allowed — it
lets claude pick its own default rather than naming one.

Refs #4436
2026-09-15 23:27:02 +02:00
atlas
4121ccf1f3 subagent: let the daemon see the tool groups it resolves --tools from
`build_config` now resolves a subagent's `--tools` from
`HIVE_TOOL_GROUPS`, the same var the harness resolves its own session
from — but the meta renderer writes that var onto the `hive-agent` unit
alone (`systemd.services.${service}.environment`), and the subagent
daemon is a separate unit. It would therefore have resolved the default
groups no matter what the agent was actually granted.

That direction is safe — the default groups add no built-ins, so the
resolution is a subset of the parent's either way, never a superset — but
it isn't what the code says it does: an agent granted `web_tools` would
spawn subagents silently without `WebFetch`/`WebSearch`, and the "same
set as the parent" property would be true only for agents whose groups
happen not to matter.

Forward the var onto the daemon's unit, read off the harness unit rather
than re-derived, so there is one place it is decided. Absent stays
absent: `or null`, which systemd drops from the unit, leaving the daemon
the same fallback the harness would take.

Refs #4416
2026-09-15 17:40:27 +02:00
atlas
34129d776c subagent: give each run its own signal URL, and drop the name argument
`goal_reached`/`need_help` took the session name as a tool argument, so
identity was an assertion by the caller and the only guard on it was
`occupancy()` — "does that name have a turn in flight", which two
concurrently running siblings both satisfy for each other. A subagent
could stop its sibling's run by naming it.

Identity moves into the URL. Each spawned run is minted an unguessable
token (`Uuid::new_v4`, the OS CSPRNG), the URL carrying it goes into that
one subagent's own `--mcp-config`, and the route resolves it back to a
session before dispatching to a handler bound to that session. Neither
tool takes a `name` any more: a subagent has no field in which to name a
sibling, and a sibling's name — which a brief may well mention — is not a
token.

One route with a path parameter, not a route per session: the `Router` is
built once at startup and subagents come and go for the daemon's whole
life. An unminted or revoked token gets a bare 404, the same answer either
way, so nothing enumerates. A run's token is revoked when the run ends
(`finish_turn`) or when a call never reached a spawn.

Two things fall out of that:

- the config file becomes one per session. A single shared path was
  already a race between two `start`s; with a per-session URL in it, the
  loser would read the winner's identity.
- `occupancy()` stops being the identity guard and is gone from the signal
  path entirely rather than kept "just in case" — a revoked token can't
  reach it, and it never answered the question it was standing in for.
  It still backs `status`, which is what it was always actually for.

Refs #4403
Refs #4413
2026-09-14 22:24:51 +02:00
atlas
b18348bc9a subagent: give a run a goal, turns toward it, and a reason it stopped
`start` takes an optional `goal`. With one set a session stops being a
single turn: when a turn ends and nothing has said to stop, the daemon
spawns another turn re-prompting the subagent toward that goal, up to
`max_turns` (default 5, per-session). Without a goal nothing changes —
one turn, one todo, same as before.

Four things end a run, each recorded distinctly and reported by `status`:
the turn ending with no goal, `goal_reached`, `need_help`, and the turn
cap. The last says so out loud rather than stopping quietly — the todo
states the harness limit was reached and the goal was never reported
reached. Every stop extends the done message rather than replacing it,
and lands in the session's report file when it has one. The path is
never inferred: it comes from `start`'s `report_file` or from the
subagent naming where it wrote.

`goal_reached` and `need_help` are the subagent's own, served on a second
route (`/signal/mcp`) that carries those two tools and nothing else, so
reporting on a run can't become starting one. `goal_reached` is built as
a label, never a gate: it is self-reported by a subagent that has just
been re-prompted with "you haven't reached the goal", which is exactly
the incentive to claim it — the same failure class as a build report
asserting the tests pass. Every surface that renders it says so.
`need_help` is the blocking signal, and shows in `status` as its own
state so a parent polling it sees the block without reading a file.

`status` also carries `turn N of M`: with 4330's last-event age, that
separates working from wedged from out of turns off one answer.

Two bugs the new tests caught: a `tokio::fs::File` was dropped without
flushing, so the report line was written to nothing, and the plain idle
answer dropped the turn counter.

Also documents `await_resume`'s third case — a closed channel with no
send, which fails open the same as `Underway` — per argus on #4411.

Refs #4403
2026-09-14 21:46:59 +02:00
atlas
6de28514cd fix(otel): let file_storage extension create its own directory
validateConfigFile runs `otelcol validate` at nix build time, in a pure
sandbox where systemd's StateDirectory= has not run yet, so the
directory named by extensions.file_storage.directory does not exist.
Setting create_directory = true lets the validator create it itself,
matching what happens at runtime once StateDirectory= has acted.

Refs #4375

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-13 21:43:14 +02:00
atlas
fbd9afa7fa otel: persist journald cursor across collector restarts
The journald receiver runs with --lines=0, so every collector start
only ships what's written after it starts, and a restart silently
loses whatever landed while it was down. Point it at a file_storage
extension so the read cursor survives a restart; start_at stays at
its 'end' default since the cursor now covers everything after the
first run.

Refs #3818

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-13 19:55:00 +02:00
damocles
670e0ccad3 docs: drop the auto-injected hyphen vale flags 2026-09-13 13:57:53 +02:00
damocles
16eec3c314 subagents: add availableToSubagents opt-in toggle for extraMcpServers 2026-09-13 13:57:53 +02:00
atlas
33aa2bdbc8 subagent daemon: scope an OOM kill to the session that lost the draw
systemd's default OOMPolicy=stop tears the whole unit down the moment the
kernel kills any process in its cgroup, so a single over-large subagent
takes the daemon and every sibling session with it — measured on a live
agent, both units show OOMPolicy=stop today, which is exactly the "every
live subagent session was cut mid-turn" symptom.

That blast radius is also what would make the preceding commit a bad
trade: deliberately putting this unit first in the OOM queue is only an
improvement if losing one subagent isn't losing all of them. OOMPolicy=
continue scopes the loss to the process the kernel actually chose, and
leaves the daemon alive to report the kill instead of vanishing and being
restarted with an empty session map.

Refs #4316
2026-09-13 13:05:17 +02:00
atlas
d249468db2 agent: make the OOM killer prefer a subagent over the agent's own turn
Both units ran at OOMScoreAdjust=0, so under container memory pressure the
kernel picked purely on footprint — and the agent's own claude is often
the fattest process in the container, which means the session supervising
the work died before the work did.

The sign is the load-bearing part and is easy to invert: a HIGHER
OOMScoreAdjust means MORE likely to be killed, because the kernel adds it
to the badness score it derives from the process's memory footprint and
then kills the highest scorer. So hive-subagent-daemon gets +500 (first in
line) and hive-agent gets -500 (last in line). Written backwards this
makes the reported bug worse rather than better, so module-eval pins the
order as an inequality.

Both values are inherited by the nested claude each unit spawns as a
child, so ordering the units orders the sessions underneath them. -500
rather than -1000 on the harness: fully exempting it would leave the
kernel nothing to kill in a container whose only large process is the
harness.

Refs #4316
2026-09-13 13:04:45 +02:00
atlas
4daab7efe4 subagent daemon: throttle at two thirds of the container's memory
The daemon spawns every nested claude as a plain child, so its cgroup is
already the "all subagents" cgroup — but it ran with MemoryHigh=infinity
and MemoryMax=infinity, so nothing slowed a subagent down before the
kernel's OOM killer stopped the unit and cut every live session with it.

MemoryHigh= and not MemoryMax=: a soft ceiling reclaims and stalls the
cgroup past two thirds of the container's cap, which turns a silent kill
into a visible throttle, while still letting a single subagent exceed its
share when the container has memory free. A hard per-agent cap would make
overprovisioning impossible, which is not wanted — most of the time
nothing in these sessions is compiling.

The fraction is taken from hyperhive.claudeMemoryMaxBytes, the container's
own effective MemoryMax= that meta.rs already bakes in per agent. When
that is null (an `infinity` or percentage cap) the unit renders no ceiling
rather than a fabricated constant, and module-eval pins both arms.

Refs #4316
2026-09-13 13:04:00 +02:00
atlas
2989c5ccdb swarm: say "no queue coordinates", never "a hive with no queue"
The swarm always has exactly one queue; a hive can only lack its
address. Reworded every prose site this PR added that stated or
implied the opposite, to name what is actually absent (coordinates,
credential, or address) instead of the queue itself.

Refs #3805
2026-09-13 11:13:17 +02:00
atlas
86652f051a swarm: wire the agents' queue coordinates and credential through the modules
The host end: `HIVE_C0RE_AGENT_QUEUE_CREDENTIAL_DIR` tells the daemon
where the reader unit put the files, and a new
`deploy.hive-controller.queue.agentNatsUrl` says where the queue is as an
agent *container* reaches it. That address defaults to the bridge one and
never to loopback — `statusPublish.natsUrl` beside it is loopback and
correct, because hive-c0re shares the host netns and an agent does not.
Paired with the swarm's token endpoint, gated together, and forwarded by
`hive_c0re::meta` as both an env var and an agent option: the harness
reads the variable at runtime, its unit is built from the option.

The agent end: `nix/agent-modules/queue.nix` declares that option pair
and, when set, has the harness unit inherit the two credentials by name.
Bare-id `LoadCredential=` is the terse form documented for inheriting
what the service manager received, and is non-fatal when the credential
is absent — which a hive whose publisher has not run yet needs.

No `HIVE_AGENT_OIDC_CA_FILE`: the meta flake already embeds the hive CA
and the swarm root into each container's trust store at build time, and
reqwest's rustls backend verifies against it.

Refs #3805
2026-09-13 11:13:17 +02:00
atlas
5af1f6a8e5 docs, mcp.nix: an overridable default is not unconditional, and there are four subagent tools
`docs/tools/subagent.md` and `docs/tools/bash.md` both described their MCP
server as injected "unconditionally". Both entries are `lib.mkDefault`, and
the module says why one line above each: "so an agent.nix can still
override/disable the entry", "so the operator's own agent.nix can override
the entry".

The word matters for the subagent one in particular. The same comment block
records the framing that it is default-on for now and should become a real
capability gate later, so "can I turn this off today?" is a question an
operator has — and "unconditionally" answers it as "patch nix/" when the
answer is one override in agent.nix.

Both pages now say default, and say what the default yields to.

The other direction on the same page: `subagentHttpPort`'s option
description and the unit comment beside it both listed three tools,
`start`/`continue`/`interrupt`. The daemon serves four. #4101, which
introduced it, is titled with the three-verb phrasing, so `status` landed
afterwards and never reached either description — while `subagent.md` had
the full set all along. The option description renders into the generated
options doc, so it is the one an operator reads.

Closes #4231.
2026-09-11 16:58:12 +02:00
atlas
68711796ef otel: forward each agent container's journal to its hive collector
An agent container writes a complete journal — 991 MB and nine days deep
on this hive — that nothing outside it can read: the host-side
per-container journal directory is an id-mapped bind mount, and journald
writes nothing into it. So the reader has to run inside the container,
and the path it would push to did not exist.

Three tiers, one vertical slice, because any two of them alone are
silent:

- the agent container gains an `opentelemetry-collector` with a
  `journald` receiver aimed at its own journal and an exporter aimed at
  the same base address every in-process producer already exports to.
- the hive collector gains `service.pipelines.logs`. Without it the
  `otlp` receiver answers 404 on `/v1/logs` — measured, and
  indistinguishable from a route that was never meant to exist.
- the swarm collector gains a per-hive `logs/<hive>` pipeline beside
  `metrics/<hive>`. Without it the push is accepted, answered 200, and
  routed nowhere.

The receiver's `directory` is stated rather than inherited, and that is
the load-bearing line: its default is the RUNTIME journal
(`/run/log/journal`), which in an agent container is empty. Left at the
default this whole path validates, starts, reports healthy and forwards
nothing. The assertion beside it covers the same silence from the other
end — a `volatile` or `none` journald storage empties the directory the
receiver reads.

Attribution follows the tier that can prove it. The forwarder stamps
`agent`, which no host-side reader could supply; `hive` is deliberately
left to the swarm tier, which upserts it from whichever receiver
accepted the sample, precisely so the label comes from something the
sender cannot write.

No `units` allowlist, unlike the swarm tier's journald receiver. That
one needs one because the host's journal also holds an operator's own
session; a container's journal is the harness and what the harness
spawns. Measured volume is 20827 entries / 6.3 MB per agent per day,
with nothing logging below `info` — so the receiver's `info` default
filters nothing and there is no bill to justify a knob.

Agent containers only, per the ruling on the issue: swarm services need
one forwarder per service container and get re-measured once this works.

Part of #3940.
2026-09-11 09:03:49 +02:00
damocles
561bd09618 subagent: close the start/continue TOCTOU race with an atomic reservation 2026-09-09 23:45:12 +02:00
damocles
c280664d74 nix: wire the independent hive-subagent-daemon systemd unit and MCP server 2026-09-09 23:45:12 +02:00
damocles
451b461afb docs: matrix gateway vhost defaults to chat.<swarm-domain>, not matrix.<domain> 2026-09-07 16:53:22 +02:00
atlas
af5dce0648 matrix: scope the daemon's token watcher to this agent's own state dir
`hive-matrix-daemon.path` globbed `/agents/*/state/matrix-token*`. Every
agent's state dir is visible from inside every container, so the
condition is satisfied by a sibling's token.

That is reachable, not cosmetic. The daemon deliberately exits 0 when it
has no token of its own — `Restart = "on-failure"` therefore does not
restart it, and the unit sits inactive, which is the state the path unit
exists for. In that state a sibling's token keeps the glob satisfied:
the path fires, the daemon exits 0, the unit deactivates, the path
re-arms, the condition is still true. systemd.path(5) activates a
`PathExists`-family condition that already holds immediately on arming,
so it repeats until the start limit stops it.

Scoped to this agent, the condition is false exactly when the daemon
would have nothing to do.

The glob is quoted in four other places, all of which would otherwise
name a pattern that no longer exists — a doc, a Rust doc-comment in
hive-c0re, a nix comment, and an assertion message an operator reads.
Each is reworded to the basename (`matrix-token*` in this agent's state
dir), which is what the assertion actually enforces via `baseNameOf`, so
they stay true wherever the directory moves.

Refs #4030.
2026-09-07 13:37:09 +02:00
atlas
e9faa3896f forge: stop the avatar sync retriggering itself into the start limit
`forge-avatar-sync.path` used `PathExists=`. systemd.path(5): a
`PathExists=` condition that already holds activates the configured unit
immediately whenever the path unit is activated. A `Type=oneshot` unit
with `RemainAfterExit=false` deactivates after each run, which re-arms
the path, which fires again because the token is still there — an
unconditional loop that ends at `StartLimitBurst`.

Measured on a container boot carrying the previous fix: five starts and
`start-limit-hit` with the glob already removed and exactly one matching
file. So the count was never one-per-token; it was the start limit, and
scoping the watch (#4025) could not have fixed it.

`PathChanged=` does not fire on an already-present path, and hive-priv
writes this file in place — `write_state_file_nofollow` opens with
`O_TRUNC` and no rename — so close-after-write still triggers it. The
token-present-at-boot case stays covered by the service's own
`wantedBy = multi-user.target`.

Same directive and same reasoning as `swarm-controller.nix`'s
queue-credential watcher, which reached it first: "`PathChanged=` requires
a write, so it cannot do that and cannot spin."

The comment claiming the storm came from many agents' tokens arriving at
once is removed with it — that model is what let the defect survive the
previous fix.

Refs #3984.
2026-09-03 20:29:48 +02:00
atlas
3b5bdcf262 forge: the avatar path unit watches this agent's token, not every agent's
The service reads `$HYPERHIVE_STATE_DIR/forge-token` — its own. The path unit
that retriggers it globbed `/agents/*/state/forge-token`, and every agent's
state dir is visible from inside every container, so a sibling's token
appearing re-fired this agent's sync.

Enough of them arrive together to trip systemd's start rate limit, so the
unit ends `start-limit-hit` after the upload has already succeeded: a red
[FAILED] on every container on every boot, for work that worked.

Measured on this container at tonight's 23:54 boot, before the change: five
`avatar uploaded (HTTP 204)` inside one second, then `Start request repeated
too quickly`. After it lands, that boot line should read one upload and no
limit.

The path is spelled the way the same file already spells it for tea-login,
223 lines up — `userName` was in scope the whole time.
2026-09-03 01:25:57 +02:00
atlas
d2747c7b77 refs: repoint seven comments that name files which have moved
Comments cite nix modules, scripts and crate source files constantly,
and nothing evaluates a comment — so when a file moves, the reference
rots silently and `nix flake check` stays green. A reader following one
finds nothing and cannot tell whether the file was renamed, deleted, or
never existed.

Seven such references, each repointed at the file that actually holds
the thing the sentence is about rather than at the directory the old
name became:

  hive-c0re/src/agent_config/limits.rs   hive-agent/src/mcp.rs
                                       -> hive-agent-mcp/src/mcp/mod.rs
  hive-agent-mcp/src/mcp/mod.rs          hive-c0re/src/limits.rs
                                       -> hive-c0re/src/agent_config/limits.rs
                                         (and the module path in the doc
                                          comment above it, which was stale
                                          in the same way)
  hive-c0re/src/forge/mod.rs             hive-c0re/src/knowledge.rs
                                       -> hive-c0re/src/workers/knowledge.rs
  nix/host-modules/hive-c0re/options.nix hive-c0re/src/hive_stats.rs
                                       -> hive-c0re/src/stats/hive_stats.rs
  nix/packages/default.nix               nix/host-modules/hive-c0re.nix
                                       -> .../hive-c0re/options.nix
  nix/agent-modules/network.nix          nix/host-modules/hive-gateway.nix
                                       -> .../hive-gateway/dnsmasq.nix
  frontend/README.md                     nix/modules/frontend.nix
                                       -> nix/packages/frontend.nix

The two `limits.rs` comments are a matched pair: each names the other's
old path, so the "keep in sync" instruction they exist to carry pointed
both ways at nothing.

Where a flat module became a directory the target is the file that
declares the named thing, not `default.nix` by reflex — the
`preBuildAgentTemplates` option is declared in `options.nix`, and the
DHCP pool that sentence is about lives in `dnsmasq.nix`.

Comments only; no behaviour change. Refs #3923, which is about whether a
gate should cover this class at all — that question is unanswered and
this does not close it.
2026-09-02 08:58:31 +02:00
iris
07b62612b0 docs: restructure into topic subdirectories, collapse duplicated index
Per mara's go-ahead on hyperhive#3902 ("getting started is good, but
terminal rendering does not go in there i think"):

Moved 21 top-level docs/*.md files into 7 new topic subdirectories
(existing web-ui/, turn-loop/, swarm/, tools/, crates/ untouched):
  getting-started/  setup.md
  agent-lifecycle/  agent-hierarchy.md, approvals.md, persistence.md
  trust-boundary/   boundary.md, security.md
  integrations/     forge.md, matrix.md, github.md, knowledge.md
  networking/       gateway.md, network.md, snapshot-store.md
  scheduler/        jobq.md, coordinator.md, ci.md, observability.md
  process/          conventions.md, gotchas.md, pr-review-gate.md
  web-ui/           terminal-rendering.md (moved into the EXISTING dir,
                    per mara's correction to the original getting-started
                    guess -- it's UI implementation detail, not onboarding)

The physical layout now matches docs/README.md's own topical headers,
which already amounted to this taxonomy -- see the scoping comment on
the issue for the two findings that motivated this (a genuine
duplication between CLAUDE.md's old "Reading paths" list and
docs/README.md's grouped one, since drifted out of sync with each
other; and the flat layout not matching the grouping we already had).

Fixed every cross-reference this moved across the whole repo (~120
files: docs/ internal links at every depth, Rust doc comments, nix
module option docs, crate READMEs) -- verified two ways: a grep sweep
confirming zero remaining references to any old path, and a script
that resolves every markdown link in docs/**/*.md + CLAUDE.md +
README.md against the filesystem and reports anything that doesn't
exist (zero broken links).

Collapsed CLAUDE.md's "Reading paths" section (the duplicate) down to
a pointer at docs/README.md, now the single index. Rewrote
docs/README.md itself to use the new subdirectory paths and added the
one doc it was missing that CLAUDE.md's old copy had (pr-review-gate.md).

Classified all 22 docs/*.md files first via a haiku subagent (mara's
suggestion) on two axes -- proposed grouping and operator-vs-
implementation focus -- before finalizing the taxonomy; spot-checked
the report and found internal inconsistencies (its classification
table disagreed with its own summary section for a few files), so this
taxonomy is my original proposal + the one correction mara gave
directly, not a blind application of the subagent's table. The
operator-focus data it gathered is still useful for a follow-up
content pass (docs skewing 'mixed' rather than pure operator-facing),
not addressed in this PR -- structure only.

nix fmt clean, both pre-push lints clean.
2026-09-02 01:55:37 +02:00
atlas
d17aa254dd nix, hive-sh4re: name modules that exist in the stale harness-base refs
harness-base.nix has never existed in this tree. Four comments named it,
or a `harness-base` module, as the place to look:

- weston-vnc.nix: the agent user is declared and home-chowned by
  nix/agent-modules/user.nix
- hive-ci.nix: the sandbox-fallback reasoning lives in
  nix/agent-modules/default.nix -- which the very next comment block in
  the same file already cites correctly
- packages/default.nix: the per-bin consumer is
  nix/agent-modules/packages.nix
- hive-sh4re/src/assets.rs: HIVE_ASSETS_DIR is set by
  hive-c0re/environment.nix and agent-modules/default.nix +
  agent-service.nix, and the package is built by nix/packages/assets.nix
  -- not the equally nonexistent nix/assets.nix

assets.rs was twice declared out of scope on the sibling PR because it
names a module rather than a file. That distinction was real and
irrelevant: neither the module nor the nix/assets.nix path it points at
exists. Reading the wording is not checking the reference.

Every replacement path was verified to exist, with a deliberately bogus
path as a control.
2026-08-30 04:19:11 +02:00
atlas
8b845896e2 forge: name the credential helper the way git resolves it
/etc/gitconfig shipped `helper = git-credential-hive-forge`. git prepends
`git-credential-` to any helper value that is not an absolute path, so
that resolves to `git-credential-git-credential-hive-forge`, which does
not exist -- no helper runs at all. The sibling github.nix has always used
the short form.

Measured rather than read off the docs, with the arms isolated from the
personal ~/.gitconfig:

  helper = git-credential-hive-forge  -> 0 credentials, and git prints
      "'credential-git-credential-hive-forge' is not a git command"
  helper = hive-forge                 -> 1 credential, clean stderr

The reason this survived: every long-lived agent has a personal
~/.gitconfig naming the helper by ABSOLUTE path, which git accepts, so
pushes keep working and the stderr line reads as noise. The system config
is masked exactly where someone would notice it and bites a fresh agent
that has no such file.
2026-08-27 14:05:07 +02:00