Watch
0
0
Fork
You've already forked hyperhive
0
Commit graph hyperhive/hive-c0re/src/job_queue
Author SHA1 Message Date
atlas
5785c0024c Make agent creation swarm-only and refuse a name placed on another hive
swarm-controller's POST /api/agents now refuses (409) a name the swarm
has already placed on a different hive: a non-Destroyed declaration in
that hive's wanted state, or a SetAgentWanted node still queued for it.
The same name on the same hive is that agent being re-created and goes
through. A wanted state that cannot be read refuses (503/500) instead of
reading as "placed nowhere". Creations are serialised from that read to
the graph insert so two concurrent creations of one name cannot both
pass.

Hive-level creation is removed: hivectl `agent create` / `request-create`,
HostRequest::Spawn / RequestSpawn, the dashboard POST /api/request-spawn
route, and ApprovalKind::Spawn with its approve/resolve arms and the
approval-carrying `templates::spawn`. The swarm path (deploy request or
wanted-state sweep -> queue_first_deploy -> templates::first_deploy) used
none of them. Old `spawn` approval rows are skipped by collect_lenient,
as `init_config` rows were in a3b672d1.

policy.rs's comment on agent_object_name stated swarm-wide name
uniqueness as a fact; it now says where it is enforced and what that
check cannot see.

Refs #4396
2026-09-29 15:47:40 +02:00
atlas
313d582c01 job_queue: drop insert_unless_live, accept extra queue passes
mara chose to accept extra queued sweep passes over adding a new
hive-jobq primitive (or a hive-c0re one-off) for "don't queue another
of this kind". Every sweep caller now plain-inserts its node; the
capacity-1 Dep::Resource per sweep kind (MatrixSweep/KnowledgeTree)
still keeps two passes of the same kind from running concurrently, it
just no longer collapses a tick that lands while one is live or queued
into the existing one.
2026-09-27 20:15:55 +02:00
atlas
1d4c77d2c8 hive-c0re: serialise the matrix and knowledge sweeps through the job queue
The matrix sweep and the /knowledge pull each had concurrent callers
(#4723 item 4). Two overlapping knowledge pulls fail on .git/index.lock
and the remote-tracking ref lock: 30 of 30 concurrent replays of the
reset/clean/pull sequence in a scratch repo errored, 0 of 10 sequential
ones did. Two overlapping matrix sweeps on a hive with no persisted
Space / chat-room id both miss the by-name lookup and both createRoom
(from reading the code, not reproduced against a homeserver). On every
boot the MatrixSweep DAG node and the main.rs loop's immediate first
call ran at once.

Every sweep now runs as a job node, and each sweep's node holds its own
capacity-1 queue resource (Resource::MatrixSweep,
Resource::KnowledgeTree), the MetaWindow pattern: the scheduler never
starts a second pass of one sweep while the first holds the resource,
and different sweeps still run side by side.

- templates::matrix_sweep / templates::knowledge_pull build the node
  with its resource; boot, the periodic loops and the swarm event all
  use them.
- JobQueue::insert_unless_live folds a submission into a live node of
  the same kind instead of queueing another. Periodic ticks fold into a
  queued or running pass. The swarm knowledge event folds into a queued
  pull only, and queues one behind a running pull, which may have
  fetched before the push.
- The main.rs matrix loop no longer sweeps immediately at startup; the
  boot MatrixSweep node is the startup pass, as KnowledgePull already
  was for knowledge.
- The executors bound each pass (10 min matrix, 5 min knowledge), since
  a hung pass would otherwise hold its resource against every later one,
  and own the sweep-health banners, so every pass reports to them.

Replaces the SweepLock version of this branch, per review.

Refs #4723
2026-09-27 20:15:55 +02:00
atlas
2252c55df8 hive-priv: create agent socket dirs on start; drop hyperhive-agents.conf
/etc/tmpfiles.d/hyperhive-agents.conf was a boot-time backstop (#2290)
that pre-created every agent's bind sources. The start preamble already
creates them for every c0re-driven start, and on this host only hive-c0re
starts agent containers. The file was also the reason the socket dir's
owner had to be declared there, which is how it spent its life at
`0777 root root` whenever the uid could not be resolved (#4742).

- hive-priv gains `EnsureAgentSocketDir { name }`, called from
  `set_nspawn_flags` in every start path. It creates
  `/run/hive-agent/<name>` `0751 root:root` with mkdirat relative to an
  O_DIRECTORY|O_NOFOLLOW fd for the parent. An existing entry has to be a
  directory (fstatat AT_SYMLINK_NOFOLLOW); anything else is refused, and a
  directory is left alone. hive-c0re's own create_dir_all went: its /run
  is read-only under ProtectSystem=strict.
- The container's `hive-agent-user-migrate` activation chowns that dir to
  the agent user and sets 0751, the same way it already handles state/ and
  harness/. It refuses a symlink or non-directory there, since `test -d`
  and chmod follow links. No host-side passwd parse, and no window where
  the dir is world-writable.
- `/run/hyperhive/agents/<name>` stays created by hive-c0re itself
  (`ensure_agent_runtime_dir`). It holds the `mcp.sock` that hive-c0re
  binds as hive-core, so it must not become root- or agent-owned.
- The `/run/hive-agent` parent is declared in hive-priv.nix, `0755
  root:root`, instead of hive-gateway's hive-core rule. hive-priv is its
  only writer now, and hive-priv's ReadWritePaths needs it to exist.
- The manager start in `ensure_root_agent` now goes through
  `converge_start_preamble` + `start_with_fallback`. It was a bare start,
  so after a reboot the manager's bind sources existed only because of the
  tmpfiles file, and its limits drop-in did not exist at all.
- Removed: `sync_tmpfiles`, `agent_uid_gid` / `parse_passwd_uid_gid`,
  `priv_client::sync_agent_tmpfiles`, `AgentTmpfilesEntry`, the tmpfiles
  body builder and their tests, plus the three call sites.
- Legacy: hive-priv unlinks the file at every start, ignoring ENOENT.
  `SyncAgentTmpfiles` stays one release as a payload-ignoring variant that
  does the same unlink and returns Ok, for an older hive-c0re.

Salvaged from #4752: the boundary.md correction that nginx only dials,
because ProtectSystem=strict makes its /run read-only.

Behaviour change: a manual `nixos-container start h-<name>` right after a
reboot, before hive-c0re has started that agent, now fails on a missing
bind source instead of starting.

Closes #4742
2026-09-27 18:55:33 +02:00
atlas
7ac6819652 hive-c0re: fail on a malformed agent name and on an unreadable container list
An agent name that is not a valid Ident made `Coordinator::agent_paths`
panic. Job payloads carry names as plain strings (the swarm's published
wanted state is one source), and a panic inside a job-queue node never
reaches `complete_growing`, so the node's resources (the deploy
window included) were held until hive-c0re restarted. `agent_paths` now
returns an error; the job-queue nodes, the admin-socket spawn and
set-limits paths, the root-agent spawn and the dashboard set-limits
handler propagate it.

`lifecycle::list().await.unwrap_or_default()` turned a failed container
list into "no agents":
- meta-update cascade: the lock bump committed and zero rebuilds fanned
  out, reported as success. The cascade is now resolved before the lock
  bump and a list failure fails the node.
- dashboard update-all: queued nothing and returned 200 "ok". Now 500
  with the error.
- container rescan: every row was emitted as removed and the cache
  emptied. Now the last snapshot stands; `hivectl status` gets an error.
- dashboard journal: answered 404 "no managed container". Now 500.
- spawn/rebuild port-collision check: silently skipped. Now fails.
- startup migration: the per-agent phases ran over nothing, and phase 3
  handed an empty agent list to `meta::sync_agents`, which renders the
  meta flake with exactly the agents it is given. Both now log the list
  failure and skip.

The hive-jobq scheduler still leaks a node's resources on any executor
panic; that root is not addressed here.

Refs #4723
2026-09-27 05:13:21 +02:00
atlas
bfd8189900 hive-c0re: stop reporting refused invites as success; re-register the config-PR hook when its secret changes
invite_user_id mapped every 403 M_FORBIDDEN to Ok(()). The membership
pre-check already skips invited/joined users, so the 403s that reach the
POST are mostly real refusals (banned target, sender without power),
including `hivectl matrix invite`. A 403 is now success only when a
membership re-read shows the user invited or joined; otherwise it is an
error carrying the status and body.

admin_room_send_and_poll read the send response's event_id with
unwrap_or_default() and, when it was missing, walked every recent event
unanchored, so an older bot reply (an earlier reset password) could be
returned as this command's result. A send response without an event_id
is now an error.

run_destroy_bookkeeping discarded fail_pending_for_agent's error; it now
warns like its neighbouring steps.

ensure_config_pr_webhook returned as soon as a hook with the target URL
existed, so a regenerated webhook-secret never reached Forgejo and every
config-PR delivery failed HMAC until the 5-minute poll caught up.
Forgejo's edit-hook API ignores `secret` and never returns it, so the
SHA-256 of the secret last registered is recorded at
forge/config-pr-webhook-secret-sha256; when it doesn't match, the
same-URL hook is deleted and recreated. The paths.rs doc claiming
re-registration on change now describes this.

Refs #4723
2026-09-27 04:41:31 +02:00
atlas
20419ccd41 swarm-controller: own the swarm-wide forge objects; hive-c0re stops creating them
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.

swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.

hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.

nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.

Refs #3782
2026-09-25 08:36:05 +02:00
atlas
89a5dd752c hive-c0re: stop minting agents' matrix accounts
The swarm mints each agent's `main` account now, so the hive's own mint
goes: `ensure_user_for`, `finish_user_provisioning`, `sync_agent`,
`sync_agent_standalone`, `token_path`, `legacy_password_path`,
`auto_reset_password` and `token_file_present`, and the calls from the startup sweep and the
rebuild bookkeeping. Both mints pinned the device `hyperhive-<agent>`, so
leaving this one would have each re-login kill the other's token.

`hivectl matrix create-user` refuses an agent's name and says where its
account comes from. Everything that still uses the hive's appservice token
stays: the hive's own account, the Space and chat room, and operator
accounts.
2026-09-25 08:31:01 +02:00
atlas
18bd8dd2c7 hive-c0re: fail a nixos-container destroy that leaves the container in place
lifecycle::destroy logged a failed `nixos-container destroy` and returned
Ok, so the DestroyContainer node went green, the agent was unregistered,
and with purge the after_ok PurgeState deleted its state while the
container config and root still existed.

Propagate the error unless the container list, read after the failure,
no longer names the container. An unreadable list fails too.
2026-09-24 12:18:34 +02:00
atlas
336ed5a010 docs+comments: say what changed instead of tagging the tracker item
The prose added by this branch named the tracker item in seventeen
places, which check-issue-refs.sh rejects: a `#N` tag is dead weight for
anyone reading the public mirror, where no issue data exists. Each one
now states the fact it was pointing at — the parent field is gone — so
the sentence stands on its own.

Two of those lines also carried a rustdoc break: `[`write`]` in
topology.rs is ambiguous between the module's own `write` fn and the
`write!` macro, which `-D rustdoc::broken-intra-doc-links` fails. Spelled
`[`write()`]`, per rustdoc's own suggestion.

The host_config.rs rewrite is two lines rather than three so the doc
block stays under check-comment-blocks.sh's 30-line ceiling.
2026-09-21 22:43:16 +02:00
atlas
d94bc2188d topology: drop the parent field and the hierarchy it fed
`topology.json` was a map of `name -> parent | null`, and that value fed
the whole agent hierarchy: `<parent>` / `<children>` recipient sentinels,
the reparenting API (CLI verb, wire verb, dashboard endpoints, DAG node),
the dashboard tree, the rebuild depth sort, and an unconditional
bind-mount grant giving every agent RW on its direct children's state.

Per the operator's ruling the field goes, and with it all of the above.
The file survives as what remains once the value is gone: the roster of
agent names, which is the set `ManageRootAgent` grants mounts over. It is
now a JSON array; `read` still accepts the old map shape and keeps its
keys, so a hive that upgrades across this does not blank its roster (and
so no capability holder loses its mounts for the length of that window).

Two sites kept their behaviour under a different recipient rather than
losing it. Both addressed `<parent>`, which the broker already resolved to
`operator` for a root agent, and every agent is now what that fallback
called a root:

- the harness's turn-failure / plugin-failure notification
  (`Surface::send_to_parent` -> `send_to_operator`), and
- the send allow-list's always-permitted escape hatch, so an agent with a
  restrictive allow-list still has a way to say it is stuck.

What is NOT preserved, deliberately: an agent with no capability no longer
sees any other agent's dirs. `ManageRootAgent`'s own grant is unchanged --
still every agent in the roster, still state RW + config RO, still no
`harness`.

The dashboard's reparenting control (the M0V3 picker) is deleted with its
CSS. The tree rendering that reads `ContainerView.parent` is left for the
frontend owner -- it degrades to a flat list with the field gone.
2026-09-21 22:08:47 +02:00
atlas
b88a5b2430 remove the list_containers and request_update_meta_inputs MCP tools
Both agent-facing tools go away end to end, with no replacement. This is
an intentional capability removal: agents can no longer enumerate their
own subtree, and can no longer queue a meta-flake input bump.

The system prompt and docs/tools/lifecycle.md land in this same commit
on purpose. A tool named in the prompt but absent from the server makes
agents confidently call something that doesn't exist, and the failure
then surfaces far from its cause.

Removed:

- MCP registrations and bodies (hive-agent-mcp), plus the now-unused
  UpdateMetaInputsArgs.
- Wire variants Request::ListDescendants,
  Request::RequestUpdateMetaInputs and Response::Containers, plus
  ContainerInfo, whose only consumer was that response.
- hive-c0re's handle_list_descendants (its whole module) and
  handle_request_update_meta_inputs, the two dispatch arms, and the
  require_group(agent, "approvals", ...) gate on the meta-inputs verb.
- The stream_enrich emoji entry and argument formatter.
- docs/tools/lifecycle.md (both tools it documented are gone), its two
  referrers, the tool-group tables and the agent-hierarchy prose.

Tool groups are kept, deliberately. ToolGroup::Lifecycle listed exactly
one tool and now lists none — it is vestigial, but the variant stays so
existing meta/capabilities.json grants still parse; retiring it is a
separate decision. ToolGroup::Approvals also listed exactly one tool,
but the group is NOT dead: check_can_cancel_approval still gates
cancel_loose_end's approval-cancel arm on it server-side.

ApprovalKind::UpdateMetaInputs stays too. Nothing in production code
produces it any more, but pre-existing approval rows may still carry it,
and the operator's own path to a meta update is unaffected — the
dashboard's POST /api/meta-update inserts the meta_update job directly,
bypassing approvals entirely.

The two format_ack tests in hive-agent-mcp that named
request_update_meta_inputs were only using it as a label string while
exercising the generic OkWarn/Ok renderer, so they are retargeted to a
surviving tool rather than deleted.

Note hive-c0re's priv_client::list_containers is a different thing (the
host-side privileged container listing behind hive-priv) and is
untouched.

Closes #4591
2026-09-20 22:47:46 +02:00
damocles
95898338fc hive-c0re: validate cascade agent names in meta_update_cascade_agents' fanout path too
extract validate_agent_names() and use it for both the parsed
agent-<name> inputs and run_meta_lock's pre-computed fanout list, so a
malformed name can't reach the new fast_forward_applied_main / lock_update
filesystem+git+forge-URL operations regardless of which of the two
sources it came from
2026-09-13 17:33:58 +02:00
damocles
c45d679a32 hive-c0re: fast-forward applied/<name>/main too, not just the one-shot relock
argus + mara (PR #4339 review): the previous commit's per-agent lock
relock is a one-shot effect on the single rebuild the cascade triggers
- applied/<name> never moves, so the next relock=true rebuild trigger
(the boot sweep, most notably) re-locks against applied/<name> and
reverts straight back to whatever it was stuck on. The fix didn't
outlive the transaction it ran in.

New forge::fast_forward_applied_main(name), sibling to the existing
reseed-only fetch_config_main_into_applied: for an applied repo that
already has a .git and just needs to catch up, force-set rather than
fast-forward-gated since there's no PR to review on this path either.
Called per cascade agent right after the relock, best-effort so one
unreachable agent repo doesn't block the others.

Once applied/<name>/main has actually moved, lock_update_for_rebuild's
override (always reads current applied/<name>/main, no ?rev pin)
naturally stays in sync on any later relock=true rebuild instead of
reverting.
2026-09-13 17:33:58 +02:00
damocles
f0ddbe49d0 hive-c0re: pull each meta-update cascade agent's own config-repo main
mara (hyperhive#4271): "whatever is on main is trusted and should be
pulled. dont make it periodic, just add it to the meta update when
choosing the agent."

A meta-input bump previously rebuilt every affected agent against
whatever applied/<name> was already locked to. Normally current, but
silently stale forever if a past deploy failed and nothing since
retried it — the manual meta-input trigger never special-cased that
either, since nothing wired it to check the config repo's live main at
all.

run_meta_lock now relocks each cascade agent's own input alongside the
originally-requested ones, against the input's declared source (the
forge URL) rather than the local applied/<name> mirror prepare_deploy
uses for a reviewed deploy — there's no PR to review on this path, so
nothing to gate. One combined lock_update call for the whole cascade:
simpler than per-agent isolation, at the cost of one broken/unreachable
agent repo failing the whole cascade relock rather than just that
agent (the separate top-level input bump above it is unaffected).

Known gap, not fixed here: applied/<name>'s own main/deployed/* tags
never advance from this path, only a real MergeConfigPr deploy does
that — so the audit trail stays as accurate as today, it just stops
being what actually got built.
2026-09-13 17:33:58 +02:00
damocles
01c0a6dcba job_queue: fix dangling rustdoc intra-doc link left by strum conversion 2026-09-12 00:06:31 +02:00
damocles
fa139342ed fix job_queue/tests.rs's 5 dead NodeKind::as_str() call sites, dead since the first wrapper removal 2026-09-12 00:06:31 +02:00
damocles
46456f75ce remove as_str() legacy wrappers, callers use .into() directly 2026-09-12 00:06:31 +02:00
damocles
299add158f convert hand-written enum as_str matches to strum derives workspace-wide 2026-09-12 00:06:31 +02:00
damocles
223ac257e0 job_queue: trim NodeKind/Resource doc comments now that coordinator.md covers them 2026-09-02 08:39:52 +02:00
iris
07b62612b0 docs: restructure into topic subdirectories, collapse duplicated index
Per mara's go-ahead on hyperhive#3902 ("getting started is good, but
terminal rendering does not go in there i think"):

Moved 21 top-level docs/*.md files into 7 new topic subdirectories
(existing web-ui/, turn-loop/, swarm/, tools/, crates/ untouched):
  getting-started/  setup.md
  agent-lifecycle/  agent-hierarchy.md, approvals.md, persistence.md
  trust-boundary/   boundary.md, security.md
  integrations/     forge.md, matrix.md, github.md, knowledge.md
  networking/       gateway.md, network.md, snapshot-store.md
  scheduler/        jobq.md, coordinator.md, ci.md, observability.md
  process/          conventions.md, gotchas.md, pr-review-gate.md
  web-ui/           terminal-rendering.md (moved into the EXISTING dir,
                    per mara's correction to the original getting-started
                    guess -- it's UI implementation detail, not onboarding)

The physical layout now matches docs/README.md's own topical headers,
which already amounted to this taxonomy -- see the scoping comment on
the issue for the two findings that motivated this (a genuine
duplication between CLAUDE.md's old "Reading paths" list and
docs/README.md's grouped one, since drifted out of sync with each
other; and the flat layout not matching the grouping we already had).

Fixed every cross-reference this moved across the whole repo (~120
files: docs/ internal links at every depth, Rust doc comments, nix
module option docs, crate READMEs) -- verified two ways: a grep sweep
confirming zero remaining references to any old path, and a script
that resolves every markdown link in docs/**/*.md + CLAUDE.md +
README.md against the filesystem and reports anything that doesn't
exist (zero broken links).

Collapsed CLAUDE.md's "Reading paths" section (the duplicate) down to
a pointer at docs/README.md, now the single index. Rewrote
docs/README.md itself to use the new subdirectory paths and added the
one doc it was missing that CLAUDE.md's old copy had (pr-review-gate.md).

Classified all 22 docs/*.md files first via a haiku subagent (mara's
suggestion) on two axes -- proposed grouping and operator-vs-
implementation focus -- before finalizing the taxonomy; spot-checked
the report and found internal inconsistencies (its classification
table disagreed with its own summary section for a few files), so this
taxonomy is my original proposal + the one correction mara gave
directly, not a blind application of the subagent's table. The
operator-focus data it gathered is still useful for a follow-up
content pass (docs skewing 'mixed' rather than pure operator-facing),
not addressed in this PR -- structure only.

nix fmt clean, both pre-push lints clean.
2026-09-02 01:55:37 +02:00
atlas
37f3c63eeb feat(#3124): converge the hive onto the agent set the swarm declares
The deploy event is a nudge with no second path: core NATS is
at-most-once, so a hive that was down when the controller published
simply never learns that an agent is meant to exist here. This adds the
repair path — one boot-time DAG node that reads this hive's own key in
the `hive-wanted` bucket and converges the agents it names.

Two semantics settled on the issue thread, and both are places where a
plausible implementation is the wrong one:

- **Absence is not a deletion order.** No bucket, no key, or an agent
  the value does not name all mean the controller has said nothing.
  Swarm-side lifecycle does not yet cover agents that predate it, so
  "converge to exactly this set" would tear down every agent the swarm
  has not adopted. `plan` only ever inspects the agents a declaration
  names.
- **An unrecognised state is inert.** `AgentState` is an open enum: a
  value this build cannot read deserialises into `Unrecognised` and is
  left alone. A closed enum would force "not `Up`" onto a state like
  `paused`, so a controller that learned a new value would take agents
  down on every hive not yet updated.

Divergence is measured against the hive's **stored power intent**, not
the container's observed running state — an agent that is down while its
intent says `Up` is already the boot reconcile's work, and a loop reading
`is_running` would insert a start DAG behind that reconcile's back on
every boot. A hive that already agrees with its declaration queues
nothing at all.

`queue_first_deploy` is extracted from the deploy-event path rather than
open-coded here, for the power-intent seed: without it `first_deploy`'s
tail `Reconcile` seeds `Wanted` from a container that exists but has not
started yet, which locks the agent to `Offline` on its first reconcile.

The read is authorised as-is: `store.get` takes async-nats' direct-get
arm (the KV bucket is created with `allow_direct`), which is exactly the
`$JS.API.DIRECT.GET.KV_hive-wanted.$KV.hive-wanted.<hive>` subject
`swarm-nats-auth` grants a hive. The fallback subject is not granted, and
a refused NATS request surfaces as a timeout rather than an error.

Nothing writes the bucket yet — the controller-side writer is the other
half of #3124, so this does not close it.
2026-09-01 13:07:54 +02:00
atlas
bfae9aa51a hive-c0re: a swarm deploy for an unknown agent provisions it
Per mara on the PR: the issue is about a *new* agent, there is no
approval because the operator clicked create at swarm level, and most of
what a hive does on create is already done by the controller.

`spawn_nodes` splits out of `spawn` the way `rebuild_nodes` already
splits out of `rebuild`: two callers want the same four nodes and
disagree only about what closes them. `first_deploy` is that subgraph
with no approval tail, and the absence is the point — that tail exists
because an operator used to approve the spawn at the hive, and asking
again after they clicked create at swarm level asks the same person the
same question twice.

The handler's predicate is "does a container exist", not "is one
running". `agents_for_meta_listing` is `nixos-container list`, so a
stopped agent still counts. `Coordinator::list_agents` looks like the
right check and is the registered-MCP-socket set — a stopped agent is
absent from it, and this would then try to create over an existing
container.

Enumeration failure drops the request rather than guessing: without the
list this cannot tell first deploy from rebuild, and guessing "new" is
the destructive direction.

Still missing, and the reason this is not the whole change: the hive
seeds its own config repo with `git init` instead of cloning the one the
controller already created.
2026-08-31 00:17:41 +02:00
damocles
e44ea9d8d4 swarm-queue-based lifecycle notices, replacing push_todo(MANAGER_AGENT) 2026-08-24 14:34:37 +02:00
atlas
d2a550e685 feat(#3255): hives stop owning the knowledge webhook, and clean up their own
A webhook has exactly one target URL, so every hive registering one
against the shared internal/knowledge repository was last-writer-wins
rather than idempotent: all but the most recent silently stopped
receiving deliveries. The swarm controller holds the single registration
and now addresses an event to each hive over the queue instead.

This is a migration, not a deletion. Not registering any more fixes
nothing on a hive that has already run — the hook it created persists on
the forge, so the contention would survive on exactly the deployments
that have it while fresh installs looked fixed. The hive that created a
hook removes it.

It removes only its OWN, matched on the full URL rather than the
/webhook/knowledge suffix. A hook with that suffix and a different base
belongs to another hive, possibly one not yet upgraded, and deleting it
would break that hive's knowledge sync until it caught up. Reaping a
neighbour's registration is the behaviour being removed here; doing it
while fixing it would only invert the direction.

The predecessor did reap by suffix, to clear loopback hooks left by an
older single-hive layout. That was safe when a hive was alone on its
forge and is not safe now. The hive-side registrars also acted as reapers
of hooks under their own path, which is why the swarm hook lives under
/webhook/forge/; removing this registrar removes that reaper too.
Intended, and stated because no reviewer would infer it from the diff.

The receive endpoint goes with it. A live HMAC-verified
/webhook/knowledge that nothing can legitimately reach would tell the
next reader that this is how a hive learns about knowledge changes.

Docs move in the same commit: docs/swarm/README.md said two hooks exist
per swarm-wide repo and neither should be deleted, which is now true for
agent-configs and wrong for internal/knowledge — a half-correct
description being worse than an uncorrected one.
2026-08-19 21:05:52 +02:00
damocles
bc40947550 job_queue: trim graph_snapshot's doc comment back under the 30-line comment-block cap 2026-08-16 17:19:04 +02:00
damocles
eae04ac2c5 address review: switch hive-c0re to hive-jobq-wire's shared parse_states/filter_nodes_by_state, note the generic shape in endpoint docs 2026-08-16 16:59:54 +02:00
iris
37161cd136 dashboard: replace <hive-jobq-graph> with a shared Preact component
Ports the shadow-DOM <hive-jobq-graph> custom element
(frontend/packages/shared/src/jobq-graph/) to a Preact component
(JobqGraph.js) shared by the dashboard and swarm-ui, per hyperhive#3310.

- JobqGraph.js: written with plain h() calls (no JSX) so the same file
  compiles unmodified under both the dashboard's text-loader CSS config
  and swarm-ui's JSX config. Exports `JobqGraph` for JSX use and
  `mountJobqGraph(container, props)` for the dashboard's non-JSX
  imperative mount, returning a `{refresh(), update()}` handle matching
  the old custom element's public surface. Same rendering contract as
  before: indented state tree, payload.label verbatim, payload.data as
  a generic key/value list, "waits on: <label>" text for Node-kind deps,
  per-state filter checkboxes, optional cancel button.
- jobq-graph.css: light-DOM adaptation of the old shadow-scoped
  stylesheet (:host -> .jg-root, otherwise unchanged).
- dashboard/src/builds.js: local mountJobqGraph() renamed to
  mountRebuildQueue() to avoid colliding with the newly-imported shared
  mountJobqGraph; cancel handling is now a plain onCancel callback
  instead of a DOM CustomEvent listener (no shadow boundary to cross
  anymore).
- dashboard + shared package.json: added preact as a dependency (matches
  swarm-ui's existing pin, 10.29.8) - the dashboard was a vanilla-JS MPA
  with no Preact/JSX pipeline before this.
- Removed the old hive-jobq-graph.js/.css entirely (confirmed via grep
  it had exactly one consumer, dashboard/src/builds.js, so this is a
  clean swap, not parallel maintenance of two implementations).
- Updated stale doc-comment references to the old element name in
  builds.html, tabs.js, swarm.js, docs/web-ui/dashboard.md, and
  hive-c0re/src/job_queue/mod.rs.

Verified: npm run build (whole frontend workspace) and npm run
typecheck (swarm-ui) both clean; cargo build/clippy/test -p hive-c0re
all clean (331 tests, 0 failures); headless-chromium screenshot of
/builds.html against a mock GET /api/jobq/graph payload confirms full
visual/behavioral parity with the old custom element (tree, filter
checkboxes, cancel buttons, error text, waits-on line, data list, live
build log panel).

This covers the dashboard-replacement half of hyperhive#3310 only. The
swarm-ui half (rendering the CreateAgent DAG on the agent-creation page)
is downstream of hyperhive#3306/#3124 landing - no swarm-ui page exists
yet to mount it in.
2026-08-16 15:18:25 +02:00
damocles
f367fa518e job_queue: submit boot-time forge/matrix/webhook/knowledge sweeps as DAG nodes 2026-08-14 02:23:05 +02:00
atlas
e15c499a31 fix(#3245): resolve the remaining intra-doc links in hive-c0re
Takes the crate from 26 rustdoc warnings to 1, on top of the ten in the
previous commit.

argus's review findings:
- agent_sockets.rs: [`write`] was still ambiguous (function vs macro).
  The previous change narrowed the qualifier and left the ambiguity;
  [`write()`] is what resolves it.
- forge/users.rs <hex> and stats/container_stats.rs <name>: unclosed
  HTML tags in prose, now backticked.

The rest of the crate, so the count actually reaches zero:
- job_queue/mod.rs: Queue::graph_snapshot -> JobQueue::graph_snapshot
  (there is no Queue type), and super::scheduler -> scheduler (mod.rs
  *is* job_queue, so super:: pointed outside it)
- job_queue/resource.rs: NodeKind -> super::model::NodeKind
- matrix.rs: password_path(name) -> password_path; and
  forge::provision_user_token -> crate::forge::provision_user_token.
  Note the path has no `users` segment: forge/mod.rs declares `mod
  users` private and re-exports it, so the canonical path comes from the
  re-export rather than the directory tree.
- socket_server/lifecycle_handlers.rs: InfraContainer ->
  hive_priv_sock::InfraContainer
- stats/otel_metrics.rs: crate::meta::otel_config is a private fn no
  path can name from another module, so it becomes prose
- main.rs: redundant explicit link target dropped

coordinator.rs:405 (CrashWatchGuard) is deliberately untouched: #3244
deletes that doc block, so fixing it here would conflict with an open PR
and repair a symbol that is about to stop existing.
2026-08-14 00:25:35 +02:00
bitburner
517acc9f33 fix(#3245): resolve broken intra-doc links in hive-c0re
Remove or fix broken documentation links that accumulate silently:
- container_view.rs: HiveEnv reference
- forge/mod.rs: READY_TIMEOUT and webhook handler links
- workers/knowledge.rs: webhook handler link
- job_queue/model.rs: Claim::deps and WireNode::data references
- stats/hive_stats.rs: read_skill_breakdown reference
- stores/audit_log.rs: global() reference
- workers/agent_sockets.rs: ambiguous agent_sockets::write reference
- coordinator.rs: systemd.services.<harness> formatting
- resource_limits.rs: ambiguous write/read references

Some broken links were to deleted functions/types; these are replaced
with prose descriptions. Others referenced items outside this crate or
were private; these are replaced with plain text references or qualified
paths as appropriate.

Fixes: #3245
2026-08-14 00:25:35 +02:00
atlas
6338939657 refactor(#2916): destroy submits a DAG instead of an imperative teardown
Destroy was a straight-line async fn with no queue node behind it, so
nothing in the graph could answer "is this container going down on
purpose?". That gap is why an imperative crash-watch suppression guard
existed: an RAII handle held for the operation's duration, a second way
to say what every other lifecycle op already says through its node.

Reuse the existing Stop node rather than teaching a new node to stop
things:

    Stop -> DestroyContainer -> (PurgeState) -> DestroyBookkeeping

Stop already declares takes_container_down honestly, so the suppression
is now derived from the graph like every other op's. It also turns the
precondition into an edge: DestroyContainer runs only under a completed
Stop, so it operates on an already-stopped container and carries
takes_container_down = false permanently. A container still alive at
that point is a real bug and stays loud instead of being absorbed by a
flag -- which matters because a wrong true silently swallows a crash
while a wrong false only costs a spurious event.

Removes suppress_crash_watch, CrashWatchSuppression, crash_suppressed,
crash_watch_suppressed and NO_NODE_LABEL. The migration call sites went
with the obsolete startup migrations, so destroy was the last caller and
intent now has exactly one home.

destroy() becomes a submit-and-return, matching every sibling endpoint
(rebuild, kill, restart, start, pause, resume) -- it was the only
lifecycle op that awaited its work. The container rescan moves into the
bookkeeping tail, so ContainerRemoved now arrives after the 200 rather
than before it.

Also drops an orphaned doc-comment in coordinator.rs: two stacked blocks
where only the second described crash_suppressed, the first documenting
a field that no longer exists. Removing the field would have re-pointed
it at recent_transient.
2026-08-14 00:24:45 +02:00
damocles
1b72ed56ff todos: reopen an acked keyed row when the caller says so 2026-08-13 12:55:20 +02:00
atlas
272944b98f docs: delete two false cost claims about the boot sweep's drains
Both said a whole-hive graceful stop costs ONE `GRACEFUL_STOP_TIMEOUT`
in total because drains overlap. That is only true for a power op. In a
rebuild subtree the brace holds the build slot across the whole subtree,
drain included, so the boot sweep's per-agent drains serialise and the
sweep costs one timeout per wave of `buildSlots`.

Deleted rather than corrected. The right cost statement depends on an
operator knob and belongs in docs/coordinator.md if it belongs anywhere;
a comment that has to hedge about a config value is the kind that goes
stale silently. A comment saying nothing beats one that lies.
2026-08-13 12:46:06 +02:00
iris
861a1f8f26 job_queue: filter jobq graph snapshot by per-node state
graph_snapshot previously filtered which whole roots got projected
based on the root node's own state, so a group root that was still
Running but had already-Done internal steps couldn't be filtered
down to just its live nodes, and a filtered-out root hid its entire
subtree even when a descendant still matched.

Apply the states filter after GraphWire::wire_snapshot instead, over
every node in the flattened tree, not just roots. The jobq-graph
client already handles an orphaned node (parent filtered out) by
promoting it to a rendered root, so this is safe on the client side
with no changes needed there.

Fixes hyperhive#3210
2026-08-12 20:55:25 +02:00
damocles
2bf0d73c8b hive-c0re: drop needless async from pause_many (clippy pedantic) 2026-08-11 23:47:09 +02:00
damocles
20a7a21053 hive-c0re/hivectl/hive-agent: pause as a job-queue DAG node (closes #3056) 2026-08-11 23:47:09 +02:00
damocles
138f6b6c10 hive-sh4re: split manager-socket constants + HelperEvent into their own topic module 2026-08-10 23:05:18 +02:00
iris
1e13b88c8c jobq graph: state filter on /api/jobq/graph + multi-select checkboxes
GET /api/jobq/graph gains a states query param (comma-separated
hive_jobq::State names): narrows the served root groups to the named
states, keeping a group whole (filtering by a root's own state, which
is already its subtree's rolled-up answer). Absent, empty, or fully
unrecognised is the identity filter, matching prior behaviour.

hive-jobq-graph.js gains a row of per-state checkboxes above the tree,
re-fetching the endpoint with the selection on toggle. Default
selection hides Done and Skipped.

Server-side filtering (not client-side hiding) so hive-jobq-graph-update's
node list, and everything downstream of it in builds.js (count pill,
live-log panel), only ever sees what's actually shown.
2026-08-10 23:05:03 +02:00
atlas
ddc017f01b wip(#3001): convert tests off the container id; drop Source + insert_group
The last of the DAG-container removal. `tests.rs` navigated by the id
`submit` returned, so removing the container removed the tests' way of
finding what they inserted; they name the roots they assert on now, which
is the same handle production uses.

Three findings the port surfaced, each a behaviour change rather than a
test fix:

- Cancelling a rebuild's head no longer drops the job. `Reconcile`'s edge
  accepts a skipped brace, and a cancel-cascade skips rather than cancels,
  so the tail stays claimable. Dropping a job means cancelling every id the
  insert returned.
- A directly-cancelled group root reads terminal while a spared tail still
  runs; the cancel used to land on a node above it, which rolled up
  Finishing instead.
- "One DAG per hive-wide op" is not expressible without a container. The
  three tests asserting it now assert that every named root is top-level,
  which is what makes the per-agent subgraphs concurrent.

Deletes two tests: one asserted only that two containers get distinct ids,
the other re-ran an existing case under a second name.

`Source`, `insert_group` and the stop path's `reason` string went dead with
the container and are removed with it.
2026-08-04 19:57:32 +02:00
atlas
aef7ead0bc wip(#3001): delete NodeKind::Dag, the container this issue is about
The variant, its label, its agent-accessor arm and its no-op executor arm are
gone, along with the module prose describing a job as "a single container node
whose subtree is the work". A job is now just its nodes: a template declares
them and names the roots it wants back.

`dag_of` becomes `root_of`. It always wrapped the graph's `root_of` and still
returns the same thing, but the old name asserted a concept that no longer
exists — with no container, the parent chain ends at whichever root the template
declared, so the honest question is "which root owns this node", not "which DAG
is this in".

One comment kept its old wording on purpose: `visible_roots` explains that the
projection it replaced keyed on the container kind rather than selecting
structurally. That is a statement about the past and stays true; it now says
"the since-removed container kind" rather than naming a type that is not there
to look up.
2026-08-04 19:57:32 +02:00
atlas
0523b4f7de wip(#3001): convert the last submit call sites; the binary compiles again
`server.rs`'s five sites move to `power::{stop,start,restart}_many` and direct
template inserts. `submit_single` routes through the `*_many` builders with a
one-element slice rather than keeping a parallel single-target shape.

`templates::rebuild` and `templates::reparent` now return the guids of the roots
they declare, so a caller that has to wait on them can name them; previously
only the void-returning form existed and every caller got an empty id list.

Two comments corrected while converting, both contradicted by the code they sit
above:

* `templates::rebuild` said its tail is edged onto "(MetaSync, Prebuild,
  Reconcile)" and that "Prebuild's roll-up carries the subtree" — the brace has
  been the middle root since the AgentWindow change.
* the restart handler described the per-agent shape as starting with SetWanted,
  while `restart_chain`'s own doc says a restart never rewrites `wanted` — that
  is the difference between restart and stop/start.

Error handling is no longer swallowed: a failed insert becomes a reported error
rather than a silently-absent id.
2026-08-04 19:57:32 +02:00
atlas
02e916feee wip(#3001): power chains name their group roots
The `*_many` entry points returned `insert_job`'s result while their closures
ended in `Vec::new()` — naming nothing, so the returned id list was always
empty. `queued_dags` would have shipped `Some([])` and hivectl's wait loop would
have had nothing to poll. Silent: it compiles, the op still runs, and no test in
isolation looks.

Each `*_chain` now returns its group root's guid and the `*_nodes` collectors
gather them, so the ids a caller gets back are the roots it can actually wait on.

`start_chain` returns *four* in the stale branch, not one: `rebuild_nodes`
chains its roots behind `SetWanted` with `after_ok` rather than nesting them
under it, so `SetWanted` rolls up only itself. Naming it alone would have
reported the start complete while the rebuild was still running — the same
under-reporting bug one level down.
2026-08-04 19:57:32 +02:00
atlas
fe52037b0d wip(#3001): rename insert -> insert_job per mara's 50056 2026-08-04 19:57:32 +02:00
atlas
102ebdc03d wip(#3001): 9 of 10 test helpers off the container id; sweep stale prose
The helpers keep shrinking the same way: resolve_id goes, 'n.id != root'
goes (the container was the only non-work node), root_of goes, and the
container-parent normalisation goes because a group root now genuinely
has parent = None. pending_kinds_filtered drops from a four-clause
multi-line filter to one line.

Also removed a doc block my earlier edit had orphaned above the renamed
helper, and swept 'under `dag`' / '`submit` returns' out of the prose.

state_of stays untouched: it reads a roll-up, which is the same question
as hivectl's queued_dags.
2026-08-04 19:57:32 +02:00
atlas
be4763678b wip(#3001): test helpers off the container id
submit() -> insert() in tests, and the two shape walkers lose their dag
param: with no container there is no per-DAG root to filter on, nothing
to exclude (every node is real work now), and a group root genuinely has
parent = None, so the container-parent normalisation goes too. Each test
builds a fresh JobQueue, so "the DAG" is "the graph".

20 errors remain, all in tests.rs, and they are the point: changing the
helper's type from u64 to () made every site that consumed the container
id light up as `expected u64, found ()`. A type error is an exhaustive
grep -- ten helpers take a dag id, not the three I had measured.

state_of(q, dag_id) is not mechanical: it read the DAG's ROLLED-UP state,
which was the container node's own. That makes it the second consumer of
the container-as-roll-up-point, alongside hivectl's queued_dags poll.
Both want the same answer, so it waits on the same ruling.
2026-08-04 19:57:32 +02:00
atlas
f04a0cee92 wip(#3001): convert remaining unblocked call sites; sweep docs
21 of 28 non-test call sites now insert directly. power.rs compiles.
The only remaining errors are server.rs's 5, which are blocked: those
sites feed the returned id into HostResponse::queued -> `queued_dags`,
a wire field hivectl polls via QueueDag. Removing the container without
answering that breaks hivectl's wait/progress loop; asked on the issue.

Also swept the deleted symbol out of prose, not just code:
- docs/coordinator.md: "the submit layer (job_queue/submit.rs)" ->
  the power layer (job_queue/power.rs), and "submits" -> "inserts".
- templates.rs module doc: points at super::power for the power ops.
- lifecycle_ops.rs module doc: says which path each op takes now.
- mod.rs's insert_group comment restated the open issue verbatim
  ("a DAG is addressed by its container node, which submit inserts
  itself"). Replaced with what is actually true for that path.

Dashboard behaviour deltas worth review: insert failures are now
logged per agent instead of swallowed, and UPDATE-ALL emits one queue
snapshot after the loop rather than one per agent.
2026-08-04 19:57:32 +02:00
atlas
7c0d9d2379 wip(#3001): remove submit layer, rescue power ops into job_queue/power.rs
TREE IS RED ON PURPOSE — there is no compiling intermediate between
deleting submit and converting every caller. Checkpoint commit so the
work is durable; do not "fix" it by restoring submit.

Done:
- JobQueue::submit -> JobQueue::insert (no source/reason/container;
  returns the ids insert_job names).
- submit.rs deleted. Its 6 pure chain builders + 3 async *_many
  gatherers were NOT wrapper code and are rescued into
  job_queue/power.rs (templates.rs documents power ops as living
  outside it, because their shape needs a live is_running read).
- Converted: meta_inputs 1, topology 2, permissions 3, auto_update 2,
  actions 3, lifecycle_handlers 3.
- Dropped source/reason at every converted site: nothing ever read
  NodeKind::Dag's fields (only `{ .. }` matches exist), so they are
  write-only. Dead reason-only locals deleted; the boot sweep's summary
  became a tracing::info! rather than being lost.

Remaining: dashboard/lifecycle_ops 7, server.rs 7, and the test suite —
tests.rs has its own submit() helper whose u64 return is used as the
handle to navigate the inserted DAG, so those need a different way to
find nodes, not a mechanical port.
2026-08-04 19:57:32 +02:00
atlas
0a14055a33 docs(#3034): give the brace one home instead of five
Measured after mara's "2/3 of this is docs, most of it duplicated": 243 of 380
added .rs lines were comments. The brace rationale was written out in full in
`model.rs`, the `templates.rs` module header, `quiesce`, `rebuild_subtree` and
`docs/coordinator.md` — five copies of one argument, which is why four docs
needed correcting earlier in this branch. Correcting every copy preserves the
thing that made them go stale.

`docs/coordinator.md` (_Braces_) is now the single home. The rest state what a
node *is* and point there. Also drops the per-operation DAG diagram from the
`templates.rs` header, which the same doc already carries, and cuts
`rebuild_subtree`'s node-by-node walkthrough down to the three choices a reader
would otherwise undo — the code below it is the source of truth for the shape.

Comments -59 lines, no behaviour change, 317 tests unchanged.
2026-08-04 16:38:30 +02:00
atlas
c2eafa7548 refactor(#3034): one quiesce builder, shared by the rebuild and the stop chain
The `Signal` -> `Drain` pair was built in three places, in three different
shapes: siblings under `SetWanted` in `stop_chain`, `Drain` nested under a
lease-holding `Signal` in `restart_chain`, and — as of this branch — a hybrid in
`rebuild_subtree` that was brace-held like the first and nested like the second.

`templates::quiesce(builder, agent, brace)` is now the one definition, returning
the `Drain` handle a caller edges its stop onto. Both nodes hang off the brace as
dep-ordered siblings and declare nothing, borrowing the lease it already holds.

That also fixes an inconsistency this branch introduced: the PR argued that a
brace makes nesting unnecessary and used it to flatten `StopForUpdate` off
`Signal`, then left `Drain` nested under `Signal` two lines away. Nesting is only
load-bearing where `Signal` is itself the lease holder.

`stop_chain`'s pair loses its own `Agent` declaration as a result — `SetWanted`
holds the lease for the subtree, so those were redundant re-entrant borrows.

`restart_chain` is left alone and says why in place: it has no brace, so `Signal`
holds the lease and the nesting under it is what keeps the grant continuous.
Giving it one would unify all three sites at the cost of an extra no-op node on
every graceful restart, which an operator would see — not something to change as
a side effect of a rebuild-shape PR.
2026-08-04 15:04:26 +02:00