Commit graph hyperhive/hive-c0re/src
Author SHA1 Message Date
atlas
fe17b5f8a7 feat(matrix): auto-create a hive chat room as a child of the hive Space
The hive Space was created empty — joining it surfaced no rooms because
Matrix doesn't auto-join a Space's children. Provision a default
"hive-chat" room on the matrix sweep, wire it bidirectionally to the
Space (m.space.child on the Space, m.space.parent on the room), and
invite @hive + every agent. The room uses a restricted join rule
allowing any Space member to join, so the operator (a Space member) can
join it from the Space hierarchy without an explicit invite.

Idempotent, mirroring ensure_hive_space: persisted chat-room-id wins,
else rediscover a non-space room named hive-chat, else createRoom. The
space-child link is re-applied each sweep (idempotent PUT) so a
recovered room reconverges its hierarchy link. Room id persisted to
matrix/chat-room-id (0600, survives destroy --purge).
2026-06-06 07:59:24 +02:00
damocles
6e39515669 feat: type-scope events vacuum to prune only stream rows (14d) + drop turn-stats vacuum 2026-06-06 07:57:27 +02:00
iris
dbb9f2a787 stats(p3): normalise cpu% against host CPU count, not process affinity
Per review: available_parallelism() respects the hive-core process's
CPU affinity, so if it's ever affinity-pinned the denominator would
under-count and inflate cpu_pct. Read the online host CPUs from
/sys/devices/system/cpu/online instead (fall back to the process count,
then 1) so the 'percent of total host CPU' definition holds regardless.
2026-06-05 23:06:33 +02:00
iris
03ea6d1bda feat(stats): per-container cpu/mem load (#1424 p3)
C0NT41N3R L04D on the SYST3M tab + GET /api/container-resources.

Backend (hive-c0re/src/container_stats.rs): reads cgroup v2 cpu.stat +
memory.{current,peak,max} for each running agent machine
(machine-h\x2d<name>.scope under machine.slice), read-only/world-
readable so no hive-priv. CPU is a two-sample (~200ms) host-normalised
percentage; one shared sleep covers all agents. Skips agents whose
scope dir is absent (= not running). Network omitted: agents share the
host netns, no per-container counter.

Frontend: a polled C0NT41N3R L04D table on SYST3M (agent / cpu / mem /
peak / limit with meter bars), reusing the ST4TS table style. Polls
/api/container-resources every 5s only while the tab is active.

Backend reviewed-in-principle by damocles (path escaping + cpu delta
math); ping for the on-host sign-off.
2026-06-05 23:06:33 +02:00
damocles
14c7b0d406 feat: group host-side /var/lib/hyperhive state into db/ forge/ matrix/ run/ subdirs with startup migration 2026-06-05 23:01:47 +02:00
iris
8e74dda4c5 feat(dashboard): ST4TS tab — hive-wide stats view
Builds the dashboard UI on the /api/stats-hive endpoint: a new ST4TS
tab showing swarm totals (active agents, turns, tokens, labelled est
cost), a busiest-first per-agent table, and a model-mix bar list.
Plain tables/CSS bars — no chart lib in the dashboard bundle; per-agent
trend charts stay on each agent's own /stats page. Fetched on tab
activation + window change (pull-only, no SSE).

Also adds conn.busy_timeout(500ms) in hive_stats read_agent per review:
turn_stats is rollback-journal, so a read landing mid-INSERT would hit
SQLITE_BUSY and silently drop that active agent — wait the write out.
2026-06-05 22:51:11 +02:00
iris
447a84e8a6 feat(c0re): /api/stats-hive — hive-wide turn-stats rollup
#1424 P2 backend. New hive_stats module aggregates every agent's
turn-stats.sqlite read-only (reusing Coordinator::kept_state_names +
agent_harness_dir, skipping missing/unreadable dbs) into swarm totals,
a busiest-first per-agent rollup, swarm model mix, and a labelled USD
cost estimate (rough model->price table; can move to a nix option
later). Exposed as GET /api/stats-hive?window=. Dashboard UI follows.
2026-06-05 22:51:11 +02:00
damocles
07e8f442cb fix: hivectl choom passes the harness --settings/--mcp-config/--system-prompt-file so claude gets settings + tools + persona 2026-06-05 22:48:34 +02:00
damocles
c59a0de01e fix: hivectl choom enters as the agent user from the state dir so claude gets creds + session 2026-06-05 21:41:50 +02:00
damocles
ad8a3a8cbf feat: hidden hivectl markdown-docs subcommand emitting the cli reference as commonmark 2026-06-05 21:07:51 +02:00
damocles
e7a68472fd fix: reject bare --room value in hivectl matrix invite with a clear hint 2026-06-05 20:51:31 +02:00
damocles
e8d5eee659 feat: hivectl matrix invite — add a user to the hive Space or a room (closes #1402) 2026-06-05 20:51:31 +02:00
damocles
782438be89 fix(#1374): make /shared writable by all agents (sticky world-writable) 2026-06-05 18:46:16 +02:00
damocles
fb726197ea fix(#1375): clean up pedantic warnings and re-enable -D warnings without pedantic bypass 2026-06-05 16:55:09 +02:00
atlas
734fe88858 fix(ci): unblock nix flake check after clippy 0.1.95 bump (#1368)
The nixpkgs bump to clippy 0.1.95 / cargo 1.95.0 added + strengthened a
large batch of lints. CI denied ALL warnings (`-D warnings`) against the
`pedantic = warn` workspace lint, so the bump hard-failed `nix flake
check` workspace-wide with zero code changes — and would recur on every
future clippy bump.

Posture fix (the durable part): CI now runs
`-D warnings -A clippy::pedantic`, so the default/correctness/style lints
stay a hard gate while the "extra, opinionated" pedantic group is
advisory only (still `warn` for local `cargo clippy` via the workspace
lints table, just non-blocking in CI). `-A` rather than `-W` so the
group drop doesn't re-enable the specific pedantic lints the workspace
allows (e.g. `must_use_candidate`).

Also fixes the genuine DEFAULT/STYLE lints the bump surfaced across the
workspace (doc_lazy_continuation, collapsible_if, ptr_arg,
match_like_matches_macro, …) via `cargo clippy --fix` + manual stragglers
(`too_many_arguments` #[allow] on the host-config constructors), and
three tests that had rotted while the CI runner was offline (#1221):
- topology::top_level_agents_in_multi_root — hardcoded unsorted expected
- rebuild_queue::depends_on_evicted_dep_counts_as_resolved — needs
  MAX_HISTORY_PER_KIND newer terminals to evict, not one
- coordinator::agent_paths doctest — illustrative pseudo-code, now `ignore`

Validated: clippy + formatting + cargo-test checks all pass.
2026-06-05 15:32:07 +02:00
damocles
f201f04d4e fix(#1329): restart hive-matrix-daemon after token write so new credential is picked up immediately 2026-06-05 15:30:21 +02:00
atlas
d043b1ed4e fix: rediscover hive Space by hardcoded name, not a room alias
Per mara's review: drop the #hive:<server> room alias (special chars) and
rediscover the canonical Space by its hardcoded plain name instead.

ensure_hive_space dedup is now:
  1. room-id file present -> reuse it
  2. else scan the admin's joined rooms for the m.space named HIVE_SPACE_NAME
     ('hive') and adopt the first match (re-persisting the file) -> recovers
     the existing space after a state wipe instead of creating a duplicate
  3. else createRoom (plain name, no alias)

find_space_by_name walks /joined_rooms and checks each room's m.room.create
type == m.space and m.room.name == 'hive'. No alias, no special-char anchor.
server_name is no longer needed by ensure_hive_space (dropped the param).
2026-06-05 13:17:16 +02:00
atlas
134a40a5e2 fix: anchor hive Space to canonical #hive alias for dedup
ensure_hive_space relied solely on the persisted room-id file. If that file
is ever lost (a full /var/lib/hyperhive wipe), the next sweep blind-creates a
new m.space — the homeserver keeps the old one, so duplicate hive spaces
accumulate (observed: multiple 'hive'/'pr1ma' rooms on the live instance).

Anchor the Space to a stable canonical alias #hive:<server>:
- fast path (room-id file present): reuse it and heal the alias mapping so
  it keeps pointing at the canonical room
- no file: resolve #hive:<server> and adopt the existing room if present,
  re-persisting the file — recovers the space after a wipe instead of
  duplicating it
- only create (with the alias) when neither yields a room

server_name is now discovered before ensure_hive_space in ensure_all and
threaded through (the alias needs it). Existing deployments heal the alias
onto their current space on the next sweep; no new room is created when the
file is present.
2026-06-05 13:17:16 +02:00
atlas
20c5039156 style: treefmt main — fix CI formatting check
Full `nix flake check` (CI) runs the treefmt formatting derivation. While the
hive-ci runner was offline (#1221), PRs merged without it, leaving 5 files
unformatted: hive-ag3nt/src/web_ui.rs, hive-c0re/src/bin/hivectl.rs,
hive-c0re/src/knowledge.rs, hive-c0re/src/matrix.rs,
hive-forge/src/verbs/attachment_get.rs. `nix fmt` output, pure formatting.
2026-06-05 13:09:56 +02:00
iris
2080fd3866 feat: live SSE updates for P3RM1SS10NS tab (capabilities + tool groups)
Add CapabilitiesChanged and ToolGroupsChanged DashboardEvent variants
so the P3RM1SS10NS tab reflects perm changes without the operator
navigating away and back.

Backend:
- DashboardEvent::CapabilitiesChanged { seq, caps, descriptions,
  assignments } — same payload shape as GET /api/capabilities
- DashboardEvent::ToolGroupsChanged { seq, groups, descriptions,
  assignments } — same payload shape as GET /api/tool-groups
- Coordinator::emit_capabilities_snapshot() and
  emit_tool_groups_snapshot() — read from the JSON files and broadcast
- rebuild_queue.rs PermChange worker: emit after each successful
  commit_capabilities / commit_tool_groups call

Frontend:
- applyCapabilitiesChanged(ev): calls renderCapabilities(root, ev)
- applyToolGroupsChanged(ev): calls renderToolGroups(root, ev)
- Both registered in MUTATION_HANDLERS
- activateTab comment updated (SSE now covers perm changes)

Docs: dashboard.md and CLAUDE.md updated.

This completes SSE coverage for all dashboard sections: SW4RM,
Y3R C4LL, SYST3M, SCH3DUL3S/reminders, and P3RM1SS10NS all
derive live updates from /dashboard/stream.
2026-06-05 12:06:32 +02:00
iris
a1e46e2b3d feat: live SSE updates for the SYST3M reminders section
Add RemindersChanged SSE event so the pending-reminders list in the
SYST3M tab updates live without polling.

Backend emission sites (every path that mutates the reminders table):
- agent_server: store_remind (remind MCP call)
- dashboard.rs: post_cancel_reminder, post_retry_reminder
- questions.rs: cancel_loose_end Reminder kind
- reminder_scheduler: after each delivery batch (any_delivered)

Coordinator gets emit_reminders_snapshot() mirroring the existing
emit_schedules_snapshot() pattern: lists PendingReminder rows from the
broker and emits DashboardEvent::RemindersChanged.

Frontend: applyRemindersChanged(ev) calls renderReminders(ev.reminders)
and is registered as reminders_changed in MUTATION_HANDLERS.

Docs: dashboard.md reminders_changed entry; CLAUDE.md file map updated.
2026-06-05 12:06:32 +02:00
iris
68108fe5f8 fix: emit SchedulesChanged from manager-server + approval paths
Agent-triggered schedule mutations (cancel_schedule, fire_schedule_now,
edit_schedule MCP tools) go through manager_server.rs, not the HTTP API
handlers. Approval-resolved SchedulePrompt inserts go through actions.rs.
Neither was emitting SchedulesChanged.

- manager_server.rs: add emit_schedules_snapshot() on Ok in
  handle_cancel_schedule, handle_fire_schedule_now, handle_edit_schedule
- actions.rs: emit_schedules_snapshot() after successful
  run_approval_schedule_prompt (covers request_schedule_prompt approval
  resolving)

Coverage is now complete: every path that writes a scheduled_prompts row
emits the SSE snapshot.
2026-06-05 12:06:32 +02:00
iris
76c4a67b1c feat: live SSE updates for the SCH3DUL3S tab
Add `SchedulesChanged` to the dashboard event channel so the
operator's schedule list updates in real time without requiring a
tab-activation or form-submit refresh.

Backend:
- `dashboard_events.rs`: new `SchedulesChanged { seq, schedules }`
  variant carrying a full `Vec<WireSchedule>` snapshot (same
  snapshot-over-diff rationale as `RebuildQueueChanged`).
- `coordinator.rs`: `emit_schedules_snapshot()` helper — queries the
  scheduled_prompts list, converts to wire shape, broadcasts the event.
- `dashboard.rs`: call `emit_schedules_snapshot()` at the end of each
  operator API handler that mutates a schedule:
  `post_schedule_new`, `post_schedule_fire_now`,
  `patch_schedule`, `post_schedule_cancel`.
- `scheduled_prompts_worker.rs`: call `emit_schedules_snapshot()`
  after each tick that fires schedules, so `last_fired_at_unix`,
  `next_fire_at_unix`, and reaped one-shots surface live.

Frontend:
- `tabs.js`: add `applySchedulesChanged(ev)` — replaces
  `schedulesState` from the snapshot and calls `renderSchedulesList()`.
  Registered in `MUTATION_HANDLERS` as `schedules_changed`.
  Tab-activation re-fetch kept as safety net for approval-path
  inserts and disconnect windows; comment updated to reflect this.

Docs:
- `docs/web-ui/dashboard.md`: document `schedules_changed` event.
- `CLAUDE.md`: add `SchedulesChanged` to the file-map entry.
2026-06-05 12:06:32 +02:00
atlas
b411a81c1a fix: point hivectl agents at the correct host socket path
hivectl's DEFAULT_HOST_SOCKET was /run/hive/host.sock, but hive-c0re binds
/run/hyperhive/host.sock (main.rs default + startup log). The doc comment
even claimed it matched hive-c0re's default — it didn't. So 'hivectl agents
restart' / 'restart-all' failed with 'connect to /run/hive/host.sock: No such
file or directory' unless the operator passed --socket explicitly.

Point the default at /run/hyperhive/host.sock.
2026-06-05 09:55:24 +02:00
atlas
3858740488 refactor: drop speculative prose markers, keep only backtick extraction
Per operator review on the PR: the prose marker list ('new password is:',
' to:', 'changed to:', etc.) was speculative — built from a misread
screenshot, not a real reply. The conduit admin bot always code-spans the
password, so the backtick anchor is the verified, complete format. Removing
the markers leaves a ~15-line function that's honest about what it parses.

If a non-code-span format ever appears, extract_new_password returns None and
the diagnostic logging in admin_room_send_and_poll records the raw body, so a
real format change is visible — far better than a speculative marker silently
mis-parsing it. Tests now cover the live format, symbol passwords, the
code-spanned-user-id error guard, and the no-codespan / non-password cases.
2026-06-05 01:53:05 +02:00
atlas
2d8a68e4d1 fix: extract matrix password from backtick code span (real conduit format)
The live conduit admin-room reply (observed directly in #admins) is:

  Successfully reset the password for user @x:server: `<password>`

The delimiter is ': ' after the user id, NOT ' to:' — so the ' to: ' marker
added earlier never matched the real format, and extract_new_password returned
None for every reply, timing out auto-recovery for every agent.

conduit always renders the password as a backtick code span, so anchor on
that directly: take the content of the first backtick pair when the message
is a password-reset success. Guard against a code-spanned matrix user id in
an error message (a real password has no whitespace and isn't @localpart:server).
Prose markers stay as a fallback for hypothetical non-code-span builds.

Tests cover the exact live format and the code-spanned-user-id error case.
2026-06-05 01:49:31 +02:00
atlas
93d28ce0b1 fix: strip code-span backticks from extracted matrix password
mara observed the tuwunel reply renders the new password as a code span,
so the plain message body carries literal backticks:

  Successfully reset password for @user:server to: `<password>`

The previous extraction stopped at the first whitespace, capturing the
surrounding backticks ("`<password>`") and producing a login string that
doesn't match the password the bot actually set — recovery would still
fail after parsing.

Strip a leading backtick after the marker and stop the token at the first
whitespace OR closing backtick. Generated passwords contain neither, so a
real password is never truncated mid-token. Two regression tests cover the
backtick-wrapped form, including trailing prose after the closing backtick.
2026-06-05 00:47:37 +02:00
atlas
defface0a5 fix: parse tuwunel 'reset password for X to: <pw>' admin reply
extract_new_password had markers for 'is:', 'changed to:', 'reset to:',
'set to:' etc. but none match tuwunel's actual reset-password reply:

  Successfully reset password for @user:server to: <password>

Here the verb ('reset') is not adjacent to 'to:', so every marker missed
and the function returned None for every polled event. The admin-room poll
then ran its full 15s loop without a match and the auto-recovery timed out
on every agent — even though the bot replied correctly and the backward
poll found the message. This is why matrix password auto-recovery kept
looping despite the poll-strategy fix.

Add a generic ' to: ' marker (with surrounding spaces) placed after the
specific verb markers and before the bare 'password:' last resort. Matrix
user ids and server names can't contain ' to: ', so it only ever anchors
on the prose delimiter. Two regression tests cover the exact production
wording.
2026-06-05 00:47:37 +02:00
atlas
54f7c0facc docs: comment on empty event_id fallback in admin-room poll 2026-06-04 20:13:47 +02:00
atlas
e20a311786 tidy: is_whitespace() already covers newline, remove redundant arm 2026-06-04 20:13:47 +02:00
atlas
254fd1f9f1 fix: admin-room poll strategy — backward fetch anchored on sent event_id
The forward-pagination approach (dir=b anchor → dir=f poll) fails in
production: commands sent as @hive time out consistently even though
tuwunel responds. Root cause is likely a pagination-token direction
incompatibility in some tuwunel builds where 'end' from dir=b cannot
be used as 'from' for dir=f.

New strategy: send the command, capture the event_id from the PUT
response, then poll dir=b&limit=20 each tick. Events come back
newest-first; walk until we hit our own event_id, then stop — anything
before that marker arrived after our command. Simpler, avoids stored
tokens entirely.

Also:
- check formatted_body in addition to body (some admin bots put content
  only in HTML formatted_body)
- add more extract_new_password patterns: 'changed to:', 'reset to:',
  'set to:', 'new password:', 'password:' to handle different tuwunel
  version response formats
- add unit tests for new patterns

Fixes #1283.
2026-06-04 20:13:47 +02:00
damocles
ab65245619 feat(#1298): add hivectl choom verb — interactive claude session in agent container 2026-06-04 20:08:04 +02:00
iris
e0ea22f3ee refactor(#1292): remove agent-owned fields from ContainerView
Now that the dashboard fetches agent-owned state directly from
GET /api/dashboard-state (via gateway), hive-c0re no longer needs
to read those fields from disk on the agent's behalf.

Removed from ContainerView:
  ctx_tokens, context_window_tokens, rate_limited,
  extra_links, status_text, status_set_at

Removed from container_view.rs:
  DashboardLink struct, build_nav_links, read_dashboard_links,
  read_status, is_rate_limited, read_last_turn, resolve_ctx_window
  (and the resolve_ctx_window unit tests)

Removed from dashboard.rs:
  GET /api/agent/{name}/links route + get_agent_links handler

dashboard JS (tabs.js):
  Merged rate_limited, ctx-window badge, and status-text rendering
  into the async dashboard-state fetch block. c0re still provides
  needs_login (auth sentinel on host), needs_update, pending_reminders,
  running, deployed_sha, parent — all genuinely host-side fields.
2026-06-04 19:39:06 +02:00
damocles
5f05caee31 fix(queue): rename queued_ids to active_ids, document circular-dep caveat 2026-06-04 17:39:31 +02:00
damocles
7118c5efdd feat(queue): link rebuild queue entries to build log rows for live streaming 2026-06-04 17:39:31 +02:00
damocles
d31a723daf feat(#550): add depends_on to queue entries for explicit dep tracking 2026-06-04 17:39:31 +02:00
atlas
1fa398e99e fix: use correct tuwunel admin room command prefix for reset-password 2026-06-04 16:43:53 +02:00
atlas
85de83bd49 fix: use correct tuwunel admin room command prefix for make-user-admin
tuwunel requires '!admin users <cmd>' — bare 'make-user-admin @user:server'
is not recognised. Update command string and doc comments.
2026-06-04 16:35:59 +02:00
atlas
0f505d5c95 fix: increase admin-room poll timeout from 5s to 15s
tuwunel can take longer than 5 seconds to process admin-room commands
during startup when the homeserver is under load. Bump both the poll
count (5→15) and the timeout message strings to match.
2026-06-04 16:28:47 +02:00
atlas
e06d3e3a8b refactor: extract admin_room_send_and_poll helper — deduplicate poll loop 2026-06-04 16:28:47 +02:00
atlas
2252eac650 fix: replace Synapse admin API in promote_user_to_admin with admin-room command 2026-06-04 16:26:59 +02:00
atlas
74253a9f71 fix: increase admin-room poll timeout from 5s to 15s
tuwunel can take longer than 5 seconds to process admin-room commands
during startup when the homeserver is under load. Bump poll count 5→15
and update timeout message strings to match.
2026-06-04 16:26:04 +02:00
atlas
a56badd004 fix: remove Synapse admin API path — use admin-room reset exclusively (tuwunel only) 2026-06-04 15:16:22 +02:00
atlas
97a78cc5d6 fix: address argus review on matrix admin-room reset — ascii lowercase, tighter markers, unit tests 2026-06-04 15:16:22 +02:00
atlas
b2913bc656 fix: admin-room fallback for matrix password reset when Synapse API absent
tuwunel 1.6.x does not implement the Synapse admin REST API.
reset_user_password() now falls back to the Matrix admin room
(#admins:<server>) when PUT /_synapse/admin/v2/users returns 404:

1. Discover admin room ID via #admins:<server> alias
2. Get current messages end-token (pagination anchor)
3. Send 'reset-password @<localpart>:<server>' as @hive admin user
4. Poll for bot response up to 5 x 1s; extract password from message
5. Persist the new password and return it

The function signature changes from Result<()> to Result<String> so the
caller can use the effective password (which may be server-generated on
the admin-room path) for subsequent login calls.

Closes #1267.
2026-06-04 15:16:22 +02:00
atlas
7e229889a6 fix: address argus review on ensure_claude_dir — narrow to EPERM, fix doc comment 2026-06-04 15:03:12 +02:00
atlas
e553e40577 fix: soft-fail chmod in ensure_claude_dir when dir is agent-owned
hive-agent-user-migrate chowns the bind-mounted claude dir to the agent
user on every container boot. After that, hive-core (a different user)
cannot chmod it — set_permissions fails with EPERM, which was propagated
as an error and caused the rebuild to fail entirely.

Fix: make the chmod best-effort. Newly created dirs (owned by hive-core)
get the 0755 mode set immediately; after the agent-migration chown the
mode is preserved so claude_has_session works correctly. Subsequent calls
that hit the EPERM path just log at DEBUG and continue.
2026-06-04 14:48:08 +02:00
damocles
41eb3f806c refactor: remove hyperhive.role option — there is only one role: agent 2026-06-04 14:31:44 +02:00
atlas
7022cd3826 fix: split WriteAgentStateFile into WriteAgentForgeToken + WriteAgentMatrixToken
Addresses mara's review: each credential type gets its own PrivRequest
variant, making the exact priv surface visible in the wire protocol.
No runtime filename dispatch — the operation name is the gate.

- WriteAgentForgeToken { agent_name, token } → state/forge-token
- WriteAgentMatrixToken { agent_name, token } → state/matrix-token
- priv_client: two typed fns (write_agent_forge_token, write_agent_matrix_token)
- forge.rs: split mint_and_persist_token into mint_and_persist_agent_token
  (priv) + mint_and_persist_core_token (direct write); drop dead token_path fn
- matrix.rs: call write_agent_matrix_token directly
2026-06-04 14:30:01 +02:00
atlas
89092caba4 fix: restrict WriteAgentStateFile to explicit filename allowlist
Addresses mara's security review: replace validate_state_filename (which
accepted any non-traversal filename) with a tight allowlist containing
only the two known credential filenames: forge-token and matrix-token.

Also addresses argus review feedback:
- drop issue tag from priv_proto.rs doc comment
- add comment explaining the path-detection heuristic in forge.rs
- add note about create_dir_all uid=0 edge case in write_agent_state_file
2026-06-04 14:30:01 +02:00