Replace brittle msg.contains("not found") string matching in
post_schedule_pause / post_schedule_resume with a typed
ScheduleNotFoundOrCancelled error that handlers downcast on directly.
pause() and resume() now return Err(ScheduleNotFoundOrCancelled(id).into())
instead of bail!("schedule {id} not found or is cancelled"); handlers call
e.downcast_ref::<ScheduleNotFoundOrCancelled>().is_some() for the 404 branch,
making the discrimination stable even if the error message wording changes.
Adds pause/resume support for scheduled prompts.
Backend:
- New paused_at_unix column on scheduled_prompts table (added via
ALTER TABLE migration so existing databases are upgraded on first
start). The due-rows index is dropped and recreated to also exclude
paused rows so the worker never fires them while paused.
- Worker's due() query gains AND paused_at_unix IS NULL filter.
- New pause(id) and resume(id) methods on ScheduledPrompts; both are
idempotent and refuse cancelled rows.
- New POST /api/schedules/{id}/pause and /api/schedules/{id}/resume
dashboard endpoints (operator-direct, no approval gate). Both emit
a schedules snapshot on success so the tab updates live.
- WireSchedule gains paused_at_unix: Option<i64> so the frontend can
render the state without an extra fetch.
Frontend:
- Paused rows render with a distinct row class + muted opacity.
- The next-fire cell shows a yellow pause glyph + tooltip with the
paused-since timestamp and the would-have-fired time.
- Actions column: pause/resume toggle button (⏸/▶) beside fire/edit/cancel.
Fire-now is disabled while paused (resume first).
- Sort order: active → paused → cancelled (paused slot keeps schedules
visible without mixing them into the active top section).
- pauseSchedule() / resumeSchedule() async functions POST to the new
endpoints and refresh the table on success.
Both remove_agent() calls now run unconditionally for maximum partial
cleanup, but any I/O error is returned as HTTP 500 instead of silently
200-ing — so the frontend's !resp.ok path fires and the operator sees a
meaningful error rather than the stale row reappearing unchanged.
Also add a clarifying comment on isStale in permissions.js explaining
that containersState is keyed from nixos-container list (which includes
stopped-but-configured containers), so a temporarily-stopped agent is
not treated as stale — only destroyed/renamed agents are absent.
The P3RM1SS10NS tab showed agents that no longer exist in the live
container roster — e.g. an agent named 'root' that was renamed or
destroyed but still had explicit entries in tool-groups.json and/or
capabilities.json. The roster-union behaviour is intentional for
temporarily-stopped agents, but stale entries from renamed/destroyed
agents are confusing.
Backend (dashboard/permissions.rs):
- New DELETE /api/permissions/{agent} handler that bypasses the live-
roster guard (intentionally — that's the point). Calls
tool_groups::remove_agent + capabilities::remove_agent to clear both
JSON files, then emits live SSE snapshots so the tab updates without
a page reload. Format-checks the agent name but does not require it to
be in the containers snapshot.
Frontend (permissions.js):
- renderCapabilities / renderToolGroups now cross-reference agentNames
against containersState (the live roster, already imported). Agents
not in the live roster get an isStale flag.
- Stale rows get a '(not running)' label and a '✕ remove' button that
calls clearStaleAgent() — a new async helper that DELETEs the stale
entry and re-fetches both perm tables.
- Non-stale agents without explicit assignments still get '(default)'.
CSS (dashboard.css):
- .perm-row-stale (reduced opacity), .perm-stale-label (muted small
text), .perm-remove-btn (small red-bordered button) + disabled state.
The per-agent and manager sockets ran two parallel dispatchers with
duplicated lifecycle handlers (agent-side topology-gated, manager-side
ungated) plus a manager-only handler set. Collapse to one parameterized
server in socket_server.rs:
- one serve() + dispatch(req, agent, privileged, coord); start() binds
the per-agent sockets (privileged=false), start_manager() binds the
manager socket (privileged=true).
- each lifecycle/config handler (start/restart/kill/update/init_config/
apply_commit) merges its dual: the topology guard (require_child /
require_new_child) runs only on the !privileged path; init_config
records the requester as parent only when !privileged. restart keeps
the orthogonal, capability-gated + audited infra-container branch.
- the agent-state queries (loose-ends / reminder count + rollup) branch
on privileged: privileged keeps any-target + the "*" hive-wide sweep
(query_agent_state-gated), non-privileged keeps the topology/cap gate.
- the privileged-only verbs (schedules / meta-inputs / get_logs) plus
the submit/schedule/watchdog helpers move into socket_server; they are
reached via dispatch_privileged_only(), which rejects the whole group
on a non-privileged socket.
- delete manager_server.rs; repoint refs; merge the test modules.
No behavior change: the topology guard still applies on every
non-privileged lifecycle call, the privileged socket still acts on any
agent, and privileged-only verbs are still rejected on agent sockets.