Agent activity and resume safety: is this agent truly idle? - #15276
teamleaderleo wants to merge 13 commits into
Conversation
AgentActivity splits the coarse working/needs-input lifecycle into what an agent is doing now (thinking, a tool with its command and start time, subagents, background work, a question, a permission) and ResumeSafety classifies whether a restart or update can interrupt it (safe, care, risky, with reasons). Pure package code with tests; the app wiring follows. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Warning Review limit reachedNext included review available in 26 seconds. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (41)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
Dogfood build of cmux DEV pr-15276-ee70bfcb.app The link opens this exact commit in the cmux dev menu bar app. The build starts on each push and the page waits until it is ready; a newer push replaces it. It signs in against production, so Cloud or backend changes still need a tagged build with a development backend. Dogfood tours of
|
This comment has been minimized.
This comment has been minimized.
AgentHookActivityState folds one session's queued hooks (prompt submit, pre and post tool use by tool_use_id, stop with background work, idle notification, session start and end) into turn facts. AgentActivityEvidence combines them with the session registry state, the Feed decision overlay and a foreground command into AgentActivitySignals. AgentForegroundCommand picks the command an agent runs through a foreground shell child, so MCP servers and background shells do not count. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CI failure attributionCI passes on Written by |
The current-work reducer built a Cloud or SSH agent row with the badge's report provenance (hook, plugin, detected) as its kind and the raw daemon state, so a blocked remote agent never raised needs_input attention. SurfaceAgentBadge now exposes agentIdentity (the adapter, shared with the sidebar slot key) and sessionState (blocked is needs_input, done is ended), and carries extra.agent_session_id from the cmux-tui catalog so the row gets its session id. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentActivityIndex joins the agent session registry to live workspace and dock panels in one main-actor turn, adds each session's hook turn facts, the Feed decision overlay and, for local panes, a foreground command from one cached process census, then classifies. Each entry carries workspace, panel, surface and pane ids, name, agent, session, pid, placement (local, ssh with its host, cloud), survives_app_relaunch, activity and resume_safety. Claude now installs an unmatched queued PostToolUse hook. agent.hook.enqueue records every admitted hook in AgentHookActivityTracker before queueing, and a Claude post-tool-use only closes the open call: no hook process runs for it. agent.list is a worker-lane read advertised in capabilities. It has no relay contract because it returns local argv, and a test pins that. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review fixes for the activity index: - A compaction, resume or unknown SessionStart keeps the running turn; only startup and clear reset it. - Placement comes from the surface (cloud attachment, owning machine, remote terminal context), so a local pane in a remote workspace no longer reports survives_app_relaunch. - An unavailable or partial process census, or a local agent without a pid, sets foregroundCommandUnknown; the classifier never calls that safe (reason process_unknown). - Failed, denied and interrupted calls close: Claude also sends PostToolUseFailure to post-tool-use, an idle prompt ends the turn, and a later call of another tool closes a pending question. Both post-tool-use hooks are async. - Compaction keeps tool_use_id and agent_id; a post without an id closes the newest same-named call. - started_at is the PreToolUse time; since moves only on a change. - Subagent hooks and late posts never reopen a finished turn; hooks decide the turn only after they have seen a boundary. - Process filtering starts at the turn start, else at the last idle point, so older shells (such as shell-wrapped MCP servers) are ignored. - Relayed question PreToolUse is ignored: remote daemons send no PostToolUse, and the Feed overlay covers remote questions. - agent.list errors when the session owners are unavailable instead of returning an empty list, and redacts tool commands and argv. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…types package-conventions-lint flags caseless all-static enums. The classifier becomes AgentActivitySignals.classify() (plus AgentActivityEvidence.classify()), with the read-only and subagent-launcher tool sets on AgentActivity.Tool. AgentForegroundCommand becomes AgentProcessTree, an instantiated census with foregroundCommandPID(agentPID:notBefore:) and a static describe(arguments:). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Cross-model review (Codex gpt-5.6-sol)
|
The CLI's redaction policy is not compiled into the app target. The socket is same-user and never relayed, and the updater shows these commands to the user, so report them as recorded. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # cmux.xcodeproj/project.pbxproj
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Repro from the #15296 auto-update dogfood, for the classifier here. Claude's API was unreachable, so it kept retrying ( #15296 will switch its relaunch gate to |
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> # Conflicts: # tests/test_claude_wrapper_hooks.py
The agent census failed closed whenever any process on the machine exited between the PID list and its record read. On a busy host that is nearly every sample, so idle agents read care (process_unknown) instead of safe. A process that already exited cannot be a live foreground command; only an unavailable census or a truncated PID list leaves one unaccounted for. DarwinProcessListing and CmuxTopProcessSnapshot now carry pidListIsComplete alongside the strict enumerationIsComplete, which hibernation and memory pressure keep using. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
DogfoodTwo DEV builds of this branch ran on a separate Mac that was busy building. A scripted Claude session was played through the real hook path (
Process fallback: with hooks reporting an ended turn and a foreground child running under the agent, the row read Fixed during dogfood (ed0f5f0)The first build reported idle agents as Not covered by this dogfood: SSH and Cloud rows ( |
The synthetic reader rebuilt its listing without the new flag, so a missing PID read as a truncated list there, the opposite of production. Tests now cover an exited PID (whole list), a truncated or unavailable census (not whole) and an empty PID list. The census JSON reports pid_list_complete. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review of ed0f5f0 (census completeness)A code-reading review of the
Fixed in ce3ce29:
Left as follow-ups:
|
|
Implemented in #15887. Sidebar lifecycle now distinguishes Running, Subagents, and deterministic Waiting, while preserving Needs input and hibernation safety. Merged after review and CI gating. |
Step 2 of #15266. Answers "is this agent truly idle?" precisely enough for restart, update and hibernation decisions, and exposes it as
agent.list.Model (
Packages/Shared/CmuxAgentChat,Model/)AgentActivity.kindis one ofidle,awaiting_input,question,permission,thinking,tool(withtool = {name, command, started_at}),subagents,background,endedorunknown. Each value comes withsinceandsource(hook,transcript,screenorprocess).ResumeSafetyissafe,careorrisky, andResumeSafetyAssessmentaddsreasons.AgentActivityClassifier.classifymaps per-pane signals in this precedence order:process)A half-typed draft makes any live state risky, but draft detection belongs to separate work, so it stays unknown for now. This is the one field set shared with the updater's relaunch gate, so there is a single classifier. It stays separate from
AgentHibernationLifecycleState, which keeps its own gate.The rest of the pure logic lives in the same package:
AgentHookActivityStatefolds one session's hooks into turn facts. Open tools are keyed bytool_use_id, and a tool inside a subagent wins over theTasklauncher.AgentActivityEvidencebuilds signals from the registry state, the hook facts, the Feed decision overlay and the foreground command.AgentForegroundCommandfinds the running command in the process tree. Only a shell child of the agent in the terminal's foreground process group counts, so MCP servers never read as a running command.AgentPanePlacementis local, ssh(host) or cloud, withsurvivesAppRelaunch.App
Sources/Agents/AgentHookActivityTracker.swiftis a lock-protected per-(surface, session) store, capped at 512 sessions. It records hook facts inagent.hook.enqueuebefore queue admission.Sources/Agents/AgentActivityIndex.swiftprovides@MainActor snapshot() async -> [AgentActivitySnapshot]. It is the one functionagent.list,cmux agentsand the updater gate read.agent.listruns on the worker lane and is advertised in capabilities. It is deliberately not on the remote relay allowlist, because entries carry local argv. A relay policy test pins that. Each entry hasworkspace_id, panel_id, surface_id, pane_id, name, agent, session_id, pid, placement{kind,host}, survives_app_relaunch, activity{kind,since,source,tool}, resume_safety{safety,reasons}.isActivityRecordOnly), so it starts no hook delivery process. The wrapper copy, the wrapper hook test and the delivery policy test are updated to match.Remote agents in current work (follow-up from #15273)
SurfaceAgentBadgegainsagentIdentity(the adapter; Claude aliases becomeclaude) andsessionState(blockedbecomesneeds_input,donebecomesended).RemoteAgentSidebarStatus.agentKeynow readsagentIdentity, with unchanged behavior.extra.agent_session_idon its badge paths.CurrentWorkReducerreports remote agents by adapter, session id and normalized state, so a blocked remote agent raisesneeds_inputattention andcmux agentsstops showingunknownfor them.Known limits
PostToolUseFailure; adding it is a one-line follow-up once older versions no longer reject unknown hook keys.Testing
verify-localproject, launch-policy, feature-flag and localization checks pass.CurrentWorkReducerTests, the CLI Claude settings test, and the end-to-end hook flow.🤖 Generated with Claude Code