Repository navigation
Agent activity and resume safety: is this agent truly idle? - #17754
azooz2003-bit wants to merge 13 commits into
Conversation
AgentActivity splits the coarse working/needs-input lifecycle into what an agent is doing now (thinking, a tool with its command and start time, subagents, background work, a question, a permission) and ResumeSafety classifies whether a restart or update can interrupt it (safe, care, risky, with reasons). Pure package code with tests; the app wiring follows. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentHookActivityState folds one session's queued hooks (prompt submit, pre and post tool use by tool_use_id, stop with background work, idle notification, session start and end) into turn facts. AgentActivityEvidence combines them with the session registry state, the Feed decision overlay and a foreground command into AgentActivitySignals. AgentForegroundCommand picks the command an agent runs through a foreground shell child, so MCP servers and background shells do not count. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The current-work reducer built a Cloud or SSH agent row with the badge's report provenance (hook, plugin, detected) as its kind and the raw daemon state, so a blocked remote agent never raised needs_input attention. SurfaceAgentBadge now exposes agentIdentity (the adapter, shared with the sidebar slot key) and sessionState (blocked is needs_input, done is ended), and carries extra.agent_session_id from the cmux-tui catalog so the row gets its session id. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentActivityIndex joins the agent session registry to live workspace and dock panels in one main-actor turn, adds each session's hook turn facts, the Feed decision overlay and, for local panes, a foreground command from one cached process census, then classifies. Each entry carries workspace, panel, surface and pane ids, name, agent, session, pid, placement (local, ssh with its host, cloud), survives_app_relaunch, activity and resume_safety. Claude now installs an unmatched queued PostToolUse hook. agent.hook.enqueue records every admitted hook in AgentHookActivityTracker before queueing, and a Claude post-tool-use only closes the open call: no hook process runs for it. agent.list is a worker-lane read advertised in capabilities. It has no relay contract because it returns local argv, and a test pins that. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review fixes for the activity index: - A compaction, resume or unknown SessionStart keeps the running turn; only startup and clear reset it. - Placement comes from the surface (cloud attachment, owning machine, remote terminal context), so a local pane in a remote workspace no longer reports survives_app_relaunch. - An unavailable or partial process census, or a local agent without a pid, sets foregroundCommandUnknown; the classifier never calls that safe (reason process_unknown). - Failed, denied and interrupted calls close: Claude also sends PostToolUseFailure to post-tool-use, an idle prompt ends the turn, and a later call of another tool closes a pending question. Both post-tool-use hooks are async. - Compaction keeps tool_use_id and agent_id; a post without an id closes the newest same-named call. - started_at is the PreToolUse time; since moves only on a change. - Subagent hooks and late posts never reopen a finished turn; hooks decide the turn only after they have seen a boundary. - Process filtering starts at the turn start, else at the last idle point, so older shells (such as shell-wrapped MCP servers) are ignored. - Relayed question PreToolUse is ignored: remote daemons send no PostToolUse, and the Feed overlay covers remote questions. - agent.list errors when the session owners are unavailable instead of returning an empty list, and redacts tool commands and argv. Refs #15266 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…types package-conventions-lint flags caseless all-static enums. The classifier becomes AgentActivitySignals.classify() (plus AgentActivityEvidence.classify()), with the read-only and subagent-launcher tool sets on AgentActivity.Tool. AgentForegroundCommand becomes AgentProcessTree, an instantiated census with foregroundCommandPID(agentPID:notBefore:) and a static describe(arguments:). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The CLI's redaction policy is not compiled into the app target. The socket is same-user and never relayed, and the updater shows these commands to the user, so report them as recorded. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # cmux.xcodeproj/project.pbxproj
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> # Conflicts: # tests/test_claude_wrapper_hooks.py
The agent census failed closed whenever any process on the machine exited between the PID list and its record read. On a busy host that is nearly every sample, so idle agents read care (process_unknown) instead of safe. A process that already exited cannot be a live foreground command; only an unavailable census or a truncated PID list leaves one unaccounted for. DarwinProcessListing and CmuxTopProcessSnapshot now carry pidListIsComplete alongside the strict enumerationIsComplete, which hibernation and memory pressure keep using. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The synthetic reader rebuilt its listing without the new flag, so a missing PID read as a truncated list there, the opposite of production. Tests now cover an exited PID (whole list), a truncated or unavailable census (not whole) and an empty PID list. The census JSON reports pid_list_complete. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 52 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (41)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Note Pull Request opener @azooz2003-bit is not an author or co-author of any commit in this PR (commit identities: All contributors have signed the CLA ✍️ ✅ |
Step 2 of #15266. Answers "is this agent truly idle?" precisely enough for restart, update and hibernation decisions, and exposes it as
agent.list.Model (
Packages/Shared/CmuxAgentChat,Model/)AgentActivity.kindis one ofidle,awaiting_input,question,permission,thinking,tool(withtool = {name, command, started_at}),subagents,background,endedorunknown. Each value comes withsinceandsource(hook,transcript,screenorprocess).ResumeSafetyissafe,careorrisky, andResumeSafetyAssessmentaddsreasons.AgentActivityClassifier.classifymaps per-pane signals in this precedence order:process)A half-typed draft makes any live state risky, but draft detection belongs to separate work, so it stays unknown for now. This is the one field set shared with the updater's relaunch gate, so there is a single classifier. It stays separate from
AgentHibernationLifecycleState, which keeps its own gate.The rest of the pure logic lives in the same package:
AgentHookActivityStatefolds one session's hooks into turn facts. Open tools are keyed bytool_use_id, and a tool inside a subagent wins over theTasklauncher.AgentActivityEvidencebuilds signals from the registry state, the hook facts, the Feed decision overlay and the foreground command.AgentForegroundCommandfinds the running command in the process tree. Only a shell child of the agent in the terminal's foreground process group counts, so MCP servers never read as a running command.AgentPanePlacementis local, ssh(host) or cloud, withsurvivesAppRelaunch.App
Sources/Agents/AgentHookActivityTracker.swiftis a lock-protected per-(surface, session) store, capped at 512 sessions. It records hook facts inagent.hook.enqueuebefore queue admission.Sources/Agents/AgentActivityIndex.swiftprovides@MainActor snapshot() async -> [AgentActivitySnapshot]. It is the one functionagent.list,cmux agentsand the updater gate read.agent.listruns on the worker lane and is advertised in capabilities. It is deliberately not on the remote relay allowlist, because entries carry local argv. A relay policy test pins that. Each entry hasworkspace_id, panel_id, surface_id, pane_id, name, agent, session_id, pid, placement{kind,host}, survives_app_relaunch, activity{kind,since,source,tool}, resume_safety{safety,reasons}.isActivityRecordOnly), so it starts no hook delivery process. The wrapper copy, the wrapper hook test and the delivery policy test are updated to match.Remote agents in current work (follow-up from #15273)
SurfaceAgentBadgegainsagentIdentity(the adapter; Claude aliases becomeclaude) andsessionState(blockedbecomesneeds_input,donebecomesended).RemoteAgentSidebarStatus.agentKeynow readsagentIdentity, with unchanged behavior.extra.agent_session_idon its badge paths.CurrentWorkReducerreports remote agents by adapter, session id and normalized state, so a blocked remote agent raisesneeds_inputattention andcmux agentsstops showingunknownfor them.Known limits
PostToolUseFailure; adding it is a one-line follow-up once older versions no longer reject unknown hook keys.Testing
verify-localproject, launch-policy, feature-flag and localization checks pass.CurrentWorkReducerTests, the CLI Claude settings test, and the end-to-end hook flow.🤖 Generated with Claude Code
Migrated from #15276 after correcting the PR author identity. The head branch and commit history are preserved.
Summary by cubic
Adds precise agent activity tracking that answers "is this agent truly idle?" precisely enough for restart, update, and hibernation decisions (step 2 of #15266). Classifies each agent pane's state into activity kinds (thinking, tool, permission, question, subagents, background, awaiting input, idle, ended, unknown) and resume safety (safe, care, risky with reasons) and exposes it via a new
agent.listcommand.New Features:
agent.listreports each live agent's activity, resume safety, placement (local/ssh/cloud), and whether it survives an app relaunch; it runs on the worker lane and is deliberately not relayed to remote workspaces since entries carry local argv (policy test pins that).blockedbecomesneeds_input,donebecomesended), and carryextra.agent_session_id.Bug Fixes:
Known limits: a failed tool reads as open, and therefore risky, until the next prompt or Stop; the foreground-command probe runs for local panes only; OpenCode sends no tool hooks and relies on the process tree.
Written for commit ce3ce29. Summary will update on new commits.