Skip to content

Agent activity and resume safety: is this agent truly idle? - #17754

Closed
azooz2003-bit wants to merge 13 commits into
mainfrom
feat/agent-activity
Closed

azooz2003-bit wants to merge 13 commits into
mainfrom
feat/agent-activity

Conversation

@azooz2003-bit

@azooz2003-bit azooz2003-bit commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Step 2 of #15266. Answers "is this agent truly idle?" precisely enough for restart, update and hibernation decisions, and exposes it as agent.list.

Model (Packages/Shared/CmuxAgentChat, Model/)

  • AgentActivity.kind is one of idle, awaiting_input, question, permission, thinking, tool (with tool = {name, command, started_at}), subagents, background, ended or unknown. Each value comes with since and source (hook, transcript, screen or process).
  • ResumeSafety is safe, care or risky, and ResumeSafetyAssessment adds reasons. AgentActivityClassifier.classify maps per-pane signals in this precedence order:
Signal Kind Safety
ended ended safe
pending permission permission risky
open question or plan approval question risky
open Task/Agent tool subagents care
open read-only tool (Read, Grep, Glob, LS, WebFetch, WebSearch, ...) tool care
any other open tool tool (with command) risky
a live foreground shell under the agent, even when hooks look idle tool (source process) risky
mid-turn, last hook was a tool end thinking safe (between tool calls)
mid-turn otherwise thinking care
background task or scheduled wakeup after Stop background care
no hook evidence at all unknown care
awaiting a human awaiting_input safe
otherwise idle safe

A half-typed draft makes any live state risky, but draft detection belongs to separate work, so it stays unknown for now. This is the one field set shared with the updater's relaunch gate, so there is a single classifier. It stays separate from AgentHibernationLifecycleState, which keeps its own gate.

The rest of the pure logic lives in the same package:

  • AgentHookActivityState folds one session's hooks into turn facts. Open tools are keyed by tool_use_id, and a tool inside a subagent wins over the Task launcher.
  • AgentActivityEvidence builds signals from the registry state, the hook facts, the Feed decision overlay and the foreground command.
  • AgentForegroundCommand finds the running command in the process tree. Only a shell child of the agent in the terminal's foreground process group counts, so MCP servers never read as a running command.
  • AgentPanePlacement is local, ssh(host) or cloud, with survivesAppRelaunch.

App

  • Sources/Agents/AgentHookActivityTracker.swift is a lock-protected per-(surface, session) store, capped at 512 sessions. It records hook facts in agent.hook.enqueue before queue admission.
  • Sources/Agents/AgentActivityIndex.swift provides @MainActor snapshot() async -> [AgentActivitySnapshot]. It is the one function agent.list, cmux agents and the updater gate read.
  • agent.list runs on the worker lane and is advertised in capabilities. It is deliberately not on the remote relay allowlist, because entries carry local argv. A relay policy test pins that. Each entry has workspace_id, panel_id, surface_id, pane_id, name, agent, session_id, pid, placement{kind,host}, survives_app_relaunch, activity{kind,since,source,tool}, resume_safety{safety,reasons}.
  • Claude gets an unmatched, queued PostToolUse hook, so a PreToolUse without a PostToolUse means a tool is in flight. The app answers it as record-only (isActivityRecordOnly), so it starts no hook delivery process. The wrapper copy, the wrapper hook test and the delivery policy test are updated to match.

Remote agents in current work (follow-up from #15273)

  • SurfaceAgentBadge gains agentIdentity (the adapter; Claude aliases become claude) and sessionState (blocked becomes needs_input, done becomes ended). RemoteAgentSidebarStatus.agentKey now reads agentIdentity, with unchanged behavior.
  • The cmux-tui catalog parser now carries extra.agent_session_id on its badge paths.
  • CurrentWorkReducer reports remote agents by adapter, session id and normalized state, so a blocked remote agent raises needs_input attention and cmux agents stops showing unknown for them.

Known limits

  • A failed tool (no PostToolUse) reads as open, and therefore risky, until the next prompt or Stop. Newer Claude versions have PostToolUseFailure; adding it is a one-line follow-up once older versions no longer reject unknown hook keys.
  • The foreground-command probe runs for local panes only.
  • OpenCode sends no tool hooks, so it relies on the process tree.

Testing

  • Two throwaway packages that build only these sources: 26 CmuxAgentChat cases (classifier plus evidence assembly), and 6 badge identity and sidebar status cases.
  • The wrapper fast-path byte-match test passes.
  • verify-local project, launch-policy, feature-flag and localization checks pass.
  • Left to CI: the app and CLI compile, the CMUXAgentLaunch, CmuxControlSocket and CmuxRemoteWorkspace package tests, CurrentWorkReducerTests, the CLI Claude settings test, and the end-to-end hook flow.

🤖 Generated with Claude Code


Migrated from #15276 after correcting the PR author identity. The head branch and commit history are preserved.


Summary by cubic

Adds precise agent activity tracking that answers "is this agent truly idle?" precisely enough for restart, update, and hibernation decisions (step 2 of #15266). Classifies each agent pane's state into activity kinds (thinking, tool, permission, question, subagents, background, awaiting input, idle, ended, unknown) and resume safety (safe, care, risky with reasons) and exposes it via a new agent.list command.

New Features:

  • agent.list reports each live agent's activity, resume safety, placement (local/ssh/cloud), and whether it survives an app relaunch; it runs on the worker lane and is deliberately not relayed to remote workspaces since entries carry local argv (policy test pins that).
  • Claude now installs queued, async PostToolUse and PostToolUseFailure hooks; the app records them at admission to close open tool calls and starts no delivery process.
  • Remote agents in current work now report adapter identity and normalized session state (blocked becomes needs_input, done becomes ended), and carry extra.agent_session_id.

Bug Fixes:

  • Process census now distinguishes a whole PID list from full completeness, so processes that exit mid-census no longer make idle agents read as risky.

Known limits: a failed tool reads as open, and therefore risky, until the next prompt or Stop; the foreground-command probe runs for local panes only; OpenCode sends no tool hooks and relies on the process tree.

Written for commit ce3ce29. Summary will update on new commits.

Review in cubic Turn on auto-fix

teamleaderleo and others added 13 commits September 28, 2026 05:13
AgentActivity splits the coarse working/needs-input lifecycle into what an
agent is doing now (thinking, a tool with its command and start time,
subagents, background work, a question, a permission) and ResumeSafety
classifies whether a restart or update can interrupt it (safe, care, risky,
with reasons). Pure package code with tests; the app wiring follows.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentHookActivityState folds one session's queued hooks (prompt submit,
pre and post tool use by tool_use_id, stop with background work, idle
notification, session start and end) into turn facts. AgentActivityEvidence
combines them with the session registry state, the Feed decision overlay
and a foreground command into AgentActivitySignals. AgentForegroundCommand
picks the command an agent runs through a foreground shell child, so MCP
servers and background shells do not count.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The current-work reducer built a Cloud or SSH agent row with the badge's
report provenance (hook, plugin, detected) as its kind and the raw daemon
state, so a blocked remote agent never raised needs_input attention.
SurfaceAgentBadge now exposes agentIdentity (the adapter, shared with the
sidebar slot key) and sessionState (blocked is needs_input, done is ended),
and carries extra.agent_session_id from the cmux-tui catalog so the row
gets its session id.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentActivityIndex joins the agent session registry to live workspace and
dock panels in one main-actor turn, adds each session's hook turn facts,
the Feed decision overlay and, for local panes, a foreground command from
one cached process census, then classifies. Each entry carries workspace,
panel, surface and pane ids, name, agent, session, pid, placement (local,
ssh with its host, cloud), survives_app_relaunch, activity and
resume_safety.

Claude now installs an unmatched queued PostToolUse hook. agent.hook.enqueue
records every admitted hook in AgentHookActivityTracker before queueing,
and a Claude post-tool-use only closes the open call: no hook process runs
for it.

agent.list is a worker-lane read advertised in capabilities. It has no
relay contract because it returns local argv, and a test pins that.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review fixes for the activity index:
- A compaction, resume or unknown SessionStart keeps the running turn;
  only startup and clear reset it.
- Placement comes from the surface (cloud attachment, owning machine,
  remote terminal context), so a local pane in a remote workspace no
  longer reports survives_app_relaunch.
- An unavailable or partial process census, or a local agent without a
  pid, sets foregroundCommandUnknown; the classifier never calls that
  safe (reason process_unknown).
- Failed, denied and interrupted calls close: Claude also sends
  PostToolUseFailure to post-tool-use, an idle prompt ends the turn, and
  a later call of another tool closes a pending question. Both
  post-tool-use hooks are async.
- Compaction keeps tool_use_id and agent_id; a post without an id closes
  the newest same-named call.
- started_at is the PreToolUse time; since moves only on a change.
- Subagent hooks and late posts never reopen a finished turn; hooks
  decide the turn only after they have seen a boundary.
- Process filtering starts at the turn start, else at the last idle
  point, so older shells (such as shell-wrapped MCP servers) are ignored.
- Relayed question PreToolUse is ignored: remote daemons send no
  PostToolUse, and the Feed overlay covers remote questions.
- agent.list errors when the session owners are unavailable instead of
  returning an empty list, and redacts tool commands and argv.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…types

package-conventions-lint flags caseless all-static enums. The classifier
becomes AgentActivitySignals.classify() (plus AgentActivityEvidence.classify()),
with the read-only and subagent-launcher tool sets on AgentActivity.Tool.
AgentForegroundCommand becomes AgentProcessTree, an instantiated census with
foregroundCommandPID(agentPID:notBefore:) and a static describe(arguments:).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The CLI's redaction policy is not compiled into the app target. The
socket is same-user and never relayed, and the updater shows these
commands to the user, so report them as recorded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	cmux.xcodeproj/project.pbxproj
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	tests/test_claude_wrapper_hooks.py
The agent census failed closed whenever any process on the machine exited
between the PID list and its record read. On a busy host that is nearly
every sample, so idle agents read care (process_unknown) instead of safe.
A process that already exited cannot be a live foreground command; only an
unavailable census or a truncated PID list leaves one unaccounted for.

DarwinProcessListing and CmuxTopProcessSnapshot now carry pidListIsComplete
alongside the strict enumerationIsComplete, which hibernation and memory
pressure keep using.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The synthetic reader rebuilt its listing without the new flag, so a missing
PID read as a truncated list there, the opposite of production. Tests now
cover an exited PID (whole list), a truncated or unavailable census (not
whole) and an empty PID list. The census JSON reports pid_list_complete.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 52 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 97660e5f-d31a-47fa-a405-b1385bcfd5c9
📥 Commits

Reviewing files that changed from the base of the PR and between a8c4861 and ce3ce29.

📒 Files selected for processing (41)
  • CLI/CMUXCLI+AgentHookAdmission.swift
  • CLI/CMUXCLI+AgentHookPayload.swift
  • CLI/CMUXCLI+ClaudeHookSettings.swift
  • CLI/cmux.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentActivity.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentActivitySignalAssembly.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentHookActivityState.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentPanePlacement.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentProcessTree.swift
  • Packages/Shared/CmuxAgentChat/Tests/CmuxAgentChatTests/AgentActivityClassifierTests.swift
  • Packages/Shared/CmuxAgentChat/Tests/CmuxAgentChatTests/AgentActivityEvidenceTests.swift
  • Packages/macOS/CMUXAgentLaunch/Sources/CMUXAgentLaunch/AgentHookDeliveryPolicy.swift
  • Packages/macOS/CMUXAgentLaunch/Tests/CMUXAgentLaunchTests/AgentHookDeliveryPolicyTests.swift
  • Packages/macOS/CmuxControlSocket/Sources/CmuxControlSocket/Wire/ControlCommandExecutionPolicy.swift
  • Packages/macOS/CmuxControlSocket/Tests/CmuxControlSocketTests/ControlCommandExecutionPolicyTests.swift
  • Packages/macOS/CmuxFoundation/Sources/CmuxFoundation/Process/DarwinProcessEnumerator.swift
  • Packages/macOS/CmuxFoundation/Sources/CmuxFoundation/Process/DarwinProcessListing.swift
  • Packages/macOS/CmuxFoundation/Tests/CmuxFoundationTests/DarwinResourceSamplingTests.swift
  • Packages/macOS/CmuxRemoteWorkspace/Tests/CmuxRemoteWorkspaceTests/RemoteCLIRelayPolicyTests.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/CmuxTuiSnapshotParser.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/RemoteAgentSidebarStatus.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/SurfaceAgentBadge+Identity.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/SurfaceCatalogModel.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Tests/CmuxSurfaceCatalogModelTests/SurfaceAgentBadgeIdentityTests.swift
  • Resources/bin/cmux-claude-wrapper
  • Sources/AgentHookDeliveryEvent.swift
  • Sources/Agents/AgentActivityIndex.swift
  • Sources/Agents/AgentHookActivityTracker.swift
  • Sources/Agents/TerminalController+AgentList.swift
  • Sources/CmuxTopProcessSampler+Enrichment.swift
  • Sources/CmuxTopProcessSampler.swift
  • Sources/CmuxTopSnapshot.swift
  • Sources/Surfaces/CurrentWorkReducer.swift
  • Sources/TerminalController+Capabilities.swift
  • Sources/TerminalController.swift
  • cmux.xcodeproj/project.pbxproj
  • cmuxCLITests/CLIClaudeHookTimeoutRegressionTests.swift
  • cmuxTests/CmuxTopProcessSnapshotCaptureCoordinatorTests.swift
  • cmuxTests/CurrentWorkReducerTests.swift
  • cmuxTests/SyntheticProcessSnapshotReader.swift
  • tests/test_claude_wrapper_hooks.py
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Note

Pull Request opener @azooz2003-bit is not an author or co-author of any commit in this PR (commit identities: teamleaderleo, claude). The CLA check will still proceed and requires every listed identity plus @azooz2003-bit to have signed.

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants