Skip to content

Agent activity and resume safety: is this agent truly idle? - #15276

Open
teamleaderleo wants to merge 13 commits into
mainfrom
feat/agent-activity
Open

teamleaderleo wants to merge 13 commits into
mainfrom
feat/agent-activity

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Step 2 of #15266. Answers "is this agent truly idle?" precisely enough for restart, update and hibernation decisions, and exposes it as agent.list.

Model (Packages/Shared/CmuxAgentChat, Model/)

  • AgentActivity.kind is one of idle, awaiting_input, question, permission, thinking, tool (with tool = {name, command, started_at}), subagents, background, ended or unknown. Each value comes with since and source (hook, transcript, screen or process).
  • ResumeSafety is safe, care or risky, and ResumeSafetyAssessment adds reasons. AgentActivityClassifier.classify maps per-pane signals in this precedence order:
Signal Kind Safety
ended ended safe
pending permission permission risky
open question or plan approval question risky
open Task/Agent tool subagents care
open read-only tool (Read, Grep, Glob, LS, WebFetch, WebSearch, ...) tool care
any other open tool tool (with command) risky
a live foreground shell under the agent, even when hooks look idle tool (source process) risky
mid-turn, last hook was a tool end thinking safe (between tool calls)
mid-turn otherwise thinking care
background task or scheduled wakeup after Stop background care
no hook evidence at all unknown care
awaiting a human awaiting_input safe
otherwise idle safe

A half-typed draft makes any live state risky, but draft detection belongs to separate work, so it stays unknown for now. This is the one field set shared with the updater's relaunch gate, so there is a single classifier. It stays separate from AgentHibernationLifecycleState, which keeps its own gate.

The rest of the pure logic lives in the same package:

  • AgentHookActivityState folds one session's hooks into turn facts. Open tools are keyed by tool_use_id, and a tool inside a subagent wins over the Task launcher.
  • AgentActivityEvidence builds signals from the registry state, the hook facts, the Feed decision overlay and the foreground command.
  • AgentForegroundCommand finds the running command in the process tree. Only a shell child of the agent in the terminal's foreground process group counts, so MCP servers never read as a running command.
  • AgentPanePlacement is local, ssh(host) or cloud, with survivesAppRelaunch.

App

  • Sources/Agents/AgentHookActivityTracker.swift is a lock-protected per-(surface, session) store, capped at 512 sessions. It records hook facts in agent.hook.enqueue before queue admission.
  • Sources/Agents/AgentActivityIndex.swift provides @MainActor snapshot() async -> [AgentActivitySnapshot]. It is the one function agent.list, cmux agents and the updater gate read.
  • agent.list runs on the worker lane and is advertised in capabilities. It is deliberately not on the remote relay allowlist, because entries carry local argv. A relay policy test pins that. Each entry has workspace_id, panel_id, surface_id, pane_id, name, agent, session_id, pid, placement{kind,host}, survives_app_relaunch, activity{kind,since,source,tool}, resume_safety{safety,reasons}.
  • Claude gets an unmatched, queued PostToolUse hook, so a PreToolUse without a PostToolUse means a tool is in flight. The app answers it as record-only (isActivityRecordOnly), so it starts no hook delivery process. The wrapper copy, the wrapper hook test and the delivery policy test are updated to match.

Remote agents in current work (follow-up from #15273)

  • SurfaceAgentBadge gains agentIdentity (the adapter; Claude aliases become claude) and sessionState (blocked becomes needs_input, done becomes ended). RemoteAgentSidebarStatus.agentKey now reads agentIdentity, with unchanged behavior.
  • The cmux-tui catalog parser now carries extra.agent_session_id on its badge paths.
  • CurrentWorkReducer reports remote agents by adapter, session id and normalized state, so a blocked remote agent raises needs_input attention and cmux agents stops showing unknown for them.

Known limits

  • A failed tool (no PostToolUse) reads as open, and therefore risky, until the next prompt or Stop. Newer Claude versions have PostToolUseFailure; adding it is a one-line follow-up once older versions no longer reject unknown hook keys.
  • The foreground-command probe runs for local panes only.
  • OpenCode sends no tool hooks, so it relies on the process tree.

Testing

  • Two throwaway packages that build only these sources: 26 CmuxAgentChat cases (classifier plus evidence assembly), and 6 badge identity and sidebar status cases.
  • The wrapper fast-path byte-match test passes.
  • verify-local project, launch-policy, feature-flag and localization checks pass.
  • Left to CI: the app and CLI compile, the CMUXAgentLaunch, CmuxControlSocket and CmuxRemoteWorkspace package tests, CurrentWorkReducerTests, the CLI Claude settings test, and the end-to-end hook flow.

🤖 Generated with Claude Code

AgentActivity splits the coarse working/needs-input lifecycle into what an
agent is doing now (thinking, a tool with its command and start time,
subagents, background work, a question, a permission) and ResumeSafety
classifies whether a restart or update can interrupt it (safe, care, risky,
with reasons). Pure package code with tests; the app wiring follows.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 26 seconds.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 1067b029-c18f-405d-967b-bfe674438bac

📥 Commits

Reviewing files that changed from the base of the PR and between 0f2d3d3 and ce3ce29.

📒 Files selected for processing (41)
  • CLI/CMUXCLI+AgentHookAdmission.swift
  • CLI/CMUXCLI+AgentHookPayload.swift
  • CLI/CMUXCLI+ClaudeHookSettings.swift
  • CLI/cmux.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentActivity.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentActivitySignalAssembly.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentHookActivityState.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentPanePlacement.swift
  • Packages/Shared/CmuxAgentChat/Sources/CmuxAgentChat/Model/AgentProcessTree.swift
  • Packages/Shared/CmuxAgentChat/Tests/CmuxAgentChatTests/AgentActivityClassifierTests.swift
  • Packages/Shared/CmuxAgentChat/Tests/CmuxAgentChatTests/AgentActivityEvidenceTests.swift
  • Packages/macOS/CMUXAgentLaunch/Sources/CMUXAgentLaunch/AgentHookDeliveryPolicy.swift
  • Packages/macOS/CMUXAgentLaunch/Tests/CMUXAgentLaunchTests/AgentHookDeliveryPolicyTests.swift
  • Packages/macOS/CmuxControlSocket/Sources/CmuxControlSocket/Wire/ControlCommandExecutionPolicy.swift
  • Packages/macOS/CmuxControlSocket/Tests/CmuxControlSocketTests/ControlCommandExecutionPolicyTests.swift
  • Packages/macOS/CmuxFoundation/Sources/CmuxFoundation/Process/DarwinProcessEnumerator.swift
  • Packages/macOS/CmuxFoundation/Sources/CmuxFoundation/Process/DarwinProcessListing.swift
  • Packages/macOS/CmuxFoundation/Tests/CmuxFoundationTests/DarwinResourceSamplingTests.swift
  • Packages/macOS/CmuxRemoteWorkspace/Tests/CmuxRemoteWorkspaceTests/RemoteCLIRelayPolicyTests.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/CmuxTuiSnapshotParser.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/RemoteAgentSidebarStatus.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/SurfaceAgentBadge+Identity.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Sources/CmuxSurfaceCatalogModel/SurfaceCatalogModel.swift
  • Packages/macOS/CmuxSurfaceCatalogModel/Tests/CmuxSurfaceCatalogModelTests/SurfaceAgentBadgeIdentityTests.swift
  • Resources/bin/cmux-claude-wrapper
  • Sources/AgentHookDeliveryEvent.swift
  • Sources/Agents/AgentActivityIndex.swift
  • Sources/Agents/AgentHookActivityTracker.swift
  • Sources/Agents/TerminalController+AgentList.swift
  • Sources/CmuxTopProcessSampler+Enrichment.swift
  • Sources/CmuxTopProcessSampler.swift
  • Sources/CmuxTopSnapshot.swift
  • Sources/Surfaces/CurrentWorkReducer.swift
  • Sources/TerminalController+Capabilities.swift
  • Sources/TerminalController.swift
  • cmux.xcodeproj/project.pbxproj
  • cmuxCLITests/CLIClaudeHookTimeoutRegressionTests.swift
  • cmuxTests/CmuxTopProcessSnapshotCaptureCoordinatorTests.swift
  • cmuxTests/CurrentWorkReducerTests.swift
  • cmuxTests/SyntheticProcessSnapshotReader.swift
  • tests/test_claude_wrapper_hooks.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Dogfood build of ee70bfcb00b3bde9d2d61e52cd0eb5f14f6899d4

cmux DEV pr-15276-ee70bfcb.app

The link opens this exact commit in the cmux dev menu bar app. The build starts on each push and the page waits until it is ready; a newer push replaces it. It signs in against production, so Cloud or backend changes still need a tagged build with a development backend.

Dogfood tours of ce3ce297

sidebar-and-chrome-tour at ce3ce297, on its merge 53bb5896 that CI built: not run (run)

skipped: the tour run could not check CI's build (a reuse error); a CI re-run or gh workflow run pr-media.yml -f pr=&lt;n&gt; tries again

Tours are picked by the paths globs in dogfood/scenarios/*.json; a Dogfood-tours: a, b line in the description picks them instead (none turns this off). Look at every frame before merging: a green tour only means no step failed.

@blacksmith-sh

This comment has been minimized.

AgentHookActivityState folds one session's queued hooks (prompt submit,
pre and post tool use by tool_use_id, stop with background work, idle
notification, session start and end) into turn facts. AgentActivityEvidence
combines them with the session registry state, the Feed decision overlay
and a foreground command into AgentActivitySignals. AgentForegroundCommand
picks the command an agent runs through a foreground shell child, so MCP
servers and background shells do not count.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

CI failure attribution

CI passes on ce3ce297f0 (run 36455673732 attempt 2).

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

teamleaderleo and others added 4 commits September 28, 2026 05:53
The current-work reducer built a Cloud or SSH agent row with the badge's
report provenance (hook, plugin, detected) as its kind and the raw daemon
state, so a blocked remote agent never raised needs_input attention.
SurfaceAgentBadge now exposes agentIdentity (the adapter, shared with the
sidebar slot key) and sessionState (blocked is needs_input, done is ended),
and carries extra.agent_session_id from the cmux-tui catalog so the row
gets its session id.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentActivityIndex joins the agent session registry to live workspace and
dock panels in one main-actor turn, adds each session's hook turn facts,
the Feed decision overlay and, for local panes, a foreground command from
one cached process census, then classifies. Each entry carries workspace,
panel, surface and pane ids, name, agent, session, pid, placement (local,
ssh with its host, cloud), survives_app_relaunch, activity and
resume_safety.

Claude now installs an unmatched queued PostToolUse hook. agent.hook.enqueue
records every admitted hook in AgentHookActivityTracker before queueing,
and a Claude post-tool-use only closes the open call: no hook process runs
for it.

agent.list is a worker-lane read advertised in capabilities. It has no
relay contract because it returns local argv, and a test pins that.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review fixes for the activity index:
- A compaction, resume or unknown SessionStart keeps the running turn;
  only startup and clear reset it.
- Placement comes from the surface (cloud attachment, owning machine,
  remote terminal context), so a local pane in a remote workspace no
  longer reports survives_app_relaunch.
- An unavailable or partial process census, or a local agent without a
  pid, sets foregroundCommandUnknown; the classifier never calls that
  safe (reason process_unknown).
- Failed, denied and interrupted calls close: Claude also sends
  PostToolUseFailure to post-tool-use, an idle prompt ends the turn, and
  a later call of another tool closes a pending question. Both
  post-tool-use hooks are async.
- Compaction keeps tool_use_id and agent_id; a post without an id closes
  the newest same-named call.
- started_at is the PreToolUse time; since moves only on a change.
- Subagent hooks and late posts never reopen a finished turn; hooks
  decide the turn only after they have seen a boundary.
- Process filtering starts at the turn start, else at the last idle
  point, so older shells (such as shell-wrapped MCP servers) are ignored.
- Relayed question PreToolUse is ignored: remote daemons send no
  PostToolUse, and the Feed overlay covers remote questions.
- agent.list errors when the session owners are unavailable instead of
  returning an empty list, and redacts tool commands and argv.

Refs #15266

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…types

package-conventions-lint flags caseless all-static enums. The classifier
becomes AgentActivitySignals.classify() (plus AgentActivityEvidence.classify()),
with the read-only and subagent-launcher tool sets on AgentActivity.Tool.
AgentForegroundCommand becomes AgentProcessTree, an instantiated census with
foregroundCommandPID(agentPID:notBefore:) and a static describe(arguments:).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Cross-model review (Codex gpt-5.6-sol)

  • Sources/Agents/AgentActivityIndex.swift:138-144 — production activity evidence never populates hasDraft, although the shared classifier treats a draft as risky. An otherwise idle/awaiting-input pane can therefore be reported safe to resume while the user has a half-typed prompt. Populate draft state when assembling local-pane evidence (or classify unavailable evidence conservatively), and add an integration test through AgentActivityIndex, not only a classifier unit test.

teamleaderleo and others added 3 commits September 28, 2026 07:40
The CLI's redaction policy is not compiled into the app target. The
socket is same-user and never relayed, and the updater shows these
commands to the user, so report them as recorded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	cmux.xcodeproj/project.pbxproj
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Repro from the #15296 auto-update dogfood, for the classifier here. Claude's API was unreachable, so it kept retrying (ConnectionRefused). Pressing Esc interrupted the turn, but no Stop hook arrived, so claude_code stayed .running until the Claude process exited, and the sidebar showed Running the whole time. With no tool open, an interrupted turn like this should read as care or safe, not risky. idle_prompt may already cover it. Evidence: #15296 (comment)

#15296 will switch its relaunch gate to AgentActivityIndex.snapshot() once this lands (safety = survivesAppRelaunch ? .safe : assessment.safety).

teamleaderleo and others added 2 commits September 28, 2026 09:54
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	tests/test_claude_wrapper_hooks.py
@teamleaderleo
teamleaderleo marked this pull request as ready for review September 28, 2026 16:17
The agent census failed closed whenever any process on the machine exited
between the PID list and its record read. On a busy host that is nearly
every sample, so idle agents read care (process_unknown) instead of safe.
A process that already exited cannot be a live foreground command; only an
unavailable census or a truncated PID list leaves one unaccounted for.

DarwinProcessListing and CmuxTopProcessSnapshot now carry pidListIsComplete
alongside the strict enumerationIsComplete, which hibernation and memory
pressure keep using.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 28, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Dogfood

Two DEV builds of this branch ran on a separate Mac that was busy building. A scripted Claude session was played through the real hook path (cmux hooks enqueue claude … from inside the DEV terminal, so CMUX_SURFACE_ID and the agent pid come from the pane), and agent.list was read after each step.

Step activity.kind resume_safety reasons
Session started idle safe idle
Prompt submitted thinking care thinking
Bash running (swift build -c release) tool risky foreground_command
Bash finished thinking safe between_tool_calls
Grep running tool care read_only_tool
Grep interrupted (PostToolUseFailure) thinking safe between_tool_calls
Task running subagents care subagents
idle_prompt after an Esc, no Stop awaiting_input safe awaiting_input
Turn ended on an API error (StopFailure) idle safe idle
Session ended row removed

Process fallback: with hooks reporting an ended turn and a foreground child running under the agent, the row read tool from source process (risky, foreground_command, argv truncated). After the child exited it read idle, safe.

Fixed during dogfood (ed0f5f0)

The first build reported idle agents as care with process_unknown on about a third of samples. The process census counted every PID that exited between listing and reading as a gap, and a host running builds has such exits on nearly every sample. A process that already exited cannot be a live command, so the agent census now fails closed only when it is unavailable or its PID list was truncated (pidListIsComplete). Hibernation and memory pressure keep the strict enumerationIsComplete. With the fix, ten back-to-back idle samples on the same busy host all read safe.

Not covered by this dogfood: SSH and Cloud rows (survives_app_relaunch), which need hooks relayed from a remote host.

The synthetic reader rebuilt its listing without the new flag, so a missing
PID read as a truncated list there, the opposite of production. Tests now
cover an exited PID (whole list), a truncated or unavailable census (not
whole) and an empty PID list. The census JSON reports pid_list_complete.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Review of ed0f5f0 (census completeness)

A code-reading review of the pidListIsComplete change found no bugs.

  • An unreadable listed PID cannot persistently hide a live foreground command. readBSDInfo falls back to sysctl KERN_PROC_PID, which also reads processes owned by other users, so a live process only fails to read while it is exiting or still being created. descendantPIDs also asks proc_listchildpids, so children of an unreadable parent are still found.
  • Every production constructor carries the right value, and the unavailable-capture path is forced false through captureIsAvailable.
  • Hibernation and memory pressure still use the strict enumerationIsComplete.

Fixed in ce3ce29:

  • SyntheticProcessSnapshotReader rebuilt its listing without the new flag, so it reported a missing PID as a truncated list, the opposite of production.
  • Added tests for an exited PID (list whole), a truncated or unavailable census (not whole), and an empty PID list.
  • The census JSON now reports pid_list_complete.
  • The comments now say an unreadable PID "exited or was still being created".

Left as follow-ups:

  • TerminalForegroundCommandCapture makes the same "is a command running" call by tty and hits the same churn.
  • The live-agent and restore indexes need their not-found semantics audited before they switch.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Implemented in #15887. Sidebar lifecycle now distinguishes Running, Subagents, and deterministic Waiting, while preserving Needs input and hibernation safety. Merged after review and CI gating.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant