Skip to content

Fix Pi and OMP fork actions in tab context menus - #8173

Merged
lawrencecchen merged 59 commits into
mainfrom
feat-pi-fork-context-menu
Jul 17, 2026
Merged

lawrencecchen merged 59 commits into
mainfrom
feat-pi-fork-context-menu

Conversation

@lawrencecchen

@lawrencecchen lawrencecchen commented Jul 15, 2026 •

Copy link
Copy Markdown
Contributor

Pi and OMP already support forking, but cmux hid native Pi and generated the stale --session <id> --fork form for both agents. Their current CLIs require --fork <id>.

This PR:

  • adds native Pi to the shared fork availability classifier used by command palette and tab context menus
  • emits pi|omp --fork <session> through the shared sanitized argv builder and built-in registry templates
  • recognizes --fork <session> and --fork=<session> during live process reconciliation
  • adds behavioral panel-menu, argv, and fork-parent regression coverage in separate failing-test/fix commits

Hermes remains hidden because its CLI has resume support but no fork command. Custom registered agents with a forkCommand and OMP remain supported through the generic registry path.

Test plan:

  • exact-head tagged cloud build
  • package argv tests and app-host regression suites in GitHub Actions
  • live Pi/OMP CLI contract probes reject nonexistent IDs as session lookups rather than unknown options

Summary by CodeRabbit

  • New Features
    • Enhanced command-palette and context-menu fork availability for Pi/OMP, with clearer “remote unsupported” vs “probe required” behavior.
  • Bug Fixes
    • Improved fork/probe decision accuracy with broader --fork …/--fork=… parsing, stricter freshness/expiry handling, and better invalidation when executables or remote/local context change, including safer cancellation behavior.
    • Updated Pi/OMP fork command formatting and ensured legacy and restored sessions migrate consistently.
  • Documentation
    • Updated Vault and in-app Pi/OMP guidance to use --fork SESSION_ID.
  • Tests
    • Improved determinism and added Pi snapshot coverage for probe-required context menu behavior.

@coderabbitai

coderabbitai Bot commented Jul 15, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Pi/OMP fork commands now use --fork SESSION_ID, legacy registrations migrate during decoding, and fork-parent detection accepts attached values. Fork probing, shared validation, command-palette caching, and SSH termination coordination were substantially reworked.

Changes

Fork capability validation

Layer / File(s) Summary
Fork command construction and registration
Packages/macOS/CMUXAgentLaunch/..., Sources/VaultAgentRegistry.swift, docs/vault.md, web/messages/*.json
Pi and OMP fork argv and registration templates now place --fork before the session identifier, with tests and documentation updated.
Fork session detection and persisted migration
Sources/VaultObservedAgentProcess.swift, Sources/VaultAgentProcessScanner+ForkMarkers.swift, Sources/SessionRestorableAgentSnapshot+Commands.swift, CLI/CMUXCLI+SessionsListForkDiagnostics.swift, cmuxTests/*Fork*
Fork parsing supports space-separated and equals-form values, marker matching supports attached values, and legacy Pi/OMP registrations migrate during snapshot decoding.
Fork probe execution and identity validation
Sources/AgentForkSupport.swift, Sources/AgentForkCapabilityProbeCache.swift, Sources/RestorableAgentSession.swift
Probe execution uses process groups and actor-based output draining, while cache keys, executable identity, environment identity, TTL, and Pi-family version probing are expanded.
Shared fork validation lifecycle
Sources/SharedLiveAgentIndex.swift
Fork validation batches request IDs, handles cancellation and deduplication, tracks executable watches, and stores identity-aware validation results.
Command-palette fork state and snapshot selection
Sources/ContentView.swift, Sources/ContentView+ForkAgentConversation.swift, cmuxTests/WorkspaceForkConversationContextMenuTests.swift
Command-palette probing tracks accepted and rejected results, fingerprints, fallback usage, timestamps, expiry, and shared-index changes when selecting forkable snapshots.

SSH subprocess termination

Layer / File(s) Summary
Locked SSH termination coordination
Sources/FileExplorerStore.swift
SSH command and download processes now mutate termination gates under lock across launch, cancellation, failure, and completion paths.

Estimated code review effort: 5 (Critical) | ~120 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CommandPalette
  participant SharedLiveAgentIndex
  participant AgentForkSupport
  participant ProbeProcess
  CommandPalette->>SharedLiveAgentIndex: request fork availability
  SharedLiveAgentIndex->>AgentForkSupport: validate probe identity
  AgentForkSupport->>ProbeProcess: spawn fork probe
  ProbeProcess-->>AgentForkSupport: output and exit status
  AgentForkSupport-->>SharedLiveAgentIndex: cache validation result
  SharedLiveAgentIndex-->>CommandPalette: publish availability
Loading

Possibly related PRs

Suggested reviewers: austinywang


Important

Pre-merge checks failed

Please resolve all errors before merging. Addressing warnings is optional.

❌ Failed checks (8 errors, 1 warning)

Check name Status Explanation Resolution
Cmux Swift Blocking Runtime ❌ Error Production diff adds timer-driven coordination in AgentForkSupport/ContentView and expands NSLock/terminationGate handling in FileExplorerStore, which the rule forbids. Move runtime coordination to actors/explicit signals and avoid new dispatch timers/locks in production code; keep such constructs only in test scaffolding.
Cmux Algorithmic Complexity ❌ Error SharedLiveAgentIndex.applyPendingForkValidations rebuilds the tail dictionary inside the batch loop, causing O(n²) rescans over pending validation batches. Build the remaining-requests dictionary once per pass (or keep a queue/index) and restore from that cached structure instead of rescanning the tail for every batch.
Cmux Swift Package Boundaries ❌ Error The diff adds reusable fork/probe, session parsing, and registry logic in app-target Sources/ instead of a SwiftPM package boundary. Move the shared fork-support, registry, and snapshot parsing/service code into an app-adjacent SwiftPM package; keep only UI/CLI composition in Sources/.
Cmux User-Facing Error Privacy ❌ Error FAIL: cmux sessions-list now outputs fork_unavailable_reason values like omp_version_unverified/opencode_version_unverified, exposing vendor names and internal verification state. Map these to generic user-safe reasons (e.g. unverified/requires_validation) and keep provider names/details in internal logs only.
Cmux Full Internationalization ❌ Error The docs page reads docs.ohMyPi via next-intl, but the copy change landed only in web/messages/en.json and ja.json; the other routing locales weren't updated. Add matching translations for the changed docs.ohMyPi.whatYouGet3 copy in every supported locale file under web/messages/, or scope the route to the locales that are actually authored.
Cmux Architecture Rethink ❌ Error SharedLiveAgentIndex adds a DispatchSourceTimer and NotificationCenter side channel plus shared watch-record caches to paper over fork-validation races, which the rule forbids. Keep fork-watch ownership in one actor/state machine and drive invalidation from that source of truth; remove timer-based retries and notification broadcasts.
Cmux No Test Or Debug Seam In Production Source ❌ Error Sources/SharedLiveAgentIndex.swift adds forkExecutableWatchOpenSuspensionForTesting and an init seam in production code, which is a test-only hook. Move the suspension hook into Tests/ or a debug-only file; keep production APIs unchanged or expose only internal state via @testable import.
Cmux No Ambient Global State ❌ Error Sources/AgentForkSupport.swift:66 adds a static cache singleton on the AgentForkSupport namespace, creating new ambient runtime state. Move the probe cache onto an injectable owning service (or an existing dependency like SharedLiveAgentIndex) and construct it at the app seam instead of a static property.
Docstring Coverage ⚠️ Warning Docstring coverage is 1.01% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (16 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly matches the main change: fixing Pi and OMP fork actions in tab context menus.
Description check ✅ Passed The description covers the summary and testing, but it omits several template sections like Demo Video, Review Trigger, and Checklist.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Cmux Swift Actor Isolation ✅ Passed No new actor-isolation debt: new shared mutable state is lock-protected or actor-isolated, and UI-facing changes stay on MainActor.
Cmux Browser Automation Off-Main ✅ Passed PR diff is fork-only; changed hunks are availability/caching, with no edits to browser.* routing, socketWorkerMethods, or processV2Command.
Cmux Expensive Synchronous Load ✅ Passed The only RestorableAgentSessionIndex.load() path is still executed inside Task.detached in SharedLiveAgentIndex.reload; the command-palette/UI changes use cached SharedLiveAgentIndex.shared...
Cmux Cache Substitution Correctness ✅ Passed Fork availability cache is UI-only and has cold fallback plus TTL/identity checks; reads verify current snapshots/fingerprints and timers prune stale entries.
Cmux No Hacky Sleeps ✅ Passed The diff only touches Swift, docs, JSON, and project files; no TypeScript/JavaScript/shell/runtime scripts were changed, so the no-hacky-sleeps rule doesn’t apply.
Cmux Swift Concurrency ✅ Passed New GCD usage is confined to OS/process/file-watch boundaries and is lifecycle-managed; no new Combine or unmanaged fire-and-forget async app-state code appears.
Cmux Swift @Concurrent ✅ Passed No changed Swift code adds invalid @concurrent, and heavy async work is either actor-bound or explicitly offloaded via Task.detached.
Cmux Swiftpm Lockfiles ✅ Passed Diff only touches two source/test files; no Package.swift, Package.resolved, .gitignore, Xcode project, or workflow changes are present.
Cmux Swift Logging ✅ Passed The only production-code hunk adds a test suspension hook in SharedLiveAgentIndex; no new print, debugPrint, dump, NSLog, Logger, or ad hoc logging appears.
Cmux Swiftui State Layout ✅ Passed New ContentView state is @State-only cache/timer bookkeeping; no new ObservableObject/@Published/@StateObject/@EnvironmentObject, GeometryReader layout change, or render-time mutation pattern.
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed PASS: The PR changes fork/probe logic and tests only; I found no added/modified WindowGroup, NSPanel, NSWindowController, or cmuxAuxiliaryWindowIdentifiers code.
Cmux Source Artifacts ✅ Passed Changed paths are hand-written source, tests, docs, and localization files; none are artifact/temp/build/cache directories or copied outputs.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat-pi-fork-context-menu

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@greptile-apps

greptile-apps Bot commented Jul 15, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes Pi and OMP fork actions in tab context menus by updating the fork argv format from --session SESSION_ID --fork to --fork SESSION_ID, and adds a robust fork-capability probe infrastructure with TTL-aware caching, async deduplication, and file-descriptor budget management.

  • Argv format change: AgentForkArgv.withForkSessionValue now emits [executable, "--fork", sessionId]; VaultAgentRegistry.migratedPersistedBuiltInRegistration handles legacy persisted sessions; VaultAgentProcessScanner+ForkMarkers recognises both exact --fork and --fork=VALUE attached-value forms.
  • New probe infrastructure: AgentForkCommandOutputRunner (actor, replaces NSLock class), AgentForkCommandOutputDrain (off-executor pipe drain), AgentForkCapabilityProbeCache (TTL LRU, 128 entries), and AgentForkExecutableIdentityResolver (in-flight deduplication with per-key timeouts) are all new; posix_spawn with POSIX_SPAWN_START_SUSPENDED avoids a fast-exit race on probe processes.
  • Pi support: Pi added to commandPaletteSnapshotForkAvailability as .requiresProbe locally / .unsupported remotely; SharedLiveAgentIndex gains fork-executable-watch records and FD-budget management (ceiling of 64 watch sources).

Confidence Score: 4/5

Safe to merge with one minor cleanup needed — test scaffolding functions in production ContentView.swift should be moved to test helpers.

The core argv change and migration path are correct, the new probe infrastructure is well-structured, and backwards compatibility is preserved. One P2 issue (test-only wrappers in production source) and the pre-existing probeWorkingDirectory priority inversion (noted in a prior thread) prevent a score of 5.

Sources/ContentView.swift (test-only wrappers), Sources/AgentForkSupport.swift (probeWorkingDirectory priority order)

Important Files Changed

Filename Overview
Sources/AgentForkSupport.swift Major refactor of fork support; ProcessTerminationGate changed to struct (safe — FileExplorerStore wraps mutations in NSLock); probeWorkingDirectory priority inverted (snapshot first, then launchCommand); new Pi-family probe/validation helpers; ForkCapabilityProbeResultCache actor added
Sources/SharedLiveAgentIndex.swift New fork-executable-watch infrastructure with identity-resolver integration, FD-budget management (64 watch sources via /dev/fd scan), and waiter-queue drain on deinit; complex but structurally sound
Sources/ContentView.swift Pi added to fork availability; new fingerprint/validation state vars; two test-only static wrapper functions shipped in production source (no-test-debug-seam-in-production-source violation)
Sources/AgentForkCommandOutputRunner.swift New actor replacing NSLock-based CommandOutputRunner; uses posix_spawn with POSIX_SPAWN_START_SUSPENDED to avoid fast-exit race; correct process-group teardown
Sources/AgentForkCommandOutputDrain.swift New actor draining probe output pipe off cooperative executor via wake-pipe + poll/read loop on concurrent GCD queue; clean design
Sources/AgentForkCapabilityProbeCache.swift New TTL-aware LRU actor cache (maxEntries=128); removeValue uses O(n) insertionOrder scan but bounded and acceptable
Sources/AgentForkExecutableIdentityResolver.swift Deduplicates in-flight identity/validation tasks with per-key timeouts and work-count ceiling of 16; correctly handles cancellation
Packages/macOS/CMUXAgentLaunch/Sources/CMUXAgentLaunch/AgentForkArgv.swift withSessionFork renamed to withForkSessionValue; emits --fork SESSION_ID instead of --session SESSION_ID --fork; backwards-compatible migration added in VaultAgentRegistry
Sources/VaultAgentProcessScanner+ForkMarkers.swift forkParentMarkerTokens extended to recognise --fork=VALUE attached-value form in addition to exact --fork token; backwards compatible
Sources/VaultAgentRegistry.swift builtInPi/builtInOmp forkCommand updated to new argv form; migratedPersistedBuiltInRegistration handles legacy sessions correctly

Reviews (48): Last reviewed commit: "Harden fork probe cache activation" | Re-trigger Greptile

@lawrencecchen lawrencecchen changed the title Show Pi fork actions in tab context menus Fix Pi and OMP fork actions in tab context menus Jul 15, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@Sources/VaultAgentProcessScanner`+ForkMarkers.swift:
- Around line 7-15: In the marker-matching logic, update the outer
markers.allSatisfy closure to construct marker + "=" once before entering
arguments.contains, then reuse that value in the argument range check. Preserve
the existing case-insensitive, literal, and anchored matching behavior.

In `@Sources/VaultAgentRegistry.swift`:
- Around line 165-170: Update migratedLegacyBuiltInForkCommand to compare
against legacy versions of both builtInOmp and builtInPi, returning the
corresponding current built-in registration when either matches the old
forkCommand template; preserve self for registrations that do not match either
legacy definition.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 22c19dd1-5d28-4225-8e67-c68075d4b98e

📥 Commits

Reviewing files that changed from the base of the PR and between 50e1a8a and 4011959.

📒 Files selected for processing (5)
  • Sources/SessionRestorableAgentSnapshot+Commands.swift
  • Sources/VaultAgentProcessScanner+ForkMarkers.swift
  • Sources/VaultAgentRegistry.swift
  • cmuxTests/ForkParentFallbackGeneralizationTests.swift
  • cmuxTests/WorkspaceForkConversationContextMenuTests.swift

Comment thread Sources/VaultAgentProcessScanner+ForkMarkers.swift
Comment thread Sources/VaultAgentRegistry.swift Outdated
@cursor

cursor Bot commented Jul 15, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
Sources/VaultAgentProcessScanner+ForkMarkers.swift (1)

7-15: 🚀 Performance & Scalability | 🔵 Trivial | 💤 Low value

Hoist marker.token + "=" string allocation outside the arguments loop.

This loop evaluates every argument against the marker. Creating the marker.token + "=" string inside the arguments.contains closure allocates a new string for every argument scanned. Hoisting it outside the closure avoids repeated per-element string construction on the process-scanning path. As per coding guidelines, avoid repeated per-element string interpolation or concatenation that creates large intermediate strings on hot paths.

♻️ Proposed fix
         return markers.allSatisfy { marker in
+            let equalsMarker = marker.acceptsAttachedValue ? marker.token + "=" : nil
             arguments.contains { argument in
                 argument.compare(marker.token, options: [.caseInsensitive, .literal]) == .orderedSame
-                    || (marker.acceptsAttachedValue && argument.range(
-                        of: marker.token + "=",
-                        options: [.anchored, .caseInsensitive, .literal]
-                    ) != nil)
+                    || (equalsMarker != nil && argument.range(
+                        of: equalsMarker!,
+                        options: [.anchored, .caseInsensitive, .literal]
+                    ) != nil)
             }
         }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/VaultAgentProcessScanner`+ForkMarkers.swift around lines 7 - 15, In
the marker-checking flow, update the closure around marker.token so the
marker.token + "=" value is constructed once before arguments.contains, then
reuse that local value for the attached-value range check. Preserve the existing
case-insensitive, literal, and anchored matching behavior.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Duplicate comments:
In `@Sources/VaultAgentProcessScanner`+ForkMarkers.swift:
- Around line 7-15: In the marker-checking flow, update the closure around
marker.token so the marker.token + "=" value is constructed once before
arguments.contains, then reuse that local value for the attached-value range
check. Preserve the existing case-insensitive, literal, and anchored matching
behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: dcd5386f-9e26-49cf-a2d0-ba165dc51f26

📥 Commits

Reviewing files that changed from the base of the PR and between eace3ac and 4568378.

📒 Files selected for processing (8)
  • Sources/SessionRestorableAgentSnapshot+Commands.swift
  • Sources/VaultAgentProcessScanner+ForkMarkers.swift
  • Sources/VaultAgentRegistry.swift
  • cmuxTests/ForkParentFallbackGeneralizationTests.swift
  • cmuxTests/WorkspaceForkConversationContextMenuTests.swift
  • docs/vault.md
  • web/messages/en.json
  • web/messages/ja.json

@cursor

cursor Bot commented Jul 15, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@cursor

cursor Bot commented Jul 15, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
Sources/ContentView.swift (1)

5831-6049: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚖️ Poor tradeoff

Duplicated reuse/reject/clear decision logic between fallback and non-fallback branches.

The .requiresProbe fallback branch (around lines 5907-5976) and the no-fallback branch (5980-6048) both independently compute cachedResultIsFresh, probeRejectionMatches, sharedProbeAccepted/sharedProbeRejected, and the reuse/clear decision, with only the snapshot source differing. Extracting a shared helper would reduce the risk of a future fix landing in only one branch.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/ContentView.swift` around lines 5831 - 6049, Refactor
refreshCommandPaletteForkableAgentAvailabilityIfNeeded so the fallback
.requiresProbe path and the no-fallback path share a helper for cached-result
freshness, rejection matching, shared probe status, reuse, and clearing
decisions. Keep snapshot-specific inputs configurable through the helper,
preserve each branch’s existing probe arguments and behavior, and update both
branches to use the shared decision logic.
Sources/AgentForkSupport.swift (1)

851-960: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Misleading name: probeFromDefaultDirectoryWhenWorkingDirectoryIsMissing actually rejects, not substitutes a default directory.

Despite the name, usesDefaultDirectoryForMissingWorkingDirectory == true causes an early return false (reject) rather than probing from any default/fallback directory. The name suggests a substitution behavior that doesn't exist. Naming is internally consistent across all three call sites, so this is cosmetic only.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/AgentForkSupport.swift` around lines 851 - 960, The parameter
probeFromDefaultDirectoryWhenWorkingDirectoryIsMissing and derived flag
usesDefaultDirectoryForMissingWorkingDirectory describe substitution but
actually reject when the working directory is missing. Rename both symbols to
clearly indicate rejection of a missing working directory, and update all three
call sites consistently without changing behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@CLI/CMUXCLI`+SessionsListForkDiagnostics.swift:
- Around line 385-411: The sessionsListPiFamilyAgent heuristic duplicates
AgentForkSupport.piFamilyProbeAgentID and may drift from runtime detection.
Reuse the shared AgentForkSupport detection logic, adapting the available
AgentHookLaunchCommandRecord data as needed, and update
sessionsListPiFamilyAgent and related diagnostics to derive Pi/OMP identity and
capability decisions from that shared implementation rather than maintaining
separate agent, launcher, and executable checks.

In `@Sources/ContentView`+ForkAgentConversation.swift:
- Around line 195-204: Remove the duplicate field-clearing implementation from
clearCommandPaletteForkableAgentCache(panelKey:) and delegate to the canonical
clearCommandPaletteForkableAgentProbeResult(for:) helper instead, preserving the
existing cache invalidation behavior through a single maintenance point.

---

Outside diff comments:
In `@Sources/AgentForkSupport.swift`:
- Around line 851-960: The parameter
probeFromDefaultDirectoryWhenWorkingDirectoryIsMissing and derived flag
usesDefaultDirectoryForMissingWorkingDirectory describe substitution but
actually reject when the working directory is missing. Rename both symbols to
clearly indicate rejection of a missing working directory, and update all three
call sites consistently without changing behavior.

In `@Sources/ContentView.swift`:
- Around line 5831-6049: Refactor
refreshCommandPaletteForkableAgentAvailabilityIfNeeded so the fallback
.requiresProbe path and the no-fallback path share a helper for cached-result
freshness, rejection matching, shared probe status, reuse, and clearing
decisions. Keep snapshot-specific inputs configurable through the helper,
preserve each branch’s existing probe arguments and behavior, and update both
branches to use the shared decision logic.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 0631e4f3-911e-49f4-b6ca-ccea9ae4202b

📥 Commits

Reviewing files that changed from the base of the PR and between eace3ac and b71b68e.

📒 Files selected for processing (9)
  • CLI/CMUXCLI+SessionsListForkDiagnostics.swift
  • Sources/AgentForkCapabilityProbeCache.swift
  • Sources/AgentForkSupport.swift
  • Sources/ContentView+ForkAgentConversation.swift
  • Sources/ContentView.swift
  • Sources/FileExplorerStore.swift
  • Sources/RestorableAgentSession.swift
  • Sources/SessionRestorableAgentSnapshot+Commands.swift
  • Sources/SharedLiveAgentIndex.swift

Comment thread CLI/CMUXCLI+SessionsListForkDiagnostics.swift
Comment thread Sources/ContentView+ForkAgentConversation.swift

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@Sources/AgentForkSupport.swift`:
- Around line 422-435: Update CommandOutputDrain.run() to capture self in the
queue’s asynchronous closure instead of capturing the non-Sendable readHandle
property directly, while continuing to pass self.readHandle to drainOutput.
Remove the readHandle.close() call from drainOutput so
CommandOutputRunner.complete() remains the sole owner responsible for closing
the handle.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: e081510b-ec64-4d97-8c99-444933c6cb38

📥 Commits

Reviewing files that changed from the base of the PR and between e111420 and a61a30a.

📒 Files selected for processing (1)
  • Sources/AgentForkSupport.swift

Comment thread Sources/AgentForkSupport.swift Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
Sources/SharedLiveAgentIndex.swift (2)

624-624: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Preserve live-panel ownership for mixed request batches.

A batch containing both a fallback request and a live-index request selects the fallback, making this flag false. That validation then survives panel removal. Mark it live-owned whenever any active request lacks a fallback.

Proposed fix
-            let validationRequiresLiveIndexPanel = fallbackSnapshot == nil
+            let validationRequiresLiveIndexPanel = activeRequests.contains {
+                $0.fallbackSnapshot == nil
+            }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/SharedLiveAgentIndex.swift` at line 624, Update
validationRequiresLiveIndexPanel in the request-batch validation logic to be
true whenever any active request has no fallbackSnapshot, rather than only when
the selected fallbackSnapshot is nil. Preserve live-panel ownership for mixed
fallback and live-index batches so validation remains after panel removal.

994-1067: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Re-check watch membership before committing the new generation Sources/SharedLiveAgentIndex.swift:994-1067 clears the existing watch and then suspends twice before installing the replacement. On @MainActor, another refresh can interleave in that window, so the later install can orphan the earlier generation after both calls passed the pre-open lookup. Re-check and merge immediately before the final assignment, or keep the old watch until the replacement is ready.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/SharedLiveAgentIndex.swift` around lines 994 - 1067, Update
updateForkExecutableWatch to re-check forkExecutableWatchRecords and watch
membership immediately before committing a newly opened watch generation, after
the detached descriptor-opening work completes. If another refresh has installed
or joined the same watchKey, close the newly opened descriptors and merge the
probe into the existing record instead of assigning a replacement; preserve
generation/key mappings and avoid orphaning the prior watch.

Sources: Coding guidelines, Path instructions

Sources/AgentForkSupport.swift (1)

1270-1328: 🚀 Performance & Scalability | 🟠 Major | 🏗️ Heavy lift

Avoid the global PID/FD sweep in probe teardown
Sources/AgentForkSupport.swift:1270-1328 still walks every PID and each process’s fd table on timeout cleanup. That’s O(processes × descriptors), and the fixed 8,192/1,024 buffers can miss the leaked holder on larger hosts. Keep an explicit descendant/ownership set and signal that instead.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/AgentForkSupport.swift` around lines 1270 - 1328, Replace the global
PID and file-descriptor scans in processIdentifiersHoldingProbeOutputPipe and
processHoldsProbeOutputPipe with an explicit descendant/ownership set maintained
during probe execution. Track processes that may own the probe output pipe, and
have timeout teardown signal only those tracked identifiers; remove the
fixed-size proc_listallpids/proc_pidinfo sweep and its missed-holder behavior.

Sources: Coding guidelines, Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@Sources/AgentForkSupport.swift`:
- Around line 1270-1328: Replace the global PID and file-descriptor scans in
processIdentifiersHoldingProbeOutputPipe and processHoldsProbeOutputPipe with an
explicit descendant/ownership set maintained during probe execution. Track
processes that may own the probe output pipe, and have timeout teardown signal
only those tracked identifiers; remove the fixed-size
proc_listallpids/proc_pidinfo sweep and its missed-holder behavior.

In `@Sources/SharedLiveAgentIndex.swift`:
- Line 624: Update validationRequiresLiveIndexPanel in the request-batch
validation logic to be true whenever any active request has no fallbackSnapshot,
rather than only when the selected fallbackSnapshot is nil. Preserve live-panel
ownership for mixed fallback and live-index batches so validation remains after
panel removal.
- Around line 994-1067: Update updateForkExecutableWatch to re-check
forkExecutableWatchRecords and watch membership immediately before committing a
newly opened watch generation, after the detached descriptor-opening work
completes. If another refresh has installed or joined the same watchKey, close
the newly opened descriptors and merge the probe into the existing record
instead of assigning a replacement; preserve generation/key mappings and avoid
orphaning the prior watch.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: eaa93e24-7e30-4857-9cb7-1fbfba715950

📥 Commits

Reviewing files that changed from the base of the PR and between a61a30a and 0b4c377.

📒 Files selected for processing (3)
  • Sources/AgentForkSupport.swift
  • Sources/SharedLiveAgentIndex.swift
  • cmuxTests/WorkspaceForkConversationContextMenuTests.swift

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (4)
Sources/AgentForkSupport.swift (4)

899-954: 🚀 Performance & Scalability | 🟠 Major | ⚡ Quick win

Hoist regex compilation out of the output parsing loop.

String.range(of:options: .regularExpression) and piFamilyVersionBoundToAgent both implicitly or explicitly compile NSRegularExpressions. Calling them inside the for rawLine in output.split(...) loop forces O(N) repeated regex compilations for every line of process output, creating unnecessary CPU overhead on parsing-heavy paths.

Precompile both regexes outside the loop to improve performance.

🚀 Proposed performance fix
     private static func piFamilyProbeVersion(
         in output: String,
         agentID: String,
         acceptsBareVersionOutput: Bool
     ) -> SemanticVersion? {
         let normalizedAgentID = agentID
             .trimmingCharacters(in: .whitespacesAndNewlines)
             .lowercased()
         guard normalizedAgentID == "pi" || normalizedAgentID == "omp" else { return nil }
 
+        let bareRegex = try? NSRegularExpression(pattern: #"^v?\d+\.\d+(?:\.\d+)?$"#)
+        let escapedAgentID = NSRegularExpression.escapedPattern(for: normalizedAgentID)
+        let boundPattern = #"(^|[^a-z0-9])"# + escapedAgentID + #"([/\s:_-]+)v?(\d+)\.(\d+)(?:\.(\d+))?($|[^a-z0-9])"#
+        let boundRegex = try? NSRegularExpression(pattern: boundPattern)
+
         var candidates: [SemanticVersion] = []
         for rawLine in output.split(whereSeparator: \.isNewline) {
             let line = String(rawLine).trimmingCharacters(in: .whitespacesAndNewlines)
             let lowercasedLine = line.lowercased()
-            let isBareVersionLine = lowercasedLine.range(
-                of: #"^v?\d+\.\d+(?:\.\d+)?$"#,
-                options: .regularExpression
-            ) != nil
+            
+            let isBareVersionLine: Bool
+            if let bareRegex {
+                let range = NSRange(lowercasedLine.startIndex..<lowercasedLine.endIndex, in: lowercasedLine)
+                isBareVersionLine = bareRegex.firstMatch(in: lowercasedLine, range: range) != nil
+            } else {
+                isBareVersionLine = false
+            }
+
             if acceptsBareVersionOutput && isBareVersionLine,
                let version = SemanticVersion.first(in: lowercasedLine) {
                 candidates.append(version)
                 continue
             }
-            if let version = piFamilyVersionBoundToAgent(
-                lowercasedLine,
-                agentID: normalizedAgentID
-            ) {
+            if let boundRegex,
+               let version = piFamilyVersionBoundToAgent(lowercasedLine, expression: boundRegex) {
                 candidates.append(version)
             }
         }
         return candidates.count == 1 ? candidates[0] : nil
     }
 
-    private static func piFamilyVersionBoundToAgent(_ line: String, agentID: String) -> SemanticVersion? {
-        let escapedAgentID = NSRegularExpression.escapedPattern(for: agentID)
-        let pattern = #"(^|[^a-z0-9])"# + escapedAgentID
-            + #"([/\s:_-]+)v?(\d+)\.(\d+)(?:\.(\d+))?($|[^a-z0-9])"#
-        guard let expression = try? NSRegularExpression(pattern: pattern) else { return nil }
+    private static func piFamilyVersionBoundToAgent(_ line: String, expression: NSRegularExpression) -> SemanticVersion? {
         let range = NSRange(line.startIndex..<line.endIndex, in: line)
         guard let match = expression.firstMatch(in: line, range: range) else { return nil }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/AgentForkSupport.swift` around lines 899 - 954, Hoist the
bare-version and agent-bound version regex compilation out of the for loop in
piFamilyProbeVersion, reusing precompiled expressions for each output line.
Update piFamilyVersionBoundToAgent to accept and use the precompiled
agent-specific expression instead of compiling an NSRegularExpression per call,
while preserving the existing matching and candidate-selection behavior.

546-549: 🚀 Performance & Scalability | 🟠 Major | ⚡ Quick win

Add @concurrent to explicitly hop heavy async work away from the caller's actor.

These nonisolated async functions perform or orchestrate heavy synchronous workloads, including file I/O (stat/realpath via forkProbeExecutableIdentity), POSIX process spawning, and waiting. Under Swift 6, functions without explicit isolation inherit the caller's actor context. If invoked from UI isolation, this synchronous I/O will severely block the main thread. As per coding guidelines, explicitly add the @concurrent attribute to ensure they predictably leave the caller's actor.

  • Sources/AgentForkSupport.swift#L546-L549: Prepend @concurrent to static func supportsFork(.
  • Sources/AgentForkSupport.swift#L963-L970: Prepend @concurrent to private static func supportsLocalForkProbe(.
  • Sources/AgentForkSupport.swift#L1331-L1336: Prepend @concurrent to private static func commandOutput(.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/AgentForkSupport.swift` around lines 546 - 549, Add the `@concurrent`
attribute to the async functions supportsFork, supportsLocalForkProbe, and
commandOutput in Sources/AgentForkSupport.swift at lines 546-549, 963-970, and
1331-1336 respectively, so their heavy synchronous work explicitly executes away
from the caller’s actor context.

Source: Coding guidelines


338-346: 🩺 Stability & Availability | 🔴 Critical | 🏗️ Heavy lift

Avoid blocking the actor's cooperative executor with system-wide kernel scans.

terminateProcessesHoldingOutputPipe runs synchronously on the CommandOutputRunner actor. It invokes processIdentifiersHoldingProbeOutputPipe, which queries proc_listallpids and then performs proc_pidinfo(PROC_PIDLISTFDS) across potentially thousands of processes. Executing this unbounded O(P * F) kernel scan synchronously blocks the actor's thread, stalling the global cooperative thread pool and degrading app stability.

Since the actor only needs to send signals and doesn't require waiting for this scan to update its state, dispatch this heavy work to a detached task. As per path instructions, avoid unbenchmarked O(N) paths without offloading or bounds.

⚡ Proposed fix
         private func terminateProcessesHoldingOutputPipe(signal: Int32) {
             guard !outputPipeHandles.isEmpty else { return }
-            for processIdentifier in AgentForkSupport.processIdentifiersHoldingProbeOutputPipe(
-                outputPipeHandles,
-                excluding: [Darwin.getpid()]
-            ) {
-                Darwin.kill(processIdentifier, signal)
-            }
+            let handles = outputPipeHandles
+            Task.detached(priority: .utility) {
+                for processIdentifier in AgentForkSupport.processIdentifiersHoldingProbeOutputPipe(
+                    handles,
+                    excluding: [Darwin.getpid()]
+                ) {
+                    Darwin.kill(processIdentifier, signal)
+                }
+            }
         }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/AgentForkSupport.swift` around lines 338 - 346, Update
terminateProcessesHoldingOutputPipe to offload
processIdentifiersHoldingProbeOutputPipe and the subsequent Darwin.kill loop to
a detached task, preserving the existing handle guard, excluded PID, and signal
behavior while avoiding synchronous kernel scanning on the CommandOutputRunner
actor.

Source: Path instructions


504-509: 🩺 Stability & Availability | 🔴 Critical | ⚡ Quick win

Handle POLLNVAL to prevent infinite CPU-spinning tight loops.

If the file descriptor becomes invalid, poll returns immediately with the POLLNVAL bit set in revents. Because POLLNVAL is not included in the (POLLIN | POLLHUP | POLLERR) mask, the bitwise AND evaluates to 0. This triggers the continue statement, causing poll to be immediately re-invoked and returning POLLNVAL again, resulting in an infinite 100% CPU tight loop.

Explicitly check for POLLNVAL and break out of the loop.

🩹 Proposed fix
                 if pollDescriptors[1].revents != 0 {
                     break
                 }
+                if pollDescriptors[0].revents & Int16(POLLNVAL) != 0 {
+                    output.removeAll(keepingCapacity: false)
+                    break
+                }
                 if pollDescriptors[0].revents & Int16(POLLIN | POLLHUP | POLLERR) == 0 {
                     continue
                 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Sources/AgentForkSupport.swift` around lines 504 - 509, Update the polling
loop around pollDescriptors[0].revents to explicitly detect the POLLNVAL bit and
break before the existing readiness-mask continue check. Preserve the current
handling for POLLIN, POLLHUP, and POLLERR while ensuring an invalid descriptor
cannot re-enter poll indefinitely.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@Sources/AgentForkSupport.swift`:
- Around line 899-954: Hoist the bare-version and agent-bound version regex
compilation out of the for loop in piFamilyProbeVersion, reusing precompiled
expressions for each output line. Update piFamilyVersionBoundToAgent to accept
and use the precompiled agent-specific expression instead of compiling an
NSRegularExpression per call, while preserving the existing matching and
candidate-selection behavior.
- Around line 546-549: Add the `@concurrent` attribute to the async functions
supportsFork, supportsLocalForkProbe, and commandOutput in
Sources/AgentForkSupport.swift at lines 546-549, 963-970, and 1331-1336
respectively, so their heavy synchronous work explicitly executes away from the
caller’s actor context.
- Around line 338-346: Update terminateProcessesHoldingOutputPipe to offload
processIdentifiersHoldingProbeOutputPipe and the subsequent Darwin.kill loop to
a detached task, preserving the existing handle guard, excluded PID, and signal
behavior while avoiding synchronous kernel scanning on the CommandOutputRunner
actor.
- Around line 504-509: Update the polling loop around pollDescriptors[0].revents
to explicitly detect the POLLNVAL bit and break before the existing
readiness-mask continue check. Preserve the current handling for POLLIN,
POLLHUP, and POLLERR while ensuring an invalid descriptor cannot re-enter poll
indefinitely.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 29175dc6-149e-484a-84d6-3e6abe11649b

📥 Commits

Reviewing files that changed from the base of the PR and between 0b4c377 and 859762e.

📒 Files selected for processing (1)
  • Sources/AgentForkSupport.swift

Comment thread Sources/SharedLiveAgentIndex.swift Outdated
private let customForkSupportProvider: (@Sendable (SessionRestorableAgentSnapshot, Bool) async -> Bool)?
private let hookStoreDirectoryProvider: @MainActor () -> String
private let dateProvider: @MainActor () -> Date
private let forkExecutableWatchOpenSuspensionForTesting: @MainActor @Sendable () async -> Void

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Test-only suspension hook in production source

forkExecutableWatchOpenSuspensionForTesting carries the ForTesting-suffix name that the no-test-debug-seam-in-production-source rule prohibits. It is stored as a private let property, injected via the production init, and called at line 1628 — but its only non-{} callers are two test sites in WorkspaceForkConversationContextMenuTests.swift. The production code path can reach the suspension point, meaning a stale test fixture could accidentally affect production timing if the wrong injected value is captured. The canonical fix is to widen the surrounding state to internal and read it from the test target via @testable import, or to isolate the suspension point behind a debug-only wrapper rather than naming it ForTesting in shipping source.

Rule Used: Flag Swift files under a production Sources path (... (source)

@lawrencecchen
lawrencecchen force-pushed the feat-pi-fork-context-menu branch 3 times, most recently from ca50106 to 4e6f415 Compare July 17, 2026 08:02
@lawrencecchen
lawrencecchen force-pushed the feat-pi-fork-context-menu branch from 4e6f415 to d08c378 Compare July 17, 2026 08:23
@lawrencecchen
lawrencecchen merged commit f4f420f into main Jul 17, 2026
6 checks passed
@lawrencecchen
lawrencecchen deleted the feat-pi-fork-context-menu branch July 17, 2026 08:47
@lawrencecchen
lawrencecchen restored the feat-pi-fork-context-menu branch July 18, 2026 10:24
ejc3 added a commit to ejc3/cmux that referenced this pull request Jul 25, 2026
manaflow-ai#8173 replaced the fork-probe reuse gate: `!cachedResultHadFallback` became
`cachedResultIsFresh`, and the fallback case is now re-verified against SharedLiveAgentIndex at
the call site instead of being refused outright. The parameter stayed in both signatures, so
these two assertions still compiled while asserting the opposite of what the product does, and
WorkspaceForkConversationContextMenuTests asserts the new contract in both directions a few
files away.

Removes the two assertions whose only purpose was the removed term, and renames the clear-side
test to say what it still covers.
ejc3 added a commit to ejc3/cmux that referenced this pull request Jul 31, 2026
manaflow-ai#8173 replaced the fork-probe reuse gate: `!cachedResultHadFallback` became
`cachedResultIsFresh`, and the fallback case is now re-verified against SharedLiveAgentIndex at
the call site instead of being refused outright. The parameter stayed in both signatures, so
these two assertions still compiled while asserting the opposite of what the product does, and
WorkspaceForkConversationContextMenuTests asserts the new contract in both directions a few
files away.

Removes the two assertions whose only purpose was the removed term, and renames the clear-side
test to say what it still covers.
austinywang added a commit that referenced this pull request Aug 4, 2026
…test host (#9572)

* cmuxTests: derive the theme reload target from a dash-free socket suffix

The CLI derives a theme reload target from the socket file name, collapsing every run of
non-alphanumerics in the slug to a dot. #6452 made this fixture's socket path unique with a raw
UUID to stop two runs colliding in /tmp, which put the UUID's dashes into the derived identifier
as dots, so the expected literal could no longer match and the test waited out its five seconds.
The stdout assertion kept passing because the derived id still has the expected value as a
prefix, which is why this read as a timeout rather than a string mismatch.

Keeps the unique suffix hex-only so the expected identifier stays a plain template instead of a
call into the CLI's own helper, which would agree by construction.

* cmuxTests: drop two palette assertions for a gate that no longer exists

#8173 replaced the fork-probe reuse gate: `!cachedResultHadFallback` became
`cachedResultIsFresh`, and the fallback case is now re-verified against SharedLiveAgentIndex at
the call site instead of being refused outright. The parameter stayed in both signatures, so
these two assertions still compiled while asserting the opposite of what the product does, and
WorkspaceForkConversationContextMenuTests asserts the new contract in both directions a few
files away.

Removes the two assertions whose only purpose was the removed term, and renames the clear-side
test to say what it still covers.

* cmuxTests: stop the remote-connection suite killing its own test host

Three separate problems, in order of blast radius.

Two assertions indexed `operations` right after asserting its count. A count assertion does not
stop execution, so on failure the next line trapped with Index out of range and took the shared
test host down, and every remaining test in the shard never ran. Measured twice in one run.

Four @mainactor tests waited on a DispatchSemaphore. configureRemoteConnection enqueues its
session transition as a main-actor Task, so blocking the main actor stopped the very work being
waited on from ever being scheduled. They now use expectations, which pump the run loop.

Fifteen fixtures passed an unresolved %C control template. The broker deliberately refuses to
own a path it cannot resolve, so no lease was ever taken and cleanup could not run; six inverted
expectations were passing vacuously as a result. They now use the resolved form ssh -G produces,
and a new test pins the unowned-template policy so the fixtures cannot quietly regress to it.

Two more read activeRemoteSessionControllerID straight after configureRemoteConnection and now
await the transition instead.

* cmuxTests: point the daemon-upload tests at the transport that replaced scp

Two tests waited on an scp invocation that no longer happens. #8434 moved the daemon upload off scp
and onto the ssh exec channel, streaming the binary into `cat >`, and did not touch these tests.
Their stubs only fulfilled inside an `executable == "/usr/bin/scp"` branch, so the expectation
could never fire, the wait spent its whole budget, and the unwrap on the next line reported nil.

Both now capture the upload from the ssh branch. The property each one is about is unchanged: the
daemon still has to land on an absolute path under the remote HOME, that path just travels inside
the remote command instead of an scp destination, so the assertion moved with it.

The scp branch is kept and fails loudly. If the upload ever returns to scp, that should be a
sentence in the failure output rather than a silent timeout, which is precisely how these two broke.

The reinstall test also now records how many capability hellos preceded the upload and requires at
least one. Retargeting alone would have let it pass on a first install, which is not the
missing-pty-capability path it is named for.

Renamed the first test off "ScpDestination" since it no longer describes what is asserted.

* cmuxTests: fix three CLI tests that could not pass, and stop one hiding why

Three separate causes, all in the fixtures rather than the product.

Two socket-selection tests replied to the CLI with a bareword. SocketClient only treats OK, OK …,
PONG, ERROR: … or JSON as a complete single-line reply, so a bareword sends it into the multiline
drain pass, where reconfiguring the receive timeout on a socket whose peer already hung up fails with
EINVAL — and the CLI reports "Invalid argument" instead of the reply it already had buffered. The
replies are now OK-framed. These were the only two barewords in the suite, which is why eleven
near-identical siblings pass.

Both now also assert which responder received the request. That is the property they exist for —
the tagged socket is chosen and the stable one is not — and unlike the stdout comparison it cannot
be made vacuous by a future change to the reply.

A fork-diagnostics fixture passed agent "project-agent", which is not in the CLI's catalog, so the
command exited before emitting any JSON. The test has never passed; it went in already red alongside
the pi-family gate it is meant to cover. It now uses grok, a catalog agent that is neither pi-family
nor one of the transcript-walking agents, so the basename gate is still what is under test.

The shared helper turned all of that into a JSON decoding failure, because it only expected a zero
exit before parsing. It now requires the exit status and a completed run, so the next fixture mistake
reports the CLI's own error text instead of a parse error.

* cmuxTests: pair the pi-basename fixture with an agent that can actually fork

The pi-family basename test asked for fork_command_available, fork_supported and
fork_startup_input_available, but its fixture stored the record under a grok
launcher pair. A captured launch command is only used when its launcher describes
the requested agent, so the grok/omo pair was dropped as untrusted, no fork argv
was built for any agent, and all four assertions failed on
agent_has_no_fork_command without ever reaching the rule under test.

Store the record under opencode instead, whose wrapper launcher is omo. The
capture is now trusted, the fork argv resolves through the omo launcher, and the
executable basename stays /tmp/pi so the disagreement between the structured
identity and the basename is still what the test measures. The omo launcher also
answers fork support before the opencode executable probe, so the result does not
depend on a /tmp/pi existing on the machine running the test.

* cmuxTests: assert the stderr-closed CLI does not crash, instead of a CLI that no longer exists

This test asserted exit 1 and a "Usage:" banner on stdout. Neither has been true since #f48922aa94:
an unknown command exits 2 with a single line and no usage dump, and that line goes to stderr — which
the test closes with 2>&-. So it could not pass, and the crash it was written for was not what it
checked.

The regression is still worth guarding. cc4a610 replaced FileHandle.standardError.write, which
raises and aborts when stderr is closed, with a raw Darwin.write that returns -1 on EBADF. The oracle
is therefore that the CLI exited on its own terms rather than dying from a signal, so ProcessRunResult
now carries terminationReason and both runners set it. Without that, a signalled process is
indistinguishable from an ordinary non-zero exit, because its terminationStatus is just the signal
number.

The command now runs under exec, so the process being waited on is the CLI rather than the shell. A
shell reports a signalled child as a normal exit with status 128+signal, which would have hidden
exactly the crash being tested.

It also pins CMUX_SOCKET_PATH and the home directory. Socket resolution otherwise consults a
machine-global marker file, and a spawn with a pristine temp home was measured reaching a real running
app — which would make the exit code depend on what is running on the machine. With the socket pinned
the unknown-command path is a single branch, so the test asserts exit 2 exactly rather than settling
for non-zero.

* cmuxTests: isolate the CLI regression suite from the machine's own cmux

A CLI spawned from this suite with a pristine temp home and a scrubbed
environment still reached a real running app. CFFIXED_USER_HOME moves the socket
directory but not socket discovery: the CLI also reads the machine-wide
/tmp/cmux-last-socket-path marker, and for an untagged debug build it scans /tmp
for cmux-debug-*.sock and connects to what it finds. Resolution runs before the
command dispatches, so even `claude-teams --help` did this. Every spawn site that
is not itself testing resolution now pins CMUX_SOCKET_PATH to a per-run path, the
three stable-variant tests write the marker inside their own temp home, and
runShell takes an explicit environment instead of handing the child everything
the test host was launched with.

Two tests bound a responder on /tmp/cmux.sock, the release app's socket path, and
UnixSocketResponder unlinks before it binds, so a run could take the control
socket away from a release app in use. The early returns meant to prevent that
raced the app, disagreed about whether a dangling symlink counts as present, and
turned the tests into silent passes. The symlink fallback case moves to the
user-scoped stable path inside its temp home. The legacy case keeps the part that
needs the real path, that /tmp/cmux.sock is classified as a stable implicit
default, and no longer creates, binds, or removes it. Three more guards tested
paths inside a freshly created temp home and could never fire, so they are gone.

stderr was pointed at the stdout pipe while about thirty tests parse stdout as
JSON or compare it to an exact reply, so one diagnostic line from the runtime
broke a content check instead of naming itself. stderr now has its own pipe,
failure messages carry both streams, and the negative checks that meant "the CLI
never said this anywhere" read both rather than silently narrowing to stdout.
Readers for both pipes start before the wait, because reading after
waitUntilExit deadlocks once a child fills a pipe buffer and that looks like a
hang inside the CLI. A launch failure is reported on stdout as well as stderr,
since five sibling suites share this runner and print only stdout.

Runs that assert nothing about latency no longer carry a 5s cap and take a 60s
guard instead, which still fails a stuck CLI rather than passing slowly. The two
browser-download tests keep their 3s and 16s caps, where the deadline is the
assertion. The two theme tests with fixed bundle identifiers now scope them per
run, since the reload notification goes out machine-wide; for the nightly one
that means scoping the socket file name too, because the identifier is derived
from it.

* cmuxTests: assert the exit code this fixture actually produces

The stderr-closed test asserted exit 2, the unknown-command code. Measured, it exits 1: the pinned
socket has no listener, so the CLI fails at connect and the top-level handler returns before the
unknown-command arm runs. That ordering makes the fixture a better exercise of what the test guards,
not a worse one, because the connect error is written to the stderr the test has closed. The run
confirmed the guard itself holds — termination reason was a normal exit, not a signal.

* cmuxTests: report stderr in sessions helper failures

* cmuxTests: preserve restore assertions after stream split

* cmuxTests: close review gaps in process and upload fixtures

* cmuxTests: align remote fixtures with streamed input and scoped identity

* cmuxTests: yield main actor while awaiting daemon upload

* cmuxTests: repair CLI regression fixtures and child lifetimes

* cmuxTests: isolate daemon bootstrap fixtures from ControlMaster

* cmuxTests: keep theme notification state nonisolated

* cmuxTests: detach live argv fixture from test host

* cmuxTests: own Go discovery in daemon reinstall fixture

* cmuxTests: make subprocess and bootstrap fixtures deterministic

* cmuxTests: remove detached fixture wall clock

* cmuxTests: use async-safe scoped locking

* cmuxTests: make off-host process work concurrent

* cmuxTests: keep blocking process wait off cooperative executor

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
austinywang added a commit that referenced this pull request Aug 27, 2026
…dLiveAgentIndex

Since #8173 the convenience supportsFork(snapshot:) probes with a throwaway
cache; the probe cache is owned by the caller. Share one cache and resolver
across both probes so the cached unsupported verdict is what is observed, and
show a fresh cache probes again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
austinywang added a commit that referenced this pull request Sep 21, 2026
…y, and shortcut routing

Merge 8d33410 on the issue-2824 branch took main's copy of
`cmuxTests/WorkspaceUnitTests.swift` wholesale and discarded the branch's
September repairs (c8bfb58, c402b9c, 2127d97); PR #12053 then landed
without them while the strict gate started counting the failures. Restore
them against current production, plus the shortcut-suite fixes:

- Fork in a remote workspace: `sshBootstrapArguments` has used `/usr/bin/ssh`
  since #9114, and agent socket propagation requires a socket that exists on
  disk (3bf87e3), so bind a real unix socket and expect the absolute path.
- Fork Conversation context actions dispatch asynchronously (#7259, #8173);
  await the fork panel before asserting.
- Git branch and pull request updates publish through
  `sidebarObservationPublisher` since #6226, not `objectWillChange`.
- Config sanitization: `addWorkspaceIfActive` rebuilds templates from the
  font-size lineage since #8543; override the lineage hook instead.
- Focus recovery: AppKit focus is authorized only for a registered, selected
  workspace whose window carries the main-window identity; use the
  `TerminalPortalTestWorkspace` fixture and the same registration/pump path
  as `WorkspaceTerminalFocusRecoverySwiftTests`.
- Shortcut routing: clear both Cloud defaults after 9c2ba78 swapped them;
  neutralize Ghostty's imported goto_split fallback (⌘] by ANSI keyCode) in
  the unshifted-symbol test and assert the digit shortcut does not match;
  wait for the async runtime start before judging keyDown forwarding (skip
  loudly without a live surface); wait for the portal to mount before the
  second-Escape check.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
teamleaderleo added a commit that referenced this pull request Sep 23, 2026
* test(terminal): restore the presented-surface fixture contract so package tests compile

`swift test --package-path Packages/macOS/CmuxTerminal` has not compiled on
main since merge 38b32bb: TerminalSurfaceRendererCallbackTests calls
`PresentedSurfaceFixture(installRendererCallbacks: false)`, but the fixture
initializer only takes `windowVisibleAtCreation`. The flag came from the
issue-2824 branch (77c46d0, 6fe70f6) and was dropped when the fixture
was reworked for the native callback lifecycle (8008b7c, 97c6088);
the merge re-applied the test call without the fixture side.

Package tests only register render callbacks through the fixture
(RendererCallbackTestSupport), so a fixture that skipped registration would
leave `cmux_test_ghostty_renderer_present` with nothing to route. Restore
the calls to `PresentedSurfaceFixture()`, the combination that was green at
0251a44. Verified locally: 294 tests in 37 suites pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(popover): stop rewriting the presentation binding during view update

`ArrowlessPopoverAnchor.updateNSView` calls `coordinator.dismiss()` on every
update where `isPresented` is already false. PR #13311 (01f0a50) made
`dismiss()` write `isPresented = false` on the no-popover path, which SwiftUI
reports as "Modifying state during view update" because updateNSView runs
inside the view update. The sidebar footer mounts two anchors whose parents
re-evaluate on every `selectedTabId` change, so every workspace switch emitted
faults and `SidebarWorkspaceSwitchLayoutFaultTests` failed with 15 of them.

Pass `resetPresentation: false` from updateNSView: the binding is already
false on that branch, so the write was redundant. `popoverDidClose` still
resets the binding when AppKit closes a live popover.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(cloud): satisfy the plus menu's sign-in gate and route Cmd+Y through a registered window

`NewCloudWorkspaceShortcutTests` never passed: the plus menu deliberately hides
its Cloud rows unless the account is signed in (#12305, mirroring the command
palette and File menu), and a test `AppDelegate()` has no account flow. The
XCTest version crashed the app host on `rows[1]`, xcodebuild restarted it, and
the lenient gate accepted the partial run; #13178 now rejects that and #13193
migrated the suite to Swift Testing, so the failures became visible.

Inject the signed-in state through the existing `isAuthenticated:` seam, add
signed-out coverage of the gate, route the Cmd+Y event through a registered
main window (as `testReboundKeyRoutesAndOldKeyDoesNot` does, since shortcut
routing bypasses events bound to windows the delegate cannot resolve), and
register a window context for the unavailable-Cloud check.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(chrome): assert the composited Bonsplit chrome contract instead of a stale alpha hex

`WorkspaceChromeColorTests` expected `bonsplitChromeHex` to return
`#1122337F` (theme with opacity as alpha). Since b6d3470 (2026-05-19)
`compositedTerminalColor` composites the theme over the window base and
returns an opaque color, and 4cbb354 made that deliberate: Bonsplit derives
its tab glyph contrast from the rendered backdrop, and an `#RRGGBBAA` hex
reproduces the white-on-white bug it fixed. The tests kept failing unnoticed
because the app-host gate only counted "unexpected" XCTest failures until the
strict check (acedf3f) reached main through #12053.

Exercise the `chromeBackgroundColor` seam that every production call site
uses with literal expectations, verify the ambient default path against the
resolver plus independent blend arithmetic so the result is deterministic
under any host appearance, and keep coverage for opaque, shared-backdrop,
pane-clear, pane-border, and explicit translucent chrome colors.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(tmux): wait for the main window to adopt the mirror pane before sending keys

`RemoteTmuxMirrorPaneInputMappingTests` required `panel.surface.uiWindow != nil`
right after selecting the workspace. A manual-I/O mirror pane spawns eagerly in
its hidden bootstrap window, which `uiWindow` deliberately excludes, and the
main window's portal adopts the pane host only once AppKit and SwiftUI get
run-loop time; `waitForLiveSurface` returns immediately for an already-live
surface, so the check ran before adoption and the four key-delivery tests
failed at line 176. The failure was hidden until the strict app-host gate.

Order the harness window front and pump the run loop until the surface and
its native view are in that window, as the other hosted-view input suites do,
and assert against the harness window instead of any window.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(browser): report the refused connection through the navigation delegate

`browserPanelRetriesDiscardedRestoreAfterConnectionRefused` waited up to 30s
for WebKit to fail a provisional load to a bound-but-unlistened loopback port.
On the hosted app-host runners that failure never arrives: the load neither
fails nor commits, so the test timed out at line 147 (no
"provisional navigation failed" log line appears for it in any shard).

Stop the in-flight load and report `NSURLErrorCannotConnectToHost` for the
attempted URL through the panel's real navigation delegate, the pattern
`BrowserFailedNavigationReloadTests` already uses. The restore bookkeeping,
error page, `restore_pending` clearing, and retry policy under test run
unchanged and every assertion is kept.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(cloud): scope reserved-workspace cleanup to its pending card and return the bind window

PR #13202 (5d616a7) imported `CloudMachineWorkspaceAdoptionTests` and
`CloudMachineWorkspaceResolutionTests` from the still-open
13141-cloud-vm-workspace branch without the production changes they assert,
and its own run skipped the app-host shards, so they landed red:

- `NewMachineSheetPresenter.closeReservedWorkspace` closed the whole
  workspace, which is a no-op for the last tab and discards user panes added
  next to a creating card. Port the branch behavior: remove only unadopted
  loading cards owned by the cancelled machine, clear that binding, keep user
  content, and give a last-tab loading workspace a local anchor first.
  `MachineCreateCoordinator` passes the created or reconciling machine id so
  cancelling machine X cannot clear a binding to machine Y (the tests now
  assert that scoping).
- `v2WorkspaceCloudVMBind` now returns `window_id` alongside the workspace
  refs, as the bind acknowledgement test expects.
- The resolution test selected "first" while the fixture's terminal key is
  "term-first"; use the key so the placement resolves as intended.
- `CloudPaneCreationRetryTests` asserted `discarded` synchronously after the
  projection returned, but the coordinator applies its generation fence behind
  `CloudOperationContext.withPhase`'s recorder await, which suspends on a cold
  per-suite process. Settle with bounded yielding before asserting.

Runtime behavior change (cancelled Cloud create cleanup); needs dogfood.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(cli): restore deferred socket connection, explicit SSH control options, and Codex stop idle observations

Branch commit c8bfb58 ("fix: repair failures exposed by strict app-host
CI", 2026-09-10) made `cmux vm dev|layout|env` finish local validation and
dry runs before opening the socket, passed the caller's explicit ssh options
into `userConfiguredControlOptions(fromSSHConfigOutput:explicitOptions:)` so a
normalized `ControlPersist=0` from `ssh -G` is not mistaken for host
customization, and published `idleObserved` for transcript-terminal prior
turns before a legacy Codex Stop so the journal drops an obsolete running
turn. The branch's later single-parent commit 38b32bb reverted
`CLI/cmux.swift` to main's version while keeping the stricter fixtures, and
PR #12053 landed that inconsistent state; the fixtures (CLIVMDevTests,
CLIVMLayoutEnvTests, the SSH sharing tests, and the Codex missed-prompt Stop
test) have failed since.

Restore the three CLI changes. `SocketClient.configureAuthentication` and
`SocketPasswordResolver` already exist on main, and the CmuxFoundation
overload landed with the branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(cli): align CLI integration fixtures with the shipped hook, SSH, and session-list contracts

The main copy of `CLINotifyProcessIntegrationRegressionTests` predates
several shipped contracts that the lenient app-host gate never enforced and
that 38b32bb reverted on the issue-2824 branch: Claude hook acks print
`{}` (#7963), Codex resume bindings require rollout evidence so fixtures
carry a `session_meta` transcript (#10100), SessionStart publishes a binding
so `/clear` counts start after the clear, fresh-terminal SSH startup commands
are script paths that the support decoder must read, cmux control-path
options follow a resolved `ssh -G` (#8308), and `ssh session list` reports the
localized "remote state unavailable" summary with `--json` detail (#9971). The
missed-prompt Codex Stop test now asserts the restored `agent.idle.observed`
for the prior turn.

Flagged for review: the three `ssh pty-attach` cases now expect only
`workspace.remote.pty_bridge` once the endpoint is established, matching
#12726's `preserveLifecycleForRecovery`; if pre-READY bridge failures were
meant to keep reconciling, the flag should be set only after READY instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(workspace): restore the app-host repairs for fork, focus recovery, and shortcut routing

Merge 8d33410 on the issue-2824 branch took main's copy of
`cmuxTests/WorkspaceUnitTests.swift` wholesale and discarded the branch's
September repairs (c8bfb58, c402b9c, 2127d97); PR #12053 then landed
without them while the strict gate started counting the failures. Restore
them against current production, plus the shortcut-suite fixes:

- Fork in a remote workspace: `sshBootstrapArguments` has used `/usr/bin/ssh`
  since #9114, and agent socket propagation requires a socket that exists on
  disk (3bf87e3), so bind a real unix socket and expect the absolute path.
- Fork Conversation context actions dispatch asynchronously (#7259, #8173);
  await the fork panel before asserting.
- Git branch and pull request updates publish through
  `sidebarObservationPublisher` since #6226, not `objectWillChange`.
- Config sanitization: `addWorkspaceIfActive` rebuilds templates from the
  font-size lineage since #8543; override the lineage hook instead.
- Focus recovery: AppKit focus is authorized only for a registered, selected
  workspace whose window carries the main-window identity; use the
  `TerminalPortalTestWorkspace` fixture and the same registration/pump path
  as `WorkspaceTerminalFocusRecoverySwiftTests`.
- Shortcut routing: clear both Cloud defaults after 9c2ba78 swapped them;
  neutralize Ghostty's imported goto_split fallback (⌘] by ANSI keyCode) in
  the unshifted-symbol test and assert the digit shortcut does not match;
  wait for the async runtime start before judging keyDown forwarding (skip
  loudly without a live surface); wait for the portal to mount before the
  second-Escape check.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(restore): keep a restore identity when the persisted surface id collides

Since #13098 (e0e77eb), `newTerminalSurfaceOutcome` treats a nil
`restoredSurfaceId` as an interactive create and routes it to the selected
pane's Cloud source. Session restore passed nil whenever the persisted panel
id was still live (duplicate-workspace or restore-into-live), so a legacy
managed-Cloud SSH workspace restored into a live manager was turned into a
remote tab create: the scaffold panel stayed startup-suppressed with no
initial command and `TabManagerSessionSnapshotTests` failed to unwrap it.

Mint a fresh UUID on collision instead of nil; it is free by construction, so
the restore keeps its identity and the old-to-new remap works as before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(terminal): align snapshot, projection, terminal, and browser fixtures with shipped contracts

Stale expectations and fixed-spin timing in app-host suites that the lenient
gate never counted:

- Cloud-projected panes restore as manual-mirror reservations with a staged
  remote identity (#12675), not a local placeholder projection.
- Restore reuses the persisted runtime id when free (aff0e32), so assert
  liveness rather than a new id; the catalog ignores writes for a Cloud
  machine without a registered provider (#11877), so register the fixture
  provider before publishing.
- Wheel sync requires an authoritative scrollbar response (bbc3edf); the
  fixture now answers like `AuthoritativeScrollbarSurfaceView`.
- Search overlay mount, first-responder focus, runtime creation, and the
  visibility-restore redraw run through deferred main-actor tasks; wait for
  them with the class's `waitUntil` helpers instead of fixed run-loop spins.
- Owning the socket path lock is definitive (959f38a): a refused inode left
  by a dead listener is replaced, so the restarted listener accepts.
- The split-divider hit band extends `dividerHitExpansion` past the divider
  (667cc43); derive the pass-through boundary from the constant.
- Browser page background blends against Ghostty's effective terminal color
  scheme, which is host dependent; read the same preference the product uses.
- Cloud Machines defaults on in dev builds (#12318); pin the toggle off for
  the default-mode palette contract.
- Prepared navigation requests keep the caller's cache policy (#13003).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(browser): align lifecycle, identity, host-view, loopback bridge, and portal rebind tests with shipped contracts

Browser app-host suites that the lenient gate never counted:

- `BrowserPanelWebViewLifecycleTests`: out-of-range discard delays are
  rejected to the default (the cmux.json loader relies on the nil), and the
  panel's own `isLoading` stays true for the indicator floor, so wait for both
  flags and assert no discard blockers before discarding.
- `browserNavigationUsesEmbeddedWebKitIdentity`: WebKit reports the native
  identity as nil or "" (#9482); accept either.
- `WindowBrowserHostViewTests`: production routes Dock-divider hits by
  yielding to AppKit so the live sidebar tracker receives them (#10902,
  e2e3818); the tests asserting the portal owns and forwards the hit were
  merged red against a design that never shipped. Realign the stale-frame
  test to the pass-through contract and remove the four own-and-forward
  tests with their fixtures. **Flagged for review:** if own-and-forward is
  still wanted, that is a hit-testing product change for its own PR.
- `testRemoteWorkspaceRuntimeBridgeAliasesMultipleLoopbackPortsFromSamePage`:
  the navigation delegate restarts main-frame loads to apply the user-agent
  policy (11c6efe); a direct `loadHTMLString` with an HTTP base skipped
  that step, so the data load was cancelled and replayed as a deferred
  request. Apply the identity first and assert the document is current.
- `portalRebindPreservesDocumentAndRoutesRefreshToTheSameWebView`: the portal
  host is a theme-frame sibling of `contentView` for non-glass windows
  (#12929); assert same-window instead of descendant-of-contentView.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(terminal): wait for asynchronous runtime creation in the remaining direct-interaction tests

The same "Expected runtime surface before ..." precondition that
9f258dd made wait for the deferred runtime start still sampled after a
fixed run-loop spin in the detach-race, close-lifecycle, repeat-key, and
repeat-IME tests, and CI on the branch head showed them failing that way.
Use the class's `waitUntil` for those four sites too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(cli): model surface.respawn for respawn-pane and give the Codex stack fixture rollout evidence

`cmux respawn-pane` has sent `surface.respawn` instead of `surface.send_text`
since #5465 (8cafcc3); the window-flag fixture's mock still rejected that
method, so the CLI exited 1. Answer `surface.respawn`, asserting the window and
surface ids, `tmux_start_command`, and that the shell-invoked command carries
the user command but never the `--window` flag, which is the test's intent.

The Codex interrupted-stack fixture's transcript had no `session_meta` line,
so `CodexSessionResumeVerifier` (#10100) found no rollout evidence and the
prompt published `surface.resume.clear`. Prepend the session line as the
other Codex fixtures do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(socket): keep mobile.panel.artifact.fetch off the local socket and advertise the served artifact reads

02c1ba4 removed `mobile.panel.artifact.fetch` from the socket worker
methods because it needs the authenticated mobile execution context, but
that half of the change was lost in a merge: the policy still routed fetch to
the worker lane, which has no handler for it, so the local socket answered
`internal_error` instead of the `method_not_found` boundary that
`TerminalControllerSocketSecurityTests` pins. Restore the removal.

`system.capabilities` never advertised `mobile.panel.artifact.stat` and
`.thumbnail` although both are served on the worker lane (04ff18e added
them only to the test's expectation); advertise those two and keep fetch out.
The remote relay allowlist is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(remote): align tmux seed transport, SSH, socket command, and port scanner fixtures with shipped behavior

- `RemoteTmuxPaneSeedTransportTests`: manual-I/O mirror panes spawn eagerly
  (#9272) and stay runtime-backed after portal churn (#9769), so the pane
  renders its assigned grid before a seed arrives and nothing was retained.
  Publish a larger tmux pane grid than any test surface applies so the
  retention under test sees the lag it exists for.
- `SSHRemoteCWDRegressionTests`: the persistent-PTY exec helper runs only
  with `protectsFromHangup: true` (cd9f34f); pass it, and widen the
  first-spawn guard.
- `SSHDeepSleepReattachTests`: a Cloud-owned workspace rejects launch
  overrides on splits (#13098); create the custom-identity pane before
  configuring the remote connection.
- `SSHConfiguredRemoteCommandHostTests`: ssh-pty-attach validates the bridge
  `daemon_version` before dialing (#12726); the mock now reports one.
- `CloudManualMirrorTransportTests`: the pane failure card uses the short
  title since 9bf6cb8.
- `SurfaceSocketCommandTests`: `vm.workspace_new` admits its optimistic
  workspace through the active main window (#13152, #13155), so bind a bare
  window to the fixture context; a receipt without a starter terminal costs
  one snapshot (6d43ea6).
- `PortScannerPublicationTests`: the forced-result acknowledgement hops off
  the main actor, so an unchanged port set may be deduped under a later
  refresh; drain publications until the retirement lands.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: align mobile artifact fetch lane assertion

* repair: close remaining full-suite contracts

* test(restore,sidebar): align two stale contracts with shipped behavior

testRemoteWorkspaceAutoResumeKeepsRemoteStartupCommand asserted the restored
remote panel's local spawn cwd equals the remote-host path. It never can:
OneShotTerminalLauncherStore.enterableWorkingDirectory rejects a path that is
not locally enterable, which is what keeps the owning shell valid (#7031).
The remote cwd does survive restore — as the panel's trusted remote directory
report, and as the `cd` prefix the resume input carries (both already asserted).
Assert it where it actually lives and pin the spawn cwd to nil.

testSidebarPullRequestsTrackFocusedPanelOnly expected a background panel's PR
to be hidden from sidebarPullRequestsInDisplayOrder(). That list is documented
as the workspace's deduplicated rows in pane/tab order, both consumers
(taskStatusSignals, the control-sidebar snapshot) want every panel, and the
sibling branch test asserts the same all-panel model. Focus scoping lives in
the `pullRequest` binding, which the test already covers. Renamed to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: give Cloud catalog tests live destination workspaces

#13196 made SurfaceCatalog.validateOwnership refuse a destination that is
not a live workspace. That rule stays. Tests that projected, restored or
opened browsers into a made-up workspace ID now failed with
destinationNotFound, timed out waiting on a provider that was never called,
or passed a later check for the wrong reason.

LiveWorkspaceFixture registers real Workspace objects and hands the catalog
a CloudWorkspaceRenameService that resolves them, the way the app's
composition root does. Its workspaces() list stays empty so the native
projection coordinator does not start mirroring into them. For tests on
SurfaceCatalog.shared, withAppRegistration registers the TabManager as a
windowless main-window context for the test body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: normalize pbxproj after live workspace fixture

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: admit agent renames of agent-owned accepted Cloud names

e6926fb (#13403) made submitCloudPanelRename run admitsTerminalRename
for automatic names. Its last clause only admitted replacing an accepted
name when the local panel still carried `.auto` provenance, but since
1e1d319 an accepted daemon name reconciles locally as `.remote` and the
owner lives in the tab's nameAuthority. Every agent title after the first
accepted one was refused, so "Failed agent rename keeps the accepted title",
"Mirroring an agent-named placement does not block its next agent title" and
"An older automatic result and old snapshots cannot replace an accepted name"
failed on their second agentName call.

Admit an automatic rename when no write is pending and the accepted tab name
is owned by the daemon's `auto` authority. User-owned names, pending writes and
local user titles on any projection still refuse it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: expect adopted machine rows to keep the pending create identity

#12919 (021f792) made New Machine creation optimistic: once a running
create's machine appears in the fleet list or catalog, its row keeps the
`pending-machine:<operation>` node ID so selection and expansion survive
adoption (see committedProjectionKeepsThePendingNodeIdentityUntilFleetAdoptsIt).
pendingRowStepsAsideOnceItsMachineHasARow still expected `machine:<id>`.

Assert the new identity and that the row is the adopted machine, not a
stand-in: the stand-in is gone, one row remains, and it shows the created
machine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: restore per-display noVNC targets and refresh stale Cloud tests

Product (#13196 regression): the 6db7593 main merge into
13192-cloud-display-ownership took main's CmuxTuiSurfaceProviders.swift and
CloudPortRoutePlan.swift, dropping eb3edd7's per-display ports.
23c807b restored the display coordinator but not these hunks, so
withPrivateBrowserURL rewrote every display to 6901 and a daemon pointer
without a discovered target fell back to display 1. Restore both hunks and
drop the duplicate port-less privateDesktopURL overload.

Tests:
- CloudDisplayCatalogTests: 178d35e (#13196) made every guest command run
  `list` as a readiness probe, so fakes dispatching on " list" answered
  creation with the list catalog. Dispatch on the create action line, and pin
  the command shape.
- CloudPortOpenRegressionTests: 178d35e (#13196) filters RFB/noVNC ports
  only when the display catalog owns them (displayPortsOwned). Assert both
  the desktop and non-desktop results.
- CloudTreeOneMachineManyWorkspacesTests: #12740 added a final Resources
  section under each machine. Expected trees include it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: publish every guest display from daemon graph updates

The 6db7593 main merge into 13192-cloud-display-ownership (#13196) put
back main's [desktopDisplayResource()] pools in CmuxTuiSurfaceProvider's
refresh, publish and delta paths. 23c807b restored the display
coordinator lifecycle but not these pools, so a display created beyond
display:1 vanished on the next daemon publish, and a delta that touched it
removed it. Restore eb3edd7's displayResources pools, the display-kind
delta check, and the injectable displayCoordinator. Drop the now-unused
desktopDisplayResource().

Adds a test that creates display:2 through the provider and asserts that
both displays and their ports survive a full publish and a display delta.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: fence the fake Cloud projection reply with its mutation cursor

aDetachedTerminalDropsItsStaleTabBeforeMoving installs a daemon graph
since c402b9c (#13403). When the placement lane drains it reconciles
against that graph, which predates the projected tab, and clears
tab_projected. The real reply always carries a mutation cursor
(CmuxTuiSnapshotParser.placedTab requires one) that fences exactly this;
the fake returned none. Give the fake a projectCursor and set it ahead
of the installed graph.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: register the restored window before relinking Cloud projections

#13196 made SurfaceCatalog.validateOwnership refuse a destination the app
cannot resolve, so SurfaceCatalog.shared no longer relinked a restored
projection into a TabManager that no main window owns. Register it for
the test body with LiveWorkspaceFixture.withAppRegistration, as #13651
does for CloudClosedPanelRestoreTests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: wait for the relay ports kick, not the first relay line

Since #8442, a relay prompt also reports shell state. Both RPCs run in
separate background children, and the zsh prompt-refresh test waited
only until the log was non-empty, so it could read report_shell_state
alone. Wait (deadline-bounded) for the ports_kick line itself, in the
zsh test and its bash sibling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): relay output of an SSH session that ends right after auth

The expect wrapper watches the first two seconds after sending the
password for a rejection with log_user 0. A session that authenticates
and exits inside that window hit the eof branch and its output was
dropped. Flush the buffered output before exiting. #13207 replaced the
test's 9 s sleep with a FIFO release, which is what exposed this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: restore the no-connection path for rejected vm dev input

4120e35 taught runVMDev that invalid input never connects: keep the
listener open, then check its backlog after the CLI exits. #13207 dropped
that again, so each rejected run waited 60 s for a mock-server
expectation that only the listener closing (after the wait) fulfills.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: drain the terminal while the SCP host-key failure runs

The pty case read the master only after the CLI exited. A pty's output
queue holds about 1 KiB and the host-key failure report is longer, so the
CLI blocked writing stderr until the 30 s timeout. Read the master on a
thread while the CLI runs. The test starts its own sshd; no runner host
dependency is involved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: fail fast when a gated Cloud call ends before its fake is entered

Every app-host restart in the 09-21/09-22 shard logs I sampled was a Swift
Testing time-limit hit in one of two suites, and each hit relaunches the
test host:

- SurfaceCatalogTests: when catalog.project threw before reaching the
  provider (destinationNotFound, fixed by 3dc91b2), each gated test
  parked in MaterializeGate.waitUntilEntered() until the 300 s limit.
  Five tests restarted the host one after another, about 25 minutes per shard.
- CloudDisplayCatalogTests: when the fake no longer recognized the create
  command (fixed by 70eaf44), create() failed before the exec started
  and `await started.result` parked until the 60 s limit.

The waits now also end when the caller's task finishes, so the next such
setup failure is an ordinary failed #require. The cancellation test also
waits on its own cancellation signal instead of a 60 s Task.sleep.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: reveal a retired terminal through the portal rebind, as the app does

Since #12607 hiding a terminal removes its hosted view from the window, and
only a bind reinstalls it. The test flipped portal visibility on the detached
view, so no size commit could ever land and the final shrink check failed
every run. Rebind like TerminalPortalReconciliation and drop the wait loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: pin key-window status in terminal focus suites

The app-host test process runs headless and is usually not the active app,
so makeKeyAndOrderFront never makes a programmatic window key. Terminal focus
paths gate on isKeyWindow (automatic first-responder apply, focus redraws,
deferred focus reapply, ensureFocus window activation), so these suites
passed only when an earlier test in the shard had activated the app.

Share the existing KeyStatusTestWindow and use it in
WorkspaceTerminalFocusRecoveryTests, WorkspaceTerminalFocusRecoverySwiftTests
and TerminalNotificationDirectInteractionTests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: report which layer holds off-plan tmux mirror geometry

offPlanGeometryWithUnchangedSizingInputsReconverges fails every run with all
three output-parity re-arms spent and the hosted view still at the perturbed
0.8 divider. A standalone bonsplit replay of the same sequence converges, and
a sibling test without a bound (portal-visible) workspace heals the same
displacement. Add the live split view's arranged widths and the split model's
imposed extent to the failure message so the next run shows whether bonsplit
refused the apply, the imposition was cleared, or the portal did not follow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: restore Cloud projections before the workspace is published

TabManager.restoreSessionSnapshot restores each workspace's surface
projections before it assigns tabs, so SurfaceCatalog.restore's ownership
check could not find the destination and silently dropped every restored
remote projection. Check ownership against the workspace being restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: repair stale and broken app-host expectations

- CLI ssh attach: count identity UUIDs exactly; #11497 added an auth token uuidgen
- device mirror directories: give the fake device trusted presence (e3d424c gate)
- font zoom mirrors: deltas from the 8pt inherited base (e6926fb arithmetic)
- Computer Use refresh: fake daemon reply was invalid JSON in a raw string (#13599)
- quit alert: compare button alignment rects, not padded frames
- remote split cwd rescue: cwd travels as CMUX_REMOTE_INITIAL_CWD since #12054
- default freestyle split: Cloud-owned splits route to Cloud since fa5dc4c

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: activate the app host before waiting for terminal focus

`focusTerminalForTesting` waits for `window.isKeyWindow` before it hands
first responder to the terminal, but `makeKeyAndOrderFront` only makes a
programmatic window key while the test host is the active app. The
app-host process starts inactive under `xcodebuild test`, so the wait
succeeded only when an earlier test in the shard happened to activate the
app. Its callers use a real `createMainWindow()` window, which cannot be
swapped for `KeyStatusTestWindow` the way the pinned focus suites were.

Activate explicitly, and give the wait a CI-appropriate timeout: the
activation and the key-window transition land on later main run-loop
turns, and the pump's one-second default expires before the window goes
key on a contended runner.

Observed on run 35743588302: `plainTerminalTextDoesNotResolveAppShortcutContext`
(shard 1/6), `keyboardCopyModeKeyClearsTerminalUnread` and
`workspaceFontSizeShortcutPreservesBackgroundTerminalUnread` (shard 2/6)
all fail at this helper's return value, and shard 6/6 shows the same
process state directly as `NSApp.keyWindow -> nil`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: give standalone terminal fixtures a live portal authority

`setVisibleInUI` and `setActive` fold their request through
`Workspace.portalRenderingEnabled(for:)`, which denies any workspace id
the app delegate cannot resolve to a selected tab. Three direct
interaction tests build their surface with `tabId: UUID()`, so once any
earlier test installs an `AppDelegate.shared` the authority denies the
portal, the hosted view is never actually made visible or active, and the
fixture stops exercising the behavior it asserts: the surface never takes
Ghostty focus and never schedules a visibility-restore redraw.

Register a real selected workspace and build the surface with its id, so
the fixture gets the same authority the app grants the selected tab.

This is the fixture, not the product: the authority check is deliberate,
and denying an unresolvable workspace is what keeps queued portal
callbacks from reviving an inactive workspace.

Fixes on run 35743588302 shard 4/6:
`testKeyDownRecoveryDoesNotReplayFocusAfterResponderMovesAway`,
`testVisibilityRestoreRefreshesSurfaceWhileTerminalIsInactive`, and
`testDirectFirstResponderFocusRefreshesCursorStateAfterForeignResponder`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: normalize project.pbxproj after merging main

Run 35763999808 failed `Fast static checks` before any macOS job could
start, so the app-host test fixes on this branch were never exercised:

    error: cmux.xcodeproj/project.pbxproj is not normalized.
    Run scripts/normalize-pbxproj.py to fix.

The branch head alone checks clean. CI builds the merge of this branch
into main, and that merge is what leaves the file unnormalized, so the
failure does not reproduce without merging main first.

The change is one build-phase entry moving back into alphabetical order.
`LiveWorkspaceFixture.swift in Sources` keeps the same UUID and the same
occurrence count, and `scripts/lint-pbxproj-test-wiring.sh` still reports
ok across 1049 test files, so no target lost a source file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Drop a cancelled fork-probe request, and name the inputs behind the last two app-host failures (#13744)

* fix: drop a cancelled fork-probe request while it waits behind an active probe

A second fork-availability request for the same panel with a different
fallback snapshot parks in `applyPendingForkValidations`' contention
branch: it restores its request to the pending queue and awaits the
active probe. That wait had no cancellation handler, unlike every other
wait in this type, so a cancelled caller kept its request in the queue
until the active probe finished. The probe's completion restarts the
single-flight refresh for whatever is still pending, and the cancelled
request rode along: its fallback was probed and its result replaced the
surviving request's validation for that panel.

Give the wait the same cancellation handler its siblings have, and drop
the waiter's own pending requests when it is cancelled, so the restart
that follows the active probe sees only live requests.

Covers cancelledSharedForkProbeRefreshPreservesSurvivingFallbackSnapshot,
which failed when the restart won the race against the cancelled task's
own cleanup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: name the input behind the off-plan re-arm and the sidebar reveal double pass

Both failures currently report only their outcome, and each has two
incompatible explanations that the message cannot tell apart.

offPlanGeometryWithUnchangedSizingInputsReconverges reports `rearms=3`
with `imposed=nil`. That is either three recovery passes that ran and
failed to impose, or a re-arm budget already spent before the
perturbation, in which case no recovery pass ran at all. Bracket the
recovery window with the existing DEBUG sizing counters and report the
budget at perturbation time plus the planned outers.

visibilityToggleKeepsAppKitTableContainerMounted reports exactly two
projections per row. Report how many the first run-loop turn produced
and which async signals landed inside the reveal window (the workspace
directory channel, workspace order, the shared agent index), since the
hidden phase queues main-queue work that can land during the reveal.

No assertion changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* test: repair the remaining app-host failures and the shard-2 host relaunches (#13759)

Most of these share one cause: the app-host test process runs headless and is
not the active app, so AppKit and WebKit withhold state the tests assumed.
NSApp.keyWindow is nil, makeKeyAndOrderFront never makes a window key,
occlusionState never carries .visible, and behind XCTest's shielding window
WebKit suspends requestAnimationFrame entirely.

The four shard-2 host relaunches were one hung await: renderMarkdown awaited
two animation frames inside callAsyncJavaScript (cd9f34f), which behind the
shielding window never fire, so four MarkdownPanelTests cases each hung until
XCTest's five-minute allowance killed the host. Every WebKit call in that file
now fails fast with a stated reason instead of hanging, renderMarkdown waits on
the viewer's own render contract, and the scroll restore the rAF await was
papering over no longer clobbers a scroll that lands after a content update.

Two real product regressions: windowless shortcut events stopped pruning the
orphaned main-window context (6d7b23f), and ticket minting read the routes
and the v2 installation identity as two separate cache reads.

No assertion is weakened; several are strengthened.

Rebased onto fix/app-host-green after 87a7016 landed a different remedy for
the same key-window cause. The three focus tests this branch fixed through a
test-target MainWindowKeyStatusPin (a swizzle of NSWindow.isKeyWindow, needed
because CmuxMainWindow is final) are exactly the three that commit fixes by
activating the app host and widening the pump's timeout. The pin is dropped:
activating the host makes isKeyWindow true for real rather than forcing the
answer, and a process-wide swizzle installed for one helper is a worse neighbour
to the rest of the shard. cmuxTests/AppDelegateMainWindowTestingSupport.swift is
no longer touched by this branch.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* test: measure Cloud title ink beyond the icon and trace measured invalidations

Real rendering at 75% and 1x reproduces disconnected icon ink misidentified as the title. Retain the alignment tolerance and verify that displaced text still fails.

Enable DEBUG tracing only during the existing sidebar and minimal-mode measured intervals. Count assertions and timing remain unchanged.

* test: preserve carrier preparation before fleet discovery

* fix: retain independent cloud carrier prewarming

---------

Co-authored-by: austinpower1258 <austinwang115@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
teamleaderleo added a commit that referenced this pull request Sep 25, 2026
…meouts (#14210)

* test(terminal): restore the presented-surface fixture contract so package tests compile

`swift test --package-path Packages/macOS/CmuxTerminal` has not compiled on
main since merge 38b32bb191: TerminalSurfaceRendererCallbackTests calls
`PresentedSurfaceFixture(installRendererCallbacks: false)`, but the fixture
initializer only takes `windowVisibleAtCreation`. The flag came from the
issue-2824 branch (77c46d0c02, 6fe70f6f68) and was dropped when the fixture
was reworked for the native callback lifecycle (8008b7c062, 97c60889b4);
the merge re-applied the test call without the fixture side.

Package tests only register render callbacks through the fixture
(RendererCallbackTestSupport), so a fixture that skipped registration would
leave `cmux_test_ghostty_renderer_present` with nothing to route. Restore
the calls to `PresentedSurfaceFixture()`, the combination that was green at
0251a448c9. Verified locally: 294 tests in 37 suites pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(popover): stop rewriting the presentation binding during view update

`ArrowlessPopoverAnchor.updateNSView` calls `coordinator.dismiss()` on every
update where `isPresented` is already false. PR #13311 (01f0a503b9) made
`dismiss()` write `isPresented = false` on the no-popover path, which SwiftUI
reports as "Modifying state during view update" because updateNSView runs
inside the view update. The sidebar footer mounts two anchors whose parents
re-evaluate on every `selectedTabId` change, so every workspace switch emitted
faults and `SidebarWorkspaceSwitchLayoutFaultTests` failed with 15 of them.

Pass `resetPresentation: false` from updateNSView: the binding is already
false on that branch, so the write was redundant. `popoverDidClose` still
resets the binding when AppKit closes a live popover.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(cloud): satisfy the plus menu's sign-in gate and route Cmd+Y through a registered window

`NewCloudWorkspaceShortcutTests` never passed: the plus menu deliberately hides
its Cloud rows unless the account is signed in (#12305, mirroring the command
palette and File menu), and a test `AppDelegate()` has no account flow. The
XCTest version crashed the app host on `rows[1]`, xcodebuild restarted it, and
the lenient gate accepted the partial run; #13178 now rejects that and #13193
migrated the suite to Swift Testing, so the failures became visible.

Inject the signed-in state through the existing `isAuthenticated:` seam, add
signed-out coverage of the gate, route the Cmd+Y event through a registered
main window (as `testReboundKeyRoutesAndOldKeyDoesNot` does, since shortcut
routing bypasses events bound to windows the delegate cannot resolve), and
register a window context for the unavailable-Cloud check.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(chrome): assert the composited Bonsplit chrome contract instead of a stale alpha hex

`WorkspaceChromeColorTests` expected `bonsplitChromeHex` to return
`#1122337F` (theme with opacity as alpha). Since b6d3470668 (2026-05-19)
`compositedTerminalColor` composites the theme over the window base and
returns an opaque color, and 4cbb354cfc made that deliberate: Bonsplit derives
its tab glyph contrast from the rendered backdrop, and an `#RRGGBBAA` hex
reproduces the white-on-white bug it fixed. The tests kept failing unnoticed
because the app-host gate only counted "unexpected" XCTest failures until the
strict check (acedf3f244) reached main through #12053.

Exercise the `chromeBackgroundColor` seam that every production call site
uses with literal expectations, verify the ambient default path against the
resolver plus independent blend arithmetic so the result is deterministic
under any host appearance, and keep coverage for opaque, shared-backdrop,
pane-clear, pane-border, and explicit translucent chrome colors.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(tmux): wait for the main window to adopt the mirror pane before sending keys

`RemoteTmuxMirrorPaneInputMappingTests` required `panel.surface.uiWindow != nil`
right after selecting the workspace. A manual-I/O mirror pane spawns eagerly in
its hidden bootstrap window, which `uiWindow` deliberately excludes, and the
main window's portal adopts the pane host only once AppKit and SwiftUI get
run-loop time; `waitForLiveSurface` returns immediately for an already-live
surface, so the check ran before adoption and the four key-delivery tests
failed at line 176. The failure was hidden until the strict app-host gate.

Order the harness window front and pump the run loop until the surface and
its native view are in that window, as the other hosted-view input suites do,
and assert against the harness window instead of any window.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(browser): report the refused connection through the navigation delegate

`browserPanelRetriesDiscardedRestoreAfterConnectionRefused` waited up to 30s
for WebKit to fail a provisional load to a bound-but-unlistened loopback port.
On the hosted app-host runners that failure never arrives: the load neither
fails nor commits, so the test timed out at line 147 (no
"provisional navigation failed" log line appears for it in any shard).

Stop the in-flight load and report `NSURLErrorCannotConnectToHost` for the
attempted URL through the panel's real navigation delegate, the pattern
`BrowserFailedNavigationReloadTests` already uses. The restore bookkeeping,
error page, `restore_pending` clearing, and retry policy under test run
unchanged and every assertion is kept.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(cloud): scope reserved-workspace cleanup to its pending card and return the bind window

PR #13202 (5d616a7f44) imported `CloudMachineWorkspaceAdoptionTests` and
`CloudMachineWorkspaceResolutionTests` from the still-open
13141-cloud-vm-workspace branch without the production changes they assert,
and its own run skipped the app-host shards, so they landed red:

- `NewMachineSheetPresenter.closeReservedWorkspace` closed the whole
  workspace, which is a no-op for the last tab and discards user panes added
  next to a creating card. Port the branch behavior: remove only unadopted
  loading cards owned by the cancelled machine, clear that binding, keep user
  content, and give a last-tab loading workspace a local anchor first.
  `MachineCreateCoordinator` passes the created or reconciling machine id so
  cancelling machine X cannot clear a binding to machine Y (the tests now
  assert that scoping).
- `v2WorkspaceCloudVMBind` now returns `window_id` alongside the workspace
  refs, as the bind acknowledgement test expects.
- The resolution test selected "first" while the fixture's terminal key is
  "term-first"; use the key so the placement resolves as intended.
- `CloudPaneCreationRetryTests` asserted `discarded` synchronously after the
  projection returned, but the coordinator applies its generation fence behind
  `CloudOperationContext.withPhase`'s recorder await, which suspends on a cold
  per-suite process. Settle with bounded yielding before asserting.

Runtime behavior change (cancelled Cloud create cleanup); needs dogfood.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(cli): restore deferred socket connection, explicit SSH control options, and Codex stop idle observations

Branch commit c8bfb580ba ("fix: repair failures exposed by strict app-host
CI", 2026-09-10) made `cmux vm dev|layout|env` finish local validation and
dry runs before opening the socket, passed the caller's explicit ssh options
into `userConfiguredControlOptions(fromSSHConfigOutput:explicitOptions:)` so a
normalized `ControlPersist=0` from `ssh -G` is not mistaken for host
customization, and published `idleObserved` for transcript-terminal prior
turns before a legacy Codex Stop so the journal drops an obsolete running
turn. The branch's later single-parent commit 38b32bb191 reverted
`CLI/cmux.swift` to main's version while keeping the stricter fixtures, and
PR #12053 landed that inconsistent state; the fixtures (CLIVMDevTests,
CLIVMLayoutEnvTests, the SSH sharing tests, and the Codex missed-prompt Stop
test) have failed since.

Restore the three CLI changes. `SocketClient.configureAuthentication` and
`SocketPasswordResolver` already exist on main, and the CmuxFoundation
overload landed with the branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(cli): align CLI integration fixtures with the shipped hook, SSH, and session-list contracts

The main copy of `CLINotifyProcessIntegrationRegressionTests` predates
several shipped contracts that the lenient app-host gate never enforced and
that 38b32bb191 reverted on the issue-2824 branch: Claude hook acks print
`{}` (#7963), Codex resume bindings require rollout evidence so fixtures
carry a `session_meta` transcript (#10100), SessionStart publishes a binding
so `/clear` counts start after the clear, fresh-terminal SSH startup commands
are script paths that the support decoder must read, cmux control-path
options follow a resolved `ssh -G` (#8308), and `ssh session list` reports the
localized "remote state unavailable" summary with `--json` detail (#9971). The
missed-prompt Codex Stop test now asserts the restored `agent.idle.observed`
for the prior turn.

Flagged for review: the three `ssh pty-attach` cases now expect only
`workspace.remote.pty_bridge` once the endpoint is established, matching
#12726's `preserveLifecycleForRecovery`; if pre-READY bridge failures were
meant to keep reconciling, the flag should be set only after READY instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(workspace): restore the app-host repairs for fork, focus recovery, and shortcut routing

Merge 8d33410585 on the issue-2824 branch took main's copy of
`cmuxTests/WorkspaceUnitTests.swift` wholesale and discarded the branch's
September repairs (c8bfb580ba, c402b9c515, 2127d97afd); PR #12053 then landed
without them while the strict gate started counting the failures. Restore
them against current production, plus the shortcut-suite fixes:

- Fork in a remote workspace: `sshBootstrapArguments` has used `/usr/bin/ssh`
  since #9114, and agent socket propagation requires a socket that exists on
  disk (3bf87e3b4d), so bind a real unix socket and expect the absolute path.
- Fork Conversation context actions dispatch asynchronously (#7259, #8173);
  await the fork panel before asserting.
- Git branch and pull request updates publish through
  `sidebarObservationPublisher` since #6226, not `objectWillChange`.
- Config sanitization: `addWorkspaceIfActive` rebuilds templates from the
  font-size lineage since #8543; override the lineage hook instead.
- Focus recovery: AppKit focus is authorized only for a registered, selected
  workspace whose window carries the main-window identity; use the
  `TerminalPortalTestWorkspace` fixture and the same registration/pump path
  as `WorkspaceTerminalFocusRecoverySwiftTests`.
- Shortcut routing: clear both Cloud defaults after 9c2ba78be4 swapped them;
  neutralize Ghostty's imported goto_split fallback (⌘] by ANSI keyCode) in
  the unshifted-symbol test and assert the digit shortcut does not match;
  wait for the async runtime start before judging keyDown forwarding (skip
  loudly without a live surface); wait for the portal to mount before the
  second-Escape check.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(restore): keep a restore identity when the persisted surface id collides

Since #13098 (e0e77eb74f), `newTerminalSurfaceOutcome` treats a nil
`restoredSurfaceId` as an interactive create and routes it to the selected
pane's Cloud source. Session restore passed nil whenever the persisted panel
id was still live (duplicate-workspace or restore-into-live), so a legacy
managed-Cloud SSH workspace restored into a live manager was turned into a
remote tab create: the scaffold panel stayed startup-suppressed with no
initial command and `TabManagerSessionSnapshotTests` failed to unwrap it.

Mint a fresh UUID on collision instead of nil; it is free by construction, so
the restore keeps its identity and the old-to-new remap works as before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(terminal): align snapshot, projection, terminal, and browser fixtures with shipped contracts

Stale expectations and fixed-spin timing in app-host suites that the lenient
gate never counted:

- Cloud-projected panes restore as manual-mirror reservations with a staged
  remote identity (#12675), not a local placeholder projection.
- Restore reuses the persisted runtime id when free (aff0e32e93), so assert
  liveness rather than a new id; the catalog ignores writes for a Cloud
  machine without a registered provider (#11877), so register the fixture
  provider before publishing.
- Wheel sync requires an authoritative scrollbar response (bbc3edfae4); the
  fixture now answers like `AuthoritativeScrollbarSurfaceView`.
- Search overlay mount, first-responder focus, runtime creation, and the
  visibility-restore redraw run through deferred main-actor tasks; wait for
  them with the class's `waitUntil` helpers instead of fixed run-loop spins.
- Owning the socket path lock is definitive (959f38a4c3): a refused inode left
  by a dead listener is replaced, so the restarted listener accepts.
- The split-divider hit band extends `dividerHitExpansion` past the divider
  (667cc431d9); derive the pass-through boundary from the constant.
- Browser page background blends against Ghostty's effective terminal color
  scheme, which is host dependent; read the same preference the product uses.
- Cloud Machines defaults on in dev builds (#12318); pin the toggle off for
  the default-mode palette contract.
- Prepared navigation requests keep the caller's cache policy (#13003).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(browser): align lifecycle, identity, host-view, loopback bridge, and portal rebind tests with shipped contracts

Browser app-host suites that the lenient gate never counted:

- `BrowserPanelWebViewLifecycleTests`: out-of-range discard delays are
  rejected to the default (the cmux.json loader relies on the nil), and the
  panel's own `isLoading` stays true for the indicator floor, so wait for both
  flags and assert no discard blockers before discarding.
- `browserNavigationUsesEmbeddedWebKitIdentity`: WebKit reports the native
  identity as nil or "" (#9482); accept either.
- `WindowBrowserHostViewTests`: production routes Dock-divider hits by
  yielding to AppKit so the live sidebar tracker receives them (#10902,
  e2e381825f); the tests asserting the portal owns and forwards the hit were
  merged red against a design that never shipped. Realign the stale-frame
  test to the pass-through contract and remove the four own-and-forward
  tests with their fixtures. **Flagged for review:** if own-and-forward is
  still wanted, that is a hit-testing product change for its own PR.
- `testRemoteWorkspaceRuntimeBridgeAliasesMultipleLoopbackPortsFromSamePage`:
  the navigation delegate restarts main-frame loads to apply the user-agent
  policy (11c6efee86); a direct `loadHTMLString` with an HTTP base skipped
  that step, so the data load was cancelled and replayed as a deferred
  request. Apply the identity first and assert the document is current.
- `portalRebindPreservesDocumentAndRoutesRefreshToTheSameWebView`: the portal
  host is a theme-frame sibling of `contentView` for non-glass windows
  (#12929); assert same-window instead of descendant-of-contentView.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(terminal): wait for asynchronous runtime creation in the remaining direct-interaction tests

The same "Expected runtime surface before ..." precondition that
9f258dd7c7 made wait for the deferred runtime start still sampled after a
fixed run-loop spin in the detach-race, close-lifecycle, repeat-key, and
repeat-IME tests, and CI on the branch head showed them failing that way.
Use the class's `waitUntil` for those four sites too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(cli): model surface.respawn for respawn-pane and give the Codex stack fixture rollout evidence

`cmux respawn-pane` has sent `surface.respawn` instead of `surface.send_text`
since #5465 (8cafcc3bac); the window-flag fixture's mock still rejected that
method, so the CLI exited 1. Answer `surface.respawn`, asserting the window and
surface ids, `tmux_start_command`, and that the shell-invoked command carries
the user command but never the `--window` flag, which is the test's intent.

The Codex interrupted-stack fixture's transcript had no `session_meta` line,
so `CodexSessionResumeVerifier` (#10100) found no rollout evidence and the
prompt published `surface.resume.clear`. Prepend the session line as the
other Codex fixtures do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(socket): keep mobile.panel.artifact.fetch off the local socket and advertise the served artifact reads

02c1ba4dd9 removed `mobile.panel.artifact.fetch` from the socket worker
methods because it needs the authenticated mobile execution context, but
that half of the change was lost in a merge: the policy still routed fetch to
the worker lane, which has no handler for it, so the local socket answered
`internal_error` instead of the `method_not_found` boundary that
`TerminalControllerSocketSecurityTests` pins. Restore the removal.

`system.capabilities` never advertised `mobile.panel.artifact.stat` and
`.thumbnail` although both are served on the worker lane (04ff18eea6 added
them only to the test's expectation); advertise those two and keep fetch out.
The remote relay allowlist is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(remote): align tmux seed transport, SSH, socket command, and port scanner fixtures with shipped behavior

- `RemoteTmuxPaneSeedTransportTests`: manual-I/O mirror panes spawn eagerly
  (#9272) and stay runtime-backed after portal churn (#9769), so the pane
  renders its assigned grid before a seed arrives and nothing was retained.
  Publish a larger tmux pane grid than any test surface applies so the
  retention under test sees the lag it exists for.
- `SSHRemoteCWDRegressionTests`: the persistent-PTY exec helper runs only
  with `protectsFromHangup: true` (cd9f34ff1b); pass it, and widen the
  first-spawn guard.
- `SSHDeepSleepReattachTests`: a Cloud-owned workspace rejects launch
  overrides on splits (#13098); create the custom-identity pane before
  configuring the remote connection.
- `SSHConfiguredRemoteCommandHostTests`: ssh-pty-attach validates the bridge
  `daemon_version` before dialing (#12726); the mock now reports one.
- `CloudManualMirrorTransportTests`: the pane failure card uses the short
  title since 9bf6cb8c94.
- `SurfaceSocketCommandTests`: `vm.workspace_new` admits its optimistic
  workspace through the active main window (#13152, #13155), so bind a bare
  window to the fixture context; a receipt without a starter terminal costs
  one snapshot (6d43ea699a).
- `PortScannerPublicationTests`: the forced-result acknowledgement hops off
  the main actor, so an unchanged port set may be deduped under a later
  refresh; drain publications until the retirement lands.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: align mobile artifact fetch lane assertion

* repair: close remaining full-suite contracts

* test(restore,sidebar): align two stale contracts with shipped behavior

testRemoteWorkspaceAutoResumeKeepsRemoteStartupCommand asserted the restored
remote panel's local spawn cwd equals the remote-host path. It never can:
OneShotTerminalLauncherStore.enterableWorkingDirectory rejects a path that is
not locally enterable, which is what keeps the owning shell valid (#7031).
The remote cwd does survive restore — as the panel's trusted remote directory
report, and as the `cd` prefix the resume input carries (both already asserted).
Assert it where it actually lives and pin the spawn cwd to nil.

testSidebarPullRequestsTrackFocusedPanelOnly expected a background panel's PR
to be hidden from sidebarPullRequestsInDisplayOrder(). That list is documented
as the workspace's deduplicated rows in pane/tab order, both consumers
(taskStatusSignals, the control-sidebar snapshot) want every panel, and the
sibling branch test asserts the same all-panel model. Focus scoping lives in
the `pullRequest` binding, which the test already covers. Renamed to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: give Cloud catalog tests live destination workspaces

#13196 made SurfaceCatalog.validateOwnership refuse a destination that is
not a live workspace. That rule stays. Tests that projected, restored or
opened browsers into a made-up workspace ID now failed with
destinationNotFound, timed out waiting on a provider that was never called,
or passed a later check for the wrong reason.

LiveWorkspaceFixture registers real Workspace objects and hands the catalog
a CloudWorkspaceRenameService that resolves them, the way the app's
composition root does. Its workspaces() list stays empty so the native
projection coordinator does not start mirroring into them. For tests on
SurfaceCatalog.shared, withAppRegistration registers the TabManager as a
windowless main-window context for the test body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: normalize pbxproj after live workspace fixture

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: admit agent renames of agent-owned accepted Cloud names

e6926fbbda (#13403) made submitCloudPanelRename run admitsTerminalRename
for automatic names. Its last clause only admitted replacing an accepted
name when the local panel still carried `.auto` provenance, but since
1e1d319dd1 an accepted daemon name reconciles locally as `.remote` and the
owner lives in the tab's nameAuthority. Every agent title after the first
accepted one was refused, so "Failed agent rename keeps the accepted title",
"Mirroring an agent-named placement does not block its next agent title" and
"An older automatic result and old snapshots cannot replace an accepted name"
failed on their second agentName call.

Admit an automatic rename when no write is pending and the accepted tab name
is owned by the daemon's `auto` authority. User-owned names, pending writes and
local user titles on any projection still refuse it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: expect adopted machine rows to keep the pending create identity

#12919 (021f792b95) made New Machine creation optimistic: once a running
create's machine appears in the fleet list or catalog, its row keeps the
`pending-machine:<operation>` node ID so selection and expansion survive
adoption (see committedProjectionKeepsThePendingNodeIdentityUntilFleetAdoptsIt).
pendingRowStepsAsideOnceItsMachineHasARow still expected `machine:<id>`.

Assert the new identity and that the row is the adopted machine, not a
stand-in: the stand-in is gone, one row remains, and it shows the created
machine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: restore per-display noVNC targets and refresh stale Cloud tests

Product (#13196 regression): the 6db759357f main merge into
13192-cloud-display-ownership took main's CmuxTuiSurfaceProviders.swift and
CloudPortRoutePlan.swift, dropping eb3edd7f98's per-display ports.
23c807b50a restored the display coordinator but not these hunks, so
withPrivateBrowserURL rewrote every display to 6901 and a daemon pointer
without a discovered target fell back to display 1. Restore both hunks and
drop the duplicate port-less privateDesktopURL overload.

Tests:
- CloudDisplayCatalogTests: 178d35e5da (#13196) made every guest command run
  `list` as a readiness probe, so fakes dispatching on " list" answered
  creation with the list catalog. Dispatch on the create action line, and pin
  the command shape.
- CloudPortOpenRegressionTests: 178d35e5da (#13196) filters RFB/noVNC ports
  only when the display catalog owns them (displayPortsOwned). Assert both
  the desktop and non-desktop results.
- CloudTreeOneMachineManyWorkspacesTests: #12740 added a final Resources
  section under each machine. Expected trees include it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: publish every guest display from daemon graph updates

The 6db759357f main merge into 13192-cloud-display-ownership (#13196) put
back main's [desktopDisplayResource()] pools in CmuxTuiSurfaceProvider's
refresh, publish and delta paths. 23c807b50a restored the display
coordinator lifecycle but not these pools, so a display created beyond
display:1 vanished on the next daemon publish, and a delta that touched it
removed it. Restore eb3edd7f98's displayResources pools, the display-kind
delta check, and the injectable displayCoordinator. Drop the now-unused
desktopDisplayResource().

Adds a test that creates display:2 through the provider and asserts that
both displays and their ports survive a full publish and a display delta.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: fence the fake Cloud projection reply with its mutation cursor

aDetachedTerminalDropsItsStaleTabBeforeMoving installs a daemon graph
since c402b9c515 (#13403). When the placement lane drains it reconciles
against that graph, which predates the projected tab, and clears
tab_projected. The real reply always carries a mutation cursor
(CmuxTuiSnapshotParser.placedTab requires one) that fences exactly this;
the fake returned none. Give the fake a projectCursor and set it ahead
of the installed graph.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: register the restored window before relinking Cloud projections

#13196 made SurfaceCatalog.validateOwnership refuse a destination the app
cannot resolve, so SurfaceCatalog.shared no longer relinked a restored
projection into a TabManager that no main window owns. Register it for
the test body with LiveWorkspaceFixture.withAppRegistration, as #13651
does for CloudClosedPanelRestoreTests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: wait for the relay ports kick, not the first relay line

Since #8442, a relay prompt also reports shell state. Both RPCs run in
separate background children, and the zsh prompt-refresh test waited
only until the log was non-empty, so it could read report_shell_state
alone. Wait (deadline-bounded) for the ports_kick line itself, in the
zsh test and its bash sibling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): relay output of an SSH session that ends right after auth

The expect wrapper watches the first two seconds after sending the
password for a rejection with log_user 0. A session that authenticates
and exits inside that window hit the eof branch and its output was
dropped. Flush the buffered output before exiting. #13207 replaced the
test's 9 s sleep with a FIFO release, which is what exposed this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: restore the no-connection path for rejected vm dev input

4120e35654 taught runVMDev that invalid input never connects: keep the
listener open, then check its backlog after the CLI exits. #13207 dropped
that again, so each rejected run waited 60 s for a mock-server
expectation that only the listener closing (after the wait) fulfills.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: drain the terminal while the SCP host-key failure runs

The pty case read the master only after the CLI exited. A pty's output
queue holds about 1 KiB and the host-key failure report is longer, so the
CLI blocked writing stderr until the 30 s timeout. Read the master on a
thread while the CLI runs. The test starts its own sshd; no runner host
dependency is involved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: fail fast when a gated Cloud call ends before its fake is entered

Every app-host restart in the 09-21/09-22 shard logs I sampled was a Swift
Testing time-limit hit in one of two suites, and each hit relaunches the
test host:

- SurfaceCatalogTests: when catalog.project threw before reaching the
  provider (destinationNotFound, fixed by 3dc91b2321), each gated test
  parked in MaterializeGate.waitUntilEntered() until the 300 s limit.
  Five tests restarted the host one after another, about 25 minutes per shard.
- CloudDisplayCatalogTests: when the fake no longer recognized the create
  command (fixed by 70eaf44819), create() failed before the exec started
  and `await started.result` parked until the 60 s limit.

The waits now also end when the caller's task finishes, so the next such
setup failure is an ordinary failed #require. The cancellation test also
waits on its own cancellation signal instead of a 60 s Task.sleep.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: reveal a retired terminal through the portal rebind, as the app does

Since #12607 hiding a terminal removes its hosted view from the window, and
only a bind reinstalls it. The test flipped portal visibility on the detached
view, so no size commit could ever land and the final shrink check failed
every run. Rebind like TerminalPortalReconciliation and drop the wait loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: pin key-window status in terminal focus suites

The app-host test process runs headless and is usually not the active app,
so makeKeyAndOrderFront never makes a programmatic window key. Terminal focus
paths gate on isKeyWindow (automatic first-responder apply, focus redraws,
deferred focus reapply, ensureFocus window activation), so these suites
passed only when an earlier test in the shard had activated the app.

Share the existing KeyStatusTestWindow and use it in
WorkspaceTerminalFocusRecoveryTests, WorkspaceTerminalFocusRecoverySwiftTests
and TerminalNotificationDirectInteractionTests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: report which layer holds off-plan tmux mirror geometry

offPlanGeometryWithUnchangedSizingInputsReconverges fails every run with all
three output-parity re-arms spent and the hosted view still at the perturbed
0.8 divider. A standalone bonsplit replay of the same sequence converges, and
a sibling test without a bound (portal-visible) workspace heals the same
displacement. Add the live split view's arranged widths and the split model's
imposed extent to the failure message so the next run shows whether bonsplit
refused the apply, the imposition was cleared, or the portal did not follow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: restore Cloud projections before the workspace is published

TabManager.restoreSessionSnapshot restores each workspace's surface
projections before it assigns tabs, so SurfaceCatalog.restore's ownership
check could not find the destination and silently dropped every restored
remote projection. Check ownership against the workspace being restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: repair stale and broken app-host expectations

- CLI ssh attach: count identity UUIDs exactly; #11497 added an auth token uuidgen
- device mirror directories: give the fake device trusted presence (e3d424c722 gate)
- font zoom mirrors: deltas from the 8pt inherited base (e6926fbbda arithmetic)
- Computer Use refresh: fake daemon reply was invalid JSON in a raw string (#13599)
- quit alert: compare button alignment rects, not padded frames
- remote split cwd rescue: cwd travels as CMUX_REMOTE_INITIAL_CWD since #12054
- default freestyle split: Cloud-owned splits route to Cloud since fa5dc4cc10

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: activate the app host before waiting for terminal focus

`focusTerminalForTesting` waits for `window.isKeyWindow` before it hands
first responder to the terminal, but `makeKeyAndOrderFront` only makes a
programmatic window key while the test host is the active app. The
app-host process starts inactive under `xcodebuild test`, so the wait
succeeded only when an earlier test in the shard happened to activate the
app. Its callers use a real `createMainWindow()` window, which cannot be
swapped for `KeyStatusTestWindow` the way the pinned focus suites were.

Activate explicitly, and give the wait a CI-appropriate timeout: the
activation and the key-window transition land on later main run-loop
turns, and the pump's one-second default expires before the window goes
key on a contended runner.

Observed on run 35743588302: `plainTerminalTextDoesNotResolveAppShortcutContext`
(shard 1/6), `keyboardCopyModeKeyClearsTerminalUnread` and
`workspaceFontSizeShortcutPreservesBackgroundTerminalUnread` (shard 2/6)
all fail at this helper's return value, and shard 6/6 shows the same
process state directly as `NSApp.keyWindow -> nil`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: give standalone terminal fixtures a live portal authority

`setVisibleInUI` and `setActive` fold their request through
`Workspace.portalRenderingEnabled(for:)`, which denies any workspace id
the app delegate cannot resolve to a selected tab. Three direct
interaction tests build their surface with `tabId: UUID()`, so once any
earlier test installs an `AppDelegate.shared` the authority denies the
portal, the hosted view is never actually made visible or active, and the
fixture stops exercising the behavior it asserts: the surface never takes
Ghostty focus and never schedules a visibility-restore redraw.

Register a real selected workspace and build the surface with its id, so
the fixture gets the same authority the app grants the selected tab.

This is the fixture, not the product: the authority check is deliberate,
and denying an unresolvable workspace is what keeps queued portal
callbacks from reviving an inactive workspace.

Fixes on run 35743588302 shard 4/6:
`testKeyDownRecoveryDoesNotReplayFocusAfterResponderMovesAway`,
`testVisibilityRestoreRefreshesSurfaceWhileTerminalIsInactive`, and
`testDirectFirstResponderFocusRefreshesCursorStateAfterForeignResponder`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: normalize project.pbxproj after merging main

Run 35763999808 failed `Fast static checks` before any macOS job could
start, so the app-host test fixes on this branch were never exercised:

    error: cmux.xcodeproj/project.pbxproj is not normalized.
    Run scripts/normalize-pbxproj.py to fix.

The branch head alone checks clean. CI builds the merge of this branch
into main, and that merge is what leaves the file unnormalized, so the
failure does not reproduce without merging main first.

The change is one build-phase entry moving back into alphabetical order.
`LiveWorkspaceFixture.swift in Sources` keeps the same UUID and the same
occurrence count, and `scripts/lint-pbxproj-test-wiring.sh` still reports
ok across 1049 test files, so no target lost a source file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* perf(tests): stop app-host unit tests waiting out real timeouts

The app-host suite spends most of its wall clock in a handful of tests that
wait out production timeouts, churn oversized fixtures, or run benchmarks.
Each one now drives the same behaviour through an injected timeout, clock,
or completion signal, with a separate cheap assertion pinning the default
value to the production/server default.

Production changes are configuration seams and one real fix:
- KeyboardShortcutSettings.resetAll only removes keys that are stored, so a
  reset no longer fans out one UserDefaults.didChangeNotification per action.
- SocketStartupWaiter / AgentRestorePreflightTimeout / BrowserDownloadWaitTimeout
  own the CLI wait windows (default plus a narrowing environment override), so
  client and handler cannot drift apart.
- PortScanner and BrowserScreenshotWebViewSnapshotter take their schedules as
  parameters instead of hardcoding them.

Benchmarks move out of the app-host unit suite behind the existing gating
pattern, keeping their correctness assertions in the unit suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: keep the scroll settle default readable from a nonisolated context

The default lives on a @MainActor enum, so every default-argument use of it
is evaluated in a nonisolated context and trips the Swift 6 isolation
warning. The value is an immutable Sendable constant, so mark it nonisolated
rather than isolating four call sites.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(tests): stop kicking the port scanner from the poll loop

Review found two defects in the slow-test pass.

1. `PortScannerPortRetirementTests` polled every 25ms and called
   `scanner.kick()` on every poll, against the 10ms compressed coalesce
   delay this PR introduced. `kick()` calls `startCoalesce()` whenever no
   burst is running, and `startCoalesce()` cancels and re-arms the timer.
   On a loaded runner the timer's jitter is the same order as the coalesce
   window, so the kick burst can cancel the timer indefinitely: no scan
   ever runs and the test burns its 20s deadline. The previous 500ms poll
   against a 200ms delay left a 2.5x margin that the compressed schedule
   removed.

   Fixed by removing the kick from the poll loop rather than by widening
   the interval, so the speed win stays and the flake vector is gone
   instead of made less likely. A kick already guarantees
   `minimumScansPerKick` scans - exactly the complete misses the
   reconciler needs to retire a port - so one kick after
   `stopListening()` is sufficient. That is the shape
   `lateBurstKickRetiresStoppedListener` already used. `onKick` is gone
   from `waitForPublication` entirely, so the poll interval is no longer
   coupled to the coalesce delay and the footgun cannot come back.

2. `scripts/test-command-palette-nucleo-ffi.sh` set
   `CMUX_COMMAND_PALETTE_SEARCH_BENCHMARKS=1` and then passed
   `-only-testing:cmuxTests/CommandPaletteNucleoFFITests`, a different
   class from the four `CommandPaletteSearchEngineTests` benchmarks this
   PR gated behind that variable, so those four ran nowhere. The script
   now names both classes and asserts one BENCH line per gated benchmark,
   since a skipped benchmark otherwise passes silently.

   Nothing referenced the script either, so even the FFI tests it names
   were not run by CI. `.github/workflows/command-palette-search-benchmarks.yml`
   is its scheduled caller, modelled on the `tmux-corpus.yml` nightly
   (same checkout/GhosttyKit/zig/rust/SPM-cache setup, same `cmux-unit`
   scheme). The script takes an optional
   `CMUX_NUCLEO_FFI_SOURCE_PACKAGES` so the workflow can reuse the cached
   package clone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: keep the benchmark caller dispatch-only while automatic CI is paused

tmux-corpus.yml and perf-activation.yml -- the two closest macOS benchmark
nightlies -- both carry "Temporarily manual-only beginning 2026-07-13 to pause
automatic CI" and have their crons removed. Adding a live cron here would
quietly reverse that reduction, which is the opposite of what this branch is
for.

The gate still has a caller, and the script's BENCH assertions still fail the
job if a gated benchmark silently skips; running it is now a deliberate act
rather than a daily cost. Drop "nightly" from the name, since it is not one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(tests): model the title-word ranking term in the search reference

testBenchmarkCorporaMatchReferencePipelineOnSmallFixture asserts the
optimized engine equals a reference pipeline rebuilt in the test. On the
query "workspace 31" the engine scored workspace.large.31 at 17598 and the
reference at 16000, so the suite failed.

The engine ranks on three terms; the reference modelled two. The missing
one is commandPaletteTitleWordScore: a query that is, or prefixes, a
title's search words scores the tokens' upper bounds plus a per-token
title bonus. For "workspace 31" that is 6799 + 6799 + 2000*2 = 17598,
exactly the observed gap.

The reference now models that term too, reimplemented rather than calling
the engine's copy, because a reference that shared the implementation
would assert nothing about it. FixtureEntry prepares the two title texts
the term needs once at construction, so the benchmark timing loops — which
build fixtures outside the timed region — do not start charging the
legacy side for preparation the engine does not repeat either.

Verified by compiling the engine sources against this reference on Linux
and comparing all 34 corpus/query pairs the parity tests use: the large
benchmark corpus goes from 1 mismatch to 0, and the command, switcher and
switcher-benchmark corpora stay at 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(tests): drop the ineffective SSH auth retry budget rewrite

persistentAttachExitsAtForegroundAuthenticationFailureLimit() rewrote the
generated startup script to shrink the 20-failure foreground-authentication
budget to 3, then asserted 3 attempts. CI recorded 20: the rewrite never
changed the behaviour it was meant to bound.

The budget the test edited lives in the shell wrapper, but the preserved
__ssh-* CLI helper is what owns authentication retries and counts them
itself, so rewriting the wrapper's literal leaves the helper executing its
own 20. The shrink was decoration over a real 20-attempt run.

The test now keeps the production budget and asserts 20. The fake sleep
already removes the backoff between attempts, which is where the wall-clock
cost actually was, so the same failing run reached the assertion in 4.25 s
against the 6.1 s this test cost before the PR. That is less than the <1 s
the PR's table claims for this row; the table is corrected in the PR body.

replacingWithinScript, scriptDecodeLevels and applyingReplacements existed
only for this rewrite and have no other callers, so they go with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(tests): close the persistent CLI client before awaiting the mock

testDefaultFreestyleSSHAttachHidesPersistentRetryLimitInCountdown failed
with "Asynchronous wait failed: Exceeded timeout of 5 seconds, with
unfulfilled expectations: cli mock socket handled", after 5.227 s.

startMockServer fulfills once the first connection is done. This test
drives a persistent SSH attach, so the client holds its connection open
across the retry countdown and the mock has no "done" to observe. The
wait could only ever end by timing out; it had been passing on the longer
window this PR narrowed.

The test now terminates the child once it has observed the progress it
actually asserts — the "Retrying in 0.1s (attempt 2)." banner — and waits
for the mock afterwards. The countdown assertions that follow are
unchanged and still read the captured output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Retrigger CI after retargeting this PR onto main

The synchronize event that pushed ed3ab7a61e fired while the base was
fix/app-host-green, a branch the #13643 squash merge had already deleted,
so GitHub could not build the merge ref and the pull_request CI workflow
never started. Only the two pull_request_target workflows ran, which left
the required ci-status check absent rather than red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Make the two new timeout holders instantiable values

CI's iOS conventions guard reported two violations this branch introduced:
AgentRestorePreflightTimeout and BrowserDownloadWaitTimeout were both
caseless enums with an all-static public surface, which the lint refuses
because such a type can never be substituted.

The refusal is right, and for this branch in particular: the point of the
change is that a test should be able to hold a timeout agreement without
spending it. A static holder cannot be narrowed, so the tests could only
assert the shipped numbers.

The restore preflight budget moves onto AgentRestorePreflightInvocation,
which is the type it bounds and already carries the environment it is read
from. Names lose the now-redundant prefix: defaultSeconds becomes
defaultTimeoutSeconds, environmentKey becomes timeoutEnvironmentKey, and
seconds(environment:) becomes timeoutSeconds(environment:).

BrowserDownloadWaitTimeout becomes a struct whose three windows are stored
properties with the shipped values as init defaults, plus a `standard`
value the handler and the CLI client both read. Both call sites take
`.standard`; the behaviour is unchanged. The regression test now also
constructs a narrow window and asserts the client still outwaits the
handler, which makes the invariant a property of the pair rather than an
assertion about 10 seconds.

Verified with swiftc 6.1.3 on Linux: both files type-check under
-swift-version 6, and the extracted behaviour matches the previous
constants exactly (10s default, 10s ceiling on the override, whitespace
trimmed, non-finite and non-positive overrides ignored; 10000ms default
window, 120000ms cap, 15.0s client response timeout). The conventions
diff lint now reports no new violations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Remove the KeyStatusTestWindow the main merge declared twice

`macOS compile admission` failed with "ambiguous use of
init(contentRect:styleMask:backing:defer:)" at
WorkspaceTerminalFocusRecoveryTests.swift:551 and WorkspaceUnitTests.swift:4890.
Neither file is touched by this branch. The cause is
cmuxTests/AppDelegateMainWindowTestingSupport.swift, where the merge of main
kept both sides' addition of the same class: `final class KeyStatusTestWindow`
appeared twice, byte-identical including its doc comment, so every construction
of it was ambiguous.

Git had no conflict to report — the two additions did not overlap — which is
the third time on this branch that a clean automerge produced code that could
not compile.

Checked the rest of the merge for the same shape: no other duplicated top-level
type in cmuxTests, and none in CLI/ or Sources/. The one other repeated name,
`optimizedResults` in CommandPaletteNucleoFixtures.swift, is a legitimate pair
of overloads taking `corpus:` and `entries:`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: expect the reconnect budget #13959 honors, not the old clamp

The new budget probe asserted that "21" and "021" clamp to 20. Main's
#13959 made a well-formed budget above 20 honored up to SSHReconnectBudget's
ceiling, so the merge ran green through git and red on CI: stdout "21".
The cases now carry their expected value, from SSHReconnectBudget where
it is a policy constant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: gate the benchmark runner behind the paid-overflow switch

Main's #13994 guard rejects a vars.MACOS_RUNNER_15 read without
CI_PAID_MACOS_OVERFLOW, since it can hold a metered WarpBuild label. This
branch's new benchmark workflow predates the guard, so the merge of main
passed git and failed workflow-guard-tests. Same form as perf-activation.yml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: close CodeRabbit gaps in the fast app-host test rewrites

- PortScanner late-burst: the runner stops listening and issues the late
  kick from inside the fifth lsof call, instead of the test task racing a
  one-second timer gap it could miss in either direction.
- Sidebar quiescence: the quiet window is measured in time and spans twice
  the sidebar's 50ms coalesce stage, so a pending coalesced or debounced
  row invalidation cannot land after the drain declared quiet.
- SSH exit prompt: waitUntilBlocked samples the launched process and all
  of its descendants, so it holds whether the shell execs the startup
  script in place or forks it; terminate() retires that tree so a forked,
  parked helper cannot hold the output pipes or leak into the test host.
- Sidebar git refresh: wait until the panel is a poll candidate again
  before each refresh; a completed read can precede the probe clearing.
- Diff fixtures pin SHA-1 objects and loose-file refs, which the
  handwritten refs and in-process blob ids assume.
- Rename the registry test to the post-list enrollment order and assert it.
- TerminalController uses BrowserDownloadWaitTimeout's shared clamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep the Cloud carrier prewarm on activation

A merge from the fix/app-host-green lineage deleted the prewarm that
syncPollingToActivationPolicy() starts (999693e015, perf: prewarm cloud
carrier before first machine), so enrollment waited for the first
listPage() round trip. The polling test was then rewritten to assert that
slower order. Neither change is part of this PR, and #13403 carries the
same removal for its owner to argue. Both files now match main, whose test
asserts the prewarm starts before fleet discovery finishes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: keep the git fixture's formats when the host sets Git defaults

GIT_DEFAULT_REF_FORMAT and GIT_DEFAULT_HASH override the init options the
handwritten-ref fixtures depend on. runGitProcess no longer passes them on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: report browser.download.wait's requested timeout as a number again

The shared timeout window left requested_timeout_ms an Optional, which the
timeout error payload coerced to Any and which put a new warning over the
Swift warning budget. The handler resolves the default and the lower bound
before clamping, as it did before, so the field is the same Int as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: settle the Global Search palette before each routing test

KeyboardShortcutSettings.resetAll() no longer posts one
UserDefaults.didChangeNotification per action, so the next test's init
no longer spends enough main-thread time for the previous test's animated
NSPopover close to finish. Two routing tests then saw isShown == true from
the palette an earlier test had opened (run 35967763918, shard 7).

Dismiss the palette and pump the run loop until it closes, bounded at 2 s,
in both suites' init, matching the helper the sibling suites already use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: route the palette benchmark workflow to a hosted runner on forks

main's fork runner routing guard (test_ci_fork_runner_routing.py) now
rejects a runs-on expression that can reach a Blacksmith label outside
manaflow-ai. Lead with the owner fork branch, as perf-activation.yml does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: drop the extra paren in the palette benchmark runs-on expression

572dc2a3aa7 kept the old closing paren and added another, so actionlint
could not parse the expression.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: set up Bun on the shard that runs agent notification semantics

The notification semantics step resolves `bun` with `command -v bun`, but
Bun was only set up on the CLI regression shard (4) while the step runs on
the focused regression shard (6). Main never reaches the step because its
unit batch fails first, so the missing Bun only shows once shard 6's unit
tests pass, as they do on this PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: scope the CmuxWebView keyDown hook to each reentry test

CmuxWebViewKeyDownReentryTests swizzled CmuxWebView.keyDown(with:) for the
rest of the process. When CmuxWebViewWebContentUndoTests ran after it in the
same app host, browserCmdZPerformsWebContentUndoWhenPageDeclinesTheChord
recursed in the test hook until the stack overflowed (seen on macOS 15 and
26 once this PR's shard layout put the two suites back to back).

Install the hook per test window, call the captured original implementation
directly instead of re-dispatching through a swapped selector, and restore
the original implementation when the window closes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: keep the OpenCode feed harness socket inside sun_path

testOpenCodeFeedPluginEmitsCompletionForBothIdleEventShapes put its Unix
socket under FileManager.temporaryDirectory plus a UUID-named root. On a
Blacksmith runner that is /private/var/folders/.../T/, and the path runs
past the 104-byte sun_path limit, so the Bun harness fails in listen() and
prints nothing. Use a short /tmp path for the socket instead.

This surfaced once the agent notification semantics step ran on this PR;
main never reaches the step because its shard 6 unit batch fails first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: wait for the relayed pwd line, not the first relay line

zshRelayPromptReportsRemotePWD broke out of its wait as soon as the relay
log was non-empty. _cmux_precmd sends report_shell_state and report_pwd as
separate background relay calls, so on a busy runner the log held only the
shell-state line when the test read it. Wait for the report_pwd line itself,
up to 5 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: tie the E2E stale-snapshot case to OWNED_MAX_AGE_MINUTES

#14341 raised OWNED_MAX_AGE_MINUTES from 20 to 45, so a 30 minute old
snapshot is now fresh enough and the "snapshot too old" case in
test_run_e2e takes the owned Mac, failing app-host-execution guards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3dfc85f17f2b54936fb044ee718387b2f4258903)

* test: pin the full paste worker to lossy rich text, per #14121

#14121 sends mixed rich and plain text through the plain-text helper, so
richPasteDoesNotUsePlainTextHelper (HTML plus a clean plain string) now gets
a successful fast-path paste instead of the full worker's error, and it has
failed on every main run since. The helper still declines rich text whose
plain export lost characters (U+FFFD or repeated '?'), because the full
worker can recover them. Assert that case instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: austinpower1258 <austinwang115@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant