Fix Cloud display ownership and connection readiness - #13196
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThe pull request adds multi-display Cloud desktop support. It adds guest display allocation, discovery, creation, routing, readiness tracking, session persistence, and ownership checks across transfers, duplication, projection, and restoration. ChangesCloud display lifecycle
Priority: ➖ Normal Estimated code review effort: 5 (Critical) | ~90 minutes Change: Feature Suggested reviewers: Merge Risk: 🟡 Moderate · up to Cloud displays can become stuck, fail during daemon startup, connect unexpectedly while hidden, or fail to recover after discovery. These issues should be resolved before merging. Important Pre-merge checks failedPlease resolve all errors before merging. Addressing warnings is optional. ❌ Failed checks (10 errors, 1 warning)
✅ Passed checks (14 passed)
Full details: Cmux Swift Actor IsolationExplanation The production diff adds actor-isolation debt covered by the rule. Resolution Mark the new value models as Full details: Cmux Swift Blocking RuntimeExplanation The PR adds production synchronization in Resolution Remove the production shell polling loop and the guest helper's sleep-based readiness waits. Use an explicit startup/readiness handshake on the control socket, or return a retryable unavailable result and let the async coordinator retry from a real state transition. Replace blocking creation and shutdown waits with completion events handled by the service's serialized event loop. Replace shared-state manual locks with one serialized owner, or document and isolate a necessary low-level process bridge if no actor/event-loop design is possible. Full details: Cmux Cache Substitution CorrectnessExplanation The PR introduces an unsynchronized cache fallback in persistence paths. Resolution Use the current Full details: Cmux No Hacky SleepsExplanation The PR adds production guest-runtime synchronization based on fixed delays. Resolution Remove the fixed shell and Python sleeps and the wall-clock health polling. Let the guest service owner publish readiness only after the control socket and display endpoints are ready, then wait for that readiness through a blocking connect/read or another event-driven, cancellation-aware mechanism with a bounded deadline. Drive supervision from process-exit and endpoint-state events, or use a tested cancellation-aware scheduler that reacts to those events, instead of a fixed 30-second polling interval. Full details: Cmux Algorithmic ComplexityExplanation The PR adds an unbounded scan to the Cloud sidebar snapshot path. Resolution Move display-creation eligibility out of Full details: Cmux Swift Package BoundariesExplanation The PR adds and expands independently testable Cloud display domain logic directly in the app target. Resolution Create a small macOS SwiftPM target named Full details: Cmux Swift LoggingExplanation The PR adds a file-scoped Full details: Cmux User-Facing Error PrivacyExplanation The new production New Display flow can expose unsanitized Cloud VM error details. Resolution Catch and classify failures from display discovery/creation before they reach the generic Full details: Cmux Full InternationalizationExplanation The PR adds user-facing Swift strings through Resolution Add real translated Full details: Cmux Architecture RethinkExplanation The PR introduces a production readiness repair path based on polling. Resolution Remove the shell ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
| private static let source = #""" | ||
| #!/usr/bin/env python3 | ||
| \"\"\"Guest-owned display catalog and supervisor; invoked by the desktop service. |
There was a problem hiding this comment.
This is a Swift raw string, so the backslashes in \"\"\" are preserved rather than decoded. On a guest without an installed helper, the generated executable fails at its first docstring with a Python syntax error. Discovery cannot succeed, leaving New Display unavailable. Remove the unnecessary escaping throughout the embedded script and validate the exact generated payload; the Python tests currently import a separate fixture whose docstrings are not escaped.
| if not self.processes.get(number) or not ready(number): | ||
| self.start_components(number, environment, runtime) |
There was a problem hiding this comment.
HTTP failure destroys desktop session
ready(number) returns false when the noVNC HTTP probe fails or times out, then start_components terminates every recorded process—including an otherwise healthy Xvnc server. A websockify crash or transient HTTP failure can therefore disconnect all guest GUI applications from X and lose unsaved work. Recover the failed component independently rather than treating an HTTP health failure as permission to replace the whole desktop session.
| bus = subprocess.run(["dbus-launch", "--sh-syntax"], env=environment, | ||
| capture_output=True, text=True, check=True).stdout | ||
| for line in bus.splitlines(): | ||
| if "=" not in line: | ||
| continue | ||
| key, value = line.split("=", 1) | ||
| environment[key] = value.strip().strip("'") |
There was a problem hiding this comment.
Recovery leaks session bus daemons
Each component restart launches a new session bus, but the daemon is never included in the processes owned by terminate. Its PID is only copied into the environment and overwritten on the next launch. Repeated recovery leaves bus daemons behind, and display cleanup does not stop them either. Track and terminate the bus with its display, or deliberately reuse the existing live bus.
|
|
||
| currentURL = restoredURL | ||
|
|
||
| if let resource = snapshot.cloudResource { restoreCloudResource(resource); return } |
There was a problem hiding this comment.
Restore discards saved browser location
A Cloud port browser navigated to /dashboard?view=jobs persists that complete URL, but this new early return restores from the catalog resource instead. Port resources contain the service root, so reopening the closed panel opens / rather than the saved page. Preserve the validated saved path, query, and fragment while resolving the host and resource ownership through the current provider.
Knowledge Base Used: Browser panels and automation
| if [ ! -x \"$path\" ]; then printf '%s' \"\(encoded)\" | base64 -d > \"$path\"; chmod 700 \"$path\"; fi | ||
| if ! pgrep -u \"$(id -u)\" -f \"$path serve\" >/dev/null 2>&1; then | ||
| nohup \"$path\" serve > \"$HOME/.cmux/display-service.log\" 2>&1 & | ||
| for i in $(seq 1 100); do [ -S /run/cmux-desktop/display-control.sock ] && break; sleep 0.1; done |
There was a problem hiding this comment.
Startup relies on fixed polling
This loop checks for the control socket every 100 ms, and wait_for_port similarly sleeps between readiness probes. These additions violate the repository requirement that production startup synchronization use an owning-subsystem signal or a dedicated cancellation-aware timeout/retry abstraction with tests. Socket-file existence also does not establish that the server accepts requests. Replace these waits with an explicit service-ready handshake and process/readiness events under a bounded, cancellation-aware deadline. This repository requirement must be satisfied before merging.
Context Used: Apply cmux's custom review rules from .github/review-bot-rules/ during review. Treat those files as the source of truth. Focus on production Swift and runtime changes, and ignore test-only scaffolding unless it makes production behavior worse. Do not... (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| @MainActor | ||
| enum CloudGuestDisplayScript { | ||
| private static let source = #""" |
There was a problem hiding this comment.
Command builder lacks scoped ownership
CloudGuestDisplayScript is a new caseless enum used solely as a static command-builder namespace. The repository's ownership rule explicitly requires this behavior to live on a constructable, injectable owner rather than a static-only namespace. Make the builder an instance owned by CloudDisplayCoordinator, or keep the implementation private to that owner, so installation policy and script construction have a scoped dependency boundary. This requirement must be satisfied before merging.
Rule Used: Flag new ambient global state in production Swift: a top-level (file-scope) func used as API, a top-level mutable var or a stub class/struct holding a global flag/once-token, a caseless enum/empty struct used purely as a static func/static let namesp... (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| var isDesktop: Bool { | ||
| model?.target.port == CmuxTuiSnapshotParser.desktopPort && remoteURL?.path == "/vnc.html" | ||
| (resourceID?.kind == .display && remoteURL == nil) || (model != nil && remoteURL?.path == "/vnc.html") |
There was a problem hiding this comment.
Filename overrides structured display identity
Once a remote URL exists, this expression ignores the resource kind and identifies a desktop solely by the /vnc.html path, now on any port. That classification controls the new connection deadline and recovery UI: a browser resource with that path can receive a false display timeout even when its page loads successfully. The repository requires reliable structured identity for correctness-critical state. Use the retained display resource identity as the authority instead of broadening the filename heuristic; this requirement must be satisfied before merging.
Rule Used: Flag correctness-critical detection/identity derived unreliably: a value the UI trusts (which agent is running, agent/session lifecycle and liveness, workspace/pane/surface identity, controls enable/route input) derived from a window/pane/terminal ti... (source)
| return surfaceOwnershipPolicy.rejection(for: resource.machine) == nil | ||
| } | ||
| if let raw = snapshot.browser?.urlString, URL(string: raw)?.path == "/vnc.html" { | ||
| return scope != .global && surfaceOwnershipPolicy.rejection(for: nil) == nil |
There was a problem hiding this comment.
Local browsers block Dock restoration
The new global-Dock check rejects every URL with the path /vnc.html, including a local noVNC page such as http://localhost:6080/vnc.html that has no Cloud ownership. Because restoreSessionSnapshot requires every saved panel to pass this check, one such browser prevents the entire window Dock—including unrelated terminals and browsers—from restoring. The matching creation check also prevents opening these local pages directly in a global Dock.
Preserve ordinary local-browser admission in global Docks, and apply Cloud provenance restrictions through the ownership policy rather than treating the filename as proof of Cloud ownership.
Knowledge Base Used: Desktop workspace experience
| if isCurrentLifecycleGeneration(generation), isRegisteredInCatalog() { publishDisplays() } | ||
| } | ||
| let snapshot = try await displayCoordinator.create() | ||
| guard isCurrentRefresh(lifecycle: generation, refresh: refresh) else { throw CancellationError() } |
There was a problem hiding this comment.
Refresh abandons successful display creation
A routine fleet refresh calls update(summary:) and refreshes the graph, incrementing refreshGeneration even when the VM and provider are unchanged. If that happens during guest creation, this guard throws after the coordinator has accepted the result and cleared its request ID. Deferred publication still adds the display to the catalog, but the requested pane never opens, and the action silently swallows the cancellation. Clicking New Display again allocates another resource instead of completing the original placement.
Use an operation-specific validity check that distinguishes provider retirement or relevant capability changes from an unrelated graph refresh.
Knowledge Base Used: Cloud services and identity
| func configure(model: CloudPortAccessModel, url: URL, resourceID: SurfaceResourceID? = nil) { | ||
| observationGeneration &+= 1 | ||
| unavailable = nil | ||
| self.resourceID = resourceID |
There was a problem hiding this comment.
Profile switching erases Cloud provenance
BrowserPanel.switchToProfile calls configure(model:url:) without the optional resource argument, so this unconditional assignment now clears the existing identity. A browser opened by navigating directly to a Cloud private address can have only a local catalog projection, leaving cloudAccess.resourceID as its sole Cloud resource identity.
After a profile switch, its session snapshot loses that identity and duplication treats it as an unowned URL. For a noVNC page in a VM-bound workspace, duplication is then rejected by the unowned-URL check, and the saved pane cannot restore under that ownership policy. Pass the current identity explicitly from the profile-switch path while retaining the intended reset for genuinely new configurations.
Knowledge Base Used: Browser panels and automation
| isAvailable = response.exitCode == 0 | ||
| } catch { | ||
| guard token == generation else { return } | ||
| snapshot = nil |
There was a problem hiding this comment.
Failed refresh forgets existing displays
After discovering display:2, a transient timeout or malformed discovery response now clears the saved snapshot. displayResources then falls back to only display:1, and the following graph refresh replaces the catalog: additional displays are either removed or replaced with daemon records that have no port. Running guest displays therefore disappear from the sidebar or become unopenable until another successful explicit discovery.
Retain the last validated display identities and endpoints separately from discovery availability rather than treating a failed request as an empty catalog.
Knowledge Base Used: Cloud services and identity
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
Comments Outside DiffThese findings sit on lines the diff does not cover, so they could not be posted inline. Each one leaves this list once its file changes.
|
| let previousPrivateAddress = info.privateAddress | ||
| refreshGeneration &+= 1 | ||
| refreshCoordinator.invalidate() | ||
| displayCoordinator.invalidate() |
There was a problem hiding this comment.
Routine refresh erases discovered displays
The registry calls update(summary:) for existing VMs even when their identity is unchanged. This unconditional invalidation clears the guest snapshot and disables New Display. The subsequent non-forced graph refresh skips guest discovery and rebuilds resources from the fallback containing only display:1, so additional displays disappear or retain daemon records without usable endpoints.
Preserve the validated catalog and creation availability across ordinary summary updates, and invalidate them only when the owning identity or authorization scope actually changes.
…wnership # Conflicts: # cmux.xcodeproj/project.pbxproj
…wnership # Conflicts: # Sources/Cloud/CloudTreeRowContentView.swift # Sources/Surfaces/SurfaceCatalog.swift # Sources/Workspace.swift # cmux.xcodeproj/project.pbxproj
The latest origin/main merge moved CloudTreeRowHoverButtons into its own file, and the conflict resolution kept main's copy, which dropped the displays-pool New Display button and its hasButtons entry. Re-apply them in the new file, and restore the blank lines the resolution stripped from SurfaceCatalog.swift so the PR diff stays limited to behavior changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The blank lines restored in the previous commit put the file seven lines over the Swift file-length budget, so drop them again. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wnership # Conflicts: # Sources/Workspace.swift
#13196 made SurfaceCatalog.validateOwnership refuse a destination that is not a live workspace. That rule stays. Tests that projected, restored or opened browsers into a made-up workspace ID now failed with destinationNotFound, timed out waiting on a provider that was never called, or passed a later check for the wrong reason. LiveWorkspaceFixture registers real Workspace objects and hands the catalog a CloudWorkspaceRenameService that resolves them, the way the app's composition root does. Its workspaces() list stays empty so the native projection coordinator does not start mirroring into them. For tests on SurfaceCatalog.shared, withAppRegistration registers the TabManager as a windowless main-window context for the test body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 6db7593 main merge into 13192-cloud-display-ownership (manaflow-ai#13196) put back main's [desktopDisplayResource()] pools in CmuxTuiSurfaceProvider's refresh, publish and delta paths. 23c807b restored the display coordinator lifecycle but not these pools, so a display created beyond display:1 vanished on the next daemon publish, and a delta that touched it removed it. Restore eb3edd7's displayResources pools, the display-kind delta check, and the injectable displayCoordinator. Drop the now-unused desktopDisplayResource(). Adds a test that creates display:2 through the provider and asserts that both displays and their ports survive a full publish and a display delta. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Product (#13196 regression): the 6db7593 main merge into 13192-cloud-display-ownership took main's CmuxTuiSurfaceProviders.swift and CloudPortRoutePlan.swift, dropping eb3edd7's per-display ports. 23c807b restored the display coordinator but not these hunks, so withPrivateBrowserURL rewrote every display to 6901 and a daemon pointer without a discovered target fell back to display 1. Restore both hunks and drop the duplicate port-less privateDesktopURL overload. Tests: - CloudDisplayCatalogTests: 178d35e (#13196) made every guest command run `list` as a readiness probe, so fakes dispatching on " list" answered creation with the list catalog. Dispatch on the create action line, and pin the command shape. - CloudPortOpenRegressionTests: 178d35e (#13196) filters RFB/noVNC ports only when the display catalog owns them (displayPortsOwned). Assert both the desktop and non-desktop results. - CloudTreeOneMachineManyWorkspacesTests: #12740 added a final Resources section under each machine. Expected trees include it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
#13196 made SurfaceCatalog.validateOwnership refuse a destination the app cannot resolve, so SurfaceCatalog.shared no longer relinked a restored projection into a TabManager that no main window owns. Register it for the test body with LiveWorkspaceFixture.withAppRegistration, as #13651 does for CloudClosedPanelRestoreTests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(terminal): restore the presented-surface fixture contract so package tests compile `swift test --package-path Packages/macOS/CmuxTerminal` has not compiled on main since merge 38b32bb: TerminalSurfaceRendererCallbackTests calls `PresentedSurfaceFixture(installRendererCallbacks: false)`, but the fixture initializer only takes `windowVisibleAtCreation`. The flag came from the issue-2824 branch (77c46d0, 6fe70f6) and was dropped when the fixture was reworked for the native callback lifecycle (8008b7c, 97c6088); the merge re-applied the test call without the fixture side. Package tests only register render callbacks through the fixture (RendererCallbackTestSupport), so a fixture that skipped registration would leave `cmux_test_ghostty_renderer_present` with nothing to route. Restore the calls to `PresentedSurfaceFixture()`, the combination that was green at 0251a44. Verified locally: 294 tests in 37 suites pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(popover): stop rewriting the presentation binding during view update `ArrowlessPopoverAnchor.updateNSView` calls `coordinator.dismiss()` on every update where `isPresented` is already false. PR #13311 (01f0a50) made `dismiss()` write `isPresented = false` on the no-popover path, which SwiftUI reports as "Modifying state during view update" because updateNSView runs inside the view update. The sidebar footer mounts two anchors whose parents re-evaluate on every `selectedTabId` change, so every workspace switch emitted faults and `SidebarWorkspaceSwitchLayoutFaultTests` failed with 15 of them. Pass `resetPresentation: false` from updateNSView: the binding is already false on that branch, so the write was redundant. `popoverDidClose` still resets the binding when AppKit closes a live popover. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cloud): satisfy the plus menu's sign-in gate and route Cmd+Y through a registered window `NewCloudWorkspaceShortcutTests` never passed: the plus menu deliberately hides its Cloud rows unless the account is signed in (#12305, mirroring the command palette and File menu), and a test `AppDelegate()` has no account flow. The XCTest version crashed the app host on `rows[1]`, xcodebuild restarted it, and the lenient gate accepted the partial run; #13178 now rejects that and #13193 migrated the suite to Swift Testing, so the failures became visible. Inject the signed-in state through the existing `isAuthenticated:` seam, add signed-out coverage of the gate, route the Cmd+Y event through a registered main window (as `testReboundKeyRoutesAndOldKeyDoesNot` does, since shortcut routing bypasses events bound to windows the delegate cannot resolve), and register a window context for the unavailable-Cloud check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(chrome): assert the composited Bonsplit chrome contract instead of a stale alpha hex `WorkspaceChromeColorTests` expected `bonsplitChromeHex` to return `#1122337F` (theme with opacity as alpha). Since b6d3470 (2026-05-19) `compositedTerminalColor` composites the theme over the window base and returns an opaque color, and 4cbb354 made that deliberate: Bonsplit derives its tab glyph contrast from the rendered backdrop, and an `#RRGGBBAA` hex reproduces the white-on-white bug it fixed. The tests kept failing unnoticed because the app-host gate only counted "unexpected" XCTest failures until the strict check (acedf3f) reached main through #12053. Exercise the `chromeBackgroundColor` seam that every production call site uses with literal expectations, verify the ambient default path against the resolver plus independent blend arithmetic so the result is deterministic under any host appearance, and keep coverage for opaque, shared-backdrop, pane-clear, pane-border, and explicit translucent chrome colors. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(tmux): wait for the main window to adopt the mirror pane before sending keys `RemoteTmuxMirrorPaneInputMappingTests` required `panel.surface.uiWindow != nil` right after selecting the workspace. A manual-I/O mirror pane spawns eagerly in its hidden bootstrap window, which `uiWindow` deliberately excludes, and the main window's portal adopts the pane host only once AppKit and SwiftUI get run-loop time; `waitForLiveSurface` returns immediately for an already-live surface, so the check ran before adoption and the four key-delivery tests failed at line 176. The failure was hidden until the strict app-host gate. Order the harness window front and pump the run loop until the surface and its native view are in that window, as the other hosted-view input suites do, and assert against the harness window instead of any window. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(browser): report the refused connection through the navigation delegate `browserPanelRetriesDiscardedRestoreAfterConnectionRefused` waited up to 30s for WebKit to fail a provisional load to a bound-but-unlistened loopback port. On the hosted app-host runners that failure never arrives: the load neither fails nor commits, so the test timed out at line 147 (no "provisional navigation failed" log line appears for it in any shard). Stop the in-flight load and report `NSURLErrorCannotConnectToHost` for the attempted URL through the panel's real navigation delegate, the pattern `BrowserFailedNavigationReloadTests` already uses. The restore bookkeeping, error page, `restore_pending` clearing, and retry policy under test run unchanged and every assertion is kept. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cloud): scope reserved-workspace cleanup to its pending card and return the bind window PR #13202 (5d616a7) imported `CloudMachineWorkspaceAdoptionTests` and `CloudMachineWorkspaceResolutionTests` from the still-open 13141-cloud-vm-workspace branch without the production changes they assert, and its own run skipped the app-host shards, so they landed red: - `NewMachineSheetPresenter.closeReservedWorkspace` closed the whole workspace, which is a no-op for the last tab and discards user panes added next to a creating card. Port the branch behavior: remove only unadopted loading cards owned by the cancelled machine, clear that binding, keep user content, and give a last-tab loading workspace a local anchor first. `MachineCreateCoordinator` passes the created or reconciling machine id so cancelling machine X cannot clear a binding to machine Y (the tests now assert that scoping). - `v2WorkspaceCloudVMBind` now returns `window_id` alongside the workspace refs, as the bind acknowledgement test expects. - The resolution test selected "first" while the fixture's terminal key is "term-first"; use the key so the placement resolves as intended. - `CloudPaneCreationRetryTests` asserted `discarded` synchronously after the projection returned, but the coordinator applies its generation fence behind `CloudOperationContext.withPhase`'s recorder await, which suspends on a cold per-suite process. Settle with bounded yielding before asserting. Runtime behavior change (cancelled Cloud create cleanup); needs dogfood. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): restore deferred socket connection, explicit SSH control options, and Codex stop idle observations Branch commit c8bfb58 ("fix: repair failures exposed by strict app-host CI", 2026-09-10) made `cmux vm dev|layout|env` finish local validation and dry runs before opening the socket, passed the caller's explicit ssh options into `userConfiguredControlOptions(fromSSHConfigOutput:explicitOptions:)` so a normalized `ControlPersist=0` from `ssh -G` is not mistaken for host customization, and published `idleObserved` for transcript-terminal prior turns before a legacy Codex Stop so the journal drops an obsolete running turn. The branch's later single-parent commit 38b32bb reverted `CLI/cmux.swift` to main's version while keeping the stricter fixtures, and PR #12053 landed that inconsistent state; the fixtures (CLIVMDevTests, CLIVMLayoutEnvTests, the SSH sharing tests, and the Codex missed-prompt Stop test) have failed since. Restore the three CLI changes. `SocketClient.configureAuthentication` and `SocketPasswordResolver` already exist on main, and the CmuxFoundation overload landed with the branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cli): align CLI integration fixtures with the shipped hook, SSH, and session-list contracts The main copy of `CLINotifyProcessIntegrationRegressionTests` predates several shipped contracts that the lenient app-host gate never enforced and that 38b32bb reverted on the issue-2824 branch: Claude hook acks print `{}` (#7963), Codex resume bindings require rollout evidence so fixtures carry a `session_meta` transcript (#10100), SessionStart publishes a binding so `/clear` counts start after the clear, fresh-terminal SSH startup commands are script paths that the support decoder must read, cmux control-path options follow a resolved `ssh -G` (#8308), and `ssh session list` reports the localized "remote state unavailable" summary with `--json` detail (#9971). The missed-prompt Codex Stop test now asserts the restored `agent.idle.observed` for the prior turn. Flagged for review: the three `ssh pty-attach` cases now expect only `workspace.remote.pty_bridge` once the endpoint is established, matching #12726's `preserveLifecycleForRecovery`; if pre-READY bridge failures were meant to keep reconciling, the flag should be set only after READY instead. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(workspace): restore the app-host repairs for fork, focus recovery, and shortcut routing Merge 8d33410 on the issue-2824 branch took main's copy of `cmuxTests/WorkspaceUnitTests.swift` wholesale and discarded the branch's September repairs (c8bfb58, c402b9c, 2127d97); PR #12053 then landed without them while the strict gate started counting the failures. Restore them against current production, plus the shortcut-suite fixes: - Fork in a remote workspace: `sshBootstrapArguments` has used `/usr/bin/ssh` since #9114, and agent socket propagation requires a socket that exists on disk (3bf87e3), so bind a real unix socket and expect the absolute path. - Fork Conversation context actions dispatch asynchronously (#7259, #8173); await the fork panel before asserting. - Git branch and pull request updates publish through `sidebarObservationPublisher` since #6226, not `objectWillChange`. - Config sanitization: `addWorkspaceIfActive` rebuilds templates from the font-size lineage since #8543; override the lineage hook instead. - Focus recovery: AppKit focus is authorized only for a registered, selected workspace whose window carries the main-window identity; use the `TerminalPortalTestWorkspace` fixture and the same registration/pump path as `WorkspaceTerminalFocusRecoverySwiftTests`. - Shortcut routing: clear both Cloud defaults after 9c2ba78 swapped them; neutralize Ghostty's imported goto_split fallback (⌘] by ANSI keyCode) in the unshifted-symbol test and assert the digit shortcut does not match; wait for the async runtime start before judging keyDown forwarding (skip loudly without a live surface); wait for the portal to mount before the second-Escape check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(restore): keep a restore identity when the persisted surface id collides Since #13098 (e0e77eb), `newTerminalSurfaceOutcome` treats a nil `restoredSurfaceId` as an interactive create and routes it to the selected pane's Cloud source. Session restore passed nil whenever the persisted panel id was still live (duplicate-workspace or restore-into-live), so a legacy managed-Cloud SSH workspace restored into a live manager was turned into a remote tab create: the scaffold panel stayed startup-suppressed with no initial command and `TabManagerSessionSnapshotTests` failed to unwrap it. Mint a fresh UUID on collision instead of nil; it is free by construction, so the restore keeps its identity and the old-to-new remap works as before. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(terminal): align snapshot, projection, terminal, and browser fixtures with shipped contracts Stale expectations and fixed-spin timing in app-host suites that the lenient gate never counted: - Cloud-projected panes restore as manual-mirror reservations with a staged remote identity (#12675), not a local placeholder projection. - Restore reuses the persisted runtime id when free (aff0e32), so assert liveness rather than a new id; the catalog ignores writes for a Cloud machine without a registered provider (#11877), so register the fixture provider before publishing. - Wheel sync requires an authoritative scrollbar response (bbc3edf); the fixture now answers like `AuthoritativeScrollbarSurfaceView`. - Search overlay mount, first-responder focus, runtime creation, and the visibility-restore redraw run through deferred main-actor tasks; wait for them with the class's `waitUntil` helpers instead of fixed run-loop spins. - Owning the socket path lock is definitive (959f38a): a refused inode left by a dead listener is replaced, so the restarted listener accepts. - The split-divider hit band extends `dividerHitExpansion` past the divider (667cc43); derive the pass-through boundary from the constant. - Browser page background blends against Ghostty's effective terminal color scheme, which is host dependent; read the same preference the product uses. - Cloud Machines defaults on in dev builds (#12318); pin the toggle off for the default-mode palette contract. - Prepared navigation requests keep the caller's cache policy (#13003). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(browser): align lifecycle, identity, host-view, loopback bridge, and portal rebind tests with shipped contracts Browser app-host suites that the lenient gate never counted: - `BrowserPanelWebViewLifecycleTests`: out-of-range discard delays are rejected to the default (the cmux.json loader relies on the nil), and the panel's own `isLoading` stays true for the indicator floor, so wait for both flags and assert no discard blockers before discarding. - `browserNavigationUsesEmbeddedWebKitIdentity`: WebKit reports the native identity as nil or "" (#9482); accept either. - `WindowBrowserHostViewTests`: production routes Dock-divider hits by yielding to AppKit so the live sidebar tracker receives them (#10902, e2e3818); the tests asserting the portal owns and forwards the hit were merged red against a design that never shipped. Realign the stale-frame test to the pass-through contract and remove the four own-and-forward tests with their fixtures. **Flagged for review:** if own-and-forward is still wanted, that is a hit-testing product change for its own PR. - `testRemoteWorkspaceRuntimeBridgeAliasesMultipleLoopbackPortsFromSamePage`: the navigation delegate restarts main-frame loads to apply the user-agent policy (11c6efe); a direct `loadHTMLString` with an HTTP base skipped that step, so the data load was cancelled and replayed as a deferred request. Apply the identity first and assert the document is current. - `portalRebindPreservesDocumentAndRoutesRefreshToTheSameWebView`: the portal host is a theme-frame sibling of `contentView` for non-glass windows (#12929); assert same-window instead of descendant-of-contentView. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(terminal): wait for asynchronous runtime creation in the remaining direct-interaction tests The same "Expected runtime surface before ..." precondition that 9f258dd made wait for the deferred runtime start still sampled after a fixed run-loop spin in the detach-race, close-lifecycle, repeat-key, and repeat-IME tests, and CI on the branch head showed them failing that way. Use the class's `waitUntil` for those four sites too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cli): model surface.respawn for respawn-pane and give the Codex stack fixture rollout evidence `cmux respawn-pane` has sent `surface.respawn` instead of `surface.send_text` since #5465 (8cafcc3); the window-flag fixture's mock still rejected that method, so the CLI exited 1. Answer `surface.respawn`, asserting the window and surface ids, `tmux_start_command`, and that the shell-invoked command carries the user command but never the `--window` flag, which is the test's intent. The Codex interrupted-stack fixture's transcript had no `session_meta` line, so `CodexSessionResumeVerifier` (#10100) found no rollout evidence and the prompt published `surface.resume.clear`. Prepend the session line as the other Codex fixtures do. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(socket): keep mobile.panel.artifact.fetch off the local socket and advertise the served artifact reads 02c1ba4 removed `mobile.panel.artifact.fetch` from the socket worker methods because it needs the authenticated mobile execution context, but that half of the change was lost in a merge: the policy still routed fetch to the worker lane, which has no handler for it, so the local socket answered `internal_error` instead of the `method_not_found` boundary that `TerminalControllerSocketSecurityTests` pins. Restore the removal. `system.capabilities` never advertised `mobile.panel.artifact.stat` and `.thumbnail` although both are served on the worker lane (04ff18e added them only to the test's expectation); advertise those two and keep fetch out. The remote relay allowlist is unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(remote): align tmux seed transport, SSH, socket command, and port scanner fixtures with shipped behavior - `RemoteTmuxPaneSeedTransportTests`: manual-I/O mirror panes spawn eagerly (#9272) and stay runtime-backed after portal churn (#9769), so the pane renders its assigned grid before a seed arrives and nothing was retained. Publish a larger tmux pane grid than any test surface applies so the retention under test sees the lag it exists for. - `SSHRemoteCWDRegressionTests`: the persistent-PTY exec helper runs only with `protectsFromHangup: true` (cd9f34f); pass it, and widen the first-spawn guard. - `SSHDeepSleepReattachTests`: a Cloud-owned workspace rejects launch overrides on splits (#13098); create the custom-identity pane before configuring the remote connection. - `SSHConfiguredRemoteCommandHostTests`: ssh-pty-attach validates the bridge `daemon_version` before dialing (#12726); the mock now reports one. - `CloudManualMirrorTransportTests`: the pane failure card uses the short title since 9bf6cb8. - `SurfaceSocketCommandTests`: `vm.workspace_new` admits its optimistic workspace through the active main window (#13152, #13155), so bind a bare window to the fixture context; a receipt without a starter terminal costs one snapshot (6d43ea6). - `PortScannerPublicationTests`: the forced-result acknowledgement hops off the main actor, so an unchanged port set may be deduped under a later refresh; drain publications until the retirement lands. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test: align mobile artifact fetch lane assertion * repair: close remaining full-suite contracts * test(restore,sidebar): align two stale contracts with shipped behavior testRemoteWorkspaceAutoResumeKeepsRemoteStartupCommand asserted the restored remote panel's local spawn cwd equals the remote-host path. It never can: OneShotTerminalLauncherStore.enterableWorkingDirectory rejects a path that is not locally enterable, which is what keeps the owning shell valid (#7031). The remote cwd does survive restore — as the panel's trusted remote directory report, and as the `cd` prefix the resume input carries (both already asserted). Assert it where it actually lives and pin the spawn cwd to nil. testSidebarPullRequestsTrackFocusedPanelOnly expected a background panel's PR to be hidden from sidebarPullRequestsInDisplayOrder(). That list is documented as the workspace's deduplicated rows in pane/tab order, both consumers (taskStatusSignals, the control-sidebar snapshot) want every panel, and the sibling branch test asserts the same all-panel model. Focus scoping lives in the `pullRequest` binding, which the test already covers. Renamed to match. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: give Cloud catalog tests live destination workspaces #13196 made SurfaceCatalog.validateOwnership refuse a destination that is not a live workspace. That rule stays. Tests that projected, restored or opened browsers into a made-up workspace ID now failed with destinationNotFound, timed out waiting on a provider that was never called, or passed a later check for the wrong reason. LiveWorkspaceFixture registers real Workspace objects and hands the catalog a CloudWorkspaceRenameService that resolves them, the way the app's composition root does. Its workspaces() list stays empty so the native projection coordinator does not start mirroring into them. For tests on SurfaceCatalog.shared, withAppRegistration registers the TabManager as a windowless main-window context for the test body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: normalize pbxproj after live workspace fixture Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: admit agent renames of agent-owned accepted Cloud names e6926fb (#13403) made submitCloudPanelRename run admitsTerminalRename for automatic names. Its last clause only admitted replacing an accepted name when the local panel still carried `.auto` provenance, but since 1e1d319 an accepted daemon name reconciles locally as `.remote` and the owner lives in the tab's nameAuthority. Every agent title after the first accepted one was refused, so "Failed agent rename keeps the accepted title", "Mirroring an agent-named placement does not block its next agent title" and "An older automatic result and old snapshots cannot replace an accepted name" failed on their second agentName call. Admit an automatic rename when no write is pending and the accepted tab name is owned by the daemon's `auto` authority. User-owned names, pending writes and local user titles on any projection still refuse it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: expect adopted machine rows to keep the pending create identity #12919 (021f792) made New Machine creation optimistic: once a running create's machine appears in the fleet list or catalog, its row keeps the `pending-machine:<operation>` node ID so selection and expansion survive adoption (see committedProjectionKeepsThePendingNodeIdentityUntilFleetAdoptsIt). pendingRowStepsAsideOnceItsMachineHasARow still expected `machine:<id>`. Assert the new identity and that the row is the adopted machine, not a stand-in: the stand-in is gone, one row remains, and it shows the created machine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: restore per-display noVNC targets and refresh stale Cloud tests Product (#13196 regression): the 6db7593 main merge into 13192-cloud-display-ownership took main's CmuxTuiSurfaceProviders.swift and CloudPortRoutePlan.swift, dropping eb3edd7's per-display ports. 23c807b restored the display coordinator but not these hunks, so withPrivateBrowserURL rewrote every display to 6901 and a daemon pointer without a discovered target fell back to display 1. Restore both hunks and drop the duplicate port-less privateDesktopURL overload. Tests: - CloudDisplayCatalogTests: 178d35e (#13196) made every guest command run `list` as a readiness probe, so fakes dispatching on " list" answered creation with the list catalog. Dispatch on the create action line, and pin the command shape. - CloudPortOpenRegressionTests: 178d35e (#13196) filters RFB/noVNC ports only when the display catalog owns them (displayPortsOwned). Assert both the desktop and non-desktop results. - CloudTreeOneMachineManyWorkspacesTests: #12740 added a final Resources section under each machine. Expected trees include it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: publish every guest display from daemon graph updates The 6db7593 main merge into 13192-cloud-display-ownership (#13196) put back main's [desktopDisplayResource()] pools in CmuxTuiSurfaceProvider's refresh, publish and delta paths. 23c807b restored the display coordinator lifecycle but not these pools, so a display created beyond display:1 vanished on the next daemon publish, and a delta that touched it removed it. Restore eb3edd7's displayResources pools, the display-kind delta check, and the injectable displayCoordinator. Drop the now-unused desktopDisplayResource(). Adds a test that creates display:2 through the provider and asserts that both displays and their ports survive a full publish and a display delta. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: fence the fake Cloud projection reply with its mutation cursor aDetachedTerminalDropsItsStaleTabBeforeMoving installs a daemon graph since c402b9c (#13403). When the placement lane drains it reconciles against that graph, which predates the projected tab, and clears tab_projected. The real reply always carries a mutation cursor (CmuxTuiSnapshotParser.placedTab requires one) that fences exactly this; the fake returned none. Give the fake a projectCursor and set it ahead of the installed graph. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: register the restored window before relinking Cloud projections #13196 made SurfaceCatalog.validateOwnership refuse a destination the app cannot resolve, so SurfaceCatalog.shared no longer relinked a restored projection into a TabManager that no main window owns. Register it for the test body with LiveWorkspaceFixture.withAppRegistration, as #13651 does for CloudClosedPanelRestoreTests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: wait for the relay ports kick, not the first relay line Since #8442, a relay prompt also reports shell state. Both RPCs run in separate background children, and the zsh prompt-refresh test waited only until the log was non-empty, so it could read report_shell_state alone. Wait (deadline-bounded) for the ports_kick line itself, in the zsh test and its bash sibling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(cli): relay output of an SSH session that ends right after auth The expect wrapper watches the first two seconds after sending the password for a rejection with log_user 0. A session that authenticates and exits inside that window hit the eof branch and its output was dropped. Flush the buffered output before exiting. #13207 replaced the test's 9 s sleep with a FIFO release, which is what exposed this. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: restore the no-connection path for rejected vm dev input 4120e35 taught runVMDev that invalid input never connects: keep the listener open, then check its backlog after the CLI exits. #13207 dropped that again, so each rejected run waited 60 s for a mock-server expectation that only the listener closing (after the wait) fulfills. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: drain the terminal while the SCP host-key failure runs The pty case read the master only after the CLI exited. A pty's output queue holds about 1 KiB and the host-key failure report is longer, so the CLI blocked writing stderr until the 30 s timeout. Read the master on a thread while the CLI runs. The test starts its own sshd; no runner host dependency is involved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: fail fast when a gated Cloud call ends before its fake is entered Every app-host restart in the 09-21/09-22 shard logs I sampled was a Swift Testing time-limit hit in one of two suites, and each hit relaunches the test host: - SurfaceCatalogTests: when catalog.project threw before reaching the provider (destinationNotFound, fixed by 3dc91b2), each gated test parked in MaterializeGate.waitUntilEntered() until the 300 s limit. Five tests restarted the host one after another, about 25 minutes per shard. - CloudDisplayCatalogTests: when the fake no longer recognized the create command (fixed by 70eaf44), create() failed before the exec started and `await started.result` parked until the 60 s limit. The waits now also end when the caller's task finishes, so the next such setup failure is an ordinary failed #require. The cancellation test also waits on its own cancellation signal instead of a 60 s Task.sleep. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: reveal a retired terminal through the portal rebind, as the app does Since #12607 hiding a terminal removes its hosted view from the window, and only a bind reinstalls it. The test flipped portal visibility on the detached view, so no size commit could ever land and the final shrink check failed every run. Rebind like TerminalPortalReconciliation and drop the wait loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: pin key-window status in terminal focus suites The app-host test process runs headless and is usually not the active app, so makeKeyAndOrderFront never makes a programmatic window key. Terminal focus paths gate on isKeyWindow (automatic first-responder apply, focus redraws, deferred focus reapply, ensureFocus window activation), so these suites passed only when an earlier test in the shard had activated the app. Share the existing KeyStatusTestWindow and use it in WorkspaceTerminalFocusRecoveryTests, WorkspaceTerminalFocusRecoverySwiftTests and TerminalNotificationDirectInteractionTests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: report which layer holds off-plan tmux mirror geometry offPlanGeometryWithUnchangedSizingInputsReconverges fails every run with all three output-parity re-arms spent and the hosted view still at the perturbed 0.8 divider. A standalone bonsplit replay of the same sequence converges, and a sibling test without a bound (portal-visible) workspace heals the same displacement. Add the live split view's arranged widths and the split model's imposed extent to the failure message so the next run shows whether bonsplit refused the apply, the imposition was cleared, or the portal did not follow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: restore Cloud projections before the workspace is published TabManager.restoreSessionSnapshot restores each workspace's surface projections before it assigns tabs, so SurfaceCatalog.restore's ownership check could not find the destination and silently dropped every restored remote projection. Check ownership against the workspace being restored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: repair stale and broken app-host expectations - CLI ssh attach: count identity UUIDs exactly; #11497 added an auth token uuidgen - device mirror directories: give the fake device trusted presence (e3d424c gate) - font zoom mirrors: deltas from the 8pt inherited base (e6926fb arithmetic) - Computer Use refresh: fake daemon reply was invalid JSON in a raw string (#13599) - quit alert: compare button alignment rects, not padded frames - remote split cwd rescue: cwd travels as CMUX_REMOTE_INITIAL_CWD since #12054 - default freestyle split: Cloud-owned splits route to Cloud since fa5dc4c Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: activate the app host before waiting for terminal focus `focusTerminalForTesting` waits for `window.isKeyWindow` before it hands first responder to the terminal, but `makeKeyAndOrderFront` only makes a programmatic window key while the test host is the active app. The app-host process starts inactive under `xcodebuild test`, so the wait succeeded only when an earlier test in the shard happened to activate the app. Its callers use a real `createMainWindow()` window, which cannot be swapped for `KeyStatusTestWindow` the way the pinned focus suites were. Activate explicitly, and give the wait a CI-appropriate timeout: the activation and the key-window transition land on later main run-loop turns, and the pump's one-second default expires before the window goes key on a contended runner. Observed on run 35743588302: `plainTerminalTextDoesNotResolveAppShortcutContext` (shard 1/6), `keyboardCopyModeKeyClearsTerminalUnread` and `workspaceFontSizeShortcutPreservesBackgroundTerminalUnread` (shard 2/6) all fail at this helper's return value, and shard 6/6 shows the same process state directly as `NSApp.keyWindow -> nil`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: give standalone terminal fixtures a live portal authority `setVisibleInUI` and `setActive` fold their request through `Workspace.portalRenderingEnabled(for:)`, which denies any workspace id the app delegate cannot resolve to a selected tab. Three direct interaction tests build their surface with `tabId: UUID()`, so once any earlier test installs an `AppDelegate.shared` the authority denies the portal, the hosted view is never actually made visible or active, and the fixture stops exercising the behavior it asserts: the surface never takes Ghostty focus and never schedules a visibility-restore redraw. Register a real selected workspace and build the surface with its id, so the fixture gets the same authority the app grants the selected tab. This is the fixture, not the product: the authority check is deliberate, and denying an unresolvable workspace is what keeps queued portal callbacks from reviving an inactive workspace. Fixes on run 35743588302 shard 4/6: `testKeyDownRecoveryDoesNotReplayFocusAfterResponderMovesAway`, `testVisibilityRestoreRefreshesSurfaceWhileTerminalIsInactive`, and `testDirectFirstResponderFocusRefreshesCursorStateAfterForeignResponder`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: normalize project.pbxproj after merging main Run 35763999808 failed `Fast static checks` before any macOS job could start, so the app-host test fixes on this branch were never exercised: error: cmux.xcodeproj/project.pbxproj is not normalized. Run scripts/normalize-pbxproj.py to fix. The branch head alone checks clean. CI builds the merge of this branch into main, and that merge is what leaves the file unnormalized, so the failure does not reproduce without merging main first. The change is one build-phase entry moving back into alphabetical order. `LiveWorkspaceFixture.swift in Sources` keeps the same UUID and the same occurrence count, and `scripts/lint-pbxproj-test-wiring.sh` still reports ok across 1049 test files, so no target lost a source file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Drop a cancelled fork-probe request, and name the inputs behind the last two app-host failures (#13744) * fix: drop a cancelled fork-probe request while it waits behind an active probe A second fork-availability request for the same panel with a different fallback snapshot parks in `applyPendingForkValidations`' contention branch: it restores its request to the pending queue and awaits the active probe. That wait had no cancellation handler, unlike every other wait in this type, so a cancelled caller kept its request in the queue until the active probe finished. The probe's completion restarts the single-flight refresh for whatever is still pending, and the cancelled request rode along: its fallback was probed and its result replaced the surviving request's validation for that panel. Give the wait the same cancellation handler its siblings have, and drop the waiter's own pending requests when it is cancelled, so the restart that follows the active probe sees only live requests. Covers cancelledSharedForkProbeRefreshPreservesSurvivingFallbackSnapshot, which failed when the restart won the race against the cancelled task's own cleanup. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: name the input behind the off-plan re-arm and the sidebar reveal double pass Both failures currently report only their outcome, and each has two incompatible explanations that the message cannot tell apart. offPlanGeometryWithUnchangedSizingInputsReconverges reports `rearms=3` with `imposed=nil`. That is either three recovery passes that ran and failed to impose, or a re-arm budget already spent before the perturbation, in which case no recovery pass ran at all. Bracket the recovery window with the existing DEBUG sizing counters and report the budget at perturbation time plus the planned outers. visibilityToggleKeepsAppKitTableContainerMounted reports exactly two projections per row. Report how many the first run-loop turn produced and which async signals landed inside the reveal window (the workspace directory channel, workspace order, the shared agent index), since the hidden phase queues main-queue work that can land during the reveal. No assertion changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * test: repair the remaining app-host failures and the shard-2 host relaunches (#13759) Most of these share one cause: the app-host test process runs headless and is not the active app, so AppKit and WebKit withhold state the tests assumed. NSApp.keyWindow is nil, makeKeyAndOrderFront never makes a window key, occlusionState never carries .visible, and behind XCTest's shielding window WebKit suspends requestAnimationFrame entirely. The four shard-2 host relaunches were one hung await: renderMarkdown awaited two animation frames inside callAsyncJavaScript (cd9f34f), which behind the shielding window never fire, so four MarkdownPanelTests cases each hung until XCTest's five-minute allowance killed the host. Every WebKit call in that file now fails fast with a stated reason instead of hanging, renderMarkdown waits on the viewer's own render contract, and the scroll restore the rAF await was papering over no longer clobbers a scroll that lands after a content update. Two real product regressions: windowless shortcut events stopped pruning the orphaned main-window context (6d7b23f), and ticket minting read the routes and the v2 installation identity as two separate cache reads. No assertion is weakened; several are strengthened. Rebased onto fix/app-host-green after 87a7016 landed a different remedy for the same key-window cause. The three focus tests this branch fixed through a test-target MainWindowKeyStatusPin (a swizzle of NSWindow.isKeyWindow, needed because CmuxMainWindow is final) are exactly the three that commit fixes by activating the app host and widening the pump's timeout. The pin is dropped: activating the host makes isKeyWindow true for real rather than forcing the answer, and a process-wide swizzle installed for one helper is a worse neighbour to the rest of the shard. cmuxTests/AppDelegateMainWindowTestingSupport.swift is no longer touched by this branch. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * test: measure Cloud title ink beyond the icon and trace measured invalidations Real rendering at 75% and 1x reproduces disconnected icon ink misidentified as the title. Retain the alignment tolerance and verify that displaced text still fails. Enable DEBUG tracing only during the existing sidebar and minimal-mode measured intervals. Count assertions and timing remain unchanged. * test: preserve carrier preparation before fleet discovery * fix: retain independent cloud carrier prewarming --------- Co-authored-by: austinpower1258 <austinwang115@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…meouts (#14210) * test(terminal): restore the presented-surface fixture contract so package tests compile `swift test --package-path Packages/macOS/CmuxTerminal` has not compiled on main since merge 38b32bb191: TerminalSurfaceRendererCallbackTests calls `PresentedSurfaceFixture(installRendererCallbacks: false)`, but the fixture initializer only takes `windowVisibleAtCreation`. The flag came from the issue-2824 branch (77c46d0c02, 6fe70f6f68) and was dropped when the fixture was reworked for the native callback lifecycle (8008b7c062, 97c60889b4); the merge re-applied the test call without the fixture side. Package tests only register render callbacks through the fixture (RendererCallbackTestSupport), so a fixture that skipped registration would leave `cmux_test_ghostty_renderer_present` with nothing to route. Restore the calls to `PresentedSurfaceFixture()`, the combination that was green at 0251a448c9. Verified locally: 294 tests in 37 suites pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(popover): stop rewriting the presentation binding during view update `ArrowlessPopoverAnchor.updateNSView` calls `coordinator.dismiss()` on every update where `isPresented` is already false. PR #13311 (01f0a503b9) made `dismiss()` write `isPresented = false` on the no-popover path, which SwiftUI reports as "Modifying state during view update" because updateNSView runs inside the view update. The sidebar footer mounts two anchors whose parents re-evaluate on every `selectedTabId` change, so every workspace switch emitted faults and `SidebarWorkspaceSwitchLayoutFaultTests` failed with 15 of them. Pass `resetPresentation: false` from updateNSView: the binding is already false on that branch, so the write was redundant. `popoverDidClose` still resets the binding when AppKit closes a live popover. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cloud): satisfy the plus menu's sign-in gate and route Cmd+Y through a registered window `NewCloudWorkspaceShortcutTests` never passed: the plus menu deliberately hides its Cloud rows unless the account is signed in (#12305, mirroring the command palette and File menu), and a test `AppDelegate()` has no account flow. The XCTest version crashed the app host on `rows[1]`, xcodebuild restarted it, and the lenient gate accepted the partial run; #13178 now rejects that and #13193 migrated the suite to Swift Testing, so the failures became visible. Inject the signed-in state through the existing `isAuthenticated:` seam, add signed-out coverage of the gate, route the Cmd+Y event through a registered main window (as `testReboundKeyRoutesAndOldKeyDoesNot` does, since shortcut routing bypasses events bound to windows the delegate cannot resolve), and register a window context for the unavailable-Cloud check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(chrome): assert the composited Bonsplit chrome contract instead of a stale alpha hex `WorkspaceChromeColorTests` expected `bonsplitChromeHex` to return `#1122337F` (theme with opacity as alpha). Since b6d3470668 (2026-05-19) `compositedTerminalColor` composites the theme over the window base and returns an opaque color, and 4cbb354cfc made that deliberate: Bonsplit derives its tab glyph contrast from the rendered backdrop, and an `#RRGGBBAA` hex reproduces the white-on-white bug it fixed. The tests kept failing unnoticed because the app-host gate only counted "unexpected" XCTest failures until the strict check (acedf3f244) reached main through #12053. Exercise the `chromeBackgroundColor` seam that every production call site uses with literal expectations, verify the ambient default path against the resolver plus independent blend arithmetic so the result is deterministic under any host appearance, and keep coverage for opaque, shared-backdrop, pane-clear, pane-border, and explicit translucent chrome colors. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(tmux): wait for the main window to adopt the mirror pane before sending keys `RemoteTmuxMirrorPaneInputMappingTests` required `panel.surface.uiWindow != nil` right after selecting the workspace. A manual-I/O mirror pane spawns eagerly in its hidden bootstrap window, which `uiWindow` deliberately excludes, and the main window's portal adopts the pane host only once AppKit and SwiftUI get run-loop time; `waitForLiveSurface` returns immediately for an already-live surface, so the check ran before adoption and the four key-delivery tests failed at line 176. The failure was hidden until the strict app-host gate. Order the harness window front and pump the run loop until the surface and its native view are in that window, as the other hosted-view input suites do, and assert against the harness window instead of any window. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(browser): report the refused connection through the navigation delegate `browserPanelRetriesDiscardedRestoreAfterConnectionRefused` waited up to 30s for WebKit to fail a provisional load to a bound-but-unlistened loopback port. On the hosted app-host runners that failure never arrives: the load neither fails nor commits, so the test timed out at line 147 (no "provisional navigation failed" log line appears for it in any shard). Stop the in-flight load and report `NSURLErrorCannotConnectToHost` for the attempted URL through the panel's real navigation delegate, the pattern `BrowserFailedNavigationReloadTests` already uses. The restore bookkeeping, error page, `restore_pending` clearing, and retry policy under test run unchanged and every assertion is kept. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cloud): scope reserved-workspace cleanup to its pending card and return the bind window PR #13202 (5d616a7f44) imported `CloudMachineWorkspaceAdoptionTests` and `CloudMachineWorkspaceResolutionTests` from the still-open 13141-cloud-vm-workspace branch without the production changes they assert, and its own run skipped the app-host shards, so they landed red: - `NewMachineSheetPresenter.closeReservedWorkspace` closed the whole workspace, which is a no-op for the last tab and discards user panes added next to a creating card. Port the branch behavior: remove only unadopted loading cards owned by the cancelled machine, clear that binding, keep user content, and give a last-tab loading workspace a local anchor first. `MachineCreateCoordinator` passes the created or reconciling machine id so cancelling machine X cannot clear a binding to machine Y (the tests now assert that scoping). - `v2WorkspaceCloudVMBind` now returns `window_id` alongside the workspace refs, as the bind acknowledgement test expects. - The resolution test selected "first" while the fixture's terminal key is "term-first"; use the key so the placement resolves as intended. - `CloudPaneCreationRetryTests` asserted `discarded` synchronously after the projection returned, but the coordinator applies its generation fence behind `CloudOperationContext.withPhase`'s recorder await, which suspends on a cold per-suite process. Settle with bounded yielding before asserting. Runtime behavior change (cancelled Cloud create cleanup); needs dogfood. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): restore deferred socket connection, explicit SSH control options, and Codex stop idle observations Branch commit c8bfb580ba ("fix: repair failures exposed by strict app-host CI", 2026-09-10) made `cmux vm dev|layout|env` finish local validation and dry runs before opening the socket, passed the caller's explicit ssh options into `userConfiguredControlOptions(fromSSHConfigOutput:explicitOptions:)` so a normalized `ControlPersist=0` from `ssh -G` is not mistaken for host customization, and published `idleObserved` for transcript-terminal prior turns before a legacy Codex Stop so the journal drops an obsolete running turn. The branch's later single-parent commit 38b32bb191 reverted `CLI/cmux.swift` to main's version while keeping the stricter fixtures, and PR #12053 landed that inconsistent state; the fixtures (CLIVMDevTests, CLIVMLayoutEnvTests, the SSH sharing tests, and the Codex missed-prompt Stop test) have failed since. Restore the three CLI changes. `SocketClient.configureAuthentication` and `SocketPasswordResolver` already exist on main, and the CmuxFoundation overload landed with the branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cli): align CLI integration fixtures with the shipped hook, SSH, and session-list contracts The main copy of `CLINotifyProcessIntegrationRegressionTests` predates several shipped contracts that the lenient app-host gate never enforced and that 38b32bb191 reverted on the issue-2824 branch: Claude hook acks print `{}` (#7963), Codex resume bindings require rollout evidence so fixtures carry a `session_meta` transcript (#10100), SessionStart publishes a binding so `/clear` counts start after the clear, fresh-terminal SSH startup commands are script paths that the support decoder must read, cmux control-path options follow a resolved `ssh -G` (#8308), and `ssh session list` reports the localized "remote state unavailable" summary with `--json` detail (#9971). The missed-prompt Codex Stop test now asserts the restored `agent.idle.observed` for the prior turn. Flagged for review: the three `ssh pty-attach` cases now expect only `workspace.remote.pty_bridge` once the endpoint is established, matching #12726's `preserveLifecycleForRecovery`; if pre-READY bridge failures were meant to keep reconciling, the flag should be set only after READY instead. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(workspace): restore the app-host repairs for fork, focus recovery, and shortcut routing Merge 8d33410585 on the issue-2824 branch took main's copy of `cmuxTests/WorkspaceUnitTests.swift` wholesale and discarded the branch's September repairs (c8bfb580ba, c402b9c515, 2127d97afd); PR #12053 then landed without them while the strict gate started counting the failures. Restore them against current production, plus the shortcut-suite fixes: - Fork in a remote workspace: `sshBootstrapArguments` has used `/usr/bin/ssh` since #9114, and agent socket propagation requires a socket that exists on disk (3bf87e3b4d), so bind a real unix socket and expect the absolute path. - Fork Conversation context actions dispatch asynchronously (#7259, #8173); await the fork panel before asserting. - Git branch and pull request updates publish through `sidebarObservationPublisher` since #6226, not `objectWillChange`. - Config sanitization: `addWorkspaceIfActive` rebuilds templates from the font-size lineage since #8543; override the lineage hook instead. - Focus recovery: AppKit focus is authorized only for a registered, selected workspace whose window carries the main-window identity; use the `TerminalPortalTestWorkspace` fixture and the same registration/pump path as `WorkspaceTerminalFocusRecoverySwiftTests`. - Shortcut routing: clear both Cloud defaults after 9c2ba78be4 swapped them; neutralize Ghostty's imported goto_split fallback (⌘] by ANSI keyCode) in the unshifted-symbol test and assert the digit shortcut does not match; wait for the async runtime start before judging keyDown forwarding (skip loudly without a live surface); wait for the portal to mount before the second-Escape check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(restore): keep a restore identity when the persisted surface id collides Since #13098 (e0e77eb74f), `newTerminalSurfaceOutcome` treats a nil `restoredSurfaceId` as an interactive create and routes it to the selected pane's Cloud source. Session restore passed nil whenever the persisted panel id was still live (duplicate-workspace or restore-into-live), so a legacy managed-Cloud SSH workspace restored into a live manager was turned into a remote tab create: the scaffold panel stayed startup-suppressed with no initial command and `TabManagerSessionSnapshotTests` failed to unwrap it. Mint a fresh UUID on collision instead of nil; it is free by construction, so the restore keeps its identity and the old-to-new remap works as before. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(terminal): align snapshot, projection, terminal, and browser fixtures with shipped contracts Stale expectations and fixed-spin timing in app-host suites that the lenient gate never counted: - Cloud-projected panes restore as manual-mirror reservations with a staged remote identity (#12675), not a local placeholder projection. - Restore reuses the persisted runtime id when free (aff0e32e93), so assert liveness rather than a new id; the catalog ignores writes for a Cloud machine without a registered provider (#11877), so register the fixture provider before publishing. - Wheel sync requires an authoritative scrollbar response (bbc3edfae4); the fixture now answers like `AuthoritativeScrollbarSurfaceView`. - Search overlay mount, first-responder focus, runtime creation, and the visibility-restore redraw run through deferred main-actor tasks; wait for them with the class's `waitUntil` helpers instead of fixed run-loop spins. - Owning the socket path lock is definitive (959f38a4c3): a refused inode left by a dead listener is replaced, so the restarted listener accepts. - The split-divider hit band extends `dividerHitExpansion` past the divider (667cc431d9); derive the pass-through boundary from the constant. - Browser page background blends against Ghostty's effective terminal color scheme, which is host dependent; read the same preference the product uses. - Cloud Machines defaults on in dev builds (#12318); pin the toggle off for the default-mode palette contract. - Prepared navigation requests keep the caller's cache policy (#13003). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(browser): align lifecycle, identity, host-view, loopback bridge, and portal rebind tests with shipped contracts Browser app-host suites that the lenient gate never counted: - `BrowserPanelWebViewLifecycleTests`: out-of-range discard delays are rejected to the default (the cmux.json loader relies on the nil), and the panel's own `isLoading` stays true for the indicator floor, so wait for both flags and assert no discard blockers before discarding. - `browserNavigationUsesEmbeddedWebKitIdentity`: WebKit reports the native identity as nil or "" (#9482); accept either. - `WindowBrowserHostViewTests`: production routes Dock-divider hits by yielding to AppKit so the live sidebar tracker receives them (#10902, e2e381825f); the tests asserting the portal owns and forwards the hit were merged red against a design that never shipped. Realign the stale-frame test to the pass-through contract and remove the four own-and-forward tests with their fixtures. **Flagged for review:** if own-and-forward is still wanted, that is a hit-testing product change for its own PR. - `testRemoteWorkspaceRuntimeBridgeAliasesMultipleLoopbackPortsFromSamePage`: the navigation delegate restarts main-frame loads to apply the user-agent policy (11c6efee86); a direct `loadHTMLString` with an HTTP base skipped that step, so the data load was cancelled and replayed as a deferred request. Apply the identity first and assert the document is current. - `portalRebindPreservesDocumentAndRoutesRefreshToTheSameWebView`: the portal host is a theme-frame sibling of `contentView` for non-glass windows (#12929); assert same-window instead of descendant-of-contentView. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(terminal): wait for asynchronous runtime creation in the remaining direct-interaction tests The same "Expected runtime surface before ..." precondition that 9f258dd7c7 made wait for the deferred runtime start still sampled after a fixed run-loop spin in the detach-race, close-lifecycle, repeat-key, and repeat-IME tests, and CI on the branch head showed them failing that way. Use the class's `waitUntil` for those four sites too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cli): model surface.respawn for respawn-pane and give the Codex stack fixture rollout evidence `cmux respawn-pane` has sent `surface.respawn` instead of `surface.send_text` since #5465 (8cafcc3bac); the window-flag fixture's mock still rejected that method, so the CLI exited 1. Answer `surface.respawn`, asserting the window and surface ids, `tmux_start_command`, and that the shell-invoked command carries the user command but never the `--window` flag, which is the test's intent. The Codex interrupted-stack fixture's transcript had no `session_meta` line, so `CodexSessionResumeVerifier` (#10100) found no rollout evidence and the prompt published `surface.resume.clear`. Prepend the session line as the other Codex fixtures do. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(socket): keep mobile.panel.artifact.fetch off the local socket and advertise the served artifact reads 02c1ba4dd9 removed `mobile.panel.artifact.fetch` from the socket worker methods because it needs the authenticated mobile execution context, but that half of the change was lost in a merge: the policy still routed fetch to the worker lane, which has no handler for it, so the local socket answered `internal_error` instead of the `method_not_found` boundary that `TerminalControllerSocketSecurityTests` pins. Restore the removal. `system.capabilities` never advertised `mobile.panel.artifact.stat` and `.thumbnail` although both are served on the worker lane (04ff18eea6 added them only to the test's expectation); advertise those two and keep fetch out. The remote relay allowlist is unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(remote): align tmux seed transport, SSH, socket command, and port scanner fixtures with shipped behavior - `RemoteTmuxPaneSeedTransportTests`: manual-I/O mirror panes spawn eagerly (#9272) and stay runtime-backed after portal churn (#9769), so the pane renders its assigned grid before a seed arrives and nothing was retained. Publish a larger tmux pane grid than any test surface applies so the retention under test sees the lag it exists for. - `SSHRemoteCWDRegressionTests`: the persistent-PTY exec helper runs only with `protectsFromHangup: true` (cd9f34ff1b); pass it, and widen the first-spawn guard. - `SSHDeepSleepReattachTests`: a Cloud-owned workspace rejects launch overrides on splits (#13098); create the custom-identity pane before configuring the remote connection. - `SSHConfiguredRemoteCommandHostTests`: ssh-pty-attach validates the bridge `daemon_version` before dialing (#12726); the mock now reports one. - `CloudManualMirrorTransportTests`: the pane failure card uses the short title since 9bf6cb8c94. - `SurfaceSocketCommandTests`: `vm.workspace_new` admits its optimistic workspace through the active main window (#13152, #13155), so bind a bare window to the fixture context; a receipt without a starter terminal costs one snapshot (6d43ea699a). - `PortScannerPublicationTests`: the forced-result acknowledgement hops off the main actor, so an unchanged port set may be deduped under a later refresh; drain publications until the retirement lands. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test: align mobile artifact fetch lane assertion * repair: close remaining full-suite contracts * test(restore,sidebar): align two stale contracts with shipped behavior testRemoteWorkspaceAutoResumeKeepsRemoteStartupCommand asserted the restored remote panel's local spawn cwd equals the remote-host path. It never can: OneShotTerminalLauncherStore.enterableWorkingDirectory rejects a path that is not locally enterable, which is what keeps the owning shell valid (#7031). The remote cwd does survive restore — as the panel's trusted remote directory report, and as the `cd` prefix the resume input carries (both already asserted). Assert it where it actually lives and pin the spawn cwd to nil. testSidebarPullRequestsTrackFocusedPanelOnly expected a background panel's PR to be hidden from sidebarPullRequestsInDisplayOrder(). That list is documented as the workspace's deduplicated rows in pane/tab order, both consumers (taskStatusSignals, the control-sidebar snapshot) want every panel, and the sibling branch test asserts the same all-panel model. Focus scoping lives in the `pullRequest` binding, which the test already covers. Renamed to match. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: give Cloud catalog tests live destination workspaces #13196 made SurfaceCatalog.validateOwnership refuse a destination that is not a live workspace. That rule stays. Tests that projected, restored or opened browsers into a made-up workspace ID now failed with destinationNotFound, timed out waiting on a provider that was never called, or passed a later check for the wrong reason. LiveWorkspaceFixture registers real Workspace objects and hands the catalog a CloudWorkspaceRenameService that resolves them, the way the app's composition root does. Its workspaces() list stays empty so the native projection coordinator does not start mirroring into them. For tests on SurfaceCatalog.shared, withAppRegistration registers the TabManager as a windowless main-window context for the test body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: normalize pbxproj after live workspace fixture Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: admit agent renames of agent-owned accepted Cloud names e6926fbbda (#13403) made submitCloudPanelRename run admitsTerminalRename for automatic names. Its last clause only admitted replacing an accepted name when the local panel still carried `.auto` provenance, but since 1e1d319dd1 an accepted daemon name reconciles locally as `.remote` and the owner lives in the tab's nameAuthority. Every agent title after the first accepted one was refused, so "Failed agent rename keeps the accepted title", "Mirroring an agent-named placement does not block its next agent title" and "An older automatic result and old snapshots cannot replace an accepted name" failed on their second agentName call. Admit an automatic rename when no write is pending and the accepted tab name is owned by the daemon's `auto` authority. User-owned names, pending writes and local user titles on any projection still refuse it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: expect adopted machine rows to keep the pending create identity #12919 (021f792b95) made New Machine creation optimistic: once a running create's machine appears in the fleet list or catalog, its row keeps the `pending-machine:<operation>` node ID so selection and expansion survive adoption (see committedProjectionKeepsThePendingNodeIdentityUntilFleetAdoptsIt). pendingRowStepsAsideOnceItsMachineHasARow still expected `machine:<id>`. Assert the new identity and that the row is the adopted machine, not a stand-in: the stand-in is gone, one row remains, and it shows the created machine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: restore per-display noVNC targets and refresh stale Cloud tests Product (#13196 regression): the 6db759357f main merge into 13192-cloud-display-ownership took main's CmuxTuiSurfaceProviders.swift and CloudPortRoutePlan.swift, dropping eb3edd7f98's per-display ports. 23c807b50a restored the display coordinator but not these hunks, so withPrivateBrowserURL rewrote every display to 6901 and a daemon pointer without a discovered target fell back to display 1. Restore both hunks and drop the duplicate port-less privateDesktopURL overload. Tests: - CloudDisplayCatalogTests: 178d35e5da (#13196) made every guest command run `list` as a readiness probe, so fakes dispatching on " list" answered creation with the list catalog. Dispatch on the create action line, and pin the command shape. - CloudPortOpenRegressionTests: 178d35e5da (#13196) filters RFB/noVNC ports only when the display catalog owns them (displayPortsOwned). Assert both the desktop and non-desktop results. - CloudTreeOneMachineManyWorkspacesTests: #12740 added a final Resources section under each machine. Expected trees include it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: publish every guest display from daemon graph updates The 6db759357f main merge into 13192-cloud-display-ownership (#13196) put back main's [desktopDisplayResource()] pools in CmuxTuiSurfaceProvider's refresh, publish and delta paths. 23c807b50a restored the display coordinator lifecycle but not these pools, so a display created beyond display:1 vanished on the next daemon publish, and a delta that touched it removed it. Restore eb3edd7f98's displayResources pools, the display-kind delta check, and the injectable displayCoordinator. Drop the now-unused desktopDisplayResource(). Adds a test that creates display:2 through the provider and asserts that both displays and their ports survive a full publish and a display delta. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: fence the fake Cloud projection reply with its mutation cursor aDetachedTerminalDropsItsStaleTabBeforeMoving installs a daemon graph since c402b9c515 (#13403). When the placement lane drains it reconciles against that graph, which predates the projected tab, and clears tab_projected. The real reply always carries a mutation cursor (CmuxTuiSnapshotParser.placedTab requires one) that fences exactly this; the fake returned none. Give the fake a projectCursor and set it ahead of the installed graph. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: register the restored window before relinking Cloud projections #13196 made SurfaceCatalog.validateOwnership refuse a destination the app cannot resolve, so SurfaceCatalog.shared no longer relinked a restored projection into a TabManager that no main window owns. Register it for the test body with LiveWorkspaceFixture.withAppRegistration, as #13651 does for CloudClosedPanelRestoreTests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: wait for the relay ports kick, not the first relay line Since #8442, a relay prompt also reports shell state. Both RPCs run in separate background children, and the zsh prompt-refresh test waited only until the log was non-empty, so it could read report_shell_state alone. Wait (deadline-bounded) for the ports_kick line itself, in the zsh test and its bash sibling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(cli): relay output of an SSH session that ends right after auth The expect wrapper watches the first two seconds after sending the password for a rejection with log_user 0. A session that authenticates and exits inside that window hit the eof branch and its output was dropped. Flush the buffered output before exiting. #13207 replaced the test's 9 s sleep with a FIFO release, which is what exposed this. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: restore the no-connection path for rejected vm dev input 4120e35654 taught runVMDev that invalid input never connects: keep the listener open, then check its backlog after the CLI exits. #13207 dropped that again, so each rejected run waited 60 s for a mock-server expectation that only the listener closing (after the wait) fulfills. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: drain the terminal while the SCP host-key failure runs The pty case read the master only after the CLI exited. A pty's output queue holds about 1 KiB and the host-key failure report is longer, so the CLI blocked writing stderr until the 30 s timeout. Read the master on a thread while the CLI runs. The test starts its own sshd; no runner host dependency is involved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: fail fast when a gated Cloud call ends before its fake is entered Every app-host restart in the 09-21/09-22 shard logs I sampled was a Swift Testing time-limit hit in one of two suites, and each hit relaunches the test host: - SurfaceCatalogTests: when catalog.project threw before reaching the provider (destinationNotFound, fixed by 3dc91b2321), each gated test parked in MaterializeGate.waitUntilEntered() until the 300 s limit. Five tests restarted the host one after another, about 25 minutes per shard. - CloudDisplayCatalogTests: when the fake no longer recognized the create command (fixed by 70eaf44819), create() failed before the exec started and `await started.result` parked until the 60 s limit. The waits now also end when the caller's task finishes, so the next such setup failure is an ordinary failed #require. The cancellation test also waits on its own cancellation signal instead of a 60 s Task.sleep. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: reveal a retired terminal through the portal rebind, as the app does Since #12607 hiding a terminal removes its hosted view from the window, and only a bind reinstalls it. The test flipped portal visibility on the detached view, so no size commit could ever land and the final shrink check failed every run. Rebind like TerminalPortalReconciliation and drop the wait loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: pin key-window status in terminal focus suites The app-host test process runs headless and is usually not the active app, so makeKeyAndOrderFront never makes a programmatic window key. Terminal focus paths gate on isKeyWindow (automatic first-responder apply, focus redraws, deferred focus reapply, ensureFocus window activation), so these suites passed only when an earlier test in the shard had activated the app. Share the existing KeyStatusTestWindow and use it in WorkspaceTerminalFocusRecoveryTests, WorkspaceTerminalFocusRecoverySwiftTests and TerminalNotificationDirectInteractionTests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: report which layer holds off-plan tmux mirror geometry offPlanGeometryWithUnchangedSizingInputsReconverges fails every run with all three output-parity re-arms spent and the hosted view still at the perturbed 0.8 divider. A standalone bonsplit replay of the same sequence converges, and a sibling test without a bound (portal-visible) workspace heals the same displacement. Add the live split view's arranged widths and the split model's imposed extent to the failure message so the next run shows whether bonsplit refused the apply, the imposition was cleared, or the portal did not follow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: restore Cloud projections before the workspace is published TabManager.restoreSessionSnapshot restores each workspace's surface projections before it assigns tabs, so SurfaceCatalog.restore's ownership check could not find the destination and silently dropped every restored remote projection. Check ownership against the workspace being restored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: repair stale and broken app-host expectations - CLI ssh attach: count identity UUIDs exactly; #11497 added an auth token uuidgen - device mirror directories: give the fake device trusted presence (e3d424c722 gate) - font zoom mirrors: deltas from the 8pt inherited base (e6926fbbda arithmetic) - Computer Use refresh: fake daemon reply was invalid JSON in a raw string (#13599) - quit alert: compare button alignment rects, not padded frames - remote split cwd rescue: cwd travels as CMUX_REMOTE_INITIAL_CWD since #12054 - default freestyle split: Cloud-owned splits route to Cloud since fa5dc4cc10 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: activate the app host before waiting for terminal focus `focusTerminalForTesting` waits for `window.isKeyWindow` before it hands first responder to the terminal, but `makeKeyAndOrderFront` only makes a programmatic window key while the test host is the active app. The app-host process starts inactive under `xcodebuild test`, so the wait succeeded only when an earlier test in the shard happened to activate the app. Its callers use a real `createMainWindow()` window, which cannot be swapped for `KeyStatusTestWindow` the way the pinned focus suites were. Activate explicitly, and give the wait a CI-appropriate timeout: the activation and the key-window transition land on later main run-loop turns, and the pump's one-second default expires before the window goes key on a contended runner. Observed on run 35743588302: `plainTerminalTextDoesNotResolveAppShortcutContext` (shard 1/6), `keyboardCopyModeKeyClearsTerminalUnread` and `workspaceFontSizeShortcutPreservesBackgroundTerminalUnread` (shard 2/6) all fail at this helper's return value, and shard 6/6 shows the same process state directly as `NSApp.keyWindow -> nil`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: give standalone terminal fixtures a live portal authority `setVisibleInUI` and `setActive` fold their request through `Workspace.portalRenderingEnabled(for:)`, which denies any workspace id the app delegate cannot resolve to a selected tab. Three direct interaction tests build their surface with `tabId: UUID()`, so once any earlier test installs an `AppDelegate.shared` the authority denies the portal, the hosted view is never actually made visible or active, and the fixture stops exercising the behavior it asserts: the surface never takes Ghostty focus and never schedules a visibility-restore redraw. Register a real selected workspace and build the surface with its id, so the fixture gets the same authority the app grants the selected tab. This is the fixture, not the product: the authority check is deliberate, and denying an unresolvable workspace is what keeps queued portal callbacks from reviving an inactive workspace. Fixes on run 35743588302 shard 4/6: `testKeyDownRecoveryDoesNotReplayFocusAfterResponderMovesAway`, `testVisibilityRestoreRefreshesSurfaceWhileTerminalIsInactive`, and `testDirectFirstResponderFocusRefreshesCursorStateAfterForeignResponder`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: normalize project.pbxproj after merging main Run 35763999808 failed `Fast static checks` before any macOS job could start, so the app-host test fixes on this branch were never exercised: error: cmux.xcodeproj/project.pbxproj is not normalized. Run scripts/normalize-pbxproj.py to fix. The branch head alone checks clean. CI builds the merge of this branch into main, and that merge is what leaves the file unnormalized, so the failure does not reproduce without merging main first. The change is one build-phase entry moving back into alphabetical order. `LiveWorkspaceFixture.swift in Sources` keeps the same UUID and the same occurrence count, and `scripts/lint-pbxproj-test-wiring.sh` still reports ok across 1049 test files, so no target lost a source file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * perf(tests): stop app-host unit tests waiting out real timeouts The app-host suite spends most of its wall clock in a handful of tests that wait out production timeouts, churn oversized fixtures, or run benchmarks. Each one now drives the same behaviour through an injected timeout, clock, or completion signal, with a separate cheap assertion pinning the default value to the production/server default. Production changes are configuration seams and one real fix: - KeyboardShortcutSettings.resetAll only removes keys that are stored, so a reset no longer fans out one UserDefaults.didChangeNotification per action. - SocketStartupWaiter / AgentRestorePreflightTimeout / BrowserDownloadWaitTimeout own the CLI wait windows (default plus a narrowing environment override), so client and handler cannot drift apart. - PortScanner and BrowserScreenshotWebViewSnapshotter take their schedules as parameters instead of hardcoding them. Benchmarks move out of the app-host unit suite behind the existing gating pattern, keeping their correctness assertions in the unit suite. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep the scroll settle default readable from a nonisolated context The default lives on a @MainActor enum, so every default-argument use of it is evaluated in a nonisolated context and trips the Swift 6 isolation warning. The value is an immutable Sendable constant, so mark it nonisolated rather than isolating four call sites. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(tests): stop kicking the port scanner from the poll loop Review found two defects in the slow-test pass. 1. `PortScannerPortRetirementTests` polled every 25ms and called `scanner.kick()` on every poll, against the 10ms compressed coalesce delay this PR introduced. `kick()` calls `startCoalesce()` whenever no burst is running, and `startCoalesce()` cancels and re-arms the timer. On a loaded runner the timer's jitter is the same order as the coalesce window, so the kick burst can cancel the timer indefinitely: no scan ever runs and the test burns its 20s deadline. The previous 500ms poll against a 200ms delay left a 2.5x margin that the compressed schedule removed. Fixed by removing the kick from the poll loop rather than by widening the interval, so the speed win stays and the flake vector is gone instead of made less likely. A kick already guarantees `minimumScansPerKick` scans - exactly the complete misses the reconciler needs to retire a port - so one kick after `stopListening()` is sufficient. That is the shape `lateBurstKickRetiresStoppedListener` already used. `onKick` is gone from `waitForPublication` entirely, so the poll interval is no longer coupled to the coalesce delay and the footgun cannot come back. 2. `scripts/test-command-palette-nucleo-ffi.sh` set `CMUX_COMMAND_PALETTE_SEARCH_BENCHMARKS=1` and then passed `-only-testing:cmuxTests/CommandPaletteNucleoFFITests`, a different class from the four `CommandPaletteSearchEngineTests` benchmarks this PR gated behind that variable, so those four ran nowhere. The script now names both classes and asserts one BENCH line per gated benchmark, since a skipped benchmark otherwise passes silently. Nothing referenced the script either, so even the FFI tests it names were not run by CI. `.github/workflows/command-palette-search-benchmarks.yml` is its scheduled caller, modelled on the `tmux-corpus.yml` nightly (same checkout/GhosttyKit/zig/rust/SPM-cache setup, same `cmux-unit` scheme). The script takes an optional `CMUX_NUCLEO_FFI_SOURCE_PACKAGES` so the workflow can reuse the cached package clone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: keep the benchmark caller dispatch-only while automatic CI is paused tmux-corpus.yml and perf-activation.yml -- the two closest macOS benchmark nightlies -- both carry "Temporarily manual-only beginning 2026-07-13 to pause automatic CI" and have their crons removed. Adding a live cron here would quietly reverse that reduction, which is the opposite of what this branch is for. The gate still has a caller, and the script's BENCH assertions still fail the job if a gated benchmark silently skips; running it is now a deliberate act rather than a daily cost. Drop "nightly" from the name, since it is not one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(tests): model the title-word ranking term in the search reference testBenchmarkCorporaMatchReferencePipelineOnSmallFixture asserts the optimized engine equals a reference pipeline rebuilt in the test. On the query "workspace 31" the engine scored workspace.large.31 at 17598 and the reference at 16000, so the suite failed. The engine ranks on three terms; the reference modelled two. The missing one is commandPaletteTitleWordScore: a query that is, or prefixes, a title's search words scores the tokens' upper bounds plus a per-token title bonus. For "workspace 31" that is 6799 + 6799 + 2000*2 = 17598, exactly the observed gap. The reference now models that term too, reimplemented rather than calling the engine's copy, because a reference that shared the implementation would assert nothing about it. FixtureEntry prepares the two title texts the term needs once at construction, so the benchmark timing loops — which build fixtures outside the timed region — do not start charging the legacy side for preparation the engine does not repeat either. Verified by compiling the engine sources against this reference on Linux and comparing all 34 corpus/query pairs the parity tests use: the large benchmark corpus goes from 1 mismatch to 0, and the command, switcher and switcher-benchmark corpora stay at 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(tests): drop the ineffective SSH auth retry budget rewrite persistentAttachExitsAtForegroundAuthenticationFailureLimit() rewrote the generated startup script to shrink the 20-failure foreground-authentication budget to 3, then asserted 3 attempts. CI recorded 20: the rewrite never changed the behaviour it was meant to bound. The budget the test edited lives in the shell wrapper, but the preserved __ssh-* CLI helper is what owns authentication retries and counts them itself, so rewriting the wrapper's literal leaves the helper executing its own 20. The shrink was decoration over a real 20-attempt run. The test now keeps the production budget and asserts 20. The fake sleep already removes the backoff between attempts, which is where the wall-clock cost actually was, so the same failing run reached the assertion in 4.25 s against the 6.1 s this test cost before the PR. That is less than the <1 s the PR's table claims for this row; the table is corrected in the PR body. replacingWithinScript, scriptDecodeLevels and applyingReplacements existed only for this rewrite and have no other callers, so they go with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(tests): close the persistent CLI client before awaiting the mock testDefaultFreestyleSSHAttachHidesPersistentRetryLimitInCountdown failed with "Asynchronous wait failed: Exceeded timeout of 5 seconds, with unfulfilled expectations: cli mock socket handled", after 5.227 s. startMockServer fulfills once the first connection is done. This test drives a persistent SSH attach, so the client holds its connection open across the retry countdown and the mock has no "done" to observe. The wait could only ever end by timing out; it had been passing on the longer window this PR narrowed. The test now terminates the child once it has observed the progress it actually asserts — the "Retrying in 0.1s (attempt 2)." banner — and waits for the mock afterwards. The countdown assertions that follow are unchanged and still read the captured output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Retrigger CI after retargeting this PR onto main The synchronize event that pushed ed3ab7a61e fired while the base was fix/app-host-green, a branch the #13643 squash merge had already deleted, so GitHub could not build the merge ref and the pull_request CI workflow never started. Only the two pull_request_target workflows ran, which left the required ci-status check absent rather than red. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Make the two new timeout holders instantiable values CI's iOS conventions guard reported two violations this branch introduced: AgentRestorePreflightTimeout and BrowserDownloadWaitTimeout were both caseless enums with an all-static public surface, which the lint refuses because such a type can never be substituted. The refusal is right, and for this branch in particular: the point of the change is that a test should be able to hold a timeout agreement without spending it. A static holder cannot be narrowed, so the tests could only assert the shipped numbers. The restore preflight budget moves onto AgentRestorePreflightInvocation, which is the type it bounds and already carries the environment it is read from. Names lose the now-redundant prefix: defaultSeconds becomes defaultTimeoutSeconds, environmentKey becomes timeoutEnvironmentKey, and seconds(environment:) becomes timeoutSeconds(environment:). BrowserDownloadWaitTimeout becomes a struct whose three windows are stored properties with the shipped values as init defaults, plus a `standard` value the handler and the CLI client both read. Both call sites take `.standard`; the behaviour is unchanged. The regression test now also constructs a narrow window and asserts the client still outwaits the handler, which makes the invariant a property of the pair rather than an assertion about 10 seconds. Verified with swiftc 6.1.3 on Linux: both files type-check under -swift-version 6, and the extracted behaviour matches the previous constants exactly (10s default, 10s ceiling on the override, whitespace trimmed, non-finite and non-positive overrides ignored; 10000ms default window, 120000ms cap, 15.0s client response timeout). The conventions diff lint now reports no new violations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Remove the KeyStatusTestWindow the main merge declared twice `macOS compile admission` failed with "ambiguous use of init(contentRect:styleMask:backing:defer:)" at WorkspaceTerminalFocusRecoveryTests.swift:551 and WorkspaceUnitTests.swift:4890. Neither file is touched by this branch. The cause is cmuxTests/AppDelegateMainWindowTestingSupport.swift, where the merge of main kept both sides' addition of the same class: `final class KeyStatusTestWindow` appeared twice, byte-identical including its doc comment, so every construction of it was ambiguous. Git had no conflict to report — the two additions did not overlap — which is the third time on this branch that a clean automerge produced code that could not compile. Checked the rest of the merge for the same shape: no other duplicated top-level type in cmuxTests, and none in CLI/ or Sources/. The one other repeated name, `optimizedResults` in CommandPaletteNucleoFixtures.swift, is a legitimate pair of overloads taking `corpus:` and `entries:`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: expect the reconnect budget #13959 honors, not the old clamp The new budget probe asserted that "21" and "021" clamp to 20. Main's #13959 made a well-formed budget above 20 honored up to SSHReconnectBudget's ceiling, so the merge ran green through git and red on CI: stdout "21". The cases now carry their expected value, from SSHReconnectBudget where it is a policy constant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: gate the benchmark runner behind the paid-overflow switch Main's #13994 guard rejects a vars.MACOS_RUNNER_15 read without CI_PAID_MACOS_OVERFLOW, since it can hold a metered WarpBuild label. This branch's new benchmark workflow predates the guard, so the merge of main passed git and failed workflow-guard-tests. Same form as perf-activation.yml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: close CodeRabbit gaps in the fast app-host test rewrites - PortScanner late-burst: the runner stops listening and issues the late kick from inside the fifth lsof call, instead of the test task racing a one-second timer gap it could miss in either direction. - Sidebar quiescence: the quiet window is measured in time and spans twice the sidebar's 50ms coalesce stage, so a pending coalesced or debounced row invalidation cannot land after the drain declared quiet. - SSH exit prompt: waitUntilBlocked samples the launched process and all of its descendants, so it holds whether the shell execs the startup script in place or forks it; terminate() retires that tree so a forked, parked helper cannot hold the output pipes or leak into the test host. - Sidebar git refresh: wait until the panel is a poll candidate again before each refresh; a completed read can precede the probe clearing. - Diff fixtures pin SHA-1 objects and loose-file refs, which the handwritten refs and in-process blob ids assume. - Rename the registry test to the post-list enrollment order and assert it. - TerminalController uses BrowserDownloadWaitTimeout's shared clamp. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep the Cloud carrier prewarm on activation A merge from the fix/app-host-green lineage deleted the prewarm that syncPollingToActivationPolicy() starts (999693e015, perf: prewarm cloud carrier before first machine), so enrollment waited for the first listPage() round trip. The polling test was then rewritten to assert that slower order. Neither change is part of this PR, and #13403 carries the same removal for its owner to argue. Both files now match main, whose test asserts the prewarm starts before fleet discovery finishes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: keep the git fixture's formats when the host sets Git defaults GIT_DEFAULT_REF_FORMAT and GIT_DEFAULT_HASH override the init options the handwritten-ref fixtures depend on. runGitProcess no longer passes them on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: report browser.download.wait's requested timeout as a number again The shared timeout window left requested_timeout_ms an Optional, which the timeout error payload coerced to Any and which put a new warning over the Swift warning budget. The handler resolves the default and the lower bound before clamping, as it did before, so the field is the same Int as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: settle the Global Search palette before each routing test KeyboardShortcutSettings.resetAll() no longer posts one UserDefaults.didChangeNotification per action, so the next test's init no longer spends enough main-thread time for the previous test's animated NSPopover close to finish. Two routing tests then saw isShown == true from the palette an earlier test had opened (run 35967763918, shard 7). Dismiss the palette and pump the run loop until it closes, bounded at 2 s, in both suites' init, matching the helper the sibling suites already use. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: route the palette benchmark workflow to a hosted runner on forks main's fork runner routing guard (test_ci_fork_runner_routing.py) now rejects a runs-on expression that can reach a Blacksmith label outside manaflow-ai. Lead with the owner fork branch, as perf-activation.yml does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: drop the extra paren in the palette benchmark runs-on expression 572dc2a3aa7 kept the old closing paren and added another, so actionlint could not parse the expression. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: set up Bun on the shard that runs agent notification semantics The notification semantics step resolves `bun` with `command -v bun`, but Bun was only set up on the CLI regression shard (4) while the step runs on the focused regression shard (6). Main never reaches the step because its unit batch fails first, so the missing Bun only shows once shard 6's unit tests pass, as they do on this PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: scope the CmuxWebView keyDown hook to each reentry test CmuxWebViewKeyDownReentryTests swizzled CmuxWebView.keyDown(with:) for the rest of the process. When CmuxWebViewWebContentUndoTests ran after it in the same app host, browserCmdZPerformsWebContentUndoWhenPageDeclinesTheChord recursed in the test hook until the stack overflowed (seen on macOS 15 and 26 once this PR's shard layout put the two suites back to back). Install the hook per test window, call the captured original implementation directly instead of re-dispatching through a swapped selector, and restore the original implementation when the window closes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: keep the OpenCode feed harness socket inside sun_path testOpenCodeFeedPluginEmitsCompletionForBothIdleEventShapes put its Unix socket under FileManager.temporaryDirectory plus a UUID-named root. On a Blacksmith runner that is /private/var/folders/.../T/, and the path runs past the 104-byte sun_path limit, so the Bun harness fails in listen() and prints nothing. Use a short /tmp path for the socket instead. This surfaced once the agent notification semantics step ran on this PR; main never reaches the step because its shard 6 unit batch fails first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: wait for the relayed pwd line, not the first relay line zshRelayPromptReportsRemotePWD broke out of its wait as soon as the relay log was non-empty. _cmux_precmd sends report_shell_state and report_pwd as separate background relay calls, so on a busy runner the log held only the shell-state line when the test read it. Wait for the report_pwd line itself, up to 5 s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: tie the E2E stale-snapshot case to OWNED_MAX_AGE_MINUTES #14341 raised OWNED_MAX_AGE_MINUTES from 20 to 45, so a 30 minute old snapshot is now fresh enough and the "snapshot too old" case in test_run_e2e takes the owned Mac, failing app-host-execution guards. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 3dfc85f17f2b54936fb044ee718387b2f4258903) * test: pin the full paste worker to lossy rich text, per #14121 #14121 sends mixed rich and plain text through the plain-text helper, so richPasteDoesNotUsePlainTextHelper (HTML plus a clean plain string) now gets a successful fast-path paste instead of the full worker's error, and it has failed on every main run since. The helper still declines rich text whose plain export lost characters (U+FFFD or repeated '?'), because the full worker can recover them. Assert that case instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: austinpower1258 <austinwang115@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Cloud display views now retain stable owning VM identity through creation, duplication, browser and Dock persistence, delayed restoration, moves, and session restore. The shared
SurfaceOwnershipPolicyrejects foreign, unknown, and stale provenance before placement or transport mutation, with a final ownership check after async work. Existing same-VM multiple views remain supported.Addresses #13192.
The screenshot does not establish a duration or transport failure. This change traces route readiness, navigation start/commit/finish/cancel, noVNC connection state, reconnect, and a bounded retryable deadline. Progress and errors are display-specific; no URLs or credentials are logged.
A scoped New Display action now discovers and creates independent guest display resources. Each resource has a stable VM/display identity, separate X/RFB/noVNC ports and session state, a persisted request receipt, and its own route. Discovery and creation are capability-gated through the existing VM-authorized exec path; unchanged images remain truthfully unavailable until the runtime is present. The existing Desktop resource and prior views are preserved.
The guest supervisor now isolates and recovers per-display desktop components, noVNC transport, and process ownership across helper restarts. URL-only navigation cannot invent undiscovered display slots, and restored display queries cannot override the authenticated host or port.
Validation on the merged branch (latest
origin/mainmerged and pushed):$autoreviewis clean with no actionable findings. PR Fix Cloud display ownership and connection readiness #13196 is open and mergeable; no merge or issue close was performed.Fleet build and launch:
e2fb4e3465efac18b4f670d1693205f9edc99f4fissue-13192-cloud-display-ownershipfb18a0d65f1eb6747f49b053cmux7s-Mac-mini.localartifacts/fleet/e2fb4e3465efac18b4f670d1693205f9edc99f4f-terminal.json40e22393a8e442c3145f3321ae86916825e589f9e09810aa87901d836ca7851eThe verified app is running from the tagged bundle
cmux DEV issue-13192-cloud-display-ownership. A supported isolated Cloud GUI recipe was unavailable, so no two-VM visual rejection video, original first-open timing capture, or two-real-guest-screen rendering proof is claimed. No production image, live user VM, billing state, or deployment was changed.Latest main merge and review-comment closeout
Merged the current
origin/mainat931792787fand pushed the resolved branch at812b4ef07d.The remaining non-outdated review threads were verified against the current code, replied to, and resolved. The closeout includes global Dock ownership/pending restores, Cloud service-port rebinding, external/cancelled navigation ownership, hidden restore activation, display-slot discovery, guest supervisor recovery/readiness, bounded discovery retries, duplicate URL preservation, sanitized New Display errors, and display-port filtering.
Latest validation passes:
git diff --check.The PR remains open and unmerged. GUI Cloud runtime proof remains unavailable because the controller has no supported isolated GUI recipe.
Latest merge/conflict and comment closeout
Merged the current
origin/mainatc29c4da25cand resolved the Cloud browser/route, SurfaceCatalog materialization, and Cloud display UI conflicts. The pushed head is42f1ab4a076daefbb396ae3501fd66fa34688ab4and GitHub reports the PR mergeable.All currently non-outdated review threads are resolved. The latest fixes include global Dock pending restore policy, display-port preservation filtering, route/duplicate URL preservation, bounded guest startup/discovery readiness, hidden restore activation, sanitized display errors, and final ownership validation after materialization.
Final focused validation after the merge: 14 guest catalog/supervisor tests, 14 desktop contract tests, Swift parsing, PBX normalization, test wiring, localization parity, Swift file-length budget, and diff checks pass. No merge or issue close was performed.
Latest CI repair: after current
origin/mainadvanced, the merged base exposed duplicateSurfaceCatalogdeclarations and a provider display-coordinator initialization conflict. Those were resolved in the pushed head23c807b50ab406e508de46c07b5e22f7df7c1ad9; the PR is currently mergeable and CI is rerunning. The latest focused guards pass locally.CI status for current head
23c807b50ab406e508de46c07b5e22f7df7c1ad9: static/workflow/Linux checks are green; macOS compile admission is still running in CI run 35560913107. All current non-outdated review threads are resolved.Merge of origin/main at dcdeab3
Merged the current
origin/main(dcdeab336e) into the branch. The first merge (11396c3382, pushed as210df4c84d) had conflicts inSources/Cloud/CloudTreeRowContentView.swift,Sources/Surfaces/SurfaceCatalog.swift,Sources/Workspace.swift, andcmux.xcodeproj/project.pbxproj. Because main movedCloudTreeRowHoverButtonsinto its own file, that resolution silently dropped this PR's displays-pool New Display hover button;e18669f557restores it inSources/Cloud/CloudTreeRowHoverButtons.swifttogether with itshasButtonsentry. The follow-up merge of the newer main (7856fa7fa0) applied cleanly, and1eda271279keepsSurfaceCatalog.swiftwithin its Swift file-length budget.Checks on the pushed head
1eda271279:check-pbxproj.sh,lint-pbxproj-test-wiring.sh,normalize-pbxproj.py --check,localization_catalog.py check(9 locales, 0 parity errors),swift_file_length_budget.py,check-package-resolved-policy.py,check-workspace-package-groups.py,git diff --check, andswiftc -frontend -parseon every Swift file the merge touched all pass.origin/mainis an ancestor of the branch. The fleet controller was unreachable from the submitting network at merge time, so the fleet build of this head is still pending; hosted CI is the current compile evidence.Merge of origin/main at 0ddc6f4
Merged the current
origin/main(0ddc6f406c, 89 commits including the relay core RPC v2 and production-app-hang work) into the branch as2bc5ca3aa8. One conflict, inSources/Workspace.swift: this PR'sacceptsRestoredSessionownership guard at the top ofrestoreSessionSnapshotcollided with main's newbeginTerminalGeometryTransition(.restore)bracket. Resolved by keeping both, with the ownership guard first so a rejected snapshot never opens a geometry transition. The Cloud tree row, localization catalog, and project file auto-merged with the earlier New Display hover-button restoration intact.Checks on
2bc5ca3aa8:check-pbxproj.sh,normalize-pbxproj.py --check,lint-pbxproj-test-wiring.sh,localization_catalog.py check,swift_file_length_budget.py,check-package-resolved-policy.py,check-workspace-package-groups.py,git diff --check, andswiftc -frontend -parseon all 94 app Swift files the merge touched pass.origin/mainis an ancestor of the branch.Merge of origin/main at 4be6fb7
Merged the current
origin/main(4be6fb7865, 185 commits) into the branch as21a652e8ba. No conflicts. Checks on the pushed head:check-pbxproj.sh,normalize-pbxproj.py --check,lint-pbxproj-test-wiring.sh,localization_catalog.py check,swift_file_length_budget.py,check-package-resolved-policy.py,check-workspace-package-groups.py, andswiftc -frontend -parseon the 34 app Swift files the merge touched all pass.origin/mainis an ancestor of the branch.