Skip to content

Coalesce and retry SSH PTY resize delivery (#6306) - #6320

Closed
austinywang wants to merge 9 commits into
mainfrom
issue-6306
Closed

austinywang wants to merge 9 commits into
mainfrom
issue-6306

Conversation

@austinywang

@austinywang austinywang commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

Closes #6306

Problem

In long-lived SSH workspaces, terminal TUI panes could enter a stale resize state where changing pane width no longer reliably caused the remote TUI to reconcile to the visible size. Output kept flowing so the terminal looked alive, but the remote PTY/TUI stayed stuck at the old geometry — only a manual Reconnect Workspace restored self-healing resize.

Root cause

The SSH PTY attach CLI forwards SIGWINCH changes via workspace.remote.pty_resize, but the send was best-effort and fire-and-forget:

_ = try? client.sendV2(method: "workspace.remote.pty_resize", params: params)

A send that raced a stale or blocked remote-session control path was silently dropped. The resize never reached the remote PTY/TUI and was never retried, so the size desynced until the session/controller path was rebuilt by a manual reconnect.

Fix

Add SSHPTYResizeCoordinator (CLI/CMUXCLI+SSHCommandSupport.swift) and route the SIGWINCH source through it (CLI/cmux.swift):

  • Coalesce rapid SIGWINCH bursts (e.g. during a divider drag) into a single delivery of the newest size.
  • Retry the latest size with bounded exponential backoff when a delivery fails, instead of dropping it.
  • Supersede in-flight retries when a newer size arrives, and dedup redundant deliveries of an already-delivered size.
  • Bounded retries (no spin); a fresh SIGWINCH after giving up still recovers.

All coordinator state is touched only on the signal source's serial queue, so no extra locking is needed; the existing socketLock still guards the shared control socket. The coordinator exposes injected send/scheduleAfter seams for deterministic testing.

Verification

  • ./scripts/reload.sh --tag issue-6306 — clean Debug build of the app + CLI.
  • Regression test testSSHPTYAttachRetriesResizeAfterDeliveryFailure (in CLINotifyProcessIntegrationRegressionTests): spawns the real ssh-pty-attach CLI against a mock control socket where every workspace.remote.pty_resize fails; asserts that a single SIGWINCH still produces a retried resize RPC with no further signals. The pre-fix best-effort send would drop the failed resize and the retry would never arrive.

Notes / tradeoffs

  • The coordinator lives in the cmux_cli module, which the cmuxTests target (@testable import cmux) cannot import, so coverage uses the repo's standard subprocess-based CLI test harness rather than an in-process unit test.
  • No user-facing strings changed (only debug-log lines), so no localization updates were required.

🤖 Generated with Claude Code


View with Codesmith Autofix with Codesmith
Need help on this PR? Tag /codesmith with what you need. Autofix is disabled.


Summary by cubic

Coalesces and retries SSH PTY resize events so remote TUIs stay in sync with the visible terminal, and ties retries to the attach lifecycle to avoid sends after teardown. Fixes #6306 by replacing best‑effort sends with coordinated, protocol‑aware backoff retries and a safe, synchronous cancel.

  • Bug Fixes

    • Added SSHPTYResizeCoordinator to coalesce rapid SIGWINCH into the latest size, retry with bounded exponential backoff, supersede/dedupe in‑flight or redundant deliveries, and run work on a serial queue; still uses socketLock.
    • Lifecycle‑safe retries: the coordinator owns the SIGWINCH source and exposes a synchronous, lock‑guarded cancel(), so pending retries don’t run after bridge EOF.
    • Only retry when the control socket is in sync: issue resize via raw client.send(command:), retry protocol rejections (ok:false / SSHPTYResizeProtocolRejection), and stop on transport failures (SSHPTYResizeTransportError); a fresh SIGWINCH later recovers.
    • Routed ssh-pty-attach resize through the coordinator and added a subprocess regression test (SSHPTYResizeRetryIntegrationTests.swift) that forces workspace.remote.pty_resize failures and asserts a single retry without extra signals.
  • Refactors

    • Moved SSHPTYResizeCoordinator into its own file (CLI/SSHPTYResizeCoordinator.swift); no behavior change.
    • Trimmed a comment in CLI/cmux.swift to meet the file‑length budget.

Written for commit a929814. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes

    • Improved SSH terminal resize event handling with intelligent coalescing of rapid resize requests and automatic retry logic for failed deliveries, ensuring more reliable terminal synchronization when connecting via SSH.
  • Tests

    • Added integration tests to verify SSH terminal resize retry behavior and resilience.

austinywang and others added 2 commits June 17, 2026 12:11
Previously workspace.remote.pty_resize was sent best-effort with try?, so a
send that raced a stale/blocked remote-session control path was silently
dropped. Output kept flowing so the terminal looked alive while the remote
PTY/TUI never received the new size, requiring a manual workspace reconnect
to recover (#6306).

Add SSHPTYResizeCoordinator which coalesces rapid SIGWINCH bursts into one
delivery of the newest size and retries the latest size with bounded
exponential backoff when delivery fails, instead of dropping it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Subprocess integration test driving the real ssh-pty-attach CLI against a
mock control socket: every workspace.remote.pty_resize fails, and a single
SIGWINCH must still produce a retried resize RPC with no further signals.

The coordinator lives in the cmux_cli module, which the cmuxTests target
(@testable import cmux) cannot import, so coverage uses the repo's standard
subprocess-based CLI test harness rather than a unit test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
cmux Ready Ready Preview, Comment Jun 17, 2026 10:07pm
cmux-staging Building Building Preview, Comment Jun 17, 2026 10:07pm

@coderabbitai

coderabbitai Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds SSHPTYResizeCoordinator, a new class that coalesces rapid SIGWINCH events and retries failed workspace.remote.pty_resize sends with bounded exponential backoff. The existing inline SIGWINCH handler in cmux.swift is replaced to delegate to this coordinator. An integration regression test verifies retry behavior when the resize RPC returns failure.

Changes

SSH PTY Resize Coalescing and Retry

Layer / File(s) Summary
SSHPTYResizeCoordinator class, state, and initializers
CLI/CMUXCLI+SSHCommandSupport.swift
Introduces SSHPTYResizeCoordinator with a Scheduler typealias, all stored state (pendingSize, lastDeliveredSize, retry/cancel flags, injected closures), a designated initializer with injectable send/scheduleAfter seams, and a convenience initializer wiring to SocketClient and DispatchQueue.
noteResize(), cancel(), and scheduleDelivery()
CLI/CMUXCLI+SSHCommandSupport.swift
Implements noteResize() to update pendingSize, validate non-zero dimensions, reset retry budget, and schedule coalesced delivery; cancel() to set isCancelled and clear pendingSize; scheduleDelivery(after:) to guard against cancellation and duplicate scheduling.
deliver() state machine with exponential backoff
CLI/CMUXCLI+SSHCommandSupport.swift
Implements deliver() which clears the scheduled flag, short-circuits on no-op resizes, attempts the send closure, reschedules when a newer size arrives mid-flight, and retries failed sends with bounded exponential backoff up to maxRetries before resetting.
SIGWINCH handler wired to coordinator
CLI/cmux.swift
Replaces the inline SIGWINCH handler (direct lock → size read → try? sendV2) with a dedicated resize DispatchQueue, a shared baseParams dict, and an SSHPTYResizeCoordinator instance whose noteResize()/cancel() are wired to the dispatch source event and cancel handlers.
Integration regression test for retry-on-failure
cmuxTests/SSHPTYResizeRetryIntegrationTests.swift, cmux.xcodeproj/project.pbxproj
Adds testSSHPTYAttachRetriesResizeAfterDeliveryFailure which configures a mock bridge that always fails workspace.remote.pty_resize, launches ssh-pty-attach, drives SIGWINCH until the first RPC arrives, asserts a retry RPC follows without new signals, and verifies at least two total resize RPCs. The four pbxproj entries wire the new test file into the cmuxTests target.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • manaflow-ai/cmux#4323: Both PRs modify the PTY resize delivery path — this PR adds client-side coalescing/retry for workspace.remote.pty_resize on SIGWINCH, while that PR refactors the daemon's PTY/websocket hub to normalize and propagate resize across attachments.
  • manaflow-ai/cmux#4562: This PR's addition of a new Swift integration test into cmux.xcodeproj is directly covered by the pbxproj test-wiring lint/CI guard introduced in that PR.

Poem

🐇 Hop hop, SIGWINCH rings the bell,
No more silent drops — we retry so well!
A coalescing dance, a backoff queue,
The PTY stays fresh the whole day through.
Two RPCs proved it — the rabbit's math is true! ✨

🚥 Pre-merge checks | ✅ 20 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (20 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely summarizes the main change: adding coalescing and retry logic for SSH PTY resize delivery, directly addressing the core problem described in issue #6306.
Linked Issues check ✅ Passed The PR implements all key coding requirements from #6306: coalesces rapid SIGWINCH bursts, retries failed resize with bounded backoff, deduplicates redundant deliveries, and includes a regression test confirming retry behavior on delivery failures.
Out of Scope Changes check ✅ Passed All code changes directly support the stated objective of fixing SSH PTY resize resilience. The SSHPTYResizeCoordinator, routing through cmux.swift, and integration test are all within scope of addressing #6306.
Cmux Swift Actor Isolation ✅ Passed SSHPTYResizeCoordinator uses queue-based serialization (serial dispatch queue) for all mutable state access, avoiding MainActor/Sendable violations. Closures capture self safely within the coordina...
Cmux Swift Blocking Runtime ✅ Passed The PR introduces SSHPTYResizeCoordinator with non-blocking asyncAfter scheduling for retry backoff, no new semaphores/sleeps/polling in production code, and test-only DispatchSemaphore scaffolding.
Cmux Expensive Synchronous Load ✅ Passed Code changes run costly operations (client.sendV2 network I/O) on dedicated background DispatchQueue via asyncAfter scheduling, not on main actor/interactive paths; ioctl syscall is lightweight; ru...
Cmux Cache Substitution Correctness ✅ Passed This PR adds SSH PTY resize retry logic via SSHPTYResizeCoordinator using transient in-memory deduplication (pendingSize, lastDeliveredSize), not caching fresh reads in persistence/history/undo/sna...
Cmux No Hacky Sleeps ✅ Passed Check does not apply: PR contains only Swift files (CLI/.swift, cmuxTests/.swift, .pbxproj); check scope is non-Swift runtime (TS/JS/shell/build). Swift sleeps covered by separate check.
Cmux Algorithmic Complexity ✅ Passed SSHPTYResizeCoordinator uses only O(1) fixed-size state (tuples, ints, bools) with no collection iteration. Backoff calculation is O(1) bit shift. Coordinator runs on dedicated serial queue, instan...
Cmux Swift Concurrency ✅ Passed The PR introduces a DispatchQueue for DispatchSourceSignal (OS boundary) and test-only DispatchQueue.global/Semaphore usage, both allowed per swift-concurrency-modernization.md. No completion-handl...
Cmux Swift @Concurrent ✅ Passed No @concurrent annotation violations found. The PR adds only synchronous code: SSHPTYResizeCoordinator uses a serial DispatchQueue for safe concurrency without async/await, so @concurrent is not ap...
Cmux Swift File And Package Boundaries ✅ Passed PR satisfies swift-file-package-boundaries.md: new file CLI/CMUXCLI+SSHCommandSupport.swift is 184 lines with clear single responsibility (SSH PTY resize coordination); CLI/cmux.swift received only...
Cmux Swift Logging ✅ Passed SSHPTYResizeCoordinator uses injected log closure (default no-op), no direct print/NSLog/debugPrint/dump; integration routes through DEBUG-guarded cliDebugLog; test file appropriate.
Cmux User-Facing Error Privacy ✅ Passed All user-facing string checks passed. The two log statements in SSHPTYResizeCoordinator are debug-only, written via cliDebugLog which is #if DEBUG guarded and only writes to debug log files, not us...
Cmux Full Internationalization ✅ Passed All strings added are either literal protocol tokens (workspace.remote.pty_resize, cols, rows), debug-only logs wrapped in #if DEBUG via cliDebugLog(), or test code—all exempt from i18n requirement...
Cmux Swiftui State Layout ✅ Passed No SwiftUI changes detected. PR modifies CLI command-line code (SSHPTYResizeCoordinator dispatch coordination and SIGWINCH handling), not SwiftUI views or state management.
Cmux Architecture Rethink ✅ Passed PR introduces SSHPTYResizeCoordinator as a proper architectural fix, not a symptom patch. Uses serial queue for state serialization, bounded retries for resilience, and reuses existing socketLock—n...
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed PR adds SSH PTY resize coordinator and regression test with no NSWindow, NSPanel, NSWindowController, SwiftUI Window, or WindowGroup code changes. Not applicable to auxiliary window close-shortcut...
Cmux Source Artifacts ✅ Passed All four changed files are legitimate hand-written source code, test files, and project configuration. No local tool output, generated logs, screenshots, temp folders, caches, build output, or othe...
Description check ✅ Passed The PR description comprehensively covers the problem, root cause, solution, verification, and trade-offs with clear technical details.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-6306

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@greptile-apps

greptile-apps Bot commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR replaces the best-effort, fire-and-forget workspace.remote.pty_resize send with SSHPTYResizeCoordinator, which coalesces rapid SIGWINCH bursts into a single delivery of the newest size and retries with bounded exponential backoff when the remote rejects the resize — fixing the stale-geometry bug from #6306.

  • SSHPTYResizeCoordinator (new file, 257 lines): serial-queue state machine with pendingSize/lastDeliveredSize coalescing, up to 6 retries at 50–800 ms backoff, and a stateLock-guarded cancel() that can safely stop in-flight retries from the bridge-EOF teardown thread.
  • cmux.swift: startSSHPTYResizeSource now returns the coordinator instead of a raw DispatchSourceSignal; both EOF paths and the function defer call resizeCoordinator.cancel().
  • SSHPTYResizeRetryIntegrationTests.swift: subprocess regression test against a mock socket that rejects every resize RPC, asserting at least two RPCs arrive (one original send + one retry) after a single SIGWINCH.

Confidence Score: 5/5

Safe to merge. The coordinator's state machine is well-bounded, cancel() correctly quiesces pending retries before teardown, and the transport/protocol error distinction prevents socket desync on retry.

The change replaces a single best-effort try? send with a clearly-scoped coordinator. All mutable coordinator state is confined to a serial queue; the only cross-thread field (cancelled) is properly guarded by stateLock. The retry loop is bounded (6 attempts), transport failures stop retries rather than desyncing the socket, and the regression test exercises the exact failure path the fix targets. The one style concern (explicit lock/unlock instead of defer) has no correctness impact on current code.

No files require special attention beyond the previously-noted asyncAfter scheduling concern in SSHPTYResizeCoordinator.swift.

Important Files Changed

Filename Overview
CLI/SSHPTYResizeCoordinator.swift New 257-line coordinator implementing coalesce + bounded-backoff retry for SIGWINCH delivery. State machine is correct; cancelled is properly guarded by stateLock for cross-thread teardown safety. Minor: explicit lock/unlock in the send closure instead of defer; and asyncAfter (flagged in a previous thread) lacks a cancellation handle.
CLI/cmux.swift Routing change: startSSHPTYResizeSource now returns SSHPTYResizeCoordinator instead of a raw DispatchSourceSignal, and resizeCoordinator.cancel() is called on both EOF paths and in the defer. No logic changes to the surrounding bridge read/write loop.
cmuxTests/SSHPTYResizeRetryIntegrationTests.swift Subprocess-based regression test matching the repo's established CLI test harness. Drives a mock control socket that rejects every pty_resize and asserts at least two resize RPCs arrive, proving the retry path. Semaphore-based synchronization is acceptable test-only scaffolding.
cmux.xcodeproj/project.pbxproj Adds SSHPTYResizeCoordinator.swift to the app CLI target and SSHPTYResizeRetryIntegrationTests.swift to the test target. Standard project file additions with no anomalies.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    SIGWINCH([SIGWINCH signal]) --> NR[noteResize\nupdate pendingSize\nreset retryCount]
    NR --> SD{deliveryScheduled?}
    SD -- No --> SCHED[scheduleDelivery\nafter coalesceDelay 20ms]
    SD -- Yes --> PEND[pendingSize updated,\nno new schedule\ncoalesced into in-flight delivery]

    SCHED --> DELIVER[deliver]
    DELIVER --> CHK{isCancelled or\nno pendingSize?}
    CHK -- Yes --> BAIL[return / no-op]
    CHK -- No --> DEDUP{lastDeliveredSize\n== pendingSize?}
    DEDUP -- Yes --> CLEAR[pendingSize = nil\nreturn dedup]
    DEDUP -- No --> SEND[send RPC\nworkspace.remote.pty_resize]

    SEND --> OK{result?}
    OK -- ok:true --> SUCCESS[lastDeliveredSize = size\nschedule newer pending if any]
    OK -- ok:false --> RETRY{retryCount < maxRetries?}
    OK -- transport error --> TFERR[keep pendingSize\nstop retrying\nawait fresh SIGWINCH]

    RETRY -- Yes --> BACKOFF[retryCount++\nbackoff 50-800ms\nscheduleDelivery]
    BACKOFF --> DELIVER
    RETRY -- No --> GIVEUP[log give-up\nreset retryCount\nawait next SIGWINCH]

    CANCEL([cancel\nbridge EOF / defer]) --> CSET[stateLock: cancelled=true\ncancel signalSource]
    CSET --> BAILCHECK[not-yet-started deliver\nblocks bail at isCancelled check]
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
    SIGWINCH([SIGWINCH signal]) --> NR[noteResize\nupdate pendingSize\nreset retryCount]
    NR --> SD{deliveryScheduled?}
    SD -- No --> SCHED[scheduleDelivery\nafter coalesceDelay 20ms]
    SD -- Yes --> PEND[pendingSize updated,\nno new schedule\ncoalesced into in-flight delivery]

    SCHED --> DELIVER[deliver]
    DELIVER --> CHK{isCancelled or\nno pendingSize?}
    CHK -- Yes --> BAIL[return / no-op]
    CHK -- No --> DEDUP{lastDeliveredSize\n== pendingSize?}
    DEDUP -- Yes --> CLEAR[pendingSize = nil\nreturn dedup]
    DEDUP -- No --> SEND[send RPC\nworkspace.remote.pty_resize]

    SEND --> OK{result?}
    OK -- ok:true --> SUCCESS[lastDeliveredSize = size\nschedule newer pending if any]
    OK -- ok:false --> RETRY{retryCount < maxRetries?}
    OK -- transport error --> TFERR[keep pendingSize\nstop retrying\nawait fresh SIGWINCH]

    RETRY -- Yes --> BACKOFF[retryCount++\nbackoff 50-800ms\nscheduleDelivery]
    BACKOFF --> DELIVER
    RETRY -- No --> GIVEUP[log give-up\nreset retryCount\nawait next SIGWINCH]

    CANCEL([cancel\nbridge EOF / defer]) --> CSET[stateLock: cancelled=true\ncancel signalSource]
    CSET --> BAILCHECK[not-yet-started deliver\nblocks bail at isCancelled check]
Loading

Reviews (4): Last reviewed commit: "Merge remote-tracking branch 'origin/mai..." | Re-trigger Greptile

Comment thread CLI/CMUXCLI+SSHCommandSupport.swift Outdated
Comment on lines +107 to +109
scheduleAfter: { delay, block in
queue.asyncAfter(deadline: .now() + delay, execute: block)
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 asyncAfter for retry backoff in a socket/terminal path

The production scheduleAfter closure drops directly into queue.asyncAfter, which is explicitly flagged by the repo's blocking-runtime rule for retry backoff in socket and terminal paths. The rule says retry backoff must use "a real cancellation-aware scheduler, timer abstraction, async sequence, callback, notification, or state transition." The cancelled flag and [weak self] guard against stale fires but asyncAfter itself has no cancellation handle — once scheduled, the closure cannot be recalled, only no-oped. A DispatchSourceTimer configured on queue with setEventHandler / cancel() would give the coordinator an actual cancellation-owning handle and is the idiomatic replacement here.

Rule Used: Flag new blocking or timing-based synchronization ... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment thread CLI/CMUXCLI+SSHCommandSupport.swift Outdated
Comment on lines +175 to +182
if retryCount < maxRetries {
retryCount += 1
let backoffMs = min(1000, 50 * (1 << min(retryCount - 1, 4)))
scheduleDelivery(after: .milliseconds(backoffMs))
} else {
log("ssh-pty resize giving up after \(maxRetries) attempts (\(size.cols)x\(size.rows)); awaiting next resize")
retryCount = 0
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Give-up log reports attempt count one less than actual deliveries made

When retryCount reaches maxRetries the give-up message prints "giving up after \(maxRetries) attempts", but by that point maxRetries + 1 delivery calls have been made (one original call at retryCount == 0 plus maxRetries retried calls). With the default maxRetries = 6 the log will say "6 attempts" while 7 deliveries were actually attempted. Since this is a debug-only log the impact is cosmetic, but the count will mislead anyone reading the log while investigating a stuck resize.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@CLI/CMUXCLI`+SSHCommandSupport.swift:
- Around line 116-124: The issue is that when a retry timer fires after a failed
resize delivery, and then a fresh user resize occurs (noteResize() call), the
old stale retry timer takes precedence over the new coalesced delivery.
Implement a generation token mechanism to fix this: add a generation counter
that increments each time noteResize() is called, then pass this generation
token when scheduling delivery in both the scheduleDelivery(after:
coalesceDelay) call in noteResize() and when scheduling retry backoff. When a
scheduled delivery fires, check that the generation token still matches the
current generation before proceeding with delivery; if it's stale, discard it so
that fresh resize events can properly supersede pending retry timers.
- Around line 62-64: The cached `lastDeliveredSize` is being used to suppress
fresh SIGWINCH delivery when the new size matches a previously delivered size,
but this causes stale remote PTY geometry to persist if the remote session is
rebuilt or reset without updating the local cache. Remove the comparison logic
that checks `lastDeliveredSize` against the current pending size to suppress
delivery (around line 146). Keep only the burst coalescing mechanism using
`deliveryScheduled` to combine rapid consecutive resize events, but ensure that
fresh SIGWINCH signals are always delivered to the remote PTY regardless of
whether they match the cached `lastDeliveredSize`, allowing the remote state to
be revalidated with each signal.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6347d0cb-aed2-4f50-868f-155ba4ca3ac0

📥 Commits

Reviewing files that changed from the base of the PR and between e26dca1 and 24efe76.

📒 Files selected for processing (3)
  • CLI/CMUXCLI+SSHCommandSupport.swift
  • CLI/cmux.swift
  • cmuxTests/CLINotifyProcessIntegrationRegressionTests.swift

Comment thread CLI/CMUXCLI+SSHCommandSupport.swift Outdated
Comment thread CLI/CMUXCLI+SSHCommandSupport.swift Outdated
Comment on lines +116 to +124
func noteResize() {
guard !cancelled else { return }
let size = sizeProvider()
guard size.cols > 0, size.rows > 0 else { return }
pendingSize = size
// A fresh, user-driven resize resets the retry budget so we keep trying
// to deliver the newest size even after a prior burst gave up.
retryCount = 0
scheduleDelivery(after: coalesceDelay)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Let fresh resizes replace a pending retry timer.

After a failure schedules backoff, Line 135 makes later noteResize() calls keep the old timer. A user resize during a 400–1000ms retry window updates pendingSize, but delivery is delayed until the stale retry fires instead of being coalesced at coalesceDelay.

Use a generation token so fresh SIGWINCH can supersede retry backoff
     private var pendingSize: (cols: Int, rows: Int)?
     private var lastDeliveredSize: (cols: Int, rows: Int)?
     private var deliveryScheduled = false
+    private var scheduledDeliveryToken = 0
+    private var scheduledDeliveryIsRetry = false
     private var retryCount = 0
     private var cancelled = false
@@
         pendingSize = size
         // A fresh, user-driven resize resets the retry budget so we keep trying
         // to deliver the newest size even after a prior burst gave up.
         retryCount = 0
-        scheduleDelivery(after: coalesceDelay)
+        scheduleDelivery(after: coalesceDelay, replacingExistingRetry: true)
@@
-    private func scheduleDelivery(after delay: DispatchTimeInterval) {
-        guard !cancelled, !deliveryScheduled else { return }
+    private func scheduleDelivery(
+        after delay: DispatchTimeInterval,
+        replacingExistingRetry: Bool = false,
+        isRetry: Bool = false
+    ) {
+        guard !cancelled else { return }
+        if deliveryScheduled {
+            guard replacingExistingRetry, scheduledDeliveryIsRetry else { return }
+        }
+        scheduledDeliveryToken += 1
+        let token = scheduledDeliveryToken
         deliveryScheduled = true
+        scheduledDeliveryIsRetry = isRetry
         scheduleAfter(delay) { [weak self] in
-            self?.deliver()
+            self?.deliver(ifCurrent: token)
         }
     }
 
-    private func deliver() {
+    private func deliver(ifCurrent token: Int) {
+        guard token == scheduledDeliveryToken else { return }
         deliveryScheduled = false
+        scheduledDeliveryIsRetry = false
         guard !cancelled, let size = pendingSize else { return }
@@
             retryCount += 1
             let backoffMs = min(1000, 50 * (1 << min(retryCount - 1, 4)))
-            scheduleDelivery(after: .milliseconds(backoffMs))
+            scheduleDelivery(after: .milliseconds(backoffMs), isRetry: true)

Also applies to: 134-139, 175-179

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@CLI/CMUXCLI`+SSHCommandSupport.swift around lines 116 - 124, The issue is
that when a retry timer fires after a failed resize delivery, and then a fresh
user resize occurs (noteResize() call), the old stale retry timer takes
precedence over the new coalesced delivery. Implement a generation token
mechanism to fix this: add a generation counter that increments each time
noteResize() is called, then pass this generation token when scheduling delivery
in both the scheduleDelivery(after: coalesceDelay) call in noteResize() and when
scheduling retry backoff. When a scheduled delivery fires, check that the
generation token still matches the current generation before proceeding with
delivery; if it's stale, discard it so that fresh resize events can properly
supersede pending retry timers.

austinywang and others added 2 commits June 17, 2026 13:11
The retry regression test pushed CLINotifyProcessIntegrationRegressionTests.swift
over the Swift file-length budget (workflow-guard-tests). Move it into a new
SSHPTYResizeRetryIntegrationTests.swift that extends the same test class, so it
still reuses the mock-socket / bundled-CLI harness helpers while keeping the
large file unchanged. Wired into the cmuxTests target.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@cmuxTests/SSHPTYResizeRetryIntegrationTests.swift`:
- Around line 148-165: The test's retry assertion can false-pass because the
loop sending up to 20 SIGWINCH signals may result in multiple resizes being
queued, and the second resizeReceived.wait could be satisfied by a late resize
from one of those earlier signals rather than from the retry-after-failure
mechanism. To tighten this test, replace the loop at line 148 with a single
SIGWINCH send (not multiple attempts), capture the resize count before
signaling, wait for the first resize after that single signal, then assert that
the count increases again without any additional SIGWINCH being sent, ensuring
the second resize definitively comes from the retry coordinator's retry logic
and not from buffered/delayed events from prior signals.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 887fb339-6d9a-4fd7-9e00-e424e072acad

📥 Commits

Reviewing files that changed from the base of the PR and between 24efe76 and a3639e0.

📒 Files selected for processing (2)
  • cmux.xcodeproj/project.pbxproj
  • cmuxTests/SSHPTYResizeRetryIntegrationTests.swift

Comment on lines +148 to +165
for _ in 0..<20 {
Darwin.kill(process.processIdentifier, SIGWINCH)
if resizeReceived.wait(timeout: .now() + 0.2) == .success {
sawFirstResize = true
break
}
}
XCTAssertTrue(sawFirstResize, "Expected ssh-pty-attach to issue an initial resize RPC after SIGWINCH")

// No further SIGWINCH is sent. A second resize attempt can only arrive
// from the coalesce/retry coordinator re-sending the latest size after
// the first delivery failed. The pre-fix best-effort send would never
// retry, so this would time out.
XCTAssertEqual(
resizeReceived.wait(timeout: .now() + 3),
.success,
"Expected ssh-pty-attach to retry the resize after a delivery failure"
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Retry assertion can false-pass because the test may send multiple user SIGWINCH events.

The loop at Line 148 can deliver more than one SIGWINCH before the first resize is observed; then the second resizeReceived.wait at Line 162 may be satisfied by a late resize from those extra signals, not by retry-after-failure. This weakens the regression guarantee.

A tighter pattern is: record resize count before signaling, send exactly one SIGWINCH (or gate to ensure only one post-install signal), wait for first resize, then assert count increases again without any additional signal.

Suggested direction
-        var sawFirstResize = false
-        for _ in 0..<20 {
-            Darwin.kill(process.processIdentifier, SIGWINCH)
-            if resizeReceived.wait(timeout: .now() + 0.2) == .success {
-                sawFirstResize = true
-                break
-            }
-        }
-        XCTAssertTrue(sawFirstResize, "Expected ssh-pty-attach to issue an initial resize RPC after SIGWINCH")
+        // Send one user signal, then prove subsequent resize comes from retry.
+        Darwin.kill(process.processIdentifier, SIGWINCH)
+        XCTAssertEqual(
+            resizeReceived.wait(timeout: .now() + 3),
+            .success,
+            "Expected initial resize RPC after SIGWINCH"
+        )

         XCTAssertEqual(
             resizeReceived.wait(timeout: .now() + 3),
             .success,
             "Expected ssh-pty-attach to retry the resize after a delivery failure"
         )
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cmuxTests/SSHPTYResizeRetryIntegrationTests.swift` around lines 148 - 165,
The test's retry assertion can false-pass because the loop sending up to 20
SIGWINCH signals may result in multiple resizes being queued, and the second
resizeReceived.wait could be satisfied by a late resize from one of those
earlier signals rather than from the retry-after-failure mechanism. To tighten
this test, replace the loop at line 148 with a single SIGWINCH send (not
multiple attempts), capture the resize count before signaling, wait for the
first resize after that single signal, then assert that the count increases
again without any additional SIGWINCH being sent, ensuring the second resize
definitively comes from the retry coordinator's retry logic and not from
buffered/delayed events from prior signals.

austinywang and others added 3 commits June 17, 2026 13:26
Autoreview flagged that scheduled retry timers could outlive the attach: the
DispatchSource cancel handler runs asynchronously, so a retry block already due
when the bridge hit EOF could still call deliver() with cancelled==false and
send workspace.remote.pty_resize on the shared socket after pty_attach_end (or
block EOF cleanup on socketLock).

Make SSHPTYResizeCoordinator.cancel() synchronous and thread-safe: guard the
cancelled flag with a lock, have the coordinator own the SIGWINCH source, and
cancel it directly. Teardown (bridge EOF / defer) now calls coordinator.cancel()
synchronously, so any not-yet-started retry observes cancellation and bails
before sending.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The lifecycle-fix comment pushed CLI/cmux.swift 3 lines over its tracked
budget (workflow-guard-tests). Condense it; the rationale already lives in the
SSHPTYResizeCoordinator doc comment.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Autoreview flagged that retrying every sendV2 failure on the shared control
socket is unsafe after a transport timeout: SocketClient.send leaves the fd open
on timeout and the v2 protocol does not match response ids, so a late resize
reply could be misattributed to a later request (pty_sessions / pty_attach_end),
desyncing the control protocol.

Issue the resize at the raw send(command:) layer so delivery outcome can be
classified: send only returns once a complete response line is consumed (socket
in sync) — an ok:false there is a protocol rejection and safe to retry (the
stale/blocked remote path from #6306). A thrown error (timeout / socket error)
means the socket may hold a pending late reply, so surface it as
SSHPTYResizeTransportError and stop retrying on that socket; a fresh SIGWINCH
recovers once it is healthy.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
austinywang and others added 2 commits June 17, 2026 14:08
Aziz file-organization policy: a major type should live in its own
TypeName.swift rather than appended to CMUXCLI+SSHCommandSupport.swift (which is
an extension CMUXCLI of SSH command-string helpers). Relocate the coordinator
and its two error types to CLI/SSHPTYResizeCoordinator.swift, wired into the
cmux-cli target. No behavior change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@lawrencecchen lawrencecchen added the stale-revisit Closed after 30+ days without activity; preserved for possible revisit or reopening. label Sep 23, 2026
@github-project-automation github-project-automation Bot moved this from Todo to Done in cmux backlog Sep 23, 2026

This branch was successfully deployed

1 active deployment
Preview – cmux — a929814e Deployed Jun 17, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stale-revisit Closed after 30+ days without activity; preserved for possible revisit or reopening.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SSH workspace TUI resize can remain stale until workspace reconnect

3 participants