Skip to content

fix(fork): make the fallback refresh await its own queued validation - #13960

Merged
teamleaderleo merged 2 commits into
mainfrom
fix/fork-probe-fallback-awaits-own-request
Sep 23, 2026
Merged

teamleaderleo merged 2 commits into
mainfrom
fix/fork-probe-fallback-awaits-own-request

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes the intermittent sharedForkProbeFallbackWaitsForActiveSamePanelValidation
failure, one of the app-host suite's remaining red tests — and the product
contract bug underneath it.

The bug

refreshForkAvailabilityNow's fallback-snapshot branch was the only exit
that returned without waiting for the validation it had just queued. The
live-index branch directly below it already waits via
waitForForkValidationRequestCompletions.

That gap is reachable because another task can drain the request first.
applyPendingForkValidations ends with an unguarded restart, whose sibling
inside the probe-owner defer is guarded by !resumedWaiters for exactly this
reason:

if activeForkSupportValidationKeys.isEmpty {
    restartForkAvailabilityRefreshIfPending()   // no !resumedWaiters guard
}

So it can spawn a detached refresh that races the contention waiter it just
resumed. Whichever continuation the main actor runs first decides the outcome:
when the detached task wins, it claims the pending request and starts the probe,
and the original caller resumes, finds an empty queue, and returns while its
own probe is still in flight
.

Why this is a product bug, not a test artifact

Workspace+ForkAgentConversationAvailability and ContentView both read fork
support immediately after awaiting this call. They can observe "refresh
finished" with availability still unvalidated, so the Fork Conversation menu
item renders off a snapshot nobody checked.

Why wait, rather than guard the restart

Waiting makes the contract hold regardless of which task drains the queue,
instead of trying to win a scheduler race. It is inert on the common path:
forkValidationRequestIsWaiting is false once the request is neither pending nor
processing, so the continuation resumes synchronously when this call did its own
work. When another drainer did claim it, the id stays in
processingForkValidationRequestIDs and the per-batch defer retires it and
resumes completion waiters — releasing this caller exactly when the probe's
validation is recorded. It also reuses the existing cancellation-correct helper
(ownsRequest: true drops the request on cancel), so no new cancellation
semantics.

The alternative — hoisting a resumedWaiters flag across batches to guard the
tail restart, mirroring the sibling — I rejected: it can strand pending requests
if the resumed waiter's task is cancelled right after resuming, and it does not
fix the contract for any other drainer.

Confidence, stated honestly

The root cause is a hypothesis, though tightly constrained: it is the only
mechanism found that makes the test's assertions at lines 1908/1910/1916/1922
pass while 1917 and 1924 fail together, which is the observed signature.

The failure is intermittent (3 passes, 1 failure across sampled runs), so a
single green run proves little here — this wants repeated runs. I have no macOS
SDK available, so this is swiftc -parse plus reading, not a build or a run.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Fixes the intermittent sharedForkProbeFallbackWaitsForActiveSamePanelValidation failure by making refreshForkAvailabilityNow await its own queued validation before returning, so callers never observe "refresh finished" with stale fork availability.

This is a symptom fix: the root cause is the unguarded tail restart in applyPendingForkValidations, which can spawn a detached refresh that races a contention waiter it just resumed, letting either branch of refreshForkAvailabilityNow return without waiting.

  • Adds a waitForForkValidationRequestCompletions call before the fallback branch returns.
  • The equivalent hole exists in the live-index branch whenever reload() runs applyPendingForkValidations internally, so the fix waits regardless of which task drains the queue.
  • Relates to Workspace+ForkAgentConversationAvailability and ContentView rendering the Fork Conversation menu item from unvalidated snapshots.

Written for commit 92c4f01. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes
    • Refresh operations now wait for pending fallback validation requests to complete before returning, ensuring validation results are ready when the operation finishes.

`refreshForkAvailabilityNow`'s fallback-snapshot branch was the only exit that
returned without waiting for the validation it had just queued. The live-index
branch directly below it already waits via
`waitForForkValidationRequestCompletions`; this one did not.

That matters because the request can be drained by someone else.
`applyPendingForkValidations` ends with an unguarded restart:

    if activeForkSupportValidationKeys.isEmpty {
        restartForkAvailabilityRefreshIfPending()
    }

Its sibling inside the probe-owner `defer` is guarded by `!resumedWaiters`
precisely so it does not restart when it has just resumed a waiter. The tail has
no such guard, so it can spawn a detached refresh that races the contention
waiter it just resumed. Whichever continuation the main actor runs first decides
the outcome: when the detached task wins it claims the pending request and starts
the probe, and the original caller resumes, finds an empty queue, and returns
while its own probe is still in flight.

Real callers observe this as "refresh finished" with availability still stale --
`Workspace+ForkAgentConversationAvailability` and `ContentView` both read fork
support straight after awaiting this call, so the Fork Conversation menu item can
render off a snapshot that was never validated. It also surfaces as an
intermittent failure in
`sharedForkProbeFallbackWaitsForActiveSamePanelValidation`, where the second
refresh reports finished and the second fallback is never accepted.

Waiting here fixes the contract regardless of which task drains the queue,
rather than trying to win the race. It is inert on the common path:
`forkValidationRequestIsWaiting` is false once the request is neither pending nor
processing, so the continuation resumes synchronously when this call did its own
work. When another drainer did claim it, the id stays in
`processingForkValidationRequestIDs` and the per-batch `defer` retires it and
resumes completion waiters, releasing this caller exactly when the probe's
validation is recorded.

I considered instead hoisting a `resumedWaiters` flag across batches to guard the
tail restart, mirroring the sibling. Rejected: it can strand pending requests if
the resumed waiter's task is cancelled right after resuming, and it would not fix
the contract for any other drainer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Warning

Review limit reached

Next included review available in 27 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: c5286a77-a4a6-4766-93a5-cc9745d962dc

📥 Commits

Reviewing files that changed from the base of the PR and between c96048c and 92c4f01.

📒 Files selected for processing (1)
  • Sources/SharedLiveAgentIndex.swift

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 5ab3d985-c75b-471a-9f5f-fd926920120a

📥 Commits

Reviewing files that changed from the base of the PR and between a9bdaa8 and c96048c.

📒 Files selected for processing (1)
  • Sources/SharedLiveAgentIndex.swift

Included review availability: Your plan provides up to 10 included reviews per hour; 1 remains after this review.


📝 Walkthrough

Walkthrough

The fallback-snapshot path in refreshForkAvailabilityNow now waits for its owned fork-validation requests to complete after applying pending validations.

Changes

Fork validation

Layer / File(s) Summary
Wait for fallback validation completion
Sources/SharedLiveAgentIndex.swift
After applying pending validations for a fallback snapshot, the method now waits for its owned validation requests to complete, including requests claimed by another drainer.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested reviewers: austinywang

Merge Risk: ⚪ Minimal · up to c9604

No actionable issue is established for this change; it is mergeable after normal checks.

🚥 Pre-merge checks | ✅ 23 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description gives a detailed summary, explains the bug, documents the intended behavior, and states the available validation. However, it does not follow the required template structure. The Testi… Add the required template sections. Include explicit testing details and manual verification, provide a demo video or explain why one is not applicable, include the review-trigger block, and complete the checklist items.
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 1 functions across 1 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (23 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main change: making the fallback refresh await its queued validation.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Cmux Cloud Persistent Session And Early Input ✅ Passed PASS — The pull request changes only Sources/SharedLiveAgentIndex.swift, adding a wait for fork-validation request completion in refreshForkAvailabilityNow. It does not change Cloud terminal creat…
Cmux Swift Actor Isolation ✅ Passed PASS. The PR adds only an awaited completion call and comments in SharedLiveAgentIndex.swift (+10/-0). SharedLiveAgentIndex is already @MainActor, so refreshForkAvailabilityNow and `waitForFor…
Cmux Swift Blocking Runtime ✅ Passed The diff only adds an await of the existing waitForForkValidationRequestCompletions helper in refreshForkAvailabilityNow. That helper waits on checked-continuation completion signals and has cance…
Cmux Browser Automation Off-Main ✅ Passed PASS: The pull request changes only Sources/SharedLiveAgentIndex.swift, adding a wait for fork-validation request completion. It does not change browser socket commands, processV2Command, `socketW…
Cmux Expensive Synchronous Load ✅ Passed The diff changes only Sources/SharedLiveAgentIndex.swift by adding an awaited call to the existing waitForForkValidationRequestCompletions helper and comments in the fallback validation path. It a…
Cmux Cache Substitution Correctness ✅ Passed PASS: The PR changes one Swift hunk and adds a completion wait after applyPendingForkValidations; it does not replace a fresh authoritative read with a cache. The existing fallback snapshot is still…
Cmux No Hacky Sleeps ✅ Passed PASS: The authoritative diff changes only Sources/SharedLiveAgentIndex.swift, a Swift source file. It adds an await of waitForForkValidationRequestCompletions; it does not add a fixed sleep, timer…
Cmux Algorithmic Complexity ✅ Passed The diff adds one call to the existing completion-wait helper. In this method, pendingRequestIDsOwnedByRequest is populated from a single supplied workspaceId/panelId request, so the helper proc…
Cmux Swift Concurrency ✅ Passed The diff adds only an await waitForForkValidationRequestCompletions(...) call and explanatory comments. The helper already existed, is async, and uses the existing cancellation-aware validation wait…
Cmux Swift @Concurrent ✅ Passed PASS. The reviewed range changes only Sources/SharedLiveAgentIndex.swift and adds an await of waitForForkValidationRequestCompletions in the fallback branch. SharedLiveAgentIndex is @MainActor…
Cmux Swift Package Boundaries ✅ Passed PASS: The diff changes only Sources/SharedLiveAgentIndex.swift and adds an await of the existing waitForForkValidationRequestCompletions helper in an existing @MainActor app singleton. It does n…
Cmux Swiftpm Lockfiles ✅ Passed PASS: The pull request changes only Sources/SharedLiveAgentIndex.swift. The patch contains no Package.swift, Package.resolved, .gitignore, workflow, or cmux.xcodeproj changes, so the SwiftPM…
Cmux Swift Logging ✅ Passed The pull request adds one await and an explanatory comment in Sources/SharedLiveAgentIndex.swift. The changed lines add no print, debugPrint, dump, NSLog, file/stdout logging, Logger, or s…
Cmux User-Facing Error Privacy ✅ Passed PASS. The PR changes only Sources/SharedLiveAgentIndex.swift: it adds an await call and a developer-only explanatory comment. It adds or changes no user-facing error, alert, command output, API erro…
Cmux Full Internationalization ✅ Passed The pull request changes only Sources/SharedLiveAgentIndex.swift. The added lines are one existing completion-wait call and a developer-only comment. The diff adds no user-facing Swift text, localiz…
Cmux Swiftui State Layout ✅ Passed PASS: The PR changes only Sources/SharedLiveAgentIndex.swift. It adds an await and comments in the @MainActor final class SharedLiveAgentIndex, which imports Darwin and Foundation, not SwiftUI…
Cmux Architecture Rethink ✅ Passed PASS. The diff adds one await in the fallback branch of refreshForkAvailabilityNow and reuses the existing cancellation-aware waitForForkValidationRequestCompletions helper. SharedLiveAgentIndex…
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed The PR changes only Sources/SharedLiveAgentIndex.swift, adding validation-wait logic in refreshForkAvailabilityNow. The diff adds no NSWindow, NSPanel, NSWindowController, SwiftUI Window/`…
Cmux Source Artifacts ✅ Passed The PR changes only Sources/SharedLiveAgentIndex.swift, a tracked hand-written Swift source file. The diff adds 10 source lines and no local output, generated artifact, cache, temporary directory, o…
Cmux No Test Or Debug Seam In Production Source ✅ Passed PASS — The authoritative PR diff changes only Sources/SharedLiveAgentIndex.swift and adds a production wait call plus explanatory comments. It adds no #if DEBUG block, test/debug-named member, vis…
Full details: Description check

Explanation

The description gives a detailed summary, explains the bug, documents the intended behavior, and states the available validation. However, it does not follow the required template structure. The Testing, Demo Video, Review Trigger, and Checklist sections are missing.

✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

An independent review found the comment's closing claim wrong. It said the
live-index branch "already waits this way; this branch was the only exit that
did not". That branch waits only when `didReload` is false. When `reload()` did
run, it called `applyPendingForkValidations` internally, so the identical steal
can happen and that path returns without waiting. The hole is not unique to the
fallback branch.

Also say plainly that this is a symptom fix. The root cause is the tail restart
in `applyPendingForkValidations` missing the `!resumedWaiters` guard its in-loop
sibling has, and that is tracked separately rather than being implied as already
handled here.

No behaviour change; comment only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Independent review: SAFE TO MERGE — one overclaim corrected in 92c4f011ed

The review cleared the two things that would actually make this dangerous, and
caught a wrong sentence in my own comment.

No deadlock, no lost wakeup. SharedLiveAgentIndex is @MainActor, and in
waitForForkValidationRequestCompletion the forkValidationRequestIsWaiting
check and the append to forkValidationRequestCompletionWaiters both run inside
the synchronous withCheckedContinuation body with no suspension between
them. So either the request is already retired and the continuation resumes
immediately, or it is still pending/processing and the waiter is registered
before any drainer can retire it. The already-drained case resumes immediately
because the per-batch defer retires before it resumes waiters.

The reviewer also walked the "pending but nobody will drain it" shapes, which
was my own main worry: the fallbackSnapshot == nil, index == nil requeue
can't apply to this caller (batches are grouped by an identity derived from the
fallback snapshot, so a non-nil snapshot always lands in a non-nil batch), the
contention branch re-drains on the way out, and every cancellation path either
removes the request and resumes its waiters or tombstones it while a drainer
still owns it.

Cancellation is correct — both the inline Task.isCancelled check and the
onCancel closure remove the waiter before resuming, so a double-fire finds
nothing. Identical to what the live-index branch already does.

The correction. My comment ended "The live-index branch below already waits
this way; this branch was the only exit that did not." That's wrong. That branch
waits only when didReload is false; when reload() did run it called
applyPendingForkValidations internally, so the same steal can happen and it
returns without waiting. The hole isn't unique to the fallback branch. Comment
fixed, and it now says plainly that this is a symptom fix.

The root cause is tracked separately as #13995: the tail restart at
SharedLiveAgentIndex.swift:1868 lacks the !resumedWaiters guard its in-loop
sibling has, so it can resume a waiter and immediately race it. I didn't fold
that in because the obvious guard can strand a pending request when the resumed
waiter's task is cancelled right after resuming — it needs its own change and
its own test.

Worth noting for whoever picks that up: there's no timeout on
waitForForkValidationRequestCompletion, unlike the reload path. Liveness
depends on some drainer eventually retiring the request. Pre-existing and shared
with the non-fallback branch, but it's the thing that would bite if #13995's fix
goes wrong.

No regression test here, and I want to be straight about that: the existing
assertion in WorkspaceForkConversationContextMenuTests already passes on
main, so nothing in the suite fails without this change. The race is genuinely
hard to schedule deterministically, which is also why the related test is
intermittent rather than reliably red.

🤖 Generated with Claude Code

@teamleaderleo teamleaderleo left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 92c4f011ed. This is a code review, not a run. I agree with the independent review above. The change is sound, and it stays with Leo for approval because it changes app runtime behavior.

What I checked by reading SharedLiveAgentIndex.swift at this head:

  • The new wait matches the existing !didReload wait in the live-index branch. Both use the same waitForForkValidationRequestCompletions helper with ownsRequest: true.
  • No lost wakeup. forkValidationRequestIsWaiting and the waiter append both run inside the synchronous withCheckedContinuation body on the main actor. If this call drained its own request, the per-batch defer in applyPendingForkValidations has already retired it (retireProcessingForkValidationRequests), so the wait resumes immediately.
  • The other drainer's paths all release the waiter. A stolen request stays in processingForkValidationRequestIDs until that drainer's defer runs resumeForkValidationRequestCompletionWaiters. The contention branch restores to pending and then re-drains recursively. The fallbackSnapshot == nil, index == nil requeue can't catch a request that carries a fallback snapshot.

What I did not verify: that this fixes the intermittent sharedForkProbeFallbackWaitsForActiveSamePanelValidation failure. There's no deterministic regression test, and one green run proves little. The wait also has no timeout, a gap the non-fallback branch already has. The root cause is tracked in #13995.

— Ophelia g1 🍄
Run: run_cmux_main_red_triage_app_host_census_and_pr_review_20260923_07d8d17b

@teamleaderleo
teamleaderleo merged commit 2a4f3f6 into main Sep 23, 2026
54 checks passed
rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 23, 2026
06c2101 ci: route streamed validation by capability instead of by lane name (manaflow-ai#14002)
1773c54 ci(e2e): start builds from main's DerivedData so test-only changes skip the app compile (manaflow-ai#14016)
c890374 ci: pin the nightly runner guards to the whole expression (manaflow-ai#13997)
8abd2e9 ci: flag condition polls bounded by a Task.yield() count (manaflow-ai#14019)
e4ca672 ci(ios): record the cmux.app upload once Apple accepts it (manaflow-ai#14014)
260b648 ci: check what the runner variables hold, not just what the workflows say (manaflow-ai#13992)
25ad5af feat(terminal): opt-in macOS text-editing gestures at the shell prompt (manaflow-ai#13921)
daf9649 test: drop six focus-history cases superseded by FocusHistoryScopeTests (manaflow-ai#13975)
11202e3 Name the workspace that workspace.reorder could not resolve (manaflow-ai#13961)
2a4f3f6 fix(fork): make the fallback refresh await its own queued validation (manaflow-ai#13960)
4b82298 ci: let test-depot run one app-host test by selector (manaflow-ai#14001)
5d1ecb8 test: give each drained write its own deadline in the short-chunks reader test (manaflow-ai#13999)

# Conflicts:
#	.github/workflows/ci-guards.yml
#	.github/workflows/ci-health-report.yml
#	.github/workflows/ios-appstore-upload.yml
#	.github/workflows/ios-streamed-validate.yml
#	.github/workflows/iroh-release-gate.yml
#	.github/workflows/nightly.yml
#	.github/workflows/test-depot.yml
#	.github/workflows/test-e2e.yml
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant