Repository navigation
Recover a Cloud machine graph stuck on an equal-cursor conflict - #15328
teamleaderleo merged 11 commits into
Conversation
Red: a full snapshot that disagrees with the applied graph at the same cursor is refused forever, and there is no equalCursorConflict state. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A full snapshot that disagreed with the delta-built graph at the same cursor was refused forever, so the machine's graph never became current again. The first conflict from a current full refresh now keeps the graph and schedules the bounded recovery read; a second conflict at the same cursor adopts the daemon's snapshot and leaves a breadcrumb. Event-feed snapshots and stale reads never adopt. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Warning Review limit reachedNext included review available in 7 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
Review: this is a self-review by the authoring agent, not a separate review subagent. This worker could not spawn one, so a separate subagent review is still owed before merge. Checked:
Fixed during self-review:
Left:
|
Catch-up merge by scripts/ci/catch_up_pr.py (RFC manaflow-ai#14631). Merged by scripts/merge-main.sh: origin/main at 436909b, the newest commit with green CI fast guards (2 newer skipped). Catch-up-previous-head: 796b6a1 Catch-up-base: 436909b
Catch-up merge by scripts/ci/catch_up_pr.py (RFC manaflow-ai#14631). Merged by scripts/merge-main.sh: origin/main at 0c753fe. Resolved conflicts: - cmux.xcodeproj/project.pbxproj: union of added entries, then normalize-pbxproj.py Catch-up-previous-head: b4abb93 Catch-up-base: 0c753fe
Only the install that arms an equal-cursor conflict schedules the recovery read, so stale or rename-fenced reads do not spend the budget. A delta that advances the cursor and a feature-flag suspend both clear the armed conflict. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Review (separate subagent, correctness first; the self-review above was not one): nothing serious. Checked:
Fixed:
Left:
|
CI failure attributionCI passes on Written by |
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # cmux.xcodeproj/project.pbxproj
|
@coderabbitai review |
|
|
Review (separate subagent) of head 8e2ff57: nothing serious.
|
|
Review (separate subagent) of head 8e2ff57: no correctness bugs. Arming and adoption are gated on a current full refresh, the rename fence still runs first, and the recovery budget is spent only by the read that armed. Test coverage is partial, so the unit tests are not enough to stand in for a dogfood:
Leaving this open until it can be dogfooded against a Cloud backend. |
# Conflicts: # cmux.xcodeproj/project.pbxproj
…aring Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d-equal-cursor-recovery # Conflicts: # cmux.xcodeproj/project.pbxproj
Catch-up merge by scripts/ci/catch_up_pr.py (RFC manaflow-ai#14631). Merged by scripts/merge-main.sh: origin/main at 62ee70e. Catch-up-previous-head: 16b4f81 Catch-up-base: 62ee70e
|
Subagent review at 8e2ff57: no correctness bugs (arming and adoption gated on a current full refresh, rename fence first, recovery budget spent only by the arming read). Findings: test catalog lifetime bug, fixed in e975b55; coverage gaps, addressed in 6a7c880 (test only). Only main merges since. Landing: merging main after #15414, then auto-merge. |
|
Merge receipt for |
762c3ed Recover a Cloud machine graph stuck on an equal-cursor conflict (manaflow-ai#15328) 524ebff ci: replay the fuzz regressions on sidebar, split and window changes (manaflow-ai#15412) 818d475 Let a user's Cloud open dial even right after a background link failure (manaflow-ai#15291) 97491a7 Let the Cloud toolbar name the machine-list failure it has (manaflow-ai#15236) 0abac32 PR media: adopt CI's build only, start when CI completes, run for every app PR (manaflow-ai#15418) 7f08715 ci(seed): keep the trusted seed on the Mac before the R2 upload (manaflow-ai#15411) 5663c13 Finish the destroy work where a Cloud machine is first found gone (manaflow-ai#15359) # Conflicts: # .github/workflows/ci.yml # .github/workflows/pr-media.yml # .github/workflows/seed-derived-data.yml # .github/workflows/test-e2e.yml
Summary
A Cloud or SSH machine's graph can wedge for the life of the app. A full snapshot and the graph the deltas built can disagree at the same cursor. When they do,
installSnapshotIfNewerrefuses the snapshot as an "equal-cursor conflict", and no later read at that cursor ever succeeds.refreshCurrentGraphnever becomes current, and opens on that machine fail until relaunch. #15205 fixes one known cause (exited terminals' tab fields). Any other strictly compared field that differs at the same revision still wedges:cwd,lifecycle, or keys the model doesn't include.Now:
cloud.state.equalCursorConflictAdoptedbreadcrumb. A full read at the current cursor is the daemon's own answer. So a wedge always breaks, and a single race never discards the graph.cloudStateand are reapplied on publish. A pending rename's predecessor is still refused outright and never arms recovery.Composes with #15205 and #15283:
git merge-treeshows no source conflict with either. The only overlap with #15283 is both PRs adding a test file toproject.pbxproj, which catch-up regenerates.Testing
cmuxTests/CloudEqualCursorConflictRecoveryTests.swift, committed failing before the fix:scripts/sync-test-wiringfor the pbxproj.scripts/verify-local.pypasses. No local native build; CI runs the tests.Changelog
Fixed: a Cloud or SSH machine whose state got out of sync no longer stays stuck until relaunch; the next refresh recovers it.
🤖 Generated with Claude Code
Summary by cubic
Recovers a Cloud or SSH machine's graph that was permanently stuck when a full snapshot disagreed with the delta-built graph at the same cursor. The first such conflict from a full refresh keeps the installed graph and schedules the bounded state-recovery read; a second conflict at the same cursor adopts the daemon's full snapshot, so the wedge always breaks.
CloudEqualCursorConflictRecoveryTestscovering the arming, adoption, fence, suspend, and clearing paths.Written for commit 7784fef. Summary will update on new commits.