Repository navigation
Repair shared native test synchronization and CLI fixtures - #13263
Conversation
Await the fake sign-in flow startup instead of assuming one Task.yield finishes the model task. Match the browser-purpose staging tunnel filename. Production behavior and the remaining assertions are unchanged.
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
All contributors have signed the CLA ✍️ ✅ |
|
Warning Review limit reachedNext included review available in 15 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (2)
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 10 included reviews per hour; 1 remains after this review. 📝 WalkthroughWalkthroughThe tests now synchronize with fake sign-in startup through checked continuations instead of task yielding. The staging tunnel test now expects the browser-specific configuration filename. ChangesSign-in test synchronization
Tunnel filename assertion
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Other Suggested reviewers: 🚥 Pre-merge checks | ✅ 24 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (24 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Exercise alternative layout targets independently; retain transport and response assertions. Scope tunnel expectations by browser role and pin the stable interface where asserted. Expect Claude structured hook acknowledgement and bind Cursor approval config to its actual fixture directory.
Use the existing app-host AsyncStream/task-group deadline convention. A missing startup throws an explicit test error; cancellation terminates the stream waiter. Preserve same-actor completion ordering.
|
Shared full-suite failure evidence from the module/contributor campaign (2026-09-20 UTC):
No checks or assertions were disabled in the owning PRs. They remain blocked until the relevant full-suite evidence is green. Shared repair should land once rather than be copied as unrelated changes into each PR. |
|
Pushed 46c12b3 to make the notification click-action delivery fixture explicitly unfocused and restore its previous focus override afterward. A focused live workspace intentionally suppresses external delivery; this test checks the click-action payload, so ambient host focus must not decide whether its delivery callback runs. Both stored and delivered action assertions remain unchanged. The predecessor 7f39cc7 passed package tests and compile admission in run 35529349027 before this push; its runtime shards had started. That is partial predecessor evidence, not validation of the new head. No local native build was run. |
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
Integrated the #86 owner's two-file fixture patch on the existing branch as
Verified the supplied patch SHA-256, exact owning head, initializer/transform source and |
Main landed equivalent fixes for the sign-in start signal, VM layout target runs, simulator rounding and tunnel naming; take main's versions and drop duplicate lines the auto-merge produced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
macOS concurrency on Blacksmith is roughly ten slots, and a run that can no longer go green keeps holding them. This is structural: `ci-status` accepts only `success` or `skipped` from each of its `needs`, so an `app-host unit tests` shard concluding `failure` fails the `macos` reusable-workflow call and the required check by construction. Across the 299 CI runs created between 2026-09-22T06:05Z and 17:00Z, 21 runs had such a shard failure, `ci-status` concluded `failure` in all 21, and their sibling macOS jobs went on to burn 1,522 macOS runner-minutes after the verdict was already fixed. The janitor gains a second rule for that shape. It stays inside the existing contract: the scheduled run still only reports, cancellation still requires a workflow_dispatch with `cleanup`, and both rules share the one `max_actions` budget, doomed runs first. The rule names one job rather than reading the whole `needs` list, because cancelling must reclaim only test shards and never a compile. `app-host-unit-tests` needs `macos-compile-admission` to have succeeded, so the compiled app-host product is published and seeded before any shard can fail and later runs still reuse it. A Linux guard failure decides `ci-status` just as firmly but lands while the macOS compile is still running, where cancelling would destroy a product other runs would have reused. The run repairing the failing job is the exception that matters, because its remaining shards are the result someone is waiting on. A pull request whose diff touches the shards' own inputs is never cancelled, and `no-janitor` covers a fix the path list cannot recognise. Replayed over the 21 real runs, this preserves every app-host repair among them (#13643, #13579, #13574, #13427, #13414, #13615, #13263, #13271) and leaves two unrelated runs eligible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
macOS concurrency on Blacksmith is roughly ten slots, and a run that can no longer go green keeps holding them. This is structural: `ci-status` accepts only `success` or `skipped` from each of its `needs`, so an `app-host unit tests` shard concluding `failure` fails the `macos` reusable-workflow call and the required check by construction. Across the 299 CI runs created between 2026-09-22T06:05Z and 17:00Z, 21 runs had such a shard failure, `ci-status` concluded `failure` in all 21, and their sibling macOS jobs went on to burn 1,522 macOS runner-minutes after the verdict was already fixed. The janitor gains a second rule for that shape. It stays inside the existing contract: the scheduled run still only reports, cancellation still requires a workflow_dispatch with `cleanup`, and both rules share the one `max_actions` budget, doomed runs first. The rule names one job rather than reading the whole `needs` list, because cancelling must reclaim only test shards and never an in-flight compile. `app-host-unit-tests` needs `macos-compile-admission` to have succeeded, so the compiled app-host product is already published before any shard can fail and there is nothing in flight to lose. A Linux guard failure decides `ci-status` just as firmly but lands while the macOS compile is still running, so a rule built on it would be discarding compiles: safe only while cross-run reuse stays broken (#13709), and destructive when #13718 lands. The run repairing the failing job is the exception that matters, because its remaining shards are the result someone is waiting on. A pull request whose diff touches the shards' own inputs is never cancelled, and `no-janitor` covers a fix the path list cannot recognise. Replayed over the 21 real runs, this preserves every app-host repair among them (#13643, #13579, #13574, #13427, #13414, #13615, #13263, #13271) and leaves two unrelated runs eligible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
macOS concurrency on Blacksmith is roughly ten slots, and a run that can no longer go green keeps holding them. This is structural: `ci-status` accepts only `success` or `skipped` from each of its `needs`, so an `app-host unit tests` shard concluding `failure` fails the `macos` reusable-workflow call and the required check by construction. Across the 299 CI runs created between 2026-09-22T06:05Z and 17:00Z, 21 runs had such a shard failure, `ci-status` concluded `failure` in all 21, and their sibling macOS jobs went on to burn 1,522 macOS runner-minutes after the verdict was already fixed. The janitor gains a second rule for that shape. It stays inside the existing contract: the scheduled run still only reports, cancellation still requires a workflow_dispatch with `cleanup`, and both rules share the one `max_actions` budget, doomed runs first. The rule names one job rather than reading the whole `needs` list, because cancelling must reclaim only test shards and never an in-flight compile. `app-host-unit-tests` needs `macos-compile-admission` to have succeeded, so the compiled app-host product is already published before any shard can fail and there is nothing in flight to lose. A Linux guard failure decides `ci-status` just as firmly but lands while the macOS compile is still running, and since #13718 fixed cross-run reuse on pull requests those compiles produce products later runs consume. The run repairing the failing job is the exception that matters, because its remaining shards are the result someone is waiting on. A pull request whose diff touches the shards' own inputs is never cancelled, and `no-janitor` covers a fix the path list cannot recognise. Replayed over the 21 real runs, this preserves every app-host repair among them (#13643, #13579, #13574, #13427, #13414, #13615, #13263, #13271) and leaves two unrelated runs eligible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`ci-status` accepts only `success` or `skipped` from each of its `needs`, so an `app-host unit tests` shard concluding `failure` fails the `macos` reusable-workflow call and the required check by construction; no later job takes it back. Across the 299 CI runs created between 2026-09-22T06:05Z and 17:00Z, 21 runs had such a shard failure, `ci-status` concluded `failure` in all 21, and their sibling macOS jobs burned 1,522 macOS runner-minutes after the verdict was already fixed. This lands as a fourth category in the queue janitor rather than a second janitor. Reclaiming macOS pool capacity is that module's charter, and putting it there means one queue threshold, one priority order, one per-sweep cancel cap and one concurrency group instead of two workflows with `actions: write` and no shared bound. The threshold gate is also the right policy on its own terms: cancelling a doomed run when the pool is idle frees nothing anybody is waiting for and still destroys the remaining shard output. The category is ordered last. Categories (a) to (c) cancel runs nobody will read -- an experiment push, a closed or superseded PR, a replaced full-suite run. A doomed run is still current and its remaining shards are still readable, so it is the most debatable of the four and is spent only after the others. That same difference is why this category needs a fix-branch exclusion the others do not. A run whose diff touches the shards' own inputs is the run whose remaining shards someone is waiting on, and is never cancelled; `no-janitor` covers a fix the path list cannot recognise, and an unreadable diff preserves the run. Replayed over the 21 real runs, this preserves every app-host repair among them (#13643, #13579, #13574, #13427, #13414, #13615, #13263, #13271, and #13408 which was cancelled by hand and had to be restarted) and leaves two unrelated runs eligible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ed (#13724) `ci-status` accepts only `success` or `skipped` from each of its `needs`, so an `app-host unit tests` shard concluding `failure` fails the `macos` reusable-workflow call and the required check by construction; no later job takes it back. Across the 299 CI runs created between 2026-09-22T06:05Z and 17:00Z, 21 runs had such a shard failure, `ci-status` concluded `failure` in all 21, and their sibling macOS jobs burned 1,522 macOS runner-minutes after the verdict was already fixed. This lands as a fourth category in the queue janitor rather than a second janitor. Reclaiming macOS pool capacity is that module's charter, and putting it there means one queue threshold, one priority order, one per-sweep cancel cap and one concurrency group instead of two workflows with `actions: write` and no shared bound. The threshold gate is also the right policy on its own terms: cancelling a doomed run when the pool is idle frees nothing anybody is waiting for and still destroys the remaining shard output. The category is ordered last. Categories (a) to (c) cancel runs nobody will read -- an experiment push, a closed or superseded PR, a replaced full-suite run. A doomed run is still current and its remaining shards are still readable, so it is the most debatable of the four and is spent only after the others. That same difference is why this category needs a fix-branch exclusion the others do not. A run whose diff touches the shards' own inputs is the run whose remaining shards someone is waiting on, and is never cancelled; `no-janitor` covers a fix the path list cannot recognise, and an unreadable diff preserves the run. Replayed over the 21 real runs, this preserves every app-host repair among them (#13643, #13579, #13574, #13427, #13414, #13615, #13263, #13271, and #13408 which was cancelled by hand and had to be restarted) and leaves two unrelated runs eligible. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Full app-host CI exposed existing fixture failures shared by unrelated PRs. This change repairs the source-proven cases while retaining the behavioral assertions and production code.
Sign-in tests wait for the fake flow’s actual startup signal instead of assuming one Task.yield finishes a model task that itself yields. A cancellable AsyncStream/task-group race follows the repository’s 10-second app-host readiness bound and throws if startup never occurs; state, URL and start-count checks remain.
Tunnel fixtures expect browser-purpose scoped paths; the role-isolation fixture explicitly selects the stable interface whose legacy paths it checks. Key separation, permissions and config isolation checks remain.
VM layout apply exercises --workspace and --name independently, matching their documented mutual exclusion. Document bytes, vm.exec/no-open, result, hint and warning checks remain.
Claude clear-session startup expects the canonical structured
{}acknowledgment. Pane-targeted clear/status checks remain.Cursor approval explicitly points at its authored config directory, so the isolated runner’s XDG_CONFIG_HOME cannot redirect lookup away from the fixture. Approval, sandbox and persistence checks remain.
Notification click-action delivery explicitly makes the app unfocused and restores the previous override. A focused live workspace intentionally suppresses external delivery; both stored and delivered payload assertions remain.
Simulator orientation coordinates tolerate
1e-12rounding from1 - coordinate, retaining all axes, exact phase/edge and required secondary touch. The naming-agent fixture injects a unique UserDefaults suite while retaining its exact configured-value assertion; runtime confirmation is pending.Validation:
git diff --checkpasses. No local native build was run; exact-head execution is gated by this PR’s full-ci run. Prior red evidence is preserved in 13230 shard4, 13232 shard3, and 13201 shard4. These are existing tests corrected to wait for observable completion or construct the documented input/environment; no assertions were deleted and no production behavior was changed.Based on main
b093335054fbf2fccad53ece5841799b6ad382f8. Nine existing test files are changed. Current head173b4cee81a5beff04fbed1a9f629726b5b8d678requests full validation with the additional simulator-rounding and unique-defaults fixtures. Earlier package/compile successes and runtime failures are predecessor evidence, not current-head proof. The unique-defaults change removes shared-domain coupling but its original runtime interference mechanism has not been independently reproduced. Other full-suite failures remain tracked in #8565 and are not claimed fixed; further verified fixture repairs belong in this PR. The distinct production Grok environment defect is repaired separately in #13271.