ci: queue main's whole full-suite run on the owned pool - #15356
Conversation
With per-job placement, main's full-suite dispatch put the jobs that did not fit on the owned pool on the retry runner. When the owned machines were busy, those jobs waited behind every overflowed pull request on Blacksmith: run 36402943637 had five jobs queued there for over an hour behind 76 others with 3 running, and the next main run waited on it in main's concurrency group, so main produced no verdict. Main now queues its whole run on the owned pool it picked (with queue rounds on). The owned queue drains in rounds, and the rescue still bounds the wait. With the rounds at 0 it splits as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
Warning Review limit reachedNext included review available in 4 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (2)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
Dogfood build of cmux DEV pr-15356-b0993281.app The link opens this exact commit in the cmux dev menu bar app. The build starts on each push and the page waits until it is ready; a newer push replaces it. It signs in against production, so Cloud or backend changes still need a tagged build with a development backend. |
A root budget of the run's machines holds every job on a pool with or without gui runners, and a pool smaller than the run (light) keeps the split, since its excess would wait past the owned-pool rescue's budget. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Merge receipt for |
7b6ad14 PR media: tour-only pushes load the earlier build of the same inputs (manaflow-ai#15361) 842d580 ci: prebuild selected Swift packages in parallel before the serial test pass (manaflow-ai#15114) 1a7467a dogfood: give the agent activity reorder tour paths globs (manaflow-ai#15360) 31b2619 Recover routed Claude sessions through cmux restore; journal closed panes (manaflow-ai#15324) 5a3ca1a ci: price mini contention in distance routing and keep SwiftPM builds on owned Macs (manaflow-ai#14804) b8afe20 Add notifications.suppressWhenAppFocused setting (manaflow-ai#14701) 3843dd6 ci: queue main's whole full-suite run on the owned pool (manaflow-ai#15356) # Conflicts: # .github/workflows/ci-guards.yml # .github/workflows/ci-macos.yml # .github/workflows/pr-media.yml # .github/workflows/test-e2e.yml
* ci: place attempt 2 like attempt 1, owned minis first A re-run's attempt 2 now takes the owned pools the way attempt 1 does, whoever started it; only attempt 3 on (the rescue's move of a stuck attempt 2) goes to retry_runner on Blacksmith. - Pickers, janitor and runs-on: the bot's retry rule is attempt > 2 (LAST_OWNED_ATTEMPT), and main/side lanes keep their pools on attempt 2. - admission-placement and late-placement run again on a re-run that re-runs them, publish the attempt they placed, and consumers use a placement only when it matches github.run_attempt. Overflow reuses the late-placement queue comparison from #15336/#15356. - admission-placement skips minis whose jobs failed in the previous attempt (one attempts/N-1/jobs read). - The rescue sweeper watches every unfinished CI re-run up to attempt 2, gives it the queue allowance, and waits for a re-run late-placement. - CI_OWNED_LIGHT_RETRY is gone: attempt 2 takes any owned tier. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: update the wiring guard for attempt 2 on the owned pools The workflow wiring tests now expect attempt 2 to be placed like attempt 1 and only the bot's attempt 3 to take the retry runner. late-placement skips the bot's attempt 3, so its labels cannot pull that attempt back onto a mini ahead of the retry clause. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: drop the removed light-retry argument and env after merging main Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: keep watching a re-run's own admission; re-run a machine-failed attempt 2 - owned_pool_rescue.watch(): "the fleet accepted the retry" no longer stops the watch while this attempt runs its own compile admission. The bot's attempt 2 re-running a failed admission on the root label counted as accepted past the refusal window, so the watch ended mid-compile and a shard stuck on an owned label afterwards was never rescued. - classify_failures.rerun_decision(): the bot re-runs a machine-failed attempt up to LAST_OWNED_ATTEMPT. An online but broken mini (full disk, failed product restore) may fail attempt 2 again, and failed-only re-runs never re-run admission-placement, so its exclusion does not apply; attempt 3 goes to Blacksmith, which ends the chain. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: serialize timing-sensitive CLI regressions on shared Macs * fix: classify carried rescue jobs from attempt start --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Summary
Testing
python3 -m unittest tests.test_ci_pr_runner_pool(227 tests) passes, including a new case with production's shape: 0 of 14 root runners free and 6 queue places, 9 needed. Main places all ten jobs on the minis; a pull request with the same load still splits.🤖 Generated with Claude Code
Summary by cubic
Queues main's entire full-suite run on the owned pool so its verdict is no longer delayed when the owned machines are busy. Previously, jobs that did not fit on the owned pool went to the retry runner and waited behind every overflowed pull request, which could hold main's verdict for over an hour.
Written for commit b099328. Summary will update on new commits.