Skip to content

ci: queue main's whole full-suite run on the owned pool - #15356

Merged
teamleaderleo merged 2 commits into
mainfrom
fix/main-owned-queue
Sep 28, 2026
Merged

teamleaderleo merged 2 commits into
mainfrom
fix/main-owned-queue

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • With per-job placement, main's full-suite dispatch sent the jobs that did not fit on the owned pool to the retry runner. When the owned machines were busy, those jobs queued behind every overflowed pull request there. Run 36402943637 had five jobs (tests-build-and-lag and four app-host shards) queued for over an hour behind 76 others with 3 running, and the next main run sat pending behind it in main's concurrency group, so main gave no verdict at all.
  • Main now queues its whole run on the owned pool it picked, when queue rounds are on. The owned queue drains in rounds, and the owned-pool rescue still bounds the wait. With the rounds at 0 (the kill switch) it splits as before, so the upstream-parity table is unchanged.
  • Pull requests keep the split.

Testing

  • python3 -m unittest tests.test_ci_pr_runner_pool (227 tests) passes, including a new case with production's shape: 0 of 14 root runners free and 6 queue places, 9 needed. Main places all ten jobs on the minis; a pull request with the same load still splits.

🤖 Generated with Claude Code


Summary by cubic

Queues main's entire full-suite run on the owned pool so its verdict is no longer delayed when the owned machines are busy. Previously, jobs that did not fit on the owned pool went to the retry runner and waited behind every overflowed pull request, which could hold main's verdict for over an hour.

  • Queues whole only on a pool that can hold the entire run; a smaller pool keeps splitting, since the excess would wait past the rescue's budget.
  • Disabled when queue rounds are at 0; main then splits as before, so the upstream-parity table is unchanged.
  • Pull requests keep the split behavior.
  • Adds a test for production's shape (0 of 14 root runners free, 6 queue places, 9 needed).

Written for commit b099328. Summary will update on new commits.

Review in cubic

With per-job placement, main's full-suite dispatch put the jobs that did
not fit on the owned pool on the retry runner. When the owned machines
were busy, those jobs waited behind every overflowed pull request on
Blacksmith: run 36402943637 had five jobs queued there for over an hour
behind 76 others with 3 running, and the next main run waited on it in
main's concurrency group, so main produced no verdict.

Main now queues its whole run on the owned pool it picked (with queue
rounds on). The owned queue drains in rounds, and the rescue still
bounds the wait. With the rounds at 0 it splits as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 28, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 4 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 34b62960-5987-4294-936d-5db03b1d3951

📥 Commits

Reviewing files that changed from the base of the PR and between 16f1270 and b099328.

📒 Files selected for processing (2)
  • scripts/ci/pr_runner_pool.py
  • tests/test_ci_pr_runner_pool.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Dogfood build of b099328199c17332d55f570628ac13c8925236a1

cmux DEV pr-15356-b0993281.app

The link opens this exact commit in the cmux dev menu bar app. The build starts on each push and the page waits until it is ready; a newer push replaces it. It signs in against production, so Cloud or backend changes still need a tagged build with a development backend.

A root budget of the run's machines holds every job on a pool with or
without gui runners, and a pool smaller than the run (light) keeps the
split, since its excess would wait past the owned-pool rescue's budget.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo
teamleaderleo merged commit 3843dd6 into main Sep 28, 2026
52 checks passed
@teamleaderleo
teamleaderleo deleted the fix/main-owned-queue branch September 28, 2026 12:25
@github-actions

Copy link
Copy Markdown
Contributor

Merge receipt for b099328199: every check was green at merge (14 verified; 18 skipped by policy). Full suite runs on main after merge.

rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 28, 2026
7b6ad14 PR media: tour-only pushes load the earlier build of the same inputs (manaflow-ai#15361)
842d580 ci: prebuild selected Swift packages in parallel before the serial test pass (manaflow-ai#15114)
1a7467a dogfood: give the agent activity reorder tour paths globs (manaflow-ai#15360)
31b2619 Recover routed Claude sessions through cmux restore; journal closed panes (manaflow-ai#15324)
5a3ca1a ci: price mini contention in distance routing and keep SwiftPM builds on owned Macs (manaflow-ai#14804)
b8afe20 Add notifications.suppressWhenAppFocused setting (manaflow-ai#14701)
3843dd6 ci: queue main's whole full-suite run on the owned pool (manaflow-ai#15356)

# Conflicts:
#	.github/workflows/ci-guards.yml
#	.github/workflows/ci-macos.yml
#	.github/workflows/pr-media.yml
#	.github/workflows/test-e2e.yml
teamleaderleo added a commit that referenced this pull request Sep 30, 2026
* ci: place attempt 2 like attempt 1, owned minis first

A re-run's attempt 2 now takes the owned pools the way attempt 1 does,
whoever started it; only attempt 3 on (the rescue's move of a stuck
attempt 2) goes to retry_runner on Blacksmith.

- Pickers, janitor and runs-on: the bot's retry rule is attempt > 2
  (LAST_OWNED_ATTEMPT), and main/side lanes keep their pools on attempt 2.
- admission-placement and late-placement run again on a re-run that
  re-runs them, publish the attempt they placed, and consumers use a
  placement only when it matches github.run_attempt. Overflow reuses the
  late-placement queue comparison from #15336/#15356.
- admission-placement skips minis whose jobs failed in the previous
  attempt (one attempts/N-1/jobs read).
- The rescue sweeper watches every unfinished CI re-run up to attempt 2,
  gives it the queue allowance, and waits for a re-run late-placement.
- CI_OWNED_LIGHT_RETRY is gone: attempt 2 takes any owned tier.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: update the wiring guard for attempt 2 on the owned pools

The workflow wiring tests now expect attempt 2 to be placed like attempt 1
and only the bot's attempt 3 to take the retry runner. late-placement
skips the bot's attempt 3, so its labels cannot pull that attempt back
onto a mini ahead of the retry clause.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: drop the removed light-retry argument and env after merging main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: keep watching a re-run's own admission; re-run a machine-failed attempt 2

- owned_pool_rescue.watch(): "the fleet accepted the retry" no longer
  stops the watch while this attempt runs its own compile admission. The
  bot's attempt 2 re-running a failed admission on the root label counted
  as accepted past the refusal window, so the watch ended mid-compile and
  a shard stuck on an owned label afterwards was never rescued.
- classify_failures.rerun_decision(): the bot re-runs a machine-failed
  attempt up to LAST_OWNED_ATTEMPT. An online but broken mini (full disk,
  failed product restore) may fail attempt 2 again, and failed-only
  re-runs never re-run admission-placement, so its exclusion does not
  apply; attempt 3 goes to Blacksmith, which ends the chain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: serialize timing-sensitive CLI regressions on shared Macs

* fix: classify carried rescue jobs from attempt start

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant