ci: never abandon a run the owned pool rescue cancelled - #14326
Merged
Merged
Conversation
Contributor
|
All contributors have signed the CLA ✍️ ✅ |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (3)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The rescue gave up 180 s after cancelling, leaving the pull request's run cancelled with no re-run. Run 36074561333 (PR #14314) hit exactly that: two cheap jobs were refused at start, the rescue cancelled the run, and the in-flight compile admission on a Mac took about 5 minutes to settle. Wait up to 20 minutes for the cancel to settle, force-cancelling again every 5, and measure the rescue against the job's own timeout (now 90 minutes) rather than the 60-minute watch, so a refusal found late in the watch is re-run too instead of being left alone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
teamleaderleo
force-pushed
the
ci/rescue-never-abandon
branch
from
September 25, 2026 00:29
533158f to
2218f0e
Compare
rustybret
pushed a commit
to rustybret/bmux
that referenced
this pull request
Sep 25, 2026
488b058 ci: count the 12vcpu macOS pool at 5 machines, the most it ran with a queue (manaflow-ai#14330) bcb162c Merge pull request manaflow-ai#14121 from manaflow-ai/issue-13640-terminal-paste-latency 24efd87 test(focus-recovery): run the automatic apply against a pinned tiny surface (manaflow-ai#14322) 57d3d7c Merge pull request manaflow-ai#14293 from manaflow-ai/issue-14290-directory-unavailable-stale e1a5ea3 ci: never abandon a run the owned pool rescue cancelled (manaflow-ai#14326) a76f47a web: contain Hexclave failures on every page (manaflow-ai#14316) fc7f80d ci: start macOS compile admission beside the fast Linux jobs (manaflow-ai#14314) 85f3d9f ci: roll a PR run over to the next pool when one is full (manaflow-ai#14323) 04e8a05 ci: let E2E runs take an owned Mac with a free slot (manaflow-ai#14311) 034025f reload: resolve the cmux-tui client before the build (manaflow-ai#14313) 7472de4 fix: pass projected resource to stale cwd resolver 9045f37 Merge remote-tracking branch 'origin/main' into issue-14290-directory-unavailable-stale 51a9004 fix: retain accepted Cloud cwd while stale 97d44fd test: retain known Cloud cwd during stale refresh f6df46e fix: retain rich text fallback for lossy paste data f3f43ed fix: retain rich text fallback for lossy paste data 5d8258b test: preserve rich paste fallback and text fidelity a9c54ba fix: keep mixed rich text paste on the fast plain-text path d284b6a test: cover fast paste for mixed plain and HTML clipboard # Conflicts: # .github/workflows/ci-guards.yml # .github/workflows/ci-macos.yml # .github/workflows/ci-owned-pool-rescue.yml # .github/workflows/ci.yml # .github/workflows/test-e2e.yml
teamleaderleo
added a commit
that referenced
this pull request
Sep 25, 2026
A scripts/ci helper any routed job could reach selected every area, macOS, web and Release included, wherever it ran. The walk that decided it also followed comments, docstrings and the routing tables that list helpers as data, so most helpers reached the `changes` job and forced everything. A ci-macos.yml edit above `jobs:`, such as a new workflow_call input, did too. - ci_helper_areas() replaces ci_helper_reaches_routed_lane(): a helper selects the areas that gate the routed jobs running it (ci-macos.yml jobs by the same rules as a job edit, ci-web.yml web, the CLI lane cli, a Linux job the areas its `if:` reads). Routing, status and other Mac jobs still run every area. Comments, docstrings and the three routing tables no longer count as running a helper. - A ci-macos.yml workflow_call input edit reaches only the jobs that read the input, unless the workflow env reads it; a comment-only edit changes no job. Replayed on the 20 CI-only PRs of 2026-09-23/24 that do not edit ci.yml or ci-macos.yml, 5 drop from every area to none or macOS+CLI (#14326, #14312, #14299, #14187: none; #14309, #14250: macOS+CLI), one of them drops web. #14318's ci-macos.yml input edit would select macOS+CLI, not Release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
teamleaderleo
added a commit
that referenced
this pull request
Sep 25, 2026
…#14339) * ci: route CI helper and ci-macos.yml edits to the lanes that run them A scripts/ci helper any routed job could reach selected every area, macOS, web and Release included, wherever it ran. The walk that decided it also followed comments, docstrings and the routing tables that list helpers as data, so most helpers reached the `changes` job and forced everything. A ci-macos.yml edit above `jobs:`, such as a new workflow_call input, did too. - ci_helper_areas() replaces ci_helper_reaches_routed_lane(): a helper selects the areas that gate the routed jobs running it (ci-macos.yml jobs by the same rules as a job edit, ci-web.yml web, the CLI lane cli, a Linux job the areas its `if:` reads). Routing, status and other Mac jobs still run every area. Comments, docstrings and the three routing tables no longer count as running a helper. - A ci-macos.yml workflow_call input edit reaches only the jobs that read the input, unless the workflow env reads it; a comment-only edit changes no job. Replayed on the 20 CI-only PRs of 2026-09-23/24 that do not edit ci.yml or ci-macos.yml, 5 drop from every area to none or macOS+CLI (#14326, #14312, #14299, #14187: none; #14309, #14250: macOS+CLI), one of them drops web. #14318's ci-macos.yml input edit would select macOS+CLI, not Release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: close the routing review's under-selection gaps - An input key with a trailing comment still opens its own input; a flow-style or unreadable one at that indent answers every job. - A Linux job reads area outputs anywhere in its block (folded or step conditions), maps outputs derived from macOS to macOS, and runs every area when it waits on another job or reads an output this cannot place. - The routing tables' imports still run a helper; only their path lists are dead ends. - A helper the Swift package lane runs selects every area, since that lane is chosen by package path. - `#!` lines are not comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The owned-pool rescue cancels a PR run and re-runs it when an owned Mac is busy or refuses a job. Today it gives up 180 s after the cancel if the run hasn't finished cancelling. The PR's run is then left cancelled with nobody to re-run it.
Rescue run 36074563873 did exactly that to CI 36074561333 (PR #14314):
glaeda-std-xcode-26.6("a fleet build holds the host lock")cancelled run, thenforce-cancelled rungave up: run 36074561333 did not finish 180s after cancel; not re-runThe in-flight
macos / macOS compile admissionon a Mac settled about 5 minutes after the cancel. The PR had to be closed and reopened to get CI.Change
CANCEL_WAIT_SECONDSgoes from 180 to 20 minutes. The force-cancel is re-sent every 5 minutes while waiting, not just once. If GitHub refuses a force-cancel (the run settled between the read and the POST), the rescue logs it and keeps waiting instead of aborting.RESCUE_GRACE_SECONDS) past the end of the watch: 60 minutes for a PR run, 150 for an E2E run (ci: let E2E runs take an owned Mac with a free slot #14311). Before, a refusal found in the watch's last few minutes was left alone ("too little of the watch left"). It is now re-run too. No cancel starts without 21 minutes of that grace left.ci-owned-pool-rescue.ymltimeout-minutesgoes from 160 to 180 (the E2E watch, plus the grace, plus 5 minutes for checkout). A test pins it toJOB_TIMEOUT_SECONDS.What's unchanged: the watch itself (60 minutes, one deadline across attempts), the pushed-head check before the re-run, and the "someone else re-ran it" check.
Not changed here
For a refused cheap job, the rescue still cancels the whole run, including a compile already running on another Mac. That throws away minutes of compile to retry a 6-second job. Letting the run finish and then re-running failed jobs would keep the compile, but the refused job would retry later. That tradeoff belongs with the pool-routing work in #14310 and #14318, which also edit this file.
Testing
tests/test_ci_owned_pool_rescue.py, 42 tests:actionlintandvalidate_test_execution_registry.pyare clean.🤖 Generated with Claude Code
Summary by CodeRabbit