Skip to content

ci: give a refused owned job one more try on the fleet before Blacksmith - #14312

Merged
teamleaderleo merged 2 commits into
mainfrom
ci/refused-retry-to-fleet
Sep 24, 2026
Merged

teamleaderleo merged 2 commits into
mainfrom
ci/refused-retry-to-fleet

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Why

Leo's rule: a job an owned runner refuses goes back to the fleet. It goes to Blacksmith only when the fleet keeps refusing or is busy. The rescue already re-runs a refused run's failed jobs (#14299), but only ever watched attempt 1.

GitHub sends no workflow_run requested event for a re-run. Run 36059281883 was re-run to attempt 2, and there is no owned-pool-rescue-36059281883-2; all 200 recent rescue runs are for attempt 1. So attempt 2 can't be watched from a new event.

What (rescue side, scripts/ci/owned_pool_rescue.py)

  • After re-running the failed jobs of attempt 1, the same watch follows attempt 2 (LAST_OWNED_ATTEMPT = 2).
    • It only acts once a job of attempt 2 asks for an owned label; changes isn't re-run, so there's no marker.
    • It stops when no owned job appears, or once every owned job has run past the 120-second refusal window ("the fleet accepted the retry"). It doesn't wait for the whole run, so it stays within the job's 70-minute timeout.
  • A job refused, or queued past the budget, on attempt 2 gets its failed jobs re-run once more, keeping what passed. Attempt 3 and later never take an owned pool, so it can't loop.

Wiring left to the per-job placement PR

Whether attempt 2 actually lands on the fleet depends on each owned-capable job's runs-on reading github.run_attempt == 2 && inputs.pr_refused_retry_runner || first, from a new picker output refused_retry_runner (the owned label) that is separate from retry_runner. The "Place each PR macOS job on a free owned mini" session is restructuring the picker and those runs-on expressions, and it knows which jobs may run owned (GUI jobs may not yet). So the wiring belongs there. Until it lands, attempt 2 takes retry_runner as today and the extra watch ends at its first look.

Verification

tests/test_ci_owned_pool_rescue.py passes locally (32 tests). The fake API now tracks attempts. New cases:

  • the fleet refuses the retry again → second re-run, attempt 3 never watched
  • the fleet accepts the retry → the watch stops once the job is past the refusal window
  • the retry is stuck queued → failed jobs re-run again and nothing that passed is lost

🤖 Generated with Claude Code


Summary by cubic

Gives a refused owned job one more try on the fleet before sending it to Blacksmith.

  • The rescue now follows the re-run (attempt 2) it triggered, watches for owned-label jobs, and re-runs failed jobs again if the fleet refuses or keeps them queued past the budget; attempt 3 and later always take retry_runner on Blacksmith. One deadline spans both attempts, so a rescue that can't fit its cancel and re-run inside it is skipped rather than risk killing a job.
  • Per-job runs-on wiring so attempt 2 can take the owned pool via a new pr_refused_retry_runner picker output is left to the placement change; until that lands, attempt 2 uses retry_runner and the extra watch ends at its first look.

Written for commit 25cdf79. Summary will update on new commits.

Review in cubic

Leo's rule: a job an owned runner refuses goes back to the fleet, and
only to Blacksmith when the fleet keeps refusing or is busy. The rescue
already re-runs a refused run's failed jobs (#14299). GitHub sends no
workflow_run `requested` event for a re-run (run 36059281883's attempt 2
started no rescue), so the watch that re-ran them now follows attempt 2
itself, for owned jobs only, until they have run past the refusal
window. A job refused, or queued past the budget, on attempt 2 gets its
failed jobs re-run once more, keeping what passed; attempt 3 and later
never take an owned pool, so it cannot loop.

Whether attempt 2 lands on the fleet is up to each job's runs-on reading
`github.run_attempt == 2 && inputs.pr_refused_retry_runner` first, which
the per-job placement change wires; until then attempt 2 takes
retry_runner and the extra watch ends at its first look.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 20 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: afd1c7d7-67d0-44cf-8186-d1577e70d764

📥 Commits

Reviewing files that changed from the base of the PR and between 0cc5be3 and 25cdf79.

📒 Files selected for processing (3)
  • docs/ci-runners.md
  • scripts/ci/owned_pool_rescue.py
  • tests/test_ci_owned_pool_rescue.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

Review of #14312:

- A re-run whose jobs include none on an owned label ends the watch at
  the first look that lists jobs; it no longer polls until the run
  finishes (179 polls over an hour in the reviewer's probe).
- One deadline covers both attempts, and a rescue that could not cancel,
  wait for the cancel and re-run inside it is not started, so the job
  is never killed between a cancel and its re-run.
- Tests for both, and the docstring now says the attempt-2 path cancels
  the run and re-runs its failed and cancelled jobs, keeping what passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@teamleaderleo
teamleaderleo merged commit 9d53af4 into main Sep 24, 2026
48 checks passed
@teamleaderleo
teamleaderleo deleted the ci/refused-retry-to-fleet branch September 24, 2026 23:53
teamleaderleo added a commit that referenced this pull request Sep 25, 2026
Resolves owned_pool_rescue.py against #14312 (one more fleet try for a
refused job). E2E runs keep their own rule: every re-run attempt takes
retry_label on Blacksmith, so the follow-on watch of attempt 2 stops.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 25, 2026
3d5b648 test: give the Cloud Desktop fixture a routable window for pane drops (manaflow-ai#14304)
9d53af4 ci: give a refused owned job one more try on the fleet before Blacksmith (manaflow-ai#14312)
379091b ci: let PR runs overflow to macOS 15 at 4 queued jobs deeper, not 12 (manaflow-ai#14319)
df6efb5 test(cloud): key the refresh URL protocol stub per request, not by address (manaflow-ai#14239)
51bd322 ci: build only the CLI product for CLI-only changes (manaflow-ai#14212)
0345a5c ci: make a changed-suites run prove a known-failure fix (manaflow-ai#14307)
370b7f6 ci: fix the owned build state save step's argument count (manaflow-ai#14309)
319adff test: wait for the SSH cleanup policy bound after a restored-attach signal (manaflow-ai#14305)
0cc5be3 fix(portal): flush the coalesced live-resize pass on its first hop (manaflow-ai#14297)
5887891 test(minimal-mode): measure the toggle only after setup stops re-rendering (manaflow-ai#14298)
5fbcc48 ci: charge newer PR runs what they took on the owned pool, not a guess (manaflow-ai#14300)

# Conflicts:
#	.github/workflows/ci-macos.yml
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Note: pr_refused_retry_runner is described here and in owned_pool_rescue.py, but nothing on main sets it yet, so a refused job still goes straight to Blacksmith on attempt 2. The wiring is in #14325, which is stacked on #14318.

teamleaderleo added a commit that referenced this pull request Sep 25, 2026
A scripts/ci helper any routed job could reach selected every area, macOS,
web and Release included, wherever it ran. The walk that decided it also
followed comments, docstrings and the routing tables that list helpers as
data, so most helpers reached the `changes` job and forced everything. A
ci-macos.yml edit above `jobs:`, such as a new workflow_call input, did too.

- ci_helper_areas() replaces ci_helper_reaches_routed_lane(): a helper
  selects the areas that gate the routed jobs running it (ci-macos.yml jobs
  by the same rules as a job edit, ci-web.yml web, the CLI lane cli, a Linux
  job the areas its `if:` reads). Routing, status and other Mac jobs still
  run every area. Comments, docstrings and the three routing tables no
  longer count as running a helper.
- A ci-macos.yml workflow_call input edit reaches only the jobs that read
  the input, unless the workflow env reads it; a comment-only edit changes
  no job.

Replayed on the 20 CI-only PRs of 2026-09-23/24 that do not edit ci.yml or
ci-macos.yml, 5 drop from every area to none or macOS+CLI (#14326, #14312,
#14299, #14187: none; #14309, #14250: macOS+CLI), one of them drops web.
#14318's ci-macos.yml input edit would select macOS+CLI, not Release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
teamleaderleo added a commit that referenced this pull request Sep 25, 2026
…#14339)

* ci: route CI helper and ci-macos.yml edits to the lanes that run them

A scripts/ci helper any routed job could reach selected every area, macOS,
web and Release included, wherever it ran. The walk that decided it also
followed comments, docstrings and the routing tables that list helpers as
data, so most helpers reached the `changes` job and forced everything. A
ci-macos.yml edit above `jobs:`, such as a new workflow_call input, did too.

- ci_helper_areas() replaces ci_helper_reaches_routed_lane(): a helper
  selects the areas that gate the routed jobs running it (ci-macos.yml jobs
  by the same rules as a job edit, ci-web.yml web, the CLI lane cli, a Linux
  job the areas its `if:` reads). Routing, status and other Mac jobs still
  run every area. Comments, docstrings and the three routing tables no
  longer count as running a helper.
- A ci-macos.yml workflow_call input edit reaches only the jobs that read
  the input, unless the workflow env reads it; a comment-only edit changes
  no job.

Replayed on the 20 CI-only PRs of 2026-09-23/24 that do not edit ci.yml or
ci-macos.yml, 5 drop from every area to none or macOS+CLI (#14326, #14312,
#14299, #14187: none; #14309, #14250: macOS+CLI), one of them drops web.
#14318's ci-macos.yml input edit would select macOS+CLI, not Release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: close the routing review's under-selection gaps

- An input key with a trailing comment still opens its own input; a
  flow-style or unreadable one at that indent answers every job.
- A Linux job reads area outputs anywhere in its block (folded or step
  conditions), maps outputs derived from macOS to macOS, and runs every
  area when it waits on another job or reads an output this cannot place.
- The routing tables' imports still run a helper; only their path lists
  are dead ends.
- A helper the Swift package lane runs selects every area, since that lane
  is chosen by package path.
- `#!` lines are not comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant