Skip to content

ci: a job stuck with no runner no longer cancels jobs running on a mini - #16479

Merged
teamleaderleo merged 3 commits into
mainfrom
ci-rescue-spare-running-minis
Oct 1, 2026
Merged

teamleaderleo merged 3 commits into
mainfrom
ci-rescue-spare-running-minis

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

The owned-pool rescue cancels a whole run when one of its jobs has waited past its budget with no runner, then re-runs the failed and cancelled jobs on Blacksmith. Since #16435 it also covers cmux-next.yml runs. On #16463 and #16468, release-compile waited for a mini while swift-test was already running on one. The rescue cancelled the attempt, which killed a swift test three minutes into its run on a mini and moved it to Blacksmith as attempt 2. That works against minis-first.

Now assess() holds a stuck job while any job of the run is running on a persistent runner (an owned label, in progress, past glaeda's setup hook). The look stays watch with waiting=True:

  • If the running job ends inside the watch, the rescue proceeds as before. For side lanes, E2E runs and attempt 2 onwards, the cancel then touches only the stuck job, and the re-run of failed and cancelled jobs keeps what finished. A picker CI run (ci.yml attempt 1) still gets its usual full re-run.
  • If the watch ends first ("watch limit reached"), nothing is cancelled. The job stays queued for the mini that frees up. A later sweeper re-adopts the run at most once, while its marker is under 150 minutes old.

Refused jobs and jobs held in setup are judged before this hold. A refusal next to a running mini job still gets its failed jobs re-run (at once on main, otherwise when the run finishes or at the watch's end). A job held in setup is still rescued at the watch's end. Those two paths can still cancel a job running on a mini. They aren't jobs that never got a runner, so they're outside this change.

Trade-offs:

  • In a picker CI run, a stuck shard usually has a sibling shard on a mini, so its rescue now waits until the mini shards finish.
  • A stuck job whose label no mini can serve, next to a sibling that outlives the watch, stays queued, where before it moved to Blacksmith.

Both follow from minis-first.

Tests:

  • test_a_stuck_side_job_never_cancels_a_sibling_running_on_a_mini, in its own first commit, fails on main with 'cancel' unexpectedly found.
  • test_a_sibling_on_a_mini_does_not_hide_a_refusal_or_a_held_job covers the ordering. It came from the exact-head review of ae11ab6, and failed there with 'watch' != 'refused'.
  • python3 tests/test_ci_owned_pool_rescue.py: 121 tests OK.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Stops the owned-pool rescue from cancelling a sibling job already running on a mini when another job of the same run is stuck waiting for a runner. The rescue previously cancelled the whole run and re-ran failed and cancelled jobs on Blacksmith, which killed in-flight mini jobs (e.g. a cmux-next swift test three minutes into its run). A stuck job now waits while any sibling runs on a persistent runner: if that job finishes within the watch, the rescue cancels only the stuck one; if the watch ends first, nothing is cancelled and the job stays queued for the mini that frees up. This applies to every watched workflow. Refusals and jobs held in glaeda's setup hook are still acted on first, so a sibling on a mini no longer hides them.

Written for commit 8370637. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes
    • Queued jobs in an owned runner pool are no longer rescued while another job in the same run is actively running outside setup. They remain queued until that job finishes; if the watch period ends first, they stay queued.
    • Refused jobs and jobs still in runner setup continue to be handled as before.

teamleaderleo and others added 2 commits October 1, 2026 13:00
Fails on main: the rescue cancels the whole cmux-next run, including a swift
test already running on a mini (#16463).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The rescue cancels the whole run to move a job that never got a runner, so a
sibling already running on a mini died with it and re-ran on Blacksmith
(#16463, #16468). The stuck job now waits until no job of the run is running
on a persistent runner; if the watch ends first it stays queued for the mini.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: bc15f6d0-e4fd-4768-825f-84e4e68d6538

📥 Commits

Reviewing files that changed from the base of the PR and between 15cf1ed and 8370637.

📒 Files selected for processing (2)
  • scripts/ci/owned_pool_rescue.py
  • tests/test_ci_owned_pool_rescue.py

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 2 remain after this review.


📝 Walkthrough

Walkthrough

Assessment now defers rescue of a stuck queued job while another job in the run is active on an owned runner outside setup. Tests cover the wait, subsequent rescue, and cases where refusals or setup-waiting jobs remain actionable.

Changes

Owned-pool rescue

Layer / File(s) Summary
Wait for active sibling jobs
scripts/ci/owned_pool_rescue.py, tests/test_ci_owned_pool_rescue.py
assess returns a waiting result while an owned-runner sibling is active outside setup. Tests cover rescue after the sibling completes, refusal precedence, and setup-waiting rescue.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Merge Risk: ⚪ Minimal · up to 83706

The change makes owned-pool rescue wait for a running sibling job before cancelling a stuck queued job. No merge-blocking risk was identified; the author documents the trade-off that some rescues may wait or remain queued.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 83706

The change protects active persistent-runner jobs without expanding cancellation privileges or weakening run eligibility checks. No introduced security finding was established. Recovery after watch timeout remains partly unverified because the cmux-next workflow that supplies the recovery marker was unavailable.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The inspected change affects rescue timing for already eligible workflow runs in the configured repository. Cancellation remains run-wide and can affect sibling jobs; the new hold reduces that exposure for queue starvation. No new cross-repository authority or sensitive sink was identified in the changed code.

Trust Boundaries and Controls

  • observed — Run eligibility continues to restrict workflow paths and event types, reject fork heads, and require the appropriate attempt. Before privileged cancellation or rerun, the controller checks attempt identity and applicable pull or branch revision identity. The PR changes assessment timing rather than these controls.

Resilience and Maintainability Implications

  • observed — Within a sweep, adoption is tracked by run ID and rerun attempt, read failures can be retried, and individual watch failures are contained. These existing controls limit repeated recovery actions and isolate failures, but do not establish distributed exclusivity between separate controllers.
🚥 Pre-merge checks | ✅ 23 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 14.29% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 7 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
Description check ⚠️ Warning The description provides detailed problem, behavior, trade-offs, and test results. However, it does not follow the repository template because it omits the required Summary, Testing, Changelog, Demo V… Restructure the description using the repository template. Add explicit Summary and Testing sections, include a Changelog line such as "none" for this internal CI change, address the Demo Video requirement or explain why it does not apply, …
✅ Passed checks (23 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Cmux Cloud Persistent Session And Early Input ✅ Passed PASS: The PR changes only scripts/ci/owned_pool_rescue.py and its tests. The changes adjust CI job rescue decisions and add CI tests. They do not change Cloud terminal creation, cmux-tui clients, ph…
Cmux Swift Actor Isolation ✅ Passed PASS: The pull request changes only Python CI rescue logic and its Python tests. It introduces no production Swift changes, so it cannot introduce or worsen Swift 6 actor-isolation mistakes.
Cmux Swift Blocking Runtime ✅ Passed PASS: The pull request changes only scripts/ci/owned_pool_rescue.py and its Python test file. The authoritative diff contains no Swift production changes, so the Swift blocking-runtime check does no…
Cmux Browser Automation Off-Main ✅ Passed PASS: The PR changes only scripts/ci/owned_pool_rescue.py and its CI tests. It introduces no browser.* socket command, WebKit/AppKit access, worker-router change, or browser policy test gap. The b…
Cmux Expensive Synchronous Load ✅ Passed PASS: The authoritative pull-request diff changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. It contains no Swift production changes or synchronous agent-history …
Cmux Cache Substitution Correctness ✅ Passed PASS: The pull request changes only Python code in scripts/ci/owned_pool_rescue.py and Python tests in tests/test_ci_owned_pool_rescue.py. It does not make a production Swift, TypeScript, or JavaS…
Cmux No Hacky Sleeps ✅ Passed PASS: The production diff changes assess() state decisions only. It adds no sleep, timer, polling primitive, or timing constant. Existing bounded watch() polling observes the real sibling job st…
Cmux Algorithmic Complexity ✅ Passed The production diff adds one linear pass over the run's jobs in scripts/ci/owned_pool_rescue.py:564 and moves the existing stuck-job name sort before the branch. It does not add a nested scan, per-t…
Cmux Swift Concurrency ✅ Passed The pull request changes only scripts/ci/owned_pool_rescue.py and its Python test file. It contains no changed Swift files and no added Swift concurrency patterns. The custom check is therefore not …
Cmux Swift @Concurrent ✅ Passed The authoritative PR diff changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. It introduces no Swift files or Swift declarations, so the @concurrent check is not…
Cmux Swift Package Boundaries ✅ Passed The review-scoped diff changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. It contains no production Swift changes, so the Swift package-boundaries check is not ap…
Cmux Swiftpm Lockfiles ✅ Passed The PR changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. It changes no Package.swift, Package.resolved, Xcode project, .gitignore, workflow, or dependency …
Cmux Swift Logging ✅ Passed PASS: The pull request changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. The diff contains no Swift files or added logging statements. The Swift logging rule is …
Cmux User-Facing Error Privacy ✅ Passed PASS: The diff changes only the internal CI rescue script and its tests. The new text describes queued jobs and persistent runners, and the script emits it only to the GitHub Actions log/step summary.…
Cmux Full Internationalization ✅ Passed PASS: The PR changes only scripts/ci/owned_pool_rescue.py and its tests. The production additions are CI control logic, developer comments, and operational Look.reason log text. They do not add Sw…
Cmux Swiftui State Layout ✅ Passed The pull request changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. The diff contains no Swift or SwiftUI files and no SwiftUI state or layout constructs. The Swi…
Cmux Architecture Rethink ✅ Passed PASS: The exact PR diff changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py, both Python files. It contains no Swift changes, so the Swift architectural-rethink cri…
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed PASS: The pull request changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. The diff adds no Swift code and no standalone cmux-owned window declarations, so the aux…
Cmux Source Artifacts ✅ Passed The diff changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. Both are tracked, hand-written Python source and test files. No logs, caches, build output, scratch di…
Cmux No Test Or Debug Seam In Production Source ✅ Passed The pull request changes only scripts/ci/owned_pool_rescue.py and tests/test_ci_owned_pool_rescue.py. It changes no Swift file under a production Sources/ path, so the custom check is not applic…
Title check ✅ Passed The title clearly describes the main change: preventing a stuck no-runner job from cancelling jobs already running on a mini.
Full details: Description check

Explanation

The description provides detailed problem, behavior, trade-offs, and test results. However, it does not follow the repository template because it omits the required Summary, Testing, Changelog, Demo Video, and Checklist sections.

Resolution

Restructure the description using the repository template. Add explicit Summary and Testing sections, include a Changelog line such as "none" for this internal CI change, address the Demo Video requirement or explain why it does not apply, and complete the Checklist with the review and testing status.

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 2
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

The hold for a stuck job returned before the refused and held-in-setup
checks, so a refusal next to a running mini job was never re-run and a held
job was not rescued at the watch's end. Judge those first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

CI fast guards failed on 8370637607 (https://github.com/manaflow-ai/cmux/actions/runs/36918849947). It does not block the merge; a red guard merged into main breaks it for every open PR.

Run canonical CMUX CI guard profile (red on main too, not this PR)

Main has failed this step since #16031 by @teamleaderleo (self-merged) (#16466). Merge main again once the fix lands there.

Agents: python3 scripts/ci/guard_attribution.py fix applies the mechanical fixes locally. This comment is updated in place on each push.

@teamleaderleo teamleaderleo changed the title ci: the owned-pool rescue never cancels a job running on a mini ci: a job stuck with no runner no longer cancels jobs running on a mini Oct 1, 2026
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Review: two independent subagent reviews.

ae11ab6: fix first. The new hold returned before the held-in-setup and refused checks. A refusal next to a running mini job was never re-run, and a held job wasn't rescued at the watch's end.

8370637 (exact head): ship. It checked:

  • The order in assess(): stuck with no mini sibling, then held, then refused, then stuck with a mini sibling, then watch.
  • That no other path rescues a stuck job while a sibling runs on a mini.
  • The PR body's claims, including the full re-run for picker CI and the trade-offs.
  • 121 tests OK.

Fixed:

  • The ordering, with test_a_sibling_on_a_mini_does_not_hide_a_refusal_or_a_held_job.
  • The PR body: picker CI's full re-run and the trade-offs are now stated.
  • The title, which overstated "never".

Left: the held-at-end and refused paths can still cancel a job running on a mini, by design. Neither is a job that never got a runner.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

CI failure attribution

CI failed on 8370637607 (run 36918850526 attempt 1): 1 unknown.

Job Verdict Why
guards / workflow-guard-tests / ci unknown no known signature; failed step: Propagate failed independent fast guard

Not re-run automatically: guards / workflow-guard-tests / ci is not a machine failure.

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 2 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="scripts/ci/owned_pool_rescue.py">

<violation number="1" location="scripts/ci/owned_pool_rescue.py:582">
P2: A stuck job held back by a mini-running sibling has no recovery path once the watch's deadline passes. The watch loop returns "stop: watch limit reached" with nothing cancelled, and later sweepers skip this run as soon as its marker passes SWEEP_MAX_AGE_SECONDS (150 min; sweep() drops marked runs whose `created` is older than `oldest`). If the pool stays busy past that point the shard waits indefinitely and is never rescued or re-run. The narrower new case this hold also enables: if the run reaches "completed" while the held job is still queued with no runner (a third sibling fails, or a newer push/concurrency cancels the run), the loop's `finished and not any(refused(job)...)` stop fires and the queued shard ends up cancelled without any re-run — an outcome the "assessed before this hold" carve-outs (refused, held-in-setup) don't cover, and one a plain budget rescue previously prevented. Consider letting a held stuck job survive the deadline (re-adopt in-progress runs regardless of marker age, or treat the hold like the refusal path and finish it within a grace window), so a sibling that finishes late still gets its shard rescued.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

if turned_away:
names = ", ".join(sorted(str(job.get("name") or job.get("id")) for job in turned_away))
return Look("refused", f"{names} refused by {job_pool(turned_away[0])} at job start or Xcode selection")
if stuck:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: A stuck job held back by a mini-running sibling has no recovery path once the watch's deadline passes. The watch loop returns "stop: watch limit reached" with nothing cancelled, and later sweepers skip this run as soon as its marker passes SWEEP_MAX_AGE_SECONDS (150 min; sweep() drops marked runs whose created is older than oldest). If the pool stays busy past that point the shard waits indefinitely and is never rescued or re-run. The narrower new case this hold also enables: if the run reaches "completed" while the held job is still queued with no runner (a third sibling fails, or a newer push/concurrency cancels the run), the loop's finished and not any(refused(job)...) stop fires and the queued shard ends up cancelled without any re-run — an outcome the "assessed before this hold" carve-outs (refused, held-in-setup) don't cover, and one a plain budget rescue previously prevented. Consider letting a held stuck job survive the deadline (re-adopt in-progress runs regardless of marker age, or treat the hold like the refusal path and finish it within a grace window), so a sibling that finishes late still gets its shard rescued.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At scripts/ci/owned_pool_rescue.py, line 582:

<comment>A stuck job held back by a mini-running sibling has no recovery path once the watch's deadline passes. The watch loop returns "stop: watch limit reached" with nothing cancelled, and later sweepers skip this run as soon as its marker passes SWEEP_MAX_AGE_SECONDS (150 min; sweep() drops marked runs whose `created` is older than `oldest`). If the pool stays busy past that point the shard waits indefinitely and is never rescued or re-run. The narrower new case this hold also enables: if the run reaches "completed" while the held job is still queued with no runner (a third sibling fails, or a newer push/concurrency cancels the run), the loop's `finished and not any(refused(job)...)` stop fires and the queued shard ends up cancelled without any re-run — an outcome the "assessed before this hold" carve-outs (refused, held-in-setup) don't cover, and one a plain budget rescue previously prevented. Consider letting a held stuck job survive the deadline (re-adopt in-progress runs regardless of marker age, or treat the hold like the refusal path and finish it within a grace window), so a sibling that finishes late still gets its shard rescued.</comment>

<file context>
@@ -566,6 +579,9 @@ def assess(jobs: Sequence[Mapping[str, Any]], *, now: dt.datetime, budget_second
     if turned_away:
         names = ", ".join(sorted(str(job.get("name") or job.get("id")) for job in turned_away))
         return Look("refused", f"{names} refused by {job_pool(turned_away[0])} at job start or Xcode selection")
+    if stuck:
+        return Look("watch", f"{names} queued on {job_pool(stuck[0])} with no runner, but "
+                             f"{len(on_mini)} job(s) of the run are running on a persistent runner", waiting=True)
</file context>

@teamleaderleo
teamleaderleo merged commit bbb73b6 into main Oct 1, 2026
48 of 53 checks passed
@teamleaderleo
teamleaderleo deleted the ci-rescue-spare-running-minis branch October 1, 2026 20:18
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Merge receipt for 8370637607, merged 2026-10-01 20:17:57 UTC

  • Not verified at merge: ci-status (failure), CI fast guards (failure), guards (10) (failure)
  • Verified: CI timing, Fast static checks, GhosttyKit release check, tests, Web complexity, web-validation
  • Skipped by policy: browser, Claude request, Claude wrapper regressions, Dogfood build #​${{ github.event.pull_request.number }}, full-suite-coverage, linux-preflight, macos, macOS admission gate, remote-daemon, suite-coverage, ui-tests, web, and 3 more
  • Full suite: runs on main after merge.

Labeled merged-unverified: if main breaks near this merge, look here first.

@github-actions github-actions Bot added the merged-unverified A judging check was not green at merge; see the merge receipt comment label Oct 1, 2026
rustybret pushed a commit to rustybret/bmux that referenced this pull request Oct 1, 2026
984179e fix: two more crossed-merge compile breaks on main (sidebar test seam, Cloud agent launch) (manaflow-ai#16490)
b1b876c fix(ios): prevent composer shortcut strip edge snapping (manaflow-ai#16128)
337861c ci: keep janitor sweeps green on refused cancellations (manaflow-ai#16488)
e0dc415 fix(ci): install Go before iOS App Store archive (manaflow-ai#16486)
c247a88 rename the duplicate node options resume test so ci's selector check passes (manaflow-ai#16481)
bbb73b6 ci: a job stuck with no runner no longer cancels jobs running on a mini (manaflow-ai#16479)
28ed45d fix(cli): repair main seed compile errors (manaflow-ai#16485)

# Conflicts:
#	.github/workflows/ios-appstore-upload.yml
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merged-unverified A judging check was not green at merge; see the merge receipt comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant