Repository navigation
ci: re-run by cause: host faults to Blacksmith, code failures back to the minis - #15045
Conversation
|
Warning Review limit reachedNext included review available in 2 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (16)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
d1d06f8 to
29a724d
Compare
… the minis github-actions[bot] re-runs a PR run only after a host fault on a mini: the owned-pool rescue after a refusal or stuck queue, and the failure attribution when every failed job is a machine failure. Those re-runs keep #15035's path (retry_runner on Blacksmith). Anyone else's re-run follows a code or test failure and goes back to the owned label (a failed-jobs re-run) or the normal pick without queueing (a full re-run). If a mini fails that re-run, the attribution re-runs it as the bot onto Blacksmith, so it never loops. 7 days: 135/138 bot re-runs followed a host fault; 139/217 other re-runs a code failure only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lacksmith Review follow-ups: the rescue sweeper lists in-progress CI runs a person re-ran (no marker of their own) and watches them on any attempt; the picker and janitor read the newest marker up to a re-run's attempt; the person case applies to pull requests only (main's dispatch has no attribution re-run or watch). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
29a724d to
808eae7
Compare
|
Merge receipt for |
|
Toolbox g1 🔔 reviewed the cause-based retry routing. The bot-only host-fault path is bounded, person-triggered reruns preserve attempt-1 ownership, and the classifier/rescue/janitor changes avoid retry loops. Full required CI, macOS, remote-daemon, guard, and classifier checks are green. Enabling ordinary squash auto-merge subject to human approval. |
a64d59b tools: ui-lab renders view code in seconds; wire-app-sources.py (manaflow-ai#15049) 4e03ed2 fix(events): harden durable replay recovery (manaflow-ai#15054) ac51546 Settings: native terminal theme gallery (manaflow-ai#14996) 867e7a0 Add native Ghostty option rows to Settings > Terminal (manaflow-ai#15005) 7f97b0d ui-tests: wait for static preflight when a reused compile skips the gate (manaflow-ai#15051) c708e0c Add a chat view for the terminal's agent session (Claude Code, Codex) (manaflow-ai#14965) b762a3d ci: the picker fetches kept bases' trees, not just checks their commits (manaflow-ai#15053) 2570eed docs: refresh and trim contributor build guidance (manaflow-ai#15050) 20019d3 ci: re-run by cause: host faults to Blacksmith, code failures back to the minis (manaflow-ai#15045) 36ee3e9 Add Warn Before Closing Workspace setting (manaflow-ai#14979) 4df2317 CI: run changed UI test classes in PRs, keep UI runs off Blacksmith, probe the GUI session (manaflow-ai#14964) 9efe05e Owned-pool sweeper: page the marker listing back to the runs it adopts (manaflow-ai#15033) 6361554 fix(events): restore durable replay across restarts (manaflow-ai#15030) 0bc5145 ci: place side lanes on the light minis one per idle side runner (manaflow-ai#15047) 9d4e92b ci: the E2E rule's queue-round reason names the owned pools the run may take (manaflow-ai#15044) 9e6e216 Dogfood the app from CI with JSON tours (manaflow-ai#14928) fd3dcf6 ci: retry the picker's kept-base fetch and record how it went (manaflow-ai#15040) 3a64e0e Reload the Ghostty config when its files change, and show config errors (manaflow-ai#14859) f412b05 test: hit-test the browser portal tab strip with its own click (manaflow-ai#15031) 64d5235 test: route the reopen-last-closed shortcut through the test's own window (manaflow-ai#15036) 7037079 ci: take the gui token in the E2E test job's step, not at job start (manaflow-ai#15037) # Conflicts: # .github/workflows/ci-macos.yml # .github/workflows/ci.yml # .github/workflows/remote-daemon.yml # .github/workflows/test-e2e.yml
Leo picked "split by cause". This builds on #15035 and changes only who goes where.
github-actions[bot]: owned-pool rescue, failure attribution (classify_failures.py, all-machine verdict)retry_runneron Blacksmith, as #15035 doesWhy the actor carries the cause. A failed-jobs re-run does not re-run
changes, soruns-oncan only see attempt 1's outputs plusgithub.run_attemptandgithub.triggering_actor. The classifiers already exist: the rescue re-runs only refusals and stuck queues, and the attribution re-runs only when every failed job is a machine failure. So a bot re-run is a host-fault re-run by construction.Measured over 7 days (ci.yml, same-repo, runs with 2 attempts):
Bound. If a mini fails a person's re-run, the attribution now re-runs that attempt as the bot, onto Blacksmith. Before, it re-ran attempt 1 only. It never re-runs its own re-run, so there is no loop.
Changes:
runs-on(ci-macos.yml, ci.yml claude-wrapper, remote-daemon PR lane):github.run_attempt > 1becomesgithub.run_attempt > 1 && github.triggering_actor == 'github-actions[bot]'. swift-package also admits a person's retry.host_fault_retry()). Any other retry picks every pool, withqueue_rounds=0.may_hold_owned_pool: they count a person's re-run as possibly owned.classify_failures.rerun_decision: also re-runs a person-started attempt.Tests: I ran the targeted files locally (picker, janitor, rescue, classify, self-hosted guard .sh, release-sdk .sh, seed routing subset, change-areas routing subset). The full suite runs on CI.
Supersedes #15032.
🤖 Generated with Claude Code
Summary by cubic
Routes CI re-runs by cause:
github-actions[bot]re-runs a pull request run only after a host fault on a mini (refusal, lost runner, timeout, restore failure) and sends it to Blacksmith'sretry_runner, while anyone else's re-run follows a code or test failure and goes back to the owned pool attempt 1 placed the job on. This changes the old behavior, which sent every re-run to Blacksmith regardless of who started it.Changes
runs-ongates the Blacksmith retry lane ongithub.run_attempt > 1 && github.triggering_actor == 'github-actions[bot]'; the picker'shost_fault_retry()drops the owned pools only for bot retries, and a person's full re-run picks like attempt 1 without queueing.classify_failures.pynow re-runs a machine failure on a person-started attempt to Blacksmith; it never re-runs the bot's own re-run, so the flow cannot loop.Measured over 7 days the classifier matches the intent: 135 of 138 bot re-runs followed a host fault, and 139 of 217 other re-runs a code failure only.
Written for commit 808eae7. Summary will update on new commits.