Repository navigation
test: wait for the SSH cleanup policy bound after a restored-attach signal - #14305
Conversation
…ignal restoredAttachSignalTerminatesForegroundAuthenticationProcessTree gave the restored attach supervisor 3s to exit after SIGINT. The foreground authentication cleanup policy allows a 2s discovery window plus a bounded force pass, and its own deadline regression caps total cleanup at 15s. Main CI run 36036314182 (shard 2/7) hit the 3s bound: the child's TERM handler ran (the signal-log expectation passed) but the post-TERM process-table snapshots were still running on a loaded runner. The next three main runs passed the same test. No product code in the cleanup path changed since #13615. Wait for the policy's 15s bound, like directSignalTerminatesPersistentAttachAuthenticationProcessTree does, and report the elapsed time when it still fails. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
All contributors have signed the CLA ✍️ ✅ |
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 10 included reviews per hour; 2 remain after this review. 📝 WalkthroughWalkthroughThe restored-attach signal test now allows up to 15 seconds for process exit after SIGINT. It records elapsed time for failure messages. The existing SIGKILL fallback and termination-status check remain. ChangesSSH authentication marker cleanup test
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~5 minutes Change: Other Merge Risk: ⚪ Minimal · up to The cleanup test allows more time for process exit while retaining its termination checks. No merge-blocking risk is apparent. 🚥 Pre-merge checks | ✅ 24 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (24 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The minis' runners are LaunchAgents in the logged-in user's Aqua session, and #14305's changed suites passed inside compile admission on cmuxs-mac-mini-5. Job 107862186541's exit 65 was that PR's own test (CMUXCLICodexUnavailableAdmissionTests), not the environment. Owned placement now goes admission, app-host shards by index, tests-build-and-lag, then the light jobs; the shards queue longest on Blacksmith. CI_PR_POOL_OWNED_GUI=0 keeps GUI jobs off the minis again, and only then does a persistent pick move the changed suites out of admission (new output owned_gui). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
3d5b648 test: give the Cloud Desktop fixture a routable window for pane drops (manaflow-ai#14304) 9d53af4 ci: give a refused owned job one more try on the fleet before Blacksmith (manaflow-ai#14312) 379091b ci: let PR runs overflow to macOS 15 at 4 queued jobs deeper, not 12 (manaflow-ai#14319) df6efb5 test(cloud): key the refresh URL protocol stub per request, not by address (manaflow-ai#14239) 51bd322 ci: build only the CLI product for CLI-only changes (manaflow-ai#14212) 0345a5c ci: make a changed-suites run prove a known-failure fix (manaflow-ai#14307) 370b7f6 ci: fix the owned build state save step's argument count (manaflow-ai#14309) 319adff test: wait for the SSH cleanup policy bound after a restored-attach signal (manaflow-ai#14305) 0cc5be3 fix(portal): flush the coalesced live-resize pass on its first hop (manaflow-ai#14297) 5887891 test(minimal-mode): measure the toggle only after setup stops re-rendering (manaflow-ai#14298) 5fbcc48 ci: charge newer PR runs what they took on the owned pool, not a guess (manaflow-ai#14300) # Conflicts: # .github/workflows/ci-macos.yml
The minis' runners are LaunchAgents in the logged-in user's Aqua session, and #14305's changed suites passed inside compile admission on cmuxs-mac-mini-5. Job 107862186541's exit 65 was that PR's own test (CMUXCLICodexUnavailableAdmissionTests), not the environment. Owned placement now goes admission, app-host shards by index, tests-build-and-lag, then the light jobs; the shards queue longest on Blacksmith. CI_PR_POOL_OWNED_GUI=0 keeps GUI jobs off the minis again, and only then does a persistent pick move the changed suites out of admission (new output owned_gui). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…14318) * ci: place each PR macOS job on a free owned mini, overflow the rest The picker took an owned pool only when a run's whole peak was free, so a full suite with 2 of 11 minis busy went to Blacksmith entirely and queued there while 9 minis sat idle. With CI_PR_POOL_OWNED_SPLIT=1, a run that does not fit takes the owned pool with the most machines free, and a new owned_jobs output names the jobs that fit: compile admission first, then the light jobs (cli-product, cli-pipe, remote-daemon, claude-wrapper). Every other attempt-1 job takes retry_runner, the Blacksmith pool on the lane's Xcode. The marker's <jobs> is now the owned machines the run holds, so the janitor's committed count covers only the jobs placed there. GUI jobs (app-host shards, tests-build-and-lag) never take an owned pool: the minis have no console session. A persistent pick turns off unit_in_admission, so the changed suites run on a Blacksmith shard. Splitting a run is sound because both sides run Xcode 26.6 build 17F113 (minis cmux15, cmuxs-mac-mini-5, cmux13s and Blacksmith 6vcpu/12vcpu macOS 26 on 2026-09-24). The product only moves from the mini to Blacksmith, and check_xcode refuses a product from a newer Xcode. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: let owned minis take GUI jobs, behind CI_PR_POOL_OWNED_GUI The minis' runners are LaunchAgents in the logged-in user's Aqua session, and #14305's changed suites passed inside compile admission on cmuxs-mac-mini-5. Job 107862186541's exit 65 was that PR's own test (CMUXCLICodexUnavailableAdmissionTests), not the environment. Owned placement now goes admission, app-host shards by index, tests-build-and-lag, then the light jobs; the shards queue longest on Blacksmith. CI_PR_POOL_OWNED_GUI=0 keeps GUI jobs off the minis again, and only then does a persistent pick move the changed suites out of admission (new output owned_gui). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: always move the changed suites out of an owned compile admission glaeda's runner hook gives compile admission the compile token, never the gui token, so app-host suites run inside it on a mini could collide with a GUI shard on the same machine. Every persistent pick now runs them on shard 8 instead, and the picker plans for that shard. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: free a run's owned machines once its owned jobs finish With per-job placement, a run's owned jobs (admission, light lanes) can finish long before its Blacksmith shards, but the janitor charged the marker's peak until the whole run completed, so idle minis read as busy and new runs overflowed. A marked run whose owned jobs all completed now holds nothing; before its first owned job exists, the marker still reserves its peak. The product-consumer guard also requires tests-build-and-lag to test its own ' lag ' key, so it cannot follow admission's placement by copy-paste. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: normalize the owned key in the transport route test; hold a run's minis until its peak finishes The CLI product and app-host shard routes differ only by their owned_jobs key, so the transport test compares them with the key normalized. The janitor now releases a marked run's owned machines only once as many owned jobs as its peak have completed: shard jobs exist only after admission. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
restoredAttachSignalTerminatesForegroundAuthenticationProcessTreegave the restored attach supervisor 3 s to exit after SIGINT. The foreground authentication cleanup it exercises allows a 2 s discovery window plus a bounded force pass, andSSHForegroundAuthenticationRetryPolicyTeststolerates up to 15 s of cleanup for a much larger tree.On a loaded runner the 3 s bound is too tight. Main runs 35984682064 and 36036314182 failed at
exited(4.8 s and 5.0 s total). In both, the child's TERM handler ran, so the signal-log expectation passed and the tree was being torn down. Across the 13 most recent main runs that reached shard 2 the test passed the other 11 times, at 2.1 to 4.0 s total. No product code in the cleanup path changed since #13615, so this is a flake, not a regression.The test now waits up to that 15 s tolerance, like
directSignalTerminatesPersistentAttachAuthenticationProcessTree, and reports the elapsed time if it still fails. The exit status and TERM-handler assertions are unchanged.Validation: not built on this Mac. This diff routes
SSHForegroundAuthenticationMarkerCleanupTestsinto the app-host changed-suites lane.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by cubic
Fixes a flaky SSH hidden attached-cleanup test by extending its exit deadline from 3 s to the cleanup policy's 15 s bound.
Written for commit 315c2f0. Summary will update on new commits.
Summary by CodeRabbit