ci: bound the SwiftPM scratch holder and cache scratch sizes - #15366
Conversation
…ory's size The holder that keeps a job's shared lock on its scratch directory exited only when the runner killed it at job end. A runner that died mid-job left it holding the lock until reboot, so that directory could never be pruned or evicted. It now exits after 65 minutes (signal.alarm), past the swift-package-tests job's 60-minute timeout. tree_stats walked every file of up to 24 GiB of scratch on every link. The lock file, which link now touches, gives a directory's last use, and its size is cached in <fingerprint>.size, measured again only once the directory was used after it and no job holds it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
Warning Review limit reachedNext included review available in 58 seconds. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
Dogfood build of cmux DEV pr-15366-7504716c.app The link opens this exact commit in the cmux dev menu bar app. The build starts on each push and the page waits until it is ready; a newer push replaces it. It signs in against production, so Cloud or backend changes still need a tagged build with a development backend. |
|
Found 1 test failure on Blacksmith runners: Failure
|
CI failure attributionCI passes on Written by |
…ocks) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Merge receipt for |
Resolve the scratch-age overlap using main’s complete implementation and tests from #15366. The cache repair now has no remaining delta from main.
0e298fb ci: wait for the product's canonical root instead of compiling beside it (manaflow-ai#15379) 3088273 ci: UI test runs adopt compile admission's product, skip the re-upload, and report progress (manaflow-ai#15331) b681e7e Keep a pending banner quiet once its pane is focused (manaflow-ai#15357) 03a2f6e Record that cloud_vm_sessions.attachment_count is cumulative (manaflow-ai#15321) 48258b4 fix(iroh-v2): check the team socket cap before opening the session (manaflow-ai#15340) 2638d56 Agent activity reorder follow-ups: group on-top check, search, subtitle (manaflow-ai#15362) 9ed83fd Dogfood journey: record whether a paused Cloud machine is asleep (manaflow-ai#15293) 7171ea8 Add app.tabBarVisibility to hide the pane tab bar when a pane has one tab (manaflow-ai#15294) 8743ec8 test: stop Computer Use onboarding tests waiting out the helper status deadline (manaflow-ai#15329) 6e4f1da ci: drain the snapshot's owned queue by what the machines finished since (manaflow-ai#15374) 9373164 ci: queue a pull request's admission for a root runner when Blacksmith's wait is longer (manaflow-ai#15376) 634a155 test: expect injected pane attention accent (manaflow-ai#15370) cd030e9 Keep a named Cloud machine's prompt name instead of flipping to its slug (manaflow-ai#15288) 24ee0ee Exit 1 when cmux terminal screen wait times out (manaflow-ai#15282) 1b857ac test: cover a live Codex turn owner keeping its turn on SessionStart (manaflow-ai#13588) 56ec600 PR media: prune media of long-closed pull requests (manaflow-ai#15364) 4898cde ci: bound the SwiftPM scratch holder and cache scratch sizes (manaflow-ai#15366) # Conflicts: # .github/workflows/ci-guards.yml # .github/workflows/ci.yml # .github/workflows/test-e2e.yml
A follow-up to #14804, from its review.
Bound the holder.
owned_spm_scratch.py linkstarts a small process that holds a shared flock on the job's SwiftPM scratch directory. Until now it exited only when the runner killed a job's leftover processes at job end. If the runner process died mid-job, the holder kept the lock until the mini rebooted, and that directory could never be pruned or evicted.holdnow exits after 65 minutes (signal.alarm), just past the swift-package-tests job's 60-minute timeout.Stop walking the whole scratch on every link.
tree_statslstat'ed every file of up to 24 GiB of scratch to size and rank each directory. Now:linktouches the directory's lock file, and that time is its last use, which decides prune order.<fingerprint>.sizebeside the directory. It is measured again only when the directory was used after the size was written.Validation.
tests/test_ci_owned_spm_scratch.py(12 tests) passes. The 3 new tests cover:Changelog
none
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by cubic
Follow-up to #14804. Bounds the SwiftPM scratch holder so a runner that dies mid-job can't leave a scratch directory locked until the mini reboots, and caches each directory's size so pruning no longer walks the whole scratch on every link.
linktouches the lock file as the directory's last use, and sizes are cached in<fingerprint>.size, re-measured only when a later job used the directory and no job holds it; a time that ties on coarse clocks also re-measures.Written for commit 7504716. Summary will update on new commits.