ci(reborn): retry crates.io network failures in the closure (CARGO_NET_RETRY) - #5115
Conversation
…T_RETRY) The 64-crate closure runs ~64 jobs that each resolve+download the dependency tree from crates.io in parallel. A transient registry hiccup (seen twice: 'download of once_cell_polyfill failed' HTTP2 framing, and 'agent-client-protocol ... curl [55] SSL_ERROR_SYSCALL') hits many jobs at once and reddens the whole run even though nothing is wrong with the code. Bump Cargo's network retry from the default 3 to 10 and fetch git deps via the git CLI, so transient download failures self-heal instead of failing the job. The heavier fix (build-once `nextest archive` + shard the run, eliminating 64x dependency resolution) remains a separate follow-up. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
Note Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughSummary by CodeRabbit
WalkthroughTwo environment variables — ChangesReborn CI Network Resilience
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~2 minutes Poem
🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Comment |
|
🚅 Deployed to the ironclaw-pr-5115 environment in ironclaw-ci-preview
|
…r-crate caches (#5118) The closure's crate-tests matrix used a per-crate cache key (`reborn-tests-<crate>`), producing ~60 caches that each bundle the cargo registry + a target dir (~0.4-0.7 GB each), plus 4 per-partition root caches (~1 GB each). That's ~30+ GB competing for GitHub's ~10 GB per-repo cache LRU, so the caches evict each other almost as fast as they're written: a live snapshot showed only 10 of 64 crates had a cache present, summing to 18 GB and actively evicting. The result is "No cache found" on most jobs -> the crates.io registry is re-downloaded from cold every run, and 64 parallel cold downloads amplify transient registry flakes (the SSL_ERROR_SYSCALL / HTTP2 reds we saw). Switch both matrices to `shared-key`: - crate-tests -> `shared-key: reborn-tests-crates` (one entry for all crate jobs; the registry + shared-dep build is downloaded/compiled once and resident, not ~60x). - root partitions -> `shared-key: reborn-tests-root` (all 4 partitions compile the identical root build, so one cache is strictly better than 4 copies). This drops the reborn-tests cache footprint from ~30+ GB / 68 entries to ~2 entries that comfortably fit the LRU, so the registry stays resident and stops being re-downloaded. Complements CARGO_NET_RETRY (#5115), which remains the safety net for the now-rare cold download. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Problem
The 64-crate closure (#5110) runs ~64 jobs that each resolve+download the dependency tree from crates.io in parallel. A transient registry SSL/HTTP2 hiccup hits many jobs at once and reddens the whole run — seen twice already:
download of once_cell_polyfill failed(HTTP2 framing)agent-client-protocol ... curl [55] OpenSSL SSL_ERROR_SYSCALL(61 jobs red in one run)Neither is a code failure — they pass on re-run.
Fix
Add to
reborn-tests.ymlenv:CARGO_NET_RETRY: "10"— Cargo retries network failures (default 3) far more aggressively, so transient download errors self-heal within the job.CARGO_NET_GIT_FETCH_WITH_CLI: "true"— use the git CLI for git deps (more robust than libgit2 under flaky networks).2-line env change, applies to all reborn-tests jobs. Zero behavior change otherwise.
Follow-up (not this PR)
The real scaling fix is build-once
cargo nextest archive→ shard the run, which eliminates the 64× parallel dependency resolution entirely. Tracked for the bake phase.Automated agent-authored.