[CI] harden aiter-whl download (digest-mismatch fallback) + raise fd limit for ATOM editable build - #1660
Conversation
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
There was a problem hiding this comment.
Pull request overview
This PR hardens the atom-sglang-test accuracy workflow by making the aiter-whl artifact download resilient to intermittent digest-mismatch failures, and by raising the open-file descriptor limit before the editable ATOM build (cargo) inside the SGLang overlay image build.
Changes:
- Make
actions/download-artifact@v8foraiter-whlnon-fatal and add a retrying REST/API-based fallback downloader. - Increase
nofilesoft limit (ulimit -n) beforepip install -e .during the Docker build to avoidToo many open files (os error 24).
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| set -uo pipefail | ||
| # actions/download-artifact@v8 intermittently fails the aiter-whl | ||
| # (~600 MB) transfer with a digest-mismatch after 5 internal retries. | ||
| # If the primary download did not land the wheel, re-fetch the same | ||
| # run-scoped artifact directly via the API with backoff. | ||
| if ls aiter-whl/amd_aiter*.whl >/dev/null 2>&1; then | ||
| echo "aiter wheel present from primary download"; exit 0 | ||
| fi | ||
| echo "Primary download-artifact failed or empty; falling back to direct API download" | ||
| mkdir -p aiter-whl | ||
| aid=$(gh api "repos/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}/artifacts" \ | ||
| --jq '.artifacts[] | select(.name=="aiter-whl") | .id' | head -1) | ||
| if [ -z "${aid:-}" ]; then echo "ERROR: no aiter-whl artifact found for run ${GITHUB_RUN_ID}"; exit 1; fi | ||
| for i in 1 2 3 4 5; do | ||
| if gh api "repos/${GITHUB_REPOSITORY}/actions/artifacts/${aid}/zip" > aiter-whl.zip 2>/dev/null \ | ||
| && unzip -o -q aiter-whl.zip -d aiter-whl \ | ||
| && ls aiter-whl/amd_aiter*.whl >/dev/null 2>&1; then | ||
| echo "Fallback download succeeded on attempt ${i}"; rm -f aiter-whl.zip; exit 0 | ||
| fi | ||
| echo "Fallback attempt ${i} failed; retrying in $((i*20))s"; sleep $((i*20)) | ||
| done | ||
| echo "ERROR: aiter wheel download failed after primary + 5 fallback retries"; exit 1 |
| - name: Download aiter wheel | ||
| id: dl_aiter | ||
| uses: actions/download-artifact@v8 | ||
| continue-on-error: true |
|
@okakarpa please review before merge. This hardens two PR-agnostic CI failures in
Static checks are green (Black/Ruff/actionlint/unit). The GPU accuracy jobs are path-skipped on this workflow-only PR, so the download-fallback/ulimit are exercised only on a full accuracy re-run. Cross-PR evidence + root causes: https://amd.atlassian.net/wiki/spaces/~pensun/pages/1809672378 |
|
@sunway513 Thanks for the enhancement. This is very helpful. We already have a unified script for handling aiter wheel downloads, and I am optimizing it in this commit: 0fd8cb7. I think I can adopt your optimization in that script after this lands, so we only need to maintain the logic in one place. |
faef936 to
6d5701b
Compare
| git clone "${GITHUB_REPO_URL}" /app/ATOM && \ | ||
| cd /app/ATOM && \ | ||
| git checkout "${GITHUB_COMMIT_SHA}" && \ | ||
| ulimit -n 1048576 || true && \ |
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
| echo "=== nofile BEFORE: soft=$(ulimit -Sn) hard=$(ulimit -Hn)" && \ | ||
| (ulimit -n 1048576 2>/dev/null || ulimit -n 65536 2>/dev/null || true) && \ | ||
| echo "=== nofile AFTER: soft=$(ulimit -Sn) hard=$(ulimit -Hn)" && \ |
What
Harden two intermittent, PR-agnostic failures in the
atom-sglang-testaccuracy jobs (.github/workflows/atom-sglang-test.yaml):Download aiter wheelflakiness.actions/download-artifact@v8intermittently fails the ~600 MBaiter-whltransfer with adigest-mismatchafter its 5 internal retries (observed onDeepSeek-R1-FP4-V2 TP8 MTP3, ~37 min then hard fail). The step is nowcontinue-on-error, followed by a resilient fallback that re-fetches the same run-scopedaiter-whlartifact via the REST API with 5 backoff retries. A flaky artifact transfer no longer red-lights the accuracy job.ATOM editable build fd exhaustion.
pip install -e .(which builds the Rust extension via cargo) fails withcargo build ... Too many open files (os error 24)when the runner's defaultnofilesoft limit is low. Raise it in the same shell before the build.Why
Both failures are unrelated to any model/PR change and reproduce on
main:unified_kv_rope/ cudagraph and download failures on in-flight PRs (e.g. adding profiling context #477) trace to these infra issues, not the PR content.main'sInstall ATOM and dependenciesstep directly.Not covered here (needs infra / image owner)
The
DeepSeek-V4-Proaccuracy job separately fails with'DeepseekV4Attention' object has no attribute 'unified_kv_rope'during SGLang cudagraph capture.unified_kv_ropeis set inatom/model_ops/attentions/deepseek_v4_attn.pyand read inatom/models/deepseek_v4.py— both consistent on currentmain, so this is a stale/mismatched ATOM in the SGLang base image. The fix is to rebuild the SGLang base image from currentmainand pin the base-image + aiter-wheel to a coherent set (rather than "latest main"). Tracked alongside ROCm/aiter#4304.Test plan
Re-run the
atom-sglang-testaccuracy matrix; the aiter-whl download should survive a digest-mismatch via the fallback, and the ATOM editable build should no longer hit the fd limit. (CI validation needs a maintainer run.)