Repository navigation
ci: increase gpu timeout budgets - #1721
Conversation
Signed-off-by: key4ng <rukeyang@gmail.com>
|
Note Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported. |
📝 WalkthroughWalkthroughUpdates GitHub Actions workflow timeouts: standalone ChangesCI Timeout Configuration Updates
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Signed-off-by: key4ng <rukeyang@gmail.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 940d292cc2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| timeout: 36 | ||
| test_timeout: 28 |
There was a problem hiding this comment.
Apply the intended SGLang chat timeout budget
For the e2e-1gpu-chat (sglang) matrix entry, these two values are forwarded directly to the reusable workflow as the job timeout and the Run E2E tests step timeout (.github/workflows/e2e-gpu-job.yml:56 and :131). The change notes say this shard needs 40 minutes overall and 32 minutes for pytest after the slow SGLang startup plus rerun case, but this leaves it at 36/28, so any run in that 28–32 minute pytest window will still be cancelled before completing. Please set this entry to the intended 40/32 budget.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/pr-test-rust.yml:
- Around line 498-499: The workflow currently sets timeout: 36 and test_timeout:
28 but the PR's objective is to raise the 1GPU SGLang chat lane to 40/32; update
the two keys so that timeout is set to 40 and test_timeout is set to 32 (i.e.,
replace the existing timeout: 36 and test_timeout: 28 entries with timeout: 40
and test_timeout: 32).
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 07fbf155-5c69-4e0a-8bde-ae2e7b3a5b42
📒 Files selected for processing (1)
.github/workflows/pr-test-rust.yml
|
The failed CI doesn't related to this fix |
Description
Problem
Two GPU CI jobs are timing out even though the underlying workers are not crashing:
Run E2E testsstep timeout after slow SGLang large-model startup and a rerun consumed most of the shard budget.Solution
Increase the affected CI timeout budgets so these jobs have enough room for slow tokenizer/report startup and large-model reloads.
Changes
GENAI_BENCH_TEST_TIMEOUT=480for the benchmark job.Test Plan
ruby -e 'require "yaml"; YAML.load_file(".github/workflows/pr-test-rust.yml"); puts "yaml ok"'pre-commit run --files .github/workflows/pr-test-rust.ymlChecklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit