Skip to content

[CI/Build] Continue Whisper GPU capacity fix on current main - #6936

Closed
andyluo7 wants to merge 7 commits into
vllm-project:mainfrom
andyluo7:fix/whisper-gpu-capacity
Closed

andyluo7 wants to merge 7 commits into
vllm-project:mainfrom
andyluo7:fix/whisper-gpu-capacity

Conversation

@andyluo7

@andyluo7 andyluo7 commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

This is a current-main continuation of #5675, preserving Akshat Nayak's five original commits and authorship so its approved, backend-generic Whisper placement fix can receive valid AMD L3 coverage. If #5675 is rebased and validated directly, this continuation can be closed in favor of the original PR.

Recent AMD L3 builds in #6926 show the Qwen3-TTS serving requests completing, followed by the 40-minute Base or 75-minute CustomVoice timeout while quality validation uses CPU Whisper. The carried fix uses the existing platform API to select the highest-index accelerator with at least 16 GiB free, and keeps the CPU fallback when no device has enough capacity. It does not hardcode ROCm, AMD, or a model name.

Current-main follow-ups in this continuation are deliberately narrow:

  • a typing update to the existing test fake required by the current mypy hook;
  • CUDA ready/merge source dependencies so a change to tests/helpers/media.py actually selects the Qwen3-TTS integration lanes. Without this routing, a green CUDA pipeline would not exercise the changed helper.

Related failure evidence: AMD CI #11286, #11289, #11292, and #11296.

Test Plan

vLLM Version: Current project dependency; no production dependency change.

vLLM-Omni Commit: Base e51fe6ec1b9a9a0e14bb1fdb296d61b6593b93c6; head f436186792a644b8a80a138593101c9049e66203.

Test Result

  • pytest tests/helpers/tests/test_media.py: 16 passed
  • pytest tests/buildkite/test_upload_pipeline.py tests/buildkite/test_skip_ci.py: 46 passed
  • Exact PR-diff rendering selects Qwen3-TTS CustomVoice in CUDA ready and both Qwen3-TTS CustomVoice and Base in CUDA merge.
  • pre-commit run --files .buildkite/cuda/test-ready.yml .buildkite/cuda/test-merge.yml tests/helpers/media.py tests/helpers/tests/test_media.py
  • python3 -m compileall -q tests/helpers/media.py tests/helpers/tests/test_media.py
  • git diff --check origin/main...HEAD

All local/static checks pass. Exact-head AMD L3 validation is requested to verify that Base and CustomVoice complete quality validation without changing their existing full-sample checks. The routed CUDA TTS jobs should also run to guard the shared helper path.

akshatvishu and others added 6 commits September 1, 2026 21:03
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7
andyluo7 marked this pull request as ready for review September 2, 2026 04:46
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@andyluo7

andyluo7 commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator Author

Self-review completed against base e51fe6ec.

  • Verified that tests/helpers/media.py matches approved PR [Perf][CI] Use Whisper on GPU when memory permits #5675 exactly; its five original commits and Akshat Nayak's authorship/signoffs are preserved.
  • The placement policy is backend-generic: reverse first-fit through the existing platform API, select an accelerator only with at least 16 GiB free, otherwise keep CPU fallback. There is no AMD, ROCm, or model-name branch.
  • The worker still has max_workers=1, and the Whisper model remains cached inside that worker; this avoids multiplying GPU-resident Whisper instances.
  • The only continuation-specific change is a narrow typing update to the test fake required by current main's mypy hook.
  • Reviewed the full diff for broad exception handling, new Any, hot-path copies/synchronization, dead branches, and new examples; none were introduced.
  • Local validation passed: all 16 tests in tests/helpers/tests/test_media.py, complete changed-file pre-commit, compileall, and git diff --check origin/main...HEAD.

@akshatvishu, this PR is intentionally a credited current-main continuation for CI validation; if you prefer to rebase #5675 directly, I am happy to close this one in its favor. @yenuo26, could you please add the label needed for exact-head merge-test coverage? I will verify the AMD Base and CustomVoice jobs reach their pytest summaries and also check the relevant CUDA result before treating this as resolved.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26

@andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7
andyluo7 requested a review from congw729 as a code owner September 2, 2026 04:58
@andyluo7

andyluo7 commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator Author

Current-head self-review addendum for f4361867:

  • Added tests/helpers/media.py as a source dependency of the CUDA Qwen3-TTS CustomVoice ready/merge steps and Base merge step. This is CI routing only; it does not alter CUDA runtime behavior.
  • Simulated an exact PR diff through upload_pipeline.py: CUDA ready retains TTS · Qwen3-TTS CustomVoice Test; CUDA merge retains both TTS · Qwen3-TTS CustomVoice Test and TTS · Qwen3-TTS Base Test. Without these dependency entries, those integration jobs would be filtered out for this helper-only change.
  • Re-ran the Buildkite uploader/skip test suites: 46 passed.
  • Re-ran changed-file pre-commit, including Buildkite schema validation: all hooks passed.

The request for an exact-head merge-test run remains current; adding ready as well would exercise the shorter CUDA CustomVoice path.

@andyluo7

andyluo7 commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

@vllm-omni-review-bot please run an automated review of the current head f436186. Please focus on the platform-generic free-memory selection, CPU fallback, worker/model reuse, and whether the added source dependencies exercise the shared helper on CUDA without changing CUDA runtime configuration. Also flag any backend/model hardcoding or missed failure paths.

@andyluo7

andyluo7 commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator Author

@yenuo26 one CI-routing clarification for the requested hardware proof: could you add amd-test, ready, and merge-test? The AMD PR bootstrap defaults to the ready/L2 YAML, so please also trigger an AMD build with DEBUG_TEST_YAML=merge,ready (or a separate DEBUG_TEST_YAML=merge run) on head f436186. That is needed to verify AMD Base and CustomVoice through their pytest summaries plus the corresponding CUDA ready/merge jobs on the same commit.

@akshatvishu

akshatvishu commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Thanks @andyluo7 ! I rebased #5675 onto current main and cherry-picked your two follow-up commits. #5675 is now at d9e0aac. Waiting for CI validation!

@andyluo7

andyluo7 commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks @akshatvishu. I verified that #5675 at d9e0aac8 is rebased onto current main and contains patch-identical versions of both continuation commits from this PR, with the original authorship and signoffs preserved. Closing this duplicate in favor of #5675 so review and hardware validation stay on the original credited PR.

@andyluo7 andyluo7 closed this Sep 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants