Repository navigation
Conversation
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
Self-review completed against base
@akshatvishu, this PR is intentionally a credited current-main continuation for CI validation; if you prefer to rebase #5675 directly, I am happy to close this one in its favor. @yenuo26, could you please add the label needed for exact-head |
|
This PR was classified as CI work. CI owner: @yenuo26 @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Current-head self-review addendum for
The request for an exact-head |
|
@vllm-omni-review-bot please run an automated review of the current head f436186. Please focus on the platform-generic free-memory selection, CPU fallback, worker/model reuse, and whether the added source dependencies exercise the shared helper on CUDA without changing CUDA runtime configuration. Also flag any backend/model hardcoding or missed failure paths. |
|
@yenuo26 one CI-routing clarification for the requested hardware proof: could you add amd-test, ready, and merge-test? The AMD PR bootstrap defaults to the ready/L2 YAML, so please also trigger an AMD build with DEBUG_TEST_YAML=merge,ready (or a separate DEBUG_TEST_YAML=merge run) on head f436186. That is needed to verify AMD Base and CustomVoice through their pytest summaries plus the corresponding CUDA ready/merge jobs on the same commit. |
|
Thanks @akshatvishu. I verified that #5675 at |
Purpose
This is a current-main continuation of #5675, preserving Akshat Nayak's five original commits and authorship so its approved, backend-generic Whisper placement fix can receive valid AMD L3 coverage. If #5675 is rebased and validated directly, this continuation can be closed in favor of the original PR.
Recent AMD L3 builds in #6926 show the Qwen3-TTS serving requests completing, followed by the 40-minute Base or 75-minute CustomVoice timeout while quality validation uses CPU Whisper. The carried fix uses the existing platform API to select the highest-index accelerator with at least 16 GiB free, and keeps the CPU fallback when no device has enough capacity. It does not hardcode ROCm, AMD, or a model name.
Current-main follow-ups in this continuation are deliberately narrow:
tests/helpers/media.pyactually selects the Qwen3-TTS integration lanes. Without this routing, a green CUDA pipeline would not exercise the changed helper.Related failure evidence: AMD CI #11286, #11289, #11292, and #11296.
Test Plan
vLLM Version: Current project dependency; no production dependency change.
vLLM-Omni Commit: Base
e51fe6ec1b9a9a0e14bb1fdb296d61b6593b93c6; headf436186792a644b8a80a138593101c9049e66203.Test Result
pytest tests/helpers/tests/test_media.py: 16 passedpytest tests/buildkite/test_upload_pipeline.py tests/buildkite/test_skip_ci.py: 46 passedpre-commit run --files .buildkite/cuda/test-ready.yml .buildkite/cuda/test-merge.yml tests/helpers/media.py tests/helpers/tests/test_media.pypython3 -m compileall -q tests/helpers/media.py tests/helpers/tests/test_media.pygit diff --check origin/main...HEADAll local/static checks pass. Exact-head AMD L3 validation is requested to verify that Base and CustomVoice complete quality validation without changing their existing full-sample checks. The routed CUDA TTS jobs should also run to guard the shared helper path.