Repository navigation
[CI/Build] Restrict ERNIE fused RoPE tests to NVIDIA CUDA - #7771
Conversation
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Self-review: I verified that this PR changes only the five NVIDIA-only skip guards in the ERNIE-Image fused RoPE test module. The implementation already rejects ROCm through |
|
This PR was classified as CI work. Routing: @yenuo26 via CI owner, CODEOWNERS; @NickCao via CODEOWNERS @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
Updated exact-head CI evidence for
This validates both the ROCm routing fix and the new explicit |
| from vllm_omni.diffusion.models.ernie_image.ernie_image_transformer import _apply_rotary_emb | ||
| from vllm_omni.platforms import current_omni_platform | ||
|
|
||
| pytestmark = [pytest.mark.core_model, pytest.mark.cuda, pytest.mark.diffusion] |
There was a problem hiding this comment.
Please modify pytest.mark.cuda to hardware_test, to mark the actual machine type and card count it uses.
There was a problem hiding this comment.
Updated in 9482260. I removed the module-level pytest.mark.cuda and added explicit hardware_test metadata: the five NVIDIA-only fused-RoPE tests use cuda/L4 with one card, while the platform-independent cache-bound test keeps one-card coverage on both cuda/L4 and rocm/MI325. I kept the current_omni_platform.is_cuda() guards because the shared AMD diffusion lane still selects cuda-marked tests. Pre-commit, compileall, and the static decorator-contract check pass. Fresh exact-head validation is also green: AMD #12282 Diffusion · Model Test passed with 50 passed, 34 skipped, 3203 deselected, and CUDA #15600 Simple · Diffusion Test passed with 5979 passed, 41 skipped, 555 deselected.
There was a problem hiding this comment.
Thanks for the guidance. The requested hardware_test update is now in 9482260, and fresh exact-head AMD #12282 plus CUDA #15600 validation passed. Could you please take another look when you have a chance?
Signed-off-by: andyluo7 <andy.luo@amd.com>
…ct#7771) Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…ct#7771) Signed-off-by: andyluo7 <andy.luo@amd.com>
Purpose
What broke
The AMD
Diffusion · Model Testselected the ERNIE-Image fused RoPE tests and failed four cases becausetry_fused_qk_rotary_emb(...)correctly returnedNoneon ROCm.Observed in AMD builds #12244 and #12243, on different MI300X workers:
Reproduction
On a ROCm worker:
pytest -sv tests/diffusion/models/ \ -m "core_model and cuda and not (cards_2 or cards_3 or cards_4 or cards_5 or cards_6 or cards_7 or cards_8)" \ --run-level core_modelRoot cause
PyTorch exposes ROCm devices through the
torch.cudanamespace, sotorch.cuda.is_available()is true on ROCm. The tests therefore ran even though the implementation intentionally enables this fused kernel only whencurrent_omni_platform.is_cuda()is true.The same tests execute successfully on NVIDIA CUDA in build #15561.
Fix
Use the project platform abstraction in all five GPU-dependent skip guards. This keeps the NVIDIA fused-kernel coverage on CUDA, skips it on ROCm, and leaves the platform-independent failed-key cache test enabled.
Declare the actual CI resources with
hardware_test: the five NVIDIA-only tests require one CUDA L4, while the cache-bound test retains one-card coverage on CUDA L4 and ROCm MI325.The production kernel guard is unchanged.
Test Plan
Local validation:
Focused pytest cannot run on this macOS host because the available Python environments do not have the project
vllm/PyTorch dependencies installed. AMD and CUDA runtime validation is delegated to CI.vLLM Version: CI image version for the Buildkite runs cited below
vLLM-Omni Commit:
94822606be6d77dcdc2e1b432cb8eb122b1a5821Test Result
mi300_1: Diffusion · Model Test: passed in 6m05s (50 passed, 34 skipped, 3203 deselected). All ten NVIDIA-only cases skipped withNVIDIA CUDA required; the platform-independent cache-bound test ran and passed.Simple · Diffusion Test: passed in 17m18s (5979 passed, 41 skipped, 555 deselected).AI assistance: Used Codex to investigate the cross-platform failure, make the focused test-only change, and draft this description. I reviewed the one-file diff and the cited CI evidence.