Skip to content

[CI/Build] Restrict ERNIE fused RoPE tests to NVIDIA CUDA - #7771

Merged
yenuo26 merged 2 commits into
vllm-project:mainfrom
andyluo7:fix/ernie-rope-nvidia-test-guard
Sep 19, 2026
Merged

yenuo26 merged 2 commits into
vllm-project:mainfrom
andyluo7:fix/ernie-rope-nvidia-test-guard

Conversation

@andyluo7

@andyluo7 andyluo7 commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

What broke

The AMD Diffusion · Model Test selected the ERNIE-Image fused RoPE tests and failed four cases because try_fused_qk_rotary_emb(...) correctly returned None on ROCm.

Observed in AMD builds #12244 and #12243, on different MI300X workers:

4 failed, 56 passed, 24 skipped, 3203 deselected

Reproduction

On a ROCm worker:

pytest -sv tests/diffusion/models/ \
  -m "core_model and cuda and not (cards_2 or cards_3 or cards_4 or cards_5 or cards_6 or cards_7 or cards_8)" \
  --run-level core_model

Root cause

PyTorch exposes ROCm devices through the torch.cuda namespace, so torch.cuda.is_available() is true on ROCm. The tests therefore ran even though the implementation intentionally enables this fused kernel only when current_omni_platform.is_cuda() is true.

The same tests execute successfully on NVIDIA CUDA in build #15561.

Fix

Use the project platform abstraction in all five GPU-dependent skip guards. This keeps the NVIDIA fused-kernel coverage on CUDA, skips it on ROCm, and leaves the platform-independent failed-key cache test enabled.

Declare the actual CI resources with hardware_test: the five NVIDIA-only tests require one CUDA L4, while the cache-bound test retains one-card coverage on CUDA L4 and ROCm MI325.

The production kernel guard is unchanged.

Test Plan

Local validation:

pre-commit run --files tests/diffusion/models/ernie_image/test_ernie_image_fused_rope.py
All applicable hooks passed.

python3 -m compileall -q tests/diffusion/models/ernie_image/test_ernie_image_fused_rope.py
Passed.

Static hardware-marker contract check
Passed.

Focused pytest cannot run on this macOS host because the available Python environments do not have the project vllm/PyTorch dependencies installed. AMD and CUDA runtime validation is delegated to CI.

vLLM Version: CI image version for the Buildkite runs cited below

vLLM-Omni Commit: 94822606be6d77dcdc2e1b432cb8eb122b1a5821

Test Result

  • AMD #12282, mi300_1: Diffusion · Model Test: passed in 6m05s (50 passed, 34 skipped, 3203 deselected). All ten NVIDIA-only cases skipped with NVIDIA CUDA required; the platform-independent cache-bound test ran and passed.
  • CUDA #15600, Simple · Diffusion Test: passed in 17m18s (5979 passed, 41 skipped, 555 deselected).
  • Local pre-commit: passed.
  • Local syntax and hardware-marker contract checks: passed.

AI assistance: Used Codex to investigate the cross-platform failure, make the focused test-only change, and draft this description. I reviewed the one-file diff and the cited CI evidence.

Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review: I verified that this PR changes only the five NVIDIA-only skip guards in the ERNIE-Image fused RoPE test module. The implementation already rejects ROCm through current_omni_platform.is_cuda(), so the test now uses the same platform contract instead of torch.cuda.is_available(), which is true on ROCm. I compared the identical AMD failures in Buildkite #12243 and #12244 with the passing NVIDIA CUDA run #15561. All applicable pre-commit hooks and syntax checks pass; local pytest was unavailable because this macOS environment has no PyTorch installation.

@andyluo7 andyluo7 added the ready label to trigger buildkite CI label Sep 18, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26 @NickCao

Routing: @yenuo26 via CI owner, CODEOWNERS; @NickCao via CODEOWNERS

@andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

andyluo7 commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator Author

Updated exact-head CI evidence for 94822606be6d77dcdc2e1b432cb8eb122b1a5821 after the hardware-marker review change:

  • AMD #12282 mi300_1: Diffusion · Model Test: passed in 6m05s (50 passed, 34 skipped, 3203 deselected). All ten NVIDIA-only ERNIE cases were collected and skipped with NVIDIA CUDA required; test_failed_runtime_key_cache_is_bounded remained collected and passed.
  • CUDA #15600 Simple · Diffusion Test: passed in 17m18s (5979 passed, 41 skipped, 555 deselected).

This validates both the ROCm routing fix and the new explicit hardware_test resource metadata at the current PR head.

from vllm_omni.diffusion.models.ernie_image.ernie_image_transformer import _apply_rotary_emb
from vllm_omni.platforms import current_omni_platform

pytestmark = [pytest.mark.core_model, pytest.mark.cuda, pytest.mark.diffusion]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please modify pytest.mark.cuda to hardware_test, to mark the actual machine type and card count it uses.

@andyluo7 andyluo7 Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated in 9482260. I removed the module-level pytest.mark.cuda and added explicit hardware_test metadata: the five NVIDIA-only fused-RoPE tests use cuda/L4 with one card, while the platform-independent cache-bound test keeps one-card coverage on both cuda/L4 and rocm/MI325. I kept the current_omni_platform.is_cuda() guards because the shared AMD diffusion lane still selects cuda-marked tests. Pre-commit, compileall, and the static decorator-contract check pass. Fresh exact-head validation is also green: AMD #12282 Diffusion · Model Test passed with 50 passed, 34 skipped, 3203 deselected, and CUDA #15600 Simple · Diffusion Test passed with 5979 passed, 41 skipped, 555 deselected.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the guidance. The requested hardware_test update is now in 9482260, and fresh exact-head AMD #12282 plus CUDA #15600 validation passed. Could you please take another look when you have a chance?

Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 18, 2026
@andyluo7
andyluo7 requested a review from yenuo26 September 18, 2026 15:41

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yenuo26
yenuo26 merged commit 232dbc8 into vllm-project:main Sep 19, 2026
8 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…ct#7771)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants