Skip to content

test(diffusion): cover residual dispatch and Qwen CUDA fallbacks - #42256

Open
Tokha233 wants to merge 2 commits into
sgl-project:mainfrom
Tokha233:fix/diffusion-residual-hip-fallback
Open

Tokha233 wants to merge 2 commits into
sgl-project:mainfrom
Tokha233:fix/diffusion-residual-hip-fallback

Conversation

@Tokha233

@Tokha233 Tokha233 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

ROCm tensors also report is_cuda, but the diffusion residual fast path contains NVIDIA PTX. Upstream #42391 already merged the HIP eligibility guard and simplified the native dispatch. This PR now adds regression coverage on top of that implementation; it changes no production kernel.

Changes

  • Exercise HIP rejection and the exact eager result for FP16/BF16/FP32 with full, row, token, and transposed-row gate layouts. A sentinel fails the test if the PTX custom op is attempted.
  • Run the real CUDA residual fast path for FP16/BF16 with those four layouts, including the new transposed Triton path, and compare with eager at atol=0, rtol=0.
  • Run the Qwen Q/K norm + RoPE/KV packing regression with native kernels available and deliberately unavailable, checking exact model output, prefix-cache ownership, and CUDA graph replay.
  • Preserve upstream's NVIDIA-only Qwen suite and newer LoRA-format tests. Remove references to the deleted failed-runtime cache and old packing eligibility API.

Validation

Synced with upstream bab04cd7913ff50d1bf1a677a2d10e5015dc4c3b.

RTX 5090 D v2, Python 3.12, PyTorch 2.13.0+cu130: 44 passed across the full residual-dispatch and Qwen CUDA test files. CUDA kernels, small-model forwards, prefix caches, and graph replay ran on the real GPU. HIP eligibility was simulated on that host; native AMD execution is still required from CI.

python -m pytest -q \
  python/sglang/multimodal_gen/test/unit/test_residual_gate_add_dispatch.py \
  python/sglang/multimodal_gen/test/unit/test_qwen_image21_cuda.py

Changed-file pre-commit and git diff --check passed. Complete source content was verified against the local Git index and working-tree changes before the remote run. Existing tolerances were not relaxed. No new inference performance claim. Codex assisted with implementation and testing.


CI States

Latest PR Test (Base): ❌ Run #37233260867
Latest PR Test (Extra): ❌ Run #37233260677
Latest PR Test (AMD ROCm 10): ❌ Run #37233260792

@Tokha233 Tokha233 changed the title [diffusion] Guard PTX residual kernels on HIP and exercise Qwen fallbacks test(diffusion): cover residual dispatch and Qwen CUDA fallbacks Oct 4, 2026
@Tokha233

Tokha233 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Resolved the conflict in d0db0cba, keeping #42391's native dispatch implementation. Removed obsolete test references and added transposed-row residual coverage. The remaining diff is tests only. Fresh real 5090 D v2 run: 44 passed, covering residual CUDA kernels, Qwen forward/cache equality and CUDA graph replay, including deliberately disabled CUDA fusion. HIP eligibility is simulated locally, not claimed as native AMD validation.

Local pre-commit and GitHub lint passed. Logs/diffs: https://github.com/Tokha233/ComfyUI-H3-SpeedKit/tree/8791760/evidence/pr-maintenance-1005 . Hosted GPU tests currently stop at the run-ci authorization gate.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion jit-kernel

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant