Repository navigation
[AMD] Deselect slow survives_real_pytest_runner from diffusion unit step - #42243
Open
michaelzhang-ai wants to merge 4 commits into
Open
michaelzhang-ai wants to merge 4 commits into
michaelzhang-ai wants to merge 4 commits into
Conversation
The AMD diffusion unit step had accumulated 27 ROCm-only failures and ran close to its 30 minute timeout: - Perf-policy tests expected AssertionError, but PerformanceValidator only warns on HIP. Patch current_platform.is_hip in their fixtures, as the other strict-path perf tests already do. - ModelOpt FP8 loader tests compared against raw e4m3fn checkpoints; on gfx94x the loader normalizes to e4m3fnuz with 2x scales. Assert the platform layout instead. - Flux3Fp8RowwiseLinear hard-coded e4m3fn, which torch._scaled_mm rejects on gfx94x. Normalize the weight to e4m3fnuz and quantize activations with the platform fp8 dtype/max. - Gate the bit-exact CUDA fusion tests in test_qwen_image21_cuda on NVIDIA, since the Q/K-norm+RoPE and KV-pack kernels refuse HIP. - Deselect survives_real_pytest_runner on AMD: it is platform-agnostic, runs in the CUDA lane, and takes ~20 min on MI300.
michaelzhang-ai
requested review from
AgainstEntropy,
BBuf,
Fridge003,
HaiShaw,
Kangyan-Zhou,
bingxche,
ispobock,
kevin-mii,
merrymercy,
mickqian,
niehen6174,
ping1jing2 and
yichiche
as code owners
October 2, 2026 16:10
This was referenced Oct 2, 2026
# Conflicts: # python/sglang/multimodal_gen/runtime/models/dits/flux3.py # python/sglang/multimodal_gen/test/unit/test_modelopt_fp8_layerwise_offload_load.py # python/sglang/multimodal_gen/test/unit/test_transformer_quant.py
Keep the fix AMD-only: revert the shared test edits and deselect the failing tests in the AMD workflow instead. They all keep running in the CUDA diffusion unit lane.
# Conflicts: # .github/workflows/pr-test-amd.yml
This was referenced Oct 5, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
multimodal-gen-test-1-gpu-amd"Run diffusion unit tests" has been failing on main since at least 2026-09-09. Example: run 37006548834 ended with27 failed, 3477 passed ... in 1734.15s (0:28:54)against a 30 minute step timeout.Most of that is now fixed on main:
current_platform.is_hipin the perf-policy tests.test_qwen_image21_cuda.pyoff NVIDIA.One problem is left.
test_performance_failure_survives_real_pytest_runner(#39206) has 9 cases, and each one runs 7 subprocess pytest attempts that re-import sglang + aiter. On MI300 that takes about 20 minutes, and the test lands in a single shard. In #42391's own AMD run (37181100988), that shard's unit step took 25m20s of its 30 minute budget. The other three shards took 2.5–5 minutes. A slow runner or a few more unit tests in that shard will push it over the timeout.Modifications
Only
.github/workflows/pr-test-amd.ymlchanges. No shared code or tests change, so NVIDIA is unaffected.The AMD diffusion unit step now also deselects
survives_real_pytest_runner. It still runs in the CUDA diffusion unit lane (base-b-test-diffusion-unit-1-gpu-h100). It exercises the platform-independent perf-failure plumbing, and it already patchesis_hipto False, so MI300 adds no extra coverage.Accuracy Tests
N/A. This PR changes CI configuration only.
Testing
-kexpression to pytest--collect-onlyover stubs of all 3132test_*names inmultimodal_gen/test/unit. Onlytest_performance_failure_survives_real_pytest_runneris newly deselected (ltx2_vae_channels_lastmatches nothing underunit/).Follow-up (not in this PR)
On MI300,
test_bf16_fusions_match_eager_prefill_and_cached_stepsdiffers from eager by 0.457 max abs on 100% of elements, which is more than bit-exactness noise. It is now skipped on ROCm, but the Triton fused residual/gate path on ROCm is worth a separate look.CI States
Latest PR Test (Base): ✅ Run #37203476973
Latest PR Test (Extra): ❌ Run #37203476810
Latest PR Test (AMD ROCm 10): ❌ Run #37203476966