Skip to content

[AMD] Deselect slow survives_real_pytest_runner from diffusion unit step - #42243

Open
michaelzhang-ai wants to merge 4 commits into
sgl-project:mainfrom
michaelzhang-ai:fix/amd-mmgen-unit-suite
Open

michaelzhang-ai wants to merge 4 commits into
sgl-project:mainfrom
michaelzhang-ai:fix/amd-mmgen-unit-suite

Conversation

@michaelzhang-ai

@michaelzhang-ai michaelzhang-ai commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

multimodal-gen-test-1-gpu-amd "Run diffusion unit tests" has been failing on main since at least 2026-09-09. Example: run 37006548834 ended with 27 failed, 3477 passed ... in 1734.15s (0:28:54) against a 30 minute step timeout.

Most of that is now fixed on main:

One problem is left. test_performance_failure_survives_real_pytest_runner (#39206) has 9 cases, and each one runs 7 subprocess pytest attempts that re-import sglang + aiter. On MI300 that takes about 20 minutes, and the test lands in a single shard. In #42391's own AMD run (37181100988), that shard's unit step took 25m20s of its 30 minute budget. The other three shards took 2.5–5 minutes. A slow runner or a few more unit tests in that shard will push it over the timeout.

Modifications

Only .github/workflows/pr-test-amd.yml changes. No shared code or tests change, so NVIDIA is unaffected.

The AMD diffusion unit step now also deselects survives_real_pytest_runner. It still runs in the CUDA diffusion unit lane (base-b-test-diffusion-unit-1-gpu-h100). It exercises the platform-independent perf-failure plumbing, and it already patches is_hip to False, so MI300 adds no extra coverage.

-              -k "not ltx2_vae_channels_last" \
+              -k "not ltx2_vae_channels_last and not survives_real_pytest_runner" \

Accuracy Tests

N/A. This PR changes CI configuration only.

Testing

  • The workflow YAML parses.
  • I fed the new -k expression to pytest --collect-only over stubs of all 3132 test_* names in multimodal_gen/test/unit. Only test_performance_failure_survives_real_pytest_runner is newly deselected (ltx2_vae_channels_last matches nothing under unit/).
  • MI300 timing comes from this PR's AMD CI. The affected shard should drop from ~25 to ~5 minutes.

Follow-up (not in this PR)

On MI300, test_bf16_fusions_match_eager_prefill_and_cached_steps differs from eager by 0.457 max abs on 100% of elements, which is more than bit-exactness noise. It is now skipped on ROCm, but the Triton fused residual/gate path on ROCm is worth a separate look.


CI States

Latest PR Test (Base): ✅ Run #37203476973
Latest PR Test (Extra): ❌ Run #37203476810
Latest PR Test (AMD ROCm 10): ❌ Run #37203476966

The AMD diffusion unit step had accumulated 27 ROCm-only failures and ran
close to its 30 minute timeout:

- Perf-policy tests expected AssertionError, but PerformanceValidator only
  warns on HIP. Patch current_platform.is_hip in their fixtures, as the
  other strict-path perf tests already do.
- ModelOpt FP8 loader tests compared against raw e4m3fn checkpoints; on
  gfx94x the loader normalizes to e4m3fnuz with 2x scales. Assert the
  platform layout instead.
- Flux3Fp8RowwiseLinear hard-coded e4m3fn, which torch._scaled_mm rejects
  on gfx94x. Normalize the weight to e4m3fnuz and quantize activations
  with the platform fp8 dtype/max.
- Gate the bit-exact CUDA fusion tests in test_qwen_image21_cuda on
  NVIDIA, since the Q/K-norm+RoPE and KV-pack kernels refuse HIP.
- Deselect survives_real_pytest_runner on AMD: it is platform-agnostic,
  runs in the CUDA lane, and takes ~20 min on MI300.
# Conflicts:
#	python/sglang/multimodal_gen/runtime/models/dits/flux3.py
#	python/sglang/multimodal_gen/test/unit/test_modelopt_fp8_layerwise_offload_load.py
#	python/sglang/multimodal_gen/test/unit/test_transformer_quant.py
Keep the fix AMD-only: revert the shared test edits and deselect the
failing tests in the AMD workflow instead. They all keep running in the
CUDA diffusion unit lane.
@michaelzhang-ai michaelzhang-ai added the run-ci CI: run the baseline test suite on this PR label Oct 4, 2026
# Conflicts:
#	.github/workflows/pr-test-amd.yml

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd diffusion SGLang Diffusion quant LLM Quantization run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant