Repository navigation
[XPU] Add diffusion server test config for Intel XPU B60 - #35267
Closed
Amrutha-M05 wants to merge 1 commit into
Closed
Amrutha-M05 wants to merge 1 commit into
Amrutha-M05 wants to merge 1 commit into
Conversation
Amrutha-M05
force-pushed
the
xpu-mmgen-test-config
branch
from
August 19, 2026 04:17
66a2f03 to
9741bbc
Compare
siju-samuel
approved these changes
Aug 19, 2026
Amrutha-M05
marked this pull request as ready for review
August 19, 2026 05:31
Amrutha-M05
requested review from
AgainstEntropy,
BBuf,
HaiShaw,
mickqian,
ping1jing2 and
yichiche
as code owners
August 19, 2026 05:31
Contributor
|
/tag-run-ci-label |
Register intel_xpu_b60 as a consistency-threshold platform and add XPU perf baselines for the 2-GPU fsdp-inference case. Depends on sgl-project#33320 for CustomOp.forward_xpu.
Amrutha-M05
force-pushed
the
xpu-mmgen-test-config
branch
from
September 2, 2026 07:59
1fead71 to
7758979
Compare
5 tasks
airMeng
pushed a commit
to airMeng/sglang
that referenced
this pull request
Sep 22, 2026
airMeng
pushed a commit
to airMeng/sglang
that referenced
this pull request
Sep 24, 2026
Contributor
Author
|
Closing, as #40664 merged this PR's changes into main. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Let the
multimodal_gendiffusion server tests run on Intel XPU with correct expectations.The diffusion server tests compare generated images against committed ground-truth images and check stage timings against committed perf baselines. Both sets of reference values are CUDA-derived, and neither is reachable from XPU today:
get_consistency_platform()has no XPU branch, so an XPU run silently falls through to theh100thresholds.BASELINE_CONFIGcomposes per-platform perf overrides for Ascend and MUSA, but there is no XPU entry.The image mismatch is not a correctness bug. XPU seeds its initial latent noise from a different RNG stream than CUDA, so the pixels cannot match bit-for-bit: the image stays semantically identical (CLIP remains high) while
ssim/psnr/mean_abs_diffdegrade. Comparing XPU output against CUDA thresholds therefore fails on a difference that carries no signal.This PR adds XPU as a first-class platform in both registries, following the existing Ascend and MUSA pattern.
Modifications
Purely additive: 4 files, 56 insertions, 0 deletions.
test/test_utils.py— registerintel_xpu_b60inCONSISTENCY_THRESHOLD_FILE_BY_PLATFORM, add theintelxpub60alias, and add thecurrent_platform.is_xpu()branch toget_consistency_platform().test/server/consistency_thresholds/intel_xpu_b60.json— XPU consistency thresholds forqwen_image_t2i_2_gpusandfsdp-inference.test/server/xpu/perf_baselines_xpu.json— XPU perf baselines forfsdp-inference, measured on Intel Arc Pro B60 (22.71 GiB per card).test/server/testcase_configs.py— chain the XPU baseline file intoBASELINE_CONFIG, alongside the existing Ascend and MUSA.update()calls.No CUDA path changes: both new files are read only when the resolved platform is XPU, and the two registry additions are new keys.
Thresholds
fsdp-inferenceqwen_image_t2i_2_gpusOnly these two cases were measured on XPU hardware. Every other case still inherits from
h100.json; that is called out in the_commentfield of the threshold file so the remaining cases can be triaged the same way rather than being assumed correct.Accuracy Tests
Verified on 2x Intel Arc Pro B60 (22.7 GiB per card):
Model
Tongyi-MAI/Z-Image-Turbo. Passes end to end in ~157 s, including the consistency check against the thresholds above. Measured CLIP 0.9779 (passes); SSIM 0.6609 / PSNR 12.0434 / mean-abs-diff 33.3489, which fail the H100 thresholds and pass the XPU ones. XPU output is deterministic across 4 runs (identical md5).The perf baselines are single-host measurements, not an averaged multi-run distribution, so they are starting reference values rather than tight regression bounds.
Checklist
cc: @siju-samuel @rbabukv
CI States
Latest PR Test (Base): ❌ Run #33606366411
Latest PR Test (Extra): ❌ Run #33606364529
Latest PR Test (AMD ROCm 7.2): ❌ Run #33606366341