Repository navigation
[XPU] Add CustomOp.forward_xpu fallback to forward_native - #33320
Closed
dayanandav wants to merge 7 commits into
Closed
dayanandav wants to merge 7 commits into
dayanandav wants to merge 7 commits into
Conversation
dispatch_forward() routes the xpu platform to self.forward_xpu, but the base class never defined it. Since dispatch_forward() runs in __init__, any subclass without its own override raises AttributeError at construction, breaking model loading on Intel GPUs: AttributeError: 'LayerNorm' object has no attribute 'forward_xpu' Route xpu to forward_native, matching the existing forward_tpu/npu/oot fallbacks. No dispatch branch is modified, so CUDA/ROCm/NPU/MUSA/CPU behavior is unchanged.
dayanandav
requested review from
AgainstEntropy,
BBuf,
HaiShaw,
mickqian,
ping1jing2 and
yichiche
as code owners
August 3, 2026 03:01
Contributor
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
Contributor
Author
|
@siju-samuel requested for review. |
Contributor
Author
|
@mickqian request for review. |
Contributor
Author
Contributor
Author
|
@mingfeima requested for review. |
siju-samuel
approved these changes
Aug 7, 2026
Contributor
Author
|
@ping1jing2 @yichiche @AgainstEntropy @BBuf requested for review. |
Contributor
Author
|
@mingfeima help to review and merge PR. |
4 of 5 tasks
Contributor
|
/tag-run-ci-label |
Amrutha-M05
added a commit
to Amrutha-M05/sglang
that referenced
this pull request
Sep 2, 2026
Register intel_xpu_b60 as a consistency-threshold platform and add XPU perf baselines for the 2-GPU fsdp-inference case. Depends on sgl-project#33320 for CustomOp.forward_xpu.
ANSHUMAN87
reviewed
Sep 3, 2026
| # PyTorch-native implementation. | ||
| return self.forward_native(*args, **kwargs) | ||
|
|
||
| def forward_xpu(self, *args, **kwargs) -> Any: |
Contributor
There was a problem hiding this comment.
IMO we should not add forward_xpu in the CustomeOp base class. The design is the feature owner will inherit this class and implement the missing APIs. If XPU wants to support, it will be added in the inherited class. Better to add in respective failed instance. Thanks!
dayanandav
marked this pull request as draft
September 15, 2026 12:48
dayanandav
marked this pull request as ready for review
September 16, 2026 04:27
Bring XPU memory reporting and component residency in line with the other backends: - get_available_gpu_memory() now reports mem_get_info() rather than total - memory_allocated(), so the caching allocator's reserve is no longer counted as free. - XpuComponentOffloadStrategy parks the preferred component after warmup only while its weights plus a stage's headroom still fit in driver-free memory, and finish_request() returns the allocator's idle reserve at the request boundary. Both thresholds are XPU_HEADROOM_FRACTION (0.05) x device total memory, so no absolute byte count goes stale across device sizes. Other backends are unaffected: the existing strategies and release paths are byte-identical, the new path is reached by strategy selection behind one is_xpu() guard, and the XPU helpers call torch.xpu directly. Applies on top of the CustomOp.forward_xpu fallback earlier in this PR, which is what puts the sglang transformer on the XPU path. Verified with the wan2_1_t2v_1.3b diffusion server case at its component offload defaults, which generates its output; that case's perf threshold is derived from CUDA baselines and is untouched here. Unit coverage in test_component_residency.py. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Contributor
Author
|
Same solution applied along with #36825 , so closing duplicate PR. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
dispatch_forward() routes the xpu platform to self.forward_xpu, but the base class never defined it. Since dispatch_forward() runs in init, any subclass without its own override raises AttributeError at construction, breaking model loading on Intel GPUs:
Route xpu to forward_native, matching the existing forward_tpu/npu/oot fallbacks. No dispatch branch is modified, so CUDA/ROCm/NPU/MUSA/CPU behavior is unchanged.
Motivation
Fixes #33292 .
CustomOp.dispatch_forward() maps the detected platform to a forward_ method and caches it in init. Every backend has a base-class fallback — hip/cpu/musa → forward_cuda, tpu/oot/npu → forward_native — so a subclass only needs forward_native to work everywhere. The xpu branch (custom_op.py:83-84) was the sole exception: it returns self.forward_xpu, which was never defined.
Because dispatch_forward() is called from init (custom_op.py:27), this fails at construction, so it breaks model loading rather than inference:
Five shipped subclasses relied on the missing method — LayerNorm (layernorm.py:329), RMSNormNoWeight (layernorm.py:316), GeluAndMul (activation.py:75), FusedLayerNormScaleShiftGateSelect01 (fused_scale_shift_gate.py:20), FusedResidualLayerNormScaleShiftGateSelect01 (fused_scale_shift_gate.py:85). All implement forward_native. Eight other subclasses define their own forward_xpu, which is why this went unnoticed. Present since #12484 (7bc1dae).
Modifications
Just introduced forward_xpu in python/sglang/multimodal_gen/runtime/layers/custom_op.py. Verified after the patch that all six branches dispatch_forward() can return resolve to a base-class method.
Accuracy Tests
This PR does not touch kernels or model forward code, so outputs are unchanged on any platform that was already binding correctly (CUDA/NPU single-GPU).
Speed Tests and Profiling
No kernel or scheduling changes, so there is no expected throughput change on platforms that were already binding correctly.
CI States
Latest PR Test (Base): Not run yet⚠️ Not enabled -- add
Latest PR Test (Extra):
run-ci-extralabel to opt in.Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.