Skip to content

[XPU] Add CustomOp.forward_xpu fallback to forward_native - #33320

Closed
dayanandav wants to merge 7 commits into
sgl-project:mainfrom
dayanandav:bug_33292
Closed

dayanandav wants to merge 7 commits into
sgl-project:mainfrom
dayanandav:bug_33292

Conversation

@dayanandav

@dayanandav dayanandav commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

dispatch_forward() routes the xpu platform to self.forward_xpu, but the base class never defined it. Since dispatch_forward() runs in init, any subclass without its own override raises AttributeError at construction, breaking model loading on Intel GPUs:

AttributeError: 'LayerNorm' object has no attribute 'forward_xpu'

Route xpu to forward_native, matching the existing forward_tpu/npu/oot fallbacks. No dispatch branch is modified, so CUDA/ROCm/NPU/MUSA/CPU behavior is unchanged.

Motivation

Fixes #33292 .
CustomOp.dispatch_forward() maps the detected platform to a forward_ method and caches it in init. Every backend has a base-class fallback — hip/cpu/musa → forward_cuda, tpu/oot/npu → forward_native — so a subclass only needs forward_native to work everywhere. The xpu branch (custom_op.py:83-84) was the sole exception: it returns self.forward_xpu, which was never defined.

Because dispatch_forward() is called from init (custom_op.py:27), this fails at construction, so it breaks model loading rather than inference:

from sglang.multimodal_gen.runtime.layers.layernorm import LayerNorm
LayerNorm(hidden_size=64)
AttributeError: 'LayerNorm' object has no attribute 'forward_xpu'

Five shipped subclasses relied on the missing method — LayerNorm (layernorm.py:329), RMSNormNoWeight (layernorm.py:316), GeluAndMul (activation.py:75), FusedLayerNormScaleShiftGateSelect01 (fused_scale_shift_gate.py:20), FusedResidualLayerNormScaleShiftGateSelect01 (fused_scale_shift_gate.py:85). All implement forward_native. Eight other subclasses define their own forward_xpu, which is why this went unnoticed. Present since #12484 (7bc1dae).

Modifications

Just introduced forward_xpu in python/sglang/multimodal_gen/runtime/layers/custom_op.py. Verified after the patch that all six branches dispatch_forward() can return resolve to a base-class method.

Accuracy Tests

This PR does not touch kernels or model forward code, so outputs are unchanged on any platform that was already binding correctly (CUDA/NPU single-GPU).

Speed Tests and Profiling

No kernel or scheduling changes, so there is no expected throughput change on platforms that were already binding correctly.


CI States

Latest PR Test (Base): Not run yet
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

dispatch_forward() routes the xpu platform to self.forward_xpu, but the
base class never defined it. Since dispatch_forward() runs in __init__,
any subclass without its own override raises AttributeError at
construction, breaking model loading on Intel GPUs:

	AttributeError: 'LayerNorm' object has no attribute 'forward_xpu'

Route xpu to forward_native, matching the existing forward_tpu/npu/oot
fallbacks. No dispatch branch is modified, so CUDA/ROCm/NPU/MUSA/CPU
behavior is unchanged.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@github-actions github-actions Bot added the diffusion SGLang Diffusion label Aug 3, 2026
@dayanandav

Copy link
Copy Markdown
Contributor Author

@siju-samuel requested for review.

@dayanandav

Copy link
Copy Markdown
Contributor Author

@mickqian request for review.

@dayanandav

Copy link
Copy Markdown
Contributor Author

@mickqian @HaiShaw requested for review.

@dayanandav

Copy link
Copy Markdown
Contributor Author

@mingfeima requested for review.

@dayanandav

Copy link
Copy Markdown
Contributor Author

@ping1jing2 @yichiche @AgainstEntropy @BBuf requested for review.

@dayanandav

Copy link
Copy Markdown
Contributor Author

@mingfeima help to review and merge PR.

@siju-samuel siju-samuel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@siju-samuel

Copy link
Copy Markdown
Contributor

/tag-run-ci-label

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 24, 2026
Amrutha-M05 added a commit to Amrutha-M05/sglang that referenced this pull request Sep 2, 2026
Register intel_xpu_b60 as a consistency-threshold platform and add XPU
perf baselines for the 2-GPU fsdp-inference case.

Depends on sgl-project#33320 for CustomOp.forward_xpu.
# PyTorch-native implementation.
return self.forward_native(*args, **kwargs)

def forward_xpu(self, *args, **kwargs) -> Any:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should not add forward_xpu in the CustomeOp base class. The design is the feature owner will inherit this class and implement the missing APIs. If XPU wants to support, it will be added in the inherited class. Better to add in respective failed instance. Thanks!

@dayanandav
dayanandav marked this pull request as draft September 15, 2026 12:48
@dayanandav
dayanandav marked this pull request as ready for review September 16, 2026 04:27
Bring XPU memory reporting and component residency in line with the other
backends:

- get_available_gpu_memory() now reports mem_get_info() rather than
  total - memory_allocated(), so the caching allocator's reserve is no
  longer counted as free.
- XpuComponentOffloadStrategy parks the preferred component after warmup
  only while its weights plus a stage's headroom still fit in driver-free
  memory, and finish_request() returns the allocator's idle reserve at the
  request boundary. Both thresholds are XPU_HEADROOM_FRACTION (0.05) x
  device total memory, so no absolute byte count goes stale across device
  sizes.

Other backends are unaffected: the existing strategies and release paths
are byte-identical, the new path is reached by strategy selection behind
one is_xpu() guard, and the XPU helpers call torch.xpu directly.

Applies on top of the CustomOp.forward_xpu fallback earlier in this PR,
which is what puts the sglang transformer on the XPU path.

Verified with the wan2_1_t2v_1.3b diffusion server case at its component
offload defaults, which generates its output; that case's perf threshold is
derived from CUDA baselines and is untouched here. Unit coverage in
test_component_residency.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dayanandav

Copy link
Copy Markdown
Contributor Author

Same solution applied along with #36825 , so closing duplicate PR.

@dayanandav dayanandav closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] [multimodal_gen] CustomOp.dispatch_forward() XPU breaking model construction on Intel GPUs

3 participants