Skip to content

[Intel][XPU] Device-agnostic fixes for XPU, scripted chunked-prefill, and DWDP - #37698

Open
dayanandav wants to merge 43 commits into
sgl-project:mainfrom
dayanandav:combine/xpu-device-generic-and-scripted-fixes
Open

dayanandav wants to merge 43 commits into
sgl-project:mainfrom
dayanandav:combine/xpu-device-generic-and-scripted-fixes

Conversation

@dayanandav

@dayanandav dayanandav commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Fixes for #31995 #24922 #36478 and Device-agnostic device timer fix for Intel XPU backend

Verification

  • pre-commit run --all-files clean (28 hooks) with the 36 changed files in the tree.
  • The DWDP argparse cases live in test/registered/unit/server_args/test_server_args.py.
  • Pending on my XPU box: Qwen3.5 before/after perf numbers, and a GDN decode run
    against the pre-PR fork.

Notes on three deliberate choices

DWDP platform gate (arg_groups/parallel_hook.py). The #31996 reference in an
earlier commit message is wrong: #31996 is closed, not merged, and
layers/moe/dwdp/page_pool.py:9 still carries a bare
from cuda.bindings import driver as cuda. The accurate justification is that the
import chain described in #31995 no longer exists on main --
model_runner.py:1331-1335 imports DwdpManager lazily inside maybe_init_dwdp
behind if get_parallel().dwdp_size <= 1: return, and
fused_moe_triton/layer.py:78 reaches the manager through
get_global_dwdp_manager, which pulls in nothing CUDA-specific. So the only
remaining route to ModuleNotFoundError: No module named 'cuda' is
--dwdp-size >= 2 on a platform without cuda-python, which is exactly what
handle_dwdp now rejects with a message. This PR changes nothing under
layers/moe/dwdp/.

ScriptedReqHandle.lock_refs contract (scripted_runtime/req_handle.py:56-59).
A root last_node now folds to 0 instead of returning the tree's permanent
sentinel. inc_lock_ref / dec_lock_ref both loop while node != self.root_node
(radix_cache.py:668,683), so the root's lock_ref is fixed at 1 for the process
lifetime and no req's lock can move it. The pre-change value on a root last_node
was therefore a constant unrelated to the req, and this is the fix for that, not a
weakening.

Deleted XPU GDN fork
(srt/hardware_backend/xpu/kernels/fla/fused_sigmoid_gating_recurrent.py).
That
fork imported the shared fused_sigmoid_gating_delta_rule_update_kernel and
launched it by keyword without stride_h0_source, which has no default in the
kernel signature, so every XPU GDN decode on main raises at launch. Removing it
cannot regress a working path. XPU now takes the shared wrapper, whose only
XPU-specific behavior is the BV cap at 16.

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ❌ Run #37654107331
Latest PR Test (Extra): ❌ Run #37654106872
Latest PR Test (AMD ROCm 10): ❌ Run #37654107341

dayanandav and others added 5 commits September 2, 2026 13:13
Fix sgl-project#24922

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fix sgl-project#36478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Device-agnostic device timer plus scripted chunked-prefill KV canary fixes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Device-agnostic solution for the Intel XPU backend.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dayanandav and others added 4 commits September 21, 2026 15:59
Resolve GDN conflicts: keep gfx95 launch tuning, clamp BV to 16 only on XPU,
union the test imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Restore the fork, drop the shared-wrapper BV clamp, re-add the is_xpu
dispatch. Pass stride_h0_source -- now a required kernel arg.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The kernel uses cache_steps as the per-request stride, so it must be the
allocated pitch, not the runtime draft count. Derive it from stride(0) as
the shared wrapper does. Point the test at the dispatched wrapper so XPU
covers the fork.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
KDA decode passes a as 4-D [B, T, H, K], where stride()[-2] is the head
stride rather than the token stride.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dayanandav

dayanandav commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

can you also provide perf numbers of qwen3.5 before and after this PR. need to verify no perf regression.

CUDA / ROCm: unchanged by construction. The only kernel change is XPU-gated:

BV, num_warps = _select_recurrent_launch_config(N, H, HV, K, V, is_kda)
if q.device.type == "xpu":
    BV = min(BV, 16)

_select_recurrent_launch_config (including the gfx950 tuning) is untouched, so BV / num_warps / grid are identical to main on CUDA and ROCm.

XPU: there is no "before" to measure. On main, gdn_triton.py dispatches XPU to a wrapper that never passed stride_h0_source, so every GDN decode raised:

TypeError: dynamic_func() missing 1 required positional argument: 'stride_h0_source'

After this PR — Qwen3.5-4B linear-attention shapes (H=16, HV=32, K=128, V=128), T=1 decode, bf16, 1x Intel XPU:

batch 1 8 32 64 128
us/call 55.1 55.8 375.0 718.8 1399.5

The launch config this PR uses on XPU is identical to the deleted wrapper's (BV=16, BK=128, num_warps=1, num_stages=3, grid=(NK, NV, N*HV)) — it de-duplicates the wrapper rather than retuning it, so there is no XPU tiling change to regress either.

@mingfeima
mingfeima marked this pull request as draft September 22, 2026 07:41
@mingfeima

Copy link
Copy Markdown
Collaborator

fused_sigmoid_gating_recurrent.py: the request was to keep the XPU fork (sglang.srt.hardware_backend.xpu.kernels.fla.fused_sigmoid_gating_recurrent) and drop the shared-wrapper change.

The current head does the opposite — deletes the fork and keeps if q.device.type == "xpu": BV = min(BV, 16) in the shared file.

cc @Xia-Weiwen

@dayanandav

dayanandav commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Done — fork restored, shared-wrapper clamp removed, elif is_xpu() dispatch re-added in gdn_triton.py.

One line beyond a plain restore: the fork's kernel launch now passes stride_h0_source. The shared kernel made it a required positional, so main's fork raises TypeError on every XPU GDN decode without it.

@dayanandav dayanandav left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request for review

@dayanandav
dayanandav marked this pull request as ready for review September 23, 2026 18:30
dayanandav and others added 5 commits September 23, 2026 19:18
Call torch.cuda.empty_cache() per shard instead of once after the
loop so the setup peak stays at ~1x local expert memory, not ~2x.

Fixes sgl-project#41077

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nto combine/dwdp-shard-release-and-embedding-abort
Internal review comment fix
…to combine/dwdp-shard-release-and-embedding-abort
Resolve conflicts in mamba_state_scatter_triton.py, test_flux_pipeline.py,
scheduling_comfyui_passthrough.py, the XPU GDN fork, and test_utils.py.
Drop hunks now redundant with main: the passthrough scheduler step-index
fix and the batch.scheduler fallback in ComfyUILatentPreparationStage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dayanandav and others added 3 commits September 28, 2026 18:18
…neric-and-scripted-fixes

# Conflicts:
#	python/sglang/multimodal_gen/runtime/pipelines_core/stages/text_encoding.py
#	python/sglang/test/scripted_runtime/http_server.py
#	python/sglang/test/scripted_runtime/req_handle.py
#	test/manual/chunked_prefill/test_scripted_regression.py
…edding-abort' into combine/xpu-device-generic-and-scripted-fixes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…neric-and-scripted-fixes

# Conflicts:
#	test/registered/kernels/ops/attention/test_fused_verify_triton_gdn.py

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion intel jit-kernel memory-pool run-ci CI: run the baseline test suite on this PR xpu intel gpu with device `torch.xpu`

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants