Repository navigation
[Intel][XPU] Device-agnostic fixes for XPU, scripted chunked-prefill, and DWDP - #37698
dayanandav wants to merge 43 commits into
Conversation
Fix sgl-project#31995 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fix sgl-project#24922 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fix sgl-project#36478 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Device-agnostic device timer plus scripted chunked-prefill KV canary fixes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Device-agnostic solution for the Intel XPU backend. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resolve GDN conflicts: keep gfx95 launch tuning, clamp BV to 16 only on XPU, union the test imports. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Restore the fork, drop the shared-wrapper BV clamp, re-add the is_xpu dispatch. Pass stride_h0_source -- now a required kernel arg. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The kernel uses cache_steps as the per-request stride, so it must be the allocated pitch, not the runtime draft count. Derive it from stride(0) as the shared wrapper does. Point the test at the dispatched wrapper so XPU covers the fork. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
KDA decode passes a as 4-D [B, T, H, K], where stride()[-2] is the head stride rather than the token stride. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CUDA / ROCm: unchanged by construction. The only kernel change is XPU-gated: BV, num_warps = _select_recurrent_launch_config(N, H, HV, K, V, is_kda)
if q.device.type == "xpu":
BV = min(BV, 16)
XPU: there is no "before" to measure. On main, After this PR — Qwen3.5-4B linear-attention shapes (
The launch config this PR uses on XPU is identical to the deleted wrapper's ( |
|
The current head does the opposite — deletes the fork and keeps cc @Xia-Weiwen |
|
Done — fork restored, shared-wrapper clamp removed, One line beyond a plain restore: the fork's kernel launch now passes |
dayanandav
left a comment
There was a problem hiding this comment.
Request for review
Call torch.cuda.empty_cache() per shard instead of once after the loop so the setup peak stays at ~1x local expert memory, not ~2x. Fixes sgl-project#41077 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nto combine/dwdp-shard-release-and-embedding-abort
Internal review comment fix
…to combine/dwdp-shard-release-and-embedding-abort
Resolve conflicts in mamba_state_scatter_triton.py, test_flux_pipeline.py, scheduling_comfyui_passthrough.py, the XPU GDN fork, and test_utils.py. Drop hunks now redundant with main: the passthrough scheduler step-index fix and the batch.scheduler fallback in ComfyUILatentPreparationStage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…neric-and-scripted-fixes # Conflicts: # python/sglang/multimodal_gen/runtime/pipelines_core/stages/text_encoding.py # python/sglang/test/scripted_runtime/http_server.py # python/sglang/test/scripted_runtime/req_handle.py # test/manual/chunked_prefill/test_scripted_regression.py
…edding-abort' into combine/xpu-device-generic-and-scripted-fixes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…neric-and-scripted-fixes # Conflicts: # test/registered/kernels/ops/attention/test_fused_verify_triton_gdn.py
…neric-and-scripted-fixes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fixes for #31995 #24922 #36478 and Device-agnostic device timer fix for Intel XPU backend
Verification
pre-commit run --all-filesclean (28 hooks) with the 36 changed files in the tree.test/registered/unit/server_args/test_server_args.py.against the pre-PR fork.
Notes on three deliberate choices
DWDP platform gate (
arg_groups/parallel_hook.py). The #31996 reference in anearlier commit message is wrong: #31996 is closed, not merged, and
layers/moe/dwdp/page_pool.py:9still carries a barefrom cuda.bindings import driver as cuda. The accurate justification is that theimport chain described in #31995 no longer exists on main --
model_runner.py:1331-1335importsDwdpManagerlazily insidemaybe_init_dwdpbehind
if get_parallel().dwdp_size <= 1: return, andfused_moe_triton/layer.py:78reaches the manager throughget_global_dwdp_manager, which pulls in nothing CUDA-specific. So the onlyremaining route to
ModuleNotFoundError: No module named 'cuda'is--dwdp-size >= 2on a platform without cuda-python, which is exactly whathandle_dwdpnow rejects with a message. This PR changes nothing underlayers/moe/dwdp/.ScriptedReqHandle.lock_refscontract (scripted_runtime/req_handle.py:56-59).A root
last_nodenow folds to 0 instead of returning the tree's permanentsentinel.
inc_lock_ref/dec_lock_refboth loopwhile node != self.root_node(
radix_cache.py:668,683), so the root'slock_refis fixed at 1 for the processlifetime and no req's lock can move it. The pre-change value on a root
last_nodewas therefore a constant unrelated to the req, and this is the fix for that, not a
weakening.
Deleted XPU GDN fork
(
srt/hardware_backend/xpu/kernels/fla/fused_sigmoid_gating_recurrent.py). Thatfork imported the shared
fused_sigmoid_gating_delta_rule_update_kernelandlaunched it by keyword without
stride_h0_source, which has no default in thekernel signature, so every XPU GDN decode on main raises at launch. Removing it
cannot regress a working path. XPU now takes the shared wrapper, whose only
XPU-specific behavior is the
BVcap at 16.🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ❌ Run #37654107331
Latest PR Test (Extra): ❌ Run #37654106872
Latest PR Test (AMD ROCm 10): ❌ Run #37654107341