Repository navigation
[XPU]Enable HiSparse hierarchical sparse KV cache on Intel XPU - #32792
Conversation
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
f18129c to
4c99140
Compare
731cf12 to
6c23caa
Compare
|
/tag-run-ci-label |
7808ffb to
021c96f
Compare
|
/rerun-failed-ci |
|
/rerun-failed-ci |
|
Hi @mingfeima , @siju-samuel I've analyzed all 21 CI failures and None are caused by this PR - they are 5 infra issues & 1 missing opt-in label. 1. XPU —
|
One textual conflict: - memory_pool_host.py: sgl-project#32792 moved the transfer_cache_dsv4_mla import into an XPU branch and widened the kvcacheio guard with _is_xpu, right where this branch had left a stray blank line. Upstream's block taken verbatim and the blank line dropped, so the import section is now main's; the only delta left in the file is the SWA capture-event helpers.
Enable HiSparse (hierarchical sparse KV cache) on Intel XPU (in addition to CUDA/ROCm), and enable the corresponding unit tests.
The two hot kernels —
load_cache_to_device_buffer_{mla,dsv4_mla}andtransfer_cache_dsv4_mla— are ported to SYCL in sgl-kernel-xpu#335. CUDA/ROCm JIT-compiles them viasglang.kernels.ops.kvcache.hisparse; the XPU port is AOT-compiled into thesgl_kernelwheel.Source changes:
pool_host/common.py: register"xpu": alloc_with_pin_memoryinALLOC_MEMORY_FUNCSso XPU uses torch's built-inpin_memory=Trueinstead ofcudaHostRegister.pool_host/mla.py,memory_pool_host.py,hisparse_memory_pool.py: widenif _is_cuda or _is_hip:to... or _is_xpu:sosgl_kernel.kvcacheio.transfer_kv_all_layer_*is imported on XPU.hisparse_coordinator.py,memory_pool_host.py: branch the kernel imports explicitly — XPU tosgl_kernel, everything else to the JIT ops.Supported and Verified following Features
page_size> 1)Not supported on XPU
copy_cache_planned_mlahas no AOT SYCL kernel, so it is bound to a raising stub; the coordinator disables prefetch at init with a warning.SGLANG_DEBUG_HISPARSE_SKIP_IO:skip_iois a JIT template parameter baked in at compile time, so the AOT ops do not accept it. Setting the env var on XPU now raises at init instead of silently producing timings that include the KV copy.Test changes:
get_device()/get_device_module()andregister_xpu_ciacross the HiSparse kernel and unit tests.Motivation
is_cuda() or is_hip()in several places and dispatches host pin-memory allocation through a CUDA-only registrar, so DSA / DeepSeek-V4-class models cannot use hierarchical sparse attention there.sgl-kernel-xpu; what was missing was the sglang-side wiring, backend gates, and device-agnostic tests.Modifications
ALLOC_MEMORY_FUNCSand widened the three_is_cuda or _is_hipimport gates to admit XPU.sgl_kernelops and CUDA/ROCm keeps the JIT ops unchanged.get_device()/get_device_module()and registered them under the XPU CI suitestage-b-test-1-gpu-xpu.Accuracy Tests
All 10 skips are ROCm-only tests gated on
not is_hip(); they skip on CUDA as well. The unit tests include the kernel-vs-naive_load_topkoracle comparison, so SYCL/CUDA numerical parity is exercised on the same test bodies.CUDA/ROCm is unaffected on this revision.
Speed Tests and Profiling
N/A — enablement only; the CUDA/ROCm hot path is untouched. SYCL kernel benchmarking is tracked in
sgl-kernel-xpu.Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-cicc: @siju-samuel @rbabukv
CI States
Latest PR Test (Base): 🚫 Run #35296443070
Latest PR Test (Extra): ❌ Run #35296442913
Latest PR Test (AMD ROCm 10): ❌ Run #35296442812