Skip to content

[XPU]Enable HiSparse hierarchical sparse KV cache on Intel XPU - #32792

Merged
mingfeima merged 4 commits into
sgl-project:mainfrom
Amrutha-M05:hisparse-xpu
Sep 21, 2026
Merged

mingfeima merged 4 commits into
sgl-project:mainfrom
Amrutha-M05:hisparse-xpu

Conversation

@Amrutha-M05

@Amrutha-M05 Amrutha-M05 commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Enable HiSparse (hierarchical sparse KV cache) on Intel XPU (in addition to CUDA/ROCm), and enable the corresponding unit tests.

The two hot kernels — load_cache_to_device_buffer_{mla,dsv4_mla} and transfer_cache_dsv4_mla — are ported to SYCL in sgl-kernel-xpu#335. CUDA/ROCm JIT-compiles them via sglang.kernels.ops.kvcache.hisparse; the XPU port is AOT-compiled into the sgl_kernel wheel.

Source changes:

  • pool_host/common.py: register "xpu": alloc_with_pin_memory in ALLOC_MEMORY_FUNCS so XPU uses torch's built-in pin_memory=True instead of cudaHostRegister.
  • pool_host/mla.py, memory_pool_host.py, hisparse_memory_pool.py: widen if _is_cuda or _is_hip: to ... or _is_xpu: so sgl_kernel.kvcacheio.transfer_kv_all_layer_* is imported on XPU.
  • hisparse_coordinator.py, memory_pool_host.py: branch the kernel imports explicitly — XPU to sgl_kernel, everything else to the JIT ops.

Supported and Verified following Features

  • Swap-in of top-k selected pages (MLA and DSv4 layouts)
  • Device-buffer LRU replacement across decode steps
  • Long-sequence host-to-device DMA path
  • Batched multi-request swap-in with padding
  • Request lifecycle: staging path and direct-to-host path
  • PD decode host-slot preallocation
  • Paged host layout (page_size > 1)

Not supported on XPU

  • Shared-index prefetch: copy_cache_planned_mla has no AOT SYCL kernel, so it is bound to a raising stub; the coordinator disables prefetch at init with a warning.
  • SGLANG_DEBUG_HISPARSE_SKIP_IO: skip_io is a JIT template parameter baked in at compile time, so the AOT ops do not accept it. Setting the env var on XPU now raises at init instead of silently producing timings that include the KV copy.

Test changes:

  • Device-agnostic device selection via get_device() / get_device_module() and register_xpu_ci across the HiSparse kernel and unit tests.

Motivation

  • Intel XPU is a first-class inference target, but HiSparse hard-gates on is_cuda() or is_hip() in several places and dispatches host pin-memory allocation through a CUDA-only registrar, so DSA / DeepSeek-V4-class models cannot use hierarchical sparse attention there.
  • The SYCL kernels already exist in sgl-kernel-xpu; what was missing was the sglang-side wiring, backend gates, and device-agnostic tests.
  • Without XPU CI coverage, XPU-specific breakage in the HiSparse paths would go undetected until a user report.

Modifications

  • Registered XPU in ALLOC_MEMORY_FUNCS and widened the three _is_cuda or _is_hip import gates to admit XPU.
  • Branched the swap-in / replay / transfer kernel imports so XPU resolves to the AOT sgl_kernel ops and CUDA/ROCm keeps the JIT ops unchanged.
  • Made the two XPU-unsupported features fail loudly at init rather than silently no-op.
  • Converted the HiSparse kernel and unit tests to get_device() / get_device_module() and registered them under the XPU CI suite stage-b-test-1-gpu-xpu.

Accuracy Tests

$ pytest test/registered/kernels/ops/kvcache/test_hisparse.py
10 passed, 8 skipped

$ pytest test/registered/unit/managers/test_hisparse_unit.py
9 passed, 2 skipped

All 10 skips are ROCm-only tests gated on not is_hip(); they skip on CUDA as well. The unit tests include the kernel-vs-naive_load_topk oracle comparison, so SYCL/CUDA numerical parity is exercised on the same test bodies.

CUDA/ROCm is unaffected on this revision.

Speed Tests and Profiling

N/A — enablement only; the CUDA/ROCm hot path is untouched. SYCL kernel benchmarking is tracked in sgl-kernel-xpu.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

cc: @siju-samuel @rbabukv


CI States

Latest PR Test (Base): 🚫 Run #35296443070
Latest PR Test (Extra): ❌ Run #35296442913
Latest PR Test (AMD ROCm 10): ❌ Run #35296442812

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@siju-samuel

Copy link
Copy Markdown
Contributor

/tag-run-ci-label

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 26, 2026
Comment thread test/registered/unit/managers/test_hisparse_unit.py Outdated
Comment thread test/registered/unit/managers/test_hisparse_unit.py Outdated
Comment thread python/sglang/srt/mem_cache/pool_host/common.py Outdated
Comment thread python/sglang/srt/managers/hisparse_coordinator.py
Comment thread test/registered/kernels/ops/kvcache/test_hisparse.py
Comment thread python/sglang/srt/mem_cache/memory_pool_host.py
Comment thread python/sglang/srt/managers/hisparse_coordinator.py Outdated
Comment thread python/sglang/srt/mem_cache/pool_host/common.py Outdated
Comment thread test/registered/unit/managers/test_hisparse_unit.py Outdated
@jianan-gu

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@mingfeima mingfeima added intel xpu intel gpu with device `torch.xpu` labels Sep 16, 2026
@mingfeima

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@Amrutha-M05

Copy link
Copy Markdown
Contributor Author

Hi @mingfeima , @siju-samuel I've analyzed all 21 CI failures and None are caused by this PR - they are 5 infra issues & 1 missing opt-in label.

1. XPU — stage-a-test-1-gpu-xpu, finish (2 jobs)

Dies at Start CI container; zero tests run.

Removing existing container: ci_sglang_xpu
Error response from daemon: cannot remove container "ci_sglang_xpu":
  could not kill container: tried to kill container, but did not receive an exit event

main fails identically (run 35301744417, on system815752), as do #38779, #38831, #39980.

2. AMD — 3 test jobs + wait-for-stage-a-amd, pr-test-amd-finish (5 jobs)

All three fail at Start CI container, before any test:

failed to convert whiteout file "opt/venv/.../.wh.pip-24.0.dist-info": operation not permitted
Error: 'docker pull rocm/sgl-dev:v0.5.19-rocm10-mi30x-20260917' failed after 6 attempts

An overlayfs fault on the MI300 runner, not disk space. Same failure on 8 other branches, including #36700, #39903, #36546, #36187.

3. CUDA base-b (8 jobs)

  • base-b-test-2-gpu-large (2) — the only real failure.
    test/registered/e2e/speculative/test_dflash_domino.py aborted in setUpClass,
    0 tests executed:
    RuntimeError: GPU(s) still not idle after waiting 30s before setUpClass:
      GPU 0 uses 4.65 GiB (no other compute processes; self pid=911977 holds 0.00 GiB)
    Ran 0 tests in 30.070s
    
    That is the _wait_for_gpu_idle_in_ci gate from [CI] Wait for GPU memory release before each test class setUpClass #31509, whose docstring names this mode:
    "Killed server processes return GPU memory asynchronously." The 4.65 GiB belongs to no live process — residue from the previous file in the shard, which passed. The test imports nothing from mem_cache / hisparse / pool_host.
  • base-b-test-1-gpu-large (2) and (3) — not real failures; the log says so:
    ##[error]Fast-fail: skipping — root cause job(s): wait-for-base-b, base-b-test-2-gpu-large (2)
    
  • base-b-test-2-gpu-large (4) and (5) — cancelled, never assigned a runner (queued 02:33 → 08:14, 5h41m). Runner starvation; the whole run was then cancelled.
  • wait-for-base-b, pr-test-finish, notify-pr-states — aggregators / cancelled.

4. MLX — stage-a-unit-test-mlx, pr-test-mlx-finish (2 jobs)

test_scheduler_mixin.py::TestOverlapLoopGracefulExit raises _StopLoop. Pre-existing on main, unchanged by my rebase; this PR adds no MLX code. Also failing on #39902, #38468, #36559, #39059.

5. call-gate / pr-gate ×2 + 2 extra aggregators (4 jobs)

Not a test failure, due to missing opt-in label:

Missing required label 'run-ci-extra'. Add the label (e.g. via /tag-and-rerun-ci extra)

Local verification on Intel XPU

Because the XPU job never starts a container, CI has no coverage of this change. Run locally on Intel BMG at this head:

test/registered/unit/managers/test_hisparse_unit.py         9 passed,  2 skipped
test/registered/kernels/ops/kvcache/test_hisparse.py       10 passed,  8 skipped
                                                     total 19 passed, 10 skipped, 0 failed

The blocker is the XPU runner (group 1), that is the job that actually covers this change, and it cannot pass until system815752 is fixed or the container name is made runner-unique.

@mingfeima
mingfeima merged commit d20cd9d into sgl-project:main Sep 21, 2026
138 of 159 checks passed
amd-danli103 added a commit to amd-danli103/sglang that referenced this pull request Sep 21, 2026
One textual conflict:

- memory_pool_host.py: sgl-project#32792 moved the transfer_cache_dsv4_mla import
  into an XPU branch and widened the kvcacheio guard with _is_xpu, right
  where this branch had left a stray blank line. Upstream's block taken
  verbatim and the blank line dropped, so the import section is now
  main's; the only delta left in the file is the SWA capture-event
  helpers.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang intel memory-pool run-ci CI: run the baseline test suite on this PR xpu intel gpu with device `torch.xpu`

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants