Skip to content

[Intel][XPU][LoRA] Enable LoRA on Intel XPU - #30345

Merged
mingfeima merged 1 commit into
sgl-project:mainfrom
siju-samuel:enable-lora-xpu
Sep 8, 2026
Merged

mingfeima merged 1 commit into
sgl-project:mainfrom
siju-samuel:enable-lora-xpu

Conversation

@AnuSajikumar6264

@AnuSajikumar6264 AnuSajikumar6264 commented Jul 7, 2026 •

Copy link
Copy Markdown
Contributor

Enable the LoRA functionality on XPU (in addition to CUDA/ROCm), and enable the corresponding unit tests.

Source changes:

  • backends (triton/chunked/torch): use torch.device(self.device) instead of a hard-coded "cuda".
  • lora_moe_runners: route XPU to the pure-torch _naive_moe_lora_align_block_size fallback.
  • rotary_embedding base.py / mrope.py: guard the XPU-only sgl_kernel imports (fused_qk_rope_with_cos_sin_cache_inplace, multimodal_rotary_embedding).
  • lora_overlap_loader: use self.device_module.current_stream() instead of torch.cuda.current_stream().

Supported and Verified following Features

  • Core dense LoRA (triton + csgmv backends)
  • cuda-graph + LoRA
  • Multi-LoRA
  • MoE-LoRA
  • Dynamic Load/ Unload
  • Pinned Adapters
  • LoRA with Overlap Loading
  • LoRA with Radix Cache
  • LoRA with TP
  • Eviction (LRU/FIFO)
  • Embedding with LoRA

Test changes:

  • Device-agnostic device selection via get_device() and register_xpu_ci across the kernel/unit and small-model LoRA tests.

Motivation

  • Intel XPU is a first-class inference target: SGLang already supports XPU for base model inference; LoRA fine-tuned models are widely used in production and should be deployable on XPU without requiring a separate code path or falling back to the generic Transformers backbone.
  • Hard-coded "cuda" strings are silent correctness bugs on XPU: Several hot paths (init_cuda_graph_batch_info, lora_overlap_loader) referenced torch.cuda directly, causing device mismatches or runtime errors when the active device is an Intel XPU, even though the surrounding logic was otherwise device-agnostic.
  • CI coverage prevents regressions across backends: Without device-agnostic test infrastructure (get_device(), register_xpu_ci, ROUGE-L tolerance on XPU), XPU-specific breakage in LoRA paths would go undetected until a user report, making the XPU support effectively untested and unreliable.

Modifications

  • Wrote a combined feature-level LoRA test suite validating Dynamic Load/Unload, Pinned Adapters, Radix Cache with LoRA, Embedding with LoRA, Multi-LoRA, and MoE-LoRA both individually and in combination across all three attention backends (csgmv, triton, and torch-native) to catch feature interaction bugs across devices and backend configurations
  • Replaced torch.cuda.current_stream() with self.device_module.current_stream() in lora_overlap_loader to prevent runtime errors on XPU using the existing device module abstraction
  • Added get_device() helper and register_xpu_ci for device-agnostic device selection across kernel and unit tests so the same test suite runs on both CUDA and XPU

Accuracy Tests

Speed Tests and Profiling

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #34091201271
Latest PR Test (Extra): ❌ Run #34091201062
Latest PR Test (AMD ROCm 7.2): ❌ Run #34091201287

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@github-actions github-actions Bot added the lora label Jul 7, 2026
@AnuSajikumar6264 AnuSajikumar6264 changed the title [LoRA][XPU] Enable LoRA on Intel XPU [LoRA][XPU][WIP] Enable LoRA on Intel XPU Jul 7, 2026
@AnuSajikumar6264 AnuSajikumar6264 changed the title [LoRA][XPU][WIP] Enable LoRA on Intel XPU [WIP][Intel][XPU][LoRA] Enable LoRA on Intel XPU Jul 7, 2026
@dayanandav dayanandav mentioned this pull request Jul 8, 2026
3 tasks done
@AnuSajikumar6264
AnuSajikumar6264 force-pushed the enable-lora-xpu branch 2 times, most recently from 1f97ee7 to 7c8d106 Compare July 20, 2026 04:00
@AnuSajikumar6264 AnuSajikumar6264 changed the title [WIP][Intel][XPU][LoRA] Enable LoRA on Intel XPU [Intel][XPU][LoRA] Enable LoRA on Intel XPU Jul 20, 2026
@AnuSajikumar6264
AnuSajikumar6264 marked this pull request as ready for review July 20, 2026 04:03
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@AnuSajikumar6264

Copy link
Copy Markdown
Contributor Author

@mingfeima Could you please help to review this PR. This enables Lora Functionality for XPU.

Comment thread test/registered/lora/test_lora_tied_lm_head.py Outdated
Comment thread test/registered/lora/test_virtual_experts_kernels.py Outdated
Comment thread test/registered/lora/test_lora_moe_vllm_sgl_logprob_diff.py Outdated
@AnuSajikumar6264
AnuSajikumar6264 force-pushed the enable-lora-xpu branch 4 times, most recently from d191e68 to 9882404 Compare August 2, 2026 17:28
@AnuSajikumar6264
AnuSajikumar6264 force-pushed the enable-lora-xpu branch 3 times, most recently from af8b191 to 4aaf9f4 Compare August 5, 2026 10:28
@AnuSajikumar6264
AnuSajikumar6264 force-pushed the enable-lora-xpu branch 3 times, most recently from 81196e4 to a050702 Compare August 12, 2026 16:26
@mingfeima

Copy link
Copy Markdown
Collaborator

@AnuSajikumar6264 could you please first check whether this pull request brings regression on CIs for other devices? a lot of red in the CI...

besides, do you have perf / accuracy test result for it? one or two prioritized model will do.

we are launching local test for this one.

@mingfeima

Copy link
Copy Markdown
Collaborator

@gaopengff please help run this one.

@AnuSajikumar6264

Copy link
Copy Markdown
Contributor Author

1. Accuracy

test/manual/lora/test_lora_backend.py::TestLoRABackend::test_all_lora_models — 1 passed
(22 m 36 s). Each case runs 4 engines (SRT+LoRA, SRT base, HF+LoRA, HF base) and compares
top input/output logprobs against HuggingFace plus ROUGE-L of the decoded text.
max_new_tokens=32, dtype=float16, 2 prompts per case (values below are prompt 0 / prompt 1).

Base model Adapter Backend Max prefill logprob diff Max decode logprob diff ROUGE-L
Llama-3.1-8B-Instruct llama-3.1-nemoguard-8b-topic-control triton 0.0164 / 0.0312 0.0164 / 0.0312 1.0 / 1.0
Llama-3.1-8B-Instruct llama-3.1-nemoguard-8b-topic-control csgmv 0.0164 / 0.0299 0.0234 / 0.0233 1.0 / 1.0
Llama-2-7b-hf wizardLM-LlaMA-LoRA-7B triton 0.0156 / 0.0312 0.0396 / 0.0312 1.0 / 1.0
Llama-2-7b-hf wizardLM-LlaMA-LoRA-7B csgmv 0.0156 / 0.0312 0.0396 / 0.0312 1.0 / 1.0

On Llama-2-7b the two backends agree exactly, digit for digit.

All diffs are far inside the test tolerance (1e-1 for the nemoguard adapter) and ROUGE-L
is exactly 1.0 everywhere, i.e. XPU reproduces the HF reference text token for token.

Other LoRA tests on XPU

Test Result
test/registered/lora/test_fused_moe_lora_kernel.py 108 passed (57.41 s)
test/registered/lora/test_lora_overlap_loading.py 9 passed (2 m 22 s)
test/registered/lora/test_moe_lora_info.py 3 passed (7.91 s) †

† needs the local torch.searchsorted workaround described in §3.


2. Performance

python -m sglang.benchmark.serving, random dataset with lengths pinned
(--random-range-ratio 1.0, ignore_eos on, so every request emits exactly output_len
tokens). Server: --disable-radix --max-loras-per-batch 8 --max-running-requests 8 --tp-size 1 --dtype float16; client --max-concurrency 8 to match server capacity,
4 adapters drawn uniformly per request.

Workload A — 512 in / 64 out, 32 requests

Config Duration (s) Total tok/s Output tok/s Mean E2E (ms) Mean TTFT (ms) Mean TPOT (ms)
base, no LoRA 15.08 1220.27 135.82 3754.94 921.60 44.97
LoRA csgmv 17.67 1041.42 115.91 4398.54 1031.19 53.45
LoRA triton 18.35 1002.46 111.58 4570.59 1077.91 55.44

Workload B — 1024 in / 128 out, 16 requests

Config Duration (s) Total tok/s Output tok/s Mean E2E (ms) Mean TTFT (ms) Mean TPOT (ms)
base, no LoRA 18.07 1019.36 113.36 9002.33 2023.99 54.95
LoRA csgmv 20.65 891.94 99.19 10289.02 2199.53 63.70
LoRA triton 21.18 869.37 96.68 10555.75 2203.80 65.76

Cost of attaching adapters — each row is the LoRA run relative to the same base model
served with no adapters
, so a negative throughput delta is the expected price of LoRA
(two extra GEMMs per adapted layer — the r×d shrink and d×r expand — plus per-batch
adapter routing), not a regression against pre-PR behaviour.

Workload Backend Δ total throughput vs base Δ mean TPOT vs base
A (512/64) csgmv −14.7 % +18.9 %
A (512/64) triton −17.8 % +23.3 %
B (1024/128) csgmv −12.5 % +15.9 %
B (1024/128) triton −14.7 % +19.7 %

The base path itself is unchanged by this PR on this box: its only base-forward edit is the
fused_qk_rope_with_cos_sin_cache_inplace is not None guard in rotary_embedding/base.py,
and sgl_kernel provides that symbol here, so both the baseline and the LoRA runs take the
same fused rotary kernel.

csgmv is the faster backend on XPU in both workloads (+3.9 % on A, +2.6 % on B over
triton), consistent with its behaviour on CUDA. All 32/16 requests succeeded in every
run, with exactly 2048 generated tokens per run.

@gaopengff gaopengff left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Generally LGTM. Next step is to make this PR's CI pass.

Comment thread test/registered/lora/test_virtual_experts_kernels.py Outdated
Make the LoRA path device-portable so it runs on Intel XPU in addition to
CUDA/ROCm, and enable the corresponding tests.

Source changes:
- backends (base/triton/chunked/torch): use torch.device(self.device) instead
  of a hard-coded "cuda" in init_cuda_graph_batch_info.
- lora_moe_runners: route XPU to the pure-torch
  _naive_moe_lora_align_block_size fallback (the moe_lora_align .cu kernel is
  CUDA-only).
- rotary_embedding base/mrope: guard the XPU-only sgl_kernel imports so a
  missing symbol no longer breaks the native model registry on XPU (which
  would silently fall back to the generic Transformers backbone).
- lora_overlap_loader: use self.device_module.current_stream() instead of
  torch.cuda.current_stream(), and fix a scheduling livelock on backends whose
  staging stream completes synchronously. The load finishes before its event is
  queried, so the adapter was never reported LOADED in the current pass and a
  sibling's load could evict it before it was ever scheduled. On CUDA the copy
  is genuinely async, so overlap behavior is unchanged.
- arg_groups/overrides: make supports_mamba_cache_extra_buffer() device-aware.
  The extra_buffer strategy is backed by FLA, which has no XPU kernels, so
  "auto" selected it and then tripped a device assert, making every
  hybrid-mamba model unusable. Gating the probe lets "auto" fall back to
  no_buffer.

Test changes:
- Device-agnostic device selection via get_device() and register_xpu_ci across
  the kernel/unit and small-model LoRA tests.
- test_lora_xpu_basics: new suite covering the LoRA serving features on XPU --
  accuracy vs HF+PEFT, multi-adapter batching, the drainer, runtime
  load/unload, pinning, pool eviction, the radix cache, and embedding models,
  each in isolation, plus feature combinations with graph capture off (the
  graph-on combinations live in test_lora_comb_xpu.py). Adapters are synthetic
  with known-distinct weights, so routing is asserted by nearest-reference
  rather than exact output.
- test_lora_qwen3_8b_logprob_diff: select the attention backend per platform.
  fa4 dispatches into the CUTLASS CUTE kernel, which cannot import off CUDA.
- test_torch_backend: pass the required output_offset_cpu argument.
- test_lora_moe_vllm_sgl_logprob_diff: split the parity test per attention
  backend so each platform's backend is covered explicitly.
- test_virtual_experts_kernels: narrow the tvm_ffi align-variant guard to
  skipIf(is_xpu()) so the CUDA-JIT variant still runs where available.
- test_lora_overlap_loading: model a genuinely in-flight async load so the
  livelock fix is exercised.
- lora_utils.run_lora_test_by_batch: compare SRT vs HF greedy outputs with the
  established ROUGE-L tolerance on XPU (exact match elsewhere), since
  kernel-level fp differences can diverge greedy decoding.

Co-Authored-By: Anupa Sajikumar <anupa.sajikumar@intel.com>
@mingfeima
mingfeima merged commit 91a45ea into sgl-project:main Sep 8, 2026
149 of 188 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

intel lora run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) xpu intel gpu with device `torch.xpu`

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants