Repository navigation
[Intel][XPU][LoRA] Enable LoRA on Intel XPU - #30345
Conversation
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
1f97ee7 to
7c8d106
Compare
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
@mingfeima Could you please help to review this PR. This enables Lora Functionality for XPU. |
7c8d106 to
dca7c91
Compare
3fe68b1 to
dca7c91
Compare
d191e68 to
9882404
Compare
af8b191 to
4aaf9f4
Compare
81196e4 to
a050702
Compare
a050702 to
7b07d51
Compare
|
@AnuSajikumar6264 could you please first check whether this pull request brings regression on CIs for other devices? a lot of red in the CI... besides, do you have perf / accuracy test result for it? one or two prioritized model will do. we are launching local test for this one. |
|
@gaopengff please help run this one. |
1. Accuracy
On Llama-2-7b the two backends agree exactly, digit for digit. All diffs are far inside the test tolerance ( Other LoRA tests on XPU
† needs the local 2. Performance
Workload A — 512 in / 64 out, 32 requests
Workload B — 1024 in / 128 out, 16 requests
Cost of attaching adapters — each row is the LoRA run relative to the same base model
The base path itself is unchanged by this PR on this box: its only base-forward edit is the
|
9bc3fa4 to
7c50b5c
Compare
gaopengff
left a comment
There was a problem hiding this comment.
Generally LGTM. Next step is to make this PR's CI pass.
7c50b5c to
7897fbf
Compare
Make the LoRA path device-portable so it runs on Intel XPU in addition to CUDA/ROCm, and enable the corresponding tests. Source changes: - backends (base/triton/chunked/torch): use torch.device(self.device) instead of a hard-coded "cuda" in init_cuda_graph_batch_info. - lora_moe_runners: route XPU to the pure-torch _naive_moe_lora_align_block_size fallback (the moe_lora_align .cu kernel is CUDA-only). - rotary_embedding base/mrope: guard the XPU-only sgl_kernel imports so a missing symbol no longer breaks the native model registry on XPU (which would silently fall back to the generic Transformers backbone). - lora_overlap_loader: use self.device_module.current_stream() instead of torch.cuda.current_stream(), and fix a scheduling livelock on backends whose staging stream completes synchronously. The load finishes before its event is queried, so the adapter was never reported LOADED in the current pass and a sibling's load could evict it before it was ever scheduled. On CUDA the copy is genuinely async, so overlap behavior is unchanged. - arg_groups/overrides: make supports_mamba_cache_extra_buffer() device-aware. The extra_buffer strategy is backed by FLA, which has no XPU kernels, so "auto" selected it and then tripped a device assert, making every hybrid-mamba model unusable. Gating the probe lets "auto" fall back to no_buffer. Test changes: - Device-agnostic device selection via get_device() and register_xpu_ci across the kernel/unit and small-model LoRA tests. - test_lora_xpu_basics: new suite covering the LoRA serving features on XPU -- accuracy vs HF+PEFT, multi-adapter batching, the drainer, runtime load/unload, pinning, pool eviction, the radix cache, and embedding models, each in isolation, plus feature combinations with graph capture off (the graph-on combinations live in test_lora_comb_xpu.py). Adapters are synthetic with known-distinct weights, so routing is asserted by nearest-reference rather than exact output. - test_lora_qwen3_8b_logprob_diff: select the attention backend per platform. fa4 dispatches into the CUTLASS CUTE kernel, which cannot import off CUDA. - test_torch_backend: pass the required output_offset_cpu argument. - test_lora_moe_vllm_sgl_logprob_diff: split the parity test per attention backend so each platform's backend is covered explicitly. - test_virtual_experts_kernels: narrow the tvm_ffi align-variant guard to skipIf(is_xpu()) so the CUDA-JIT variant still runs where available. - test_lora_overlap_loading: model a genuinely in-flight async load so the livelock fix is exercised. - lora_utils.run_lora_test_by_batch: compare SRT vs HF greedy outputs with the established ROUGE-L tolerance on XPU (exact match elsewhere), since kernel-level fp differences can diverge greedy decoding. Co-Authored-By: Anupa Sajikumar <anupa.sajikumar@intel.com>
7897fbf to
2c5f8d3
Compare
Enable the LoRA functionality on XPU (in addition to CUDA/ROCm), and enable the corresponding unit tests.
Source changes:
Supported and Verified following Features
Test changes:
Motivation
Modifications
Accuracy Tests
Speed Tests and Profiling
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #34091201271
Latest PR Test (Extra): ❌ Run #34091201062
Latest PR Test (AMD ROCm 7.2): ❌ Run #34091201287