[CPU] enable CI for PRs, add Dockerfile and auto build task #6458

ZailiWang · 2025-05-20T08:59:42Z

Motivation

The PR is for enabling the docker env setup and CI processes for running SGLang on Xeon CPU servers.

Modifications

Add the dockerfile, and yml files for CPU part of CI process and auto docker image build-up.
The test files are updated to enable device specific tests, as well as the CPU test cases.

Checklist

Format your code according to the Code Formatting with Pre-Commit.
Add unit tests as outlined in the Running Unit Tests.
Update documentation / docstrings / example tutorials as needed, according to Writing Documentation.
Provide throughput / latency benchmark results and accuracy evaluation results as needed, according to Benchmark and Profiling and Accuracy Results.
For reviewers: If you haven't made any contributions to this PR and are only assisting with merging the main branch, please remove yourself as a co-author when merging the PR.
Please feel free to join our Slack channel at https://slack.sglang.ai to discuss your PR.

docker/Dockerfile.xeon

mingfeima · 2025-05-20T12:13:39Z

can docker file accepts arguements? if yes, we can change TORCH_VERSION as an arguement and set it default to be 2.6

ZailiWang · 2025-05-20T13:45:01Z

can docker file accepts arguements? if yes, we can change TORCH_VERSION as an arguement and set it default to be 2.6

Yeah, by setting ARG VER_TORCH=2.6.0 in the Dockerfile (L5 in the added Dockerfile in this PR), it is a settable argument and can be changed by e.g. --build-arg VER_TORCH=2.7.0 in docker build command.

zhyncs · 2025-05-23T09:43:31Z

ref #6408 (comment)

mingfeima · 2025-05-28T02:01:56Z

@ZailiWang can we satisfy:

skip sgl-kernel cpu tests when CPU doesn't have AMX
skip certain tests when vllm is not installed (since some test cases, e.g. rope, uses reference from sglang)

mingfeima · 2025-05-28T02:03:55Z

@ZailiWang change the tile and provide some details of what this PR is for?

ZailiWang · 2025-05-28T03:10:35Z

@ZailiWang can we satisfy:

skip sgl-kernel cpu tests when CPU doesn't have AMX

skip certain tests when vllm is not installed (since some test cases, e.g. rope, uses reference from sglang)

Hi @zhyncs ,
We have drafted this PR for CPU CI and docker env build-up. Yet there are still something that we are unsure. Your comment and guidance is highly appreciated!

For satisfying this

skip sgl-kernel cpu tests when CPU doesn't have AMX

We added a Check AMX Support step in CPU CI process and CPU test cases are triggered only when the check is passed here. Is this good or would you comment with some better solution?

For CPU docker build-up, we created release-docker-xeon.yml here by imitating "release-docker" files for other devices. Please correct if anything still required to be updated in this file.
In 2 CPU test cases (test_rope.py and test_topk.py), vllm implementation is still invoked when calculating the reference results. Since in SGLang we do not install vllm as a dependency now, such errors would be thrown:

File "/sgl-workspace/sglang/python/sglang/srt/layers/quantization/utils.py", line 14, in <module>
    from vllm._custom_ops import scaled_fp8_quant
ModuleNotFoundError: No module named 'vllm'

File "/sgl-workspace/sglang/python/sglang/srt/layers/rotary_embedding.py", line 87, in __init__
    from vllm._custom_ops import rotary_embedding
ModuleNotFoundError: No module named 'vllm'

Hence these 2 test cases are not added into run_suite.py yet.

ZailiWang · 2025-06-05T01:56:12Z

In 2 CPU test cases (test_rope.py and test_topk.py), vllm implementation is still invoked when calculating the reference results. Since in SGLang we do not install vllm as a dependency now, such errors would be thrown:

We have another PR to mitigate vllm dependency issue now. Will add back the 2 test cases when #6614 merged.

zhyncs · 2025-06-05T04:25:36Z

.github/workflows/pr-test-xeon.yml

+jobs:
+  build-test:
+    if: github.event_name == 'pull_request'
+    runs-on: spr-node


Currently this has not been set up in the sglang repo as a self-hosted runner.

.github/workflows/pr-test-xeon.yml

…ect#6458) Co-authored-by: diwei sun <[email protected]> Co-authored-by: Yineng Zhang <[email protected]>

Merge branch 'sgl_20250610_sync_tag047 of [email protected]:Theta/SGLang.git into main https://code.alipay.com/Theta/SGLang/pull_requests/52 Reviewed-by: 剑川 <[email protected]> * [Bugfix] Fix slice operation when chunk size mismatch (sgl-project#6697) * [Bugfix] Fix ChatCompletion endpoint of mini_lb when stream is set (sgl-project#6703) * [CI] Fix setup of disaggregation with different tp (sgl-project#6706) * [PD] Remove Unnecessary Exception Handling for FastQueue.get() (sgl-project#6712) * Fuse routed_scaling_factor in DeepSeek (sgl-project#6710) * Overlap two kernels in DeepSeek with communication (sgl-project#6711) * Minor refactor two-batch overlap (sgl-project#6682) * Speed up when having padding tokens two-batch overlap (sgl-project#6668) * [Feature] Support Flashinfer fp8 blockwise GEMM kernel on Blackwell (sgl-project#6479) * Fix LoRA bench (sgl-project#6719) * temp * Fix PP for Qwen3 MoE (sgl-project#6709) * [feat] triton kernel for get_last_loc (sgl-project#6676) * [fix] more mem for draft_extend cuda_graph (sgl-project#6726) * [PD] bug fix: Update status if nixl receiver send a a dummy req. (sgl-project#6720) * Tune memory arguments on B200 (sgl-project#6718) * Add DeepSeek-R1-0528 function call chat template (sgl-project#6725) * refactor(tool call): Fix BaseFormatDetector tool_index issue and refactor `parse_streaming_increment` (sgl-project#6715) * Add draft extend CUDA graph for Triton backend (sgl-project#6705) * refactor apply_w8a8_block_fp8_linear in fp (sgl-project#6545) * [PD] Support completion endpoint (sgl-project#6729) * PD Rust LB (PO2) (sgl-project#6437) * Super tiny enable sole usage of expert distribution metrics and update doc (sgl-project#6680) * Support picking variants of EPLB algorithms (sgl-project#6728) * Support tuning DeepEP configs (sgl-project#6742) * [test] add ut and bm for get_last_loc (sgl-project#6746) * Fix mem_fraction_static for AMD CI (sgl-project#6748) * [fix][RL] Fix DeepSeekV3ForCausalLM.post_load_weights for multiple update weight (sgl-project#6265) * Improve EPLB logical to physical dispatch map (sgl-project#6727) * Update DeepSeek-R1-0528 function call chat template (sgl-project#6765) * [PD] Optimize time out logic and add env var doc for mooncake (sgl-project#6761) * Fix aiohttp 'Chunk too big' in bench_serving (sgl-project#6737) * Support sliding window in triton backend (sgl-project#6509) * Fix shared experts fusion error (sgl-project#6289) * Fix one bug in the grouped-gemm triton kernel (sgl-project#6772) * update llama4 chat template and pythonic parser (sgl-project#6679) * feat(tool call): Enhance Llama32Detector for improved JSON parsing in non-stream (sgl-project#6784) * Support token-level quantization for EP MoE (sgl-project#6782) * Temporarily lower mmlu threshold for triton sliding window backend (sgl-project#6785) * ci: relax test_function_call_required (sgl-project#6786) * Add intel_amx backend for Radix Attention for CPU (sgl-project#6408) * Fix incorrect LoRA weight loading for fused gate_up_proj (sgl-project#6734) * fix(PD-disaggregation): Can not get local ip (sgl-project#6792) * [FIX] mmmu bench serving result display error (sgl-project#6525) (sgl-project#6791) * Bump torch to 2.7.0 (sgl-project#6788) * chore: bump sgl-kernel v0.1.5 (sgl-project#6794) * Improve profiler and integrate profiler in bench_one_batch_server (sgl-project#6787) * chore: upgrade sgl-kernel v0.1.5 (sgl-project#6795) * [Minor] Always append newline after image token when parsing chat message (sgl-project#6797) * Update CI tests for Llama4 models (sgl-project#6421) * [Feat] Enable PDL automatically on Hopper architecture (sgl-project#5981) * chore: update blackwell docker (sgl-project#6800) * misc: cache is_hopper_arch (sgl-project#6799) * Remove contiguous before Flashinfer groupwise fp8 gemm (sgl-project#6804) * Correctly abort the failed grammar requests & Improve the handling of abort (sgl-project#6803) * [EP] Add cuda kernel for moe_ep_pre_reorder (sgl-project#6699) * Add draft extend CUDA graph for flashinfer backend (sgl-project#6805) * Refactor CustomOp to avoid confusing bugs (sgl-project#5382) * Tiny log prefill time (sgl-project#6780) * Tiny fix EPLB assertion about rebalancing period and recorder window size (sgl-project#6813) * Add simple utility to dump tensors for debugging (sgl-project#6815) * Fix profiles do not have consistent names (sgl-project#6811) * Speed up rebalancing when using non-static dispatch algorithms (sgl-project#6812) * [1/2] Add Kernel support for Cutlass based Fused FP4 MoE (sgl-project#6093) * [Router] Fix k8s Service Discovery (sgl-project#6766) * Add CPU optimized kernels for topk and rope fusions (sgl-project#6456) * fix new_page_count_next_decode (sgl-project#6671) * Fix wrong weight reference in dynamic EPLB (sgl-project#6818) * Minor add metrics to expert location updater (sgl-project#6816) * [Refactor] Rename `n_share_experts_fusion` as `num_fused_shared_experts` (sgl-project#6735) * [FEAT] Add transformers backend support (sgl-project#5929) * [fix] recover auto-dispatch for rmsnorm and rope (sgl-project#6745) * fix ep_moe_reorder kernel bugs (sgl-project#6858) * [Refactor] Multimodal data processing for VLM (sgl-project#6659) * Decoder-only Scoring API (sgl-project#6460) * feat: add dp-rank to KV events (sgl-project#6852) * Set `num_fused_shared_experts` as `num_shared_experts` when shared_experts fusion is not disabled (sgl-project#6736) * Fix one missing arg in DeepEP (sgl-project#6878) * Support LoRA in TestOpenAIVisionServer and fix fused kv_proj loading bug. (sgl-project#6861) * support 1 shot allreduce in 1-node and 2-node using mscclpp (sgl-project#6277) * Fix Qwen3MoE missing token padding optimization (sgl-project#6820) * Tiny update error hints (sgl-project#6846) * Support layerwise rebalancing experts (sgl-project#6851) * Tiny allow profiler API to auto create directory (sgl-project#6865) * Support Blackwell DeepEP docker images (sgl-project#6868) * [EP] Add cuda kernel for moe_ep_post_reorder (sgl-project#6837) * [theta]merge 0605 * oai: fix openAI client error with single request via batch api (sgl-project#6170) * [PD] Fix potential perf spike caused by tracker gc and optimize doc (sgl-project#6764) * Use deepgemm instead of triton for fused_qkv_a_proj_with_mqa (sgl-project#6890) * [CUTLASS-FP4-MOE] Introduce CutlassMoEParams class for easy initialization of Cutlass Grouped Gems Metadata (sgl-project#6887) * bugfix(OAI): Fix image_data processing for jinja chat templates (sgl-project#6877) * [CPU] enable CI for PRs, add Dockerfile and auto build task (sgl-project#6458) * AITER backend extension and workload optimizations (sgl-project#6838) * [theta]merge * [theta]merge * [Feature] Support Flashinfer fmha on Blackwell (sgl-project#6930) * Fix a bug in abort & Improve docstrings for abort (sgl-project#6931) * Tiny support customize DeepEP max dispatch tokens per rank (sgl-project#6934) * Sync the changes on cuda graph runners (sgl-project#6932) * [PD] Optimize transfer queue forward logic for dummy rank (sgl-project#6922) * [Refactor] image data process in bench_serving (sgl-project#6879) * [fix] logical_to_all_physical_map index 256 is out of bounds in EP parallel. (sgl-project#6767) * Add triton fused moe kernel config for E=257 on B200 (sgl-project#6939) * [sgl-kernel] update deepgemm (sgl-project#6942) * chore: bump sgl-kernel v0.1.6 (sgl-project#6943) * Minor compile fused topk (sgl-project#6944) * [Bugfix] pipeline parallelism and Eagle Qwen2 (sgl-project#6910) * Tiny re-introduce profile id logging (sgl-project#6912) * Add triton version as a fused_moe_triton config search key to avoid performace decrease in different Triton version (sgl-project#5955) * reduce torch.zeros overhead in moe align block size kernel (sgl-project#6369) * chore: upgrade sgl-kernel v0.1.6 (sgl-project#6945) * add fbgemm moe grouped gemm kernel benchmark (sgl-project#6924) * [Docker] Add docker file for SGL Router (sgl-project#6915) * Disabling mixed chunked prefill when eagle is enabled (sgl-project#6874) * Add canary for EPLB rebalancing (sgl-project#6895) * Refactor global_server_args_dict (sgl-project#6866) * Fuse routed scaling factor in topk_reduce kernel (sgl-project#6220) * Update server timeout time in AMD CI. (sgl-project#6953) * [misc] add is_cpu() (sgl-project#6950) * Add H20 fused MoE kernel tuning configs for DeepSeek-R1/V3 (sgl-project#6885) * Add a CUDA kernel for fusing mapping and weighted sum for MoE. (sgl-project#6916) * chore: bump sgl-kernel v0.1.6.post1 (sgl-project#6955) * chore: upgrade sgl-kernel v0.1.6.post1 (sgl-project#6957) * [DeepseekR1-FP4] Add Support for nvidia/DeepSeekR1-FP4 model (sgl-project#6853) * Revert "Fuse routed scaling factor in topk_reduce kernel (sgl-project#6220)" (sgl-project#6968) * [AMD] Add more tests to per-commit-amd (sgl-project#6926) * chore: bump sgl-kernel v0.1.7 (sgl-project#6963) * Slightly improve the sampler to skip unnecessary steps (sgl-project#6956) * rebase h20 fused_moe config (sgl-project#6966) * Fix CI and triton moe Configs (sgl-project#6974) * Remove unnecessary kernels of num_token_non_padded (sgl-project#6965) * Extend cuda graph capture bs for B200 (sgl-project#6937) * Fuse routed scaling factor in deepseek (sgl-project#6970) * Sync cuda graph runners (sgl-project#6976) * Fix draft extend ut stability with flush cache (sgl-project#6979) * Fix triton sliding window test case (sgl-project#6981) * Fix expert distribution dumping causes OOM (sgl-project#6967) * Minor remove one kernel for DeepSeek (sgl-project#6977) * [perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 (sgl-project#6929) * Enable more unit tests for AMD CI. (sgl-project#6983) * Use torch.compile to fuse flash attention decode metadata preparation (sgl-project#6973) * Eliminate stream sync to speed up LoRA batch init (sgl-project#6960) * support qwen3 emebedding (sgl-project#6990) * Fix torch profiler bugs for bench_offline_throughput.py (sgl-project#6557) * chore: upgrade flashinfer v0.2.6.post1 jit (sgl-project#6958) * cleanup tmp dir (sgl-project#7007) * chore: update pr test xeon (sgl-project#7008) * Fix cutlass MLA gets almost zero accuracy (sgl-project#6998) * Update amd nightly models CI. (sgl-project#6992) * feat: add direct routing strategy to DP worker (sgl-project#6884) * Fallback to lower triton version for unfound fused moe configs (sgl-project#7013) * Fix torchvision version for Blackwell (sgl-project#7015) * Simplify prepare_extend_after_decode (sgl-project#6987) * Migrate to assertEqual (sgl-project#6741) * Fix torch version in blackwell dockerfile (sgl-project#7017) * chore: update pr test xeon (sgl-project#7018) * Update default settings for blackwell (sgl-project#7023) * Support both approximate and exact expert distribution collection (sgl-project#6964) * Add decode req pool (sgl-project#6980) * [theta]merge 0610 * [theta]merge 0610 * [CI] Add CI workflow for sgl-router docker build (sgl-project#7027) * Fix fused_moe triton configs (sgl-project#7029) * CPU: map changes from developing branch in sgl-kernel (sgl-project#6833) * chore: bump v0.4.7 (sgl-project#7038) * Update README.md (sgl-project#7040)

DiweiSun and others added 7 commits May 11, 2025 19:52

Create Dockerfil.xeon

e04c2d9

Rename Dockerfil.xeon to Dockerfile.xeon

0b05058

enable device auto detection for cuda/cpu/intel-gpu/rocm

96b7fee

Delete docker/Dockerfile.xeon

a6328b3

lint format fix

f405aac

Merge branch 'main' into molly/cpu_ut

47f46e3

add Dockerfile for Xeon

c397a5f

ZailiWang requested review from ByronHsu, HaiShaw and zhyncs as code owners May 20, 2025 08:59

ZailiWang commented May 20, 2025

View reviewed changes

docker/Dockerfile.xeon Outdated Show resolved Hide resolved

Update run_suite.py

6db5e06

mingfeima requested changes May 20, 2025

View reviewed changes

docker/Dockerfile.xeon Outdated Show resolved Hide resolved

mingfeima marked this pull request as draft May 20, 2025 11:58

mingfeima mentioned this pull request May 20, 2025

[Feature] RFC for adding CPU support for SGLang #2807

Closed

8 tasks

Merge branch 'sgl-project:main' into main

16e3f07

mingfeima added intel cpu cpu backend performance optimization labels May 21, 2025

ZailiWang and others added 3 commits May 21, 2025 11:02

add autotask yml file

9c8b3b8

Merge branch 'main' into main

7bde01a

Merge branch 'main' into main

0856e0c

chunyuan-w mentioned this pull request May 23, 2025

[CPU] Bind threads and numa node for each TP rank #6549

Merged

ZailiWang added 2 commits May 23, 2025 16:05

Merge branch 'main' into main

3f9c509

Merge branch 'main' into main

850e128

zhyncs added the high priority label May 23, 2025

ZailiWang and others added 2 commits May 23, 2025 17:51

replace setup.py

e97287b

Merge branch 'main' into main

ef631d2

mingfeima marked this pull request as ready for review May 28, 2025 02:03

mingfeima requested review from Ying1123 and merrymercy as code owners May 28, 2025 02:03

ZailiWang changed the title ~~Add Dockerfile for Xeon~~ [CPU] enable CI for PRs, add Dockerfile and auto build task May 28, 2025

ZailiWang added 2 commits May 28, 2025 10:47

fix lint error

232b5c7

Merge branch 'main' of https://github.com/ZailiWang/sglang

2bace09

Merge branch 'main' into main

aedeae2

zhyncs self-assigned this May 28, 2025

zhyncs had a problem deploying to prod May 29, 2025 07:26 — with GitHub Actions Error

zhyncs added 2 commits May 30, 2025 01:12

Merge branch 'main' into main

5c79c04

Merge branch 'main' into main

9b0087b

zhyncs had a problem deploying to prod June 4, 2025 04:16 — with GitHub Actions Error

zhyncs reviewed Jun 5, 2025

View reviewed changes

.github/workflows/pr-test-xeon.yml Outdated Show resolved Hide resolved

zhyncs reviewed Jun 5, 2025

View reviewed changes

.github/workflows/pr-test-xeon.yml Outdated Show resolved Hide resolved

zhyncs reviewed Jun 5, 2025

View reviewed changes

.github/workflows/pr-test-xeon.yml Outdated Show resolved Hide resolved

zhyncs added 4 commits June 5, 2025 13:42

Merge branch 'main' into main

a3c804b

upd

fd13d4b

upd

9ee7a2c

upd

1cf96d7

zhyncs approved these changes Jun 5, 2025

View reviewed changes

zhyncs merged commit 562f279 into sgl-project:main Jun 5, 2025

jianan-gu pushed a commit to jianan-gu/sglang that referenced this pull request Jun 12, 2025

[CPU] enable CI for PRs, add Dockerfile and auto build task (sgl-proj…

cc09f6d

…ect#6458) Co-authored-by: diwei sun <[email protected]> Co-authored-by: Yineng Zhang <[email protected]>

xwu-intel pushed a commit to xwu-intel/sglang that referenced this pull request Jun 17, 2025

[CPU] enable CI for PRs, add Dockerfile and auto build task (sgl-proj…

fbd7446

…ect#6458) Co-authored-by: diwei sun <[email protected]> Co-authored-by: Yineng Zhang <[email protected]>

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[CPU] enable CI for PRs, add Dockerfile and auto build task #6458

[CPU] enable CI for PRs, add Dockerfile and auto build task #6458

Uh oh!

ZailiWang commented May 20, 2025 •

edited

Loading

Uh oh!

Uh oh!

Uh oh!

mingfeima commented May 20, 2025

Uh oh!

ZailiWang commented May 20, 2025 •

edited

Loading

Uh oh!

zhyncs commented May 23, 2025

Uh oh!

mingfeima commented May 28, 2025

Uh oh!

mingfeima commented May 28, 2025

Uh oh!

ZailiWang commented May 28, 2025 •

edited

Loading

Uh oh!

ZailiWang commented Jun 5, 2025 •

edited

Loading

Uh oh!

zhyncs Jun 5, 2025

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

[CPU] enable CI for PRs, add Dockerfile and auto build task #6458

[CPU] enable CI for PRs, add Dockerfile and auto build task #6458

Uh oh!

Conversation

ZailiWang commented May 20, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Motivation

Modifications

Checklist

Uh oh!

Uh oh!

Uh oh!

mingfeima commented May 20, 2025

Uh oh!

ZailiWang commented May 20, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

zhyncs commented May 23, 2025

Uh oh!

mingfeima commented May 28, 2025

Uh oh!

mingfeima commented May 28, 2025

Uh oh!

ZailiWang commented May 28, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

ZailiWang commented Jun 5, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

zhyncs Jun 5, 2025

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

ZailiWang commented May 20, 2025 •

edited

Loading

ZailiWang commented May 20, 2025 •

edited

Loading

ZailiWang commented May 28, 2025 •

edited

Loading

ZailiWang commented Jun 5, 2025 •

edited

Loading