Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1510 commits
Select commit Hold shift + click to select a range
e332e1b
[Fix] Don't write conv state from the fused KDA verify kernel (#39524)
mmangkad Sep 22, 2026
9eda772
[Test] Handle tied top-k indices in graph-pool logprob regression (#4…
ch-wan Sep 22, 2026
bc30fa1
[AMD][Fix] AgentX HIP TPOT regression when SGLANG_SIMULATE_ACC_LEN is…
yichiche Sep 22, 2026
15eba3b
Feat: Add TensorCast storage as a new HiCache backend (#27265)
zhou-yuhan Sep 22, 2026
a1b2b97
[diffusion] CI: restore public Qwen-Image 2.1 TP2 E2E coverage (#40507)
mickqian Sep 22, 2026
5f9c6b9
[diffusion] fix: separate a use-scoped layerwise release from release…
mickqian Sep 22, 2026
90cf471
[AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA …
Jacob0226 Sep 22, 2026
1d02549
[Test] Set DP size in the mocked Metal profiler test (#40667)
ch-wan Sep 22, 2026
27f796c
[sgl-router] Fix readiness, IPv6 discovery, logging, and model valida…
sherlockwu Sep 22, 2026
59a723e
[sgl-router] refactor - SLO ordering for bucket selection (#40292)
sherlockwu Sep 22, 2026
56fee88
fix(moe): support Llama4 NVFP4 router input weights on SM120 (#35504)
janbernloehr Sep 22, 2026
15ba54b
perf(engine): avoid timed waits for Engine responses (#39486)
jthomson04 Sep 22, 2026
b44e248
[AMD] [GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950 …
Raiden-Makoto Sep 22, 2026
018b73c
[PD] Pack draft KV head slices for DCP transfers (#40500)
kpham-sgl Sep 22, 2026
e1daf68
[AMD] [GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark…
Raiden-Makoto Sep 22, 2026
877a293
[Benchmark] Optionally clear HiCache storage between cases (#40659)
metamergebot Sep 22, 2026
a9f02b0
[NPU] Fix xgrammar apply_vocab_mask device dispatch to use torch.ops.…
zhujianwei-ops Sep 22, 2026
2032f3a
[Router] Abort the engine when a client disconnects mid-request (#39461)
Kangyan-Zhou Sep 22, 2026
9b59fc5
[ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP…
YAMY1234 Sep 22, 2026
b01961e
[LFM2-VL] Add DSpark speculative decoding (#40651)
tugot17 Sep 22, 2026
264da63
[AMD] Update ROCm AITER pin to acf8fdf9 (#39965)
kangwangamd Sep 22, 2026
095e451
[AMD] [GLM-5.3-Flash Day 0] Route mHC through AITER on gfx950 (#38545)
Raiden-Makoto Sep 22, 2026
04c0913
[HiSparse] Add MHA hisparse support for MiniMax M3 (#31446)
bingps Sep 22, 2026
a0781f2
[Docs] GLM-5.3/5.3-Flash cookbooks: enable reasoning/tool-call parser…
b8zhong Sep 22, 2026
bc22e1d
[DSpark] Fix draft CUDA graph stream explosion (#40658)
kpham-sgl Sep 22, 2026
4c81cd1
[KDA] Fix missing beta sigmoid in PTX prefill (#40685)
mmangkad Sep 22, 2026
8ac19cc
[AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbo…
kevin-mii Sep 22, 2026
c948114
[DOC] Update quickstart guide to use `sglang serve` for launching the…
jshn9515 Sep 22, 2026
1815e49
[observability] Fix negative queue_time for retracted requests (#39312)
surajm20061998 Sep 22, 2026
771c9d7
[PD] Validate Mooncake EFA allocator compatibility (#39973)
nvpohanh Sep 22, 2026
6412ad8
[AMD] Pack Qwen3.5 GDN input projections on ROCm (#39902)
zijiecode Sep 22, 2026
d00adfd
[AMD] Fix deferred Kimi-K3 forget gate in fused in-projection (#39525)
tomjen12 Sep 22, 2026
367e370
avoid host sync in DSpark prefill slot expansion (#40111)
inkcherry Sep 22, 2026
c2f14bf
[AMD] Use exact CU share for gfx950 segment-plan headroom (#39503)
chuyeh Sep 22, 2026
790551c
[AMD] Skip full-vocab softmax in EAGLE topk==1 draft on ROCm (#35872)
yichiche Sep 22, 2026
70a2bb2
[AMD][DI] Keep loopback in UCX_NET_DEVICES (#40641)
Lzy17 Sep 22, 2026
3f00fb7
[AMD] Register unified KV page-zeroing test in PR CI (#40123)
michaelzhang-ai Sep 22, 2026
debbb5c
[AMD] Drop the redundant scale zero-fill before AITER per-tensor FP8 …
siliangchen-amd Sep 22, 2026
9d58189
[diffusion] attention: add fp8_fa_sm120 FP8 backend for SM120 GPUs (#…
emre570 Sep 22, 2026
8ef6d31
[PD] Preserve abort ACKs until in-flight KV transfers drain (#40645)
cctry Sep 22, 2026
cb8dab0
[NPU] [DOC] Remove duplicated features in npu docs (#40714)
amote-i Sep 22, 2026
861b11f
[PD] Simplify late-abort quiescent ack branch to else (#40711)
ShangmingCai Sep 22, 2026
2524b61
[diffusion] feat: allow a component use retain its layerwise resident…
mickqian Sep 22, 2026
ddebc52
[kv-shard 3/4] Enable Control plane (#38468)
Shunkangz Sep 22, 2026
bc3a63e
[Fix] Missing SWA eviction during decode preallocation (#40309)
wookjeHan Sep 22, 2026
244db08
[NPU][CI] Constrain evalscope dependency versions (#40133)
pllimax Sep 22, 2026
4a1b69a
[diffusion] docs: correct the resident-layer help text to match its s…
mickqian Sep 22, 2026
c79510c
[DSv4.1] Score prefill consumer index layers on candidate blocks with…
yuan-luo Sep 22, 2026
6fd98c9
[Qwen3.8-Next] Pipeline-parallel serving and PD-prefill MTP for Qwen4…
YAMY1234 Sep 22, 2026
a7b96fd
[ci] run cpu ci for renderer-only changes (#40636)
sagearc Sep 22, 2026
8ab21c8
[ci] publish renderer image (#40639)
sagearc Sep 22, 2026
b4701d6
[rust-renderer] decouple renderer sampling from protocols (#40747)
sagearc Sep 22, 2026
66454d7
[Feature] Support --tokenizer-worker-num > 1 in the offline Engine AP…
peilii Sep 22, 2026
91c329c
Take the model config out of the parallel group build, and finish ret…
ch-wan Sep 22, 2026
6f4c2b9
[HiCache] Demote internal-node mamba states on write_back eviction (#…
ShangmingCai Sep 22, 2026
db73f35
Fix the Inkling per-expert sync test and collect it in the weekly CPU…
ch-wan Sep 22, 2026
5c18780
[Router] Keep e2e workers inside the job's CUDA_VISIBLE_DEVICES allot…
Kangyan-Zhou Sep 22, 2026
720617b
[AMD] Reuse KV gather indices across ASM context prefill layers (#39901)
zijiecode Sep 22, 2026
4cbf290
[Fix] Handle chunked paged MQA metadata in DSV4.1 eager forwards (#40…
Oasis-Git Sep 22, 2026
d6cc283
[AMD] [GLM-5.3-Flash Day 0] Enable zero-RoPE TileLang DSA on gfx950 (…
Raiden-Makoto Sep 22, 2026
077c319
[AMD] [GLM-5.3-Flash Day 0] Enable speculative decoding (MTP) on ROCm…
Jacob0226 Sep 22, 2026
3afdde5
[AMD] [GLM-5.3-Flash Day 0] Enable the k-pool DSA indexer on gfx950 (…
Jacob0226 Sep 22, 2026
542c817
[Docs] Fix benchmark table column overflow in cookbook deployment pan…
JustinTong0323 Sep 22, 2026
c08ffef
docs: sync LMSYS SGLang blog cards (#40655)
sglang-bot Sep 22, 2026
d72629e
Add GB200/GB300 hardware to Qwen3.5 (#40770)
faradawn Sep 22, 2026
9a53f75
[mem_cache] Remove the experimental C++ radix tree (#40775)
hnyls2002 Sep 22, 2026
01275aa
fix(grpc): expose native response timeout as a server argument (#40644)
jain-ria Sep 22, 2026
c19dc43
[AMD] [GLM-5.3-Flash Day 0] Load the MXFP4 MTP draft layer (#39779)
Jacob0226 Sep 22, 2026
b77833c
[MM] Keep scheduler padding in packed token arrays (#40357)
Jialin Sep 22, 2026
19ee4b5
fix(function_call): buffer complete DeepSeek DSML invokes (#39632)
kflansburg Sep 22, 2026
d34f7b2
[mem_cache] Clean up SWA/Mamba radix cache leftovers and drop SGLANG_…
hnyls2002 Sep 22, 2026
7b977ce
[Fix] Decide the MoE padded-row bound from the layer scatter mode (#4…
mmangkad Sep 22, 2026
4ce2354
[HiCache] Remove the unused HiRadixCache (#40787)
hnyls2002 Sep 22, 2026
06008c1
[dsv4.1]Optimize FP4 indexer by skipping invisible tiles (#40431)
shiyu7 Sep 22, 2026
0cd8be3
ci: stop Runner Utilization Report from draining the shared API quota…
alisonshao Sep 23, 2026
78980a3
[misc] Remove deprecated endpoints, env vars and aliases past two rel…
hnyls2002 Sep 23, 2026
40048f6
[dLLM] feat: support DiffusionGemma serving (#34061)
rwang5203 Sep 23, 2026
2e35685
[sgl-router] Launch reorg routing with existing policy options (#40766)
sherlockwu Sep 23, 2026
2bc435d
[HiCache] fix: bound the controller reset join so a stalled storage t…
alphabetc1 Sep 23, 2026
c2bc456
[chore] point agents at the cookbook before test configs (#40818)
mickqian Sep 23, 2026
224a247
[HiSparse] ci: add cross-directory rerun test group (#40749)
alphabetc1 Sep 23, 2026
525f140
[NPU]fix ci hicache oom (#40824)
zhaozx-cn Sep 23, 2026
1d59ce7
fix: partition selective CI reruns into matrix jobs (#40760)
alphabetc1 Sep 23, 2026
973fb44
[Diffusion] Support MiniMax-H3 PDD(Parallel Decoding Distillation) in…
IPostYellow Sep 23, 2026
9d26765
Fix NIXL transfer of MXFP8 KV block scales (#40792)
ekzhang Sep 23, 2026
ced0a2c
[Hisparse] fix: account for MiniMax HiSparse full-pool memory (#40743)
alphabetc1 Sep 23, 2026
5b8d8b2
[Fix] Keep Inkling automatic tool grammar active across the response …
jshanson7 Sep 23, 2026
58988be
[AMD] Add diffusion (Wan2.2) extras to gfx1151 Docker image (#40844)
yichiche Sep 23, 2026
28be39f
[deepep_v2] support GLM-5.3-Flash (Glm5NextForConditionalGeneration) …
whn09 Sep 23, 2026
a3d11f9
[diffusion] feat: support permanent lifetime for layerwise resident l…
mickqian Sep 23, 2026
66ce8c5
[HiCache] ci: add HiCache and unified radix rerun group (#40831)
alphabetc1 Sep 23, 2026
aa0feca
[AMD][DI][CI] Use a node-local model cache on the SPUR cluster (#40813)
Lzy17 Sep 23, 2026
eb5e8c8
Speculative Decoding with NGRAM support for XPU (#31362)
ANSHUMAN87 Sep 23, 2026
abca3b2
[Intel GPU] Add DeepSeek-V2-Lite-Chat-FP8 gsm8k e2e accuracy nightly …
polisettyvarma Sep 23, 2026
86cb4a0
[CI] Update GLM-5.3-Flash H200/B200 test args (#40862)
mmangkad Sep 23, 2026
0b0f947
[CPU] Add fused_sigmod_mul_cpu operators to the Meta Muse Glimmer mod…
nzr-niu Sep 23, 2026
d4dcce1
[npu] decoding procedure optimization on qwen3.5/3.6 (#35958)
MatsueYu Sep 23, 2026
48c3854
[ROCm] feat: enable aiter allreduce fusion for GLM models (#39790)
RuibinCheung Sep 23, 2026
aa0b65c
[AMD] Fix DeepSeek-V4 accuracy by not passing num_token_non_padded to…
At1a8 Sep 23, 2026
de123f3
[3/N] elastic-ep: Recapture decode CUDA graphs after scale-up (#33723)
zackyoray Sep 23, 2026
401d5ae
[AMD] Drop the unreachable vLLM fallback from ROCm FP8 activation qua…
siliangchen-amd Sep 23, 2026
a89f849
[ROCm] Fuse the MLA q absorb into the RoPE + KV-write kernel on gfx95…
xiaobochen-amd Sep 23, 2026
4cd63da
[Unified Cache] Dedup replicated MLA/DSA KV in the UMBP direct linker…
TianDi101 Sep 23, 2026
58f622e
Reduce decode bootstrap latency with request-owned speculative KV (#3…
inkcherry Sep 23, 2026
2d25767
[NPU] Enable piecewise CUDA graph support on NPU (#28417)
hanwlax Sep 23, 2026
abef3ef
[Diffusion] Fuse rounded SwiGLU for quantized MiniMax-H3 MLPs (#40378)
BBuf Sep 23, 2026
3fdd63a
[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache c…
iridiumine Sep 23, 2026
16e353d
[NPU] Skip fused gmm1+swiglu for swiglu_limit (SiLU-with-clamp) check…
iridiumine Sep 23, 2026
172b1b4
[Diffusion] Accelerate Cosmos3 Edge on Hopper with lossless fusions (…
BBuf Sep 23, 2026
4e60d70
[Diffusion] Fuse Joy Image Edit QKV concatenation and avoid QK copies…
BBuf Sep 23, 2026
8993f79
[Diffusion] Fuse lossless SenseNova RoPE for 5% faster H200 inference…
BBuf Sep 23, 2026
c3e9852
[Diffusion] Remove unused standalone benchmarks and deduplicate kerne…
BBuf Sep 23, 2026
5c154c2
Revert " [NPU] Enable piecewise CUDA graph support on NPU" (#40895)
iforgetmyname Sep 23, 2026
7fef014
feat(kv-hints): add kv hint envelope to request transport (#38891)
linhu-nv Sep 23, 2026
f2eebd5
[RL] Fix Kimi K3 expert-count lookup for routed-expert capture (#40700)
ByronHsu Sep 23, 2026
f2f223e
Allow attention layers to opt out of the prefill wrapper (#40683)
metamergebot Sep 23, 2026
4bb5611
[Fix] Give the full prefill CUDA graph replay view the captured bucke…
metamergebot Sep 23, 2026
890a960
[Metrics] Log forward and forward+idle occupancy over total wall time…
metamergebot Sep 23, 2026
2909f84
[JIT] Add an occupancy-preserving L1 carveout preference (#40767)
metamergebot Sep 23, 2026
9544585
[mem_cache] Remove unreachable RadixCache paths in KV canary and HiCa…
hnyls2002 Sep 23, 2026
0e80c73
[Fix] Avoid duplicate residual in LongCat MoE shortcut (#40799)
ch-wan Sep 23, 2026
9ebe422
[Fix] Reduce Nemotron MTP attention outputs once (#40800)
ch-wan Sep 23, 2026
379e8f9
[Fix] Capture complete Nemotron auxiliary hidden states (#40801)
ch-wan Sep 23, 2026
701cf7e
[Refactor] Run Nemotron-H DP attention through the standard layer com…
ch-wan Sep 23, 2026
e31cbc0
[Fix] Stop deferring the last layer's FFN all-reduce in five models (…
ch-wan Sep 23, 2026
7dd9640
[Refactor] Compare token layouts instead of group sizes when selectin…
ch-wan Sep 23, 2026
e1048a2
[Refactor] Let LayerCommunicator own the FFN exit in Qwen3-MoE, DeepS…
ch-wan Sep 23, 2026
ab0c31b
[Refactor] Move eleven more MoE models to LayerCommunicator.ffn_exit …
ch-wan Sep 23, 2026
c39a1c0
[mem_cache] Free the rows below the SWA evict floor on all-SWA reques…
hnyls2002 Sep 23, 2026
f766397
[Refactor] Trim server configuration and runtime context comments (#4…
ch-wan Sep 23, 2026
4fa2c9c
MiniMax-M3: run the sparse prefill main attention through AITER Gluon…
zcnrex Sep 23, 2026
f2a1366
MiniMax-M3: wave64 histogram-select decode top-k, and raise kMaxNumBl…
zcnrex Sep 23, 2026
79fec59
[Refactor] Read parallel placement in consumers (#40638)
ch-wan Sep 23, 2026
6fe4b66
[RL] Keep pause_generation and weight updates from deadlocking each o…
yueming-yuan Sep 23, 2026
0010f56
[Test] Remove obsolete configuration migration guards (#40976)
ch-wan Sep 23, 2026
3f69789
ci: reinstall torch/triton left incomplete by a cancelled job (#40819)
alisonshao Sep 23, 2026
208f6f7
[Spec] Support DFLASH for Kimi K3 (#40794)
chromecast56 Sep 23, 2026
542a043
[RL] Release the weight-checker snapshot once compare passes (#37284)
yueming-yuan Sep 23, 2026
ffac53d
[Doc] Add H200 recipes to MiMo-V2.6 cookbook (#40969)
zijiexia Sep 23, 2026
ec75d3d
MiniMax-M3: allocate the lightning-indexer K cache in fp8 on gfx95 (#…
zcnrex Sep 24, 2026
81cb895
[Fix] Recover from stale torch extension locks in every `cpp_extensio…
hnyls2002 Sep 24, 2026
621136e
[mem_cache] Remove unused helpers in mem_cache, storage backends, and…
hnyls2002 Sep 24, 2026
81b8166
Revert "[AMD] Fix DeepSeek-V4 accuracy by not passing num_token_non_p…
At1a8 Sep 24, 2026
03fcbe1
[NPU] Support batch invariant FIA graphs for deterministic inference …
hanwlax Sep 24, 2026
82cdd72
[NPU] [DOC] Remove --enforce-shared-experts-fusion from npu docs (#41…
amote-i Sep 24, 2026
8b5d77c
[AMD] Add a Triton packed sparse decode path for QSA on ROCm (#38876)
yichiche Sep 24, 2026
e98b2f7
[AMD] Speed up Wan2.2 DiT FP8 attention per-tensor quantization (#34695)
yichiche Sep 24, 2026
cc6e8b0
[Fix] Derive per-runner hybrid SWA layer ids on ModelLayerInfo instea…
hnyls2002 Sep 24, 2026
f4b9038
[RL] Keep DSA cuda-graph state and the graph pool intact across TMS p…
yueming-yuan Sep 24, 2026
f4d9d2e
[RL] Add RL weight-update sessions and support updating spec draft ru…
yueming-yuan Sep 24, 2026
5c44214
[XPU] Qwen3.8-flash-next enablement (#37213)
Xia-Weiwen Sep 24, 2026
7798386
[Moe] Honor swiglu_limit clamped activation in flashinfer_cutlass run…
Dovis01 Sep 24, 2026
e2f4fed
[AMD] Critical fix enabling Qwen3.8 FP8: restore dropped fused shared…
yichiche Sep 24, 2026
be495a6
[Fix] Use cached prefix lengths for FlashInfer full-attention ragged …
Jiminator Sep 24, 2026
8eedf61
[AMD][DI][CI] Say which image the MI355X nightly ran on (#41026)
Lzy17 Sep 24, 2026
c3685df
[Sampling] Add selected/support sampling logprob modes (#40932)
nanjiangwill Sep 24, 2026
0dc8b29
[ROCm][DSA] Enable AITER fused FP8 indexer writer (#40710)
fanxingran Sep 24, 2026
7a9feac
[Docs] DeepSeek-V4 MI355X Pro Official PD pairs with DSpark and UMBP …
ichbinblau Sep 24, 2026
efff836
[PD] Add decode host receive for custom transfer backends (#40238)
cctry Sep 24, 2026
90663cc
[DSV4] fix: size the C4 state ring by the page it is addressed by (#4…
alphabetc1 Sep 24, 2026
32290dd
[AMD] Small-M MXFP4 fused-MoE kernel for gfx950 (Qwen) (#40204)
zijiecode Sep 24, 2026
9f6fc55
[AMD][Diffusion] FlyDSL fused norm kernels on wave32 targets (gfx1250…
yctseng0211 Sep 24, 2026
43af9fc
[AMD][DI][CI] Move MI355X disagg nightly to ROCm 10 (#41053)
yctseng0211 Sep 24, 2026
a164c6d
[AMD] Register mem-cache unit tests in PR CI (#40812)
michaelzhang-ai Sep 24, 2026
cae4927
[DSV4] Budget the ratio-2 pair state pool in DSV4PoolConfigurator (#4…
hnyls2002 Sep 24, 2026
1417345
[Experimental] Preserve speculative decoding during prefill across DP…
Oasis-Git Sep 24, 2026
d65503f
[NPU] Remove the LLaDA2.0-mini basic-function test case (#40905)
pllimax Sep 24, 2026
81f2b43
[ci] pr-gate: add generic require-label input and support pull_reques…
AgainstEntropy Sep 24, 2026
41812af
[NPU] Fix DSV4 hard-coding kv dtype (#40510)
cx22757 Sep 24, 2026
174a5f3
[feature] add per-item candidate token scoring and calibration (#40826)
mickqian Sep 24, 2026
261cb82
[Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (#40384)
BBuf Sep 24, 2026
c85df5b
[Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (#40405)
BBuf Sep 24, 2026
bb1c98b
Fix TBO child batch missing dp_spec_prefill_coordination_applied (#41…
mmangkad Sep 24, 2026
d94d784
[AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces gr…
jiejingzhangamd Sep 24, 2026
26c6333
[Diffusion] Fuse lossless LingBot World FP32 normalization (#40425)
BBuf Sep 24, 2026
0a59830
[AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tok…
AMD-yanfeiwang Sep 24, 2026
a69583f
[AMD] Add tuned dsv4 shape (#40996)
akao-amd Sep 24, 2026
c6565c1
[Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% l…
BBuf Sep 24, 2026
ce06a14
[Diffusion] migrate the whole _register_configs from registry.py to t…
ping1jing2 Sep 24, 2026
3177d10
Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3,…
b8zhong Sep 24, 2026
28ec670
[HiCache] Make host reclamation independent of transfer order (#40512)
paulzhang-tm Sep 24, 2026
ea5baf4
[Refactor] Retire the model-specific Kimi K3 kernel namespace (#40922)
BBuf Sep 24, 2026
ec7eb6b
[diffusion] update code owner (#41130)
niehen6174 Sep 24, 2026
4142235
[NPU] Update CANN version to 9.1.0 (#40524)
huangxiaojun15 Sep 24, 2026
cd11037
[PD] Enable deferred decode-side KV release by default (#41023)
ShangmingCai Sep 24, 2026
17a7484
[chore] surface the cookbook to users who pip install sglang (#40866)
mickqian Sep 24, 2026
752801e
fix(sampling): validate sampling_seed is an int within int64 range (#…
Sunt-ing Sep 24, 2026
96bc99c
Merge remote-tracking branch 'origin/main' into pr-35990
niehen6174 Sep 24, 2026
77173ff
[HiCache] Demote SWA KV to host on write_back eviction instead of dro…
ShangmingCai Sep 24, 2026
8ec65e8
[AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (#40387)
XinyuJiangCMU Sep 24, 2026
5bb24e3
[Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and …
zhendonghua Sep 24, 2026
984994e
fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 ima…
trevor-m Sep 24, 2026
78b382b
Support unified memory decode host pools (#39478)
ZYHowell Sep 24, 2026
86b3558
[Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotun…
mmangkad Sep 24, 2026
f404db9
[DP attention] Publish DP buffer sizes from a ForwardBatch (#40858)
hanming-lu Sep 24, 2026
1b03d31
[Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enable…
kpham-sgl Sep 24, 2026
6a14b80
[AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (#41109)
ChangLiu0709 Sep 24, 2026
928e683
[Fix] Complete the deferred FFN all-reduce before a pipeline-parallel…
ch-wan Sep 24, 2026
07a3735
[Fix] Complete the deferred FFN all-reduce before deepstack addition …
ch-wan Sep 24, 2026
5df1667
[Refactor] Pass each layer stack's output through a communicator exit…
ch-wan Sep 24, 2026
62ae032
[Refactor] Share the MoE output all-reduce between models (#41097)
ch-wan Sep 24, 2026
7c9f74c
[Fix] Step-3.5: stop dense layers from summing their output twice und…
ch-wan Sep 24, 2026
a80055d
[Fix] Broadcast requests along attention CP before attention TP (#41083)
ch-wan Sep 24, 2026
5bb850d
[Refactor] Split prepare_attn into a reduction step and per-quant-for…
ch-wan Sep 24, 2026
1446e24
[AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode …
1am9trash Sep 24, 2026
b129504
[qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for ta…
Qiaolin-Yu Sep 24, 2026
0b53305
MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx95…
zcnrex Sep 24, 2026
36f5998
MoE: small-batch sorting path with fused mxfp8 quantisation (#36559)
zcnrex Sep 24, 2026
961404b
[DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (#41090)
alphabetc1 Sep 24, 2026
182f62d
[Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordinat…
hnyls2002 Sep 24, 2026
a1eb691
[Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (#41…
Qiaolin-Yu Sep 24, 2026
0c578d9
[DSV4] Size compressed pools from one per-ratio table in DSV4PoolConf…
hnyls2002 Sep 24, 2026
6bbd689
[Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (#39929)
Apexsf Sep 24, 2026
3ec8630
Fix mixed chunk prefill with DP speculative coordination (#41179)
Oasis-Git Sep 24, 2026
e047e50
Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++…
caiocbr Sep 24, 2026
466e985
[HiCache] Batch buffer-only KV backups within each flush (#40960)
metamergebot Sep 24, 2026
7c5bdb9
[PD] Add a `none` decode retraction backup and subclass seams in the …
metamergebot Sep 24, 2026
2f2f9d1
[Score API] Setwise Scoring Support (#38965)
sundar24295s Sep 24, 2026
a2025b8
[DSV4] Account for FlashMLA physical KV page padding in memory budget…
alphabetc1 Sep 25, 2026
8ca8211
[mem_cache] Drop `is_insert` from `cache_finished_req`; release rows …
hnyls2002 Sep 25, 2026
9dc4c5d
[AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cac…
yichiche Sep 25, 2026
6103a8c
[diffusion] feat: add opt-in SRT prompt enhancement to image and vide…
mickqian Sep 25, 2026
37ebcac
Support XQA backend for SpecDec verify (#32269)
akhilg-nv Sep 25, 2026
7325b38
[PD] Honor gracefully_exit in disaggregation event loops and keep non…
metamergebot Sep 25, 2026
e0b4d66
[Perf] Lazy-load built-in model definitions and nixl_ep at startup (#…
metamergebot Sep 25, 2026
3fc7a66
[CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (#41207)
mmangkad Sep 25, 2026
cec70a4
[Fix] Keep the target's DP sync slot in draft scopes (#41062)
mmangkad Sep 25, 2026
16d1c93
[feat] add a system one compatible /v1/systemone route (#41208)
rwang5203 Sep 25, 2026
3e7e652
fix(openai): reject request-supplied chat_template by default (#28135)
Sunt-ing Sep 25, 2026
ec070ec
[AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (#41159)
RolaoDenthu Sep 25, 2026
9c8340a
[DeepEP v2] Let a model package supply its per-rank prefill dispatch …
metamergebot Sep 25, 2026
0fb699a
[diffusion] model: support Ming-Image Design and Design-Layer (#41067)
mickqian Sep 25, 2026
cbe1377
[HiCache] Give trailing sidecar storage transfers a contiguous prefix…
alphabetc1 Sep 25, 2026
434c2e3
[HiCache] fix: Drain pending backups before internal Mamba write-back…
alphabetc1 Sep 25, 2026
515f5be
[AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (#35619)
kkHuang-amd Sep 25, 2026
51c92e8
[Fix] Complete the all-reduce when the flashinfer fused norm declines…
ch-wan Sep 25, 2026
1409f46
[Fix] Plan NextN / MTP draft layers as one-layer models and fix the B…
ch-wan Sep 25, 2026
402df23
[Fix] Stop counting a deferred FFN sum more than once: replicated TP1…
ch-wan Sep 25, 2026
9d7f44b
[Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and com…
ch-wan Sep 25, 2026
963e9fb
[Refactor] Leave the FFN reduction to the next layer under attention …
ch-wan Sep 25, 2026
b7f6d04
[Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 on…
ch-wan Sep 25, 2026
8f5a636
[Refactor] Build prepare_mlp and the layout moves from named steps (#…
ch-wan Sep 25, 2026
26a3214
[Refactor] Take a layer's last-layer fact from its scatter-mode plan …
ch-wan Sep 25, 2026
3450d68
[Refactor] Decide an FFN exit's completion once and declare the group…
ch-wan Sep 25, 2026
62122c8
[XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter …
arathi-hlab Sep 25, 2026
a749d84
[Diffusion] Fix AttributeError in grouped forward_batch by installing…
CjhHa1 Sep 25, 2026
3cf274f
[diffusion] Format the H3 ComfyUI split error as one f-string
niehen6174 Sep 25, 2026
9b1d9bb
Merge remote-tracking branch 'origin/main' into pr-35990
niehen6174 Sep 25, 2026
0573d93
[diffusion] Load ComfyUI only when a DiT graph is built
niehen6174 Sep 25, 2026
f5ae8ec
[diffusion] Set comfyui_mode on the component-loader stub
niehen6174 Sep 26, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
21 changes: 21 additions & 0 deletions .claude/rules/cookbook-first.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Serving a model? Read its cookbook page first

Before running, deploying, benchmarking, or reproducing a specific model, read
its page under `docs/cookbook/`:

- `docs/cookbook/autoregressive/` — LLM and VLM serving
- `docs/cookbook/diffusion/` — image, video, 3D, robotics
- `docs/cookbook/omni/`, `docs/cookbook/vla/`, `docs/cookbook/specbundle/`,
`docs/cookbook/base/`

That page carries the deployment we recommend to users — GPU count, parallelism
degrees, quantization, and the flags that matter for that model — so it is the
answer to "how should this be served", and the baseline any tuning starts from.

Configs under `test/` are a different thing: pinned, reproducible CI setups for
correctness and regression checks, not deployment advice. Reach for them when
you need an exactly reproducible run, and say which cookbook flags you diverged
from — a measurement taken on a configuration nobody deploys describes nothing.

Keep it true in both directions: when a change alters a model's recommended
deployment, update that model's cookbook page in the same PR.
9 changes: 9 additions & 0 deletions .claude/rules/unit-test-admission.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
paths:
- "test/**/*.py"
- "python/sglang/multimodal_gen/test/**/*.py"
---

# Unit Test Admission Criteria
Expand Down Expand Up @@ -68,5 +69,13 @@ One strong case beats several weak ones: each additional case must guard a
distinct failure mode. Ask "which bug escapes if I delete this case?" -- no
answer means delete it.

New cases join an existing file in the same subsystem by default. Create a new
file only when it needs a different fixture, dependency, owner, or CI contract;
every file pays a separate interpreter-import cost in the CPU gate.

Suite cadence is part of admission: if a failing run cannot be attributed to a
single PR's diff, the test belongs in a nightly or weekly suite rather than a
per-commit lane.

Test mechanics (placement, CI registration, fixtures) live in
[`write-sglang-test`](../skills/write-sglang-test/SKILL.md).
55 changes: 46 additions & 9 deletions .claude/skills/add-jit-kernel/SKILL.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .claude/skills/babysit-pr-to-pass-ci/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ Examples:
$babysit-pr-to-pass-ci
$babysit-pr-to-pass-ci 12345
$babysit-pr-to-pass-ci https://github.com/sgl-project/sglang/pull/12345 pr-test-extra.yml
$babysit-pr-to-pass-ci 12345 --only pr-test-amd-rocm720.yml
$babysit-pr-to-pass-ci 12345 --only pr-test-amd.yml
```

## Start or continue the durable goal
Expand Down
40 changes: 40 additions & 0 deletions .claude/skills/ci-test-audit/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
---
name: ci-test-audit
description: Audit the existing test tree and CI configuration for improvements, using a catalog of patterns previously applied in this repo. Use when asked to audit tests or CI, shrink CI time or cost, find redundant or misplaced tests, clean up a test group, or review whether a CI change follows established practice.
---

# CI / Test Audit

Audits what is already in `test/`, `.github/` and `scripts/ci/`. For writing a new test see
[write-sglang-test](../write-sglang-test/SKILL.md); for how the pipeline dispatches and
gates work see [ci-workflow-guide](../ci-workflow-guide/SKILL.md).

## How to use it

Read [action-items.md](action-items.md) first. It is a catalog of patterns that have
been applied to this repo, each with a way to spot it and example PRs. The catalog is the substance of this skill; everything below is only scaffolding.

The user says what to audit -- a test group, a stage, a workflow file, or a complaint
like "this suite is slow" -- and how thoroughly. Read that scope and decide which
patterns apply.

An audit is only useful if it is specific. "This suite could be trimmed" is not a
finding; "`TestFooLargePage` is a strict subset of `TestFooRetractLargePage`, drop it"
is. Every finding needs the file, the pattern it matches, and what to do.

## Reporting

Group findings by confidence in the claim, not by pattern id. Lead with the ones where
the evidence is in the file you just read; keep the speculative ones separate and say
what you would need to check. Skip anything you are not reasonably sure about -- a long
list of maybes costs more to triage than it saves.

Propose, do not apply. Deleting a test, moving a registration to another stage, and
changing a threshold are all decisions for the user. Land them as separate changes,
since a trim and a threshold change fail for different reasons.

## Keeping the catalog current

Example PRs age. When a pattern shows up in newer work, or a new pattern recurs, update
`action-items.md`: the pattern text should stay general, the examples are
replaceable evidence.
261 changes: 261 additions & 0 deletions .claude/skills/ci-test-audit/action-items.md

Large diffs are not rendered by default.

40 changes: 24 additions & 16 deletions .claude/skills/ci-workflow-guide/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,11 @@
---
name: ci-workflow-guide
description: Guide to SGLang CI workflow orchestration — stage ordering, fast-fail, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched and gated across stages.
description: Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched and gated across stages.
---

# SGLang CI Workflow Orchestration Guide

This skill covers the CI **infrastructure** layer — how tests are dispatched, gated, and fast-failed across stages. For test authoring (templates, fixtures, registration, model selection), see the [write-sglang-test skill](../write-sglang-test/SKILL.md).
This skill covers the CI **infrastructure** layer — how tests are dispatched, gated, and aborted on failure across stages. For test authoring (templates, fixtures, registration, model selection), see the [write-sglang-test skill](../write-sglang-test/SKILL.md).

---

Expand All @@ -24,9 +24,10 @@ This skill covers the CI **infrastructure** layer — how tests are dispatched,
| `.github/workflows/pr-test.yml` | Main workflow — all stages, jobs, conditions, matrix definitions |
| `.github/workflows/pr-test-extra.yml` | Extra workflow — gated by BOTH `run-ci` and `run-ci-extra` labels |
| `.github/workflows/pr-gate.yml` | PR gating: draft check, `run-ci` label, per-user rate limiting |
| `.github/actions/check-pr-test-health/action.yml` | Cross-job fast-fail: queries API for any failed job |
| `.github/actions/check-pr-test-health/action.yml` | Cross-job fail-fast: queries API for any failed job |
| `.github/actions/wait-for-jobs/action.yml` | Stage gating: polls API until stage jobs complete |
| `.github/actions/check-maintenance/action.yml` | Maintenance mode check |
| `.github/scripts/ci-labels.cjs` | Resolves the four CI control labels into dispatch axes |
| `test/run_suite.py` | Suite runner: collects, filters, partitions, executes tests |
| `python/sglang/test/ci/ci_register.py` | Test registration (AST-parsed markers), LPT auto-partition |
| `python/sglang/test/ci/ci_utils.py` | `run_unittest_files()`: execution, retry, continue-on-error |
Expand Down Expand Up @@ -113,21 +114,21 @@ This skill covers the CI **infrastructure** layer — how tests are dispatched,
└─────────────────────────────────────┘
```

**Every stage test job** includes a `check-pr-test-health` step after checkout — if any job in the run has already failed, the job fast-fails (red X) with a root cause annotation.
**Every stage test job** includes a `check-pr-test-health` step after checkout — if any job in the run has already failed, the job fails fast (red X) with a root cause annotation.

**Scheduled runs** skip `wait-for-base-*` jobs, running all stages in parallel. Fast-fail is also disabled.
**Scheduled runs** skip `wait-for-base-*` jobs, running all stages in parallel. Fail-fast is also disabled.

---

## Fast-Fail Layers
## Fail-Fast Layers

4 layers of fast-fail, from fine to coarse:
4 layers of fail-fast, from fine to coarse:

| Layer | Mechanism | Granularity | Disabled on schedule? |
|-------|-----------|-------------|----------------------|
| **1. Test method → file** | `unittest -f` (failfast) | One test method fails → entire test file stops immediately | Yes |
| **2. File → suite** | `run_unittest_files()` default | One test file fails → entire suite stops (`--continue-on-error` off) | Yes |
| **3. Job → job (same stage)** | `check-pr-test-health` action | One job fails → other waiting jobs in same stage fast-fail (red X) | Yes |
| **3. Job → job (same stage)** | `check-pr-test-health` action | One job fails → other waiting jobs in same stage fail-fast (red X) | Yes |
| **4. Stage → stage (cross-stage)** | `wait-for-base-*` + `needs` | Base A fails → base B/C jobs skip entirely (never get a runner) | Yes (wait jobs skipped) |

- **Layer 1**: `-f` flag appended to all `python3 -m pytest` / `unittest` invocations in `ci_utils.py`
Expand All @@ -142,12 +143,17 @@ This skill covers the CI **infrastructure** layer — how tests are dispatched,
| Aspect | PR (`pull_request`) | Scheduled (`cron`, every 6h) | Manual dispatch (`workflow_dispatch`) |
|--------|---------------------|------------------------------|--------------------------------------|
| **Stage ordering** | Sequential: A → B → C via `wait-for-base-*` | Parallel (all at once) | Single target stage only |
| **Cross-job fast-fail** | Yes (`check-pr-test-health`) | Yes | Yes |
| **Cross-job fail-fast** | Yes (`check-pr-test-health`) | Yes | Yes |
| **continue-on-error** | No (stop at first failure within suite) | Yes (run all tests) | No |
| **Retry** | Enabled | Enabled | Enabled |
| **max_parallel** | 3 (default), 14 if `high priority` label | 14 | 3 (default), 14 if `high priority` |
| **max_parallel** | 3 (default), 14 if `max-concurrency` label | 14 | 3 (default), 14 if `max-concurrency` |
| **PR gate** | Yes (draft, label, rate limit) | Skipped | Skipped |
| **Concurrency** | `cancel-in-progress: true` per branch | Queue (no cancel) | Isolated per stage+SHA |
| **Concurrency** | `cancel-in-progress: true` per PR | Queue (no cancel) | Isolated per stage+SHA |

Four labels relax these limits for one PR: `bypass-fail-fast`, `parallel-stages`,
`max-concurrency`, and `highest-priority` (all three). `.github/scripts/ci-labels.cjs`
resolves them; the [contribution guide](https://docs.sglang.io/developer_guide/contribution_guide.html#ci-control-labels)
describes what each one does.

---

Expand All @@ -158,7 +164,7 @@ This skill covers the CI **infrastructure** layer — how tests are dispatched,
**How it works:**
1. Calls `listJobsForWorkflowRun` to list all jobs in the current run
2. Matches jobs by exact name or prefix (for matrix jobs, e.g., `base-b-test-1-gpu-small (3)`)
3. If any matched job has `conclusion === 'failure'` → fail immediately (fast-fail)
3. If any matched job has `conclusion === 'failure'` → fail immediately (fail-fast)
4. If all matched jobs are completed and count matches `expected_count` → success
5. Otherwise → sleep `poll-interval-seconds` (default: 60s) and retry
6. Timeout after `max-wait-minutes` (240 min for base-a, 480 min for base-b)
Expand All @@ -179,7 +185,7 @@ This skill covers the CI **infrastructure** layer — how tests are dispatched,

---

## Cross-Job Fast-Fail (`check-pr-test-health` action)
## Cross-Job Fail-Fast (`check-pr-test-health` action)

Composite action called after checkout in every stage test job (21 jobs total across `pr-test.yml`, `pr-test-multimodal-gen.yml`, `pr-test-sgl-kernel.yml`, `pr-test-jit-kernel.yml`).

Expand All @@ -189,7 +195,7 @@ Composite action called after checkout in every stage test job (21 jobs total ac
3. If root cause failures found → calls `core.setFailed()` with the list of root cause job names
4. If none → does nothing (step succeeds)

**Cascade filtering**: When job A fast-fails due to health check, it also has `conclusion: failure`. Without filtering, job B would list both the original failure AND job A's fast-fail. The filter checks each failed job's `steps` array — if the failing step name contains `check-pr-test-health` or `Check PR test health`, it's excluded from the root cause list.
**Cascade filtering**: When job A fails fast due to the health check, it also has `conclusion: failure`. Without filtering, job B would list both the original failure AND job A's fail-fast. The filter checks each failed job's `steps` array — if the failing step name contains `check-pr-test-health` or `Check PR test health`, it's excluded from the root cause list.

**Usage pattern:**
```yaml
Expand All @@ -210,11 +216,11 @@ steps:

**Visual effect**: Job shows **red X** (failure) with error annotation showing root cause job names. Subsequent steps are naturally skipped (default `if: success()` is false after a failed step). No per-step `if` guards needed.

**No stage filtering**: Checks ALL jobs in the run, not just the current stage. Any failure anywhere triggers fast-fail.
**No stage filtering**: Checks ALL jobs in the run, not just the current stage. Any failure anywhere triggers fail-fast.

**Error message example:**
```
Fast-fail: skipping — root cause job(s): base-b-test-1-gpu-small (0), base-b-test-1-gpu-small (1)
Fail-fast: skipping — root cause job(s): base-b-test-1-gpu-small (0), base-b-test-1-gpu-small (1)
```

---
Expand Down Expand Up @@ -388,6 +394,8 @@ group: pr-test-{event_name}-{branch}-{pr_sha}-{stage}
| `/rerun-failed-ci` | Reruns failed jobs in the latest workflow run |
| `/tag-and-rerun-ci` | Adds `run-ci` label + reruns failed |
| `/tag-and-rerun-ci extra` | Adds both `run-ci` and `run-ci-extra` labels + reruns failed |
| `/run-full-ci` | Short form of `/tag-and-rerun-ci extra` (baseline + extra). Alias: `/rerun-full-ci` |
| `/run-extra-ci` | Adds both labels + reruns **only** `PR Test Extra`, leaving baseline runs alone. Alias: `/rerun-extra-ci` |
| `/rerun-test <test-file> [<test-file> ...]` | Reruns specific test file(s) via `rerun-test.yml`. A file arg containing a glob metacharacter (`*`, `?`, `[...]`) expands against `test/registered/` and the multimodal test dir to every matching `test_*.py` (e.g. `/rerun-test test_*backend*.py` — wrap in backticks so GitHub doesn't italicize the `*`); matches are deduped, grouped by dispatch shape, and can't carry a `::test` selector. No match → single ⛔ reply, nothing dispatched. Each reply echoes its originating command (`Results for …`) so concurrent commands stay distinguishable |
| `/rerun-group <group> [<group> ...]` | Expands registered test groups, then reuses `/rerun-test` |

Expand Down
Loading
Loading