Skip to content

[PD] Share one head-slice helper across mooncake, mori, and nixl - #39660

Merged
hnyls2002 merged 12 commits into
mainfrom
lsyin/head-slice-ut
Sep 25, 2026
Merged

hnyls2002 merged 12 commits into
mainfrom
lsyin/head-slice-ut

Conversation

@hnyls2002

@hnyls2002 hnyls2002 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Background

Step 1 of #34510: the same common semantic is implemented more completely in one backend than another.

Heterogeneous-TP head-index arithmetic exists in four places -- compute_head_slice_params in common/staging_buffer.py (already on main, already used by the staging path), plus hand-rolled copies in mooncake, nixl and mori. mooncake and nixl divide by the GQA replication factor; mori does not, and so maps decode ranks onto the wrong KV heads whenever decode_tp_size > total_kv_heads.

This PR routes the three copies through the shared one. compute_head_slice_params itself is unchanged; only the call sites change. No Transport abstraction, no wire-format change.

Modifications

mooncake -- send_kvcache_slice drops 41 lines of inline arithmetic, including its if attn_tp_size > dst_attn_tp_size direction branch, and calls the shared helper.

nixl -- the gather branch does the same. Its scatter branch already called the helper.

mori -- same call site, and this one changes behavior. Its scatter branch read:

src_head_start = (dst_tp_rank * dst_heads_per_rank) % src_heads_per_rank

Under GQA replication (decode_tp_size > total_kv_heads) consecutive decode ranks share a KV head -- QKVParallelLinear maps by tp_rank // num_kv_head_replicas -- so a plain modulo hands ranks 1..r-1 of each group a head they do not own. With prefill TP=1, decode TP=4 and 2 KV heads the correct map is [0, 0, 1, 1]; mori produced [0, 1, 0, 1]. Wrong head indices do not fail loudly: the transfer delivers the wrong channels and the only symptom is a garbled accuracy score.

Equivalence. The old inline code was extracted verbatim from main and compared against the shared helper over a sweep of src_tp/dst_tp in {1,2,4,8,16,32}, total_kv_heads in {1..8,16,32,64,128}, ranks up to 2x tp to cover modulo wraparound:

call site combinations mismatches
mooncake send_kvcache_slice, both directions 46128 0
nixl gather branch, its only reachable domain 62496 0
mori gather branch 14880 0
mori scatter branch 31248 932 -- the fix above

Unit test

test/registered/unit/disaggregation/test_head_slice_params.py, registered CPU-only (est_time=10), so it costs no GPU runner time. Covers gather (4->2), scatter (2->4), GQA replication (1->4) and equal TP. Expected values are derived by hand from the head-distribution rules, never by calling the implementation, so a bug there cannot make both sides agree.

On the Step 1 "backend-parameterized regression test" convention: after this PR there is one implementation rather than three, so the test targets the shared helper directly instead of parameterizing over backends.

Test suite

The silent-failure mode above is why the 8-gpu-h20 suite carried a {plain, staging} x {gather, scatter} matrix. With the indices guarded on CPU, mooncake's send_kvcache_slice has no direction branch left -- gather and scatter differ only in the values the helper returns -- so MooncakeMHADecodeLargerTP and StagingPrefillLargerTP are dropped from test_disaggregation_different_tp.py (8 classes -> 6).

Every transfer path keeps an end-to-end test:

path remaining coverage
send_kvcache_slice gather MooncakeMHAPrefillLargerTP (MHA 4->2)
send_kvcache_slice scatter GDNHybridHeteroTP (GDN 1->4)
_do_staging_transfer, both directions StagingDecodeLargerTP, StagingRadixPrefillLargerTP
send_kvcache whole-block MooncakePrefillLargerTP, MooncakeDecodeLargerTP (MLA)

CI time

test_disaggregation_different_tp.py is the longest file in the 8-gpu-h20 suite -- 934s of
a 1728s suite on the 2026-09-25 scheduled run -- so it sets that stage's wall-clock floor.

samples mean
main, 8 classes 934 / 995 / 1173 / 1179s 1070s
this PR, 6 classes 650 / 779s 715s

About -355s (-33%) on the means. Run-to-run spread on this runner is wide (129s between
two runs of the same code here, 245s across the four main samples), so pairing the worst
case for this PR against the best for main still gives -155s (-17%); that is the floor,
not the expectation.

main samples are the h20 partitions of scheduled runs 36127993962 / 36071364371 /
35991484462 / 35932079543; this PR's are /rerun-test on the same runner class. The
registered est_time=1171 is left alone -- scripts/ci/update_est_time.py refreshes it
from stats p90 once this lands.

The new unit test is CPU-only (est_time=10, measured 6s) and adds no GPU runner time.

Not included

Mamba/linear-attention state transfer slices on state dimensions rather than KV heads, and has its own duplication (utils.py:862 and mori/conn.py) with the same plain-modulo shape. Different axis, left alone here.

Accuracy Tests

On the CUDA side the h20 disaggregation suite covers the mooncake and nixl call sites.

The mori change is on the AMD path. TestMoriTransferEngineTPMismatchE2E exercises prefill TP=2 / decode TP=4, but on Llama-3.2-1B-Instruct (8 KV heads) dst_replication = max(1, 4 // 8) = 1, so the old and new maps agree and that test is green either way -- it guards against regression on the non-replicating path, not the fix. Reaching the divergent path needs decode_tp > total_kv_heads; no shape in the AMD matrix has it today. The CPU unit test is therefore the only automated guard on the corrected mapping.

Speed Tests and Profiling

Not applicable -- the arithmetic is per-transfer setup, not a data path.

Checklist


CI States

Latest PR Test (Base): 🚫 Run #36199005947
Latest PR Test (Extra): ❌ Run #36199005896
Latest PR Test (AMD ROCm 10): 🚫 Run #36199006038

@hnyls2002 hnyls2002 changed the title [misc] Guard heterogeneous-TP head slicing with a CPU unit test [PD] Share one head-slice helper across mooncake and nixl, guarded by a CPU unit test Sep 15, 2026
@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 15, 2026
@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_head_slice_params.py test_disaggregation_different_tp.py test_disaggregation_nixl.py test_mori_transfer_engine_e2e.py test_nixl_backend_basic.py

@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_head_slice_params.py test_disaggregation_different_tp.py test_disaggregation_nixl.py test_mori_transfer_engine_e2e.py test_nixl_backend_basic.py:

🚀 ubuntu-latest (2 tests): ✅ View workflow run

cd test/ && python3 registered/unit/disaggregation/test_head_slice_params.py
cd test/ && python3 registered/unit/disaggregation/test_nixl_backend_basic.py

🚀 8-gpu-h20 (2 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_different_tp.py
cd test/ && python3 registered/disaggregation/test_disaggregation_nixl.py

⛔ test_mori_transfer_engine_e2e.py: test/registered/amd/disaggregation/test_mori_transfer_engine_e2e.py is registered for AMD (suite stage-b-test-large-8-gpu-mi35x-disaggregation-amd), not for CUDA or CPU; rerun-test.yml has no AMD job. Rerun it with /rerun-failed-ci, or dispatch the AMD workflow manually.

@BBuf BBuf changed the title [PD] Share one head-slice helper across mooncake and nixl, guarded by a CPU unit test [PD] Share one head-slice helper across mooncake, mori, and nixl Sep 16, 2026
@BBuf

BBuf commented Sep 16, 2026

Copy link
Copy Markdown
Collaborator

Reviewed the refactor itself carefully and it is faithful: I extracted the three old inline blocks and compute_head_slice_params and they agree on both directions for mooncake and on the reachable domain for nixl's gather branch, with the mori scatter change being the documented GQA-replication fix. No objection to the production diff.

The concern is what is left guarding it.

The stated justification for dropping the two e2e classes no longer holds.

The PR body argues the {plain, staging} x {gather, scatter} matrix can shrink because the arithmetic is now guarded on CPU:

With the indices guarded on CPU, mooncake's send_kvcache_slice has no direction branch left [...] so MooncakeMHADecodeLargerTP and StagingPrefillLargerTP are dropped

but the "Test cleanup" section then says:

The newly added CPU head-slice unit-test file was removed in follow-up CI cleanup.

test/registered/unit/disaggregation/test_head_slice_params.py is not on the branch. So the net effect of the PR on the test suite is +0 / -160: two e2e classes removed, and the CPU guard that was supposed to replace them removed as well. What is left is the manual equivalence sweep in the PR description, which is not executable and will not run again.

This matters more than usual for this particular change, because the PR itself documents the failure mode:

Wrong head indices do not fail loudly: the transfer delivers the wrong channels and the only symptom is a garbled accuracy score.

Could the unit test be restored? It is CPU-only at est_time=10, costs no GPU runner time, and was the whole argument for shrinking the GPU matrix. If it was dropped because of a registry/CI issue rather than on purpose, I am happy to help get it landed.

Second: the mori fix ships untested on every lane.

The mori scatter branch is the one real behaviour change here ([0, 1, 0, 1] -> [0, 0, 1, 1] for prefill TP=1 / decode TP=4 / 2 KV heads). Per the PR that path needs AMD CI, PR Test (AMD ROCm 10) is currently red, and test_mori_transfer_engine_e2e.py could not be dispatched from /rerun-test. With the CPU unit test also gone, this fix has no automated coverage at all on any lane. Was the AMD failure triaged as unrelated?

(Review by Claude Opus 5, run by @BBuf.)

@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test registered/unit/disaggregation/test_head_slice_params.py registered/disaggregation/test_disaggregation_different_tp.py registered/disaggregation/test_disaggregation_nixl.py

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test registered/unit/disaggregation/test_head_slice_params.py registered/disaggregation/test_disaggregation_different_tp.py registered/disaggregation/test_disaggregation_nixl.py:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/disaggregation/test_head_slice_params.py

🚀 8-gpu-h20 (2 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_different_tp.py
cd test/ && python3 registered/disaggregation/test_disaggregation_nixl.py

@hnyls2002
hnyls2002 merged commit efd9a40 into main Sep 25, 2026
86 of 125 checks passed
@hnyls2002
hnyls2002 deleted the lsyin/head-slice-ut branch September 25, 2026 23:13
arbi-dev added a commit to arbicity/sglang-turbo that referenced this pull request Oct 7, 2026
…5.21 (#12)

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21

The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream
v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the
commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in
the post-port commits #9 (hybrid sliding-window through the kv-cache
plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory
reads attention_backends()).

Upstream's structure wins and the seam is re-applied on it:
- server_args.py is now a thin record over arg_groups/: the
  --kv-cache-dtype choices are hoisted into arg_groups/choices.py as
  KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from
  server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a
  plugin backend move to arg_groups/model_hook.py.
- ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's
  resolution hooks (arg_groups/resolution_hooks.py) are the official slot,
  and turbo-attn's plugin now pairs the flags there.
- Plugin dtype reads go through the resolved model bag (get_model()), not
  the raw ServerArgs record, which v0.5.21 leaves as operator input.
- HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it
  holds and stands in as the per-layer cost, so upstream's new draft-SWA
  and unified-pool formulas apply unchanged; layer ids come from
  kvc.layer_info, as upstream's own SWA pool takes them.
- Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21
  a PP draft may load its own embedding (pp_draft_embedding).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
sam123456598 added a commit to vessl-ai/sglang that referenced this pull request Oct 8, 2026
* [feature] add per-item candidate token scoring and calibration (sgl-project#40826)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants