Repository navigation
[DeepSeek V4.1] Add DeepSelect JIT kernel. - #40556
Conversation
973eb23 to
1c14325
Compare
The AOT build compiled all 86 explicit instantiations of DeepSelect's top-k into sgl_kernel and resolved a TopkSelectConfig at runtime through nested BOOL_SWITCH macros. A JIT build targets one architecture and one call signature, so the axes those macros expanded -- value dtype, index dtype, sorted_value, sorted_index, return_value, the top-k bucket and the cluster size -- become template arguments chosen in Python, and the compiler instantiates only the one or two kernels a module can reach. What is left in C++ is the decision that depends on the launch shape. - Vendor the upstream kernel headers under jit/csrc/deepselect/vendor, dropping the instantiations, api.cpp and dispatch_utils.h. - Add entry.cuh: tensor validation, the TopkSelectArgs block, and the TopkNormal / TopkCluster host entry points. - Add ops/deep_select.py: one module per (value dtype, top-k bucket, normal or cluster), with the configuration tables it selects from. - Promote DeviceCacheMap out of host::runtime::details and add get_max_smem_per_block, which the kernels' shared-memory budget needs. - Add aligned_new_empty for the padded row strides DeepSelect requires. Verified on an H200: every (dtype, bucket, cluster) module compiles for sm_90a, and 18 correctness cases match torch.topk -- both wave tunings, the 8-CTA cluster path, variable row lengths via `end`, and the indices-only path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Please fix lint |
Done. |
|
/rerun-test test/registered/kernels/ops/attention/test_deep_select.py |
|
Results for 🚀 |
|
/rerun-group kernels |
|
Results for 🚀 🚀 🚀 🚀 🚀 🚀 🚀 ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ ⛔ |
get_max_smem_per_block (sgl-project#40556) reads cudaDevAttrMaxSharedMemoryPerBlock[Optin], which the HIP alias block did not define, so every JIT module including runtime.cuh failed to build on ROCm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Could you help to compare with https://github.com/yiakwy-xpu-ml-framework-team/flash-float-jit-kernels top v3 in old days. |
|
The new topk has been added to sglang at #40556 |
|
@yiakwy-xpu-ml-framework-team pls change the v2 baseline to JIT v2 (the AOT version is deprecated).
|
We have a record of how to optimizing topk previously. Kindly check it out https://github.com/yiakwy-xpu-ml-framework-team/flash-float-jit-kernels/blob/main/AGENTS.md. Let me know if I can help ! D |
| // tma gather4 (https://docs.nvidia.com/cuda/parallel-thread-execution/#data-movement-and-conversion-instructions-cp-async-bulk-tensor) | ||
| // Please pay attention that the coordinates of TMA gather4 are int32, which may lead to overflow under some scenarios | ||
| CUTE_DEVICE | ||
| void tma_gather4(const void* desc_ptr, transac_bar_t &mbar_ptr, void* smem_ptr, int col_idx, int4 row_idxs, int64_t cache_hint) { |
There was a problem hiding this comment.
Do we have scenarios (single instruction test), where this outperforms te orignal method ?
|
@DarkSharpness I noticed your great efforts in V2 optimizaiton ( Could you help to migrate the V3 to the repo ? Oh just find that @DarkSharpness we have talked about this in April. I guess we can do it again how do you suggest ? |
…5.21 (#12) * [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384) * [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405) * Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096) * [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059) * [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425) * [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878) Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> * [AMD] Add tuned dsv4 shape (sgl-project#40996) * [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388) * [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612) * Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> * [HiCache] Make host reclamation independent of transfer order (sgl-project#40512) * [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922) * [diffusion] update code owner (sgl-project#41130) * [NPU] Update CANN version to 9.1.0 (sgl-project#40524) * [PD] Enable deferred decode-side KV release by default (sgl-project#41023) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [chore] surface the cookbook to users who pip install sglang (sgl-project#40866) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960) Signed-off-by: Ting Sun <suntcrick@gmail.com> * [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> * [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805) * Support unified memory decode host pools (sgl-project#39478) Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> * [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858) Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> * [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155) * [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109) * [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079) * [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081) * [Refactor] Share the MoE output all-reduce between models (sgl-project#41097) * [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082) * [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083) * [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084) * [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120) * [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041) * MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> * MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559) Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090) * [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162) * [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046) * [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049) * [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929) Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179) * Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442) Co-authored-by: Caio Rocha <caiorocha@microsft.com> * [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960) Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> * [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> * [Score API] Setwise Scoring Support (sgl-project#38965) * [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091) Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988) * [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907) * [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * Support XQA backend for SpecDec verify (sgl-project#32269) * [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793) Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> * [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> * [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * [feat] add a system one compatible /v1/systemone route (sgl-project#41208) * fix(openai): reject request-supplied chat_template by default (sgl-project#28135) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159) * [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201) Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> * [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456) * [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092) * [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619) * [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193) * [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194) * [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195) * [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196) * [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197) * [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198) * [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191) * [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199) * [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200) * [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224) * [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417) Co-authored-by: mickqian <mickqian@users.noreply.github.com> * [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486) * Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132) * [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064) Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: jacky.cheng <yichiche@amd.com> * fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248) Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> * Fix multimodal feature offload races (sgl-project#40621) Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986) * [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246) Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> * [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215) * [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216) * [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379) * [KVCache] Support lmcache unified radix cache (sgl-project#38652) Signed-off-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389) Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> * dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280) * [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660) Co-authored-by: BBuf <1182563586@qq.com> * [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284) * [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286) * [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285) * [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490) * [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291) Co-authored-by: DarkSharpness <2040703891@qq.com> * [diffusion] model: support Anima Base v1.0 (sgl-project#41011) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261) * [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297) * [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505) * [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150) Co-authored-by: jacky.cheng <yichiche@amd.com> * [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Restore logical kernel groups and test organization (sgl-project#41243) * [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188) * [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223) Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> * [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556) Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252) * [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253) * [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254) * [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255) * [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256) * [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257) * [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321) * [AMD] Fix jit broken on rocm env (sgl-project#41356) * Make sliding-window caching and speculative batch padding extensible (sgl-project#41325) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> * [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377) * dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019) Co-authored-by: x <x> * [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [mem_cache] Never free the protected prefix on request release (sgl-project#41312) * [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345) * [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416) Co-authored-by: aimicahchen <aimicahchen@tencent.com> * [kv-shard 3/4] Enable Control Plane B (sgl-project#39964) Co-authored-by: Zhangheng <hzh0425@apache.org> * [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166) * [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305) Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> * [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281) * [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192) * [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684) * [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990) * [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418) Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> * [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272) Signed-off-by: rockdu <kangrdu@gmail.com> * [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> * [bench] Take each request's prompt length from the server (sgl-project#39889) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> * [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266) * Add MiniMax arch fallback to auto parser resolution (sgl-project#40930) * [HiCache] Add the page-unified KV load-back JIT kernel (sgl-project#39726) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> * [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458) * [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: kpham-sgl <khoa.pham@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Port chat_parsing core (sgl-project#40477) Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> * [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417) * [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418) * [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419) * [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420) * [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421) * [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422) * [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423) * [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424) * [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425) * [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426) * [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427) * [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428) * [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429) * [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430) * [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431) * [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432) * [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433) * [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434) * [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435) * [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436) * [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437) * [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438) * [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439) * [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440) * [Refactor] Build the boundary into any stage with one construction (sgl-project#41441) * [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442) * [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443) * [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075) * [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814) * [diffusion] feat: support multiple task types for pipelines (sgl-project#38762) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [CI] fix CI regression on xeon (sgl-project#41002) * [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378) * [AMD] Register Triton data movement tests in PR CI (sgl-project#41137) * [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> * [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> * [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664) Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899) Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Rust server unify datapath for mm and generate requests (sgl-project#39679) * feat(npu): Support returning indexer top-k results (sgl-project#39060) * Fix chat template cache key order (sgl-project#41517) * docs: add prefill context parallelism guide and design draft (sgl-project#39354) * [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220) Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> * [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164) * [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165) * [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539) * [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943) * [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> * [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161) Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> * [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133) * [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144) * [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339) * [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446) * [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643) * [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235) * [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987) * [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597) * dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> * [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588) Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> * [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469) * [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196) Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118) Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> * [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590) Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> * [diffusion] model: support flux 3 action robot policies (sgl-project#41066) * [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627) Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> * [npu]support NPU 910C L2 memcache offload (sgl-project#41527) * [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600) * Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Refactor] Group communicator fusion and CP adapters (sgl-project#41547) * [Refactor] Centralize decoder output access (sgl-project#41548) * [Refactor] Carry residual state across stage boundaries (sgl-project#41549) * [Refactor] Capture auxiliary states at residual reads (sgl-project#41550) * [Refactor] Select reduction fusion at the consumer (sgl-project#41551) * [Refactor] Construct independent decoder stage boundaries (sgl-project#41552) * [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553) * [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554) * [Refactor] Rename the module to layer_boundary (sgl-project#41555) * [Refactor] Group layer boundary unit tests (sgl-project#41556) * [Refactor] Document layer boundary contracts and integration (sgl-project#41557) * [Rust] Extract a transport-neutral frontend core (sgl-project#39385) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618) * [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804) Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528) * [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059) * [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> * [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802) Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21 The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in the post-port commits #9 (hybrid sliding-window through the kv-cache plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory reads attention_backends()). Upstream's structure wins and the seam is re-applied on it: - server_args.py is now a thin record over arg_groups/: the --kv-cache-dtype choices are hoisted into arg_groups/choices.py as KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a plugin backend move to arg_groups/model_hook.py. - ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's resolution hooks (arg_groups/resolution_hooks.py) are the official slot, and turbo-attn's plugin now pairs the flags there. - Plugin dtype reads go through the resolved model bag (get_model()), not the raw ServerArgs record, which v0.5.21 leaves as operator input. - HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it holds and stands in as the per-layer cost, so upstream's new draft-SWA and unified-pool formulas apply unchanged; layer ids come from kvc.layer_info, as upstream's own SWA pool takes them. - Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21 a PP draft may load its own embedding (pp_draft_embedding). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> Signed-off-by: Ting Sun <suntcrick@gmail.com> Signed-off-by: chunxiaozheng <1179548172@qq.com> Signed-off-by: rockdu <kangrdu@gmail.com> Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com> Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com> Co-authored-by: Alan Kao <akao@amd.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca> Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: 黄孝君 <dingfangsu23@gmail.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Ting SUN <suntcrick@gmail.com> Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu> Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> Co-authored-by: ChangLiu0709 <cliu1004@amd.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Thomas Wang <thomawan@amd.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: Chunan Zeng <zcnrex@gmail.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com> Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com> Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com> Co-authored-by: Caio Rocha <caiorocha@microsft.com> Co-authored-by: metamergebot <metamergebot@gmail.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com> Co-authored-by: jacky.cheng <yichiche@amd.com> Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com> Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> Co-authored-by: Richard Wang <wangrichard08@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: Xinyi Song <xinyis10@illinois.edu> Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: ashwini rathi <ashwini.rathi@intel.com> Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com> Co-authored-by: Even Zhou <even.y.zhou@outlook.com> Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: cctry <csycfl@gmail.com> Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> Co-authored-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai> Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com> Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: WMC <tnwilly@gmail.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com> Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com> Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> Co-authored-by: Ankith Averineni <saverine@amd.com> Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com> Co-authored-by: Yuan Luo <yuan.luo@hotmail.com> Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com> Co-authored-by: CuzMi <simon.weijie@gmail.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com> Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com> Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com> Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com> Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com> Co-authored-by: flb_ <floatlibai@gmail.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> Co-authored-by: Peng Wu <peng@thinkingmachines.ai> Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com> Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com> Co-authored-by: karverma-amd <karan.verma@amd.com> Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: triple-mu <gpu@163.com> Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com> Co-authored-by: Hrithvik Alex <halex623@gmail.com> Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com> Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> Co-authored-by: sglang-bot <sglangbot@gmail.com> Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> Co-authored-by: Jialin Ouyang <jialino@meta.com> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> Co-authored-by: James <445169590@qq.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com> Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com> Co-authored-by: Jiacong Fang <zldrobit@126.com>
* [feature] add per-item candidate token scoring and calibration (sgl-project#40826) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384) * [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405) * Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096) * [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059) * [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425) * [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878) Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> * [AMD] Add tuned dsv4 shape (sgl-project#40996) * [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388) * [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612) * Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> * [HiCache] Make host reclamation independent of transfer order (sgl-project#40512) * [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922) * [diffusion] update code owner (sgl-project#41130) * [NPU] Update CANN version to 9.1.0 (sgl-project#40524) * [PD] Enable deferred decode-side KV release by default (sgl-project#41023) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [chore] surface the cookbook to users who pip install sglang (sgl-project#40866) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960) Signed-off-by: Ting Sun <suntcrick@gmail.com> * [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> * [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805) * Support unified memory decode host pools (sgl-project#39478) Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> * [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858) Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> * [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155) * [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109) * [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079) * [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081) * [Refactor] Share the MoE output all-reduce between models (sgl-project#41097) * [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082) * [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083) * [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084) * [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120) * [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041) * MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> * MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559) Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090) * [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162) * [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046) * [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049) * [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929) Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179) * Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442) Co-authored-by: Caio Rocha <caiorocha@microsft.com> * [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960) Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> * [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> * [Score API] Setwise Scoring Support (sgl-project#38965) * [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091) Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988) * [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907) * [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * Support XQA backend for SpecDec verify (sgl-project#32269) * [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793) Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> * [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> * [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * [feat] add a system one compatible /v1/systemone route (sgl-project#41208) * fix(openai): reject request-supplied chat_template by default (sgl-project#28135) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159) * [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201) Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> * [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456) * [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092) * [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619) * [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193) * [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194) * [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195) * [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196) * [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197) * [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198) * [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191) * [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199) * [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200) * [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224) * [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417) Co-authored-by: mickqian <mickqian@users.noreply.github.com> * [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486) * Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132) * [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064) Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: jacky.cheng <yichiche@amd.com> * fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248) Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> * Fix multimodal feature offload races (sgl-project#40621) Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986) * [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246) Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> * [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215) * [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216) * [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379) * [KVCache] Support lmcache unified radix cache (sgl-project#38652) Signed-off-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389) Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> * dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280) * [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660) Co-authored-by: BBuf <1182563586@qq.com> * [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284) * [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286) * [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285) * [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490) * [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291) Co-authored-by: DarkSharpness <2040703891@qq.com> * [diffusion] model: support Anima Base v1.0 (sgl-project#41011) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261) * [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297) * [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505) * [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150) Co-authored-by: jacky.cheng <yichiche@amd.com> * [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Restore logical kernel groups and test organization (sgl-project#41243) * [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188) * [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223) Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> * [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556) Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252) * [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253) * [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254) * [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255) * [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256) * [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257) * [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321) * [AMD] Fix jit broken on rocm env (sgl-project#41356) * Make sliding-window caching and speculative batch padding extensible (sgl-project#41325) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> * [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377) * dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019) Co-authored-by: x <x> * [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [mem_cache] Never free the protected prefix on request release (sgl-project#41312) * [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345) * [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416) Co-authored-by: aimicahchen <aimicahchen@tencent.com> * [kv-shard 3/4] Enable Control Plane B (sgl-project#39964) Co-authored-by: Zhangheng <hzh0425@apache.org> * [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166) * [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305) Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> * [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281) * [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192) * [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684) * [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990) * [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418) Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> * [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272) Signed-off-by: rockdu <kangrdu@gmail.com> * [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> * [bench] Take each request's prompt length from the server (sgl-project#39889) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> * [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266) * Add MiniMax arch fallback to auto parser resolution (sgl-project#40930) * [HiCache] Add the page-unified KV load-back JIT kernel (sgl-project#39726) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> * [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458) * [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: kpham-sgl <khoa.pham@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Port chat_parsing core (sgl-project#40477) Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> * [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417) * [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418) * [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419) * [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420) * [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421) * [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422) * [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423) * [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424) * [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425) * [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426) * [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427) * [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428) * [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429) * [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430) * [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431) * [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432) * [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433) * [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434) * [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435) * [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436) * [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437) * [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438) * [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439) * [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440) * [Refactor] Build the boundary into any stage with one construction (sgl-project#41441) * [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442) * [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443) * [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075) * [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814) * [diffusion] feat: support multiple task types for pipelines (sgl-project#38762) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [CI] fix CI regression on xeon (sgl-project#41002) * [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378) * [AMD] Register Triton data movement tests in PR CI (sgl-project#41137) * [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> * [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> * [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664) Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899) Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Rust server unify datapath for mm and generate requests (sgl-project#39679) * feat(npu): Support returning indexer top-k results (sgl-project#39060) * Fix chat template cache key order (sgl-project#41517) * docs: add prefill context parallelism guide and design draft (sgl-project#39354) * [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220) Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> * [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164) * [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165) * [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539) * [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943) * [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> * [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161) Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> * [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133) * [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144) * [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339) * [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446) * [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643) * [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235) * [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987) * [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597) * dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> * [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588) Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> * [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469) * [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196) Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118) Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> * [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590) Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> * [diffusion] model: support flux 3 action robot policies (sgl-project#41066) * [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627) Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> * [npu]support NPU 910C L2 memcache offload (sgl-project#41527) * [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600) * Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Refactor] Group communicator fusion and CP adapters (sgl-project#41547) * [Refactor] Centralize decoder output access (sgl-project#41548) * [Refactor] Carry residual state across stage boundaries (sgl-project#41549) * [Refactor] Capture auxiliary states at residual reads (sgl-project#41550) * [Refactor] Select reduction fusion at the consumer (sgl-project#41551) * [Refactor] Construct independent decoder stage boundaries (sgl-project#41552) * [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553) * [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554) * [Refactor] Rename the module to layer_boundary (sgl-project#41555) * [Refactor] Group layer boundary unit tests (sgl-project#41556) * [Refactor] Document layer boundary contracts and integration (sgl-project#41557) * [Rust] Extract a transport-neutral frontend core (sgl-project#39385) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618) * [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804) Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528) * [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059) * [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> * [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802) Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> --------- Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> Signed-off-by: Ting Sun <suntcrick@gmail.com> Signed-off-by: chunxiaozheng <1179548172@qq.com> Signed-off-by: rockdu <kangrdu@gmail.com> Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com> Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com> Co-authored-by: Alan Kao <akao@amd.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca> Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: 黄孝君 <dingfangsu23@gmail.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Ting SUN <suntcrick@gmail.com> Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu> Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> Co-authored-by: ChangLiu0709 <cliu1004@amd.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Thomas Wang <thomawan@amd.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: Chunan Zeng <zcnrex@gmail.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com> Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com> Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com> Co-authored-by: Caio Rocha <caiorocha@microsft.com> Co-authored-by: metamergebot <metamergebot@gmail.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com> Co-authored-by: jacky.cheng <yichiche@amd.com> Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com> Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> Co-authored-by: Richard Wang <wangrichard08@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: Xinyi Song <xinyis10@illinois.edu> Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: ashwini rathi <ashwini.rathi@intel.com> Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com> Co-authored-by: Even Zhou <even.y.zhou@outlook.com> Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: cctry <csycfl@gmail.com> Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> Co-authored-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai> Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com> Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: WMC <tnwilly@gmail.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com> Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com> Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> Co-authored-by: Ankith Averineni <saverine@amd.com> Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com> Co-authored-by: Yuan Luo <yuan.luo@hotmail.com> Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com> Co-authored-by: CuzMi <simon.weijie@gmail.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com> Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com> Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com> Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com> Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com> Co-authored-by: flb_ <floatlibai@gmail.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> Co-authored-by: Peng Wu <peng@thinkingmachines.ai> Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com> Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com> Co-authored-by: karverma-amd <karan.verma@amd.com> Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: triple-mu <gpu@163.com> Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com> Co-authored-by: Hrithvik Alex <halex623@gmail.com> Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com> Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> Co-authored-by: sglang-bot <sglangbot@gmail.com> Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> Co-authored-by: Jialin Ouyang <jialino@meta.com> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> Co-authored-by: James <445169590@qq.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com> Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com> Co-authored-by: Jiacong Fang <zldrobit@126.com>

Motivation
#39421.
This PR moves DeepSelect to the existing JIT infrastructure so only the required call signature is compiled for the local CUDA target.
Modifications
jit/csrc/deepselect/vendor.indices, normal and cluster kernels, sorting, offsets, and variable lengths.
drop includes of the deleted AOT
topk_select.hfiles.handling.
ScopedCudaDevice, runtimeC++ SM check, or
override_jit_cuda_archis added.Accuracy Tests
Speed Tests and Profiling
The complete candidate path was benchmarked on NVIDIA H20 using:
topk_blocks=2048.The measured region includes:
Final Top-K 512
Final Top-K 2048
Across all 16 shapes:
1.91x–3.77x.2.79x.3.77xfor 6 rows, 128K sequence length, and final Top-K 512.The AOT Top-K kernel was also compared directly with the external DeepSelect implementation at width 131072 and Top-K 2048. The latency difference was approximately 0.3% to 0.4%, indicating that moving the kernel into
sgl-kerneldoes not introduce a kernel-level performance regression.Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #36169401421
Latest PR Test (Extra): ❌ Run #36169401278
Latest PR Test (AMD ROCm 10): ❌ Run #36169401651