Repository navigation
Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) - #39816
Conversation
…/ GLM-5.x The FlashInfer MNNVL CuTe DSL AR fusion landed inside a Qwen-named module whose contents were almost entirely architecture-agnostic, reachable only by Qwen3.5, and gated by an env var that shadowed the existing backend selector. Extract the mechanism and wire the DeepSeek-V3 family, which covers GLM-5.x unchanged (GlmMoeDsaForCausalLM subclasses DeepseekV2ForCausalLM). Shared core (layers/moe/cutedsl_ar_fusion.py, replacing qwen35_flashinfer_fusion.py -- no alias shim, importers are repointed): - fused_norm_gamma() resolves the gamma from the norm module itself (GemmaRMSNorm.gemma_weight, RMSNorm.weight), so the fusion needs no per-architecture subclass at all; one communicator serves both families. - install_cutedsl_fusion() / prepare_cutedsl_fusion() replace the install loop and the stringly-typed model reach each model was hand-rolling. - should_use_finalize() (consumer) and should_defer_moe_finalize() (producer) are now separate questions. Consuming is a property of the fused kernel and topology; producing additionally needs a runner that can defer and a successor that will consume. Conflating them made the last layer refuse its predecessor's handoff. - Both live on the base LayerCommunicator returning False, so no call site probes with hasattr. Gating: --flashinfer-allreduce-fusion-backend cute-dsl is now the single switch; SGLANG_FLASHINFER_MNNVL_CUTEDSL_AR_FUSION is deleted along with the override pass that had to suppress the backend arg when both were set. The legacy TRTLLM/MNNVL workspace, its group tagging and its dispatch all stand down under cute-dsl (uses_cutedsl_ar_fusion()). Two bugs the extraction exposed, both live in the Qwen module today: - prepare_attn's handoff branch returned without publishing attn_inputs, so any model using qkv_latent_func asserted in fetch_qkv_latent(). The tail is now _finish_prepare_attn(), shared by both paths. - FlashInfer resolves its routing profile by exact (tp, hidden, top_k, dtype) and ships GB300 H=8192/K=10 only, so every other shape raised "No MNNVL CuTe DSL profile supports this static shape". _config_for_shape() rebuilds one profile at the running shape from the shipped presets and their measured crossovers; H=8192/K=10 still gets DEFAULT_CONFIG unchanged. GLM-5.2-NVFP4, TP8, B300: GSM8K 200q = 0.945 accuracy, 0.000 invalid.
Cut the prose that restated the code or argued for the diff, and condense the blocks that carried one non-recoverable fact down to that fact: the FlashInfer HT kernel's shard-split arithmetic, the exact-shape profile resolution, the TP1-replicated shared-expert double count, and the reason prepare_attn's tail cannot be short-circuited. Comment-only: the AST with docstrings stripped is unchanged for all 8 files.
SGLANG_FLASHINFER_MNNVL_CUTEDSL_AR_FUSION_MAX_INSTANCES had no valid setting. The registry keys workspaces by full signature and one model configuration exists per process, so the count is an invariant rather than a knob; it is now a hardcoded guard with the same fail-closed behavior at the same call site. The optional-argument form of get_flashinfer_mnnvl_cutedsl_ar_fusion() went with it: the no-argument lookup only had to disambiguate between several workspaces, and its sole caller always supplied all five. The arguments are now required. Added in #35758 alongside the gate env var and never released, so no deprecation alias.
Removes machinery that no caller reaches, most of it inherited from #35758. Workspace module: - The registry (signature dataclass, dict, RLock, best-fit search) held at most one entry once the instance cap became an invariant; it collapses to a single module global. - destroy() was never called and _destroyed was therefore always False, so supports() carried a check that could not fire. - moe_finalize_all_reduce_rms_norm / all_reduce_residual_rms_norm took optional norm_output / residual_output that only the unit test supplied; production always let them allocate. Fusion core: - The service re-validated shapes, dtypes and contiguity on every eligible layer of every forward. The handoff builds its own views, and the workspace already refuses an M it cannot serve, so this was duplicated work on the hot path. - supports() also duplicated the workspace's range check, and is_prepared was subsumed by it (supports() is False before prepare()). Qwen3.5 final norm: - SGLANG_TRACE_QWEN35_FINAL_NORM and SGLANG_QWEN35_NATIVE_FINAL_NORM were bring-up scaffolding: print() plus torch.cuda.synchronize() around the norm, and a forward_native A/B switch. Both env vars are deleted; like the other two they shipped in #35758 and were never released.
…plain all-reduce above it Deferring the MoE finalize trades the GEMM2 in-epilogue reduction for an HBM round trip of the [M*top_k, hidden] permuted output, so it only pays at small M. Gate it on a new SGLANG_MOE_DEFERRED_FINALIZE_MAX_TOKENS bound (default 192, 0 disables) applied in both DeepseekV2MoE and the CuTe DSL communicator, which must agree or the layer skips an all-reduce waiting for a handoff that never arrives. Above the bound the next layer's input norm can still absorb the plain post-MoE all-reduce, so keep skipping it in the MoE layer and fuse it in prepare_attn. That path needs a real successor communicator, tracked separately from handoff_has_consumer since a model's final norm closes a handoff but performs no all-reduce. Measured on GLM-5.2 H=6144/top_k=8 TP8 B300 with the cute-dsl backend: -8.4% TPOT at M=16, neutral at 192, +9.3% at 512.
Covers the refactor, the deferred-finalize bound, and the TP8 and TP4 sweeps behind it. Doubles as the PR description.
The linter reads remove/delete as one action, which is right for deleting a file and wrong for eliminating a kernel launch. Complying blindly produced "deletes one kernel launch". Also re-wraps a paragraph an earlier edit mangled.
Same measurements, converted to output tokens per second. The speedup column still comes from the median inter-token latency, which is the outlier-immune quantity the ratios were computed from.
Stock keeps its absolute throughput. The other arms now report their gain over it, which is the quantity a reader wants.
…sion-shared-core # Conflicts: # python/sglang/srt/models/deepseek_v2.py # python/sglang/srt/server_args.py
They were swept into "Delete dead surface in the CuTe DSL AR fusion", which says nothing about them. They cover DSA sparse-MLA decode under decode context parallel, which no commit here touches, and they register CI time on a 4-gpu-b200 and a 1-gpu-large runner.
The only call sites are inside CuteDSLFusionLayerCommunicator, which defines its own should_use_finalize, so the base was never the one that resolved. The Qwen3.5 sites that used to reach it through hasattr now call should_defer_moe_finalize. The should_defer_moe_finalize stub next to it stays: deepseek_v2 calls that one on a plain LayerCommunicator.
experts_can_defer_finalize and handoff_has_consumer were never read apart -- both reads, the gate in should_defer_moe_finalize and the count in the install log, took their conjunction. install_cutedsl_fusion records the AND as may_defer_moe_finalize. can_defer_finalize() stays on the left of the and, so it is still called once for every fusion layer.
Both Qwen3.5 decoder layers take their communicator from _layer_communicator_class(config, is_nextn), which returns the fusion class under _use_mnnvl_cutedsl_fusion(config, is_nextn) -- the same predicate that gates the block the check sat in. The check could not fire. The unsupported_layers check above it, which tests something the layer classes do not determine, stays.
prepare_attn() consumed a pending post-MoE all-reduce through can_absorb_post_moe_all_reduce(), which gates on successor_absorbs_all_reduce -- a property of the outgoing side. The penultimate layer therefore skipped its all-reduce on the strength of a successor, and the last layer, having none of its own, declined the handoff it was owed. The predicate splits in two: can_consume_post_moe_all_reduce() for the incoming side, which a layer owes its predecessor whatever follows it, and can_absorb_post_moe_all_reduce() for the outgoing side, which keeps the successor requirement. That walk found a second half. LayerCommunicator refuses to fuse when moe_ep_size > 1 and moe_tp_size > 1, because skipping the post-experts reduction drops both the EP and the TP leg and one fused collective cannot restore both; the declined fallback reduces over _MOE_TP alone. Both overrides return True before reaching super(), so that guard never ran. should_defer_moe_finalize() was safe only because should_use_finalize() happens to require moe_ep_size == 1. _common_eligible() now carries the refusal, so both fusion patterns inherit it. Pure EP, where one collective over the full TP group is the whole reduction, stays allowed. The last-layer regression drives the real prepare_attn() with a tagged tensor and fails if the call falls through to the unfused path, so reverting the call site alone -- with the split left intact -- is caught.
dsv2_flashinfer_moe_dual_stream_graph is a registered custom op whose schema returns a Tensor, and forward_normal_dual_stream() can now return a MoeFinalizeHandoff. ForwardFlags.scoped() overrides only the flags it is given and leaves the rest alone, so the decoder's defer_moe_finalize reached inside the op even though the op republishes fuse_mlp_allreduce and mlp_reduce_scatter as scalar operands precisely because the caller's scope is not to be relied on. The op pins the flag off and finalizes locally. The skipped all-reduce still reaches the next layer through _sglang_needs_allreduce_fusion, which is the path the decoder already takes above the deferred-finalize bound. The op is CUDA-only, so its regression is a GPU suite that invokes the registered op under an outer deferral scope. Without the pin it reproduces "RuntimeError: Unable to cast ... to Tensor"; with it the op returns the right tensor, leaves the caller's scope untouched, and still carries its two operand flags inward. A CPU test pins the scoped() behaviour that makes the pin necessary.
SGLANG_FLASHINFER_MNNVL_CUTEDSL_AR_FUSION was deleted without a _DEPRECATED_ENVS entry, and environ.py names that registry the one place deprecations belong. Silence is worse than usual here: with the suppression block gone and Qwen3_5MoeForConditionalGeneration in _FLASHINFER_ALLREDUCE_FUSION_ARCHS, an existing command does not merely lose the CuTe DSL fusion, it auto-enables the legacy mnnvl backend instead. The three other switches deleted alongside it get entries too. The cookbook still told users to set it. Six quantized recipes in qwen3.8.jsx move the env entry to --flashinfer-allreduce-fusion-backend cute-dsl, and the prose that told them not to pass that flag alongside it -- advice this branch inverts -- is corrected. Two BF16 recipes set the switch on a shape that cannot run it, which predates this branch: the HT kernel accepts only tp in (2, 4, 8, 16) and both are TP32, and the fusion needs the deferred finalize, which FusedMoE.supports_deferred_finalize offers for NVFP4 and block-FP8 weights only. They lose the selection rather than carry a broken recipe into the new spelling, and the paragraph now says why.
The value is new in this branch, so no alias is owed.
Audited against .claude/rules and fixed what did not comply. no-dataclasses: MoeFinalizeHandoff is a msgspec.Struct. The rule grandfathers existing dataclasses but asks for migration while editing the file, and this branch creates it. no-getattr-defensive: the functools.partial probe in the AR+RMSNorm predicate narrows with isinstance instead of getattr defaults. general-code-style: the three fusion-service calls pass keywords; resolve_max_m() and prepare_cutedsl_fusion() take server_args and max_running_requests rather than the whole ModelRunner, which the rule names as the anti-pattern; the four predicates with no external caller are protected. should_defer_moe_finalize stays public -- it overrides LayerCommunicator and both models call it. comment-style: five comment blocks ran past two lines, and the env var's block broke mid-phrase. All are within the limit and break at a clause boundary, keeping the provenance of the 192 default. unit-test-admission: the ForwardFlags.scoped() case moves to the runtime-context suite, whose subsystem it belongs to. The custom-op suite moves out of test/registered/unit/, where scripts/lint/ check_registered_tests.py allows CPU registrations only, and joins the DeepSeek model tests under e2e/models. hasattr() on the all-reduce tag stays: neither remedy the rule offers fits a tensor tagged across modules, and LayerCommunicator tests it the same way, so the two halves of the handshake stay symmetric.
The test asserted an operator contract with mocked fusions and launched nothing, so e2e was the wrong kind: every other file under e2e/models brings up a server, the 23 that call no launcher directly inheriting DefaultServerBase. It only landed there because the taxonomy checker rejects CUDA registrations under unit/. It does not need a GPU. The op is registered for CUDA only, so dispatch by device fails on a CPU runner, but redispatching to the CUDA key runs the real registered implementation -- its scope management and its Tensor schema -- while the stubbed MoE keeps the arithmetic on CPU tensors. Verified with CUDA_VISIBLE_DEVICES='' and torch.cuda.is_available() false, and verified to fail with the pin removed, raising the same "Unable to cast ... to Tensor" the dispatcher raised originally. That keeps the guard on the boundary that actually failed, which neither the comment nor an AST assertion can do, and costs no GPU slot. The ForwardFlags.scoped() case stays in the runtime-context suite.
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
The cookbook targets the released sglang package and the pinned lmsysorg/sglang:qwen38 image, neither of which carries this flag yet.
…sion-shared-core # Conflicts: # python/sglang/srt/layers/flashinfer_mnnvl_cutedsl.py
kpham-sgl
left a comment
There was a problem hiding this comment.
remove ai-gen comments and trim unit tests
| # Deferring costs an HBM round trip of [M*top_k, hidden], so it pays only | ||
| # at small M; 0 disables. Measured neutral at 192 on GLM-5.2 H=6144 TP8. | ||
| SGLANG_MOE_DEFERRED_FINALIZE_MAX_TOKENS = EnvInt(192) |
There was a problem hiding this comment.
I think we are only turning on moe deferred finalize for decode only right? Btw this number feels a little hardcoded / unintuitive.
There was a problem hiding this comment.
Yes. The 192 was measured by @b8zhong on B300, and I confirmed on GB300 that it's reasonable (the crossover is around 192–223), so I think we can keep it
There was a problem hiding this comment.
Should we generalize this style for trtllm and mnnvl ar fusion too? Can do it step-by-step in later PRs yeah
There was a problem hiding this comment.
Yes we can consider this
ch-wan
left a comment
There was a problem hiding this comment.
The shared CuTe DSL service fits communicator-owned boundary processing. The inline comments cover integration with the FFN exit now on main, an overly broad EP x TP fusion restriction, and one non-blocking capability-interface cleanup.
Please preserve two existing safeguards during integration: TP1 shared-expert output must be added only once (reducing routed_partial + replicated_shared would multiply the shared contribution by TP), and Qwen3.5's terminal handoff consumer already performs finalize + all-reduce + residual + norm, so a common exit must not apply the final norm again.
Reviewed head 21ba00e86a against main 81b81664a6. Validation was static inspection plus isolated Python control-flow checks; GPU numerical correctness and performance were not rerun.
| communicate_fn = self._communicate_with_all_reduce_and_layer_norm_fn | ||
| if isinstance(communicate_fn, functools.partial): | ||
| norm_fn = communicate_fn.func | ||
| residual_input_mode = communicate_fn.keywords.get("residual_input_mode") | ||
| else: | ||
| norm_fn = communicate_fn | ||
| residual_input_mode = None |
There was a problem hiding this comment.
Non-blocking follow-up: could this eligibility check use an explicit capability exposed by the selected communication path, instead of inspecting functools.partial.func and its bound keyword arguments?
The property needed here is whether the selected path can perform the pending reduction, residual addition, and normalization with the required layouts. Function identity is an implementation detail; wrapping or reorganizing an equivalent communication path can silently disable this fusion. The current check still works with main's selector, so this can be a separate cleanup rather than a blocker for the backend extension.
…sion-shared-core Route the DeepSeek-V2 deferred MoE finalize through LayerCommunicator.ffn_exit (#40870): FfnExit now publishes defer_moe_finalize, treats a deferral as a fusion, and passes a MoeFinalizeHandoff through finish(). Move the Qwen3.8 cookbook cells off the retired SGLANG_FLASHINFER_MNNVL_CUTEDSL_AR_FUSION onto --flashinfer-allreduce-fusion-backend cutedsl.
- Drop the Qwen3.8 cookbook edits the merge re-applied: the cookbook targets the released sglang package and the pinned qwen38 image, neither of which carries the cutedsl flag yet. - Derive the workspace M bound from the published config instead of reading server_args (runtime_context.cutedsl_moe_max_num_tokens). - Fuse the plain all-reduce for hybrid EP x MoE-TP whenever both legs merge into one TP reduction, as the base communicator does; the deferred finalize keeps its EP=1 restriction. - Trim the fusion comments to the facts the code cannot show, and state where the 192-token deferral bound was measured. - Cut the unit tests to the cases that each guard a distinct failure.
- Prepare every installed fusion service from BaseRunner by scanning the module tree, so a wrapper holding the model as a submodule (Kimi-K2.5, KimiVL, DeepSeek-VL2, ...) no longer passes the check with an unprepared workspace; the per-model pre-capture hooks and the NextN stub are gone. - Fuse DeepSeek dense layers too, keeping the local reduction of a TP1-replicated MLP, since selecting cutedsl turns the legacy fusion off. - Count a layer as deferring only when its shared experts are unfused, the only case in which the dual-stream path produces a handoff. - Skip moe_output_buffer_ctx when the layer input is a handoff. - Drop the PP=1 rejection: a stage's last layer has no successor and keeps its reduction. Verified at PP2 x TP2. - Read SGLANG_MOE_DEFERRED_FINALIZE_MAX_TOKENS at MoE construction instead of through an lru_cache, so envs.override works; moe/utils.py matches main. - Remove the EP x MoE-TP clause: attn_cp_size == 1 already forces moe_dp_size == 1, so the legs always merge. - Make the deferred-finalize handoff a msgspec.Struct, split the GB300 preset selection out of _retargeted_config, make module-internal helpers protected, and call _finish_prepare_attn and finalize by keyword. - Tests: validate retargeted HT presets with FlashInfer's own kernel checks (grid taken from the preset), restore coverage for the wrapper call contract, handoff aliasing, early shared load and resolve_max_m, and add a regression for the nested wrapper.
88d16b7 to
e827f19
Compare
…5.21 (#12) * [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384) * [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405) * Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096) * [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059) * [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425) * [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878) Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> * [AMD] Add tuned dsv4 shape (sgl-project#40996) * [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388) * [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612) * Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> * [HiCache] Make host reclamation independent of transfer order (sgl-project#40512) * [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922) * [diffusion] update code owner (sgl-project#41130) * [NPU] Update CANN version to 9.1.0 (sgl-project#40524) * [PD] Enable deferred decode-side KV release by default (sgl-project#41023) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [chore] surface the cookbook to users who pip install sglang (sgl-project#40866) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960) Signed-off-by: Ting Sun <suntcrick@gmail.com> * [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> * [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805) * Support unified memory decode host pools (sgl-project#39478) Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> * [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858) Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> * [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155) * [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109) * [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079) * [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081) * [Refactor] Share the MoE output all-reduce between models (sgl-project#41097) * [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082) * [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083) * [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084) * [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120) * [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041) * MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> * MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559) Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090) * [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162) * [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046) * [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049) * [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929) Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179) * Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442) Co-authored-by: Caio Rocha <caiorocha@microsft.com> * [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960) Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> * [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> * [Score API] Setwise Scoring Support (sgl-project#38965) * [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091) Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988) * [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907) * [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * Support XQA backend for SpecDec verify (sgl-project#32269) * [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793) Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> * [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> * [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * [feat] add a system one compatible /v1/systemone route (sgl-project#41208) * fix(openai): reject request-supplied chat_template by default (sgl-project#28135) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159) * [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201) Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> * [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456) * [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092) * [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619) * [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193) * [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194) * [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195) * [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196) * [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197) * [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198) * [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191) * [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199) * [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200) * [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224) * [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417) Co-authored-by: mickqian <mickqian@users.noreply.github.com> * [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486) * Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132) * [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064) Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: jacky.cheng <yichiche@amd.com> * fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248) Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> * Fix multimodal feature offload races (sgl-project#40621) Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986) * [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246) Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> * [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215) * [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216) * [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379) * [KVCache] Support lmcache unified radix cache (sgl-project#38652) Signed-off-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389) Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> * dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280) * [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660) Co-authored-by: BBuf <1182563586@qq.com> * [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284) * [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286) * [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285) * [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490) * [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291) Co-authored-by: DarkSharpness <2040703891@qq.com> * [diffusion] model: support Anima Base v1.0 (sgl-project#41011) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261) * [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297) * [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505) * [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150) Co-authored-by: jacky.cheng <yichiche@amd.com> * [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Restore logical kernel groups and test organization (sgl-project#41243) * [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188) * [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223) Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> * [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556) Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252) * [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253) * [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254) * [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255) * [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256) * [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257) * [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321) * [AMD] Fix jit broken on rocm env (sgl-project#41356) * Make sliding-window caching and speculative batch padding extensible (sgl-project#41325) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> * [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377) * dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019) Co-authored-by: x <x> * [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [mem_cache] Never free the protected prefix on request release (sgl-project#41312) * [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345) * [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416) Co-authored-by: aimicahchen <aimicahchen@tencent.com> * [kv-shard 3/4] Enable Control Plane B (sgl-project#39964) Co-authored-by: Zhangheng <hzh0425@apache.org> * [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166) * [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305) Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> * [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281) * [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192) * [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684) * [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990) * [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418) Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> * [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272) Signed-off-by: rockdu <kangrdu@gmail.com> * [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> * [bench] Take each request's prompt length from the server (sgl-project#39889) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> * [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266) * Add MiniMax arch fallback to auto parser resolution (sgl-project#40930) * [HiCache] Add the page-unified KV load-back JIT kernel (sgl-project#39726) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> * [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458) * [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: kpham-sgl <khoa.pham@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Port chat_parsing core (sgl-project#40477) Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> * [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417) * [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418) * [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419) * [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420) * [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421) * [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422) * [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423) * [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424) * [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425) * [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426) * [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427) * [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428) * [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429) * [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430) * [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431) * [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432) * [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433) * [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434) * [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435) * [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436) * [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437) * [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438) * [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439) * [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440) * [Refactor] Build the boundary into any stage with one construction (sgl-project#41441) * [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442) * [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443) * [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075) * [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814) * [diffusion] feat: support multiple task types for pipelines (sgl-project#38762) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [CI] fix CI regression on xeon (sgl-project#41002) * [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378) * [AMD] Register Triton data movement tests in PR CI (sgl-project#41137) * [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> * [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> * [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664) Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899) Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Rust server unify datapath for mm and generate requests (sgl-project#39679) * feat(npu): Support returning indexer top-k results (sgl-project#39060) * Fix chat template cache key order (sgl-project#41517) * docs: add prefill context parallelism guide and design draft (sgl-project#39354) * [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220) Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> * [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164) * [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165) * [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539) * [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943) * [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> * [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161) Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> * [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133) * [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144) * [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339) * [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446) * [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643) * [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235) * [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987) * [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597) * dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> * [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588) Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> * [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469) * [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196) Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118) Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> * [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590) Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> * [diffusion] model: support flux 3 action robot policies (sgl-project#41066) * [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627) Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> * [npu]support NPU 910C L2 memcache offload (sgl-project#41527) * [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600) * Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Refactor] Group communicator fusion and CP adapters (sgl-project#41547) * [Refactor] Centralize decoder output access (sgl-project#41548) * [Refactor] Carry residual state across stage boundaries (sgl-project#41549) * [Refactor] Capture auxiliary states at residual reads (sgl-project#41550) * [Refactor] Select reduction fusion at the consumer (sgl-project#41551) * [Refactor] Construct independent decoder stage boundaries (sgl-project#41552) * [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553) * [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554) * [Refactor] Rename the module to layer_boundary (sgl-project#41555) * [Refactor] Group layer boundary unit tests (sgl-project#41556) * [Refactor] Document layer boundary contracts and integration (sgl-project#41557) * [Rust] Extract a transport-neutral frontend core (sgl-project#39385) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618) * [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804) Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528) * [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059) * [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> * [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802) Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21 The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in the post-port commits #9 (hybrid sliding-window through the kv-cache plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory reads attention_backends()). Upstream's structure wins and the seam is re-applied on it: - server_args.py is now a thin record over arg_groups/: the --kv-cache-dtype choices are hoisted into arg_groups/choices.py as KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a plugin backend move to arg_groups/model_hook.py. - ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's resolution hooks (arg_groups/resolution_hooks.py) are the official slot, and turbo-attn's plugin now pairs the flags there. - Plugin dtype reads go through the resolved model bag (get_model()), not the raw ServerArgs record, which v0.5.21 leaves as operator input. - HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it holds and stands in as the per-layer cost, so upstream's new draft-SWA and unified-pool formulas apply unchanged; layer ids come from kvc.layer_info, as upstream's own SWA pool takes them. - Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21 a PP draft may load its own embedding (pp_draft_embedding). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> Signed-off-by: Ting Sun <suntcrick@gmail.com> Signed-off-by: chunxiaozheng <1179548172@qq.com> Signed-off-by: rockdu <kangrdu@gmail.com> Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com> Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com> Co-authored-by: Alan Kao <akao@amd.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca> Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: 黄孝君 <dingfangsu23@gmail.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Ting SUN <suntcrick@gmail.com> Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu> Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> Co-authored-by: ChangLiu0709 <cliu1004@amd.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Thomas Wang <thomawan@amd.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: Chunan Zeng <zcnrex@gmail.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com> Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com> Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com> Co-authored-by: Caio Rocha <caiorocha@microsft.com> Co-authored-by: metamergebot <metamergebot@gmail.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com> Co-authored-by: jacky.cheng <yichiche@amd.com> Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com> Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> Co-authored-by: Richard Wang <wangrichard08@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: Xinyi Song <xinyis10@illinois.edu> Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: ashwini rathi <ashwini.rathi@intel.com> Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com> Co-authored-by: Even Zhou <even.y.zhou@outlook.com> Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: cctry <csycfl@gmail.com> Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> Co-authored-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai> Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com> Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: WMC <tnwilly@gmail.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com> Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com> Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> Co-authored-by: Ankith Averineni <saverine@amd.com> Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com> Co-authored-by: Yuan Luo <yuan.luo@hotmail.com> Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com> Co-authored-by: CuzMi <simon.weijie@gmail.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com> Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com> Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com> Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com> Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com> Co-authored-by: flb_ <floatlibai@gmail.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> Co-authored-by: Peng Wu <peng@thinkingmachines.ai> Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com> Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com> Co-authored-by: karverma-amd <karan.verma@amd.com> Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: triple-mu <gpu@163.com> Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com> Co-authored-by: Hrithvik Alex <halex623@gmail.com> Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com> Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> Co-authored-by: sglang-bot <sglangbot@gmail.com> Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> Co-authored-by: Jialin Ouyang <jialino@meta.com> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> Co-authored-by: James <445169590@qq.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com> Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com> Co-authored-by: Jiacong Fang <zldrobit@126.com>
* [feature] add per-item candidate token scoring and calibration (sgl-project#40826) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384) * [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405) * Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096) * [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059) * [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425) * [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878) Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> * [AMD] Add tuned dsv4 shape (sgl-project#40996) * [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388) * [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612) * Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> * [HiCache] Make host reclamation independent of transfer order (sgl-project#40512) * [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922) * [diffusion] update code owner (sgl-project#41130) * [NPU] Update CANN version to 9.1.0 (sgl-project#40524) * [PD] Enable deferred decode-side KV release by default (sgl-project#41023) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [chore] surface the cookbook to users who pip install sglang (sgl-project#40866) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960) Signed-off-by: Ting Sun <suntcrick@gmail.com> * [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> * [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805) * Support unified memory decode host pools (sgl-project#39478) Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> * [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858) Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> * [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155) * [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109) * [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079) * [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081) * [Refactor] Share the MoE output all-reduce between models (sgl-project#41097) * [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082) * [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083) * [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084) * [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120) * [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041) * MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> * MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559) Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090) * [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162) * [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046) * [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049) * [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929) Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179) * Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442) Co-authored-by: Caio Rocha <caiorocha@microsft.com> * [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960) Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> * [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> * [Score API] Setwise Scoring Support (sgl-project#38965) * [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091) Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988) * [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907) * [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * Support XQA backend for SpecDec verify (sgl-project#32269) * [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793) Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> * [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> * [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * [feat] add a system one compatible /v1/systemone route (sgl-project#41208) * fix(openai): reject request-supplied chat_template by default (sgl-project#28135) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159) * [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201) Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> * [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456) * [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092) * [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619) * [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193) * [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194) * [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195) * [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196) * [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197) * [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198) * [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191) * [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199) * [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200) * [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224) * [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417) Co-authored-by: mickqian <mickqian@users.noreply.github.com> * [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486) * Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132) * [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064) Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: jacky.cheng <yichiche@amd.com> * fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248) Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> * Fix multimodal feature offload races (sgl-project#40621) Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986) * [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246) Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> * [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215) * [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216) * [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379) * [KVCache] Support lmcache unified radix cache (sgl-project#38652) Signed-off-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389) Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> * dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280) * [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660) Co-authored-by: BBuf <1182563586@qq.com> * [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284) * [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286) * [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285) * [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490) * [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291) Co-authored-by: DarkSharpness <2040703891@qq.com> * [diffusion] model: support Anima Base v1.0 (sgl-project#41011) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261) * [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297) * [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505) * [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150) Co-authored-by: jacky.cheng <yichiche@amd.com> * [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Restore logical kernel groups and test organization (sgl-project#41243) * [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188) * [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223) Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> * [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556) Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252) * [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253) * [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254) * [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255) * [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256) * [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257) * [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321) * [AMD] Fix jit broken on rocm env (sgl-project#41356) * Make sliding-window caching and speculative batch padding extensible (sgl-project#41325) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> * [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377) * dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019) Co-authored-by: x <x> * [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [mem_cache] Never free the protected prefix on request release (sgl-project#41312) * [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345) * [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416) Co-authored-by: aimicahchen <aimicahchen@tencent.com> * [kv-shard 3/4] Enable Control Plane B (sgl-project#39964) Co-authored-by: Zhangheng <hzh0425@apache.org> * [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166) * [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305) Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> * [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281) * [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192) * [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684) * [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990) * [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418) Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> * [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272) Signed-off-by: rockdu <kangrdu@gmail.com> * [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> * [bench] Take each request's prompt length from the server (sgl-project#39889) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> * [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266) * Add MiniMax arch fallback to auto parser resolution (sgl-project#40930) * [HiCache] Add the page-unified KV load-back JIT kernel (sgl-project#39726) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> * [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458) * [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: kpham-sgl <khoa.pham@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Port chat_parsing core (sgl-project#40477) Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> * [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417) * [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418) * [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419) * [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420) * [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421) * [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422) * [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423) * [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424) * [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425) * [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426) * [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427) * [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428) * [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429) * [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430) * [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431) * [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432) * [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433) * [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434) * [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435) * [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436) * [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437) * [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438) * [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439) * [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440) * [Refactor] Build the boundary into any stage with one construction (sgl-project#41441) * [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442) * [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443) * [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075) * [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814) * [diffusion] feat: support multiple task types for pipelines (sgl-project#38762) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [CI] fix CI regression on xeon (sgl-project#41002) * [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378) * [AMD] Register Triton data movement tests in PR CI (sgl-project#41137) * [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> * [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> * [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664) Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899) Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Rust server unify datapath for mm and generate requests (sgl-project#39679) * feat(npu): Support returning indexer top-k results (sgl-project#39060) * Fix chat template cache key order (sgl-project#41517) * docs: add prefill context parallelism guide and design draft (sgl-project#39354) * [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220) Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> * [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164) * [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165) * [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539) * [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943) * [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> * [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161) Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> * [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133) * [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144) * [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339) * [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446) * [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643) * [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235) * [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987) * [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597) * dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> * [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588) Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> * [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469) * [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196) Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118) Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> * [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590) Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> * [diffusion] model: support flux 3 action robot policies (sgl-project#41066) * [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627) Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> * [npu]support NPU 910C L2 memcache offload (sgl-project#41527) * [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600) * Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Refactor] Group communicator fusion and CP adapters (sgl-project#41547) * [Refactor] Centralize decoder output access (sgl-project#41548) * [Refactor] Carry residual state across stage boundaries (sgl-project#41549) * [Refactor] Capture auxiliary states at residual reads (sgl-project#41550) * [Refactor] Select reduction fusion at the consumer (sgl-project#41551) * [Refactor] Construct independent decoder stage boundaries (sgl-project#41552) * [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553) * [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554) * [Refactor] Rename the module to layer_boundary (sgl-project#41555) * [Refactor] Group layer boundary unit tests (sgl-project#41556) * [Refactor] Document layer boundary contracts and integration (sgl-project#41557) * [Rust] Extract a transport-neutral frontend core (sgl-project#39385) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618) * [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804) Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528) * [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059) * [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> * [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802) Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> --------- Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> Signed-off-by: Ting Sun <suntcrick@gmail.com> Signed-off-by: chunxiaozheng <1179548172@qq.com> Signed-off-by: rockdu <kangrdu@gmail.com> Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com> Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com> Co-authored-by: Alan Kao <akao@amd.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca> Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: 黄孝君 <dingfangsu23@gmail.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Ting SUN <suntcrick@gmail.com> Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu> Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> Co-authored-by: ChangLiu0709 <cliu1004@amd.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Thomas Wang <thomawan@amd.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: Chunan Zeng <zcnrex@gmail.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com> Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com> Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com> Co-authored-by: Caio Rocha <caiorocha@microsft.com> Co-authored-by: metamergebot <metamergebot@gmail.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com> Co-authored-by: jacky.cheng <yichiche@amd.com> Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com> Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> Co-authored-by: Richard Wang <wangrichard08@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: Xinyi Song <xinyis10@illinois.edu> Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: ashwini rathi <ashwini.rathi@intel.com> Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com> Co-authored-by: Even Zhou <even.y.zhou@outlook.com> Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: cctry <csycfl@gmail.com> Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> Co-authored-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai> Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com> Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: WMC <tnwilly@gmail.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com> Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com> Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> Co-authored-by: Ankith Averineni <saverine@amd.com> Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com> Co-authored-by: Yuan Luo <yuan.luo@hotmail.com> Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com> Co-authored-by: CuzMi <simon.weijie@gmail.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com> Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com> Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com> Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com> Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com> Co-authored-by: flb_ <floatlibai@gmail.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> Co-authored-by: Peng Wu <peng@thinkingmachines.ai> Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com> Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com> Co-authored-by: karverma-amd <karan.verma@amd.com> Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: triple-mu <gpu@163.com> Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com> Co-authored-by: Hrithvik Alex <halex623@gmail.com> Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com> Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> Co-authored-by: sglang-bot <sglangbot@gmail.com> Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> Co-authored-by: Jialin Ouyang <jialino@meta.com> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> Co-authored-by: James <445169590@qq.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com> Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com> Co-authored-by: Jiacong Fang <zldrobit@126.com>
Accuracy Tests
sgl-eval run gsm8k, 1319 examples, 8x GB300, TP8, NVFP4RadixArk/GLM-5.3-NVFP4RadixArk/Qwen3.8-2.4T-A95B-NVFP4Both fused patterns also match a torch reference to 0.031 (bf16 rounding) across all four M routes at hidden 7168 and 6144, including with HT forced off. Technically, TP for BS = 512 is not recommended, but it might be useful in agg. prefill
Speed Tests and Profiling
nvidia/GLM-5.2-NVFP4, 8x B300, 1024 in / 1024 out, decode tok/s (batch / median ITL). All arms:SGLANG_FLASHINFER_AUTOTUNE_CACHE=0.Base:
SGLANG_MOE_DEFERRED_FINALIZE_MAX_TOKENS=0SGLANG_MOE_DEFERRED_FINALIZE_MAX_TOKENS=192)--flashinfer-allreduce-fusion-backend cutedsl SGLANG_MOE_DEFERRED_FINALIZE_MAX_TOKENS=0--flashinfer-allreduce-fusion-backend cutedslTP8
TP4
Also confirmed on 4x GB300 at TP4, where the defer/no-defer crossover sits between batch 192 and 224, making the shipped 192 the lower edge of the optimal
[192, 223].CI States
Latest PR Test (Base): 🚫 Run #35991532729
Latest PR Test (Extra): 🚫 Run #35991532242
Latest PR Test (AMD ROCm 10): ⏳ Run #35991532652