Repository navigation
[Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts - #37462
Conversation
43257a1 to
45cdb63
Compare
kpham-sgl
left a comment
There was a problem hiding this comment.
Thank you for the contribution! Some very high level comments before I review further
- Please point your agents to these rules we have in
.claude/rules/. Some to point out
- Avoid extensive AI comments and docstrings
- Avoid using getattr/hasattr
- Avoid writing unnecessary unit tests
- Can you add an E2E test with a small target for this new algorithm? Make sure to use existing spec decoding test kit?
|
Can you also resolve conflicts @mrusanovsky? Thanks! |
nvpohanh
left a comment
There was a problem hiding this comment.
[by Codex] Five inline review findings are attached.
45cdb63 to
c87e8f2
Compare
Reading the rest of getattr/hasattr : all removed:
Remaining Comments and docstrings : docstring lines in non-test source down from 476 to 270. Removed Tests : 48 CPU to 37, 16 GPU to 12. Back to 40 and 13 after @nvpohanh's findings needed guards. Also fixed : E2E test : Run on an H100: That accept length is the mixin's figure over 200 short 5-shot completions, not the block-weighted Registered
One gap left: no small-target drafter exists, all of ours target Qwen3-8B, so it runs on In addition, I resolved conflicts. Rebased on |
nvpohanh
left a comment
There was a problem hiding this comment.
[by Codex] One inline review finding is attached.
| # is correct: the argmax draft is still verified losslessly, just target-only. | ||
| # Checked after the selector branch because the two heads are mutually | ||
| # exclusive: a selector worker returns above and never reads `lilicorr`. | ||
| if self.lilicorr is not None and _LILICORR_SAMPLING_ENABLED: |
There was a problem hiding this comment.
[by Codex] Severity: functional | Confidence: High
Issue
This return treats LiLiCorr sampling as supported on NPU. The selector path explicitly disables sampling there, but LiLiCorr later sets _selector_sample and _selector_sampling_accept calls accept_sampling, which unconditionally launches chain_speculative_sampling_triton. NPU tensors cannot run that kernel, so LILICORR_SAMPLING=1 with a non-greedy request fails during verification instead of taking the existing fallback.
Fix
Gate LiLiCorr sampling on the same supported-device condition as the selector, or add a supported NPU accept implementation. Add a focused NPU-path test for the fallback.
There was a problem hiding this comment.
Fixed. The selector gates both publish sites on _selector_sampling_enabled; LiLiCorr's didn't, and the accept branch keys on _selector_sample is not None alone, so a non-greedy request on NPU reached chain_speculative_sampling_triton.
Gating only the publish would be worse than the crash: the draft would still sample while verify treated the drawn token as a point mass, making acceptance p(x) instead of min(1, p(x)/q(x)) with a residual from the wrong q : not distribution-preserving. So sampling_enabled is threaded into the sampler as done for the selector, deciding which body compiles. Unsupported devices take the argmax commit and the existing verify, with the warning that path already emits.
Fallback test added in test_dflash_logits.py, stub-based so it runs on CPU; the dispatch test now also asserts the flag reaches the builder. On CUDA the effective flag is unchanged : re-ran on an H100 and τ is identical to before: 6.8158 sampled, 6.3944 argmax, T=0 gate 7.557260716552241 on 51999 blocks.
### What does this PR do? Type of change: new feature Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a new `projector_type` on the existing `dflash` mode — plus three DFlash-wide improvements that apply to every variant, and an optional composition with DFlash2's grouped convolutions. A DFlash drafter is trained on per-position marginals rather than on the joint block distribution, so its drafted tokens are individually plausible yet jointly incoherent. LiLiCorr keeps the top-`k` candidates the backbone already produces at each block position, scores transitions between adjacent candidates with a small two-layer transformer, and commits a path through the lattice greedily. Serving is unchanged in kind: verify still checks every drafted token against the target, so the emitted distribution is untouched and only acceptance length moves. - Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530) (arXiv:2608.20530) - Blog: https://research.nvidia.com/labs/nemotron/lilicorr/ - **Companion PR — serving support:** [sgl-project/sglang#37462](sgl-project/sglang#37462) This PR is the **training** half. It trains the drafters and exports them; the companion PR above is what serves the resulting checkpoints, and is what the comparison table below was measured through. **What is in the commits** | | | | --- | --- | | LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`, conversion routing, config fields, export | | Three DFlash-wide features | fp32 master weights for the draft, draft activation checkpointing, and a DDP hang fix — all default-off or behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and `dflash2` alike | | Optional grouped convolutions | composes LiLiCorr with DFlash2's `DFlashGroupedConv`; see the dependency note below | | Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` | | CPU unit tests, CHANGELOG, one launcher example | | **⚠️ The convolutions depend on the DFlash2 branch, and cannot run until it merges.** `modeling_lilicorr.py` imports `DFlashGroupedConv` from `modeling_dflash2`, which today exists only on `haoguo/dflash2-support`. The class is **imported rather than copied on purpose** — it is the only way the two variants cannot drift apart arithmetically — but the consequence is that the convolutional recipe cannot run against `main` as it stands. So the import is **deferred into `_install_sublayer_convs`** rather than taken at module scope. Everything else in this PR, including the plain LiLiCorr reranker, has no DFlash2 dependency at all and works on `main` today; an eager import would have made the whole plugin unimportable for the sake of one optional feature. Requesting the convolutions without DFlash2 present raises an `ImportError` naming the two config keys to remove, rather than failing at import time. **This PR carries two of @h-guo18's commits, with authorship and sign-off preserved.** Both are independent of DFlash2 itself and both are needed here: - `1419d47e`, the no-op sublayer seam. Without it `DFlashDecoderLayer.forward` never calls the wrappers the convolutions install onto, so the modules would be built, counted and exported while computing nothing. It is arithmetically an identity on its own. - `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a top-level `rope_theta` and a `rope_parameters` dict; the real base lives in the dict while the class default (10,000 for Qwen3) stays visible as the flat attribute. Reading the flat field first builds a draft whose RoPE base is 100× off a Qwen3-8B target's, which trains and exports without complaint. Both the training-side enforcement and the exporter's `_get_rope_theta` are affected on `main` today. Both are @h-guo18's work and belong to their branches; they are carried here only so that this PR stands on its own. **If those branches land first, this PR can be rebased onto them and the two commits dropped**, and they can equally be split out now if that is easier to review. The same applies to `dflash_fp32_master_weights`, which is also in flight on `haoguo/dflash-fp32-master-weights`. The field name is shared deliberately so that there is only ever one knob rather than two spellings of it, and both versions default to off. Whichever lands first, this PR can be rebased onto it. ### Usage Train with the shipped recipe: ```python from modelopt.recipe import load_recipe config = load_recipe("general/speculative_decoding/lilicorr.yaml") # Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified), # DFlash decay objective at gamma 7.0, fp32 master weights for the draft. ``` Or convert directly: ```python import modelopt.torch.speculative as mtsp config = { "dflash_block_size": 16, "dflash_loss_objective": "decay", "dflash_loss_decay_factor": 7.0, "dflash_fp32_master_weights": True, "dflash_lilicorr_w_ce": 0.25, "dflash_lilicorr_w_margin": 0.0, "dflash_lilicorr_w_pen": 0.25, "dflash_architecture_config": { "num_hidden_layers": 5, "projector_type": "lilicorr", "lilicorr_candidate_topk": 8, # Optional, and all-or-nothing: adding these two keys wraps every draft # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant. # "conv_kernel_size": 2, # "conv_group_size": 16, }, } mtsp.convert(model, [("dflash", config)]) ``` ### Results Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one matched contract** — the same corpus, schedule and block geometry for every arm, so no row carries a training advantage. Training data is NVIDIA's [Nemotron Post-Training Dataset v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2) with the multilingual split excluded, generated from the target with **thinking disabled**; **6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7; **8 nodes × 8 H100, global batch size 64** (one sequence per device, no gradient accumulation). All six were then exported and served through SGLang on a **single H100 80GB**, `tp_size 1`, at concurrency 1, greedy, `fa3`, mean of two replicates, with the whole node held exclusive per benchmark. Speedup is output tokens/s against an autoregressive baseline measured in the same allocation. Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second fastest**: | benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino | DFlash | |---|---|---|---|---|---|---| | gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 / 5.06x | 7.225 / 4.87x | 6.341 / 4.59x | | math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 / 6.49x | 8.976 / 6.25x | 7.909 / 5.88x | | aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 / 5.91x | 8.066 / 5.77x | 7.126 / 5.44x | | humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081 / 3.95x | 6.864 / 3.73x | 6.156 / 3.68x | | mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x | 5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x | | livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x | 7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x | | alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x | 3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x | | mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 / 2.80x | 3.948 / 2.84x | 3.478 / 2.67x | **Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a 164-prompt slice, where DFlash2 is ahead by 0.5%. `DFlash` is the deliberately head-free control; every head clears it by +7.60% to +21.67% on acceptance, which is the check that a head actually loaded. Reproducing the `LiLiCorr+conv` column additionally needs the DFlash2 variant. Acceptance length is bit-reproducible under greedy decoding and its replicate spread here was 0.00% on every benchmark; throughput has a ~0.2% floor. ### What `dflash_fp32_master_weights` does, and what it is worth Today the draft is cast to the frozen base model's dtype — bf16 — before the optimizer is built. AdamW then allocates its moments with `zeros_like(p)`, so the **optimizer state becomes bf16 too**. That is the problem: bf16 has too few mantissa bits to represent the small updates Adam's second moment accumulates, so those updates round away and the effective step size decays on its own, independently of the learning-rate schedule. The flag is standard mixed precision instead: the draft's master weights stay in fp32 while the matmuls run in bf16. It requires a bf16 autocast around the forward, which HF `Trainer` supplies under `TrainingArguments.bf16`. Paths that do not go through the Trainer — evaluation, `pseudo_speculative_generate`, a plain `convert()` and forward — currently need the caller to supply it, and no shipped recipe exercises those (`estimate_ar: false`, `do_eval: false`). Making the draft supply its own autocast is a follow-up, held back from here on review because it touches every DFlash variant and wants e2e coverage of the existing recipes. Compute speed is unchanged. The cost is memory, about 12 bytes per parameter for the weight plus Adam's two moments instead of 6, plus a doubled gradient all-reduce under DDP, since fp32 parameters mean fp32 gradients. Under FSDP2 that second cost is what `MixedPrecisionPolicy(reduce_dtype=...)` exists to control. It is worth **7 to 14 percent of acceptance length**, measured at the end of training on gsm8k, and it helps every projector type: | arm | bf16 | fp32 | Δ acceptance length | | --- | ---: | ---: | ---: | | LiLiCorr | 6.8670 | 7.5573 | **+10.05%** | | DFlash2 | 6.7396 | 7.2518 | **+7.60%** | | Domino | 6.5854 | 7.2252 | **+9.71%** | | DSpark | 6.4621 | 7.3752 | **+14.13%** | | DFlash | 5.9030 | 6.3412 | **+7.42%** | Every arm in the comparison table above was trained with it on, and **both shipped recipes set it `true`**, so the documented path gets it. It defaults to **off**, so no existing DFlash, Domino or DSpark run changes behaviour. Both shipped LiLiCorr recipes set it `true`, which is the arithmetic their numbers were trained with. Flipping the default is a reasonable follow-up once the autocast above is in. The draft is drawn in fp32 and, under this flag, kept there; an unpromoted run rounds the same draw to the base model's dtype. So the bf16 and fp32 rows of the table above start from the same initialization at the precision each trains in, rather than from two different draws. A unit test pins that. The flag also survives a resume. `modify()` runs under `from_pretrained` with the base model still on meta and cannot place the draft at all, so `restore_draft_precision` re-applies the dtype, the device and the rotary buffer once the weights are loaded and before the Trainer builds the optimizer — the last point that can still decide the Adam moment dtype. It also reloads the draft's tensors at the dtype they were saved in, since checkpoints store the draft in fp32 while the base is bf16 and `dtype="auto"` gives every tensor one dtype. @h-guo18 has the same field in flight on `haoguo/dflash-fp32-master-weights`, plus an HF-format-resume fix this PR does not have. The name is shared deliberately so there is only ever one knob; whichever lands first, the other should be dropped rather than merged. ### Testing - **257 CPU unit tests pass** across `tests/unit/torch/speculative/`, including the existing DFlash, Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr specifically: conversion routing, head geometry, the required-field validation, the three-term objective and its absolute weights, gradient reach into both the head and the drafter body, and the export contract. - Both recipes load and validate through `modelopt.recipe.load_recipe`. - The three DFlash-wide changes are covered behaviourally: the fp32 flag is checked on the optimizer's moment dtypes rather than only on parameters, since the moments are the point of the change, and on the initialization described above; activation checkpointing is asserted to leave draft gradients bit-identical with the flag on and off; and the rotary buffer is asserted present after `modify()` on a real device while still deferred on meta, which is the case the laziness existed for. - The resume path has its own test: after a `save_pretrained` / `from_pretrained` round trip, `restore_draft_precision` is asserted to return the draft to fp32 with its stored weights intact and its Adam moments in fp32. Without it the draft comes back in the base dtype with the flag still set, which is the failure it exists to prevent. - `TestDFlashLazyRotaryEmb` was updated rather than left passing: it asserted the rotary buffer does *not* exist after convert, and the DDP fix deliberately changes that on non-meta devices. The replacement pins the refined invariant in both directions. - The published checkpoints were trained with this arithmetic, verified rather than assumed: a fingerprint over draft initialisation, loss and gradients is compared against the pre-review tree for both `dflash` and `lilicorr`. Loss and gradients are **bitwise identical**. Initialisation moves, by less than bf16 resolution, and that is the single-dtype change described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — every addition is opt-in. The new `projector_type` is selected only by config, `dflash_fp32_master_weights` defaults to off, and the activation-checkpointing and DDP fixes preserve behaviour. No existing default changes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted from https://github.com/sgl-project/SpecForge/...` headers for the DFlash backbone and loss they derive from (Apache-2.0), matching the attribution already on `hf_dflash.py` in this repo. The two commits described above are @h-guo18's, cherry-picked with authorship and sign-off preserved. - Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the updated rotary test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` once opened. ### Additional Information The convolutional recipe is the memory worst case: at an 8B target, combined with fp32 master weights, it may need `training.gradient_checkpointing: true` to fit on 80 GiB, and it fits without at 4B. Checkpointing is mathematically neutral — same objective, same data order, same resulting model — but it trades step time for memory, so a run using it is not step-time-comparable with one that does not. The recipe header says so. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added LiLiCorr speculative decoding with candidate-lattice reranking, configurable objectives, metrics, export support, and optional grouped convolutions. * Added FP32 master-weight support with improved mixed-precision behavior and gradient checkpointing. * Added LiLiCorr training recipes and a Qwen3-8B launcher configuration. * **Bug Fixes** * Improved rotary-embedding configuration handling and corrected DFlash distributed-training hangs. * Added validation for invalid LiLiCorr configurations and improved exported reranking metadata. * **Documentation** * Expanded guidance for FP32 master weights, training workflows, and LiLiCorr configuration. * **Tests** * Expanded coverage across training, evaluation, generation, export, and checkpoint workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
DFlash is trained on per-position marginals rather than on the joint block distribution, so its drafted tokens are individually plausible yet jointly incoherent. LiLiCorr keeps the top-k candidates DFlash already produces at each block position and scores transitions between adjacent candidates with a small two-layer transformer, then commits a path through the resulting lattice greedily, left to right, instead of taking the per-slot argmax. Verify is untouched, so outputs remain distributionally identical to the target model's. This adds no SpeculativeAlgorithm, no worker subclass and no registration. It rides --speculative-algorithm DFLASH and is selected by the checkpoint declaring architectures: ["LiLiCorrDraftModel"], exactly as DFlash2DraftModel selects the candidate selector. Of the 11 files touched, only two already existed: dflash_worker_v2.py gains four dispatch seams (+39) and the speculative-decoding docs gain a subsection (+19). Nothing is deleted or modified anywhere. Two Triton kernels carry most of the new non-model code, and both replace a composition of existing ops that costs extra passes over the vocabulary: lilicorr_topk_lse returns an exact per-row top-k and the full-vocabulary log-partition from one read of [n, V], where topk plus a logsumexp epilogue would read it three times and materialize two more [n, V] temporaries; lilicorr_greedy_path folds the whole left-to-right commit into one launch instead of roughly three per slot. Both fall back to a value-identical torch implementation off CUDA. Tests: 43 CPU tests, plus 16 GPU tests pinning both kernels against those torch references on device (tile-boundary vocabulary sizes, bf16 logits, the narrow-vocabulary fallback, and the greedy commit at k = 1..16 including tie-breaking). Paper: https://arxiv.org/abs/2608.20530 Blog: https://research.nvidia.com/labs/nemotron/lilicorr/
Condense multi-line comments and docstrings per comment-style.md (no code changes; AST-identical modulo docstrings). Drop the non-unit-stride walk test and the full-vocab log_softmax test, and fold the two sampled-path parity tests into one parametrized case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LILICORR_SAMPLING becomes SGLANG_ENABLE_LILICORR_SAMPLING and SGLANG_LILICORR_REQUIRE_SAMPLING moves to an EnvBool; EnvBool already rejects non-boolean values, so the local parser goes. Register lilicorr_topk_lse and lilicorr_sample_path in the speculative kernel inventory, and drop the hasattr guard in the E2E test teardown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Remove file/class docstrings, section banners, return-shape recaps and comments restating the code; keep only non-obvious design constraints. No code changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| return vals.float(), ids.to(torch.int64), lse | ||
|
|
||
|
|
||
| def lilicorr_topk_lse( |
There was a problem hiding this comment.
Reuse DFlash2's compute_candidates (radix top-k + TP gather) and add only the lse, instead of a new top-k kernel and TP combine. It also keeps quantized lm_head support, which the eager path here breaks.
There was a problem hiding this comment.
Done for the structure: compute_candidates body is now candidate_topk() in models/dflash.py and LiLiCorr calls it, so there is one projection (through quant_method), one TP gather, and our TP combine is gone.
One difference: with with_partition=True the top-k and lse come from a single fused pass instead of radix top-k plus a separate logsumexp. The second pass re-reads the [N, V] logits, which dominates at batch: 143 vs 680 us per step at c=32 (56 vs 69 us at c=1; H100, bf16, V=151,936, CUDA graph), and about -1.5% output throughput end to end at c=32. The selector keeps _radix_topk.
…ader Resolve SGLANG_ENABLE_LILICORR_SAMPLING and the REQUIRE guard only when a LiLiCorr head is loaded, instead of at import on every DFLASH server. Move the grouped-convolution coverage check from LiLiCorrDraftModel into DFlashDraftModel.load_weights, since it validates backbone weights. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| ) | ||
|
|
||
| anchor_state = self._project_anchor(anchor_hidden, anchor_valid) | ||
| # Materialized rather than a stride-0 broadcast, which measured -1.75pp. |
There was a problem hiding this comment.
The -1.75pp was measured on the compiled body (see 56a1572); without compile you measured +0.73% for broadcast. Compile is gone, so pass self._attn_bias.unsqueeze(0) straight to SDPA and drop the per-step expand/reshape copy, or re-measure without compile if you want to keep it.
There was a problem hiding this comment.
Done, passing self._attn_bias.unsqueeze(0) to SDPA. Re-measured without compile, it's neutral: -0.004% / -0.04% tok/s at c=1, +0.46% / -0.19% at c=32 (two A/B runs on one node, the second with the arms swapped), acceptance identical.
Review, kpham-sgl: reuse DFlash2's compute_candidates (radix top-k + TP gather) and add only the lse, rather than carrying a second top-k kernel and a second TP combine. He also noted the eager path breaks a quantized lm_head, which it did: lilicorr_candidates read lm_head.weight directly, and a packed weight is [V, H/pack], so the matmul could not even be shaped. Lifts compute_candidates' body into candidate_topk(), which gains an opt-in full-vocabulary log-partition: a logsumexp over the projection it already computes, combined across shards by a second logsumexp. The padded tail is already -inf and logsumexp ignores it, so no extra mask. Deletes lilicorr_topk_lse and its tiled scan/select kernels, the torch reference, _combine_across_ranks, and the static logits buffer they needed. The projection now goes through quant_method, so the worker screens LiLiCorr on the selector's gate and a quantized head folds instead of dropping to an eager path that could not run. Also drops LiLiCorrRMSNorm for layers.layernorm.RMSNorm, per review: the DFlash backbone already runs that module inside this same draft graph.
…e logits logsumexp(logits.float()) upcasts the whole [N, V] logits tensor, so it pays a second full-size allocation and read. Take the row max from the top-k, which comes back sorted, and let only the reduction be fp32 via sum(dtype=). Measured on an H100 at [480, 151936] bf16: 895 us -> 359 us, and 53 -> 40 us at [15, 151936]. logsumexp on the native dtype is 426 us but carries 6e-2 of error, which a log-prob cannot take; this form is 2.6e-4, under the bf16 logits' own ~4e-3.
|
Addressed the latest review, replies inline. One small addition, 17 lines in |
Keeps DFlashDraftModel.load_weights identical to main; the check stays scoped to LiLiCorr drafts as originally written. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
/rerun-tests test/registered/core/test_basic_sanity_dflash.py test/registered/core/test_basic_sanity_dspark.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_mixed_chunk.py test/registered/spec/test_gemma4_dflash_31b_extra.py test/registered/spec/dflash/test_muse_glimmer_dflash_assistant_gsm8k.py test/registered/e2e/speculative/test_dflash_domino.py test/registered/e2e/models/test_nvidia_nemotron_3_nano.py test/registered/e2e/models/test_glm53_flash_b200.py test/registered/cuda_graph/piecewise/test_pcg_with_speculative_decoding_dflash.py test/registered/kernels/ops/speculative/test_lilicorr_cuda.py test/registered/kernels/ops/speculative/test_dflash_domino.py test/registered/unit/spec/test_dflash_logits.py test/registered/unit/spec/test_dflash_overlap_hostsync.py test/registered/unit/spec/test_dflash_extra_buffer_lazy.py test/registered/unit/spec/test_dflash_domino.py test/registered/unit/spec/test_oot_dflash_hooks.py test/registered/unit/spec/test_draft_construction_isolation.py test/registered/unit/spec/test_dspark_target_hidden_projection.py test/registered/unit/models/test_glm5_next_dflash_capture.py |
|
|
/rerun-tests test/registered/core/test_basic_sanity_dflash.py test/registered/core/test_basic_sanity_dspark.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_mixed_chunk.py test/registered/spec/test_gemma4_dflash_31b_extra.py test/registered/spec/dflash/test_muse_glimmer_dflash_assistant_gsm8k.py test/registered/e2e/speculative/test_dflash_domino.py test/registered/e2e/models/test_nvidia_nemotron_3_nano.py test/registered/e2e/models/test_glm53_flash_b200.py test/registered/cuda_graph/piecewise/test_pcg_with_speculative_decoding_dflash.py test/registered/kernels/ops/speculative/test_lilicorr_cuda.py test/registered/kernels/ops/speculative/test_dflash_domino.py test/registered/unit/spec/test_dflash_logits.py test/registered/unit/spec/test_dflash_overlap_hostsync.py test/registered/unit/spec/test_dflash_extra_buffer_lazy.py test/registered/unit/spec/test_dflash_domino.py test/registered/unit/spec/test_oot_dflash_hooks.py test/registered/unit/spec/test_draft_construction_isolation.py test/registered/unit/spec/test_dspark_target_hidden_projection.py test/registered/unit/models/test_glm5_next_dflash_capture.py |
|
Results for 🚀 🚀 🚀 🚀 🚀 |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
/rerun-tests test/registered/unit/spec/test_dflash_logits.py |
|
Results for 🚀 |
…5.21 (#12) * [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384) * [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405) * Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096) * [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059) * [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425) * [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878) Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> * [AMD] Add tuned dsv4 shape (sgl-project#40996) * [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388) * [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612) * Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> * [HiCache] Make host reclamation independent of transfer order (sgl-project#40512) * [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922) * [diffusion] update code owner (sgl-project#41130) * [NPU] Update CANN version to 9.1.0 (sgl-project#40524) * [PD] Enable deferred decode-side KV release by default (sgl-project#41023) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [chore] surface the cookbook to users who pip install sglang (sgl-project#40866) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960) Signed-off-by: Ting Sun <suntcrick@gmail.com> * [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> * [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805) * Support unified memory decode host pools (sgl-project#39478) Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> * [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858) Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> * [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155) * [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109) * [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079) * [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081) * [Refactor] Share the MoE output all-reduce between models (sgl-project#41097) * [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082) * [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083) * [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084) * [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120) * [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041) * MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> * MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559) Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090) * [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162) * [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046) * [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049) * [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929) Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179) * Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442) Co-authored-by: Caio Rocha <caiorocha@microsft.com> * [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960) Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> * [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> * [Score API] Setwise Scoring Support (sgl-project#38965) * [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091) Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988) * [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907) * [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * Support XQA backend for SpecDec verify (sgl-project#32269) * [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793) Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> * [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> * [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * [feat] add a system one compatible /v1/systemone route (sgl-project#41208) * fix(openai): reject request-supplied chat_template by default (sgl-project#28135) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159) * [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201) Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> * [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456) * [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092) * [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619) * [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193) * [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194) * [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195) * [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196) * [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197) * [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198) * [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191) * [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199) * [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200) * [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224) * [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417) Co-authored-by: mickqian <mickqian@users.noreply.github.com> * [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486) * Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132) * [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064) Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: jacky.cheng <yichiche@amd.com> * fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248) Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> * Fix multimodal feature offload races (sgl-project#40621) Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986) * [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246) Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> * [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215) * [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216) * [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379) * [KVCache] Support lmcache unified radix cache (sgl-project#38652) Signed-off-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389) Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> * dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280) * [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660) Co-authored-by: BBuf <1182563586@qq.com> * [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284) * [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286) * [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285) * [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490) * [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291) Co-authored-by: DarkSharpness <2040703891@qq.com> * [diffusion] model: support Anima Base v1.0 (sgl-project#41011) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261) * [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297) * [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505) * [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150) Co-authored-by: jacky.cheng <yichiche@amd.com> * [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Restore logical kernel groups and test organization (sgl-project#41243) * [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188) * [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223) Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> * [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556) Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252) * [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253) * [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254) * [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255) * [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256) * [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257) * [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321) * [AMD] Fix jit broken on rocm env (sgl-project#41356) * Make sliding-window caching and speculative batch padding extensible (sgl-project#41325) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> * [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377) * dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019) Co-authored-by: x <x> * [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [mem_cache] Never free the protected prefix on request release (sgl-project#41312) * [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345) * [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416) Co-authored-by: aimicahchen <aimicahchen@tencent.com> * [kv-shard 3/4] Enable Control Plane B (sgl-project#39964) Co-authored-by: Zhangheng <hzh0425@apache.org> * [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166) * [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305) Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> * [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281) * [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192) * [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684) * [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990) * [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418) Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> * [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272) Signed-off-by: rockdu <kangrdu@gmail.com> * [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> * [bench] Take each request's prompt length from the server (sgl-project#39889) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> * [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266) * Add MiniMax arch fallback to auto parser resolution (sgl-project#40930) * [HiCache] Add the page-unified KV load-back JIT kernel (sgl-project#39726) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> * [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458) * [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: kpham-sgl <khoa.pham@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Port chat_parsing core (sgl-project#40477) Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> * [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417) * [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418) * [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419) * [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420) * [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421) * [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422) * [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423) * [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424) * [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425) * [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426) * [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427) * [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428) * [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429) * [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430) * [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431) * [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432) * [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433) * [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434) * [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435) * [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436) * [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437) * [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438) * [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439) * [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440) * [Refactor] Build the boundary into any stage with one construction (sgl-project#41441) * [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442) * [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443) * [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075) * [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814) * [diffusion] feat: support multiple task types for pipelines (sgl-project#38762) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [CI] fix CI regression on xeon (sgl-project#41002) * [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378) * [AMD] Register Triton data movement tests in PR CI (sgl-project#41137) * [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> * [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> * [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664) Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899) Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Rust server unify datapath for mm and generate requests (sgl-project#39679) * feat(npu): Support returning indexer top-k results (sgl-project#39060) * Fix chat template cache key order (sgl-project#41517) * docs: add prefill context parallelism guide and design draft (sgl-project#39354) * [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220) Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> * [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164) * [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165) * [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539) * [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943) * [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> * [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161) Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> * [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133) * [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144) * [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339) * [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446) * [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643) * [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235) * [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987) * [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597) * dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> * [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588) Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> * [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469) * [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196) Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118) Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> * [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590) Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> * [diffusion] model: support flux 3 action robot policies (sgl-project#41066) * [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627) Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> * [npu]support NPU 910C L2 memcache offload (sgl-project#41527) * [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600) * Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Refactor] Group communicator fusion and CP adapters (sgl-project#41547) * [Refactor] Centralize decoder output access (sgl-project#41548) * [Refactor] Carry residual state across stage boundaries (sgl-project#41549) * [Refactor] Capture auxiliary states at residual reads (sgl-project#41550) * [Refactor] Select reduction fusion at the consumer (sgl-project#41551) * [Refactor] Construct independent decoder stage boundaries (sgl-project#41552) * [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553) * [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554) * [Refactor] Rename the module to layer_boundary (sgl-project#41555) * [Refactor] Group layer boundary unit tests (sgl-project#41556) * [Refactor] Document layer boundary contracts and integration (sgl-project#41557) * [Rust] Extract a transport-neutral frontend core (sgl-project#39385) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618) * [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804) Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528) * [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059) * [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> * [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802) Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21 The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in the post-port commits #9 (hybrid sliding-window through the kv-cache plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory reads attention_backends()). Upstream's structure wins and the seam is re-applied on it: - server_args.py is now a thin record over arg_groups/: the --kv-cache-dtype choices are hoisted into arg_groups/choices.py as KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a plugin backend move to arg_groups/model_hook.py. - ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's resolution hooks (arg_groups/resolution_hooks.py) are the official slot, and turbo-attn's plugin now pairs the flags there. - Plugin dtype reads go through the resolved model bag (get_model()), not the raw ServerArgs record, which v0.5.21 leaves as operator input. - HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it holds and stands in as the per-layer cost, so upstream's new draft-SWA and unified-pool formulas apply unchanged; layer ids come from kvc.layer_info, as upstream's own SWA pool takes them. - Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21 a PP draft may load its own embedding (pp_draft_embedding). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> Signed-off-by: Ting Sun <suntcrick@gmail.com> Signed-off-by: chunxiaozheng <1179548172@qq.com> Signed-off-by: rockdu <kangrdu@gmail.com> Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com> Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com> Co-authored-by: Alan Kao <akao@amd.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca> Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: 黄孝君 <dingfangsu23@gmail.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Ting SUN <suntcrick@gmail.com> Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu> Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> Co-authored-by: ChangLiu0709 <cliu1004@amd.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Thomas Wang <thomawan@amd.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: Chunan Zeng <zcnrex@gmail.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com> Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com> Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com> Co-authored-by: Caio Rocha <caiorocha@microsft.com> Co-authored-by: metamergebot <metamergebot@gmail.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com> Co-authored-by: jacky.cheng <yichiche@amd.com> Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com> Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> Co-authored-by: Richard Wang <wangrichard08@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: Xinyi Song <xinyis10@illinois.edu> Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: ashwini rathi <ashwini.rathi@intel.com> Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com> Co-authored-by: Even Zhou <even.y.zhou@outlook.com> Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: cctry <csycfl@gmail.com> Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> Co-authored-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai> Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com> Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: WMC <tnwilly@gmail.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com> Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com> Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> Co-authored-by: Ankith Averineni <saverine@amd.com> Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com> Co-authored-by: Yuan Luo <yuan.luo@hotmail.com> Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com> Co-authored-by: CuzMi <simon.weijie@gmail.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com> Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com> Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com> Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com> Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com> Co-authored-by: flb_ <floatlibai@gmail.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> Co-authored-by: Peng Wu <peng@thinkingmachines.ai> Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com> Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com> Co-authored-by: karverma-amd <karan.verma@amd.com> Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: triple-mu <gpu@163.com> Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com> Co-authored-by: Hrithvik Alex <halex623@gmail.com> Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com> Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> Co-authored-by: sglang-bot <sglangbot@gmail.com> Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> Co-authored-by: Jialin Ouyang <jialino@meta.com> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> Co-authored-by: James <445169590@qq.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com> Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com> Co-authored-by: Jiacong Fang <zldrobit@126.com>
* [feature] add per-item candidate token scoring and calibration (sgl-project#40826) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384) * [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405) * Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096) * [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059) * [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425) * [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878) Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> * [AMD] Add tuned dsv4 shape (sgl-project#40996) * [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388) * [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612) * Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> * [HiCache] Make host reclamation independent of transfer order (sgl-project#40512) * [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922) * [diffusion] update code owner (sgl-project#41130) * [NPU] Update CANN version to 9.1.0 (sgl-project#40524) * [PD] Enable deferred decode-side KV release by default (sgl-project#41023) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [chore] surface the cookbook to users who pip install sglang (sgl-project#40866) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960) Signed-off-by: Ting Sun <suntcrick@gmail.com> * [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> * [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805) * Support unified memory decode host pools (sgl-project#39478) Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> * [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858) Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> * [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155) * [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109) * [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079) * [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081) * [Refactor] Share the MoE output all-reduce between models (sgl-project#41097) * [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082) * [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083) * [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084) * [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120) * [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041) * MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> * MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559) Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090) * [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162) * [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046) * [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049) * [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929) Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> * Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179) * Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442) Co-authored-by: Caio Rocha <caiorocha@microsft.com> * [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960) Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> * [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> * [Score API] Setwise Scoring Support (sgl-project#38965) * [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091) Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988) * [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907) * [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * Support XQA backend for SpecDec verify (sgl-project#32269) * [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793) Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> * [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061) Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> * [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * [feat] add a system one compatible /v1/systemone route (sgl-project#41208) * fix(openai): reject request-supplied chat_template by default (sgl-project#28135) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159) * [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201) Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> * [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456) * [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092) * [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619) * [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193) * [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194) * [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195) * [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196) * [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197) * [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198) * [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191) * [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199) * [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200) * [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224) * [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417) Co-authored-by: mickqian <mickqian@users.noreply.github.com> * [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486) * Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132) * [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064) Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: jacky.cheng <yichiche@amd.com> * fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131) Signed-off-by: Ting Sun <suntcrick@gmail.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> * [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248) Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> * Fix multimodal feature offload races (sgl-project#40621) Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986) * [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246) Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> * [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215) * [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216) * [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379) * [KVCache] Support lmcache unified radix cache (sgl-project#38652) Signed-off-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389) Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> * dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280) * [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660) Co-authored-by: BBuf <1182563586@qq.com> * [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284) * [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286) * [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285) * [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490) * [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291) Co-authored-by: DarkSharpness <2040703891@qq.com> * [diffusion] model: support Anima Base v1.0 (sgl-project#41011) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261) * [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297) * [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505) * [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150) Co-authored-by: jacky.cheng <yichiche@amd.com> * [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Refactor] Restore logical kernel groups and test organization (sgl-project#41243) * [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> * [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188) * [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223) Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> * [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556) Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252) * [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253) * [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254) * [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255) * [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256) * [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257) * [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321) * [AMD] Fix jit broken on rocm env (sgl-project#41356) * Make sliding-window caching and speculative batch padding extensible (sgl-project#41325) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> * [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328) Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> * [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377) * dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019) Co-authored-by: x <x> * [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [mem_cache] Never free the protected prefix on request release (sgl-project#41312) * [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345) * [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416) Co-authored-by: aimicahchen <aimicahchen@tencent.com> * [kv-shard 3/4] Enable Control Plane B (sgl-project#39964) Co-authored-by: Zhangheng <hzh0425@apache.org> * [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166) * [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305) Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> * [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281) * [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192) * [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684) * [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990) * [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418) Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> * [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272) Signed-off-by: rockdu <kangrdu@gmail.com> * [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> * [bench] Take each request's prompt length from the server (sgl-project#39889) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> * [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266) * Add MiniMax arch fallback to auto parser resolution (sgl-project#40930) * [HiCache] Add the page-unified KV load-back JIT kernel (sgl-project#39726) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> * [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458) * [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: kpham-sgl <khoa.pham@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * Port chat_parsing core (sgl-project#40477) Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> * [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417) * [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418) * [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419) * [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420) * [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421) * [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422) * [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423) * [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424) * [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425) * [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426) * [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427) * [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428) * [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429) * [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430) * [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431) * [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432) * [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433) * [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434) * [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435) * [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436) * [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437) * [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438) * [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439) * [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440) * [Refactor] Build the boundary into any stage with one construction (sgl-project#41441) * [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442) * [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443) * [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075) * [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814) * [diffusion] feat: support multiple task types for pipelines (sgl-project#38762) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [CI] fix CI regression on xeon (sgl-project#41002) * [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378) * [AMD] Register Triton data movement tests in PR CI (sgl-project#41137) * [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> * [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387) Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> * [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664) Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899) Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Rust server unify datapath for mm and generate requests (sgl-project#39679) * feat(npu): Support returning indexer top-k results (sgl-project#39060) * Fix chat template cache key order (sgl-project#41517) * docs: add prefill context parallelism guide and design draft (sgl-project#39354) * [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220) Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> * [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164) * [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165) * [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539) * [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943) * [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226) Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> * [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161) Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> * [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133) * [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144) * [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339) * [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446) * [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643) * [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235) * [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987) * [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597) * dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> * [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588) Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> * [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469) * [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196) Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> * [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118) Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> * [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590) Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> * [diffusion] model: support flux 3 action robot policies (sgl-project#41066) * [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627) Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> * [npu]support NPU 910C L2 memcache offload (sgl-project#41527) * [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600) * Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608) Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> * [Refactor] Group communicator fusion and CP adapters (sgl-project#41547) * [Refactor] Centralize decoder output access (sgl-project#41548) * [Refactor] Carry residual state across stage boundaries (sgl-project#41549) * [Refactor] Capture auxiliary states at residual reads (sgl-project#41550) * [Refactor] Select reduction fusion at the consumer (sgl-project#41551) * [Refactor] Construct independent decoder stage boundaries (sgl-project#41552) * [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553) * [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554) * [Refactor] Rename the module to layer_boundary (sgl-project#41555) * [Refactor] Group layer boundary unit tests (sgl-project#41556) * [Refactor] Document layer boundary contracts and integration (sgl-project#41557) * [Rust] Extract a transport-neutral frontend core (sgl-project#39385) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618) * [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804) Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> * [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528) * [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> * [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059) * [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> * [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802) Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: hnyls2002 <lsyincs@gmail.com> --------- Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal> Signed-off-by: Ting Sun <suntcrick@gmail.com> Signed-off-by: chunxiaozheng <1179548172@qq.com> Signed-off-by: rockdu <kangrdu@gmail.com> Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Xiaoyu Zhang <1182563586@qq.com> Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com> Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com> Co-authored-by: Alan Kao <akao@amd.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca> Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai> Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: 黄孝君 <dingfangsu23@gmail.com> Co-authored-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Ting SUN <suntcrick@gmail.com> Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu> Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com> Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com> Co-authored-by: yhzhuang <yhzhuang@fb.com> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com> Co-authored-by: Cheng Wan <cheng.wan@radixark.ai> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: hanminglu <hanminglu@fb.com> Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com> Co-authored-by: Khoa Pham <khoa.pham@radixark.ai> Co-authored-by: ChangLiu0709 <cliu1004@amd.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Thomas Wang <thomawan@amd.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: Chunan Zeng <zcnrex@gmail.com> Co-authored-by: mikevin920 <mikevin920@yahoo.com> Co-authored-by: Kevin Mi <mikevin920@gmail.com> Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com> Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com> Co-authored-by: Kevin Mi <kevin.mi@radixark.ai> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com> Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com> Co-authored-by: fengtinglei <fengtinglei@bytedance.com> Co-authored-by: Liangsheng Yin <lsyincs@gmail.com> Co-authored-by: Yuwei An <ayw.sirius19@gmail.com> Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com> Co-authored-by: Caio Rocha <caiorocha@microsft.com> Co-authored-by: metamergebot <metamergebot@gmail.com> Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: metamergebot <metamergebot@users.noreply.github.com> Co-authored-by: Zhiqiang Xie <zqx@meta.com> Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com> Co-authored-by: jacky.cheng <yichiche@amd.com> Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com> Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com> Co-authored-by: Shiyan Deng <dsy842974287@meta.com> Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com> Co-authored-by: Richard Wang <wangrichard08@gmail.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: Xinyi Song <xinyis10@illinois.edu> Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: ashwini rathi <ashwini.rathi@intel.com> Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com> Co-authored-by: Even Zhou <even.y.zhou@outlook.com> Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com> Co-authored-by: cctry <csycfl@gmail.com> Co-authored-by: cctry <cctry@fb.com> Co-authored-by: Yongji Wu <yongji@meta.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com> Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai> Co-authored-by: chunxiaozheng <1179548172@qq.com> Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai> Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com> Co-authored-by: DarkSharpness <2040703891@qq.com> Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com> Co-authored-by: WMC <tnwilly@gmail.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Kan Wu <kan.wu@radixark.ai> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com> Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com> Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com> Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com> Co-authored-by: jasonjk-park <jasonjk@fb.com> Co-authored-by: Ankith Averineni <saverine@amd.com> Co-authored-by: aimicahchen <aimicahchen@tencent.com> Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com> Co-authored-by: Zhangheng <hzh0425@apache.org> Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com> Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com> Co-authored-by: Yuan Luo <yuan.luo@hotmail.com> Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com> Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com> Co-authored-by: CuzMi <simon.weijie@gmail.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com> Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com> Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com> Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com> Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com> Co-authored-by: quitenode <quitenode@users.noreply.github.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Amrutha M <amrutha.m@intel.com> Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com> Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com> Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com> Co-authored-by: Ma Mingfei <mingfei.ma@intel.com> Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com> Co-authored-by: agent <agent@local> Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com> Co-authored-by: flb_ <floatlibai@gmail.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com> Co-authored-by: Peng Wu <peng@thinkingmachines.ai> Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com> Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com> Co-authored-by: karverma-amd <karan.verma@amd.com> Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com> Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com> Co-authored-by: triple-mu <gpu@163.com> Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com> Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com> Co-authored-by: Hrithvik Alex <halex623@gmail.com> Co-authored-by: Hrithvik Alex <hrithvik@baseten.co> Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com> Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru> Co-authored-by: sglang-bot <sglangbot@gmail.com> Co-authored-by: zelong huang <yzhu@ubiquant.com> Co-authored-by: zelong518 <zelonghuang05@gmail.com> Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com> Co-authored-by: Jialin Ouyang <jialino@meta.com> Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com> Co-authored-by: James <445169590@qq.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com> Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com> Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com> Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com> Co-authored-by: Jiacong Fang <zldrobit@126.com>
Motivation
DFlash is trained on per-position marginals rather than on the joint block distribution, so its
drafted tokens are individually plausible yet jointly incoherent — slot 3 may be a fine
continuation of the prompt while contradicting the token DFlash itself chose for slot 2.
LiLiCorr keeps the top-
kcandidates DFlash already produces at each block position and scorestransitions between adjacent candidates with a small two-layer transformer, then commits a path
through the resulting lattice greedily, left to right, instead of taking the per-slot argmax. One
network pass produces every vector, the pairwise scores are a single batched matmul, and only the
greedy walk stays sequential. Verify is untouched, so outputs remain distributionally identical to
the target model's — only acceptance length and throughput change.
This PR is the serving half. It ships no checkpoints and no training code; the companion PR
above trains the drafters and exports them in the format this loader reads.
Performance
How these were produced. Six drafters for a Qwen3-8B target, all trained in NVIDIA
ModelOpt on one matched contract — the same corpus, schedule and block geometry for every arm, so
no row carries a training advantage. Training data is NVIDIA's Nemotron Post-Training Dataset v2
with the multilingual split excluded, generated from the target with thinking disabled;
6 epochs; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7;
8 nodes × 8 H100, global batch size 64 (one sequence per device, no gradient accumulation).
All six were then exported and served through the code in this PR on a single H100 80GB,
tp_size 1, at concurrency 1, greedyT=0,--attention-backend fa3,draft_length 15, mean oftwo replicates, with the whole node held exclusive per benchmark so nothing else shared it.
Speedup is output tokens/s against an autoregressive baseline measured in the same allocation,
because a denominator borrowed from another node carries that node's clock into every ratio.
Three of the five alternatives —
DFlash,DSpark,Domino— already live in this repo.Cells are
acceptance length / speedup-vs-AR; ★ fastest, ☆ second fastest:Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight
benchmarks. Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a
164-prompt slice, where DFlash2 is ahead by 0.5%.
DFlashis the deliberately head-free control. Against it, the reranked drafter is worth+9.1% to +14.6% output tokens/s — the number to read if the question is "what does this buy over
what I already run":
Every head also clears that control by +7.60% to +21.67% on acceptance, on every arm and every
benchmark, which is the check that a head actually loaded rather than silently falling back.
Acceptance length is bit-reproducible under greedy decoding; throughput is not. The replicate
spread on acceptance was 0.00% on every benchmark; on throughput it is about 0.2% within one
allocation, and about 1% across allocations.
ⓘ Reproducing these needs
--attention-backend fa3. Backends that publishseq_lens_cpueverystep pay a device-to-host sync per block, worth about 8% here. That is a property of DFlash decoding
rather than of LiLiCorr, and the new docs subsection explains which backends are affected and why.
Backends also differ in attention numerics, so acceptance measured on different backends does not
belong in one table.
Modifications
No new
SpeculativeAlgorithm, no worker subclass, no registration, no CLI flag, and no newentry in any choices list. Two environment variables gate the opt-in sampled commit described
below; nothing else reads the environment. LiLiCorr rides
--speculative-algorithm DFLASHand isselected by the checkpoint declaring
architectures: ["LiLiCorrDraftModel"], exactly asDFlash2DraftModelselects the candidate selector. If you go looking for the registration, that iswhy there isn't one.
10 files, +3,076 / −1. The single deleted line is one
ifin the draft loop that gained abranch; everything else is an addition.
Only four files already existed — +164 / −1 between them:
__init__assignments, the fold branch in_maybe_build_draft_sampler, the guarded anchor publication inside_append_target_hidden_to_draft_kv_by_loc, and a fourth arm on the existing three-way draft-loop branchDFlashDraftModeldeclareslilicorr: Optional[nn.Module] = Nonebeside the existingcandidate_selector, and stores therms_norm_epsit already computesFor scale,
selector(DFlash2 — the closest precedent, also a head on the same draft model, alsocheckpoint-selected) appears on 43 lines of that worker;
lilicorrappears on 30.New, self-contained:
srt/models/lilicorr.py(632) — the head,LiLiCorrDraftModel(DFlashDraftModel),EntryClass, and the two weight-coverage checkskernels/ops/speculative/lilicorr.py(520) — three Triton kernels; see "Why the custom kernels" belowsrt/speculative/lilicorr_utils.py(656) — head geometry, the candidate lattice, the eager draft seam and the CUDA-graph-folded draft sampler, in one module besidedflash_utils.pytest/registered/unit/spec/test_lilicorr.py(805) — 46 CPU teststest/registered/kernels/ops/speculative/test_lilicorr_cuda.py(216) — 18 GPU tests pinning the Triton kernels against their torch referencestest/registered/e2e/speculative/test_dflash_lilicorr.py(84) — the E2E server test, registereddisabled=until a drafter is publishedThe head runs inside the draft CUDA graph, and that is the whole optimization
The head is a long tail of small kernels, so what it costs is host dispatch, not FLOPs. It
therefore rides the seam DFlash already has for exactly this:
_maybe_build_draft_samplerrunsbefore
init_cuda_graphs, and the sampler is registered ondraft_model_runner.capture_tail_hooksthroughmake_draft_sampler_capture_hook— the samemechanism DFlash2's
_SelectorDraftSamplerand DSpark'sDsparkDraftSampleruse, andLiLiCorrDraftSamplerhas the same shape as the latter: static buffers sized from the capturebuckets, host-side staging of the per-row sampling params before the replay, an in-graph philox
draw, and the drafted tokens written into a static
outthe worker reads after the replay.Nothing in the head is capture-hostile: static shapes, no collectives at
tp=1, no host syncs, afixed decode trip count, and every parameter-derived buffer built by
materialize_inference_buffersbefore capture.There is no
torch.compileanywhere in this PR. An earlier revision compiled the capturedbody, on the theory that the graph removes the launches' host cost while leaving their number and
memory traffic alone. Measured end to end on one H100 — Qwen3-8B, gsm8k, greedy,
fa3,tp_size 1,both arms as sequential cells in one allocation so the node term cancels — it is a null:
Four pairs straddling zero, against a 1.12% spread for the same arm measured in two different
allocations. The fold, by contrast, is load-bearing: in the same allocation and on the same tree,
forcing the head out of the graph with
SGLANG_DFLASH_EAGER_DRAFT_SAMPLER=1costs −8.4%.So there is no dynamo surface here at all — no compile prewarm, no recompile-limit handling, no
torch._inductor.configmutation. The one operational note that remains is that the head should runon the folded path: an eager fallback exists for the steps the graph cannot serve and is correct but
costs a large fraction of throughput, and because acceptance is identical either way that shows up
only in tokens/s.
build_lilicorr_draft_samplerwarns with a reason when it declines to fold, andthe worker warns once if a decode step lands on the eager head.
Two capture gaps remain, both shared with the existing DFlash heads. Prefill and extend run eager —
the eager seam is the third arm of the same
elifchain that already carries the selector's andplain DFlash's. And
tp>1declines to fold, because the candidate top-k needs a packed all-gatherinside the graph.
Convolutional drafters are supported, and the support is a refusal
The leading column of the table above is LiLiCorr composed with the grouped dynamic convolutions
that wrap each DFlash sublayer. It serves through this PR, and that is worth stating precisely,
because it is not a feature this PR adds.
The convolution belongs to the DFlash backbone, not to the reranker:
LiLiCorrDraftModelextends
DFlashDraftModeland therefore inherits the__init__that builds it fromconv_kernel_sizeandconv_group_sizeindflash_config. Both the plain and the convolutionaldrafter already load and run through the same path, with no new flag and no new code.
What this PR contributes is the refusal. Both geometry keys default to
0, so a checkpoint whoseconfig lost them builds no convolution modules at all — and the backbone loader then drops every
convolution tensor it cannot resolve. That is correct behaviour for rotary caches and the worst
possible behaviour here: all 20 tensors of a five-layer drafter disappear in silence, the draft
serves as its convolution-free parent, and the only symptom is an acceptance length that is lower
than it should be and entirely believable.
check_conv_weight_coveragerefuses any checkpoint whoseconvolution tensors and built modules disagree, in either direction. Five of the CPU tests cover it:
the matched case, the convolution-free case, both silent-drop directions, and a partial checkpoint.
Sampling the committed path (opt-in)
At
T > 0the target samples, but the head committed the per-slot argmax and verify was handed apoint mass, so each slot was paid
p(argmax)rather than the overlap of the two distributions.LILICORR_SAMPLING=1draws each slot fromsoftmax(scores / T)over that slot's candidates, at therequest's own temperature, and publishes the proposal it drew from into the same buffers the DFlash2
selector already fills — so the accept path is untouched and verify stays exact. Greedy requests
sharing the batch take the argmax and report a one-hot proposal. On Qwen3-8B it is worth about 6%
acceptance and 6% end-to-end at
T = 1, at under 1% of the per-block wall.Off by default, and with it unset the captured body is
selectitself rather than somethingequivalent to it, so
T = 0output is unchanged bit for bit.SGLANG_LILICORR_REQUIRE_SAMPLING=1refuses to start if the mode was asked for but is not enabled.Both are read once at import: the sampler is replayed from a captured draft graph, so a per-call
read would be a host-side branch inside a replay — happy to move them to
ServerArgsfieldsresolved at init if you would rather.
The sampled commit is additionally gated on the same supported-device condition as the selector's,
on both the folded and the eager path, because it publishes into an accept kernel that does not run
everywhere. Where it is unsupported the arm falls back to the argmax commit and the existing verify.
Why the custom kernels
kernels/ops/speculative/lilicorr.pyis the single largest new file after the head itself, so it isworth saying up front what it buys. All three kernels dispatch to a value-identical torch
implementation off CUDA, which is also what makes the head exercisable in a CPU unit test.
lilicorr_topk_lse— exact per-row top-k and the full-vocab log-partition from one pass over[n, V]. The head scores log-probs normalized over the whole vocabulary, so it needs the partitionas well as the top-k. Composing that from existing ops (
torch.topkfollowed by alogsumexpepilogue) reads the vocabulary logits three times and materializes two more
[n, V]temporaries,and at a 150k vocabulary that read is the single largest cost in the block. Fusing them makes it one
read and no temporaries. The tile pre-selection is exact, not approximate — the argument is in the
function docstring and worth reading before trusting the kernel: if a top-k element's tile were not
among the k tiles with the largest maxima, k other tiles would each hold an element at least as
large, contradicting its rank.
lilicorr_greedy_path— the whole left-to-right commit in one kernel. The torch form issuesroughly three kernels per slot over
[bs, k]tensors, so at 15 slots the decode is dozens offew-microsecond launches and is pure launch overhead rather than arithmetic. Same recurrence, and
tl.argmaxbreaks ties toward the lower index asTensor.argmaxdoes, so the committed path isidentical.
lilicorr_sample_path— that same walk with a sampled commit, emitting the per-slot proposal theverify needs to accept it by rejection sampling. A row with
greedy_maskset takestl.argmaxonthe same fp32 node values as the greedy kernel, so one captured graph serves greedy and sampling
batches and the greedy rows walk a bit-identical path.
Correctness
Verify is unmodified, so this change cannot alter the distribution of emitted tokens. Every
drafted token is still checked against the target model exactly as before; a worse drafter would show
up as lower acceptance, never as different output. That is the primary correctness argument, and it
is why the section above measures acceptance rather than task scores.
Unit tests: 46 CPU tests (
test_lilicorr.py), one worker-stub test intest_dflash_logits.py, and 18 GPU tests (test_lilicorr_cuda.py) pinning all three Tritonkernels against their torch reference implementations on device — vocabulary sizes straddling the
kernel's tile boundary in both directions, bf16 logits as production uses, the narrow-vocabulary
fallback, and the greedy commit at k = 1…16 including tie-breaking. Verified on an H100: 75/75 pass
across the three files. The registry validates —
validate_all_suites()is clean over 2,487registered tests — and every pinned pre-commit hook passes over the PR range.
E2E:
test/registered/e2e/speculative/test_dflash_lilicorr.pyserves the head through--speculative-algorithm DFLASHwith only the draft path changed, and asserts GSM8K accuracy plustwo acceptance floors. It is registered
disabled=because no drafter is published yet, but it hasbeen run against one on an H100:
1 passed in 177.19s, GSM8K 0.950, accept length 5.0392 against ahead-free DFLASH control reading 4.1694 on the same eval.
Checklist
register_cpu_ci), 18 GPU kernel-parity tests (register_cuda_ci), and an E2E server test on the DFlash fixture (register_cuda_ci,disabled=pending a published drafter)Design notes
The choices most likely to prompt a question, each also commented at the site.
srt/speculative/lilicorr_utils.py, not as a field onDFlashDraftConfig. The shared DFLASH config therefore carries no knowledge of this head, at thecost of one extra
dflash_configread at model build, anddflash_utils.pystays at zerochanges.
[batch*heads, S, S]rather than broadcast. Thebroadcast form is mathematically identical and saves a copy, but a stride-0 mask measured
−1.75pp. Please measure before changing it back.
F.rms_normrather thansglang.srt.layers.layernorm.RMSNorm, and a hand-rolled attentionmodule carrying
nn.MultiheadAttention's parameter layout. Same math and the same state_dictkeys in both cases. The custom-op norm is not capturable here, and the module's Python control flow
and
need_weightsbookkeeping are what make it uncapturable.candidate_topkmust be a power of two, refused at config parse. The surplus direction is live,not hypothetical: one projection is an
Identitywhen the head is as wide as the draft, so aconfig omitting the head width would otherwise silently drop it and serve a different architecture
than the one trained. The power-of-two constraint comes from the tiled top-k holding its selected
tiles in a single Triton lane group; supporting other widths would mean serving the head on the
slow reference path, so it is refused at load instead of silently deoptimized.
at different namespaces: the convolution tensors are backbone parameters named
layers.*.{attention,mlp}_conv.*, none of which is underlilicorr., so the head check cannotsee them. See "Convolutional drafters" above for why a missing check is dangerous rather than
merely untidy.
DFlashDraftInputV2. Carrying it there sothat filter and merge keep it aligned with the batch is the tidier design, but it measured
−4.67% acceptance at concurrency 32, and the drift it prevents is not detectable: reordering
cannot happen at concurrency 1 and happens constantly at 32, where acceptance reads 7.5573 and
7.5669. The measurement is recorded at the site in
set_anchor.Known gaps
ModelOpt PR linked at the top adds the training and checkpoint export, including the optional
convolutions.
tp>1serves the head eagerly. Folding it needs the candidate top-k's packed all-gatherinside the draft graph.
CI States
Latest PR Test (Base): ❌ Run #36355573025
Latest PR Test (Extra): ❌ Run #36355572926
Latest PR Test (AMD ROCm 10): ❌ Run #36355573123