Skip to content

[Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts - #37462

Merged
kpham-sgl merged 38 commits into
sgl-project:mainfrom
mrusanovsky:lilicorr-support
Sep 27, 2026
Merged

kpham-sgl merged 38 commits into
sgl-project:mainfrom
mrusanovsky:lilicorr-support

Conversation

@mrusanovsky

@mrusanovsky mrusanovsky commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

DFlash is trained on per-position marginals rather than on the joint block distribution, so its
drafted tokens are individually plausible yet jointly incoherent — slot 3 may be a fine
continuation of the prompt while contradicting the token DFlash itself chose for slot 2.

LiLiCorr keeps the top-k candidates DFlash already produces at each block position and scores
transitions between adjacent candidates with a small two-layer transformer, then commits a path
through the resulting lattice greedily, left to right, instead of taking the per-slot argmax. One
network pass produces every vector, the pairwise scores are a single batched matmul, and only the
greedy walk stays sequential. Verify is untouched, so outputs remain distributionally identical to
the target model's — only acceptance length and throughput change.

This PR is the serving half. It ships no checkpoints and no training code; the companion PR
above trains the drafters and exports them in the format this loader reads.

Performance

How these were produced. Six drafters for a Qwen3-8B target, all trained in NVIDIA
ModelOpt on one matched contract — the same corpus, schedule and block geometry for every arm, so
no row carries a training advantage. Training data is NVIDIA's Nemotron Post-Training Dataset v2
with the multilingual split excluded, generated from the target with thinking disabled;
6 epochs; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7;
8 nodes × 8 H100, global batch size 64 (one sequence per device, no gradient accumulation).

All six were then exported and served through the code in this PR on a single H100 80GB,
tp_size 1, at concurrency 1, greedy T=0, --attention-backend fa3, draft_length 15, mean of
two replicates, with the whole node held exclusive per benchmark so nothing else shared it.
Speedup is output tokens/s against an autoregressive baseline measured in the same allocation,
because a denominator borrowed from another node carries that node's clock into every ratio.

Three of the five alternatives — DFlash, DSpark, Domino — already live in this repo.

Cells are acceptance length / speedup-vs-AR; ★ fastest, ☆ second fastest:

benchmark LiLiCorr+conv LiLiCorr DSpark DFlash2 Domino DFlash
gsm8k ★ 7.715 / 5.26x ☆ 7.557 / 5.22x 7.375 / 4.86x 7.252 / 5.06x 7.225 / 4.87x 6.341 / 4.59x
math500 ★ 9.241 / 6.54x ☆ 9.064 / 6.52x 9.012 / 6.15x 8.999 / 6.49x 8.976 / 6.25x 7.909 / 5.88x
aime25 ★ 8.285 / 6.03x ☆ 8.156 / 6.03x 8.043 / 5.61x 7.967 / 5.91x 8.066 / 5.77x 7.126 / 5.44x
humaneval ★ 7.393 / 4.01x 7.077 / 3.93x 7.163 / 3.72x ☆ 7.081 / 3.95x 6.864 / 3.73x 6.156 / 3.68x
mbpp_sanitized ★ 5.999 / 4.18x ☆ 5.849 / 4.13x 5.888 / 3.95x 5.685 / 4.05x 5.679 / 3.91x 5.027 / 3.70x
livecodebench ★ 7.975 / 5.40x ☆ 7.754 / 5.33x 7.775 / 5.10x 7.601 / 5.26x 7.553 / 5.04x 6.808 / 4.88x
alpaca_eval ☆ 3.697 / 2.69x ★ 3.656 / 2.70x 3.588 / 2.52x 3.467 / 2.58x 3.627 / 2.59x 3.222 / 2.46x
mtbench ★ 4.014 / 2.94x ☆ 3.939 / 2.93x 3.957 / 2.78x 3.748 / 2.80x 3.948 / 2.84x 3.478 / 2.67x

Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight
benchmarks. Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a
164-prompt slice, where DFlash2 is ahead by 0.5%.

DFlash is the deliberately head-free control. Against it, the reranked drafter is worth
+9.1% to +14.6% output tokens/s — the number to read if the question is "what does this buy over
what I already run":

benchmark gsm8k mbpp_sanitized math500 livecodebench mtbench humaneval
throughput vs head-free DFlash +14.6% +12.9% +11.1% +10.5% +9.9% +9.1%

Every head also clears that control by +7.60% to +21.67% on acceptance, on every arm and every
benchmark, which is the check that a head actually loaded rather than silently falling back.

Acceptance length is bit-reproducible under greedy decoding; throughput is not. The replicate
spread on acceptance was 0.00% on every benchmark; on throughput it is about 0.2% within one
allocation, and about 1% across allocations.

ⓘ Reproducing these needs --attention-backend fa3. Backends that publish seq_lens_cpu every
step pay a device-to-host sync per block, worth about 8% here. That is a property of DFlash decoding
rather than of LiLiCorr, and the new docs subsection explains which backends are affected and why.
Backends also differ in attention numerics, so acceptance measured on different backends does not
belong in one table.

Modifications

No new SpeculativeAlgorithm, no worker subclass, no registration, no CLI flag, and no new
entry in any choices list. Two environment variables gate the opt-in sampled commit described
below; nothing else reads the environment. LiLiCorr rides --speculative-algorithm DFLASH and is
selected by the checkpoint declaring architectures: ["LiLiCorrDraftModel"], exactly as
DFlash2DraftModel selects the candidate selector. If you go looking for the registration, that is
why there isn't one.

10 files, +3,076 / −1. The single deleted line is one if in the draft loop that gained a
branch; everything else is an addition.

Only four files already existed — +164 / −1 between them:

file added what
srt/speculative/dflash_worker_v2.py +95 / −1 five dispatch seams: two __init__ assignments, the fold branch in _maybe_build_draft_sampler, the guarded anchor publication inside _append_target_hidden_to_draft_kv_by_loc, and a fourth arm on the existing three-way draft-loop branch
srt/models/dflash.py +2 DFlashDraftModel declares lilicorr: Optional[nn.Module] = None beside the existing candidate_selector, and stores the rms_norm_eps it already computes
docs/.../speculative_decoding.mdx +34 two subsections under "DFlash Decoding"
test/registered/unit/spec/test_dflash_logits.py +33 one worker-stub test: the sampled arm's device gate

For scale, selector (DFlash2 — the closest precedent, also a head on the same draft model, also
checkpoint-selected) appears on 43 lines of that worker; lilicorr appears on 30.

New, self-contained:

  • srt/models/lilicorr.py (632) — the head, LiLiCorrDraftModel(DFlashDraftModel), EntryClass, and the two weight-coverage checks
  • kernels/ops/speculative/lilicorr.py (520) — three Triton kernels; see "Why the custom kernels" below
  • srt/speculative/lilicorr_utils.py (656) — head geometry, the candidate lattice, the eager draft seam and the CUDA-graph-folded draft sampler, in one module beside dflash_utils.py
  • test/registered/unit/spec/test_lilicorr.py (805) — 46 CPU tests
  • test/registered/kernels/ops/speculative/test_lilicorr_cuda.py (216) — 18 GPU tests pinning the Triton kernels against their torch references
  • test/registered/e2e/speculative/test_dflash_lilicorr.py (84) — the E2E server test, registered disabled= until a drafter is published

The head runs inside the draft CUDA graph, and that is the whole optimization

The head is a long tail of small kernels, so what it costs is host dispatch, not FLOPs. It
therefore rides the seam DFlash already has for exactly this: _maybe_build_draft_sampler runs
before init_cuda_graphs, and the sampler is registered on
draft_model_runner.capture_tail_hooks through make_draft_sampler_capture_hook — the same
mechanism DFlash2's _SelectorDraftSampler and DSpark's DsparkDraftSampler use, and
LiLiCorrDraftSampler has the same shape as the latter: static buffers sized from the capture
buckets, host-side staging of the per-row sampling params before the replay, an in-graph philox
draw, and the drafted tokens written into a static out the worker reads after the replay.

Nothing in the head is capture-hostile: static shapes, no collectives at tp=1, no host syncs, a
fixed decode trip count, and every parameter-derived buffer built by
materialize_inference_buffers before capture.

There is no torch.compile anywhere in this PR. An earlier revision compiled the captured
body, on the theory that the graph removes the launches' host cost while leaving their number and
memory traffic alone. Measured end to end on one H100 — Qwen3-8B, gsm8k, greedy, fa3, tp_size 1,
both arms as sequential cells in one allocation so the node term cancels — it is a null:

requests concurrency compiled no compile Δ acceptance Δ output tok/s
1319 1 6.9951 / 731.5 6.9947 / 735.6 −0.005% +0.56%
1319 1 6.9951 / 739.7 6.9947 / 735.0 −0.005% −0.62%
1319 32 7.0085 / 10179.5 7.0022 / 10131.9 −0.090% −0.47%
320 32 7.0650 / 9410.8 7.0576 / 9489.0 −0.105% +0.83%

Four pairs straddling zero, against a 1.12% spread for the same arm measured in two different
allocations. The fold, by contrast, is load-bearing: in the same allocation and on the same tree,
forcing the head out of the graph with SGLANG_DFLASH_EAGER_DRAFT_SAMPLER=1 costs −8.4%.

So there is no dynamo surface here at all — no compile prewarm, no recompile-limit handling, no
torch._inductor.config mutation. The one operational note that remains is that the head should run
on the folded path: an eager fallback exists for the steps the graph cannot serve and is correct but
costs a large fraction of throughput, and because acceptance is identical either way that shows up
only in tokens/s. build_lilicorr_draft_sampler warns with a reason when it declines to fold, and
the worker warns once if a decode step lands on the eager head.

Two capture gaps remain, both shared with the existing DFlash heads. Prefill and extend run eager —
the eager seam is the third arm of the same elif chain that already carries the selector's and
plain DFlash's. And tp>1 declines to fold, because the candidate top-k needs a packed all-gather
inside the graph.

Convolutional drafters are supported, and the support is a refusal

The leading column of the table above is LiLiCorr composed with the grouped dynamic convolutions
that wrap each DFlash sublayer. It serves through this PR, and that is worth stating precisely,
because it is not a feature this PR adds.

The convolution belongs to the DFlash backbone, not to the reranker: LiLiCorrDraftModel
extends DFlashDraftModel and therefore inherits the __init__ that builds it from
conv_kernel_size and conv_group_size in dflash_config. Both the plain and the convolutional
drafter already load and run through the same path, with no new flag and no new code.

What this PR contributes is the refusal. Both geometry keys default to 0, so a checkpoint whose
config lost them builds no convolution modules at all — and the backbone loader then drops every
convolution tensor it cannot resolve. That is correct behaviour for rotary caches and the worst
possible behaviour here: all 20 tensors of a five-layer drafter disappear in silence, the draft
serves as its convolution-free parent, and the only symptom is an acceptance length that is lower
than it should be and entirely believable. check_conv_weight_coverage refuses any checkpoint whose
convolution tensors and built modules disagree, in either direction. Five of the CPU tests cover it:
the matched case, the convolution-free case, both silent-drop directions, and a partial checkpoint.

Sampling the committed path (opt-in)

At T > 0 the target samples, but the head committed the per-slot argmax and verify was handed a
point mass, so each slot was paid p(argmax) rather than the overlap of the two distributions.
LILICORR_SAMPLING=1 draws each slot from softmax(scores / T) over that slot's candidates, at the
request's own temperature, and publishes the proposal it drew from into the same buffers the DFlash2
selector already fills — so the accept path is untouched and verify stays exact. Greedy requests
sharing the batch take the argmax and report a one-hot proposal. On Qwen3-8B it is worth about 6%
acceptance and 6% end-to-end at T = 1, at under 1% of the per-block wall.

Off by default, and with it unset the captured body is select itself rather than something
equivalent to it, so T = 0 output is unchanged bit for bit.
SGLANG_LILICORR_REQUIRE_SAMPLING=1 refuses to start if the mode was asked for but is not enabled.
Both are read once at import: the sampler is replayed from a captured draft graph, so a per-call
read would be a host-side branch inside a replay — happy to move them to ServerArgs fields
resolved at init if you would rather.

The sampled commit is additionally gated on the same supported-device condition as the selector's,
on both the folded and the eager path, because it publishes into an accept kernel that does not run
everywhere. Where it is unsupported the arm falls back to the argmax commit and the existing verify.

Why the custom kernels

kernels/ops/speculative/lilicorr.py is the single largest new file after the head itself, so it is
worth saying up front what it buys. All three kernels dispatch to a value-identical torch
implementation off CUDA, which is also what makes the head exercisable in a CPU unit test.

lilicorr_topk_lse — exact per-row top-k and the full-vocab log-partition from one pass over
[n, V]. The head scores log-probs normalized over the whole vocabulary, so it needs the partition
as well as the top-k. Composing that from existing ops (torch.topk followed by a logsumexp
epilogue) reads the vocabulary logits three times and materializes two more [n, V] temporaries,
and at a 150k vocabulary that read is the single largest cost in the block. Fusing them makes it one
read and no temporaries. The tile pre-selection is exact, not approximate — the argument is in the
function docstring and worth reading before trusting the kernel: if a top-k element's tile were not
among the k tiles with the largest maxima, k other tiles would each hold an element at least as
large, contradicting its rank.

lilicorr_greedy_path — the whole left-to-right commit in one kernel. The torch form issues
roughly three kernels per slot over [bs, k] tensors, so at 15 slots the decode is dozens of
few-microsecond launches and is pure launch overhead rather than arithmetic. Same recurrence, and
tl.argmax breaks ties toward the lower index as Tensor.argmax does, so the committed path is
identical.

lilicorr_sample_path — that same walk with a sampled commit, emitting the per-slot proposal the
verify needs to accept it by rejection sampling. A row with greedy_mask set takes tl.argmax on
the same fp32 node values as the greedy kernel, so one captured graph serves greedy and sampling
batches and the greedy rows walk a bit-identical path.

Correctness

Verify is unmodified, so this change cannot alter the distribution of emitted tokens. Every
drafted token is still checked against the target model exactly as before; a worse drafter would show
up as lower acceptance, never as different output. That is the primary correctness argument, and it
is why the section above measures acceptance rather than task scores.

Unit tests: 46 CPU tests (test_lilicorr.py), one worker-stub test in
test_dflash_logits.py, and 18 GPU tests (test_lilicorr_cuda.py) pinning all three Triton
kernels against their torch reference implementations on device — vocabulary sizes straddling the
kernel's tile boundary in both directions, bf16 logits as production uses, the narrow-vocabulary
fallback, and the greedy commit at k = 1…16 including tie-breaking. Verified on an H100: 75/75 pass
across the three files. The registry validates — validate_all_suites() is clean over 2,487
registered tests — and every pinned pre-commit hook passes over the PR range.

E2E: test/registered/e2e/speculative/test_dflash_lilicorr.py serves the head through
--speculative-algorithm DFLASH with only the draft path changed, and asserts GSM8K accuracy plus
two acceptance floors. It is registered disabled= because no drafter is published yet, but it has
been run against one on an H100: 1 passed in 177.19s, GSM8K 0.950, accept length 5.0392 against a
head-free DFLASH control reading 4.1694 on the same eval.

Checklist

  • Format your code according to Format code with pre-commit — every pinned hook passes over the PR range
  • Add unit tests according to Run and add unit tests — 46 CPU tests (register_cpu_ci), 18 GPU kernel-parity tests (register_cuda_ci), and an E2E server test on the DFlash fixture (register_cuda_ci, disabled= pending a published drafter)
  • Update documentation according to Write documentations — a subsection under "DFlash Decoding"
  • Provide accuracy and speed benchmark results — see Performance
  • Follow the SGLang code style guidance

Design notes

The choices most likely to prompt a question, each also commented at the site.

  1. The head geometry is parsed in srt/speculative/lilicorr_utils.py, not as a field on
    DFlashDraftConfig.
    The shared DFLASH config therefore carries no knowledge of this head, at the
    cost of one extra dflash_config read at model build, and dflash_utils.py stays at zero
    changes.
  2. The lattice attention bias is materialized [batch*heads, S, S] rather than broadcast. The
    broadcast form is mathematically identical and saves a copy, but a stride-0 mask measured
    −1.75pp. Please measure before changing it back.
  3. F.rms_norm rather than sglang.srt.layers.layernorm.RMSNorm, and a hand-rolled attention
    module carrying nn.MultiheadAttention's parameter layout. Same math and the same state_dict
    keys in both cases. The custom-op norm is not capturable here, and the module's Python control flow
    and need_weights bookkeeping are what make it uncapturable.
  4. The weight loader raises in both directions — missing head tensors and surplus ones — and
    candidate_topk must be a power of two, refused at config parse. The surplus direction is live,
    not hypothetical: one projection is an Identity when the head is as wide as the draft, so a
    config omitting the head width would otherwise silently drop it and serve a different architecture
    than the one trained. The power-of-two constraint comes from the tiled top-k holding its selected
    tiles in a single Triton lane group; supporting other widths would mean serving the head on the
    slow reference path, so it is refused at load instead of silently deoptimized.
  5. The convolution coverage check is separate from the head coverage check, because the two look
    at different namespaces: the convolution tensors are backbone parameters named
    layers.*.{attention,mlp}_conv.*, none of which is under lilicorr., so the head check cannot
    see them. See "Convolutional drafters" above for why a missing check is dangerous rather than
    merely untidy.
  6. The anchor is worker-owned rather than carried on DFlashDraftInputV2. Carrying it there so
    that filter and merge keep it aligned with the batch is the tidier design, but it measured
    −4.67% acceptance at concurrency 32, and the drift it prevents is not detectable: reordering
    cannot happen at concurrency 1 and happens constantly at 32, where acceptance reads 7.5573 and
    7.5669. The measurement is recorded at the site in set_anchor.

Known gaps

  • Checkpoints are not yet published. The docs example carries a placeholder path. The companion
    ModelOpt PR linked at the top adds the training and checkpoint export, including the optional
    convolutions.
  • tp>1 serves the head eagerly. Folding it needs the candidate top-k's packed all-gather
    inside the draft graph.

CI States

Latest PR Test (Base): ❌ Run #36355573025
Latest PR Test (Extra): ❌ Run #36355572926
Latest PR Test (AMD ROCm 10): ❌ Run #36355573123

@kpham-sgl kpham-sgl left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for the contribution! Some very high level comments before I review further

  1. Please point your agents to these rules we have in .claude/rules/. Some to point out
  • Avoid extensive AI comments and docstrings
  • Avoid using getattr/hasattr
  • Avoid writing unnecessary unit tests
  1. Can you add an E2E test with a small target for this new algorithm? Make sure to use existing spec decoding test kit?

@kpham-sgl

Copy link
Copy Markdown
Collaborator

Can you also resolve conflicts @mrusanovsky? Thanks!

@nvpohanh nvpohanh left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Codex] Five inline review findings are attached.

Comment thread python/sglang/srt/speculative/lilicorr_components/lilicorr_candidates.py Outdated
Comment thread python/sglang/srt/speculative/lilicorr_components/lilicorr_candidates.py Outdated
Comment thread python/sglang/srt/speculative/lilicorr_components/lilicorr_draft_sampler.py Outdated
Comment thread python/sglang/srt/speculative/lilicorr_components/lilicorr_config.py Outdated
Comment thread python/sglang/kernels/ops/speculative/lilicorr.py
@mrusanovsky

Copy link
Copy Markdown
Contributor Author

Thank you for the contribution! Some very high level comments before I review further

  1. Please point your agents to these rules we have in .claude/rules/. Some to point out
  • Avoid extensive AI comments and docstrings
  • Avoid using getattr/hasattr
  • Avoid writing unnecessary unit tests
  1. Can you add an E2E test with a small target for this new algorithm? Make sure to use existing spec decoding test kit?

Reading the rest of .claude/rules/ turned up several issues that were addressed.

getattr/hasattr : all removed:

  • worker.lilicorr probed the draft model. Now reads DFlashDraftModel.lilicorr, declared None
    next to candidate_selector.
  • LiLiCorrHead re-derived rms_norm_eps with a default. The base __init__ already computes it.
  • resolve_vocab_shard probed for shard_indices. Now narrows on VocabParallelEmbedding.
  • torch._dynamo / torch._inductor knobs assigned directly.

Remaining hasattr is in tearDownClass, required by write-sglang-test.

Comments and docstrings : docstring lines in non-test source down from 476 to 270. Removed
multi-line rationale on private helpers. Kept module docstrings, tensor shape contracts, and
provenance of measured constants.

Tests : 48 CPU to 37, 16 GPU to 12. Back to 40 and 13 after @nvpohanh's findings needed guards.
Deleted a tautology against if topk & (topk-1): raise, a case asserting a property of the test's
own fake, and one made unfalsifiable by an API change.

Also fixed : LiLiCorrConfig was a frozen @dataclass, now msgspec.Struct.
candidate_token_ids renamed to candidate_tokens per speculative-naming Rule 5.
build_lilicorr_draft_sampler and target_input_embeddings no longer take the whole worker.

E2E test : test/registered/spec/dflash/test_dflash_lilicorr.py, same fixture shape as
test_dflash.py with GSM8KMixin. Launch args are plain DFLASH with only the draft path changed.

Run on an H100: 1 passed in 177.19s, GSM8K 0.950, accept length 5.0392. The head-free control
reads 4.1694 on the same eval, so the floor is 4.6. est_time=300 measured at 177 s.

That accept length is the mixin's figure over 200 short 5-shot completions, not the block-weighted
acceptance in the PR body. Different quantities, noted at the site.

Registered disabled= since no drafter is published, same as test_spec_eagle_cpu.py. Draft path
is the literal <unpublished>, so enabling it without an id fails immediately.

SpecDecodingMixin is mixed in too, with both floors measured against the head-free control on the
same test: acceptance 8.192 vs 7.314, speed 883.8 vs 788.0 tok/s on an H100. Floors are 7.7 and 600.
The speed one is deliberately loose since throughput isn't reproducible.

One gap left: no small-target drafter exists, all of ours target Qwen3-8B, so it runs on
1-gpu-large like test_uno.py.

In addition, I resolved conflicts. Rebased on main. Only dflash_worker_v2.py conflicted, against the NPU DFlash2 change on the same __init__ lines. Both kept.

@kpham-sgl kpham-sgl self-assigned this Sep 8, 2026
@nvpohanh

nvpohanh commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator

@mrusanovsky could you fix lint issue?
https://github.com/sgl-project/sglang/actions/runs/34270114821/job/102209278752?pr=37462

@nvpohanh nvpohanh left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Codex] One inline review finding is attached.

# is correct: the argmax draft is still verified losslessly, just target-only.
# Checked after the selector branch because the two heads are mutually
# exclusive: a selector worker returns above and never reads `lilicorr`.
if self.lilicorr is not None and _LILICORR_SAMPLING_ENABLED:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Codex] Severity: functional | Confidence: High

Issue

This return treats LiLiCorr sampling as supported on NPU. The selector path explicitly disables sampling there, but LiLiCorr later sets _selector_sample and _selector_sampling_accept calls accept_sampling, which unconditionally launches chain_speculative_sampling_triton. NPU tensors cannot run that kernel, so LILICORR_SAMPLING=1 with a non-greedy request fails during verification instead of taking the existing fallback.

Fix

Gate LiLiCorr sampling on the same supported-device condition as the selector, or add a supported NPU accept implementation. Add a focused NPU-path test for the fallback.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. The selector gates both publish sites on _selector_sampling_enabled; LiLiCorr's didn't, and the accept branch keys on _selector_sample is not None alone, so a non-greedy request on NPU reached chain_speculative_sampling_triton.

Gating only the publish would be worse than the crash: the draft would still sample while verify treated the drawn token as a point mass, making acceptance p(x) instead of min(1, p(x)/q(x)) with a residual from the wrong q : not distribution-preserving. So sampling_enabled is threaded into the sampler as done for the selector, deciding which body compiles. Unsupported devices take the argmax commit and the existing verify, with the warning that path already emits.

Fallback test added in test_dflash_logits.py, stub-based so it runs on CPU; the dispatch test now also asserts the flag reaches the builder. On CUDA the effective flag is unchanged : re-ran on an H100 and τ is identical to before: 6.8158 sampled, 6.3944 argmax, T=0 gate 7.557260716552241 on 51999 blocks.

h-guo18 added a commit to NVIDIA/Model-Optimizer that referenced this pull request Sep 9, 2026
### What does this PR do?

Type of change: new feature

Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a
new `projector_type` on the
existing `dflash` mode — plus three DFlash-wide improvements that apply
to every variant, and an
optional composition with DFlash2's grouped convolutions.

A DFlash drafter is trained on per-position marginals rather than on the
joint block distribution,
so its drafted tokens are individually plausible yet jointly incoherent.
LiLiCorr keeps the top-`k`
candidates the backbone already produces at each block position, scores
transitions between adjacent
candidates with a small two-layer transformer, and commits a path
through the lattice greedily.
Serving is unchanged in kind: verify still checks every drafted token
against the target, so the
emitted distribution is untouched and only acceptance length moves.

- Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel
Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530)
(arXiv:2608.20530)
- Blog: https://research.nvidia.com/labs/nemotron/lilicorr/
- **Companion PR — serving support:**
[sgl-project/sglang#37462](sgl-project/sglang#37462)

This PR is the **training** half. It trains the drafters and exports
them; the companion PR above is
what serves the resulting checkpoints, and is what the comparison table
below was measured through.

**What is in the commits**

| | |
| --- | --- |
| LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`,
conversion routing, config fields, export |
| Three DFlash-wide features | fp32 master weights for the draft, draft
activation checkpointing, and a DDP hang fix — all default-off or
behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and
`dflash2` alike |
| Optional grouped convolutions | composes LiLiCorr with DFlash2's
`DFlashGroupedConv`; see the dependency note below |
| Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` |
| CPU unit tests, CHANGELOG, one launcher example | |

**⚠️ The convolutions depend on the DFlash2 branch, and cannot run until
it merges.**

`modeling_lilicorr.py` imports `DFlashGroupedConv` from
`modeling_dflash2`, which today exists only
on `haoguo/dflash2-support`. The class is **imported rather than copied
on purpose** — it is the only
way the two variants cannot drift apart arithmetically — but the
consequence is that the
convolutional recipe cannot run against `main` as it stands.

So the import is **deferred into `_install_sublayer_convs`** rather than
taken at module scope.
Everything else in this PR, including the plain LiLiCorr reranker, has
no DFlash2 dependency at all
and works on `main` today; an eager import would have made the whole
plugin unimportable for the sake
of one optional feature. Requesting the convolutions without DFlash2
present raises an `ImportError`
naming the two config keys to remove, rather than failing at import
time.

**This PR carries two of @h-guo18's commits, with authorship and
sign-off preserved.** Both are
independent of DFlash2 itself and both are needed here:

- `1419d47e`, the no-op sublayer seam. Without it
`DFlashDecoderLayer.forward` never calls the
wrappers the convolutions install onto, so the modules would be built,
counted and exported while
  computing nothing. It is arithmetically an identity on its own.
- `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a
top-level `rope_theta` and a
`rope_parameters` dict; the real base lives in the dict while the class
default (10,000 for Qwen3)
stays visible as the flat attribute. Reading the flat field first builds
a draft whose RoPE base is
100× off a Qwen3-8B target's, which trains and exports without
complaint. Both the training-side
enforcement and the exporter's `_get_rope_theta` are affected on `main`
today.

Both are @h-guo18's work and belong to their branches; they are carried
here only so that this PR
stands on its own. **If those branches land first, this PR can be
rebased onto them and the two
commits dropped**, and they can equally be split out now if that is
easier to review.

The same applies to `dflash_fp32_master_weights`, which is also in
flight on
`haoguo/dflash-fp32-master-weights`. The field name is shared
deliberately so that there is only
ever one knob rather than two spellings of it, and both versions default
to off. Whichever lands
first, this PR can be rebased onto it.

### Usage

Train with the shipped recipe:

```python
from modelopt.recipe import load_recipe

config = load_recipe("general/speculative_decoding/lilicorr.yaml")
# Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified),
# DFlash decay objective at gamma 7.0, fp32 master weights for the draft.
```

Or convert directly:

```python
import modelopt.torch.speculative as mtsp

config = {
    "dflash_block_size": 16,
    "dflash_loss_objective": "decay",
    "dflash_loss_decay_factor": 7.0,
    "dflash_fp32_master_weights": True,
    "dflash_lilicorr_w_ce": 0.25,
    "dflash_lilicorr_w_margin": 0.0,
    "dflash_lilicorr_w_pen": 0.25,
    "dflash_architecture_config": {
        "num_hidden_layers": 5,
        "projector_type": "lilicorr",
        "lilicorr_candidate_topk": 8,
        # Optional, and all-or-nothing: adding these two keys wraps every draft
        # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant.
        # "conv_kernel_size": 2,
        # "conv_group_size": 16,
    },
}
mtsp.convert(model, [("dflash", config)])
```

### Results

Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one
matched contract** — the
same corpus, schedule and block geometry for every arm, so no row
carries a training advantage.
Training data is NVIDIA's
[Nemotron Post-Training Dataset
v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)
with the multilingual split excluded, generated from the target with
**thinking disabled**;
**6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash
decay objective at gamma 7;
**8 nodes × 8 H100, global batch size 64** (one sequence per device, no
gradient accumulation).

All six were then exported and served through SGLang on a **single H100
80GB**, `tp_size 1`, at
concurrency 1, greedy, `fa3`, mean of two replicates, with the whole
node held exclusive per
benchmark. Speedup is output tokens/s against an autoregressive baseline
measured in the same
allocation.

Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second
fastest**:

| benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino |
DFlash |
|---|---|---|---|---|---|---|
| gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 /
5.06x | 7.225 / 4.87x | 6.341 / 4.59x |
| math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 /
6.49x | 8.976 / 6.25x | 7.909 / 5.88x |
| aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 /
5.91x | 8.066 / 5.77x | 7.126 / 5.44x |
| humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081
/ 3.95x | 6.864 / 3.73x | 6.156 / 3.68x |
| mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x |
5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x |
| livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x |
7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x |
| alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x |
3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x |
| mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 /
2.80x | 3.948 / 2.84x | 3.478 / 2.67x |

**Against every other approach in the table, LiLiCorr with convolutions
is the fastest on all eight
benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the
exception is humaneval, a
164-prompt slice, where DFlash2 is ahead by 0.5%.

`DFlash` is the deliberately head-free control; every head clears it by
+7.60% to +21.67% on
acceptance, which is the check that a head actually loaded. Reproducing
the `LiLiCorr+conv` column
additionally needs the DFlash2 variant.

Acceptance length is bit-reproducible under greedy decoding and its
replicate spread here was 0.00%
on every benchmark; throughput has a ~0.2% floor.

### What `dflash_fp32_master_weights` does, and what it is worth

Today the draft is cast to the frozen base model's dtype — bf16 — before
the optimizer is built.
AdamW then allocates its moments with `zeros_like(p)`, so the
**optimizer state becomes bf16 too**.
That is the problem: bf16 has too few mantissa bits to represent the
small updates Adam's second
moment accumulates, so those updates round away and the effective step
size decays on its own,
independently of the learning-rate schedule.

The flag is standard mixed precision instead: the draft's master weights
stay in fp32 while the
matmuls run in bf16. It requires a bf16 autocast around the forward,
which HF `Trainer` supplies
under `TrainingArguments.bf16`. Paths that do not go through the Trainer
— evaluation,
`pseudo_speculative_generate`, a plain `convert()` and forward —
currently need the caller to
supply it, and no shipped recipe exercises those (`estimate_ar: false`,
`do_eval: false`). Making
the draft supply its own autocast is a follow-up, held back from here on
review because it touches
every DFlash variant and wants e2e coverage of the existing recipes.

Compute speed is unchanged. The cost is memory, about 12 bytes per
parameter for the weight plus
Adam's two moments instead of 6, plus a doubled gradient all-reduce
under DDP, since fp32
parameters mean fp32 gradients. Under FSDP2 that second cost is what
`MixedPrecisionPolicy(reduce_dtype=...)` exists to control.

It is worth **7 to 14 percent of acceptance length**, measured at the
end of training on gsm8k, and
it helps every projector type:

| arm | bf16 | fp32 | Δ acceptance length |
| --- | ---: | ---: | ---: |
| LiLiCorr | 6.8670 | 7.5573 | **+10.05%** |
| DFlash2 | 6.7396 | 7.2518 | **+7.60%** |
| Domino | 6.5854 | 7.2252 | **+9.71%** |
| DSpark | 6.4621 | 7.3752 | **+14.13%** |
| DFlash | 5.9030 | 6.3412 | **+7.42%** |

Every arm in the comparison table above was trained with it on, and
**both shipped recipes set it
`true`**, so the documented path gets it.

It defaults to **off**, so no existing DFlash, Domino or DSpark run
changes behaviour. Both shipped
LiLiCorr recipes set it `true`, which is the arithmetic their numbers
were trained with. Flipping
the default is a reasonable follow-up once the autocast above is in.

The draft is drawn in fp32 and, under this flag, kept there; an
unpromoted run rounds the same draw
to the base model's dtype. So the bf16 and fp32 rows of the table above
start from the same
initialization at the precision each trains in, rather than from two
different draws. A unit test
pins that.

The flag also survives a resume. `modify()` runs under `from_pretrained`
with the base model still
on meta and cannot place the draft at all, so `restore_draft_precision`
re-applies the dtype, the
device and the rotary buffer once the weights are loaded and before the
Trainer builds the
optimizer — the last point that can still decide the Adam moment dtype.
It also reloads the draft's
tensors at the dtype they were saved in, since checkpoints store the
draft in fp32 while the base is
bf16 and `dtype="auto"` gives every tensor one dtype.

@h-guo18 has the same field in flight on
`haoguo/dflash-fp32-master-weights`, plus an
HF-format-resume fix this PR does not have. The name is shared
deliberately so there is only ever
one knob; whichever lands first, the other should be dropped rather than
merged.

### Testing

- **257 CPU unit tests pass** across `tests/unit/torch/speculative/`,
including the existing DFlash,
Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr
specifically: conversion
routing, head geometry, the required-field validation, the three-term
objective and its absolute
weights, gradient reach into both the head and the drafter body, and the
export contract.
- Both recipes load and validate through `modelopt.recipe.load_recipe`.
- The three DFlash-wide changes are covered behaviourally: the fp32 flag
is checked on the
optimizer's moment dtypes rather than only on parameters, since the
moments are the point of the
change, and on the initialization described above; activation
checkpointing is asserted to leave
draft gradients bit-identical with the flag on and off; and the rotary
buffer is asserted present
after `modify()` on a real device while still deferred on meta, which is
the case the laziness
  existed for.
- The resume path has its own test: after a `save_pretrained` /
`from_pretrained` round trip,
`restore_draft_precision` is asserted to return the draft to fp32 with
its stored weights intact
and its Adam moments in fp32. Without it the draft comes back in the
base dtype with the flag
  still set, which is the failure it exists to prevent.
- `TestDFlashLazyRotaryEmb` was updated rather than left passing: it
asserted the rotary buffer does
*not* exist after convert, and the DDP fix deliberately changes that on
non-meta devices. The
  replacement pins the refined invariant in both directions.
- The published checkpoints were trained with this arithmetic, verified
rather than assumed: a
fingerprint over draft initialisation, loss and gradients is compared
against the pre-review tree
for both `dflash` and `lilicorr`. Loss and gradients are **bitwise
identical**. Initialisation
moves, by less than bf16 resolution, and that is the single-dtype change
described above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — every addition is opt-in. The
new `projector_type` is
selected only by config, `dflash_fp32_master_weights` defaults to off,
and the
activation-checkpointing and DDP fixes preserve behaviour. No existing
default changes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in
`CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted
from
https://github.com/sgl-project/SpecForge/...` headers for the DFlash
backbone and loss they derive
from (Apache-2.0), matching the attribution already on `hf_dflash.py` in
this repo. The two
commits described above are @h-guo18's, cherry-picked with authorship
and sign-off preserved.
- Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the
updated rotary test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once opened.

### Additional Information

The convolutional recipe is the memory worst case: at an 8B target,
combined with fp32 master
weights, it may need `training.gradient_checkpointing: true` to fit on
80 GiB, and it fits without at
4B. Checkpointing is mathematically neutral — same objective, same data
order, same resulting model —
but it trades step time for memory, so a run using it is not
step-time-comparable with one that does
not. The recipe header says so.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added LiLiCorr speculative decoding with candidate-lattice reranking,
configurable objectives, metrics, export support, and optional grouped
convolutions.
* Added FP32 master-weight support with improved mixed-precision
behavior and gradient checkpointing.
* Added LiLiCorr training recipes and a Qwen3-8B launcher configuration.
* **Bug Fixes**
* Improved rotary-embedding configuration handling and corrected DFlash
distributed-training hangs.
* Added validation for invalid LiLiCorr configurations and improved
exported reranking metadata.
* **Documentation**
* Expanded guidance for FP32 master weights, training workflows, and
LiLiCorr configuration.
* **Tests**
* Expanded coverage across training, evaluation, generation, export, and
checkpoint workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
DFlash is trained on per-position marginals rather than on the joint block
distribution, so its drafted tokens are individually plausible yet jointly
incoherent. LiLiCorr keeps the top-k candidates DFlash already produces at
each block position and scores transitions between adjacent candidates with a
small two-layer transformer, then commits a path through the resulting lattice
greedily, left to right, instead of taking the per-slot argmax. Verify is
untouched, so outputs remain distributionally identical to the target model's.

This adds no SpeculativeAlgorithm, no worker subclass and no registration. It
rides --speculative-algorithm DFLASH and is selected by the checkpoint
declaring architectures: ["LiLiCorrDraftModel"], exactly as DFlash2DraftModel
selects the candidate selector. Of the 11 files touched, only two already
existed: dflash_worker_v2.py gains four dispatch seams (+39) and the
speculative-decoding docs gain a subsection (+19). Nothing is deleted or
modified anywhere.

Two Triton kernels carry most of the new non-model code, and both replace a
composition of existing ops that costs extra passes over the vocabulary:
lilicorr_topk_lse returns an exact per-row top-k and the full-vocabulary
log-partition from one read of [n, V], where topk plus a logsumexp epilogue
would read it three times and materialize two more [n, V] temporaries;
lilicorr_greedy_path folds the whole left-to-right commit into one launch
instead of roughly three per slot. Both fall back to a value-identical torch
implementation off CUDA.

Tests: 43 CPU tests, plus 16 GPU tests pinning both kernels against those
torch references on device (tile-boundary vocabulary sizes, bf16 logits, the
narrow-vocabulary fallback, and the greedy commit at k = 1..16 including
tie-breaking).

Paper: https://arxiv.org/abs/2608.20530
Blog: https://research.nvidia.com/labs/nemotron/lilicorr/
kpham-sgl and others added 4 commits September 27, 2026 01:31
Condense multi-line comments and docstrings per comment-style.md (no code
changes; AST-identical modulo docstrings). Drop the non-unit-stride walk
test and the full-vocab log_softmax test, and fold the two sampled-path
parity tests into one parametrized case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LILICORR_SAMPLING becomes SGLANG_ENABLE_LILICORR_SAMPLING and
SGLANG_LILICORR_REQUIRE_SAMPLING moves to an EnvBool; EnvBool already
rejects non-boolean values, so the local parser goes. Register
lilicorr_topk_lse and lilicorr_sample_path in the speculative kernel
inventory, and drop the hasattr guard in the E2E test teardown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Remove file/class docstrings, section banners, return-shape recaps and
comments restating the code; keep only non-obvious design constraints.
No code changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread python/sglang/srt/models/lilicorr.py Outdated
return vals.float(), ids.to(torch.int64), lse


def lilicorr_topk_lse(

@kpham-sgl kpham-sgl Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reuse DFlash2's compute_candidates (radix top-k + TP gather) and add only the lse, instead of a new top-k kernel and TP combine. It also keeps quantized lm_head support, which the eager path here breaks.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done for the structure: compute_candidates body is now candidate_topk() in models/dflash.py and LiLiCorr calls it, so there is one projection (through quant_method), one TP gather, and our TP combine is gone.

One difference: with with_partition=True the top-k and lse come from a single fused pass instead of radix top-k plus a separate logsumexp. The second pass re-reads the [N, V] logits, which dominates at batch: 143 vs 680 us per step at c=32 (56 vs 69 us at c=1; H100, bf16, V=151,936, CUDA graph), and about -1.5% output throughput end to end at c=32. The selector keeps _radix_topk.

…ader

Resolve SGLANG_ENABLE_LILICORR_SAMPLING and the REQUIRE guard only when a
LiLiCorr head is loaded, instead of at import on every DFLASH server. Move
the grouped-convolution coverage check from LiLiCorrDraftModel into
DFlashDraftModel.load_weights, since it validates backbone weights.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread python/sglang/srt/speculative/lilicorr_utils.py Outdated
Comment thread python/sglang/srt/speculative/lilicorr_utils.py Outdated
Comment thread python/sglang/srt/models/lilicorr.py Outdated
)

anchor_state = self._project_anchor(anchor_hidden, anchor_valid)
# Materialized rather than a stride-0 broadcast, which measured -1.75pp.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The -1.75pp was measured on the compiled body (see 56a1572); without compile you measured +0.73% for broadcast. Compile is gone, so pass self._attn_bias.unsqueeze(0) straight to SDPA and drop the per-step expand/reshape copy, or re-measure without compile if you want to keep it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, passing self._attn_bias.unsqueeze(0) to SDPA. Re-measured without compile, it's neutral: -0.004% / -0.04% tok/s at c=1, +0.46% / -0.19% at c=32 (two A/B runs on one node, the second with the arms swapped), acceptance identical.

Review, kpham-sgl: reuse DFlash2's compute_candidates (radix top-k + TP
gather) and add only the lse, rather than carrying a second top-k kernel and
a second TP combine. He also noted the eager path breaks a quantized
lm_head, which it did: lilicorr_candidates read lm_head.weight directly, and
a packed weight is [V, H/pack], so the matmul could not even be shaped.

Lifts compute_candidates' body into candidate_topk(), which gains an opt-in
full-vocabulary log-partition: a logsumexp over the projection it already
computes, combined across shards by a second logsumexp. The padded tail is
already -inf and logsumexp ignores it, so no extra mask.

Deletes lilicorr_topk_lse and its tiled scan/select kernels, the torch
reference, _combine_across_ranks, and the static logits buffer they needed.
The projection now goes through quant_method, so the worker screens LiLiCorr
on the selector's gate and a quantized head folds instead of dropping to an
eager path that could not run.

Also drops LiLiCorrRMSNorm for layers.layernorm.RMSNorm, per review: the
DFlash backbone already runs that module inside this same draft graph.
…e logits

logsumexp(logits.float()) upcasts the whole [N, V] logits tensor, so it pays a
second full-size allocation and read. Take the row max from the top-k, which
comes back sorted, and let only the reduction be fp32 via sum(dtype=).

Measured on an H100 at [480, 151936] bf16: 895 us -> 359 us, and 53 -> 40 us
at [15, 151936]. logsumexp on the native dtype is 426 us but carries 6e-2 of
error, which a log-prob cannot take; this form is 2.6e-4, under the bf16
logits' own ~4e-3.
@mrusanovsky

Copy link
Copy Markdown
Contributor Author

Addressed the latest review, replies inline.

One small addition, 17 lines in models/lilicorr.py only: LiLiCorr now serves --speculative-num-draft-tokens below the trained block, like base DFlash / DFlash2 / Domino already do. LiLiCorrDraftModel.set_block_size used to refuse a mismatch; it now has the head use a prefix of its trained slots (every relative offset a shorter lattice needs was trained). Nothing else changes, it is bit-identical at the trained size (same greedy acceptance fingerprint on gsm8k), and it serves correctly at 8 and 4 tokens.

Keeps DFlashDraftModel.load_weights identical to main; the check stays
scoped to LiLiCorr drafts as originally written.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kpham-sgl

Copy link
Copy Markdown
Collaborator

/rerun-tests test/registered/core/test_basic_sanity_dflash.py test/registered/core/test_basic_sanity_dspark.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_mixed_chunk.py test/registered/spec/test_gemma4_dflash_31b_extra.py test/registered/spec/dflash/test_muse_glimmer_dflash_assistant_gsm8k.py test/registered/e2e/speculative/test_dflash_domino.py test/registered/e2e/models/test_nvidia_nemotron_3_nano.py test/registered/e2e/models/test_glm53_flash_b200.py test/registered/cuda_graph/piecewise/test_pcg_with_speculative_decoding_dflash.py test/registered/kernels/ops/speculative/test_lilicorr_cuda.py test/registered/kernels/ops/speculative/test_dflash_domino.py test/registered/unit/spec/test_dflash_logits.py test/registered/unit/spec/test_dflash_overlap_hostsync.py test/registered/unit/spec/test_dflash_extra_buffer_lazy.py test/registered/unit/spec/test_dflash_domino.py test/registered/unit/spec/test_oot_dflash_hooks.py test/registered/unit/spec/test_draft_construction_isolation.py test/registered/unit/spec/test_dspark_target_hidden_projection.py test/registered/unit/models/test_glm5_next_dflash_capture.py

@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Rebase Required Before Re-run

A major update has landed on main. Your PR is diverged relative to required base commit 3fc7a669bfe3.

Re-run was not dispatched. What to do:

  • Rebase your branch onto the latest main and push again
  • Follow issue #21065 for context
  • CI-fix PRs may request the bypass-maintenance label to skip this check

@kpham-sgl

Copy link
Copy Markdown
Collaborator

/rerun-tests test/registered/core/test_basic_sanity_dflash.py test/registered/core/test_basic_sanity_dspark.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_mixed_chunk.py test/registered/spec/test_gemma4_dflash_31b_extra.py test/registered/spec/dflash/test_muse_glimmer_dflash_assistant_gsm8k.py test/registered/e2e/speculative/test_dflash_domino.py test/registered/e2e/models/test_nvidia_nemotron_3_nano.py test/registered/e2e/models/test_glm53_flash_b200.py test/registered/cuda_graph/piecewise/test_pcg_with_speculative_decoding_dflash.py test/registered/kernels/ops/speculative/test_lilicorr_cuda.py test/registered/kernels/ops/speculative/test_dflash_domino.py test/registered/unit/spec/test_dflash_logits.py test/registered/unit/spec/test_dflash_overlap_hostsync.py test/registered/unit/spec/test_dflash_extra_buffer_lazy.py test/registered/unit/spec/test_dflash_domino.py test/registered/unit/spec/test_oot_dflash_hooks.py test/registered/unit/spec/test_draft_construction_isolation.py test/registered/unit/spec/test_dspark_target_hidden_projection.py test/registered/unit/models/test_glm5_next_dflash_capture.py

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-tests test/registered/core/test_basic_sanity_dflash.py test/registered/core/test_basic_sanity_dspark.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_mixed_chunk.py test/registered/spec/test_gemma4_dflash_31b_extra.py test/registered/spec/dflash/test_muse_glimmer_dflash_assistant_gsm8k.py test/registered/e2e/speculative/test_dflash_domino.py test/registered/e2e/models/test_nvidia_nemotron_3_nano.py test/registered/e2e/models/test_glm53_flash_b200.py test/registered/cuda_graph/piecewise/test_pcg_with_speculative_decoding_dflash.py test/registered/kernels/ops/speculative/test_lilicorr_cuda.py test/registered/kernels/ops/speculative/test_dflash_domino.py test/registered/unit/spec/test_dflash_logits.py test/registered/unit/spec/test_dflash_overlap_hostsync.py test/registered/unit/spec/test_dflash_extra_buffer_lazy.py test/registered/unit/spec/test_dflash_domino.py test/registered/unit/spec/test_oot_dflash_hooks.py test/registered/unit/spec/test_draft_construction_isolation.py test/registered/unit/spec/test_dspark_target_hidden_projection.py test/registered/unit/models/test_glm5_next_dflash_capture.py:

🚀 1-gpu-5090 (4 tests): ✅ View workflow run

cd test/ && python3 registered/core/test_basic_sanity_dflash.py
cd test/ && python3 registered/spec/dflash/test_dflash.py
cd test/ && python3 registered/unit/spec/test_dflash_overlap_hostsync.py
cd test/ && python3 registered/unit/spec/test_dflash_extra_buffer_lazy.py

🚀 1-gpu-h100 (6 tests): ✅ View workflow run

cd test/ && python3 registered/core/test_basic_sanity_dspark.py
cd test/ && python3 registered/spec/test_spec_mixed_chunk.py
cd test/ && python3 registered/spec/dflash/test_muse_glimmer_dflash_assistant_gsm8k.py
cd test/ && python3 registered/cuda_graph/piecewise/test_pcg_with_speculative_decoding_dflash.py
cd test/ && python3 registered/kernels/ops/speculative/test_lilicorr_cuda.py
cd test/ && python3 registered/kernels/ops/speculative/test_dflash_domino.py

🚀 2-gpu-h100 (2 tests): ✅ View workflow run

cd test/ && python3 registered/spec/test_gemma4_dflash_31b_extra.py
cd test/ && python3 registered/e2e/speculative/test_dflash_domino.py

🚀 4-gpu-b200 (2 tests): ✅ View workflow run

cd test/ && python3 registered/e2e/models/test_nvidia_nemotron_3_nano.py
cd test/ && python3 registered/e2e/models/test_glm53_flash_b200.py

🚀 ubuntu-latest (6 tests): ❌ View workflow run

cd test/ && python3 registered/unit/spec/test_dflash_logits.py
cd test/ && python3 registered/unit/spec/test_dflash_domino.py
cd test/ && python3 registered/unit/spec/test_oot_dflash_hooks.py
cd test/ && python3 registered/unit/spec/test_draft_construction_isolation.py
cd test/ && python3 registered/unit/spec/test_dspark_target_hidden_projection.py
cd test/ && python3 registered/unit/models/test_glm5_next_dflash_capture.py

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kpham-sgl

Copy link
Copy Markdown
Collaborator

/rerun-tests test/registered/unit/spec/test_dflash_logits.py

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-tests test/registered/unit/spec/test_dflash_logits.py:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/spec/test_dflash_logits.py

@kpham-sgl
kpham-sgl merged commit 78eee88 into sgl-project:main Sep 27, 2026
93 of 110 checks passed
arbi-dev added a commit to arbicity/sglang-turbo that referenced this pull request Oct 7, 2026
…5.21 (#12)

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21

The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream
v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the
commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in
the post-port commits #9 (hybrid sliding-window through the kv-cache
plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory
reads attention_backends()).

Upstream's structure wins and the seam is re-applied on it:
- server_args.py is now a thin record over arg_groups/: the
  --kv-cache-dtype choices are hoisted into arg_groups/choices.py as
  KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from
  server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a
  plugin backend move to arg_groups/model_hook.py.
- ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's
  resolution hooks (arg_groups/resolution_hooks.py) are the official slot,
  and turbo-attn's plugin now pairs the flags there.
- Plugin dtype reads go through the resolved model bag (get_model()), not
  the raw ServerArgs record, which v0.5.21 leaves as operator input.
- HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it
  holds and stands in as the per-layer cost, so upstream's new draft-SWA
  and unified-pool formulas apply unchanged; layer ids come from
  kvc.layer_info, as upstream's own SWA pool takes them.
- Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21
  a PP draft may load its own embedding (pp_draft_embedding).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
sam123456598 added a commit to vessl-ai/sglang that referenced this pull request Oct 8, 2026
* [feature] add per-item candidate token scoring and calibration (sgl-project#40826)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation jit-kernel speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants