Skip to content

[Score API] Setwise scoring: CausalLM support (batched + --enable-mis) - #41188

Merged
Qiaolin-Yu merged 4 commits into
sgl-project:mainfrom
sundar24295s:suramach/setwise-causallm
Sep 26, 2026
Merged

Qiaolin-Yu merged 4 commits into
sgl-project:mainfrom
sundar24295s:suramach/setwise-causallm

Conversation

@sundar24295s

@sundar24295s sundar24295s commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Setwise Scoring: CausalLM Support

Follow-up to [Score API] Setwise Scoring Support (#38965) (merged). That PR
added setwise scoring for SequenceClassification models; this PR extends the same
token_indices_to_pool primitive to generative (CausalLM) models.

Summary

Extends setwise scoring to CausalLM models. A generative reranker can now
read label-token logprobs at every occurrence of a caller-chosen extraction
token
in each query + item sequence, instead of only at the last token — the
generative analogue of the SequenceClassification setwise path. A candidate set is
packed into one item with one extraction token per candidate; the LM head is read at
each of those positions in a single prefill, so scores comes back as one
[N_i x num_labels] matrix per item (nested per item), where num_labels = len(label_token_ids).

The unifying signal is unchanged: token_indices_to_pool still names the readout
positions. For SequenceClassification the pooler pools the classification head there;
for CausalLM the LM head is applied there and the requested label_token_ids
logprobs are gathered — one row per anchor.

Modes: both batched (default) and fused --enable-mis are supported,
matching the SequenceClassification setwise paths.

Motivation

Many rerankers are decoder-only models that express a candidate's score as the
probability the model assigns to specific label tokens (e.g. yes/no, a rating
digit) right after each candidate. Without setwise support such a caller must issue
one request per candidate, paying the shared query-prefix cost N times. This PR packs
the whole candidate block into one prompt with one anchor per candidate and returns
all N per-candidate label distributions from a single prefill.

API

Engine (in-process)

result = engine.score(
    query="",                        # optional shared prefix
    items=[prompt_with_anchors],     # one item = one candidate set
    label_token_ids=[yes_id, no_id], # required for CausalLM
    score_extraction_token_id=anchor_id,
    apply_softmax=False,
)
# result.scores -> [num_items][N_i x num_labels]   (nested per item)
#   each row = per-anchor label-token scores: softmax over labels when
#   apply_softmax=True, else exp(logprob) probabilities (the existing pointwise
#   CausalLM conversion, applied per anchor).

HTTP (POST /v1/score)

{
  "query": "",
  "items": ["Rank the candidates. C0. C1. C2. Scores:<|object_ref_start|><|object_ref_start|><|object_ref_start|>"],
  "label_token_ids": [9454, 2753],
  "score_extraction_token": "<|object_ref_start|>",
  "apply_softmax": false
}

No new request/response fields: this reuses score_extraction_token /
score_extraction_token_id and the existing nested ScoringResponse.scores union.
label_token_ids is required for CausalLM (as for pointwise CausalLM scoring).

Semantics & validation

  • Both execution modes. Batched scores each item as an independent query + item
    sequence; fused --enable-mis packs the items into one sequence with a
    block-diagonal mask that isolates each set, and splits the flat result back per
    item. Both return nested [num_items][Nᵢ x num_labels].
  • The SequenceClassification is_score_and_pool_model allow-list is unchanged and
    still governs the non-generation path; generation skips that check.
  • All other setwise constraints are shared: every item must contain at least one
    extraction token, the shared query prefix must not contain it, and radix cache /
    chunked prefill must be off with no --allow-auto-truncate (pooling positions are
    full-prompt coordinates). All raise a clean 400 before inference.
  • return_pooled_hidden_states stays rejected for CausalLM (no task head).

How it works

token_indices_to_pool already flows to every ForwardBatch. For SequenceClassification
the pooler consumes it; for CausalLM there is no task head, so the LM head is applied
at those positions instead — mirroring the existing multi-item (--enable-mis)
delimiter path, but pooling AT the anchor (no delimiter - 1 shift, no discarded
row). The forward dispatch prioritizes the anchor readout over the delimiter path, so
it serves both modes; the block-diagonal MIS mask is applied independently by the
attention backend (gated on enable_mis), so a batched request never triggers it
while a fused one does.

Implementation

Area File Change
LM-head readout layers/logits_processor.py compute_logprobs_at_positions() — reads the LM head AT token_indices_to_pool and gathers label_token_ids logprobs (one row per anchor); a forward() branch dispatches to it for prefill-only requests carrying the field, taking precedence over the MIS delimiter path so it serves both modes.
Logprob bookkeeping managers/scheduler_components/logprob_result_processor.py Generalized "position-based scoring" (_scoring_positions): setwise anchors take precedence over MIS delimiters (a fused setwise request carries both), so the expected input-logprob count is len(token_indices_to_pool).
Request plumbing managers/io_struct.py token_indices_to_pool on GenerateReqInput (per-item split in __getitem__) and TokenizedGenerateReqInput.
managers/tokenizer_manager.py, managers/scheduler.py Thread the field into the tokenized generate request and the Req.
Score engine managers/tokenizer_manager_score_mixin.py Validation allows CausalLM setwise (batched and --enable-mis); score_request sets token_indices_to_pool + logprob_start_len=0; the batched result branch reads per-request input_token_ids_logprobs, and the fused _process_multi_item_extraction_results gains a generation branch that reads the flat input_token_ids_logprobs and splits per item via per_item_anchor_counts. Both reuse _convert_logprobs_to_scores (per-item labels + temperature).
Docs entrypoints/openai/protocol.py, entrypoints/engine_score_mixin.py score_extraction_token(_id) docs note CausalLM support.
Tests test/registered/unit/managers/test_setwise_score_mixin.py, test/registered/e2e/scoring/test_setwise_scoring.py CPU result-grouping + validation tests and a CausalLM HF-parity / MIS set-isolation E2E suite.

Tests

# CPU unit tests (no GPU / model needed)
python test/registered/unit/managers/test_setwise_score_mixin.py -v
python test/registered/unit/layers/test_pooler_score_and_pool.py -v

# E2E (GPU); CausalLM setwise uses a decoder-only checkpoint
export TEST_CAUSAL_LM_MODEL=<path-to>/Qwen3-0.6B
python test/registered/e2e/scoring/test_setwise_scoring.py -v
  • Unit (test_setwise_score_mixin.py): a _GenHarness (is_generation=True)
    plus a TestGenerationSetwiseResults suite covering nested grouping (batched and
    fused/MIS) from input_token_ids_logprobs, softmax on/off, missing-label → 0.0,
    empty-logprobs / row-count-mismatch errors, and a pointwise-unchanged regression;
    validation tests assert CausalLM setwise is allowed.
  • E2E (test_setwise_scoring.py): TestGenerationSetwiseScoringHFParity asserts
    numerical parity against a HuggingFace reference (gather log_softmax(LM-head logits) at each anchor, exponentiate for apply_softmax=False), single- and
    multi-item. TestGenerationSetwiseMISScoring asserts block-diagonal set
    isolation
    — the property only a real fused forward exercises. Shape/grouping/error
    cases live in the CPU unit tests, not duplicated in the GPU suite.

Validation

Environment: H100 80GB, flashinfer 0.6.18, sglang-kernel 0.4.7, Qwen3-0.6B
(CausalLM base, anchor <|object_ref_start|>), engine float16 on flashinfer.

Suite Cases Type Result
test_setwise_score_mixin.py 50 CPU unit PASS
test_pooler_score_and_pool.py 12 CPU unit PASS
test_token_scoring.py 10 CPU unit PASS
test_setwise_scoring.py::TestGenerationSetwiseScoringHFParity 2 E2E (HF parity, single + multi-item) PASS
test_setwise_scoring.py::TestGenerationSetwiseMISScoring 1 E2E (fused set isolation) PASS
lint: ruff / ruff-format / isort / codespell / registered-tests — lint PASS

HTTP smoke test (launch_server + curl /v1/score)

  • Setwise, 1 item × 3 anchors, apply_softmax=false → nested [1][3 x 2] probability
    matrix.
  • Setwise, 1 item × 2 anchors, apply_softmax=true → nested [1][2 x 2], each row
    sums to 1.0.
  • Pointwise (no score_extraction_token), 2 items → flat [2 x 2] — backward
    compatibility intact.

Backward compatibility

Purely additive. Pointwise CausalLM scoring is unchanged (score_extraction_token_id
unset → last-token output_token_ids_logprobs path, flat [num_rows x num_labels]).
SequenceClassification setwise behavior is untouched — the only shared-path change is
the logprob-count bookkeeping, which now recognizes token_indices_to_pool in
addition to MIS delimiters.
}


CI States

Latest PR Test (Base): ❌ Run #36188053776
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ⏳ Run #36188053694

Extend setwise scoring (score_extraction_token) to CausalLM models. The LM head
is read at each extraction-token anchor and the requested label_token_ids
logprobs are gathered, returning one [N_i x num_labels] matrix per item (nested).
token_indices_to_pool stays the unifying readout signal; for CausalLM the LM head
is applied there instead of the pooler head.

Both execution modes are supported: batched (each item an independent query+item
sequence) and fused --enable-mis (block-diagonal mask isolating sets, flat result
split per item via per_item_anchor_counts). The forward dispatch prioritizes the
anchor readout over the MIS delimiter path so it serves both; the block-diagonal
mask is applied independently by the attention backend (gated on enable_mis).

Stacked on sgl-project#38965 (SequenceClassification setwise).
@Qiaolin-Yu Qiaolin-Yu self-assigned this Sep 25, 2026
@sundar24295s

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@sundar24295s sundar24295s added the run-ci-extra CI: also run the extra suite (requires run-ci) label Sep 25, 2026
@sundar24295s

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@Qiaolin-Yu Qiaolin-Yu added the bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) label Sep 25, 2026
@Qiaolin-Yu Qiaolin-Yu added release-highlight Candidate PR for release note highlight highest-priority CI: all three control labels, plus never batch-cancelled or stale-closed and removed run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 25, 2026
@Qiaolin-Yu

Qiaolin-Yu commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

failed cis are unrelated

@Qiaolin-Yu
Qiaolin-Yu merged commit 3ed56a3 into sgl-project:main Sep 26, 2026
183 of 287 checks passed
arbi-dev added a commit to arbicity/sglang-turbo that referenced this pull request Oct 7, 2026
…5.21 (#12)

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21

The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream
v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the
commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in
the post-port commits #9 (hybrid sliding-window through the kv-cache
plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory
reads attention_backends()).

Upstream's structure wins and the seam is re-applied on it:
- server_args.py is now a thin record over arg_groups/: the
  --kv-cache-dtype choices are hoisted into arg_groups/choices.py as
  KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from
  server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a
  plugin backend move to arg_groups/model_hook.py.
- ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's
  resolution hooks (arg_groups/resolution_hooks.py) are the official slot,
  and turbo-attn's plugin now pairs the flags there.
- Plugin dtype reads go through the resolved model bag (get_model()), not
  the raw ServerArgs record, which v0.5.21 leaves as operator input.
- HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it
  holds and stands in as the per-layer cost, so upstream's new draft-SWA
  and unified-pool formulas apply unchanged; layer ids come from
  kvc.layer_info, as upstream's own SWA pool takes them.
- Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21
  a PP draft may load its own embedding (pp_draft_embedding).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
sam123456598 added a commit to vessl-ai/sglang that referenced this pull request Oct 8, 2026
* [feature] add per-item candidate token scoring and calibration (sgl-project#40826)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) highest-priority CI: all three control labels, plus never batch-cancelled or stale-closed release-highlight Candidate PR for release note highlight run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants