Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1086 commits
Select commit Hold shift + click to select a range
cfbc5af
[BugFix] lora_base_layer / routed_experts order in expert param mappi…
HollowMan6 Aug 17, 2026
ceb340e
fix: prevent PyNvVideoCodec decoder slot limit bypass via ClassVar sh…
jperezdealgaba Aug 17, 2026
402547d
[Bugfix][CI] Release the shared ColBERT engine before `test_colbert_h…
stefankoncarevic Aug 17, 2026
75dde08
[Perf][MoE] Optimize deepep_v2 receiver CPU Overhead (#51114)
LucasWilkinson Aug 17, 2026
c1e4387
[ROCm][CI] Restore Torch defaults and type DSV4 scratch buffers (#52566)
AndreasKaratzas Aug 17, 2026
3fc2893
[Bugfix] Account for local DP workers in startup thread allocation (#…
cr-zhao Aug 17, 2026
9633933
Relax CuPy constraint to only exclude 14.1.0 (#44284)
khluu Aug 17, 2026
d1e3eee
[Spec decode] Support Kimi-K3 DCP with DSpark (#52188)
wzhao18 Aug 17, 2026
f08a95f
[Rust Frontend][gRPC] Advertise LoRA capabilities (#52031)
connorcarpenter15 Aug 17, 2026
8878ebd
[ROCm][CI] Expand AITER W4A4 MoE Coverage (#52647)
micah-wil Aug 17, 2026
e68fb75
[ROCm][AMD][Installation] add LMCache kv-connector installation and …
hongxiayang Aug 17, 2026
455edc0
[ModelRunnerV2] Support prompt embeds (#42963)
gcanlin Aug 17, 2026
60c3a31
[CI][AMD] Improve Kubernetes failure diagnostics (#52264)
AndreasKaratzas Aug 17, 2026
5ae2d38
[Perf][Structured Output] Skip unused request-local reasoners (#52573)
BugenZhao Aug 17, 2026
49fb2ee
[ROCm][CI] add Aiter ops tests (#52208)
divakar-amd Aug 17, 2026
c296cf8
[Bugfix] Add forward_xpu to XDRotaryEmbedding for HunyuanOCR on XPU (…
jbyczkow Aug 17, 2026
58aa1e3
[Bugfix][SM120][MLA] Disable dense prefill for FlashInfer sparse MLA …
tommy-asai-sonarsource Aug 17, 2026
0e8989b
[ROCm] gaurd on_gfx1250 call with rocm platform (#52625)
jikunshang Aug 18, 2026
0db502c
[Kimi-K3] support DCP partial prefix cache hit (#50493)
GirasoleY Aug 18, 2026
c296851
[MoE] Refine FlashInfer one-sided All2All integration (#51924)
bobboli Aug 18, 2026
cdb8545
[Kernel][Perf] Support Qwen head ratios in fused GDN MTP (#52539)
BabyDrangoner Aug 18, 2026
69d3335
[Bugfix][Quantization] Guard the MXFP8 FlashInfer path on FlashInfer …
LH-and-FPGA Aug 18, 2026
f4b161d
[XPU][UT] Fix OOM and skip graph case (#49287)
mayuyuace Aug 18, 2026
d5f5de7
[Bugfix] Fix DeepSeek V4 mHC broadcast buffer for weight sync (#52626)
HollowMan6 Aug 18, 2026
e0e5a7f
[Rust Frontend] Fix GLM-5.2 chat template rendering parity (#51426)
WoosukKwon Aug 18, 2026
101c447
[Bugfix] Handle DeepseekV4ForCausalLM in benchmark_moe get_model_para…
SayHelloToWorld Aug 18, 2026
aa99034
[Cohere] Misc changes to cohere model definitions (#50156)
kkt-cohere Aug 18, 2026
b0e9cff
[XPU] update xpu-manager to v2.1.0 (#52569)
yma11 Aug 18, 2026
c89d692
[Build] Propagate vLLM version to Rust binaries (#52593)
BugenZhao Aug 18, 2026
d785eb5
[Test] Add pause/resume E2E tests (#52144)
floatlibai Aug 18, 2026
5fa8ca9
[ROCm][CI] Move ROCm AITER quantization tests (#40938)
AndreasKaratzas Aug 18, 2026
2687fec
Replicated embedding and norm fusion for DSV3 flat model (#48484)
jeejeelee Aug 18, 2026
b01728b
[Bugfix] Return 4xx for client-caused errors in /detokenize (#52622)
rajathpi Aug 18, 2026
5c9ff53
[Bugfix] Accept logprobs=-1 in the Completion API (#46175)
he-yufeng Aug 18, 2026
e8ad285
[Bugfix][ROCm] Fix a few int4/int8 quantization errors (#52112)
qli88 Aug 18, 2026
eab1cff
Harden DeepSeek V3.2 fused kernel grids (#52381)
yimdev Aug 18, 2026
41f179b
[Rust Frontend] Simplify data-parallel size ownership (#52575)
BugenZhao Aug 18, 2026
be3f614
[CI] Register CPU CI "VLLM_CPU_CI_ENV" environment variable (#52633)
taneem-ibrahim Aug 18, 2026
d29dc3a
[Bugfix][Gemma4] Align parser enable_thinking default with template (…
lxy-alexander Aug 18, 2026
689be2b
Upgrade Flashinfer version to 0.6.17 (#52681)
wzhao18 Aug 18, 2026
3bb9c18
[Multimodal] Reorganize video decoder backends (#49155)
Isotr0py Aug 18, 2026
241ff8c
[Model] Enable LoRA support for tower and connector in LlavaNextForCo…
gangula-karthik Aug 18, 2026
f9f066d
[Bugfix][PaliGemma] Remove stale image embedding scaling (#52692)
ActiveSky Aug 18, 2026
88b2bff
[MOE] Standardize and abstract fused shared expert optimization selec…
fxmarty-amd Aug 18, 2026
bca7bea
Remove VLLM_TEST_FORCE_FP8_MARLIN to replace with linear_backend/moe_…
mgoin Aug 18, 2026
ddbf826
[ROCm] Gate Torch FP8 scaled-MM on architecture support (#51021)
sstamenk Aug 18, 2026
ad5e71b
[ROCm][Perf] Enable fused KDA decode on gfx942 (MI325X) (#52293)
mpashkovskii Aug 18, 2026
01e56ca
[Bugfix][MLA] Do not use Dense MHA for GLM-5.2 (#52512)
WoosukKwon Aug 18, 2026
d75136c
[Rust Frontend] Wait for all utility calls to finish (#52671)
connorcarpenter15 Aug 18, 2026
90984dd
[CI] Upgrade huggingface-hub to 1.28.0 (#52797)
AndreasKaratzas Aug 18, 2026
6948a43
[Bugfix] Detect all attention-spelling variants in ModelConfig.is_hyb…
mganczarenko Aug 18, 2026
9842d70
[DBO][CI] Increase the coverage of prefill DBO in test_dbo.py (#48628)
SageMoore Aug 18, 2026
8d6b183
[CI] Standardize test job labels by device (#52659)
khluu Aug 18, 2026
6a391a9
[Rust Frontend][RL] add routed expert prompt offset (#52703)
biswapanda Aug 18, 2026
7ddb507
[Bugfix][DP] Don't assume the engines started when forwarding a wake …
aoshen02 Aug 18, 2026
12f64b3
[Bugfix][Structured Output] Stop XGrammar token batches at terminatio…
sfeng33 Aug 18, 2026
5d8a4cf
[ROCm] Pad non-aligned AITER MLA heads (#51647)
LiuYinfeng01 Aug 18, 2026
6066bb3
[ROCM][CI] Attention test speedup (#52763)
stefankoncarevic Aug 18, 2026
d3fafe0
[Quantization] Remove the dead ocp_mx_scheme branch from moe_kernel_q…
xuebwang-amd Aug 18, 2026
5f7a20b
[nv] add pcp support in dsv3.2 (#52046)
GirasoleY Aug 18, 2026
203926c
[ROCm][CI] Add AMD CI Pull-Request Commands (#52822)
AndreasKaratzas Aug 18, 2026
aa6abec
[CI][ROCm] Prevent Git maintenance races during shallow fetches (#52810)
AndreasKaratzas Aug 18, 2026
8f4a7f4
[ROCm][CI] Gating more ROCm tests (#44969)
AndreasKaratzas Aug 18, 2026
ef47a89
[Core] Make prefix-cache NONE_HASH deterministic by default (#51875)
russellb Aug 18, 2026
a9f4afb
[Bugfix] Fix DeepSeek V4 mHC broadcast buffer for dummy load (#51368)
HollowMan6 Aug 19, 2026
f1178f3
Revert DSv4 eager workspace reuse (#52836)
WoosukKwon Aug 19, 2026
b05ae5d
[CI][Bugfix] Complete DeepSeek-V4 FSE test fixture contract (#52842)
khluu Aug 19, 2026
d36cc42
[Bugfix][Elastic EP] Reject scale below the minimum data parallel siz…
Etelis Aug 19, 2026
a2257f9
[Elastic EP] Reduce eager-mode reconfiguration downtime (#51885)
itayalroy Aug 19, 2026
b1d9337
[EPD] Allow KV consumers to omit MM embeddings (#52697)
zhenwei-intel Aug 19, 2026
f485081
[XPU] Support EC connector KV Offloading on XPU (#49532)
chaojun-zhang Aug 19, 2026
e4d61d0
[CPU] Add AMX-only high-performance MLA backend for DeepSeek V2/V3/R1…
bigPYJ1151 Aug 19, 2026
f936a26
[Bugfix][XPU] Fix Mamba state pointer overflow (#48109)
Oxygen56 Aug 19, 2026
03b87dc
[XPU] upgrade requirements/test/xpu.txt (#52672)
mayuyuace Aug 19, 2026
aabc1a0
[Bugfix][Benchmark] Check readiness before tokenizer init in rust vll…
tlrmchlsmth Aug 19, 2026
deeeae7
[MM] Keep more metadata tensors on CPU (#52827)
njhill Aug 19, 2026
e575b5f
[XPU][CI] fix hf runner (#52730)
mayuyuace Aug 19, 2026
8e46acc
[Attention] Vectorize sparse MLA mask loads (#52217)
MatthewBonanni Aug 19, 2026
3130573
[Model] Add GraniteSWA and GraniteMoeSWA via existing Granite (#52706)
daviswer Aug 19, 2026
63ff748
[Attention][MLA] FlashMLA sparse: DCP on the fp8_ds_mla mixed-batch p…
drakosha Aug 19, 2026
08afae2
[ModelRunner v2] Enable MRV2 for pooling models by default (#48290)
taneem-ibrahim Aug 19, 2026
86a89a9
[Bugfix][Frontend] Run the serve arg checks for `vllm launch` too (#5…
vineethsaivs Aug 19, 2026
5a4c8d9
[Bugfix][LoRA] Add embedding_modules for Qwen3.5 CausalLM (#48850)
Agoni-02 Aug 19, 2026
ee11730
[BugFix] Revert incorrect MM keep_on_cpu=True changes (#52881)
njhill Aug 19, 2026
b09bd69
[Model][NVIDIA] Route DSA models to the CUDA non-compiled path (#52861)
WoosukKwon Aug 19, 2026
842dd8f
[Pooling] Use semantic task validation errors (#52867)
taneem-ibrahim Aug 19, 2026
9a9aa2b
[Bugfix] Redact api_key in non-default args log (#52523)
Andy365-365 Aug 19, 2026
93eea4f
[XPU][CI] downgrade sentencepiece (#52904)
mayuyuace Aug 19, 2026
340b7e4
fix: reject string schemas that mix pattern/format with length bounds…
he-yufeng Aug 19, 2026
eac636a
[Frontend] Move api_server.py out openai folder (#52131)
noooop Aug 19, 2026
cba0676
[Platform] Fill in the missing backend parameter for torch.compile (#…
wangxiyuan Aug 19, 2026
b160cab
[Bugfix] Restore model info caching for package backends (#52690)
haoyangqian Aug 19, 2026
58302b4
[Doc] Fix group numbering in Case 3 of hybrid_kv_cache_manager.md (#5…
qwerqwerqwe8688-jpg Aug 19, 2026
2f54100
[CI] Fix docs build (#52937)
hmellor Aug 19, 2026
58de6cb
Add NemotronH_Omni_Reasoning_V3 as a supported Nemotron architecture …
Naveassaf Aug 19, 2026
92bdee0
[Bugfix][Frontend] Return all choices from /inference/v1/generate whe…
qgallouedec Aug 19, 2026
c2e7242
[Bugfix][LoRA] Guard None group members in expand_packed_lora (partia…
eilamc14 Aug 19, 2026
2b7fcbf
[Kernel] SM120: stop routing misaligned-M blockwise FP8 GEMMs to the …
lucifer1004 Aug 19, 2026
db92053
[Core] Skip broadcasting mm tensor data to workers for prefix-cache-c…
sseanliu Aug 19, 2026
2d7f42b
[Build] Add InstantTensor to CUDA dependencies (#52801)
mgoin Aug 19, 2026
be06873
[Bugfix] compressed-tensors: restore int8 grouped WNA16 MoE support (…
y0hnn Aug 19, 2026
160f7f0
[ROCm][CI] Extended Fused MoE and FP8 MoE test support (#41100)
AndreasKaratzas Aug 19, 2026
17dbd42
[ROCm] Add UE8M0 scale packing for Triton silu_mul_quant (#37835)
AndreasKaratzas Aug 19, 2026
525b7bb
[Bugfix][CPU] Enable C++ causal_conv1d GDN path and float32 SSM cache…
dineshchitlangia Aug 19, 2026
583a002
[ROCm] [Bugfix] Fix Triton fused shared expert alignment (#51632)
akii96 Aug 19, 2026
c676232
[Bugfix][Quantization] Fix OCP MX MoE emulation silently skipping mxf…
xuebwang-amd Aug 19, 2026
54dd98b
[CT] Support Humming for WNA16 MoE (#48918)
yiliu30 Aug 19, 2026
e9e1630
[Model] Support bidirectional (encoder-only) attention for DeepSeek e…
Lossfull Aug 19, 2026
3a386cf
[ROCm] Give EngineCore cleanup grace after request abort (#52281)
AndreasKaratzas Aug 19, 2026
0c8c3f4
[Bugfix][V1] Sync mamba_block_size via EngineCoreReadyResponse (#50809)
lxyxinyi Aug 19, 2026
9b5f345
[LoRA] Avoid false target matches for unsupported module types (#52313)
linitra24 Aug 19, 2026
f76d71d
[CI/Build] Fix CPU platform pre-commit formatting (#52981)
mgoin Aug 19, 2026
755492e
Revert "[Kernel] Gemma-4 FA4 FP8 Kernel" (#52987)
ywang96 Aug 19, 2026
480d4f0
[ROCm][CI] Enable modular OAI Triton MoE tests (#46434)
AndreasKaratzas Aug 19, 2026
cb58bb9
[CI] Harden RemoteVLLMServer GPU cleanup checks (#52282)
AndreasKaratzas Aug 19, 2026
541c6d6
[Bugfix][Quantization] Support CT block FP8 with Marlin (#52966)
mgoin Aug 19, 2026
d591d1d
[Bugfix] Add Kimi K3 MoE support to benchmark_moe.py (#50082)
vanshbhatia-amd Aug 19, 2026
c205726
[CI] Fix and extend PR/issue auto-labeling (#51459)
jcotant-inferact Aug 19, 2026
823ec22
[ROCm]: Bump triton 3.7 commit (#52819)
Rohan138 Aug 19, 2026
5bf0dbd
[Bugfix] vLLM crashes at startup when DeepEP v2 is used with `--enfor…
SageMoore Aug 20, 2026
0a111cc
[kv_offload] fix(metrics): rename kv_offload_tiering_block_{queries,h…
ronensc Aug 20, 2026
d4f4d3f
[Model Runner V2][Spec Decode] Fix draft logits cache column stride i…
TheEpicDolphin Aug 20, 2026
bf2866f
[KV Connector] Add decode offloading to Mooncake Store consumers (#52…
chengy-sysu Aug 20, 2026
58e5ee0
[refactor] consolidate cp attn ops (#52839)
GirasoleY Aug 20, 2026
6a96207
[Distributed] Enable FlashInfer all-reduce by default (#52998)
WoosukKwon Aug 20, 2026
fbb4c04
[ROCm][CI] Speed up `test_rocm_aiter_qk_norm_rope_kvcache_fusion` (#5…
micah-wil Aug 20, 2026
c0233fc
[Model] Remove unused DeepseekV32Indexer forward (#53021)
WoosukKwon Aug 20, 2026
76fb6d2
[Bugfix][Core] Reserve the KV null block when validating max_model_le…
92hyungjun Aug 20, 2026
f85e060
[Doc] Update Gaudi HPU committers (#52726)
pmanczak Aug 20, 2026
e85dd21
fix: report stop_sequence stop_reason in Anthropic Messages API (#45807)
he-yufeng Aug 20, 2026
4f66bc3
[CI][ROCm] Standardize AMD test job labels by device (#52976)
AndreasKaratzas Aug 20, 2026
11baa0e
[Attention] Avoid redundant mask compute in GDN metadata build (#52078)
xyang16 Aug 20, 2026
9b2aef5
[Bugfix] Video loading: sample over presentable frames, not header sa…
AmitMY Aug 20, 2026
cd50353
[ROCm] [Bugfix] Preserve CPU query offsets during capture (#51585)
akii96 Aug 20, 2026
d626108
[ROCm][Perf] Fuse DeepSeek-V4 mHC post/pre and RMSNorm with AITER (#5…
shen-shanshan Aug 20, 2026
754e1c3
[CI][XPU] Skip test_fused_shared_expert.py on XPU (#53035)
mayuyuace Aug 20, 2026
a1c5b1f
[Core][V1] Support trace_decode_token_ids for deterministic decode re…
zllion Aug 20, 2026
d66300a
[Bugfix][EPD] Fix encoder round-robin fan-out (#52491)
AnkitNakhawa Aug 20, 2026
16cfe72
[Bugfix][Rust Frontend] Reject n > 1 in the `/inference/v1/generate` …
qgallouedec Aug 20, 2026
a34fd69
[CI][Bugfix] Update distributed DP API server test path (#52939)
khluu Aug 20, 2026
14617c2
[Bugfix] DeepEP-V2: expert_tokens_meta must be None on the decode/cud…
dmvevents Aug 20, 2026
4666a8b
[Refactor] Remove InputPreprocessor (#53064)
DarkLight1337 Aug 20, 2026
963fcfa
[Rust][Benchmark] Align speed-bench CLI flags with Python and add fla…
esmeetu Aug 20, 2026
c0a25c0
[Bugfix][Security] Guard _load_ov2_processor with resolve_trust_remot…
jperezdealgaba Aug 20, 2026
5d4d470
[CI] Fix nonexistent dependency for data-parallel example test select…
taneem-ibrahim Aug 20, 2026
c8de519
[Kernel][Kimi] fused vision q/k roper kernel (#50400)
lengrongfu Aug 20, 2026
38e9cef
[Bugfix] Return HTTP 400 instead of 501 for unknown chat roles in Dee…
JC-ut0 Aug 20, 2026
44cf3f0
[XPU][Tests] Make tests device-agnostic (#51968)
pmanczak Aug 20, 2026
6b68db4
[ROCm] Fix DeepSeek V4 indexer numerics and coverage (#50803)
AndreasKaratzas Aug 20, 2026
30e2394
[Bugfix] Record non-ImportError attention backend probe failures inst…
Eoin-Houstoun Aug 20, 2026
727274a
[Rust Frontend] Fix Qwen parser auto-detection (#51169)
sagearc Aug 20, 2026
5b1e7a8
[BugFix][Mooncake] Fix Mooncake saves from sparse Mamba block tables …
ZeldaHuang Aug 20, 2026
1eab6fe
[CI] replace shellcheck script with shellcheck-py hook (#52572)
wjabbour Aug 20, 2026
6e85feb
[3/N] Harden Transformers modelling backend multi-modal path (#51827)
hmellor Aug 20, 2026
f0c14b4
Fix weight tying (#51665)
hmellor Aug 20, 2026
6df7adc
[Bugfix][GDN] Reset speculative decode count for an empty draft sched…
khluu Aug 20, 2026
6259572
[Docs] Use incremental builds for C++ changes in `AGENTS.md` (#53098)
gau-nernst Aug 20, 2026
df13769
[Model] Add tower and connector LoRA support for LFM2-VL (#51498)
zupengwang Aug 20, 2026
4b7cb94
[Docker] Update to nixl-1.3.2 (#51777)
sandeep-maddipatla Aug 20, 2026
de216b6
[Bugfix] Skip MM processor cache inserts larger than capacity (#53016)
Prudhvivuda Aug 20, 2026
bd8865a
[Kernel] Add FlashInfer TRTLLM MXFP8 linear backend (#52204)
seonjinn Aug 20, 2026
cb09dd7
[Core][Multimodal] Skip redundant placeholder scan when token match s…
Yiqin-17 Aug 20, 2026
01af92e
[Feature][Model Runner V2] Support extract_hidden_states speculation …
zupengwang Aug 20, 2026
2dd1722
Fix Transformers modelling backend `RMSNormFuser.fuse` performance (#…
hmellor Aug 20, 2026
1fe3a15
Reduce `AutoWeightsLoader` kwargs (#53106)
hmellor Aug 20, 2026
ae25628
upgrade tpu-inference to v0.27.0 (#53088)
meiyeh123 Aug 20, 2026
00f7f25
[Misc] Don't allow language-model-only used with encoder CG together …
Isotr0py Aug 20, 2026
7c8b68b
[Bugfix][MiniMax-M3] Keep FP8 query allocation stable across CUDA gra…
kyleliang-nv Aug 20, 2026
d56bbf3
[Bugfix] Support MistralCommonBackend tokenizers in structured output…
thanhpt1110 Aug 20, 2026
3b829cf
[CI/Build][ROCm] Keep the CUDA-only kernel tests out of the ROCm run …
stefankoncarevic Aug 20, 2026
4f6885f
[DSV4][Kernel] Fuse shared experts into MegaMoE (#53040)
gcanlin Aug 20, 2026
bfb6c13
[Bugfix][MoE] Tune FlashInfer experts to scheduler token limit (#52989)
mgoin Aug 20, 2026
54ba80d
[CI][Docker] Pin manylinux2_28-builder:cuda13.0 to the release/2.13 i…
atalman Aug 20, 2026
2f41c89
Fix seed loss when batch contains unseeded requests (#51866)
MKQuantum Aug 20, 2026
7cfb97e
[Frontend][Core][Spec Decode] Per-request acceptance stats in OpenAI …
matthewkotila Aug 20, 2026
0a5a551
[CI][Docker] Pin remaining manylinux builder images (#53172)
khluu Aug 20, 2026
f32b17b
[Rust Frontend] Support `--generation-config vllm` (#53044)
BugenZhao Aug 21, 2026
d29f7f5
[Bugfix] Load untied Gemma LM head weights (#53170)
khluu Aug 21, 2026
91a893d
[Bugfix][Spec Decode] Scope DSpark backend inheritance to DeepSeek V4…
mgoin Aug 21, 2026
83c5d59
[Rust Frontend] Replace external `protoc` with pure Rust lib `protox`…
BugenZhao Aug 21, 2026
c6e19b3
[Bugfix][Structured Output] Avoid spurious FSM errors after speculati…
chaunceyjiang Aug 21, 2026
df3b342
[Rust Frontend] Add HY3 unified parser and local XGrammar structural-…
BugenZhao Aug 21, 2026
5df31ea
[Spec Decode] Enable adaptive verification on DSv4 + sm90 (#52795)
LucasWilkinson Aug 21, 2026
2adf4b9
[Rust Frontend] Fix Kimi K3 reasoning_effort="none" handling (#53043)
reidliu41 Aug 21, 2026
b00f475
[MM] Remove text components from ProcessorInputs (#53093)
DarkLight1337 Aug 21, 2026
2785c72
[Kimi-K3] Extend GEMM-RS to GEMM-AR (#53053)
gau-nernst Aug 21, 2026
ba07e4a
[Bugfix] Fix batch-invariant fp32 matmul OOR on SM89 for N=1 (#52960)
vhagor Aug 21, 2026
df344bf
[Rust][Benchmark] Load HF datasets from parquet shards via hf-hub, fi…
esmeetu Aug 21, 2026
b389ac2
[Spec Decode] DFlash2: local convolution + candidate selector (#52816)
SubSir Aug 21, 2026
f8e0602
Support kimi k3 nvfp4 checkpoint (#53132)
wzhao18 Aug 21, 2026
9fd750f
[XPU][INC] Add int4 w4a8 (dynamic int8 activation) backend for INC li…
tthakkal Aug 21, 2026
0a21947
[XPU][CI]Add parallelism for long-running Intel GPU cases (#52257)
zxd1997066 Aug 21, 2026
36bad1b
[XPU][CI]Add more cases in intel GPU CI and reorganize to align non-x…
zxd1997066 Aug 21, 2026
e85d1b6
[Bugfix][LoRA] Use an explicit capability flag for tower connector Lo…
linitra24 Aug 21, 2026
6feafb8
[XPU] follow cuda path for mrope on XPU (#53201)
yma11 Aug 21, 2026
a60c66e
[Bugfix] Fix int32 index overflow in LoRA punica kernels at long cont…
ShuaiShao93 Aug 21, 2026
c8438a3
[Doc] Fix dead link in KV transfer README (#53230)
hagaikwa-redhat Aug 21, 2026
cda3868
[Bugfix] Deterministic MoE combine (reduce_scatterv) under VLLM_BATCH…
shijuzhao Aug 21, 2026
5ee84d3
[CI][Bugfix] Use a prompt that survives offload-resume rounding in ma…
frgossen Aug 21, 2026
2740c81
[Kernel] Add b12x FP4 MoE backend (#52018)
lukealonso Aug 21, 2026
574e6a0
[Bugfix][Attention] Normalize FlashInfer prefill LSE before merging (…
yimdev Aug 21, 2026
6f74337
[Bugfix][Spec Decode]Preserve user --speculative-config overrides for…
elwhyjay Aug 21, 2026
6d8cd88
[Docs] document cache salting for prefix cache timing side-channel mi…
russellb Aug 21, 2026
08b73d0
[Frontend] Prevent Kimi K3 reserved markers in response text (#52889)
BugenZhao Aug 21, 2026
18aa245
using existing uvicorn configuration for dp supervisor (#52473)
Gregory-Pereira Aug 21, 2026
fbb17e7
[Bugfix] Fix six quantization exception messages split across positio…
ErenAta16 Aug 21, 2026
1183f04
[Bugfix] test_batch_inference_correctness now uses batch invariance (…
morrison-turnansky Aug 21, 2026
c1d7a38
[Bugfix] Fix HYV3 shared_mlp prefix for compressed-tensors ignore mat…
DCoEngine Aug 21, 2026
cd7b7c2
[ROCm] Cpu offload for ROCm 7.13+ to align the hipMemcpyBatchAsync pa…
hongxiayang Aug 21, 2026
72aedcc
[DCP] Default query replication for GLM sparse attention (#50382)
LucasWilkinson Aug 21, 2026
27ec8ac
[Bugfix] Fix MTP draft model using local cache path instead of S3 URL…
SoluMilken Aug 21, 2026
d53b1c2
Remove native Hunyuan V1 and VL implementations (#53272)
xianbaoqian Aug 21, 2026
463aa5e
[ROCm][Perf] Kimi-K3 Fused kernels for KDA prefill (#52606)
kliuae Aug 21, 2026
6dcc5d7
[Bugfix] Include device index in compile cache paths (#38962)
nascheme Aug 21, 2026
b41dc4e
[Bugfix] Return HTTP 500 for non-streaming generate errors (#49195)
waynehacking8 Aug 21, 2026
e00a034
(security) fix: enforce decoder prompt-length validation for skip-che…
jperezdealgaba Aug 21, 2026
47cd1c8
[Core] Add dynamo_timed tracing for print_readable (#40834)
frgossen Aug 21, 2026
1baf372
[Frontend] Use VLLMValidationError for batch request URL validation (…
tanchao Aug 21, 2026
fe76112
[ROCm][Perf] Optimize DeepSeek V4 C4A top-k with AITER (#52882)
Fangzhou-Ai Aug 21, 2026
7a2fdba
[K3 Perf] Fuse MXFP4 top-k finalization into latent-tail, ~5% E2E lat…
yewentao256 Aug 21, 2026
ba53da6
Revert "[Bugfix][MoE] Tune FlashInfer experts to scheduler token limi…
vllm-agent Aug 21, 2026
d9e0ace
[Bugfix][Attention] Fall back to native FlashInfer decode when XQA ca…
stecasta Aug 21, 2026
d6c2fec
[Cleanup][MLA] Remove FlashInfer DSpark DCP support (#53139)
GirasoleY Aug 21, 2026
e6f35d3
[DSv4 Perf] Adaptive topk width for dsv4, making #50004 back (#52823)
yewentao256 Aug 21, 2026
f15ea66
[ROCm][Test] Use platform FP8 dtype in ModelOpt FP8_PB_WO test (#53268)
djramic Aug 21, 2026
c0ff334
Revert compile-cache device index regression on CPU (#53304)
khluu Aug 21, 2026
88eb946
[ROCm][CI] Stabilize MI355 FusedMoE test group (#53025)
AndreasKaratzas Aug 21, 2026
b5f7fcc
[ROCm][CI] Add float16 dtype and unsupported head size tests for page…
divakar-amd Aug 21, 2026
a0af854
[ROCm][CI] Stabilize MI355 FlyDSL MoE accuracy test (#53024)
AndreasKaratzas Aug 21, 2026
29b7c2f
[ROCm][CI] aiter kernel ops - enable rope test (#52854)
divakar-amd Aug 21, 2026
3e47a9a
[CI] Fix MultiConnector accuracy test lifecycle (#53023)
AndreasKaratzas Aug 21, 2026
0a3a4f7
[Bugfix][KV Connector][NIXL] Support PCP producers (#52779)
LucasWilkinson Aug 21, 2026
592e06f
Revert "[ROCm][Perf] Kimi-K3 Fused kernels for KDA prefill" (#53294)
khluu Aug 21, 2026
e3f6026
Revert "Remove native Hunyuan V1 and VL implementations" (#53296)
khluu Aug 21, 2026
f37d586
[CI/Build][ROCm] Run the TileLang HIP symbol checks in their own inte…
stefankoncarevic Aug 21, 2026
f94dcde
Fix Cohere ChatV2 citation and tool handling issues (#52175)
andrewbcohere Aug 21, 2026
a556f3f
Forward Anthropic vllm_xargs to sampling params (#53308)
vMaroon Aug 21, 2026
a34f2ab
[Test] Add focused hybrid MTP prefix-cache regressions (#53189)
mgoin Aug 22, 2026
9ff7041
[Bugfix][Spec Decode] Use group geometry for FlashAttention metadata …
mgoin Aug 22, 2026
8bdc70e
[6/N][KV-Cache Layout Refactor] Standardize KV cache layout (#51718)
LucasWilkinson Aug 22, 2026
7f4a1b7
[Bugfix][ROCM] Fix the MXFP8 block scale exponent (#53110)
stefankoncarevic Aug 22, 2026
b2db227
[Bugfix][R3] Unwrap UniformTypeKVCacheSpecs when selecting the routed…
HollowMan6 Aug 22, 2026
e9d1398
[Bugfix][Kimi K3] Enable deferred MoE finalization before weight load…
zyongye Aug 22, 2026
7ca49fb
[Refactor][Model Runner V2][Multimodal] Move the encoder-only path ou…
gty111 Aug 22, 2026
0b19ebc
[MM] Simplify _apply_hf_processor_main (#53275)
DarkLight1337 Aug 22, 2026
040700a
[MM] Address comments on #53275 (#53364)
DarkLight1337 Aug 22, 2026
1070454
[CI][ARM64] Restore known-good CUDA 13 builder image (#53374)
khluu Aug 22, 2026
704f12a
[Pooling] Report input throughput for batched requests (#53213)
taneem-ibrahim Aug 22, 2026
da329cc
[Bugfix] Fix speculative decoding for short_conv (LFM2) models (#50272)
zwischenraum Aug 22, 2026
e0446ec
Merge current main into the DCP correctness branch
drakosha Aug 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
166 changes: 166 additions & 0 deletions .buildkite/amd-disagg/cluster.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,166 @@
#!/bin/bash
# =============================================================================
# cluster.sh — single, generic config for the vLLM disaggregated (P/D) launcher.
# -----------------------------------------------------------------------------
# This script is *sourced* by vllm_disagg.sh.
#
# Every value is `${VAR:-default}`, so the environment always wins:
# environment variable > built-in default below
#
# So you can override anything inline:
# PREFILL_IP=10.0.0.1 DECODE_IP=10.0.0.2 ./vllm_disagg.sh prefill
#
# Site-specific values (model dir, IPs, NIC list, partition) are the defaults
# in the "site defaults" sections below — edit those for a new cluster.
# Model-SPECIFIC perf flags live in models.yaml, NOT here.
# =============================================================================

_CLUSTER_SH_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"

# ----------------------------------------------------------------- model / mode
# MODEL_NAME indexes into models.yaml; MODEL_DIR is the parent dir holding it.
# MODEL_PATH (resolved by the launcher) = ${MODEL_DIR}/${MODEL_NAME}. [site]
export MODEL_NAME="${MODEL_NAME:-DeepSeek-V3}"
export MODEL_DIR="${MODEL_DIR:-/data/models2}"

# Shared NFS root (5 TB, visible on every node): model weights + per-run logs. [site]
export SHARED_MOUNT="${SHARED_MOUNT:-/data}"

# Parallelism mode (launcher derives PARALLEL_MODE tp|ep from this):
# WIDE_EP_MODE=0 tp : each node is an independent TP server (TP8, 1P1D)
# WIDE_EP_MODE=1 ep : data-parallel + expert-parallel across xP/yD nodes (wideep)
export WIDE_EP_MODE="${WIDE_EP_MODE:-0}"

# ----------------------------------------------------------------- topology
# xP prefill nodes + yD decode nodes. IPADDRS is the ordered, comma-separated
# node list (prefill IPs first, then decode IPs). NODE_RANK is this node's global
# 0-based rank; under SLURM it defaults to $SLURM_PROCID. Leave IPADDRS empty to
# use the PREFILL_IP/DECODE_IP fallback defaults below (1P1D only).
export xP="${xP:-1}"
export yD="${yD:-1}"
export IPADDRS="${IPADDRS:-}"
export NODE_RANK="${NODE_RANK:-${SLURM_PROCID:-}}"
export PREFILL_IP="${PREFILL_IP:-10.0.0.1}"
export DECODE_IP="${DECODE_IP:-10.0.0.2}"

# Per-node GPU count and TP degree
export GPUS_PER_NODE="${GPUS_PER_NODE:-8}"
export TP_SIZE="${TP_SIZE:-${GPUS_PER_NODE}}"

# ----------------------------------------------------------------- ports
# TP-mode (WIDE_EP_MODE=0) server ports.
export PREFILL_PORT="${PREFILL_PORT:-8100}"
export DECODE_PORT="${DECODE_PORT:-8200}"

# EP-mode (WIDE_EP_MODE=1) ports: API serve port, DP RPC port, KV transfer port, and
# per-node MoRIIO local ping port (must differ from PROXY_PING_PORT).
export SERVE_PORT="${SERVE_PORT:-20005}"
export RPC_PORT="${RPC_PORT:-13345}"
export KV_PORT="${KV_PORT:-9711}"
export LOCAL_PING_PORT="${LOCAL_PING_PORT:-61555}"

# MoRIIO proxy: HTTP port clients/benchmark hit, plus the connector control ports.
# PROXY_PING_PORT MUST be 36367 — the proxy hardcodes its zmq service-discovery
# socket on that port; prefill/decode register to PROXY_IP:PROXY_PING_PORT.
export PROXY_IP="${PROXY_IP:-${PREFILL_IP}}"
export PROXY_PORT="${PROXY_PORT:-10001}"
export PROXY_PING_PORT="${PROXY_PING_PORT:-36367}"
export HANDSHAKE_PORT="${HANDSHAKE_PORT:-6301}"
export NOTIFY_PORT="${NOTIFY_PORT:-61005}"

export PROXY_SCRIPT="${PROXY_SCRIPT:-/app/vllm/examples/disaggregated/disaggregated_serving/moriio_toy_proxy_server.py}"

# MoRIIO KV transfer direction (injected into --kv-transfer-config by the launcher):
# 0 -> omit read_mode (default; MoRIIO write mode: prefill pushes to decode)
# 1 -> "read_mode": true (decode pulls KV from prefill; matches upstream disagg)
export MORIIO_READ_MODE="${MORIIO_READ_MODE:-0}"

# ----------------------------------------------------------------- router / gateway
# Selection for client (bench/accuracy) traffic:
# proxy -> the in-container MoRIIO proxy started by the launcher
# vllm-router -> an external `vllm/vllm-router` container started by the SLURM job
# on the rank-0 node (default)
# Both use the SAME MoRIIO discovery mechanism (prefill/decode register to
# PROXY_IP:PROXY_PING_PORT=36367); only the client HTTP front door differs.
export ROUTER_TYPE="${ROUTER_TYPE:-vllm-router}"
export ROUTER_PORT="${ROUTER_PORT:-30000}"
export ROUTER_POLICY="${ROUTER_POLICY:-round_robin}"
export VLLM_ROUTER_IMAGE="${VLLM_ROUTER_IMAGE:-vllm/vllm-router:nightly}"
# Single client-facing port bench/accuracy target: the router port when routing,
# else the proxy port. Env override always wins.
if [[ "${ROUTER_TYPE}" == "vllm-router" ]]; then
export GATEWAY_PORT="${GATEWAY_PORT:-${ROUTER_PORT}}"
else
export GATEWAY_PORT="${GATEWAY_PORT:-${PROXY_PORT}}"
fi

# Where per-run logs / benchmark results are written. A $SLURM_JOB_ID subdir is
# appended so each CI run is self-scoped (falls back to 'local' off-SLURM). [site]
_LOG_BASE="${LOG_BASE:-/data/${USER:-$(id -un)}/disagg_logs}"
export LOG_PATH="${LOG_PATH:-${_LOG_BASE}/${SLURM_JOB_ID:-local}}"

# ----------------------------------------------------------------- vLLM runtime
# Engine/platform/transport-level env (NOT model-specific). Model-architecture
# AITER kernel toggles live in models.yaml under each model's `env:` block.
export VLLM_USE_V1="${VLLM_USE_V1:-1}"
export HSA_NO_SCRATCH_RECLAIM="${HSA_NO_SCRATCH_RECLAIM:-1}"

#export HF_HUB_OFFLINE="${HF_HUB_OFFLINE:-1}"
#export TRANSFORMERS_OFFLINE="${TRANSFORMERS_OFFLINE:-1}"
#
#export HOME=/tmp
export HF_HOME="${HF_HOME:-/tmp/hf_home}"
export XDG_CACHE_HOME="${XDG_CACHE_HOME:-/tmp/.cache}"
export VLLM_ENGINE_READY_TIMEOUT_S="${VLLM_ENGINE_READY_TIMEOUT_S:-3600}"

# ----------------------------------------------------------------- scale
# DP/EP group formation timeout across nodes (seconds).
export DISTRIBUTED_TIMEOUT_SECONDS="${DISTRIBUTED_TIMEOUT_SECONDS:-7200}"

# ----------------------------------------------------------------- RDMA / NCCL
# AMD Pensando AINIC RoCE fabric: 8 NICs exposed as ionic_0..7 (netdevs eth2..9),
# each rail on its own /24. GID index 1 + traffic class 104
_IB_DEVICES="${IB_DEVICES:-ionic_0,ionic_1,ionic_2,ionic_3,ionic_4,ionic_5,ionic_6,ionic_7}"
_IB_GID_INDEX="${NCCL_IB_GID_INDEX:-1}"
export IB_DEVICES="${IB_DEVICES:-${_IB_DEVICES}}"
export NCCL_IB_HCA="${NCCL_IB_HCA:-${_IB_DEVICES}}"
export NCCL_IB_GID_INDEX="${NCCL_IB_GID_INDEX:-${_IB_GID_INDEX}}"
export NCCL_IB_DISABLE="${NCCL_IB_DISABLE:-0}"
# In-box RCCL net transport (no external plugin); pin bootstrap to the VPC iface.
export NCCL_NET_PLUGIN="${NCCL_NET_PLUGIN:-none}"
export NCCL_SOCKET_IFNAME="${NCCL_SOCKET_IFNAME:-eth0}"
export GLOO_SOCKET_IFNAME="${GLOO_SOCKET_IFNAME:-${NCCL_SOCKET_IFNAME}}"
export NCCL_CROSS_NIC="${NCCL_CROSS_NIC:-0}"
export NCCL_PXN_DISABLE="${NCCL_PXN_DISABLE:-0}"
export NCCL_NET_DISABLE_INTRA="${NCCL_NET_DISABLE_INTRA:-1}"
export NCCL_IB_TC="${NCCL_IB_TC:-104}"
export NCCL_IB_FIFO_TC="${NCCL_IB_FIFO_TC:-192}"
export NCCL_IB_QPS_PER_CONNECTION="${NCCL_IB_QPS_PER_CONNECTION:-1}"
export NCCL_IB_TIMEOUT="${NCCL_IB_TIMEOUT:-22}"
export NCCL_IB_RETRY_CNT="${NCCL_IB_RETRY_CNT:-12}"

# MoRI uses the same NIC set as NCCL.
export MORI_RDMA_DEVICES="${MORI_RDMA_DEVICES:-${_IB_DEVICES}}"
export MORI_IB_GID_INDEX="${MORI_IB_GID_INDEX:-${_IB_GID_INDEX}}"
export MORI_SHMEM_HEAP_SIZE="${MORI_SHMEM_HEAP_SIZE:-16G}"

# Pin to gfx950 to avoid jit compilation failures with other archs on this cluster.
export MORI_GPU_ARCHS="gfx950"

# ----------------------------------------------------------------- benchmark
export BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS:-1024/128 2048/128}"
export BENCHMARK_CON="${BENCHMARK_CON:-32 64}"
export NUM_PROMPTS_FACTOR="${NUM_PROMPTS_FACTOR:-2}"
export BENCHMARK_MIN_PROMPTS="${BENCHMARK_MIN_PROMPTS:-32}"

# ----------------------------------------------------------------- accuracy
# `accuracy` role runs lm_eval (local-completions backend) against the proxy.
export ACCURACY_TASKS="${ACCURACY_TASKS:-gsm8k}"
export ACCURACY_NUM_CONCURRENT="${ACCURACY_NUM_CONCURRENT:-64}"
export ACCURACY_MAX_RETRIES="${ACCURACY_MAX_RETRIES:-3}"
export ACCURACY_METRIC="${ACCURACY_METRIC:-exact_match}"
export ACCURACY_THRESHOLD="${ACCURACY_THRESHOLD:-0.90}"

# ----------------------------------------------------------------- SLURM (submit)
# Used by run-slurm-disagg-test.sh on the login node (harmless to export here). [site]
export SLURM_PARTITION="${SLURM_PARTITION:-default}"
97 changes: 97 additions & 0 deletions .buildkite/amd-disagg/models.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
# =============================================================================
# Model catalog for the vLLM disaggregated (P/D) inference CI.
# -----------------------------------------------------------------------------
# Loaded by vllm_disagg.sh for the selected MODEL_NAME. WIDE_EP_MODE selects the
# flag set (0 -> tp keys, 1 -> ep keys):
# tp -> base_flags + prefill.tp / decode.tp
# ep -> base_flags + prefill.ep / decode.ep
#
# Keep ONLY model-specific perf flags here. The launcher owns everything
# topological / mode-structural:
# - tp: --host/--port, --tensor-parallel-size, --kv-transfer-config
# - ep: --tp 1, --data-parallel-size[-local], --data-parallel-address/-rpc-port,
# --data-parallel-start-rank/--headless (children), --enable-expert-parallel,
# --all2all-backend mori, --no-enable-prefix-caching,
# --api-server-count + --kv-transfer-config (masters only)
# Do NOT put any of those in this file.
#
# Format: a top-level `models:` list. The launcher selects the entry whose
# `model:` matches MODEL_NAME. Each entry:
# model : name (must match MODEL_NAME; also the dir under MODEL_DIR)
# env : model/arch-specific environment variables. The launcher
# exports these BEFORE `vllm serve`, but only if they are
# not already set, so precedence is:
# caller/inline env > models.yaml env:
# (Transport/fabric/engine env stays in cluster.sh,
# e.g. MORI_RDMA_DEVICES, NCCL_IB_HCA, VLLM_USE_V1.)
# base_flags : always applied (both roles, both modes)
# prefill/decode.<tp|ep> : role + mode specific perf flags
# experimental_flags : optional extra flags (omit if empty)
# =============================================================================

models:
- model: DeepSeek-V3
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MLA: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS: "0"
base_flags: "--trust-remote-code --kv-cache-dtype fp8"
prefill:
tp: "--gpu-memory-utilization 0.85"
ep: "--gpu-memory-utilization 0.85 --enforce-eager"
decode:
tp: "--gpu-memory-utilization 0.85"
#ep: '--gpu-memory-utilization 0.75 --enforce-eager'
ep: '--gpu-memory-utilization 0.75 --compilation-config {"cudagraph_mode":"PIECEWISE","custom_ops":["+quant_fp8"]}'

- model: MiniMax-M3-MXFP8
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
VLLM_USE_BREAKABLE_CUDAGRAPH: "0"
VLLM_ROCM_QUICK_REDUCE_QUANTIZATION: "INT6"
base_flags: "--trust-remote-code --attention-backend TRITON_ATTN --block-size 128 --language-model-only --kv-cache-dtype fp8"
prefill:
tp: "--gpu-memory-utilization 0.85 --enforce-eager"
decode:
tp: "--gpu-memory-utilization 0.85"

- model: DeepSeek-R1-MXFP4
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MLA: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
base_flags: "--trust-remote-code --kv-cache-dtype fp8"
prefill:
tp: "--gpu-memory-utilization 0.85"
decode:
tp: "--gpu-memory-utilization 0.85"

- model: Kimi-K2.5-MXFP4
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
base_flags: "--trust-remote-code"
prefill:
tp: "--gpu-memory-utilization 0.85"
decode:
tp: "--gpu-memory-utilization 0.85"

- model: Kimi-K2.6-MXFP4
env:
VLLM_ROCM_USE_AITER: "1"
AMDGCN_USE_BUFFER_OPS: "1"
VLLM_ROCM_USE_AITER_MLA: "1"
VLLM_ROCM_QUICK_REDUCE_QUANTIZATION: "INT4"
VLLM_ROCM_USE_SKINNY_GEMM: "0"
VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS: "1"
base_flags: "--trust-remote-code --kv-cache-dtype fp8 --mm-encoder-tp-mode data --block-size 1 --attention-backend ROCM_AITER_MLA"
prefill:
tp: "--gpu-memory-utilization 0.9"
decode:
tp: "--gpu-memory-utilization 0.9"
Loading