Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1526 commits
Select commit Hold shift + click to select a range
19d61b1
[Bugfix] Initialize DeepGemmQuantScaleFMT oracle lazily; bound QuantF…
BabyDrangoner Aug 12, 2026
fe889ac
[K3 Perf] Flash kda out kernel for prefill, 1.1~1.4x kernel performan…
yewentao256 Aug 12, 2026
7f7a32c
[Spec Decode] DSpark confidence-scheduled verification (#47808)
LucasWilkinson Aug 12, 2026
e62abc3
[Bugfix][Kernel] Fix persistent top-k histogram reuse after short row…
fxfxfxfxfxfxfxfx Aug 12, 2026
34735ac
[Bugfix] Pin DeepEP by its full commit hash (#52028)
tlrmchlsmth Aug 12, 2026
025d56a
[Build] Update DeepGEMM pin to deepseek-ai nv_dev tip (#52035)
zyongye Aug 12, 2026
23f360e
[CI Bug] Fix ci moe test (#52009)
yewentao256 Aug 12, 2026
98f86b9
[ROCm] [bugfix] Chunked prefill paged decode masked load perf (#50017)
afriedri Aug 12, 2026
caf9e8f
[CI] Force source builds for hybrid dependencies (#52043)
AndreasKaratzas Aug 13, 2026
5936bac
[Bugfix] Report FULL_ATTENTION for uniform-base UniformTypeKVCacheSpe…
yifjiang Aug 13, 2026
b369f10
[Bugfix][ROCm][CI] Restore the DeepSeek-V4 input GEMM override point …
stefankoncarevic Aug 13, 2026
8eb35c5
[CI] Mirror external test assets in vLLM S3 (#52064)
khluu Aug 13, 2026
61826c1
[Hardware][Power] Unqualized MoE Backend for Power (VSX) (#51624)
Akashcodes732 Aug 13, 2026
6e502b6
fix and test EPLB balancedness calculation (#51813)
jdebache Aug 13, 2026
f9a0f62
[KV Offload] Expose data-parallel topology to offloading backends (#5…
ziqifan617 Aug 13, 2026
15227b9
[Bugfix][KV Offload] Emit self-describing CPU events at KV-group bloc…
ziqifan617 Aug 13, 2026
9a276d6
[Mypy Fix] Mypy fix for "vllm/model_executor/models/[cC][dD]" (#52003)
yewentao256 Aug 13, 2026
f9538af
[Core] Clearer comments in `BlockPool.free_blocks()` (#52076)
njhill Aug 13, 2026
79f3183
[Bugfix] Bound KV block zeroing launch geometry (#52058)
LucasWilkinson Aug 13, 2026
2ac1f68
[BugFix] Reserve the bonus query slot in DFlash scheduling budget (#5…
HF-001 Aug 13, 2026
3d204df
Revert "[Perf][ROCm] Dual-stream decode with hipgraphs" (#52024)
simondanielsson Aug 13, 2026
6334491
[XPU] update UMD to 26.27 (#50513)
yma11 Aug 13, 2026
903d2ef
[Attention][MLA] Fuse Kimi-K3 chunked-context K/V packing (#51772)
zyongye Aug 13, 2026
5658391
Remove NIXL reinstall step (#51882)
ovidiusm Aug 13, 2026
b908a21
[Model][LoRA] Add tower/connector LoRA support for Ultravox (#48215)
arthurgao2003 Aug 13, 2026
7bc1660
[Misc] Use VLLMValidationError in pooling input validation (#51931)
frank-suwen Aug 13, 2026
89c8401
[Model] Skip unused Jina V5 output layers (#52037)
BabyDrangoner Aug 13, 2026
373592e
[CPU] Ship triton-cpu wheel and fix several hardcoded pin_memory=True…
bigPYJ1151 Aug 13, 2026
399f974
[Bugfix][R3] Size monolithic routing replay buffer for DP (#50874)
TomerBN-Nvidia Aug 13, 2026
7bbbf7c
[Core] Configure custom encoder cache managers from VllmConfig (#51251)
hotTea123 Aug 13, 2026
10bcad2
Update CODEOWNERS (#52123)
DarkLight1337 Aug 13, 2026
50ba4bc
[Feature] Mask Replay (#49577)
vx120 Aug 13, 2026
f962616
[ROCm][Perf] Kimi-K3 Remove prefill pipeline stall in chunk KDA (#51862)
kliuae Aug 13, 2026
1c3633a
[CI/Build][CPU] Shrink triton-cpu-build layer by dropping build artif…
bigPYJ1151 Aug 13, 2026
903da60
[Docs] Fix `WhisperEncoderLayer.forward` docstring in `dots3_note` (#…
hmellor Aug 13, 2026
b8baa31
Hardware-agnostic model definition via HF transformer backend (1/N) (…
bohnstingl Aug 13, 2026
37c3bdf
fix(security): enforce audio decode duration limit in NanoNemotronVL …
jperezdealgaba Aug 13, 2026
443fa5a
[Doc] Fix stale rejection_sample_method and synthetic_acceptance_rate…
qwerqwerqwe8688-jpg Aug 13, 2026
5fee0a8
chore: Upstream Cohere parser fixes + tests (#51998)
jasonozuzu-cohere Aug 13, 2026
d0ae25e
[Bugfix] Preserve Anthropic disable_parallel_tool_use (#52021)
taneem-ibrahim Aug 13, 2026
c4e9692
[XPU] Add tuned Mamba SSU configs for Intel Arc Pro B70 (#50534)
pmanczak Aug 13, 2026
015660d
[Misc] Add missing return type annotations in outputs.py (#52145)
vineetatiwari27 Aug 13, 2026
152c913
[Frontend] Log output token IDs at DEBUG level (#52098)
ruirui6946 Aug 13, 2026
8e1131e
[Model] [Quantization] Add Ling hybrid MXFP4 routed experts support (…
zexplorerhj Aug 13, 2026
95c9144
[LoRA][Gemma4] Support vision tower LoRA (#42662)
linitra24 Aug 13, 2026
170592a
[Bugfix] Disable sequence parallelism for Dots3 NOTE (#52172)
KurodaKanbei Aug 13, 2026
2d24355
[Bugfix] Fix packed GDN decode launch for large batch-head grids (#52…
mgoin Aug 13, 2026
f3c1638
[ROCm] Enable V2 model runner for Kimi-K3 on ROCm (#51653)
vllmellm Aug 13, 2026
7553aac
[Frontend] Add routed-experts prompt offset (#51906)
aoshen02 Aug 13, 2026
80d6d55
Standardise weight tying on `ParallelLMHead.tie_weights` (#52147)
hmellor Aug 13, 2026
64ca614
[Bugfix] Fix `--data-parallel-start-rank 0` being treated as unset in…
syedalijaseem Aug 13, 2026
96acd47
[Bugfix][MiniCPM-V] Fix AssertionError in get_dummy_mm_data when pass…
mayuyuace Aug 13, 2026
6014f9e
[Kimi-K3] Add GEMM-RS for sequence parallelism (#52079)
gau-nernst Aug 13, 2026
b245d8e
Apply logit softcapping in Transformers modelling backend (#52173)
hmellor Aug 13, 2026
83d4c61
[Bugfix] Declare SupportsEagle3 on KimiLinearForCausalLM (#52171)
nickus Aug 13, 2026
69d4c3a
Auto-ping Cohere on related issues (#52091)
DarkLight1337 Aug 13, 2026
b96bcd0
[ROCm][CI] Solidify entrypoint LLM lifecycle (#51280)
AndreasKaratzas Aug 13, 2026
48825ac
[Quantization] Remove dead `QuantizationConfig.is_mxfp4_quant` (#51793)
fxmarty-amd Aug 13, 2026
e6b2a8a
[Bugfix][Structured Output] Mask request stop tokens in xgrammar unti…
yzong-rh Aug 13, 2026
c5b7c06
[ROCm] Defer `tilelang` import through its import `from vllm.tilelang…
fxmarty-amd Aug 13, 2026
11c3fa4
[Bugfix][ROCm][CI] Give the AITER MLA decode metadata stub its MLA di…
stefankoncarevic Aug 13, 2026
73b8394
[Platform] Add check_runner_kv_caches_multi_layer (#51633)
wangxiyuan Aug 13, 2026
2e2ffd1
[CI Failure] Fix CUDA wheel build for the Kimi K3 fused MLA kernel (#…
mgoin Aug 13, 2026
b652ded
[Attention] Fix FlashInfer SM12x prefill with sinks (#52148)
askliar Aug 13, 2026
6355051
[Bugfix] Correct prompt lengths for timed_traces benchmark (#45423)
s3woz Aug 13, 2026
51def78
[Bugfix] Reapply 50869 (#52223)
benchislett Aug 13, 2026
71b0da7
[Bugfix] Fix .../mrope.py::apply_interleaved_rope() when torch.compil…
bastefaniak Aug 13, 2026
f80b66f
[Model Runner V2][Spec Decode] Add KV cache support for multi-layer M…
TheEpicDolphin Aug 13, 2026
38f097f
[Bugfix] Reject NUL byte in structured_outputs.regex (#51796)
ECMGit Aug 14, 2026
827a2af
[Kernel] Gemma-4 FA4 FP8 Kernel (#48666)
jhaotingc Aug 14, 2026
b216db3
[Bugfix] Reject negative token ids as out-of-vocabulary (#51795)
ECMGit Aug 14, 2026
fe4c5dc
[XPU] [Bugfix] process ragged weights in xpu linear backend (#52118)
zufangzhu Aug 14, 2026
3c79b1a
[Kernel] Add B12X dense linear backends (#52016)
lukealonso Aug 14, 2026
d18bb7b
[UT] fix device of test_outputs.py (#52237)
mayuyuace Aug 14, 2026
1be3628
[Kernel][Perf] Add fused CUDA post-conv MTP decode kernel for Qwen3.5…
Jie-Fang Aug 14, 2026
c05d75a
[Bugfix][Refactor] Keep Qwen3Next layer boundaries sequence parallel …
kzwrime Aug 14, 2026
653cc6f
[Bugfix][NIXL] Include transfer mode (push/pull) in the compatibility…
tzulingk Aug 14, 2026
59d1af5
Detect ROCm wheel variant from environment for precompiled wheels. (#…
aarushjain29 Aug 14, 2026
6adad08
Add Muse Glimmer model support (#51655)
xianbaoqian Aug 14, 2026
ac7509e
[Bugfix][CPU][RISC-V] Fix build: make FP32Vec copy constructors non-e…
velonica0 Aug 14, 2026
8e6d8e4
[XPU][CI/Release][3/N] Add xpu wheel release to release pipeline (#52…
jikunshang Aug 14, 2026
bda4c3e
[CPU] Fold the MXFP4 block scale in 2 instructions instead of 4 (#51583)
ccaadaro Aug 14, 2026
d4c24e6
[CI] Increase extended generation test timeout (#52252)
LucasWilkinson Aug 14, 2026
3c8676a
[PP][XPU]Overlap async-scheduling PP sampled-token broadcast with com…
yisustc Aug 14, 2026
624999a
[XPU]bump up vllm_xpu_kernels to 0.1.13.2 (#52138)
jikunshang Aug 14, 2026
103c419
[Perf][Frontend] Vectorize Cohere binary embedding bit-packing (#52277)
fangchenli Aug 14, 2026
b8165e5
[Frontend] Consolidate entrypoint exception handler (#52261)
noooop Aug 14, 2026
66728fe
[MRV2][Multimodal] Enable encoder cuda graph for model runner v2 (#49…
Isotr0py Aug 14, 2026
20405bf
[Bugfix] Fix Cosmos3-Edge processor after transformers 5.15 release (…
bastefaniak Aug 14, 2026
69e0e58
[Doc] Update model support information (#52289)
jeejeelee Aug 14, 2026
aa31003
[Bugfix][Helm] Fix chart resource references (#51664)
iwannagotobed Aug 14, 2026
63a9a50
[Attention][DSA] Take the native decode path for MTP=3 on SM90 (#52164)
zobinHuang Aug 14, 2026
57bd0ed
[5/N][KV-Cache Layout Refactor] Backend-published KV packing via cust…
LucasWilkinson Aug 14, 2026
1f7427b
[UT][XPU] fix b12x UT (#52265)
mayuyuace Aug 14, 2026
cdc4824
[Misc] Remove `override_attention_dtype` (#48684)
wangxiyuan Aug 14, 2026
03a8d0b
[Model][Spec Decode] Tap the pre-norm AttnRes mixture as the Kimi K3 …
rchalamala Aug 14, 2026
e078a22
[Bugfix] Widen flashinfer.comm import guard so a failed import doesn'…
shanjiaz Aug 14, 2026
83ded8d
[Test][LoRA] Speed up the LoRA test job (#52331)
stefankoncarevic Aug 14, 2026
9b0ab5d
[ROCm][AMD] Enable preshuffled sparse indexing for 16-token blocks (#…
jamesETsmith Aug 14, 2026
3e3ceb1
[Perf] Avoid more GPU<->CPU syncs in multimodal encoders (#52369)
njhill Aug 14, 2026
f473870
[Rust Frontend][gRPC] Add RL lifecycle control (#51316)
Aug 14, 2026
c794754
[ROCm]Remove special-case SiTU support model-specific gating (#50597)
stacyroberts Aug 14, 2026
925ea7e
[CI][Test] Seed the DeepEP v2 MoE workers, not just the parent (#50589)
guanxingithub Aug 14, 2026
d87ef45
[CI] Shard Quantization job into 4 parallel shards (≤30 min target) (…
khluu Aug 14, 2026
549cef0
[CI] Shard extended pooling model tests (#52322)
khluu Aug 14, 2026
694db07
[CI] Shard multimodal extended generation 2 (#52323)
khluu Aug 14, 2026
81e81da
[CI] Shard MoE refactor B200 eval (#52327)
khluu Aug 14, 2026
bb4b448
[CI] Shard Humming H100 eval (#52326)
khluu Aug 14, 2026
9df9b0b
[ROCm]: Drop pybind11 from Dockerfile.rocm to prevent version mismatc…
Rohan138 Aug 14, 2026
d6f17f3
[MRV2] Support attention-free models (#52374)
njhill Aug 14, 2026
7b544ec
[Spec Decode][Perf] Fuse the MTP trailing all-reduce; local-argmax dr…
zhou9402 Aug 15, 2026
615d4cf
[Core] Check for GPU<->CPU syncs during CI (#43107)
njhill Aug 15, 2026
acb0f1d
[Bugfix][Spec Decode] DSpark: inherit the target's attention backend …
zyongye Aug 15, 2026
44fc57d
[ROCm][CI] Select CPU platform for native no-GPU jobs (#49515)
AndreasKaratzas Aug 15, 2026
4215646
[Rust Frontend][gRPC] Preserve skip_special_tokens decoding option (#…
biswapanda Aug 15, 2026
5cecfc0
[Bugfix] Fix modelscope usage (#52431)
DarkLight1337 Aug 15, 2026
ac2ae87
[Frontend] Support count_reasoning_tokens in the Streaming Parser En…
chaunceyjiang Aug 15, 2026
97388c4
[Bugfix] Make DSV4 sparse MLA work end-to-end for plain decode, MTP, …
lucifer1004 Aug 15, 2026
d480199
[CI/Build] Add warning for unsupported global PTX architecture reques…
shanewidanagama Aug 15, 2026
ed0f475
[Bugfix][Model] Kimi-K3 MegaMoE: pass situ_beta/situ_linear_beta to f…
UranusSeven Aug 15, 2026
c94cdd0
[Bugfix][Sampling] Clear empty side on thinking-budget asymmetric SWA…
hsusul Aug 15, 2026
fa9d67f
[EC Connector] Added Build Connector Worker Meta for EC Connector (#4…
omerpaz95 Aug 15, 2026
edd4c81
[Bugfix][DSv4] Revert adaptive C128A metadata packing (#51318)
tobymao Aug 16, 2026
6593754
[Bugfix][Spec Decode] Keep EAGLE cache registration on the partial-ha…
mispa-ms Aug 16, 2026
8efa13b
[Bugfix] Pick the DeepSeek V4 eager cudagraph region per model runner…
njhill Aug 16, 2026
41f12a0
[Bugfix] Raise `VLLMValidationError` from structured output validator…
jeffreywang88 Aug 16, 2026
70aaec8
[Bugfix][Anthropic] Return 4xx for client-caused errors in /v1/messag…
SayHelloToWorld Aug 16, 2026
1b079c4
[Bugfix][Model Runner V2][Spec Decode] Fix off-by-one in bad_words dr…
jyan-R Aug 16, 2026
84530eb
[Bugfix][Multimodal] Keep Gemma 4 video frame counts on CPU (#52441)
chaunceyjiang Aug 16, 2026
4d2a68d
[Bugfix][Spec Decode][Structured Output] DSpark: fix the grammar bitm…
oops-oom Aug 16, 2026
836aac9
[Perf][DSV4] Optimize sparse top-k metadata kernels for higher prefil…
chaunceyjiang Aug 16, 2026
fe1c317
[Bugfix][ROCm] Skip FP8 MLA prefill PS-metadata build for chunked-con…
shantipriya-amd Aug 16, 2026
9409f59
[Core] Add CuMemAllocator.discard() for tag-selective GPU memory rele…
andakai Aug 16, 2026
83f591d
[Perf][DSV4] Optimize global top-k index kernel with compile-time con…
chaunceyjiang Aug 16, 2026
6914d60
[ROCm][Perf] gfx942: use FlyDSL fp8 MQA logits kernel (ROCm/aiter#391…
akii96 Aug 16, 2026
7d7b6f2
[Refactor] Remove dead code for quantization (#52221)
yewentao256 Aug 16, 2026
1f0e0bf
[Bugfix][Attention] Temporarily disable FA4 head-dim 256 (#52050)
taneem-ibrahim Aug 16, 2026
ef43e31
[ROCm][DSV4][Perf] Optimize Triton sparse-MLA decode on gfx950 (#52212)
Fangzhou-Ai Aug 16, 2026
fdab2b1
[ModelRunner v2] Support Transformers pooling model (#52425)
taneem-ibrahim Aug 16, 2026
6b0b850
[CI] Fit small KV-offload evals within shared memory (#52496)
taneem-ibrahim Aug 16, 2026
e3c1cb5
[CI/Build] Avoid duplicate runner startup for multimodal test (#52417)
Isotr0py Aug 16, 2026
eee538d
[Bugfix][V1][Multimodal] Ignore stale same-step encoder cache evictio…
gty111 Aug 16, 2026
dc9ae4b
[Bugfix][Mooncake] Reference GPU blocks for in-flight store jobs and …
chengy-sysu Aug 16, 2026
a18c9b5
[Kimi-K3][Perf] Update FlashKDA for automatic K2 V-split (#52458)
BabyDrangoner Aug 17, 2026
0ad04cf
[ROCm][CI] Enable ViT CUDA graph tests on AMD gfx950 GPUs (#52256)
shen-shanshan Aug 17, 2026
967e104
[Config] Unify indexer cache dtype under attention_config.indexer_kv_…
zyongye Aug 17, 2026
a6a2a93
[Bugfix][Frontend] Guard remaining before-validators against non-obje…
Kaif10 Aug 17, 2026
502af5e
[Doc] Add MatrixHub as a model loading source (#50492)
yitingdc Aug 17, 2026
6664d39
[Performance][MRV2] Cache logits-processing request state (#52329)
positive666 Aug 17, 2026
292187d
[Bugfix][DSv4] Keep indexer scoring in breakable graphs (#52492)
LucasWilkinson Aug 17, 2026
7ea4b40
[Hardware][NVIDIA] Add GB10 fused-MoE fp8 tuning configs (E=256, E=51…
pavelzak Aug 17, 2026
71b578b
[ROCm][CI] Use the same-build wheel in Python-only CI (#49514)
AndreasKaratzas Aug 17, 2026
311b351
[ROCm][CI] Avoid forcing FlashAttention in the ColPali pooling test (…
AndreasKaratzas Aug 17, 2026
53e211d
[CI/Build] Reduce more duplicate runner startup in tests (#52570)
Isotr0py Aug 17, 2026
93550cc
[Frontend] Consolidate entrypoint middleware (#52309)
noooop Aug 17, 2026
a02cfcc
[Bugfix][Mamba] Fix overlapping state copy race (#50729)
AndreasKaratzas Aug 17, 2026
5fd7a88
[CI/Build] Fix accident pre-commit breakage due to concurrent merge (…
Isotr0py Aug 17, 2026
0ff370b
docs: fix incorrect --custom-skip-chat-template flag reference (#52588)
theamalsebastian Aug 17, 2026
c05d923
[Doc] [ROCm] Update installation documentation (#52303)
tjtanaa Aug 17, 2026
bb23362
[CI] Shard Humming A100 eval (#52325)
khluu Aug 17, 2026
cc7cf71
[XPU] Enable Kimi K3 KDA kernel tests on XPU (#51809)
pmanczak Aug 17, 2026
95901ce
fix(pooling): validate BGE-M3 combined task ownership (#51823)
030611 Aug 17, 2026
f27ae25
[Bugfix][CPU] Take an attention group's query head count from its lay…
ganeshr10 Aug 17, 2026
70afded
[K3] support recoverssm for K3 (#51855)
ZJY0516 Aug 17, 2026
1d3a8b9
[ROCm][Bugfix] Fix Triton W4A16 bug in determining if transpose is re…
qli88 Aug 17, 2026
017e9f4
Promote `prefix_cache_retention_interval` to an argument and change t…
tlrmchlsmth Aug 17, 2026
4ab5e50
[Refactor] Simplify B12X linear kernels and warmup (#52368)
mgoin Aug 17, 2026
7075dda
Support DSpark configs with `architectures=DSparkDraftModel` + `model…
mgoin Aug 17, 2026
49905ad
[3/N][Feat][Perf] Add new warmup infrastructure for JITs. Add provide…
LopezCastroRoberto Aug 17, 2026
cfbc5af
[BugFix] lora_base_layer / routed_experts order in expert param mappi…
HollowMan6 Aug 17, 2026
ceb340e
fix: prevent PyNvVideoCodec decoder slot limit bypass via ClassVar sh…
jperezdealgaba Aug 17, 2026
402547d
[Bugfix][CI] Release the shared ColBERT engine before `test_colbert_h…
stefankoncarevic Aug 17, 2026
75dde08
[Perf][MoE] Optimize deepep_v2 receiver CPU Overhead (#51114)
LucasWilkinson Aug 17, 2026
c1e4387
[ROCm][CI] Restore Torch defaults and type DSV4 scratch buffers (#52566)
AndreasKaratzas Aug 17, 2026
3fc2893
[Bugfix] Account for local DP workers in startup thread allocation (#…
cr-zhao Aug 17, 2026
9633933
Relax CuPy constraint to only exclude 14.1.0 (#44284)
khluu Aug 17, 2026
d1e3eee
[Spec decode] Support Kimi-K3 DCP with DSpark (#52188)
wzhao18 Aug 17, 2026
f08a95f
[Rust Frontend][gRPC] Advertise LoRA capabilities (#52031)
Aug 17, 2026
8878ebd
[ROCm][CI] Expand AITER W4A4 MoE Coverage (#52647)
micah-wil Aug 17, 2026
e68fb75
[ROCm][AMD][Installation] add LMCache kv-connector installation and …
hongxiayang Aug 17, 2026
455edc0
[ModelRunnerV2] Support prompt embeds (#42963)
gcanlin Aug 17, 2026
60c3a31
[CI][AMD] Improve Kubernetes failure diagnostics (#52264)
AndreasKaratzas Aug 17, 2026
5ae2d38
[Perf][Structured Output] Skip unused request-local reasoners (#52573)
BugenZhao Aug 17, 2026
49fb2ee
[ROCm][CI] add Aiter ops tests (#52208)
divakar-amd Aug 17, 2026
c296cf8
[Bugfix] Add forward_xpu to XDRotaryEmbedding for HunyuanOCR on XPU (…
jbyczkow Aug 17, 2026
58aa1e3
[Bugfix][SM120][MLA] Disable dense prefill for FlashInfer sparse MLA …
tommy-asai-sonarsource Aug 17, 2026
0e8989b
[ROCm] gaurd on_gfx1250 call with rocm platform (#52625)
jikunshang Aug 18, 2026
0db502c
[Kimi-K3] support DCP partial prefix cache hit (#50493)
GirasoleY Aug 18, 2026
c296851
[MoE] Refine FlashInfer one-sided All2All integration (#51924)
bobboli Aug 18, 2026
cdb8545
[Kernel][Perf] Support Qwen head ratios in fused GDN MTP (#52539)
BabyDrangoner Aug 18, 2026
69d3335
[Bugfix][Quantization] Guard the MXFP8 FlashInfer path on FlashInfer …
LH-and-FPGA Aug 18, 2026
f4b161d
[XPU][UT] Fix OOM and skip graph case (#49287)
mayuyuace Aug 18, 2026
d5f5de7
[Bugfix] Fix DeepSeek V4 mHC broadcast buffer for weight sync (#52626)
HollowMan6 Aug 18, 2026
e0e5a7f
[Rust Frontend] Fix GLM-5.2 chat template rendering parity (#51426)
WoosukKwon Aug 18, 2026
101c447
[Bugfix] Handle DeepseekV4ForCausalLM in benchmark_moe get_model_para…
SayHelloToWorld Aug 18, 2026
aa99034
[Cohere] Misc changes to cohere model definitions (#50156)
kkt-cohere Aug 18, 2026
b0e9cff
[XPU] update xpu-manager to v2.1.0 (#52569)
yma11 Aug 18, 2026
c89d692
[Build] Propagate vLLM version to Rust binaries (#52593)
BugenZhao Aug 18, 2026
d785eb5
[Test] Add pause/resume E2E tests (#52144)
floatlibai Aug 18, 2026
5fa8ca9
[ROCm][CI] Move ROCm AITER quantization tests (#40938)
AndreasKaratzas Aug 18, 2026
2687fec
Replicated embedding and norm fusion for DSV3 flat model (#48484)
jeejeelee Aug 18, 2026
b01728b
[Bugfix] Return 4xx for client-caused errors in /detokenize (#52622)
rajathpi Aug 18, 2026
5c9ff53
[Bugfix] Accept logprobs=-1 in the Completion API (#46175)
he-yufeng Aug 18, 2026
e8ad285
[Bugfix][ROCm] Fix a few int4/int8 quantization errors (#52112)
qli88 Aug 18, 2026
eab1cff
Harden DeepSeek V3.2 fused kernel grids (#52381)
yimdev Aug 18, 2026
41f179b
[Rust Frontend] Simplify data-parallel size ownership (#52575)
BugenZhao Aug 18, 2026
be3f614
[CI] Register CPU CI "VLLM_CPU_CI_ENV" environment variable (#52633)
taneem-ibrahim Aug 18, 2026
d29dc3a
[Bugfix][Gemma4] Align parser enable_thinking default with template (…
lxy-alexander Aug 18, 2026
689be2b
Upgrade Flashinfer version to 0.6.17 (#52681)
wzhao18 Aug 18, 2026
3bb9c18
[Multimodal] Reorganize video decoder backends (#49155)
Isotr0py Aug 18, 2026
241ff8c
[Model] Enable LoRA support for tower and connector in LlavaNextForCo…
gangula-karthik Aug 18, 2026
f9f066d
[Bugfix][PaliGemma] Remove stale image embedding scaling (#52692)
ActiveSky Aug 18, 2026
88b2bff
[MOE] Standardize and abstract fused shared expert optimization selec…
fxmarty-amd Aug 18, 2026
bca7bea
Remove VLLM_TEST_FORCE_FP8_MARLIN to replace with linear_backend/moe_…
mgoin Aug 18, 2026
ddbf826
[ROCm] Gate Torch FP8 scaled-MM on architecture support (#51021)
sstamenk Aug 18, 2026
ad5e71b
[ROCm][Perf] Enable fused KDA decode on gfx942 (MI325X) (#52293)
mpashkovskii Aug 18, 2026
01e56ca
[Bugfix][MLA] Do not use Dense MHA for GLM-5.2 (#52512)
WoosukKwon Aug 18, 2026
d75136c
[Rust Frontend] Wait for all utility calls to finish (#52671)
Aug 18, 2026
90984dd
[CI] Upgrade huggingface-hub to 1.28.0 (#52797)
AndreasKaratzas Aug 18, 2026
6948a43
[Bugfix] Detect all attention-spelling variants in ModelConfig.is_hyb…
mganczarenko Aug 18, 2026
9842d70
[DBO][CI] Increase the coverage of prefill DBO in test_dbo.py (#48628)
SageMoore Aug 18, 2026
8d6b183
[CI] Standardize test job labels by device (#52659)
khluu Aug 18, 2026
6a391a9
[Rust Frontend][RL] add routed expert prompt offset (#52703)
biswapanda Aug 18, 2026
7ddb507
[Bugfix][DP] Don't assume the engines started when forwarding a wake …
aoshen02 Aug 18, 2026
12f64b3
[Bugfix][Structured Output] Stop XGrammar token batches at terminatio…
sfeng33 Aug 18, 2026
5d8a4cf
[ROCm] Pad non-aligned AITER MLA heads (#51647)
LiuYinfeng01 Aug 18, 2026
6066bb3
[ROCM][CI] Attention test speedup (#52763)
stefankoncarevic Aug 18, 2026
d3fafe0
[Quantization] Remove the dead ocp_mx_scheme branch from moe_kernel_q…
xuebwang-amd Aug 18, 2026
5f7a20b
[nv] add pcp support in dsv3.2 (#52046)
GirasoleY Aug 18, 2026
203926c
[ROCm][CI] Add AMD CI Pull-Request Commands (#52822)
AndreasKaratzas Aug 18, 2026
aa6abec
[CI][ROCm] Prevent Git maintenance races during shallow fetches (#52810)
AndreasKaratzas Aug 18, 2026
8f4a7f4
[ROCm][CI] Gating more ROCm tests (#44969)
AndreasKaratzas Aug 18, 2026
ef47a89
[Core] Make prefix-cache NONE_HASH deterministic by default (#51875)
russellb Aug 18, 2026
a9f4afb
[Bugfix] Fix DeepSeek V4 mHC broadcast buffer for dummy load (#51368)
HollowMan6 Aug 19, 2026
f1178f3
Revert DSv4 eager workspace reuse (#52836)
WoosukKwon Aug 19, 2026
b05ae5d
[CI][Bugfix] Complete DeepSeek-V4 FSE test fixture contract (#52842)
khluu Aug 19, 2026
d36cc42
[Bugfix][Elastic EP] Reject scale below the minimum data parallel siz…
Etelis Aug 19, 2026
a2257f9
[Elastic EP] Reduce eager-mode reconfiguration downtime (#51885)
itayalroy Aug 19, 2026
b1d9337
[EPD] Allow KV consumers to omit MM embeddings (#52697)
zhenwei-intel Aug 19, 2026
f485081
[XPU] Support EC connector KV Offloading on XPU (#49532)
chaojun-zhang Aug 19, 2026
e4d61d0
[CPU] Add AMX-only high-performance MLA backend for DeepSeek V2/V3/R1…
bigPYJ1151 Aug 19, 2026
f936a26
[Bugfix][XPU] Fix Mamba state pointer overflow (#48109)
Oxygen56 Aug 19, 2026
03b87dc
[XPU] upgrade requirements/test/xpu.txt (#52672)
mayuyuace Aug 19, 2026
aabc1a0
[Bugfix][Benchmark] Check readiness before tokenizer init in rust vll…
tlrmchlsmth Aug 19, 2026
deeeae7
[MM] Keep more metadata tensors on CPU (#52827)
njhill Aug 19, 2026
e575b5f
[XPU][CI] fix hf runner (#52730)
mayuyuace Aug 19, 2026
8e46acc
[Attention] Vectorize sparse MLA mask loads (#52217)
MatthewBonanni Aug 19, 2026
3130573
[Model] Add GraniteSWA and GraniteMoeSWA via existing Granite (#52706)
daviswer Aug 19, 2026
63ff748
[Attention][MLA] FlashMLA sparse: DCP on the fp8_ds_mla mixed-batch p…
drakosha Aug 19, 2026
993d3d3
scheduler: skip speculative decoding when all scheduled requests need…
malaiwah Aug 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
166 changes: 166 additions & 0 deletions .buildkite/amd-disagg/cluster.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,166 @@
#!/bin/bash
# =============================================================================
# cluster.sh — single, generic config for the vLLM disaggregated (P/D) launcher.
# -----------------------------------------------------------------------------
# This script is *sourced* by vllm_disagg.sh.
#
# Every value is `${VAR:-default}`, so the environment always wins:
# environment variable > built-in default below
#
# So you can override anything inline:
# PREFILL_IP=10.0.0.1 DECODE_IP=10.0.0.2 ./vllm_disagg.sh prefill
#
# Site-specific values (model dir, IPs, NIC list, partition) are the defaults
# in the "site defaults" sections below — edit those for a new cluster.
# Model-SPECIFIC perf flags live in models.yaml, NOT here.
# =============================================================================

_CLUSTER_SH_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"

# ----------------------------------------------------------------- model / mode
# MODEL_NAME indexes into models.yaml; MODEL_DIR is the parent dir holding it.
# MODEL_PATH (resolved by the launcher) = ${MODEL_DIR}/${MODEL_NAME}. [site]
export MODEL_NAME="${MODEL_NAME:-DeepSeek-V3}"
export MODEL_DIR="${MODEL_DIR:-/data/models2}"

# Shared NFS root (5 TB, visible on every node): model weights + per-run logs. [site]
export SHARED_MOUNT="${SHARED_MOUNT:-/data}"

# Parallelism mode (launcher derives PARALLEL_MODE tp|ep from this):
# WIDE_EP_MODE=0 tp : each node is an independent TP server (TP8, 1P1D)
# WIDE_EP_MODE=1 ep : data-parallel + expert-parallel across xP/yD nodes (wideep)
export WIDE_EP_MODE="${WIDE_EP_MODE:-0}"

# ----------------------------------------------------------------- topology
# xP prefill nodes + yD decode nodes. IPADDRS is the ordered, comma-separated
# node list (prefill IPs first, then decode IPs). NODE_RANK is this node's global
# 0-based rank; under SLURM it defaults to $SLURM_PROCID. Leave IPADDRS empty to
# use the PREFILL_IP/DECODE_IP fallback defaults below (1P1D only).
export xP="${xP:-1}"
export yD="${yD:-1}"
export IPADDRS="${IPADDRS:-}"
export NODE_RANK="${NODE_RANK:-${SLURM_PROCID:-}}"
export PREFILL_IP="${PREFILL_IP:-10.0.0.1}"
export DECODE_IP="${DECODE_IP:-10.0.0.2}"

# Per-node GPU count and TP degree
export GPUS_PER_NODE="${GPUS_PER_NODE:-8}"
export TP_SIZE="${TP_SIZE:-${GPUS_PER_NODE}}"

# ----------------------------------------------------------------- ports
# TP-mode (WIDE_EP_MODE=0) server ports.
export PREFILL_PORT="${PREFILL_PORT:-8100}"
export DECODE_PORT="${DECODE_PORT:-8200}"

# EP-mode (WIDE_EP_MODE=1) ports: API serve port, DP RPC port, KV transfer port, and
# per-node MoRIIO local ping port (must differ from PROXY_PING_PORT).
export SERVE_PORT="${SERVE_PORT:-20005}"
export RPC_PORT="${RPC_PORT:-13345}"
export KV_PORT="${KV_PORT:-9711}"
export LOCAL_PING_PORT="${LOCAL_PING_PORT:-61555}"

# MoRIIO proxy: HTTP port clients/benchmark hit, plus the connector control ports.
# PROXY_PING_PORT MUST be 36367 — the proxy hardcodes its zmq service-discovery
# socket on that port; prefill/decode register to PROXY_IP:PROXY_PING_PORT.
export PROXY_IP="${PROXY_IP:-${PREFILL_IP}}"
export PROXY_PORT="${PROXY_PORT:-10001}"
export PROXY_PING_PORT="${PROXY_PING_PORT:-36367}"
export HANDSHAKE_PORT="${HANDSHAKE_PORT:-6301}"
export NOTIFY_PORT="${NOTIFY_PORT:-61005}"

export PROXY_SCRIPT="${PROXY_SCRIPT:-/app/vllm/examples/disaggregated/disaggregated_serving/moriio_toy_proxy_server.py}"

# MoRIIO KV transfer direction (injected into --kv-transfer-config by the launcher):
# 0 -> omit read_mode (default; MoRIIO write mode: prefill pushes to decode)
# 1 -> "read_mode": true (decode pulls KV from prefill; matches upstream disagg)
export MORIIO_READ_MODE="${MORIIO_READ_MODE:-0}"

# ----------------------------------------------------------------- router / gateway
# Selection for client (bench/accuracy) traffic:
# proxy -> the in-container MoRIIO proxy started by the launcher
# vllm-router -> an external `vllm/vllm-router` container started by the SLURM job
# on the rank-0 node (default)
# Both use the SAME MoRIIO discovery mechanism (prefill/decode register to
# PROXY_IP:PROXY_PING_PORT=36367); only the client HTTP front door differs.
export ROUTER_TYPE="${ROUTER_TYPE:-vllm-router}"
export ROUTER_PORT="${ROUTER_PORT:-30000}"
export ROUTER_POLICY="${ROUTER_POLICY:-round_robin}"
export VLLM_ROUTER_IMAGE="${VLLM_ROUTER_IMAGE:-vllm/vllm-router:nightly}"
# Single client-facing port bench/accuracy target: the router port when routing,
# else the proxy port. Env override always wins.
if [[ "${ROUTER_TYPE}" == "vllm-router" ]]; then
export GATEWAY_PORT="${GATEWAY_PORT:-${ROUTER_PORT}}"
else
export GATEWAY_PORT="${GATEWAY_PORT:-${PROXY_PORT}}"
fi

# Where per-run logs / benchmark results are written. A $SLURM_JOB_ID subdir is
# appended so each CI run is self-scoped (falls back to 'local' off-SLURM). [site]
_LOG_BASE="${LOG_BASE:-/data/${USER:-$(id -un)}/disagg_logs}"
export LOG_PATH="${LOG_PATH:-${_LOG_BASE}/${SLURM_JOB_ID:-local}}"

# ----------------------------------------------------------------- vLLM runtime
# Engine/platform/transport-level env (NOT model-specific). Model-architecture
# AITER kernel toggles live in models.yaml under each model's `env:` block.
export VLLM_USE_V1="${VLLM_USE_V1:-1}"
export HSA_NO_SCRATCH_RECLAIM="${HSA_NO_SCRATCH_RECLAIM:-1}"

#export HF_HUB_OFFLINE="${HF_HUB_OFFLINE:-1}"
#export TRANSFORMERS_OFFLINE="${TRANSFORMERS_OFFLINE:-1}"
#
#export HOME=/tmp
export HF_HOME="${HF_HOME:-/tmp/hf_home}"
export XDG_CACHE_HOME="${XDG_CACHE_HOME:-/tmp/.cache}"
export VLLM_ENGINE_READY_TIMEOUT_S="${VLLM_ENGINE_READY_TIMEOUT_S:-3600}"

# ----------------------------------------------------------------- scale
# DP/EP group formation timeout across nodes (seconds).
export DISTRIBUTED_TIMEOUT_SECONDS="${DISTRIBUTED_TIMEOUT_SECONDS:-7200}"

# ----------------------------------------------------------------- RDMA / NCCL
# AMD Pensando AINIC RoCE fabric: 8 NICs exposed as ionic_0..7 (netdevs eth2..9),
# each rail on its own /24. GID index 1 + traffic class 104
_IB_DEVICES="${IB_DEVICES:-ionic_0,ionic_1,ionic_2,ionic_3,ionic_4,ionic_5,ionic_6,ionic_7}"
_IB_GID_INDEX="${NCCL_IB_GID_INDEX:-1}"
export IB_DEVICES="${IB_DEVICES:-${_IB_DEVICES}}"
export NCCL_IB_HCA="${NCCL_IB_HCA:-${_IB_DEVICES}}"
export NCCL_IB_GID_INDEX="${NCCL_IB_GID_INDEX:-${_IB_GID_INDEX}}"
export NCCL_IB_DISABLE="${NCCL_IB_DISABLE:-0}"
# In-box RCCL net transport (no external plugin); pin bootstrap to the VPC iface.
export NCCL_NET_PLUGIN="${NCCL_NET_PLUGIN:-none}"
export NCCL_SOCKET_IFNAME="${NCCL_SOCKET_IFNAME:-eth0}"
export GLOO_SOCKET_IFNAME="${GLOO_SOCKET_IFNAME:-${NCCL_SOCKET_IFNAME}}"
export NCCL_CROSS_NIC="${NCCL_CROSS_NIC:-0}"
export NCCL_PXN_DISABLE="${NCCL_PXN_DISABLE:-0}"
export NCCL_NET_DISABLE_INTRA="${NCCL_NET_DISABLE_INTRA:-1}"
export NCCL_IB_TC="${NCCL_IB_TC:-104}"
export NCCL_IB_FIFO_TC="${NCCL_IB_FIFO_TC:-192}"
export NCCL_IB_QPS_PER_CONNECTION="${NCCL_IB_QPS_PER_CONNECTION:-1}"
export NCCL_IB_TIMEOUT="${NCCL_IB_TIMEOUT:-22}"
export NCCL_IB_RETRY_CNT="${NCCL_IB_RETRY_CNT:-12}"

# MoRI uses the same NIC set as NCCL.
export MORI_RDMA_DEVICES="${MORI_RDMA_DEVICES:-${_IB_DEVICES}}"
export MORI_IB_GID_INDEX="${MORI_IB_GID_INDEX:-${_IB_GID_INDEX}}"
export MORI_SHMEM_HEAP_SIZE="${MORI_SHMEM_HEAP_SIZE:-16G}"

# Pin to gfx950 to avoid jit compilation failures with other archs on this cluster.
export MORI_GPU_ARCHS="gfx950"

# ----------------------------------------------------------------- benchmark
export BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS:-1024/128 2048/128}"
export BENCHMARK_CON="${BENCHMARK_CON:-32 64}"
export NUM_PROMPTS_FACTOR="${NUM_PROMPTS_FACTOR:-2}"
export BENCHMARK_MIN_PROMPTS="${BENCHMARK_MIN_PROMPTS:-32}"

# ----------------------------------------------------------------- accuracy
# `accuracy` role runs lm_eval (local-completions backend) against the proxy.
export ACCURACY_TASKS="${ACCURACY_TASKS:-gsm8k}"
export ACCURACY_NUM_CONCURRENT="${ACCURACY_NUM_CONCURRENT:-64}"
export ACCURACY_MAX_RETRIES="${ACCURACY_MAX_RETRIES:-3}"
export ACCURACY_METRIC="${ACCURACY_METRIC:-exact_match}"
export ACCURACY_THRESHOLD="${ACCURACY_THRESHOLD:-0.90}"

# ----------------------------------------------------------------- SLURM (submit)
# Used by run-slurm-disagg-test.sh on the login node (harmless to export here). [site]
export SLURM_PARTITION="${SLURM_PARTITION:-default}"
97 changes: 97 additions & 0 deletions .buildkite/amd-disagg/models.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
# =============================================================================
# Model catalog for the vLLM disaggregated (P/D) inference CI.
# -----------------------------------------------------------------------------
# Loaded by vllm_disagg.sh for the selected MODEL_NAME. WIDE_EP_MODE selects the
# flag set (0 -> tp keys, 1 -> ep keys):
# tp -> base_flags + prefill.tp / decode.tp
# ep -> base_flags + prefill.ep / decode.ep
#
# Keep ONLY model-specific perf flags here. The launcher owns everything
# topological / mode-structural:
# - tp: --host/--port, --tensor-parallel-size, --kv-transfer-config
# - ep: --tp 1, --data-parallel-size[-local], --data-parallel-address/-rpc-port,
# --data-parallel-start-rank/--headless (children), --enable-expert-parallel,
# --all2all-backend mori, --no-enable-prefix-caching,
# --api-server-count + --kv-transfer-config (masters only)
# Do NOT put any of those in this file.
#
# Format: a top-level `models:` list. The launcher selects the entry whose
# `model:` matches MODEL_NAME. Each entry:
# model : name (must match MODEL_NAME; also the dir under MODEL_DIR)
# env : model/arch-specific environment variables. The launcher
# exports these BEFORE `vllm serve`, but only if they are
# not already set, so precedence is:
# caller/inline env > models.yaml env:
# (Transport/fabric/engine env stays in cluster.sh,
# e.g. MORI_RDMA_DEVICES, NCCL_IB_HCA, VLLM_USE_V1.)
# base_flags : always applied (both roles, both modes)
# prefill/decode.<tp|ep> : role + mode specific perf flags
# experimental_flags : optional extra flags (omit if empty)
# =============================================================================

models:
- model: DeepSeek-V3
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MLA: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS: "0"
base_flags: "--trust-remote-code --kv-cache-dtype fp8"
prefill:
tp: "--gpu-memory-utilization 0.85"
ep: "--gpu-memory-utilization 0.85 --enforce-eager"
decode:
tp: "--gpu-memory-utilization 0.85"
#ep: '--gpu-memory-utilization 0.75 --enforce-eager'
ep: '--gpu-memory-utilization 0.75 --compilation-config {"cudagraph_mode":"PIECEWISE","custom_ops":["+quant_fp8"]}'

- model: MiniMax-M3-MXFP8
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
VLLM_USE_BREAKABLE_CUDAGRAPH: "0"
VLLM_ROCM_QUICK_REDUCE_QUANTIZATION: "INT6"
base_flags: "--trust-remote-code --attention-backend TRITON_ATTN --block-size 128 --language-model-only --kv-cache-dtype fp8"
prefill:
tp: "--gpu-memory-utilization 0.85 --enforce-eager"
decode:
tp: "--gpu-memory-utilization 0.85"

- model: DeepSeek-R1-MXFP4
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MLA: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
base_flags: "--trust-remote-code --kv-cache-dtype fp8"
prefill:
tp: "--gpu-memory-utilization 0.85"
decode:
tp: "--gpu-memory-utilization 0.85"

- model: Kimi-K2.5-MXFP4
env:
VLLM_ROCM_USE_AITER: "1"
VLLM_ROCM_USE_AITER_MOE: "1"
VLLM_ROCM_USE_AITER_RMSNORM: "1"
base_flags: "--trust-remote-code"
prefill:
tp: "--gpu-memory-utilization 0.85"
decode:
tp: "--gpu-memory-utilization 0.85"

- model: Kimi-K2.6-MXFP4
env:
VLLM_ROCM_USE_AITER: "1"
AMDGCN_USE_BUFFER_OPS: "1"
VLLM_ROCM_USE_AITER_MLA: "1"
VLLM_ROCM_QUICK_REDUCE_QUANTIZATION: "INT4"
VLLM_ROCM_USE_SKINNY_GEMM: "0"
VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS: "1"
base_flags: "--trust-remote-code --kv-cache-dtype fp8 --mm-encoder-tp-mode data --block-size 1 --attention-backend ROCM_AITER_MLA"
prefill:
tp: "--gpu-memory-utilization 0.9"
decode:
tp: "--gpu-memory-utilization 0.9"
Loading
Loading