Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
578 commits
Select commit Hold shift + click to select a range
247470f
[CI] Add PyTorch stable ABI audit check (#48164)
cleonard530 Jul 28, 2026
25ace8f
[CI] Increase Qwen3.5 MTP GSM8K generation length (#49881)
ZJY0516 Jul 28, 2026
bf9f230
[Rust Frontend] Fix finish reason for named tool choices (#49496)
reidliu41 Jul 28, 2026
912d6b6
[Rust Frontend] Align sampling validation with Python (#47494)
reidliu41 Jul 28, 2026
d2bfc6f
[Build] Fix DeepEP CUDA driver stub linking (#50103)
khluu Jul 28, 2026
35efdf6
[Elastic EP] Async preparation (#47288)
itayalroy Jul 28, 2026
948107a
[Bugfix] Enhance extra_config handling for layer name suffix matching…
xin3he Jul 28, 2026
98e91a9
[PD][NixlPush] Skip extra `add_remote_agent` step in D->P handshake (…
NickLucche Jul 28, 2026
9b9fc40
add epilogue hook to flex attention (#45841)
liangel-02 Jul 28, 2026
601fa9a
[KV Connector] Support NIXL heterogeneous P/D block sizes for hybrid …
njhill Jul 28, 2026
62d8db7
[Bugfix] Add missing `vllm/models/kimi_k3/__init__.py` (#50131)
hmellor Jul 28, 2026
94100b5
[CI] Wire untethered test files into CI jobs (#49340)
njhill Jul 28, 2026
b6cbba8
[Bugfix][Kernel] Fix batch invariance in RMSNorm kernels by pinning b…
oops-oom Jul 28, 2026
1e81853
[Bugfix][KV Offload] Keep Mamba block span unscaled under DCP (#49964)
jongukc Jul 28, 2026
0d0504b
[Core] Warm up runner-owned Triton kernels before the first request (…
njhill Jul 28, 2026
30217b0
[Bugfix][KV Offload][P2P] Scope serve state to fetch rounds (#49877)
Etelis Jul 28, 2026
4fb483c
[Docs] Expand llm-d integration page (#45432)
Ibrahim2595 Jul 28, 2026
6453fc0
[Bugfix] Don't reuse engine core payload buffer while zmq is sending …
njhill Jul 28, 2026
ba702e9
[Attention] Skip sparse indexer scoring for dense short prefills (#48…
qianlihuang Jul 28, 2026
01661cc
[Rust][Benchmark] Make `vllm bench serve` Rust delegation opt-in (#50…
BugenZhao Jul 28, 2026
4f56321
[ROCm] Cache fp32 upcast of static e8m0 weight scale in AITER scaled_…
jiacao-amd Jul 28, 2026
05a0814
[ROCm] Fix and optimize GPT-J-style MRoPE (#49906)
AndreasKaratzas Jul 28, 2026
8a7b3c2
[compressed-tensors] update `find_matched_target` order to prioritize…
brian-dellabetta Jul 28, 2026
1db989b
[Bugfix][Multimodal] Fix video temporal padding estimates (#49030)
labAxiaoming Jul 28, 2026
6c7e679
[ROCm][Bugfix] Sanitize AITER paged-MQA logits before sparse top-k fo…
shen-shanshan Jul 28, 2026
118bcde
[BugFix] Fix clang spinloop mwaitx include (#45532)
johnnyychiu Jul 28, 2026
d552a68
[Rust Frontend] Extract shared tracing setup logic into `vllm-tracing…
BugenZhao Jul 28, 2026
0b6aa3c
[Bugfix][Spec Decode] Size DFlash query buffers for cudagraph-padded …
siddhant-bharti Jul 28, 2026
2899dca
[Model] Add Kimi K3 support: Rust frontend [1/2] (#50104)
WoosukKwon Jul 28, 2026
bb3b61f
perf: dispatch non-grouped bias-less topk routing methods to fused pa…
jdebache Jul 28, 2026
a07fac7
[Perf] Zero-copy torch.Tensor pickling in shm_broadcast MessageQueue …
mrn3088 Jul 28, 2026
e7f6a39
[Test] Make EPD correctness tests configurable for XPU (#50110)
zhenwei-intel Jul 28, 2026
5369f7b
[MXFP8][ROCm] Fix MXFP8 MoE backend selection (#49747)
fxmarty-amd Jul 28, 2026
176256b
[ROCm][CI] Stabilize ROCm audio streaming test (#50163)
AndreasKaratzas Jul 28, 2026
fe65aa6
[CI][NIXL] Fix flaky DP+EP test port conflict (#50171)
divakar-amd Jul 29, 2026
56f31af
[Bugfix] Fix /wake_up crash on hybrid models (Mamba/DeltaNet) (#41602)
kevglynn Jul 29, 2026
6fbbcf2
[BugFix] Stop dummy runs from writing mamba state through stale block…
njhill Jul 29, 2026
7f4c52f
[CI] Add comment-based Buildkite triggers (#50132)
khluu Jul 29, 2026
54ab69b
[CI] Allow comment-triggered builds past pipeline filters (#50197)
khluu Jul 29, 2026
17a74b7
[Model] Add Inkling compressed-tensors dynamic FP8 support (#48876)
krishnateja95 Jul 29, 2026
7398a30
[ROCm][CI] Stabilize ngram and suffix correctness test (#50190)
AndreasKaratzas Jul 29, 2026
30c2718
[CompressedTensors] FP4 Qutlass Integration (#43229)
kylesayrs Jul 29, 2026
32a423a
Integrate CuTeDSL MoE for ReLU2 NVFP4 (#49580)
danielafrimi Jul 29, 2026
f37f03d
[KV Connector] Support NIXL P/D for hybrid MLA+SSM models (#49762)
njhill Jul 29, 2026
58f9659
[Frontend][Core] Standardize request error handling with VLLMError hi…
zqzten Jul 29, 2026
0bb548b
[CI][ROCm] Stabilize Qwen2-VL LoRA test (#50161)
AndreasKaratzas Jul 29, 2026
dc1be79
Add CachePolicyFactory for pluggable/external eviction policies (#49114)
philippesic Jul 29, 2026
db7a79c
[CPU] Fix s390x builds and update torch version in dockerfile (#50144)
R3hankhan123 Jul 29, 2026
6f91edf
[Test] dynamic_shapes_compilation (#49974)
JaredforReal Jul 29, 2026
6f00a1a
fused_moe: add VLLM_TRITON_USE_TD tensor-descriptor path (#42436)
afierka-intel Jul 29, 2026
7c6729b
[Model] Add Kimi K3 support: model files and kernels [1/N] (#50089)
ZJY0516 Jul 29, 2026
7de49ba
[XPU][UT][CI] add xpu config to run gpt-oss accuracy in ut and ci (#4…
zufangzhu Jul 29, 2026
100d655
[CI] Allow PR comment acknowledgements (#50211)
khluu Jul 29, 2026
65a1a16
[CPU] Fix FP8 attention scratchpad sizing (#50194)
tianmu-li Jul 29, 2026
6370e53
[Frontend] Reuse prefill token ids on the decode chat path for disagg…
eicherseiji Jul 29, 2026
f5a7cce
[Model] Add Kimi K3 support: Python frontend [2/2] (#50093)
BugenZhao Jul 29, 2026
ad5d29d
[Model] Support Qwen3.5 text-only dense and MoE models (#50210)
PerkzZheng Jul 29, 2026
32e657e
[BugFix] eagle draft max position embeddings (#49343)
JaredforReal Jul 29, 2026
df2735e
[Misc][Minimax-M3]add default video_processor (#50092)
lengrongfu Jul 29, 2026
5b29c95
[XPU] upgrade to torch 2.13 (#48677)
yma11 Jul 29, 2026
dad7a63
[EC Connector] Add has_pending_push_work (#49582)
omerpaz95 Jul 29, 2026
5b14019
[CI] Fix MXFP8 MOE backend selection tests on gfx942 (#50222)
fxmarty-amd Jul 29, 2026
c44e191
[Rust Frontend] Add --limit-mm-per-prompt support (#49604)
cinnamonica02 Jul 29, 2026
9a4e5f9
[CI/Perf] Fix malformed serving benchmark config (#43538)
fallintoplace Jul 29, 2026
542a8fa
[KV Offload] Move CPUOffloadingSpec onto SharedOffloadRegion (#50094)
Change72 Jul 29, 2026
aeaa50a
[Bugfix][Multimodal] Include media IO config in MM cache hash (#49975)
guan404ming Jul 29, 2026
f51193b
[Kernel][Mamba] Fused-kernel support for align-mode DS-conv state mig…
sungsooha Jul 29, 2026
72297d8
[XPU] Route weightless RMSNorm to _C dispatch (#47121)
yintong-lu Jul 29, 2026
625871b
[CI][Test] Fix pooling truncation test after VLLMError hierarchy chan…
stefankoncarevic Jul 29, 2026
242c591
[Rust Frontend] Send multimodal tensors in auxiliary frames (#49341)
reidliu41 Jul 29, 2026
e0cfa52
[Bugfix][Frontend] Return transcription and translation verbose as fl…
wskr00 Jul 29, 2026
43eaefb
[ModelRunner V2] Enable sequence pooling for embedding and classifica…
taneem-ibrahim Jul 29, 2026
a0c092e
[BugFix] Fix `num_output_placeholders` preemption underflow (#48245)
njhill Jul 29, 2026
d6247d7
[Spec Decode][Perf] Replicate DSpark Markov head across TP ranks (#49…
mgoin Jul 29, 2026
381b691
[ROCm][CI] Fix Kimi K3 KDA on ROCm (#50262)
stefankoncarevic Jul 29, 2026
8255369
[KV Connector] Fix NIXL mamba state pairing for multi-slot block tabl…
njhill Jul 29, 2026
f98061c
fix(step3p5-mtp): honor exclude_modules for the MTP head via prefix (…
chanh Jul 29, 2026
5c7a7f9
[docs] Add documentation for pynvvideocodec video decoding backend (#…
brandonpelfrey Jul 29, 2026
9347745
[torch.compile] Compile `CustomOp.forward_native` for ReLU^2 to avoid…
roikoren755 Jul 29, 2026
82642d7
[Perf] RMSNorm uncontiguous support, 1.2~3.1x kernel performance impr…
yewentao256 Jul 29, 2026
2ecd864
Revert "[Misc][Minimax-M3]add default video_processor (#50092)" (#50313)
njhill Jul 29, 2026
48fc2e2
feat(grpc): add KV event source discovery (#50033)
Jul 29, 2026
5fa0154
[Rubin] Enable NVLink all-reduce paths on SM107 (#49647)
zaristei Jul 29, 2026
435c4da
[ROCm][CI] Avoid Ray worker startup env race (#50311)
AndreasKaratzas Jul 29, 2026
fa2a258
[CI][ROCm] Fix AMD nightly distributed regressions (#50304)
AndreasKaratzas Jul 29, 2026
1cb3fe5
[CI Bugfix] Temp disable Humming wNa8 INT8 H100 CI (#50329)
mgoin Jul 29, 2026
9d1aa4d
[Frontend] Add diarized_json support for MOSS-Transcribe-Diarize (#48…
wskr00 Jul 29, 2026
48aa8d8
[Bugfix] Prevent stale multiproc RPC deadlines from becoming unbounde…
bugkeep Jul 29, 2026
b28c178
Fix: FusedMoE AssertionError with Speculative Decoding on Quark-Quant…
vecheruk-amd Jul 30, 2026
451227c
[Bugfix][Kernel] Fix integer overflow in libtorch_stable/activation_k…
molly-ting Jul 30, 2026
0a31372
[Frontend] Add detokenization streaming derender for disaggregated se…
hickeyma Jul 30, 2026
437e0b7
[BugFix] Fix P/D preemption race condition (#50297)
njhill Jul 30, 2026
4e582c5
[PD][Bugfix] Rebase KV lease deadlines onto worker clock (#50326)
njhill Jul 30, 2026
e5f48df
[Quantization][Autoround][XPU] Add W4A16(moe) / MXFP4(linear/moe) Sup…
lkk12014402 Jul 30, 2026
1ad5182
[CI/Build] Limit wheel size check to CUDA 13 (#50357)
tlrmchlsmth Jul 30, 2026
b889166
[ROCm] [CI] Support cached K/V (key/value=None) in Triton prefix-pref…
stefankoncarevic Jul 30, 2026
a7a204c
Add FlashMLA H100 tests to CI, fix them after #32810 (#50322)
janeyx99 Jul 30, 2026
0028fc8
[XPU][CI]Add back skipped V1 test (#50207)
zxd1997066 Jul 30, 2026
f1e8fd2
[ROCm] Add AITER FP8 ViT encoder attention (#49937)
LiuYinfeng01 Jul 30, 2026
8122a10
[DOC][CPU] remove tcmalloc warning from CPU docs (#50308)
fadara01 Jul 30, 2026
445d3aa
[CI] Retry failed steps on new PR commits (#50318)
khluu Jul 30, 2026
e04a30a
[Frontend] Lazily initialize chat media connectors (#49914)
AndreasKaratzas Jul 30, 2026
48a077e
[CI] Improve comment-triggered authorization and retries (#50414)
khluu Jul 30, 2026
165ed33
[CI][ROCm] Stabilize LLM GC teardown check (#50340)
AndreasKaratzas Jul 30, 2026
61c1d09
[CI] Stabilize speculator memory teardown (#50284)
AndreasKaratzas Jul 30, 2026
0c64be8
[Test][ROCm] Account for gfx950 FP8 RMSNorm rounding (#49839)
AndreasKaratzas Jul 30, 2026
aeeb36b
[New model] Kimi K3 (#50000)
ZJY0516 Jul 30, 2026
072a472
[CPU] Bump up CPU kernels to latest version (#50387)
bigPYJ1151 Jul 30, 2026
89d97d9
docs(security): document Ray cluster trust model and env var propagat…
jperezdealgaba Jul 30, 2026
38a267c
[MyPy][1/N] Fix mypy errors in some tests/ directories and enforce fo…
hickeyma Jul 30, 2026
e2efe79
[ROCm]Migrating Deepseek V3.2 to vllm/models/deepseek_v32/ (#47207)
stacyroberts Jul 30, 2026
1a20d23
[PARSER][Mistral] unified engine-based parser for reasoning and tool …
juliendenize Jul 30, 2026
59e831c
[Compilation]Fuse Transformers Residual Add + RMSNorm (#48757)
BadrBasowid Jul 30, 2026
30b4e7f
[rl] Stateful Trainer Send: IPC [2/N] (#48981)
hao-aaron Jul 30, 2026
5b95890
[FlexAttention] Avoid encoder block-mask compile explosion (#50339)
AndreasKaratzas Jul 30, 2026
f388dd6
[XPU][CI] skip kimi-k3 test (#50447)
jikunshang Jul 30, 2026
904fae8
[DSv4 Perf] Fix redundant memory allocation and copy for dsv4 pp buff…
yewentao256 Jul 30, 2026
837eae6
[DSv4 Perf] Remove redundant full kernel for dsv4, 1.88x kernel perfo…
yewentao256 Jul 30, 2026
61cacd2
[Bugfix][MoE] Write Humming results to the supplied output buffer (#5…
netanel-haber Jul 30, 2026
5f8f728
[Build] Fix CUDA release wheel builds (#50243)
khluu Jul 30, 2026
12a34a6
[ROCm][DSV4] B-preshuffle the attention fp8 projections (#46720)
cagrikymk Jul 30, 2026
629a938
[Frontend] Preserve bare Inkling text in Python and Rust parsers (#50…
BugenZhao Jul 30, 2026
30e333c
[Bugfix] Shut down private Tensorizer engines (#49840)
AndreasKaratzas Jul 30, 2026
bdc98bf
[CI] Initialize fused gated RMSNorm weights (#50377)
AndreasKaratzas Jul 30, 2026
3f90c7e
[ROCm] Pass pointers to FlyDSL MoE kernels (#50378)
AndreasKaratzas Jul 30, 2026
c27b080
[compile] Fix fake kernel return dtype (#50444)
zou3519 Jul 30, 2026
70bd109
[Quantization] Honor `--linear-backend` for ModelOpt W4A16 (#50273)
netanel-haber Jul 30, 2026
0eec856
Add Humming indexed-MoE regression test (#50468)
mgoin Jul 30, 2026
45b60e3
[Kernel][Helion] Disable unsafe B200 RMS reduction warp specializatio…
yushangdi Jul 30, 2026
e9096fc
[Rust Frontend] Improve startup failure and readiness logs (#50406)
BugenZhao Jul 30, 2026
68fb303
[Bugfix] Preserve Marlin runtime tensor storage across weight reload …
RyanClark2k Jul 30, 2026
7fe5312
[CI] Retire the v1 PR label rule, add mrv2 (#50475)
jcotant-inferact Jul 30, 2026
dec13a3
[Model Runner V2][Spec Decode] Add multi-layer MTP speculator (#48892)
TheEpicDolphin Jul 30, 2026
8700f86
[CI] Fix `tests/entrypoints/multimodal/openai/chat_completion/test_au…
NickLucche Jul 30, 2026
0bff0ce
[Kimi K3 Bug] Fix deepgemm support for kimi k3 (#50458)
yewentao256 Jul 31, 2026
3333d7c
[ROCm]: bump AITER to 0.1.19 (#49361)
Rohan138 Jul 31, 2026
d6938b7
[Bugfix][Rust Frontend] Select earliest-completing stop string (#50200)
samlaf Jul 31, 2026
553fcb8
[CI] Retry Hugging Face processor loading (#49908)
AndreasKaratzas Jul 31, 2026
4f1da84
Enable gfx1250 ROCm architecture (#46516)
jpvillam-amd Jul 31, 2026
f1899b2
[Bugfix][ROCm] AITER MLA: size MTP verification decode metadata for r…
chaeminlim-mb Jul 31, 2026
ab98034
[Frontend][Bugfix] Use default tool call IDs for Kimi K3 for conversa…
BugenZhao Jul 31, 2026
60399d4
[CI] Retry Buildkite API rate limits (#50481)
khluu Jul 31, 2026
d91f7af
[Hardware][AMD][Kernel][CI][Bugfix] Fix ROCm DeepEP FP8 max (#50467)
mawong-amd Jul 31, 2026
5d5f22e
[ROCm][CI] Use larger atol value for INT3 in test_quick_all_reduce.py…
music-dino Jul 31, 2026
4689c7d
[ROCm] Add tuned selective_state_update float16 config for AMD Instin…
vanshbhatia-amd Jul 31, 2026
6724051
[CI/Build][AMD] Install triton_kernels via CMake (#50328)
rjrock Jul 31, 2026
ef0d084
[XPU] Fix FP8 block scale layout for MLA compatibility (#50349)
majian4work Jul 31, 2026
1d8be5c
[XPU] [BugFix] Add deepseek_v4_fp8 to xpu supported_quantization list…
xwu-intel Jul 31, 2026
541128b
[KV Offload] Enable single-copy MLA layout for CPUOffloadingSpec (#50…
Change72 Jul 31, 2026
2773ec3
[ROCm][CI] Use explicit wvSplitKrc skinny-GEMM test tolerance for bf1…
stefankoncarevic Jul 31, 2026
b49eaf2
[DSv4] Remove sparse-MLA q-head padding for FlashInfer >=0.6.14 (#48047)
majunze2001 Jul 31, 2026
bebf918
[Bugfix][Model] Reject encoder-backbone jina-embeddings-v5 checkpoint…
woosebastian Jul 31, 2026
0f17394
[Model Runner V2] Enable encoder token classification (#50293)
taneem-ibrahim Jul 31, 2026
1180b60
[Multimodal] Expose mm hash algothrim selection to cli args (#49686)
Isotr0py Jul 31, 2026
0351e9a
[XPU][CI]Adjust source_file_dependencies for NixlConnector PD accurac…
zxd1997066 Jul 31, 2026
3ee2bd1
Fix duplicate HunyuanVL image boundary tokens (#49691)
Mi-Jiazhi Jul 31, 2026
10e6b40
[CPU][BugFix] Remove redundant kv cache write (#50437)
fadara01 Jul 31, 2026
f727951
[Bugfix] Re-land MiniMax M3 default video processor (#50305)
taneem-ibrahim Jul 31, 2026
34bb795
[CI] Add M3 MSA tests to CI (#49143)
gau-nernst Jul 31, 2026
5d7647a
[UT] add skipif for rocm aiter sampler UT (#50530)
mayuyuace Jul 31, 2026
482cfc2
[XPU] Unify XPU RMSNorm kernels with vllm_c and drop redundant XPU-sp…
chaojun-zhang Jul 31, 2026
88bc8fb
[CPU][s390x] Optimize inference perf and add oneDNN INT8 GEMM for s39…
R3hankhan123 Jul 31, 2026
0e9b500
[chore] clean-up weight prepack for INT8 MoE (#50116)
fadara01 Jul 31, 2026
c911120
[ROCm][CI] Update Transformers AR+RMS fusion expectation (#50517)
AndreasKaratzas Jul 31, 2026
6e311c6
[MoE Refactor] Rename FusedMoE to FusedMoEFactory (#44941)
bnellnm Jul 31, 2026
17beffd
[Misc] Clarify mono audio requirement (#50141)
NickLucche Jul 31, 2026
92643d6
K3 DSpark AR fusion (#50242)
jeejeelee Jul 31, 2026
b2fb83e
[Attention]: Use KVCacheSpec for AttentionMetadataBuilder type hints …
hickeyma Jul 31, 2026
2c4d348
[KV Connector] Add per-layer canonical KV page mappings for paralleli…
Etelis Jul 31, 2026
f5ffc59
[Renderer] Warm up the renderer properly. (#50408)
noooop Jul 31, 2026
0b5b49d
[Bugfix][Frontend] Raise VLLMValidationError for user-facing errors i…
latent-9 Jul 31, 2026
7fdc1ab
[CI] Remove default_torch_num_threads workaround from llava-onevision…
oguzhankir Jul 31, 2026
864a87f
[Kernel][CI] `--jit-monitor-mode error` e2e tests for kernel warmup i…
NickLucche Jul 31, 2026
82ae416
[2/N][Attention] Enable masked MHA for sparse MLA prefills (#48770)
MatthewBonanni Jul 31, 2026
5233368
[ROCm][CI] Fall back to lossless Kimi K3 MXFP4 emulation on gfx942 (#…
AndreasKaratzas Jul 31, 2026
df71917
[DSv4 Perf] Optimize workspace reuse for eager break, 3.9% E2E TTFT i…
yewentao256 Jul 31, 2026
94e9ef0
[Bugfix] Don't transpose fused MoE quantization scales in `RoutedExpe…
hmellor Jul 31, 2026
a0cd2b6
[Bugfix] Universally align block table width to 128 tokens (#50302)
MatthewBonanni Jul 31, 2026
c036cb2
[Doc] Add BgeM3EmbeddingModel to embedding supported models (#50571)
LG-0927 Jul 31, 2026
10ad649
Update torch version to 2.13.0+cpu (#50412)
ylangtsou Jul 31, 2026
7c08664
Upgrade tpu-inference to v0.26.0 (#50522)
meiyeh123 Jul 31, 2026
8d8a4e0
[ROCm][CI] Restore Mistral tool-parser compatibility after unificatio…
AndreasKaratzas Jul 31, 2026
d87d2ca
[Compressed-Tensors] Support Kimi-K3 quantized models (#50500)
kylesayrs Jul 31, 2026
e67a2e0
[ROCm][Quark][7/N] Use MXFP4 linear kernel abstraction for `emulation…
fxmarty-amd Jul 31, 2026
aef85ae
[Bugfix][TurboQuant] Add KV quant mode for turboquant (#50533)
skavulya Jul 31, 2026
963a658
Bump Helion to 1.4.0 (#50307)
yushangdi Jul 31, 2026
726ef43
[chore] delete useless code (#49424)
andyxning Jul 31, 2026
9a7ae4b
[chore] log process manager shutdown with more details (#49437)
andyxning Jul 31, 2026
4fdd3e0
[ROCm][ViT] Detect Triton-AMD kernels at their new aiter location (#4…
Lafunamor Jul 31, 2026
e8b358b
[Bugfix][Responses] Add tests for Chat Completions Responses API Rend…
yzong-rh Jul 31, 2026
9032991
[UX] Reduce startup log noise (#50590)
mgoin Jul 31, 2026
0c4dc7c
[graceful shutdown] fix http server start firstly before app signal h…
andyxning Jul 31, 2026
454ea5b
[MoE Refactor] Combine CompressedTensorsWNA16MarlinMoEMethod with Com…
bnellnm Jul 31, 2026
e3be896
Enable ModelOpt FP8 emulation on SM80 (#50019)
mikekg Jul 31, 2026
b40d859
[Kernel][Helion] Add numerics checks to benchmark script (#48968)
yushangdi Aug 1, 2026
fcdc7c2
[CI] Organize speculative decoding E2E tests by coverage (#50330)
mgoin Aug 1, 2026
c4a4a42
[Bugfix][Test] Fix monolithic routing replay test buffer capacity (#5…
Amir-19 Aug 1, 2026
eb6453d
[Build] Update pin to build ABI stable FA2 (#50474)
janeyx99 Aug 1, 2026
62195e9
[Rust][Benchmark] Prevent invalid token IDs in random benchmarks (#50…
reidliu41 Aug 1, 2026
f7097a9
[Bugfix][CI] Prevent common ops imports from initializing CUDA (#50639)
AndreasKaratzas Aug 1, 2026
124154a
Add @shen-shanshan to CODEOWNERS (#50655)
shen-shanshan Aug 1, 2026
6c91de3
[Bugfix][Parser] Forward model_config to nested reasoning parsers (#5…
chaunceyjiang Aug 1, 2026
9c110fa
[Frontend] Cohere chat v2 api support (#47189)
andrewbcohere Aug 1, 2026
4ee9702
Add Understanding the Latency Metrics docs (#50600)
mgoin Aug 1, 2026
77469c9
[ROCm][MLA] Mask the AITER MLA small-head verify flatten causally (#5…
yudigege86 Aug 1, 2026
81a42d3
[Frontend] Add cache_salt support to Anthropic Messages API (#49498)
aeon-x Aug 1, 2026
3986b96
(feat): optionally disable lookup on PD decode (#50498)
majunze2001 Aug 1, 2026
652ba59
[Model Runner V2] Enable encoder token embedding (#50574)
taneem-ibrahim Aug 1, 2026
ab06486
[Bugfix][Kernel] Fix dangling temporary in AWQ gemm torch::stable::su…
wentian-byte Aug 1, 2026
63e78ce
[Benchmark] Add probe requests to vllm bench serve (#49611)
guan404ming Aug 1, 2026
39f55ff
[Core] Offload raw-prompt preprocessing to renderer thread pool in As…
almogtavor Aug 1, 2026
03c782e
[model registry] some simple typos (#50673)
andyxning Aug 1, 2026
127a7ce
Add @hongxiayang as code owner for AMD-specific model files and ROCm …
hongxiayang Aug 1, 2026
dc818c1
[GPT-OSS] Strict tool call and constrained decoding for Harmony (#45560)
yzong-rh Aug 1, 2026
38a466e
[DSV4] Implement Sequence Parallelism (#46789)
WoosukKwon Aug 1, 2026
c67fe49
[Bugfix][Doc] Fix references to FusedMoE in doc (#50701)
bnellnm Aug 2, 2026
e2fa285
[1/N] Unify multiple-path encoder cuda graph support (#49934)
Isotr0py Aug 2, 2026
0601850
[Bugfix][Models] Accept Qwen3_5MoeTextConfig in Qwen3_5MoeProcessingI…
loulanyue Aug 2, 2026
96add73
[Elastic EP] Fix non-contiguous weight transfers (#50641)
itayalroy Aug 2, 2026
55c98e3
[Model Runner v2] Enable BGE M3 pooling embed token_classify (#50661)
taneem-ibrahim Aug 2, 2026
c666810
[ROCm][Bugfix][Kimi-K3] Preserve MoE correction bias in FP32 (#50761)
Fangzhou-Ai Aug 2, 2026
0055b8b
[Attention][MiniMax-M3] Add MSA speculative decode verification (#50032)
jasonlizhengjian Aug 2, 2026
0033211
[Test][V1] Add sleep/wake correctness regression test for hybrid GDN/…
chun-wan Aug 2, 2026
9a4fd57
[Model] Support jina-embeddings-v5-text-nano (EuroBERT encoder backbo…
omkar-droid Aug 3, 2026
e42c230
[Bugfix] serving_llama70B_tp4 benchmark was silently running at tenso…
wjabbour Aug 3, 2026
5c4fe4b
[Bugfix][KV Connector] Propagate EAGLE state across merged Mooncake s…
ivanium Aug 3, 2026
5e35a6f
cpu_model_runner.py: skip the warm up if CompilationMode.NONE (#50547)
yamt Aug 3, 2026
f5bb701
[Bugfix][Frontend] Constrain Anthropic cache_salt to non-empty (#50764)
omkar-droid Aug 3, 2026
539455e
[Bugfix][Build] Fix DeepGEMM CUDA 12.9 FP8 header visibility (#51003)
khluu Aug 4, 2026
14a4751
[Bugfix][CPU] Fix macOS build: std::sqrt is not constexpr under libc+…
harjothkhara Aug 4, 2026
160356c
[Kernel] Add support for Flashinfer Mamba SSU algorithm selection (#5…
amitz-nv Aug 4, 2026
b047571
[CI] Stabilize GLM-5.2 PCP evaluation (#51015)
khluu Aug 4, 2026
4b25076
[Bugfix][Humming] Preserve ModelOpt FP8 weight dimensions (#51093)
netanel-haber Aug 5, 2026
689fa52
Bump Flashinfer version to 0.6.16.post3 (#50892)
wzhao18 Aug 9, 2026
85ea133
[INC] fix w4a4 model (#50807)
mayuyuace Aug 3, 2026
5a76724
[Kimi-K3] Add option to shard the shared expert instead of replicatin…
tlrmchlsmth Aug 3, 2026
4efaa99
[BugFix][K3] Skip moe_intermediate padding when EP is enabled (#51131)
ZeldaHuang Aug 5, 2026
31ffc88
[Bugfix][Model] Add missing fused_qkv_a_proj to Kimi-Linear packed_mo…
JianDan0212 Aug 6, 2026
3f9fc83
[Bugfix] Keep mamba align prefill chunks block-aligned past last_cach…
ivanium Aug 6, 2026
d0d9691
Fix ROCm architecture import on non-ROCm platforms (#51357)
xwu-intel Aug 7, 2026
a60fb53
[CI] Fix Batch Invariance (B200) (#51417)
ZJY0516 Aug 7, 2026
4dbf890
[K3] Allow tpu to import kimi_k3.common (#51529)
majunze2001 Aug 9, 2026
525832b
[Model] Add K-EXAONE-2.0-750B-A37B (#50524)
lkm2835 Aug 3, 2026
e50f7d3
[Kimi][MM] disable kimi_vit's dynamic torch.compile for TPU (#51196)
lk-chen Aug 8, 2026
4bdc8a7
[Docs] Fix two docs build warnings (#51014)
hmellor Aug 4, 2026
7ce42a9
Support quantized DSpark Markov heads (#50424)
askliar Aug 3, 2026
9fc29ec
[CI] Parallelize release image publishing (#51735)
khluu Aug 10, 2026
26363ce
[CI] Isolate Ubuntu 22.04 test dependency cache
khluu Aug 11, 2026
6e448d0
[CI] Limit Arctic import check to x86 test images
khluu Aug 11, 2026
2265540
Merge tag 'v0.27.1' into update/v0.27.1
swjeong9 Aug 13, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
102 changes: 102 additions & 0 deletions .buildkite/check-torch-abi.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,102 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project

"""Audit vLLM compiled libraries for PyTorch stable ABI compliance."""

import fnmatch
import sys
from pathlib import Path

from torch_abi_audit import inspect_package
from torch_abi_audit.report import ExtensionReport, PackageReport

# Temporary allowlist of extensions not yet on the stable ABI.
# Shrink and remove over time.
ALLOWED_UNSTABLE_LIBRARIES: tuple[str, ...] = (
"_flashkda_C.abi3.so",
"vllm_flash_attn/_vllm_fa3_C.abi3.so",
"third_party/deep_gemm/_C*.so",
)


def _relative_path(lib: ExtensionReport, package_root: Path) -> str:
try:
return lib.path.relative_to(package_root).as_posix()
except ValueError:
return lib.path.name


def _is_torch_unstable(lib: ExtensionReport) -> bool:
return lib.error is None and lib.torch.uses_torch and not lib.torch.stable


def _matches_allowlist(rel_path: str, patterns: tuple[str, ...]) -> bool:
return any(fnmatch.fnmatch(rel_path, pattern) for pattern in patterns)


def _iter_libs(report: PackageReport) -> tuple[ExtensionReport, ...]:
return (*report.extensions, *report.bundled_libs)


def _collect_unstable(report: PackageReport) -> list[str]:
return sorted(
_relative_path(lib, report.root)
for lib in _iter_libs(report)
if _is_torch_unstable(lib)
)


def _find_stale_allowlist_entries(
report: PackageReport, patterns: tuple[str, ...]
) -> list[str]:
"""Allowlist patterns that match a built library which is no longer unstable."""
stale: list[str] = []
for pattern in patterns:
for lib in _iter_libs(report):
if lib.error is not None:
continue
if not fnmatch.fnmatch(_relative_path(lib, report.root), pattern):
continue
if not _is_torch_unstable(lib):
stale.append(pattern)
break
return stale


def check_torch_abi(
package: str = "vllm",
patterns: tuple[str, ...] = ALLOWED_UNSTABLE_LIBRARIES,
) -> int:
report = inspect_package(package)
if report.error:
print(f"error: failed to inspect {package!r}: {report.error}", file=sys.stderr)
return 2

unstable = _collect_unstable(report)
unexpected = [
rel_path for rel_path in unstable if not _matches_allowlist(rel_path, patterns)
]
stale = _find_stale_allowlist_entries(report, patterns)

if unexpected or stale:
if unexpected:
print(
"Not allowed: torch-unstable libraries outside "
f"ALLOWED_UNSTABLE_LIBRARIES: {', '.join(unexpected)}",
file=sys.stderr,
)
if stale:
print(
"Not allowed: stale ALLOWED_UNSTABLE_LIBRARIES entries: "
f"{', '.join(stale)}",
file=sys.stderr,
)
return 1

print("Torch stable ABI check passed.")
return 0


if __name__ == "__main__":
print(">>> Auditing vLLM extension modules for PyTorch stable ABI compliance")
sys.exit(check_torch_abi())
1 change: 1 addition & 0 deletions .buildkite/ci_config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ run_all_patterns:
- "setup.py"
- "csrc/"
- "cmake/"
- ".buildkite/check-torch-abi.py"
run_all_exclude_patterns:
- "docker/Dockerfile."
- "csrc/cpu/"
Expand Down
6 changes: 5 additions & 1 deletion .buildkite/hardware_tests/cpu.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@ steps:
- tests/kernels/quantization/test_cpu_fp8_scaled_mm.py
- tests/kernels/mamba/cpu/test_cpu_gdn_ops.py
- tests/kernels/mamba/test_cpu_short_conv.py
- tests/kernels/mamba/test_causal_conv1d.py
- tests/kernels/mamba/test_mamba_ssm.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 30m "
Expand All @@ -28,7 +30,9 @@ steps:
pytest -x -v -s tests/kernels/test_onednn.py
pytest -x -v -s tests/kernels/test_awq_int4_to_int8.py
pytest -x -v -s tests/kernels/quantization/test_cpu_fp8_scaled_mm.py
pytest -x -v -s tests/kernels/mamba/cpu/test_cpu_gdn_ops.py"
pytest -x -v -s tests/kernels/mamba/cpu/test_cpu_gdn_ops.py
pytest -x -v -s tests/kernels/mamba/test_causal_conv1d.py
pytest -x -v -s tests/kernels/mamba/test_mamba_ssm.py"

# Note: SDE can't be downloaded from CI host because of AWS WAF
# - label: CPU-Compatibility Tests
Expand Down
26 changes: 26 additions & 0 deletions .buildkite/intel_jobs/benchmarks_intel.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
group: Benchmarks
depends_on:
- image-build-xpu
steps:
- label: Benchmarks CLI Test
key: benchmarks-cli-test
timeout_in_minutes: 40
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
REPO: "vllm-ci-test-repo"
VLLM_TEST_DEVICE: "xpu"
source_file_dependencies:
- vllm/
- tests/benchmarks/
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
pytest -v -s benchmarks/'
82 changes: 81 additions & 1 deletion .buildkite/intel_jobs/engine_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,44 @@ group: Engine Intel
depends_on:
- image-build-xpu
steps:
- label: Engine
key: engine
timeout_in_minutes: 40
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
REPO: "vllm-ci-test-repo"
VLLM_TEST_DEVICE: "xpu"
source_file_dependencies:
- vllm/compilation/
- vllm/config/
- vllm/engine/
- vllm/entrypoints/logger.py
- vllm/envs.py
- vllm/logger.py
- vllm/logging_utils/
- vllm/platforms/
- vllm/sequence.py
- vllm/triton_utils/
- vllm/utils/
- tests/engine
- tests/test_sequence
- tests/test_config
- tests/test_logger
- tests/test_vllm_port
- tests/jit_monitor/test_hooks.py
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
pytest -v -s engine/test_arg_utils.py test_sequence.py test_logger.py test_vllm_port.py jit_monitor/test_hooks.py'

- label: Engine (1 GPU)
timeout_in_minutes: 30
device: intel_gpu
Expand All @@ -18,8 +56,50 @@ steps:
source_file_dependencies:
- vllm/v1/engine/
- tests/v1/engine/
- tests/test_config/
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
pytest -v -s v1/engine --ignore v1/engine/test_preprocess_error_handling.py'
pytest -v -s v1/engine --ignore v1/engine/test_preprocess_error_handling.py &&
VLLM_XPU_ENABLE_XPU_GRAPH=1 pytest -v -s test_config.py'

- label: V1 e2e (2 GPUs)
timeout_in_minutes: 30
device: intel_gpu
agent_tags:
label: production
gpu: 2+
mem: 16+
no_plugin: true
working_dir: "."
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
REPO: "vllm-ci-test-repo"
VLLM_TEST_DEVICE: "xpu"
source_file_dependencies:
- vllm/compilation/
- vllm/config/
- vllm/distributed/
- vllm/engine/
- vllm/envs.py
- vllm/forward_context.py
- vllm/inputs/
- vllm/logger.py
- vllm/logging_utils/
- vllm/model_executor/
- vllm/multimodal/
- vllm/platforms/
- vllm/sampling_params.py
- vllm/transformers_utils/
- vllm/triton_utils/
- vllm/utils/
- vllm/v1/
- tests/v1/e2e/spec_decode/
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
pytest -v -s
v1/e2e/spec_decode/draft_model/test_draft_model.py::test_draft_model_tensor_parallelism
v1/e2e/spec_decode/draft_model/test_draft_model.py::test_draft_model_engine_args_tensor_parallelism'
1 change: 1 addition & 0 deletions .buildkite/intel_jobs/kernels_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -22,4 +22,5 @@ steps:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
pytest -v -s ir &&
pytest -v -s kernels/ir'
35 changes: 30 additions & 5 deletions .buildkite/intel_jobs/misc_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -125,13 +125,13 @@ steps:
pytest -v -s v1/kv_offload &&
pytest -v -s v1/kv_connector/unit/test_offloading_connector.py'

- label: NixlConnector PD accuracy (2 GPUs)
- label: NixlConnector PD accuracy (4 GPUs)
timeout_in_minutes: 60
num_devices: 2
num_devices: 4
device: intel_gpu
agent_tags:
label: production
gpu: 2+
gpu: 4+
mem: 16+
no_plugin: true
working_dir: "."
Expand All @@ -140,15 +140,18 @@ steps:
REPO: "vllm-ci-test-repo"
VLLM_TEST_DEVICE: "xpu"
source_file_dependencies:
- vllm/distributed/kv_transfer/kv_connector/v1/nixl/
- vllm/distributed/kv_transfer/kv_connector/
- vllm/v1/worker/kv_connector_model_runner_mixin.py
- tests/v1/kv_connector/nixl_integration/
- vllm/platforms/xpu.py
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
bash v1/kv_connector/nixl_integration/run_xpu_disagg_accuracy_test.sh'
bash v1/kv_connector/nixl_integration/run_xpu_disagg_accuracy_test.sh &&
PREFILLER_TP_SIZE=2 DECODER_TP_SIZE=1 bash v1/kv_connector/nixl_integration/run_xpu_disagg_accuracy_test.sh &&
PREFILLER_TP_SIZE=1 DECODER_TP_SIZE=2 bash v1/kv_connector/nixl_integration/run_xpu_disagg_accuracy_test.sh &&
PREFILLER_TP_SIZE=2 DECODER_TP_SIZE=2 bash v1/kv_connector/nixl_integration/run_xpu_disagg_accuracy_test.sh'

- label: Regression
key: regression
Expand Down Expand Up @@ -259,3 +262,25 @@ steps:
pytest -v -s detokenizer &&
pytest -v -s -m "not cpu_test" ./multimodal &&
pytest -v -s utils_ --ignore=utils_/test_mem_utils.py'

- label: Fusion Unit Tests
timeout_in_minutes: 30
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
REPO: "vllm-ci-test-repo"
VLLM_TEST_DEVICE: "xpu"
source_file_dependencies:
- vllm/compilation/
- tests/compile/passes/test_qk_norm_rope_fusion.py
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
pytest -v -s compile/passes/test_qk_norm_rope_fusion.py'
33 changes: 33 additions & 0 deletions .buildkite/intel_jobs/model_executor_intel.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
group: Model Executor Intel
depends_on:
- image-build-xpu
steps:
- label: Model Executor (Intel)
key: model-executor-intel
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 24+
no_plugin: true
working_dir: "."
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
REPO: "vllm-ci-test-repo"
VLLM_TEST_DEVICE: "xpu"
source_file_dependencies:
- vllm/engine/arg_utils.py
- vllm/config/model.py
- vllm/model_executor
- tests/model_executor
commands:
- >-
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'apt-get update && apt-get install -y curl libsodium23 &&
pip3 install tensorizer==2.10.1 &&
pip3 install runai-model-streamer[s3,gcs,azure]\>=0.15.7 &&
export VLLM_WORKER_MULTIPROC_METHOD=spawn &&
export PYTHONFAULTHANDLER=1 &&
cd tests &&
pytest -v -s model_executor -m "not slow_test" --ignore="model_executor/layers/test_rocm_unquantized_gemm.py" --deselect="tests/model_executor/model_loader/test_reload.py::test_kv_scale_reload"'
Loading
Loading