Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
725 commits
Select commit Hold shift + click to select a range
ded6676
[Bugfix] Seed RayExecutorV2 TCPStore port by DP rank to avoid collisi…
eicherseiji Jun 30, 2026
00ebf19
[Bugfix][Quant] Raise actionable error instead of bare assert for gro…
ArsalanShakil Jun 30, 2026
db808b3
[Model Runner V2][Spec Decode] Implement block verification for rejec…
TheEpicDolphin Jun 30, 2026
c231d1f
fix(security): bound tokenizer work when explicit truncation_side is …
jperezdealgaba Jun 30, 2026
7cf7cbc
[Bugfix] MiniCPM-V 4.6: fix grid rows/cols swap in placeholder genera…
tc-mb Jun 30, 2026
dc148dc
[CI][Bugfix] Fix `Hybrid SSM NixlConnector PD prefix cache test (2 GP…
NickLucche Jun 30, 2026
d8f483d
[Spec Decode] Fix hidden-state extraction block size for hybrid verif…
imargulis Jun 30, 2026
9e84ec8
[Refactor] Remove dead minimax allreduce rms kernel (#46842)
yewentao256 Jun 30, 2026
fcaa84e
[BugFix] Gate MRV2 mixed sparse-MLA warmup on `max_num_seqs` > 1 (#47…
njhill Jun 30, 2026
e840f0d
[Platform] Replace `torch.cuda.Event` with `torch.Event` (#47140)
jikunshang Jun 30, 2026
245888f
[Feature] Detect all2all peer fault with fault tolerance backend and …
fangyuchu Jun 30, 2026
f41e8dd
[ROCm][CI] Move PyTorch Compilation Unit Tests to MI300(gfx942) (#47065)
charlifu Jun 30, 2026
7a341fa
[XPU] Support ZE_AFFINITY_MASK passthrough in xpu_disagg_acc_test (#4…
zhenwei-intel Jun 30, 2026
27d5f78
[CI] Move distributed small LM eval to B200 (#47048)
LucasWilkinson Jun 30, 2026
25671cb
[Parser][Bugfix] Ensure tool call or other special tokens don't leak …
bbrowning Jun 30, 2026
727971f
Add Medusa speculative decoding e2e test (#41396)
puririshi98 Jun 30, 2026
a773253
[Bugfix] Restore part of bugfix #42650 after accidental deletion in #…
JeanPaulShapo Jun 30, 2026
3cecee4
[Model Runner V2][Spec Decode] Fix stale values in idx_mapping from C…
TheEpicDolphin Jun 30, 2026
3a9784b
[Feature] DP supervisor using rust frontend (#47076)
yewentao256 Jun 30, 2026
953bba4
[PERF] Extend NCCL symmetric memory to AllGather and ReduceScatter (#…
WoosukKwon Jun 30, 2026
c8f9c15
[ROCm][V1][MLA] Clone prefill backend state per metadata builder (#46…
AndreasKaratzas Jun 30, 2026
20434c4
[Feat] Improve Triton JIT diagnostics (#46621)
LopezCastroRoberto Jun 30, 2026
11b26c5
[Bugfix][Tool Parser] PoolsideV1: fix logprobs AttributeError on Resp…
joerowell Jun 30, 2026
248d1fb
[Feat][1/N] CuTeDSL warmup infrastructure, FA4 MLA (#46182)
LopezCastroRoberto Jun 30, 2026
345b28f
[Hardware][AMD][CI] Bump timeouts of various test groups on AMD CI (#…
mawong-amd Jun 30, 2026
c8d2f3c
[Bugfix] compressed-tensors: allow int8 grouped WNA16 MoE on Marlin (…
joerowell Jun 30, 2026
6829473
[Bugfix] Align OpenCV video metadata timeline (#47099)
VectorPeak Jun 30, 2026
2824282
[Bugfix][Frontend] Normalize constrained Harmony recipients (#45657)
tarjan1 Jun 30, 2026
ac521f6
[Bugfix][Structured Outputs] Reject degenerate `structured_outputs` t…
Sunt-ing Jun 30, 2026
92c7fac
[Perf] Restore zero-init of swizzled NVFP4 scale buffer to recover Bl…
qiching Jun 30, 2026
b1190d0
[Refactor][GPT-OSS] Harmony Responses API Refactor to use HarmonyPars…
yzong-rh Jun 30, 2026
9294dd2
fix(reasoning): guard rfind in ernie45 streaming </response> branch (…
hclsys Jul 1, 2026
f098ee7
[GLM5] Support FlashMLA FP8 KV cache (Hopper & Blackwell) (#47090)
WoosukKwon Jul 1, 2026
a264e41
[Distributed] Default FlashInfer allreduce to mnnvl on single node (#…
WoosukKwon Jul 1, 2026
3406e8f
[Bugfix][Frontend][gpt-oss] Return raw output when Harmony parser end…
Achyuthan-S Jul 1, 2026
9969466
[Spec Decode] Support SWA + DFlash for MiMo (#46104)
benchislett Jul 1, 2026
3c1396b
[Hardware][AMD][CI] Toggle test coredumps on ROCm debug agent (#47222)
mawong-amd Jul 1, 2026
c5200d3
[Attention][DSA] support dcp for FLASHINFER_MLA_SPARSE (#46076)
ZJY0516 Jul 1, 2026
9a08a51
fix: skip cooperative top-K on SM120 (#47164)
lucifer1004 Jul 1, 2026
aeb35b9
[Rust Frontend] Add error context in tool parser failures (#46512)
cinnamonica02 Jul 1, 2026
93d8f83
[Core] Pluggable sleep-mode backend abstraction (RFC #34303) (#44074)
matteso1 Jul 1, 2026
df802a8
[CPU] Remove speculative decoding stream overrides from CPUModelRunne…
jmamou Jul 1, 2026
c3b1f9e
[ROCm][CI] Enable LoRA TP Distributed Test Group In AMD CI (#47193)
micah-wil Jul 1, 2026
b446792
[ROCm][Bugfix] Fix Triton "out of resource: shared memory" Error In O…
micah-wil Jul 1, 2026
89e9920
[CPU][Perf]Added tanh AOR for faster gelu activations. (#44639)
almayne Jul 1, 2026
5b431b9
[Rust Frontend] Coerce completion `max_tokens: null` to default (#47166)
blasrodri Jul 1, 2026
697c34b
[Bugfix] Fix beam search candidate indexing when logprobs count varie…
chaunceyjiang Jul 1, 2026
4470ae8
Remove mantis (#46806)
xianbaoqian Jul 1, 2026
a461070
[Core] Make sleep-mode backend capability flags communicator-agnostic…
matteso1 Jul 1, 2026
8f82be5
[CI/Build] Fix LoRA testing (#47242)
jeejeelee Jul 1, 2026
f651a8a
[XPU][UT]Enable ut qk_norm_rope_fusion (#42486)
Yejing-Lai Jul 1, 2026
77a9c5a
Weight sync refactor + move sparse nccl engine (#44353)
hao-aaron Jul 1, 2026
ed41aa2
[ROCm][DSV4] Use aiter mHC pre/post as the default ROCm path (#43950)
Fangzhou-Ai Jul 1, 2026
dee5da1
[Test] Run SageMaker handler-override tests in-process via TestClient…
Jyothirmaikottu Jul 1, 2026
fa4bec9
[Bugfix] Fix pooled Whisper sliding-window KV sizing (#47071)
andylolu2 Jul 1, 2026
aa8bb55
[ROCm][Perf][Bugfix] DSv4 indexer: use platform FP8 dtype (fnuz) for …
akii96 Jul 1, 2026
e7d0fcb
[CI] Fix various failures on `main` (#47197)
hmellor Jul 1, 2026
024b06b
[Bugfix] Expose usage field in GenerateResponse for disaggregated ser…
AIvashov Jul 1, 2026
cc56379
[Model] Support Hy3 token suffix and JSON Schema array types (#47192)
stevenkuang-tencent Jul 1, 2026
a22e0df
[Model] Remove AyaVision, MusicFlamingo (#47263)
hmellor Jul 1, 2026
4e5ca89
[ROCm][MiniMax-M3] Cross-layer lightning-indexer top-k sharing (#47269)
Fangzhou-Ai Jul 1, 2026
5c4db60
docs(security): document gRPC interface as insecure for private use o…
jperezdealgaba Jul 1, 2026
a78c156
Migrate GPTBigCode and Starcoder2 to the Transformers modeling backen…
hmellor Jul 1, 2026
f1cf6b0
[CI] Fix segfault in tracing test (#47299)
njhill Jul 1, 2026
13c49f9
[xpu][lora]: Align LoRA implementation with Punica GPU: fix _apply_ex…
chaojun-zhang Jul 1, 2026
c638f92
[Rust Frontend] Split engine core DTOs into separate modules (#47265)
BugenZhao Jul 1, 2026
63fcce4
[Bugfix] Fix GraniteMoeShared weight loading broken by #41184 (#47031)
mganczarenko Jul 1, 2026
f5a8d73
[Spec Decode] DSpark (#46995)
benchislett Jul 1, 2026
c8bdcc0
[Bench][BugFix] Fix empty decoder prompt for Cohere ASR in throughput…
mganczarenko Jul 1, 2026
00eb7ce
[Bugfix] Prevent padding placeholders from reaching embeddings (#47029)
qianlihuang Jul 1, 2026
5fd4421
[ROCm][P/D] MoRIIO toy proxy: support JSON Content-Type for OpenAI cl…
lcskrishna Jul 1, 2026
8cfeb84
[ModelRunner V2] Warmup cross-attn properly in encoder-decoder case (…
njhill Jul 1, 2026
4787f2d
[Bugfix] Don't read KV cache past `seq_len` in triton paged attn kern…
njhill Jul 1, 2026
d322943
[DSV4] Better MXFP8 quantization kernel (#47229)
zyongye Jul 1, 2026
fa24813
[MoE] Plumb gemm1_alpha/beta/clamp_limit into TRT-LLM FP8 MoE (#45723)
zyongye Jul 1, 2026
e91f5f8
[CI] Remove torch_nightly mirror tags (superseded by TORCH_NIGHTLY fu…
atalman Jul 1, 2026
e196268
[Docker] Remove unused Dockerfile.nightly_torch (#47338)
atalman Jul 1, 2026
2b753ad
[Spec Decode] DSpark speculators checkpoint support (#47093)
mgoin Jul 2, 2026
7fe7fa9
[CI][Bugfix] Rerun test_engine_log_metrics_ray on Ray GCS startup tim…
peizhang56 Jul 2, 2026
d0a2584
[Misc] Use functions instead of PTX for the PDL instruction (#46984)
jeejeelee Jul 2, 2026
1360c42
[UX] Include NVTX in cuda.txt (#47319)
jeejeelee Jul 2, 2026
d63c8e9
[BugFix][Spec Decode] Compact shared topk indices buffer after first …
TheEpicDolphin Jul 2, 2026
09663ab
[ROCm][MLA] Fuse MLA q/kv RMSNorm + FP8 per-token quant in the FP8 at…
xaguilar-amd Jul 2, 2026
2665ed7
[Bugfix][Kernel] Correct FlashInfer CUTLASS MoE tuning token bound (#…
Aneureka Jul 2, 2026
8357226
[XPU][CI] Split test_punica_ops into separate pytest invocations for …
chaojun-zhang Jul 2, 2026
3af8789
[Feature] Universal speculative decoding for heterogeneous vocabulari…
wan-danfeng Jul 2, 2026
b0b8a28
[Model] Add LLaVA-OneVision-2 (LlavaOnevision2ForConditionalGeneratio…
chengzheng345 Jul 2, 2026
08a8a4a
feat(rust): expose profiler control routes in Rust frontend (#46306)
pranavthakur0-0 Jul 2, 2026
25fcb65
[Rust Frontend] Use enum-backed domain types for engine outputs and s…
BugenZhao Jul 2, 2026
84b9c27
Update DeepGEMM tag to point to latest nv-dev branch for sm120 suppor…
mgoin Jul 2, 2026
de2a8fc
[ROCm] [PyTorch] Move to stable abi since ROCm upgraded to torch 2.11…
tjtanaa Jul 2, 2026
a2f7130
[ModelRunner V2] Enable by default for all dense models (#44443)
yewentao256 Jul 2, 2026
3e158ae
[ModelRunner V2] Fix Mamba2 crash on non-spec-decode (#47428)
njhill Jul 2, 2026
a47f38f
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in b…
WoosukKwon Jul 2, 2026
178fd56
support GLM-5.2 gate use FP32 (#47410)
zRzRzRzRzRzRzR Jul 2, 2026
ec0ffaa
[Rust Frontend] Improve scheduler stats logging parity (#47435)
BugenZhao Jul 2, 2026
320ee28
[Model Runner V2][Perf] Warm up GLM-5.2 DSA indexer prefill metadata …
chaunceyjiang Jul 2, 2026
443e68c
[Bugfix] Fix pooled Whisper encoder sliding-window kernel size (#47437)
njhill Jul 2, 2026
e392bf7
[BugFix][MRV2] Ensure all req slots are accounted for when scheduling…
njhill Jul 2, 2026
258f8de
[Bugfix][Tool Parser] poolside_v1: accept tool calls without newline …
joerowell Jul 2, 2026
d715b3a
Delete PagedAttention (#47361)
mgoin Jul 2, 2026
d29125c
Xqa decode kernels (#43232)
DanBlanaru Jul 2, 2026
e24d1b2
Fix Transformers modeling backend usage stats (#47472)
hmellor Jul 2, 2026
407f406
[ROCm][CI] Adding metadata (#47477)
AndreasKaratzas Jul 3, 2026
6768fbc
[ROCm][CI] Adding qwen3 dp4 eplb (#47480)
AndreasKaratzas Jul 3, 2026
442ccc6
[ROCm][CI] Adding extract hs 2gpu (#47482)
AndreasKaratzas Jul 3, 2026
4c3c64f
Add Laguna XS.2.1 DFlash drafter support (#46853)
adamkbaranowski Jul 3, 2026
34bf7b4
[CI] intel CI: add quantization and awq case for xpu (#46456)
wendyliu235 Jul 3, 2026
276b837
[ModelRunner V2][BugFix] Free all model refs on shutdown (#47483)
njhill Jul 3, 2026
d85601c
[CI] Pin modelscope version to fix test breakage (#47465)
njhill Jul 3, 2026
41de138
[BugFix] Derive FlashInfer Q dtype from resolved per-group builder st…
mgoin Jul 3, 2026
979f551
[Bugfix][Gemma4] Keep image bidirectional attention within the slidin…
lucianommartins Jul 3, 2026
1aeabec
[Bugfix][Rust Frontend] Tolerate out-of-vocab prompt ids in detokeniz…
Sunt-ing Jul 3, 2026
9b8e765
[Rust Frontend] Recover buffered text from incomplete tool calls at E…
reidliu41 Jul 3, 2026
3f0b773
[XPU][CI]Mv huggingface cache to larger disk in Intel GPU CI (#47405)
zxd1997066 Jul 3, 2026
bd8d902
[CPU][Build] Enable oneDNN ITT task collection by default for CPU pri…
eparshut Jul 3, 2026
2dfaae7
[XPU][CI]Fix dependency typo in Intel GPU CI (#47510)
zxd1997066 Jul 3, 2026
fbc9ba6
New stable abi cleanup (#46656)
cleonard530 Jul 3, 2026
6429d5f
[Rust Frontend] add repetition_detection support to sampling params (…
yangyang-cs95 Jul 3, 2026
b790c84
[CI] Enable sccache for Rust build under CUDA/ROCm (#45246)
BugenZhao Jul 3, 2026
1f486d9
Add Triton Backend for Unlimited-OCR R-SWA (#47102)
andakai Jul 3, 2026
4875b44
[Doc] Fix VLM2Vec benchmark chat template path (#47517)
kalyanamdewri Jul 3, 2026
bbdcbe4
Move Roberta remaining nn.Embedding to VocabParallelEmbedding (#47452)
maxdebayser Jul 3, 2026
400a9c3
[Rust Frontend] Bump llm-multimodal version (#47530)
Isotr0py Jul 3, 2026
18f658b
[Bugfix][Frontend] Fix batch chat endpoint corrupting logprobs when r…
fenghourun Jul 3, 2026
a14f57a
[Frontend] Refine the entrypoint class's inheritance hierarchy. (#47498)
noooop Jul 3, 2026
978de83
[Bugfix][CPU] Ship examples/ in the CPU release image (#47447)
AgenticSpark Jul 3, 2026
d7192cf
[CI Bugfix] Lazily import Qwen warmup dependencies (#47539)
LopezCastroRoberto Jul 3, 2026
3775d5f
[ROCm][CI] Adding test groups for parity with upstream (#47479)
AndreasKaratzas Jul 3, 2026
8651f04
[Rust Frontend] Speed up chat roundtrip tests (#47523)
BugenZhao Jul 3, 2026
f63dca6
[ROCm] Fix encoder-decoder cross-attention KV layout aliasing (#47035)
djramic Jul 3, 2026
f006e5a
[CI][AMD] Allow git operations on previously created work trees (#47554)
tpopp Jul 3, 2026
576bf75
[AMD][EPLB] Enable EPLB for Quark OCP MXFP4 MoE (#47220)
okorzh-amd Jul 3, 2026
3799501
[Bugfix][Multimodal] Normalize direct PIL image inputs (#47566)
Sunt-ing Jul 3, 2026
d6d39c1
[GLM4V] Avoid GLM4V processor init during startup metadata reads (#47…
labAxiaoming Jul 3, 2026
fb5291b
[Frontend] [Parser] Port DeepSeek V4 to streaming parser engine frame…
bbrowning Jul 4, 2026
ab3b6d9
[Frontend] Limit `SO_REUSEPORT` to multi-worker serving (#47529)
BugenZhao Jul 4, 2026
67ff0ae
Support nvfp4 kv with kv-cache-dtype-skip-layers sliding_window (#42890)
sychen52 Jul 4, 2026
07516fd
[MRV2][SD] Make Dynamic SD comatible with Full Cuda Graphs (#45953)
ekagra-ranjan Jul 4, 2026
f329ce4
[ROCm][CI][Bugfix] Use VllmRunner for `voxtral_realtime` tests to avo…
shen-shanshan Jul 4, 2026
4c3c17d
[ROCm] Disable persistent sparse-MLA kernel for chunked-prefill conti…
Rohan138 Jul 4, 2026
26eb872
[Bugfix] Fix CPU split-KV scratchpad sizing (#45844)
gausah01 Jul 4, 2026
e7c9df9
[Bugfix][Structured Output][Spec Decode] Constrain bitmask and trim g…
yuyue0225sc Jul 4, 2026
1a308c4
[XPU] Add W8A8 FP8 linear kernel with multi-granularity quant support…
chaojun-zhang Jul 4, 2026
6eac8e0
[Misc] Preserve cross-encoder pooling extra kwargs (#47082)
taneem-ibrahim Jul 4, 2026
fa1fa96
[Misc] Forward request-level prompt extras for cross-encoder scoring …
taneem-ibrahim Jul 4, 2026
2f21224
[Misc] Update request-extras parity for batch chat completion (#47333)
taneem-ibrahim Jul 4, 2026
1d354c6
[Misc] Validate Pooling cache_salt Values (#46966)
taneem-ibrahim Jul 4, 2026
f1445f6
[CI] Bump `huggingface-hub` from `v1.10.2` to `v1.22.0` (#47551)
hmellor Jul 4, 2026
0cd6f76
[Bugfix][Frontend][gpt-oss] Recover raw tail when Harmony parser ends…
yzong-rh Jul 4, 2026
2a9113f
[Perf] Remove redundant op for GLM 5.2 (#47198)
yewentao256 Jul 4, 2026
d2afe39
[Bugfix][Frontend] Preserve default sampling params in batch chat (#4…
Sunt-ing Jul 4, 2026
4a6bf3c
[ROCm][CI] Fix Kernels and Kernels attention test failures (#47519)
cpersson-amd Jul 4, 2026
91b5647
[Bugfix][Model] Allow Run:ai memory_limit sentinel values (#47337)
Sunt-ing Jul 5, 2026
34b560b
[Bugfix][Gemma4] Fix FA4 mm_prefix mask: add sliding window and absol…
lucianommartins Jul 5, 2026
9226613
[Bugfix][Pooling] Forward instruction to Jina reranker scoring prompt…
Sunt-ing Jul 5, 2026
fa4321d
[Bugfix][TurboQuant] Preserve KV cache dtype in backend shape (#47609)
LucasWilkinson Jul 5, 2026
b6cc46e
[Feature] Support sequence parallel without the need for DP, 1.9%~5.0…
yewentao256 Jul 5, 2026
fb2face
[Bugfix][Model] Fix crash loading Mamba/Mamba2 checkpoints without an…
Sunt-ing Jul 5, 2026
8974ed8
[Bugfix][Voxtral Realtime] Fix token feedback timeout silent hang (#4…
Sunt-ing Jul 5, 2026
cc1d020
[MRV2] Enable mm prefix bidi attention support on MRV2 (#46942)
Isotr0py Jul 5, 2026
b712181
[ROCm][Test] Fix test_per_token_group_quant_fp8 tolerance for 1-ULP F…
spandantiwari Jul 5, 2026
78a04c2
[XPU] Fix CUDA API shims breaking Torch Dynamo during AOT compile (#4…
lslusarczyk Jul 6, 2026
d2ec433
[XPU] Fix Eagle3 initialization on XPU (#43957)
chaojun-zhang Jul 6, 2026
95a248f
[Attention Backend] HPC_ATTN backend support mtp and dynamic schedule…
thisjiang Jul 6, 2026
f2aaf59
[Feature] Support MTP speculative decoding for Bailing hybrid models …
alex101-ops Jul 6, 2026
6569df6
[Test][LoRA] Use lightweight CPU reference and skip heavy cleanup in …
chaojun-zhang Jul 6, 2026
6971582
[Test][XPU] Skip fork in kv_sharing_fast_prefill test on XPU (#47406)
Liangliang-Ma Jul 6, 2026
394edc8
[XPU] limit max-num-seqs in test_lmeval.py for XPU (#47682)
mayuyuace Jul 6, 2026
f1073c0
[CPU][BugFix] Multiple fixes to w4a8_int8 CPU MoE path (#46739)
fadara01 Jul 6, 2026
e9cc1fd
[CI/Build][CPU] Remove global extra index (#47687)
bigPYJ1151 Jul 6, 2026
d9c1767
[INC][ARK] Direct Register Custom Op for ARK (#46361)
Zhenzhong1 Jul 6, 2026
16f8110
[Bugfix][CPU][RISC-V] Fix VLEN detection for RVV attention path (#47532)
I3eg1nner Jul 6, 2026
e433634
[Performance][Hardware][RISC-V] Reduce LMUL pressure in INT4 LUT dequ…
I3eg1nner Jul 6, 2026
990c2a0
[RISC-V] Enable BF16 on VLEN=256 hardware (#45243)
velonica0 Jul 6, 2026
98ba9b9
[Frontend] Support OpenAI Responses API namespace tools (#47024)
zhongjing123 Jul 6, 2026
8f0e75e
[ROCm][CI] Adding nixl multiconn (#47481)
AndreasKaratzas Jul 6, 2026
fb265fc
[ROCm][CI] Increasing parallelism in Basic Models Tests (Extra Initia…
AndreasKaratzas Jul 6, 2026
2fa1056
[Core][DP] Rotate load-balancer tie-break to avoid systematic engine …
mayuyuace Jul 6, 2026
cdab283
[XPU][CI]Add agent tags for Basic Models Tests (Initialization) in In…
zxd1997066 Jul 6, 2026
d039c17
[Bugfix] Recycle post-final-norm hidden in GLM MTP (single norm) (#47…
zhou9402 Jul 6, 2026
344609a
[CI/Build] Fix pre-commit check (#47695)
bigPYJ1151 Jul 6, 2026
736f1a5
[XPU] Route mm_prefix models to Triton attention backend (#47688)
zhenwei-intel Jul 6, 2026
3d7f357
[Doc] docs: fix note formatting for pooling models (#47701)
llsj14 Jul 6, 2026
26c754d
[XPU][Bugfix] Do not transpose weight_scale_inv at load time (#47116)
majian4work Jul 6, 2026
90ce3a0
[bugfix] fix MOSS-Audio deepstack_input_embeds initialization in PP (…
yma11 Jul 6, 2026
ba22152
fix(security): block request-level GPU video backend selection withou…
jperezdealgaba Jul 6, 2026
40cc2e8
[Bugfix] Return HTTP 422 for unprocessable image URLs instead of 500 …
akinsella Jul 6, 2026
740f379
[ROCm][AITER] Directly Implement AITER Custom All-reduce in CudaCommu…
BadrBasowid Jul 6, 2026
98e4726
[fix][run_batch]: respect proxy env vars when downloading media URLs …
mayuyuace Jul 6, 2026
f676808
[CI] Use TTY for AMD CI tests for colored buildkite logs (#47730)
njhill Jul 6, 2026
8b79971
attention: pass None for unused args in unified attention TD path (#4…
afierka-intel Jul 6, 2026
8f4c69b
[Rust Frontend] Cache metric handles for scheduler & request stats (#…
BugenZhao Jul 6, 2026
7a90eb9
[Bugfix] [Gemma4] Fix Gemma4 MTP draft model layers ignoring quant_co…
ayush1399 Jul 6, 2026
07f9baf
Revert "[Platform] Replace `torch.cuda.Event` with `torch.Event` (#47…
jikunshang Jul 6, 2026
641cb59
[Doc] Clarify fastokens availability (#45813)
LiJzd Jul 6, 2026
373eb31
[Bugfix][Core] Fix num_output_placeholders underflow with async sched…
Sunt-ing Jul 6, 2026
51ee564
[CI] Skip test for checkpoint that was deleted (#47748)
hmellor Jul 6, 2026
095adf1
[Bugfix] Fix int32 overflow in triton_decode_attention page offsets (…
ivanium Jul 6, 2026
598d511
[Bugfix][Distributed] Delegate MNNVL allreduce one-shot selection (#4…
jesco-absolut Jul 6, 2026
b1c6dba
[Refactor] Remove multiple dead code (#47329)
yewentao256 Jul 6, 2026
8d8ec38
[Bugfix][Spec Decode] Add missing draft_id_to_target_id to DSparkDeep…
Laurent-Zhang Jul 6, 2026
f70caef
[Perf] Cache `token_to_req_indices` for dsv4, 5x~6x kernel performanc…
yewentao256 Jul 6, 2026
5ad1117
[perf]Add fused Kimi image preprocessing (#47416)
Kevin-XiongC Jul 6, 2026
5bce653
Make the Transformers modeling backend as fast as native vLLM (#47187)
hmellor Jul 6, 2026
3ee9eea
[macOS][CPU][Installation] Fix the broken installation of vllm 0.24.0…
WindChimeRan Jul 6, 2026
24dd2ae
[Bugfix] Preserve FP8 indexer WK pairs across incremental load_weight…
lcheng321 Jul 6, 2026
9fde043
[Kernel][Helion][1/N] Add Helion kernel for silu_and_mul_per_block_qu…
xiaohongchen1991 Jul 6, 2026
b136cc2
[Bugfix][Model] Add stability window to DiffusionGemma to match HF st…
NathanielMcVicar Jul 6, 2026
b1384f5
Enable B12x backend for non-gated MoEs (like Nemotron) (#43328)
askliar Jul 6, 2026
ae098ab
[CI] Fix some errors on `main` (#47726)
hmellor Jul 6, 2026
04adc88
[Bugfix]Fix DeepSeek-V4 fp8_ds_mla KV cache reshape (#47716)
ACEEE-1222 Jul 6, 2026
d891b9b
[Quantization] add humming moe backend to all dense/moe oracles (#41652)
jinzhen-lin Jul 6, 2026
567a784
[Bugfix] Fix dp mtp hang (#40589)
SherryC41 Jul 6, 2026
482e552
[Bugfix][ROCm] Fix memory access fault in AITER MLA backend for DPA+F…
simondanielsson Jul 6, 2026
8484ca5
[ROCm][CI] Adding Rust parity (#47478)
AndreasKaratzas Jul 6, 2026
5769a73
[ROCm][CI][Bugfix] Fix flaky parallel tool-call streaming (test asser…
akii96 Jul 6, 2026
86db6c3
[Frontend] add per-request timing `metrics` field to response body of…
nv-nedelman-1 Jul 7, 2026
69f3150
[XPU] Fix PP accuracy on XPU device (#47253)
yisustc Jul 7, 2026
445321f
[Bugfix] [Quantization] Fix loading for CT DSV2 (#47780)
kylesayrs Jul 7, 2026
a46c932
[Rust Frontend] Add DeepSeek V3.2 roundtrip fixture (#47619)
reidliu41 Jul 7, 2026
9dd2465
feat(cpu): add CPU support for Mamba ShortConv (#35059)
rahulssv-ibm Jul 7, 2026
a4f019f
fix(distributed): propagate distributed_timeout_seconds to NCCL devic…
jialoop-git Jul 7, 2026
700e882
Add TorchCodec as a video decoding backend (#46609)
NicolasHug Jul 7, 2026
39a1d32
[Rust Frontend] Avoid extra copies for multimodal tensors (#47581)
reidliu41 Jul 7, 2026
34e6dfc
[Rust Frontend] Stamp `arrival_time` at the frontend entry (#47787)
tahsintunan Jul 7, 2026
c64c356
[Perf] Bound DiffusionGemma sampler transient via request-tiled logit…
guan404ming Jul 7, 2026
32ab064
[UX] Add `model_class_overrides` for development and debugging (#47148)
jeejeelee Jul 7, 2026
2f71b2b
[ROCm] Align mixed encoder-decoder KV cache views in V2 runner (#47685)
AndreasKaratzas Jul 7, 2026
cbe9c40
[Bugfix] Forward callable hf_overrides to the draft model config (#45…
HumphreySun98 Jul 7, 2026
6db31c8
[XPU][CI]Adjust memory request for tests in Intel GPU CI (#47758)
zxd1997066 Jul 7, 2026
dd5c299
[ROCm][Bugfix] Convert ModelOpt FP8 per-channel weights to e4m3fnuz o…
micah-wil Jul 7, 2026
e040899
[KV Offloading] Add basic offloading metrics (#45958)
Srinivasoo7 Jul 7, 2026
8e61b64
fix(security): add resource bounds validation to derender endpoints (…
jperezdealgaba Jul 7, 2026
1e823dc
[docs update] Update usage of `hf` cli for cache list and removal (#4…
ariG23498 Jul 7, 2026
b4cfbc2
[Bugfix][Core] Fix host memory leak from undrained new_block_ids (#44…
Sunt-ing Jul 7, 2026
ba50b97
[Bugfix] Match the mapped filename in find_loaded_library (#47586)
lucifer1004 Jul 7, 2026
e55cc59
[Rust Frontend][CI] Unblock more end-to-end test cases (#47735)
BugenZhao Jul 7, 2026
5d23ca4
[Kernel] Applies routed_scaling_factor internally (#47408)
jeejeelee Jul 7, 2026
066f02a
[MoE] FI autotuning: max bucket = max token count [e.g. `DP_size*MNBT…
netanel-haber Jul 7, 2026
c5b6623
[Bugfix][Spec Decode] Skip uniform spec-decode padding for diffusion …
kl527 Jul 7, 2026
c85d720
[HARDWARE][POWER] optimize math functions of VSX power (#47321)
Rukhaiya2004 Jul 7, 2026
d3e69fd
[Perf] Use blocking CUDA events to avoid busy polling cuda driver loc…
GirasoleY Jul 7, 2026
b3e85be
fix: use configured max_logprobs instead of hardcoded 20 in derender …
jperezdealgaba Jul 7, 2026
8172b98
Remove standalone flash-attn dependency from Qwen2.5-Omni audio
pratapyash May 11, 2026
ad18e76
Fix Qwen2.5-Omni audio tower tests
pratapyash May 11, 2026
7714f06
Use packed native Qwen2.5-Omni audio attention
pratapyash Jun 4, 2026
7b02a58
refactor(qwen2_5_omni): port audio-tower hygiene from the full featur…
pratapyash Jul 7, 2026
04ea911
fix(config): fail loud when cudagraph_mm_encoder is set under the v2 …
pratapyash Jul 7, 2026
bdbc4f4
test(qwen2_5_omni): assert the packed qkv k-bias contract at construc…
pratapyash Jul 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
6 changes: 3 additions & 3 deletions .buildkite/ci_config_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,17 +2,17 @@ name: vllm_intel_ci
job_dirs:
- ".buildkite/intel_jobs"
run_all_patterns:
- ".buildkite/ci_config_intel.yaml"
- ".buildkite/scripts/hardware_ci/run-intel-test.sh"
- "docker/Dockerfile"
- "docker/Dockerfile.xpu"
- "CMakeLists.txt"
- "requirements/common.txt"
- "requirements/xpu.txt"
- "requirements/build/cuda.txt"
- "requirements/test/cuda.txt"
- "setup.py"
- "csrc/"
- "cmake/"
run_all_exclude_patterns:
- "docker/Dockerfile."
- "csrc/cpu/"
- "csrc/rocm/"
- "cmake/hipify.py"
Expand Down
2 changes: 2 additions & 0 deletions .buildkite/hardware_tests/amd.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ steps:
# differ ci_base is rebuilt and pushed automatically.
- label: "AMD: :docker: ensure ci_base"
key: ensure-ci-base-amd
soft_fail: false
depends_on: []
device: amd_cpu
no_plugin: true
Expand All @@ -26,6 +27,7 @@ steps:

- label: "AMD: :docker: build test image and artifacts"
key: image-build-amd
soft_fail: false
depends_on:
- ensure-ci-base-amd
device: amd_cpu
Expand Down
14 changes: 10 additions & 4 deletions .buildkite/hardware_tests/cpu.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -17,12 +17,14 @@ steps:
- tests/kernels/test_awq_int4_to_int8.py
- tests/kernels/quantization/test_cpu_fp8_scaled_mm.py
- tests/kernels/mamba/cpu/test_cpu_gdn_ops.py
- tests/kernels/mamba/test_cpu_short_conv.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 30m "
pytest -x -v -s tests/kernels/attention/test_cpu_attn.py
pytest -x -v -s tests/kernels/moe/test_cpu_fused_moe.py
pytest -x -v -s tests/kernels/moe/test_cpu_quant_fused_moe.py
pytest -x -v -s tests/kernels/mamba/test_cpu_short_conv.py
pytest -x -v -s tests/kernels/test_onednn.py
pytest -x -v -s tests/kernels/test_awq_int4_to_int8.py
pytest -x -v -s tests/kernels/quantization/test_cpu_fp8_scaled_mm.py
Expand Down Expand Up @@ -53,7 +55,7 @@ steps:
- tests/models/language/pooling/
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 40m "
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 50m "
pytest -x -v -s tests/models/language/generation -m cpu_model
pytest -x -v -s tests/models/language/pooling -m cpu_model"

Expand All @@ -68,13 +70,15 @@ steps:
- vllm/v1/sample/ops/topk_topp_triton.py
- vllm/v1/sample/ops/topk_topp_sampler.py
- tests/v1/sample/test_topk_topp_sampler.py
- tests/v1/e2e/test_cpu_linear_attn_chunked_prefix.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 45m "
uv pip install git+https://github.com/triton-lang/triton-cpu.git@270e696d
VLLM_USE_V2_MODEL_RUNNER=1 pytest -x -v -s tests/models/language/generation/test_granite.py -m cpu_model
# TODO: move to CPU-Kernel Tests once triton-cpu has a pre-built wheel
pytest -x -v -s tests/v1/sample/test_topk_topp_sampler.py::TestTritonTopkTopp"
pytest -x -v -s tests/v1/sample/test_topk_topp_sampler.py::TestTritonTopkTopp
pytest -x -v -s tests/v1/e2e/test_cpu_linear_attn_chunked_prefix.py"

- label: CPU-Quantization Model Tests
depends_on: []
Expand All @@ -89,11 +93,13 @@ steps:
- vllm/model_executor/layers/fused_moe/experts/cpu_moe.py
- tests/quantization/test_compressed_tensors.py
- tests/quantization/test_cpu_wna16.py
- tests/quantization/test_cpu_w8a8.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 45m "
pytest -x -v -s tests/quantization/test_compressed_tensors.py::test_compressed_tensors_w8a8_logprobs
pytest -x -v -s tests/quantization/test_cpu_wna16.py"
pytest -x -v -s tests/quantization/test_cpu_wna16.py
pytest -x -v -s tests/quantization/test_cpu_w8a8.py"

- label: CPU-Distributed Tests (PP+TP)
depends_on: []
Expand Down Expand Up @@ -136,7 +142,7 @@ steps:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 45m "
pytest -x -v -s tests/models/multimodal/generation --ignore=tests/models/multimodal/generation/test_pixtral.py -m cpu_model --num-shards=$$BUILDKITE_PARALLEL_JOB_COUNT --shard-id=$$BUILDKITE_PARALLEL_JOB"
parallelism: 3
parallelism: 4

- label: "Arm CPU Test"
depends_on: []
Expand Down
12 changes: 12 additions & 0 deletions .buildkite/hardware_tests/intel_xpu_ci/test-intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,10 @@ steps:
timeout_in_minutes: 30
optional: true
device: intel_gpu
agent_tags:
label: production
gpu: 2+
mem: 24+
no_plugin: true
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
Expand All @@ -38,6 +42,10 @@ steps:
timeout_in_minutes: 30
optional: true
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 24+
no_plugin: true
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
Expand All @@ -55,6 +63,10 @@ steps:
timeout_in_minutes: 30
optional: true
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
env:
REGISTRY: "public.ecr.aws/q9t5s3a7"
Expand Down
5 changes: 3 additions & 2 deletions .buildkite/image_build/image_build_arm64.sh
Original file line number Diff line number Diff line change
Expand Up @@ -21,12 +21,13 @@ else
exit 0
fi

# build (Grace/GH200 is the arm64 GPU target; sm_90)
# build for arm64 GPU targets: Grace/GH200 (sm_90) and DGX Spark/GB10
# (sm_121, family-covered by 12.0 under CUDA 13)
docker build --file docker/Dockerfile \
--platform linux/arm64 \
--build-arg max_jobs=16 \
--build-arg nvcc_threads=4 \
--build-arg torch_cuda_arch_list="9.0" \
--build-arg torch_cuda_arch_list="9.0 12.0" \
--build-arg USE_SCCACHE=1 \
--build-arg buildkite_commit="$BUILDKITE_COMMIT" \
--tag "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-arm64 \
Expand Down
5 changes: 5 additions & 0 deletions .buildkite/intel_jobs/basic_correctness.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,10 @@ steps:
- label: XPU Sleep Mode
timeout_in_minutes: 30
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand All @@ -19,4 +23,5 @@ steps:
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
export VLLM_WORKER_MULTIPROC_METHOD=spawn &&
pytest -v -s basic_correctness/test_cpu_offload.py &&
pytest -v -s basic_correctness/test_mem.py::test_end_to_end'
4 changes: 4 additions & 0 deletions .buildkite/intel_jobs/engine_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,10 @@ steps:
- label: Engine (1 GPU)
timeout_in_minutes: 30
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand Down
6 changes: 5 additions & 1 deletion .buildkite/intel_jobs/expert_parallelism_intel.yaml
Original file line number Diff line number Diff line change
@@ -1,11 +1,15 @@
group: Expert Parallelism
depends_on:
depends_on:
- image-build-xpu
steps:
- label: EPLB Algorithm
key: eplb-algorithm
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand Down
4 changes: 4 additions & 0 deletions .buildkite/intel_jobs/kernels_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,10 @@ steps:
- label: vLLM IR Tests
timeout_in_minutes: 30
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand Down
30 changes: 28 additions & 2 deletions .buildkite/intel_jobs/lora_intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,10 @@ steps:
- label: LoRA Runtime + Utils
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 24+
no_plugin: true
working_dir: "."
env:
Expand Down Expand Up @@ -34,6 +38,10 @@ steps:
- label: LoRA Fused/MoE Kernels
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand All @@ -54,6 +62,10 @@ steps:
- label: LoRA Punica Kernels
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand All @@ -69,11 +81,17 @@ steps:
'cd tests &&
export VLLM_WORKER_MULTIPROC_METHOD=spawn &&
set -o pipefail &&
pytest -v -s lora/test_punica_ops.py --deselect="tests/lora/test_punica_ops.py::test_kernels_hidden_size[expand-0-xpu:0-dtype0-3-43264-32-4-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels[shrink-0-xpu:0-dtype1-1-2049-64-128-16]" --deselect="tests/lora/test_punica_ops.py::test_kernels[shrink-0-xpu:0-dtype0-1-2049-128-1-32]" --deselect="tests/lora/test_punica_ops.py::test_kernels[shrink-0-xpu:0-dtype0-1-2049-256-1-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels[shrink-0-xpu:0-dtype0-1-2049-256-8-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels[expand-0-xpu:0-dtype0-3-2049-128-8-16]" --deselect="tests/lora/test_punica_ops.py::test_kernels[shrink-0-xpu:0-dtype0-1-2049-128-8-32]" --deselect="tests/lora/test_punica_ops.py::test_kernels[expand-0-xpu:0-dtype1-1-2049-256-128-32]" --deselect="tests/lora/test_punica_ops.py::test_kernels_hidden_size[shrink-0-xpu:0-dtype0-3-64256-32-4-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels_hidden_size[shrink-0-xpu:0-dtype1-2-29696-32-4-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels_hidden_size[shrink-0-xpu:0-dtype1-3-49408-32-4-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels_hidden_size[shrink-0-xpu:0-dtype0-2-16384-32-4-4]" --deselect="tests/lora/test_punica_ops.py::test_kernels_hidden_size[expand-0-xpu:0-dtype0-2-51328-32-4-4]"'
pytest -v -s lora/test_punica_ops.py::test_kernels &&
pytest -v -s lora/test_punica_ops.py::test_kernels_hidden_size &&
pytest -v -s lora/test_punica_ops.py::test_add_lora_fused_moe_early_exit'

- label: LoRA Punica FP8/XPU Ops
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand All @@ -94,6 +112,10 @@ steps:
- label: LoRA Models
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 2+
mem: 24+
no_plugin: true
working_dir: "."
env:
Expand All @@ -108,15 +130,19 @@ steps:
bash .buildkite/scripts/hardware_ci/run-intel-test.sh
'cd tests &&
export VLLM_WORKER_MULTIPROC_METHOD=spawn &&
(pytest -v -s lora/test_mixtral.py --deselect="tests/lora/test_mixtral.py::test_mixtral_lora[4]" || true) &&
pytest -v -s lora/test_quant_model.py --deselect="tests/lora/test_quant_model.py::test_quant_model_lora[model0]" --deselect="tests/lora/test_quant_model.py::test_quant_model_lora[model1]" --deselect="tests/lora/test_quant_model.py::test_quant_model_tp_equality[model0]" &&
pytest -v -s lora/test_transformers_model.py &&
pytest -v -s lora/test_chatglm3_tp.py &&
pytest -v -s lora/test_llama_tp.py::test_llama_lora &&
pytest -s -v lora/test_minicpmv_tp.py'

- label: LoRA Multimodal
timeout_in_minutes: 45
device: intel_gpu
agent_tags:
label: production
gpu: 1+
mem: 16+
no_plugin: true
working_dir: "."
env:
Expand Down
Loading
Loading