Skip to content

[Core] Skip JIT monitor when JIT warmup is disabled - #58590

Merged
vllm-bot merged 1 commit into
vllm-project:mainfrom
njhill:fix-jit-monitor-eager
Sep 24, 2026
Merged

vllm-bot merged 1 commit into
vllm-project:mainfrom
njhill:fix-jit-monitor-eager

Conversation

@njhill

@njhill njhill commented Sep 24, 2026

Copy link
Copy Markdown
Member

Follow-on from #58197

With enable_jit_warmup=False (e.g. under enforce_eager), runtime JIT compilation is expected, so the post-warmup JIT monitor would only produce noisy warnings (or spurious errors in error mode).

With enable_jit_warmup=False (e.g. under enforce_eager), runtime JIT
compilation is expected, so the post-warmup JIT monitor would only
produce noisy warnings (or spurious errors in error mode).

Co-authored-by: Kimi Code <noreply@moonshot.cn>
Signed-off-by: Nick Hill <nickhill123@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@njhill
njhill requested a review from mgoin September 24, 2026 16:48
@njhill

njhill commented Sep 24, 2026

Copy link
Copy Markdown
Member Author

/ci run

@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 24, 2026
@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90994 for commit 0798baa77473.

@mgoin

mgoin commented Sep 24, 2026

Copy link
Copy Markdown
Member

My only caveat is it might still be useful to get the "warning" when a jit is compiling at runtime for debugging purposes, but obviously we shouldn't error

@njhill
njhill enabled auto-merge (squash) September 24, 2026 17:18
@vllm-bot
vllm-bot merged commit e30559b into vllm-project:main Sep 24, 2026
155 of 158 checks passed
@njhill
njhill deleted the fix-jit-monitor-eager branch September 24, 2026 19:14
kristobalus added a commit to kristobalus/vllm that referenced this pull request Sep 25, 2026
* [Bugfix] Fix external LB DP rank handling when replicas share nodes (#53743)

Signed-off-by: Tony Lin <tony.lin@intel.com>

* [docs] Fix legacy hf CLI references (vllm) (#57958)

Signed-off-by: Wauplin <lucainp@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NIXL] Fix DCP pulls across MLA cache regions (#57389)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [ROCm] Refactor tuned gemms (#55001)

Signed-off-by: Andy Friedrich <afriedri@amd.com>
Signed-off-by: afriedri <afriedri@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] unskip InternViT test for transformers v5 compatibility (#55767)

Signed-off-by: sahil <sahil@example.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [MM] Move get_dummy_processor_inputs into MM processor (#57967)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [ROCm][CI] Use ROCm backend for DeepSeek V4.1 ViT test (#57931)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Feature] Add first-class KV hints request envelope for programmatic KV management (#53423)

Signed-off-by: Karen Chung <karenc@nvidia.com>

* [Docs] Add an Engram feature page explaining Engram usage in vLLM (#57910)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [ROCm][Perf] Route the fused shared-expert gate GEMM through the platform dispatcher (#54185)

Signed-off-by: Mikko Tukiainen <Mikko.Tukiainen@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Engram] Drop redundant VLLM_PLE_CPU_OFFLOAD env var (#57937)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [Bugfix][MoE] Reject hash routing for unsupported monolithic backends (#57867)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][ROCm] Reject unsupported EP for monolithic AITER MXFP4 MoE (#57866)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Docker] Use zstd for CI images and offer a Docker Hub variant (#55608)

Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Kimi-K3][AMD] Return KDA and MLA projection outputs directly (#50592)

Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Co-authored-by: Liuyinfeng01 <199041580+LiuYinfeng01@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][Build] Harden triton-cpu sleef submodule fetch in CPU image build (#57871)

Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>

* [Bugfix][Kernel] Skip the fused silu-mul block-quant fast path when a swiglu clamp is set (#57984)

Signed-off-by: Garrett Goon <garrett@primeintellect.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Frontend] Add reusable TP1 initialized-engine snapshots (#51360)

Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: elehayym <52448798+Yuzu23@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] MRV2 pooling shutdown model ref (#57737)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [CI] Split (H200) LM Eval Large Models into per-model jobs (#57965)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Scheduler] Soften Long Prefill Tokens Threshhold (#57951)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Perf][Attention] Reduce GLM sparse MLA preparation overhead (#57458)

Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>

* [Bugfix] Annotate MTP draft KV cache groups positionally on the hybrid grouping path (#55390)

Signed-off-by: Navjot Singh <navjot.singh@shopify.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][GDN] Fix stateless first-chunk classification (#51565)

Signed-off-by: taking-lying-flat <1615405@qq.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com>

* [CI] Add pre-commit check that new tests are tethered to Buildkite jobs (#54867)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Turner <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [XPU] Fix Nemotron FP8 LM-eval config: drop CUDA-only moe_backend and wire to new Buildkite job (#49685)

Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [Docker] Expose bundled vllm-rs on PATH (#57606)

Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Signed-off-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>

* [Fast Start] Cache the MTP draft model in a separate daemon group (#57312)

Signed-off-by: liusy58 <mg21330037@smail.nju.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Rust Frontend] Add MiMo V2.5 parser support (#57933)

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Bugfix][Spec Decode] Cap DFlash/DSpark profiling query batch (#56448)

Signed-off-by: wangyicong <wangyicong@bytedance.com>

* [RL][Sleep] Retain frozen weights across level-2 sleep (#57891)

Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>

* [ROCm][Bugfix] Explicitly reject FSE=1 with DPA+ETP deployment for DeepSeek-V4 (#57919)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Spec Decode] Enable async scheduling for DFlash (#58065)

* [Frontend][Rust] Add mm-processor benchmark for Rust frontend (#51922)

Signed-off-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Signed-off-by: gangula-karthik <gkarthik923@gmail.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Rust Frontend] Introduce parser-owned output grammar interfaces (#55269)

Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Security] Reject min_tokens that exceeds the filled max_tokens default (#57731)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>

* [Bugfix][Frontend] Validate mixed prompt embedding mask lengths (#57006)

Signed-off-by: 子华 <huaxi.shx@alibaba-inc.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Qwen2.5-VL] Honor video fps for temporal M-RoPE (#47736)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Rust Frontend] Build full-output grammars from initialized reasoning parsers (#57340)

Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Codex <noreply@openai.com>

* [ROCm][Perf] Avoid extra reshape kernel in Qwen GDN output norm (#47842)

Signed-off-by: Mikko Tukiainen <Mikko.Tukiainen@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Kernel] Add opt-in load-time MXFP4 dequantization (#50814)

Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>

* [Rust Frontend] Separate multimodal instrumentation from request timing (#58084)

Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Kernel][DSV4.1] Fuse MXFP8 wo_b GEMM with sequence-parallel reduce-scatter (#57428)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: Canlin <canlinguosdu@gmail.com>
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [CI] Emit a kernel symbol map from the csrc build (opt-in, for test selection) (#58097)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [perf] wire FA and FlashMLA for sm90 GLM5Next NoPE SparseMLA (#55385)

Signed-off-by: JaredforReal <w13431838023@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [CPU] Add device-memory-utilization CLI alias (#56547)

Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Louie Tsai <louie.tsai@intel.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Rust Frontend] Construct model-owned vision processors through specs (#58109)

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [BugFix][Core] Make the structured-output grammar poll non-blocking (#55931)

Signed-off-by: ubwzwd <ubwzwd@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com>

* [XPU][CI]Remove model_runner_v2 test from Intel GPU CI (#58050)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Bugfix][Structured Outputs] Reject empty `structural_tag` at request validation (#47450)

Signed-off-by: linnea-lin-00638949 <15521435947@163.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com>

* [Build] Fix CUDA 12 KV connector dependency selection (#57945)

Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>

* [Bugfix][GLM-5.3-Flash] Run the dense MLP layers on the sequence-parallel shard (#58061)

Signed-off-by: Jared Wen <w13431838023@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Structured Output] Disallow MRV1 + PP>1 + async sched + structured output (#56250)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: CNE Pierre FICHEPOIL <pierre-1.fichepoil@gendarmerie.interieur.gouv.fr>

* [Feat][XPU] VLLM_BATCH_INVARIANT support for Dense/MoE models (#55881)

Signed-off-by: Tony Lin <tony.lin@intel.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [SpecDecode] Restore residual-logits comments in _resample_kernel (#58166)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Narrow AuxOutput KV restrictions to known PD connectors (#58150)

Signed-off-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix] prioritize architecture capability before DeepGEMM availability check (#58073)

Signed-off-by: Tony Lin <tony.lin@intel.com>

* [Bugfix][Attention] Avoid NaN in the Triton softcap for large attention logits (#56579)

Signed-off-by: Kushal Dabbe <72650064+kushaldabbe@users.noreply.github.com>
Co-authored-by: opencode <noreply@opencode.ai>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][Engram] Fall back when /dev/shm is absent before sharing tables (#57914)

Signed-off-by: Juntian Liu <juntianl@inferact.ai>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] hadacore_transform: respect inplace parameter to fix garbage outputs with QuIP transforms (#43462)

Signed-off-by: Gilles Turpin <turpingilles15@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Bugfix][ROCm] Dispatch the QuantFP8 CUDA fallback on the class (#58136)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][DSv4.1][Perf] Fuse the inverse RoPE into the sparse decode reduce (#57435)

Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Fast Start] Support data parallelism in the weight cache daemon  (#57386)

Signed-off-by: liusy58 <mg21330037@smail.nju.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix] batch_invariant: keep non-AllReduce collectives enabled on NCCL >= 2.31 (#58179)

Signed-off-by: Guanxin Li <38149783+guanxingithub@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [MRV2] Release weight offloader on shutdown (#57834)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][V1] Honor enable_jit_warmup for V2 kernel warmup (#55146)

Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm][CI] Add GELU activation for AiterExperts in the modular-kernel coverage (#58030)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][V1] Read ModelState max_model_len from model config (#58149)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Model][Spec Decode] Defer disposable GLM MTP head (#55442)

Signed-off-by: Luca Motz <luca.motz@icloud.com>

* [Refactor] Remove dead code multiple places (#58002)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Core] structured generation mode for DiffusionGemma model (Jev-like) (#57250)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Razorback16 <razorback16@protonmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Razorback16 <razorback16@protonmail.com>

* [Kernel] Remove AllSpark INT8 W8A16 GEMM backend (#58001)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [CI] Build the torch-nightly image on Ubuntu 24.04 (#58204)

* [Bugfix] Backport Inductor custom-op pattern matching fix (#58189)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Core] Disable JIT warmup in eager mode (#58197)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [CI] Split LM Eval TurboQuant KV Cache into per-config jobs (#57113)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.ai>

* [CI] Shard (H200 MIG 18GB) Spec Decode Draft Model across whole-directory replicas (#58193)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Perf] Remove CPU-GPU sync in heterogeneous vocabulary speculative decoding (#57396)

* [ROCm][Build][The Rock] Bump Triton version to 3.8.x tip-of-tree with source build in The Rock image (#58006)

Signed-off-by: Randall Smith <Randall.Smith@amd.com>

* [Bugfix][ROCm] Use the platform FP8 range in the concat MLA q test (#58153)

Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][Model][Bugfix] Enable GLM-5.2-MXFP4 on the deepseek_v32 path and fix sparse attention correctness (#51915)

Signed-off-by: Jack Hu <Jack.Hu@amd.com>
Signed-off-by: Jack Hu <jack.hu@amd.com>
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com>
Co-authored-by: James E T Smith <jamesETsmith@users.noreply.github.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Douglas Lehr <Doug.Lehr@amd.com>

* [ROCm] Use silu_and_mul_with_clamp's torch._C op (#52052)

Signed-off-by: Tres Popp <tres.popp@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix] Skip VllmConfig re-validation for with_hf_config submodel views (#58212)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Roger Wang <rogerw@inferact.ai>

* [EPD] Support metadata-only audio inputs (#57887)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [ROCm][DSv4][Perf] Fuse the inverse RoPE into the sparse decode reduce (#57451)

Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>

* [ROCm][Compile] Support BF16 AsyncTP fusion (#58098)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [CI][Bugfix] Update IPC test caller for #57312's _apply_entries signature (#58107)

Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [ROCm][CI] Stage G gating (#50922)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Bugfix] Set worker runtime threads before profiling and compilation (#55891)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [XPU] Wire up SYCL apply_rotary_emb kernel in ApplyRotaryEmb (#55721)

Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com>
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Compilation] Fix QuTLASS compilation with PyTorch 2.13 (#58173)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Spec decode] Support variable-length decode for Kimi-K3 adaptive ver (#52988)

Signed-off-by: Albert Cheng <albecheng@nvidia.com>
Signed-off-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-bia01.bia.clusters.nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [XPU][CI] Deselect tests/v1/spec_decode/test_mtp.py::test_glm_mtp_defers_lm_head (#58237)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>

* [MoE] Use GateLinear for all MoE models (#58234)

Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>

* [Bugfix][Frontend] Keep length finish_reason for max_tokens-truncated streaming tool calls (#46303)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [ROCm][Perf] Use wvSplitK for single-output GEMMs (#53283)

Signed-off-by: tangzzycc <3081129260@qq.com>

* [Tests] Select V2 for diffusion scheduler unit tests (#58272)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Quantization] Select per-token NVFP4 MoE backends explicitly (#57176)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the three TurboQuant evaluation groups on MI355 (#58282)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Mooncake] Address review nits from #56855 (#57174)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Yifan Qiao <17067717+ivanium@users.noreply.github.com>

* [Refactor][Quantization] Make FP8 and MLA weight transforms reusable pure functions (#57732)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [ROCm][Compile] Fuse AITER static FP8 attention output (#58099)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [ROCm][Bugfix] Register MRV2 sampler JIT warmups (#58092)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Attention] Avoid CPU-GPU sync in DCP sequence lengths (#58169)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Extend groupwise rms_norm scale tolerance to CUDA (#58252)

Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>

* [Perf][ROCm][Attention] Narrow the Triton prefill-attention KV tile on RDNA3/RDNA4 (#58225)

Signed-off-by: Jipeng Li <jipengli@amd.com>
Co-authored-by: GitHub Copilot CLI <noreply@github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][Bugfix] Keep zero MiniMax MXFP8 activation blocks finite (#58089)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Include Python tooling in ROCm CI artifacts (#58271)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [MRV2] Miscellaneous code cleanup (#57980)

* [Bugfix][SM120][MLA] Support NoPE sparse MLA (GLM-5.3-Flash) on the FlashInfer SM120 backend (#55277)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [XPU] upgrade to PyTorch 2.14 (#56013)

Signed-off-by: Yan Ma <yan.ma@intel.com>
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [ROCm][Test] Check GDN prefill numerics and output ownership (#58091)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix] Disable prefix caching for encoder-only before model config hooks (#58287)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [Bugfix][Tool Parser] Migrate Granite to the streaming Parser Engine (#49648)

Signed-off-by: Nikhil Kulkarni <nikhilkulkarni1755@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [DSpark] Support pipeline-parallel targets in aggregated serving (#56956)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Feature][Frontend] Add DeepSeek-V4 FIM completion rendering (#44229)

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [Perf][MoE] Skip top-k slots routed to non-local experts in TritonExp… (#58051)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Shuolei Wang <948904026@qq.com>

* [CPU] Adds support for fp32 attention sinks (#56252)

Signed-off-by: Ankit Jaiswal <ankit.jaiswal@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>

* [Quantization] Enable humming wNaM asymmetric quant (zero_point) with compressed-tensors (#46528)

* [Quantization][Bugfix] Bump humming-kernels to 0.1.16 (#58054)

Signed-off-by: jinzhen.ljz <jinzhen.ljz@antgroup.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Quark] Remove quark-specific silent online quantization (#51800)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm][Perf] Extend QK-norm/RoPE/KV-cache fusion to MRoPE (#50212)

Signed-off-by: Vorapol Assavasangthong <Vorapol.Assavasangthong@amd.com>
Co-authored-by: Santosh Hiremath <Santosh.Hiremath@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>

* [ROCm][CI] Validate Mooncake and NIXL prefill/decode accuracy (#58095)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [CI][ROCm] Add an MI355 Kimi-K3 unit test group (#58012)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][NIXL] Restore successful push completion reporting (#58188)

Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][ROCm] Fix startup OOM in AITER MLA FP8 prefill workspace sizing (#57923)

Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>

* Doc: add DiffusionGemma to supported models (#46466)

Signed-off-by: Bruce <Bruce798858117@gmail.com>
Signed-off-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf] Use breakable CUDA graphs (no torch.compile) by default under VLLM_BATCH_INVARIANT so the tuned matmul configs see the runtime M (#57586)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>

* [CI] Select one GPU for the H200 initialized snapshot E2E step (#58351)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Pick a KV block size supported by every attention backend (#49845)

Signed-off-by: Divy <divy@coralbricks.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Docs] Fix docstring typos (output_dytpe, kwrags, Abbrivations) (#55936)

Signed-off-by: simpleqt <89645338+simpleqt@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Test] Cover MoRI graph replay and output lifetime (#58093)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][Bugfix] Fix TileLang mHC fused RMSNorm on 64-wide wavefronts (#58419)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][Bugfix] Limit MRV2 sampler JIT warmup registration to ROCm (#58465)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Disable JIT warmup by default in VllmRunner (#58452)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][MRV2] Align dummy idx_mapping dtype to avoid runtime jit (#58462)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [5/12][ci-selector][CI] Skip the Proton GPU test when another CUPTI tool is injected (#58455)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Share BF16 baselines across quantization comparison tests (#58469)

Signed-off-by: Aarushi Jain <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][DSA] Bound DeepSelect sentinel columns in the sparse top-k remap (#58215)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Test][Determinism] Cover chunked prefill in the batch-invariance suite (#55612)

Signed-off-by: Bob Ok <49168652+blipbyte@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [CI][ROCM] Add the Fusion E2E TP2 Quick group on MI355, and the AITER MLA fix it needs (#58369)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [Scheduler] Tune --long-prefill-token-threshold adaptiveness (#58459)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>

* [PCP] Support prefill context parallelism with data parallelism (#57075)

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Give turboquant boundary layers a layout-compatible backend (#54988)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Resubmit PR 48666 - Gemma4 FP8 KV FA4 head dim 512 backend selection (#53175)

Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>

* [Bugfix][KV Offload] Retain offload event metadata through batch translation (#57453)

Signed-off-by: Kapil Arya <kapila@nvidia.com>
Signed-off-by: Kapil Arya <kapil.arya.17@gmail.com>
Co-authored-by: Or Ozeri <or@ozery.com>

* Fix full logprobs in token-in/token-out responses (#58488)

Signed-off-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][CPU][MoE] Fix out-of-bounds write and segfault when router weights are fp32 (#56168)

Signed-off-by: Farzad Abdolhosseini <farzad@elastix.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Rust Frontend] Recognize new frontend-owned serve args as unsupported or no-op (#58330)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Rust Frontend] Accept custom chat roles for HF templates (#58311)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Dependency] Upgrade FlashInfer version to 0.7.0 (#58069)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Bugfix][Spec Decode] Separate DSpark width from MTP stage validation (#54631)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Rust Frontend] Pass vision preprocessing context for Nemotron-H (#57634)

Pass the remaining engine context-length budget to model-owned vision processors through VisionPreprocessingContext. Preserve Nemotron batched engine fields and recognize llm_config as a text_config alias.

Use the merged upstream llm-multimodal revision f0985ef65967615db2c79279aa07818499301bfd.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Rust Frontend] Support `--sse-keep-alive-interval` (#58306)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Compile][CI] Honor Triton cache overrides and add AMD timeout headroom (#58474)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [CPU][GDN] Support NIXL DS convolution-state layout (#53300)

Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [ROCm][CI] Add the MI355 TurboQuant t3nc mirror (#58432)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [EPD][Model Loader] Skip language-model checkpoint shards for `--mm-encoder-only` (#58086)

Signed-off-by: grYe99 <guorongye99@gmail.com>
Co-authored-by: grYe99 <guorongye99@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [Perf] Use Conv3dLayer for MiniMax M3 patch embedding (#58512)

Signed-off-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Feature][Frontend] Request JSON body debug logging on `--enable-log-requests` flag (#58163)

Signed-off-by: talora <talora@nvidia.com>

* [CPU] Use pre-built triton (#58140)

Signed-off-by: jiang1.li <jiang1.li@intel.com>

* [Bugfix][V1] Reject encoder-cache hits with mismatched embedding counts (#57696)

Signed-off-by: jackLei0901 <42642542+jackLei0901@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Capture prefill kernels for mixed FULL graphs (#58275)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Bugfix][CI] Fix the flaky sharded-sampling tests, and the engine teardown need (#58342)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

---------

Signed-off-by: Tony Lin <tony.lin@intel.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Andy Friedrich <afriedri@amd.com>
Signed-off-by: afriedri <afriedri@amd.com>
Signed-off-by: sahil <sahil@example.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Karen Chung <karenc@nvidia.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Mikko Tukiainen <Mikko.Tukiainen@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Signed-off-by: Garrett Goon <garrett@primeintellect.ai>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Navjot Singh <navjot.singh@shopify.com>
Signed-off-by: taking-lying-flat <1615405@qq.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Turner <doubleujabbour@gmail.com>
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Signed-off-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Signed-off-by: liusy58 <mg21330037@smail.nju.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Signed-off-by: gangula-karthik <gkarthik923@gmail.com>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: 子华 <huaxi.shx@alibaba-inc.com>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: Canlin <canlinguosdu@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: JaredforReal <w13431838023@gmail.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Louie Tsai <louie.tsai@intel.com>
Signed-off-by: ubwzwd <ubwzwd@gmail.com>
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Signed-off-by: linnea-lin-00638949 <15521435947@163.com>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Signed-off-by: Jared Wen <w13431838023@gmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: Kushal Dabbe <72650064+kushaldabbe@users.noreply.github.com>
Signed-off-by: Juntian Liu <juntianl@inferact.ai>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Gilles Turpin <turpingilles15@gmail.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>
Signed-off-by: Guanxin Li <38149783+guanxingithub@users.noreply.github.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Razorback16 <razorback16@protonmail.com>
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
Signed-off-by: Jack Hu <Jack.Hu@amd.com>
Signed-off-by: Jack Hu <jack.hu@amd.com>
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com>
Signed-off-by: Tres Popp <tres.popp@amd.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com>
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com>
Signed-off-by: Albert Cheng <albecheng@nvidia.com>
Signed-off-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-bia01.bia.clusters.nvidia.com>
Signed-off-by: tangzzycc <3081129260@qq.com>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: Jipeng Li <jipengli@amd.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: Yan Ma <yan.ma@intel.com>
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Signed-off-by: Nikhil Kulkarni <nikhilkulkarni1755@gmail.com>
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Shuolei Wang <948904026@qq.com>
Signed-off-by: Ankit Jaiswal <ankit.jaiswal@amd.com>
Signed-off-by: jinzhen.ljz <jinzhen.ljz@antgroup.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vorapol Assavasangthong <Vorapol.Assavasangthong@amd.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: Bruce <Bruce798858117@gmail.com>
Signed-off-by: Misha Goin <mgoin64@gmail.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Divy <divy@coralbricks.ai>
Signed-off-by: simpleqt <89645338+simpleqt@users.noreply.github.com>
Signed-off-by: Aarushi Jain <Aarushi.Jain2@amd.com>
Signed-off-by: Bob Ok <49168652+blipbyte@users.noreply.github.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>
Signed-off-by: Kapil Arya <kapila@nvidia.com>
Signed-off-by: Kapil Arya <kapil.arya.17@gmail.com>
Signed-off-by: Farzad Abdolhosseini <farzad@elastix.ai>
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Signed-off-by: grYe99 <guorongye99@gmail.com>
Signed-off-by: OpenAI Codex <codex@openai.com>
Signed-off-by: talora <talora@nvidia.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jackLei0901 <42642542+jackLei0901@users.noreply.github.com>
Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: Tony Lin <tony.lin@intel.com>
Co-authored-by: Lucain <lucainp@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: afriedri <afriedri@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Sahil Patel <91423311+Sip4818@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
Co-authored-by: djramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Karen Chung <karenc@nvidia.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Mikko Tukiainen <mikko.tukiainen@amd.com>
Co-authored-by: Nils Matteson <nilsmatteson@icloud.com>
Co-authored-by: yinfengLiu <yinfeliu@amd.com>
Co-authored-by: Liuyinfeng01 <199041580+LiuYinfeng01@users.noreply.github.com>
Co-authored-by: vllm-agent <claw@inferact.ai>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: Garrett Goon <44747910+garrett361@users.noreply.github.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: elehayym <52448798+Yuzu23@users.noreply.github.com>
Co-authored-by: Thang Nguyen <69278249+Thangnguyenvn98@users.noreply.github.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
Co-authored-by: Navjot Singh <navjot.singh@uwaterloo.ca>
Co-authored-by: cherry77-cloud <1615405@qq.com>
Co-authored-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Chaojun Zhang <chaojun.zhang@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: siyu <mg21330037@smail.nju.edu.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Yicong Wang <wangyicong@bytedance.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Jeff (Junze) Ma <93145857+majunze2001@users.noreply.github.com>
Co-authored-by: karthik <56480632+gangula-karthik@users.noreply.github.com>
Co-authored-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Co-authored-by: Juan Pérez de Algaba <124347725+jperezdealgaba@users.noreply.github.com>
Co-authored-by: shaohuaxi <huaxi.shx@alibaba-inc.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
Co-authored-by: Jared Wen <w13431838023@gmail.com>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: Louie Tsai <louie.tsai@intel.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: ubwzwd <ubwzwd@gmail.com>
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: xiangdong <40376367+zxd1997066@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: linyafeng <15521435947@163.com>
Co-authored-by: CNE Pierre FICHEPOIL <pierre-1.fichepoil@gendarmerie.interieur.gouv.fr>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Kushal <72650064+kushaldabbe@users.noreply.github.com>
Co-authored-by: opencode <noreply@opencode.ai>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Gilles Turpin <turpingilles15@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: stefankoncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Guanxin Li <38149783+guanxingithub@users.noreply.github.com>
Co-authored-by: pengyihang <1017861497@qq.com>
Co-authored-by: Divakar Verma <137818590+divakar-amd@users.noreply.github.com>
Co-authored-by: hcl <chenglunhu@gmail.com>
Co-authored-by: lucamotz <luca.motz@icloud.com>
Co-authored-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Razorback16 <razorback16@protonmail.com>
Co-authored-by: Andrey Talman <atalman@fb.com>
Co-authored-by: Kimi <noreply@moonshot.ai>
Co-authored-by: Michael Lapshin <55516685+MichaelLapshin@users.noreply.github.com>
Co-authored-by: rasmith <Randall.Smith@amd.com>
Co-authored-by: Jack Hu <jack.hu@amd.com>
Co-authored-by: James E T Smith <jamesETsmith@users.noreply.github.com>
Co-authored-by: Douglas Lehr <Doug.Lehr@amd.com>
Co-authored-by: Tres <tpopp@users.noreply.github.com>
Co-authored-by: Roger Wang <rogerw@inferact.ai>
Co-authored-by: Tianyu Guo <guoty@inferact.ai>
Co-authored-by: Michał Ganczarenko <michal.gancz…
arbi-dev added a commit to arbicity/vllm-turbo that referenced this pull request Oct 7, 2026
* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [SpecDecode] Add LiLiCorr drafter (#57934)

Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>

* [MRV2] Minor model_runner.py code cleanup (#58610)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix generative scoring body cancellation (#57729)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] Preserve reranker tokenization with document limits (#57666)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Kernel][Perf] Register-resident path for per-token-group 8-bit quant (#55330)

Signed-off-by: chao.huan <chao.huan@nio.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf][Kernel] Vectorized flat abs-max for dynamic per-tensor FP8 quantization (#58194)

Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Core][Logging] Fix JSON logging process decoration (#57957)

Signed-off-by: Mark McLoughlin <markmc@redhat.com>

* [Bugfix] Stop allocator fragmentation from shrinking the KV cache during memory profiling (#58430)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Keep logprobs of parser-suppressed streaming chunks (#58583)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Yanxiao Zhao <39199723+sdpkjc@users.noreply.github.com>
Co-authored-by: Prakhar Agarwal <270064960+agarwalprakhar2511@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix][GLM-5.3-Flash] SM90 sparse MLA: index_kpool mismatch leads to corruption via unread query token (#58704)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Benchmark] Record model_id in bench latency/throughput --output-json (#58112)

Signed-off-by: yashasvi <yashasvi@ibm.com>

* [Bugfix][GLM-5.3-Flash] kpool corruption with speculative decoding (#58454)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [GLM5.3 Perf] Optimize glm 5.3 metadata op, 1.6~4.8x kernel level performance improvement (#58450)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Bugfix][KV Connector] Reap expired NIXL leases behind a heartbeated head (#58292)

Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>

* [ROCm][Kimi-K3] Optimize low-concurrency speculative KDA (#58045)

Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Cover the AITER MQA logits dispatch on gfx950 (#58724)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Add quantized MoE serving test for gfx950 (#58748)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Test AMD DeepSeek V4 MoE routing against a PyTorch reference (#58740)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: one-pass sampler statistics kernel (#58226)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [ROCm][CI] Expand single-GPU coverage on MI355 DPX (#57599)

Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [CI] [MRV2] Restore MRV2 pp dp coverage (#57735)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][CI] Report subprocess test skips as skips, not passes (#58701)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Run the MLA attention+quant fusion test on ROCm (#58717)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm] Bump torch 2.13, triton 3.8, torchaudio, torchvision (#50605)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: jpvillam <juan.villamizar@amd.com>

* [Feature][Frontend] Add granite_thinking_parser reasoning parser for Granite 4.2 (#55957)

Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Respect max_output_tokens in the Harmony tool-call loop (#58551)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Du Bin <8174807+dubin555@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Frontend] Document 404 response for `/generative_scoring` (#58788)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Perf][Frontend] Defer reasoning usage recounts for non-continuous chat streams (#56067)

Signed-off-by: Cheng Rui <286040359@qq.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] Default missing detail for Responses API input images (#57241)

Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>

* [watermarking] golden tests for backwards compatibility (#56809)

Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Don't drop the rest of the allocator config when toggling expandable segments (#57982)

Signed-off-by: Oxana Korzh <okorzh@amd.com>

* [AuxOutput] Only require Model Runner V2 on GPU platform (#58205)

Signed-off-by: Linkun Chen <github@lkchen.net>

* [Pooling] Preserve BERT-family heads for raw logits (#57664)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* fix: perf: use startswith(x, i) instead of string slicing to avoid O(N^2) (#52580)

Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states (#48419)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Co-authored-by: Claude <noreply@anthropic.com>

* [Bugfix][Reasoning] Count Kimi K3 reasoning tokens (#58372)

Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [CI] Split (H200 MIG/MI300) Basic Correctness into named jobs (#57054)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Frontend] Fix Inkling tool name leaking into content after reasoning (#58792)

Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Core] Bound draft-token RPC waits by the execute-model timeout (#58779)

Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Co-authored-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>

* [Perf][DSv4.1] Shard the Engram wkv projection across TP ranks (#58678)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>

* [RL] Add sharding-aware NCCL M2N weight transfer (#51520)

Signed-off-by: Ke Wen <kwen@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][ROCm] AMD-Quark mixed-precision DeepSeek-V4.1 support (#57071)

Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix][Frontend][Rust Frontend] Update DeepSeek V4.1 Flash reasoning effort mappings (#58316)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>

* [CI] Only isolate the registry tests that need a fresh process (#58764)

Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Spec Decode] Enable Gemma4 DSpark adaptive verification with FlashInfer (#57263)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [CI] Split (B200) Miscellaneous Kernels into mHC, FLA Ops and Misc named jobs (#58609)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Security] Gate per-request multimodal processor kwargs (#58830)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Elastic EP] Fix EPLB load statistics during scaling (#58473)

Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>

* [Refactor] Remove dead tests utils (#58803)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights (#58720)

Signed-off-by: Willian <willian@willian.email>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix] Fix Anthropic Thinking Disabled with P/D (#58786)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Bugfix][Frontend] Detect Anthropic inline-system merge against the resolved chat template (#58754)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Mypy] Fix mypy typing for Qwen and Qianfan models (#58046)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [CI] Reduce CUDA graph mode test overhead (#58749)

* [GLM5.3 Bug] Fix sparse indexer attn topk backend selection (#58594)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [CI] Stabilize batch submission in full CUDA graph tests (#58810)

* [Bugfix][DSV4.1] Avoid host sync in ViT CUDA graph replay metadata (#58499)

* [Security] Harden message sanitization (#58832)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [Kernel][DSV4.1] Fuse MoE finalize into the TP all-reduce + mHC boundary (#58586)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [PCP][DCP] Support DCP target model with non-DCP Dspark (#56723)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Bump FlashKDA to keep the recurrent state in fp32 (#58846)

Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector (#57251)

Signed-off-by: Vincent <vincexxchan@gmail.com>

* [mooncake] support CUSTOM_MEM_POOL in vllm (#49300)

Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: bruce.xu <bruce.xb@alibaba-inc.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Qwen3.8-Flash-Next] Avoid memory fragmentation in QSA indexer logits workspace (#57105)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix] Disable sequence parallelism / async TP under batch invariance and add a TP regression test (#56377)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Kimi-K3] Make VLLM_ROCM_USE_AITER_MOE_SITUV2 select a4w4/a8w4/a16w4 (#58201)

Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>

* [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL (#50047)

Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Co-authored-by: xijiade.aihemaiti <3146335281@qq.com>

* [Bugfix] V1: clear stale allowed_token_ids mask in InputBatch.condense (#43931)

Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Perf][DSv4.1] Fuse small-batch WO-A with inverse RoPE and MXFP8 quant on SM100/SM103 (#58634)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Attention][CPU] Use zentorch SDPA for CPU MLA prefill (#54967)

Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [CI] Split (H200 MIG 35GB) Spec Decode Speculators + MTP into 4 named jobs (#57237)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Thang Nguyen <thang.nguyen@inferact.ai>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][Perf] Enable layer-aware CSA2 multi-stream overlap for DeepSeek-V4.1-Flash (#57407)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Perf][Pooling] Avoid blocking seq_lens GPU-to-CPU copy for pooling in FlashInfer metadata builder (#57214)

Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>

* [Bugfix] Support repsonse_format + tool_choice=auto (#56086)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: pablopupo <145598901+pablopupo@users.noreply.github.com>
Co-authored-by: hubunt <150658615+hubunt@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Fast Start] Add `/health` endpoint for the weight cache daemon (#58552)

Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Perf][MoE] Use fused MiniMax2 routing with non-unit routed scaling (#58880)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>

* [ROCm][Bugfix] Fall back to default GEMM for CPU tensors on ROCm builds (#58923)

Signed-off-by: fai <fangzhouai@gmail.com>

* [Bugfix][KV Connector] Retry Mooncake bootstrap registration on timeout (reopens #55763) (#58919)

* [Frontend] Switch Python Harmony dependency to oss-harmony (#55128)

Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm] Bump AITER to v0.1.23 (#58867)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Count Responses reasoning tokens per tool round (#58927)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [KimiViT][Perf] Fuse per-layer QK RoPE into one in-place kernel (#58651)

* [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models (#58239)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [Bugfix][Frontend] Sample batched chat completions from the adjusted requests (#58929)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Use a fresh parser per choice in non-streaming chat completions (#58939)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][EPD] Skip sampling for encoder-only async steps (#58490)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [CPU] Build CPU wheels on Ubuntu 22.04 with AMX-FP8 support (#58515)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>

* [Test][ROCm] Stabilize the mixed OLMoE LoRA test (#58945)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [Model] Enable LoRA support for RobertaForSequenceClassification (#58884)

Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Co-authored-by: jz-yolo <jz-yolo@users.noreply.github.com>

* [Bugfix][Frontend] Apply Harmony adjust_request in batched chat completions (#58958)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Minimax-M3] Add Encoder CUDA graph support (#58673)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <codex@openai.com>

* [Perf][MRV2] Reuse Mamba/GDN metadata across KV cache groups (#58762)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Xin Yang <xyangx@amazon.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [Bugfix] Fix the two multimodal root tests that fail on main (OpenPangu-VL embed merge, MiMo sink test fixture) (#58900)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Benchmark] Add Responses API backend to vllm bench serve (#54628)

Signed-off-by: QHarshil <harshil_c@hotmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [CI] Allowlist-shrink batch 1: wire 17 root-level tests + drop 5 stale watermarking entries into misc.yaml (#58055)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [CPU] Use accelerator memory API in DiffusionGemma (#58964)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [CPU] Add video inferencing via torchcodec on s390x (#58693)

Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Skills] Update kernel-microbenchmark to include ROCm (#58646)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Misc] Name each backend and its kernel block sizes in block-size errors (#58557)

Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: khluu <khluu000@gmail.com>

* [CI] Split Dynamic Shapes out of (H200 MIG 35GB) PyTorch Compilation + (MI300) mirror (#58451)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Release the profiling KV cache held by QSA key views (#58961)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CPU] Include vLLM Recipes tooling in release image to deploy models using vLLM Recipes (#58796)

Signed-off-by: louie-tsai <louie.tsai@intel.com>

* Revert "[CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)" (#59011)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Perf][Qwen4Exp] Fuse HC down projection and SiLU on NVIDIA (#58957)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix][Model] Gemma4: register aliased embedding scalars as buffers (#54213)

Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>

* [Bugfix][Spec Decode] Implement get_top_tokens() on the ROCm DeepSeek V4 MTP drafter (#57568)

Signed-off-by: BaoYunkai <ybao@amd.com>

* [Perf][Qwen3.8] Reduce PLE metadata construction overhead (#58114)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm][Refactor] Move DeepSeek-V4/V4.1 multi-stream overlap gate to ROCm platform (#58983)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Multimodal] Avoid extra d2d for encoder cudagraph with fused input norm (#56711)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Logging] Preserve application log record factories (#58747)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix] Use the correct repository revision for secondary artifact loaders (#57461)

Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Perf] Kimi-K3 Enable sharded latent MoE up-projection under EP (#54956)

Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Feature] Add fixed-token prefill scoring (#54335)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [LoRA] Support variable num_labels for sequence classification (#57766)

Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [ROCm][Perf] Replace torch.topk in DSA candidate block selection (#58208)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [ROCm]Keep LMCache OpenTelemetry on the image's 1.40 stack (#59056)

Signed-off-by: Micah Williamson <micah.williamson@amd.com>

* [Bugfix][Kimi-K3] Refresh DSpark context KV cache pointers after the KV cache is re-bound (#58814)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [KVConnector][NIXL] Support packed MLA KV layouts in pipeline-parallel push prefill (#50499)

Signed-off-by: zixi-qi <zixi@inferact.ai>

* [Quant] Use canonical N-first weight format for CT WNA16 MoE (#52798)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: HDCharles <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Relax packed_qk_rope_ correctness test to one ULP (#59008)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Pin OpenTelemetry to LMCache's cap in the ROCm images (#59051)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Increase timeout for Entrypoints Unit (#59076)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][MLA] Add an AITER ASM round-robin decode route for DCP multi-token verify (#56861)

Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix] Don't sync-police or retry FlashInfer all-reduce workspace creation in eager mode (#58498)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix resumable request + async scheduling handoff race (#58259)

Signed-off-by: Yifan Zong <yzong@redhat.com>

* [Model Runner V2][Spec Decode] Support spec decode with draft model (#43091)

Signed-off-by: Icey <1790571317@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][MRV2] Allow FULL decode graphs for one-token prompt tails (#58400)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Bugfix][Scheduler] Preserve logprobs across streaming continuations (#57447)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>

* [Mypy] Fix mypy typing for Voxtral and vision models (#58251)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Core] Rework scheduler `skipped_waiting` queue (#58947)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Bugfix][MRV2][Spec Decode] Reject draft slots that were never proposed (#58784)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][CI] Add missing test coverage for upstream parity (#50519)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][Scheduler] Refresh max tokens for streaming continuations (#57676)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Spec Decode] Avoid triton recompiles in the acceptance estimator (#57107)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>

* [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8664f67d8ae1538e5e5cec2b1ce3f258dd)

* [XPU][CI]Skip test_abort_timeout_on_prefiller in nightly (#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db1ca080793dc2351144cb39513fd7984ca)

* [Bugfix][Engram] Keep THP tables private when resolving shared memory (#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c352f079c2cb8f46752fd0317a2745050bf)

* [ROCm][CI] Expand MI355 mirrors and route MIG-sized jobs to DPX (#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd64f5ef75dbe1471ee88c009003b3d03ec)

* [Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse retention (#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bddbeab6b4a0d5861dfcb171bf61ce2109d6)

* [Core][BugFix] Tag prefix-cache extra keys by source (#51899)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e7ed5b6f372f0898a67e48d3365fa28139)

* [Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e097b905127b1d5f09a7e93bd1095e87a6c)

* [Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf53724e06cef7f07628ae8fd4cc6c96f11a)

* [Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f668608155bc8a8c74cb97e3ddd0d3053085)

* [Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc757cffc11b55689d9cc33d844d9606516)

* [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a4297921c699f9aa58f9c6395e95ede718)

* [Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa130fac1e1252330a680fbed45ad05eba64f)

* [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)

(cherry picked from commit 5463fe4962785cdc3383477bf3af6533e7647dfd)

* [Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87cca25690fef6bd12667fd0d26690d949af)

* [Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec22ad9bb098393021b956feaa97081fc85)

* [Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae4d746f64c22677348d1cc120eea6f7439)

* [Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e615c20ee4ff0cd82aac592c7d6310aa7c3)

* [Core] Include the LoRA path in prefix-cache block hashes (#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e075ed461cff922aa710d387281f651940)

* [Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360c9eebb3b4b460ec49f6606ef0a08fcae2)

* [CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d7332db897885fdc0e6ed61b28c969d743)

* [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba9b7d18f9746d989e389a51a94b0f96f6f)

* [Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a6963664537ed21172e2ec12e96e3a2dcd3718c)

* [Bugfix][HiSparse] Size the KV cache from the groups HiSparse allocates (#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2e4489126f1c21868699f3890021f01b73)

* [Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with a saturated GPU pool (#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8ac1f8a7e8fb50cf1c56ff96649fcadccc6)

* [CI] Drop a test that depends on #57930 from the #59309 backport

Resolving the #59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from #57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream #59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>

* [Bugfix] Fix minimax-m3 multimodal processor compatability with Transformers v5.18 (#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f160a2c0abcb5902acc3c91a14c38a4af173)

* [Misc] Add Transformers version upper bound in requirements (#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298457dde7b4554b3b4e20b238c0ac2c3a65)

* Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0

Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2eb #28, 76a3e6d72, 6d4c75611 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a46, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 18.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c19, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - per-group BlockPools, sampler reserve, padded-page block fill, drain
    hook / on_kv_manager_created, KVBlockZeroer clamp: conflict with
    upstream's single-pool allocator; no turbo-attn consumer for the hooks
  - rotary fast-path registry: upstream guards the import (1f60771c7,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Signed-off-by: chao.huan <chao.huan@nio.com>
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: errmakov <ide404@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Signed-off-by: yashasvi <yashasvi@ibm.com>
Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Signed-off-by: Cheng Rui <286040359@qq.com>
Signed-off-by: Yifan Zong <yzong@redhat.com>
Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Linkun Chen <github@lkchen.net>
Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Ke Wen <kwen@nvidia.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: zixi-qi <zixi@inferact.ai>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: Willian <willian@willian.email>
Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Vincent <vincexxchan@gmail.com>
Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: QHarshil <harshil_c@hotmail.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>
Signed-off-by: BaoYunkai <ybao@amd.com>
Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Stefan Kon…
arbi-dev added a commit to arbicity/vllm-turbo that referenced this pull request Oct 7, 2026
* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [SpecDecode] Add LiLiCorr drafter (#57934)

Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>

* [MRV2] Minor model_runner.py code cleanup (#58610)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix generative scoring body cancellation (#57729)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] Preserve reranker tokenization with document limits (#57666)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Kernel][Perf] Register-resident path for per-token-group 8-bit quant (#55330)

Signed-off-by: chao.huan <chao.huan@nio.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf][Kernel] Vectorized flat abs-max for dynamic per-tensor FP8 quantization (#58194)

Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Core][Logging] Fix JSON logging process decoration (#57957)

Signed-off-by: Mark McLoughlin <markmc@redhat.com>

* [Bugfix] Stop allocator fragmentation from shrinking the KV cache during memory profiling (#58430)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Keep logprobs of parser-suppressed streaming chunks (#58583)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Yanxiao Zhao <39199723+sdpkjc@users.noreply.github.com>
Co-authored-by: Prakhar Agarwal <270064960+agarwalprakhar2511@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix][GLM-5.3-Flash] SM90 sparse MLA: index_kpool mismatch leads to corruption via unread query token (#58704)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Benchmark] Record model_id in bench latency/throughput --output-json (#58112)

Signed-off-by: yashasvi <yashasvi@ibm.com>

* [Bugfix][GLM-5.3-Flash] kpool corruption with speculative decoding (#58454)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [GLM5.3 Perf] Optimize glm 5.3 metadata op, 1.6~4.8x kernel level performance improvement (#58450)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Bugfix][KV Connector] Reap expired NIXL leases behind a heartbeated head (#58292)

Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>

* [ROCm][Kimi-K3] Optimize low-concurrency speculative KDA (#58045)

Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Cover the AITER MQA logits dispatch on gfx950 (#58724)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Add quantized MoE serving test for gfx950 (#58748)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Test AMD DeepSeek V4 MoE routing against a PyTorch reference (#58740)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: one-pass sampler statistics kernel (#58226)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [ROCm][CI] Expand single-GPU coverage on MI355 DPX (#57599)

Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [CI] [MRV2] Restore MRV2 pp dp coverage (#57735)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][CI] Report subprocess test skips as skips, not passes (#58701)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Run the MLA attention+quant fusion test on ROCm (#58717)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm] Bump torch 2.13, triton 3.8, torchaudio, torchvision (#50605)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: jpvillam <juan.villamizar@amd.com>

* [Feature][Frontend] Add granite_thinking_parser reasoning parser for Granite 4.2 (#55957)

Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Respect max_output_tokens in the Harmony tool-call loop (#58551)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Du Bin <8174807+dubin555@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Frontend] Document 404 response for `/generative_scoring` (#58788)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Perf][Frontend] Defer reasoning usage recounts for non-continuous chat streams (#56067)

Signed-off-by: Cheng Rui <286040359@qq.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] Default missing detail for Responses API input images (#57241)

Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>

* [watermarking] golden tests for backwards compatibility (#56809)

Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Don't drop the rest of the allocator config when toggling expandable segments (#57982)

Signed-off-by: Oxana Korzh <okorzh@amd.com>

* [AuxOutput] Only require Model Runner V2 on GPU platform (#58205)

Signed-off-by: Linkun Chen <github@lkchen.net>

* [Pooling] Preserve BERT-family heads for raw logits (#57664)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* fix: perf: use startswith(x, i) instead of string slicing to avoid O(N^2) (#52580)

Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states (#48419)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Co-authored-by: Claude <noreply@anthropic.com>

* [Bugfix][Reasoning] Count Kimi K3 reasoning tokens (#58372)

Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [CI] Split (H200 MIG/MI300) Basic Correctness into named jobs (#57054)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Frontend] Fix Inkling tool name leaking into content after reasoning (#58792)

Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Core] Bound draft-token RPC waits by the execute-model timeout (#58779)

Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Co-authored-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>

* [Perf][DSv4.1] Shard the Engram wkv projection across TP ranks (#58678)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>

* [RL] Add sharding-aware NCCL M2N weight transfer (#51520)

Signed-off-by: Ke Wen <kwen@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][ROCm] AMD-Quark mixed-precision DeepSeek-V4.1 support (#57071)

Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix][Frontend][Rust Frontend] Update DeepSeek V4.1 Flash reasoning effort mappings (#58316)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>

* [CI] Only isolate the registry tests that need a fresh process (#58764)

Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Spec Decode] Enable Gemma4 DSpark adaptive verification with FlashInfer (#57263)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [CI] Split (B200) Miscellaneous Kernels into mHC, FLA Ops and Misc named jobs (#58609)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Security] Gate per-request multimodal processor kwargs (#58830)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Elastic EP] Fix EPLB load statistics during scaling (#58473)

Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>

* [Refactor] Remove dead tests utils (#58803)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights (#58720)

Signed-off-by: Willian <willian@willian.email>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix] Fix Anthropic Thinking Disabled with P/D (#58786)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Bugfix][Frontend] Detect Anthropic inline-system merge against the resolved chat template (#58754)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Mypy] Fix mypy typing for Qwen and Qianfan models (#58046)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [CI] Reduce CUDA graph mode test overhead (#58749)

* [GLM5.3 Bug] Fix sparse indexer attn topk backend selection (#58594)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [CI] Stabilize batch submission in full CUDA graph tests (#58810)

* [Bugfix][DSV4.1] Avoid host sync in ViT CUDA graph replay metadata (#58499)

* [Security] Harden message sanitization (#58832)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [Kernel][DSV4.1] Fuse MoE finalize into the TP all-reduce + mHC boundary (#58586)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [PCP][DCP] Support DCP target model with non-DCP Dspark (#56723)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Bump FlashKDA to keep the recurrent state in fp32 (#58846)

Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector (#57251)

Signed-off-by: Vincent <vincexxchan@gmail.com>

* [mooncake] support CUSTOM_MEM_POOL in vllm (#49300)

Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: bruce.xu <bruce.xb@alibaba-inc.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Qwen3.8-Flash-Next] Avoid memory fragmentation in QSA indexer logits workspace (#57105)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix] Disable sequence parallelism / async TP under batch invariance and add a TP regression test (#56377)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Kimi-K3] Make VLLM_ROCM_USE_AITER_MOE_SITUV2 select a4w4/a8w4/a16w4 (#58201)

Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>

* [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL (#50047)

Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Co-authored-by: xijiade.aihemaiti <3146335281@qq.com>

* [Bugfix] V1: clear stale allowed_token_ids mask in InputBatch.condense (#43931)

Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Perf][DSv4.1] Fuse small-batch WO-A with inverse RoPE and MXFP8 quant on SM100/SM103 (#58634)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Attention][CPU] Use zentorch SDPA for CPU MLA prefill (#54967)

Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [CI] Split (H200 MIG 35GB) Spec Decode Speculators + MTP into 4 named jobs (#57237)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Thang Nguyen <thang.nguyen@inferact.ai>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][Perf] Enable layer-aware CSA2 multi-stream overlap for DeepSeek-V4.1-Flash (#57407)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Perf][Pooling] Avoid blocking seq_lens GPU-to-CPU copy for pooling in FlashInfer metadata builder (#57214)

Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>

* [Bugfix] Support repsonse_format + tool_choice=auto (#56086)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: pablopupo <145598901+pablopupo@users.noreply.github.com>
Co-authored-by: hubunt <150658615+hubunt@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Fast Start] Add `/health` endpoint for the weight cache daemon (#58552)

Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Perf][MoE] Use fused MiniMax2 routing with non-unit routed scaling (#58880)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>

* [ROCm][Bugfix] Fall back to default GEMM for CPU tensors on ROCm builds (#58923)

Signed-off-by: fai <fangzhouai@gmail.com>

* [Bugfix][KV Connector] Retry Mooncake bootstrap registration on timeout (reopens #55763) (#58919)

* [Frontend] Switch Python Harmony dependency to oss-harmony (#55128)

Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm] Bump AITER to v0.1.23 (#58867)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Count Responses reasoning tokens per tool round (#58927)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [KimiViT][Perf] Fuse per-layer QK RoPE into one in-place kernel (#58651)

* [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models (#58239)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [Bugfix][Frontend] Sample batched chat completions from the adjusted requests (#58929)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Use a fresh parser per choice in non-streaming chat completions (#58939)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][EPD] Skip sampling for encoder-only async steps (#58490)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [CPU] Build CPU wheels on Ubuntu 22.04 with AMX-FP8 support (#58515)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>

* [Test][ROCm] Stabilize the mixed OLMoE LoRA test (#58945)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [Model] Enable LoRA support for RobertaForSequenceClassification (#58884)

Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Co-authored-by: jz-yolo <jz-yolo@users.noreply.github.com>

* [Bugfix][Frontend] Apply Harmony adjust_request in batched chat completions (#58958)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Minimax-M3] Add Encoder CUDA graph support (#58673)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <codex@openai.com>

* [Perf][MRV2] Reuse Mamba/GDN metadata across KV cache groups (#58762)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Xin Yang <xyangx@amazon.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [Bugfix] Fix the two multimodal root tests that fail on main (OpenPangu-VL embed merge, MiMo sink test fixture) (#58900)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Benchmark] Add Responses API backend to vllm bench serve (#54628)

Signed-off-by: QHarshil <harshil_c@hotmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [CI] Allowlist-shrink batch 1: wire 17 root-level tests + drop 5 stale watermarking entries into misc.yaml (#58055)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [CPU] Use accelerator memory API in DiffusionGemma (#58964)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [CPU] Add video inferencing via torchcodec on s390x (#58693)

Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Skills] Update kernel-microbenchmark to include ROCm (#58646)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Misc] Name each backend and its kernel block sizes in block-size errors (#58557)

Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: khluu <khluu000@gmail.com>

* [CI] Split Dynamic Shapes out of (H200 MIG 35GB) PyTorch Compilation + (MI300) mirror (#58451)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Release the profiling KV cache held by QSA key views (#58961)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CPU] Include vLLM Recipes tooling in release image to deploy models using vLLM Recipes (#58796)

Signed-off-by: louie-tsai <louie.tsai@intel.com>

* Revert "[CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)" (#59011)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Perf][Qwen4Exp] Fuse HC down projection and SiLU on NVIDIA (#58957)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix][Model] Gemma4: register aliased embedding scalars as buffers (#54213)

Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>

* [Bugfix][Spec Decode] Implement get_top_tokens() on the ROCm DeepSeek V4 MTP drafter (#57568)

Signed-off-by: BaoYunkai <ybao@amd.com>

* [Perf][Qwen3.8] Reduce PLE metadata construction overhead (#58114)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm][Refactor] Move DeepSeek-V4/V4.1 multi-stream overlap gate to ROCm platform (#58983)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Multimodal] Avoid extra d2d for encoder cudagraph with fused input norm (#56711)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Logging] Preserve application log record factories (#58747)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix] Use the correct repository revision for secondary artifact loaders (#57461)

Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Perf] Kimi-K3 Enable sharded latent MoE up-projection under EP (#54956)

Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Feature] Add fixed-token prefill scoring (#54335)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [LoRA] Support variable num_labels for sequence classification (#57766)

Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [ROCm][Perf] Replace torch.topk in DSA candidate block selection (#58208)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [ROCm]Keep LMCache OpenTelemetry on the image's 1.40 stack (#59056)

Signed-off-by: Micah Williamson <micah.williamson@amd.com>

* [Bugfix][Kimi-K3] Refresh DSpark context KV cache pointers after the KV cache is re-bound (#58814)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [KVConnector][NIXL] Support packed MLA KV layouts in pipeline-parallel push prefill (#50499)

Signed-off-by: zixi-qi <zixi@inferact.ai>

* [Quant] Use canonical N-first weight format for CT WNA16 MoE (#52798)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: HDCharles <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Relax packed_qk_rope_ correctness test to one ULP (#59008)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Pin OpenTelemetry to LMCache's cap in the ROCm images (#59051)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Increase timeout for Entrypoints Unit (#59076)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][MLA] Add an AITER ASM round-robin decode route for DCP multi-token verify (#56861)

Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix] Don't sync-police or retry FlashInfer all-reduce workspace creation in eager mode (#58498)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix resumable request + async scheduling handoff race (#58259)

Signed-off-by: Yifan Zong <yzong@redhat.com>

* [Model Runner V2][Spec Decode] Support spec decode with draft model (#43091)

Signed-off-by: Icey <1790571317@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][MRV2] Allow FULL decode graphs for one-token prompt tails (#58400)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Bugfix][Scheduler] Preserve logprobs across streaming continuations (#57447)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>

* [Mypy] Fix mypy typing for Voxtral and vision models (#58251)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Core] Rework scheduler `skipped_waiting` queue (#58947)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Bugfix][MRV2][Spec Decode] Reject draft slots that were never proposed (#58784)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][CI] Add missing test coverage for upstream parity (#50519)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][Scheduler] Refresh max tokens for streaming continuations (#57676)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Spec Decode] Avoid triton recompiles in the acceptance estimator (#57107)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>

* [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8664f67d8ae1538e5e5cec2b1ce3f258dd)

* [XPU][CI]Skip test_abort_timeout_on_prefiller in nightly (#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db1ca080793dc2351144cb39513fd7984ca)

* [Bugfix][Engram] Keep THP tables private when resolving shared memory (#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c352f079c2cb8f46752fd0317a2745050bf)

* [ROCm][CI] Expand MI355 mirrors and route MIG-sized jobs to DPX (#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd64f5ef75dbe1471ee88c009003b3d03ec)

* [Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse retention (#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bddbeab6b4a0d5861dfcb171bf61ce2109d6)

* [Core][BugFix] Tag prefix-cache extra keys by source (#51899)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e7ed5b6f372f0898a67e48d3365fa28139)

* [Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e097b905127b1d5f09a7e93bd1095e87a6c)

* [Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf53724e06cef7f07628ae8fd4cc6c96f11a)

* [Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f668608155bc8a8c74cb97e3ddd0d3053085)

* [Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc757cffc11b55689d9cc33d844d9606516)

* [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a4297921c699f9aa58f9c6395e95ede718)

* [Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa130fac1e1252330a680fbed45ad05eba64f)

* [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)

(cherry picked from commit 5463fe4962785cdc3383477bf3af6533e7647dfd)

* [Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87cca25690fef6bd12667fd0d26690d949af)

* [Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec22ad9bb098393021b956feaa97081fc85)

* [Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae4d746f64c22677348d1cc120eea6f7439)

* [Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e615c20ee4ff0cd82aac592c7d6310aa7c3)

* [Core] Include the LoRA path in prefix-cache block hashes (#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e075ed461cff922aa710d387281f651940)

* [Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360c9eebb3b4b460ec49f6606ef0a08fcae2)

* [CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d7332db897885fdc0e6ed61b28c969d743)

* [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba9b7d18f9746d989e389a51a94b0f96f6f)

* [Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a6963664537ed21172e2ec12e96e3a2dcd3718c)

* [Bugfix][HiSparse] Size the KV cache from the groups HiSparse allocates (#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2e4489126f1c21868699f3890021f01b73)

* [Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with a saturated GPU pool (#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8ac1f8a7e8fb50cf1c56ff96649fcadccc6)

* [CI] Drop a test that depends on #57930 from the #59309 backport

Resolving the #59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from #57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream #59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>

* [Bugfix] Fix minimax-m3 multimodal processor compatability with Transformers v5.18 (#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f160a2c0abcb5902acc3c91a14c38a4af173)

* [Misc] Add Transformers version upper bound in requirements (#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298457dde7b4554b3b4e20b238c0ac2c3a65)

* Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0

Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2eb #28, 76a3e6d72, 6d4c75611 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a46, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 27.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner
  - per-group KV pool for hybrid models, re-expressed on v0.31's single
    backing allocation: the O(1) Mamba/GDN state groups get their own
    BlockPool sized for max_num_seqs, a region of the allocation after the
    attention pool; their pages are not unified with attention pages;
    per-pool admission, events and usage; CoW copies and warmup/profiling
    block ids per pool; VLLM_SAMPLER_RESERVE_MIB headroom

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c19, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - padded-page block fill (the split pool removes the page unification
    it worked around), drain hook / on_kv_manager_created, KVBlockZeroer
    clamp: no turbo-attn consumer
  - rotary fast-path registry: upstream guards the import (1f60771c7,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Signed-off-by: chao.huan <chao.huan@nio.com>
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: errmakov <ide404@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Signed-off-by: yashasvi <yashasvi@ibm.com>
Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Signed-off-by: Cheng Rui <286040359@qq.com>
Signed-off-by: Yifan Zong <yzong@redhat.com>
Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Linkun Chen <github@lkchen.net>
Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Ke Wen <kwen@nvidia.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: zixi-qi <zixi@inferact.ai>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: Willian <willian@willian.email>
Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Vincent <vincexxchan@gmail.com>
Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: QHarshil <harshil_c@hotmail.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Yannick Schnider <Yannick.Schnider1@ib…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants