Skip to content

[CPU] Gate the AVX10.2 paths on compiler support - #58133

Merged
bigPYJ1151 merged 5 commits into
vllm-project:mainfrom
ganeshr10:cpu-avx10-2-gcc-guard
Sep 24, 2026
Merged

bigPYJ1151 merged 5 commits into
vllm-project:mainfrom
ganeshr10:cpu-avx10-2-gcc-guard

Conversation

@ganeshr10

@ganeshr10 ganeshr10 commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fixes #58363.

cmake/cpu_extension.cmake requires gcc/g++ >= 12.3 for the x86 backend, but csrc/cpu/sgl-kernels has used AVX10.2 since #49942, and AVX10.2 landed in GCC 15. On anything older the build does not lose a fast path, it fails outright:

common.h:537: error: parameter to builtin not valid: avx10.2
gemm_fp8_w8a8.cpp:775: error: attribute 'target' argument 'avx10.2' is unknown
gemm_fp8_w8a8.cpp:777: error: '_mm256_cvtph_hf8' was not declared in this scope

Five objects fail, so a source build on any current distro toolchain is broken — Ubuntu 24.04 (gcc 13), Debian 12 (gcc 12), RHEL 9 (gcc 11). docker/Dockerfile.cpu is unaffected because it installs gcc-15 from a PPA, which is why CI does not see this.

This probes for -mavx10.2 and defines VLLM_CPU_HAS_AVX10_2 accordingly, guarding the two target("avx10.2") helpers and their two call sites. The AVX512 loops directly below each call site already handle the work. On GCC 15+ nothing changes.

The probe mirrors the -mamx-fp8 handling a few lines above, which gates "on actual flag support rather than compiler version alone" — a version check would also have been wrong here, since AVX10.2 is 15 while AVX10.1 is 14.

Test Plan

Build main plus this patch for CPU on GCC 12.4, and confirm the probe result reaches the guarded translation unit.

Test Result

-mavx10.2 rejected by GCC 12.3, 12.4 and 13.4. Configure reports:

-- Performing Test COMPILER_SUPPORTS_AVX10_2_FLAG - Failed
-- AVX10.2 disabled: compiler does not support -mavx10.2 (compiler: GNU 12.3.0)

In the generated compile_commands.json, gemm_fp8_w8a8.cpp appears exactly once, under the _C target, with no CPU_CAPABILITY_AVX10_2 but with CPU_CAPABILITY_AMXBF16 present. So the target-scoped define does reach the guarded unit, and _C_AVX512/_C_AVX2 never compile it — which is why PRIVATE on _C is sufficient.

Full CPU build on GCC 12.4 completes with no errors: vllm 0.29.1rc1.dev527+g5fc8113af, torch 2.13.0+cpu, and import vllm._C succeeds. Unpatched main fails the five objects above on the same machine.

The AVX10.2 target attribute, the fp8 convert intrinsics and the matching
__builtin_cpu_supports argument all arrived in GCC 15, so on the GCC 12.3+
cpu_extension.cmake advertises they are hard errors rather than a missing
fast path. Probe for -mavx10.2 and fall back to the AVX512 loops.

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Change-Id: Icddc900a9efc9b64cc7d966e7021f2c487a2d249

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Comment thread cmake/cpu_extension.cmake
Define CPU_CAPABILITY_AVX10_2 alongside CPU_CAPABILITY_AMXFP8 instead of
adding a directory-wide definition, so it reaches only the target that
carries the sgl-kernels sources.

Signed-off-by: R <Ganesh.R@amd.com>
Change-Id: I29184d8225c83e5b0e5cf927802ec40a51ae2784
Comment thread cmake/cpu_extension.cmake Outdated
The previous commit dropped four spaces from the AMX-FP8 block while adding
the AVX10.2 one beside it, putting an unrelated whitespace hunk in the diff.
Both now match the file's existing style.

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@bigPYJ1151

Copy link
Copy Markdown
Member

/ci run

@github-actions

Copy link
Copy Markdown

❌ This PR is 1 commit behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@bigPYJ1151

Copy link
Copy Markdown
Member

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90890 for commit 01d15ae65769.

@bigPYJ1151
bigPYJ1151 merged commit f34a0e0 into vllm-project:main Sep 24, 2026
47 checks passed
kristobalus added a commit to kristobalus/vllm that referenced this pull request Sep 25, 2026
* [Bugfix] Fix external LB DP rank handling when replicas share nodes (#53743)

Signed-off-by: Tony Lin <tony.lin@intel.com>

* [docs] Fix legacy hf CLI references (vllm) (#57958)

Signed-off-by: Wauplin <lucainp@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NIXL] Fix DCP pulls across MLA cache regions (#57389)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [ROCm] Refactor tuned gemms (#55001)

Signed-off-by: Andy Friedrich <afriedri@amd.com>
Signed-off-by: afriedri <afriedri@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] unskip InternViT test for transformers v5 compatibility (#55767)

Signed-off-by: sahil <sahil@example.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [MM] Move get_dummy_processor_inputs into MM processor (#57967)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [ROCm][CI] Use ROCm backend for DeepSeek V4.1 ViT test (#57931)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Feature] Add first-class KV hints request envelope for programmatic KV management (#53423)

Signed-off-by: Karen Chung <karenc@nvidia.com>

* [Docs] Add an Engram feature page explaining Engram usage in vLLM (#57910)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [ROCm][Perf] Route the fused shared-expert gate GEMM through the platform dispatcher (#54185)

Signed-off-by: Mikko Tukiainen <Mikko.Tukiainen@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Engram] Drop redundant VLLM_PLE_CPU_OFFLOAD env var (#57937)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [Bugfix][MoE] Reject hash routing for unsupported monolithic backends (#57867)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][ROCm] Reject unsupported EP for monolithic AITER MXFP4 MoE (#57866)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Docker] Use zstd for CI images and offer a Docker Hub variant (#55608)

Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Kimi-K3][AMD] Return KDA and MLA projection outputs directly (#50592)

Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Co-authored-by: Liuyinfeng01 <199041580+LiuYinfeng01@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][Build] Harden triton-cpu sleef submodule fetch in CPU image build (#57871)

Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>

* [Bugfix][Kernel] Skip the fused silu-mul block-quant fast path when a swiglu clamp is set (#57984)

Signed-off-by: Garrett Goon <garrett@primeintellect.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Frontend] Add reusable TP1 initialized-engine snapshots (#51360)

Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: elehayym <52448798+Yuzu23@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] MRV2 pooling shutdown model ref (#57737)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [CI] Split (H200) LM Eval Large Models into per-model jobs (#57965)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Scheduler] Soften Long Prefill Tokens Threshhold (#57951)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Perf][Attention] Reduce GLM sparse MLA preparation overhead (#57458)

Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>

* [Bugfix] Annotate MTP draft KV cache groups positionally on the hybrid grouping path (#55390)

Signed-off-by: Navjot Singh <navjot.singh@shopify.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][GDN] Fix stateless first-chunk classification (#51565)

Signed-off-by: taking-lying-flat <1615405@qq.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com>

* [CI] Add pre-commit check that new tests are tethered to Buildkite jobs (#54867)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Turner <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [XPU] Fix Nemotron FP8 LM-eval config: drop CUDA-only moe_backend and wire to new Buildkite job (#49685)

Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [Docker] Expose bundled vllm-rs on PATH (#57606)

Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Signed-off-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>

* [Fast Start] Cache the MTP draft model in a separate daemon group (#57312)

Signed-off-by: liusy58 <mg21330037@smail.nju.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Rust Frontend] Add MiMo V2.5 parser support (#57933)

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Bugfix][Spec Decode] Cap DFlash/DSpark profiling query batch (#56448)

Signed-off-by: wangyicong <wangyicong@bytedance.com>

* [RL][Sleep] Retain frozen weights across level-2 sleep (#57891)

Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>

* [ROCm][Bugfix] Explicitly reject FSE=1 with DPA+ETP deployment for DeepSeek-V4 (#57919)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Spec Decode] Enable async scheduling for DFlash (#58065)

* [Frontend][Rust] Add mm-processor benchmark for Rust frontend (#51922)

Signed-off-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Signed-off-by: gangula-karthik <gkarthik923@gmail.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Rust Frontend] Introduce parser-owned output grammar interfaces (#55269)

Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Security] Reject min_tokens that exceeds the filled max_tokens default (#57731)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>

* [Bugfix][Frontend] Validate mixed prompt embedding mask lengths (#57006)

Signed-off-by: 子华 <huaxi.shx@alibaba-inc.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Qwen2.5-VL] Honor video fps for temporal M-RoPE (#47736)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Rust Frontend] Build full-output grammars from initialized reasoning parsers (#57340)

Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Codex <noreply@openai.com>

* [ROCm][Perf] Avoid extra reshape kernel in Qwen GDN output norm (#47842)

Signed-off-by: Mikko Tukiainen <Mikko.Tukiainen@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Kernel] Add opt-in load-time MXFP4 dequantization (#50814)

Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>

* [Rust Frontend] Separate multimodal instrumentation from request timing (#58084)

Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Kernel][DSV4.1] Fuse MXFP8 wo_b GEMM with sequence-parallel reduce-scatter (#57428)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: Canlin <canlinguosdu@gmail.com>
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [CI] Emit a kernel symbol map from the csrc build (opt-in, for test selection) (#58097)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [perf] wire FA and FlashMLA for sm90 GLM5Next NoPE SparseMLA (#55385)

Signed-off-by: JaredforReal <w13431838023@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [CPU] Add device-memory-utilization CLI alias (#56547)

Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Louie Tsai <louie.tsai@intel.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Rust Frontend] Construct model-owned vision processors through specs (#58109)

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [BugFix][Core] Make the structured-output grammar poll non-blocking (#55931)

Signed-off-by: ubwzwd <ubwzwd@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com>

* [XPU][CI]Remove model_runner_v2 test from Intel GPU CI (#58050)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Bugfix][Structured Outputs] Reject empty `structural_tag` at request validation (#47450)

Signed-off-by: linnea-lin-00638949 <15521435947@163.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com>

* [Build] Fix CUDA 12 KV connector dependency selection (#57945)

Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>

* [Bugfix][GLM-5.3-Flash] Run the dense MLP layers on the sequence-parallel shard (#58061)

Signed-off-by: Jared Wen <w13431838023@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Structured Output] Disallow MRV1 + PP>1 + async sched + structured output (#56250)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: CNE Pierre FICHEPOIL <pierre-1.fichepoil@gendarmerie.interieur.gouv.fr>

* [Feat][XPU] VLLM_BATCH_INVARIANT support for Dense/MoE models (#55881)

Signed-off-by: Tony Lin <tony.lin@intel.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [SpecDecode] Restore residual-logits comments in _resample_kernel (#58166)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Narrow AuxOutput KV restrictions to known PD connectors (#58150)

Signed-off-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix] prioritize architecture capability before DeepGEMM availability check (#58073)

Signed-off-by: Tony Lin <tony.lin@intel.com>

* [Bugfix][Attention] Avoid NaN in the Triton softcap for large attention logits (#56579)

Signed-off-by: Kushal Dabbe <72650064+kushaldabbe@users.noreply.github.com>
Co-authored-by: opencode <noreply@opencode.ai>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][Engram] Fall back when /dev/shm is absent before sharing tables (#57914)

Signed-off-by: Juntian Liu <juntianl@inferact.ai>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] hadacore_transform: respect inplace parameter to fix garbage outputs with QuIP transforms (#43462)

Signed-off-by: Gilles Turpin <turpingilles15@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Bugfix][ROCm] Dispatch the QuantFP8 CUDA fallback on the class (#58136)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][DSv4.1][Perf] Fuse the inverse RoPE into the sparse decode reduce (#57435)

Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Fast Start] Support data parallelism in the weight cache daemon  (#57386)

Signed-off-by: liusy58 <mg21330037@smail.nju.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix] batch_invariant: keep non-AllReduce collectives enabled on NCCL >= 2.31 (#58179)

Signed-off-by: Guanxin Li <38149783+guanxingithub@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [MRV2] Release weight offloader on shutdown (#57834)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][V1] Honor enable_jit_warmup for V2 kernel warmup (#55146)

Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm][CI] Add GELU activation for AiterExperts in the modular-kernel coverage (#58030)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][V1] Read ModelState max_model_len from model config (#58149)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Model][Spec Decode] Defer disposable GLM MTP head (#55442)

Signed-off-by: Luca Motz <luca.motz@icloud.com>

* [Refactor] Remove dead code multiple places (#58002)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Core] structured generation mode for DiffusionGemma model (Jev-like) (#57250)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Razorback16 <razorback16@protonmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Razorback16 <razorback16@protonmail.com>

* [Kernel] Remove AllSpark INT8 W8A16 GEMM backend (#58001)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [CI] Build the torch-nightly image on Ubuntu 24.04 (#58204)

* [Bugfix] Backport Inductor custom-op pattern matching fix (#58189)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Core] Disable JIT warmup in eager mode (#58197)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [CI] Split LM Eval TurboQuant KV Cache into per-config jobs (#57113)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.ai>

* [CI] Shard (H200 MIG 18GB) Spec Decode Draft Model across whole-directory replicas (#58193)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Perf] Remove CPU-GPU sync in heterogeneous vocabulary speculative decoding (#57396)

* [ROCm][Build][The Rock] Bump Triton version to 3.8.x tip-of-tree with source build in The Rock image (#58006)

Signed-off-by: Randall Smith <Randall.Smith@amd.com>

* [Bugfix][ROCm] Use the platform FP8 range in the concat MLA q test (#58153)

Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][Model][Bugfix] Enable GLM-5.2-MXFP4 on the deepseek_v32 path and fix sparse attention correctness (#51915)

Signed-off-by: Jack Hu <Jack.Hu@amd.com>
Signed-off-by: Jack Hu <jack.hu@amd.com>
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com>
Co-authored-by: James E T Smith <jamesETsmith@users.noreply.github.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Douglas Lehr <Doug.Lehr@amd.com>

* [ROCm] Use silu_and_mul_with_clamp's torch._C op (#52052)

Signed-off-by: Tres Popp <tres.popp@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix] Skip VllmConfig re-validation for with_hf_config submodel views (#58212)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Roger Wang <rogerw@inferact.ai>

* [EPD] Support metadata-only audio inputs (#57887)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [ROCm][DSv4][Perf] Fuse the inverse RoPE into the sparse decode reduce (#57451)

Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>

* [ROCm][Compile] Support BF16 AsyncTP fusion (#58098)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [CI][Bugfix] Update IPC test caller for #57312's _apply_entries signature (#58107)

Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [ROCm][CI] Stage G gating (#50922)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Bugfix] Set worker runtime threads before profiling and compilation (#55891)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [XPU] Wire up SYCL apply_rotary_emb kernel in ApplyRotaryEmb (#55721)

Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com>
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Compilation] Fix QuTLASS compilation with PyTorch 2.13 (#58173)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Spec decode] Support variable-length decode for Kimi-K3 adaptive ver (#52988)

Signed-off-by: Albert Cheng <albecheng@nvidia.com>
Signed-off-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-bia01.bia.clusters.nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [XPU][CI] Deselect tests/v1/spec_decode/test_mtp.py::test_glm_mtp_defers_lm_head (#58237)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>

* [MoE] Use GateLinear for all MoE models (#58234)

Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>

* [Bugfix][Frontend] Keep length finish_reason for max_tokens-truncated streaming tool calls (#46303)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [ROCm][Perf] Use wvSplitK for single-output GEMMs (#53283)

Signed-off-by: tangzzycc <3081129260@qq.com>

* [Tests] Select V2 for diffusion scheduler unit tests (#58272)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Quantization] Select per-token NVFP4 MoE backends explicitly (#57176)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the three TurboQuant evaluation groups on MI355 (#58282)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Mooncake] Address review nits from #56855 (#57174)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Yifan Qiao <17067717+ivanium@users.noreply.github.com>

* [Refactor][Quantization] Make FP8 and MLA weight transforms reusable pure functions (#57732)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [ROCm][Compile] Fuse AITER static FP8 attention output (#58099)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [ROCm][Bugfix] Register MRV2 sampler JIT warmups (#58092)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Attention] Avoid CPU-GPU sync in DCP sequence lengths (#58169)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Extend groupwise rms_norm scale tolerance to CUDA (#58252)

Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>

* [Perf][ROCm][Attention] Narrow the Triton prefill-attention KV tile on RDNA3/RDNA4 (#58225)

Signed-off-by: Jipeng Li <jipengli@amd.com>
Co-authored-by: GitHub Copilot CLI <noreply@github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][Bugfix] Keep zero MiniMax MXFP8 activation blocks finite (#58089)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Include Python tooling in ROCm CI artifacts (#58271)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [MRV2] Miscellaneous code cleanup (#57980)

* [Bugfix][SM120][MLA] Support NoPE sparse MLA (GLM-5.3-Flash) on the FlashInfer SM120 backend (#55277)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [XPU] upgrade to PyTorch 2.14 (#56013)

Signed-off-by: Yan Ma <yan.ma@intel.com>
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [ROCm][Test] Check GDN prefill numerics and output ownership (#58091)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix] Disable prefix caching for encoder-only before model config hooks (#58287)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [Bugfix][Tool Parser] Migrate Granite to the streaming Parser Engine (#49648)

Signed-off-by: Nikhil Kulkarni <nikhilkulkarni1755@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [DSpark] Support pipeline-parallel targets in aggregated serving (#56956)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Feature][Frontend] Add DeepSeek-V4 FIM completion rendering (#44229)

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [Perf][MoE] Skip top-k slots routed to non-local experts in TritonExp… (#58051)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Shuolei Wang <948904026@qq.com>

* [CPU] Adds support for fp32 attention sinks (#56252)

Signed-off-by: Ankit Jaiswal <ankit.jaiswal@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>

* [Quantization] Enable humming wNaM asymmetric quant (zero_point) with compressed-tensors (#46528)

* [Quantization][Bugfix] Bump humming-kernels to 0.1.16 (#58054)

Signed-off-by: jinzhen.ljz <jinzhen.ljz@antgroup.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Quark] Remove quark-specific silent online quantization (#51800)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm][Perf] Extend QK-norm/RoPE/KV-cache fusion to MRoPE (#50212)

Signed-off-by: Vorapol Assavasangthong <Vorapol.Assavasangthong@amd.com>
Co-authored-by: Santosh Hiremath <Santosh.Hiremath@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>

* [ROCm][CI] Validate Mooncake and NIXL prefill/decode accuracy (#58095)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [CI][ROCm] Add an MI355 Kimi-K3 unit test group (#58012)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][NIXL] Restore successful push completion reporting (#58188)

Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][ROCm] Fix startup OOM in AITER MLA FP8 prefill workspace sizing (#57923)

Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>

* Doc: add DiffusionGemma to supported models (#46466)

Signed-off-by: Bruce <Bruce798858117@gmail.com>
Signed-off-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf] Use breakable CUDA graphs (no torch.compile) by default under VLLM_BATCH_INVARIANT so the tuned matmul configs see the runtime M (#57586)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>

* [CI] Select one GPU for the H200 initialized snapshot E2E step (#58351)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Pick a KV block size supported by every attention backend (#49845)

Signed-off-by: Divy <divy@coralbricks.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Docs] Fix docstring typos (output_dytpe, kwrags, Abbrivations) (#55936)

Signed-off-by: simpleqt <89645338+simpleqt@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Test] Cover MoRI graph replay and output lifetime (#58093)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][Bugfix] Fix TileLang mHC fused RMSNorm on 64-wide wavefronts (#58419)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][Bugfix] Limit MRV2 sampler JIT warmup registration to ROCm (#58465)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Disable JIT warmup by default in VllmRunner (#58452)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][MRV2] Align dummy idx_mapping dtype to avoid runtime jit (#58462)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [5/12][ci-selector][CI] Skip the Proton GPU test when another CUPTI tool is injected (#58455)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Share BF16 baselines across quantization comparison tests (#58469)

Signed-off-by: Aarushi Jain <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][DSA] Bound DeepSelect sentinel columns in the sparse top-k remap (#58215)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Test][Determinism] Cover chunked prefill in the batch-invariance suite (#55612)

Signed-off-by: Bob Ok <49168652+blipbyte@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [CI][ROCM] Add the Fusion E2E TP2 Quick group on MI355, and the AITER MLA fix it needs (#58369)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [Scheduler] Tune --long-prefill-token-threshold adaptiveness (#58459)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>

* [PCP] Support prefill context parallelism with data parallelism (#57075)

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Give turboquant boundary layers a layout-compatible backend (#54988)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Resubmit PR 48666 - Gemma4 FP8 KV FA4 head dim 512 backend selection (#53175)

Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>

* [Bugfix][KV Offload] Retain offload event metadata through batch translation (#57453)

Signed-off-by: Kapil Arya <kapila@nvidia.com>
Signed-off-by: Kapil Arya <kapil.arya.17@gmail.com>
Co-authored-by: Or Ozeri <or@ozery.com>

* Fix full logprobs in token-in/token-out responses (#58488)

Signed-off-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][CPU][MoE] Fix out-of-bounds write and segfault when router weights are fp32 (#56168)

Signed-off-by: Farzad Abdolhosseini <farzad@elastix.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Rust Frontend] Recognize new frontend-owned serve args as unsupported or no-op (#58330)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Rust Frontend] Accept custom chat roles for HF templates (#58311)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Dependency] Upgrade FlashInfer version to 0.7.0 (#58069)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Bugfix][Spec Decode] Separate DSpark width from MTP stage validation (#54631)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Rust Frontend] Pass vision preprocessing context for Nemotron-H (#57634)

Pass the remaining engine context-length budget to model-owned vision processors through VisionPreprocessingContext. Preserve Nemotron batched engine fields and recognize llm_config as a text_config alias.

Use the merged upstream llm-multimodal revision f0985ef65967615db2c79279aa07818499301bfd.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Rust Frontend] Support `--sse-keep-alive-interval` (#58306)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [Compile][CI] Honor Triton cache overrides and add AMD timeout headroom (#58474)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [CPU][GDN] Support NIXL DS convolution-state layout (#53300)

Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [ROCm][CI] Add the MI355 TurboQuant t3nc mirror (#58432)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [EPD][Model Loader] Skip language-model checkpoint shards for `--mm-encoder-only` (#58086)

Signed-off-by: grYe99 <guorongye99@gmail.com>
Co-authored-by: grYe99 <guorongye99@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [Perf] Use Conv3dLayer for MiniMax M3 patch embedding (#58512)

Signed-off-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Feature][Frontend] Request JSON body debug logging on `--enable-log-requests` flag (#58163)

Signed-off-by: talora <talora@nvidia.com>

* [CPU] Use pre-built triton (#58140)

Signed-off-by: jiang1.li <jiang1.li@intel.com>

* [Bugfix][V1] Reject encoder-cache hits with mismatched embedding counts (#57696)

Signed-off-by: jackLei0901 <42642542+jackLei0901@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Capture prefill kernels for mixed FULL graphs (#58275)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Bugfix][CI] Fix the flaky sharded-sampling tests, and the engine teardown need (#58342)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

---------

Signed-off-by: Tony Lin <tony.lin@intel.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Andy Friedrich <afriedri@amd.com>
Signed-off-by: afriedri <afriedri@amd.com>
Signed-off-by: sahil <sahil@example.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Karen Chung <karenc@nvidia.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Mikko Tukiainen <Mikko.Tukiainen@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Signed-off-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Signed-off-by: Garrett Goon <garrett@primeintellect.ai>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Navjot Singh <navjot.singh@shopify.com>
Signed-off-by: taking-lying-flat <1615405@qq.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Turner <doubleujabbour@gmail.com>
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Signed-off-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Signed-off-by: liusy58 <mg21330037@smail.nju.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Signed-off-by: gangula-karthik <gkarthik923@gmail.com>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: 子华 <huaxi.shx@alibaba-inc.com>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: Canlin <canlinguosdu@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: JaredforReal <w13431838023@gmail.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Louie Tsai <louie.tsai@intel.com>
Signed-off-by: ubwzwd <ubwzwd@gmail.com>
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Signed-off-by: linnea-lin-00638949 <15521435947@163.com>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Signed-off-by: Jared Wen <w13431838023@gmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: Kushal Dabbe <72650064+kushaldabbe@users.noreply.github.com>
Signed-off-by: Juntian Liu <juntianl@inferact.ai>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Gilles Turpin <turpingilles15@gmail.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>
Signed-off-by: Guanxin Li <38149783+guanxingithub@users.noreply.github.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Razorback16 <razorback16@protonmail.com>
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
Signed-off-by: Jack Hu <Jack.Hu@amd.com>
Signed-off-by: Jack Hu <jack.hu@amd.com>
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com>
Signed-off-by: Tres Popp <tres.popp@amd.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com>
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com>
Signed-off-by: Albert Cheng <albecheng@nvidia.com>
Signed-off-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-bia01.bia.clusters.nvidia.com>
Signed-off-by: tangzzycc <3081129260@qq.com>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: Jipeng Li <jipengli@amd.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: Yan Ma <yan.ma@intel.com>
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Signed-off-by: Nikhil Kulkarni <nikhilkulkarni1755@gmail.com>
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Shuolei Wang <948904026@qq.com>
Signed-off-by: Ankit Jaiswal <ankit.jaiswal@amd.com>
Signed-off-by: jinzhen.ljz <jinzhen.ljz@antgroup.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vorapol Assavasangthong <Vorapol.Assavasangthong@amd.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: Bruce <Bruce798858117@gmail.com>
Signed-off-by: Misha Goin <mgoin64@gmail.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Divy <divy@coralbricks.ai>
Signed-off-by: simpleqt <89645338+simpleqt@users.noreply.github.com>
Signed-off-by: Aarushi Jain <Aarushi.Jain2@amd.com>
Signed-off-by: Bob Ok <49168652+blipbyte@users.noreply.github.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>
Signed-off-by: Kapil Arya <kapila@nvidia.com>
Signed-off-by: Kapil Arya <kapil.arya.17@gmail.com>
Signed-off-by: Farzad Abdolhosseini <farzad@elastix.ai>
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Signed-off-by: grYe99 <guorongye99@gmail.com>
Signed-off-by: OpenAI Codex <codex@openai.com>
Signed-off-by: talora <talora@nvidia.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jackLei0901 <42642542+jackLei0901@users.noreply.github.com>
Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: Tony Lin <tony.lin@intel.com>
Co-authored-by: Lucain <lucainp@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: afriedri <afriedri@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Sahil Patel <91423311+Sip4818@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
Co-authored-by: djramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Karen Chung <karenc@nvidia.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Mikko Tukiainen <mikko.tukiainen@amd.com>
Co-authored-by: Nils Matteson <nilsmatteson@icloud.com>
Co-authored-by: yinfengLiu <yinfeliu@amd.com>
Co-authored-by: Liuyinfeng01 <199041580+LiuYinfeng01@users.noreply.github.com>
Co-authored-by: vllm-agent <claw@inferact.ai>
Co-authored-by: kimi-no-na-wa <kimi-no-na-wa@users.noreply.github.com>
Co-authored-by: Garrett Goon <44747910+garrett361@users.noreply.github.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: elehayym <52448798+Yuzu23@users.noreply.github.com>
Co-authored-by: Thang Nguyen <69278249+Thangnguyenvn98@users.noreply.github.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
Co-authored-by: Navjot Singh <navjot.singh@uwaterloo.ca>
Co-authored-by: cherry77-cloud <1615405@qq.com>
Co-authored-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Chaojun Zhang <chaojun.zhang@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: siyu <mg21330037@smail.nju.edu.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Yicong Wang <wangyicong@bytedance.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Jeff (Junze) Ma <93145857+majunze2001@users.noreply.github.com>
Co-authored-by: karthik <56480632+gangula-karthik@users.noreply.github.com>
Co-authored-by: Karthik Gangula <gangula-karthik@users.noreply.github.com>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Co-authored-by: Juan Pérez de Algaba <124347725+jperezdealgaba@users.noreply.github.com>
Co-authored-by: shaohuaxi <huaxi.shx@alibaba-inc.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
Co-authored-by: Jared Wen <w13431838023@gmail.com>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: Louie Tsai <louie.tsai@intel.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: ubwzwd <ubwzwd@gmail.com>
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: xiangdong <40376367+zxd1997066@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: linyafeng <15521435947@163.com>
Co-authored-by: CNE Pierre FICHEPOIL <pierre-1.fichepoil@gendarmerie.interieur.gouv.fr>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Kushal <72650064+kushaldabbe@users.noreply.github.com>
Co-authored-by: opencode <noreply@opencode.ai>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Gilles Turpin <turpingilles15@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: stefankoncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Guanxin Li <38149783+guanxingithub@users.noreply.github.com>
Co-authored-by: pengyihang <1017861497@qq.com>
Co-authored-by: Divakar Verma <137818590+divakar-amd@users.noreply.github.com>
Co-authored-by: hcl <chenglunhu@gmail.com>
Co-authored-by: lucamotz <luca.motz@icloud.com>
Co-authored-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Razorback16 <razorback16@protonmail.com>
Co-authored-by: Andrey Talman <atalman@fb.com>
Co-authored-by: Kimi <noreply@moonshot.ai>
Co-authored-by: Michael Lapshin <55516685+MichaelLapshin@users.noreply.github.com>
Co-authored-by: rasmith <Randall.Smith@amd.com>
Co-authored-by: Jack Hu <jack.hu@amd.com>
Co-authored-by: James E T Smith <jamesETsmith@users.noreply.github.com>
Co-authored-by: Douglas Lehr <Doug.Lehr@amd.com>
Co-authored-by: Tres <tpopp@users.noreply.github.com>
Co-authored-by: Roger Wang <rogerw@inferact.ai>
Co-authored-by: Tianyu Guo <guoty@inferact.ai>
Co-authored-by: Michał Ganczarenko <michal.gancz…
ganeshr10 added a commit to ganeshr10/vllm that referenced this pull request Sep 28, 2026
Catches up 227 commits, including the AVX10.2 compiler gate (vllm-project#58133) that
supersedes the local workaround this branch was carrying. No conflicts.

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Change-Id: I6f489663afc326fc716e7bd70b906ceea29a183c
arbi-dev added a commit to arbicity/vllm-turbo that referenced this pull request Oct 7, 2026
* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [SpecDecode] Add LiLiCorr drafter (#57934)

Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>

* [MRV2] Minor model_runner.py code cleanup (#58610)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix generative scoring body cancellation (#57729)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] Preserve reranker tokenization with document limits (#57666)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Kernel][Perf] Register-resident path for per-token-group 8-bit quant (#55330)

Signed-off-by: chao.huan <chao.huan@nio.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf][Kernel] Vectorized flat abs-max for dynamic per-tensor FP8 quantization (#58194)

Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Core][Logging] Fix JSON logging process decoration (#57957)

Signed-off-by: Mark McLoughlin <markmc@redhat.com>

* [Bugfix] Stop allocator fragmentation from shrinking the KV cache during memory profiling (#58430)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Keep logprobs of parser-suppressed streaming chunks (#58583)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Yanxiao Zhao <39199723+sdpkjc@users.noreply.github.com>
Co-authored-by: Prakhar Agarwal <270064960+agarwalprakhar2511@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix][GLM-5.3-Flash] SM90 sparse MLA: index_kpool mismatch leads to corruption via unread query token (#58704)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Benchmark] Record model_id in bench latency/throughput --output-json (#58112)

Signed-off-by: yashasvi <yashasvi@ibm.com>

* [Bugfix][GLM-5.3-Flash] kpool corruption with speculative decoding (#58454)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [GLM5.3 Perf] Optimize glm 5.3 metadata op, 1.6~4.8x kernel level performance improvement (#58450)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Bugfix][KV Connector] Reap expired NIXL leases behind a heartbeated head (#58292)

Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>

* [ROCm][Kimi-K3] Optimize low-concurrency speculative KDA (#58045)

Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Cover the AITER MQA logits dispatch on gfx950 (#58724)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Add quantized MoE serving test for gfx950 (#58748)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Test AMD DeepSeek V4 MoE routing against a PyTorch reference (#58740)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: one-pass sampler statistics kernel (#58226)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [ROCm][CI] Expand single-GPU coverage on MI355 DPX (#57599)

Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [CI] [MRV2] Restore MRV2 pp dp coverage (#57735)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][CI] Report subprocess test skips as skips, not passes (#58701)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Run the MLA attention+quant fusion test on ROCm (#58717)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm] Bump torch 2.13, triton 3.8, torchaudio, torchvision (#50605)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: jpvillam <juan.villamizar@amd.com>

* [Feature][Frontend] Add granite_thinking_parser reasoning parser for Granite 4.2 (#55957)

Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Respect max_output_tokens in the Harmony tool-call loop (#58551)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Du Bin <8174807+dubin555@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Frontend] Document 404 response for `/generative_scoring` (#58788)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Perf][Frontend] Defer reasoning usage recounts for non-continuous chat streams (#56067)

Signed-off-by: Cheng Rui <286040359@qq.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] Default missing detail for Responses API input images (#57241)

Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>

* [watermarking] golden tests for backwards compatibility (#56809)

Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Don't drop the rest of the allocator config when toggling expandable segments (#57982)

Signed-off-by: Oxana Korzh <okorzh@amd.com>

* [AuxOutput] Only require Model Runner V2 on GPU platform (#58205)

Signed-off-by: Linkun Chen <github@lkchen.net>

* [Pooling] Preserve BERT-family heads for raw logits (#57664)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* fix: perf: use startswith(x, i) instead of string slicing to avoid O(N^2) (#52580)

Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states (#48419)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Co-authored-by: Claude <noreply@anthropic.com>

* [Bugfix][Reasoning] Count Kimi K3 reasoning tokens (#58372)

Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [CI] Split (H200 MIG/MI300) Basic Correctness into named jobs (#57054)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Frontend] Fix Inkling tool name leaking into content after reasoning (#58792)

Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Core] Bound draft-token RPC waits by the execute-model timeout (#58779)

Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Co-authored-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>

* [Perf][DSv4.1] Shard the Engram wkv projection across TP ranks (#58678)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>

* [RL] Add sharding-aware NCCL M2N weight transfer (#51520)

Signed-off-by: Ke Wen <kwen@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][ROCm] AMD-Quark mixed-precision DeepSeek-V4.1 support (#57071)

Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix][Frontend][Rust Frontend] Update DeepSeek V4.1 Flash reasoning effort mappings (#58316)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>

* [CI] Only isolate the registry tests that need a fresh process (#58764)

Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Spec Decode] Enable Gemma4 DSpark adaptive verification with FlashInfer (#57263)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [CI] Split (B200) Miscellaneous Kernels into mHC, FLA Ops and Misc named jobs (#58609)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Security] Gate per-request multimodal processor kwargs (#58830)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Elastic EP] Fix EPLB load statistics during scaling (#58473)

Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>

* [Refactor] Remove dead tests utils (#58803)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights (#58720)

Signed-off-by: Willian <willian@willian.email>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix] Fix Anthropic Thinking Disabled with P/D (#58786)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Bugfix][Frontend] Detect Anthropic inline-system merge against the resolved chat template (#58754)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Mypy] Fix mypy typing for Qwen and Qianfan models (#58046)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [CI] Reduce CUDA graph mode test overhead (#58749)

* [GLM5.3 Bug] Fix sparse indexer attn topk backend selection (#58594)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [CI] Stabilize batch submission in full CUDA graph tests (#58810)

* [Bugfix][DSV4.1] Avoid host sync in ViT CUDA graph replay metadata (#58499)

* [Security] Harden message sanitization (#58832)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [Kernel][DSV4.1] Fuse MoE finalize into the TP all-reduce + mHC boundary (#58586)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [PCP][DCP] Support DCP target model with non-DCP Dspark (#56723)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Bump FlashKDA to keep the recurrent state in fp32 (#58846)

Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector (#57251)

Signed-off-by: Vincent <vincexxchan@gmail.com>

* [mooncake] support CUSTOM_MEM_POOL in vllm (#49300)

Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: bruce.xu <bruce.xb@alibaba-inc.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Qwen3.8-Flash-Next] Avoid memory fragmentation in QSA indexer logits workspace (#57105)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix] Disable sequence parallelism / async TP under batch invariance and add a TP regression test (#56377)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Kimi-K3] Make VLLM_ROCM_USE_AITER_MOE_SITUV2 select a4w4/a8w4/a16w4 (#58201)

Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>

* [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL (#50047)

Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Co-authored-by: xijiade.aihemaiti <3146335281@qq.com>

* [Bugfix] V1: clear stale allowed_token_ids mask in InputBatch.condense (#43931)

Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Perf][DSv4.1] Fuse small-batch WO-A with inverse RoPE and MXFP8 quant on SM100/SM103 (#58634)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Attention][CPU] Use zentorch SDPA for CPU MLA prefill (#54967)

Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [CI] Split (H200 MIG 35GB) Spec Decode Speculators + MTP into 4 named jobs (#57237)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Thang Nguyen <thang.nguyen@inferact.ai>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][Perf] Enable layer-aware CSA2 multi-stream overlap for DeepSeek-V4.1-Flash (#57407)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Perf][Pooling] Avoid blocking seq_lens GPU-to-CPU copy for pooling in FlashInfer metadata builder (#57214)

Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>

* [Bugfix] Support repsonse_format + tool_choice=auto (#56086)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: pablopupo <145598901+pablopupo@users.noreply.github.com>
Co-authored-by: hubunt <150658615+hubunt@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Fast Start] Add `/health` endpoint for the weight cache daemon (#58552)

Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Perf][MoE] Use fused MiniMax2 routing with non-unit routed scaling (#58880)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>

* [ROCm][Bugfix] Fall back to default GEMM for CPU tensors on ROCm builds (#58923)

Signed-off-by: fai <fangzhouai@gmail.com>

* [Bugfix][KV Connector] Retry Mooncake bootstrap registration on timeout (reopens #55763) (#58919)

* [Frontend] Switch Python Harmony dependency to oss-harmony (#55128)

Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm] Bump AITER to v0.1.23 (#58867)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Count Responses reasoning tokens per tool round (#58927)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [KimiViT][Perf] Fuse per-layer QK RoPE into one in-place kernel (#58651)

* [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models (#58239)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [Bugfix][Frontend] Sample batched chat completions from the adjusted requests (#58929)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Use a fresh parser per choice in non-streaming chat completions (#58939)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][EPD] Skip sampling for encoder-only async steps (#58490)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [CPU] Build CPU wheels on Ubuntu 22.04 with AMX-FP8 support (#58515)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>

* [Test][ROCm] Stabilize the mixed OLMoE LoRA test (#58945)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [Model] Enable LoRA support for RobertaForSequenceClassification (#58884)

Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Co-authored-by: jz-yolo <jz-yolo@users.noreply.github.com>

* [Bugfix][Frontend] Apply Harmony adjust_request in batched chat completions (#58958)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Minimax-M3] Add Encoder CUDA graph support (#58673)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <codex@openai.com>

* [Perf][MRV2] Reuse Mamba/GDN metadata across KV cache groups (#58762)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Xin Yang <xyangx@amazon.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [Bugfix] Fix the two multimodal root tests that fail on main (OpenPangu-VL embed merge, MiMo sink test fixture) (#58900)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Benchmark] Add Responses API backend to vllm bench serve (#54628)

Signed-off-by: QHarshil <harshil_c@hotmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [CI] Allowlist-shrink batch 1: wire 17 root-level tests + drop 5 stale watermarking entries into misc.yaml (#58055)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [CPU] Use accelerator memory API in DiffusionGemma (#58964)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [CPU] Add video inferencing via torchcodec on s390x (#58693)

Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Skills] Update kernel-microbenchmark to include ROCm (#58646)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Misc] Name each backend and its kernel block sizes in block-size errors (#58557)

Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: khluu <khluu000@gmail.com>

* [CI] Split Dynamic Shapes out of (H200 MIG 35GB) PyTorch Compilation + (MI300) mirror (#58451)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Release the profiling KV cache held by QSA key views (#58961)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CPU] Include vLLM Recipes tooling in release image to deploy models using vLLM Recipes (#58796)

Signed-off-by: louie-tsai <louie.tsai@intel.com>

* Revert "[CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)" (#59011)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Perf][Qwen4Exp] Fuse HC down projection and SiLU on NVIDIA (#58957)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix][Model] Gemma4: register aliased embedding scalars as buffers (#54213)

Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>

* [Bugfix][Spec Decode] Implement get_top_tokens() on the ROCm DeepSeek V4 MTP drafter (#57568)

Signed-off-by: BaoYunkai <ybao@amd.com>

* [Perf][Qwen3.8] Reduce PLE metadata construction overhead (#58114)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm][Refactor] Move DeepSeek-V4/V4.1 multi-stream overlap gate to ROCm platform (#58983)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Multimodal] Avoid extra d2d for encoder cudagraph with fused input norm (#56711)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Logging] Preserve application log record factories (#58747)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix] Use the correct repository revision for secondary artifact loaders (#57461)

Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Perf] Kimi-K3 Enable sharded latent MoE up-projection under EP (#54956)

Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Feature] Add fixed-token prefill scoring (#54335)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [LoRA] Support variable num_labels for sequence classification (#57766)

Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [ROCm][Perf] Replace torch.topk in DSA candidate block selection (#58208)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [ROCm]Keep LMCache OpenTelemetry on the image's 1.40 stack (#59056)

Signed-off-by: Micah Williamson <micah.williamson@amd.com>

* [Bugfix][Kimi-K3] Refresh DSpark context KV cache pointers after the KV cache is re-bound (#58814)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [KVConnector][NIXL] Support packed MLA KV layouts in pipeline-parallel push prefill (#50499)

Signed-off-by: zixi-qi <zixi@inferact.ai>

* [Quant] Use canonical N-first weight format for CT WNA16 MoE (#52798)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: HDCharles <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Relax packed_qk_rope_ correctness test to one ULP (#59008)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Pin OpenTelemetry to LMCache's cap in the ROCm images (#59051)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Increase timeout for Entrypoints Unit (#59076)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][MLA] Add an AITER ASM round-robin decode route for DCP multi-token verify (#56861)

Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix] Don't sync-police or retry FlashInfer all-reduce workspace creation in eager mode (#58498)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix resumable request + async scheduling handoff race (#58259)

Signed-off-by: Yifan Zong <yzong@redhat.com>

* [Model Runner V2][Spec Decode] Support spec decode with draft model (#43091)

Signed-off-by: Icey <1790571317@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][MRV2] Allow FULL decode graphs for one-token prompt tails (#58400)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Bugfix][Scheduler] Preserve logprobs across streaming continuations (#57447)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>

* [Mypy] Fix mypy typing for Voxtral and vision models (#58251)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Core] Rework scheduler `skipped_waiting` queue (#58947)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Bugfix][MRV2][Spec Decode] Reject draft slots that were never proposed (#58784)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][CI] Add missing test coverage for upstream parity (#50519)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][Scheduler] Refresh max tokens for streaming continuations (#57676)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Spec Decode] Avoid triton recompiles in the acceptance estimator (#57107)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>

* [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8664f67d8ae1538e5e5cec2b1ce3f258dd)

* [XPU][CI]Skip test_abort_timeout_on_prefiller in nightly (#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db1ca080793dc2351144cb39513fd7984ca)

* [Bugfix][Engram] Keep THP tables private when resolving shared memory (#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c352f079c2cb8f46752fd0317a2745050bf)

* [ROCm][CI] Expand MI355 mirrors and route MIG-sized jobs to DPX (#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd64f5ef75dbe1471ee88c009003b3d03ec)

* [Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse retention (#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bddbeab6b4a0d5861dfcb171bf61ce2109d6)

* [Core][BugFix] Tag prefix-cache extra keys by source (#51899)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e7ed5b6f372f0898a67e48d3365fa28139)

* [Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e097b905127b1d5f09a7e93bd1095e87a6c)

* [Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf53724e06cef7f07628ae8fd4cc6c96f11a)

* [Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f668608155bc8a8c74cb97e3ddd0d3053085)

* [Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc757cffc11b55689d9cc33d844d9606516)

* [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a4297921c699f9aa58f9c6395e95ede718)

* [Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa130fac1e1252330a680fbed45ad05eba64f)

* [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)

(cherry picked from commit 5463fe4962785cdc3383477bf3af6533e7647dfd)

* [Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87cca25690fef6bd12667fd0d26690d949af)

* [Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec22ad9bb098393021b956feaa97081fc85)

* [Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae4d746f64c22677348d1cc120eea6f7439)

* [Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e615c20ee4ff0cd82aac592c7d6310aa7c3)

* [Core] Include the LoRA path in prefix-cache block hashes (#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e075ed461cff922aa710d387281f651940)

* [Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360c9eebb3b4b460ec49f6606ef0a08fcae2)

* [CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d7332db897885fdc0e6ed61b28c969d743)

* [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba9b7d18f9746d989e389a51a94b0f96f6f)

* [Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a6963664537ed21172e2ec12e96e3a2dcd3718c)

* [Bugfix][HiSparse] Size the KV cache from the groups HiSparse allocates (#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2e4489126f1c21868699f3890021f01b73)

* [Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with a saturated GPU pool (#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8ac1f8a7e8fb50cf1c56ff96649fcadccc6)

* [CI] Drop a test that depends on #57930 from the #59309 backport

Resolving the #59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from #57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream #59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>

* [Bugfix] Fix minimax-m3 multimodal processor compatability with Transformers v5.18 (#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f160a2c0abcb5902acc3c91a14c38a4af173)

* [Misc] Add Transformers version upper bound in requirements (#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298457dde7b4554b3b4e20b238c0ac2c3a65)

* Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0

Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2eb #28, 76a3e6d72, 6d4c75611 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a46, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 18.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c19, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - per-group BlockPools, sampler reserve, padded-page block fill, drain
    hook / on_kv_manager_created, KVBlockZeroer clamp: conflict with
    upstream's single-pool allocator; no turbo-attn consumer for the hooks
  - rotary fast-path registry: upstream guards the import (1f60771c7,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Signed-off-by: chao.huan <chao.huan@nio.com>
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: errmakov <ide404@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Signed-off-by: yashasvi <yashasvi@ibm.com>
Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Signed-off-by: Cheng Rui <286040359@qq.com>
Signed-off-by: Yifan Zong <yzong@redhat.com>
Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Linkun Chen <github@lkchen.net>
Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Ke Wen <kwen@nvidia.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: zixi-qi <zixi@inferact.ai>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: Willian <willian@willian.email>
Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Vincent <vincexxchan@gmail.com>
Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: QHarshil <harshil_c@hotmail.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>
Signed-off-by: BaoYunkai <ybao@amd.com>
Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Stefan Kon…
arbi-dev added a commit to arbicity/vllm-turbo that referenced this pull request Oct 7, 2026
* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [SpecDecode] Add LiLiCorr drafter (#57934)

Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>

* [MRV2] Minor model_runner.py code cleanup (#58610)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix generative scoring body cancellation (#57729)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] Preserve reranker tokenization with document limits (#57666)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Kernel][Perf] Register-resident path for per-token-group 8-bit quant (#55330)

Signed-off-by: chao.huan <chao.huan@nio.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf][Kernel] Vectorized flat abs-max for dynamic per-tensor FP8 quantization (#58194)

Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Core][Logging] Fix JSON logging process decoration (#57957)

Signed-off-by: Mark McLoughlin <markmc@redhat.com>

* [Bugfix] Stop allocator fragmentation from shrinking the KV cache during memory profiling (#58430)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Keep logprobs of parser-suppressed streaming chunks (#58583)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Yanxiao Zhao <39199723+sdpkjc@users.noreply.github.com>
Co-authored-by: Prakhar Agarwal <270064960+agarwalprakhar2511@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix][GLM-5.3-Flash] SM90 sparse MLA: index_kpool mismatch leads to corruption via unread query token (#58704)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Benchmark] Record model_id in bench latency/throughput --output-json (#58112)

Signed-off-by: yashasvi <yashasvi@ibm.com>

* [Bugfix][GLM-5.3-Flash] kpool corruption with speculative decoding (#58454)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [GLM5.3 Perf] Optimize glm 5.3 metadata op, 1.6~4.8x kernel level performance improvement (#58450)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Bugfix][KV Connector] Reap expired NIXL leases behind a heartbeated head (#58292)

Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>

* [ROCm][Kimi-K3] Optimize low-concurrency speculative KDA (#58045)

Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Cover the AITER MQA logits dispatch on gfx950 (#58724)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Add quantized MoE serving test for gfx950 (#58748)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Test AMD DeepSeek V4 MoE routing against a PyTorch reference (#58740)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: one-pass sampler statistics kernel (#58226)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [ROCm][CI] Expand single-GPU coverage on MI355 DPX (#57599)

Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [CI] [MRV2] Restore MRV2 pp dp coverage (#57735)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][CI] Report subprocess test skips as skips, not passes (#58701)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Run the MLA attention+quant fusion test on ROCm (#58717)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm] Bump torch 2.13, triton 3.8, torchaudio, torchvision (#50605)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: jpvillam <juan.villamizar@amd.com>

* [Feature][Frontend] Add granite_thinking_parser reasoning parser for Granite 4.2 (#55957)

Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Respect max_output_tokens in the Harmony tool-call loop (#58551)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Du Bin <8174807+dubin555@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Frontend] Document 404 response for `/generative_scoring` (#58788)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Perf][Frontend] Defer reasoning usage recounts for non-continuous chat streams (#56067)

Signed-off-by: Cheng Rui <286040359@qq.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] Default missing detail for Responses API input images (#57241)

Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>

* [watermarking] golden tests for backwards compatibility (#56809)

Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Don't drop the rest of the allocator config when toggling expandable segments (#57982)

Signed-off-by: Oxana Korzh <okorzh@amd.com>

* [AuxOutput] Only require Model Runner V2 on GPU platform (#58205)

Signed-off-by: Linkun Chen <github@lkchen.net>

* [Pooling] Preserve BERT-family heads for raw logits (#57664)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* fix: perf: use startswith(x, i) instead of string slicing to avoid O(N^2) (#52580)

Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states (#48419)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Co-authored-by: Claude <noreply@anthropic.com>

* [Bugfix][Reasoning] Count Kimi K3 reasoning tokens (#58372)

Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [CI] Split (H200 MIG/MI300) Basic Correctness into named jobs (#57054)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Frontend] Fix Inkling tool name leaking into content after reasoning (#58792)

Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Core] Bound draft-token RPC waits by the execute-model timeout (#58779)

Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Co-authored-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>

* [Perf][DSv4.1] Shard the Engram wkv projection across TP ranks (#58678)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>

* [RL] Add sharding-aware NCCL M2N weight transfer (#51520)

Signed-off-by: Ke Wen <kwen@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][ROCm] AMD-Quark mixed-precision DeepSeek-V4.1 support (#57071)

Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix][Frontend][Rust Frontend] Update DeepSeek V4.1 Flash reasoning effort mappings (#58316)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>

* [CI] Only isolate the registry tests that need a fresh process (#58764)

Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Spec Decode] Enable Gemma4 DSpark adaptive verification with FlashInfer (#57263)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [CI] Split (B200) Miscellaneous Kernels into mHC, FLA Ops and Misc named jobs (#58609)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Security] Gate per-request multimodal processor kwargs (#58830)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Elastic EP] Fix EPLB load statistics during scaling (#58473)

Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>

* [Refactor] Remove dead tests utils (#58803)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights (#58720)

Signed-off-by: Willian <willian@willian.email>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix] Fix Anthropic Thinking Disabled with P/D (#58786)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Bugfix][Frontend] Detect Anthropic inline-system merge against the resolved chat template (#58754)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Mypy] Fix mypy typing for Qwen and Qianfan models (#58046)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [CI] Reduce CUDA graph mode test overhead (#58749)

* [GLM5.3 Bug] Fix sparse indexer attn topk backend selection (#58594)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [CI] Stabilize batch submission in full CUDA graph tests (#58810)

* [Bugfix][DSV4.1] Avoid host sync in ViT CUDA graph replay metadata (#58499)

* [Security] Harden message sanitization (#58832)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [Kernel][DSV4.1] Fuse MoE finalize into the TP all-reduce + mHC boundary (#58586)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [PCP][DCP] Support DCP target model with non-DCP Dspark (#56723)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Bump FlashKDA to keep the recurrent state in fp32 (#58846)

Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector (#57251)

Signed-off-by: Vincent <vincexxchan@gmail.com>

* [mooncake] support CUSTOM_MEM_POOL in vllm (#49300)

Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: bruce.xu <bruce.xb@alibaba-inc.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Qwen3.8-Flash-Next] Avoid memory fragmentation in QSA indexer logits workspace (#57105)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix] Disable sequence parallelism / async TP under batch invariance and add a TP regression test (#56377)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Kimi-K3] Make VLLM_ROCM_USE_AITER_MOE_SITUV2 select a4w4/a8w4/a16w4 (#58201)

Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>

* [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL (#50047)

Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Co-authored-by: xijiade.aihemaiti <3146335281@qq.com>

* [Bugfix] V1: clear stale allowed_token_ids mask in InputBatch.condense (#43931)

Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Perf][DSv4.1] Fuse small-batch WO-A with inverse RoPE and MXFP8 quant on SM100/SM103 (#58634)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Attention][CPU] Use zentorch SDPA for CPU MLA prefill (#54967)

Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [CI] Split (H200 MIG 35GB) Spec Decode Speculators + MTP into 4 named jobs (#57237)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Thang Nguyen <thang.nguyen@inferact.ai>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][Perf] Enable layer-aware CSA2 multi-stream overlap for DeepSeek-V4.1-Flash (#57407)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Perf][Pooling] Avoid blocking seq_lens GPU-to-CPU copy for pooling in FlashInfer metadata builder (#57214)

Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>

* [Bugfix] Support repsonse_format + tool_choice=auto (#56086)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: pablopupo <145598901+pablopupo@users.noreply.github.com>
Co-authored-by: hubunt <150658615+hubunt@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Fast Start] Add `/health` endpoint for the weight cache daemon (#58552)

Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Perf][MoE] Use fused MiniMax2 routing with non-unit routed scaling (#58880)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>

* [ROCm][Bugfix] Fall back to default GEMM for CPU tensors on ROCm builds (#58923)

Signed-off-by: fai <fangzhouai@gmail.com>

* [Bugfix][KV Connector] Retry Mooncake bootstrap registration on timeout (reopens #55763) (#58919)

* [Frontend] Switch Python Harmony dependency to oss-harmony (#55128)

Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm] Bump AITER to v0.1.23 (#58867)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Count Responses reasoning tokens per tool round (#58927)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [KimiViT][Perf] Fuse per-layer QK RoPE into one in-place kernel (#58651)

* [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models (#58239)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [Bugfix][Frontend] Sample batched chat completions from the adjusted requests (#58929)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Use a fresh parser per choice in non-streaming chat completions (#58939)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][EPD] Skip sampling for encoder-only async steps (#58490)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [CPU] Build CPU wheels on Ubuntu 22.04 with AMX-FP8 support (#58515)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>

* [Test][ROCm] Stabilize the mixed OLMoE LoRA test (#58945)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [Model] Enable LoRA support for RobertaForSequenceClassification (#58884)

Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Co-authored-by: jz-yolo <jz-yolo@users.noreply.github.com>

* [Bugfix][Frontend] Apply Harmony adjust_request in batched chat completions (#58958)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Minimax-M3] Add Encoder CUDA graph support (#58673)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <codex@openai.com>

* [Perf][MRV2] Reuse Mamba/GDN metadata across KV cache groups (#58762)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Xin Yang <xyangx@amazon.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [Bugfix] Fix the two multimodal root tests that fail on main (OpenPangu-VL embed merge, MiMo sink test fixture) (#58900)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Benchmark] Add Responses API backend to vllm bench serve (#54628)

Signed-off-by: QHarshil <harshil_c@hotmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [CI] Allowlist-shrink batch 1: wire 17 root-level tests + drop 5 stale watermarking entries into misc.yaml (#58055)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [CPU] Use accelerator memory API in DiffusionGemma (#58964)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [CPU] Add video inferencing via torchcodec on s390x (#58693)

Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Skills] Update kernel-microbenchmark to include ROCm (#58646)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Misc] Name each backend and its kernel block sizes in block-size errors (#58557)

Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: khluu <khluu000@gmail.com>

* [CI] Split Dynamic Shapes out of (H200 MIG 35GB) PyTorch Compilation + (MI300) mirror (#58451)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Release the profiling KV cache held by QSA key views (#58961)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CPU] Include vLLM Recipes tooling in release image to deploy models using vLLM Recipes (#58796)

Signed-off-by: louie-tsai <louie.tsai@intel.com>

* Revert "[CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)" (#59011)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Perf][Qwen4Exp] Fuse HC down projection and SiLU on NVIDIA (#58957)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix][Model] Gemma4: register aliased embedding scalars as buffers (#54213)

Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>

* [Bugfix][Spec Decode] Implement get_top_tokens() on the ROCm DeepSeek V4 MTP drafter (#57568)

Signed-off-by: BaoYunkai <ybao@amd.com>

* [Perf][Qwen3.8] Reduce PLE metadata construction overhead (#58114)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm][Refactor] Move DeepSeek-V4/V4.1 multi-stream overlap gate to ROCm platform (#58983)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Multimodal] Avoid extra d2d for encoder cudagraph with fused input norm (#56711)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Logging] Preserve application log record factories (#58747)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix] Use the correct repository revision for secondary artifact loaders (#57461)

Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Perf] Kimi-K3 Enable sharded latent MoE up-projection under EP (#54956)

Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Feature] Add fixed-token prefill scoring (#54335)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [LoRA] Support variable num_labels for sequence classification (#57766)

Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [ROCm][Perf] Replace torch.topk in DSA candidate block selection (#58208)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [ROCm]Keep LMCache OpenTelemetry on the image's 1.40 stack (#59056)

Signed-off-by: Micah Williamson <micah.williamson@amd.com>

* [Bugfix][Kimi-K3] Refresh DSpark context KV cache pointers after the KV cache is re-bound (#58814)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [KVConnector][NIXL] Support packed MLA KV layouts in pipeline-parallel push prefill (#50499)

Signed-off-by: zixi-qi <zixi@inferact.ai>

* [Quant] Use canonical N-first weight format for CT WNA16 MoE (#52798)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: HDCharles <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Relax packed_qk_rope_ correctness test to one ULP (#59008)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Pin OpenTelemetry to LMCache's cap in the ROCm images (#59051)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Increase timeout for Entrypoints Unit (#59076)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][MLA] Add an AITER ASM round-robin decode route for DCP multi-token verify (#56861)

Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix] Don't sync-police or retry FlashInfer all-reduce workspace creation in eager mode (#58498)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix resumable request + async scheduling handoff race (#58259)

Signed-off-by: Yifan Zong <yzong@redhat.com>

* [Model Runner V2][Spec Decode] Support spec decode with draft model (#43091)

Signed-off-by: Icey <1790571317@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][MRV2] Allow FULL decode graphs for one-token prompt tails (#58400)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Bugfix][Scheduler] Preserve logprobs across streaming continuations (#57447)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>

* [Mypy] Fix mypy typing for Voxtral and vision models (#58251)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Core] Rework scheduler `skipped_waiting` queue (#58947)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Bugfix][MRV2][Spec Decode] Reject draft slots that were never proposed (#58784)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][CI] Add missing test coverage for upstream parity (#50519)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][Scheduler] Refresh max tokens for streaming continuations (#57676)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Spec Decode] Avoid triton recompiles in the acceptance estimator (#57107)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>

* [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8664f67d8ae1538e5e5cec2b1ce3f258dd)

* [XPU][CI]Skip test_abort_timeout_on_prefiller in nightly (#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db1ca080793dc2351144cb39513fd7984ca)

* [Bugfix][Engram] Keep THP tables private when resolving shared memory (#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c352f079c2cb8f46752fd0317a2745050bf)

* [ROCm][CI] Expand MI355 mirrors and route MIG-sized jobs to DPX (#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd64f5ef75dbe1471ee88c009003b3d03ec)

* [Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse retention (#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bddbeab6b4a0d5861dfcb171bf61ce2109d6)

* [Core][BugFix] Tag prefix-cache extra keys by source (#51899)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e7ed5b6f372f0898a67e48d3365fa28139)

* [Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e097b905127b1d5f09a7e93bd1095e87a6c)

* [Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf53724e06cef7f07628ae8fd4cc6c96f11a)

* [Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f668608155bc8a8c74cb97e3ddd0d3053085)

* [Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc757cffc11b55689d9cc33d844d9606516)

* [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a4297921c699f9aa58f9c6395e95ede718)

* [Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa130fac1e1252330a680fbed45ad05eba64f)

* [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)

(cherry picked from commit 5463fe4962785cdc3383477bf3af6533e7647dfd)

* [Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87cca25690fef6bd12667fd0d26690d949af)

* [Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec22ad9bb098393021b956feaa97081fc85)

* [Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae4d746f64c22677348d1cc120eea6f7439)

* [Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e615c20ee4ff0cd82aac592c7d6310aa7c3)

* [Core] Include the LoRA path in prefix-cache block hashes (#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e075ed461cff922aa710d387281f651940)

* [Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360c9eebb3b4b460ec49f6606ef0a08fcae2)

* [CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d7332db897885fdc0e6ed61b28c969d743)

* [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba9b7d18f9746d989e389a51a94b0f96f6f)

* [Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a6963664537ed21172e2ec12e96e3a2dcd3718c)

* [Bugfix][HiSparse] Size the KV cache from the groups HiSparse allocates (#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2e4489126f1c21868699f3890021f01b73)

* [Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with a saturated GPU pool (#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8ac1f8a7e8fb50cf1c56ff96649fcadccc6)

* [CI] Drop a test that depends on #57930 from the #59309 backport

Resolving the #59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from #57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream #59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>

* [Bugfix] Fix minimax-m3 multimodal processor compatability with Transformers v5.18 (#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f160a2c0abcb5902acc3c91a14c38a4af173)

* [Misc] Add Transformers version upper bound in requirements (#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298457dde7b4554b3b4e20b238c0ac2c3a65)

* Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0

Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2eb #28, 76a3e6d72, 6d4c75611 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a46, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 27.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner
  - per-group KV pool for hybrid models, re-expressed on v0.31's single
    backing allocation: the O(1) Mamba/GDN state groups get their own
    BlockPool sized for max_num_seqs, a region of the allocation after the
    attention pool; their pages are not unified with attention pages;
    per-pool admission, events and usage; CoW copies and warmup/profiling
    block ids per pool; VLLM_SAMPLER_RESERVE_MIB headroom

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c19, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - padded-page block fill (the split pool removes the page unification
    it worked around), drain hook / on_kv_manager_created, KVBlockZeroer
    clamp: no turbo-attn consumer
  - rotary fast-path registry: upstream guards the import (1f60771c7,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Signed-off-by: chao.huan <chao.huan@nio.com>
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: errmakov <ide404@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Signed-off-by: yashasvi <yashasvi@ibm.com>
Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Signed-off-by: Cheng Rui <286040359@qq.com>
Signed-off-by: Yifan Zong <yzong@redhat.com>
Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Linkun Chen <github@lkchen.net>
Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Ke Wen <kwen@nvidia.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: zixi-qi <zixi@inferact.ai>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: Willian <willian@willian.email>
Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Vincent <vincexxchan@gmail.com>
Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: QHarshil <harshil_c@hotmail.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Yannick Schnider <Yannick.Schnider1@ib…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/build cpu Related to CPU backends

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Installation]: CPU source build fails on GCC < 15: sgl-kernels use AVX10.2 unconditionally

2 participants