Skip to content

[Core] Bound UniProc EngineCore startup threads to available CPUs - #58946

Merged
njhill merged 5 commits into
vllm-project:mainfrom
AndreasKaratzas:akaratza_uniproc_startup_threads
Sep 30, 2026
Merged

njhill merged 5 commits into
vllm-project:mainfrom
AndreasKaratzas:akaratza_uniproc_startup_threads

Conversation

@AndreasKaratzas

@AndreasKaratzas AndreasKaratzas commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

In Buildkite CI #91445, an MI355 process with a 12-CPU cgroup quota loaded weights with 192 PyTorch threads. Weight loading took 38 minutes and the job timed out after its accuracy checks passed. The quota-aware budget added by #49919 covered multiprocessing workers, but a UniProc worker runs in EngineCore and did not receive that budget.

For process-managed UniProc, cap startup threads at the smaller of the current PyTorch per-process limit and #49919's available-CPU count divided by the number of local EngineCore processes. Set OMP_NUM_THREADS and #49919's existing vLLM ownership marker only while starting those processes, then restore the parent's settings even if launch fails. An explicit user OMP_NUM_THREADS remains authoritative. The worker keeps #49919's existing transition to one PyTorch thread after weight loading.

This differs from open #45124, which proposes a fixed, tunable four-thread cap inside UniProc for serving. Here the cap is the #49919 affinity/cgroup startup budget, applied before subprocess launch; serving returns to one thread. The launch-only scope is needed because fork inherits the API parent's PyTorch count while spawn reads its OMP setting. Restoring the parent prevents a later launch from mistaking vLLM's generated OMP value for a user override and skipping a fresh CPU budget.

Serving evaluations compare one and four EngineCore threads on a ROCm gfx942 host with 12-CPU server affinity. Both variants started with OMP_NUM_THREADS=4; toggling #49919's existing marker changed only the EngineCore's post-load thread count. These comparisons isolate the existing serving policy rather than exercise the new automatic startup cap.

Model and workload One EngineCore thread Four EngineCore threads Result
Qwen3-0.6B, 96 GSM8K prompts at 8 RPS, two runs each 96/96 completed per run; E2E p50 211/221 ms 96/96 completed per run; E2E p50 224/225 ms Similar throughput (7.88 RPS each); extracted GSM8K answer accuracy was 14/96 in three runs and 15/96 in one.
Qwen3-0.6B, 96 JSON-schema requests at 8 RPS, two runs each 96/96 schema-valid per run; E2E p50 188/179 ms 96/96 schema-valid per run; E2E p50 195/180 ms Small difference relative to run variation.
SmolVLM-256M-Instruct, default compilation, four 224x224 images/request, 256 requests/run in A-B-B-A order 256/256 completed per run; mean 5.60 RPS 256/256 completed per run; mean 5.84 RPS Four threads had a small, order-sensitive 4.3% throughput advantage.
SmolVLM-256M-Instruct, default compilation, one image/request, 256 requests/run 256/256 completed per run; mean 19.07 RPS 256/256 completed per run; mean 18.99 RPS No stable gain from four threads.

The deterministic four-image SmolVLM request returned the same HTTP 200 answer under both thread counts. A preliminary 64-request eager multimodal comparison showed lower concurrent text-probe latency with four threads (98–105 ms versus 199–202 ms). Two longer eager pairs, each with 256 four-image requests per setting, did not reproduce that gap: probe medians were 150–155 ms with one thread and 141–160 ms with four. Four-image throughput was 5.52/5.63 versus 5.93/5.61 requests/s (one versus four); all 1,024 image requests and 273 probes succeeded. The 96-prompt GSM8K run used a 128-token cap reached by 83 prompts in every run, so its low exact-match score is a consistency check, not an accuracy claim. A single saturated 256-request JSON pair favored one thread (38.90 versus 33.08 RPS; 256/256 schema-valid under each setting), but needs repetition before assigning an effect size.

AI assistance was used for the investigation, implementation, tests, and evaluation analysis.

Apply the existing quota-aware startup thread policy before launching UniProc engine processes, then restore the parent settings. Preserve explicit user settings and the runtime thread transition.

Assisted-by: OpenAI Codex
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
@mergify mergify Bot added the rocm Related to AMD ROCm label Sep 28, 2026
@github-project-automation github-project-automation Bot moved this to Todo in AMD Sep 28, 2026
@AndreasKaratzas
AndreasKaratzas marked this pull request as ready for review September 28, 2026 00:17

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@AndreasKaratzas

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91497 for commit d4537e27d4ca.

Apply the existing quota-aware startup thread budget when launching
process-managed UniProc EngineCores. Limit the temporary OMP and PyTorch
settings to process start, then restore the parent on success or failure.

Preserve explicit user OMP settings and the existing one-thread serving
transition. Remove the provisional Multiproc max_cpu_threads argument.

Assisted-by: OpenAI Codex
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
@AndreasKaratzas AndreasKaratzas changed the title [Core][ROCm] Respect CPU limits when starting UniProc workers [Core] Bound UniProc EngineCore startup threads to available CPUs Sep 29, 2026
@AndreasKaratzas

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

❌ This PR is 1 commit behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@AndreasKaratzas

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91749 for commit 9c600e5e9aec.

@khluu khluu added this to the v0.31.0 cherry picks milestone Sep 30, 2026
@AndreasKaratzas

Copy link
Copy Markdown
Member Author

@njhill njhill left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @AndreasKaratzas. Thinking about refactoring this a bit but can do it as a follow-on.

@njhill
njhill merged commit 866fa13 into vllm-project:main Sep 30, 2026
155 of 156 checks passed
@AndreasKaratzas
AndreasKaratzas deleted the akaratza_uniproc_startup_threads branch September 30, 2026 06:06
khluu pushed a commit that referenced this pull request Oct 1, 2026
…8946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa13)
charlesli640 added a commit to vllm-project/tpu-inference that referenced this pull request Oct 4, 2026
…Small-4 hang)

vllm-project/vllm#58946 (in vLLM LKG 72e7874f) wraps EngineCore process
start in torch.set_num_threads(n) plus OMP_NUM_THREADS=n when the executor
is a UniProcExecutor, which TPU forces on a single host. On TPU the
processes forked after that (EngineCore, then the DPScheduler workers) hang
in libgomp: tests/e2e/mistral_small_4_correctness.py on tpu7x never
finishes its first step and the CI job runs until it is canceled.

Bisected on tpu-inference-dev with only the vLLM commit changed:
vLLM 5faf81a429 (parent of #58946) passes in 6 min, vLLM 866fa130fa
(#58946) hangs.

Replace vllm.v1.engine.utils._configure_uniproc_startup_threads with a
no-op when TPU selects the UniProc executor, restoring the pre-#58946
behavior.

Signed-off-by: Charles Li <licharles@google.com>
arbi-dev added a commit to arbicity/vllm-turbo that referenced this pull request Oct 7, 2026
* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [SpecDecode] Add LiLiCorr drafter (#57934)

Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>

* [MRV2] Minor model_runner.py code cleanup (#58610)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix generative scoring body cancellation (#57729)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] Preserve reranker tokenization with document limits (#57666)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Kernel][Perf] Register-resident path for per-token-group 8-bit quant (#55330)

Signed-off-by: chao.huan <chao.huan@nio.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf][Kernel] Vectorized flat abs-max for dynamic per-tensor FP8 quantization (#58194)

Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Core][Logging] Fix JSON logging process decoration (#57957)

Signed-off-by: Mark McLoughlin <markmc@redhat.com>

* [Bugfix] Stop allocator fragmentation from shrinking the KV cache during memory profiling (#58430)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Keep logprobs of parser-suppressed streaming chunks (#58583)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Yanxiao Zhao <39199723+sdpkjc@users.noreply.github.com>
Co-authored-by: Prakhar Agarwal <270064960+agarwalprakhar2511@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix][GLM-5.3-Flash] SM90 sparse MLA: index_kpool mismatch leads to corruption via unread query token (#58704)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Benchmark] Record model_id in bench latency/throughput --output-json (#58112)

Signed-off-by: yashasvi <yashasvi@ibm.com>

* [Bugfix][GLM-5.3-Flash] kpool corruption with speculative decoding (#58454)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [GLM5.3 Perf] Optimize glm 5.3 metadata op, 1.6~4.8x kernel level performance improvement (#58450)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Bugfix][KV Connector] Reap expired NIXL leases behind a heartbeated head (#58292)

Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>

* [ROCm][Kimi-K3] Optimize low-concurrency speculative KDA (#58045)

Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Cover the AITER MQA logits dispatch on gfx950 (#58724)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Add quantized MoE serving test for gfx950 (#58748)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Test AMD DeepSeek V4 MoE routing against a PyTorch reference (#58740)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: one-pass sampler statistics kernel (#58226)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [ROCm][CI] Expand single-GPU coverage on MI355 DPX (#57599)

Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [CI] [MRV2] Restore MRV2 pp dp coverage (#57735)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][CI] Report subprocess test skips as skips, not passes (#58701)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Run the MLA attention+quant fusion test on ROCm (#58717)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm] Bump torch 2.13, triton 3.8, torchaudio, torchvision (#50605)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: jpvillam <juan.villamizar@amd.com>

* [Feature][Frontend] Add granite_thinking_parser reasoning parser for Granite 4.2 (#55957)

Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Respect max_output_tokens in the Harmony tool-call loop (#58551)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Du Bin <8174807+dubin555@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Frontend] Document 404 response for `/generative_scoring` (#58788)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Perf][Frontend] Defer reasoning usage recounts for non-continuous chat streams (#56067)

Signed-off-by: Cheng Rui <286040359@qq.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] Default missing detail for Responses API input images (#57241)

Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>

* [watermarking] golden tests for backwards compatibility (#56809)

Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Don't drop the rest of the allocator config when toggling expandable segments (#57982)

Signed-off-by: Oxana Korzh <okorzh@amd.com>

* [AuxOutput] Only require Model Runner V2 on GPU platform (#58205)

Signed-off-by: Linkun Chen <github@lkchen.net>

* [Pooling] Preserve BERT-family heads for raw logits (#57664)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* fix: perf: use startswith(x, i) instead of string slicing to avoid O(N^2) (#52580)

Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states (#48419)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Co-authored-by: Claude <noreply@anthropic.com>

* [Bugfix][Reasoning] Count Kimi K3 reasoning tokens (#58372)

Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [CI] Split (H200 MIG/MI300) Basic Correctness into named jobs (#57054)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Frontend] Fix Inkling tool name leaking into content after reasoning (#58792)

Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Core] Bound draft-token RPC waits by the execute-model timeout (#58779)

Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Co-authored-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>

* [Perf][DSv4.1] Shard the Engram wkv projection across TP ranks (#58678)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>

* [RL] Add sharding-aware NCCL M2N weight transfer (#51520)

Signed-off-by: Ke Wen <kwen@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][ROCm] AMD-Quark mixed-precision DeepSeek-V4.1 support (#57071)

Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix][Frontend][Rust Frontend] Update DeepSeek V4.1 Flash reasoning effort mappings (#58316)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>

* [CI] Only isolate the registry tests that need a fresh process (#58764)

Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Spec Decode] Enable Gemma4 DSpark adaptive verification with FlashInfer (#57263)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [CI] Split (B200) Miscellaneous Kernels into mHC, FLA Ops and Misc named jobs (#58609)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Security] Gate per-request multimodal processor kwargs (#58830)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Elastic EP] Fix EPLB load statistics during scaling (#58473)

Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>

* [Refactor] Remove dead tests utils (#58803)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights (#58720)

Signed-off-by: Willian <willian@willian.email>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix] Fix Anthropic Thinking Disabled with P/D (#58786)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Bugfix][Frontend] Detect Anthropic inline-system merge against the resolved chat template (#58754)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Mypy] Fix mypy typing for Qwen and Qianfan models (#58046)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [CI] Reduce CUDA graph mode test overhead (#58749)

* [GLM5.3 Bug] Fix sparse indexer attn topk backend selection (#58594)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [CI] Stabilize batch submission in full CUDA graph tests (#58810)

* [Bugfix][DSV4.1] Avoid host sync in ViT CUDA graph replay metadata (#58499)

* [Security] Harden message sanitization (#58832)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [Kernel][DSV4.1] Fuse MoE finalize into the TP all-reduce + mHC boundary (#58586)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [PCP][DCP] Support DCP target model with non-DCP Dspark (#56723)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Bump FlashKDA to keep the recurrent state in fp32 (#58846)

Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector (#57251)

Signed-off-by: Vincent <vincexxchan@gmail.com>

* [mooncake] support CUSTOM_MEM_POOL in vllm (#49300)

Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: bruce.xu <bruce.xb@alibaba-inc.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Qwen3.8-Flash-Next] Avoid memory fragmentation in QSA indexer logits workspace (#57105)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix] Disable sequence parallelism / async TP under batch invariance and add a TP regression test (#56377)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Kimi-K3] Make VLLM_ROCM_USE_AITER_MOE_SITUV2 select a4w4/a8w4/a16w4 (#58201)

Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>

* [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL (#50047)

Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Co-authored-by: xijiade.aihemaiti <3146335281@qq.com>

* [Bugfix] V1: clear stale allowed_token_ids mask in InputBatch.condense (#43931)

Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Perf][DSv4.1] Fuse small-batch WO-A with inverse RoPE and MXFP8 quant on SM100/SM103 (#58634)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Attention][CPU] Use zentorch SDPA for CPU MLA prefill (#54967)

Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [CI] Split (H200 MIG 35GB) Spec Decode Speculators + MTP into 4 named jobs (#57237)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Thang Nguyen <thang.nguyen@inferact.ai>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][Perf] Enable layer-aware CSA2 multi-stream overlap for DeepSeek-V4.1-Flash (#57407)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Perf][Pooling] Avoid blocking seq_lens GPU-to-CPU copy for pooling in FlashInfer metadata builder (#57214)

Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>

* [Bugfix] Support repsonse_format + tool_choice=auto (#56086)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: pablopupo <145598901+pablopupo@users.noreply.github.com>
Co-authored-by: hubunt <150658615+hubunt@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Fast Start] Add `/health` endpoint for the weight cache daemon (#58552)

Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Perf][MoE] Use fused MiniMax2 routing with non-unit routed scaling (#58880)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>

* [ROCm][Bugfix] Fall back to default GEMM for CPU tensors on ROCm builds (#58923)

Signed-off-by: fai <fangzhouai@gmail.com>

* [Bugfix][KV Connector] Retry Mooncake bootstrap registration on timeout (reopens #55763) (#58919)

* [Frontend] Switch Python Harmony dependency to oss-harmony (#55128)

Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm] Bump AITER to v0.1.23 (#58867)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Count Responses reasoning tokens per tool round (#58927)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [KimiViT][Perf] Fuse per-layer QK RoPE into one in-place kernel (#58651)

* [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models (#58239)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [Bugfix][Frontend] Sample batched chat completions from the adjusted requests (#58929)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Use a fresh parser per choice in non-streaming chat completions (#58939)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][EPD] Skip sampling for encoder-only async steps (#58490)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [CPU] Build CPU wheels on Ubuntu 22.04 with AMX-FP8 support (#58515)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>

* [Test][ROCm] Stabilize the mixed OLMoE LoRA test (#58945)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [Model] Enable LoRA support for RobertaForSequenceClassification (#58884)

Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Co-authored-by: jz-yolo <jz-yolo@users.noreply.github.com>

* [Bugfix][Frontend] Apply Harmony adjust_request in batched chat completions (#58958)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Minimax-M3] Add Encoder CUDA graph support (#58673)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <codex@openai.com>

* [Perf][MRV2] Reuse Mamba/GDN metadata across KV cache groups (#58762)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Xin Yang <xyangx@amazon.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [Bugfix] Fix the two multimodal root tests that fail on main (OpenPangu-VL embed merge, MiMo sink test fixture) (#58900)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Benchmark] Add Responses API backend to vllm bench serve (#54628)

Signed-off-by: QHarshil <harshil_c@hotmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [CI] Allowlist-shrink batch 1: wire 17 root-level tests + drop 5 stale watermarking entries into misc.yaml (#58055)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [CPU] Use accelerator memory API in DiffusionGemma (#58964)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [CPU] Add video inferencing via torchcodec on s390x (#58693)

Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Skills] Update kernel-microbenchmark to include ROCm (#58646)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Misc] Name each backend and its kernel block sizes in block-size errors (#58557)

Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: khluu <khluu000@gmail.com>

* [CI] Split Dynamic Shapes out of (H200 MIG 35GB) PyTorch Compilation + (MI300) mirror (#58451)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Release the profiling KV cache held by QSA key views (#58961)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CPU] Include vLLM Recipes tooling in release image to deploy models using vLLM Recipes (#58796)

Signed-off-by: louie-tsai <louie.tsai@intel.com>

* Revert "[CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)" (#59011)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Perf][Qwen4Exp] Fuse HC down projection and SiLU on NVIDIA (#58957)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix][Model] Gemma4: register aliased embedding scalars as buffers (#54213)

Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>

* [Bugfix][Spec Decode] Implement get_top_tokens() on the ROCm DeepSeek V4 MTP drafter (#57568)

Signed-off-by: BaoYunkai <ybao@amd.com>

* [Perf][Qwen3.8] Reduce PLE metadata construction overhead (#58114)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm][Refactor] Move DeepSeek-V4/V4.1 multi-stream overlap gate to ROCm platform (#58983)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Multimodal] Avoid extra d2d for encoder cudagraph with fused input norm (#56711)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Logging] Preserve application log record factories (#58747)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix] Use the correct repository revision for secondary artifact loaders (#57461)

Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Perf] Kimi-K3 Enable sharded latent MoE up-projection under EP (#54956)

Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Feature] Add fixed-token prefill scoring (#54335)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [LoRA] Support variable num_labels for sequence classification (#57766)

Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [ROCm][Perf] Replace torch.topk in DSA candidate block selection (#58208)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [ROCm]Keep LMCache OpenTelemetry on the image's 1.40 stack (#59056)

Signed-off-by: Micah Williamson <micah.williamson@amd.com>

* [Bugfix][Kimi-K3] Refresh DSpark context KV cache pointers after the KV cache is re-bound (#58814)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [KVConnector][NIXL] Support packed MLA KV layouts in pipeline-parallel push prefill (#50499)

Signed-off-by: zixi-qi <zixi@inferact.ai>

* [Quant] Use canonical N-first weight format for CT WNA16 MoE (#52798)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: HDCharles <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Relax packed_qk_rope_ correctness test to one ULP (#59008)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Pin OpenTelemetry to LMCache's cap in the ROCm images (#59051)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Increase timeout for Entrypoints Unit (#59076)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][MLA] Add an AITER ASM round-robin decode route for DCP multi-token verify (#56861)

Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix] Don't sync-police or retry FlashInfer all-reduce workspace creation in eager mode (#58498)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix resumable request + async scheduling handoff race (#58259)

Signed-off-by: Yifan Zong <yzong@redhat.com>

* [Model Runner V2][Spec Decode] Support spec decode with draft model (#43091)

Signed-off-by: Icey <1790571317@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][MRV2] Allow FULL decode graphs for one-token prompt tails (#58400)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Bugfix][Scheduler] Preserve logprobs across streaming continuations (#57447)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>

* [Mypy] Fix mypy typing for Voxtral and vision models (#58251)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Core] Rework scheduler `skipped_waiting` queue (#58947)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Bugfix][MRV2][Spec Decode] Reject draft slots that were never proposed (#58784)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][CI] Add missing test coverage for upstream parity (#50519)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][Scheduler] Refresh max tokens for streaming continuations (#57676)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Spec Decode] Avoid triton recompiles in the acceptance estimator (#57107)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>

* [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8664f67d8ae1538e5e5cec2b1ce3f258dd)

* [XPU][CI]Skip test_abort_timeout_on_prefiller in nightly (#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db1ca080793dc2351144cb39513fd7984ca)

* [Bugfix][Engram] Keep THP tables private when resolving shared memory (#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c352f079c2cb8f46752fd0317a2745050bf)

* [ROCm][CI] Expand MI355 mirrors and route MIG-sized jobs to DPX (#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd64f5ef75dbe1471ee88c009003b3d03ec)

* [Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse retention (#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bddbeab6b4a0d5861dfcb171bf61ce2109d6)

* [Core][BugFix] Tag prefix-cache extra keys by source (#51899)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e7ed5b6f372f0898a67e48d3365fa28139)

* [Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e097b905127b1d5f09a7e93bd1095e87a6c)

* [Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf53724e06cef7f07628ae8fd4cc6c96f11a)

* [Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f668608155bc8a8c74cb97e3ddd0d3053085)

* [Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc757cffc11b55689d9cc33d844d9606516)

* [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a4297921c699f9aa58f9c6395e95ede718)

* [Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa130fac1e1252330a680fbed45ad05eba64f)

* [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)

(cherry picked from commit 5463fe4962785cdc3383477bf3af6533e7647dfd)

* [Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87cca25690fef6bd12667fd0d26690d949af)

* [Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec22ad9bb098393021b956feaa97081fc85)

* [Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae4d746f64c22677348d1cc120eea6f7439)

* [Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e615c20ee4ff0cd82aac592c7d6310aa7c3)

* [Core] Include the LoRA path in prefix-cache block hashes (#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e075ed461cff922aa710d387281f651940)

* [Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360c9eebb3b4b460ec49f6606ef0a08fcae2)

* [CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d7332db897885fdc0e6ed61b28c969d743)

* [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba9b7d18f9746d989e389a51a94b0f96f6f)

* [Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a6963664537ed21172e2ec12e96e3a2dcd3718c)

* [Bugfix][HiSparse] Size the KV cache from the groups HiSparse allocates (#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2e4489126f1c21868699f3890021f01b73)

* [Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with a saturated GPU pool (#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8ac1f8a7e8fb50cf1c56ff96649fcadccc6)

* [CI] Drop a test that depends on #57930 from the #59309 backport

Resolving the #59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from #57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream #59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>

* [Bugfix] Fix minimax-m3 multimodal processor compatability with Transformers v5.18 (#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f160a2c0abcb5902acc3c91a14c38a4af173)

* [Misc] Add Transformers version upper bound in requirements (#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298457dde7b4554b3b4e20b238c0ac2c3a65)

* Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0

Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2eb #28, 76a3e6d72, 6d4c75611 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a46, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 18.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c19, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - per-group BlockPools, sampler reserve, padded-page block fill, drain
    hook / on_kv_manager_created, KVBlockZeroer clamp: conflict with
    upstream's single-pool allocator; no turbo-attn consumer for the hooks
  - rotary fast-path registry: upstream guards the import (1f60771c7,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Signed-off-by: chao.huan <chao.huan@nio.com>
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: errmakov <ide404@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Signed-off-by: yashasvi <yashasvi@ibm.com>
Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Signed-off-by: Cheng Rui <286040359@qq.com>
Signed-off-by: Yifan Zong <yzong@redhat.com>
Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Linkun Chen <github@lkchen.net>
Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Ke Wen <kwen@nvidia.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: zixi-qi <zixi@inferact.ai>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: Willian <willian@willian.email>
Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Vincent <vincexxchan@gmail.com>
Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: QHarshil <harshil_c@hotmail.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>
Signed-off-by: BaoYunkai <ybao@amd.com>
Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Stefan Kon…
arbi-dev added a commit to arbicity/vllm-turbo that referenced this pull request Oct 7, 2026
* [Bugfix][XPU] store the pointer raw bit pattern instead of its numeric value (#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Attention][CPU] Run Zen CPU encoder attention on zentorch SDPA (#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [MRV2] Validate MRV2 entrypoint logits processors (#57728)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][Rust Frontend] Prevent MM timing from enabling debug tracing (#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [XPU][UT] Align HF and vLLM inputs for Qwen2 embedding test by preventing Sentence Transformers from applying chat template (#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>

* [CPU] Gate the AVX10.2 paths on compiler support (#58133)

Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Perf][Frontend] Offload streaming derender detokenization (#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>

* [Multimodal] Reuse the supplied tokenizer in the MiniMax-M3 VL processor (#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>

* [XPU] enable XPU GRAPH by default (#51600)

Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>

* [ROCm][DSv4.1][Perf] Emit MXFP8 from the sparse decode reduce and run wo_a as a grouped FP8 GEMM (#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Remove `.gemini/` and `CLAUDE.md` (#58541)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Pooling] Fix JinaVL label configuration and restore multimodal tests (#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>

* [Chore] Use Transformers v5 names and drop redundant processor `use_fast` (#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Refactor] Remove dead or duplicate tests (#58446)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][Attention] Bound FlashInfer prefill dequantization scratch (#57918)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>

* fix(config): apply presence_penalty/frequency_penalty from override-generation-config (#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Resolve the Hub revision once per repo (#56092)

Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Frontend] Remove the slow tokenizer mode (#58545)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Revert "[DSpark] Support pipeline-parallel targets in aggregated serving (#56956)" (#58484)

* [transformer] RMSNorm matching for alternative rsqrt (#54461)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

* [Bugfix] Count unsplit Idefics3 image patches (#48760)

Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>

* [Bugfix] Keep JIT warmup under enforce-eager when fault tolerance is on (#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Core] Skip JIT monitor when JIT warmup is disabled (#58590)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Fast Start] Wait for weight cache daemon readiness (#58370)

* [Bugfix][Quantization] Add Humming to the W4A8 (INT4xFP8) MoE oracle (#58427)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Refactor] Move auxiliary files out of the repository root (#58572)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Cleanup] Remove online quantization support in `fp8.py` in favor of online shorthands (#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm] Fix misrouting race-condition in multi-decode P/D disagg with mori-io (#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: constrained reads over the request's logprob_token_ids (#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead (#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config (#58535)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [DFlash] Capture the context K/V precompute in the draft CUDA graph (#57632)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix][Outlines] Fix EOS termination and unconstrained masks after rejected drafts (#58612)

* [Bugfix][KV Cache] Fix incremental multimodal block hashing (#51694)

Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>

* [XPU][CI] enable prompt embeds tests on XPU (#58283)

Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [CI] Report to CRCR after all jobs finish, gated on the build's long pole (#58628)

* [PD][PushConnector] Record last activity of remotes on the D side (#52245)

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>

* [BUGFIX] fix ovis2_5 multimodal tokens (#52623)

Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>

* [Bugfix][Core] Keep every multimodal feature in the partial-block KV event (#58288)

Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355 (#58558)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Quantization] Give LM heads standard linear metadata (#58444)

Signed-off-by: mgoin <mgoin64@gmail.com>

* [Bugfix][Mamba] Restore prompt-tail prefix-cache hits with MTP (#58368)

Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>

* [Perf] Parallelize registered CUDA Triton kernel warmup at startup (#58582)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>

* [KV Connector] Fix DecodeBench fp8 fill values and add a startup fill mode (#58472)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* Release prompt_embeds tensor when its InputBatch slot is freed (#57988)

Signed-off-by: khushali9 <khushali.desai9@gmail.com>

* [Bugfix][KVConnector] Finalize saves on steps without a forward (#57775)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Kimi Code <noreply@moonshot.ai>

* [Bugfix][Frontend] Count reasoning tokens for Harmony, DeepSeek-V3 and Step3 parsers (#58626)

Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] GLM-5.3-Flash: launch the kpool paged MQA logits in the varlen mode its schedule was built with (#55270)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Bugfix] Accept EOS after grammar finish in outlines backend; reject json_object at validation (#57743)

Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>

* [Perf] Batch Mamba2 prefill SSM state saves, removing GPU<->CPU syncs (#49371)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [CI] Run DFlash2 NVFP4 acceptance test on B200; skip it on H200 35GB MIG (#58496)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Bugfix][MRV2] Treat padded prompt tails as spec-decode rows for hybrid models (#58434)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Kimi-K3][Perf] Dispatch GEMM for vision patch embedder (#58527)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Minimax-M3][Perf] Use triton_mrope for vision tower + int64 offset fix for triton_mrope (#58526)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Quantization] Refresh online NVFP4 scales before reload post-processing (#57954)

Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Perf][DSv4.1] Restore the fused query RMSNorm + MXFP8 quantization path (#57679)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

* [Bugfix][LogitsProcessor] Validate ':' separator in custom logits processor FQCN (#56020)

Signed-off-by: 100milliongold <gadian88@gmail.com>

* [ROCm][CI][AITER Coverage] Harden MoE sorting-backend/dispatch env-var test matrix (#58393)

Signed-off-by: Divakar Verma <divakar.verma@amd.com>

* [gRPC] Fix ping tolerance so long non-streaming RPCs are not dropped (#55102)

Signed-off-by: Wei Gong <wei@together.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [Perf][Rust Frontend] Make histogram observations lock-free (#58574)

Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm][Perf] MXFP8 GEMM on native 32x32 block scales for gfx950 (#58510)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Keep pinned PLE prefetch ids out of the CUDA graph pool (#58489)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][Engram] Serialize offloaded lookups and pack host tables into huge pages (#56926)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Quantization] Fix MXFP8 startup crash on layers below mm_mxfp8 shape limits (#54223)

Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [ROCm] Cut 69 wasted contiguous copies per decode step from the skinny GEMM path (#58566)

Signed-off-by: lifulu <fululi12@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][Build] Filter crate tags from vLLM version detection (#57744)

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>

* [Perf][Distributed] Add low-SM multimem reduce-scatter for SM100/SM103 (#55072)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>

* [PP][XPU]Add the flag to control microbatch feature on MRV2+PP (#55145)

Signed-off-by: yisheng <yi.sheng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

* [Feature] Triton kernel dispatcher (#43048)

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

* [ROCm] Fix CI runtime and tests for MI355 DPX (#58244)

Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Mahesh Kunreddi <mahesh.kunreddi@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [CI][ROCm] Prevent Model Executor apt stalls (#58607)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Qwen4Exp][ROCm] PLE n-gram table CPU offload (#57497)

Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: opencode+deepseek-v4-flash+vllm <opencode+deepseek-v4-flash+vllm@example.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [MM] Add Triton kernel for mm_input_normal. (#56798)

Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Structured Outputs] Parse Lark grammars natively in the xgrammar backend (#58321)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>

* [ROCm] Credit ROCm/aiter for the block32 GEMM's packed kernel and in-launch split-K (#58659)

Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Frontend] Handle Disable Thinking in /v1/messages (#58613)

Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: jryberg <johan.ryberg@security.ntt>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Bugfix] Stop leaking the internal field name in the max_tokens validation error (#58336)

Signed-off-by: shallow10 <495593563@qq.com>

* [Bugfix][KV Cache][MLA] Align packed block strides for V3.2 sparse MLA (#55528)

Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][DSv4] Fuse inverse RoPE + FP8 quant into FlashInfer sparse MLA (#58621)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [UX][Frontend] Introduce `vllm preload` cli for fast restart (#56680)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [MoE] Defer the TRTLLM-Gen top-k finalize on the modular path (#58635)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Docs] Add return annotation to `fused_mm_input_norm_triton` (#58687)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [CI] Shard (H100) Helion Kernels five ways (#58645)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Core] Model console logging as CLI configuration (#57205)

Add `--logging-config` CLI argument which can be supplied as
JSON or using dotted arguments. The `--log-level` argument
is provided for convenience, and `--log-config-file` is deprecated
in favor of `--logging-config.pylogging_config_file`.

Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Co-authored-by: AI Assistant <noreply@openai.com>

* [ROCm][CI] Pass weight_shape in MXFP8 block32 linear tests (#58698)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [SpecDecode] Add LiLiCorr drafter (#57934)

Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>

* [MRV2] Minor model_runner.py code cleanup (#58610)

Signed-off-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix generative scoring body cancellation (#57729)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Pooling] Preserve reranker tokenization with document limits (#57666)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Kernel][Perf] Register-resident path for per-token-group 8-bit quant (#55330)

Signed-off-by: chao.huan <chao.huan@nio.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Perf][Kernel] Vectorized flat abs-max for dynamic per-tensor FP8 quantization (#58194)

Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>

* [Core][Logging] Fix JSON logging process decoration (#57957)

Signed-off-by: Mark McLoughlin <markmc@redhat.com>

* [Bugfix] Stop allocator fragmentation from shrinking the KV cache during memory profiling (#58430)

Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Keep logprobs of parser-suppressed streaming chunks (#58583)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Yanxiao Zhao <39199723+sdpkjc@users.noreply.github.com>
Co-authored-by: Prakhar Agarwal <270064960+agarwalprakhar2511@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix][GLM-5.3-Flash] SM90 sparse MLA: index_kpool mismatch leads to corruption via unread query token (#58704)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Benchmark] Record model_id in bench latency/throughput --output-json (#58112)

Signed-off-by: yashasvi <yashasvi@ibm.com>

* [Bugfix][GLM-5.3-Flash] kpool corruption with speculative decoding (#58454)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [GLM5.3 Perf] Optimize glm 5.3 metadata op, 1.6~4.8x kernel level performance improvement (#58450)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Bugfix][KV Connector] Reap expired NIXL leases behind a heartbeated head (#58292)

Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>

* [ROCm][Kimi-K3] Optimize low-concurrency speculative KDA (#58045)

Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Cover the AITER MQA logits dispatch on gfx950 (#58724)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Add quantized MoE serving test for gfx950 (#58748)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Test AMD DeepSeek V4 MoE routing against a PyTorch reference (#58740)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Perf] DiffusionGemma: one-pass sampler statistics kernel (#58226)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [ROCm][CI] Expand single-GPU coverage on MI355 DPX (#57599)

Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [CI] [MRV2] Restore MRV2 pp dp coverage (#57735)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

* [Bugfix][CI] Report subprocess test skips as skips, not passes (#58701)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Run the MLA attention+quant fusion test on ROCm (#58717)

Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm] Bump torch 2.13, triton 3.8, torchaudio, torchvision (#50605)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: jpvillam <juan.villamizar@amd.com>

* [Feature][Frontend] Add granite_thinking_parser reasoning parser for Granite 4.2 (#55957)

Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Respect max_output_tokens in the Harmony tool-call loop (#58551)

Signed-off-by: errmakov <ide404@gmail.com>
Co-authored-by: Du Bin <8174807+dubin555@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix][Frontend] Document 404 response for `/generative_scoring` (#58788)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>

* [Perf][Frontend] Defer reasoning usage recounts for non-continuous chat streams (#56067)

Signed-off-by: Cheng Rui <286040359@qq.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Bugfix] Default missing detail for Responses API input images (#57241)

Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>

* [watermarking] golden tests for backwards compatibility (#56809)

Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Bugfix] Don't drop the rest of the allocator config when toggling expandable segments (#57982)

Signed-off-by: Oxana Korzh <okorzh@amd.com>

* [AuxOutput] Only require Model Runner V2 on GPU platform (#58205)

Signed-off-by: Linkun Chen <github@lkchen.net>

* [Pooling] Preserve BERT-family heads for raw logits (#57664)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* fix: perf: use startswith(x, i) instead of string slicing to avoid O(N^2) (#52580)

Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states (#48419)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Co-authored-by: Claude <noreply@anthropic.com>

* [Bugfix][Reasoning] Count Kimi K3 reasoning tokens (#58372)

Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [CI] Split (H200 MIG/MI300) Basic Correctness into named jobs (#57054)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>

* [Bugfix][Frontend] Fix Inkling tool name leaking into content after reasoning (#58792)

Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>

* [Core] Bound draft-token RPC waits by the execute-model timeout (#58779)

Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Co-authored-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>

* [Perf][DSv4.1] Shard the Engram wkv projection across TP ranks (#58678)

Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>

* [RL] Add sharding-aware NCCL M2N weight transfer (#51520)

Signed-off-by: Ke Wen <kwen@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen@inferact.ai>

* [Bugfix][ROCm] AMD-Quark mixed-precision DeepSeek-V4.1 support (#57071)

Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix][Frontend][Rust Frontend] Update DeepSeek V4.1 Flash reasoning effort mappings (#58316)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>

* [CI] Only isolate the registry tests that need a fresh process (#58764)

Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Spec Decode] Enable Gemma4 DSpark adaptive verification with FlashInfer (#57263)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [CI] Split (B200) Miscellaneous Kernels into mHC, FLA Ops and Misc named jobs (#58609)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>

* [Security] Gate per-request multimodal processor kwargs (#58830)

Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Elastic EP] Fix EPLB load statistics during scaling (#58473)

Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>

* [Refactor] Remove dead tests utils (#58803)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [Perf][MoE] Index expert mapping lookups in RoutedExperts.load_weights (#58720)

Signed-off-by: Willian <willian@willian.email>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>

* [Bugfix] Fix Anthropic Thinking Disabled with P/D (#58786)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Bugfix][Frontend] Detect Anthropic inline-system merge against the resolved chat template (#58754)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>

* [Mypy] Fix mypy typing for Qwen and Qianfan models (#58046)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [CI] Reduce CUDA graph mode test overhead (#58749)

* [GLM5.3 Bug] Fix sparse indexer attn topk backend selection (#58594)

Signed-off-by: yewentao256 <zhyanwentao@126.com>

* [CI] Stabilize batch submission in full CUDA graph tests (#58810)

* [Bugfix][DSV4.1] Avoid host sync in ViT CUDA graph replay metadata (#58499)

* [Security] Harden message sanitization (#58832)

Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>

* [Kernel][DSV4.1] Fuse MoE finalize into the TP all-reduce + mHC boundary (#58586)

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [PCP][DCP] Support DCP target model with non-DCP Dspark (#56723)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>

* [Kernel] Bump FlashKDA to keep the recurrent state in fp32 (#58846)

Signed-off-by: Simon Veitner <sveitner@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector (#57251)

Signed-off-by: Vincent <vincexxchan@gmail.com>

* [mooncake] support CUSTOM_MEM_POOL in vllm (#49300)

Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: bruce.xu <bruce.xb@alibaba-inc.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Qwen3.8-Flash-Next] Avoid memory fragmentation in QSA indexer logits workspace (#57105)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix] Disable sequence parallelism / async TP under batch invariance and add a TP regression test (#56377)

Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Kimi-K3] Make VLLM_ROCM_USE_AITER_MOE_SITUV2 select a4w4/a8w4/a16w4 (#58201)

Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>

* [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL (#50047)

Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Co-authored-by: xijiade.aihemaiti <3146335281@qq.com>

* [Bugfix] V1: clear stale allowed_token_ids mask in InputBatch.condense (#43931)

Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [Perf][DSv4.1] Fuse small-batch WO-A with inverse RoPE and MXFP8 quant on SM100/SM103 (#58634)

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [Attention][CPU] Use zentorch SDPA for CPU MLA prefill (#54967)

Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Perf][Attention] Remove D2H sync from FlashInfer SM90 sparse MLA plan under async scheduling (#58684)

Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>

* [CI] Split (H200 MIG 35GB) Spec Decode Speculators + MTP into 4 named jobs (#57237)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Thang Nguyen <thang.nguyen@inferact.ai>
Co-authored-by: Kimi Code <noreply@moonshot.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][Perf] Enable layer-aware CSA2 multi-stream overlap for DeepSeek-V4.1-Flash (#57407)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Perf][Pooling] Avoid blocking seq_lens GPU-to-CPU copy for pooling in FlashInfer metadata builder (#57214)

Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>

* [Bugfix] Support repsonse_format + tool_choice=auto (#56086)

Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: pablopupo <145598901+pablopupo@users.noreply.github.com>
Co-authored-by: hubunt <150658615+hubunt@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Fast Start] Add `/health` endpoint for the weight cache daemon (#58552)

Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>

* [Perf][MoE] Use fused MiniMax2 routing with non-unit routed scaling (#58880)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>

* [ROCm][Bugfix] Fall back to default GEMM for CPU tensors on ROCm builds (#58923)

Signed-off-by: fai <fangzhouai@gmail.com>

* [Bugfix][KV Connector] Retry Mooncake bootstrap registration on timeout (reopens #55763) (#58919)

* [Frontend] Switch Python Harmony dependency to oss-harmony (#55128)

Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm] Bump AITER to v0.1.23 (#58867)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [Bugfix][Frontend] Count Responses reasoning tokens per tool round (#58927)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [KimiViT][Perf] Fuse per-layer QK RoPE into one in-place kernel (#58651)

* [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models (#58239)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>

* [Bugfix][Frontend] Sample batched chat completions from the adjusted requests (#58929)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][Frontend] Use a fresh parser per choice in non-streaming chat completions (#58939)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Bugfix][EPD] Skip sampling for encoder-only async steps (#58490)

Signed-off-by: Tianyu Guo <guoty@inferact.ai>

* [CPU] Build CPU wheels on Ubuntu 22.04 with AMX-FP8 support (#58515)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>

* [Test][ROCm] Stabilize the mixed OLMoE LoRA test (#58945)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>

* [Model] Enable LoRA support for RobertaForSequenceClassification (#58884)

Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Co-authored-by: jz-yolo <jz-yolo@users.noreply.github.com>

* [Bugfix][Frontend] Apply Harmony adjust_request in batched chat completions (#58958)

Signed-off-by: sfeng33 <4florafeng@gmail.com>

* [Minimax-M3] Add Encoder CUDA graph support (#58673)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <codex@openai.com>

* [Perf][MRV2] Reuse Mamba/GDN metadata across KV cache groups (#58762)

Signed-off-by: Luca Motz <luca.motz@icloud.com>
Co-authored-by: Xin Yang <xyangx@amazon.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>

* [Bugfix] Fix the two multimodal root tests that fail on main (OpenPangu-VL embed merge, MiMo sink test fixture) (#58900)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Benchmark] Add Responses API backend to vllm bench serve (#54628)

Signed-off-by: QHarshil <harshil_c@hotmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>

* [CI] Allowlist-shrink batch 1: wire 17 root-level tests + drop 5 stale watermarking entries into misc.yaml (#58055)

Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [CPU] Use accelerator memory API in DiffusionGemma (#58964)

Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [CPU] Add video inferencing via torchcodec on s390x (#58693)

Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>

* [Skills] Update kernel-microbenchmark to include ROCm (#58646)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Kimi <noreply@moonshot.cn>

* [Misc] Name each backend and its kernel block sizes in block-size errors (#58557)

Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: khluu <khluu000@gmail.com>

* [CI] Split Dynamic Shapes out of (H200 MIG 35GB) PyTorch Compilation + (MI300) mirror (#58451)

Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Bugfix][Qwen4Exp] Release the profiling KV cache held by QSA key views (#58961)

Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [CPU] Include vLLM Recipes tooling in release image to deploy models using vLLM Recipes (#58796)

Signed-off-by: louie-tsai <louie.tsai@intel.com>

* Revert "[CI] Shard (H200 MIG 35GB / MI355 DPX) Entrypoints Integration (Pooling) into named jobs (#58652)" (#59011)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Perf][Qwen4Exp] Fuse HC down projection and SiLU on NVIDIA (#58957)

Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>

* [Bugfix][Model] Gemma4: register aliased embedding scalars as buffers (#54213)

Signed-off-by: Yannick Schnider <Yannick.Schnider1@ibm.com>

* [Bugfix][Spec Decode] Implement get_top_tokens() on the ROCm DeepSeek V4 MTP drafter (#57568)

Signed-off-by: BaoYunkai <ybao@amd.com>

* [Perf][Qwen3.8] Reduce PLE metadata construction overhead (#58114)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [ROCm][Refactor] Move DeepSeek-V4/V4.1 multi-stream overlap gate to ROCm platform (#58983)

Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>

* [Multimodal] Avoid extra d2d for encoder cudagraph with fused input norm (#56711)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>

* [Bugfix][Logging] Preserve application log record factories (#58747)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix] Use the correct repository revision for secondary artifact loaders (#57461)

Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [ROCm][Perf] Kimi-K3 Enable sharded latent MoE up-projection under EP (#54956)

Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Feature] Add fixed-token prefill scoring (#54335)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [LoRA] Support variable num_labels for sequence classification (#57766)

Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>

* [ROCm][Perf] Replace torch.topk in DSA candidate block selection (#58208)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* [ROCm]Keep LMCache OpenTelemetry on the image's 1.40 stack (#59056)

Signed-off-by: Micah Williamson <micah.williamson@amd.com>

* [Bugfix][Kimi-K3] Refresh DSpark context KV cache pointers after the KV cache is re-bound (#58814)

Signed-off-by: Oxana Korzh <okorzh@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [KVConnector][NIXL] Support packed MLA KV layouts in pipeline-parallel push prefill (#50499)

Signed-off-by: zixi-qi <zixi@inferact.ai>

* [Quant] Use canonical N-first weight format for CT WNA16 MoE (#52798)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: HDCharles <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [CI][Bugfix] Relax packed_qk_rope_ correctness test to one ULP (#59008)

Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][CI] Pin OpenTelemetry to LMCache's cap in the ROCm images (#59051)

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [ROCm][CI] Increase timeout for Entrypoints Unit (#59076)

Signed-off-by: Djordje Ramic <djoramic@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>

* [ROCm][MLA] Add an AITER ASM round-robin decode route for DCP multi-token verify (#56861)

Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Bugfix] Don't sync-police or retry FlashInfer all-reduce workspace creation in eager mode (#58498)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [Bugfix] Fix resumable request + async scheduling handoff race (#58259)

Signed-off-by: Yifan Zong <yzong@redhat.com>

* [Model Runner V2][Spec Decode] Support spec decode with draft model (#43091)

Signed-off-by: Icey <1790571317@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>

* [Perf][MRV2] Allow FULL decode graphs for one-token prompt tails (#58400)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>

* [Bugfix][Scheduler] Preserve logprobs across streaming continuations (#57447)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>

* [Mypy] Fix mypy typing for Voxtral and vision models (#58251)

Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>

* [Core] Rework scheduler `skipped_waiting` queue (#58947)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>

* [Bugfix][MRV2][Spec Decode] Reject draft slots that were never proposed (#58784)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>

* [ROCm][CI] Add missing test coverage for upstream parity (#50519)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>

* [Bugfix][Scheduler] Refresh max tokens for streaming continuations (#57676)

Signed-off-by: 0xsensei <prblmslvr.aditya@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

* [Perf][Spec Decode] Avoid triton recompiles in the acceptance estimator (#57107)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>

* [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8664f67d8ae1538e5e5cec2b1ce3f258dd)

* [XPU][CI]Skip test_abort_timeout_on_prefiller in nightly (#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db1ca080793dc2351144cb39513fd7984ca)

* [Bugfix][Engram] Keep THP tables private when resolving shared memory (#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c352f079c2cb8f46752fd0317a2745050bf)

* [ROCm][CI] Expand MI355 mirrors and route MIG-sized jobs to DPX (#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd64f5ef75dbe1471ee88c009003b3d03ec)

* [Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse retention (#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bddbeab6b4a0d5861dfcb171bf61ce2109d6)

* [Core][BugFix] Tag prefix-cache extra keys by source (#51899)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e7ed5b6f372f0898a67e48d3365fa28139)

* [Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e097b905127b1d5f09a7e93bd1095e87a6c)

* [Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)

Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf53724e06cef7f07628ae8fd4cc6c96f11a)

* [Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f668608155bc8a8c74cb97e3ddd0d3053085)

* [Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc757cffc11b55689d9cc33d844d9606516)

* [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a4297921c699f9aa58f9c6395e95ede718)

* [Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa130fac1e1252330a680fbed45ad05eba64f)

* [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)

(cherry picked from commit 5463fe4962785cdc3383477bf3af6533e7647dfd)

* [Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87cca25690fef6bd12667fd0d26690d949af)

* [Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec22ad9bb098393021b956feaa97081fc85)

* [Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae4d746f64c22677348d1cc120eea6f7439)

* [Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e615c20ee4ff0cd82aac592c7d6310aa7c3)

* [Core] Include the LoRA path in prefix-cache block hashes (#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e075ed461cff922aa710d387281f651940)

* [Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360c9eebb3b4b460ec49f6606ef0a08fcae2)

* [CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d7332db897885fdc0e6ed61b28c969d743)

* [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba9b7d18f9746d989e389a51a94b0f96f6f)

* [Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a6963664537ed21172e2ec12e96e3a2dcd3718c)

* [Bugfix][HiSparse] Size the KV cache from the groups HiSparse allocates (#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2e4489126f1c21868699f3890021f01b73)

* [Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with a saturated GPU pool (#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8ac1f8a7e8fb50cf1c56ff96649fcadccc6)

* [CI] Drop a test that depends on #57930 from the #59309 backport

Resolving the #59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from #57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream #59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>

* [Bugfix] Fix minimax-m3 multimodal processor compatability with Transformers v5.18 (#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f160a2c0abcb5902acc3c91a14c38a4af173)

* [Misc] Add Transformers version upper bound in requirements (#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298457dde7b4554b3b4e20b238c0ac2c3a65)

* Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0

Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2eb #28, 76a3e6d72, 6d4c75611 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a46, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 27.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner
  - per-group KV pool for hybrid models, re-expressed on v0.31's single
    backing allocation: the O(1) Mamba/GDN state groups get their own
    BlockPool sized for max_num_seqs, a region of the allocation after the
    attention pool; their pages are not unified with attention pages;
    per-pool admission, events and usage; CoW copies and warmup/profiling
    block ids per pool; VLLM_SAMPLER_RESERVE_MIB headroom

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c19, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - padded-page block fill (the split pool removes the page unification
    it worked around), drain hook / on_kv_manager_created, KVBlockZeroer
    clamp: no turbo-attn consumer
  - rotary fast-path registry: upstream guards the import (1f60771c7,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: fai <fangzhouai@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Jellow <49915976+CZT0@users.noreply.github.com>
Signed-off-by: Jellow <dvdx@foxmail.com>
Signed-off-by: Lin, Fanli <fanli.lin@intel.com>
Signed-off-by: Fanli Lin <fanli.lin@intel.com>
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
Signed-off-by: Milosz Grunwald <milosz.grunwald@intel.com>
Signed-off-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: khushali9 <khushali.desai9@gmail.com>
Signed-off-by: Samyabrata Maji <116789799+sammaji@users.noreply.github.com>
Signed-off-by: SIDDARTHA REDDY <75976672+SIDDARTHAREDDY8@users.noreply.github.com>
Signed-off-by: samuelkim7 <samuelmwkim@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Signed-off-by: 100milliongold <gadian88@gmail.com>
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
Signed-off-by: Wei Gong <wei@together.ai>
Signed-off-by: lifulu <fululi12@amd.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: yisheng <yi.sheng@intel.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Mathew Odden <modden@redhat.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Lingpeng Jin <103567126+valarLip@users.noreply.github.com>
Signed-off-by: jryberg <johan.ryberg@security.ntt>
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Signed-off-by: shallow10 <495593563@qq.com>
Signed-off-by: lz <145014769+200lz@users.noreply.github.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Thang Nguyen <thangnguyenvn647@gmail.com>
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Signed-off-by: Andrii Skliar <andreyws96@gmail.com>
Signed-off-by: chao.huan <chao.huan@nio.com>
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2-redhat@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
Signed-off-by: errmakov <ide404@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Signed-off-by: yashasvi <yashasvi@ibm.com>
Signed-off-by: GokayAI <60583610+gokay-ai@users.noreply.github.com>
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Signed-off-by: Sheral Kumar <shekumar@amd.com>
Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Signed-off-by: jpvillam <juan.villamizar@amd.com>
Signed-off-by: Yousaf shah <yousaf.shah@gmail.com>
Signed-off-by: Cheng Rui <286040359@qq.com>
Signed-off-by: Yifan Zong <yzong@redhat.com>
Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai>
Signed-off-by: Simon Veitner <sveitner@redhat.com>
Signed-off-by: Oxana Korzh <okorzh@amd.com>
Signed-off-by: Linkun Chen <github@lkchen.net>
Signed-off-by: Ricardo-M-L <ricardoporsche001@icloud.com>
Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Baljinder Hothi <baljinder.hothi@cohere.com>
Signed-off-by: VS Chandra Mourya <219748331+vschandramourya@users.noreply.github.com>
Signed-off-by: Shuolei Wang <shuoleiwang123@gmail.com>
Signed-off-by: Ke Wen <kwen@nvidia.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Signed-off-by: zhec <chengyunfei@ruc.edu.cn>
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: zixi-qi <zixi@inferact.ai>
Signed-off-by: Juan Pérez de Algaba <jperezde@redhat.com>
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: Willian <willian@willian.email>
Signed-off-by: Ashraf Bhuiyan <mbhuiyan@redhat.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Signed-off-by: Vincent <vincexxchan@gmail.com>
Signed-off-by: bruce.xu <bruce.xb@alibaba-inc.com>
Signed-off-by: LioEinaudi <zhao3024667639@gmail.com>
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: Yannik Hinteregger <37209495+YannikHinteregger@users.noreply.github.com>
Signed-off-by: Varshith <kvarshithgowda@gmail.com>
Signed-off-by: Rakul Chauhan <rakul.chauhan@amd.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com>
Signed-off-by: frankwang28 <frank.wbb@hotmail.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
Signed-off-by: Xun Sun <UNIDY2002@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Anton Peganov <apeganov@nvidia.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Signed-off-by: Tianyu Guo <guoty@inferact.ai>
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com>
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: jz-yolo <jz-yolo@users.noreply.github.com>
Signed-off-by: Jane Zhu <jane.zhu@slack-corp.com>
Signed-off-by: Luca Motz <luca.motz@icloud.com>
Signed-off-by: QHarshil <harshil_c@hotmail.com>
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com>
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Signed-off-by: Eugenio "Jay" Zuccarelli <11176606+jayzuccarelli@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Yannick Schnider <Yannick.Schnider1@ib…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rocm Related to AMD ROCm

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants