Conversation
…P8 UE8M0 packed path to group_size 128 (vllm-project#51359) Signed-off-by: BabyDrangoner <148877251+BabyDrangoner@users.noreply.github.com> Signed-off-by: kiroxu <148877251+BabyDrangoner@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…ce improvement (vllm-project#51311) Signed-off-by: yewentao256 <zhyanwentao@126.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…47808) Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com> Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com> Signed-off-by: Lucas Wilkinson <wilkinson.lucas@gmail.com> Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: OpenAI Codex <codex@openai.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: OpenAI Codex <noreply@openai.com>
vllm-project#49139) Signed-off-by: fxfxfxfxfxfxfxfx <227935476@qq.com> Co-authored-by: Michael Goin <mgoin64@gmail.com>
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
…2035) Signed-off-by: Yongye Zhu <zyy1102000@gmail.com> Co-authored-by: Kimi Code <noreply@moonshot.ai>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
…roject#50017) Signed-off-by: Andy Friedrich <afriedri@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…cs groups instead of UNKNOWN (vllm-project#51218) Signed-off-by: Yifan Jiang <19356972+yifjiang@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
…llm-project#51821) Signed-off-by: Stefan Koncarevic <Stefan.Koncarevic@amd.com> Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: khluu <khluu000@gmail.com> Co-authored-by: OpenAI Codex <noreply@openai.com>
…t#51624) Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com> Co-authored-by: Akash kaothalkar <akash.kaothalkar@ibm.com> Co-authored-by: Antigravity <antigravity@google.com>
Signed-off-by: jdebache <jdebache@nvidia.com>
…lm-project#51879) Signed-off-by: Ziqi Fan <ziqif@nvidia.com>
…k granularity (vllm-project#51614) Signed-off-by: Ziqi Fan <ziqif@nvidia.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…lm-project#51256) Signed-off-by: HF-001 <1670186653@qq.com>
…#52024) Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com> Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Yan Ma <yan.ma@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…ct#51772) Signed-off-by: Yongye Zhu <zyy1102000@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
…ject#48215) Signed-off-by: arthurgao2003 <arthurgao2003@gmail.com> Signed-off-by: Arthur Gao <arthurgao2003@gmail.com> Signed-off-by: 高杨懿 <15312196+gao-yangyi@user.noreply.gitee.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: 高杨懿 <15312196+gao-yangyi@user.noreply.gitee.com> Co-authored-by: OpenAI Codex <codex@openai.com> Co-authored-by: OpenAI Codex <noreply@openai.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…ect#51931) Signed-off-by: frank-suwen <suwenw2@outlook.com>
Signed-off-by: kiroxu <148877251+BabyDrangoner@users.noreply.github.com> Co-authored-by: Codex <noreply@openai.com>
…vllm-project#52092) Signed-off-by: jiang1.li <jiang1.li@intel.com>
…ct#50874) Signed-off-by: tbarnatan <tbarnatan@nvidia.com> Signed-off-by: zjy0516 <riverclouds.zhu@qq.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: zjy0516 <riverclouds.zhu@qq.com>
…project#51251) Signed-off-by: hotTea <958436561@qq.com> Signed-off-by: HMCCMH <chenminghaoscu@163.com> Co-authored-by: HMCCMH <chenminghaoscu@163.com>
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
…uantize_input (vllm-project#52603) Signed-off-by: xuebwang-amd <xuebwang@amd.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
…-project#52810) Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com> Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…ject#51875) Signed-off-by: Russell Bryant <rbryant@redhat.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…oject#51368) Signed-off-by: Hollow Man <hollowman@opensuse.org>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai> Co-authored-by: OpenAI Codex <codex@openai.com>
…ject#52842) Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com> Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com> Co-authored-by: Thomas Parnell <tpa@zurich.ibm.com> Co-authored-by: Claude <noreply@anthropic.com>
vllm-project#52702) Signed-off-by: Itay Etelis <itay.etelis@ibm.com> Co-authored-by: Itay Etelis <itay.etelis@ibm.com>
…#51885) Signed-off-by: Itay Alroy <ialroy@nvidia.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
…vllm-project#52616) Signed-off-by: jiang1.li <jiang1.li@intel.com>
Signed-off-by: Oxygen56 <jiangth99@163.com>
Signed-off-by: mayuyuace <qiming1.zhang@intel.com> Signed-off-by: Qiming Zhang <qiming1.zhang@intel.com>
…m-bench (vllm-project#51863) Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mayuyuace <qiming1.zhang@intel.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…roject#52706) Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
…ath + MTP (vllm-project#46514) Signed-off-by: Mikhail Kostryukov <mike@triptrack.net> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
… ≤1 output token When scheduled_new_reqs is non-empty, scheduled_running_reqs is empty, and every new request has max_tokens <= 1, set num_spec_tokens_to_schedule = 0. Speculative decoding is useless for 1-token outputs; draft generation + verification adds measurable latency. Measured on RTX 5090 (31.4 GB), MTP=6: - 1-token-prompt request latency 141ms→127ms (−14ms) - 2000-token prefill benchmark +2.5% (7445→7635 tok/s) - TG for normal requests unchanged (189.8 tok/s) (cherry picked from commit 43e8693) Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
|
Important Review skippedToo many files! This PR contains 2995 files, which is 2895 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. Usage-priced reviews support at most 300 files. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: ⛔ Files ignored due to path filters (5)
📒 Files selected for processing (2995)
You can disable this status message by setting the |
|
Closing in favour of #439 — my error, apologies. I branched this from #439 is the same change branched correctly from It also documents an interaction I hit afterwards: on the V2 model runner in your newer image, a scheduler-selected spec depth of 0 raises in |
Supersedes #437 — same one-line-scope change, but rebased onto current
main, DCO signed-off, andruff check/ruff formatclean. #437 could never go green: its commit was unsigned (signoff-commithook) and the condition tripped ruff's "combineifstatements usingand" plus a format diff. Both are fixed here.What
When a batch contains only new requests (no running ones) and every one of them
has
max_tokens <= 1, setnum_spec_tokens_to_schedule = 0. Speculativedecoding cannot help a 1-token output — the draft pass and verification are pure
overhead.
Why it matters
Beyond the obvious
max_tokens=1API calls, this is the shape of everyprefill-throughput benchmark and every embedding-style/classification request,
so the overhead shows up in a lot of measurements.
Measured on RTX 5090 (31.4 GiB), Qwen3.8-27B EXL3, MTP=6:
The guard is strictly conservative: it only fires when no running request is
scheduled, so an in-flight multi-token generation can never lose its draft
tokens.
Note on CI
pre-run-checkgates on aready/verifiedlabel or an author with 4+merged PRs, so this will sit red until a maintainer applies a label — nothing in
the diff can satisfy it. Everything within my control (DCO, ruff, format,
rebase) is green.