Upstream sync 10/N: merge 457a5f34db (31 commits, conflict-free) - #1257
Conversation
…ct#51556) Signed-off-by: cherry77-cloud <1615405@qq.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com> Co-authored-by: OpenAI Codex <noreply@openai.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…1733) Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
…vllm-project#51666) Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com> Signed-off-by: Andreas Karatzas <akaratza@amd.com> Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Yan Ma <yan.ma@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…ject#44201) Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
…#51473) Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>
… graph (vllm-project#46849) Signed-off-by: Yizhou Liu <liu_yizhou@outlook.com>
Signed-off-by: lkm2835 <lkm2835@gmail.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
…llm-project#50826) Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…project#50977) Signed-off-by: David Holtz <56723830+dmholtz@users.noreply.github.com>
…roject#51766) Signed-off-by: Dao Le <Dao007forever@gmail.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…lm-project#51622) Signed-off-by: Alex <jihui.huang@daocloud.io>
Signed-off-by: shujialiu <liushujia0122@163.com>
…te filtering for JIT warmup, and migrate Inkling FA4 (vllm-project#49315) Signed-off-by: LopezCastroRoberto <rocastro@redhat.com> Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com> Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…raising (vllm-project#51627) Signed-off-by: Kush Zingade <kush.zingade@gmail.com>
…riton launcher) (vllm-project#51770) Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
…benchmark (vllm-project#51308) Signed-off-by: louie-tsai <louie.tsai@intel.com>
Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com> Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…ject#51145) Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
…1768) Signed-off-by: Woosuk Kwon <woosuk@inferact.ai> Co-authored-by: OpenAI Codex <codex@openai.com>
…51819) Signed-off-by: Andrii Skliar <askliar@nvidia.com> Co-authored-by: Andrii Skliar <askliar@nvidia.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…llm-project#51726) Signed-off-by: yewentao256 <zhyanwentao@126.com>
…49444) Signed-off-by: pmanczak <pawel.manczak@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…1812) Signed-off-by: zjy0516 <riverclouds.zhu@qq.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…t#51806) Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…ant KV cache (vllm-project#47896) Signed-off-by: root <root@smci355-ccs-aus-m02-09.cs-aus.dcgpu> Co-authored-by: root <root@smci355-ccs-aus-m02-09.cs-aus.dcgpu> Co-authored-by: Andreas Karatzas <akaratza@amd.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…project#50020) Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
BenchmarksSix-model sweep run on the Strix Halo board via Custom Regression run 3512149. No regression — decode and TTFT both flat across the set.
Largest decode move is −0.6%, inside the 1% gate. All six report Decode is derived from Baselines are the nightly dashboard run of 2026-09-05 ( Scope. The run built Environment: host AI assistance was used to run this verification. |
Summary
Conflict-free upstream batch: the 31 commits between
dc5101fb1band457a5f34db(2026-08-11 04:43 UTC .. 2026-08-11 16:00 UTC).git mergereported no conflicts; nothing in this merge is a manual resolution.111 files, +11076/−1080.
Findings
Boundary probed in a ladder: k=25 clean; k=100 through the tip (k=720) all conflict. Stepping forward from 25, k=26..31 are clean and k=32 conflicts.
Deleted files. One,
vllm/models/inkling/nvidia/ops/fa4_warmup.py, removed by upstream's6c95a641e9("[2/N][Feat][Perf] Add new warmup infrastructure for JITs…", vllm-project#49315) which migrates Inkling FA4 to the new warmup infrastructure. Verified this is the NVIDIA copy only and that nothing imports it: nofa4_warmupreferences remain anywhere undervllm/models/inkling/nvidia/, and the directory now containsfa4_rel_attention.py,lamport.py,mm_towers.py,norm.py,qkvr_prep.py,sconv.pyand__init__.py. The AMD copy the fork depends on,vllm/models/inkling/amd/ops/fa4_warmup.py, is untouched and its importervllm/models/inkling/amd/attention.py:42still resolves.The next commit,
c65aa2ee6c(Bump Transformers version to 5.15.0, vllm-project#51668), conflicts in two requirements files and gets its own PR.Merge commit only — do not squash or rebase.
AI assistance was used to prepare this merge.
Test plan
pre-commit run --all-filesgreen on the merged treefa4_warmupconfirmed unreferenced; AMD copy verified intact