Repository navigation
test: keep Hunyuan LoRA CPU path on native linear - #8166
Conversation
|
This PR matches CODEOWNERS paths: /tests/. Code owners: @NickCao @yenuo26 Routing: @NickCao via CODEOWNERS; @yenuo26 via CODEOWNERS @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot: supersededThe CI failure noted on |
Omni ReviewBot routing recordAssigned Direct under experiment |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
PR description
This is a test-only change to the existing Hunyuan Image 3 PEFT QKV forward check. The test still builds a real QKVParallelLinear, loads a disk PEFT adapter through DiffusionLoRAManager, and compares wrapped.apply(x) to an explicit row-index oracle for TP sizes 1, 2, and 4. The new piece stubs dispatch_unquantized_gemm with torch.nn.functional.linear so CPU tensors are not sent into the process-wide ROCm unquantized GEMM custom operator. User-visible serving behavior is unchanged; the intended effect is that this CPU core_model test can keep its LoRA/QKV numerical contract on AMD images.
Change flow
flowchart TD
A["[EXISTING] CPU PEFT QKV test TP 1/2/4"]:::existing
B["[CHANGED] test_hunyuan_image3_peft_qkv_forward"]:::changed
C["[NEW] cpu_unquantized_gemm via F.linear"]:::new
D["[EXISTING] QKVParallelLinear and LoRA manager"]:::existing
E["[EXISTING] Oracle x@A.T@B[rows].T * 0.5"]:::existing
A --> B
B --> C
C --> D
D --> E
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
CI at
1214805ebc44(2026-09-26T20:31:16.574327+00:00): required check(s) blocking:buildkite/vllm-omni(failed), and -5 more. Observed Buildkite:buildkite/vllm-omni-amd-ci(failed),buildkite/vllm-omni-npu-ci(failed),buildkite/vllm-omni(failed), and 1 more.
See inline comments below.
| # This numerical contract intentionally runs on CPU. Accelerator builds | ||
| # select the unquantized GEMM from the process-wide platform, which would | ||
| # otherwise dispatch CPU tensors to the ROCm-only custom operator. | ||
| monkeypatch.setattr(linear, "dispatch_unquantized_gemm", lambda: cpu_unquantized_gemm) |
There was a problem hiding this comment.
[P1] Changed PEFT QKV CPU test has no runtime result
The only new runtime behavior is the dispatch_unquantized_gemm stub at this line, added so wrapped.apply(x) stays on native CPU linear math. The PR Validation section ran py_compile, Ruff format, and git diff --check, and deferred exact-head AMD runtime until after open. Those checks cannot tell whether the monkeypatch name is the one QKVParallelLinear.apply actually looks up, or whether an import-time GEMM cache still selects vllm::rocm_unquantized_gemm. Frozen CI at 1214805ebc4466b8cf04b0713240d2b3464cd9b9 (2026-09-26T20:31:16Z) has required buildkite/vllm-omni failed and operator-watched AMD/NPU/Intel lanes failed; .buildkite step commands were not readable in this sparse checkout, so I am not naming a guard selector. Run tests/diffusion/lora/test_lora_manager.py::test_hunyuan_image3_peft_qkv_forward (parametrized TP 1/2/4) and record that result against the CPU core_model lane before treating the ROCm CPU crash as fixed.
Omni ReviewBot: finding feedback[p1] Changed PEFT QKV CPU test has no runtime result — The only new runtime behavior is the This finding appears in the bot's COMMENT review, but GitHub could not place it as an inline diff comment. If you are the PR author and disagree, react 👎 to this comment. The disagreement will be shown to the maintainer; it does not approve or merge the PR. |
Signed-off-by: andyluo7 <andy.luo@amd.com>
1214805 to
8f2553e
Compare
|
Self-review at |
Omni ReviewBot: supersededThe CI failure noted on |
Summary
Root cause
The test is intentionally CPU-only, but vLLM selects unquantized GEMM from the process-wide platform. In the ROCm image that selects vllm::rocm_unquantized_gemm even though the test tensors are on CPU, producing NotImplementedError for the CPU backend for TP sizes 1, 2, and 4.
Example failure: https://buildkite.com/vllm/vllm-omni-amd-ci/builds/12820/table?jid=01a0dab5-3687-400c-b435-24f5786b8c98&tab=output
Validation
Signed-off-by: andyluo7 andy.luo@amd.com