Skip to content

[compressed-tensors] update find_matched_target order to prioritize fused name matches over class match - #49483

Merged
mgoin merged 7 commits into
vllm-project:mainfrom
neuralmagic:bdellabe/ct-find-match-order
Jul 28, 2026
Merged

mgoin merged 7 commits into
vllm-project:mainfrom
neuralmagic:bdellabe/ct-find-match-order

Conversation

@brian-dellabetta

@brian-dellabetta brian-dellabetta commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Resolves vllm-project/llm-compressor#2730

Currently on main, the order of target resolution is

  1. match name
  2. match class
  3. match fused name

This PR flips 2 and 3, matching the logic of compressed-tensors. This is helpful when users have mixed-precision quantization configs and aren't including all fused names in their targets, and have a last config group targeting Linear.

Test Plan

This script will fail on main, but works on this branch. See config.json here for the case described above (targets unfused layer names in first group, targets Linear in second group):

import sys
import os

os.environ.setdefault("VLLM_DISABLED_KERNELS", "MarlinMxfp8LinearKernel")

model_name = "INCModel/Qwen3-0.6B-MXFP4-MXFP8"

from vllm import LLM

llm = LLM(
    model=model_name,
    revision="1a0c02ca35d68d86aaa08afc7dc1c177793a3825",
    gpu_memory_utilization=0.9,
    enforce_eager=True,
    # dtype="bfloat16",
)

print("✓ Model loaded successfully!")

# Test generation
prompts = ["Hello, how are you?"]
outputs = llm.generate(prompts)

for output in outputs:
    print(f"Prompt: {output.prompt}")
    print(f"Generated text: {output.outputs[0].text}")

print("✓ Generation test passed!")

Test Result

Above snippet succeeds for branch

✓ Model loaded successfully!
WARNING 07-22 19:52:21 [model.py:1546] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
Rendering prompts: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 32.83it/s]
Processed prompts: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00,  3.83it/s, est. speed input: 23.02 toks/s, output: 61.37 toks/s]
Prompt: Hello, how are you?
Generated text:  I'm a student, and I want to study in a university. I'm
✓ Generation test passed!

but fails with opaque error on main because targets are resolving incorrectly

(EngineCore pid=3083133)   File "/home/brian-dellabetta/projects/vllm/vllm/model_executor/models/utils.py", line 337, in _load_module
(EngineCore pid=3083133)     yield from map(
(EngineCore pid=3083133)   File "/home/brian-dellabetta/projects/vllm/vllm/model_executor/layers/linear.py", line 1010, in load_weights
(EngineCore pid=3083133)     param.weight_loader(param, loaded_weight, shard_id)
(EngineCore pid=3083133)   File "/home/brian-dellabetta/projects/vllm/vllm/model_executor/layers/linear.py", line 753, in weight_loader
(EngineCore pid=3083133)     param_data = param.data
(EngineCore pid=3083133)                  ^^^^^^^^^^
(EngineCore pid=3083133)   File "/home/brian-dellabetta/projects/.venv/lib64/python3.12/site-packages/torch/nn/modules/module.py", line 1968, in __getattr__
(EngineCore pid=3083133)     raise AttributeError(
(EngineCore pid=3083133) AttributeError: 'MergedColumnParallelLinear' object has no attribute 'data'

Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing (anything written below this line will be removed by GitHub Actions)

Signed-off-by: Brian Dellabetta <bdellabe@redhat.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@brian-dellabetta brian-dellabetta changed the title [compressed-tensors] update find_matched_target order to prioritize fused over class match [compressed-tensors] update find_matched_target order to prioritize fused name matches over class match Jul 22, 2026
@mgoin
mgoin enabled auto-merge (squash) July 22, 2026 20:05
@github-actions github-actions Bot added the ready ONLY add when PR is ready to merge/full CI is needed label Jul 22, 2026
@mgoin
mgoin merged commit 8a7b3c2 into vllm-project:main Jul 28, 2026
108 checks passed
@brian-dellabetta
brian-dellabetta deleted the bdellabe/ct-find-match-order branch July 28, 2026 18:06
xin3he added a commit to xin3he/vllm that referenced this pull request Aug 12, 2026
Signed-off-by: Xin He <xin3.he@intel.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

quantization ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Mix-precision vLLM loading failure (MXFP4+MXFP8)

4 participants