Skip to content

[Core] Default uses_xdrope_dim so the Omni runners work with newer vLLM - #7635

Closed
MrlixiangWE wants to merge 1 commit into
vllm-project:mainfrom
MrlixiangWE:fix-7452-runner-rope-compat
Closed

MrlixiangWE wants to merge 1 commit into
vllm-project:mainfrom
MrlixiangWE:fix-7452-runner-rope-compat

Conversation

@MrlixiangWE

@MrlixiangWE MrlixiangWE commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

vLLM 719284f (vllm-project/vllm#56078) unified XD-RoPE into M-RoPE and removed uses_xdrope_dim from the model runner, so on a vLLM containing that commit the Omni runners raise AttributeError where they read it.

#6923 has since guarded four of the six read sites with getattr. Put the default in OmniGPUModelRunner.__init__ and read the flag plainly everywhere, so one line carries the invariant for all six sites, including the two NPU runners that inherit this constructor and that #6923 did not cover. The pinned version's value is preserved. Part of #7452.

Propose

  1. OmniGPUModelRunner.__init__ reads the flag with a default of 0.
  2. The six read sites read self.uses_xdrope_dim directly — worker/gpu_model_runner.py 699, 1173, 1710, worker/gpu_generation_model_runner.py 815, platforms/npu/worker/npu_model_runner.py 336, platforms/npu/worker/npu_generation_model_runner.py 856.
  3. A CPU test builds both GPU runners through the real parent constructor and the installed vLLM's own config classes, admits a request through _update_states, and checks the position buffers and _preprocess output for plain, image and cached input, following whichever layout the installed version selects. One case drops the attribute after the parent constructor returns, so the pinned version covers the removal too.

Test Plan

vLLM Version: 0.29.0 (98dff2a) and 0.28.1rc1.dev633+g442d36031, which contains vllm-project/vllm#56078. PyTorch 2.13.0+cu130.

vLLM-Omni Commit: bbeec58, base 2ab5d17.

python -m pytest -q tests/worker/test_runner_rope_compat.py
python -m pytest -q tests/worker
pytest -sv tests/model_executor -m 'core_model and cpu'
pytest -sv tests/diffusion -m 'core_model and cpu'
pytest -sv tests/entrypoints tests/engine -m 'core_model and cpu'
pytest -sv tests/ -m 'core_model and cpu' --ignore=tests/diffusion --ignore=tests/model_executor --ignore=tests/entrypoints --ignore=tests/engine

Test Result

The same test file on both versions, this branch against main:

vLLM This branch main
0.29.0, the current pin 15 passed 1 failed, 14 passed
g442d36031 15 passed 15 failed

Every failure on main is AttributeError: '<Runner>' object has no attribute 'uses_xdrope_dim', raised at test_runner_rope_compat.py 170 and 178 because #6923's guards keep _update_states and _preprocess working. Dropping only the default from this branch, plain reads kept, raises in the request path instead:

vllm_omni/worker/gpu_ar_model_runner.py:462: in _update_states
E   AttributeError: 'GPUARModelRunner' object has no attribute 'uses_xdrope_dim'
vllm_omni/worker/gpu_model_runner.py:699: AttributeError
Lane This branch main
tests/worker 351 passed 336 passed
Simple · Model Executor 2413 passed, 3 skipped 2413 passed, 3 skipped
Simple · Diffusion 5913 passed, 33 skipped 5913 passed, 33 skipped
Simple · Engine&Entrypoints 3004 passed, 5 skipped, 1 xfailed 3004 passed, 5 skipped, 1 xfailed
Simple · Other 3609 passed, 26 skipped 3594 passed, 26 skipped

No lane fails on either side; the added tests account for the difference in the Other lane.

The NPU read sites are covered by inheritance — OmniNPUModelRunner defines no __init__, both NPU subclasses call super().__init__ first — and are not executed here. mrope_num_dims, the replacement the newer vLLM introduced, is not read in vllm_omni. Changed-file hooks pass.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/ar_runtime.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, semantic router, CODEOWNERS; @fake0fan via module of the changed files, semantic router; @Gaohan123 via module of the changed files, semantic router

@MrlixiangWE, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@hsliuustc0106 hsliuustc0106 added bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner labels Sep 16, 2026
@MrlixiangWE
MrlixiangWE force-pushed the fix-7452-runner-rope-compat branch 3 times, most recently from b834253 to bbeec58 Compare September 17, 2026 17:36
@MrlixiangWE

Copy link
Copy Markdown
Contributor Author

Self-reviewed against current main; all fixed in bbeec58.

#6923 now guards four of the six read sites with getattr, so this PR as written added a second mechanism for one invariant, with the __init__ line dead on the GPU path and nothing saying so. The default stays and those four sites read plainly again — one enforcement point, and it is what makes the change provable: dropping the default raises inside _update_states (gpu_ar_model_runner.py:462 into gpu_model_runner.py:699), where a served request hits it, instead of on a test assertion.

The two NPU read sites #6923 did not touch are what this still buys over main. They reach the default by inheritance — OmniNPUModelRunner defines no __init__, and both NPU subclasses call super().__init__ first — and their runtime path is not executed here.

One case now drops the attribute after the parent constructor returns, so the pinned version covers the removal as well: 15 passed on 0.29.0 and on g442d36031, against 1 and 15 failed on main with the same test file. mrope_num_dims, the replacement the newer vLLM introduced, is not read anywhere in vllm_omni, so this stops the crash and does not add N-dimensional M-RoPE.

@MrlixiangWE MrlixiangWE left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review pass over this head. The first item asks for a smaller diff; the rest are the test file.

Comment thread vllm_omni/worker/gpu_model_runner.py Outdated
if self.uses_mrope:
positions = self.mrope_positions.gpu[:, :num_input_tokens]
elif getattr(self, "uses_xdrope_dim", 0) > 0:
elif self.uses_xdrope_dim > 0:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep #6923's four getattr guards and ship only the constructor default. This repo guards attributes precisely because tests build runners with object.__new__ — gpu_ar_model_runner.py:283 says so in a comment, and tests/worker/test_omni_gpu_model_runner.py builds OmniGPUModelRunner that way and drives _preprocess, staying green only because the fixture hand-sets uses_xdrope_dim = 0. Reverting the guards makes that hand-set load-bearing for every such fixture, and the newer-vLLM fix holds without those four hunks, so they are scope this PR does not need.

# XD-RoPE path stays for versions that still provide it. Defaulting it
# here is the single place the six read sites depend on, including the
# two NPU runners that inherit this constructor.
self.uses_xdrope_dim = getattr(self, "uses_xdrope_dim", 0)

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On a vLLM that no longer sets the attribute this default is always 0, so _init_xdrope_positions never runs and _preprocess falls through to the 1-D self.positions. That is correct there — uses_xdrope_dim is gone from ModelConfig and mrope_num_dims replaces it, and an xdrope_section config reports uses_mrope True on that version, which the target-version run exercises — but nothing says so at runtime. Log once when the parent did not provide the flag, so a version that folds XD-RoPE differently is visible instead of silently linear.

Comment thread tests/worker/test_runner_rope_compat.py Outdated
output = scheduled(new=[new])
# Request admission reads the flag at gpu_model_runner.py:696. Startup
# profiling reads it earlier still, in `_dummy_run`, which this CPU fixture
# does not reach.

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment points at :696, which is the _init_mrope_positions call; the flag is read at :699. The read site that executes first in a real start is _dummy_run (gpu_model_runner.py:1173, gpu_generation_model_runner.py:815), during startup profiling, and this suite reaches neither — nor the two NPU sites. Fix the line number and say which of the six sites the suite covers.

expected = torch.arange(4)
buffer = None
else:
if runner.uses_mrope:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Assert the layout in this branch too. The else asserts layout == "xdrope", but this branch asserts nothing, so a version that reports uses_mrope True for the xdrope_section config would run the three xdrope cells through the M-RoPE assertions and pass while uses_xdrope_dim > 0 is never taken. One assert layout == "mrope" closes it.

Comment thread tests/worker/test_runner_rope_compat.py Outdated
assert runner.model.calls == [(route, [1, 2, 3, 4], mm_features, grid)]
else:
assert runner.model.calls == [(route, [1, 2, 3, 4], mm_features)]
assert buffer.cpu.shape == (runner.model.positions.shape[0], runner.max_num_tokens + 1)

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This compares the fixture with itself: runner.model.positions.shape[0] is the row count PositionModel(...) was constructed with. The number that matters is the runner's buffer width, a literal 3 in the pinned version's __init__ and exactly what vllm-project/vllm#56078 changes. Compare against len(mrope_section) for M-RoPE and runner.uses_xdrope_dim for XD-RoPE, so a width change fails here with a readable reason.

Comment thread tests/worker/test_runner_rope_compat.py Outdated

import vllm_omni.worker.gpu_model_runner as omni

for name in ("empty", "zeros", "ones", "full", "tensor"):

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Patch PIN_MEMORY instead of replacing five global torch factories for the whole file. It is a module-level import in vllm.v1.worker.gpu_model_runner and drives all the pinned allocations in that constructor, so patching it covers the runner without routing the test's own torch.tensor and assert_close through a wrapper. The current set is also partial — arange, as_tensor and from_numpy are unwrapped — so an upstream __init__ that pins through any of them turns the shared CPU lane red for an unrelated reason.

grid = [[1, 2, 2]] if input_kind == "image" else []
assert runner.model.calls == [(route, [1, 2, 3, 4], mm_features, grid)]
else:
assert runner.model.calls == [(route, [1, 2, 3, 4], mm_features)]

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

input_kind is a dead axis on the xdrope layout: mm_features here is the same list object passed to NewRequestData, so the comparison short-circuits on identity and the text, image and cached cells assert the same thing. The grid extraction the axis exists to vary lives in _init_mrope_positions only. Parametrize input_kind for the mrope layout, or assert something the XD-RoPE path derived from the request.

vLLM 719284f (vllm-project/vllm#56078) unified XD-RoPE into M-RoPE and removed
`uses_xdrope_dim` from the model runner, so on a vLLM containing that commit the
Omni runners raise AttributeError where they read it.

Read the flag with a default in `OmniGPUModelRunner.__init__`. vllm-project#6923 guards four
of the six read sites for runners built with `object.__new__`, which this repo
does in tests; those guards stay, and this default is what covers the paths they
do not, including the two NPU runners that inherit this constructor. The pinned
version's value is preserved, and a vLLM that never sets the flag says so once
in the log rather than silently serving 1-D positions.

The test constructs both GPU runners through the real parent constructor and the
installed vLLM's config classes, admits a request through `_update_states`, and
checks the position buffers and `_preprocess` output. One case drops the
attribute after the parent constructor returns, so the pinned version also
covers what the removal does.

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: zijianc2 <157244773+zijianc2@users.noreply.github.com>
@MrlixiangWE
MrlixiangWE force-pushed the fix-7452-runner-rope-compat branch from bbeec58 to 17ea395 Compare September 18, 2026 04:14
@MrlixiangWE

Copy link
Copy Markdown
Contributor Author

All fixed in 17ea395.

#6923's four guards stay: this repo builds runners with object.__new__ in tests and guards attributes for exactly that reason, with a comment saying so at gpu_ar_model_runner.py:283, so reverting them was scope this change does not need. The constructor default remains and now logs once when the parent never set the flag, so a version that folds XD-RoPE differently is visible instead of silently linear. The comment names the read site correctly and says which of the six the suite covers. In the tests: the M-RoPE branch asserts its layout, the buffer width comes from the runner rather than from the fixture, pinning is turned off at PIN_MEMORY instead of through five global torch factories, and input_kind is parametrized only where it varies.

11 passed in tests/worker/test_runner_rope_compat.py, tests/worker 347 passed, ruff clean.

@MrlixiangWE

Copy link
Copy Markdown
Contributor Author

@amy-why-3459 @linyueqian @yenuo26 Could one of you take a look at this when you have time, and add ready if it looks right? It gives uses_xdrope_dim a default so the Omni runners start on newer vLLM (#7452).

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

[P3] The Propose section is stale relative to the diff. It says "the six read sites read self.uses_xdrope_dim directly", but the head keeps getattr(self, "uses_xdrope_dim", 0) at four of the six sites (vllm_omni/worker/gpu_model_runner.py:706, :1180, :1717, vllm_omni/worker/gpu_generation_model_runner.py:815); only the two NPU sites read the attribute plainly (platforms/npu/worker/npu_model_runner.py:336, npu_generation_model_runner.py:856), which matches the in-diff comment ("The read sites keep their own getattr for runners built with object.new in tests"). The listed line numbers (699/1173/1710) are also base-relative. Functionally equivalent either way since every read defaults to 0, but the body's stated mechanism ("one line carries the invariant for all six sites") no longer describes what the diff does — worth a body refresh so reviewers don't hunt for a plain-read rewrite that isn't there.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 17ea395 — approve.

The fix is sound: the default in OmniGPUModelRunner.__init__ preserves the parent's value when vLLM still sets the flag and supplies 0 where #56078 removed it, and every runner subclass (including the two NPU runners, which define no __init__ of their own) inherits that constructor, so all six read sites are covered. Tree-wide grep at this head confirms the read-site census; the gpu_ar_model_runner.py:462 frame in the body's trace is the super()._update_states() call site, not an additional read. The getattr guards kept at the four GPU sites are justified by the object.__new__-built runners in test_omni_gpu_model_runner.py, and info_once is available on the pinned vLLM's logger.

The new CPU compat test directly simulates the upstream removal and the author's evidence shows it passing on both 0.29.0 and the pre-removal-adjacent dev build, with full lanes green on both sides. One non-blocking note posted separately: the Propose section of the body still describes an all-plain-read revision that the diff superseded.

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 25, 2026
@MrlixiangWE

Copy link
Copy Markdown
Contributor Author

Closing this. The vLLM 0.30 rebase removed every uses_xdrope_dim read in vllm_omni (#7820 the four GPU runner sites, #8027 the two NPU ones), so the AttributeError this PR guards against can no longer happen on main. Merged onto current main 8a49c65 with vLLM 0.30.0, its own test_runner_rope_compat.py now fails the two xdrope-text cases (assert 'xdrope' == 'mrope', 2 failed / 9 passed).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner cuda-test Used to trigger vllm-omni cuda CI separately. ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants