Skip to content

[Rebase] Rebase to vLLM 0.30.0 - #7820

Merged
Gaohan123 merged 19 commits into
mainfrom
dev/vllm-align
Sep 22, 2026
Merged

Gaohan123 merged 19 commits into
mainfrom
dev/vllm-align

Conversation

@tzhouam

@tzhouam tzhouam commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Adapt vLLM-Omni to the released vLLM 0.30.0 (ced6857afa0ea7b2e3f0846a62e1394e90f15607) and merge current main (ba5b7485a7f31aa2ddd90e328e873d189f23ceaf).

CUDA CI and production, ROCm, and XPU use their official v0.30.0 images; Intel CI uses the matching release. Remove the temporary commit-wheel reinstall and dependency-repair block. Ascend pins retain their separately released dependency version.

The rebase adapts scheduler/connector output contracts, ModelOpt quantization, multimodal hashing, worker interfaces, and benchmark keyword forwarding. Worker v2 uses the released RopeState constructor, forwards dummy-run state-slot arguments, supplies the execution state's CUDA graph statistics field, and converts the new compact sampling masks without the removed argument. Regression tests exercise real upstream constructors and both compact and bitmask sampling-mask paths. Grammar rejection in the generation scheduler now runs before terminal cleanup: rejected requests emit ERROR, release resources, and cannot rearm a resumable session. Regression coverage exercises both incomplete and complete prompts.

Installation and quickstart source instructions target v0.30.0, while published Omni 0.28 examples retain their matching vLLM versions. Scheduler error completion and grammar validation share helpers; generation output construction is isolated in its own helper. Restore the tested transformers>=5.13.0,<5.15 range requested in review; it satisfies upstream vLLM 0.30.0 requirements.

Validation

The release adaptation passed 650 distinct tests across compatibility suites and targeted reruns, with 24 skipped. After the review refactor, all 234 scheduler and duplex tests pass again. After the worker API fixes, all 71 worker v2 CPU tests pass (2 GPU tests deselected), including regressions reproduced against the actual vLLM 0.30.0 constructors; Ruff lint/format and whitespace checks pass. CUDA and ROCm wheel URLs resolve; documentation snippet markers are unchanged. Markdown lint has zero introduced diagnostics (46 also present on the previous head). The broad sweep initially found two generation lifecycle fixture failures; the complete scheduler suite passed after adapting the fixture (212 passed). Buildkite configuration/package discovery: 149 passed; Mooncake: 18 passed; latest-main duplex orchestrator: 22 passed. Ruff and diff whitespace checks pass.

Tests use an isolated Python 3.12 environment with released vLLM 0.30.0 and editable Omni, sharing host torch/CUDA libraries. The official Docker image was not built locally; full Buildkite GPU validation on the new head is still required.

Full rebase CI

Fresh full Build #3066 is scheduled on worker-fix head ed108e4da8825cc179340b11be3a65f585b0e218; GPU results are pending.

a4bc3e66840835cd6847165e340916e0

Build #3065 exposed six TTS/Entrypoint L4 startup failures from the removed RopeState has_delta argument, plus two Simple Other jobs failing on the changed sampling-mask constructor. These worker v2 API mismatches are addressed here, including the subsequent dummy-run and execution-state failures that startup had masked.

The Anima documentation placeholder path also fails on main #3063. Cosmos3 V2V fails after main's benchmark migration (#7737) sends MP4 uploads through an API path expecting image frames; this inherited main defect is outside the rebase-specific fixes. Qwen-Image correctness passes (SSIM 0.999998, PSNR infinity), but its 5051ms latency versus a 3747ms reference exceeds the unchanged 20% threshold. Its cause remains unresolved and requires another GPU run. Earlier Qwen3-Omni H100 function/audio and MiniMax H3 accuracy candidates pass on #3065.

tzhouam and others added 5 commits September 15, 2026 11:44
Carries forward the adaptation work produced by rebase-agent run
rebase-20260915-111941, which was left uncommitted when that run failed
at pre-commit and never pushed. Targets upstream vLLM APIs beyond
v0.29.0: uses_mrope on NewRequestData.from_request, KVConnectorBlockState,
and the associated scheduler, runner and patch plumbing.

Scope: 22 files under vllm_omni/, tests/ and benchmarks/. The same run
also left 1,637 SPDX-header rewrites and 103 markdownlint auto-fixes
(MD026/MD029/MD034, table padding, list indentation) in the working
tree; both are pre-commit auto-fix residue with no rebase content, and
neither is carried here.

Salvaged from: /data/zhoutaichang/rebase/vllm-omni (working tree, untouched)
Original pin: a529c1a748d3 (ancestor of releases/v0.30.0)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	vllm_omni/worker/gpu_generation_model_runner.py
#	vllm_omni/worker/gpu_model_runner.py
releases/v0.30.0 is cut (f2aad6aa7) but not tagged, so there is no
vllm/vllm-openai:v0.30.0 image. CI therefore rides the newest published
release image (v0.29.0) and reinstalls vLLM from the pinned commit wheel
-- the arrangement this file's own header prescribed at the v0.29.0
release boundary, and the same cycle 0.27 and 0.28 went through.

The dependency repair is lighter this time because the base image and the
target wheel agree on torch: v0.29.0 and f2aad6aa7 both specify torch
2.13.0 / torchaudio 2.11.0 / torchvision 0.28.0. The torch pin is kept as
the ordering guard its comment describes, not as a version migration --
the ordering is what stopped FlashInfer binding to a torch that the vLLM
reinstall then removed during 0.27 (21 CI jobs lost on build #2947).

The one real runtime delta is FlashInfer 0.6.18 -> 0.6.18.post1, per
upstream requirements/cuda.txt at the pin; flashinfer-jit-cache moves with
it, and 0.6.18.post1 was confirmed present on both the flashinfer.ai and
cu130 indexes. numba is still 0.65.0 at this pin, so the numpy 2.2.6
re-pin is unchanged.

Revert to the release-boundary form (base tag v0.30.0, no wheel pin, no
repair steps) once v0.30.0 is tagged and its official image ships.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adapts vLLM-Omni to upstream vLLM releases/v0.30.0 @ f2aad6aa7074, from the
v0.29.0 baseline 98dff2a81d74 (761 upstream commits).

Produced by the copilot repo-rebase-v3 pipeline running module agents on the
abstracted provider backend (cursor / Grok 4.6 High Fast). All 8 modules
completed and the wave gate passed clean; the local test loop reported 30
suites passed, 1 infrastructure failure (simple_diffusion_test timeout),
3 skipped.

Touches the areas upstream changed between 0.29 and 0.30: the AR and
generation schedulers, the GPU model runners, async entrypoint plumbing,
input preprocessing, diffusion KV backend and the benchmark patch shim,
with their regression tests.

Deliberately EXCLUDED from this commit: the pre-commit hooks' repo-wide
auto-fixes — 1,554 SPDX header rewrites and 133 markdownlint reformats
(MD026/MD029/MD034, list indentation) across docs/examples/recipes/apps.
None of it is rebase content, it is unrelated to this change, and carrying
it caused merge conflicts in an earlier attempt. Held in stashes.

Known not addressed here: markdownlint fails on main as well (570 errors on
the baseline vs 565 here), so it is pre-existing rather than introduced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-19T05:37:58.881984Z ee2f9be PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ee2f9be754

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread requirements/common.txt Outdated
# Floor is 5.13.0 so transformers.integrations.hub_kernels LayerRepository
# entries always pass version=/revision= required by kernels>=0.16 (#6971).
transformers >= 5.13.0, < 5.15
transformers >= 5.10.4

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Restore the supported Transformers version range

The relaxed requirement now accepts both Transformers 5.10.4–5.12.x and 5.15+, even though the adjacent dependency documentation states that versions before 5.13 omit the version/revision arguments required by the pinned kernels==0.16.1 integration and that 5.15 has a known model-construction regression. Environments whose resolver selects one of those newly allowed versions can therefore fail while loading hub-backed kernels or constructing models; retain the >=5.13.0,<5.15 compatibility interval unless the corresponding code and kernel pin are updated.

Useful? React with 👍 / 👎.

Comment on lines +218 to +219
except (OSError, PermissionError, RuntimeError) as exc:
pytest.skip(f"kernels-hub flash-attn2 unavailable in this env: {exc}")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Do not skip arbitrary kernel RuntimeErrors

On machines where the optional package is installed, any RuntimeError raised by kernel construction or execution—including genuine regressions such as invalid launch parameters, shape handling failures, or backend implementation errors—is now classified as an environmental skip. This makes the execution test pass without exercising its output checks; restrict the skip to the specific known incompatibility/cache errors, and allow unrelated runtime failures to fail the test. The same broad catch is repeated for FlashAttention 3 below.

Useful? React with 👍 / 👎.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/index.md, docs/design/module/entrypoints.md.

Module owners: @david6666666 @princepride @Isotr0py

Routing: @david6666666 via module of the changed files, module named in the PR description, CODEOWNERS; @princepride via module of the changed files, module named in the PR description, CODEOWNERS; @Isotr0py via module of the changed files, module named in the PR description

@tzhouam, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 19, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 236ef7a8fcd4 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: three questions on the performance claim

@tzhouam this PR reads as a performance or value claim:

  • claim: The one runtime delta is FlashInfer 0.6.18 → 0.6.18.post1 (upstream requirements/cuda.txt at the pin), with flashinfer-jit-cache moved in lockstep; .post1 was confirmed present on both…

Before the full evidence checklist, three short questions:

  1. Bottleneck — what is the current bottleneck, and which profile, trace or per-stage measurement shows it?
  2. Value — what does the change buy the user or the system (latency, throughput, memory, cost), and at which workload?
  3. A/B or ablation — is there a same-workload, same-head/config comparison that isolates each main claim on its own? For stacked optimizations, one number per item rather than a blended delta.

When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution.

tzhouam and others added 2 commits September 22, 2026 18:10
Adapt RoPE construction, dummy-run state slots, required execution state fields, and compact sampling masks. Cover the upstream constructor contracts and both mask representations with CPU regressions.

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Sep 22, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123
Gaohan123 enabled auto-merge (squash) September 22, 2026 14:19
@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 22, 2026
@Gaohan123
Gaohan123 disabled auto-merge September 22, 2026 15:23
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 22, 2026 15:23
@Gaohan123
Gaohan123 merged commit 59db6a4 into main Sep 22, 2026
12 of 14 checks passed
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
VLa1111 added a commit to VLa1111/vllm-omni that referenced this pull request Oct 5, 2026
… the PR head

The server path in run_e2e.sh cannot drive MammothModa2 T2I on this head: only
model_extras.build_text_to_image_prompt applies the AR t2i scaffold and
omni_task metadata, and no serving endpoint builds it, so every HTTP request
dies in the DiT stage with "AR stage produced no visual-token hidden states"
(verified also with an invalid eol_token_id probe: the request's
additional_information is never consumed over HTTP). Add run_e2e_offline.py,
which drives the same 10-cell matrix through Omni + model_extras with one
process per cell, a device-memory sampler that starts before the model load,
and 4-wide concurrent waves for the batch-4 row, emitting collect.py-compatible
artifacts.

Re-measured on the RTX PRO 6000 Blackwell box (vLLM upgraded 0.29.0 -> 0.30.0
as required by the head's Rebase to vLLM 0.30.0 vllm-project#7820): 10/10 cells, no
failures. Tiling bounds the 1536x1536 device peak 67,766 -> 61,656 MiB
(reference: 67,354 -> 61,814) and is a no-op at 1024; slicing alone is a no-op
at batch 1, and at batch 4 gives ~2.8-2.9x per-image throughput with a flat
peak (67,200 MiB at 1536 vs the 67,766 single-image baseline). Full analysis in
BENCHMARK_REPORT.md; runner provenance in CHANGES.md.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants