Repository navigation
[Rebase] Rebase to vLLM 0.30.0 - #7820
Conversation
Carries forward the adaptation work produced by rebase-agent run rebase-20260915-111941, which was left uncommitted when that run failed at pre-commit and never pushed. Targets upstream vLLM APIs beyond v0.29.0: uses_mrope on NewRequestData.from_request, KVConnectorBlockState, and the associated scheduler, runner and patch plumbing. Scope: 22 files under vllm_omni/, tests/ and benchmarks/. The same run also left 1,637 SPDX-header rewrites and 103 markdownlint auto-fixes (MD026/MD029/MD034, table padding, list indentation) in the working tree; both are pre-commit auto-fix residue with no rebase content, and neither is carried here. Salvaged from: /data/zhoutaichang/rebase/vllm-omni (working tree, untouched) Original pin: a529c1a748d3 (ancestor of releases/v0.30.0) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # vllm_omni/worker/gpu_generation_model_runner.py # vllm_omni/worker/gpu_model_runner.py
releases/v0.30.0 is cut (f2aad6aa7) but not tagged, so there is no vllm/vllm-openai:v0.30.0 image. CI therefore rides the newest published release image (v0.29.0) and reinstalls vLLM from the pinned commit wheel -- the arrangement this file's own header prescribed at the v0.29.0 release boundary, and the same cycle 0.27 and 0.28 went through. The dependency repair is lighter this time because the base image and the target wheel agree on torch: v0.29.0 and f2aad6aa7 both specify torch 2.13.0 / torchaudio 2.11.0 / torchvision 0.28.0. The torch pin is kept as the ordering guard its comment describes, not as a version migration -- the ordering is what stopped FlashInfer binding to a torch that the vLLM reinstall then removed during 0.27 (21 CI jobs lost on build #2947). The one real runtime delta is FlashInfer 0.6.18 -> 0.6.18.post1, per upstream requirements/cuda.txt at the pin; flashinfer-jit-cache moves with it, and 0.6.18.post1 was confirmed present on both the flashinfer.ai and cu130 indexes. numba is still 0.65.0 at this pin, so the numpy 2.2.6 re-pin is unchanged. Revert to the release-boundary form (base tag v0.30.0, no wheel pin, no repair steps) once v0.30.0 is tagged and its official image ships. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adapts vLLM-Omni to upstream vLLM releases/v0.30.0 @ f2aad6aa7074, from the v0.29.0 baseline 98dff2a81d74 (761 upstream commits). Produced by the copilot repo-rebase-v3 pipeline running module agents on the abstracted provider backend (cursor / Grok 4.6 High Fast). All 8 modules completed and the wave gate passed clean; the local test loop reported 30 suites passed, 1 infrastructure failure (simple_diffusion_test timeout), 3 skipped. Touches the areas upstream changed between 0.29 and 0.30: the AR and generation schedulers, the GPU model runners, async entrypoint plumbing, input preprocessing, diffusion KV backend and the benchmark patch shim, with their regression tests. Deliberately EXCLUDED from this commit: the pre-commit hooks' repo-wide auto-fixes — 1,554 SPDX header rewrites and 133 markdownlint reformats (MD026/MD029/MD034, list indentation) across docs/examples/recipes/apps. None of it is rebase content, it is unrelated to this change, and carrying it caused merge conflicts in an earlier attempt. Held in stashes. Known not addressed here: markdownlint fails on main as well (570 errors on the baseline vs 565 here), so it is pre-existing rather than introduced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ee2f9be754
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| # Floor is 5.13.0 so transformers.integrations.hub_kernels LayerRepository | ||
| # entries always pass version=/revision= required by kernels>=0.16 (#6971). | ||
| transformers >= 5.13.0, < 5.15 | ||
| transformers >= 5.10.4 |
There was a problem hiding this comment.
Restore the supported Transformers version range
The relaxed requirement now accepts both Transformers 5.10.4–5.12.x and 5.15+, even though the adjacent dependency documentation states that versions before 5.13 omit the version/revision arguments required by the pinned kernels==0.16.1 integration and that 5.15 has a known model-construction regression. Environments whose resolver selects one of those newly allowed versions can therefore fail while loading hub-backed kernels or constructing models; retain the >=5.13.0,<5.15 compatibility interval unless the corresponding code and kernel pin are updated.
Useful? React with 👍 / 👎.
| except (OSError, PermissionError, RuntimeError) as exc: | ||
| pytest.skip(f"kernels-hub flash-attn2 unavailable in this env: {exc}") |
There was a problem hiding this comment.
Do not skip arbitrary kernel RuntimeErrors
On machines where the optional package is installed, any RuntimeError raised by kernel construction or execution—including genuine regressions such as invalid launch parameters, shape handling failures, or backend implementation errors—is now classified as an environmental skip. This makes the execution test pass without exercising its output checks; restrict the skip to the specific known incompatibility/cache errors, and allow unrelated runtime failures to fail the test. The same broad catch is repeated for FlashAttention 3 below.
Useful? React with 👍 / 👎.
|
This PR appears to belong to: docs/design/module/diffusion/index.md, docs/design/module/entrypoints.md. Module owners: @david6666666 @princepride @Isotr0py Routing: @david6666666 via module of the changed files, module named in the PR description, CODEOWNERS; @princepride via module of the changed files, module named in the PR description, CODEOWNERS; @Isotr0py via module of the changed files, module named in the PR description @tzhouam, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot: three questions on the performance claim@tzhouam this PR reads as a performance or value claim:
Before the full evidence checklist, three short questions:
When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution. |
Adapt RoPE construction, dummy-run state slots, required execution state fields, and compact sampling masks. Cover the upstream constructor contracts and both mask representations with CPU regressions. Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
… the PR head The server path in run_e2e.sh cannot drive MammothModa2 T2I on this head: only model_extras.build_text_to_image_prompt applies the AR t2i scaffold and omni_task metadata, and no serving endpoint builds it, so every HTTP request dies in the DiT stage with "AR stage produced no visual-token hidden states" (verified also with an invalid eol_token_id probe: the request's additional_information is never consumed over HTTP). Add run_e2e_offline.py, which drives the same 10-cell matrix through Omni + model_extras with one process per cell, a device-memory sampler that starts before the model load, and 4-wide concurrent waves for the batch-4 row, emitting collect.py-compatible artifacts. Re-measured on the RTX PRO 6000 Blackwell box (vLLM upgraded 0.29.0 -> 0.30.0 as required by the head's Rebase to vLLM 0.30.0 vllm-project#7820): 10/10 cells, no failures. Tiling bounds the 1536x1536 device peak 67,766 -> 61,656 MiB (reference: 67,354 -> 61,814) and is a no-op at 1024; slicing alone is a no-op at batch 1, and at batch 4 gives ~2.8-2.9x per-image throughput with a flat peak (67,200 MiB at 1536 vs the 67,766 single-image baseline). Full analysis in BENCHMARK_REPORT.md; runner provenance in CHANGES.md. Co-Authored-By: Claude Code <noreply@anthropic.com>
Purpose
Adapt vLLM-Omni to the released vLLM 0.30.0 (
ced6857afa0ea7b2e3f0846a62e1394e90f15607) and merge current main (ba5b7485a7f31aa2ddd90e328e873d189f23ceaf).CUDA CI and production, ROCm, and XPU use their official v0.30.0 images; Intel CI uses the matching release. Remove the temporary commit-wheel reinstall and dependency-repair block. Ascend pins retain their separately released dependency version.
The rebase adapts scheduler/connector output contracts, ModelOpt quantization, multimodal hashing, worker interfaces, and benchmark keyword forwarding. Worker v2 uses the released RopeState constructor, forwards dummy-run state-slot arguments, supplies the execution state's CUDA graph statistics field, and converts the new compact sampling masks without the removed argument. Regression tests exercise real upstream constructors and both compact and bitmask sampling-mask paths. Grammar rejection in the generation scheduler now runs before terminal cleanup: rejected requests emit ERROR, release resources, and cannot rearm a resumable session. Regression coverage exercises both incomplete and complete prompts.
Installation and quickstart source instructions target v0.30.0, while published Omni 0.28 examples retain their matching vLLM versions. Scheduler error completion and grammar validation share helpers; generation output construction is isolated in its own helper. Restore the tested
transformers>=5.13.0,<5.15range requested in review; it satisfies upstream vLLM 0.30.0 requirements.Validation
The release adaptation passed 650 distinct tests across compatibility suites and targeted reruns, with 24 skipped. After the review refactor, all 234 scheduler and duplex tests pass again. After the worker API fixes, all 71 worker v2 CPU tests pass (2 GPU tests deselected), including regressions reproduced against the actual vLLM 0.30.0 constructors; Ruff lint/format and whitespace checks pass. CUDA and ROCm wheel URLs resolve; documentation snippet markers are unchanged. Markdown lint has zero introduced diagnostics (46 also present on the previous head). The broad sweep initially found two generation lifecycle fixture failures; the complete scheduler suite passed after adapting the fixture (212 passed). Buildkite configuration/package discovery: 149 passed; Mooncake: 18 passed; latest-main duplex orchestrator: 22 passed. Ruff and diff whitespace checks pass.
Tests use an isolated Python 3.12 environment with released vLLM 0.30.0 and editable Omni, sharing host torch/CUDA libraries. The official Docker image was not built locally; full Buildkite GPU validation on the new head is still required.
Full rebase CI
Fresh full Build #3066 is scheduled on worker-fix head
ed108e4da8825cc179340b11be3a65f585b0e218; GPU results are pending.Build #3065 exposed six TTS/Entrypoint L4 startup failures from the removed RopeState
has_deltaargument, plus two Simple Other jobs failing on the changed sampling-mask constructor. These worker v2 API mismatches are addressed here, including the subsequent dummy-run and execution-state failures that startup had masked.The Anima documentation placeholder path also fails on main #3063. Cosmos3 V2V fails after main's benchmark migration (#7737) sends MP4 uploads through an API path expecting image frames; this inherited main defect is outside the rebase-specific fixes. Qwen-Image correctness passes (SSIM 0.999998, PSNR infinity), but its 5051ms latency versus a 3747ms reference exceeds the unchanged 20% threshold. Its cause remains unresolved and requires another GPU run. Earlier Qwen3-Omni H100 function/audio and MiniMax H3 accuracy candidates pass on #3065.