Skip to content

[Rebase] Rebase to vLLM 0.29.0 - #7230

Merged
Gaohan123 merged 31 commits into
mainfrom
dev/vllm-align
Sep 10, 2026
Merged

Gaohan123 merged 31 commits into
mainfrom
dev/vllm-align

Conversation

@tzhouam

@tzhouam tzhouam commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Adapt vLLM-Omni to the vLLM 0.29 release API, pinned to 98dff2a81d747d1dba01a47f939f48c3526d4206 in the CUDA CI image. Preserve Omni's multimodal, staged-generation, and diffusion behavior while importing or inheriting upstream implementations wherever their contracts match.

Tests

No rebase introduced errors:
image

Change inventory and reasons

This inventory is measured at PR head 1a7c6e51f0239b7bd759909a53294e1b6b135079, against merge base 087ac46318c42d1fa1e4de13a8053f94683898bd, using git diff --numstat 087ac463...1a7c6e51f. It matches GitHub: 116 files, +3,091 / −1,665 = 4,756 changed lines; net +1,426 lines.

The table covers 100% of changed lines, exceeding 95%. Each file is assigned once to its primary aspect; counts include associated tests, comments, and blank lines, not just executable code. Multi-purpose files are not split into subjective hunk-level counts.

Aspect Files Added Deleted Changed lines Reason for the changes / reuse approach
Other model processors and weight loading 18 530 221 751 Migrate Bagel, CosyVoice3, GLM-TTS, OmniVoice, StepAudio2, Voxtral and smaller model integrations to active token/media processing APIs. Preserve model-specific image pathways, voice-cloning prompts and stage outputs. Import upstream AutoWeightsLoader/WeightsMapper, remove obsolete and pass-through hooks, and tokenize OmniVoice prompts only once. Includes shared API/loader and OmniVoice regression tests.
Hunyuan backbone and integration 3 643 28 671 vLLM removed the native Hunyuan classes that Omni's image model subclasses. Retain a 625-line compatibility backbone for the existing integration, using upstream execution layers; omit unused parent weight-loading methods and remove inactive processor hooks. A full Transformers-backend replacement is not implemented or validated in this PR.
Renderer and shared processor plumbing 6 374 267 641 The engine now processes inputs through renderers, so the old preprocessor customization was inactive. Construct a subclass of the selected upstream renderer and extend only singleton routing for no-media processor kwargs and Omni metadata. Preserve concrete-renderer overrides and tokenizer-less stages. Share common custom-processor plumbing and test sync/async routing and construction.
GLM image processor migration 2 103 459 562 Replace obsolete text-processing hooks with the active token/media path. Preserve target-shape generation suffixes, source-image placeholders, processed grids and image-to-image behavior. Most deletions remove inactive or duplicated legacy processing; regression tests cover target prompt formatting.
KV layout and attention migration 17 461 43 504 vLLM replaced per-spec block-stride flags and shared_by tensors with explicit physical layouts and strides. Resolve/propagate diffusion layouts outside vLLM's normal engine-core path, validate backend compatibility and construct correctly addressed tensors. Import KVCacheLayout, KVCacheTensor and compute_layout_strides; retain Omni-specific integration.
Qwen2.5/Qwen3 Omni migration 8 126 243 369 Move existing video pre-sampling and per-video audio-mask adjustments to the active hook. Upstream Qwen bypasses generic pre/post hooks, so a narrow wrapper is still needed. Qwen3 inherits this wrapper and upstream Qwen3 audio processing; remove duplicate prompt-update code. Adapt weight mapping and related model APIs.
MiniCPM migration 2 102 126 228 Adapt media processing, token placeholders and output fields to the new API while retaining long-audio chunking and Omni behavior. Delegate prompt batching to upstream MiniCPMVMultiModalProcessor, filtering the Omni-only use_tts kwarg.
Ming migration 2 77 85 162 Move model-specific prompt rewriting out of the removed processing hook; use token-ID placeholder targets and the new media-processing signature. Use upstream weight loading with stage exclusions.
Runtime, scheduler, platform and worker compatibility 12 150 10 160 Preserve the v1 model-runner contract, initialize the new shutdown timeout, adapt sampler state/structured-output handling and embedding-related calls, keep frontend-only logging flags out of stage validation, and expose platform-detection errors. Includes scheduler/config/platform regressions.
Rotary embedding compatibility 4 109 17 126 Adapt rotary broadcasting and preserve Qwen-Image complex-FP32 frequency precision until rotation, avoiding an observed accuracy regression. Includes dtype, layout and reference-comparison tests.
Serving and API compatibility 19 66 59 125 Follow upstream serving/protocol/launcher module moves and callback/video-decoding interfaces; remove superseded imports and update serving fixtures. These are mostly small compatibility edits across callers.
Docker runtime installation 1 104 20 124 Install the exact pinned vLLM wheel over the configured v0.28 base image and align its runtime dependencies/native-library loading. This is the current pinned-image strategy, not a claim that a matching official release image remains unavailable.
MiniMax H3 migration 1 40 44 84 Preserve H3's segmented presentation using processed image/video grids under the new API. Keep model-specific placeholder discovery and media reprocessing where cache misses alone cannot reconstruct the complete presentation.
Buildkite HF-token handling 3 60 12 72 Preserve Kubernetes-secret credentials when scheduled-build environment values would override HF_TOKEN; restore the conventional variable from a dedicated alias. This is an independent CI reliability fix carried by the branch, with tests.
Other diffusion test compatibility 11 42 22 64 Update fixtures/mocks for changed upstream signatures and dependency interfaces, including SANA GEMM dispatch and CPU-safe distributed/attention setup.
Image accuracy test compatibility 3 51 1 52 Match the Qwen-Image FlashAttention availability probe to the Diffusers API and adjust dependency-sensitive accuracy-test setup without lowering thresholds.
Audio accuracy test corrections 3 45 3 48 Measure dots.tts PCM quality at its native 48 kHz rate, with assertion regressions. This corrects the measurement rather than lowering the HNR threshold.
Quantization test compatibility 1 8 5 13 Adapt the ModelOpt NVFP4 NaN-clamp test to the updated upstream interface.
Total / coverage 116 3,091 1,665 4,756 (100%) All changed files accounted for once.

Why the diff is still large

  • Tests contribute +1,101 / −163 = 1,264 changed lines (26.6%), already included in the categories above, not additional to them.
  • Deletions count toward review size. GLM and preprocessing remove code that already exists on the base branch; deleting it reduces maintenance but still enlarges the visible deletion count.
  • The largest retained copy is Hunyuan's 625-line backbone. The whole diff is not copied upstream code: it also includes API adaptations, tests, build configuration and removal of old implementations.
  • Independent CI and accuracy-test fixes are explicitly identified above rather than presented as mandatory model-API changes.

Why model-specific processor changes remain

We do not need a duplicate processor implementation for every model. Generic parsing, caching, loading and compatible model processing should come from upstream. Local overrides retain only behavior absent from upstream: stage-specific inputs, generation prompt formatting, voice-cloning conventions, image pathways or per-video audio settings.

For example, Qwen's old _call_hf_processor customization is moved to _apply_hf_processor_main because the old hook is no longer invoked. The override prepares Omni kwargs, delegates actual processing to super(), and restores the per-video mask. Qwen3 inherits that implementation rather than copying it. Upstream Qwen's custom implementation does not call the generic _preprocess_hf_mm_data/_postprocess_hf_mm_data hooks, preventing a smaller hook-only override at the pinned version.

The branch also removes obsolete _hf_processor_applies_updates/_call_hf_processor overrides, Voxtral's pass-through override, and Qwen3's duplicate prompt-update override. The renderer now uses actual upstream-class inheritance rather than a proxy with a manually maintained method allow-list.

Validation and limitations

Evidence is commit-scoped; earlier results are not certification of the current head.

  • At cleanup commit ebd306a38, focused CPU suites passed 453 tests, 7 skipped, including renderer/input routing, model processing, stage initialization, weight mapping, GLM formatting and Hunyuan diffusion tests. Processor imports, Ruff and whitespace checks passed. GPU/end-to-end CI was not rerun for that cleanup.
  • Subsequent commit records report focused validation: Voxtral tests (62); Qwen3/Qwen2.5 tests (20 + 81); OmniVoice tests (106) plus shared-processor/TTS tests (162); renderer-head input tests (32), stage-init tests (55) and entrypoint tests (548). These are separate reported runs and must not be summed as unique full-suite coverage. The latest renderer commit also records pre-commit, including mypy, and import checks.
  • Historical Buildkite 3028, on 325217e0a and the earlier vLLM pin, passed the investigated original rebase-only cases: VoiceDesign checkpoint, SANA cache-backend setup, MiniMax I2VA accuracy, Qwen3 speaker output, Qwen-Image accuracy and dots.tts PCM measurement. This evidence does not validate the current head or pin. Fixes already present in base 087ac463 are not counted as new PR changes in the inventory.
  • At this description refresh, GitHub reports Buildkite release 2834 and documentation as pending, CodeQL queued, and DCO action required for missing sign-off on 736c2f53a. GitHub also reports merge conflicts. These statuses can change; consult live checks before merging.
  • Accuracy thresholds are unchanged. Full current-head GPU/end-to-end validation remains outstanding; this description does not claim all checks are green or that the PR is merge-ready.

Remaining reuse opportunities (not implemented here)

  • Replace the local Hunyuan MLP with upstream LlamaMLP after parameter/loading/parallelism equivalence tests.
  • Delegate diffusion layout selection to upstream resolve_kv_cache_layout while preserving backend validation and explicitly testing environment/connector policy differences.
  • Investigate constructing Omni's FP32-router Hunyuan MoE block directly instead of creating and replacing the compatibility block.
  • A full Hunyuan Transformers-backend migration requires independent model-parity validation; there is no native Hunyuan class available to simply import from the pinned vLLM version.

@tzhouam tzhouam added high priority high priority issue, needs to be done asap diffusion codes related to diffusion models labels Sep 8, 2026
@tzhouam
tzhouam force-pushed the dev/vllm-align branch 4 times, most recently from e51ff28 to 8fbefe8 Compare September 8, 2026 02:33
@tzhouam
tzhouam marked this pull request as ready for review September 8, 2026 03:00
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md.

Module owners: @tzhouam @gcanlin

@tzhouam, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 0ae9bdd8c2ee produced:

  • Priority: high. Prompt maintainer attention is suggested.

  • Quality: low. Mechanical checks suggest considering rework or declining this PR before review effort is spent.

Quality evidence:

  • Q3: touches 14 modules with no linked issue
  • Q6: 120 changed files

These are automated triage suggestions only — the final decision belongs to the maintainers.

…formatter

Nightly 3034 (Simple · Model Executor Test) failed all six
test_generation_prompt_ids_preserve_hf_target_scaffold cases with

    AttributeError: type object 'GlmImageProcessor' has no attribute '_build_prompt_with_target_shape'

The 0.29 rewrite of GlmImageMultiModalProcessor.apply reached for
GlmImageProcessor._build_prompt_with_target_shape, a private transformers
method. The test suite stubs transformers.models.glm_image.processing_glm_image
with a bare class to import glm_image_ar without the real package, so in the
full CPU run the private attribute is not there; the same dependency would
break at runtime on any transformers build that renames or removes it.

Format the scaffold in Omni instead: _build_target_shape_scaffold mirrors the
HF formula (identical in transformers 5.13 and 5.14: target grid, plus the
preview grid for text-to-image, then the image bos token) using only the
processor's public grid_bos_token / grid_eos_token / bos_token attributes.
The test keeps its hard-coded expected suffix, which is what HF produces for
512x768, and no longer imports the HF class.

Validation: tests/model_executor/models/glm_image passes (25 tests) both alone
and inside the full `tests/model_executor -m "core_model and cpu"` run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MpXe2LeRrUQGjZfKFaocMZ
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
@tzhouam tzhouam self-assigned this Sep 10, 2026
@tzhouam tzhouam added the ready label to trigger buildkite CI label Sep 10, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please also modify the installation docs

Comment thread docker/Dockerfile.ci Outdated
# (e.g. +gaeee7ef93.cu130) which pip/uv refuse to install from a PEP 503 package index.
# The cu130 wheel index lives at /<commit>/cu130/vllm/ but the actual wheel files
# are stored at the top level /<commit>/<wheel>.
ENV VLLM_PRECOMPILED_WHEEL_COMMIT=98dff2a81d747d1dba01a47f939f48c3526d4206

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove this

Drop the old-base wheel reinstall and dependency repair block. Update source installation and configuration references to vLLM 0.29 while retaining explicitly matched 0.28 prebuilt Omni examples until its next release.

Validation: official image and CUDA/ROCm release artifacts checked; 6 documentation tests passed; diff whitespace checks passed. Markdown lint reports existing snippet/anchor issues. Container build not run: no builder available.
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 10, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 10, 2026 03:23
@Gaohan123
Gaohan123 disabled auto-merge September 10, 2026 04:08
@Gaohan123
Gaohan123 merged commit aff7d64 into main Sep 10, 2026
10 of 14 checks passed
chickeyton added a commit to chickeyton/vllm-omni that referenced this pull request Sep 15, 2026
Brings 80 upstream commits onto the refactor, including the vLLM 0.29.0
rebase (vllm-project#7230), server-side VAD for turn-based models (vllm-project#6618), the non-beta
Realtime event names (vllm-project#7339) and the MiniCPM-o chat-template fix (vllm-project#7344).

Thirty conflicts. The pre-framework duplex stack that this branch removes
(runtime_adapter, runtime_bridge, session_runner, chat_fallback, protocol,
capability, realtime_output/session/state and their tests) stays removed;
upstream's edits to those files land in the refactored modules instead:

- vllm-project#7339 renamed four audio/transcript events in realtime_output.py, a file
  this branch deletes, so the merge would have left the new emitter on the old
  names while the merged clients/duplex.py maps only the new ones -- no audio
  deltas would reach DuplexClient, with no conflict to show for it. The four
  wire_type values in engine/duplex/events.py are renamed, along with the
  tests, docs and example that assert them.
- orchestrator.py: upstream's close_duplex_sessions kwarg and
  _is_duplex_session_request helper are this branch's release_owners and
  OrchestratorRequestState.session_owned under other names (upstream keys off
  req_state.duplex_identity, which the refactor replaced). Upstream's callers
  are retargeted -- including one that merged cleanly and would have called a
  helper that no longer exists. Upstream's status_code/error_type kwargs and
  its cleaned-up-request guard are kept.
- omni_engine_base.py: upstream extracted config resolution into
  config/resolver.py and dropped load_and_resolve_stage_configs from
  entrypoints/utils.py, which the refactored base class imports, so the merge
  was import-broken. Upstream's delta was three-way merged onto the base class
  (it is the old AsyncOmniEngine monolith post-split, which git cannot pair),
  keeping the diffusion/cache helpers the refactor still calls.
- api_server.py: upstream moved the duplex warmup into duplex/warmup.py. That
  extraction is adopted and the two architecture-specific lookups inside it
  are ported (engine plugin rather than serving adapter, clients.duplex rather
  than experimental.fullduplex).
- async_omni.py: upstream re-added five EngineClient properties that the
  refactor already provides on AsyncOmniBase/OmniBase; its InputProcessor
  rename is taken instead.
- vad.py: kept. Upstream's new server_vad.py is a separate Silero VAD for
  turn-based models, not a move of the energy VAD engine/duplex/turn_detection
  imports.
- Upstream's two duplex reaper-loop tests are ported onto
  DuplexSessionManager.reaper_loop, which is where that loop lives now; its
  pacing and its retry-after-failure were otherwise untested.
- CI: upstream's named source_file_dependencies groups are adopted, extended
  with the refactored modules and suites, and the jobs that ran the deleted
  turn-based chat e2e are dropped.

Known gap, not a merge artifact: vllm-project#6618 wires Silero server VAD into duplex
sessions through the old serving handler. This branch's serving layer is
transport-only and turn detection lives in engine/duplex/turn_detection.py, so
that wiring is not carried over and needs its own port.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
NumberWan added a commit to NumberWan/vllm-omni that referenced this pull request Sep 16, 2026
QwenImageTransformerBlock. That matches the Diffusers helper in unit
tests but drops Omni vs Diffusers pipeline PSNR below 27. Use the
pre-vllm-project#7230 RotaryEmbedding path in forward again; keep the helper for CPU tests.

Fixes vllm-project#7494

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
tzhouam added a commit to JiusiServe/InferMatrixCopilot that referenced this pull request Sep 18, 2026
## Change

Point the vllm-omni adapter's wheel pick and commit assignment at
`releases/v0.30.0`.

The v0.29.0 campaign merged as vllm-project/vllm-omni#7230, so the
adapter should now track the next release branch.

## Note for the reviewer

`releases/v0.30.0` is **cut but not yet tagged** (`f2aad6aa7`, branched
2026-09-15), and when this was written it carried no cherry-picks beyond
the main commit it was branched from. This tracks the branch head; the
pin moves to the tag commit once v0.30.0 is released, which is also when
`docker/Dockerfile.ci` in the target repo can drop its restored
commit-wheel block and go back to a published release image.

Tag cadence has been 15 days (0.28 Aug 24, 0.29 Sep 8), so expect the
re-pin around Sep 22.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants