Skip to content

[Core] Add native Omni and TTS Model Runner V2 on vLLM 0.29 - #4582

Open
Sy0307 wants to merge 102 commits into
vllm-project:mainfrom
Sy0307:sy03/mrv2-fixes-dev-20260621
Open

Sy0307 wants to merge 102 commits into
vllm-project:mainfrom
Sy0307:sy03/mrv2-fixes-dev-20260621

Conversation

@Sy0307

@Sy0307 Sy0307 commented Jun 20, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Run the native Omni/TTS Model Runner V2 pipeline on vLLM 0.29.0. This PR adds request-owned model state and an inter-stage data plane, then fixes output materialization, chunk conditioning, and shutdown ownership. The September 17 integration starts from the published PR head ea9cae74702017a7cbf980813df362b6aef1bc1b; it targets 0.29 directly, without a compatibility layer for older vLLM releases.

Execution and lifecycle

  • Add MRv2 model state with GPU-resident intermediate buffers and native chunk delivery, readiness, retry, terminal, and abort handling.
  • Preserve scheduler capacity for parked requests, owned request snapshots, and absolute Thinker/Talker prefill/decode coordinates.
  • Use the 0.29 renderer, serving protocols, weight loader, batch state, DP synchronization, and profiler contracts.
  • Split nested generation payloads per request without slicing a selected waveform a second time. CPU results own their storage; CPU-only output avoids unnecessary CUDA completion/copy work.
  • Always release the output materializer, data plane, and upstream runner during shutdown, including error paths.
  • Retain the full projected Talker prompt across chunked prefill and preserve unconsumed decode spans.
  • Preserve request seed isolation, no-dispatcher MTP sampling uniforms, codec validity, and the 0.29 structured-output min-tokens restore mask.

Optimizations included

Change Scope
Cache Qwen3-TTS BOS/EOS projections Preserve projection dtype; invalidate on weight loading.
Compact reference-waveform cache Owned float32 arrays, actual byte accounting, 1024 entries / 512 MiB; retain file access and invalidation rules.
Batch Code2Wav CPU codec staging One pinned packed input transfer instead of per-request transfers.
Cache MOSS resamplers and use NumPy waveform preparation Bounded 16-kernel cache; exact float32 CPU regression comparisons.
Deliver already-completed generation output before another forward Opt-in VLLM_OMNI_DRAIN_READY_GENERATION=1, default off; FIFO and UniProc with TP/PP/DP all 1.
Correct benchmark audio accounting Streaming errors fail requests; first nonempty PCM defines audio TTFP; optional WAV archiving occurs after timing.

The inherited high-concurrency TTS profile had prefix graphs enabled at the original head; this follow-up explicitly disables them pending quality validation. Experimental B2/B4 decoder defaults, MTP prefix computation, fused/Triton kernels, padded/GPU reference encoding, progressive chunking, EDF scheduling, and MPS activation are not enabled by this update. Their historical performance results are not measurements of this tree, and some change numerical results or have unresolved quality tradeoffs.

Activation

The recommended serving profiles qwen3_omni_moe.yaml, qwen3_tts.yaml, qwen3_tts_high_concurrency.yaml, and moss_tts_local.yaml on CUDA now default to model_runner: v2. Existing _mrv2.yaml aliases remain valid. NPU/XPU/ROCm/MUSA platform sections explicitly retain V1; platform runner overrides are applied before support checks. For CUDA V1 comparisons, set the top-level model_runner: v1. Thinker-only, forced-aligner, and Mori profiles retain their own settings. Internal codec CUDA graphs remain available; structured generation outputs use eager outer-runner dispatch.

These are user-facing files under vllm_omni/deploy, not CI pipeline YAMLs. The model pipeline selects its default when --deploy-config is omitted. See the recommended profile table and launch commands. MOSS coverage here is Local-Transformer-v1.5; Delay, Realtime and Nano retain their separate settings.

Validation — H200, vLLM 0.29.0

  • MOSS Local recommendation follow-up (bbcf6f8bc): 370 config tests passed, 3 skipped, all changed-file hooks passed, including mypy. No new performance claim or C64 rerun.
  • Default-profile follow-up (b0ecf406a): 365 config tests passed, 3 skipped; all changed-file pre-commit hooks, including mypy, pass. This does not clear the broader historical mypy debt below.
  • CPU regression suite: 2116 passed, 2 skipped, 123 deselected. Covers runner V1/V2, inputs, engine, Qwen/MOSS models, reference cache, speech serving, environment inventory, and benchmark accounting.
  • Targeted API/reasoning/worker rerun: 430 passed, 1 deselected (overlaps the broad suite; counts are not additive).
  • CUDA contract suite: 100 passed, 5 deselected (async metadata/output, Qwen decoder graphs, MOSS codec graphs).
  • Native vllm serve --omni + vllm bench serve --omni: Qwen3-TTS Base V1 and V2, Qwen3-Omni V2, and MOSS local V2 each completed 16/16 requests with nonempty audio (8 at C1, 8 at C4; no MPS).
  • MOSS V2 with VLLM_OMNI_DRAIN_READY_GENERATION=1 also completed 16/16 requests with nonempty audio; the server log confirms activation. The final remote source matches all 71/71 changed local files by SHA-256.
  • Initial Qwen3-TTS V1 validation had one Missing hidden_states['last'] failure. An instrumented run, a clean uninstrumented retry, and the final clean V1 run each passed 16 requests. The initial failure has no conclusive root cause; these retries do not establish its elimination.
  • Changed-file Ruff/format, markdown, SPDX, test marks, import/CUDA ratchets, and whitespace checks pass. The required mypy hook remains failing: 208 diagnostics versus 214 on the original PR head with the same file selection, with zero new diagnostics after normalizing source line numbers. The integration commit explicitly skips only mypy-3.10; this is not a claim of fully passing pre-commit.
  • Original-head DCO reports 54 historical sign-off problems, including inherited commits. This update adds a properly signed commit without rewriting existing authors/history. DCO must be rechecked on the updated head.

Not ready to merge: the current codec-EOS failure, historical type/DCO gates, full CI, human review, and matched quality/performance qualification remain. These C1/C4 functional checks do not establish ASR/SIM equivalence or C64/C128 throughput. Prior H20/H200 benchmark numbers describe older trees and are deliberately not advertised as this update's measured performance.

Performance evidence: historical, not current-head claims

Evidence report and machine-readable records retain configurations, per-run aggregates and provenance hashes.

Archived workload/configuration Audio-s/s Boundary
Qwen3-TTS Base, 1 H200, C64, B1/no MPS baseline 75.10 Archived source/configuration
TTS copylite experiment, B1/no MPS 88.15 Unsafe shallow-copy experiment not imported; PR retains owned snapshots
TTS compact cache + B2 103.52 B2 not enabled by default; numerical/quality qualification incomplete
TTS compact cache + B2 + MPS, final repeat 147.57 Includes experimental graph/runtime settings absent from current defaults
Qwen3-Omni, 2 H200, C64, V1 → V2 82.98 → 115.04 Historical matched-runner comparison; TTFP p95 worsened 757.6 → 911.6 ms
Qwen3-Omni, 2 H200, C128, V1 → V2 84.39 → 113.85 Historical matched-runner comparison

These full SeedTTS measurements used vLLM 0.29 and repeated measured rounds after full warmup. Omni sampled ASR WER was 2.15% versus 3.08%; quality equivalence was not established. Historical numbers are not substituted for measurements of this head.

Current C64 verification paused on a failure

At source b0ecf406a, the first full SeedTTS EN1088/C64 warmup completed 1087/1088 requests. Request index48 reached the 1024 codec-token budget without EOS; the server raised Qwen3TTSCodecLimitError and the client saw an incomplete streaming transfer. The cause is unresolved. The automation stopped before the remaining warmups, formal rounds or historical-configuration remeasurement. Its 74.68 audio-s/s is a failed warmup, not a valid performance result. The earlier C16/64-request screening is retained only as a functional check and is not used to claim steady-state throughput.

Current CustomVoice and VoiceDesign MRv2 smoke checks each completed 16/16 nonempty audio using seed-tts-text (8 at C1, 8 at C4). An initial CustomVoice test used the wrong Base-task dataset and returned400; that attempt is preserved and excluded from pass counts.

AI assistance

Codex assisted with implementation, experiment consolidation, tests, and this description. Contributor human review is still required before merge; this text does not attest that human review is complete.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@Sy0307

Sy0307 commented Jun 20, 2026

Copy link
Copy Markdown
Collaborator Author

PTAK @tzhouam

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed.

if cur_rows > prev_seen_rows:
code_predictor_codes = code_predictor_codes[prev_seen_rows:]
seen_codec_rows[request_id] = cur_rows
else:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This duplicates already-emitted cumulative rows when the next payload has no new valid codec rows. Example: previous filtered rows = 1, next raw payload = the same valid row plus only invalid placeholder rows, so cur_rows == prev_seen_rows; this branch keeps the old row and bumps seen_codec_rows. In the cumulative case this should emit no rows and leave the seen count unchanged.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f0010f7. The Qwen3-Omni async talker->code2wav path now filters invalid codec rows, tracks cumulative codec rows, and flushes buffered tail rows on terminal empty/invalid steps before sending EOF. That prevents stage2 from receiving only meta without codes.audio. Added regression coverage in tests/model_executor/stage_input_processors/test_qwen3_omni_streaming_helpers.py. H20 validation: related pytest suite 224 passed; text+audio async_chunk e2e produced txt+wav for 1/1 and 2/2 prompts with no Traceback, overflow, No free indices, or make_omni_output failure.

)

original_max_num_running_reqs = self.max_num_running_reqs
reserved_running_slots = self._get_async_chunk_reserved_running_slots() if self.chunk_transfer_adapter else 0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This same reservation is still missing from OmniGenerationScheduler. It also calls chunk_transfer_adapter.process_pending_chunks() and then admits waiting requests with len(self.running) < self.max_num_running_reqs, so parked generation-stage requests in waiting_for_chunk_running_requests / _held_non_active can be over-admitted before restore_queues() puts them back.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f0010f7. The generation scheduler now distinguishes a terminal empty async-chunk prompt that still carries connector-delivered codes.audio, schedules one placeholder step for code2wav, and avoids pending-finish before the payload runs. Code2Wav also prefers the connector payload over scheduler placeholder input ids. Added scheduler/model regression coverage. H20 validation: related pytest suite 224 passed; Qwen3-Omni text+audio async_chunk e2e passed 1 prompt and 2 concurrent prompts, each producing txt+wav, with no Traceback, overflow, No free indices, or make_omni_output failure.

@Sy0307
Sy0307 force-pushed the sy03/mrv2-fixes-dev-20260621 branch 2 times, most recently from e21edd6 to f0010f7 Compare June 28, 2026 12:15
@hsliuustc0106 hsliuustc0106 added bug Something isn't working omni code related to omni models labels Jul 17, 2026 — with ChatGPT Codex Connector
@Sy0307
Sy0307 force-pushed the sy03/mrv2-fixes-dev-20260621 branch from f0010f7 to fca469f Compare July 17, 2026 08:10
zyz111222 and others added 8 commits September 9, 2026 19:03
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
…init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307
Sy0307 requested a review from tjtanaa as a code owner September 9, 2026 15:04
@Sy0307

Sy0307 commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator Author

Update: the local MRV2 architecture candidate has now been rebased onto main 2f1845fa and pushed to this PR. The implementation/results publication is 60125b15; the latest tip 621d27cf additionally records Qwen3-Omni validation; the measured source was 07da1fed before DCO/ancestry and formatting/SPDX-only publication changes. The published branch retains the prior PR head aa8fbba as an ancestor, so the update is a fast-forward and preserves previous contributors' commits.

Latest Qwen3-TTS V1 vs V2 results

Qwen3-TTS-12Hz-1.7B-Base / Base voice_clone, H200, closed-loop c64. Each runner used 8 c1 requests plus 3 × 1,088 full SeedTTS English warmup requests, followed by 2 × 1,088 measured requests. Complete reference inputs were preserved. Stage0 max_num_seqs=64; Stage1 max_num_seqs=10, decoder B4. V2 completion events enabled, experimental async-index staging disabled. No profiling during measurement.

The final column is (V2 / V1 − 1) × 100%. Positive is better for audio throughput; negative is better for latency/gaps.

Metric V1 V2 V2 relative to V1
Audio throughput (audio-s/s) ↑ 64.22 89.23 +38.9%
TTFA mean (ms) ↓ 447.2 447.1 -0.0%
TTFA P95 (ms) ↓ 1098.4 1040.5 -5.3%
E2E mean (s) ↓ 4.082 2.945 -27.8%
Simulated playback gap (s/request) ↓ 1.703 1.104 -35.2%
Failures / total requests, including warmup 0 / 5,448 0 / 5,448 N/A (zero denominator)

Audio throughput is decoded audio seconds / total measured wall time, not requests/s. TTFA excludes the standalone WAV header. P95 is computed from the combined measured request samples, not the average of two P95s. Playback gaps simulate immediate playback at first audio with no startup buffer; they are not an actual player's underrun measurements. Both runners still exhibit gaps.

The throughput advantage is consistent with the old-base, equally warmed comparison: 63.08 / 86.66 audio-s/s (V1 / V2), +37.4%, versus +38.9% after rebase. Mean TTFA is effectively equal in this latest pair. V2 remains better on measured E2E and playback gaps.

However, absolute TTFA P95 increased for both runners versus the old-base pair (old: 1,011 / 923 ms, new: 1,098 / 1,041 ms). This is a single ordered cross-check, not evidence that tail latency is unchanged or that any specific main commit caused the shift. The earlier lightly warmed c64 P95 regression remains a separate condition and is not erased by this result.

Changes since the previous push (aa8fbba)

  • Earlier native AR output publication: a bounded worker materializes GPU-completed native output independently of the engine's lazy result consumption. Scheduler/connector state consumption stays on the owner thread; metadata reservations precede submission, and shutdown drains pending work. Single-rank native AR uses this path, retaining the existing multiprocess fallback.
  • Generation completion-driven scheduling: input arrival and oldest-output completion wake the engine, allowing it to return to input processing after submitting a batch instead of immediately blocking in future.result(). Queue capacity, FIFO consumption, abort-before-state-update, and the generation execute_model → sample_tokens output-construction contract are retained. This is opt-in.
  • Separate generation execution batch limits from request-state capacity, avoiding capacity changes simply to bound native execution. Retained-state codecs keep their admission contract.
  • Batched runner state/index handling: reuse existing device offsets and batch state snapshots where the contracts allow it. A separate pinned-buffer async-index experiment is included behind a disabled-by-default switch and was not enabled in the table above.
  • Full request snapshot fast path: preserve all fields, aliases, and mutable-container isolation while avoiding generic deepcopy dispatch for plain scalar trees. This is shared orchestrator work, not a claimed MRV2-exclusive or reference-pruning gain.
  • Main compatibility: migrate connector changes into main's split runtime/payload-transport modules, preserve its TTS adapters, and retain explicit decode-span accumulation inside the MRV2 data plane. Regression tests cover output completion races, ordered consumption, abort handling, payload lifecycle, and seed/EOS error propagation.

Nsight separately confirmed that Stage1 main-thread event waiting moved to a background observer (about 3.33 s → 2 ms in the respective diagnostic windows). This moves waiting off the owner thread; it does not remove GPU execution time and must not be presented as an equivalent end-to-end speedup.

Validation and limitations

  • 798 CPU tests passed, one missing-HF-config case skipped, one GPU case deselected; 5 targeted TTS seed/EOS tests passed.
  • CI-aligned whole-repository pre-commit passed with the skip list already used by .github/workflows/pre-commit.yml. The stricter local invocation exposed type-checking backlog; it is not reported as passing. Formatting/SPDX fixes were retained and checked to preserve Python ASTs.
  • All 5,448 requests per runner succeeded in this pair, including warmup. This does not establish that historical intermittent EOS caps are fixed. No retries, forced EOS, or token-limit increases were used.
  • No foreign process was detected on the measured GPU by the ownership samples. Shared CPU load was not controlled. No randomized/interleaved repetitions or confidence intervals; small changes should not be treated as established improvements.
  • Stage1 has not been jointly tuned with higher Stage0 capacities. More max_num_seqs is not automatically better for first audio or continuous playback.
  • Qwen3-Omni was tested separately on the same published implementation; see the validation result below. TTS results must not be generalized to it.

Checked-in summaries and methodology: benchmarks/results/mrv2_20260909.

CI follow-up: pre-commit passed. The DCO check currently fails on 54 imported upstream commits whose author/sign-off identities do not match. All new Sy03 publication commits have sign-offs. The PR API still reports base SHA 715b8b8 although the main ref has advanced, so the check includes upstream history in its commit range. I have not rewritten other contributors' signatures; this needs resolution before merge.

Qwen3-Omni follow-up — V1 audio baseline is currently invalid

Tested Qwen3-Omni-30B-A3B-Instruct on 2 H200s, thinker on one GPU and talker/Code2Wav on the other. This is a small probe: 8 short sentence prompts repeated, 32 c8 warmup requests, 6 mixed text/image/audio smoke requests, then 64 requests each at c8 and c16. It is not the SeedTTS workload or a full capacity/quality evaluation.

  • V1 with the default decoder CUDA graph failed during warmup: model_outputs had one entry for a batch of two requests.
  • With decoder eager, V1 returned text and termination but no valid PCM in all 32 warmup requests. A corrected client skipped zero-frame WAV chunks and waited for stream completion; the absence of audio remained. These failures are retained and the cause is not fixed here.
  • V2 completed all 166 requests with valid finite audio, text and termination, and no reported length finish reason. The table below uses decoder eager for both runners; thinker/talker graph settings stay unchanged.
Metric V1 V2 V2 relative to V1
c8 audio throughput (audio-s/s) ↑ Invalid audio baseline 46.14 N/A
c8 TTFA mean (ms) ↓ Invalid audio baseline 190.8 N/A
c8 TTFA P95 (ms) ↓ Invalid audio baseline 250.4 N/A
c8 E2E mean (s) ↓ Invalid audio baseline 0.649 N/A
c16 audio throughput (audio-s/s) ↑ Invalid audio baseline 65.69 N/A
c16 TTFA mean (ms) ↓ Invalid audio baseline 290.1 N/A
c16 TTFA P95 (ms) ↓ Invalid audio baseline 508.1 N/A
c16 E2E mean (s) ↓ Invalid audio baseline 0.857 N/A

No percentage speedup is claimed for Omni: dividing by a failed V1 run would be misleading. V2's default-graph run also completed, but is not mixed into the matched eager comparison. No WER/MOS or full codec-trajectory equivalence check was performed, and playback gaps were not measured in this Omni probe. GPU ownership samples detected no foreign process; the short phases only provide a few samples. All owned services have been cleaned up.

Omni validation and failure record.

Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com>
@yiwen101

Copy link
Copy Markdown

@Sy0307 Thanks for the scheduling and output-delivery improvements.

@hsliuustc0106 I further investigated why MRV1 on this PR is slower than main, and found two separate causes.

Without a request seed

The PR adds an unnecessary CPU wait during output processing. After the fix, MRV1 throughput to roughly main's level without request seed

MRV1 throughput, requests/s:

Concurrency Main PR before fix PR + copy fix
1 2.013 1.580 2.082
4 4.858 3.883 4.955

With request seed : The slowdown remains.

Main's speech-code-prediction path has two distinct behaviors:

  1. Its CUDA graph replay does not consume the newly supplied request generator, so individual request's seed are not followed.
  2. Despite that, supplying a seed still disables batching: main processes the speech-code-prediction rows one at a time.

The PR keeps batching and honors each request's generator by bypassing the outer graph which reduce throughput. Temporarily restoring graph replay recovered the seeded throughput, identifying that bypass path as a major source of the slowdown.

Concurrency Main (not request seeded, not batched) PR + copy fix (request seeded, batched) PR + copy fix + graph replay (not request seeded, batched)
1 2.007 0.890 2.023
4 2.643 1.161 4.697

I have not push the restored graph replay change yet, as I hope to confirm if main's supplied request generator behavior intentional? Should this speech-code-prediction step honor each request's seed?

Preserve per-request output ownership, chunked prefill conditioning, and shutdown cleanup. Consolidate reference preparation, projection caching, packed codec staging, and opt-in completed-generation draining.

Validation: H200 vLLM 0.29 CPU/CUDA contracts and Qwen TTS V1/V2, Omni V2, MOSS V2 serving smoke. Existing mypy debt explicitly retained: 208 diagnostics versus 214 at parent with no new diagnostics; only mypy-3.10 skipped for this commit. Full CI and quality/performance qualification remain pending.
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 changed the title [Feature][MR V2] Add native Qwen3 Omni and TTS execution [Core] Add native Omni and TTS Model Runner V2 on vLLM 0.29 Sep 17, 2026
Apply validated platform runner overrides before support checks so non-CUDA profiles retain V1. Keep the high-concurrency TTS prefix graph disabled pending matched quality validation.

Signed-off-by: Sy03 <1370724210@qq.com>
Record archived full-workload TTS and Omni evidence with configuration and provenance. Mark current C64 validation paused after a codec-EOS limit failure in the first full warmup; do not promote the short C16 screening or failed warmup as current-head performance.

Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Select MRv2 in the MOSS Local CUDA deployment profile and document the recommended Qwen TTS, Qwen Omni, and MOSS Local serving YAMLs. Preserve platform-specific V1 choices and distinguish deployment recommendations from CI and performance qualification.

Signed-off-by: Sy03 <1370724210@qq.com>
…nner selection

- Declare supports_native_mrv2_data_plane on both MOSS-TTS-Local stages so
  the recommended moss_tts_local.yaml (model_runner: v2, async_chunk) engages
  the native MRv2 data plane instead of silently falling back to the legacy
  chunk-transfer path. Verified end to end on 1x H200 with vLLM 0.29: server
  ready, native Stage-1 receiver markers present, 4/4 speech requests OK.
- Write the runner selection into stage engine args unconditionally; reject
  engine_extras entries named use_v2_model_runner or
  supports_native_mrv2_data_plane instead of letting them silently win.
- Extend the recommended-profile platform test to prove the resolved runner
  reaches every stage's engine args, and that a v2 recommendation always
  carries native-plane support declarations.
- Map the upstream 'Cannot load local files without
  --allowed-local-media-path' refusal to an actionable HTTP 400 instead of an
  internal-error traceback; add a regression test.
- Document cold-start duration and the file:// ref-audio opt-in for the
  recommended CUDA profiles.

mypy-3.10 hook skipped on commit: it reports the same 11 pre-existing errors
in serving_speech.py at the base bbcf6f8; this change neither adds nor
removes any of them (verified by base/head hook run comparison).

Signed-off-by: Sy03 <1370724210@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 11 days

@Sy0307 this pull request has had no human commit, comment or review since 2026-09-13. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as stale.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

PR description

This PR lands native Omni/TTS Model Runner V2 for vLLM 0.29: request-owned model state, an inter-stage data plane, generation/AR scheduler capacity for parked chunks, and output materialization fixes. Recommended CUDA deploy profiles (qwen3_tts, qwen3_tts_high_concurrency, qwen3_omni_moe, moss_tts_local) now default to model_runner: v2, while NPU/XPU/ROCm/MUSA platform overlays keep V1. User-visible effect: CUDA Omni/TTS serves use the V2 runner path by default, with opt-in generation drain and connector wakeup behind env flags.

Change flow

flowchart TD
  A["[EXISTING] Omni/TTS serve or offline request"]:::existing
  B["[CHANGED] CUDA deploy profiles default model_runner=v2"]:::changed
  R["[REMOVED] CUDA default model_runner=v1 selection"]:::removed
  C["[NEW] worker_v2 runners, model state, native data plane"]:::new
  D["[CHANGED] Omni AR/generation schedulers + stage engine"]:::changed
  E["[CHANGED] Qwen/MOSS stage processors and Code2Wav handoff"]:::changed
  F["[EXISTING] Speech/audio client outputs"]:::existing
  A --> B
  B --> R
  B --> C --> D --> E --> F
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

CI at cb4c63abb270 (2026-10-02T13:26:36.353141+00:00): required check(s) blocking: pre-commit (missing), buildkite/vllm-omni (missing).

See inline comments below.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

],
)
def test_qwen3_mrv2_profiles_are_explicit_opt_in(default_name: str, mrv2_name: str):
assert load_deploy_config(_DEPLOY_DIR / default_name).model_runner == "v1"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Stale config test still requires default profiles to be V1

test_qwen3_mrv2_profiles_are_explicit_opt_in asserts qwen3_tts.yaml, qwen3_omni_moe.yaml, and qwen3_tts_high_concurrency.yaml load as model_runner == "v1", but those deploy files now set top-level model_runner: v2 (and _mrv2.yaml aliases document that the base profile already defaults to MRv2 on CUDA). The same head also adds test_recommended_native_runner_platform_defaults in tests/config/test_config_factory.py, which correctly expects CUDA V2. As written, the opt-in test fails on this head and contradicts the PR’s default-cutover intent, so the reported “config tests passed” evidence cannot apply to cb4c63abb270.


assert default_extra["decode_batch_max_size"] == 1
assert default_extra["decode_batch_bucket_frames"] == []
assert default_extra["code_predictor_prefix_graph_seq_lens"] == [2, 3, 4, 5, 6, 7, 8]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] High-concurrency MRv1 pin test references a removed YAML key

test_qwen3_tts_mrv2_retunes_do_not_change_default_mrv1_profile reads default_extra["code_predictor_prefix_graph_seq_lens"] from qwen3_tts_high_concurrency.yaml, but that key is absent there (only code_predictor_prefix_graphs: false remains; the seq-lens list exists only on qwen3_tts_high_concurrency_mrv2.yaml). This raises KeyError before later assertions run. Update the pin to the current default profile contract, or drop the obsolete seq-lens expectation.

)


class _OmniConnectorPayloadTransportMixin(_OmniConnectorRuntimeMixin):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Diff exceeds size budget without a split or exemption

Host-computed diff_size: non_test_added_lines=34426 over budget=1000 across 339 counted files (282 exempt). Largest contributors include vllm_omni/distributed/omni_connectors/model_runner/omni_connector_payload_transport.py (+1579), vllm_omni/worker_v2/model_states/omni_model_state.py (+1311), docs/design/feature/mrv2_performance_evidence_20260917.json (+1193), omni_connector_runtime.py (+1151), and omni_ar_model_runner.py (+838). The PR body explains the MRv2 integration but does not give a split plan or concrete size exemption. Please publish a follow-up split (for example transport/runtime vs runner/state vs deploy defaults) or an explicit exemption before merge.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@Sy0307 this PR is labeled ready + high priority, but CI is failing on the latest commit:

Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working high priority high priority issue, needs to be done asap omni code related to omni models

Projects

None yet

Development

Successfully merging this pull request may close these issues.