Repository navigation
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
PTAK @tzhouam |
| if cur_rows > prev_seen_rows: | ||
| code_predictor_codes = code_predictor_codes[prev_seen_rows:] | ||
| seen_codec_rows[request_id] = cur_rows | ||
| else: |
There was a problem hiding this comment.
This duplicates already-emitted cumulative rows when the next payload has no new valid codec rows. Example: previous filtered rows = 1, next raw payload = the same valid row plus only invalid placeholder rows, so cur_rows == prev_seen_rows; this branch keeps the old row and bumps seen_codec_rows. In the cumulative case this should emit no rows and leave the seen count unchanged.
There was a problem hiding this comment.
Fixed in f0010f7. The Qwen3-Omni async talker->code2wav path now filters invalid codec rows, tracks cumulative codec rows, and flushes buffered tail rows on terminal empty/invalid steps before sending EOF. That prevents stage2 from receiving only meta without codes.audio. Added regression coverage in tests/model_executor/stage_input_processors/test_qwen3_omni_streaming_helpers.py. H20 validation: related pytest suite 224 passed; text+audio async_chunk e2e produced txt+wav for 1/1 and 2/2 prompts with no Traceback, overflow, No free indices, or make_omni_output failure.
| ) | ||
|
|
||
| original_max_num_running_reqs = self.max_num_running_reqs | ||
| reserved_running_slots = self._get_async_chunk_reserved_running_slots() if self.chunk_transfer_adapter else 0 |
There was a problem hiding this comment.
This same reservation is still missing from OmniGenerationScheduler. It also calls chunk_transfer_adapter.process_pending_chunks() and then admits waiting requests with len(self.running) < self.max_num_running_reqs, so parked generation-stage requests in waiting_for_chunk_running_requests / _held_non_active can be over-admitted before restore_queues() puts them back.
There was a problem hiding this comment.
Fixed in f0010f7. The generation scheduler now distinguishes a terminal empty async-chunk prompt that still carries connector-delivered codes.audio, schedules one placeholder step for code2wav, and avoids pending-finish before the payload runs. Code2Wav also prefers the connector payload over scheduler placeholder input ids. Added scheduler/model regression coverage. H20 validation: related pytest suite 224 passed; Qwen3-Omni text+audio async_chunk e2e passed 1 prompt and 2 concurrent prompts, each producing txt+wav, with no Traceback, overflow, No free indices, or make_omni_output failure.
e21edd6 to
f0010f7
Compare
f0010f7 to
fca469f
Compare
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
…init component (vllm-project#7328) Signed-off-by: ZhengWG <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
|
Update: the local MRV2 architecture candidate has now been rebased onto main Latest Qwen3-TTS V1 vs V2 resultsQwen3-TTS-12Hz-1.7B-Base / Base voice_clone, H200, closed-loop c64. Each runner used 8 c1 requests plus 3 × 1,088 full SeedTTS English warmup requests, followed by 2 × 1,088 measured requests. Complete reference inputs were preserved. Stage0 max_num_seqs=64; Stage1 max_num_seqs=10, decoder B4. V2 completion events enabled, experimental async-index staging disabled. No profiling during measurement. The final column is (V2 / V1 − 1) × 100%. Positive is better for audio throughput; negative is better for latency/gaps.
Audio throughput is decoded audio seconds / total measured wall time, not requests/s. TTFA excludes the standalone WAV header. P95 is computed from the combined measured request samples, not the average of two P95s. Playback gaps simulate immediate playback at first audio with no startup buffer; they are not an actual player's underrun measurements. Both runners still exhibit gaps. The throughput advantage is consistent with the old-base, equally warmed comparison: 63.08 / 86.66 audio-s/s (V1 / V2), +37.4%, versus +38.9% after rebase. Mean TTFA is effectively equal in this latest pair. V2 remains better on measured E2E and playback gaps. However, absolute TTFA P95 increased for both runners versus the old-base pair (old: 1,011 / 923 ms, new: 1,098 / 1,041 ms). This is a single ordered cross-check, not evidence that tail latency is unchanged or that any specific main commit caused the shift. The earlier lightly warmed c64 P95 regression remains a separate condition and is not erased by this result. Changes since the previous push (
|
| Metric | V1 | V2 | V2 relative to V1 |
|---|---|---|---|
| c8 audio throughput (audio-s/s) ↑ | Invalid audio baseline | 46.14 | N/A |
| c8 TTFA mean (ms) ↓ | Invalid audio baseline | 190.8 | N/A |
| c8 TTFA P95 (ms) ↓ | Invalid audio baseline | 250.4 | N/A |
| c8 E2E mean (s) ↓ | Invalid audio baseline | 0.649 | N/A |
| c16 audio throughput (audio-s/s) ↑ | Invalid audio baseline | 65.69 | N/A |
| c16 TTFA mean (ms) ↓ | Invalid audio baseline | 290.1 | N/A |
| c16 TTFA P95 (ms) ↓ | Invalid audio baseline | 508.1 | N/A |
| c16 E2E mean (s) ↓ | Invalid audio baseline | 0.857 | N/A |
No percentage speedup is claimed for Omni: dividing by a failed V1 run would be misleading. V2's default-graph run also completed, but is not mixed into the matched eager comparison. No WER/MOS or full codec-trajectory equivalence check was performed, and playback gaps were not measured in this Omni probe. GPU ownership samples detected no foreign process; the short phases only provide a few samples. All owned services have been cleaned up.
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com>
|
@Sy0307 Thanks for the scheduling and output-delivery improvements. @hsliuustc0106 I further investigated why MRV1 on this PR is slower than main, and found two separate causes. Without a request seed The PR adds an unnecessary CPU wait during output processing. After the fix, MRV1 throughput to roughly main's level without request seed MRV1 throughput, requests/s:
With request seed : The slowdown remains. Main's speech-code-prediction path has two distinct behaviors:
The PR keeps batching and honors each request's generator by bypassing the outer graph which reduce throughput. Temporarily restoring graph replay recovered the seeded throughput, identifying that bypass path as a major source of the slowdown.
I have not push the restored graph replay change yet, as I hope to confirm if main's supplied request generator behavior intentional? Should this speech-code-prediction step honor each request's seed? |
Preserve per-request output ownership, chunked prefill conditioning, and shutdown cleanup. Consolidate reference preparation, projection caching, packed codec staging, and opt-in completed-generation draining. Validation: H200 vLLM 0.29 CPU/CUDA contracts and Qwen TTS V1/V2, Omni V2, MOSS V2 serving smoke. Existing mypy debt explicitly retained: 208 diagnostics versus 214 at parent with no new diagnostics; only mypy-3.10 skipped for this commit. Full CI and quality/performance qualification remain pending. Signed-off-by: Sy03 <1370724210@qq.com>
Apply validated platform runner overrides before support checks so non-CUDA profiles retain V1. Keep the high-concurrency TTS prefix graph disabled pending matched quality validation. Signed-off-by: Sy03 <1370724210@qq.com>
Record archived full-workload TTS and Omni evidence with configuration and provenance. Mark current C64 validation paused after a codec-EOS limit failure in the first full warmup; do not promote the short C16 screening or failed warmup as current-head performance. Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Select MRv2 in the MOSS Local CUDA deployment profile and document the recommended Qwen TTS, Qwen Omni, and MOSS Local serving YAMLs. Preserve platform-specific V1 choices and distinguish deployment recommendations from CI and performance qualification. Signed-off-by: Sy03 <1370724210@qq.com>
…nner selection - Declare supports_native_mrv2_data_plane on both MOSS-TTS-Local stages so the recommended moss_tts_local.yaml (model_runner: v2, async_chunk) engages the native MRv2 data plane instead of silently falling back to the legacy chunk-transfer path. Verified end to end on 1x H200 with vLLM 0.29: server ready, native Stage-1 receiver markers present, 4/4 speech requests OK. - Write the runner selection into stage engine args unconditionally; reject engine_extras entries named use_v2_model_runner or supports_native_mrv2_data_plane instead of letting them silently win. - Extend the recommended-profile platform test to prove the resolved runner reaches every stage's engine args, and that a v2 recommendation always carries native-plane support declarations. - Map the upstream 'Cannot load local files without --allowed-local-media-path' refusal to an actionable HTTP 400 instead of an internal-error traceback; add a regression test. - Document cold-start duration and the file:// ref-audio opt-in for the recommended CUDA profiles. mypy-3.10 hook skipped on commit: it reports the same 11 pre-existing errors in serving_speech.py at the base bbcf6f8; this change neither adds nor removes any of them (verified by base/head hook run comparison). Signed-off-by: Sy03 <1370724210@qq.com>
Omni ReviewBot: no human activity for 11 days@Sy0307 this pull request has had no human commit, comment or review since 2026-09-13. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
Omni ReviewBot attempt recordReview attempt ended as stale. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
PR description
This PR lands native Omni/TTS Model Runner V2 for vLLM 0.29: request-owned model state, an inter-stage data plane, generation/AR scheduler capacity for parked chunks, and output materialization fixes. Recommended CUDA deploy profiles (qwen3_tts, qwen3_tts_high_concurrency, qwen3_omni_moe, moss_tts_local) now default to model_runner: v2, while NPU/XPU/ROCm/MUSA platform overlays keep V1. User-visible effect: CUDA Omni/TTS serves use the V2 runner path by default, with opt-in generation drain and connector wakeup behind env flags.
Change flow
flowchart TD
A["[EXISTING] Omni/TTS serve or offline request"]:::existing
B["[CHANGED] CUDA deploy profiles default model_runner=v2"]:::changed
R["[REMOVED] CUDA default model_runner=v1 selection"]:::removed
C["[NEW] worker_v2 runners, model state, native data plane"]:::new
D["[CHANGED] Omni AR/generation schedulers + stage engine"]:::changed
E["[CHANGED] Qwen/MOSS stage processors and Code2Wav handoff"]:::changed
F["[EXISTING] Speech/audio client outputs"]:::existing
A --> B
B --> R
B --> C --> D --> E --> F
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
CI at
cb4c63abb270(2026-10-02T13:26:36.353141+00:00): required check(s) blocking:pre-commit(missing),buildkite/vllm-omni(missing).
See inline comments below.
🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!
| ], | ||
| ) | ||
| def test_qwen3_mrv2_profiles_are_explicit_opt_in(default_name: str, mrv2_name: str): | ||
| assert load_deploy_config(_DEPLOY_DIR / default_name).model_runner == "v1" |
There was a problem hiding this comment.
[P1] Stale config test still requires default profiles to be V1
test_qwen3_mrv2_profiles_are_explicit_opt_in asserts qwen3_tts.yaml, qwen3_omni_moe.yaml, and qwen3_tts_high_concurrency.yaml load as model_runner == "v1", but those deploy files now set top-level model_runner: v2 (and _mrv2.yaml aliases document that the base profile already defaults to MRv2 on CUDA). The same head also adds test_recommended_native_runner_platform_defaults in tests/config/test_config_factory.py, which correctly expects CUDA V2. As written, the opt-in test fails on this head and contradicts the PR’s default-cutover intent, so the reported “config tests passed” evidence cannot apply to cb4c63abb270.
|
|
||
| assert default_extra["decode_batch_max_size"] == 1 | ||
| assert default_extra["decode_batch_bucket_frames"] == [] | ||
| assert default_extra["code_predictor_prefix_graph_seq_lens"] == [2, 3, 4, 5, 6, 7, 8] |
There was a problem hiding this comment.
[P1] High-concurrency MRv1 pin test references a removed YAML key
test_qwen3_tts_mrv2_retunes_do_not_change_default_mrv1_profile reads default_extra["code_predictor_prefix_graph_seq_lens"] from qwen3_tts_high_concurrency.yaml, but that key is absent there (only code_predictor_prefix_graphs: false remains; the seq-lens list exists only on qwen3_tts_high_concurrency_mrv2.yaml). This raises KeyError before later assertions run. Update the pin to the current default profile contract, or drop the obsolete seq-lens expectation.
| ) | ||
|
|
||
|
|
||
| class _OmniConnectorPayloadTransportMixin(_OmniConnectorRuntimeMixin): |
There was a problem hiding this comment.
[P2] Diff exceeds size budget without a split or exemption
Host-computed diff_size: non_test_added_lines=34426 over budget=1000 across 339 counted files (282 exempt). Largest contributors include vllm_omni/distributed/omni_connectors/model_runner/omni_connector_payload_transport.py (+1579), vllm_omni/worker_v2/model_states/omni_model_state.py (+1311), docs/design/feature/mrv2_performance_evidence_20260917.json (+1193), omni_connector_runtime.py (+1151), and omni_ar_model_runner.py (+838). The PR body explains the MRv2 integration but does not give a split plan or concrete size exemption. Please publish a follow-up split (for example transport/runtime vs runner/state vs deploy defaults) or an explicit exemption before merge.
Purpose
Run the native Omni/TTS Model Runner V2 pipeline on vLLM 0.29.0. This PR adds request-owned model state and an inter-stage data plane, then fixes output materialization, chunk conditioning, and shutdown ownership. The September 17 integration starts from the published PR head
ea9cae74702017a7cbf980813df362b6aef1bc1b; it targets 0.29 directly, without a compatibility layer for older vLLM releases.Execution and lifecycle
Optimizations included
VLLM_OMNI_DRAIN_READY_GENERATION=1, default off; FIFO and UniProc with TP/PP/DP all 1.The inherited high-concurrency TTS profile had prefix graphs enabled at the original head; this follow-up explicitly disables them pending quality validation. Experimental B2/B4 decoder defaults, MTP prefix computation, fused/Triton kernels, padded/GPU reference encoding, progressive chunking, EDF scheduling, and MPS activation are not enabled by this update. Their historical performance results are not measurements of this tree, and some change numerical results or have unresolved quality tradeoffs.
Activation
The recommended serving profiles
qwen3_omni_moe.yaml,qwen3_tts.yaml,qwen3_tts_high_concurrency.yaml, andmoss_tts_local.yamlon CUDA now default tomodel_runner: v2. Existing_mrv2.yamlaliases remain valid. NPU/XPU/ROCm/MUSA platform sections explicitly retain V1; platform runner overrides are applied before support checks. For CUDA V1 comparisons, set the top-levelmodel_runner: v1. Thinker-only, forced-aligner, and Mori profiles retain their own settings. Internal codec CUDA graphs remain available; structured generation outputs use eager outer-runner dispatch.These are user-facing files under
vllm_omni/deploy, not CI pipeline YAMLs. The model pipeline selects its default when--deploy-configis omitted. See the recommended profile table and launch commands. MOSS coverage here is Local-Transformer-v1.5; Delay, Realtime and Nano retain their separate settings.Validation — H200, vLLM 0.29.0
bbcf6f8bc): 370 config tests passed, 3 skipped, all changed-file hooks passed, including mypy. No new performance claim or C64 rerun.b0ecf406a): 365 config tests passed, 3 skipped; all changed-file pre-commit hooks, including mypy, pass. This does not clear the broader historical mypy debt below.vllm serve --omni+vllm bench serve --omni: Qwen3-TTS Base V1 and V2, Qwen3-Omni V2, and MOSS local V2 each completed 16/16 requests with nonempty audio (8 at C1, 8 at C4; no MPS).VLLM_OMNI_DRAIN_READY_GENERATION=1also completed 16/16 requests with nonempty audio; the server log confirms activation. The final remote source matches all 71/71 changed local files by SHA-256.Missing hidden_states['last']failure. An instrumented run, a clean uninstrumented retry, and the final clean V1 run each passed 16 requests. The initial failure has no conclusive root cause; these retries do not establish its elimination.mypy-3.10; this is not a claim of fully passing pre-commit.Not ready to merge: the current codec-EOS failure, historical type/DCO gates, full CI, human review, and matched quality/performance qualification remain. These C1/C4 functional checks do not establish ASR/SIM equivalence or C64/C128 throughput. Prior H20/H200 benchmark numbers describe older trees and are deliberately not advertised as this update's measured performance.
Performance evidence: historical, not current-head claims
Evidence report and machine-readable records retain configurations, per-run aggregates and provenance hashes.
These full SeedTTS measurements used vLLM 0.29 and repeated measured rounds after full warmup. Omni sampled ASR WER was 2.15% versus 3.08%; quality equivalence was not established. Historical numbers are not substituted for measurements of this head.
Current C64 verification paused on a failure
At source
b0ecf406a, the first full SeedTTS EN1088/C64 warmup completed 1087/1088 requests. Request index48 reached the 1024 codec-token budget without EOS; the server raisedQwen3TTSCodecLimitErrorand the client saw an incomplete streaming transfer. The cause is unresolved. The automation stopped before the remaining warmups, formal rounds or historical-configuration remeasurement. Its 74.68 audio-s/s is a failed warmup, not a valid performance result. The earlier C16/64-request screening is retained only as a functional check and is not used to claim steady-state throughput.Current CustomVoice and VoiceDesign MRv2 smoke checks each completed 16/16 nonempty audio using
seed-tts-text(8 at C1, 8 at C4). An initial CustomVoice test used the wrong Base-task dataset and returned400; that attempt is preserved and excluded from pass counts.AI assistance
Codex assisted with implementation, experiment consolidation, tests, and this description. Contributor human review is still required before merge; this text does not attest that human review is complete.