Skip to content

[Perf][Qwen3-TTS] Enable event-driven orchestration by default - #7088

Merged
Sy0307 merged 16 commits into
mainfrom
perf/qwen3-tts-event-driven-default
Sep 21, 2026
Merged

Sy0307 merged 16 commits into
mainfrom
perf/qwen3-tts-event-driven-default

Conversation

@Sy0307

@Sy0307 Sy0307 commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Enable the event-driven orchestration loop by default for the qwen3_tts pipeline.

The default remains unchanged for other pipelines. An explicit
VLLM_OMNI_EVENT_DRIVEN_ORCH=0 or 1 always takes precedence, so operators can
still select the legacy polling loop for A/B testing or rollback.

The PR also adds three low-risk diagnostics/fixes:

  • Event Driven runs record both the dispatch queue size at the monitor boundary
    and the per-window high-water mark. The high-water mark is updated when a
    reader enqueues output, so short bursts are not hidden by a dispatcher that
    drains them before the next 1-second sample.
  • Omni AR workers now wire vLLM's native CudaProfilerWrapper when
    profiler_config.profiler="cuda". Before this, /start_profile succeeded but
    /stop_profile failed with Profiling is not enabled, and Nsight Systems
    produced no capture for Omni AR stages.
  • Omni AR/TTS profiling documentation now covers stage-level profiler config,
    warmup, /start_profile//stop_profile, and Nsight Systems capture mode.

Pipeline defaults are kept local to each AsyncOmniEngine/Orchestrator; the
implementation does not mutate os.environ, so initializing Qwen3-TTS cannot
leak the default into another model in the same process.

Motivation

The event-driven loop avoids the legacy 1 ms per-replica polling cadence and
reduces output wakeup delay for streaming TTS. This is especially relevant to
Qwen3-TTS high-concurrency serving, where the Talker emits frequent codec
chunks and Code2Wav consumes them asynchronously.

Related: #4680, #5221

Scope

Only qwen3_tts receives the pipeline-specific default. Other async-chunk
pipelines are intentionally left opt-in because their streaming state machines,
duplex behavior, and output ordering need separate validation.

A parallel dispatcher is not included here. No existing PR was found that
provides a correctness-safe parallel dispatcher for the shared mutable
orchestrator state. Profile evidence shows that current Qwen3-TTS C64
process_outputs and route_output are small compared with the legacy polling
cost. The current event-driven implementation still serializes stateful routing,
preserving request/chunk ordering. A future implementation must keep
same-request ordering and cleanup serialized while parallelizing only
independent requests.

The existing multi-API/lane-sharding work is tracked separately in draft #6923;
it should not be conflated with this model-default change.

Validation

Real C64 A/B

  • H200, Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
  • qwen3_tts_high_concurrency.yaml: stage 0 max_num_seqs=64, stage 1 max_num_seqs=10
  • 512 requests, max concurrency 64, 2 warmups, bundled seed_tts_smoke text prompts
  • Event-driven off: 3 valid runs, all 512/512, request throughput 17.83-18.67 req/s
  • Event-driven on: 1 valid run, 512/512, request throughput 18.61 req/s
  • Valid on-run vs first off-run: mean E2E 3274.91 ms vs 3429.75 ms; P99 E2E 4011.66 ms vs 5392.31 ms
  • C64 Event Driven monitor: dispatch queue high-water 0; stage output queues high-water 0

Host/GPU profile

  • Torch host-stack window completed successfully for Talker and Code2Wav.
  • Nsight Systems window completed successfully after the CUDA profiler wiring fix.
  • Nsight cuda_api_sum: cudaEventSynchronize was 51.3%,
    cudaGraphLaunch 29.7%, and cudaStreamSynchronize 8.4% of CUDA API time.
  • Stage0 had approximately 27.8% GPU busy union in the narrow torch window;
    Code2Wav stage1 had approximately 1.1% GPU busy union and long host waits.
  • This points to stage1 input/batch readiness and connector scheduling as the
    next investigation area, not a proven need for parallel output processing.
  • The trace analyzer also showed stage0 host gaps in engine-step locks and CUDA
    graph replay; these are diagnostic candidates, not production latency claims.

Tests and checks

  • Engine/event-driven/monitor/process tests: 46 passed in the final focused run
  • Worker CUDA profiler regression test: 12 passed
  • Ruff check and format check: passed
  • Python compile check: passed
  • GitHub Build 3.11/3.12: passed
  • GitHub pre-commit: passed
  • GitHub CodeQL: passed
  • GitHub Buildkite omni-release: passed
  • GitHub Read the Docs: passed

The C64 benchmark is a real serving A/B, but the performance delta is based on
few valid on runs and should be treated as directional rather than a stable
guaranteed percentage. A second on-arm restart did not produce a valid result
because stage initialization failed in the shared test environment and was not
counted. The benchmark also exposed an existing streaming continuity issue
under this C64 configuration; that is independent of this default-selection
change.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: qwen-tts.

Model owners: @FayeSpica

@Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 5, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit afc4decaca68 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

return value.strip().lower() in ("1", "true", "yes", "on")


def _event_driven_orch_default_for_pipeline(pipeline_model_type: str | None) -> bool:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

??? why this is only related to qwen3-tts? it seems we can valiate this for other models as well without any changes to the codebase

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Could you extend the validation to other streaming pipelines using VLLM_OMNI_EVENT_DRIVEN_ORCH=0/1? The switch is already generic, so these checks should not require model-code changes. A useful first set would be Qwen3-Omni (qwen3_omni_moe) and one two-stage TTS pipeline such as Fish Speech (fish_qwen3_omni), CosyVoice3, or MOSS-TTS. MiniCPM-o 4.5 would also help cover session/duplex behavior before considering a broader default.

For each tested configuration, please include:

  • Matched off/on runs with the same hardware, model, prompts, concurrency, and warmups; record the commit, commands, and raw results. For Qwen3-TTS, please repeat the on arm: its one valid throughput result currently falls within the off-arm range.
  • Time to first audio, inter-chunk gaps, P99 end-to-end latency, throughput, and completion/error counts at low and high concurrency.
  • Checks for missing/duplicate/reordered chunks, audio continuity, cancellation, and terminal cleanup. Please also provide an off/on reproduction of the reported C64 continuity issue to support its being independent of this change.

This would help establish which pipelines can safely share the default and which should remain opt-in.

@hsliuustc0106 hsliuustc0106 added tts code related to tts models core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request labels Sep 5, 2026
Comment thread docs/configuration/environment_variables.md Outdated
Comment thread vllm_omni/engine/orchestrator.py
@Sy0307 Sy0307 added the ready label to trigger buildkite CI label Sep 15, 2026
@Sy0307
Sy0307 force-pushed the perf/qwen3-tts-event-driven-default branch from 9ae3e2c to 72669b3 Compare September 15, 2026 18:17
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 15, 2026
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 16, 2026
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 16, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving at 98966bb9. The default change is narrow and the precedence is right: _event_driven_orch_enabled now returns the pipeline default only when VLLM_OMNI_EVENT_DRIVEN_ORCH is unset, any explicit value still parses as before, and _event_driven_orch_default_for_pipeline is true for qwen3_tts alone. The default is computed once in OmniEngineBase._set_pipeline_runtime_config, handed to the orchestrator at construction, and read back by the frontend drain through the engine attribute, so the orchestrator loop and the async_omni_base drain cannot disagree for a production engine; the getattr fallback to False only matters for an engine object that never ran that init path, and even then both modes consume the same janus queue with a bounded wait, so the failure would be latency, not a hang. Nothing writes os.environ, so initialising Qwen3-TTS cannot leak the default into another model in the same process, which was the property I most wanted to hold.

The two side changes check out. The dispatch-queue gauge is two integer stores on the orchestrator loop with no lock, the high-water mark is a max rather than a sum, and the post-dequeue update means a peak is not double counted; one small window-rollover artefact is inline as a suggestion. The CudaProfilerWrapper wiring uses the same vllm.profiler.wrapper import the diffusion worker already relies on, and /stop_profile reaches WorkerProfiler.stop through the existing profile(is_start=False) RPC, so the Profiling is not enabled failure described in the PR is closed rather than moved.

Two things before this merges. First, main has moved under vllm_omni/entrypoints/async_omni_base.py since this branch's base (the #7006 abort-path move), so please rebase onto current main and push; the branch carries ready, so a push does not start a build by itself, ping here and I will re-fire the lane. Second, hsliuustc0106's question about validating the switch on other streaming pipelines is still open on the thread; my approval covers the code as scoped (qwen3_tts only, others opt-in), and whether the default should widen is his call to settle with you.

Validation: static read of the worktree against origin/main plus a two-model panel; no PR code executed. The general Buildkite lane is green at this head, and GitHub Actions pre-commit, wheel build and DCO are green.

Comment thread vllm_omni/engine/orchestrator_monitor.py Outdated
@linyueqian

Copy link
Copy Markdown
Collaborator

Adding to the rebase ask above: #6849 merged as f1a6e7ce and this branch now has a content conflict in tests/engine/test_single_stage_mode.py on top of the async_omni_base.py overlap from #7006. When you rebase, note that main now passes typed stage configs through startup, so if the pipeline default you thread through OmniEngineBase._set_pipeline_runtime_config reads anything from the stage configs, it should go through the typed fields (stage_pipeline_config.model_type) with the legacy fallback rather than engine_args. The approval carries over to an unchanged diff, and I will re-fire the lane after the push.

@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 21, 2026
@Sy0307

Sy0307 commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

@linyueqian Addressed the main-sync request in ea7de7e by merging main at 8efd3c2 and resolving the typed-config test imports while preserving main's shared abort path. The pipeline default reads PipelineConfig.model_type directly in OmniEngineBase._set_pipeline_runtime_config; it does not read stage engine_args, so no stage-field fallback is needed here. CI was retriggered after the push: https://buildkite.com/vllm/vllm-omni/builds/15749.

Self-review at ea7de7e: checked the merge resolution, pipeline-default propagation, explicit environment override, and dispatch-queue rollover semantics. On H200 with vLLM 0.29.0, 259 related unit regressions passed, and the same Qwen3-TTS CustomVoice streaming smoke passed separately with VLLM_OMNI_EVENT_DRIVEN_ORCH=1 and =0. These are functional checks, not the cross-model performance/continuity matrix requested above; that request remains open, and the default remains limited to Qwen3-TTS.

@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 21, 2026
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 21, 2026
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 21, 2026
@Sy0307
Sy0307 merged commit 473da2c into main Sep 21, 2026
14 checks passed
y-null pushed a commit to y-null/vllm-omni that referenced this pull request Sep 22, 2026
…project#7088)

Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: y-null <y-null@users.noreply.github.com>
y-null pushed a commit to y-null/vllm-omni that referenced this pull request Sep 22, 2026
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…project#7088)

Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request ready label to trigger buildkite CI tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants