Repository navigation
[Feat]Handle-Based Runner-Owned VocoderCUDAGraphManager - #7042
sphinxkkkbc wants to merge 8 commits into
Conversation
b76b591 to
7cc6430
Compare
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
7cc6430 to
0a7bd62
Compare
…s; 3. Remove redundant scalar config validation; 4. Clarify Routine ownership without hard restrictions Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
b077ee8 to
fadafee
Compare
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
fadafee to
1735f71
Compare
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to be related to model: qwen-tts. Model owners: @FayeSpica @sphinxkkkbc, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
776f763 to
688e32a
Compare
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
688e32a to
1e536c7
Compare
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
2d4b7df to
5d237b8
Compare
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
5d237b8 to
12da164
Compare
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
will open new PR to supersede this |
Purpose
This PR introduces runner-owned CUDA Graph management for vocoder and Code2Wav models executed by
GPUGenerationModelRunner, and migrates the existingQwen3TTSCode2WavCUDA Graph implementation to the new framework.The ownership boundary is:
This PR refines the design proposed in #4924 with model-owned callable Targets and narrow runtime Handles. Instead of binding the complete Manager to the model, the model declares stable graphable Targets. After capture, the Manager assembles a graph/eager runtime callable and binds an opaque
VocoderGraphHandleto each successfully prepared Target. The Handle exposes only__call__()and does not expose Manager resources or lifecycle operations to model code.For models that do not declare
supports_vocoder_cudagraph,GPUGenerationModelRunnerretains the existing upstream CUDA Graph path unchanged. For models that declare the capability, runner-owned vocoder Targets currently replace upstream root-model CUDA Graph ownership for that stage.File Diff
Happy to split this into two PRs(framework + migration) since the diff is large.
Design Overview
The runtime call path remains model-driven:
The model decides when and in what order Targets are invoked. Within an individual Target, the Manager-assembled runtime callable transparently selects graph replay or eager fallback.
Each Target binds exactly one callable contract:
A model may expose multiple independent Targets. Qwen3-TTS currently exposes:
Target ordering and composition remain ordinary model code. The Manager does not need to understand model topology.
Relationship to #4924
This PR implements the same core lifecycle proposed in #4924:
Capture uses reusable static buffers:
prepare_for_capture()is responsible for making the reusable static buffers ready for an independent capture invocation. No generic post-capture reset or finalize hook is required.The Routine boundary is limited to graph adaptation:
The Routine interface additionally supports:
Key Refinements Beyond #4924
1. Target / Handle Binding
#4924 describes model code calling a named Manager runtime API directly:
This PR instead keeps a stable callable Target in the model:
The Manager captures the Target's Descriptors, assembles its runtime callable, and binds a narrow
VocoderGraphHandleback to the same Target.This introduces an explicit two-phase handoff:
Model code does not hold or inspect the Manager.
The Target remains an ordinary callable:
2. Multiple Graphable Targets
Capture planning is Target-scoped rather than model-wide.
Each Target owns:
This allows one model forward to compose independently managed graphable call sites without requiring the Manager to understand model topology.
For example:
3. Generation-Stage Configuration
This PR adds a stage-local
vocoder_cudagraphconfiguration namespace forLLM_GENERATIONstages.Upstream graph settings remain the global source of truth for whether CUDA Graphs are enabled. The new namespace expresses vocoder-specific policy that cannot be represented safely by one generic
cudagraph_capture_sizeslist.Configuration is divided into:
Example:
The Manager validates and interprets framework-owned fields. Models and Targets declare their supported extension keys and own their value semantics and Descriptor construction. Unknown keys and Target IDs fail during startup.
The
targetsmapping is an override container rather than an allow-list: declared Targets remain enabled by default unlessenabled: falseis configured.Chunk-delivery semantics such as
async_chunk,codec_chunk_frames, andinitial_codec_chunk_framesremain in the existing pipeline and connector configuration and are not duplicated undervocoder_cudagraph.4. Explicit Runtime Lifecycle Semantics
The implementation defines:
A graph coverage miss safely calls the Target's eager runnable.
Once graph input preparation or replay has started, errors are recorded and propagated rather than automatically retrying eager execution, because graph/static state or the CUDA stream may already have been modified.
Qwen3-TTS Migration
This PR removes the model-owned segmented CUDA Graph wrapper and migrates its behavior into runner-owned Targets and model-specific Routines.
The migration covers:
Stateful Suffix Transitions
Stateful suffix transition validity is derived from the connector-resolved transition table shared by model routing and graph Descriptor construction.
Conceptually:
If no valid transition exists, the request remains on eager execution rather than inventing a graph specialization outside the configured state domain.
The transition table also defines the logical source and target extents independently of any physical graph buffer capacity.
Eager / Graph State Semantics
Graph-specific output lifetime is handled by the graph framework.
Stateful Qwen3-TTS Targets use:
After graph replay, the Manager materializes replay results independently from GraphEntry-owned static output storage before returning them to the model.
This keeps eager and graph suffix state handling aligned:
The model therefore keeps ordinary request-state semantics:
Graph-specific request-side ping-pong or persistent backing-buffer ownership is not required.
Graph-owned static buffers remain internal to the Routine / GraphEntry lifecycle, while semantic request state remains model-owned.
Capture Coverage
Startup Descriptors are derived from the resolved stage, connector, and decoder configuration.
With the default configuration, stateful CUDA Graph capture uses batch size 1. Larger runtime batches remain valid and fall back to the existing batched eager implementation unless larger capture batch buckets are configured.
Sequence Diagram
sequenceDiagram autonumber participant Runner as GPUGenerationModelRunner participant Manager as VocoderCUDAGraphManager participant Model as Vocoder Model participant Routine as Target Routine participant Graph as GraphEntry participant Runtime as Runtime Callable participant Handle as VocoderGraphHandle participant Target as Model Target participant Stats as StatsSink Runner->>Model: load model Runner->>Manager: create(vllm_config, device) Manager->>Manager: parse framework options Manager->>Model: get_vocoder_cudagraph_targets() Model-->>Manager: resolved Targets with owned Descriptors Manager->>Manager: validate/filter resolved Targets Note over Manager,Target: Targets still delegate to eager callables loop each selected Target and Descriptor Manager->>Routine: allocate_buffers(descriptor) Routine-->>Manager: static buffers loop warmups Manager->>Routine: prepare_for_capture(buffers) Manager->>Routine: forward_for_capture(buffers) end Manager->>Routine: prepare_for_capture(buffers) Manager->>Graph: capture forward_for_capture(buffers) Manager->>Graph: retain graph, buffers, and captured output end Note over Manager,Target: Capture completes before Handles are bound loop each prepared Target Manager->>Stats: recorder_for(target_id) Stats-->>Manager: write-only Recorder Manager->>Runtime: assemble Routine hooks, Entries, policy, Recorder Manager->>Handle: wrap runtime callable alt assembly/bind succeeds Manager->>Target: bind Handle as delegate Manager->>Manager: retain ManagedTarget else assembly/bind fails Manager->>Target: leave/restore eager delegate Manager->>Manager: release this Target's prepared entries end end Runner->>Model: runtime forward Model->>Target: target(runtime inputs) Target->>Handle: __call__(runtime inputs) Handle->>Runtime: invoke opaque callable Runtime->>Routine: validate_runtime_inputs Note over Runtime,Routine: validation failures propagate alt matching GraphEntry Runtime->>Routine: copy_runtime_inputs Runtime->>Graph: replay Runtime->>Routine: output_after_replay Runtime->>Runtime: clone output if configured Runtime->>Stats: record_graph_hit Runtime-->>Model: ordinary materialized result Model->>Model: commit semantic request state else no matching GraphEntry Runtime->>Stats: record_fallback Runtime->>Routine: eager_call Runtime-->>Model: ordinary eager result Model->>Model: commit semantic request state end Runner->>Manager: shutdown / clear Manager->>Target: restore eager delegates Manager->>Manager: release Entries and stats stateCurrent Scope and Limitations
LLM_GENERATIONvocoder/Code2Wav stages.LLM_AR, diffusion, and upstream attention/KV-cache-aware CUDA Graph dispatch are outside the scope of this PR.Test Plan
Unit Test:
Qwen3-TTS Perf and Acc Test (
async_chunkenabled):vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-Base \ --deploy-config vllm-omni/vllm_omni/deploy/qwen3_tts.yaml \ --omni --port 8091 python3 vllm-omni/benchmarks/tts/bench_tts.py \ --host 127.0.0.1 \ --port 8091 \ --model Qwen/Qwen3-TTS-12Hz-1.7B-Base \ --locale en \ --dataset-path /path/to/seedtts_testset \ --num-prompts 32 \ --concurrency 16 \ --wer-eval \ --output-dir vllm-omni/results/qwen3-tts-12hz-1.7b-baseBenchmark Environment
vLLM Version:
0.28.0
vLLM-Omni Commit:
5d237b8bc9f6829b5cb0e907da4982be4ad4e044The performance and accuracy results below were collected from this benchmark revision.
Test Result
241 passed, 1 skipped
Result Summary
The migrated implementation shows no performance regression for this workload. Request throughput, E2E latency, RTF, and TTFP are comparable to or slightly better than
main.All 32 requests completed successfully in both runs. Mean WER remained below 1% in both cases, indicating no observed quality regression. The WER difference on this 32-sample benchmark should not be interpreted as a quality improvement caused by the CUDA Graph lifecycle refactor.
BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)