Skip to content

[Feat]Handle-Based Runner-Owned VocoderCUDAGraphManager - #7042

Closed
sphinxkkkbc wants to merge 8 commits into
vllm-project:mainfrom
sphinxkkkbc:feat/runner_owned_graph_lifecycle
Closed

sphinxkkkbc wants to merge 8 commits into
vllm-project:mainfrom
sphinxkkkbc:feat/runner_owned_graph_lifecycle

Conversation

@sphinxkkkbc

@sphinxkkkbc sphinxkkkbc commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Related issues: #4924, #4571

This PR introduces runner-owned CUDA Graph management for vocoder and Code2Wav models executed by GPUGenerationModelRunner, and migrates the existing Qwen3TTSCode2Wav CUDA Graph implementation to the new framework.

The ownership boundary is:

  • the runner-owned Manager controls CUDA Graph policy, capture lifecycle, graph-pool usage, memory accounting, runtime coverage statistics, and cleanup;
  • model-specific Routines define graph shapes, static buffers, capture preparation, runtime input copies, and replay-output materialization;
  • model code remains authoritative for request routing, Target composition, and semantic state transitions, including request-cache updates around graphable call sites.

This PR refines the design proposed in #4924 with model-owned callable Targets and narrow runtime Handles. Instead of binding the complete Manager to the model, the model declares stable graphable Targets. After capture, the Manager assembles a graph/eager runtime callable and binds an opaque VocoderGraphHandle to each successfully prepared Target. The Handle exposes only __call__() and does not expose Manager resources or lifecycle operations to model code.

For models that do not declare supports_vocoder_cudagraph, GPUGenerationModelRunner retains the existing upstream CUDA Graph path unchanged. For models that declare the capability, runner-owned vocoder Targets currently replace upstream root-model CUDA Graph ownership for that stage.

File Diff

Area Additions Deletions
Generic framework production +924 -2
Generic framework tests +839 0
Framework total +1763 -2
Migration production +1865 -1525
Migration tests +1119 -580
Migration total +2984 -2105

Happy to split this into two PRs(framework + migration) since the diff is large.

Design Overview

The runtime call path remains model-driven:

GPUGenerationModelRunner
  -> _model_forward()
  -> model.forward()
  -> model-owned Target
  -> bound Handle
  -> graph replay or the Target's eager callable

The model decides when and in what order Targets are invoked. Within an individual Target, the Manager-assembled runtime callable transparently selects graph replay or eager fallback.

Each Target binds exactly one callable contract:

Target
  ├─ stable target_id
  ├─ one model-specific Routine / eager runnable
  ├─ immutable startup capture Descriptors
  └─ one delegate
       ├─ before capture: Routine.eager_call
       └─ after capture: VocoderGraphHandle

A model may expose multiple independent Targets. Qwen3-TTS currently exposes:

qwen3_tts.icl_prefix
qwen3_tts.xvec_prefix
qwen3_tts.suffix
qwen3_tts.stateless

Target ordering and composition remain ordinary model code. The Manager does not need to understand model topology.

Relationship to #4924

This PR implements the same core lifecycle proposed in #4924:

Model declares graphable Targets/Routines
  -> Runner discovers them
  -> Manager resolves capture policy
  -> Manager allocates static buffers
  -> Manager performs warmup and capture
  -> Manager owns GraphEntries and memory accounting
  -> Runtime resolves a model-specific Descriptor
  -> Graph hit replays the captured entry
  -> Coverage miss preserves eager behavior

Capture uses reusable static buffers:

allocate buffers

warmup:
  prepare_for_capture
  forward_for_capture

final capture:
  prepare_for_capture
  forward_for_capture

prepare_for_capture() is responsible for making the reusable static buffers ready for an independent capture invocation. No generic post-capture reset or finalize hook is required.

The Routine boundary is limited to graph adaptation:

Routine
  -> graph-shape resolution
  -> static-buffer allocation/preparation
  -> runtime input copies
  -> replay-output materialization

Model
  -> request routing
  -> semantic state transitions
  -> request-cache commit

The Routine interface additionally supports:

  • Target-scoped runtime key and Descriptor resolution;
  • runtime input validation;
  • optional lazy Descriptor construction;
  • output lifetime and cloning semantics;
  • structured graph-hit, fallback, and replay-error recording.

Key Refinements Beyond #4924

1. Target / Handle Binding

#4924 describes model code calling a named Manager runtime API directly:

manager.replay_or_none("routine_name", runtime_inputs)

This PR instead keeps a stable callable Target in the model:

self._suffix_target(...)

The Manager captures the Target's Descriptors, assembles its runtime callable, and binds a narrow VocoderGraphHandle back to the same Target.

This introduces an explicit two-phase handoff:

Model -> Target / Routine / Descriptors -> Manager
Model <- bound callable Handle          <- Manager

Model code does not hold or inspect the Manager.

The Target remains an ordinary callable:

before capture
  -> Routine.eager_call

after successful capture/binding
  -> VocoderGraphHandle

after Manager cleanup
  -> Routine.eager_call

2. Multiple Graphable Targets

Capture planning is Target-scoped rather than model-wide.

Each Target owns:

one Routine
one eager callable contract
one Descriptor namespace

This allows one model forward to compose independently managed graphable call sites without requiring the Manager to understand model topology.

For example:

Qwen3-TTS
  ├─ ICL prefix Target
  ├─ x-vector prefix Target
  ├─ suffix Target
  └─ stateless decode Target

3. Generation-Stage Configuration

This PR adds a stage-local vocoder_cudagraph configuration namespace for LLM_GENERATION stages.

Upstream graph settings remain the global source of truth for whether CUDA Graphs are enabled. The new namespace expresses vocoder-specific policy that cannot be represented safely by one generic cudagraph_capture_sizes list.

Configuration is divided into:

  1. Manager-owned policy, including graph memory limits, runtime statistics, per-Target enablement, lazy capture, and graph capacity;
  2. model-shared configuration, interpreted by the model and potentially affecting multiple Targets, such as Qwen3-TTS capture batch sizes;
  3. Target-specific configuration, interpreted by an individual Target and affecting its Descriptor construction.

Example:

vocoder_cudagraph:
  # Manager-owned policy.
  max_memory_bytes: 2147483648
  log_stats: true

  # Qwen3-TTS shared configuration.
  capture_batch_sizes: [1, 2, 4]
  decode_batch_max_size: 4

  targets:
    qwen3_tts.suffix:
      # Manager-owned per-Target policy.
      enabled: true
      enable_lazy_capture: true
      max_extra_graphs: 4

    qwen3_tts.stateless:
      enabled: true

      # Stateless-Target-specific configuration.
      capture_bucket_sizes: [150, 325]

The Manager validates and interprets framework-owned fields. Models and Targets declare their supported extension keys and own their value semantics and Descriptor construction. Unknown keys and Target IDs fail during startup.

The targets mapping is an override container rather than an allow-list: declared Targets remain enabled by default unless enabled: false is configured.

Chunk-delivery semantics such as async_chunk, codec_chunk_frames, and initial_codec_chunk_frames remain in the existing pipeline and connector configuration and are not duplicated under vocoder_cudagraph.

4. Explicit Runtime Lifecycle Semantics

The implementation defines:

  • capture-before-bind ordering;
  • Target-local setup failure isolation;
  • reversible binding during cleanup;
  • failed-Descriptor suppression;
  • Manager-owned lazy-capture serialization;
  • graph replay locking for mutable static buffers;
  • graph coverage miss versus replay-error semantics;
  • runtime graph-coverage statistics.

A graph coverage miss safely calls the Target's eager runnable.

Once graph input preparation or replay has started, errors are recorded and propagated rather than automatically retrying eager execution, because graph/static state or the CUDA stream may already have been modified.

Qwen3-TTS Migration

This PR removes the model-owned segmented CUDA Graph wrapper and migrates its behavior into runner-owned Targets and model-specific Routines.

The migration covers:

  • stateless Code2Wav decoding;
  • ICL prefix decoding;
  • x-vector prefix decoding;
  • stateful suffix transitions;
  • graph batch and frame padding;
  • request-local decoder-cache updates;
  • connector-derived async chunk transitions;
  • eager fallback for uncovered runtime shapes.

Stateful Suffix Transitions

Stateful suffix transition validity is derived from the connector-resolved transition table shared by model routing and graph Descriptor construction.

Conceptually:

previous suffix state
  -> transition table
  -> valid target state
  -> matching suffix Descriptor

If no valid transition exists, the request remains on eager execution rather than inventing a graph specialization outside the configured state domain.

The transition table also defines the logical source and target extents independently of any physical graph buffer capacity.

Eager / Graph State Semantics

Graph-specific output lifetime is handled by the graph framework.

Stateful Qwen3-TTS Targets use:

clone_output=True

After graph replay, the Manager materializes replay results independently from GraphEntry-owned static output storage before returning them to the model.

This keeps eager and graph suffix state handling aligned:

eager:
  compute next state
  -> model commits next state

graph:
  replay static graph
  -> clone/materialize replay result
  -> model commits next state

The model therefore keeps ordinary request-state semantics:

caches["suffix_quantized"] = next_quantized
caches["suffix_conv"] = next_conv
caches["suffix_frames"] = target_frames

Graph-specific request-side ping-pong or persistent backing-buffer ownership is not required.

Graph-owned static buffers remain internal to the Routine / GraphEntry lifecycle, while semantic request state remains model-owned.

Capture Coverage

Startup Descriptors are derived from the resolved stage, connector, and decoder configuration.

With the default configuration, stateful CUDA Graph capture uses batch size 1. Larger runtime batches remain valid and fall back to the existing batched eager implementation unless larger capture batch buckets are configured.

Sequence Diagram

sequenceDiagram
    autonumber
    participant Runner as GPUGenerationModelRunner
    participant Manager as VocoderCUDAGraphManager
    participant Model as Vocoder Model
    participant Routine as Target Routine
    participant Graph as GraphEntry
    participant Runtime as Runtime Callable
    participant Handle as VocoderGraphHandle
    participant Target as Model Target
    participant Stats as StatsSink

    Runner->>Model: load model
    Runner->>Manager: create(vllm_config, device)
    Manager->>Manager: parse framework options
    Manager->>Model: get_vocoder_cudagraph_targets()
    Model-->>Manager: resolved Targets with owned Descriptors
    Manager->>Manager: validate/filter resolved Targets

    Note over Manager,Target: Targets still delegate to eager callables

    loop each selected Target and Descriptor
        Manager->>Routine: allocate_buffers(descriptor)
        Routine-->>Manager: static buffers

        loop warmups
            Manager->>Routine: prepare_for_capture(buffers)
            Manager->>Routine: forward_for_capture(buffers)
        end

        Manager->>Routine: prepare_for_capture(buffers)
        Manager->>Graph: capture forward_for_capture(buffers)
        Manager->>Graph: retain graph, buffers, and captured output
    end

    Note over Manager,Target: Capture completes before Handles are bound

    loop each prepared Target
        Manager->>Stats: recorder_for(target_id)
        Stats-->>Manager: write-only Recorder
        Manager->>Runtime: assemble Routine hooks, Entries, policy, Recorder
        Manager->>Handle: wrap runtime callable

        alt assembly/bind succeeds
            Manager->>Target: bind Handle as delegate
            Manager->>Manager: retain ManagedTarget
        else assembly/bind fails
            Manager->>Target: leave/restore eager delegate
            Manager->>Manager: release this Target's prepared entries
        end
    end

    Runner->>Model: runtime forward
    Model->>Target: target(runtime inputs)
    Target->>Handle: __call__(runtime inputs)
    Handle->>Runtime: invoke opaque callable
    Runtime->>Routine: validate_runtime_inputs

    Note over Runtime,Routine: validation failures propagate

    alt matching GraphEntry
        Runtime->>Routine: copy_runtime_inputs
        Runtime->>Graph: replay
        Runtime->>Routine: output_after_replay
        Runtime->>Runtime: clone output if configured
        Runtime->>Stats: record_graph_hit
        Runtime-->>Model: ordinary materialized result
        Model->>Model: commit semantic request state
    else no matching GraphEntry
        Runtime->>Stats: record_fallback
        Runtime->>Routine: eager_call
        Runtime-->>Model: ordinary eager result
        Model->>Model: commit semantic request state
    end

    Runner->>Manager: shutdown / clear
    Manager->>Target: restore eager delegates
    Manager->>Manager: release Entries and stats state
Loading

Current Scope and Limitations

  • Runtime coverage statistics are currently emitted only when the Manager is cleared.
  • Runner-owned vocoder CUDA Graphs currently replace the upstream root-model CUDA Graph path for the stage; coexistence is not yet implemented, but is not considered fundamentally incompatible.
  • The initial implementation applies only to LLM_GENERATION vocoder/Code2Wav stages.
  • This PR migrates Qwen3-TTS; other vocoder implementations remain unchanged.
  • Only configured or derived startup Descriptors are pre-captured. Uncovered valid runtime shapes use lazy capture when enabled, otherwise eager fallback.
  • LLM_AR, diffusion, and upstream attention/KV-cache-aware CUDA Graph dispatch are outside the scope of this PR.

Test Plan

Unit Test:

pytest -v \
  tests/config/test_omni_config.py \
  tests/model_executor/models/interfaces/test_vocoder_cudagraph.py \
  tests/model_executor/models/qwen3_tts/test_qwen3_tts_code2wav.py \
  tests/model_executor/models/qwen3_tts/test_qwen3_tts_incremental_decode.py \
  tests/model_executor/models/qwen3_tts/test_vocoder_cudagraph.py \
  tests/worker/test_generation_vocoder_cudagraph_runner.py \
  tests/worker/test_vocoder_cudagraph_manager.py \
  tests/worker/test_vocoder_cudagraph_lazy_capture.py

Qwen3-TTS Perf and Acc Test (async_chunk enabled):

vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-Base \
    --deploy-config vllm-omni/vllm_omni/deploy/qwen3_tts.yaml \
    --omni --port 8091

python3 vllm-omni/benchmarks/tts/bench_tts.py \
  --host 127.0.0.1 \
  --port 8091 \
  --model Qwen/Qwen3-TTS-12Hz-1.7B-Base \
  --locale en \
  --dataset-path /path/to/seedtts_testset \
  --num-prompts 32 \
  --concurrency 16 \
  --wer-eval \
  --output-dir vllm-omni/results/qwen3-tts-12hz-1.7b-base

Benchmark Environment

vLLM Version:
0.28.0

vLLM-Omni Commit:
5d237b8bc9f6829b5cb0e907da4982be4ad4e044

The performance and accuracy results below were collected from this benchmark revision.

Test Result

241 passed, 1 skipped


Metrics main This PR ∆
Request throughput (req/s) 4.99 5.07 +1.60%
Mean E2E latency (ms) 2765.15 2750.44 -0.53%
P99 E2E latency (ms) 4322.88 4217.63 -2.44%
Mean RTF 0.70 0.70 0.00%
Mean TTFP (ms) 864.04 819.93 -5.11%
Mean WER 0.83% 0.26% -68.67%
Successful requests 32 / 32 32 / 32 —

Result Summary

The migrated implementation shows no performance regression for this workload. Request throughput, E2E latency, RTF, and TTFP are comparable to or slightly better than main.

All 32 requests completed successfully in both runs. Mean WER remained below 1% in both cases, indicating no observed quality regression. The WER difference on this 32-sample benchmark should not be interpreted as a quality improvement caused by the CUDA Graph lifecycle refactor.


BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.

(anything written below this line will be removed by GitHub Actions)

@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch 3 times, most recently from b76b591 to 7cc6430 Compare September 4, 2026 13:37
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch from 7cc6430 to 0a7bd62 Compare September 4, 2026 13:50
@hsliuustc0106 hsliuustc0106 added enhancement New feature or request tts code related to tts models labels Sep 5, 2026
…s; 3. Remove redundant scalar config validation; 4. Clarify Routine ownership without hard restrictions

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch 2 times, most recently from b077ee8 to fadafee Compare September 5, 2026 12:20
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: qwen-tts.

Model owners: @FayeSpica

@sphinxkkkbc, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 5, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 12da164093b2 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch from 776f763 to 688e32a Compare September 6, 2026 03:12
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch from 688e32a to 1e536c7 Compare September 6, 2026 09:49
@sphinxkkkbc
sphinxkkkbc marked this pull request as draft September 6, 2026 10:17
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch 2 times, most recently from 2d4b7df to 5d237b8 Compare September 6, 2026 13:26
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/runner_owned_graph_lifecycle branch from 5d237b8 to 12da164 Compare September 6, 2026 13:45
@sphinxkkkbc
sphinxkkkbc marked this pull request as ready for review September 6, 2026 13:46
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@sphinxkkkbc sphinxkkkbc closed this Sep 7, 2026
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

will open new PR to supersede this

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants