Skip to content

[Perf][PersonaPlex] Serve 16 concurrent realtime duplex sessions on one GPU - #8192

Merged
linyueqian merged 6 commits into
vllm-project:mainfrom
linyueqian:perf/personaplex-multisession-pr
Sep 28, 2026
Merged

linyueqian merged 6 commits into
vllm-project:mainfrom
linyueqian:perf/personaplex-multisession-pr

Conversation

@linyueqian

@linyueqian linyueqian commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Part of RFC #7389. Makes PersonaPlex serve many concurrent realtime sessions on the unified duplex path. On one 141 GB Hopper-class GPU the number of paced sessions that stay realtime goes from 1 to 16.

Builds on #7695 (PersonaPlex on the unified duplex plugin), now merged.

What was wrong (measured on #7695)

sessions client RTF (per session) frame deficit / 375
1 1.01 4
2 1.44 69
4 1.87 ~140
8 stage 0 dies (device-side assert)
  1. Crash when a new session joins a live one. make_omni_output truncated the token hidden states to the number of cached audio-code rows (personaplex_talker.py). A step that mixes a new session's multi-row prefill with a live session's one-row append then indexes the truncated tensor with token-space logits_indices, and stage 0 dies for every session. This reproduces at N=2 whenever the two land in one step.
  2. Sessions do not share a step. Each live session appends one frame per 80 ms tick as a one-token resumable request. With stage 0 async_scheduling: true the next step is scheduled before a session's previous frame is parked, so a session can only join every other step. At N=4 the per-step batch was 1 or 3, never 4.
  3. Per-session serial encode. prepare_append ran a batch-1 streaming Mimi encode per session, followed by a device-to-host copy, inside the runner's per-request loop. Each session also owned a full codec instance, built lazily (weight init on the step loop) on its first append.
  4. Unused depformer work. The duplex path ran all 16 depformer steps; only the first 8 (agent) codebooks are consumed downstream.

Changes (model-local; no runner, scheduler or engine/duplex/ changes)

  • personaplex_talker.py: keep every token row in make_omni_output; add preprocess_batch, which hands every live duplex append of the step to the stage 0 runtime.
  • duplex/stage0.py: one shared streaming encoder with max_sessions rows; each live (session, epoch) leases a row. encode_appends encodes all new appends of a step in one call with active set only for rows that have a new, not yet encoded (epoch, seq), so a chunked first prefill or a retried step never advances a row twice. prepare_append consumes the cached codes (and falls back to a single-row encode on the same encoder). Closing a session resets only its row. In a cancel overlap (old and new epoch of one session in the same step) only the newest epoch is admitted.
  • personaplex_mimi.py / personaplex_temporal.py: a recycled Mimi transformer row restarts at position 0 (_RingKV.reset_row), so it matches a fresh stream exactly instead of carrying its predecessor's absolute RoPE positions.
  • personaplex_depformer.py: num_steps argument; the duplex path runs the 8 vocoded steps. The loop is causal, so the codes are identical to the first 8 of a full run.
  • deploy/personaplex.yaml: stage 0 async_scheduling: false, with the reason in a comment. max_sessions stays 2 (the lifecycle e2e expects the third open to be refused); the comment documents how to raise it.

Test Plan

  • CPU: pytest tests/model_executor/models/personaplex tests/engine/duplex tests/entrypoints/duplex tests/model_executor/common tests/model_executor/models/minicpmo_4_5/duplex tests/engine/test_duplex_import_boundary.py -m "not cuda"
  • GPU: pytest tests/model_executor/models/personaplex -m cuda
  • GPU e2e lifecycle on the shipped deploy: python tests/e2e/online_serving/personaplex_realtime_duplex.py --model <ckpt> --input-wav <wav>
  • GPU load: same driver with --sessions N --load-frames 375, deploy with max_sessions/max_num_seqs raised to 64 and stage 0 gpu_memory_utilization: 0.75

New tests: hidden rows kept for mixed prefill and live rows; depformer num_steps prefix equality (with teacher forcing); one encoder call for several sessions; rows without a new append stay untouched; dedupe across chunked prefill; over-capacity appends left to prepare_append; cancel overlap in both orders; close resets only its own row; recycled Mimi transformer row equals a fresh stream after the ring wrapped.

Test Result

vLLM 0.30.0, torch 2.13, one 141 GB Hopper-class GPU, paced 24 kHz clients, 375 frames per session, same input for every run.

  • CPU suites: 763 passed, 1 skipped. CUDA-marked PersonaPlex tests: 4 passed. ruff (pinned) clean.
  • Lifecycle e2e on the shipped personaplex.yaml: ok (two sessions, third refused with resource_exhausted, slot reuse, audible output).

Load sweep (a pass means every session within the driver's 4-frame deficit bound):

sessions before after frames per stage 0 step
1 RTF 1.01, pass RTF 1.01, pass 1
2 RTF 1.44, fail RTF 1.02, pass 1-2
4 RTF 1.87, fail RTF 1.02, pass mostly 4
8 crash RTF 1.02, pass mostly 8
12 RTF 1.02, pass mostly 12
16 RTF 1.03, pass 16
24 RTF 1.37, fail 24

The N=4..24 rows were measured with a temporary per-step batch counter in the talker (not part of this diff); the final tree was re-run at N=1 and N=2 on the shipped deploy (RTF 1.01 / 1.02, pass).

Past 16 sessions a 24-row step exceeds the 80 ms budget. The next items are CUDA graphs for the depformer (#7734) and the encoder, and batching the stage 1 decoder; they are out of scope here.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md, docs/design/module/entrypoints.md.

Module owners: @alex-jw-brooks @NickCao @tzhouam

Routing: @alex-jw-brooks via module named in the PR description, CODEOWNERS; @NickCao via module named in the PR description; @tzhouam via module of the changed files, semantic router, CODEOWNERS

@linyueqian, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit cad8c1e3dc1d produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

…ws share a step

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
…x path

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
…ed streaming Mimi encoder

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
… share a step; fix recycled-row reset and cancel overlap

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian
linyueqian force-pushed the perf/personaplex-multisession-pr branch from ea8b884 to 9231919 Compare September 27, 2026 03:04
@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 27, 2026
…ded codebooks

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 27, 2026

@Sy0307 Sy0307 left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please narrow the initialization exception handler and cover the full processing order for cancellation overlap, as detailed in the two inline comments.

Comment on lines +194 to +197
try:
state = self._session_state(session_id, epoch)
except RuntimeError:
continue

@Sy0307 Sy0307 Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Catching every RuntimeError also suppresses _shared_codec() initialization failures, including OOM, causing repeated initialization attempts within the same batch. With codec initialization failures injected into the actual runtime on CPU, 16 appends caused 17 attempts before the PR propagated the error, versus one attempt on main; this was fault injection, not an actual GPU OOM. Please catch only a dedicated capacity-exhaustion exception, propagate other initialization errors immediately, and add a failure-path test.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, thanks. I added a PersonaPlexStage0CapacityError, and that's now the only thing encode_appends swallows, so a codec init failure like OOM surfaces right away instead of getting retried for every append in the step. test_codec_init_failure_propagates_after_one_attempt injects a failing codec factory with 16 appends and checks it's called exactly once. Fixed in cad8c1e.

Comment on lines +419 to +420
runtime.encode_appends([old, new] if old_first else [new, old])
restarted = runtime.prepare_append(new, prompt_len=18, request_id="req-e1")

@Sy0307 Sy0307 Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After encode_appends([old, new]), this test only calls prepare_append() for the new epoch, leaving the case where the old request reaches per-request preparation uncovered. In a component-level reproduction, preparing the old request reallocated its stale epoch's state slot and raised a capacity-exhaustion error at capacity 1, although this ordering has not been proven reachable through the real scheduler. Please add a runner-level test covering the full processing order, or demonstrate that stale requests are filtered before preparation and update the test accordingly.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right that the old request can get there. Rather than argue it's unreachable through the scheduler, I made it safe: _session_state now refuses an epoch that's older than one already live for the same session (PersonaPlexStage0StaleEpochError) instead of giving it a new row. The talker hands that request a neutral row, since the engine throws its output away anyway, and record_sample skips it until its finish comes in.

The updated test runs the whole step in all four encode/prepare orders at capacity 1 and 2, including sampling and the late finish. There's also a talker-level test that sends both requests through preprocess_batch and preprocess in the same step. Also in cad8c1e.

…erseded epochs

encode_appends caught every RuntimeError, so a codec init failure (e.g. OOM)
was retried once per append in the step. Capacity exhaustion now raises its
own PersonaPlexStage0CapacityError and that is the only error skipped there.

In a cancel overlap the aborted epoch's request could still reach
prepare_append and re-lease a row for its stale epoch, failing the step at
capacity. _session_state now refuses an epoch older than a live epoch of the
same session (PersonaPlexStage0StaleEpochError); the talker gives that request
a neutral row and record_sample skips it until its finish arrives.

Tests cover both encode orders and both prepare orders at capacity 1 and 2,
the talker path for both requests in one step, and a failing codec factory.

Claude-Session: https://claude.ai/code/session_0148vEXqDMtc7hnR7dSeLusB
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian
linyueqian requested a review from Sy0307 September 27, 2026 20:11
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict under experiment vllm-omni-strict-5050-20260829.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as stale.

@hsliuustc0106 hsliuustc0106 added high priority high priority issue, needs to be done asap enhancement New feature or request labels Sep 28, 2026

@Sy0307 Sy0307 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 28, 2026
@linyueqian
linyueqian enabled auto-merge (squash) September 28, 2026 04:25
@linyueqian
linyueqian merged commit f5bdb20 into vllm-project:main Sep 28, 2026
7 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants