Repository navigation
[Perf][PersonaPlex] Serve 16 concurrent realtime duplex sessions on one GPU - #8192
Conversation
|
This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md, docs/design/module/entrypoints.md. Module owners: @alex-jw-brooks @NickCao @tzhouam Routing: @alex-jw-brooks via module named in the PR description, CODEOWNERS; @NickCao via module named in the PR description; @tzhouam via module of the changed files, semantic router, CODEOWNERS @linyueqian, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
…ws share a step Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
…x path Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
…ed streaming Mimi encoder Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
… share a step; fix recycled-row reset and cancel overlap Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
ea8b884 to
9231919
Compare
…ded codebooks Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
| try: | ||
| state = self._session_state(session_id, epoch) | ||
| except RuntimeError: | ||
| continue |
There was a problem hiding this comment.
Catching every RuntimeError also suppresses _shared_codec() initialization failures, including OOM, causing repeated initialization attempts within the same batch. With codec initialization failures injected into the actual runtime on CPU, 16 appends caused 17 attempts before the PR propagated the error, versus one attempt on main; this was fault injection, not an actual GPU OOM. Please catch only a dedicated capacity-exhaustion exception, propagate other initialization errors immediately, and add a failure-path test.
There was a problem hiding this comment.
Good catch, thanks. I added a PersonaPlexStage0CapacityError, and that's now the only thing encode_appends swallows, so a codec init failure like OOM surfaces right away instead of getting retried for every append in the step. test_codec_init_failure_propagates_after_one_attempt injects a failing codec factory with 16 appends and checks it's called exactly once. Fixed in cad8c1e.
| runtime.encode_appends([old, new] if old_first else [new, old]) | ||
| restarted = runtime.prepare_append(new, prompt_len=18, request_id="req-e1") |
There was a problem hiding this comment.
After encode_appends([old, new]), this test only calls prepare_append() for the new epoch, leaving the case where the old request reaches per-request preparation uncovered. In a component-level reproduction, preparing the old request reallocated its stale epoch's state slot and raised a capacity-exhaustion error at capacity 1, although this ordering has not been proven reachable through the real scheduler. Please add a runner-level test covering the full processing order, or demonstrate that stale requests are filtered before preparation and update the test accordingly.
There was a problem hiding this comment.
You're right that the old request can get there. Rather than argue it's unreachable through the scheduler, I made it safe: _session_state now refuses an epoch that's older than one already live for the same session (PersonaPlexStage0StaleEpochError) instead of giving it a new row. The talker hands that request a neutral row, since the engine throws its output away anyway, and record_sample skips it until its finish comes in.
The updated test runs the whole step in all four encode/prepare orders at capacity 1 and 2, including sampling and the late finish. There's also a talker-level test that sends both requests through preprocess_batch and preprocess in the same step. Also in cad8c1e.
…erseded epochs encode_appends caught every RuntimeError, so a codec init failure (e.g. OOM) was retried once per append in the step. Capacity exhaustion now raises its own PersonaPlexStage0CapacityError and that is the only error skipped there. In a cancel overlap the aborted epoch's request could still reach prepare_append and re-lease a row for its stale epoch, failing the step at capacity. _session_state now refuses an epoch older than a live epoch of the same session (PersonaPlexStage0StaleEpochError); the talker gives that request a neutral row and record_sample skips it until its finish arrives. Tests cover both encode orders and both prepare orders at capacity 1 and 2, the talker path for both requests in one step, and a failing codec factory. Claude-Session: https://claude.ai/code/session_0148vEXqDMtc7hnR7dSeLusB Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Omni ReviewBot routing recordAssigned Strict under experiment |
Omni ReviewBot attempt recordReview attempt ended as stale. |
Purpose
Part of RFC #7389. Makes PersonaPlex serve many concurrent realtime sessions on the unified duplex path. On one 141 GB Hopper-class GPU the number of paced sessions that stay realtime goes from 1 to 16.
Builds on #7695 (PersonaPlex on the unified duplex plugin), now merged.
What was wrong (measured on #7695)
make_omni_outputtruncated the token hidden states to the number of cached audio-code rows (personaplex_talker.py). A step that mixes a new session's multi-row prefill with a live session's one-row append then indexes the truncated tensor with token-spacelogits_indices, and stage 0 dies for every session. This reproduces at N=2 whenever the two land in one step.async_scheduling: truethe next step is scheduled before a session's previous frame is parked, so a session can only join every other step. At N=4 the per-step batch was 1 or 3, never 4.prepare_appendran a batch-1 streaming Mimi encode per session, followed by a device-to-host copy, inside the runner's per-request loop. Each session also owned a full codec instance, built lazily (weight init on the step loop) on its first append.Changes (model-local; no runner, scheduler or
engine/duplex/changes)personaplex_talker.py: keep every token row inmake_omni_output; addpreprocess_batch, which hands every live duplex append of the step to the stage 0 runtime.duplex/stage0.py: one shared streaming encoder withmax_sessionsrows; each live(session, epoch)leases a row.encode_appendsencodes all new appends of a step in one call withactiveset only for rows that have a new, not yet encoded(epoch, seq), so a chunked first prefill or a retried step never advances a row twice.prepare_appendconsumes the cached codes (and falls back to a single-row encode on the same encoder). Closing a session resets only its row. In a cancel overlap (old and new epoch of one session in the same step) only the newest epoch is admitted.personaplex_mimi.py/personaplex_temporal.py: a recycled Mimi transformer row restarts at position 0 (_RingKV.reset_row), so it matches a fresh stream exactly instead of carrying its predecessor's absolute RoPE positions.personaplex_depformer.py:num_stepsargument; the duplex path runs the 8 vocoded steps. The loop is causal, so the codes are identical to the first 8 of a full run.deploy/personaplex.yaml: stage 0async_scheduling: false, with the reason in a comment.max_sessionsstays 2 (the lifecycle e2e expects the third open to be refused); the comment documents how to raise it.Test Plan
pytest tests/model_executor/models/personaplex tests/engine/duplex tests/entrypoints/duplex tests/model_executor/common tests/model_executor/models/minicpmo_4_5/duplex tests/engine/test_duplex_import_boundary.py -m "not cuda"pytest tests/model_executor/models/personaplex -m cudapython tests/e2e/online_serving/personaplex_realtime_duplex.py --model <ckpt> --input-wav <wav>--sessions N --load-frames 375, deploy withmax_sessions/max_num_seqsraised to 64 and stage 0gpu_memory_utilization: 0.75New tests: hidden rows kept for mixed prefill and live rows; depformer
num_stepsprefix equality (with teacher forcing); one encoder call for several sessions; rows without a new append stay untouched; dedupe across chunked prefill; over-capacity appends left toprepare_append; cancel overlap in both orders; close resets only its own row; recycled Mimi transformer row equals a fresh stream after the ring wrapped.Test Result
vLLM 0.30.0, torch 2.13, one 141 GB Hopper-class GPU, paced 24 kHz clients, 375 frames per session, same input for every run.
ruff(pinned) clean.personaplex.yaml: ok (two sessions, third refused withresource_exhausted, slot reuse, audible output).Load sweep (a pass means every session within the driver's 4-frame deficit bound):
The N=4..24 rows were measured with a temporary per-step batch counter in the talker (not part of this diff); the final tree was re-run at N=1 and N=2 on the shipped deploy (RTF 1.01 / 1.02, pass).
Past 16 sessions a 24-row step exceeds the 80 ms budget. The next items are CUDA graphs for the depformer (#7734) and the encoder, and batching the stage 1 decoder; they are out of scope here.