Repository navigation
Conversation
|
This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md. Module owners: @gcanlin @Sy0307 @tzhouam Routing: @gcanlin via module of the changed files, CODEOWNERS; @Sy0307 via module of the changed files, CODEOWNERS; @tzhouam via module of the changed files, CODEOWNERS @BeatSeat, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
04152eb to
e4297a5
Compare
amy-why-3459
left a comment
There was a problem hiding this comment.
The caching change is small and removes redundant tokenizer loads, but the cached failure result needs fixing (see inline comment).
Non-blocking clarification: the first cache miss still loads the tokenizer synchronously on the event loop. Please describe this as reducing repeated-load blocking rather than fully preventing event-loop blocking or cold-session handshake timeouts.
Validation: exercised the original loader methods in isolation with a mocked AutoTokenizer. A failed first load was cached as None, so the next call did not retry; clearing the cache allowed recovery. Successful loads were reused and invalid model paths bypassed the loader. I did not independently reproduce the A100 benchmark. No additional code-redundancy concerns found.
Current-head CI: pre-commit, DCO, and Python 3.11/3.12 builds passed; documentation is pending. No Buildkite status was reported.
…reduce event loop blocking During duplex session initialization, MiniCPMO45NativeDuplexServingAdapter synchronously calls _load_native_tokenizer() 3 times per session (6 times in concurrent session tests). MiniCPM-o 4.5 has a 152k-token vocabulary with ~11.4MB tokenizer.json and dynamic code execution. Repeated synchronous loading on the asyncio main thread causes significant event-loop stalls, which leads to handshake timeouts on slow CI runners (vllm-project#6719). Wrap native tokenizer loading with @functools.lru_cache(maxsize=4) keyed on model_path, matching existing patterns in hunyuan_image3 and indextts2. Keep exception handling in the outer wrapper so transient load failures are not cached, allowing subsequent calls to retry. Empirical benchmark shows a 6.5x speedup for 2 concurrent sessions (2.50s -> 0.38s). Signed-off-by: BeatSeat <wendavid552@gmail.com>
e4297a5 to
64a1f3a
Compare
|
Now the session cold start only costs 0.5s, far below 20s. |
|
The caching part is already handled on upstream/main, where the old adapter is replaced by the new duplex plugin. |
|
Closing in favor of #7413. Upstream main has migrated the duplex engine to the new plugin framework, where () implements instance-level tokenizer caching () and offloads via , resolving the event-loop stalls. The old adapter touched in this PR is no longer present. |
Summary
During MiniCPM-o 4.5 duplex session initialization,
MiniCPMO45NativeDuplexServingAdapterinvokes_load_native_tokenizer()3 separate times synchronously on the main asyncio event loop:_native_stage0_stop_token_ids(model_config)_native_scheduler_token_id(model_config)_apply_first_append_context_tokens(...)openbmb/MiniCPM-o-4_5has a 152k-token vocabulary with an 11.4MBtokenizer.jsonand dynamic code execution. Repeated synchronous loading blocks the main asyncio thread (6 times in concurrent 2-session tests), leading to handshake timeouts in CI runners under slower or virtualized disk I/O (#6719).This PR wraps native tokenizer loading with
@functools.lru_cache(maxsize=4)keyed onmodel_path: str, matching existing caching patterns inhunyuan_image3andindextts2.Empirical Benchmark (A100 with
openbmb/MiniCPM-o-4_5)@lru_cache(After)Backward Compatibility & Testing
_load_native_tokenizer(cls, model_config: Any)interface so callers passingSimpleNamespace(model=...)or test mocks continue to work seamlessly.ruff checkandruff format.Fixes #6719