Skip to content

[Performance][MiniCPM-o] Cache native tokenizer in duplex adapter to reduce event loop blocking (#6719) - #7524

Closed
BeatSeat wants to merge 2 commits into
vllm-project:mainfrom
BeatSeat:perf/minicpmo45-cache-native-tokenizer-6719
Closed

BeatSeat wants to merge 2 commits into
vllm-project:mainfrom
BeatSeat:perf/minicpmo45-cache-native-tokenizer-6719

Conversation

@BeatSeat

Copy link
Copy Markdown
Contributor

Summary

During MiniCPM-o 4.5 duplex session initialization, MiniCPMO45NativeDuplexServingAdapter invokes _load_native_tokenizer() 3 separate times synchronously on the main asyncio event loop:

  1. _native_stage0_stop_token_ids(model_config)
  2. _native_scheduler_token_id(model_config)
  3. _apply_first_append_context_tokens(...)

openbmb/MiniCPM-o-4_5 has a 152k-token vocabulary with an 11.4MB tokenizer.json and dynamic code execution. Repeated synchronous loading blocks the main asyncio thread (6 times in concurrent 2-session tests), leading to handshake timeouts in CI runners under slower or virtualized disk I/O (#6719).

This PR wraps native tokenizer loading with @functools.lru_cache(maxsize=4) keyed on model_path: str, matching existing caching patterns in hunyuan_image3 and indextts2.

Empirical Benchmark (A100 with openbmb/MiniCPM-o-4_5)

Scenario Uncached (Before) Cached with @lru_cache (After) Time Saved
1 Session (3 calls) 1.235 s 0.391 s 0.844 s (-68.4%)
2 Sessions / CI Test (6 calls) 2.495 s 0.385 s 2.111 s (-84.6%)
Warm Session (3 calls) 1.250 s 0.019 ms 1.250 s (~100%)

Backward Compatibility & Testing

  • Preserves the _load_native_tokenizer(cls, model_config: Any) interface so callers passing SimpleNamespace(model=...) or test mocks continue to work seamlessly.
  • Linted with ruff check and ruff format.

Fixes #6719

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md.

Module owners: @gcanlin @Sy0307 @tzhouam

Routing: @gcanlin via module of the changed files, CODEOWNERS; @Sy0307 via module of the changed files, CODEOWNERS; @tzhouam via module of the changed files, CODEOWNERS

@BeatSeat, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@BeatSeat
BeatSeat force-pushed the perf/minicpmo45-cache-native-tokenizer-6719 branch from 04152eb to e4297a5 Compare September 14, 2026 13:44

@amy-why-3459 amy-why-3459 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The caching change is small and removes redundant tokenizer loads, but the cached failure result needs fixing (see inline comment).

Non-blocking clarification: the first cache miss still loads the tokenizer synchronously on the event loop. Please describe this as reducing repeated-load blocking rather than fully preventing event-loop blocking or cold-session handshake timeouts.

Validation: exercised the original loader methods in isolation with a mocked AutoTokenizer. A failed first load was cached as None, so the next call did not retry; clearing the cache allowed recovery. Successful loads were reused and invalid model paths bypassed the loader. I did not independently reproduce the A100 benchmark. No additional code-redundancy concerns found.

Current-head CI: pre-commit, DCO, and Python 3.11/3.12 builds passed; documentation is pending. No Buildkite status was reported.

Comment thread vllm_omni/model_executor/models/minicpmo_4_5/duplex/adapter.py
…reduce event loop blocking

During duplex session initialization, MiniCPMO45NativeDuplexServingAdapter
synchronously calls _load_native_tokenizer() 3 times per session (6 times
in concurrent session tests). MiniCPM-o 4.5 has a 152k-token vocabulary
with ~11.4MB tokenizer.json and dynamic code execution.

Repeated synchronous loading on the asyncio main thread causes significant
event-loop stalls, which leads to handshake timeouts on slow CI runners
(vllm-project#6719).

Wrap native tokenizer loading with @functools.lru_cache(maxsize=4) keyed
on model_path, matching existing patterns in hunyuan_image3 and indextts2.
Keep exception handling in the outer wrapper so transient load failures
are not cached, allowing subsequent calls to retry. Empirical benchmark
shows a 6.5x speedup for 2 concurrent sessions (2.50s -> 0.38s).

Signed-off-by: BeatSeat <wendavid552@gmail.com>
@BeatSeat
BeatSeat force-pushed the perf/minicpmo45-cache-native-tokenizer-6719 branch from e4297a5 to 64a1f3a Compare September 14, 2026 14:02
@BeatSeat BeatSeat changed the title [Performance][MiniCPM-o] Cache native tokenizer in duplex adapter to prevent event loop blocking (#6719) [Performance][MiniCPM-o] Cache native tokenizer in duplex adapter to reduce event loop blocking (#6719) Sep 14, 2026
@amy-why-3459 amy-why-3459 added the ready label to trigger buildkite CI label Sep 15, 2026
@BeatSeat

Copy link
Copy Markdown
Contributor Author

Now the session cold start only costs 0.5s, far below 20s.

@NickCao

NickCao commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

The caching part is already handled on upstream/main, where the old adapter is replaced by the new duplex plugin.

@BeatSeat

Copy link
Copy Markdown
Contributor Author

Closing in favor of #7413. Upstream main has migrated the duplex engine to the new plugin framework, where () implements instance-level tokenizer caching () and offloads via , resolving the event-loop stalls. The old adapter touched in this PR is no longer present.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

4 participants