[Feature][310p] support Prefix Mamba Cache on 310p - #9514
Conversation
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for Prefix Mamba Cache on the 310p NPU backend. The changes ensure that Mamba-based models can effectively utilize prefix caching by adjusting block table allocation logic and refining how kernel block sizes are determined during the model runner's input batch initialization. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces support for Prefix Mamba Cache on the 310p NPU by refactoring the may_reinitialize_input_batch method in NPUModelRunner310. The changes ensure that the number of blocks required for Mamba is correctly calculated when prefix caching is enabled and that cp_kv_cache_interleave_size is passed to NPUInputBatch. A new unit test has been added to verify these changes. Feedback includes a recommendation to use self.model_config.max_model_len to avoid stale values when Prefill Context Parallel (PCP) is active and a suggestion to optimize Mamba block allocation by avoiding sequence-length-based allocation when prefix caching is disabled.
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
❌ This pull request cannot be evaluated by MergifyDetailsfiles are inaccessible |
SP-1 A1a.0 investigation revealed что upstream `vllm-ascend` уже merged большинство fixes планировавшихся как наша работа (PyTorch fallback decode/chunk, patch_qwen3_5). PR vllm-project#9514 (merged 2026-05-30) addresses vllm-project#7306 MambaSpec dtype. Decision: rebase jh/main onto upstream/main HEAD (a8658e9) + reduce SP-1 scope to verify-only (skip ABI/stub kernel — defer to SP-3 on real signature basis). Backup tag: pre-rebase-2026-05-30 (732f7f8) для recovery. Refs: - huawei-ascend/docs/investigations/2026-05-30-sp1-scope-shift.md - PRs vllm-project#7109, vllm-project#7398, vllm-project#7306, vllm-project#9514
SP-1 A1a.0 investigation revealed что upstream `vllm-ascend` уже merged большинство fixes планировавшихся как наша работа (PyTorch fallback decode/chunk, patch_qwen3_5). PR vllm-project#9514 (merged 2026-05-30) addresses vllm-project#7306 MambaSpec dtype. Decision: rebase jh/main onto upstream/main HEAD (a8658e9) + reduce SP-1 scope to verify-only (skip ABI/stub kernel — defer to SP-3 on real signature basis). Backup tag: pre-rebase-2026-05-30 (732f7f8) для recovery. Refs: - huawei-ascend/docs/investigations/2026-05-30-sp1-scope-shift.md - PRs vllm-project#7109, vllm-project#7398, vllm-project#7306, vllm-project#9514
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: liyishi <1252651434@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: yilunh <hanyilun1@huawei.com>
|
trace RFC: #9711 |
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: shenqiangqiang <2416602906@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
### What this PR does / why we need it?
- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.
### Does this PR introduce _any_ user-facing change?
NA
### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL │ total │ 1965.7806 ms │ 1955.5967 ms │ 1975.9645 ms │ 1965.7806 ms │ 1970.8726 ms │ 1973.9277 ms │ 1975.7609 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT │ total │ 1835.7014 ms │ 1825.5388 ms │ 1845.864 ms │ 1835.7014 ms │ 1840.7827 ms │ 1843.8315 ms │ 1845.6607 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT │ total │ 32.5198 ms │ 32.5145 ms │ 32.5251 ms │ 32.5198 ms │ 32.5225 ms │ 32.5241 ms │ 32.525 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL │ total │ 21.5905 ms │ 0.0252 ms │ 35.3068 ms │ 29.8905 ms │ 34.5709 ms │ 35.0401 ms │ 35.2815 ms │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens │ total │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 10250.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens │ total │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 5.0 │ 2 │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput │ total │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │ 2 │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric │ Stage │ Value │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration │ total │ 3931.8488 ms │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests │ total │ 0 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests │ total │ 2 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency │ total │ 0.9999 │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency │ total │ 1 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput │ total │ 0.5087 req/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens │ total │ 20500 │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens │ total │ 10 │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput │ total │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput │ total │ 2.5433 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput │ total │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}
=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id hbm_hit_rate hbm(hit/query) externel_hit_rate externel(hit/query)
-------------------------------------------------------------------------------------------------------
0 78.67% 16128/20500 0% 0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests
- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
What this PR does / why we need it?
Does this PR introduce any user-facing change?
NA
How was this patch tested?
CI and local tests