Skip to content

[Feature][310p] support Prefix Mamba Cache on 310p - #9514

Merged
wangxiyuan merged 7 commits into
vllm-project:mainfrom
Tflowers-0129:518-prefix-mamba-cache
May 30, 2026
Merged

wangxiyuan merged 7 commits into
vllm-project:mainfrom
Tflowers-0129:518-prefix-mamba-cache

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented May 25, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

  • Prefix Mamba Cache Support: Updated the NPU model runner to correctly handle and expand block tables when prefix caching is enabled for Mamba models.
  • Refactored Block Size Calculation: Streamlined the calculation of kernel block sizes and maximum blocks per request by iterating directly over KV cache specifications.
  • Configuration Integration: Included cp_kv_cache_interleave_size in the NPUInputBatch initialization to ensure consistent configuration across the 310p backend.

Does this PR introduce any user-facing change?

NA

How was this patch tested?

[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行

CI and local tests

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@Tflowers-0129
Tflowers-0129 requested a review from wangxiyuan as a code owner May 25, 2026 06:44
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for Prefix Mamba Cache on the 310p NPU backend. The changes ensure that Mamba-based models can effectively utilize prefix caching by adjusting block table allocation logic and refining how kernel block sizes are determined during the model runner's input batch initialization.

Highlights

  • Prefix Mamba Cache Support: Updated the NPU model runner to correctly handle and expand block tables when prefix caching is enabled for Mamba models.
  • Refactored Block Size Calculation: Streamlined the calculation of kernel block sizes and maximum blocks per request by iterating directly over KV cache specifications.
  • Configuration Integration: Included cp_kv_cache_interleave_size in the NPUInputBatch initialization to ensure consistent configuration across the 310p backend.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for Prefix Mamba Cache on the 310p NPU by refactoring the may_reinitialize_input_batch method in NPUModelRunner310. The changes ensure that the number of blocks required for Mamba is correctly calculated when prefix caching is enabled and that cp_kv_cache_interleave_size is passed to NPUInputBatch. A new unit test has been added to verify these changes. Feedback includes a recommendation to use self.model_config.max_model_len to avoid stale values when Prefill Context Parallel (PCP) is active and a suggestion to optimize Mamba block allocation by avoiding sequence-length-based allocation when prefix caching is disabled.

Comment thread vllm_ascend/_310p/model_runner_310p.py
Comment thread vllm_ascend/_310p/model_runner_310p.py
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Comment thread vllm_ascend/patch/worker/patch_mamba_utils.py Outdated
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@Tflowers-0129
Tflowers-0129 requested a review from MengqingCao May 28, 2026 11:54
@mergify

mergify Bot commented May 28, 2026

Copy link
Copy Markdown

❌ This pull request cannot be evaluated by Mergify

Details

files are inaccessible

@wangxiyuan
wangxiyuan merged commit a9585c8 into vllm-project:main May 30, 2026
57 checks passed
adeepn added a commit to jethome-iot/vllm-ascend that referenced this pull request May 30, 2026
SP-1 A1a.0 investigation revealed что upstream `vllm-ascend` уже merged большинство
fixes планировавшихся как наша работа (PyTorch fallback decode/chunk, patch_qwen3_5).
PR vllm-project#9514 (merged 2026-05-30) addresses vllm-project#7306 MambaSpec dtype.

Decision: rebase jh/main onto upstream/main HEAD (a8658e9) + reduce SP-1 scope to
verify-only (skip ABI/stub kernel — defer to SP-3 on real signature basis).

Backup tag: pre-rebase-2026-05-30 (732f7f8) для recovery.

Refs:
- huawei-ascend/docs/investigations/2026-05-30-sp1-scope-shift.md
- PRs vllm-project#7109, vllm-project#7398, vllm-project#7306, vllm-project#9514
adeepn added a commit to jethome-iot/vllm-ascend that referenced this pull request May 31, 2026
SP-1 A1a.0 investigation revealed что upstream `vllm-ascend` уже merged большинство
fixes планировавшихся как наша работа (PyTorch fallback decode/chunk, patch_qwen3_5).
PR vllm-project#9514 (merged 2026-05-30) addresses vllm-project#7306 MambaSpec dtype.

Decision: rebase jh/main onto upstream/main HEAD (a8658e9) + reduce SP-1 scope to
verify-only (skip ABI/stub kernel — defer to SP-3 on real signature basis).

Backup tag: pre-rebase-2026-05-30 (732f7f8) для recovery.

Refs:
- huawei-ascend/docs/investigations/2026-05-30-sp1-scope-shift.md
- PRs vllm-project#7109, vllm-project#7398, vllm-project#7306, vllm-project#9514
SOMEONEUNSEEN pushed a commit to SOMEONEUNSEEN/vllm-ascend that referenced this pull request Jun 2, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query)
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests

- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: liyishi <1252651434@qq.com>
yilunh998 pushed a commit to yilunh998/vllm-ascend that referenced this pull request Jun 2, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query)
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests

- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: yilunh <hanyilun1@huawei.com>
@Tflowers-0129

Copy link
Copy Markdown
Collaborator Author

trace RFC: #9711

2416602906 pushed a commit to 2416602906/vllm-ascend that referenced this pull request Jun 8, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset:
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query)
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests

- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: shenqiangqiang <2416602906@qq.com>
LostFox11 pushed a commit to LostFox11/vllm-ascend that referenced this pull request Jun 15, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
LostFox11 pushed a commit to LostFox11/vllm-ascend that referenced this pull request Jun 15, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
HaoxinZong pushed a commit to HaoxinZong/vllm-ascend that referenced this pull request Jun 27, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
pisceskkk pushed a commit to pisceskkk/vllm-ascend that referenced this pull request Jun 29, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
### What this PR does / why we need it?

- Prefix Mamba Cache Support: Updated the NPU model runner to correctly
handle and expand block tables when prefix caching is enabled for Mamba
models.
- Refactored Block Size Calculation: Streamlined the calculation of
kernel block sizes and maximum blocks per request by iterating directly
over KV cache specifications.
- Configuration Integration: Included cp_kv_cache_interleave_size in the
NPUInputBatch initialization to ensure consistent configuration across
the 310p backend.

### Does this PR introduce _any_ user-facing change?

NA

### How was this patch tested?
```
[2026-05-28 10:29:25,318] [ais_bench.benchmark.openicl.icl_inferencer.icl_gen_perf_inferencer] [INFO] Performance task finished, results saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat
05/28 10:29:25 - AISBench - INFO - time elapsed: 14.55s
Running tasks: 100%|██████████| 1/1 [00:27<00:00, 27.50s/it]
05/28 10:29:27 - AISBench - INFO - Performance evaluation tasks completed.
05/28 10:29:27 - AISBench - INFO - Loading detail perf data of model='vllm-api-stream-chat' dataset='gsm8kdataset' ...
05/28 10:29:27 - AISBench - INFO - Starting request timeline processing...
05/28 10:29:27 - AISBench - INFO - Data preprocessing completed in 0.0004s
05/28 10:29:27 - AISBench - INFO - Generating timeline traces for 2 requests...
05/28 10:29:28 - AISBench - INFO - Generated timeline trace chunks in 0.0806s
05/28 10:29:28 - AISBench - INFO - Generating concurrency traces...
05/28 10:29:28 - AISBench - INFO - Generated concurrency trace chunks in 0.0012s
05/28 10:29:28 - AISBench - INFO - Creating figure layout...
05/28 10:29:28 - AISBench - INFO - Figure layout created in 0.0926s
05/28 10:29:28 - AISBench - INFO - Writing to /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html...
05/28 10:29:28 - AISBench - INFO - HTML written in 0.0294s
05/28 10:29:28 - AISBench - INFO - Completed! Total execution time: 0.2050s
05/28 10:29:28 - AISBench - INFO - The gsm8kdataset_plot has been saved in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat/gsm8kdataset_plot.html
05/28 10:29:28 - AISBench - INFO - Converting perf results of stage ...
05/28 10:29:28 - AISBench - INFO - Finish Converting!
05/28 10:29:28 - AISBench - INFO - Start calculating metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating common metrics ...
05/28 10:29:28 - AISBench - INFO - Start calculating add units ...
05/28 10:29:28 - AISBench - INFO - Finish calculating perf data!
05/28 10:29:28 - AISBench - INFO - Summarizing performance results...
05/28 10:29:28 - AISBench - INFO - Performance Results of task: vllm-api-stream-chat/gsm8kdataset: 
╒══════════════════════════╤═════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤════════════════╤═════╕
│ Performance Parameters   │ Stage   │ Average        │ Min            │ Max            │ Median         │ P75            │ P90            │ P99            │  N  │
╞══════════════════════════╪═════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪════════════════╪═════╡
│ E2EL                     │ total   │ 1965.7806 ms   │ 1955.5967 ms   │ 1975.9645 ms   │ 1965.7806 ms   │ 1970.8726 ms   │ 1973.9277 ms   │ 1975.7609 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TTFT                     │ total   │ 1835.7014 ms   │ 1825.5388 ms   │ 1845.864 ms    │ 1835.7014 ms   │ 1840.7827 ms   │ 1843.8315 ms   │ 1845.6607 ms   │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ TPOT                     │ total   │ 32.5198 ms     │ 32.5145 ms     │ 32.5251 ms     │ 32.5198 ms     │ 32.5225 ms     │ 32.5241 ms     │ 32.525 ms      │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ ITL                      │ total   │ 21.5905 ms     │ 0.0252 ms      │ 35.3068 ms     │ 29.8905 ms     │ 34.5709 ms     │ 35.0401 ms     │ 35.2815 ms     │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ InputTokens              │ total   │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │ 10250.0        │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokens             │ total   │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │ 5.0            │  2  │
├──────────────────────────┼─────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼────────────────┼─────┤
│ OutputTokenThroughput    │ total   │ 2.5436 token/s │ 2.5304 token/s │ 2.5568 token/s │ 2.5436 token/s │ 2.5502 token/s │ 2.5542 token/s │ 2.5565 token/s │  2  │
╘══════════════════════════╧═════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧════════════════╧═════╛
╒══════════════════════════╤═════════╤═══════════════════╕
│ Common Metric            │ Stage   │ Value             │
╞══════════════════════════╪═════════╪═══════════════════╡
│ Benchmark Duration       │ total   │ 3931.8488 ms      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Requests           │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Failed Requests          │ total   │ 0                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Success Requests         │ total   │ 2                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Concurrency              │ total   │ 0.9999            │
├──────────────────────────┼─────────┼───────────────────┤
│ Max Concurrency          │ total   │ 1                 │
├──────────────────────────┼─────────┼───────────────────┤
│ Request Throughput       │ total   │ 0.5087 req/s      │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Input Tokens       │ total   │ 20500             │
├──────────────────────────┼─────────┼───────────────────┤
│ Prefill Token Throughput │ total   │ 5583.6968 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Total generated tokens   │ total   │ 10                │
├──────────────────────────┼─────────┼───────────────────┤
│ Input Token Throughput   │ total   │ 5213.8322 token/s │
├──────────────────────────┼─────────┼───────────────────┤
│ Output Token Throughput  │ total   │ 2.5433 token/s    │
├──────────────────────────┼─────────┼───────────────────┤
│ Total Token Throughput   │ total   │ 5216.3756 token/s │
╘══════════════════════════╧═════════╧═══════════════════╛
05/28 10:29:28 - AISBench - INFO - Performance Result files locate in /home/c00939328/outputs/default/20260528_102900/performances/vllm-api-stream-chat.
2026-05-28 10:29:30,430 - INFO - [完成] 全量数据集测试完成,结果保存在aisbench_result.csv
{0: 110702} {0: 0}
{0: 31488} {0: 0}

=======================================================================================================
POD: 127.0.0.1:2077
=======================================================================================================
engine_id       hbm_hit_rate         hbm(hit/query)       externel_hit_rate    externel(hit/query) 
-------------------------------------------------------------------------------------------------------
0               78.67%               16128/20500          0%                   0/0                 
=======================================================================================================
cp: cannot create regular file 'outputs/default/20260528_102900': No such file or directory
2026-05-28 10:29:36,519 - INFO - 成功将 aisbench.log 的内容追加到 aisbench_all.log
2026-05-28 10:29:36,527 - INFO - 文件已存在,读取现有数据
2026-05-28 10:29:36,533 - INFO - 成功追加新行
```
CI and local tests


- vLLM version: v0.20.2
- vLLM main:
vllm-project/vllm@39910f2

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants