Repository navigation
Conversation
A3 performance results: two FIA + merge vs. per-request concat + BSND FIAMeasured commit The comparison is between two FIA implementations of MLA prefix prefill:
The benchmark extracts both branch bodies directly from the recorded source and uses identical Q/K/V, cache, projection weights and sequence metadata. To compare the same DeepSeek input, it explicitly selects the BSND body for QK=192/V=128 despite the production dispatch condition; this CANN build accepted that shape and passed the numerical checks. The production condition and kernel arguments were not changed. This is not a comparison with ATB or an end-to-end comparison of Completed 18 shape cases and an independent 8-case rerun. The table below shows the rerun; the PR description now includes the complete 18-case sweep as well. B = request count, Q/P = query/prefix tokens per request, H = heads per rank; a ratio above 1 favors two FIA.
Result: for Q=100/P=768, the two-FIA branch wins from B=2 in this sweep, reaching 3.55x at B=8 and 10.80x at B=32 in the rerun. At B=1, concatenation is faster: about 23% lower latency for Q=100/P=768 and 16% lower for Q=1024/P=8192. With Q=1024/P=8192, B=4 favors BSND by about 4%, while B=8 is near parity; a 1-2% gap should not be treated as a significant stable win. These are prefix-branch ratios, not whole-model speedups. The source structure is consistent with the trend: per-request concatenation and B FIA submissions become expensive for short-query batches, while the TND path batches requests into two FIA calls plus one merge. B=1 avoids the second call and merge with BSND. This interpretation has not been confirmed by kernel-level profiling. The results support retaining the batched two-FIA path; a B=1 specialization is a possible follow-up, subject to device/version and full-model validation. Times include cache gather, BF16 Numerical checks: all 18 sweep cases and all 8 rerun cases passed full-output pairwise comparison ( |
This reverts commit 1344e1e.
NZ prefix-cache validation on A3 (2026-09-24)Commit Environment: machine 151, Ascend910_9362 (A3), CANN 9.1.0, PyTorch 2.10.0, torch_npu 2.10.0.post6. The patch changes only the two cache-gather calls; it does not reintroduce the reverted quantization API changes. Operator-path validation and performance: 18 cases plus 6 independent reruns passed. The fixture uses the actual The independent rerun below measures median amortized wall time for cache gather, BF16 projection, attention, copies and merge. QK=192, V=128, latent=512, page=128, H=4; BF16/eager. It uses 20 warmups and 9 rounds of 30 calls per method with randomized method order. NZ BSND explicitly selects the existing per-request branch for the same inputs, without changing production dispatch. NZ here describes cache storage; the two attention calls still use TND inputs after gathering.
NZ layout restoration adds about 4%-13% to the two-FIA prefix branch in this rerun. With NZ disabled, the helper adds no layout conversion; the first sweep measured ND before/after differences within approximately -0.34% to +0.65%. Short-query batched two-FIA remains faster than per-request BSND. These are branch measurements and do not establish whole-model speedup. Full model accuracy: ran the original
NZ minus ND: -0.606520 percentage points. Final numerical predictions differ on 581 questions (96 correct-to-incorrect, 88 incorrect-to-correct). These full-model runs exercise other NZ-dependent cache/decode paths as well, so model-level differences cannot be attributed to the prefix gather alone. Evaluation runtime and token/s are observations from accuracy runs with variable generation lengths, not a controlled serving performance benchmark. The host is shared; another process was observed on device 3 at the end of the NZ run, so model-run timing is not a device-exclusive measurement. Device-prefix cache hits: ND 1319/1319, NZ 1319/1319; cached tokens: ND 1012992, NZ 1012992. Host-cache reload hits: ND 0, NZ 0. This does not establish host-cache eviction/reload correctness when those counts are zero. A2/A5 and CANN 9.2 are not hardware-validated by this run. Full logs, inputs, source hashes, timing samples and per-question predictions were retained. |
Kimi-K3 accuracy validation with FIA + NZ — 2026-09-24Completed one full accuracy run on the latest-main integration snapshot recorded below. GSM8K met the requested target; GPQA Diamond finished at 184/198 (92.9293%), one correct answer short of the requested 93.4% target.
Exact source snapshot
Environment and serving configuration Four Ascend nodes, TP=64 / DP=4, container Enabled Evaluation settings
Validation checks Both evaluations exited with code 0. GSM8K had no invalid answers; independently checking the saved per-question outputs reproduced 50/50. GPQA completed all 198 requests with zero API or review errors. All 198 responses had No baseline was run, and the evaluation was not repeated with adjusted settings to reach the target. This run shows GPQA slightly below the requested threshold; it does not establish an accuracy delta versus main or attribute that difference to this PR. Full logs, dataset hashes, actual server arguments, predictions, reviews, and the 14 incorrect cases are retained with the validation artifacts. |
Kimi-K3 128K/1K cached-prefix TTFT — main vs this PRFour-node end-to-end serving measurement: mean TTFT changed from 8.629 s to 7.401 s (14.22% lower; main/PR TTFT ratio 1.166×) across three runs of 32 requests per variant. The observed average improved, but the small sample and visible run-to-run variation should be considered when interpreting the size of the gain.
The observed mean-TTFT difference is preliminary and sensitive to run-to-run variation. P99 TTFT did not improve. Additional measured metrics are included for completeness:
All six measured runs completed 32/32 requests, with exactly 128000 input tokens and 1000 output tokens per request, and no request errors. Each measured request reported 127872 cached tokens / 128000 = 99.9% cache hit. The dataset is named Snapshots and configuration
Workload and timing method Used the existing Each measured round starts with cache flush, four 128000/1 priming requests, then four unmeasured 128000/1000 stabilization requests to populate the reusable cache state on all four DP ranks. The initial main-side check with only the original 1-token priming reached 98.92% overall cache hit; it was retained as a setup diagnostic and excluded before the matched hot-cache protocol was run on either variant. All six included rounds have identical per-request cache counts and output lengths. Reported TTFT comes from the streaming benchmark client and includes serving/queueing overhead. All three rounds are retained; the main per-run mean ranges from 5.236–10.383 s, and PR from 5.867–10.418 s. This is a small end-to-end comparison with visible run-to-run variation, not an isolated kernel timing. Full per-request timings, cache reports, actual server arguments, logs, and environment attestations are retained. This performance comparison uses the newer main snapshot above. The previous accuracy results used main |
|
/tag-and-rerun-ci |
Motivation
MLA prefix-cache prefill currently requires ATB
npu_ring_mla, even whenASCEND_USE_FIA=1. This prevents the prefix path from running in environments where RingMLA is unavailable.Allow the existing
ASCEND_USE_FIAflag to select CANN FIA V2 for this path. With the flag unset or disabled, retain the original ATB calls and RingMLA mask.Honor
SGLANG_NPU_FORWARD_NATIVE_GEMMA_RMS_NORMinGemma3RMSNormon NPU.Modifications
Gemma3RMSNorm.forward_npu. The flag remains disabled by default;Gemma4RMSNormalready falls back to native on NPU.gather_mla_cache_pagesin the two-FIA branch, restoring logical token order for PA-NZ storage before projection and attention. Ordinary ND cache behavior is unchanged.npu_attention_updatein FP32 before restoring the query dtype.Enable the new path with:
export ASCEND_USE_FIA=1This reuses the existing backend-wide flag; its default remains disabled. FIA V2 is used because the TorchNPU 26.1 implementation dispatches to CANN FIA V4 on A2/A3 and V5 on Ascend 950. The approach follows the FIA + AttentionUpdate composition used by vLLM Ascend.
To select the native Gemma normalization implementation on NPU:
export SGLANG_NPU_FORWARD_NATIVE_GEMMA_RMS_NORM=1Accuracy Tests
Prefix dispatch refactor validation (2026-09-28)
Commit
ee3dfe96a1bcc93bcdf3d6f4c04eb489b659de28reorganizes the MLA prefix path into one completeif self.use_fia/elseblock. For each fixed value ofself.use_fia, the old and new Python ASTs contain the same executable statements in the same order, including operator arguments, NZ cache handling, empty-prefix handling and output padding. Ruff 0.15.1 formatting and lint (F401,F821,UP037), Python syntax parsing andgit diff --checkpassed. No new unit tests or NPU model evaluations were run for this structural change; the numerical and performance results below retain their recorded commit attribution.NZ cache validation (2026-09-24)
Commit
0c8ebbce034c11b5e0b9c2fec3193def9de2aa74passed the original full HiCache MLA GSM8K case on A3/151 with FIA enabled in both runs: NZ disabled 473/1319 (35.860500%), NZ enabled 465/1319 (35.253980%). Both exceed the 34% case threshold. The 18 operator-path cases and 6 independent reruns also passed, with fixed NZ outputs bitwise equal to the ND baseline. See the NZ validation section under Speed Tests and Profiling for the complete environment, comparison, numerical checks, performance table and limitations. Python syntax parsing andgit diff --checkpassed; no new unit tests were added.Latest validation (2026-09-24)
At commit
ef8a80be18d4d1e6efa84252cca07dd31284850a, Python syntax parsing andgit diff --checkpassed for the three changed files. The quantization implementation and its test file match the PR base; the remaining PR changes concern MLA prefix attention and Gemma normalization. No additional unit tests or model evaluations were run for this update. The accuracy and performance measurements below remain attributed to their explicitly recorded earlier commits.Gemma native fallback validation (2026-09-22)
Commit
ec72a1dd7c4292477bdd8e317a55f87c44f04357changes onlylayernorm.py. Python syntax compilation andgit diff --checkpassed. No additional unit tests or NPU model evaluations were run for this change. The model accuracy measurements below belong to their explicitly recorded earlier commits.Validation of updated PR head 3ba07dd on machine 216 (2026-09-21)
After the author merged community main, tested the actual PR head
3ba07ddc43d55782e2195f5cd2457297b1eef9fd, which includes main62ba9648482e1b0a253187c8dc6ac5d23bbbd403. Main is already an ancestor of the PR;git merge-treesucceeds and GitHub reports MERGEABLE. This rerun uses the published PR commit directly, rather than the earlier temporary integration snapshot.Ran its original
test/registered/npu/basic_function/HiCache/test_npu_hicache_mla.pyon machine 216, physical A3 devices 0–3, with CANN 9.1.0 / torch_npu 2.10.0.post4 / PyTorch 2.10.0 / Transformers 5.12.1. Reused the same full DeepSeek-V2-Lite-W8A8 checkpoint and GSM8K dataset. Both modes used TP=4, mem_fraction_static=0.8, HiCache enabled (current defaults: ratio=2.0, host memory fraction=0.8), all 1319 questions, 5-shot, temperature=0, max_new_tokens=512, concurrency=128, and seed=42. The PR smoke shortcut was disabled. No additional unit tests were added or run.FIA minus baseline: -3 correct answers / -0.227445 percentage points. 681 final numerical predictions differed: 111 changed from correct to incorrect, and 108 from incorrect to correct. This single pair does not establish lossless replacement. The existing backend-wide ASCEND_USE_FIA flag also affects other attention paths, so this does not isolate the new prefix branch alone.
Source hashes (including files updated by the merged main), dataset, devices, model path, evaluation settings, and actual server argument dictionaries were checked for consistency. Device-prefix cache hits: 1319/1319 unset and 1319/1319 enabled, 1012992 and 1012992 cached tokens respectively. Host-cache reload hits: 0 and 0; this case does not force host-cache eviction/reload. Full logs, metrics and per-question answers were retained.
Observed accuracy-evaluation runtime / output throughput: unset 198.626 s / 794.967 token/s, enabled 207.717 s / 755.773 token/s. These are not controlled performance measurements. A2/A5 and CANN 9.2 remain untested. Earlier validation records are preserved below.
Latest-main integration rerun on machine 216 (2026-09-21)
Community main
176dbcb85d3b7737564e4911947035cc0af65f0aand PR head1344e1e2a5ca3e07bf6140b8c606e38465fbc55cmerged without conflicts in bothgit merge-treeand an actual isolated merge. GitHub reports MERGEABLE; the CI gate is blocked by the missingrun-cilabel, independently of merge conflicts. Tested the resulting temporary merge2b388ce85710f74030b266f8446a01b91c556d48without pushing it to the PR branch.During the run, main advanced to
ab03a8e7eb82d35907bfbeb645bf0af8b4ce290e(#40201, import/startup changes). A finalgit merge-treecheck against that main also succeeded and GitHub still reports MERGEABLE. It does not modify this PR's NPU attention/quantization files or the HiCache MLA case. The accuracy results below belong to the fixed merge snapshot above; the full model test was not repeated on this newer main commit.Ran the current merged-source
test_npu_hicache_mla.pyon machine 216, four Ascend A3 devices (physical 0–3), CANN 9.1.0 / torch_npu 2.10.0.post4 / PyTorch 2.10.0 / Transformers 5.12.1. Used the same full DeepSeek-V2-Lite-W8A8 checkpoint as the previous runs, copied from machine 209, with TP=4,mem_fraction_static=0.8, HiCache enabled, all 1319 GSM8K questions, 5-shot, temperature 0, max 512 new tokens, concurrency 128, and seed 42 for both modes. The updated main test no longer explicitly setshicache_ratio=1.2; this rerun preserved its defaults, resolving to ratio 2.0 / host_memory_fraction 0.8. No additional unit tests were added or run for this rerun.FIA minus baseline: +11 correct answers / +0.833965 percentage points. This single paired run does not establish lossless replacement or isolate the new prefix branch:
ASCEND_USE_FIAaffects other existing attention paths too. 661 final numerical predictions differed; 102 changed from correct to incorrect and 113 from incorrect to correct.Both runs used matching source, dataset, devices, model path, evaluation parameters and actual server argument dictionaries. All 1319 requests in each run reused 768 device-cached prefix tokens; neither run forced host-cache reload. Full logs, exact metrics, source hashes and per-question answers were retained. Observed evaluation runtime / output throughput: unset 227.836 s / 691.215 token/s, enabled 205.004 s / 768.774 token/s. Other devices on the host had existing workloads; these observations are not a controlled performance benchmark. A2, A5 and CANN 9.2 remain untested.
Historical validation on machine 209 (2026-09-18)
Ran the existing
test/registered/npu/basic_function/HiCache/test_npu_hicache_mla.pyat commit1344e1e2a5ca3e07bf6140b8c606e38465fbc55con four Ascend910_9362 (A3) devices with CANN 9.1.0, PyTorch 2.10.0, torch_npu 2.10.0.post4 and Transformers 5.12.1.Used the full
vllm-ascend/DeepSeek-V2-Lite-W8A8checkpoint, adapting its local filesystem path. Each run used TP=4, HiCache enabled,hicache_ratio=1.2,mem_fraction_static=0.8, all 1319 GSM8K questions, 5-shot prompts, temperature 0,max_new_tokens=512and concurrency 128. The PR-pipeline single-request smoke shortcut was disabled.--random-seed 42--random-seed 42All four runs passed the case's 34% accuracy threshold. FIA scored 0.606520 percentage points lower in the first pair and 0.833965 percentage points lower in the fixed-seed pair. These measurements do not establish accuracy parity or lossless replacement; the ATB default and opt-in gate are retained.
In the fixed-seed pair, 686 final numerical predictions differed: 123 questions changed from correct to incorrect and 112 from incorrect to correct. Both actual server argument dictionaries matched, including seed 42. The first pair used automatically generated server seeds; both pairs are reported. The ATB runs themselves differed on 625 predictions between these two runs, so per-question changes cannot all be attributed to FIA from this experiment alone.
All 1319 requests in each run reused 768 device-cached prefix tokens, exercising prefix prefill with HiCache attached. This workload did not force host-cache eviction/reload. Source hashes, full server logs, exact metrics and per-question answers were retained.
ASCEND_USE_FIAis an existing backend-wide flag, so this model comparison includes every path affected by the flag rather than isolating only the new prefix branch.Additional validation performed before the model runs:
forward_extendprefix branch on Ascend910_9362 (A3), CANN 9.1.0, PyTorch 2.10.0 and torch_npu 2.10.0.post4: 18/18 cases passed. Cache gather, KV projection, FIA, merge and padding execute on NPU; the reference attention executes on CPU. The branch is loaded unchanged from source with unrelated imports isolated; this is not a complete serving/model test.All NPU branch cases also check output device, shape, dtype, finite values and zero padding. A2/A5 and CANN 9.2 have not been hardware tested.
Speed Tests and Profiling
NZ prefix-cache validation on A3 (2026-09-24)
Commit
0c8ebbce034c11b5e0b9c2fec3193def9de2aa74uses the existinggather_mla_cache_pageshelper from #39589 for both latent and RoPE prefix reads in the two-FIA branch. It restores logical token order before projection and TND attention when the KV cache uses PA-NZ storage. The helper is byte-identical to #39589 after newline normalization. ND behavior is preserved.Environment: machine 151, Ascend910_9362 (A3), CANN 9.1.0, PyTorch 2.10.0, torch_npu 2.10.0.post6. The patch changes only the two cache-gather calls; it does not reintroduce the reverted quantization API changes.
Operator-path validation and performance: 18 cases plus 6 independent reruns passed. The fixture uses the actual
NPUMLATokenToKVPool._set_fia_nz_kv_bufferand its index helper, executing real NPU scatter with randomized token write order. Both latent and RoPE readback are bitwise equal to the logical ND cache. Across all cases, the fixed NZ branch and fixed ND branch outputs are bitwise equal to the original ND branch output. Sampled CPU FP32 reference checks also pass (maximum RMS relative error 0.2995%). The unfixed NZ branch is retained as a negative control and shows RMS relative errors of 6.55%-119.16% on these synthetic cases.The independent rerun below measures median amortized wall time for cache gather, BF16 projection, attention, copies and merge. QK=192, V=128, latent=512, page=128, H=4; BF16/eager. It uses 20 warmups and 9 rounds of 30 calls per method with randomized method order. NZ BSND explicitly selects the existing per-request branch for the same inputs, without changing production dispatch. NZ here describes cache storage; the two attention calls still use TND inputs after gathering.
NZ layout restoration adds about 4%-13% to the two-FIA prefix branch in this rerun. With NZ disabled, the helper adds no layout conversion; the first sweep measured ND before/after differences within approximately -0.34% to +0.65%. Short-query batched two-FIA remains faster than per-request BSND. These are branch measurements and do not establish whole-model speedup.
Full model accuracy: ran the original
test_npu_hicache_mla.pyon the same physical devices 0-3 with DeepSeek-V2-Lite-W8A8, TP=4, HiCache enabled, radix enabled, seed=42, all 1319 GSM8K questions, 5-shot, temperature=0, max_new_tokens=512 and concurrency=128. Both runs setASCEND_USE_FIA=1; onlySGLANG_USE_FIA_NZchanges. The source, model path, dataset, devices, launch arguments and evaluation settings match.NZ minus ND: -0.606520 percentage points. Final numerical predictions differ on 581 questions (96 correct-to-incorrect, 88 incorrect-to-correct). These full-model runs exercise other NZ-dependent cache/decode paths as well, so model-level differences cannot be attributed to the prefix gather alone. Evaluation runtime and token/s are observations from accuracy runs with variable generation lengths, not a controlled serving performance benchmark. The host is shared; another process was observed on device 3 at the end of the NZ run, so model-run timing is not a device-exclusive measurement.
Device-prefix cache hits: ND 1319/1319, NZ 1319/1319; cached tokens: ND 1012992, NZ 1012992. Host-cache reload hits: ND 0, NZ 0. This does not establish host-cache eviction/reload correctness when those counts are zero. A2/A5 and CANN 9.2 are not hardware-validated by this run. Full logs, inputs, source hashes, timing samples and per-question predictions were retained.
Controlled prefix-branch comparison on A3 (2026-09-22)
Measured commit
ec72a1dd7c4292477bdd8e317a55f87c44f04357on 2026-09-22, machine 151: one Ascend910_9362 (A3), CANN 9.1.0 (B070 container), PyTorch 2.10.0, torch_npu 2.10.0.post6, BF16, NZ disabled, eager execution. The main shape uses QK=192 (NoPE=128 + RoPE=64), V=128, latent=512, page size=128 and H=4 per rank (the DeepSeek-V2-Lite TP=4 head shape); no TP communication is executed.The comparison is between two FIA implementations of MLA prefix prefill:
npu_attention_update.if layer.qk_head_dim == layer.v_head_dimbranch, concatenating prefix/current K/V and issuing one BSND FIA call per request. B requests therefore issue B FIA calls.The benchmark extracts both branch bodies directly from the recorded source and uses identical Q/K/V, cache, projection weights and sequence metadata. To compare the same DeepSeek input, it explicitly selects the BSND body for QK=192/V=128 despite the production dispatch condition; this CANN build accepted that shape and passed the numerical checks. The production condition and kernel arguments were not changed. This is not a comparison with ATB or an end-to-end comparison of
ASCEND_USE_FIAenabled/disabled.Times include cache gather, BF16
kv_b_proj, FIA, concatenation/contiguous copies, dtype conversions and merge. They exclude the rest of the model, QKV preparation, scheduling, decode and TP communication, and do not reproduce W8A8 projection. Each value is the median amortized wall time per complete prefix branch, with synchronization at each measurement block. Method order is randomized with a fixed seed. The first sweep uses 12 warmups and 7 rounds of 20 calls per method; an independent rerun uses 20 warmups and 9 rounds of 30 calls. Individually synchronized calls and NPU event spans were also recorded; event spans can include host submission gaps.Independent rerun (8 cases). B = requests, Q = new query tokens per request, P = cached prefix tokens per request, H = heads per rank. A BSND / two-FIA time ratio above 1 favors the PR's two-FIA path.
Result: for Q=100/P=768, the two-FIA branch wins from B=2 in this sweep, reaching 3.55x at B=8 and 10.80x at B=32 in the rerun. At B=1, concatenation is faster: about 23% lower latency for Q=100/P=768 and 16% lower for Q=1024/P=8192. With Q=1024/P=8192, B=4 favors BSND by about 4%, while B=8 is near parity; a 1-2% gap should not be treated as a significant stable win. These are prefix-branch ratios, not whole-model speedups.
The source structure is consistent with the trend: per-request concatenation and B FIA submissions become expensive for short-query batches, while the TND path batches requests into two FIA calls plus one merge. B=1 avoids the second call and merge with BSND. This interpretation has not been confirmed by kernel-level profiling. The results support retaining the batched two-FIA path; a B=1 specialization is a possible follow-up, subject to device/version and full-model validation.
Complete first sweep: 18 cases
The mixed B=10 case uses query lengths
[101,870,876,874,869,864,904,863,837,768]and prefix lengths[768,0,0,0,0,0,0,0,0,0], matching an earlier observed batch shape with synthetic tensor values. It favors two FIA by 2.84x. B=128/Q=100 contains 12,800 query tokens and is an extended sweep point, not a claim that default chunk limits produce that batch.Numerical checks: all 18 sweep cases and all 8 rerun cases passed full-output pairwise comparison (
atol=0.02,rtol=0.02) and sampled CPU FP32 full-attention reference checks. Maximum full-output absolute difference between methods was 0.001953125; maximum per-case pairwise RMS relative difference was 0.3022%. Maximum sampled RMS relative error against CPU FP32 was 0.2995% for two FIA and 0.2296% for BSND. The CPU reference uses the same projected K/V. This checks the sampled operator paths; it is not a new GSM8K accuracy run. Hardware conclusions here are limited to this A3/CANN/torch_npu configuration.Historical accuracy-run timings
The fixed-seed full GSM8K runs measured 193.304 s / 814.391 output tokens per second with FIA unset, and 196.474 s / 801.729 output tokens per second with FIA enabled. The first pair measured 225.281 s / 697.903 and 195.270 s / 802.929 respectively. These are accuracy-run measurements with differing generation lengths and dynamic batches, not a controlled performance benchmark; no speedup is claimed. The FIA path adds a separate merge and FP32 intermediate outputs and remains opt-in.
Checklist
CI States
Latest PR Test (Base): ✅ Run #36403090220
Latest PR Test (Extra): ❌ Run #36403089905
Latest PR Test (AMD ROCm 10): ❌ Run #36403090380