[Perf][SM70] Optimize Qwen3.8 Flash Next prefill - #351
Conversation
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
SM70 QSA indexer hotspot updateCommit The cached SM70 PTX showed that both QSA selected attention and index scoring
The exact 8192-token/top-10 NVFP4 MoE tuning screen was also closed: default The matched TP4 token-hash/quality gate and one Nsight Systems capture are |
|
Source audit update against the repaired #345 stack:
Decision under the project policy: the large prefill gain with bounded numerical drift and narrow fallback-safe gates is acceptable as default-on; greedy identity is not required. The prior remote failure came from #345 before its #359 fixes, not from this PR. I will resync/re-run CI after #345 lands, then merge if the final tree remains identical to the audited projection. |
|
SM70 prefill update (commit 8f151b5):
|
Objective
Improve and validate Qwen3.8 Flash Next NVFP4 TP4 prefill performance on four V100-SXM2-32GB GPUs. Decode/MTP optimization is explicitly out of scope.
Base
codex/v100-qwen38-flash-next-nvfp4-20260826-140311383ff458dbc8af1dda79a1a29bf700019030f054Planned gates
Status
Draft campaign boundary created; no prefill source change yet.