Skip to content

[Doc][SM70] Plan NVFP4 DFlash2 17 ms campaign - #294

Merged
yangzhuxinyzx merged 2 commits into
mainfrom
codex/v100-dflash2-nvfp4-17ms-20260825-152808
Aug 25, 2026
Merged

yangzhuxinyzx merged 2 commits into
mainfrom
codex/v100-dflash2-nvfp4-17ms-20260825-152808

Conversation

@yangzhuxinyzx

@yangzhuxinyzx yangzhuxinyzx commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Purpose

Merge the reproducible planning record for reducing the complete batch-one NVFP4 + DFlash2 speculative round from the accepted 18.465--18.603 ms baseline to 17.0 ms or lower on TP4 V100.

This PR does not claim that the optimization or 17 ms result already exists. It freezes the matched workload, trace-first sequence, promotion gates, rejected paths, and rollback policy so the implementation campaign does not repeat stale experiments.

Promotion policy

  • A retained matched performance win becomes default-on after engine-contract admission, rollback, numerical, official-sampling quality, and endpoint gates pass.
  • Valid reduction-order changes do not need bitwise output or greedy-token identity. They must remain finite, stay within a dtype-appropriate absolute/relative error envelope, and preserve acceptance and scored quality.
  • Model/checkpoint labels are benchmark evidence only and never runtime activation predicates.

Test Result

  • Rebased by merge onto main@cadcf1d899b6d7511f815e7ee939b1e4676aff19.
  • Removed private filesystem paths from the committed record.
  • Changed-file markdown, typo, configuration, and documentation gates pass.
  • Source implementation, current trace, and 17 ms measurement remain explicitly pending.

Define the reproducible short and practical contracts, acceptance and quality gates, long-context decay checks, and trace-first implementation sequence.\n\nAssisted-by: OpenAI Codex

Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor Author

审计结论:不合并,保留 campaign 文档分支后关闭。

当前 PR 只有 123 行计划/控制文档,没有源码实现、current-source baseline、graph-node trace、17 ms 结果、质量门禁或 rollback 测试;标题描述的性能提升尚未发生。把未执行的优化计划合并到 main 会让主线控制文档把目标误读成已完成能力。

这里没有可修复后直接合并的代码。最优路径是以后从最新 main 新开实现 PR:先给出同条件 baseline/trace,做一个最小可归因改动,再补 focused 数值、CUDA Graph、接受率和质量证据。现有 branch/commit 仍保留为研究计划,不丢失。

…4-merge-20260825-172317

Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
@yangzhuxinyzx yangzhuxinyzx changed the title [Kernel][SM70] Reduce NVFP4 DFlash2 rounds to 17 ms [Doc][SM70] Plan NVFP4 DFlash2 17 ms campaign Aug 25, 2026
@yangzhuxinyzx
yangzhuxinyzx marked this pull request as ready for review August 25, 2026 17:27
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor Author

按新的合并标准复审通过:该 PR 作为计划文档合并,不声称已有 17 ms 实现。已更新到当前 main,移除私有路径,明确性能胜出后默认开启,并将数值门禁改为 dtype 适配误差与采样质量,不再要求 bitwise/greedy 一致。变更文件文档与静态门禁通过。

@yangzhuxinyzx
yangzhuxinyzx merged commit 9368f49 into main Aug 25, 2026
@yangzhuxinyzx
yangzhuxinyzx deleted the codex/v100-dflash2-nvfp4-17ms-20260825-152808 branch August 26, 2026 07:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant