Skip to content

perf(flydsl): select validated Kimi-K3 KDA schedule - #14

Closed
JohnQinAMD wants to merge 1 commit into
perf/kimi-k3-kda-group64-gfx950from
perf/kimi-k3-wave40-kda-schedule
Closed

JohnQinAMD wants to merge 1 commit into
perf/kimi-k3-kda-group64-gfx950from
perf/kimi-k3-wave40-kda-schedule

Conversation

@JohnQinAMD

Copy link
Copy Markdown
Owner

Summary

Select the validated gfx950 Kimi-K3 KDA group64 default schedule:

  • rows_per_wave: 1 -> 2;
  • cu_count: 240 -> 256;
  • keep weight_cache_modifier=2.

The kernel, quantization format, ABI, and explicit override controls are
unchanged. A regression test records the selected wrapper defaults.

This is stacked on perf/kimi-k3-kda-group64-gfx950.

Correctness and performance

The existing group64 campaign passed all changed-input checks bitwise. Two
independent schedule runs saved 0.035622 and 0.032652 ms/token
respectively; their mean is 0.034137 ms/token.

This schedule is intentionally not assigned standalone endpoint credit. It is
one member of the indivisible Wave40 bundle, whose three exact
TP8/B1/8K+1K/FP8-KV/no-spec trials improved the immutable Wave38 parent from
12.194182 to 11.886058 ms/token (82.006322 to
84.132180 tok/s/GPU). GSM8K first-100 was 100/100.

Validation

  • Ruff: passed;
  • Python compile: passed;
  • focused wrapper/schedule tests in the exact Wave40 image: 10 passed;
  • no NVIDIA source or dispatch path changes.

@JohnQinAMD

Copy link
Copy Markdown
Owner Author

Superseded by #11. The validated rows-per-wave 2 / CU256 / cache-modifier 2 schedule is now frozen directly in the standalone group64 PR, together with the extended contract tests and current endpoint evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant