Repository navigation
[AMD] Add DeepSeek-V4.1-Flash MI35x nightly accuracy test - #41476
Merged
HaiShaw merged 4 commits intoOct 4, 2026
Merged
Conversation
michaelzhang-ai
requested review from
Fridge003,
HaiShaw,
Kangyan-Zhou,
bingxche,
ispobock and
merrymercy
as code owners
September 27, 2026 16:30
michaelzhang-ai
force-pushed
the
amd/nightly-dsv41-flash
branch
from
September 29, 2026 09:47
8e1eb0c to
89dd06a
Compare
This was referenced Sep 30, 2026
michaelzhang-ai
force-pushed
the
amd/nightly-dsv41-flash
branch
from
October 1, 2026 00:41
89dd06a to
65a6c10
Compare
DeepSeek-V4.1-Flash has no AMD test registered, so nothing in CI exercises the MI35x FP4 path the cookbook recipe publishes for it. Add a GSM8K few-shot completion eval on MI35x that launches the server with the cookbook's MI350X High-Throughput recipe (TP4 + EP4, DSpark, decode graphs up to the 256-request cap, decoder SWA bounded replay, the recipe's AITER env), and wire it into the nightly as nightly-4-gpu-mi35x-deepseek-v41-flash. FP4 experts need gfx95x, so there is no MI30x job.
michaelzhang-ai
force-pushed
the
amd/nightly-dsv41-flash
branch
from
October 1, 2026 00:41
65a6c10 to
ec8c4cf
Compare
The first full run (1319 questions, rocm10, MI35x, cookbook High-Throughput cell) scored 0.911. Set the threshold to 0.89, 0.02 under that, in place of the provisional 0.86.
Run 36797490320 took 1088 s end to end (about 16 min of server startup, then ~1 min of GSM8K), so lower est_time from 5400 to 1200.
michaelzhang-ai
force-pushed
the
amd/nightly-dsv41-flash
branch
from
October 1, 2026 13:16
dfcb339 to
d2a7ead
Compare
1 task done
michaelzhang-ai
force-pushed
the
amd/nightly-dsv41-flash
branch
from
October 1, 2026 13:46
dd33270 to
be1b03e
Compare
Delete nightly-8-gpu-mi35x-deepseek-v32 and -v32-mtp from the AMD nightly workflow: the job definitions, their job_select options, and their check-all-jobs entries. Neither runs on the schedule or by manual dispatch anymore. The accuracy and perf tests stay in the tree, and the MI30x V3.2 jobs are unchanged.
michaelzhang-ai
force-pushed
the
amd/nightly-dsv41-flash
branch
from
October 1, 2026 13:47
be1b03e to
208ef1c
Compare
This was referenced Oct 2, 2026
HaiShaw
approved these changes
Oct 4, 2026
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
DeepSeek-V4.1-Flash has no AMD test registered, so nothing in CI exercises the MI35x FP4 path that the cookbook recipe publishes for it.
Modifications
New test
test/registered/amd/accuracy/mi35x/test_deepseek_v41_flash_eval_mi35x.py: a GSM8K 8-shot completion eval (1319 questions) withdeepseek-ai/DeepSeek-V4.1-Flash.docs/src/snippets/configs/deepseek-ai/deepseek-v4_1.jsx, as dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308 left it:--tp 4 --ep-size 4 --mem-fraction-static 0.8 --speculative-algorithm DSPARK --speculative-dspark-block-size 5 --cuda-graph-max-bs-decode 256 --max-running-requests 256 --enable-decoder-swa-bounded-replay. The attention and MoE backends are left to resolve on HIP.AITER_FLYDSL_FORCE_REDUCE=1 AITER_BF16_FP8_MOE_BOUND=0.SGLANG_USE_AITERandSGLANG_MOE_PADDINGcome from the ROCm image.AITER_BF16_FP8_MOE_BOUND=0stays until [FlyDSL] [Bugfix] Route a16w4 SiLU MoE to CK-Tile and apply swiglu_limit ROCm/aiter#5802 is in the AITER pin.reasoning_content.New nightly job
nightly-4-gpu-mi35x-deepseek-v41-flashinnightly-test-amd.yml(dropdown option, job,check-all-jobs.needs), suitenightly-amd-4-gpu-mi35x-deepseek-v41-flash. It runs onlinux-mi35x-gpu-8and uses four GPUs, like the MiniMax TP4 jobs.The suite is added to
NIGHTLY_SUITESintest/run_suite.py.The DeepSeek-V3.2 MI35x nightly jobs are removed. This means
nightly-8-gpu-mi35x-deepseek-v32andnightly-8-gpu-mi35x-deepseek-v32-mtp: their job definitions, theirjob_selectoptions and theircheck-all-jobs.needsentries. They no longer run on the schedule or by manual dispatch. This offsets the MI35x runner time the new job adds. DeepSeek-V3.2 has been replaced by V4, and both jobs passed on all three ROCm images for the last 8 scheduled nightlies (48/48), at about 100 runner-minutes per night. The accuracy and perf tests stay in the tree, and the MI30x V3.2 jobs are unchanged.This is MI35x only because the FP4 experts need gfx95x. The recipe publishes no MI30x cell.
Status and threshold
The gfx950 series this needs (dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels #41019, dsv4.1-amd: gfx950 sparse decode attention and sorted top-k #41020, dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels #41021, dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308) is all on
main, and this branch is rebased on top of it.Threshold 0.89, from measured runs. The 4-GPU DeepSeek-V4.1-Flash job ran manually on all three ROCm images (MI35x, all 1319 questions):
The threshold sits 0.02 under the lowest score (rocm10), replacing the provisional 0.86. For comparison, dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308 measured 0.870–0.890 with the same cell on 200 questions.
The rocm10 run took about 9 minutes longer to start the server, likely because the weights weren't yet cached on that runner.
est_time=1200covers that case. Container setup and dependency install add about 6–7 minutes to every job. So the new job adds about 15–24 minutes per ROCm image per night onlinux-mi35x-gpu-8, and uses four of its eight GPUs.A Low-Latency (DSpark, decode cap 128) variant and a perf step are left for follow-ups.
Checklist
nightly-test-amd.ymlwithjob_select=nightly-4-gpu-mi35x-deepseek-v41-flashon this branch for rocm10, rocm724 and rocm720, and set the threshold from the results (0.911 / 0.917 / 0.920 → threshold 0.89)CI States
Latest PR Test (Base): ❌ Run #36871380250
Latest PR Test (Extra): ❌ Run #36871379298
Latest PR Test (AMD ROCm 10): ❌ Run #36871380180