Skip to content

[AMD] Add DeepSeek-V4.1-Flash MI35x nightly accuracy test - #41476

Merged
HaiShaw merged 4 commits into
sgl-project:mainfrom
michaelzhang-ai:amd/nightly-dsv41-flash
Oct 4, 2026
Merged

HaiShaw merged 4 commits into
sgl-project:mainfrom
michaelzhang-ai:amd/nightly-dsv41-flash

Conversation

@michaelzhang-ai

@michaelzhang-ai michaelzhang-ai commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

DeepSeek-V4.1-Flash has no AMD test registered, so nothing in CI exercises the MI35x FP4 path that the cookbook recipe publishes for it.

Modifications

  • New test test/registered/amd/accuracy/mi35x/test_deepseek_v41_flash_eval_mi35x.py: a GSM8K 8-shot completion eval (1319 questions) with deepseek-ai/DeepSeek-V4.1-Flash.

    • Server args are the MI350X High-Throughput cell of docs/src/snippets/configs/deepseek-ai/deepseek-v4_1.jsx, as dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308 left it: --tp 4 --ep-size 4 --mem-fraction-static 0.8 --speculative-algorithm DSPARK --speculative-dspark-block-size 5 --cuda-graph-max-bs-decode 256 --max-running-requests 256 --enable-decoder-swa-bounded-replay. The attention and MoE backends are left to resolve on HIP.
    • Env is the recipe's AITER_FLYDSL_FORCE_REDUCE=1 AITER_BF16_FP8_MOE_BOUND=0. SGLANG_USE_AITER and SGLANG_MOE_PADDING come from the ROCm image. AITER_BF16_FP8_MOE_BOUND=0 stays until [FlyDSL] [Bugfix] Route a16w4 SiLU MoE to CK-Tile and apply swiglu_limit ROCm/aiter#5802 is in the AITER pin.
    • It uses the completion harness, like the DeepSeek-V4-Flash MI35x test, so V4.1's thinking mode can't route the answer into reasoning_content.
  • New nightly job nightly-4-gpu-mi35x-deepseek-v41-flash in nightly-test-amd.yml (dropdown option, job, check-all-jobs.needs), suite nightly-amd-4-gpu-mi35x-deepseek-v41-flash. It runs on linux-mi35x-gpu-8 and uses four GPUs, like the MiniMax TP4 jobs.

  • The suite is added to NIGHTLY_SUITES in test/run_suite.py.

  • The DeepSeek-V3.2 MI35x nightly jobs are removed. This means nightly-8-gpu-mi35x-deepseek-v32 and nightly-8-gpu-mi35x-deepseek-v32-mtp: their job definitions, their job_select options and their check-all-jobs.needs entries. They no longer run on the schedule or by manual dispatch. This offsets the MI35x runner time the new job adds. DeepSeek-V3.2 has been replaced by V4, and both jobs passed on all three ROCm images for the last 8 scheduled nightlies (48/48), at about 100 runner-minutes per night. The accuracy and perf tests stay in the tree, and the MI30x V3.2 jobs are unchanged.

This is MI35x only because the FP4 experts need gfx95x. The recipe publishes no MI30x cell.

Status and threshold

Checklist

  • Dispatch nightly-test-amd.yml with job_select=nightly-4-gpu-mi35x-deepseek-v41-flash on this branch for rocm10, rocm724 and rocm720, and set the threshold from the results (0.911 / 0.917 / 0.920 → threshold 0.89)

CI States

Latest PR Test (Base): ❌ Run #36871380250
Latest PR Test (Extra): ❌ Run #36871379298
Latest PR Test (AMD ROCm 10): ❌ Run #36871380180

DeepSeek-V4.1-Flash has no AMD test registered, so nothing in CI
exercises the MI35x FP4 path the cookbook recipe publishes for it.

Add a GSM8K few-shot completion eval on MI35x that launches the server
with the cookbook's MI350X High-Throughput recipe (TP4 + EP4, DSpark,
decode graphs up to the 256-request cap, decoder SWA bounded replay, the
recipe's AITER env), and wire it into the nightly as
nightly-4-gpu-mi35x-deepseek-v41-flash. FP4 experts need gfx95x, so
there is no MI30x job.
@michaelzhang-ai
michaelzhang-ai force-pushed the amd/nightly-dsv41-flash branch from 65a6c10 to ec8c4cf Compare October 1, 2026 00:41
The first full run (1319 questions, rocm10, MI35x, cookbook High-Throughput
cell) scored 0.911. Set the threshold to 0.89, 0.02 under that, in place of
the provisional 0.86.
Run 36797490320 took 1088 s end to end (about 16 min of server startup, then ~1 min of GSM8K), so lower est_time from 5400 to 1200.
Delete nightly-8-gpu-mi35x-deepseek-v32 and -v32-mtp from the AMD
nightly workflow: the job definitions, their job_select options, and
their check-all-jobs entries. Neither runs on the schedule or by manual
dispatch anymore. The accuracy and perf tests stay in the tree, and the
MI30x V3.2 jobs are unchanged.
@HaiShaw
HaiShaw merged commit f238531 into sgl-project:main Oct 4, 2026
101 of 114 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants