Repository navigation
[AMD][SPEC]Fix DeepSeek-V4 benchmark accuracy gap by supporting temperature sampling in EAGLE verify - #39253
[AMD][SPEC]Fix DeepSeek-V4 benchmark accuracy gap by supporting temperature sampling in EAGLE verify#39253At1a8 wants to merge 5 commits into
Conversation
|
What traffic is used for |
| target_predict = torch.argmax(next_token_logits, dim=-1) | ||
| target_predict = select_target_predict( | ||
| next_token_logits, sampling_info, verify_input.draft_token_num | ||
| ) |
There was a problem hiding this comment.
Please do:
elif _is_hip:
target_predict = select_target_predict(
next_token_logits, sampling_info, verify_input.draft_token_num
)
Hi @HaiShaw , as far as I know, For now, I'd like to introduce a simpler implementation(this PR) as an initial step. Once the functionality is in place, we can further optimize and improve the implementation in follow-up work. |
|
@amd-bot ci-status |
CI Status for PR #39253Merge verdict: Do not merge on green alone. PR CI is incomplete (NVIDIA Caution The new code path is not verified by CI. The change in Caution Required NVIDIA downstream jobs did not complete: Changed files: Executed CI failure attribution: AMD: 5 failures (0 related) · Others: 2 failures (0 related) · plus cancelled B200/B300 + queued XPU (incomplete, not counted) AMD Executed Failures
Other Executed Failures
The six Details / what to do before merge
Generated by amd-bot using Claude Code CLI |
|
Please check #37134 |
Confirmed that #37134 also fixes the accuracy issue in DeepSeek-V4 MTP. Thanks @HaiShaw @xiaobochen-amd
|
This PR co-work with @yuttian1 and @amd-danli103
Motivation
On ROCm, EAGLE verify silently ignores the requested temperature.
In
eagle_sample,_is_hipandis_all_greedyuse the same path, so every HIP request takes the argmax branch — a workaround for the unregisteredtree_speculative_sampling_target_onlykernel that became a behaviour change with no error or warning.Greedy then falls into repetition loops on open-ended tasks: on Simple-QA 1323 of 4326 requests (30.6%) end without a parseable answer, versus 39 (0.9%) once temperature is honoured — most of the 8.97-point gap below.
Modifications
Add
select_target_predict(), which returns the target model's own token at each verify position:argmax(logits)(unchanged behaviour)SGLANG_SPEC_TEMPERATURE_SAMPLING_TARGET_VERIFY, hip gated) →y ~ softmax(logits / temperature)verify_tree_greedy_funcaccepts a draft token only when it equals this choice, so the committed token is always the target's own — argmax reproduces greedy exactly, sampling reproduces temperature sampling exactly. Draft probabilities are never used.Gated by
SGLANG_SPEC_TEMPERATURE_SAMPLING_TARGET_VERIFY(default off) and_is_hip; other backends(nv) unchanged.top_p/top_kfall back to argmax with a warning (renorm ops unregistered on ROCm). Accept length stays at 2.5+/4.Accuracy Tests
Official data copy from: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
Gap to official (target greedy) = Official - MTP314 (target greedy)
Gap to official (target temp sampling) = Official - MTP314 (target temp sampling)
These data measured on gfx950.
Speed Tests and Profiling
No significant performance difference between two different sampling way.
SGLANG_SPEC_TEMPERATURE_SAMPLING_TARGET_VERIFY=1Server command
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #34818468430
Latest PR Test (Extra): ❌ Run #34818468496
Latest PR Test (AMD ROCm 10): ❌ Run #34818468812