Repository navigation
Conversation
Lzy17
requested review from
Qiaolin-Yu,
Ying1123,
hnyls2002 and
merrymercy
as code owners
September 24, 2026 15:42
Collaborator
Author
|
Confirmed on MI355X, ROCm 10 (
One note on why zeros are the right value rather than a workaround: the graph |
This was referenced Sep 25, 2026
GLM-5.2 EP16 with MTP dies about a minute into serving:
File "speculative/eagle_worker_v2.py", line 932, in draft_forward
torch.stack(draft_probs_list, dim=1)
TypeError: expected Tensor as element 0 in argument 0, but got NoneType
The prefill worker sends topk_p and topk_index but no proposal distribution,
so build_eagle_disagg_draft_input leaves draft_probs None, and the eager draft
loop seeds its list with that field. ROCm auto-enables rejection sampling for
EAGLE topk=1, so the seed is required there.
The MTP tests that pass all capture CUDA graphs, and that runner allocates a
zeroed draft_probs buffer. This recipe sets --disable-cuda-graph and takes the
eager path, which had no equivalent.
Zeros, matching the graph path: the sampler rejects q == 0 and resamples that
position from the target.
Lzy17
force-pushed
the
ci/pd-eagle-draft-probs-seed
branch
from
September 30, 2026 07:16
317b4fe to
6fe3e6f
Compare
This was referenced Oct 5, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
GLM-5.2 EP16 with MTP dies about a minute into serving:
Cause
build_eagle_disagg_draft_inputis the only constructor ofEagleDraftInputon the PD decode path. The prefill worker sends
topk_pandtopk_indexbutno proposal distribution, so it leaves
draft_probsunset, and the eager draftloop seeds its list with that field:
The list is stacked at the end of the loop, so a
Noneseed is a guaranteedTypeErroras soon as the first draft step completes.The graph path never hit this because
EAGLEDraftCudaGraphRunnerallocates azeroed
draft_probsbuffer under the same condition(
eagle_draft_cuda_graph_runner.py:197). The eager path had no equivalent.Every MTP test that passes today captures graphs; this recipe sets
--disable-cuda-graph.Scope: this is not AMD-only
The tag is on the PR because that is where the nightly caught it, but the
defect is in common code and reachable on any backend:
speculative_use_rejection_samplingis a documented user-facing flag(
fields/spec.py:116), defaultFalse. Any CUDA user who sets it explicitlyand runs PD + EAGLE without CUDA graphs hits the identical crash. ROCm only
differs in that
_should_auto_enable_hip_rejection_samplingturns it on bydefault, so AMD reaches the path without opting in — which is why the nightly
found it first.
topk == 1and a DSA seed is required butdsa_topk_indicesisNone, thebatch is now marked
cuda_graph_compatible=Falseand falls back to eager.That is exactly the path with the missing seed.
Fix
Allocate zeros under the same condition the graph runner already uses, so the
two paths agree. Zeros are the correct seed rather than a placeholder: the
sampler rejects
q == 0and resamples that position from the targetdistribution, which is the behaviour the zeroed graph buffer has always
produced.
Rebased onto #32196; the two changes are independent additions to the same
function.
Verification
MI355X nightly, ROCm 10, GLM-5.2 EP16 + MTP, PD 1P1D with
--disable-cuda-graph: crashes onmain, serves to completion with thispatch, GSM8K 0.949.
CI States
Latest PR Test (Base): ✅ Run #36804310116
Latest PR Test (Extra): ❌ Run #36804309955
Latest PR Test (AMD ROCm 10): ❌ Run #36804310105