Repository navigation
[Spec Decode][ROCm] Keep draft models at checkpoint precision under a4w4 and quantized quick reduce - #59884
Draft
Fangzhou-Ai wants to merge 1 commit into
Draft
[Spec Decode][ROCm] Keep draft models at checkpoint precision under a4w4 and quantized quick reduce#59884Fangzhou-Ai wants to merge 1 commit into
Fangzhou-Ai wants to merge 1 commit into
Conversation
Run speculative draft construction and forward inside draft_model_scope(), applied once in initialize_model for every proposer. The ROCm a4w4 MoE opt-in (VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4) and quick reduce's quantized all-reduce now skip draft models, so they cannot lower acceptance; the target is unchanged. Co-authored-by: Cursor Agent Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Keep speculative-decoding draft models at their checkpoint precision when two ROCm opt-ins that lower precision are enabled:
VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4=1(FP4 MoE activations) andVLLM_ROCM_QUICK_REDUCE_QUANTIZATION(quantized TP all-reduce). The target model behaves as before.Claims
draft_model_scope(). The scope is applied once, ininitialize_model, so it covers every proposer on both model runners (MRV1 and MRV2), including CUDA-graph capture and breakable CUDA graphs.FusedMoEConfigno longer enables a4w4 for draft MoE layers. DeepSeek V4.1's DSpark draft keeps the checkpoint's a8w4 (MXFP8 activations), and the target still runs a4w4.QuickAllReduce.should_quick_allreducereturnsFalseinside a draft forward. Draft all-reduces fall back to the next all-reduce backend (AITER custom or custom all-reduce), which does no quantization. This also covers AITER's fused quick-reduce + RMSNorm path, which uses the same check.Validation
Unit tests (
.venv+ this worktree on MI355X; GPU needed only for the AITER test):test_initialize_model_scopes_only_draft_modelschecks that the scope is active while a draft is built and run. It also checks that the target, and an ngram config that aliases the target, never get the scope.test_aiter_moe_a4w4_dsv4_skips_draft_modelchecks that a4w4 is on outside the scope and off inside it.test_quick_allreduce_rejects_draft_model_inputchecks that an INT4 quick reduce of an eligible size is declined inside the scope.End-to-end: pending. Planned on MI355X: DeepSeek-V4.1-Flash with DSpark (k=5), TP4, a4w4 + INT4 quick reduce, using the InferenceX GSM8K harness (lm-eval chat completions, 5-shot). I'll compare real acceptance length and throughput with and without this PR. For reference, InferenceX evals of the same setup measured AL 3.94 with neither opt-in, 3.89 with INT4 quick reduce, and 3.90 with a4w4 + INT4 quick reduce (InferenceX#3690).
Details
Motivation. A lower-precision draft never changes outputs, because rejection sampling corrects it. It only lowers acceptance, so evals don't catch it, and the cost shows up only as lost speculative speedup. Both opt-ins are global today. On ROCm the DSpark draft reuses
DeepseekV4DecoderLayer, so a4w4 also turns the draft's MoE into FP4 activations, and the draft's attention/MoE all-reduces go through quick reduce's INT4 codec once batches pass the size threshold. Benchmark rules that require drafts to run as shipped (e.g. InferenceX's draft-precision check, which rejected InferenceX#3697 for this) then rule out these knobs entirely, even though their target-side gain is legitimate.Design.
initialize_modelidentifies the draft with the existing idiommodel_config is speculative_config.draft_model_config, also used by_resolve_and_verify_engram_config. It excludes the case where that object is the target's config (ngram/custom proposers).FusedMoEConfig.__post_init__, and that config is built only at layer construction.forwardis wrapped in the scope, because quick reduce is chosen at runtime, and at capture time under CUDA graphs. Proposers and graph wrappers (CUDAGraphWrapper,BreakableCUDAGraphWrapper) all call the draft module, so the wrappedforwardalways runs. The wrapper sits outside any compiled inner model, so the flag is set when the all-reduce op runs.set_current_vllm_configinvllm.config, which bothvllm.distributedand the MoE config already import.Tradeoffs and limitations.
@support_torch_compile, itsforwardisn't wrapped, since Dynamo would trace the wrapper. Its construction is still scoped. Current drafts compile only their inner model, and ROCm DeepSeek V4/V4.1 drafts aren't compiled.in_draft_model().Pull Request Checklist
I used vLLM's
/pr-checklistskill. (Mandatory for agents, optional for humans).AI assistance was used during the creation of this PR.
Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.
Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI. (E2E pending; see Validation.)
Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.
Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.
Not a duplicate: open-PR searches for quick reduce / a4w4 / draft precision found nothing addressing this.
AI assistance: prepared with an AI coding agent (Cursor); the human submitter must review every changed line before marking this ready.