Skip to content

[Spec Decode][ROCm] Keep draft models at checkpoint precision under a4w4 and quantized quick reduce - #59884

Draft
Fangzhou-Ai wants to merge 1 commit into
vllm-project:mainfrom
Fangzhou-Ai:rocm-draft-shipped-precision
Draft

Fangzhou-Ai wants to merge 1 commit into
vllm-project:mainfrom
Fangzhou-Ai:rocm-draft-shipped-precision

Conversation

@Fangzhou-Ai

@Fangzhou-Ai Fangzhou-Ai commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Keep speculative-decoding draft models at their checkpoint precision when two ROCm opt-ins that lower precision are enabled: VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4=1 (FP4 MoE activations) and VLLM_ROCM_QUICK_REDUCE_QUANTIZATION (quantized TP all-reduce). The target model behaves as before.

Claims

  • Draft models load and run inside a new draft_model_scope(). The scope is applied once, in initialize_model, so it covers every proposer on both model runners (MRV1 and MRV2), including CUDA-graph capture and breakable CUDA graphs.
  • FusedMoEConfig no longer enables a4w4 for draft MoE layers. DeepSeek V4.1's DSpark draft keeps the checkpoint's a8w4 (MXFP8 activations), and the target still runs a4w4.
  • QuickAllReduce.should_quick_allreduce returns False inside a draft forward. Draft all-reduces fall back to the next all-reduce backend (AITER custom or custom all-reduce), which does no quantization. This also covers AITER's fused quick-reduce + RMSNorm path, which uses the same check.
  • Without these opt-ins nothing changes: no other code reads the scope.

Validation

Unit tests (.venv + this worktree on MI355X; GPU needed only for the AITER test):

pytest tests/model_executor/model_loader/test_mtp_validation.py \
  tests/distributed/test_rocm_quick_reduce.py \
  tests/kernels/moe/test_rocm_aiter_moe.py -k "a4w4 or draft or scopes or mtp"
# 17 passed
  • test_initialize_model_scopes_only_draft_models checks that the scope is active while a draft is built and run. It also checks that the target, and an ngram config that aliases the target, never get the scope.
  • test_aiter_moe_a4w4_dsv4_skips_draft_model checks that a4w4 is on outside the scope and off inside it.
  • test_quick_allreduce_rejects_draft_model_input checks that an INT4 quick reduce of an eligible size is declined inside the scope.
  • With either gate reverted, the matching test fails.

End-to-end: pending. Planned on MI355X: DeepSeek-V4.1-Flash with DSpark (k=5), TP4, a4w4 + INT4 quick reduce, using the InferenceX GSM8K harness (lm-eval chat completions, 5-shot). I'll compare real acceptance length and throughput with and without this PR. For reference, InferenceX evals of the same setup measured AL 3.94 with neither opt-in, 3.89 with INT4 quick reduce, and 3.90 with a4w4 + INT4 quick reduce (InferenceX#3690).

Details

Motivation. A lower-precision draft never changes outputs, because rejection sampling corrects it. It only lowers acceptance, so evals don't catch it, and the cost shows up only as lost speculative speedup. Both opt-ins are global today. On ROCm the DSpark draft reuses DeepseekV4DecoderLayer, so a4w4 also turns the draft's MoE into FP4 activations, and the draft's attention/MoE all-reduces go through quick reduce's INT4 codec once batches pass the size threshold. Benchmark rules that require drafts to run as shipped (e.g. InferenceX's draft-precision check, which rejected InferenceX#3697 for this) then rule out these knobs entirely, even though their target-side gain is legitimate.

Design.

  • Detection. initialize_model identifies the draft with the existing idiom model_config is speculative_config.draft_model_config, also used by _resolve_and_verify_engram_config. It excludes the case where that object is the target's config (ngram/custom proposers).
  • Construction. The scope covers draft construction. a4w4 is decided in FusedMoEConfig.__post_init__, and that config is built only at layer construction.
  • Forward. The draft's forward is wrapped in the scope, because quick reduce is chosen at runtime, and at capture time under CUDA graphs. Proposers and graph wrappers (CUDAGraphWrapper, BreakableCUDAGraphWrapper) all call the draft module, so the wrapped forward always runs. The wrapper sits outside any compiled inner model, so the flag is set when the all-reduce op runs.
  • Placement. The scope lives next to set_current_vllm_config in vllm.config, which both vllm.distributed and the MoE config already import.

Tradeoffs and limitations.

  • Draft speed. Draft all-reduces lose quick reduce's bandwidth savings. Draft reductions are small, about 50 KiB per running request per step for DeepSeek V4.1, so the effect should be minor, and end-to-end numbers will confirm it.
  • Compiled top-level drafts. If a draft's top-level class is itself @support_torch_compile, its forward isn't wrapped, since Dynamo would trace the wrapper. Its construction is still scoped. Current drafts compile only their inner model, and ROCm DeepSeek V4/V4.1 drafts aren't compiled.
  • Scope of knobs. Only these two opt-ins check the scope. Other knobs that lower precision can opt in later with in_draft_model().

Pull Request Checklist
  • I used vLLM's /pr-checklist skill. (Mandatory for agents, optional for humans).

  • AI assistance was used during the creation of this PR.

  • Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.

  • Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI. (E2E pending; see Validation.)

  • Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.

  • Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.

Not a duplicate: open-PR searches for quick reduce / a4w4 / draft precision found nothing addressing this.

AI assistance: prepared with an AI coding agent (Cursor); the human submitter must review every changed line before marking this ready.

Run speculative draft construction and forward inside draft_model_scope(),
applied once in initialize_model for every proposer. The ROCm a4w4 MoE opt-in
(VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4) and quick reduce's quantized all-reduce
now skip draft models, so they cannot lower acceptance; the target is
unchanged.

Co-authored-by: Cursor Agent
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

quantization rocm Related to AMD ROCm

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

1 participant