Skip to content

[qwen 3.8 next] Fuse NEXTN verify and draft graph input preparation - #41175

Merged
Qiaolin-Yu merged 5 commits into
mainfrom
qiaolin/qwen-next-verify-prep
Oct 2, 2026
Merged

Qiaolin-Yu merged 5 commits into
mainfrom
qiaolin/qwen-next-verify-prep

Conversation

@Qiaolin-Yu

@Qiaolin-Yu Qiaolin-Yu commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

NEXTN transitions construct many small integer tensors for verify commit and draft replay. Fuse commit outputs, accepted-step metadata, chain/cache indices, and draft length/position/MRoPE preparation.

Modifications

Use the compact unread QSA verify-mask contract, preserve stream lifetime registration and custom-position/multimodal fallbacks, and pass precomputed outputs through the existing batch/result structures.

This PR is independently based on main. QSA KV preparation is tracked separately in #40972.

Accuracy Tests

Added focused tests for this change:

python -m pytest -q test/registered/unit/layers/test_next_verify_prep.py

Independent branch test rerun: 203 passed on NVIDIA B200. Each PR was tested from its own main-based source snapshot.

Combined implementation validation with official sgl-eval 0.1.0, temperature 1, top-p 1, maximum 131072 tokens, 8 repeats over 30 questions, concurrency 128, thinking enabled, no acceptance simulation:

Mode Correct / total Accuracy Length truncations Request errors
Normal decode 229/240 95.42% 9 0
NEXTN, 3 steps / top-k 1 / 4 draft tokens 228/240 95.00% 8 0

Real MTP acceptance length was approximately 2.146, including the bonus token, estimated by weighting the complete rounded log intervals by request count.

Audited all 8 prediction files per mode against native metrics, generation settings and token totals. These are integration results with the optimization series enabled, not an isolated accuracy A/B for this PR.

Speed Tests and Profiling

Benchmark scope: the E2E results below were measured with the full optimization series combined. Reproducing that implementation requires all PRs listed below, including this PR. They are not standalone performance results for this PR.

Full optimization series used for integration measurements

  • #40972: Fuse QSA KV preparation and sparse block expansion
  • #41166: Fuse small CUDA graph input buffer copies
  • #41167: Reduce repeated overlap batch snapshot work
  • #41168: Skip speculative KV allocation transfers when no pages are needed
  • #41169: Fuse simulated speculative acceptance tensor preparation
  • #41170: Fuse speculative relay payload stores and avoid single-request gathers
  • #41171: Fuse hybrid speculative state commits across state pools
  • #41172: Fuse exact greedy chain verification with argmax finalization
  • #41173: Fuse QSA graph replay metadata across draft steps
  • #41174: Fuse Mamba tracking index lookup and skip uniform-position transfers
  • #41175: Fuse NEXTN verify and draft graph input preparation (this PR)

Each split PR is independently based on main; the list above describes the combined benchmark source, not a required merge order. The combined AIME26 results in Accuracy Tests also use this full series (with acceptance simulation disabled).

Evidence specific to this change

Removes multiple graph-external metadata launches. Individual cumulative stages had small, mixed E2E changes; no single isolated main-versus-PR speedup is claimed.

The combined E2E table below must not be used to infer this PR's individual throughput gain.

Full-series E2E performance (all PRs above combined)

Hardware and workload: NVIDIA B200 x4 / TP4, nvidia/Qwen3.8-Flash-Next-NVFP4, 8192 input / 1024 output tokens, 5 requests at concurrency 1, plan stream disabled:

Mode Median TPOT across 3 runs Decode throughput derived from TPOT Median fwd occupancy across 3 runs
Normal decode 4.1270 ms 242.30 token/s 98.90%
NEXTN, 3 steps / top-k 1 / 4 draft tokens 1.6936 ms 590.44 token/s 97.79%

Forward occupancy was collected with SGLang's device timer. Per-run values were 98.90%, 98.91%, and 98.89% for normal decode, and 97.75%, 97.79%, and 97.80% for NEXTN. These are full-series integration measurements with plan stream disabled; the NEXTN runs use simulated acceptance length 3.3, as in the performance table.

Every run completed 5 requests and generated 5120 output tokens. Decode throughput is 1000 / mean TPOT in milliseconds; the table reports the median of three runs. NEXTN performance uses simulated acceptance length 3.3. These measurements do not establish an isolated main-versus-this-PR speedup.

Checklist

  • Pre-commit formatting and checks pass.
  • Focused regression tests included.
  • Independent branch test rerun complete.
  • Current combined accuracy validation complete.

CI States

Latest PR Test (Base): 🚫 Run #36941656114
Latest PR Test (Extra): ❌ Run #36941655865
Latest PR Test (AMD ROCm 10): ⏳ Run #36941656069

@YAMY1234 YAMY1234 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, with two non-blocking nits below.

Comment thread python/sglang/srt/speculative/eagle_worker_common.py Outdated
Comment thread python/sglang/kernels/ops/speculative/eagle.py
@Qiaolin-Yu Qiaolin-Yu added run-ci CI: run the baseline test suite on this PR bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) labels Oct 1, 2026
@Qiaolin-Yu

Copy link
Copy Markdown
Collaborator Author

@Qiaolin-Yu

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/kernels/ops/attention/qsa/test_qsa_fused_kv_prepare.py

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/kernels/ops/attention/qsa/test_qsa_fused_kv_prepare.py:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/kernels/ops/attention/qsa/test_qsa_fused_kv_prepare.py

@Qiaolin-Yu

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/e2e/models/test_qwen4_exp_models.py

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/e2e/models/test_qwen4_exp_models.py:

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/models/test_qwen4_exp_models.py

@Qiaolin-Yu
Qiaolin-Yu merged commit 67eab57 into main Oct 2, 2026
98 of 159 checks passed
@Qiaolin-Yu
Qiaolin-Yu deleted the qiaolin/qwen-next-verify-prep branch October 2, 2026 00:17
nvpohanh added a commit to ajit283/sglang that referenced this pull request Oct 2, 2026
sgl-project#41175 made `_draft_extend_for_decode` read
`batch_result.prepared_draft_extend_inputs`, so the mocked batch result in
test_eagle_draft_sampling.py raised AttributeError after merging main.
Set it to None so the test takes the CPU fallback path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) jit-kernel run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants