Repository navigation
[qwen 3.8 next] Fuse NEXTN verify and draft graph input preparation - #41175
Merged
Merged
Conversation
Qiaolin-Yu
marked this pull request as ready for review
September 24, 2026 21:33
Qiaolin-Yu
requested review from
BBuf,
DarkSharpness,
Fridge003,
HaiShaw,
HydraQYH,
Ying1123,
celve,
hebiao064,
hnyls2002,
ispobock,
merrymercy,
xiezhq-hermann and
yuan-luo
as code owners
September 24, 2026 21:33
This was referenced Sep 24, 2026
7 of 10 tasks
YAMY1234
approved these changes
Sep 30, 2026
YAMY1234
left a comment
Collaborator
There was a problem hiding this comment.
LGTM, with two non-blocking nits below.
Collaborator
Author
|
/rerun-test test/registered/kernels/ops/attention/qsa/test_qsa_fused_kv_prepare.py |
Contributor
|
Results for 🚀 |
Collaborator
Author
|
/rerun-test test/registered/e2e/models/test_qwen4_exp_models.py |
Contributor
|
Results for 🚀 |
6 tasks done
nvpohanh
added a commit
to ajit283/sglang
that referenced
this pull request
Oct 2, 2026
sgl-project#41175 made `_draft_extend_for_decode` read `batch_result.prepared_draft_extend_inputs`, so the mocked batch result in test_eagle_draft_sampling.py raised AttributeError after merging main. Set it to None so the test takes the CPU fallback path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
NEXTN transitions construct many small integer tensors for verify commit and draft replay. Fuse commit outputs, accepted-step metadata, chain/cache indices, and draft length/position/MRoPE preparation.
Modifications
Use the compact unread QSA verify-mask contract, preserve stream lifetime registration and custom-position/multimodal fallbacks, and pass precomputed outputs through the existing batch/result structures.
This PR is independently based on
main. QSA KV preparation is tracked separately in #40972.Accuracy Tests
Added focused tests for this change:
Independent branch test rerun: 203 passed on NVIDIA B200. Each PR was tested from its own
main-based source snapshot.Combined implementation validation with official
sgl-eval0.1.0, temperature 1, top-p 1, maximum 131072 tokens, 8 repeats over 30 questions, concurrency 128, thinking enabled, no acceptance simulation:Real MTP acceptance length was approximately 2.146, including the bonus token, estimated by weighting the complete rounded log intervals by request count.
Audited all 8 prediction files per mode against native metrics, generation settings and token totals. These are integration results with the optimization series enabled, not an isolated accuracy A/B for this PR.
Speed Tests and Profiling
Benchmark scope: the E2E results below were measured with the full optimization series combined. Reproducing that implementation requires all PRs listed below, including this PR. They are not standalone performance results for this PR.
Full optimization series used for integration measurements
Each split PR is independently based on
main; the list above describes the combined benchmark source, not a required merge order. The combined AIME26 results in Accuracy Tests also use this full series (with acceptance simulation disabled).Evidence specific to this change
Removes multiple graph-external metadata launches. Individual cumulative stages had small, mixed E2E changes; no single isolated main-versus-PR speedup is claimed.
The combined E2E table below must not be used to infer this PR's individual throughput gain.
Full-series E2E performance (all PRs above combined)
Hardware and workload: NVIDIA B200 x4 / TP4,
nvidia/Qwen3.8-Flash-Next-NVFP4, 8192 input / 1024 output tokens, 5 requests at concurrency 1, plan stream disabled:Forward occupancy was collected with SGLang's device timer. Per-run values were 98.90%, 98.91%, and 98.89% for normal decode, and 97.75%, 97.79%, and 97.80% for NEXTN. These are full-series integration measurements with plan stream disabled; the NEXTN runs use simulated acceptance length 3.3, as in the performance table.
Every run completed 5 requests and generated 5120 output tokens. Decode throughput is 1000 / mean TPOT in milliseconds; the table reports the median of three runs. NEXTN performance uses simulated acceptance length 3.3. These measurements do not establish an isolated main-versus-this-PR speedup.
Checklist
CI States
Latest PR Test (Base): 🚫 Run #36941656114
Latest PR Test (Extra): ❌ Run #36941655865
Latest PR Test (AMD ROCm 10): ⏳ Run #36941656069