Repository navigation
[qwen 3.8 next] Fuse hybrid speculative state commits across state pools - #41171
Open
Qiaolin-Yu wants to merge 1 commit into
Open
Qiaolin-Yu wants to merge 1 commit into
Qiaolin-Yu wants to merge 1 commit into
Conversation
Qiaolin-Yu
marked this pull request as ready for review
September 24, 2026 21:32
Qiaolin-Yu
requested review from
BBuf,
DarkSharpness,
HaiShaw,
HydraQYH,
celve,
hanming-lu,
hebiao064,
yizhang2077 and
yuan-luo
as code owners
September 24, 2026 21:32
This was referenced Sep 24, 2026
7 of 10 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Target verification commits accepted GDN, convolution, and PLE states through multiple launches. Combine compatible state pools and active/tracking writes while preserving tracking overwrite precedence.
Modifications
Cache validated state geometry with storage/layout invalidation, preserve integer state bits and strided convolution windows, and retain unsupported-path fallbacks.
This PR is independently based on
main. QSA KV preparation is tracked separately in #40972.Accuracy Tests
Added focused tests for this change:
Independent branch test rerun: 33 passed, 6 subtests passed on NVIDIA B200. Each PR was tested from its own
main-based source snapshot.Combined implementation validation with official
sgl-eval0.1.0, temperature 1, top-p 1, maximum 131072 tokens, 8 repeats over 30 questions, concurrency 128, thinking enabled, no acceptance simulation:Real MTP acceptance length was approximately 2.146, including the bonus token, estimated by weighting the complete rounded log intervals by request count.
Audited all 8 prediction files per mode against native metrics, generation settings and token totals. These are integration results with the optimization series enabled, not an isolated accuracy A/B for this PR.
Speed Tests and Profiling
Benchmark scope: the E2E results below were measured with the full optimization series combined. Reproducing that implementation requires all PRs listed below, including this PR. They are not standalone performance results for this PR.
Full optimization series used for integration measurements
Each split PR is independently based on
main; the list above describes the combined benchmark source, not a required merge order. The combined AIME26 results in Accuracy Tests also use this full series (with acceptance simulation disabled).Evidence specific to this change
At serving-sized batch size 1, the state-commit microbenchmark decreased from approximately 10.00 to 5.63 microseconds without tracking and 10.88 to 7.75 microseconds with tracking. The active/tracking consolidation alone did not produce a measurable E2E throughput improvement in the cumulative run.
The combined E2E table below must not be used to infer this PR's individual throughput gain.
Full-series E2E performance (all PRs above combined)
Hardware and workload: NVIDIA B200 x4 / TP4,
nvidia/Qwen3.8-Flash-Next-NVFP4, 8192 input / 1024 output tokens, 5 requests at concurrency 1, plan stream disabled:Forward occupancy was collected with SGLang's device timer. Per-run values were 98.90%, 98.91%, and 98.89% for normal decode, and 97.75%, 97.79%, and 97.80% for NEXTN. These are full-series integration measurements with plan stream disabled; the NEXTN runs use simulated acceptance length 3.3, as in the performance table.
Every run completed 5 requests and generated 5120 output tokens. Decode throughput is 1000 / mean TPOT in milliseconds; the table reports the median of three runs. NEXTN performance uses simulated acceptance length 3.3. These measurements do not establish an isolated main-versus-this-PR speedup.
Checklist
CI States
Latest PR Test (Base): ❌ Run #36056860534
Latest PR Test (Extra): ❌ Run #36056860291
Latest PR Test (AMD ROCm 10): ❌ Run #36056860516