Skip to content

[qwen 3.8 next] Fuse hybrid speculative state commits across state pools - #41171

Open
Qiaolin-Yu wants to merge 1 commit into
mainfrom
qiaolin/qwen-next-state-commit
Open

Qiaolin-Yu wants to merge 1 commit into
mainfrom
qiaolin/qwen-next-state-commit

Conversation

@Qiaolin-Yu

@Qiaolin-Yu Qiaolin-Yu commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Target verification commits accepted GDN, convolution, and PLE states through multiple launches. Combine compatible state pools and active/tracking writes while preserving tracking overwrite precedence.

Modifications

Cache validated state geometry with storage/layout invalidation, preserve integer state bits and strided convolution windows, and retain unsupported-path fallbacks.

This PR is independently based on main. QSA KV preparation is tracked separately in #40972.

Accuracy Tests

Added focused tests for this change:

python -m pytest -q test/registered/unit/layers/test_mamba_state_scatter_triton.py
python -m pytest -q test/registered/unit/layers/test_next_state_commit.py

Independent branch test rerun: 33 passed, 6 subtests passed on NVIDIA B200. Each PR was tested from its own main-based source snapshot.

Combined implementation validation with official sgl-eval 0.1.0, temperature 1, top-p 1, maximum 131072 tokens, 8 repeats over 30 questions, concurrency 128, thinking enabled, no acceptance simulation:

Mode Correct / total Accuracy Length truncations Request errors
Normal decode 229/240 95.42% 9 0
NEXTN, 3 steps / top-k 1 / 4 draft tokens 228/240 95.00% 8 0

Real MTP acceptance length was approximately 2.146, including the bonus token, estimated by weighting the complete rounded log intervals by request count.

Audited all 8 prediction files per mode against native metrics, generation settings and token totals. These are integration results with the optimization series enabled, not an isolated accuracy A/B for this PR.

Speed Tests and Profiling

Benchmark scope: the E2E results below were measured with the full optimization series combined. Reproducing that implementation requires all PRs listed below, including this PR. They are not standalone performance results for this PR.

Full optimization series used for integration measurements

  • #40972: Fuse QSA KV preparation and sparse block expansion
  • #41166: Fuse small CUDA graph input buffer copies
  • #41167: Reduce repeated overlap batch snapshot work
  • #41168: Skip speculative KV allocation transfers when no pages are needed
  • #41169: Fuse simulated speculative acceptance tensor preparation
  • #41170: Fuse speculative relay payload stores and avoid single-request gathers
  • #41171: Fuse hybrid speculative state commits across state pools (this PR)
  • #41172: Fuse exact greedy chain verification with argmax finalization
  • #41173: Fuse QSA graph replay metadata across draft steps
  • #41174: Fuse Mamba tracking index lookup and skip uniform-position transfers
  • #41175: Fuse NEXTN verify and draft graph input preparation

Each split PR is independently based on main; the list above describes the combined benchmark source, not a required merge order. The combined AIME26 results in Accuracy Tests also use this full series (with acceptance simulation disabled).

Evidence specific to this change

At serving-sized batch size 1, the state-commit microbenchmark decreased from approximately 10.00 to 5.63 microseconds without tracking and 10.88 to 7.75 microseconds with tracking. The active/tracking consolidation alone did not produce a measurable E2E throughput improvement in the cumulative run.

The combined E2E table below must not be used to infer this PR's individual throughput gain.

Full-series E2E performance (all PRs above combined)

Hardware and workload: NVIDIA B200 x4 / TP4, nvidia/Qwen3.8-Flash-Next-NVFP4, 8192 input / 1024 output tokens, 5 requests at concurrency 1, plan stream disabled:

Mode Median TPOT across 3 runs Decode throughput derived from TPOT Median fwd occupancy across 3 runs
Normal decode 4.1270 ms 242.30 token/s 98.90%
NEXTN, 3 steps / top-k 1 / 4 draft tokens 1.6936 ms 590.44 token/s 97.79%

Forward occupancy was collected with SGLang's device timer. Per-run values were 98.90%, 98.91%, and 98.89% for normal decode, and 97.75%, 97.79%, and 97.80% for NEXTN. These are full-series integration measurements with plan stream disabled; the NEXTN runs use simulated acceptance length 3.3, as in the performance table.

Every run completed 5 requests and generated 5120 output tokens. Decode throughput is 1000 / mean TPOT in milliseconds; the table reports the median of three runs. NEXTN performance uses simulated acceptance length 3.3. These measurements do not establish an isolated main-versus-this-PR speedup.

Checklist

  • Pre-commit formatting and checks pass.
  • Focused regression tests included.
  • Independent branch test rerun complete.
  • Current combined accuracy validation complete.

CI States

Latest PR Test (Base): ❌ Run #36056860534
Latest PR Test (Extra): ❌ Run #36056860291
Latest PR Test (AMD ROCm 10): ❌ Run #36056860516

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant