Repository navigation
Conversation
BBuf
requested review from
DarkSharpness,
HaiShaw,
HydraQYH,
celve and
yuan-luo
as code owners
September 17, 2026 08:56
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
On 4×GB300, the small DeepSeek-V4.1 Flash target router occasionally has long latency while overlapping with mHC statistics and shared-expert work. Using ordinary stream dependencies for this router improves BS=1 decode throughput without changing routing arithmetic or removing those overlap streams.
Modifications
Limit the launch-policy change to SM103,
sqrtsoftplus, 384 experts, top-6, and 1–8 rows. BS=1 with DSPARK block size 5 verifies 6 target rows. The draft router and larger batches keep their existing PDL behavior.The existing
use_pdlvalue controls bothlaunch_pdland the kernel'sgdc_wait/gdc_launch_dependentsinstructions. Set it to false for the measured shape, so all three are disabled together and CUDA's ordinary stream ordering supplies the dependency. Scoring, top-k selection, tie-breaking, normalization and packing are unchanged. No new flag, environment variable, or CUDA stream.This is an independent PR on
dsv4.1, not stacked on #39704.Accuracy Tests
On GB300, the full
test/registered/kernels/ops/moe/test_moe_fused_gate.pyfile passes: 52 passed, including 6 new graph-replay cases. The new cases compare PDL and ordinary launches bitwise for weights, expert IDs and packed outputs, with changing in-graph producers/consumers, FP32/BF16 bias inputs, equal-score ties and varying live/padded rows. The test can exercise both launch variants on Hopper CI by mocking only the dispatch architecture; Triton still targets the actual GPU.PYTHONPATH="$PWD/python" CUDA_VISIBLE_DEVICES=0 \ python -m pytest -q test/registered/kernels/ops/moe/test_moe_fused_gate.pyFull GSM8K/AIME/GPQA evaluations have not been rerun for this PR. Kernel equality is not presented as a new task-accuracy result. Pre-commit checks pass for both touched files.
Speed Tests and Profiling
Base:
00d7d516d047292263aff7ed3dd2449668d0f7c7(dsv4.1). Candidate:f4a4da0c16bdc4cd0749890126aceb6c0a2e4f5a(this PR only).Pooled median improvement: +3.62%. Individual-run medians improve by 2.94% and 4.36%. All four runs have median reported acceptance length 5.5054 (range 5.4468–5.5956).
The 40-request ranges are 1131.47–1216.70 tps for base and 1159.78–1251.30 tps for this PR. These results describe this fixed synthetic workload; no BS=32/64 gain is claimed for this small-row policy.
Hardware/software: 4×NVIDIA GB300 (SM103), TP4 / EP1; original DeepSeek-V4.1 Flash checkpoint; PyTorch 2.13.0+cu130, Triton 3.7.1, FlashInfer 0.6.18, sglang-kernel 0.4.7, sgl-deep-gemm 0.2.0, CUTLASS DSL 4.6.2.
BS=1, fixed seed-42 random-token input of 4096 tokens, 1024 output tokens, temperature 0 and ignore-EOS. DSPARK block size 5; acceptance is simulated with target length 5.5, not measured natural acceptance. Each fresh server has 1 discarded warm-up request and 20 measured requests; all measured requests are included. Order: base → candidate → candidate → base. No profiler during speed measurements.
Throughput is measured by the streaming client:
(final completion tokens − first event tokens) / (last event timestamp − first event timestamp). This excludes TTFT and the first streamed token group; it is not a server-log throughput value or whole-request throughput. The same prompt, cache flushes, launch parameters and dependencies are used for both variants.Server command and exact client workload
Run from either checkout with the same environment;
MODELpoints to the original Flash checkpoint.Save the following as
client_repro.py, then runpython client_repro.py "$MODEL". Restart the server for each repeat. The input-ID SHA256 assertion detects a mismatched tokenizer/workload.Checklist
CI States
Latest PR Test (Base): ❌ Run #35301367356
Latest PR Test (Extra): ❌ Run #35301367205
Latest PR Test (AMD ROCm 10): ❌ Run #35301367271