Repository navigation
[Perf] Optimize DeepSeek V4.1 Flash Hopper paths and Blackwell prefill selection - #41251
Conversation
|
Here is the launch config used for the results in this PR, at SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=0 \
PYTHONPATH="$PWD/python" \
python -m sglang.launch_server \
--model-path /path/to/DeepSeek-V4.1-Flash \
--served-model-name dsv41 \
--trust-remote-code \
--tp 8 \
--ep-size 1 \
--mem-fraction-static 0.8 \
--attention-backend dsv4 \
--moe-runner-backend marlin \
--enable-decoder-swa-bounded-replay \
--cuda-graph-max-bs-decode 64 \
--reasoning-parser auto \
--tool-call-parser auto \
--host 0.0.0.0 \
--port 30000Checkpoint revision: Benchmark: Request body: {"chat_template_kwargs":{"thinking":false},"stream_options":{"include_usage":true}}The key settings to match are EP1 + Marlin + GPU-resident Engram. The benchmark-infra H200 recipe currently specifies EP8 + |
42b9b99 to
f3476d4
Compare
Motivation
DeepSeek-V4.1-Flash on Hopper spends substantial decode time in group32 dense GEMMs, capacity-sized indexer work, mHC statistics, and padded Marlin MoE projections. This PR tunes those paths. On Blackwell, flattened prefill selection also launches separate mask, sort, slot-map and raw-offset operations; the added integer kernel combines them.
Changes
Pinned validation
Checkpoint
dba1be0a40aa45a94ad051997016db3960a90277; PyTorch 2.13.0+cu130 and sgl-kernel 0.4.8; sgl-bench commita9da34ad1f997ca05878d858d3d01970e4a49af9; sgl-eval 0.1.2. Both machines were allocated exclusively for these comparisons. All paired datasets, request counts, input/output lengths and API errors were validated from the request records. Timings exclude profiling and correctness instrumentation.H200: upstream main versus original PR runtime
8×H200, TP8/EP1, Marlin MoE, static memory fraction 0.8, bounded SWA replay, decode graph maximum batch64, Engram host table disabled. Main
ef867fa40dversus PR22e352d873;94d21e3162differs from the measured PR only by a test mock accessor. Random8192 input /1024 output, seed42, temperature0, thinking disabled,64 warmup requests and cache flush. Request counts30/40/80/80/160/320 at concurrency1/4/8/16/32/64; three alternating paired sweeps. Ratios below are medians of the three paired ratios.The final Blackwell addition was checked separately against the original PR on the same H200 node: one complete paired six-point sweep and full GSM8K/AIME24 at4096/32768 output caps. User-speed ratios are0.9990–1.0029 and output-throughput/GPU ratios0.9989–1.0013. This additional check is one paired sweep, not the three-sweep main comparison above. Full GSM8K is1277/1319 on both arms; AIME24 is456/480→457/480, with23 truncations each. Hopper continues to use its Torch fallback; equal aggregate scores do not imply identical generations.
GB300: added flattened prefill fusion versus ordinary PR
4×NVIDIA GB300/SM103, ARM64, TP4/EP1, flashinfer_mxfp4 MoE and flashinfer_cutedsl FP8 GEMM. Ordinary PR
94d21e3162versus candidate179b1ab2d3(identical tested file hashes). Both arms use16384 for chunked-prefill-size and max-prefill-tokens. Three rotated paired sweeps; seed42, temperature0, thinking disabled,16 warmups and cache flush. Each shape uses64 requests at concurrency1 and128 at concurrency16.Both concurrency16 shapes improve in all three pairs. Each concurrency1 shape has one regressing pair, so this is a small prefill improvement, not a claim of universal or statistically proven acceleration. In the8192-input/1024-output suite, median output-throughput/GPU ratios are only1.0009–1.0081. The standalone selection helper improves5.45–14.33×; that microbenchmark is not an end-to-end speedup.
Accuracy
Full GSM8K1319 problems, thinking disabled; full AIME24 is30 problems×16 repeats, thinking enabled. Seed42, temperature0,32 evaluation threads, no API errors. Repeats are not480 independent questions. Results are paired by question ID, problem text and expected answer.
GB300 accuracy uses2048 prefill budgets, because upstream main's FlashMLA scheduling metadata exceeds its48KB shared-memory implementation limit on a GSM8K batch at the default16K budget. Performance comparisons use the same16K budget on every arm and are kept separate. AIME24 candidate versus ordinary PR changes by−0.625 percentage points; truncations40→42,14 correct→wrong and11 wrong→correct. There is also one additional wrong nontruncated answer. The higher-cap paired run still has AIME24 −0.625 percentage points (456→453 correct; truncations22→24;7 correct→wrong,4 reverse), and GSM8K −0.1516 percentage points (1279→1277; neither arm truncated). These scores do not prove identical generation or statistical accuracy equivalence.
Correctness coverage
_finish_flat_indexer_topk_kernelevents per rank; all four original trace files are retained with source hashes.Memory and weight updates
The Hopper derived caches trade GPU memory for throughput. In the original weight-load measurement, memory increased72.71→75.43GB/GPU and max_total_num_tokens fell16,012,800→14,223,104 at the same static memory fraction. These are the original configuration's measurements, not newly measured final-commit memory numbers.
Online weight updates require disabling both derived caches with
SGLANG_OPT_HOPPER_BLOCK_FP8_BF16=0andSGLANG_OPT_DEEPGEMM_HC_PRENORM=0; otherwise updates fail before writing weights. Other TP/EP layouts are not covered by the reported serving comparisons.CI States
Latest PR Test (Base): ✅ Run #37035950158
Latest PR Test (Extra): ❌ Run #37035949974
Latest PR Test (AMD ROCm 10): ❌ Run #37035950474