Repository navigation
Conversation
Carry the twelve model-specific selections from sgl-project/sglang#39186, include their tuner input, and provide a graph-replay comparison against the shipped fallback. Five sampled batches improve operator latency; full-model validation and the remaining buckets still need review. Signed-off-by: Kevin Mi <mikevin920@yahoo.com>
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
Measured on MI355X (DeepSeek-V4.1-Flash TP4, DSpark EP4) with AITER_BF16_FP8_MOE_BOUND=0, as the AMD DeepSeek-V4 tests and cookbook set it: - the stage-1 LDS-DMA drain: GSM8K 0.885 unpatched vs 0.905 patched (200 questions), throughput within noise, and its race test passes unpatched; it stays upstream as ROCm/aiter#5561; - the bf16 SiLU route is never taken at bound 0; the fix is ROCm/aiter#5802; - the tuned FMoE CSV shows no end-to-end gain; it stays upstream as ROCm/aiter#5562. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@kevin-mii Measurements for this CSV on MI355X, plus three tiers it is missing. Setup. What changes. Without the CSV every tier logs Decode tok/s change (all requests decoding), 95% CI:
So at batch 1-8 without speculation it is within noise, matching your 86ab3ad1bc measurement. At larger batches, and with DSpark even at B=8 (verify M = 6 x B = 48), it is worth about 4%. Missing tiers. DSpark verify M at larger batches falls into tiers 1024/2048, which the CSV does not cover, so those still use the default. I tuned them with
With these rows added, DSpark at B=256 (verify M = 768 -> tier 1024) decodes +3.6% [+3.5, +3.7] faster than with the current CSV (end to end +1.5%). B=8/64/128 are unchanged (their tiers already exist). Could you add these rows to this PR? The tier-16 row's absolute timing comes from the tuner's synthetic routing and is not comparable to the existing rows. The choice was only validated through the end-to-end runs above. |
Motivation
Upstream the DSV4.1 Flash EP4 a8w4 FMoE rows carried by sgl-project/sglang#39186. Without this table the exact gfx950 shape falls back to heuristic FlyDSL kernels. Removing the downstream CSV reduced performance in the operator comparison below, so it is restored downstream while this PR is reviewed.
Technical Details
aiter/configs/model_configs/, plus the corresponding untuned shape input.op_tests/op_benchmarks/bench_dsv41_fmoe.pyfor a reproducible default-versus-tuned graph comparison. It excludes the candidate CSV from the baseline even when running from this branch and preserves the other model tables.AITER_CONFIG_FMOE_FILEmerging; no matching shape keys exist in the current shipped model/default tables.The rows and their
us/ error columns are imported from the original downstream tuning data; those columns are historical, not the measurements below. This PR is based onmainat0138f88b2.Test Plan
For the measurements below, apply the stage-1 synchronization fix in #5561 to this branch, then run:
The benchmark uses seeded synthetic FP4 weights and bf16 activations. Timing is the median of five GPU-event measurements, each covering 100 graph replays after warmup. Both configurations use the same synchronization fix and inputs. The Triton environment flag is a local toolchain workaround.
Test Result
MI350X VF (gfx950), ROCm 7.2.4, PyTorch 2.11.0+rocm7.2:
A separate one-token smoke rerun measured 47.17 -> 37.61 us (1.25x). Black, Ruff, CSV uniqueness and normal config merge checks pass.
Numerical limitation: the benchmark checks finiteness and reports differences; it is not an independent accuracy oracle. Tuned versus default maximum absolute differences were 0.0273-0.0410 on the synthetic inputs. An initial strict comparison at 32 tokens (
atol=0.02, rtol=0.05) rejected 13/163840 elements. The two paths differ in intermediate quantization/reduction, so this requires review rather than asserting equivalence.Not executed: a fresh exhaustive tuning sweep, full-model accuracy/throughput, the other seven buckets, or other architectures. Keep this PR in draft until numerical tolerances, all buckets, and the paired synchronization fix are reviewed. These operator timings are not an end-to-end model speedup.
A second comparison on the installed SGLang-pinned AITER
4ad998328also favored the CSV: M=1: 43.05 -> 37.17 us (1.16x), M=32: 291.22 -> 262.54 us (1.11x), M=128: 343.24 -> 316.97 us (1.08x), M=512: 419.93 -> 399.23 us (1.05x), M=4096: 1359.34 -> 1239.05 us (1.10x). This installation does not include the synchronization patch; the same installation was used for both sides. These are a8w4 measurements withAITER_BF16_FP8_MOE_BOUND=0, not the default small-batch bf16 route.Submission Checklist