Repository navigation
Conversation
Co-authored-by: Codex <noreply@openai.com> Signed-off-by: ahmed xijiaat <52128022+xijiaat@users.noreply.github.com>
Contributor
Author
|
Closing this draft after H20 (SM90) validation. The current implementation is not a performance win in the tested shapes, and its original numerical suite is not fully passing.
Serving dispatch was never enabled. I will retain the branch and experimental results, but will not pursue integration of this version without a materially better implementation. Thank you. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Draft a compensated BF16 mHC prenorm primitive for Hopper, adapted from SGLang's
hc_mix_stats_bf16x3in sgl-project/sglang#39664 and sgl-project/sglang#41251. It splits the FP32 mixing weight into three BF16 components, accumulates their projections separately in FP32, and computes the input squared norm in the same pass.This initial draft contains a kernel, correctness tests and a comparison benchmark. Serving dispatch is intentionally not wired yet. No kernel or end-to-end speedup, numerical equivalence, or production-readiness claim is made. Integration and default selection depend on real SM90 measurements against the current fused mHC path, not only against a standalone GEMM.
Implementation
[splits, tokens, mix]projection and[splits, tokens]squared-norm partials, compatible with the existing prenorm epilogue layout.The prepared weights occupy 6 bytes per original FP32 element, in addition to any retained original weight. The caller must rebuild them after weight updates. Different accumulation orders can change outputs; batch invariance is not promised. SGLang's bundled serving gains are not an estimate for this isolated kernel or for vLLM.
Duplicate check and provenance
Checked #57448 and public open/closed PR searches for
mHC,compensated,bf16x3,SM90, and Hopper on 2026-10-09. No open PR implementing this compensated mHC projection was found.The SGLang kernel's Apache-2.0 attribution is retained. Reviewed source:
python/sglang/kernels/ops/layernorm/mhc.pyon SGLang main around5d8e98b7fdd83907616b1ff8ca68229f2ae95b34. This adapts its numerical approach to vLLM's output layout and caller-owned workspace; it does not copy SGLang's runtime dispatch or Sinkhorn implementation.Validation completed
CPU-only Linux container, Python 3.12.3, PyTorch 2.13.0+cu130, Triton 3.7.1. Source under test was overlaid into an isolated source checkout; dependencies came from the existing vLLM v0.28.0 image, not a newly built wheel. Container used
runc, no GPU device nodes, emptyCUDA_VISIBLE_DEVICES,NVIDIA_VISIBLE_DEVICES=void, no network, 2 CPU cores and 4 GiB RAM.torch.cuda.is_available()was false.VLLM_TARGET_DEVICE=cpu PYTHONPATH=$PWD PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 \ .venv/bin/python -m pytest --confcutdir=tests/kernels/mhc -o addopts= \ -q tests/kernels/mhc/test_mhc_kernels.py -k bf16x33 passed, 9 GPU cases skipped, 320 deselected. CPU checks cover random, cancellation and scaled FP32 weight decomposition, including preserving the original weight. Source-only version/extension warnings and torch deprecation warnings were emitted.
18 SM90 offline compilations passed: explicit
GPUTarget("cuda", 90, 32), K=5120/20480, splits=1/4/16, BLOCK_M=32/64/128, N=24, BLOCK_N=32, BLOCK_K=64, four warps and three stages. PTX/cubin were generated without loading or launching a CUDA kernel. This proves compilation only, not GPU numerical correctness or speed.All applicable pre-commit hooks passed, including Ruff, mypy, SPDX and Buildkite test tethering;
git diff --checkpassed. The existing mHC CI job exercises the suite on B200; actual SM90 execution remains a validation gap.Before ready for review
benchmarks/kernels/benchmark_mhc_bf16x3.pyon H20/H200 and retain all configurations, not only winners.AI assistance: implemented and checked with OpenAI Codex. This draft was opened at the human author's explicit request to stage the implementation while GPU resources are being arranged. GPU execution, model evaluations, serving benchmarks and human review are pending; this is not ready to merge.