Repository navigation
dsv4.1: mHC computation and compensated projections - #39664
Merged
Merged
Conversation
hnyls2002
requested review from
BBuf,
DarkSharpness,
HaiShaw,
HydraQYH,
celve and
yuan-luo
September 15, 2026 22:30
hnyls2002
added this pull request to stack #39667
September 15, 2026 22:32
This was referenced Sep 15, 2026
hnyls2002
removed this pull request from stack #39667
September 15, 2026 22:43
hnyls2002
added this pull request to stack #39669
September 15, 2026 22:43
hnyls2002
removed this pull request from stack #39669
September 15, 2026 22:56
hnyls2002
force-pushed
the
dsv4.1-hopper
branch
from
September 15, 2026 22:57
5cf6459 to
fe97034
Compare
hnyls2002
requested review from
Fridge003,
Qiaolin-Yu,
hebiao064,
ispobock and
merrymercy
September 15, 2026 22:57
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 15, 2026 22:57
e6a5980 to
124c47f
Compare
hnyls2002
added this pull request to stack #39672
September 15, 2026 22:58
hnyls2002
force-pushed
the
dsv4.1-hopper
branch
from
September 15, 2026 23:38
fe97034 to
e005f05
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 15, 2026 23:38
124c47f to
e996f4e
Compare
hnyls2002
force-pushed
the
dsv4.1-hopper
branch
from
September 16, 2026 00:27
e005f05 to
7f33ed3
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 00:27
e996f4e to
63f7d97
Compare
hnyls2002
force-pushed
the
dsv4.1-hopper
branch
from
September 16, 2026 01:12
7f33ed3 to
5112b92
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 01:12
63f7d97 to
111265c
Compare
hnyls2002
force-pushed
the
dsv4.1-hopper
branch
from
September 16, 2026 01:52
5112b92 to
0834e67
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 01:52
111265c to
6dcb649
Compare
hnyls2002
force-pushed
the
dsv4.1-hopper
branch
from
September 16, 2026 03:45
0834e67 to
d225e3b
Compare
…_quant next to the kernels; align compress kernel names
… from hc_mult; share the slice count
Collaborator
Author
|
/rerun-test test/registered/kernels/ops/layernorm/test_mhc_kernels.py test/registered/kernels/ops/layernorm/test_hc_combine.py |
Contributor
|
Results for 🚀 |
This was referenced Sep 17, 2026
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
mhc.py:hc_mix_statscomputes the per-row mixing projections and sum of squares over a fixed number of K slices withtf32x3dots, so a row's result is bitwise identical whether computed alone or in a batch;hc_mix_stats_sinkhornfuses the slice reduction with the sinkhorn normalisation thathc_split_sinkhornperforms.hc_mix_stats_sinkhorn_bf16x3(Triton, three bf16 weight components, one activation read) andhc_mix_stats_sinkhorn_deepgemm(tf32 high part plus fp32 residual through DeepGEMM), withsplit_bf16_hc_weight/split_tf32_hc_weightpreparing the weight parts. Both accumulate over_HC_MIX_COMPENSATED_SLICESand share the sinkhorn reduce kernel; sizes come from the input shapes andhc_mult.hc_split_sinkhornanswers an empty batch directly instead of launching a zero-sized grid, which the TileLang backend rejects.Changes to existing kernels
hc_split_sinkhorn: the empty-batch return only. Its existing caller already returns before reaching it with no tokens, so results on existing models are unchanged; the guard is for callers that do not pre-check.Verification
hc_split_sinkhornport and the reference implementation.test_mhc_kernels.pyandtest_hc_combine.pycover the untouched kernels of the module.CI States
Latest PR Test (Base): ❌ Run #35165253652
Latest PR Test (Extra): ❌ Run #35165253438
Latest PR Test (AMD ROCm 10): ❌ Run #35165253625