Skip to content

[Quant] Serve 32-wide-K ue8m0 block-FP8 linears through the FlashInfer MXFP8 GEMMs - #40039

Merged
hnyls2002 merged 8 commits into
mainfrom
lsyin/dsv4-quant-block-fp8-mxfp8
Sep 18, 2026
Merged

hnyls2002 merged 8 commits into
mainfrom
lsyin/dsv4-quant-block-fp8-mxfp8

Conversation

@hnyls2002

@hnyls2002 hnyls2002 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Block-FP8 weights with a 32-wide K block and ue8m0 scales are MXFP8 operands; Fp8LinearMethod serves them through the FlashInfer CUTLASS / CuTe-DSL MXFP8 GEMMs when that backend is selected, accepts a prequantized Mxfp8SwizzledInput, and dispatches the block kernel on the effective block size. No behavior change for existing FP8 checkpoints; consumer is #38798.


CI States

Latest PR Test (Base): 🚫 Run #35330783063
Latest PR Test (Extra): ❌ Run #35330782851
Latest PR Test (AMD ROCm 10): 🚫 Run #35330783022

@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026
@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_flashinfer_trtllm_fp8_fallback.py test_fp8_utils.py test_fp8_blockwise_linear_backends.py

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_flashinfer_trtllm_fp8_fallback.py test_fp8_utils.py test_fp8_blockwise_linear_backends.py:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/layers/quantization/test_flashinfer_trtllm_fp8_fallback.py

🚀 1-gpu-h100 (2 tests): ❌ View workflow run

cd test/ && python3 registered/quant/test_fp8_utils.py
cd test/ && python3 registered/unit/layers/quantization/test_fp8_blockwise_linear_backends.py

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/layers/quantization/test_fp8_blockwise_linear_backends.py

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/layers/quantization/test_fp8_blockwise_linear_backends.py

@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_fp8_utils.py

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_fp8_utils.py:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/quant/test_fp8_utils.py

@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_hicache_storage_lora.py

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_hicache_storage_lora.py:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/hicache/test_hicache_storage_lora.py

…k-fp8-mxfp8

# Conflicts:
#	python/sglang/srt/layers/quantization/fp8.py
@hnyls2002
hnyls2002 merged commit 1b200ff into main Sep 18, 2026
75 of 108 checks passed
@hnyls2002
hnyls2002 deleted the lsyin/dsv4-quant-block-fp8-mxfp8 branch September 18, 2026 09:51
hnyls2002 added a commit that referenced this pull request Sep 18, 2026
zozyo pushed a commit to Phala-Network/sglang that referenced this pull request Sep 22, 2026
…r MXFP8 GEMMs (sgl-project#40039)

(cherry picked from commit 1b200ff)
(cherry picked from commit e533126)
mickqian added a commit to TyGu888/sglang that referenced this pull request Oct 6, 2026
Resolve the conflict with the online MXFP8 branch from sgl-project#37903: unserialized
checkpoints still use MXFP8OnlineLinearMethod, serialized ones now get
ComfyMXFP8LinearMethod.

Adapt the override to the scale_u8 argument that sgl-project#40039 added to
Fp8LinearMethod._process_mxfp8_linear_weight_scale. Only scales read from a
comfy checkpoint skip the interleave; an explicit scale_u8 is converted from
block-FP8 in row-major order and still goes through SRT. Update the unit test
expectations accordingly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek npu run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant