Skip to content

Revert "[AMD] Fix DeepSeek V4 Pro c128 state tensor dtype mismatch error and c4_sparse_raw_indices attribute error in cuda graph phase" - #27919

Merged
HaiShaw merged 1 commit into
sgl-project:mainfrom
At1a8:revert-27529-fangyuan/fix_c128_kv_dtype_mismatch
Jun 11, 2026
Merged

HaiShaw merged 1 commit into
sgl-project:mainfrom
At1a8:revert-27529-fangyuan/fix_c128_kv_dtype_mismatch

Conversation

@At1a8

@At1a8 At1a8 commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Reverts #27529, since it cause


CI States

Latest PR Test (Base): ❌ Run #27345325269
Latest PR Test (Extra): ❌ Run #27345324822

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request simplifies the DeepSeek-V4 JIT compression kernels (c128 and c4) by removing the BufFloat template parameter and unifying the KV buffer type with the input type (InFloat). This removes redundant template parameters, simplifies memory loading and writing routines, and defers float casting to the computation phase. The Python bindings and compressor layers are updated accordingly to remove the unused buffer dtype parameter and associated casting hotfixes. No review comments were provided for this pull request.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@HaiShaw
HaiShaw merged commit 6e885c8 into sgl-project:main Jun 11, 2026
86 of 98 checks passed
@At1a8
At1a8 deleted the revert-27529-fangyuan/fix_c128_kv_dtype_mismatch branch June 12, 2026 04:11
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
…ror and c4_sparse_raw_indices attribute error in cuda graph phase" (sgl-project#27919)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants