Skip to content

Add kernels for DeepSeek v4.1 - #221

Merged
interestingLSY merged 2 commits into
mainfrom
v41-open-source
Sep 10, 2026
Merged

interestingLSY merged 2 commits into
mainfrom
v41-open-source

Conversation

@interestingLSY

Copy link
Copy Markdown
Collaborator

No description provided.

- Restructure csrc into csrc/kernels (per-architecture / per-operation tree) with
  hierarchical namespaces and a .cpp API layer, and rename MODEL1 to DeepSeek-V4.
- Sync the vendored kerutils library with the upstream version.
- Add the sm100 sparse prefill kernels (h_q = 64 / 128, d_qk = 512 / 576).
- Add the sm100 sparse decode kernel for h_q = 64 with the DeepSeek-V4.1 KV cache
  formats (V41 fp8 and the V41_FP4 extra cache) and the shared dequantization helpers.
- Add the fused norm + RoPE + attn + RoPE + cast operator and its permute kernels.
- Build sm_100a / sm_103a targets and register the new sources.
- Extend the tests with the V4.1 KV cache layouts and the fused kernel test, and
  document the new KV cache formats in the README and the Python API.
@interestingLSY
interestingLSY merged commit 07a1089 into main Sep 10, 2026
zyongye added a commit to zyongye/FlashMLA that referenced this pull request Sep 10, 2026
Sync deepseek-ai/FlashMLA main through commit
07a1089, including the DeepSeek V4.1
kernels from deepseek-ai#221 and the reorganized csrc/kernels tree.

The upstream history is represented as one commit because the target repository's
DCO check requires a valid sign-off on every commit in the pull request.

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
lucifer1004 added a commit to lucifer1004/flashinfer that referenced this pull request Sep 10, 2026
DeepSeek-V4.1-Flash stores each KV token as 528B: 512B of FP8 E4M3 covering
the full 512-wide K (rope lanes quantized, no BF16 rope segment) plus a 16B
footer of 16 UE8M0 scales over 32-wide groups (the deepseek-ai/FlashMLA#221
layout, SM100-only there). Add ModelType::DSV4_1 on top of the ScaleSpec
groundwork: GLM53_NOPE-style geometry (D_ROPE=0) composed with a 32-wide
UE8M0 footer.

Kernel side the format is simpler than DSV4 (one 512B bulk per row, a single
unified block-scaled FP8 QK pass, no rope MMA paths). The two new mechanics:

- the 16B footer row generalizes the scale gather from one uint64_t to a
  width-selected load (uint4) in both the prefill IO helper and the decode
  kernel's staged gather;
- the 32-wide groups constrain the XV W-fold (one scale group per W buffer),
  so an 8-warp XV split would floor NT_PER_WARP_XV to 0: decode runs a
  4-math-warp tile at BI=64 (the DOTS3_SWA halving precedent) and prefill is
  SG-only on the BI=32 producer/consumer tile. The ComputeTraits/SmemLayout
  aliases now take the XV warp count so the NT_PER_WARP_XV assert evaluates at
  the warp count the XV MMA actually runs at.

Selection is explicit everywhere (d_qk=512 collides with DSV4, the 528B
payload with GLM53_NOPE): the functional API gains kv_scale_format=
"ue8m0_g32", the TRTLLM-compat entry kv_cache_format="fp8_dsv41", and the
decode-dsv4 FFI takes an explicit model_type (-1 keeps the legacy width
inference, matching the paged-entry kAuto convention). Dual-cache decode is
supported like DSV4; dual-cache prefill (MG_DUAL) is intentionally not
instantiated. Calibration keys a new dsv4_1 family at topk 512.

Tests: decode matrix (dedicated/runtime-H heads, partial-tail topk, sinks),
SG prefill, dual-cache through the TRTLLM entry, poisoned-slot-0 NaN safety
on both gather paths, dispatch/plan/envelope coverage. 666 passed on RTX PRO
6000 (SM120).

Signed-off-by: lucifer1004 <13583761+lucifer1004@users.noreply.github.com>
zyongye added a commit to vllm-project/FlashMLA that referenced this pull request Sep 10, 2026
* Sync DeepSeek upstream through V4.1 kernels (deepseek-ai#221)

Sync deepseek-ai/FlashMLA main through commit
07a1089, including the DeepSeek V4.1
kernels from deepseek-ai#221 and the reorganized csrc/kernels tree.

The upstream history is represented as one commit because the target repository's
DCO check requires a valid sign-off on every commit in the pull request.

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>

* Restore stable ABI after upstream sync

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>

* Restore NVFP4 compatibility after upstream sync

Port the V3.2 NVFP4 cache path alongside upstream's generalized V4.1 kernels. Restore optional output forwarding and add quantization and API regression coverage.

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>

* Restore api.cpp comments and document the vLLM fork's changes

Address review on #20: keep the FLASH_MLA_ENABLE_DENSE_BWD and
PyInit__flashmla_C comments that the upstream sync dropped from
csrc/api/api.cpp, and add a README section summarizing what this fork
carries on top of deepseek-ai/FlashMLA (stable ABI, registered
operators, optional output buffers, SM90 dense FP8 extension, NVFP4 KV
cache, robustness fixes, tests).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>

* Add a detailed per-file diff against upstream to the README

Extend the "Changes in the vLLM fork" section with a per-area table of
every file that differs from deepseek-ai/FlashMLA@07a1089 and what
changed in it, plus the full diffstat and the command to regenerate it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>

---------

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant