Skip to content

Fix DeepSeek V4 checkpoint loading and mixed quantization - #1

Merged
machiabeli merged 9 commits into
machiabeli:feat/deepseek-v4from
Thump604:deepseek-v4-support-fixes
Apr 24, 2026
Merged

machiabeli merged 9 commits into
machiabeli:feat/deepseek-v4from
Thump604:deepseek-v4-support-fixes

Conversation

@Thump604

Copy link
Copy Markdown

This adds checkpoint-driven fixes on top of the DeepSeek V4 support branch.

Changes:

  • Match the DeepSeek V4 RoPE path more closely, including inverse output rotation for the post-attention rope dims.
  • Pass attention sinks into SDPA.
  • Decode e8m0 block scales for FP8 and FP4 paths.
  • Unpack routed expert FP4 weights instead of treating the int8 tensors as ordinary int8 weights.
  • Fix ratio-4 overlap compressor shapes from the real checkpoint, for example ape (4, 1024) and wkv/wgate (1024, 4096).
  • Make mixed_3_6 preserve DeepSeek-specific attention, compressor, indexer, embedding, shared-expert, and output-critical paths at higher precision.

Validation:

PYTHONPATH=/Users/David/code/mlx-lm-deepseek-v4-review \
  /opt/ai-runtime/venv-live/bin/python -m pytest -q tests/test_models.py
# 80 passed, 1 skipped, 3 warnings, 51 subtests passed

I am leaving this draft because full checkpoint conversion/load validation is still pending, and sparse compressed attention plus learned indexer parity still needs review against the reference implementation before this should be treated as complete production support.

@Thump604

Copy link
Copy Markdown
Author

I pushed the follow-up branch updates here: https://github.com/Thump604/mlx-lm/tree/deepseek-v4-support-fixes

New commits since my last note:

  • 6908736 casts DeepSeek V4 attention sinks to the query dtype before MLX SDPA. Without this, quantized generation fails with Type of sinks must promote to output type bfloat16.
  • 9c990f4 fixes quantized grouped wo_a output projection. The dense path can reshape wo_a.weight, but a quantized QuantizedLinear stores packed input columns, so the output projection must slice output rows per group and call mx.quantized_matmul instead of reshaping packed weights as dense tensors.

Validation on the branch:

PYTHONPATH=/Users/David/code/mlx-lm-deepseek-v4-review /opt/ai-runtime/venv-live/bin/python -m pytest -q tests/test_models.py tests/test_tokenizers.py
87 passed, 1 skipped, 5 warnings, 57 subtests passed

Conversion evidence from the official deepseek-ai/DeepSeek-V4-Flash revision 6e763230a9d263eca2023f1d4a5ce1bfe126cf48:

  • Q3 mixed artifact uploaded: https://huggingface.co/Thump604/DeepSeek-V4-Flash-MLX-Q3-mixed-gs128-affine
  • Q3 recipe: mixed_3_6, affine, group size 128, effective 3.808 bpw, 28 shards, indexed tensor size 135,346,422,876 bytes
  • Q3 lazy-loads on a 128 GB Mac Studio, but I do not consider it generation-qualified locally. A one-token smoke crossed my memory safety boundary with heavy swap activity and was killed.
  • Q2 mixed artifact uploaded: https://huggingface.co/Thump604/DeepSeek-V4-Flash-MLX-Q2-mixed-gs128-affine
  • Q2 recipe: mixed_2_6, affine, group size 128, effective 2.992 bpw, 23 shards, indexed tensor size 106,355,393,628 bytes
  • Q2 raw generation smoke completed with --max-tokens 2 --max-kv-size 1024: output the name, prompt 4 tokens at 7.488 tok/s, generation 2 tokens at 19.182 tok/s, 54.59s real, 74.5GB max RSS, 106.94GB peak footprint, zero swaps.

I still treat full sparse compressed-attention/indexer parity as unproven. The branch now has conversion, lazy-load, and Q2 smoke-generation evidence, but not a production-quality or long-context claim.

@machiabeli
machiabeli merged commit 45665f8 into machiabeli:feat/deepseek-v4 Apr 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants