Skip to content

llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized - #25871

Merged
ggerganov merged 3 commits into
ggml-org:masterfrom
fairydreaming:force-fa-quant-kv-cache
Jul 31, 2026
Merged

llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized#25871
ggerganov merged 3 commits into
ggml-org:masterfrom
fairydreaming:force-fa-quant-kv-cache