Skip to content

llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 - #25370

Merged
fairydreaming merged 3 commits into
ggml-org:masterfrom
fairydreaming:dsv4-remove-kq-bias
Jul 10, 2026
Merged

llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4#25370
fairydreaming merged 3 commits into
ggml-org:masterfrom
fairydreaming:dsv4-remove-kq-bias

Commits

Commits on Jul 6, 2026

Commits on Jul 7, 2026