Repository navigation
perf(cuda): adaptive split-KV sizing for attention_row decode (+4-5% over fixed-256; +70% deep-ctx) - #1350
Merged
background
wait
wait-all
cancel
parallel
Loading