Skip to content

Fix SWA cache loc slicing for all attention backends - #29460

Merged
merrymercy merged 2 commits into
mainfrom
lmzheng/fix-swa-trtllm-mha
Jun 28, 2026
Merged

merrymercy merged 2 commits into
mainfrom
lmzheng/fix-swa-trtllm-mha

Conversation

@merrymercy

@merrymercy merrymercy commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Move swa_out_cache_loc slicing into KVWriteLoc.__post_init__ so it applies to all attention backends, not just TRT-LLM MHA.
  • swa_out_cache_loc is computed once at metadata-init time from the full (possibly padded) out_cache_loc. Piecewise CUDA graphs later narrow out_cache_loc to real_num_tokens per layer via radix_attention.py, but swa_out_cache_loc on the metadata is never narrowed. This causes a shape mismatch when KVWriteLoc bundles them together for set_kv_buffer.
  • The fix auto-slices swa_loc to match loc's length at KVWriteLoc construction time, fixing the issue for all 8 backends (flashinfer, triton, flashattention, trtllm_mha, aiter, xpu, torch_native, musa).

Original commits

  • 53f21e2fe

CI States

Latest PR Test (Base): ❌ Run #28273505004
Latest PR Test (Extra): ❌ Run #28273504905

Slice swa_out_cache_loc to match cache_loc length before passing to
set_kv_buffer. Without this, the SWA location tensor can be longer than
the KV tensors when sliding-window attention is active, causing shape
mismatches.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@merrymercy

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added blackwell SM100/SM120 run-ci CI: run the baseline test suite on this PR labels Jun 26, 2026
Move the swa_out_cache_loc slicing from the TRT-LLM MHA backend into
KVWriteLoc.__post_init__ so it applies to all attention backends.
@merrymercy merrymercy changed the title Fix SWA cache loc slicing in TRT-LLM MHA backend Fix SWA cache loc slicing for all attention backends Jun 27, 2026
@merrymercy
merrymercy merged commit e6cbc8f into main Jun 28, 2026
158 of 177 checks passed
@merrymercy
merrymercy deleted the lmzheng/fix-swa-trtllm-mha branch June 28, 2026 03:38
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

blackwell SM100/SM120 run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant