[AMD] Add unified kv attention support in dpsk-v4 - #27380
Conversation
Sync unified-KV attention branch with main (109 commits). Conflicts resolved in: - deepseek_v4_backend_hip_radix.py: keep unified methods (_attach_unified_kv_decode_streams/_forward_unified_kv) AND main's new get_swa_out_cache_loc (sgl-project#27091 unified full->SWA translation). - deepseek_v4.py: non-unified SWA-store now uses backend.get_swa_out_cache_loc (main sgl-project#27091) instead of the removed pool get_cached_swa_loc; unified path keeps its own ring addressing. - deepseek_v4_memory_pool.py: keep get_unified_kv; drop _should_cache_swa/ cached_loc pool caches (removed by main sgl-project#27091). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
/rerun-test registered/attention/unittests/dsv4/test_deepseek_v4.py |
|
Results for 🚀 🚀 |
This comment was marked as outdated.
This comment was marked as outdated.
|
@amd-bot ci-status |
This comment was marked as outdated.
This comment was marked as outdated.
|
@amd-bot ci-status |
CI Status for PR #27380Merge verdict: No failing job is caused by this PR — every red X is infrastructure (missing Caution The new unified-KV attention path (the whole point of this PR — Changed files: 16 files, +2418/−84 — all under AMD: 0 real failures (2 infra/cascade) · Others: 0 related failures (7 unrelated/infra/cascade) — all AMD stage-b test shards (27 jobs) passed; stage-c was gate-skipped. AMD CI Failures
Other CI Failures
Details / what to do before merge
Generated by amd-bot using Claude Code CLI |
Reverts d214a07. The KVBlockSize==1 assert was caused by the wrong container image, not the aiter indexer. The box now runs rocm/sgl-dev:v0.5.13-rocm720-mi35x-20260612 (triton 3.6.0, torch 2.9.1), which supports the aiter JIT gluon block-64 paged-mqa-logits kernel. Returning to the proven PR sgl-project#27380 config.
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>


Co-authored-by: @RolaoDenthu
Motivation
Add a unified-KV attention backend for DeepSeek-V4 on amd code path, porting ATOM's sparse attention kernels, which has great perf. It is gated by
SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton; when disabled, the default path is unchanged.Modifications
Memory layout: change to a new class
DeepSeekV4UnifiedKVPool.Compressed-KV store with unified kv layout
SWA-KV store with unified kv layout
Attention kernels
Per-forward metadata
Accuracy Tests
GSM200:
Speed Tests and Profiling
Server cmd:
Client cmd:
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #27188164054
Latest PR Test (Extra): ❌ Run #27188163931