Update stable FA3 API (flash_api_stable.cpp) to be inline with the unstable ABI port (flash_api.cpp) - #152
Conversation
61d7af3 to
c993fa0
Compare
| std::optional<int> page_size, | ||
| int max_seqlen_k_new, // 0 means we're not appending new KV | ||
| int64_t batch_size, | ||
| int64_t max_seqlen_q, |
There was a problem hiding this comment.
StableIValues API does not support to<int> but does support to<int64_t> so we update the signatures here as well as the boxed kernels boxed_mha_... below. They get narrowed back to int when assigned to params.<var> in this function below.
| torch::headeronly::ScalarType qkv_dtype, | ||
| Tensor seqused_k, // b | ||
| std::optional<Tensor> cu_seqlens_q_, // b+1 | ||
| std::optional<Tensor> cu_seqlens_k_, // b+1 |
There was a problem hiding this comment.
Similarly, StableIValue doesn't have specialization for std::optional<const Tensor&> so it is changed here to std::optional. Tensor data is never copied though as Tensor is essentially just a pointer with a reference count.
|
|
||
| auto opts = q.options(); |
There was a problem hiding this comment.
No longer needed since we use new_empty below.
| // if (out_.has_value()) { out_.value().copy_(out); } | ||
| } | ||
|
|
||
| return {out, softmax_lse}; | ||
| } | ||
|
|
||
| #ifndef FLASHATTENTION_DISABLE_PYBIND | ||
| void boxed_mha_fwd( |
There was a problem hiding this comment.
Everything below here mirrors the the registration in hopper/flash_api_torch_lib.cpp
| float softmax_scale, | ||
| int window_size_left, | ||
| int window_size_right, | ||
| int attention_chunk, |
There was a problem hiding this comment.
attention_chunk was removed in flash_api.cpp
| params.deterministic = deterministic; | ||
| } | ||
|
|
||
| template <int Arch, int Split, bool PagedKVNonTMA, bool PackGQA, bool Has_softcap> |
There was a problem hiding this comment.
Major refactoring here to align with flash_api.cpp
| "Tensor q," | ||
| "Tensor k," | ||
| "Tensor v," | ||
| "Tensor(k_new!)? k_new = None," |
There was a problem hiding this comment.
The default parameters here have been removed to align better with legacy code (see total file changes) and in particular, so that softmax_scale would be a required parameter like it is in flash_api.cpp.
…_stable.cpp. Updated flash_api_stable.cpp to be inline with the current flash_api.cpp and flash_api_torch_lib.cpp. And updated static_switch to use a stdexcept instead of a c10/util/Exception.h because importing that header was casuing one unstable symbol (from TORCH_CHECK). Signed-off-by: Chris Leonard <chleonar@redhat.com>
…_api.cpp. Returned the CMakeLists.txt and static_switch.h files back to main, will migrate them in a different PR. Signed-off-by: Chris Leonard <chleonar@redhat.com>
… to see in github. Will move back in next commit Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
…iff is easy to see in github. Will move back in next commit Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
c827b79 to
9252523
Compare
|
IIUC, are you actually porting our changes to FA3 (in our fork) to the stable API implementation (from upstream code)? |
* Fused Bwd (vllm-project#137) * Fused with Good perf and stride fixed Fix fused bugs isolate failing case fix bug bring back test cases rm split impl in fused use exp2 is global variable now try oom fix save make fused the default limit to reproduce failure return default to split fix head size bug use exp2 back to true * new grid * BLK_SLICE_FACTOR = 1 * add tflops * new commit * test in parrallel * strides added by jusson * disable alibi * fix bugs again * default to fused * add bwd options for varlen * backend filter * default to jingning and batch 4 * best fwd config * fix TRITON_PRINT_AUTOTUNING flag bug * tune * Tuning fwd prefill * add if else * use flag * Minor mask fix * FLIP GRID * use best config for default * print when autotuning * test bfloat16 * fix k and v stride bugs * skip bfloat16 * test kvpacked * disable internal tests * pick default config based on arch * Add alibi in the new bwd kernel (vllm-project#139) * enable alibi for jinging kernel enable alibi for jinging kernel match * save bad configs * fix alibi and causal bug * disable autotune by default * auto tune when benching is good * set best config * remove env var * Update amd_tests.yml * upgrad to triton==3.3.0 * increase shm * use 64 x 64 for now * save * handle 1d alibi * Add fp8 to fused kernel (vllm-project#140) * fp8 stuff find test case compute delta fp8 basic fp8 config passing non causal path works * isolate bad case * fix fp8 bug * didnot fix fp8 bug * back to failing test * fp8 tests passing * skip * skip ref tests --------- Co-authored-by: Aliasger Zaidy <aliasger.zaidy@amd.com> * head, seq, batch (vllm-project#141) * Fix keys (vllm-project#144) * save * rm keys * fix keys * use GHA_RENDER_DEVICES * normal docker * Pad LSE (vllm-project#148) * add round multiple * fix fwd * backward fix * use rounded lse flag * passing ROUNDED_LSE * default is new rounded mode * rename to fused_atmoics and fused_no_atomics * add test for torch_compile * add varlen torch compile test * add old one kernel for ref * fix varlen mismatch bug * fix shape issue in varlen but mismatch * sync torch compile kernel launch * simple varlen test * add debug code * rm old * ignore old impls * DEBUG flag works in interface only * ref uses the righ shape for lse * rm oldest bwd kernel * fix typo * fix varlen bug * fix bug. Get info from q for now * simple shape and stride checkout * add more tests * test kvcache * kvcache safe * match case * fix segfault due to bad return_softmax * run bench * run seperate for the main functions * just output benchmark * default csv format and time stamp files * non verbsoe bench * Sliding Window Forward (vllm-project#151) * Compress SWA work test case set up debug inputs add fwd ref one mask ref fwd first pass save ref doesnot work for bigger seqlens save new version some causal cases failing found bad cases working new attn new atten works new attn_fwd works reorg n_extra_tokens use seqlen_delta_qk ref fwd works add sliding window to bwd ref test kvcache decode ref work with everything except sliding window add debug code for 12 failing sliding window cases for decode attention_decode_forward_ref_impl mostly works except for alibi fix alibi in attention_decode_forward_ref_impl ref works with normal, varlen & kvcache move stuff around figure out masking old attn inner two inner functions remove load_fn do Lk - Lq like ref unify IS_CAUSAL code in epilogue clean up add args rm inference stuff simplify compute_masking simpler compute mask stub out returning front masking variables remove pointer pass compute ptrs inloop compute block min and max window stub inside inner mask loop trying to use attn_fwd_mask causes issues fix compiler bug when front masking gen specifc types add sliding window and debug statements use identity for v add more taste cases add comments save use k_max_token for clarity disable debug configs basic NON-CAUSAL SLIDING WINDOW non causal sliding window works on the all the shapes non sliding window working in fwd clean up fused bwd seperate old fwd_prefill move configs to utils.py * fix bwd ref bug * skip local cases so that fa output * no sliding window causal green * add backward test skip for sliding window * clean reduce in fwd_kvcache. no is_CASUAL branching * add kvcache masking * kvcache working * fix some bugs in test.py * clean up * Fix Device Segfault (vllm-project#152) * Compress segfault work fix backward segfault rework offset ignore .profile ignore .analysis save * assert the kernel launch device and tensor devices are the same * fix failing asserts * add asserts to fwd * Fix SDMASK bug * Log triton, torch and fa version * Fix fp8 import issues * fix docs (vllm-project#154) * Sliding Window block classification logic (vllm-project#155) * add aiter code * remove aiter stuff * sliding window non causal masking works * causal and sliding window block masking * extract common * clean up typo * helper for swa * ignore .amd * fix last block bug * Enable FA V3 (vllm-project#157) * Compress PA work narrow pa test ref works on most cases inplace ref with new_kv inplace paged attention add pa ref save pa basic paged works save fix swa + causal in pa. Also new_kv only on pa path passing build fa v3 import interface from fa v3 copy fa tests use v3 api clean up rename to match old test support different head sizes remove fp8 basisc passing v3 cases test_flash_attn_varlen_output v3 working isolate bad case for kvcache case passing save use decode is seqused/ cacheseql is given use decode if not varlen basci kvcache v3 working kvcache enable more cases detect kvcache case if seqused_q is non and sequese_k is not None skip failing test find fp8 failing case mha fp8 works fix fp8 MQA/GQA bug clean up more clean up clean up more don't need fp8 dead code remove train code with fp8 stuff fp8 working in kvcache paged + fp8 seems to be working new_kv allowed * clean up * skip hopper race test * clean up more * fix paged + alibi * similar inner paged api * unify _attn_fwd_inner * AITER integration (vllm-project#159) * clean up v2 interface * assert fp8 scale shapes * rotary working * move rotary to impl layers * remove einops * enable rotarry in v3 * create interface * fix descale assert * unify bwd * lint from aiter * clean fp8 api * add api change * assert shapes for v2 * remove ref and bench.py * remove metadata class and clean up * bwd_prefill * one bwd.py * rename * lint * add bwd_change (vllm-project#156) * Tune FP8 Perf (vllm-project#160) * check cu count for gfx942 * create get_cu_count * update repo root * update forward tune * clean up load * use float8_e4m3fnuz * save * show bwd mode * recommend fp8 * use torch.float32 for fp8 kernel * add both best fp16 and fp8 config * tune fp8 backward * descale factors should be b, hk * fp8 bwd working on all primus configs * tune bwd configs * fa v3 tests passing * better warning * clean up bwd launcher * v3 passing * tune more * improve perf * clean up * lint * clean * start tuning gfx950 * tune non causal path * fix bug * save * Skip configs where BLOCK_M2 % BLOCK_N2 != 0 * skip more * stop tuning * fix varlen bug * fix dropout & causal/swa segfault * update the to machine new changes * save * fix more bugs * remove random seed * clean up * update readme * print tensor stats for debug * disable sliding window tests * add rdna configs * fix k partial bug * fix block_size_n bug * fix type check bug --------- Co-authored-by: Aliasger Zaidy <aliasger.zaidy@amd.com> Co-authored-by: Tianxing Wu <tianxing.wu@amd.com>
@Harry-Chen Yes, the stable API that was created in upstream got out of sync with the current changes in vLLM's FA3 API, so this PR is to get them back in sync. |
Great. Could you also create a PR in the vllm repo to change the ref? (I'm not sure whether it could fetch a commit from a PR). |
@Harry-Chen Once this commit lands and I update the My original plan was: PR 1: Sync with the unstable API. I can combine the two PRs in this repository if that would be easier. However, I think it would be preferable for vllm to reference a commit that has already landed on the main branch here, which means this PR would need to merge before the corresponding vllm change. If we'd like to validate everything in vllm before merging this PR, I could also open a vllm PR that points directly to this branch and use that to run CI. After verification, we could merge this PR, update the vllm PR to reference the merged commit on main, and then merge the vllm change. What do you think about this? Hopefully this makes sense, but please let me know if you need me to clear anything up. |
Yes this is actually what I meant. I do not have permission in this repo, so I need to find other maintainers and show "hey this is compiling and working as expected" to get it merged, then in the main repo. I think one PR in this repo is totally fine, since there should not be other users other than vllm itself. |
…in static_switch to change TORCH_CHECK to STD_TORCH_CHECK because some files in the binary were using it and it caused one unstable symbol Signed-off-by: Chris Leonard <chleonar@redhat.com>
Incorporates vllm-project#160 (revert of vllm-project#152, the torch stable-ABI port that broke FLASH_ATTN_MLA_SPARSE). Net change: hopper stable-ABI files only; no cute-DSL / FP8 changes.
| @@ -2,11 +2,7 @@ | |||
| * Copyright (c) 2024, Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao. | |||
There was a problem hiding this comment.
This commit was created just to show the diff between flash_api.cpp (old) and flash_api_stable.cpp (updates)
Updated
flash_api_stable.cppto be inline with the currentflash_api.cppandflash_api_torch_lib.cpp.To view the diff between
flash_api_stable.cppandflash_api.cpp, checkout commit 7da1d8c, where theflash_api_stable.cppwas copied toflash_api.cppin order to see the diff (i.e.cp flash_api_stable.cpp flash_api.cpp)Note: The changes in
flash_api_stable.cppare to align it withflash_api.cppso looking at the commit above is probably more useful than looking at the Files Changed.A follow-up PR to update
CMakeLists.txtto use the flash_api_stable instead of flash_api is planned.cc @Harry-Chen @janeyx99
vLLM Test Plan/Result
(built using
flash_api_stable.cppinstead offlash_api.cppandflash_api_torch_lib.cpp)Test Result
The same test are passing on this branch that are passing on main (and the same failures too)
Migration progress of vLLM using the Audit Python extension torch-abi-audit: