Skip to content

[Feat] FA4 compile API - #150

Merged
LucasWilkinson merged 2 commits into
vllm-project:mainfrom
LopezCastroRoberto:feat/compile_api
Jun 22, 2026
Merged

LucasWilkinson merged 2 commits into
vllm-project:mainfrom
LopezCastroRoberto:feat/compile_api

Conversation

@LopezCastroRoberto

Copy link
Copy Markdown

No description provided.

@LopezCastroRoberto
LopezCastroRoberto marked this pull request as ready for review June 19, 2026 16:52
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
Comment thread flash_attn/cute/interface.py Outdated
Comment thread flash_attn/cute/interface.py Outdated
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>

@LucasWilkinson LucasWilkinson left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM thanks for doing this!

@LucasWilkinson
LucasWilkinson merged commit 2c839c3 into vllm-project:main Jun 22, 2026
1 check passed
MatthewBonanni added a commit to MatthewBonanni/flash-attention that referenced this pull request Jul 7, 2026
Sync with Dao-AILab/flash-attention upstream (20 new commits, up to 5835c73).

Notable upstream changes:
- Parallelize splitkv alignment templated kernels, remove flag (Dao-AILab#2683, Dao-AILab#2680)
- [Cute,Bwd,Sm100] add sparse MLA (Deepseek v4) backward kernels (Dao-AILab#2621)
- Fix compatibility with CuTe DSL 4.6.0+ (Dao-AILab#2648, Dao-AILab#2676, Dao-AILab#2679, Dao-AILab#2684)
- SM120 Pack-GQA + graceful SplitKV fallback (Dao-AILab#2656, Dao-AILab#2671)
- hd256/sm100 stride/contiguity fixes (Dao-AILab#2670, Dao-AILab#2666, Dao-AILab#2686)
- int32 overflow fix in SM100 LPT tile scheduler (Dao-AILab#2662)
- _flash_attn_fwd now returns a 4-tuple (out, lse, p, row_max) (Dao-AILab#2674)
- q_subtile_factor default -> identity (Dao-AILab#2660); FP8 causal hd128 ex2_emu_freq tune (Dao-AILab#2642)
- [AMD ROCm] RDNA backward + CK unified workspace (Dao-AILab#2675); FA3 uv install (Dao-AILab#2458)

Conflict resolutions:
- Kept fork's FA3 (hopper/) entirely; dropped upstream's FA3 changes
  (Dao-AILab#2674 stable-API/hopper, Dao-AILab#2458 flash_attn_3 package) per downstream policy.
- setup.py: kept fork's CMake-based build.
- csrc CK (mha_bwd.cpp) + generate_kernels/launch_template: adopted upstream's
  splitkv-align + CK unified-workspace refactor, kept fork's fwd_sparse kernels.
- interface.py/flash_fwd.py: adopted upstream's _flash_attn_fwd 4-tuple return
  and q_subtile_factor=1 default while preserving the fork's output_scale FP8
  fused-quant, output_quant_key, compile API, and mDynamicCausal paths.
- cute_dsl_utils.py: kept both the fork's None-guard and vllm main's
  _cute_tensor fast path (vllm-project#150).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@MatthewBonanni MatthewBonanni mentioned this pull request Jul 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants