[bugfix]: bump FA4 pin to the CuTe DSL 4.6 compatible rev - #1564
Conversation
The 2026-07-05 CI image rebuild resolved the unpinned transitive
nvidia-cutlass-dsl to 4.6.0. The FA4 cute overlay pin (940cd968) is
cutlass-4.5-era: its nvvm.fmax call signature no longer matches, so the
FA4 JIT crashes on every no-grad attention call. With FASTVIDEO_FA4=1
forced in CI, every full-suite lane fails on every PR regardless of
diff. Upstream fixed it in 82d6441e ('Fix compatibility issues with
CuTe DSL 4.6.0+', PR 2648) - bump both pins to that rev. Merging this
auto-rebuilds the image (infra-build-image.yml watches docker/).
Merge Protections🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI
🔴 PR merge requirementsWaiting for
This rule is failing.
|
There was a problem hiding this comment.
Code Review
This pull request updates the flash-attn-4 (CuTe) reference to a newer revision (82d6441eec5d4dfec120153db2c0145ae855a083) compatible with CuTe DSL 4.6 in both the Dockerfile and pyproject.toml. The review feedback correctly points out that several comments in both files still refer to the older cutlass-4.5 compatibility and should be updated to reflect the new version.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| # flashinfer/quack pull in. After the wheel install we overlay this cutlass-4.5-safe | ||
| # upstream cute (flash-attn-4) so the image runs FA4 instead of the FA2 fallback. | ||
| ARG FA4_CUTE_REF=940cd9680f3315f2f06b43ab5bea2c2cf2d96806 | ||
| ARG FA4_CUTE_REF=82d6441eec5d4dfec120153db2c0145ae855a083 |
| # FA4 cute, pinned to a cutlass-4.5-compatible revision. torch.compile support | ||
| # comes from FastVideo's own custom_op wrappers. | ||
| flash-attn-4 = { git = "https://github.com/Dao-AILab/flash-attention.git", rev = "940cd9680f3315f2f06b43ab5bea2c2cf2d96806", subdirectory = "flash_attn/cute" } | ||
| flash-attn-4 = { git = "https://github.com/Dao-AILab/flash-attention.git", rev = "82d6441eec5d4dfec120153db2c0145ae855a083", subdirectory = "flash_attn/cute" } |
Problem
The 2026-07-05 CI image rebuild resolved the unpinned transitive
nvidia-cutlass-dslto 4.6.0. Our FA4 cute overlay pin (940cd968) is cutlass-4.5-era — itsnvvm.fmaxcall signature no longer matches, so the FA4 JIT crashes on every no-grad attention call. WithFASTVIDEO_FA4=1forced in CI (#1540), every full-suite lane fails on every PR regardless of diff (#1505/#1461/#1509 all reproduced the identical fmax TypeError twice; the same crash hit the Kandinsky SSIM lane).Proof it's the image, not the code: the exact lora-training lane content passes on a GB200 (no FA4 image) at origin/main.
Solution
Bump both FA4 pins —
FA4_CUTE_REFin docker/Dockerfile and theflash-attn-4rev in pyproject.toml — to upstream82d6441e("Fix compatibility issues with CuTe DSL 4.6.0+", Dao-AILab/flash-attention#2648). Merging auto-rebuilds the image via infra-build-image.yml.Risk & rollback
Two-line pin change to upstream's dedicated compat fix; revert one commit. After merge + image rebuild (~1h) I re-trigger the affected PRs' full suites.