Skip to content

[Cute,Sm120] fix Spark forward and backward regressions - #2474

Open
CuriousCaliBoi wants to merge 1 commit into
Dao-AILab:mainfrom
CuriousCaliBoi:fix/sm120-spark-cpasync-epilogue
Open

CuriousCaliBoi wants to merge 1 commit into
Dao-AILab:mainfrom
CuriousCaliBoi:fix/sm120-spark-cpasync-epilogue

Conversation

@CuriousCaliBoi

Copy link
Copy Markdown

Summary

  • restore the intended SM80-style control flow for FlashAttentionForwardSm120 so DGX Spark does not fall into the TMA epilogue path at runtime
  • initialize the SM120 backward dQ_single_wg state and keep the shared SM80/SM120 backward launcher on concrete softmax-scale values so the Spark backward path compiles cleanly again
  • add a regression test that verifies the SM120 forward object keeps its effective arch pinned to Arch.sm_80

Test plan

  • Force a fresh Spark compile with cache disabled and run a full FA4 forward+backward smoke test on NVIDIA GB10 / compute capability 12.1
  • Benchmark forward on Spark with B=8, S=8192, H=32, D=64, causal=False: FA4 46.52 ms / 94.5 TFLOPS, SDPA 50.18 ms / 87.6 TFLOPS
  • Instantiate FlashAttentionForwardSm120 directly and verify the effective arch stays Arch.sm_80

Made with Cursor

Restore the intended SM80-style control flow for SM120 forward and initialize the SM120 backward config so FA4 compiles and runs end to end on DGX Spark. Keep the shared SM80/SM120 backward launcher on concrete softmax-scale values to avoid DSL type errors, and add a regression test for the SM120 control-flow selection.

Made-with: Cursor

Copy link
Copy Markdown
Contributor

Validated this on DGX Spark / GB10 (SM121).

Env:

  • driver 580.126.09, CUDA 13.0.88
  • PyTorch 2.11.0+cu130
  • nvidia-cutlass-dsl 4.4.2
  • commit b417fd2

Results:

  • current main ba59def reproduces the TMA-O crash (tma_atom_O=None)
  • [Cute,Fwd,Sm120] Fix FlashAttentionForwardSm120 runtime errors on SM120 #2484 fixes forward but still fails backward with dQ_single_wg unbound
  • this branch passes dense fwd+bwd smoke for fp16/bf16, D=64/128, causal/non-causal
  • this branch passes varlen fwd+bwd smoke for fp16/bf16, D=64, causal/non-causal
  • python -m pytest -q tests/cute/test_sm120.py: 1 passed when run against the FA4 editable package path

Max diffs vs PyTorch SDPA stayed in the expected fp16/bf16 range; worst observed was dv=0.003906 for bf16 D=128 causal.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants