Repository navigation
[Refactor] Chunk MAGI-2 Torch reference attention by query rows - #7282
Closed
yeahdongcn wants to merge 1 commit into
Closed
yeahdongcn wants to merge 1 commit into
yeahdongcn wants to merge 1 commit into
Conversation
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
Collaborator
|
This PR touches tests/diffusion/, vllm_omni/diffusion/ (2 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be: Could one of you take a look when you get a chance? Thanks! |
Contributor
Author
|
Superseded by #8498, which carries this change as its second commit. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Chunk the Torch reference path of MAGI-2 attention along the query axis, keeping every chunk's full key range.
query_chunk_sizequery rows (default 512) at a time during inference.The bound concerns score/probability buffers only, not total memory: K/V storage and intermediates retained for autograd are not bounded by the query chunk size.
Scope
Two files only: the reference implementation and its tests. No backend selection changes, MUSA-specific path, new environment variables, native kernels, caches, RoPE, EP or sampler changes. Existing callers need no changes.
This is an independent extraction of the chunked-reference idea in #7156. It does not depend on the standalone FA3 adapter or on the MUSA consumer. The latter's pending numerical checks are not bypassed or reclassified here.
Validation
Commit:
f644b504d7f1f2018fe8be21900845c876bd0d2d, one signed-off commit based on upstream5378b77be.[512,512,1]. Separate ragged tests verify every chunk keeps the full per-sequence key range.5.2.0-server). FP32/BF16/FP16 chunk sizes 1 and 3 match the unchunked reference on the same GPU, with/without softcap and with GQA/two sinks. Tolerance remainsrtol=1e-5, atol=2e-6.Runtime: torch/torch_musa
2.11.0.post1+musa5.2.0, torchada0.1.83, vLLM0.28.0, vllm-musa0.1.28. Exact source installed editable with--no-deps --no-build-isolation; MUSA entrypoint imports torchada first. No dependencies changed.For MUSA, run the new test file with
-m musafrom a torchada-first entrypoint with one visible GPU.Not run: CUDA hardware, GPU autograd, compile/graph, full MAGI-2 checkpoint/video, peak-memory or performance benchmarks. Structural buffer bounds and functional parity are not an E2E throughput/latency or full-resolution memory claim. This PR remains Draft.