Repository navigation
[MiniMax-M3] Fix UnboundLocalError in the sparse backend on the MSA path - #41367
Open
David-Wu1119 wants to merge 1 commit into
Open
David-Wu1119 wants to merge 1 commit into
David-Wu1119 wants to merge 1 commit into
Conversation
MiniMaxSparseAttnBackend.__init__ calls get_parallel(), imported at module level, when use_msa is set. A later branch of the same method does `from sglang.srt.runtime_context import get_parallel`, which makes the name local to the whole of __init__, so the MSA branch read an unbound local and raised UnboundLocalError before reaching that import. Drop the redundant function-level import and add a CPU regression test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
David-Wu1119
requested review from
Fridge003,
HaiShaw,
Qiaolin-Yu,
hebiao064,
ispobock and
merrymercy
as code owners
September 26, 2026 20:26
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
MiniMaxSparseAttnBackend.__init__callsget_parallel()whenuse_msais set (self.num_q_heads = ... // get_parallel().attn_tp_size).get_parallelis imported at module level. #36527 later addedfrom sglang.srt.runtime_context import get_parallelinside theindex_cache_enabledbranch of the same method. That import binds the name inside__init__, so Python treatsget_parallelas a local for the whole method. The earlier MSA branch then reads an unbound local:use_msais true on CUDA/ROCm whenever MSA is available and the config matches (block size 128, page size 128, top-k blocks 4/8/16/32, and so on), so a MiniMax-M3 start that takes the MSA path fails in the backend constructor. This is true whether or not the index cache is enabled, because the scoping is decided when the function is compiled. You can check it without a GPU:Modifications
get_parallelis already in scope, and the index-cache branch keeps working with it.test/registered/unit/layers/attention/test_minimax_sparse_backend_scoping.py(CPU suite), in the style oftest_dsa_head_gate_guard.py. No CPU runner constructs this backend, so the test checks thatget_parallelis not a local of__init__. It fails onmain('get_parallel' unexpectedly found in (...)) and passes here.Accuracy Tests
Not applicable. No kernel or math changes.
Speed Tests and Profiling
Not applicable.
Checklist
cc @zcnrex (author of #36527)
🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ❌ Run #36269536562
Latest PR Test (Extra): ❌ Run #36269536712
Latest PR Test (AMD ROCm 10): ❌ Run #36269536465