Skip to content

[MiniMax-M3] Fix UnboundLocalError in the sparse backend on the MSA path - #41367

Open
David-Wu1119 wants to merge 1 commit into
sgl-project:mainfrom
David-Wu1119:fix/minimax-sparse-get-parallel-shadowing
Open

David-Wu1119 wants to merge 1 commit into
sgl-project:mainfrom
David-Wu1119:fix/minimax-sparse-get-parallel-shadowing

Conversation

@David-Wu1119

@David-Wu1119 David-Wu1119 commented Sep 26, 2026 •

Copy link
Copy Markdown

Motivation

MiniMaxSparseAttnBackend.__init__ calls get_parallel() when use_msa is set (self.num_q_heads = ... // get_parallel().attn_tp_size). get_parallel is imported at module level. #36527 later added from sglang.srt.runtime_context import get_parallel inside the index_cache_enabled branch of the same method. That import binds the name inside __init__, so Python treats get_parallel as a local for the whole method. The earlier MSA branch then reads an unbound local:

UnboundLocalError: cannot access local variable 'get_parallel' where it is not associated with a value

use_msa is true on CUDA/ROCm whenever MSA is available and the config matches (block size 128, page size 128, top-k blocks 4/8/16/32, and so on), so a MiniMax-M3 start that takes the MSA path fails in the backend constructor. This is true whether or not the index cache is enabled, because the scoping is decided when the function is compiled. You can check it without a GPU:

from sglang.srt.layers.attention.minimax_sparse_backend import MiniMaxSparseAttnBackend
"get_parallel" in MiniMaxSparseAttnBackend.__init__.__code__.co_varnames  # main: True, this PR: False

Modifications

  • Remove the redundant function-level import. The module-level get_parallel is already in scope, and the index-cache branch keeps working with it.
  • Add test/registered/unit/layers/attention/test_minimax_sparse_backend_scoping.py (CPU suite), in the style of test_dsa_head_gate_guard.py. No CPU runner constructs this backend, so the test checks that get_parallel is not a local of __init__. It fails on main ('get_parallel' unexpectedly found in (...)) and passes here.
python test/registered/unit/layers/attention/test_minimax_sparse_backend_scoping.py   # main: FAILED (errors=1), this PR: OK
pre-commit run --files <both files>                                                  # isort, ruff, ruff format, codespell, CI-registry checks pass

Accuracy Tests

Not applicable. No kernel or math changes.

Speed Tests and Profiling

Not applicable.

Checklist

cc @zcnrex (author of #36527)

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ❌ Run #36269536562
Latest PR Test (Extra): ❌ Run #36269536712
Latest PR Test (AMD ROCm 10): ❌ Run #36269536465

MiniMaxSparseAttnBackend.__init__ calls get_parallel(), imported at module
level, when use_msa is set. A later branch of the same method does
`from sglang.srt.runtime_context import get_parallel`, which makes the name
local to the whole of __init__, so the MSA branch read an unbound local and
raised UnboundLocalError before reaching that import. Drop the redundant
function-level import and add a CPU regression test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 26, 2026 20:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants