Repository navigation
[MiniMax-M3] Drop the local get_parallel import that shadows the module-level one - #38911
Closed
huthvincent wants to merge 1 commit into
Closed
huthvincent wants to merge 1 commit into
huthvincent wants to merge 1 commit into
Conversation
…le-level one `MiniMaxSparseAttnBackend.__init__` reads `get_parallel()` under `if self.use_msa:`, against the module-level import at the top of the file. A function-local `from sglang.srt.runtime_context import get_parallel` added later in the same function makes the name local to all of `__init__`, so the earlier read compiles to LOAD_FAST_CHECK and raises UnboundLocalError whenever `use_msa` is true -- whether or not the local import executes, because the scoping is static rather than dynamic. The module-level import already provides the same symbol from the same module, and `runtime_context` documents itself as safe to import at module level for exactly this reason. `get_spec`, imported at the same site and used later in `__init__` with no local shadow, is the existing precedent in this file.
huthvincent
requested review from
Fridge003,
HaiShaw,
Qiaolin-Yu,
hebiao064,
ispobock and
merrymercy
as code owners
September 10, 2026 15:02
3 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
MiniMaxSparseAttnBackend.__init__cannot construct when MSA is active. It raisesat
python/sglang/srt/layers/attention/minimax_sparse_backend.pyline 237, which is insideif self.use_msa:. Becauseattn_backend_wrapperconstructs the backend with notry/exceptand nofallback, an affected launch dies before the first token.
The cause is a scoping collision introduced two days ago by #36527.
get_parallelis imported at modulelevel (lines 25-28). #36527 added a function-local
from sglang.srt.runtime_context import get_parallelat line 297, inside
if self.index_cache_enabled:, in the same__init__that already used the name at line237. Python binds a name imported anywhere in a function body as local to that entire body, so line 237's
local slot is empty when it is read.
This happens whether or not line 297 executes, because the scoping is static. It is also why the two uses
have never collided in testing: line 297's guard
self.index_cache_enabledrequiresis_hip() and is_gfx95_supported()(lines 286-291), so on AMD line 237 is skipped and the local import runsharmlessly, while on SM100 line 237 runs and the local import does not.
The gate at lines 210-217 is satisfied by the documented default Blackwell recipe, with no non-default flag:
not envs.SGLANG_DISABLE_MSA.get()SGLANG_DISABLE_MSA = EnvBool(False),python/sglang/srt/environ.pyline 1585msa_available()(10, 0)or(10, 3)and an importablefmha_sm100.docs/cookbook/autoregressive/MiniMax/MiniMax-M3.mdxline 90 states thatfmha_sm100"ships pre-installed in the M3 dev image (lmsysorg/sglang:dev-minimax-m3) so the Blackwell recipe above engages it automatically with no extra setup"self.kv_pool.page_size == self.block_size_kpython/sglang/srt/arg_groups/model_overrides/minimax_m3.pylines 73 and 79-81: on SM100 the backend defaults tofa4(ortrtllm_mhaforfp8_e4m3KV), and thenif page_resolved is None and backend_resolved in ("fa4", "trtllm_mha"): overrides["page_size"] = 128. Passing no--page-sizeis what triggers it.not _main_kv_is_fp8 or _msa_fp8_ok_main_kv_is_fp8false;fp8_e4m3KV withtrtllm_mhasets_msa_fp8_ok. Both SM100 sub-recipes pass.block_size_kandtopk_blockscome from the checkpoint'ssparse_attention_config.Modifications
Two lines deleted in one file:
python/sglang/srt/layers/attention/minimax_sparse_backend.py.Nothing else is touched: no refactor, no rename, no formatting, no adjacent fix.
Three reasons this is the right deletion rather than a rename or an added local import above line 237:
TYPE_CHECKINGblock at lines 54-55 imports onlyModelRunner, so the module-level import is notconditional.
python/sglang/srt/runtime_context.pylines 66-67 say why a module-level import is safe here:"Imported lazily so this module has no import-time dependencies: any module can import get_parallel at
module level without risking an import cycle."
get_specis imported at the same site (line 27) and used at line 260 with no local shadow. That is theexisting precedent in this file for relying on the module-level import.
Both lines have to go, not just line 297: deleting the import alone leaves a blank line at the top of the
ifbody, whichruff-formatv0.15.1 removes, sopre-commit run --all-fileswould fail. With both linesdeleted,
ruff format --checkreports "1 file already formatted" andruff check --select F401,F821,UP037reports "All checks passed!".Accuracy Tests
No model output changes. The patch deletes a redundant import and changes no computation, no dtype and no
kernel selection.
What we executed, and on what. We lifted
__init__'s code object out of the file at12771786f23190b1845db33366eba09cb5eacf41and ran it with its globals stubbed, so the bytecode was thisrepository's and unmodified. As shipped it raises the
UnboundLocalErrorabove at line 237; with the twolines deleted, the same harness raises nothing.
What we did not execute. We have no B200 or B300, so we have not observed the failure on the
configuration that reaches it.
msa_available()requires compute capability(10, 0)or(10, 3); ourhardware is
sm_89. We have not launched MiniMax-M3.Speed Tests and Profiling
We did not measure this. The analysis is static — read from the code and the commit history, with no
profiling and no benchmark run on our side. The check described below is what we are asking you to run.
There is no performance claim here in either direction. The patch removes an import statement.
How to check it
Two ways, and the first needs no GPU, no model and no SGLang import.
1. Read the scoping directly out of the file. On a checkout of
main:On
mainthis printsis_local: True is_global: Falseand['LOAD_FAST_CHECK']—LOAD_FAST_CHECKis theopcode that raises
UnboundLocalErroron an empty local slot. On this branch it printsis_local: False is_global: Trueand['LOAD_GLOBAL']. If it printedLOAD_GLOBALonmain, this reportwould be wrong.
2. On Blackwell hardware.
sglang serve --model-path MiniMaxAI/MiniMax-M3insidelmsysorg/sglang:dev-minimax-m3, no other flags. Expected onmain:UnboundLocalErrorduring attentionbackend construction. Expected on this branch: normal startup.
The strongest objection to this report
Stated in full rather than answered, because it is the reason to be sceptical and it is not disposed of:
The one fact that bears on it without answering it: #36527 merged with zero review comments, and the shadow
is two days old.
Checklist
ruff format --checkandruff check --select F401,F821,UP037both clean on the changed file at v0.15.1, the version pinned in.pre-commit-config.yaml.on SM100. Happy to add whichever you would accept.
made. Section "Speed Tests and Profiling" says so explicitly rather than implying one.
CI States
Latest PR Test (Base): ❌ Run #34493043676
Latest PR Test (Extra): ❌ Run #34493043211
Latest PR Test (AMD ROCm 10): ❌ Run #34493043560