Repository navigation
Conversation
janbernloehr
requested review from
Alisehen,
AniZpZ,
BBuf,
Edwardf0t1,
FlamingoPg,
Fridge003,
HaiShaw,
OrangeRedeng,
Ying1123,
b8zhong,
ch-wan,
ispobock,
merrymercy and
mmangkad
as code owners
August 19, 2026 09:56
janbernloehr
force-pushed
the
jbernloehr/support-llama4-nvfp4-rout
branch
from
August 20, 2026 07:22
11dc473 to
47b6b56
Compare
nvpohanh
reviewed
Aug 31, 2026
nvpohanh
left a comment
Collaborator
There was a problem hiding this comment.
[by Codex] One inline review finding.
|
|
||
| # Activations are already quantized and all-gathered before the runner on | ||
| # this path, so the multiply is no longer expressible here. | ||
| if getattr(dispatch_output, "hidden_states_scale", None) is not None: |
Collaborator
There was a problem hiding this comment.
[by Codex] Severity: style | Confidence: High
Please avoid getattr here. Both StandardDispatchOutput and FlashinferDispatchOutput define hidden_states_scale; annotate this helper with their explicit union (or a small protocol) and read dispatch_output.hidden_states_scale directly so the supported structure is type-checked.
Collaborator
|
/tag-and-rerun-ci |
Collaborator
|
@janbernloehr could you fix the conflicts? |
janbernloehr
force-pushed
the
jbernloehr/support-llama4-nvfp4-rout
branch
from
September 4, 2026 09:31
bb02d04 to
cecbbd4
Compare
b8zhong
reviewed
Sep 4, 2026
b8zhong
reviewed
Sep 4, 2026
b8zhong
approved these changes
Sep 4, 2026
janbernloehr
force-pushed
the
jbernloehr/support-llama4-nvfp4-rout
branch
from
September 4, 2026 12:23
cecbbd4 to
24f643b
Compare
b8zhong
enabled auto-merge (squash)
September 4, 2026 12:31
Collaborator
|
/rerun-failed-ci |
Collaborator
|
/rerun-failed-ci |
2 similar comments
Collaborator
|
/rerun-failed-ci |
Collaborator
|
/rerun-failed-ci |
nvpohanh
approved these changes
Sep 14, 2026
Collaborator
|
/rerun-failed-ci |
1 similar comment
Collaborator
|
/rerun-failed-ci |
Collaborator
|
/rerun-failed-ci |
Collaborator
mmangkad
approved these changes
Sep 22, 2026
kpham-sgl
disabled auto-merge
September 22, 2026 03:18
detain
added a commit
to detain/sglang
that referenced
this pull request
Sep 28, 2026
…input Upstream sgl-project#35504 made _run_flashinfer_cutlass read runner_config.apply_router_weight_on_input via _prescale_router_weight_on_input. The TestSmallmRuntimeFallback SimpleNamespace stub (from the SM120 small-row NVFP4 kernel pick) predates that field, so all four fallback cases raised AttributeError after the upstream merge. Add it with MoeRunnerConfig's default (False). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Fixes #34192.
Llama 4 constructs its MoE layers with
apply_router_weight_on_input=True. For ModelOpt NVFP4 on SM120/SM121, SGLang selects the FlashInfer CUTLASS MoE runner, which previously rejected this configuration during server warmup.For top-1 routing, the required semantics can be implemented exactly by multiplying each token's hidden states by its router weight before activation quantization and passing unit final scales to CUTLASS. This follows the existing pattern in the AITER MoE runner.
The fix remains intentionally restricted to top-1 routing. With top-k greater than one, each selected expert would require a differently scaled copy of the input.
Modifications
token_final_scaleswhen pre-scaling is active.token_final_scalespassed through this CUTLASS path to float32, as required by the binding.apply_router_weight_on_input=True.The normal
apply_router_weight_on_input=Falsepath is unchanged apart from enforcing the CUTLASS binding's float32 scale requirement.Accuracy Tests
Added:
The tests compare the transformed hidden states and final scales against the expected top-1 formulation and verify that unsupported configurations fail explicitly.
Local validation:
Result: passed.
All commit-time pre-commit checks passed, including Python AST validation, isort, Ruff, Black, codespell, and registered-test validation.
The focused pytest test could not be executed in the local development environment because PyTorch is not installed. No real-checkpoint accuracy test was run for this checkout.
The equivalent patch and its SM120/SM121 validation evidence are documented in #34192. That validation used dummy weights for end-to-end liveness, so it is not claimed as a model-accuracy result.
Speed Tests and Profiling
No speed test or profiling was run for this checkout.
The equivalent patch validation documented in #34192 completed both SM120 scenarios without errors and did not show a decode throughput regression. Those results used dummy weights and a separate validated build, so they are included only as prior evidence, not as performance results for this commit.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ✅ Run #35553178662
Latest PR Test (Extra): ❌ Run #35553178439
Latest PR Test (AMD ROCm 10): ❌ Run #35553178589