[Spec] Multi-layer mamba scatter cleanup; fix positional call bug by hnyls2002 · Pull Request #25030 · sgl-project/sglang

hnyls2002 · 2026-05-12T03:20:28Z

Summary

Align MultiLayerEagleWorker's mamba-scatter to the same form used in EAGLEWorker / EAGLEWorkerV2 after [Spec] Mamba scatter cleanup; fix multi-layer positional bug; dflash naming #25029.
Fix pre-existing positional-call bug where model was being passed into the mamba_track_indices slot.

Changes

Drop the num_accept_tokens = num_correct_drafts + 1 alias; use num_correct_drafts directly and inline the + 1 into cumsum.
Replace last_token_indices_per_req - first_token_indices_per_req with accepted_indices[cum - 1] - accepted_indices_offset, eliminating one cat and one index_select per step. Equivalent because first_token_indices_per_req[i] == i * draft_token_num is a verify-kernel-enforced invariant (current token is always the first accepted slot).
Else-branch returns num_correct_drafts directly instead of num_accept_tokens - 1.
Convert call to update_mamba_state_after_mtp_verify from 2-positional to keyword form with explicit mamba_track_indices=None, mamba_steps_to_track=None — fixes pre-existing bug where only 2 args were passed against a 4-arg signature, causing TypeError whenever the hybrid_gdn_config branch fires on MultiLayerEagleWorker.

Follows up on #25029.

gemini-code-assist · 2026-05-12T03:20:31Z

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

hnyls2002 · 2026-05-12T03:41:39Z

/rerun-test test_mimo_models.py test_step3p5_flash_chain_mtp.py

github-actions · 2026-05-12T03:42:03Z

🚀 8-gpu-h200 (2 tests): ✅ View workflow run

cd test/ && python3 registered/8-gpu-models/test_mimo_models.py
cd test/ && python3 registered/8-gpu-models/test_step3p5_flash_chain_mtp.py

… bug

…l-project#25030)

…ack) Brings in upstream sgl-project/sglang main commits since 096ad02 (merge base, Laguna-XS.2 model support). Total: 28 upstream commits composed. Custom-stack files preserved intact (entirely-ours, byte-identical to origin/main): - Blackwell CuTe kernel suite (warp_decode_cute, g1_attention_cute, gated_norm_cute, layersplit_cute, fused_store_index_cache) - TurboQuant 2.5-bit dense KV cache path - HIGGS 2-bit dense KV cache path (with split-K decode) - NVFP4 IndexCache dispatcher (active gate) - quantization_config_dispatch (HF-config-driven runtime routing) - All custom server-args flags and runtime methods preserved Verification: - 200+ merged Python files compile cleanly - Dispatcher symbol presence verified - HIGGS pool / TurboQuant pool classes present at expected lines - compressed_tensors_w4a4_nvfp4_moe imports clean - All custom server-args flags present (enable_higgs_dense_2bit_kv_cache, enable_turboquant_dense_kv_cache, turboquant_dense_kv_preset, indexer_quantization_declared, higgs_mla_decode_num_splits, etc.) Manual-merged shared files (auto-merge gave broken/mixed output; cleaned up post-merge): - python/sglang/srt/disaggregation/mooncake/conn.py: upstream's PR#24932 refactored maybe_send_extra into a state-types-loop. Replayed our LayerSplit NSA state-index-length-mismatch check inside the SWA/NSA branch of the new loop body. - sgl-kernel/python/sgl_kernel/__init__.py: upstream's PR#23449 (Apple Silicon Metal kernel) wrapped the entire module body in `if darwin/arm64: from sgl_kernel.metal import * else: ...`. The auto-merge duplicated the file body; rewrote cleanly with upstream's structure and re-injected our `g1_gate_forward`, `warp_decode_cute_moe_forward`, and `warp_decode_cute_moe_packed_forward` imports plus `g1_gate_forward` in _DEBUG_EXPORT_NAMES. - python/sglang/srt/managers/scheduler_output_processor_mixin.py: line 628 still referenced `result.num_accepted_drafts` (renamed by PR sgl-project#25038 to `num_correct_drafts`). Renamed in place. - python/sglang/srt/observability/scheduler_metrics_mixin.py: a block around the spec-decode logging path had mixed old/new names from auto-merge (lines 553/557/560). Renamed `spec_num_accepted_tokens` -> `spec_num_accept_tokens` and local `num_accepted_drafts` -> `num_correct_drafts` to match the rest of the file. - test/test_smc_info.py: stub Req mock used the old field names `spec_accepted_drafts` and `update_spec_acceptance_histogram`. Renamed to `spec_num_correct_drafts` and `update_spec_correct_drafts_histogram` per PR sgl-project#24081. Auto-merge cleanly integrated upstream changes to: - server_args.py (new fields: prefill_only_disable_kv_cache, weight_loader_drop_cache_after_load, prefill_delayer_queue_min_ratio, prefill_delayer_max_delay_ms, speculative_draft_window_size, etc.) - mem_cache/memory_pool.py (new NoOpMHATokenToKVPool) - model_executor/model_runner_kv_cache_mixin.py (NoOpMHATokenToKVPool pool factory + _validate_prefill_only_disable_kv_cache_pool_family) - layers/attention/nsa_backend.py (spec rename num_accepted_drafts -> num_correct_drafts; num_accepted_tokens -> num_accept_tokens) - layers/attention/nsa/nsa_indexer.py (new _apply_q_scale_and_softmax_scale compile method; torch.mm replaces deep_gemm wrapper) - 28+ disaggregation/spec/runner files with mostly clean upstream-side-only integration. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> ----- upstream commit subjects (28) ----- fd3eb77 [Cookbook]: add Laguna-XS.2 (Poolside) (sgl-project#24730) 6be1a45 Fix swa component host hit (sgl-project#25085) 693f497 [NPU] use causal_conv1d_update_v2 for performance (sgl-project#24595) 1efe9e2 [Bug Fix] Reject incompatible combination of --disable-cuda-graph-padding and --enable-torch-compile (sgl-project#23903) 8d27ce7 Optimize uvicorn startup command (sgl-project#25041) b35fd5f [fix] skip legacy minicpmv conv template for MiniCPM-V 4.6 (sgl-project#24998) 7582237 [Tiny Fix] Disable BCG when inner layer_model unresolved (sgl-project#25021) ca3bc05 Deepseek-v4-Pro share expert tp1 (sgl-project#24949) a72d3ae [Spec] Multi-layer mamba scatter cleanup; fix positional call bug (sgl-project#25030) 7128533 Revert "Migrate Intel CPU cases to the test/registered." (sgl-project#25044) 1f985c5 [Spec] Rename `accepted_indices` -> `accept_indices`; drop `_token_id` suffix per Rule 5 (sgl-project#25038) ecf5d84 Migrate Intel CPU cases to the test/registered. (sgl-project#22670) d7f4761 [PD] Refactor hybrid state transfer (sgl-project#24932) 91907b7 [UnifiedTree]: Fix Unified HiCache tombstone lock release replay (sgl-project#24972) 4ad63ad [Spec] Rename `accepted_drafts` -> `correct_drafts` for unambiguous naming (sgl-project#24081) 6bfb365 [PD] Rate limit prefill inflight polling warnings (sgl-project#24967) 6bb79c1 [Linear Attn] Add CUSTOM enum and plugin extensibility for kernel backends (sgl-project#24937) cfc41d5 Fix kimi k2.5 mla eagle + dp attention (sgl-project#25033) 0f3932c [Fix] Qwen3-ASR config: set thinker_config before super().__init__ (sgl-project#24187) f526e3f [Spec] Mamba scatter cleanup; fix multi-layer positional bug; dflash naming (sgl-project#25029) 10375a1 [NIXL][XPU] Fix uint64 overflow for mismatched P/D TP sizes (e.g. prefill_tp=1, decode_tp=2) (sgl-project#24648) 0a37d24 [diffusion] hardware: support sage attention backend on MUSA (attn backend, 21/N) (sgl-project#24752) 5495026 [HiCache] feat: default storage prefetch timeout (sgl-project#23309) 186eb42 Feat: Support SWA (Sliding Window Attention) for EAGLE-3 drafter (sgl-project#24664) a75b79e Feat: Support newer EAGLE-3 drafters (sgl-project#24663) f3a8189 [Spec] Internal rename per N2 v2 naming rule (sgl-project#25014) bfc2eda [MUSA] Use MUSA-optimized operators in piecewise CUDA graph (sgl-project#23633) 74d70af [Apple Silicon] Add Metal kernel support in sgl-kernel (sgl-project#23449)

…l-project#25030)

hnyls2002 requested review from Qiaolin-Yu, Ying1123 and merrymercy as code owners May 12, 2026 03:20

Base automatically changed from lsyin/spec-mamba-scatter-cleanup to main May 12, 2026 03:36

hnyls2002 requested review from fzyzcjy, hanming-lu, hebiao064, iforgetmyname, ping1jing2, sufeng-buaa, yizhang2077 and yuan-luo as code owners May 12, 2026 03:36

hnyls2002 force-pushed the lsyin/spec-multi-layer-mamba-cleanup branch from 50760f4 to 799a703 Compare May 12, 2026 03:38

hnyls2002 mentioned this pull request May 12, 2026

[Spec] Rename accepted_drafts -> correct_drafts for unambiguous naming #24081

Merged

multi-layer mamba scatter: align to eagle_worker form; fix positional…

12341f5

… bug

hnyls2002 force-pushed the lsyin/spec-multi-layer-mamba-cleanup branch from 799a703 to 12341f5 Compare May 12, 2026 05:37

hnyls2002 merged commit a72d3ae into main May 12, 2026
64 of 74 checks passed

hnyls2002 deleted the lsyin/spec-multi-layer-mamba-cleanup branch May 12, 2026 05:42

LucQueen pushed a commit to LucQueen/sglang that referenced this pull request May 12, 2026

[Spec] Multi-layer mamba scatter cleanup; fix positional call bug (sg…

d9d2592

…l-project#25030)

xjpang pushed a commit to xjpang/sglang that referenced this pull request May 13, 2026

[Spec] Multi-layer mamba scatter cleanup; fix positional call bug (sg…

4140c93

…l-project#25030)

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Spec] Multi-layer mamba scatter cleanup; fix positional call bug#25030

[Spec] Multi-layer mamba scatter cleanup; fix positional call bug#25030
hnyls2002 merged 1 commit into
mainfrom
lsyin/spec-multi-layer-mamba-cleanup

hnyls2002 commented May 12, 2026 •

edited

Loading

Uh oh!

gemini-code-assist Bot commented May 12, 2026

Uh oh!

hnyls2002 commented May 12, 2026

Uh oh!

github-actions Bot commented May 12, 2026 •

edited

Loading

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant

Conversation

hnyls2002 commented May 12, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

Changes

Uh oh!

gemini-code-assist Bot commented May 12, 2026

Uh oh!

hnyls2002 commented May 12, 2026

Uh oh!

github-actions Bot commented May 12, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant

hnyls2002 commented May 12, 2026 •

edited

Loading

github-actions Bot commented May 12, 2026 •

edited

Loading