Skip to content

docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state - #36611

Merged
zijiexia merged 2 commits into
sgl-project:mainfrom
huangzhilin-hzl:molou/qwen38-sm90-triton-verify
Aug 27, 2026
Merged

zijiexia merged 2 commits into
sgl-project:mainfrom
huangzhilin-hzl:molou/qwen38-sm90-triton-verify

Conversation

@huangzhilin-hzl

@huangzhilin-hzl huangzhilin-hzl commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • pin Triton as the linear-attention verify backend for the H200 BF16 and FP8 low-latency recipes
  • keep FlashInfer as the prefill and decode backend

Motivation

The H200 low-latency recipes combine NEXTN with --mamba-ssm-dtype bfloat16. On SM90, the FlashInfer MTP verify path requires fp32 SSM state, so auto-selecting it can fail during target verify CUDA graph capture with:

AssertionError: initial_state must be float32, got torch.bfloat16

#34592 made SGLang honor an explicitly selected Triton verify backend while intentionally leaving auto selection unchanged. The verified Cookbook commands therefore need to opt into that backend explicitly.

Validation

  • node docs/scripts/check_cookbook_configs.mjs
  • checked that only the H200 BF16/FP8 low-latency cells receive the new flag
  • H20 SM90, TP8/EP8, Qwen/Qwen3.8-Flash-Next-FP8: server startup, target verify CUDA graph capture, /health, and a Chat Completions smoke request succeeded
  • runtime dispatcher resolved to decode=FlashInferGDNKernel, extend=FlashInferGDNKernel, verify=TritonGDNKernel

CI States

Latest PR Test (Base): ✅ Run #33038758235
Latest PR Test (Extra): ❌ Run #33038758119
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 27, 2026
@huangzhilin-hzl
huangzhilin-hzl marked this pull request as ready for review August 27, 2026 04:15
@huangzhilin-hzl huangzhilin-hzl changed the title docs(cookbook): use Triton GDN verify for Qwen3.8 Flash Next on SM90 docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state Aug 27, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f897085998

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

"--chunked-prefill-size 8192",
"--linear-attn-prefill-backend flashinfer",
"--linear-attn-decode-backend flashinfer",
"--linear-attn-verify-backend triton",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Apply Triton verify when Playground enables MTP on H200

When a user starts from either H200 high-throughput recipe and selects NEXTN/MTP in the Playground, the generic speculative handler adds only the --speculative-* flags, while that base recipe retains FlashInfer decode with BF16 SSM state and no verify override. The generated command therefore still auto-selects FlashInfer verify and hits the SM90 initial_state must be float32 failure this change is meant to prevent; add the Triton override for that customization path as well (for example, by carrying it in all H200 bases or applying it conditionally with the MTP option).

AGENTS.md reference: docs/AGENTS.md:L23-L24

Useful? React with 👍 / 👎.

@yuyu5333

Copy link
Copy Markdown
Contributor

I have encountered this issue too, and it was resolved by adding --linear-attn-verify-backend triton.

@zijiexia zijiexia self-assigned this Aug 27, 2026
@zijiexia
zijiexia merged commit 536f570 into sgl-project:main Aug 27, 2026
93 of 97 checks passed
TobyMint added a commit to TobyMint/sglang that referenced this pull request Sep 1, 2026
- qwen3_5_mtp: captured prefill pads embeddings while target hidden
  states keep real height; graft the real rows into an equal-height
  slot before cat+fc instead of letting cat broadcast-fail or misalign
  (upstream qwen3_5_mtp forward padding fix).
- gdn_backend: honor --linear-attn-verify-backend. The dispatcher
  re-derived the verify kernel with the auto rule and ignored the
  stored choice, so an explicit triton selection (required when
  --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state
  requirement, upstream sgl-project#36611) had no effect.
- linear/utils: raise when --enable-deterministic-inference combines
  with a FlashInfer GDN prefill (upstream _validate_gdn_linear_attn_backends).

Validated on H20 TP1 with NEXTN (steps 3 / topk 1 / draft 4): boots,
gsm8k 200q = 0.975 (bf16 baseline 0.980), accept length 3.55-3.60.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants