docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state - #36611
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f897085998
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "--chunked-prefill-size 8192", | ||
| "--linear-attn-prefill-backend flashinfer", | ||
| "--linear-attn-decode-backend flashinfer", | ||
| "--linear-attn-verify-backend triton", |
There was a problem hiding this comment.
Apply Triton verify when Playground enables MTP on H200
When a user starts from either H200 high-throughput recipe and selects NEXTN/MTP in the Playground, the generic speculative handler adds only the --speculative-* flags, while that base recipe retains FlashInfer decode with BF16 SSM state and no verify override. The generated command therefore still auto-selects FlashInfer verify and hits the SM90 initial_state must be float32 failure this change is meant to prevent; add the Triton override for that customization path as well (for example, by carrying it in all H200 bases or applying it conditionally with the MTP option).
AGENTS.md reference: docs/AGENTS.md:L23-L24
Useful? React with 👍 / 👎.
|
I have encountered this issue too, and it was resolved by adding |
- qwen3_5_mtp: captured prefill pads embeddings while target hidden states keep real height; graft the real rows into an equal-height slot before cat+fc instead of letting cat broadcast-fail or misalign (upstream qwen3_5_mtp forward padding fix). - gdn_backend: honor --linear-attn-verify-backend. The dispatcher re-derived the verify kernel with the auto rule and ignored the stored choice, so an explicit triton selection (required when --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state requirement, upstream sgl-project#36611) had no effect. - linear/utils: raise when --enable-deterministic-inference combines with a FlashInfer GDN prefill (upstream _validate_gdn_linear_attn_backends). Validated on H20 TP1 with NEXTN (steps 3 / topk 1 / draft 4): boots, gsm8k 200q = 0.975 (bf16 baseline 0.980), accept length 3.55-3.60.
Summary
Motivation
The H200 low-latency recipes combine NEXTN with
--mamba-ssm-dtype bfloat16. On SM90, the FlashInfer MTP verify path requires fp32 SSM state, so auto-selecting it can fail during target verify CUDA graph capture with:#34592 made SGLang honor an explicitly selected Triton verify backend while intentionally leaving auto selection unchanged. The verified Cookbook commands therefore need to opt into that backend explicitly.
Validation
node docs/scripts/check_cookbook_configs.mjsQwen/Qwen3.8-Flash-Next-FP8: server startup, target verify CUDA graph capture,/health, and a Chat Completions smoke request succeededdecode=FlashInferGDNKernel, extend=FlashInferGDNKernel, verify=TritonGDNKernelCI States
Latest PR Test (Base): ✅ Run #33038758235
Latest PR Test (Extra): ❌ Run #33038758119
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.