Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/src/snippets/configs/Qwen/qwen3.8-flash-next.jsx
Original file line number Diff line number Diff line change
Expand Up @@ -202,6 +202,7 @@ export const config = {
"--chunked-prefill-size 8192",
"--linear-attn-prefill-backend flashinfer",
"--linear-attn-decode-backend flashinfer",
"--linear-attn-verify-backend triton",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Apply Triton verify when Playground enables MTP on H200

When a user starts from either H200 high-throughput recipe and selects NEXTN/MTP in the Playground, the generic speculative handler adds only the --speculative-* flags, while that base recipe retains FlashInfer decode with BF16 SSM state and no verify override. The generated command therefore still auto-selects FlashInfer verify and hits the SM90 initial_state must be float32 failure this change is meant to prevent; add the Triton override for that customization path as well (for example, by carrying it in all H200 bases or applying it conditionally with the MTP option).

AGENTS.md reference: docs/AGENTS.md:L23-L24

Useful? React with 👍 / 👎.

"--mamba-ssm-dtype bfloat16",
"--speculative-algorithm NEXTN",
"--speculative-num-steps 3",
Expand Down Expand Up @@ -443,6 +444,7 @@ export const config = {
"--chunked-prefill-size 8192",
"--linear-attn-prefill-backend flashinfer",
"--linear-attn-decode-backend flashinfer",
"--linear-attn-verify-backend triton",
"--mamba-ssm-dtype bfloat16",
"--speculative-algorithm NEXTN",
"--speculative-num-steps 3",
Expand Down
Loading