Skip to content

Fix online-softmax kernel to avoid infinite thinking ! - #6

Closed
mrexodia wants to merge 2 commits into
MiaAI-Lab:mainfrom
mrexodia:fix-kernel
Closed

mrexodia wants to merge 2 commits into
MiaAI-Lab:mainfrom
mrexodia:fix-kernel

Conversation

@mrexodia

@mrexodia mrexodia commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Closes #5

image

I never saw anywhere near the claimed 64 tok/s personally. This PR 'regresses' back to 31 tok/s, but better slow than broken 🤷🏻‍♂️

@mrexodia

Copy link
Copy Markdown
Contributor Author

Backported sgl-project/sglang#36556 now, performance improved a bit more now!

@mrexodia mrexodia changed the title Fix online-softmax kernel to avoid infinite thinking Fix online-softmax kernel to avoid infinite thinking ! Aug 26, 2026
@MiaAI-Lab

Copy link
Copy Markdown
Owner

Leaving this open, not merging.

The part that actually ships is a backport of sgl-project/sglang#36556 (enable FlashInfer TRTLLM sparse decode on SM12x, drop this recipe's Triton varlen fallback). We already decided not to take 36556 on this tree.

On this cluster the current Triton path is the measured 64 tok/s single-stream / 117 tok/s at ×2 (NEXTN 3/1/4). The PR body itself reports a regression to ~31 tok/s; independent 36556 e2e on 2× Spark was ~34–36 tok/s. Until someone shows TRTLLM matching or beating those numbers on this recipe (same NEXTN graphs, same NVFP4 KV, same 1M YaRN launch), swapping kernels would be a throughput hit, not a free !!!!!! fix.

The online-softmax rewrite in qsa_fa_fallback.py also does not land in the image: the new Dockerfile no longer COPYs that file, so it cannot be the thing that stops the loop.

The thinking-off workaround stays in the README for agent !!!!!! until there is an A/B on this pair. Happy to revisit with a bench against current main.

@mrexodia

Copy link
Copy Markdown
Contributor Author

To clarify: I have never measured 64 tok/s single-stream, even with the original broken code. Could you share the request that measured this?

@MiaAI-Lab

Copy link
Copy Markdown
Owner

Closing without merging.

This PR’s runtime change is a backport of sgl-project/sglang#36556: enable FlashInfer TRT-LLM sparse decode on all SM12x (is_sm120_supported is major==12, so SM121/GB10 is included) and drop this recipe’s Triton varlen fallback. That is the path that silently emits token id 0 at long context on SM121 (HTTP 200, 32/32 !). Upstream narrowed it in sgl-project/sglang#36806 (TRT-LLM on SM100 + exact SM120 only) and added the SM121 Triton packed-varlen kernel in sgl-project/sglang#36845.

That pair is what landed on main here: 6acd773

Live check on this 2×Spark boot after rebuild:

  • Thinking + tools (get_weather) 8/8 on both fp8 and NVFP4 KV, 0 token id 0.
  • fp8 KV NIAH: 16k / 32k / 64k@50% / 120k@50% exact.
  • NVFP4 KV: tools still good; 64k same-turn mid/end still !×64 (separate leftover, not a reason to take 36556).

Merging this PR would put GB10 back on the known-corrupt TRT-LLM route. Conflicts with main anyway. The !!!!!! agent loop from #5 is addressed by #36845 on main, not by 36556.

@MiaAI-Lab

Copy link
Copy Markdown
Owner

Closed: do not merge #36556 on SM121. Fix is sglang#36806+#36845 on main (6acd773).

@MiaAI-Lab MiaAI-Lab closed this Aug 28, 2026
pchar pushed a commit to pchar/Qwen3.8-Flash-Next-Dual-DGX-Sparks that referenced this pull request Sep 22, 2026
…at GMU 0.75 / MNT 8192 / MTP0 (TASK-73.06)
pchar pushed a commit to pchar/Qwen3.8-Flash-Next-Dual-DGX-Sparks that referenced this pull request Sep 22, 2026
…89, host mem leak -> memwatch kill), rollback verified, sweep exhausted (TASK-73.06)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

First request in deepseek-harness doom loops

2 participants