Repository navigation
Conversation
|
Backported sgl-project/sglang#36556 now, performance improved a bit more now! |
!
|
Leaving this open, not merging. The part that actually ships is a backport of sgl-project/sglang#36556 (enable FlashInfer TRTLLM sparse decode on SM12x, drop this recipe's Triton varlen fallback). We already decided not to take 36556 on this tree. On this cluster the current Triton path is the measured 64 tok/s single-stream / 117 tok/s at ×2 (NEXTN 3/1/4). The PR body itself reports a regression to ~31 tok/s; independent 36556 e2e on 2× Spark was ~34–36 tok/s. Until someone shows TRTLLM matching or beating those numbers on this recipe (same NEXTN graphs, same NVFP4 KV, same 1M YaRN launch), swapping kernels would be a throughput hit, not a free The online-softmax rewrite in The thinking-off workaround stays in the README for agent |
|
To clarify: I have never measured 64 tok/s single-stream, even with the original broken code. Could you share the request that measured this? |
|
Closing without merging. This PR’s runtime change is a backport of sgl-project/sglang#36556: enable FlashInfer TRT-LLM sparse decode on all SM12x ( That pair is what landed on Live check on this 2×Spark boot after rebuild:
Merging this PR would put GB10 back on the known-corrupt TRT-LLM route. Conflicts with |
|
Closed: do not merge #36556 on SM121. Fix is sglang#36806+#36845 on main (6acd773). |
…at GMU 0.75 / MNT 8192 / MTP0 (TASK-73.06)
…89, host mem leak -> memwatch kill), rollback verified, sweep exhausted (TASK-73.06)
Closes #5
I never saw anywhere near the claimed 64 tok/s personally. This PR 'regresses' back to 31 tok/s, but better slow than broken 🤷🏻♂️