Skip to content

Fix Nightly NV CI - #33564

Merged
Fridge003 merged 5 commits into
mainfrom
brayden-fix-longcat-ngram-oe-vocab-size
Aug 6, 2026
Merged

Fridge003 merged 5 commits into
mainfrom
brayden-fix-longcat-ngram-oe-vocab-size

Conversation

@b8zhong

@b8zhong b8zhong commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator

NgramEmbedding passed exclusive_oe_embedder_size_sums[-1] (a 0-dim CUDA tensor) as num_embeddings, so the fused Triton embedding kernel got a tensor where it wants a tl.constexpr int and failed to compile. Wrap it in int().

Fixes nightly test_longcat_flash_lite_fp8.py:

TypeError("cannot convert 61341792 of type <class 'torch.Tensor'> to tensor")

Verified with a 2-rank repro at the real LongCat-Flash-Lite config: crashes before, passes after, output bit-identical to the non-Triton masked path.


CI States

Latest PR Test (Base): 🚫 Run #31061386671
Latest PR Test (Extra): ❌ Run #31061386535

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@b8zhong

b8zhong commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/8-gpu-models/test_longcat_flash_lite_fp8.py

@github-actions

github-actions Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/8-gpu-models/test_longcat_flash_lite_fp8.py:

🚀 8-gpu-h200 (1 test): ❌ View workflow run

cd test/ && python3 registered/8-gpu-models/test_longcat_flash_lite_fp8.py

🚀 8-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/8-gpu-models/test_longcat_flash_lite_fp8.py

@b8zhong
b8zhong force-pushed the brayden-fix-longcat-ngram-oe-vocab-size branch from 7f2b9d9 to 4ee81f1 Compare August 4, 2026 15:23
@b8zhong

b8zhong commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/8-gpu-models/test_glm52_fp8.py

@github-actions

github-actions Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/8-gpu-models/test_glm52_fp8.py:

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/8-gpu-models/test_glm52_fp8.py

🚀 8-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/8-gpu-models/test_glm52_fp8.py

@b8zhong b8zhong changed the title [LongCat] Fix n-gram over-embedding vocab size passed as a tensor Fix Nightly CI Aug 4, 2026
@b8zhong b8zhong changed the title Fix Nightly CI Fix Nightly NV CI Aug 4, 2026
LongcatFlashConfig normalizes architectures to LongcatFlashForCausalLM, which
sits in _DEEPSEEK_FAMILY_ARCHS, so _deepseek_moe_quant_resolution forced
moe_runner_backend=flashinfer_trtllm on SM100. LongCat picks its top-12 from
n_routed_experts + zero_expert_num (256 + 128) logits and splits the identity
experts out afterwards, which trtllm-gen's internal routing cannot express --
it only sees the 256 routed logits. The nightly B200 run hit the guard:

    assert TopKOutputChecker.format_is_bypassed(topk_output)

Exclude the LongCat archs from that override so they fall back to auto
(deep_gemm/triton), the path that scored gsm8k 0.935 on H200.

Also skip the TP8+EP8+deepep variant: DeepEP's low-latency dispatch asserts
num_topk <= kNumMaxTopK (11 in internode_ll.cu) and LongCat moe_topk is 12,
so the variant has never been runnable since #30975 made --moe-a2a-backend
take effect for LongCat.
@b8zhong

b8zhong commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/8-gpu-models/test_longcat_flash_lite_fp8.py

@github-actions

github-actions Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/8-gpu-models/test_longcat_flash_lite_fp8.py:

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/8-gpu-models/test_longcat_flash_lite_fp8.py

🚀 8-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/8-gpu-models/test_longcat_flash_lite_fp8.py

@b8zhong b8zhong added the run-ci CI: run the baseline test suite on this PR label Aug 6, 2026
@Fridge003
Fridge003 merged commit 28848bf into main Aug 6, 2026
82 of 112 checks passed
@Fridge003
Fridge003 deleted the brayden-fix-longcat-ngram-oe-vocab-size branch August 6, 2026 01:06
Fridge003 pushed a commit that referenced this pull request Aug 6, 2026
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
sagearc pushed a commit to sagearc/sglang that referenced this pull request Aug 13, 2026
sgl-project#33779)

Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
saturn-acc pushed a commit to saturn-acc/sglang that referenced this pull request Aug 16, 2026
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Atituiset pushed a commit to Atituiset/sglang that referenced this pull request Sep 10, 2026
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants