Skip to content

GLM-5.3-Flash support - #36507

Merged
Fridge003 merged 65 commits into
mainfrom
xinyuan/glm-5.3-flash-support
Sep 6, 2026
Merged

Fridge003 merged 65 commits into
mainfrom
xinyuan/glm-5.3-flash-support

Conversation

@JustinTong0323

@JustinTong0323 JustinTong0323 commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator

Support GLM-5.3-Flash.


CI States

Latest PR Test (Base): 🚫 Run #34024595258
Latest PR Test (Extra): 🚫 Run #34024595202
Latest PR Test (AMD ROCm 7.2): 🚫 Run #34024595305

Squashed from zai/glm-5-next (PR #4) onto main @ 27c3636.

Hybrid MoE architecture: MLA attention, DSA sparse attention with KPool
indexer, KDA linear attention, mHC, and native NEXTN draft layer for
speculative decoding (fixed 5/1/6 and adaptive). Multimodal serving for
image and video (up to 240K visual tokens), with encoder disaggregation
(EPD) for ultra-long-video workloads. FP8 weights with BF16 KV cache;
verified on 4x GB300 (TP4/EP4) and 8x H100 (TP8/EP8).

Co-authored-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk>
Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: zanes-ops <zanes@nvidia.com>
@Fridge003

Copy link
Copy Markdown
Collaborator

@ormandj

ormandj commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Three follow-ups auto-closed when this PR merged and its base branch was deleted. Their replacements are now open against main:

Previous PR Replacement Change
#38161 #38212 Restore hybrid HiCache DSA indexes and separate divergent compressed prefixes
#38162 #38213 Opt-in speculative attention-metadata fusion
#37539 #38214 Precompile vision attention and MLP activations before KV allocation

The patches are unchanged (matching stable Git patch IDs), applied to main 938dc5621d. Author CPU checks on the new heads passed 41 tests and 12 subtests. GPU receipts remain tied to the original revisions; these main-based heads have not had new GPU runs. #37813 and the original PRs now point to the replacements.

@linkedlist771

Copy link
Copy Markdown

#37744 is automatically closed, reopen the port PR based on the main branch as #38254.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.