Conversation
dblate
marked this pull request as ready for review
August 31, 2026 09:46
dblate
requested review from
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
August 31, 2026 09:46
Contributor
Author
|
/tag-and-rerun-ci |
Contributor
Author
|
补充验证结果:在 H20(CUDA 12.9)开发空间运行 LayerSplit MLA page-first fallback,结果为 当前 CI 仍未开始真实测试,原因是缺少上游要求的 |
dblate
force-pushed
the
glm52-cp-mooncake
branch
from
August 31, 2026 10:21
85d886f to
cff50f6
Compare
dblate
marked this pull request as draft
September 1, 2026 01:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Make GLM-5.2 LayerSplit context-parallel cache backup compatible with page-first HiCache storage and Mooncake L3. CP ranks own distinct token shards, so each rank must publish its cache under rank-scoped keys.
Depends on #37231 for the per-layer LF-to-PF transfer operator.
Modifications
Accuracy Tests
1 passed.1 passed.24 passed, 1 skipped.git diff --checkand Python AST parsing passed for changed files.Speed Tests and Profiling
No end-to-end serving benchmark was run for this runtime-only change. The underlying LF-to-PF kernel microbenchmark is reported in #37231.
Checklist
CI States
Latest PR Test (Base): ❌ Run #33382072906
Latest PR Test (Extra): ❌ Run #33382072655
Latest PR Test (AMD ROCm 7.2): ❌ Run #33382072894