Repository navigation
Conversation
YAMY1234
requested review from
ByronHsu,
Duyi-Wang,
HaiShaw,
ShangmingCai,
Ying1123,
hnyls2002,
merrymercy,
sogalin and
xiezhq-hermann
as code owners
August 28, 2026 17:44
Collaborator
Author
|
/tag-and-rerun-ci |
This was referenced Sep 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation / 动机
A low-concurrency disaggregated configuration (prefill TP1/DEP4, decode TP8, attention TP gather disabled) showed strong native batch-size-1 CUDA graph performance, but longer workloads could permanently stop making request progress.
低并发 PD 分离配置(prefill TP1/DEP4、decode TP8、关闭 attention TP gather)可以获得明显的原生 batch-size-1 CUDA graph 性能收益,但长时间 workload 会永久停止请求推进。
The hang came from three synchronization/state-machine bugs:
该 hang 由三个同步与状态机问题共同造成:
对应地:decode KV polling 与请求广播可能在不同 TP rank 上进入不同 collective epoch;异构 TP 的 Mooncake fanout 在 defer 后可能重放已完成 destination;空 staging ring 可能保留旧 head,而新 watermark subscriber 又可能错过 allocator 当前状态。
Changes / 修改
Serialize decode polling and request broadcast through one ordered coordinator, using fixed-shape collective payloads and task-bound results.
Track completed fanout destinations across requeue so staging transfers are idempotent.
Reset an empty staging ring to a fresh round and bootstrap new subscribers with the current watermark.
Add focused regression tests for collective ordering, fixed-shape polling, fanout retry, ring wrap, and subscriber bootstrap.
通过单一有序 coordinator 串行化 decode polling 与请求广播,并使用固定形状 collective payload 和任务绑定结果。
在 requeue 期间记录已完成的 fanout destination,保证 staging transfer 幂等。
空 ring 进入新 round 时重置状态,并用当前 watermark 初始化新 subscriber。
为 collective ordering、固定形状 polling、fanout retry、ring wrap 和 subscriber bootstrap 增加定向回归测试。
Validation / 验证
Linux targeted tests: 18/18 passed.
Exact long-duration workload: 807/807 requests, 0 errors, 0 cancellations, 3630.00 s.
TTFT and inter-token-latency coverage: 100% / 100%.
P50/P90 output-token throughput per user: 467.70/447.07 → 640.33/685.20 tok/s/user (+36.9% / +53.3%).
No model forward or sampling math is changed.
Linux 定向测试:18/18 通过。
精确长时间 workload:807/807 requests、0 errors、0 cancellations、3630.00 秒。
TTFT 与 inter-token-latency coverage:100% / 100%。
P50/P90 每用户输出 token 吞吐:467.70/447.07 → 640.33/685.20 tok/s/user(+36.9% / +53.3%)。
未修改模型 forward 或 sampling 数学逻辑。
CI States
Latest PR Test (Base): ⏳ Run #33196082438
Latest PR Test (Extra): ❌ Run #33196082037
Latest PR Test (AMD ROCm 7.2): ⏳ Run #33196083112