-
Notifications
You must be signed in to change notification settings - Fork 9.3k
[CP] Add CP-v2 support for Kimi-Linear #31661
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
Fridge003
wants to merge
35
commits into
sgl-project:main
from
Fridge003:codex/kimi-linear-cp-v2-main
Closed
Changes from all commits
Commits
Show all changes
35 commits
Select commit
Hold shift + click to select a range
340cc02
docs: design Kimi-Linear CP-v2 transitions
Fridge003 f3f6bf4
docs: plan Kimi-Linear CP-v2 implementation
Fridge003 1d54fc2
test: specify Kimi-Linear CP-v2 entry gather
Fridge003 407354f
feat: gather Kimi-Linear CP input for KDA
Fridge003 3f452cd
test: specify KDA to MLA CP split
Fridge003 630f45d
feat: shard Kimi-Linear state before MLA
Fridge003 d5921b9
test: specify MLA to KDA CP gather
Fridge003 3c213d6
feat: gather Kimi-Linear residual before KDA
Fridge003 7b24010
test: specify Kimi decoder CP communicator wiring
Fridge003 2d7e74d
feat: wire CP-v2 transitions into Kimi layers
Fridge003 7f9362e
test: complete Kimi decoder fixture
Fridge003 92ce1af
test: isolate Kimi MLA layer wiring
Fridge003 70a8767
test: require CP-v2 for Kimi-Linear
Fridge003 898ac51
feat: enable CP-v2 for Kimi-Linear
Fridge003 cd71780
test: require Kimi input embedding accessor
Fridge003 8754d7e
feat: expose Kimi input embeddings to CP-v2
Fridge003 78c2083
test: cover Kimi CP-v2 no-op and zigzag round trip
Fridge003 02f464b
test: require KDA heads to use global TP
Fridge003 4f6220a
fix: partition KDA backend heads over global TP
Fridge003 bf733d9
test: require unfused KDA projection to use global TP
Fridge003 f0364b5
fix: shard unfused KDA projections over global TP
Fridge003 01b92d0
test: require KDA cache state to use global TP
Fridge003 e2ecec6
fix: shard KDA state cache over global TP
Fridge003 55f2a64
test: cover Kimi CP-v2 embedding keyword
Fridge003 b41d63a
fix: accept CP-v2 input embeddings in Kimi
Fridge003 5af9dd8
test: cover FlashInfer MLA CP-v2 dispatch
Fridge003 34ced26
feat: support FlashInfer MLA in CP-v2
Fridge003 5ff97cc
test: match FlashInfer MLA output layout
Fridge003 37c791f
fix: run Kimi KDA and MLP on global TP batches
Fridge003 697fff4
docs: target merged CP-v2 base
Fridge003 c422342
fix: address Kimi CP-v2 review feedback
Fridge003 8e9d306
docs: remove implementation planning files
Fridge003 c0f5526
feat: shard KDA heads over CP ranks
Fridge003 44e30b2
test: avoid CUDA stream in Kimi CP CPU test
Fridge003 fd7edf5
test: mock optional FlashInfer MLA wrapper
Fridge003 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,112 @@ | ||
| # Copyright 2023-2026 SGLang Team | ||
| # Licensed under the Apache License, Version 2.0 (the "License"); | ||
| # you may not use this file except in compliance with the License. | ||
| # You may obtain a copy of the License at | ||
| # | ||
| # http://www.apache.org/licenses/LICENSE-2.0 | ||
| # | ||
| # Unless required by applicable law or agreed to in writing, software | ||
| # distributed under the License is distributed on an "AS IS" BASIS, | ||
| # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| # See the License for the specific language governing permissions and | ||
| # limitations under the License. | ||
| # ============================================================================== | ||
|
|
||
| """CP-v2 token-layout transitions for Kimi-Linear decoder layers.""" | ||
|
|
||
| from __future__ import annotations | ||
|
|
||
| from typing import TYPE_CHECKING, Any, Optional, Tuple | ||
|
|
||
| import torch | ||
|
|
||
| from sglang.srt.layers.cp.utils import get_cp_strategy, is_cp_v2_active | ||
|
|
||
| if TYPE_CHECKING: | ||
| from sglang.srt.model_executor.forward_batch_info import ForwardBatch | ||
|
|
||
|
|
||
| class KimiLinearCPV2LayerCommunicator: | ||
| """Convert Kimi-Linear layer inputs between KDA and MLA token layouts.""" | ||
|
|
||
| def __init__( | ||
| self, | ||
| *, | ||
| is_kda_layer: bool, | ||
| previous_is_kda_layer: Optional[bool], | ||
| is_last_layer: bool = False, | ||
| ) -> None: | ||
| self._is_last_layer = is_last_layer | ||
| is_first_layer = previous_is_kda_layer is None | ||
| # CP-v2 enters the model sharded. Every MLP exits with a full TP batch. | ||
| self._gather_before_attn = is_kda_layer and is_first_layer | ||
| self._shard_before_attn = not is_kda_layer and not is_first_layer | ||
| self._gather_before_mlp = not is_kda_layer | ||
|
|
||
| def prepare_attn( | ||
| self, | ||
| hidden_states: torch.Tensor, | ||
| residual: Optional[torch.Tensor], | ||
| forward_batch: ForwardBatch, | ||
| stream: Optional[Any] = None, | ||
| ) -> Tuple[torch.Tensor, Optional[torch.Tensor]]: | ||
| if not is_cp_v2_active(forward_batch): | ||
| return hidden_states, residual | ||
|
|
||
| strategy = get_cp_strategy() | ||
| assert strategy is not None | ||
| if self._gather_before_attn: | ||
| if stream is None: | ||
| stream = torch.cuda.current_stream() | ||
| hidden_states = strategy.gather_hidden_states( | ||
| hidden_states, forward_batch, stream | ||
| ) | ||
| if residual is not None: | ||
| residual = strategy.gather_hidden_states( | ||
| residual, forward_batch, stream | ||
| ) | ||
| elif self._shard_before_attn: | ||
| hidden_states = strategy.shard_hidden_states(hidden_states, forward_batch) | ||
| if residual is not None: | ||
| residual = strategy.shard_hidden_states(residual, forward_batch) | ||
| return hidden_states, residual | ||
|
|
||
| def prepare_mlp( | ||
| self, | ||
| hidden_states: torch.Tensor, | ||
| residual: Optional[torch.Tensor], | ||
| forward_batch: ForwardBatch, | ||
| stream: Optional[Any] = None, | ||
| ) -> Tuple[torch.Tensor, Optional[torch.Tensor]]: | ||
| """Gather MLA outputs so normalization and MLP run with global TP.""" | ||
| if not self._gather_before_mlp or not is_cp_v2_active(forward_batch): | ||
| return hidden_states, residual | ||
|
|
||
| strategy = get_cp_strategy() | ||
| assert strategy is not None | ||
| if stream is None: | ||
| stream = torch.cuda.current_stream() | ||
| hidden_states = strategy.gather_hidden_states( | ||
| hidden_states, forward_batch, stream | ||
| ) | ||
| if residual is not None: | ||
| residual = strategy.gather_hidden_states(residual, forward_batch, stream) | ||
| return hidden_states, residual | ||
|
|
||
| def postprocess_layer( | ||
| self, | ||
| hidden_states: torch.Tensor, | ||
| residual: Optional[torch.Tensor], | ||
| forward_batch: ForwardBatch, | ||
| stream: Optional[Any] = None, | ||
| ) -> Tuple[torch.Tensor, Optional[torch.Tensor]]: | ||
| """Shard the full TP output for the model-boundary CP gather.""" | ||
| if not self._is_last_layer or not is_cp_v2_active(forward_batch): | ||
| return hidden_states, residual | ||
|
|
||
| strategy = get_cp_strategy() | ||
| assert strategy is not None | ||
| hidden_states = strategy.shard_hidden_states(hidden_states, forward_batch) | ||
| if residual is not None: | ||
| residual = strategy.shard_hidden_states(residual, forward_batch) | ||
| return hidden_states, residual | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -41,6 +41,7 @@ | |
| { | ||
| "Qwen3MoeForCausalLM", | ||
| "DeepseekV3ForCausalLM", | ||
| "KimiLinearForCausalLM", | ||
| } | ||
| ) | ||
|
|
||
|
|
||
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The
streamparameter inpostprocess_layeris unused. It is recommended to remove it from the method signature to keep the API clean and maintainable.