Repository navigation
Extract ToModelProto for GraphViewer and Function - #1350
Closed
daquexian (daquexian) wants to merge 9 commits into
Closed
daquexian (daquexian) wants to merge 9 commits into
daquexian (daquexian) wants to merge 9 commits into
Conversation
Contributor
|
/azp run |
|
Azure Pipelines successfully started running 22 pipeline(s). |
Contributor
Author
|
Ke Zhang (@linkerzhang) I have updated the branch :) |
Contributor
|
/azp run |
|
Azure Pipelines successfully started running 22 pipeline(s). |
Contributor
|
/azp run |
|
Azure Pipelines successfully started running 22 pipeline(s). |
|
This issue has been automatically marked as stale due to inactivity and will be closed in 7 days if no further activity occurs. If further support is needed, please provide an update and/or more details. |
|
This issue has been automatically closed due to inactivity. Please reactivate if further support is needed. |
Dmitri Smirnov (yuslepukhin)
pushed a commit
that referenced
this pull request
Mar 17, 2026
Copilot AI
pushed a commit
that referenced
this pull request
Sep 16, 2026
## Summary Plan contiguous GQA FlashDecode split-KV launches from fixed KV-cache capacity during CUDA graph capture and replay, while retaining live-sequence-length planning for ordinary eager execution. ## Why CUDA graphs freeze launch geometry and workspace addresses at capture time. Planning NumSplits from the current live sequence length can become stale as the cache grows, preventing adaptive split-KV behavior from remaining valid across replay. Graph-enabled warmup now reserves capacity-sized workspace, capture uses the fixed-capacity plan, and eager decode avoids redundant capacity heuristic work. ## Behavior - Capture/replay uses fixed cache capacity for stable NumSplits and workspace sizing - Graph warmup reserves replay-sized workspace before capture - Eager execution continues to tune from the live sequence length - Active memset size remains limited to the launch plan - Debug output reports the resolved NumSplits ## Validation Focused host tests cover head sizes 64, 128, and 256; local-window and sequence-tail behavior; non-decode inputs; capture planning; and ordinary eager routing. For an SM108 configuration with live length 129 and capacity 4097, capture selected 17/17/22 splits for head sizes 64/128/256, while eager retained live-length plans. Independent review found one redundant eager heuristic computation, which is fixed in this commit. No CUDA kernel timing is claimed because this Windows host did not have nvcc. The change preserves eager routing and targets graph planning correctness and replay-stable adaptive split-KV behavior. Based on the mechanisms validated in justinchuby/onnx-genai#1340 and #1350. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description:
As requested in #1220 (comment), extract
ToModelProtomethod foronnxruntime::GraphViewerandonnxruntime::FunctionMotivation and Context
Avoid duplicate codes in execution providers
Initial PR for NNAPI execution provider #1220 (comment)