Skip to content

Extract ToModelProto for GraphViewer and Function - #1350

Closed
daquexian (daquexian) wants to merge 9 commits into
microsoft:masterfrom
daquexian:add_toproto
Closed

daquexian (daquexian) wants to merge 9 commits into
microsoft:masterfrom
daquexian:add_toproto

Conversation

@daquexian

Copy link
Copy Markdown
Contributor

Description:
As requested in #1220 (comment), extract ToModelProto method for onnxruntime::GraphViewer and onnxruntime::Function

Motivation and Context

@daquexian
daquexian (daquexian) requested a review from a team as a code owner July 5, 2019 14:40
@daquexian daquexian (daquexian) changed the title Extract ToModelProto for GraphViewer` and Function Extract ToModelProto for GraphViewer and Function Jul 5, 2019
Comment thread onnxruntime/core/graph/graph_utils.cc
@linkerzhang

Copy link
Copy Markdown
Contributor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 22 pipeline(s).

@daquexian

Copy link
Copy Markdown
Contributor Author

Ke Zhang (@linkerzhang) I have updated the branch :)

@linkerzhang

Copy link
Copy Markdown
Contributor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 22 pipeline(s).

@linkerzhang

Copy link
Copy Markdown
Contributor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 22 pipeline(s).

@stale

stale Bot commented Jul 3, 2020

Copy link
Copy Markdown

This issue has been automatically marked as stale due to inactivity and will be closed in 7 days if no further activity occurs. If further support is needed, please provide an update and/or more details.

@stale stale Bot added the wontfix label Jul 3, 2020
@stale

stale Bot commented Jul 11, 2020

Copy link
Copy Markdown

This issue has been automatically closed due to inactivity. Please reactivate if further support is needed.

@stale stale Bot closed this Jul 11, 2020
Dmitri Smirnov (yuslepukhin) pushed a commit that referenced this pull request Mar 17, 2026
Copilot AI pushed a commit that referenced this pull request Sep 16, 2026
## Summary

Plan contiguous GQA FlashDecode split-KV launches from fixed KV-cache capacity during CUDA graph capture and replay, while retaining live-sequence-length planning for ordinary eager execution.

## Why

CUDA graphs freeze launch geometry and workspace addresses at capture time. Planning NumSplits from the current live sequence length can become stale as the cache grows, preventing adaptive split-KV behavior from remaining valid across replay. Graph-enabled warmup now reserves capacity-sized workspace, capture uses the fixed-capacity plan, and eager decode avoids redundant capacity heuristic work.

## Behavior

- Capture/replay uses fixed cache capacity for stable NumSplits and workspace sizing
- Graph warmup reserves replay-sized workspace before capture
- Eager execution continues to tune from the live sequence length
- Active memset size remains limited to the launch plan
- Debug output reports the resolved NumSplits

## Validation

Focused host tests cover head sizes 64, 128, and 256; local-window and sequence-tail behavior; non-decode inputs; capture planning; and ordinary eager routing. For an SM108 configuration with live length 129 and capacity 4097, capture selected 17/17/22 splits for head sizes 64/128/256, while eager retained live-length plans. Independent review found one redundant eager heuristic computation, which is fixed in this commit.

No CUDA kernel timing is claimed because this Windows host did not have nvcc. The change preserves eager routing and targets graph planning correctness and replay-stable adaptive split-KV behavior.

Based on the mechanisms validated in justinchuby/onnx-genai#1340 and #1350.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants