Skip to content

Reduce CPU Concat ops in Gemma4 audio encoder - #270

Merged
justinchuby merged 1 commit into
mainfrom
audio-concat-reduction
May 6, 2026
Merged

Reduce CPU Concat ops in Gemma4 audio encoder#270
justinchuby merged 1 commit into
mainfrom
audio-concat-reduction

Conversation

@justinchuby

Copy link
Copy Markdown
Member

Replace dynamic Shape+Concat with static value_ints in Reshape. Reduces CPU Concat from 49 to 24 (51%).

Replace dynamic Shape+Concat shape construction with static
value_ints in Reshape ops. ONNX Reshape treats 0 as 'copy from
input dim', eliminating the need for runtime Shape extraction.

Changes:
- Q/K/V reshape [B, T, H*D] → [B, T, H, D]: use [0, 0, H, D]
- Output reshape [B, T, H, D] → [B, T, H*D]: use [0, 0, -1]
- Subsample reshape [B, T', F', C] → [B, T', F'*C]: use [0, 0, -1]

Reduces audio encoder CPU Concat nodes from 49 to 24 (51% reduction).
Remaining Concats are from causal window mask construction that
genuinely requires dynamic shapes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
@github-actions

github-actions Bot commented May 6, 2026

Copy link
Copy Markdown

Performance Comparison

Comparing 496d2612f5e5e3

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0%
bert (feature-extraction) num_nodes 60 60 +0.0%
falcon model_size_bytes 364 KB 364 KB +0.0%
falcon num_nodes 66 66 +0.0%
gemma2 model_size_bytes 428 KB 428 KB +0.0%
gemma2 num_nodes 107 107 +0.0%
gpt2 model_size_bytes 388 KB 388 KB +0.0%
gpt2 num_nodes 53 53 +0.0%
llama model_size_bytes 425 KB 425 KB +0.0%
llama num_nodes 61 61 +0.0%
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0%
llama (static-cache) num_nodes 58 58 +0.0%
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0%
mamba (ssm-text-generation) num_nodes 98 98 +0.0%
phi3 model_size_bytes 421 KB 421 KB +0.0%
phi3 num_nodes 59 59 +0.0%
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0%
phi3 (static-cache) num_nodes 56 56 +0.0%
qwen2 model_size_bytes 425 KB 425 KB +0.0%
qwen2 num_nodes 61 61 +0.0%
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0%
qwen2 (static-cache) num_nodes 58 58 +0.0%
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0%
qwen3_5_moe (hybrid-text-generation) num_nodes 275 275 +0.0%
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0%
qwen3_5_text (hybrid-text-generation) num_nodes 129 129 +0.0%
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0%
qwen3_5_vl (hybrid-qwen-vl) num_nodes 408 408 +0.0%
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0%
t5 (seq2seq) num_nodes 166 166 +0.0%
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0%
whisper (speech-to-text) num_nodes 128 128 +0.0%

No performance regressions.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR reduces CPU-side Concat nodes in the Gemma4 audio encoder graph by replacing dynamic Shape+Concat-constructed reshape shapes with static Constant(value_ints=...) shapes that use ONNX Reshape’s 0 sentinel (“copy from input dim”). This aligns with the PR goal of reducing CPU ops on CUDA EP.

Changes:

  • Replace dynamic Shape+Concat shape construction in ConvSubsampling flattening with Reshape(..., Constant([0, 0, -1])).
  • Replace dynamic Shape+Concat shape construction for Q/K/V and output context reshapes with Reshape(..., Constant([0, 0, num_heads, head_dim])) and Constant([0, 0, -1]).

@codecov

codecov Bot commented May 6, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented May 6, 2026

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing 496d2612f5e5e3

Model Sub-model Changes Status
bert (feature-extraction) model 0
falcon model 0
gemma2 model 0
gemma4 (gemma4) decoder 0
gemma4 (gemma4) embedding 0
gemma4 (gemma4) vision_encoder 0
gemma4_text model 0
gpt2 model 0
llama model 0
llama (static-cache) model 0
mamba (ssm-text-generation) model 0
phi3 model 0
phi3 (static-cache) model 0
qwen model 0
qwen (static-cache) model 0
qwen2 model 0
qwen2 (static-cache) model 0
qwen2_moe model 0
qwen2_moe (static-cache) model 0
qwen3 model 0
qwen3 (static-cache) model 0
qwen3_5_moe (hybrid-text-generation) model 0
qwen3_5_text (hybrid-text-generation) model 0
qwen3_5_vl (hybrid-qwen-vl) decoder 0
qwen3_5_vl (hybrid-qwen-vl) embedding 0
qwen3_5_vl (hybrid-qwen-vl) vision_encoder 0
qwen3_moe model 0
qwen3_moe (static-cache) model 0
qwen3_next (hybrid-text-generation) model 0
t5 (seq2seq) decoder 0
t5 (seq2seq) encoder 0
whisper (speech-to-text) decoder 0
whisper (speech-to-text) encoder 0

No architecture changes detected.


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

@justinchuby
justinchuby merged commit b1d3f1f into main May 6, 2026
25 of 27 checks passed
@justinchuby
justinchuby deleted the audio-concat-reduction branch May 6, 2026 15:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants