Fix Gemma4 GenAI regression: disable past_present_share_buffer for dual head_dim - #274
Merged
Conversation
Performance Comparison
|
37 tasks
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…_dim Models with different head_dim across layers (e.g., Gemma4 with head_dim=256 for sliding and global_head_dim=512 for full-attention) cannot use shared KV buffers. Dynamically detect dual head_dim from config and set past_present_share_buffer=false. Implementation: - GenaiConfigGenerator: add _search_overrides dict applied in generate() - auto_export: check config.global_head_dim != config.head_dim 134 tests pass (15 gemma4 + 119 ort_genai). Signed-off-by: Justin Chu <justinchu@microsoft.com>
justinchuby
force-pushed
the
fix-attention-bias-dtype
branch
from
May 6, 2026 16:01
c531042 to
2b511ae
Compare
rui-ren
approved these changes
May 6, 2026
Contributor
|
👍 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fix CPU regression where Gemma4 GenAI generation fails with:
Root Cause
past_present_share_bufferwas set totruein genai_config.json because thedefaultEP maps tocpu(which hassupports_past_present_share_buffer=True). However, Gemma4 has dual head_dim (256 for sliding-window, 512 for full-attention), and shared buffers require uniform head_dim across all KV cache layers.Fix
Force
past_present_share_buffer=falsein genai_config.json for all Gemma4 model types.Testing