Skip to content

mtmd: fix gemma 4 projector pre_norm#23822

Merged
ngxson merged 1 commit into
masterfrom
xsn/fix_gemma4_prenorm
May 28, 2026
Merged

mtmd: fix gemma 4 projector pre_norm#23822
ngxson merged 1 commit into
masterfrom
xsn/fix_gemma4_prenorm

Conversation

@ngxson
Copy link
Copy Markdown
Contributor

@ngxson ngxson commented May 28, 2026

Overview

The early pre-release version of gemma 4 uses post norm for multimodal projector, but it was later swapped to pre-norm and I did not notice about that.

This should fix some vision-related problems with gemma 4.

Ref python code: https://github.com/huggingface/transformers/blob/1656d90b774d94c30af24113e60e926fc2f39072/src/transformers/models/gemma4/modeling_gemma4.py#L2068-L2092

Test:

image image image

Requirements

@ngxson ngxson requested a review from a team as a code owner May 28, 2026 14:42
@ngxson
Copy link
Copy Markdown
Contributor Author

ngxson commented May 28, 2026

asking for 2nd approval @ggml-org/maintainers 🙏

@ngxson ngxson merged commit c8914ad into master May 28, 2026
27 checks passed
gabe-l-hart added a commit to gabe-l-hart/llama.cpp that referenced this pull request May 28, 2026
* origin/master: (32 commits)
hexagon: basic/generic op fusion support and RMS_NORM+MUL fusion (ggml-org#23835)
mtmd-debug: add color and rainbow mode (ggml-org#23829)
mtmd: fix gemma 4 projector pre_norm (ggml-org#23822)
opencl: move backend info printing into its own function (ggml-org#23702)
ci : run ui publish on ubuntu-slim (ggml-org#23818)
ui: fix audio and video modality detection (ggml-org#23756)
ci : releases use Github-hosted builds for the UI (ggml-org#23823)
app : improve help output (ggml-org#23805)
mtmd: n_head_kv defaults to n_head (ggml-org#23782)
mtmd: fix gemma 4 audio rms norm eps (ggml-org#23815)
ci : change Vulkan builds to Release to reduce ccache (ggml-org#23820)
arg: Add LLAMA_ARG_API_KEY_FILE environment variable for --api-key-file (ggml-org#23167)
test-llama-archs: fix table format [no release] (ggml-org#23810)
ggml: auto apply iGPU flag CUDA/HIP if integrated device (ggml-org#23007)
mmvq Optim: add MMVQ_PARAMETERS_TURING(mmvq_parameter_table_id) for … (ggml-org#23729)
CUDA: route batch>=4 quantized matmul to MMQ on AMD MFMA hardware (ggml-org#23227)
server: minor tweaks to use more cpp features (ggml-org#23785)
hexagon: minor refresh for HMX FA and MM (ggml-org#23796)
vulkan: fast path for walsh-hadamard transform (ggml-org#23687)
chat : add Granite 4.1 chat template (ggml-org#23518)
...
fewtarius pushed a commit to fewtarius/llama.cpp that referenced this pull request May 30, 2026
turbo-tan pushed a commit to turbo-tan/llama.cpp-tq3 that referenced this pull request Jun 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants