Skip to content

Keep lm-head quantization scales for nemotron --mtp export - #26903

Merged
CISC merged 1 commit into
ggml-org:masterfrom
ynankani:ynankani/nemotron_nvfp4_mtp
Aug 11, 2026
Merged

Keep lm-head quantization scales for nemotron --mtp export#26903
CISC merged 1 commit into
ggml-org:masterfrom
ynankani:ynankani/nemotron_nvfp4_mtp

Conversation

@ynankani

Copy link
Copy Markdown
Contributor

Overview

This PR handles the --mtp export for the Nemotron Nano model with quantized lm_head. Performance is not yet optimal and similar to #26725 this PR also depends on PR #26623 being merged first.
I tested this PR on the top of PR #26623, and the performance looks good.

Additional information

Requirements

Signed-off-by: ynankani <ynankani@nvidia.com>
@ynankani
ynankani marked this pull request as ready for review August 11, 2026 12:38
@ynankani
ynankani requested a review from CISC as a code owner August 11, 2026 12:38
@ruixiang63

Copy link
Copy Markdown
Member

@CISC Could you take a look? This is must for NVFP4 quantized model.

@CISC
CISC merged commit 5d16e81 into ggml-org:master Aug 11, 2026
5 checks passed
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 12, 2026
Ooooze pushed a commit to AtomicBot-ai/atomic-llama-cpp-turboquant-nightly that referenced this pull request Aug 14, 2026
…g#26903)

Signed-off-by: ynankani <ynankani@nvidia.com>
(cherry picked from commit 5d16e81)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants