Skip to content

Support mixed-precision and hybrid model configurations - #1

Closed
martingiovannigeyer wants to merge 4 commits into
mainfrom
codex/fix-mixed-precision-compressed-tensors
Closed

martingiovannigeyer wants to merge 4 commits into
mainfrom
codex/fix-mixed-precision-compressed-tensors

Conversation

@martingiovannigeyer

@martingiovannigeyer martingiovannigeyer commented Jul 17, 2026 •

Copy link
Copy Markdown
Owner

Motivation

Support mixed-precision compressed-tensors checkpoints and the hybrid models covered by this patch without checkpoint-specific runtime forks.

Changes

  • parse and preserve each compressed-tensors config groups own format
  • retain exact ignore scoping and targeted ParallelLMHead quantization
  • fix Ling lightning-attention and Nemotron Puzzle initialization
  • fix compressed NVFP4 MoE W13 loading for FlashInfer CUTLASS by forwarding the per-layer [up, gate] ordering contract through the compressed-tensors wrapper

The last change is intentionally narrow. The NVIDIA ModelOpt method already exposes the same ordering contract; the compressed-tensors wrapper previously hid it and loaded [gate, up] into a CUTLASS kernel expecting [up, gate].

Validation

  • 9 focused CPU regression tests pass in the pinned SGLang CUDA 13 runtime
  • Qwen3.6 27B Unsloth NVFP4 repeated 1,000/1,000 outputs on RTX PRO 6000 and RTX 5090
  • before the W13 fix, Qwen3.6 35B Unsloth NVFP4 agreement was roughly 42.5% with both the first patch commit and the three-commit patch head
  • after 2b0dbfd9b8, 35B agreement is 95.7-95.8% on RTX PRO 6000 and 96.2% on RTX 5090
  • full 1,000-request repeats completed on both GPUs with no MTP/speculative decoding
  • Ling and Nemotron initialization regressions remain covered

No new checkpoint format, kernel, runtime flag, or forward hot path is introduced.

@martingiovannigeyer martingiovannigeyer changed the title Fix mixed-precision compressed-tensors loading Support mixed-precision and hybrid model configurations Jul 18, 2026
@martingiovannigeyer

Copy link
Copy Markdown
Owner Author

Superseded by two scoped PRs: one for mixed-precision/NVFP4 support and one for the independent Ling/Nemotron initialization fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant