Skip to content

feat(models): add quantization to nebius mappings - #2994

Merged
smakosh merged 1 commit into
theopenco:mainfrom
AmineAce:feat/quant-nebius
Jul 11, 2026
Merged

smakosh merged 1 commit into
theopenco:mainfrom
AmineAce:feat/quant-nebius

Conversation

@AmineAce

@AmineAce AmineAce commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Adds quantization to Nebius provider mappings across the catalog, per #2946.
Checked every Nebius mapping against Nebius Token Factory's model catalog. Added the field only where the endpoint details panel explicitly showed a "Quantization" value.
Confirmed:

Qwen3-235B-A22B-Instruct-2507 → fp8
Qwen3-32B → fp8
Qwen2.5-VL-72B-Instruct → fp8
Qwen3-30B-A3B-Instruct-2507 → fp8
Qwen3-Next-80B-A3B-Thinking → fp8
Qwen3.5-397B-A17B → fp4
Llama-3_1-Nemotron-Ultra-253B-v1 → fp8
Llama-3.3-70B-Instruct → fp8
gpt-oss-120b → fp4
MiniMax-M2.5 → fp4
gemma-3-27b-it → fp8

Checked but not currently listed on Nebius (20, left unchanged)

QwQ-32B
Qwen3-235B-A22B-Thinking-2507
Qwen3-14B
Qwen3-30B-A3B
Qwen2.5-Coder-7B-fast
Qwen2.5-32B-Instruct
Qwen2.5-72B-Instruct
Qwen2-VL-72B-Instruct
Qwen3-Coder-480B-A35B-Instruct
Qwen3-Coder-30B-A3B-Instruct
Qwen3-30B-A3B-Thinking-2507
DeepSeek-V3
DeepSeek-R1-0528
DeepSeek-V3.2
Meta-Llama-3.1-8B-Instruct
Meta-Llama-3.1-405B-Instruct
Kimi-K2-Instruct
Kimi-K2.5
GLM-5
Hermes-3-Llama-405B

Ran pnpm format; only these 5 files changed.

Summary by CodeRabbit

  • Model Updates
    • Added explicit quantization metadata for supported model variants across Alibaba, Google, Meta, MiniMax, and OpenAI providers.
    • Identified FP8 and FP4 configurations for improved model capability and compatibility reporting.

@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 7ea4e874-450e-4b70-88fa-1a0a6887ff7e

📥 Commits

Reviewing files that changed from the base of the PR and between 22d1af7 and 16a9e45.

📒 Files selected for processing (5)
  • packages/models/src/models/alibaba.ts
  • packages/models/src/models/google.ts
  • packages/models/src/models/meta.ts
  • packages/models/src/models/minimax.ts
  • packages/models/src/models/openai.ts

Walkthrough

Selected Nebius provider mappings now explicitly declare fp8 or fp4 quantization across Alibaba, Google, Meta, Minimax, and OpenAI model definitions.

Changes

Provider quantization metadata

Layer / File(s) Summary
FP8 provider metadata
packages/models/src/models/alibaba.ts, packages/models/src/models/google.ts, packages/models/src/models/meta.ts
Selected Nebius Qwen, Gemma, and Llama provider entries now include quantization: "fp8".
FP4 provider metadata
packages/models/src/models/alibaba.ts, packages/models/src/models/minimax.ts, packages/models/src/models/openai.ts
Nebius provider entries for Qwen35, MiniMax, and GPT-OSS now include quantization: "fp4".

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

Suggested reviewers: steebchen, smakosh

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: adding quantization metadata to Nebius model mappings.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@smakosh
smakosh added this pull request to the merge queue Jul 11, 2026
Merged via the queue into theopenco:main with commit f9e7e3e Jul 11, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants