Repository navigation
feat(models): add kimi-k3 - #3092
Conversation
Add Moonshot's Kimi K3 flagship model (released 2026-07-16): - 1M-token context window (1,048,576) with matching max output - $3.00/M input, $0.30/M cached input, $15.00/M output - Always-on thinking configured via the native top-level reasoning_effort field (currently only "max"), unlike the K2-era binary thinking toggle; the moonshot handler now forwards the effort as-is for K3 and collapses disable requests onto the provider default - K3 documents max_completion_tokens (not max_tokens), so translate - Vision, tools (all tool_choice modes), strict JSON schema output - Curated into the coding category; marked closed-source (API-only at launch, no open weights published) Claude-Session: https://claude.ai/code/session_01XETTwBHG4TF4gHDFR4VU2G
WalkthroughAdds Kimi K3 to Moonshot model metadata and model-directory categories. Its request builder maps token limits to ChangesKimi K3 Moonshot integration
Estimated code review effort: 3 (Moderate) | ~20 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
packages/actions/src/prepare-request-body.ts (1)
2013-2013: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueDefensively support future Kimi K3 variants.
Consider using
.startsWith("kimi-k3")instead of an exact match. This ensures that future variants (e.g.,kimi-k3-turboorkimi-k3.5) automatically inherit themax_completion_tokensmapping and nativereasoning_effortrouting without requiring further updates to this conditional logic, similar to howgpt-5models are handled above.💡 Proposed refactor
- const isKimiK3 = usedInternalModel === "kimi-k3"; + const isKimiK3 = usedInternalModel.startsWith("kimi-k3");🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/actions/src/prepare-request-body.ts` at line 2013, Update the isKimiK3 check in the request-body preparation logic to use a prefix match for model names beginning with “kimi-k3” rather than an exact equality check, so variants such as kimi-k3-turbo receive the same token mapping and reasoning routing.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@packages/actions/src/prepare-request-body.ts`:
- Line 2013: Update the isKimiK3 check in the request-body preparation logic to
use a prefix match for model names beginning with “kimi-k3” rather than an exact
equality check, so variants such as kimi-k3-turbo receive the same token mapping
and reasoning routing.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 8ac59a0d-cadd-410d-b08f-55f5bfa233c7
📒 Files selected for processing (4)
packages/actions/src/prepare-request-body.spec.tspackages/actions/src/prepare-request-body.tspackages/models/src/models/moonshot.tspackages/shared/src/components/models-directory/model-category-filters.ts
Summary
Adds Moonshot's Kimi K3 flagship model to the catalogue, released today (2026-07-16). Docs: pricing, quickstart, thinking effort, tool choice.
Model definition (
packages/models)kimi-k3on themoonshotprovider (sameapi.moonshot.ai/v1endpoint)max_completion_tokenscan be set up to 1048576)reasoning: truewithreasoningEfforts: ["max"]— K3 always thinks; effort is set via the top-levelreasoning_effortfield which currently accepts only"max"(more levels announced as coming)tool_choicemodes (auto/none/required/named function — K3 lifts the K2-era forced-tool-choice restriction), strict JSON schema output, streaming with separatereasoning_contentdeltassupportedParametersexcludes temperature/top_p/penalties (K3 fixes them at 1.0/0.95/0), so the gateway strips them instead of forwardingGateway request mapping (
packages/actions)K3's API differs from the K2 generation in two ways, handled in the
moonshotcase ofprepareRequestBody:reasoning_effortis forwarded natively for K3 (K2-era models get the binarythinking: { type }toggle, which K3 does not accept); disable requests (none/minimal) collapse onto the provider default since K3 cannot turn thinking offmax_tokensis translated tomax_completion_tokens(the only output-cap parameter K3 documents)Unit tests added for all three behaviors.
Models directory (
packages/shared)kimi-k3to the curated coding categorymoonshotfamily otherwise defaults to open-sourceTesting
pnpm exec vitest run packages/actions packages/models packages/shared/src/model-categories.spec.ts— all pass (including 3 new K3 tests)pnpm build— 17/17 tasks successfulLLM_MOONSHOT_API_KEYin local env); behavior is per the official K3 docs and should be smoke-tested once deployed with a keyhttps://claude.ai/code/session_01XETTwBHG4TF4gHDFR4VU2G
Summary by CodeRabbit