Skip to content

feat(models): sync Novita mappings and pricing - #3657

Merged
steebchen merged 2 commits into
mainfrom
feat/novita-model-mappings
Aug 17, 2026
Merged

steebchen merged 2 commits into
mainfrom
feat/novita-model-mappings

Conversation

@steebchen

@steebchen steebchen commented Aug 17, 2026 •

Copy link
Copy Markdown
Member

Problem

Nine models the user asked about are served by Novita, but only five were in the
catalogue, and three of those carried values Novita no longer charges or accepts.

What changed

Everything below is per Novita mapping. Prices, contextSize and maxOutput
come from Novita's /openai/v1/models rate card (price_per_m / 10000 = USD/M);
capability flags were probed against the live deployments.

Added

Model externalId in / out / cached ($/M) ctx max out
qwen3.8-max qwen/qwen3.8-max 2.00 / 6.00 / 0.25 1,000,000 131,072
kimi-k2.7-code moonshotai/kimi-k2.7-code 0.95 / 4.00 / 0.19 262,144 262,144
glm-5.2 zai-org/glm-5.2 1.40 / 4.40 / 0.26 1,048,576 131,072

Capabilities probed live:

  • qwen3.8-max — vision ✅, json_object + json_schema ✅, tool_choice
    required/named 400 → supportedToolChoices: ["auto","none"], developer role
    400 → supportsDeveloperRole: false. Every reasoning_effort tier returns 200
    but none changes behaviour (thinking stays on even for none), so no tier is
    declared and reasoning_effort is not forwarded.
  • kimi-k2.7-code — vision ✅, JSON ✅; tool_choice required/named and the
    developer role both 400. none/minimal are accepted but do not disable
    thinking, so only low..max are declared.
  • glm-5.2 — vision explicitly unsupported (model features vision not support),
    all four tool-choice modes ✅, none reliably disables thinking. JSON mode is
    false: about half the responses put the object in reasoning_content with
    empty content and the rest markdown-fence it — json_schema behaves the same.

Corrected

  • kimi-k2.6 — input 0.95 → 0.80, output 4.00 → 3.40. Novita cut both rates;
    the cache rate (0.16) was already right.
  • glm-5.1 — input 1.40 → 1.38.
  • glm-5 and glm-5.1 — maxOutput 131100 → 131072. The deployment rejects
    the advertised value outright: max_tokens (current value: 131100) must be between 0 and 131072.
  • glm-5.1 — the mapping excluded reasoning_effort from supportedParameters
    on the basis that Novita disabled thinking whenever it was forwarded. That is
    no longer true: none disables and high reasons heavily, so the full tier
    list is declared and the stale note removed. JSON mode consistently
    markdown-fences the object, so healStreamingJsonOutput is now set.
  • hy3 — named function tool_choice 400s (required works), so
    supportedToolChoices: ["auto","none","required"]. Pricing, context and the
    jsonOutputSchema-only JSON support were already correct — Novita rejects
    json_object for this model with Supported formats: json_schema.

Deliberately not added

kimi-k2.5 is served by Novita but is not mapped here. About 40% of requests
to that deployment never respond — the connection stays open until it times out.
Measured 4/10 hangs on plain chat and 4/8 on tool-result round trips, while
kimi-k2.6 and kimi-k2.7-code on the same key were unaffected. A mapping
fenced off with stability: "unstable" + test: "skip" would still surface a
model we cannot serve, so nothing ships for it until Novita fixes the deployment.

Verification

Billing — one request per mapping through the gateway with x-no-fallback,
each hand-computed and matched against log.cost exactly (reasoning tokens are
recorded but correctly not billed on top of completion_tokens):

Mapping prompt × in + completion × out log.cost
qwen3.8-max 69×2e-6 + 77×6e-6 0.0006
kimi-k3 101×3e-6 + 260×15e-6 0.004203
kimi-k2.6 25×0.8e-6 + 274×3.4e-6 0.0009516
kimi-k2.7-code 26×0.95e-6 + 125×4e-6 0.0005247
glm-5 22×1e-6 + 241×3.2e-6 0.0007932
glm-5.1 30×1.38e-6 + 247×4.4e-6 0.0011282
glm-5.2 31×1.4e-6 + 225×4.4e-6 0.0010334
hy3 27×0.14e-6 + 249×0.58e-6 0.0001482

Prompt caching also verified: a repeated 4,011-token prompt reported
cached_tokens: 3968 as a subset of prompt_tokens, matching the cost engine's
assumption.

Specs — model-metadata, providers, realtime-models, compliance,
helpers and costs.spec.ts all pass. pnpm format and a full pnpm build
are clean.

e2e — TEST_MODELS="novita/qwen3.8-max,novita/kimi-k3,novita/kimi-k2.6,novita/kimi-k2.7-code,novita/glm-5,novita/glm-5.1,novita/glm-5.2,novita/hy3" FULL_MODE=true pnpm test:e2e
→ 196 passed, 1 failed. The one failure was a harness log-flush race
(Timed out waiting for log with request ID) on glm-5.2 non-streaming tool
calls, not a model failure: re-running the same suite scoped to
TEST_MODELS="novita/glm-5.2" FULL_MODE=true was fully green (106 passed, 0
failed).

Summary by CodeRabbit

  • New Features
    • Added Novita support for Qwen 3.8 Max, Kimi K2.7 Code, and GLM 5.2.
    • Added reasoning, vision, streaming, tool use, and structured JSON capabilities where supported.
    • Expanded tool-selection options for Tencent Hy3.
  • Updates
    • Reduced Novita pricing for Kimi K2.6.
    • Updated GLM model context, output limits, reasoning options, pricing, and streaming behavior.
    • Improved compatibility for tool choices and supported request parameters.

Add the missing Novita mappings for qwen3.8-max, kimi-k2.5, kimi-k2.7-code
and glm-5.2, and reconcile the existing ones against Novita's rate card and
live deployments.

Prices, context and max-output come from Novita's /openai/v1/models endpoint;
capability flags were probed against the live deployments.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 17, 2026 15:46

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: e98881a0-432b-4936-9207-cbc3d29cbe24

📥 Commits

Reviewing files that changed from the base of the PR and between 3c14741 and bd838fc.

📒 Files selected for processing (4)
  • packages/models/src/models/alibaba.ts
  • packages/models/src/models/moonshot.ts
  • packages/models/src/models/tencent.ts
  • packages/models/src/models/zai.ts

Included review availability: Your plan includes up to 4 reviews per rolling hour; 3 remain after this review.


Walkthrough

This change adds and updates Novita provider mappings for Qwen, Kimi, Tencent Hy3, and GLM models. It adjusts pricing, context and output limits, reasoning options, JSON behavior, tool choices, streaming behavior, and developer-role support.

Changes

Novita model mappings

Layer / File(s) Summary
Model capabilities and pricing
packages/models/src/models/alibaba.ts, packages/models/src/models/moonshot.ts, packages/models/src/models/zai.ts
Adds Novita mappings for qwen3.8-max, kimi-k2.7-code, and glm-5.2. Updates pricing and capability settings for kimi-k2.6, glm-5.1, and glm-5.
Tool-choice support
packages/models/src/models/tencent.ts
Adds auto, none, and required tool-choice modes to the Novita Tencent Hy3 configuration.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk: ⚪ Minimal · up to bd838

This PR synchronizes Novita model mappings, pricing, capabilities, and limits while keeping the known unstable deployment out of routing; no actionable merge-blocking risk remains after normal checks and review.

Possibly related PRs

Suggested reviewers: smakosh, vicovaro

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main changes to Novita model mappings and pricing.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/novita-model-mappings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 51e0fa3ff2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/models/src/models/moonshot.ts Outdated
@steebchen
steebchen enabled auto-merge (squash) August 17, 2026 16:01
About 40% of requests to novita's kimi-k2.5 deployment never respond, so
ship nothing rather than a mapping fenced off by stability/test flags.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@steebchen
steebchen merged commit 0d3da39 into main Aug 17, 2026
18 checks passed
@steebchen
steebchen deleted the feat/novita-model-mappings branch August 17, 2026 16:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants