Repository navigation
feat(models): sync Novita mappings and pricing - #3657
Conversation
Add the missing Novita mappings for qwen3.8-max, kimi-k2.5, kimi-k2.7-code and glm-5.2, and reconcile the existing ones against Novita's rate card and live deployments. Prices, context and max-output come from Novita's /openai/v1/models endpoint; capability flags were probed against the live deployments. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (4)
Included review availability: Your plan includes up to 4 reviews per rolling hour; 3 remain after this review. WalkthroughThis change adds and updates Novita provider mappings for Qwen, Kimi, Tencent Hy3, and GLM models. It adjusts pricing, context and output limits, reasoning options, JSON behavior, tool choices, streaming behavior, and developer-role support. ChangesNovita model mappings
Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to This PR synchronizes Novita model mappings, pricing, capabilities, and limits while keeping the known unstable deployment out of routing; no actionable merge-blocking risk remains after normal checks and review. Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 51e0fa3ff2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
About 40% of requests to novita's kimi-k2.5 deployment never respond, so ship nothing rather than a mapping fenced off by stability/test flags. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Problem
Nine models the user asked about are served by Novita, but only five were in the
catalogue, and three of those carried values Novita no longer charges or accepts.
What changed
Everything below is per Novita mapping. Prices,
contextSizeandmaxOutputcome from Novita's
/openai/v1/modelsrate card (price_per_m / 10000= USD/M);capability flags were probed against the live deployments.
Added
externalIdqwen3.8-maxqwen/qwen3.8-maxkimi-k2.7-codemoonshotai/kimi-k2.7-codeglm-5.2zai-org/glm-5.2Capabilities probed live:
qwen3.8-max— vision ✅,json_object+json_schema✅,tool_choicerequired/named 400 →supportedToolChoices: ["auto","none"], developer role400 →
supportsDeveloperRole: false. Everyreasoning_efforttier returns 200but none changes behaviour (thinking stays on even for
none), so no tier isdeclared and
reasoning_effortis not forwarded.kimi-k2.7-code— vision ✅, JSON ✅;tool_choicerequired/named and thedeveloper role both 400.
none/minimalare accepted but do not disablethinking, so only
low..maxare declared.glm-5.2— vision explicitly unsupported (model features vision not support),all four tool-choice modes ✅,
nonereliably disables thinking. JSON mode isfalse: about half the responses put the object inreasoning_contentwithempty
contentand the rest markdown-fence it —json_schemabehaves the same.Corrected
kimi-k2.6— input0.95 → 0.80, output4.00 → 3.40. Novita cut both rates;the cache rate (0.16) was already right.
glm-5.1— input1.40 → 1.38.glm-5andglm-5.1—maxOutput131100 → 131072. The deployment rejectsthe advertised value outright:
max_tokens (current value: 131100) must be between 0 and 131072.glm-5.1— the mapping excludedreasoning_effortfromsupportedParameterson the basis that Novita disabled thinking whenever it was forwarded. That is
no longer true:
nonedisables andhighreasons heavily, so the full tierlist is declared and the stale note removed. JSON mode consistently
markdown-fences the object, so
healStreamingJsonOutputis now set.hy3— named functiontool_choice400s (requiredworks), sosupportedToolChoices: ["auto","none","required"]. Pricing, context and thejsonOutputSchema-only JSON support were already correct — Novita rejectsjson_objectfor this model withSupported formats: json_schema.Deliberately not added
kimi-k2.5is served by Novita but is not mapped here. About 40% of requeststo that deployment never respond — the connection stays open until it times out.
Measured 4/10 hangs on plain chat and 4/8 on tool-result round trips, while
kimi-k2.6andkimi-k2.7-codeon the same key were unaffected. A mappingfenced off with
stability: "unstable"+test: "skip"would still surface amodel we cannot serve, so nothing ships for it until Novita fixes the deployment.
Verification
Billing — one request per mapping through the gateway with
x-no-fallback,each hand-computed and matched against
log.costexactly (reasoning tokens arerecorded but correctly not billed on top of
completion_tokens):log.costqwen3.8-maxkimi-k3kimi-k2.6kimi-k2.7-codeglm-5glm-5.1glm-5.2hy3Prompt caching also verified: a repeated 4,011-token prompt reported
cached_tokens: 3968as a subset ofprompt_tokens, matching the cost engine'sassumption.
Specs —
model-metadata,providers,realtime-models,compliance,helpersandcosts.spec.tsall pass.pnpm formatand a fullpnpm buildare clean.
e2e —
TEST_MODELS="novita/qwen3.8-max,novita/kimi-k3,novita/kimi-k2.6,novita/kimi-k2.7-code,novita/glm-5,novita/glm-5.1,novita/glm-5.2,novita/hy3" FULL_MODE=true pnpm test:e2e→ 196 passed, 1 failed. The one failure was a harness log-flush race
(
Timed out waiting for log with request ID) onglm-5.2non-streaming toolcalls, not a model failure: re-running the same suite scoped to
TEST_MODELS="novita/glm-5.2" FULL_MODE=truewas fully green (106 passed, 0failed).
Summary by CodeRabbit