Repository navigation
feat(models): add InclusionAI Ling-3.0-flash - #3515
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
🚧 Files skipped from review as they are similar to previous changes (2)
WalkthroughAdds InclusionAI Ling-3.0-flash model definitions for DeepInfra and Novita. Maps reasoning controls to ChangesLing-3.0-flash reasoning flow
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant Client
participant Gateway
participant prepareRequestBody
participant DeepInfraOrNovita
Client->>Gateway: Ling-3.0-flash request
Gateway->>prepareRequestBody: reasoning and tool parameters
prepareRequestBody->>DeepInfraOrNovita: chat_template_kwargs and translated tools
DeepInfraOrNovita-->>Gateway: upstream completion
Gateway-->>Client: gateway response
Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
6247d03 to
05067fd
Compare
05067fd to
ff15ccb
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@apps/gateway/src/api.spec.ts`:
- Line 8670: Replace the any type for capturedBody in
apps/gateway/src/api.spec.ts lines 8670-8670 and 8845-8845 with Record<string,
unknown> | undefined or a focused request-body interface, preserving the
existing known-field assertions in both tests.
In `@packages/actions/src/prepare-request-body.spec.ts`:
- Around line 1368-1374: Rename the parameterized test description in
prepare-request-body.spec.ts to state that tool_choice is omitted when no choice
is requested, matching the existing toBeUndefined assertion; do not change the
test implementation.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 738160e7-f56d-4475-ab53-4e6297d55328
📒 Files selected for processing (5)
apps/gateway/src/api.spec.tspackages/actions/src/prepare-request-body.spec.tspackages/actions/src/prepare-request-body.tspackages/models/src/models.tspackages/models/src/models/inclusionai.ts
Adds the Ling-3.0-flash mapping (DeepInfra + NovitaAI) with chatTemplateThinkingKey reasoning handling plus unit tests covering thinking translation, tool calling, and JSON output on both providers. Co-Authored-By: Claude <noreply@anthropic.com>
ff15ccb to
f1a7bdb
Compare
There was a problem hiding this comment.
🧹 Nitpick comments (4)
apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts (1)
340-346: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winAdd a pinned-provider case for the accepting mapping.
This test passes
undefinedas the provider, soprovidersToCheckcontains every Ling mapping. The assertion therefore does not prove that the DeepInfra mapping alone satisfies the newchatTemplateThinkingKeybranch. Add a case that pins"deepinfra"to isolate the new condition, mirroring the pinned"novita"rejection case.💚 Proposed additional test
it("allows reasoning.max_tokens on chatTemplateThinkingKey mappings (budget is dropped to a binary toggle)", () => { expect(() => validateModelCapabilities(lingModel, lingModel.id, undefined, { reasoning_max_tokens: 2048, }), ).not.toThrow(); }); + + it("allows reasoning.max_tokens when deepinfra is pinned", () => { + expect(() => + validateModelCapabilities(lingModel, lingModel.id, "deepinfra", { + reasoning_max_tokens: 2048, + }), + ).not.toThrow(); + });🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts` around lines 340 - 346, Add a companion test near the existing reasoning_max_tokens acceptance test that passes "deepinfra" as the provider to validate the accepting chatTemplateThinkingKey mapping in isolation. Keep the same lingModel, model ID, and reasoning_max_tokens input, and mirror the pinned-provider structure used by the novita rejection case.apps/gateway/src/api.spec.ts (2)
8948-9048: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winReuse the helper for the native OpenAI-path test.
This test re-implements the whole setup that
exerciseLingMessagesalready performs: the API key insert, the provider key loop, the fetch spy, the canned completion, and thefinallyrestore. Only the endpoint, the request body, and one extra assertion differ. Parameterize the endpoint and body inexerciseLingMessages, then call it here. That removes about 80 duplicated lines and keeps both lanes in sync when the mock shape changes.As per coding guidelines, "Apply DRY principles for reusable code".
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/api.spec.ts` around lines 8948 - 9048, Refactor the native OpenAI-path test to reuse exerciseLingMessages instead of duplicating API-key setup, provider-key creation, fetch mocking, canned completion, and cleanup. Parameterize exerciseLingMessages with the endpoint and request body needed by both message and chat-completions tests, then invoke it here while preserving the DeepInfra, model, enable_thinking, and absent reasoning_effort assertions.Source: Coding guidelines
8711-8713: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueReturn the routed provider's upstream id in the canned response.
The condition keys the response
modelonoptions.modelinstead of the routed provider. When a bare-id test runs against Novita, the mock reports"inclusionAI/Ling-3.0-flash"while the request went to Novita. No current assertion reads the responsemodel, so the tests still pass, but the fixture contradicts the routed provider.upstreamModels[provider]is correct for every call site.♻️ Proposed simplification
- model: options.model - ? "inclusionAI/Ling-3.0-flash" - : upstreamModels[provider], + model: upstreamModels[provider],🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/api.spec.ts` around lines 8711 - 8713, Update the canned response model assignment in the relevant api test fixture to always return upstreamModels[provider], removing the options.model conditional so the response reflects the routed provider’s upstream ID for every call site.apps/gateway/src/chat/tools/validate-model-capabilities.ts (1)
263-272: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueRename the flag to match the other capability checks.
The variable now means "the mapping supports a budget or a chat-template thinking toggle". Every other check in this function uses a
supportsXname. Rename it for consistency and to avoid confusion with thereasoningMaxTokensmapping field read inside the callback.♻️ Proposed rename
- const reasoningMaxTokens = providersToCheck.some( + const supportsReasoningMaxTokens = providersToCheck.some( (provider) => (provider as ProviderModelMapping).reasoningMaxTokens === true || (provider as ProviderModelMapping).chatTemplateThinkingKey !== undefined, );- if (!reasoningMaxTokens) { + if (!supportsReasoningMaxTokens) {🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/gateway/src/chat/tools/validate-model-capabilities.ts` around lines 263 - 272, Rename the local variable reasoningMaxTokens to a supportsReasoningMaxTokens-style name that reflects the combined budget or chat-template toggle capability, and update all references to it within the surrounding validation logic. Leave the provider mapping field access unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@apps/gateway/src/api.spec.ts`:
- Around line 8948-9048: Refactor the native OpenAI-path test to reuse
exerciseLingMessages instead of duplicating API-key setup, provider-key
creation, fetch mocking, canned completion, and cleanup. Parameterize
exerciseLingMessages with the endpoint and request body needed by both message
and chat-completions tests, then invoke it here while preserving the DeepInfra,
model, enable_thinking, and absent reasoning_effort assertions.
- Around line 8711-8713: Update the canned response model assignment in the
relevant api test fixture to always return upstreamModels[provider], removing
the options.model conditional so the response reflects the routed provider’s
upstream ID for every call site.
In `@apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts`:
- Around line 340-346: Add a companion test near the existing
reasoning_max_tokens acceptance test that passes "deepinfra" as the provider to
validate the accepting chatTemplateThinkingKey mapping in isolation. Keep the
same lingModel, model ID, and reasoning_max_tokens input, and mirror the
pinned-provider structure used by the novita rejection case.
In `@apps/gateway/src/chat/tools/validate-model-capabilities.ts`:
- Around line 263-272: Rename the local variable reasoningMaxTokens to a
supportsReasoningMaxTokens-style name that reflects the combined budget or
chat-template toggle capability, and update all references to it within the
surrounding validation logic. Leave the provider mapping field access unchanged.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: e867fad8-f608-45af-81b1-6b31c3e0ee7e
📒 Files selected for processing (7)
apps/gateway/src/api.spec.tsapps/gateway/src/chat/tools/validate-model-capabilities.spec.tsapps/gateway/src/chat/tools/validate-model-capabilities.tspackages/actions/src/prepare-request-body.spec.tspackages/actions/src/prepare-request-body.tspackages/models/src/models.tspackages/models/src/models/inclusionai.ts
🚧 Files skipped from review as they are similar to previous changes (4)
- packages/actions/src/prepare-request-body.ts
- packages/models/src/models/inclusionai.ts
- packages/actions/src/prepare-request-body.spec.ts
- packages/models/src/models.ts
DeepInfra serves this model at $0.06/$0.18 per 1M tokens (cached $0.012),
not $0.045/$0.10/$0.008. Confirmed against DeepInfra's model listing and
the usage.estimated_cost returned on live completions, which matches the
corrected rates exactly. The old figures would have under-billed output by
44%.
Novita also rejects response_format for this model outright ("does not
support feature: structured-outputs") for both json_object and json_schema,
so mark the mapping jsonOutput: false and drop response_format from its
supportedParameters. This was failing two scoped e2e cases; with the flag
corrected the gateway rejects such requests up front and routes JSON traffic
for the bare model id to DeepInfra instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Summary
Adds InclusionAI Ling-3.0-flash, a 124B-parameter native hybrid-reasoning MoE (5.1B active), mapped on DeepInfra and NovitaAI.
inclusionAI/Ling-3.0-flashinclusionai/ling-3.0-flashreasoning_effort: "none"References: https://deepinfra.com/inclusionAI/Ling-3.0-flash · https://novita.ai/models/model-detail/inclusionai-ling-3.0-flash
What changed
packages/models: newinclusionai.tswith theling-3.0-flashdefinition (DeepInfra + NovitaAI mappings);models.tsregisters it and the newchatTemplateThinkingKeymapping field.packages/actions:prepare-request-body.tstranslatesreasoning_effortthroughchatTemplateThinkingKeyintochat_template_kwargs: { enable_thinking: ... }./v1/messagespath and bareling-3.0-flashrouting.Reasoning handling
Provider reasoning controls are model-specific — DeepSeek-style
reasoning_efforttiers, chat-template flags like Ling'senable_thinking,requiresEnableThinking-style booleans.chatTemplateThinkingKeyonProviderModelMappingdeclares the chat-template flag declaratively, so the translation lives in one place inprepare-request-bodyinstead of every provider branch hand-rolling quirks.Ling thinks by default; its only control is the vLLM chat-template flag. The gateway maps (effective on DeepInfra; see below for Novita):
reasoning_effort: "none"→chat_template_kwargs: { enable_thinking: false }chat_template_kwargs: { enable_thinking: true }Novita caveat (verified live 2026-08-10): Novita's Ling backend accepts
chat_template_kwargs.enable_thinking(HTTP 200) but ignores it —reasoning_contentis returned forfalse/true/unset alike, so thinking cannot be disabled there.chatTemplateThinkingKeyis therefore declared on the DeepInfra mapping only, andreasoning_effort: "none"is only advertised where it works.reasoning_effortis not in Ling'ssupportedParameters, so it never reaches the provider raw (DeepSeek V3.2-on-Novita precedent). Anthropic clients get the same behavior — the gateway's existing/v1/messagesthinking bridge (thinking-to-reasoning) already maps Anthropic thinking controls to the unified reasoning field; this PR'schatTemplateThinkingKeytranslation then applies downstream of it.Client → gateway:
{ "model": "deepinfra/ling-3.0-flash", "reasoning_effort": "none", "messages": [{ "role": "user", "content": "Hello!" }] }Gateway → provider:
{ "model": "inclusionAI/Ling-3.0-flash", "chat_template_kwargs": { "enable_thinking": false }, "messages": [{ "role": "user", "content": "Hello!" }] }Verification
pnpm build(17/17),pnpm format,tsc --noEmit(gateway, actions, models) — all clean.packages/actions+packages/models, 750 total): 16 new Ling unit cases (8 test definitions × 2 providers) — thinking translation, tool calling,json_object/json_schemapassthrough, and their composition.api.spec.ts, 163 passed): 10 new Ling integration cases — 8 on the Anthropic/v1/messagespath (thinking disabled/adaptive →enable_thinking, tools intact) and 2 for bareling-3.0-flashresolution on both API lanes.Live smoke (needs
LLM_DEEPINFRA_API_KEY+LLM_NOVITA_AI_API_KEY; dev stack on :4001; pin withx-no-fallback: true; vary prompts between retries — the gateway caches responses in Redis):Expect: thinking deltas in 1, none in 2, a
tool_usecompletion in 3, valid JSON in 4. In 5, the bareling-3.0-flashid resolves through standard routing to the cheapest available provider (deepinfra in the test harness; production routing uses weighted scoring/pinning, so the provider may differ) withenable_thinking: falseapplied. 6 was verified live on 2026-08-10: Novita acceptschat_template_kwargs.enable_thinking(HTTP 200) but ignores it —reasoning_contentis returned forfalse/true/unset alike. Thinking cannot be disabled on Novita;chatTemplateThinkingKeyis declared on DeepInfra only. Scoped e2e:TEST_MODELS="deepinfra/ling-3.0-flash,novita/ling-3.0-flash" FULL_MODE=true pnpm test:e2e.Review corrections (2026-08-12)
Two catalogue errors were found while running the scoped e2e suite against both providers and verifying every declared value against the live APIs.
1. DeepInfra pricing was wrong — a billing defect. The mapping declared $0.045 in / $0.10 out / $0.008 cached per 1M. DeepInfra actually charges $0.06 / $0.18 / $0.012. Their model listing reports
cents_per_input_token: 6e-06,cents_per_output_token: 1.8e-05andrate_per_input_token_cached: 0.2, and theusage.estimated_costreturned on live completions matches the corrected rates to the last digit while matching the declared ones at no size:estimated_cost9.54e-060.00548412The declared figures would have under-billed output by 44%. Note these are exactly Novita's rates — Novita's own numbers were correct as declared.
2. Novita does not support JSON output. The mapping declared
jsonOutput: true, but Novita rejectsresponse_formatfor this model outright, for both forms:Novita's model metadata agrees — its
featureslist is["serverless","function-calling","reasoning"]with no structured-outputs entry. This was failing two scoped e2e cases (JSON outputandJSON output streaming). The mapping is nowjsonOutput: falsewithresponse_formatdropped fromsupportedParameters, so the gateway returns a clean 400 up front for a pinned request and routes JSON traffic for the bareling-3.0-flashid to DeepInfra.Verified as correct, left unchanged: Novita's prices ($0.06/$0.18/$0.012 — its listing reports
600/1800/120per M, the same 1/10,000 USD unit that the existing deepseek-v3.2 mapping uses), both context sizes, and the reasoning design. DeepInfra's 262,144 context was double-checked empirically because DeepInfra's listing understates it as131072— the deployment served a 257,844-token prompt fine and only rejected at 274,085 with"longer than the model's context length (262144 tokens)". The PR's asymmetricchatTemplateThinkingKey-on-DeepInfra-only design was also reconfirmed live: on DeepInfraenable_thinking: falseyieldsreasoning_tokens: 0(vs 35 unset), while Novita still returnsreasoning_contentwith the flag set tofalse.Scoped e2e after the fixes:
TEST_MODELS="deepinfra/ling-3.0-flash,novita/ling-3.0-flash" FULL_MODE=true pnpm test:e2e→ 29 files passed / 2 skipped, 115 tests passed, 0 failed.pnpm buildandpnpm formatclean; fullpnpm test:unitgreen (4,453 passed).Summary by CodeRabbit
New Features
Bug Fixes