Skip to content

feat(models): add InclusionAI Ling-3.0-flash - #3515

Merged
steebchen merged 2 commits into
theopenco:mainfrom
vicovaro:feat/ling-3.0-flash
Aug 12, 2026
Merged

steebchen merged 2 commits into
theopenco:mainfrom
vicovaro:feat/ling-3.0-flash

Conversation

@vicovaro

@vicovaro vicovaro commented Aug 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds InclusionAI Ling-3.0-flash, a 124B-parameter native hybrid-reasoning MoE (5.1B active), mapped on DeepInfra and NovitaAI.

DeepInfra NovitaAI
External model id inclusionAI/Ling-3.0-flash inclusionai/ling-3.0-flash
Input $0.06 / 1M tokens $0.06 / 1M tokens
Output $0.18 / 1M tokens $0.18 / 1M tokens
Cached input $0.012 / 1M tokens $0.012 / 1M tokens
Context / Max output 262,144 / 32,768 262,144 / 32,768
Reasoning On by default; disable via reasoning_effort: "none" Always on — backend ignores thinking controls
Tools / JSON / streaming ✓ / ✓ / ✓ ✓ / ✗ / ✓

References: https://deepinfra.com/inclusionAI/Ling-3.0-flash · https://novita.ai/models/model-detail/inclusionai-ling-3.0-flash

What changed

  • packages/models: new inclusionai.ts with the ling-3.0-flash definition (DeepInfra + NovitaAI mappings); models.ts registers it and the new chatTemplateThinkingKey mapping field.
  • packages/actions: prepare-request-body.ts translates reasoning_effort through chatTemplateThinkingKey into chat_template_kwargs: { enable_thinking: ... }.
  • Tests: OpenAI-path unit tests (thinking, tools, JSON output) + gateway integration tests covering the Anthropic /v1/messages path and bare ling-3.0-flash routing.

Reasoning handling

Provider reasoning controls are model-specific — DeepSeek-style reasoning_effort tiers, chat-template flags like Ling's enable_thinking, requiresEnableThinking-style booleans. chatTemplateThinkingKey on ProviderModelMapping declares the chat-template flag declaratively, so the translation lives in one place in prepare-request-body instead of every provider branch hand-rolling quirks.

Ling thinks by default; its only control is the vLLM chat-template flag. The gateway maps (effective on DeepInfra; see below for Novita):

  • reasoning_effort: "none" → chat_template_kwargs: { enable_thinking: false }
  • any other effort → chat_template_kwargs: { enable_thinking: true }
  • absent → nothing sent (provider default: thinking on)

Novita caveat (verified live 2026-08-10): Novita's Ling backend accepts chat_template_kwargs.enable_thinking (HTTP 200) but ignores it — reasoning_content is returned for false/true/unset alike, so thinking cannot be disabled there. chatTemplateThinkingKey is therefore declared on the DeepInfra mapping only, and reasoning_effort: "none" is only advertised where it works.

reasoning_effort is not in Ling's supportedParameters, so it never reaches the provider raw (DeepSeek V3.2-on-Novita precedent). Anthropic clients get the same behavior — the gateway's existing /v1/messages thinking bridge (thinking-to-reasoning) already maps Anthropic thinking controls to the unified reasoning field; this PR's chatTemplateThinkingKey translation then applies downstream of it.

Client → gateway:

{ "model": "deepinfra/ling-3.0-flash", "reasoning_effort": "none", "messages": [{ "role": "user", "content": "Hello!" }] }

Gateway → provider:

{ "model": "inclusionAI/Ling-3.0-flash", "chat_template_kwargs": { "enable_thinking": false }, "messages": [{ "role": "user", "content": "Hello!" }] }

Verification

  • Build/lint: pnpm build (17/17), pnpm format, tsc --noEmit (gateway, actions, models) — all clean.
  • Unit tests (packages/actions + packages/models, 750 total): 16 new Ling unit cases (8 test definitions × 2 providers) — thinking translation, tool calling, json_object / json_schema passthrough, and their composition.
  • Gateway integration (api.spec.ts, 163 passed): 10 new Ling integration cases — 8 on the Anthropic /v1/messages path (thinking disabled/adaptive → enable_thinking, tools intact) and 2 for bare ling-3.0-flash resolution on both API lanes.

Live smoke (needs LLM_DEEPINFRA_API_KEY + LLM_NOVITA_AI_API_KEY; dev stack on :4001; pin with x-no-fallback: true; vary prompts between retries — the gateway caches responses in Redis):

# 1. default thinking
curl -N http://localhost:4001/v1/chat/completions -H "Authorization: Bearer ***" -H "x-no-fallback: true" \
  -d '{"model":"deepinfra/ling-3.0-flash","stream":true,"messages":[{"role":"user","content":"hi"}]}'

# 2. thinking disabled
curl -N http://localhost:4001/v1/chat/completions -H "Authorization: Bearer ***" -H "x-no-fallback: true" \
  -d '{"model":"deepinfra/ling-3.0-flash","stream":true,"reasoning_effort":"none","messages":[{"role":"user","content":"hi"}]}'

# 3. function calling
curl -N http://localhost:4001/v1/chat/completions -H "Authorization: Bearer ***" -H "x-no-fallback: true" \
  -d '{"model":"deepinfra/ling-3.0-flash","stream":true,"messages":[{"role":"user","content":"What is the weather in Paris?"}],"tools":[{"type":"function","function":{"name":"get_weather","description":"Get the weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}]}'

# 4. JSON output
curl -N http://localhost:4001/v1/chat/completions -H "Authorization: Bearer ***" -H "x-no-fallback: true" \
  -d '{"model":"deepinfra/ling-3.0-flash","stream":true,"messages":[{"role":"user","content":"Return JSON: {\"city\":\"Paris\"}"}],"response_format":{"type":"json_object"}}'

# 5. non-prefixed call (standard routing — no provider prefix)
curl -N http://localhost:4001/v1/chat/completions -H "Authorization: Bearer ***" -H "x-no-fallback: true" \
  -d '{"model":"ling-3.0-flash","stream":true,"reasoning_effort":"none","messages":[{"role":"user","content":"hi"}]}'

# 6. NovitaAI — VERIFIED 2026-08-10: backend ignores thinking controls (reproducer)
curl -N http://localhost:4001/v1/chat/completions -H "Authorization: Bearer ***" -H "x-no-fallback: true" \
  -d '{"model":"novita/ling-3.0-flash","stream":true,"reasoning_effort":"none","messages":[{"role":"user","content":"hi"}]}'

Expect: thinking deltas in 1, none in 2, a tool_use completion in 3, valid JSON in 4. In 5, the bare ling-3.0-flash id resolves through standard routing to the cheapest available provider (deepinfra in the test harness; production routing uses weighted scoring/pinning, so the provider may differ) with enable_thinking: false applied. 6 was verified live on 2026-08-10: Novita accepts chat_template_kwargs.enable_thinking (HTTP 200) but ignores it — reasoning_content is returned for false/true/unset alike. Thinking cannot be disabled on Novita; chatTemplateThinkingKey is declared on DeepInfra only. Scoped e2e: TEST_MODELS="deepinfra/ling-3.0-flash,novita/ling-3.0-flash" FULL_MODE=true pnpm test:e2e.

Review corrections (2026-08-12)

Two catalogue errors were found while running the scoped e2e suite against both providers and verifying every declared value against the live APIs.

1. DeepInfra pricing was wrong — a billing defect. The mapping declared $0.045 in / $0.10 out / $0.008 cached per 1M. DeepInfra actually charges $0.06 / $0.18 / $0.012. Their model listing reports cents_per_input_token: 6e-06, cents_per_output_token: 1.8e-05 and rate_per_input_token_cached: 0.2, and the usage.estimated_cost returned on live completions matches the corrected rates to the last digit while matching the declared ones at no size:

prompt / completion tokens returned estimated_cost at $0.06/$0.18 at declared $0.045/$0.10
21 / 46 9.54e-06 9.54e-06 ✓ 5.55e-06 ✗
91,378 / 8 0.00548412 0.00548412 ✓ 0.00411281 ✗

The declared figures would have under-billed output by 44%. Note these are exactly Novita's rates — Novita's own numbers were correct as declared.

2. Novita does not support JSON output. The mapping declared jsonOutput: true, but Novita rejects response_format for this model outright, for both forms:

$ curl .../chat/completions -d '{"model":"inclusionai/ling-3.0-flash", ..., "response_format":{"type":"json_object"}}'
{"code":400,"reason":"INVALID_REQUEST_BODY","message":"model: inclusionai/ling-3.0-flash does not support feature: structured-outputs"}

$ ... "response_format":{"type":"json_schema", ...}
{"code":400,"reason":"INVALID_REQUEST_BODY","message":"model features structured outputs not support"}

Novita's model metadata agrees — its features list is ["serverless","function-calling","reasoning"] with no structured-outputs entry. This was failing two scoped e2e cases (JSON output and JSON output streaming). The mapping is now jsonOutput: false with response_format dropped from supportedParameters, so the gateway returns a clean 400 up front for a pinned request and routes JSON traffic for the bare ling-3.0-flash id to DeepInfra.

Verified as correct, left unchanged: Novita's prices ($0.06/$0.18/$0.012 — its listing reports 600/1800/120 per M, the same 1/10,000 USD unit that the existing deepseek-v3.2 mapping uses), both context sizes, and the reasoning design. DeepInfra's 262,144 context was double-checked empirically because DeepInfra's listing understates it as 131072 — the deployment served a 257,844-token prompt fine and only rejected at 274,085 with "longer than the model's context length (262144 tokens)". The PR's asymmetric chatTemplateThinkingKey-on-DeepInfra-only design was also reconfirmed live: on DeepInfra enable_thinking: false yields reasoning_tokens: 0 (vs 35 unset), while Novita still returns reasoning_content with the flag set to false.

Scoped e2e after the fixes: TEST_MODELS="deepinfra/ling-3.0-flash,novita/ling-3.0-flash" FULL_MODE=true pnpm test:e2e → 29 files passed / 2 skipped, 115 tests passed, 0 failed. pnpm build and pnpm format clean; full pnpm test:unit green (4,453 passed).

Summary by CodeRabbit

  • New Features

    • Added support for the InclusionAI Ling 3.0 Flash model through DeepInfra and Novita.
    • Added configurable reasoning controls, including enabled, disabled, and provider-default behavior.
    • Added support for tool use and JSON object or schema responses.
    • Added provider-aware model routing and capability handling.
  • Bug Fixes

    • Improved translation of reasoning settings into provider requests.
    • Prevented unsupported reasoning parameters from being forwarded.
    • Improved validation for reasoning budget requests based on provider capabilities.

@coderabbitai

coderabbitai Bot commented Aug 9, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4d47c484-dc26-4913-a085-a194c9f5bc5e

📥 Commits

Reviewing files that changed from the base of the PR and between f1a7bdb and ab676e5.

📒 Files selected for processing (2)
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/models/src/models/inclusionai.ts
🚧 Files skipped from review as they are similar to previous changes (2)
  • packages/models/src/models/inclusionai.ts
  • packages/actions/src/prepare-request-body.spec.ts

Walkthrough

Adds InclusionAI Ling-3.0-flash model definitions for DeepInfra and Novita. Maps reasoning controls to chat_template_kwargs.enable_thinking, preserves supported request features, and adds request-body, capability, and gateway integration coverage.

Changes

Ling-3.0-flash reasoning flow

Layer / File(s) Summary
Model mappings and catalog
packages/models/src/models.ts, packages/models/src/models/inclusionai.ts
Adds chatTemplateThinkingKey to provider mappings and registers Ling-3.0-flash configurations for DeepInfra and Novita.
Request-body reasoning translation
packages/actions/src/prepare-request-body.ts, packages/actions/src/prepare-request-body.spec.ts
Maps "none" to enable_thinking: false and other reasoning requests to true. Tests cover provider defaults, tools, tool choices, and JSON response formats.
Reasoning capability validation
apps/gateway/src/chat/tools/validate-model-capabilities.ts, apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts
Accepts chat-template thinking mappings during reasoning budget validation and tests unsupported provider configurations.
Gateway routing and upstream requests
apps/gateway/src/api.spec.ts
Tests Anthropic and OpenAI request handling, bare-model routing, tool translation, and omission of upstream reasoning_effort.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant Gateway
  participant prepareRequestBody
  participant DeepInfraOrNovita
  Client->>Gateway: Ling-3.0-flash request
  Gateway->>prepareRequestBody: reasoning and tool parameters
  prepareRequestBody->>DeepInfraOrNovita: chat_template_kwargs and translated tools
  DeepInfraOrNovita-->>Gateway: upstream completion
  Gateway-->>Client: gateway response
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely identifies the main change: adding InclusionAI Ling-3.0-flash support.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@vicovaro
vicovaro force-pushed the feat/ling-3.0-flash branch 5 times, most recently from 6247d03 to 05067fd Compare August 10, 2026 17:51
@vicovaro
vicovaro marked this pull request as ready for review August 10, 2026 18:47
@vicovaro
vicovaro force-pushed the feat/ling-3.0-flash branch from 05067fd to ff15ccb Compare August 10, 2026 18:48

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/gateway/src/api.spec.ts`:
- Line 8670: Replace the any type for capturedBody in
apps/gateway/src/api.spec.ts lines 8670-8670 and 8845-8845 with Record<string,
unknown> | undefined or a focused request-body interface, preserving the
existing known-field assertions in both tests.

In `@packages/actions/src/prepare-request-body.spec.ts`:
- Around line 1368-1374: Rename the parameterized test description in
prepare-request-body.spec.ts to state that tool_choice is omitted when no choice
is requested, matching the existing toBeUndefined assertion; do not change the
test implementation.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 738160e7-f56d-4475-ab53-4e6297d55328

📥 Commits

Reviewing files that changed from the base of the PR and between cd835f8 and 05067fd.

📒 Files selected for processing (5)
  • apps/gateway/src/api.spec.ts
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/actions/src/prepare-request-body.ts
  • packages/models/src/models.ts
  • packages/models/src/models/inclusionai.ts

Comment thread apps/gateway/src/api.spec.ts Outdated
Comment thread packages/actions/src/prepare-request-body.spec.ts
Adds the Ling-3.0-flash mapping (DeepInfra + NovitaAI) with chatTemplateThinkingKey
reasoning handling plus unit tests covering thinking translation, tool calling,
and JSON output on both providers.

Co-Authored-By: Claude <noreply@anthropic.com>
@vicovaro
vicovaro force-pushed the feat/ling-3.0-flash branch from ff15ccb to f1a7bdb Compare August 10, 2026 19:34

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (4)
apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts (1)

340-346: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add a pinned-provider case for the accepting mapping.

This test passes undefined as the provider, so providersToCheck contains every Ling mapping. The assertion therefore does not prove that the DeepInfra mapping alone satisfies the new chatTemplateThinkingKey branch. Add a case that pins "deepinfra" to isolate the new condition, mirroring the pinned "novita" rejection case.

💚 Proposed additional test
 	it("allows reasoning.max_tokens on chatTemplateThinkingKey mappings (budget is dropped to a binary toggle)", () => {
 		expect(() =>
 			validateModelCapabilities(lingModel, lingModel.id, undefined, {
 				reasoning_max_tokens: 2048,
 			}),
 		).not.toThrow();
 	});
+
+	it("allows reasoning.max_tokens when deepinfra is pinned", () => {
+		expect(() =>
+			validateModelCapabilities(lingModel, lingModel.id, "deepinfra", {
+				reasoning_max_tokens: 2048,
+			}),
+		).not.toThrow();
+	});
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts` around lines
340 - 346, Add a companion test near the existing reasoning_max_tokens
acceptance test that passes "deepinfra" as the provider to validate the
accepting chatTemplateThinkingKey mapping in isolation. Keep the same lingModel,
model ID, and reasoning_max_tokens input, and mirror the pinned-provider
structure used by the novita rejection case.
apps/gateway/src/api.spec.ts (2)

8948-9048: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Reuse the helper for the native OpenAI-path test.

This test re-implements the whole setup that exerciseLingMessages already performs: the API key insert, the provider key loop, the fetch spy, the canned completion, and the finally restore. Only the endpoint, the request body, and one extra assertion differ. Parameterize the endpoint and body in exerciseLingMessages, then call it here. That removes about 80 duplicated lines and keeps both lanes in sync when the mock shape changes.

As per coding guidelines, "Apply DRY principles for reusable code".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/api.spec.ts` around lines 8948 - 9048, Refactor the native
OpenAI-path test to reuse exerciseLingMessages instead of duplicating API-key
setup, provider-key creation, fetch mocking, canned completion, and cleanup.
Parameterize exerciseLingMessages with the endpoint and request body needed by
both message and chat-completions tests, then invoke it here while preserving
the DeepInfra, model, enable_thinking, and absent reasoning_effort assertions.

Source: Coding guidelines


8711-8713: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Return the routed provider's upstream id in the canned response.

The condition keys the response model on options.model instead of the routed provider. When a bare-id test runs against Novita, the mock reports "inclusionAI/Ling-3.0-flash" while the request went to Novita. No current assertion reads the response model, so the tests still pass, but the fixture contradicts the routed provider. upstreamModels[provider] is correct for every call site.

♻️ Proposed simplification
-								model: options.model
-									? "inclusionAI/Ling-3.0-flash"
-									: upstreamModels[provider],
+								model: upstreamModels[provider],
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/api.spec.ts` around lines 8711 - 8713, Update the canned
response model assignment in the relevant api test fixture to always return
upstreamModels[provider], removing the options.model conditional so the response
reflects the routed provider’s upstream ID for every call site.
apps/gateway/src/chat/tools/validate-model-capabilities.ts (1)

263-272: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Rename the flag to match the other capability checks.

The variable now means "the mapping supports a budget or a chat-template thinking toggle". Every other check in this function uses a supportsX name. Rename it for consistency and to avoid confusion with the reasoningMaxTokens mapping field read inside the callback.

♻️ Proposed rename
-		const reasoningMaxTokens = providersToCheck.some(
+		const supportsReasoningMaxTokens = providersToCheck.some(
 			(provider) =>
 				(provider as ProviderModelMapping).reasoningMaxTokens === true ||
 				(provider as ProviderModelMapping).chatTemplateThinkingKey !==
 					undefined,
 		);
-		if (!reasoningMaxTokens) {
+		if (!supportsReasoningMaxTokens) {
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/chat/tools/validate-model-capabilities.ts` around lines 263
- 272, Rename the local variable reasoningMaxTokens to a
supportsReasoningMaxTokens-style name that reflects the combined budget or
chat-template toggle capability, and update all references to it within the
surrounding validation logic. Leave the provider mapping field access unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@apps/gateway/src/api.spec.ts`:
- Around line 8948-9048: Refactor the native OpenAI-path test to reuse
exerciseLingMessages instead of duplicating API-key setup, provider-key
creation, fetch mocking, canned completion, and cleanup. Parameterize
exerciseLingMessages with the endpoint and request body needed by both message
and chat-completions tests, then invoke it here while preserving the DeepInfra,
model, enable_thinking, and absent reasoning_effort assertions.
- Around line 8711-8713: Update the canned response model assignment in the
relevant api test fixture to always return upstreamModels[provider], removing
the options.model conditional so the response reflects the routed provider’s
upstream ID for every call site.

In `@apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts`:
- Around line 340-346: Add a companion test near the existing
reasoning_max_tokens acceptance test that passes "deepinfra" as the provider to
validate the accepting chatTemplateThinkingKey mapping in isolation. Keep the
same lingModel, model ID, and reasoning_max_tokens input, and mirror the
pinned-provider structure used by the novita rejection case.

In `@apps/gateway/src/chat/tools/validate-model-capabilities.ts`:
- Around line 263-272: Rename the local variable reasoningMaxTokens to a
supportsReasoningMaxTokens-style name that reflects the combined budget or
chat-template toggle capability, and update all references to it within the
surrounding validation logic. Leave the provider mapping field access unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: e867fad8-f608-45af-81b1-6b31c3e0ee7e

📥 Commits

Reviewing files that changed from the base of the PR and between 05067fd and f1a7bdb.

📒 Files selected for processing (7)
  • apps/gateway/src/api.spec.ts
  • apps/gateway/src/chat/tools/validate-model-capabilities.spec.ts
  • apps/gateway/src/chat/tools/validate-model-capabilities.ts
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/actions/src/prepare-request-body.ts
  • packages/models/src/models.ts
  • packages/models/src/models/inclusionai.ts
🚧 Files skipped from review as they are similar to previous changes (4)
  • packages/actions/src/prepare-request-body.ts
  • packages/models/src/models/inclusionai.ts
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/models/src/models.ts

DeepInfra serves this model at $0.06/$0.18 per 1M tokens (cached $0.012),
not $0.045/$0.10/$0.008. Confirmed against DeepInfra's model listing and
the usage.estimated_cost returned on live completions, which matches the
corrected rates exactly. The old figures would have under-billed output by
44%.

Novita also rejects response_format for this model outright ("does not
support feature: structured-outputs") for both json_object and json_schema,
so mark the mapping jsonOutput: false and drop response_format from its
supportedParameters. This was failing two scoped e2e cases; with the flag
corrected the gateway rejects such requests up front and routes JSON traffic
for the bare model id to DeepInfra instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@steebchen
steebchen merged commit 4ba75b3 into theopenco:main Aug 12, 2026
10 checks passed
@vicovaro
vicovaro deleted the feat/ling-3.0-flash branch August 12, 2026 15:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants