feat: add DeepInfra as inference provider - #2301
Conversation
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds DeepInfra (deepinfra.com) as a new provider with OpenAI-compatible API. Includes provider mappings for 15 existing models: - DeepSeek V4-Pro, V4-Flash, V3.2 - MiMo V2.5, V2.5-Pro - Kimi K2.6, K2.5 - GLM 5.1, 5, 4.7-Flash - MiniMax M2.5 - Qwen 3.5-397B, 3.6-35B, 3-Max Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (5)
🚧 Files skipped from review as they are similar to previous changes (4)
WalkthroughAdds DeepInfra provider support: provider metadata, endpoint/base-url and header routing, env/workflow wiring, and model entries mapping DeepInfra-hosted variants for DeepSeek, Moonshot, and Zai models. ChangesDeepInfra provider integration
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/actions/src/get-provider-endpoint.ts`:
- Around line 288-290: The deepinfra branch currently hardcodes url =
"https://api.deepinfra.com/v1/openai" and ignores the optional
LLM_DEEPINFRA_BASE_URL; update the case "deepinfra" in get-provider-endpoint.ts
to use envValueOrDefault("LLM_DEEPINFRA_BASE_URL",
"https://api.deepinfra.com/v1/openai") (same pattern used for xiaomi and
google-ai-studio) so the url variable respects the environment override while
falling back to the default.
In `@packages/models/src/models/xiaomi.ts`:
- Around line 140-154: Update the DeepInfra model entry for providerId
"deepinfra" and modelName "XiaomiMiMo/MiMo-V2.5": change the contextSize value
from 256000 to 1000000 so it matches the official MiMo-V2.5 1M token context
window; verify the change in the object where contextSize is currently set and
keep all other fields unchanged.
In `@packages/models/src/models/zai.ts`:
- Around line 60-74: Update the DeepInfra GLM model entries to correct their
context and output limits: for the object with modelName "zai-org/GLM-5.1" set
contextSize to 200000 and maxOutput to 128000; for the object with modelName
"zai-org/GLM-5" set contextSize to 202752 and maxOutput to 131072; and for the
object with modelName "zai-org/GLM-4.7-Flash" set contextSize to 203000 and
maxOutput to 128000. Ensure you update the corresponding objects in the models
array (matching providerId "deepinfra" and the modelName strings) so the values
replace the previous 198000 and 65536 entries.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 2c22ddec-43eb-4d02-bc8f-2fb5729efb40
📒 Files selected for processing (9)
packages/actions/src/get-provider-endpoint.tspackages/actions/src/get-provider-headers.tspackages/models/src/models/alibaba.tspackages/models/src/models/deepseek.tspackages/models/src/models/minimax.tspackages/models/src/models/moonshot.tspackages/models/src/models/xiaomi.tspackages/models/src/models/zai.tspackages/models/src/providers.ts
| case "deepinfra": | ||
| url = "https://api.deepinfra.com/v1/openai"; | ||
| break; |
There was a problem hiding this comment.
Optional baseUrl not respected.
The deepinfra provider definition declares LLM_DEEPINFRA_BASE_URL as optional, but this code ignores it and always uses the hardcoded default. Other providers with optional baseUrl (e.g., xiaomi at lines 231–237, google-ai-studio at lines 139–145) use envValueOrDefault to respect the environment variable when set.
🔧 Proposed fix to respect optional baseUrl
case "deepinfra":
- url = "https://api.deepinfra.com/v1/openai";
+ url =
+ envValueOrDefault(
+ "deepinfra",
+ "baseUrl",
+ "https://api.deepinfra.com/v1/openai",
+ ) ?? "https://api.deepinfra.com/v1/openai";
break;📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| case "deepinfra": | |
| url = "https://api.deepinfra.com/v1/openai"; | |
| break; | |
| case "deepinfra": | |
| url = | |
| envValueOrDefault( | |
| "deepinfra", | |
| "baseUrl", | |
| "https://api.deepinfra.com/v1/openai", | |
| ) ?? "https://api.deepinfra.com/v1/openai"; | |
| break; |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/actions/src/get-provider-endpoint.ts` around lines 288 - 290, The
deepinfra branch currently hardcodes url = "https://api.deepinfra.com/v1/openai"
and ignores the optional LLM_DEEPINFRA_BASE_URL; update the case "deepinfra" in
get-provider-endpoint.ts to use envValueOrDefault("LLM_DEEPINFRA_BASE_URL",
"https://api.deepinfra.com/v1/openai") (same pattern used for xiaomi and
google-ai-studio) so the url variable respects the environment override while
falling back to the default.
| { | ||
| providerId: "deepinfra", | ||
| modelName: "XiaomiMiMo/MiMo-V2.5", | ||
| inputPrice: "0.4e-6", | ||
| cachedInputPrice: "0.08e-6", | ||
| outputPrice: "2e-6", | ||
| requestPrice: "0", | ||
| contextSize: 256000, | ||
| maxOutput: 131072, | ||
| streaming: true, | ||
| reasoning: true, | ||
| vision: true, | ||
| tools: true, | ||
| jsonOutput: true, | ||
| }, |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Description: Compare context sizes for mimo-v2.5 across providers
rg -A3 -B3 'id: "mimo-v2\.5"' packages/models/src/models/xiaomi.ts
echo "---"
rg 'contextSize.*mimo-v2\.5' packages/models/src/models/xiaomi.tsRepository: theopenco/llmgateway
Length of output: 309
🏁 Script executed:
#!/bin/bash
# Get the full xiaomi.ts file to understand the structure
wc -l packages/models/src/models/xiaomi.tsRepository: theopenco/llmgateway
Length of output: 105
🏁 Script executed:
#!/bin/bash
# Find the xiaomi provider entry and show context sizes
rg -n "providerId.*xiaomi|contextSize" packages/models/src/models/xiaomi.ts | head -30Repository: theopenco/llmgateway
Length of output: 461
🏁 Script executed:
#!/bin/bash
# Search for all mimo-v2.5 entries across providers
rg -B5 "XiaomiMiMo/MiMo-V2\.5|mimo-v2\.5" packages/models/src/models/xiaomi.ts | head -60Repository: theopenco/llmgateway
Length of output: 710
🌐 Web query:
DeepInfra mimo-v2.5 context window size limit
💡 Result:
The MiMo-V2.5 model series, available on DeepInfra, supports a maximum context window size of 1 million (1,000,000) tokens [1][2][3]. This 1M-token capacity applies to both the standard MiMo-V2.5 and the MiMo-V2.5-Pro variants [1][2][4]. While the models are architected to support this length, users should note that performance and coherence across the full extent of the context window can vary depending on the specific task and implementation [5][6]. Some base versions of these models may have lower default context limits (e.g., 256K) depending on the specific deployment or configuration [1][2][3].
Citations:
- 1: https://deepinfra.com/XiaomiMiMo/MiMo-V2.5
- 2: https://deepinfra.com/XiaomiMiMo/MiMo-V2.5-Pro
- 3: https://mimo.xiaomi.com/mimo-v2-5
- 4: https://www.fonearena.com/blog/481186/xiaomi-mimo-2-5-features.html
- 5: https://www.reddit.com/r/LocalLLaMA/comments/1t7zto6/testing_mimov25iq3_s_with_1048576_context/
- 6: https://www.qoo10.co.id/en/gadget/91191/mimo-v2-5-opens-to-the-public-xiaomis-1m-token-ai-still-demands-serious-hardware/
Fix contextSize for DeepInfra's mimo-v2.5 to match official specifications.
DeepInfra's official documentation confirms that MiMo-V2.5 supports a 1M token context window, matching the model description and the xiaomi provider entry (both contextSize: 1000000). The DeepInfra provider entry at line 147 incorrectly specifies contextSize: 256000 and should be updated to 1000000.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/models/src/models/xiaomi.ts` around lines 140 - 154, Update the
DeepInfra model entry for providerId "deepinfra" and modelName
"XiaomiMiMo/MiMo-V2.5": change the contextSize value from 256000 to 1000000 so
it matches the official MiMo-V2.5 1M token context window; verify the change in
the object where contextSize is currently set and keep all other fields
unchanged.
| { | ||
| providerId: "deepinfra", | ||
| modelName: "zai-org/GLM-5.1", | ||
| inputPrice: "1.05e-6", | ||
| cachedInputPrice: "0.205e-6", | ||
| outputPrice: "3.5e-6", | ||
| requestPrice: "0", | ||
| contextSize: 198000, | ||
| maxOutput: 65536, | ||
| streaming: true, | ||
| reasoning: true, | ||
| vision: false, | ||
| tools: true, | ||
| jsonOutput: true, | ||
| }, |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Verify DeepInfra model names, pricing, and limits
# Check if DeepInfra API key is available
if [ -z "$LLM_DEEPINFRA_API_KEY" ]; then
echo "Warning: LLM_DEEPINFRA_API_KEY not set. Set it to query the API."
echo "Falling back to documentation check."
fi
# Query DeepInfra models API to list available GLM models
echo "=== Querying DeepInfra models API ==="
curl -s -H "Authorization: Bearer ${LLM_DEEPINFRA_API_KEY:-dummy}" \
https://api.deepinfra.com/v1/openai/models | \
jq -r '.data[] | select(.id | contains("GLM") or contains("glm")) | "\(.id) - context: \(.context_length // "unknown")"'
echo ""
echo "=== Expected models to find ==="
echo "- zai-org/GLM-5.1"
echo "- zai-org/GLM-5"
echo "- zai-org/GLM-4.7-Flash"Repository: theopenco/llmgateway
Length of output: 371
🌐 Web query:
DeepInfra GLM-5.1 GLM-5 GLM-4.7-Flash pricing context limits
💡 Result:
As of May 15, 2026, the pricing and context limits for the specified Z AI (Zhipu) models on DeepInfra are as follows: Model Pricing and Context Limits on DeepInfra Model | Input Price (per 1M tokens) | Output Price (per 1M tokens) | Context Limit --- | --- | --- | --- GLM-5.1 | $1.05 [1][2] | $3.50 [1][2] | Not explicitly stated* GLM-5 | $0.60 [3] | $2.08 [3] | 202,752 tokens [4][3] GLM-4.7-Flash | $0.06 [5][6] | $0.40 [5][6] | 203,000 tokens [6] *Note: While specific context limits for GLM-5.1 are not explicitly listed in the provided DeepInfra documentation, the model is part of the same generation as GLM-5, which supports a 202,752-token context window [3]. Additional Context - GLM-5.1 is positioned as a flagship model for agentic engineering and coding tasks [2][7]. - GLM-5 is a Mixture of Experts (MoE) model designed for long-context reasoning and complex systems engineering [8][3]. - GLM-4.7-Flash is optimized for cost-efficiency and low latency, making it suitable for production-scale deployments [6]. - DeepInfra provides these models as part of its dedicated inference cloud [9]. Pricing is subject to change, and users are encouraged to check the official DeepInfra pricing page for the most current rates [10].
Citations:
- 1: https://inworld.ai/models/deepinfra-glm-5-1
- 2: https://deepinfra.com/zai-org/GLM-5.1
- 3: https://deepinfra.com/zai-org/GLM-5
- 4: https://custom.typingmind.com/tools/estimate-llm-usage-costs/deepinfra/glm-5
- 5: https://preview-2025.dev.inworld.ai/models/deepinfra-glm-4-7-flash
- 6: https://deepinfra.com/blog/glm-4-7-flash-api-benchmarks
- 7: https://www.deeplearning.ai/the-batch/z-ais-glm-5-1-evaluates-interim-results-and-may-change-its-approach-hundreds-of-times-before-it-delivers-final-output/
- 8: https://deepinfra.com/blog/glm-5-api-benchmarks
- 9: https://siliconangle.com/2026/05/04/deepinfra-lands-107m-funding-build-dedicated-inference-cloud-open-source-models/
- 10: https://deepinfra.com/pricing
🌐 Web query:
DeepInfra GLM-5 GLM-4.7-Flash max output tokens completion tokens
💡 Result:
On DeepInfra, the maximum output token limits for the specified models are as follows: GLM-4.7-Flash: The maximum output token limit is 128,000 tokens [1]. Note that some third-party integrations may report a lower limit of 16,384 tokens [2], but official Z.AI documentation specifies 128K [1]. GLM-5: The maximum output token limit is 131,072 tokens [3][4]. For context, these models are part of the GLM series hosted on DeepInfra, which supports a context window of 202,752 tokens for these models [5][6]. While "completion tokens" typically refer to the output generated by the model, the "max output tokens" parameter defines the hard limit for this generation [3][1].
Citations:
- 1: https://docs.z.ai/guides/llm/glm-4.7
- 2: https://www.typingmind.com/guide/deepinfra/zai-org-GLM-4.7-Flash
- 3: https://deepinfra.ai/zai-org/GLM-5
- 4: https://deepinfra.com/zai-org/GLM-5
- 5: https://deepinfra.com/zai-org/GLM-5/api
- 6: https://deepinfra.com/zai-org/GLM-4.7/versions
🌐 Web query:
GLM-5.1 context window maximum tokens ZAI
💡 Result:
The GLM-5.1 model, developed by Z.ai, supports a maximum context window of 200,000 tokens [1][2][3]. Additionally, the model supports a maximum output length of 128,000 tokens [2][4][3].
Citations:
- 1: https://glm-ai.chat/faq/
- 2: https://apxml.com/models/glm-51
- 3: https://docs.b.ai/llmservice/models/glm-5-1/
- 4: https://www.together.ai/models/glm-51
🏁 Script executed:
# Search the codebase for other DeepInfra provider entries to compare context sizes and max outputs
rg -A 15 'providerId.*deepinfra' packages/models/src/models/ -iRepository: theopenco/llmgateway
Length of output: 13234
Correct context and max output limits for all three DeepInfra GLM models.
The pricing values are correct, but the context window and max output limits are significantly undersized:
- GLM-5.1 (lines 60-74): Should have
contextSize: 200000(not 198000) andmaxOutput: 128000(not 65536) - GLM-5 (lines 211-225): Should have
contextSize: 202752(not 198000) andmaxOutput: 131072(not 65536) - GLM-4.7-Flash (lines 712-726): Should have
contextSize: 203000(not 198000) andmaxOutput: 128000(not 65536)
The max output reduction from 128K-131K to 65K represents a 50% limit decrease, which significantly impacts usability for long-form generation tasks.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/models/src/models/zai.ts` around lines 60 - 74, Update the DeepInfra
GLM model entries to correct their context and output limits: for the object
with modelName "zai-org/GLM-5.1" set contextSize to 200000 and maxOutput to
128000; for the object with modelName "zai-org/GLM-5" set contextSize to 202752
and maxOutput to 131072; and for the object with modelName
"zai-org/GLM-4.7-Flash" set contextSize to 203000 and maxOutput to 128000.
Ensure you update the corresponding objects in the models array (matching
providerId "deepinfra" and the modelName strings) so the values replace the
previous 198000 and 65536 entries.
Remove models that fail E2E tests (JSON/tool calls/reasoning issues): MiMo V2.5/V2.5-Pro, GLM-5/4.7-Flash, MiniMax M2.5, Qwen3-Max, Qwen3.5-397B, Qwen3.6-35B, Kimi K2.6. Fix DeepInfra endpoint URL double /v1 path issue. Remaining 5 models all pass 100%: DeepSeek V4-Flash, V4-Pro, V3.2, Kimi K2.5, GLM-5.1. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
♻️ Duplicate comments (2)
packages/actions/src/get-provider-endpoint.ts (1)
288-290:⚠️ Potential issue | 🟠 Major | ⚡ Quick winOptional baseUrl not respected.
The
deepinfraprovider definition declaresLLM_DEEPINFRA_BASE_URLas optional, but this code ignores it and always uses the hardcoded default. Other providers with optional baseUrl (e.g.,xiaomiat lines 231–237,google-ai-studioat lines 139–145) useenvValueOrDefaultto respect the environment variable when set.🔧 Proposed fix to respect optional baseUrl
case "deepinfra": - url = "https://api.deepinfra.com/v1/openai"; + url = + envValueOrDefault( + "deepinfra", + "baseUrl", + "https://api.deepinfra.com/v1/openai", + ) ?? "https://api.deepinfra.com/v1/openai"; break;🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/actions/src/get-provider-endpoint.ts` around lines 288 - 290, In the case "deepinfra" branch the hardcoded URL is used instead of honoring the optional environment variable; update the branch in get-provider-endpoint (case "deepinfra") to call envValueOrDefault("LLM_DEEPINFRA_BASE_URL", "https://api.deepinfra.com/v1/openai") and assign that to url (same pattern as used for xiaomi and google-ai-studio), leaving the break intact so the env-provided base URL is respected when present.packages/models/src/models/zai.ts (1)
60-74:⚠️ Potential issue | 🟠 Major | ⚡ Quick winCorrect context and max output limits for DeepInfra GLM-5.1.
The pricing values are correct, but the context window and max output limits are significantly undersized:
- GLM-5.1: Should have
contextSize: 200000(not 198000) andmaxOutput: 128000(not 65536)The max output reduction from 128K to 65K represents a 50% limit decrease, which significantly impacts usability for long-form generation tasks.
🔧 Proposed fix
{ providerId: "deepinfra", modelName: "zai-org/GLM-5.1", inputPrice: "1.05e-6", cachedInputPrice: "0.205e-6", outputPrice: "3.5e-6", requestPrice: "0", - contextSize: 198000, - maxOutput: 65536, + contextSize: 200000, + maxOutput: 128000, streaming: true, reasoning: true, vision: false, tools: true, jsonOutput: true, },🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/models/src/models/zai.ts` around lines 60 - 74, Update the DeepInfra GLM-5.1 model entry in the models list (the object with providerId "deepinfra" and modelName "zai-org/GLM-5.1" in zai.ts) to use the correct limits: set contextSize to 200000 and maxOutput to 128000; leave pricing and other flags unchanged so only the two numeric fields are corrected.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Duplicate comments:
In `@packages/actions/src/get-provider-endpoint.ts`:
- Around line 288-290: In the case "deepinfra" branch the hardcoded URL is used
instead of honoring the optional environment variable; update the branch in
get-provider-endpoint (case "deepinfra") to call
envValueOrDefault("LLM_DEEPINFRA_BASE_URL",
"https://api.deepinfra.com/v1/openai") and assign that to url (same pattern as
used for xiaomi and google-ai-studio), leaving the break intact so the
env-provided base URL is respected when present.
In `@packages/models/src/models/zai.ts`:
- Around line 60-74: Update the DeepInfra GLM-5.1 model entry in the models list
(the object with providerId "deepinfra" and modelName "zai-org/GLM-5.1" in
zai.ts) to use the correct limits: set contextSize to 200000 and maxOutput to
128000; leave pricing and other flags unchanged so only the two numeric fields
are corrected.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: f4f41707-e5be-49aa-8d51-9206ea397e30
📒 Files selected for processing (3)
packages/actions/src/get-provider-endpoint.tspackages/models/src/models/moonshot.tspackages/models/src/models/zai.ts
🚧 Files skipped from review as they are similar to previous changes (1)
- packages/models/src/models/moonshot.ts
## Summary Add DeepInfra provider icon (dot-grid pattern) to the provider icons component. This was missing from PR #2301 which added DeepInfra as a provider. ## Changes - Added `DeepInfraIcon` SVG component to `packages/shared/src/components/provider-icons.tsx` - Added `deepinfra` entry to `ProviderIcons` map - Added `deepinfra` entry to `providerLogoUrls` map ## Test plan - [x] `pnpm --filter shared build` passes - [x] Icon renders correctly with `currentColor` fill (adapts to dark/light mode) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for the DeepInfra provider with integrated icon and branding display. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/theopenco/llmgateway/pull/2319?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Summary
Add DeepInfra (https://deepinfra.com) as a new inference provider. DeepInfra provides an OpenAI-compatible API hosting flagship models with competitive pricing. This PR adds DeepInfra as an alternative provider for models we already support.
Implementation Checklist
Provider Setup
packages/models/src/providers.tspackages/actions/src/get-provider-endpoint.tspackages/actions/src/get-provider-headers.ts/chat/completionsnot/v1/chat/completions)Models (5 models, all passing E2E 100%)
E2E Test Results
Technical Notes
Models Removed (failed E2E)
Tested but removed due to provider-side limitations:
Summary by CodeRabbit