Skip to content

feat(providers): add Cortecs as an OpenAI-compatible provider - #43872

Merged
krrish-berri-2 merged 1 commit into
mainfrom
litellm_cortecs_provider
Oct 1, 2026
Merged

krrish-berri-2 merged 1 commit into
mainfrom
litellm_cortecs_provider

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • Registers cortecs as a JSON-configured OpenAI-compatible provider
  • Wires chat completions, Responses, and Anthropic Messages to https://api.cortecs.ai/v1
  • Adds Cortecs to the Admin UI Add Model form and the endpoint support matrix

User Flow

Before: a proxy admin who adds a Cortecs model gets a provider error and no working deployment

  1. They add model: cortecs/gpt-6-sol with their Cortecs key to the proxy config and start the proxy
  2. The proxy logs LLM Provider NOT provided for that model at boot
  3. POST http://localhost:4000/v1/chat/completions (or /v1/responses, /v1/messages) returns 400 "no healthy deployments"

After: the same config serves Cortecs on all three endpoints

  1. They add model: cortecs/gpt-6-sol with CORTECS_API_KEY to the proxy config and start the proxy
  2. The model loads with no provider error, and Cortecs shows up in the Add Model provider list
  3. POST http://localhost:4000/v1/chat/completions, /v1/responses, and /v1/messages are forwarded to https://api.cortecs.ai/v1

Relevant issues

Supersedes #26971 and #24801 by @markoarnauto, credited as co-author

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy (litellm.proxy.proxy_cli, DB-free, master_key=sk-1234) with one model cortecs-gpt -> cortecs/gpt-6-sol, CORTECS_API_KEY=sk-not-a-real-key. No real Cortecs key exists, so success here means the request reaches api.cortecs.ai and comes back with Cortecs's own 401 auth_error instead of litellm failing to resolve the provider.

Proxy config (/home/ubuntu/cortecs_proxy_config.yaml):

model_list:
  - model_name: cortecs-gpt
    litellm_params:
      model: cortecs/gpt-6-sol
      api_key: os.environ/CORTECS_API_KEY
general_settings:
  master_key: sk-1234
  dangerously_permit_weak_or_unset_master_key: true

Run with DATABASE_URL/DIRECT_URL unset, CORTECS_API_KEY=sk-not-a-real-key exported:

.venv/bin/python -m litellm.proxy.proxy_cli --config /home/ubuntu/cortecs_proxy_config.yaml --port 4000 --num_workers 1

Before (9dda4d8)

On main, cortecs/gpt-6-sol resolves no provider. The router logs litellm.BadRequestError: LLM Provider NOT provided. ... You passed model=cortecs/gpt-6-sol at boot and the deployment is never created, so every route returns "no healthy deployments". (Proxy run from a git worktree at this sha so litellm imports the main checkout.)

Chat completions

  1. Request
curl -s localhost:4000/v1/chat/completions \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"model":"cortecs-gpt","messages":[{"role":"user","content":"Say hello"}]}'
  1. Response: HTTP 400
{"error":{"message":"litellm.BadRequestError: You passed in model=cortecs-gpt. There are no healthy deployments for this model\n\nLiteLLM: model group 'cortecs-gpt' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}

Responses

  1. Request
curl -s localhost:4000/v1/responses \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"model":"cortecs-gpt","input":"Say hello"}'
  1. Response: HTTP 400
{"error":{"message":"litellm.BadRequestError: You passed in model=cortecs-gpt. There are no healthy deployments for this model\n\nLiteLLM: model group 'cortecs-gpt' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}

Messages

  1. Request
curl -s localhost:4000/v1/messages \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -H 'anthropic-version: 2023-06-01' \
  -d '{"model":"cortecs-gpt","max_tokens":32,"messages":[{"role":"user","content":"Say hello"}]}'
  1. Response: HTTP 400
{"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: You passed in model=cortecs-gpt. There are no healthy deployments for this model\n\nLiteLLM: model group 'cortecs-gpt' failed with the error above. No fallback was attempted."}}

After (a84b627)

Each endpoint reaches https://api.cortecs.ai/v1/<path> and returns Cortecs's own 401 auth_error, which is the response a customer would see with a bad key. (Each capture used a freshly restarted proxy, since the first 401 puts the deployment in router cooldown.)

Chat completions

  1. Request
curl -s localhost:4000/v1/chat/completions \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"model":"cortecs-gpt","messages":[{"role":"user","content":"Say hello"}]}'
  1. Response: HTTP 401
{"error":{"message":"litellm.AuthenticationError: AuthenticationError: CortecsException - Unauthorized\n\nLiteLLM: model group 'cortecs-gpt' failed with the error above. No fallback was attempted.","type":"authentication_error","param":null,"code":"401"}}

Responses

  1. Request
curl -s localhost:4000/v1/responses \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"model":"cortecs-gpt","input":"Say hello"}'
  1. Response: HTTP 401, with Cortecs's upstream error body passed through
{"error":{"message":"litellm.AuthenticationError: AuthenticationError: CortecsException - {\"error\":{\"message\":\"Unauthorized\",\"type\":\"auth_error\",\"param\":\"None\",\"code\":\"401\"}}\n\nLiteLLM: model group 'cortecs-gpt' failed with the error above. No fallback was attempted.","type":"authentication_error","param":null,"code":"401"}}

Messages

  1. Request
curl -s localhost:4000/v1/messages \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -H 'anthropic-version: 2023-06-01' \
  -d '{"model":"cortecs-gpt","max_tokens":32,"messages":[{"role":"user","content":"Say hello"}]}'
  1. Response: HTTP 401
{"type":"error","error":{"type":"authentication_error","message":"litellm.AuthenticationError: AuthenticationError: CortecsException - {\"error\":{\"message\":\"Unauthorized\",\"type\":\"auth_error\",\"param\":\"None\",\"code\":\"401\"}}\n\nLiteLLM: model group 'cortecs-gpt' failed with the error above. No fallback was attempted."}}

Add Model form lists Cortecs

  1. Request
curl -s localhost:4000/public/providers/fields \
  -H 'Authorization: Bearer sk-1234' \
  | jq '.[] | select(.litellm_provider == "cortecs")'
  1. Response
{
  "provider": "CORTECS",
  "provider_display_name": "Cortecs",
  "litellm_provider": "cortecs",
  "credential_fields": [
    {"key": "api_base", "label": "API Base", "placeholder": "https://api.cortecs.ai/v1", "required": false, "field_type": "text"},
    {"key": "api_key", "label": "API Key", "placeholder": null, "required": true, "field_type": "password"}
  ],
  "default_model_placeholder": "cortecs/gpt-6-sol"
}

Unit tests: tests/unit/llms/openai_like/test_cortecs_provider.py covers env and explicit credential resolution, the Add Model form entry, the endpoint matrix, and respx-captured chat, Responses, and Messages requests. Mutation check: changing base_url to /v2 turns 4 tests red, and dropping /v1/responses from supported_endpoints turns the Responses test red

Type

🆕 New Feature

Caveats (if any)

Medium

  • No Cortecs key yet, so proof stops at Cortecs's 401
  • No cost map entries: Cortecs prices in EUR from a live catalog
    • Spend tracking reads $0 until per-model pricing is added

Low

  • Embeddings, image, audio, and OCR not wired
    • JSON-configured providers do not route embedding() today
  • Docs page docs/providers/cortecs lands in a separate litellm-docs PR

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/5f14fda38083472abc8ac2d183c1e19a
Open in Devin Desktop: https://app.devin.ai/desktop/session/5f14fda38083472abc8ac2d183c1e19a?variant=devin
Requested by: @krrish-berri-2

Co-authored-by: markoarnauto <7702545+markoarnauto@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds a new LLM provider integration.

The PR should not merge until Cortecs usage can be accounted for when deployment pricing is omitted.

Findings

  1. P1 Cortecs usage records zero spend ▶
  2. P2 Streaming routes lack test coverage ▶

Summary

Registers Cortecs as a JSON-configured OpenAI-compatible provider for chat completions, Responses, and Messages, and adds it to the Add Model form and endpoint matrices.

  • Provider resolution and non-streaming request forwarding are covered by mocked tests.
  • Deployments without explicit pricing record no spend, and the new endpoint tests do not cover streaming.

Reviews (1) · Last reviewed commit: "feat(providers): add Cortecs as an OpenA..."

Comment on lines +183 to 189
"cortecs": {
"base_url": "https://api.cortecs.ai/v1",
"api_key_env": "CORTECS_API_KEY",
"api_base_env": "CORTECS_API_BASE",
"supported_endpoints": ["/v1/chat/completions", "/v1/responses", "/v1/messages"]
},
"pinstripes": {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Cortecs usage records zero spend

If a Cortecs deployment has no custom pricing, a successful call to cortecs/gpt-6-sol has no matching provider-specific cost entry. Cost calculation returns no cost, and the proxy records $0 spend despite reported token usage. Key and team budgets based on accumulated spend therefore cannot cap this usage. Add pricing or require explicit pricing for billable Cortecs deployments.

Comment on lines +156 to +160
monkeypatch.setattr(litellm, "disable_aiohttp_transport", True)
monkeypatch.setattr(litellm, "in_memory_llm_clients_cache", LLMClientCache())
with respx.mock() as upstream:
route: Final = upstream.post("https://api.cortecs.ai/v1/messages").respond(
200,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Streaming routes lack test coverage

The new request tests cover only non-streaming calls, although Cortecs is registered for chat completions, Responses, and Messages. A broken streaming URL, transport, or event parser on any of these routes would pass this suite. Please add mocked streaming tests for the advertised endpoints.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_cortecs_provider (a84b627) with main (9dda4d8)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (b800528) during the generation of this report, so 9dda4d8 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@krrish-berri-2
krrish-berri-2 merged commit ae60fd1 into main Oct 1, 2026
98 of 100 checks passed
@krrish-berri-2
krrish-berri-2 deleted the litellm_cortecs_provider branch October 1, 2026 02:49
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants