Skip to content

feat(models): add Qwen3 Embedding 8B and BGE-M3 via DeepInfra - #3188

Merged
steebchen merged 1 commit into
theopenco:mainfrom
vicovaro:feat/deepinfra-qwen3-embedding-rerank
Jul 24, 2026
Merged

steebchen merged 1 commit into
theopenco:mainfrom
vicovaro:feat/deepinfra-qwen3-embedding-rerank

Conversation

@vicovaro

@vicovaro vicovaro commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds two embedding model definitions served through the DeepInfra provider:

Qwen3 Embedding 8B

  • 32K context / 8B params / 36 layers / 4096-dim output (configurable 32–8192 via MRL)
  • Works immediately via the existing /v1/embeddings endpoint
  • Priced at $0.010/1M tokens on DeepInfra
  • No. 1 on the MTEB multilingual leaderboard as of June 2025 (score: 70.58), supports 100+ languages
  • Reference: https://deepinfra.com/Qwen/Qwen3-Embedding-8B

BGE-M3

  • 8K context / 567M params / 24 layers / 1024-dim output (configurable 32–8192 via MRL)
  • Dense, sparse, and multi-vector retrieval across 100+ languages
  • Priced at $0.010/1M tokens on DeepInfra
  • MIT license
  • Reference: https://deepinfra.com/BAAI/bge-m3

DeepInfra embeddings path fix

  • DeepInfra's base URL is https://api.deepinfra.com/v1/openai, so the embeddings handler was producing …/v1/openai/v1/embeddings which 404s. This PR adds a DeepInfra-specific branch that uses /embeddings directly.

Files changed

File Change
packages/models/src/models/alibaba.ts +12: DeepInfra provider mapping for Qwen3 Embedding 8B
packages/models/src/models/baai.ts +27: BGE-M3 definition (new family)
packages/models/src/models.ts +2: baai import + models array spread
apps/gateway/src/embeddings/embeddings.ts +19: DeepInfra embeddings path fix

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

DeepInfra embedding routing and model metadata now support rerank capabilities. Gateway schemas and output validation recognize rerank, BAAI and Alibaba models are registered, and chat completion e2e filtering excludes rerank models.

Changes

DeepInfra rerank support

Layer / File(s) Summary
Rerank contracts and model catalog
packages/models/src/models.ts, packages/models/src/models/baai.ts, packages/models/src/models/alibaba.ts, packages/models/src/model-metadata.spec.ts
Model metadata supports rerank capabilities and outputs, exports the BAAI model catalog, adds DeepInfra to qwen3-embedding-8b, and registers reranker and embedding models.
Gateway rerank modality exposure
apps/gateway/src/lib/validate-model-output.ts, apps/gateway/src/models/models.ts
Gateway validation, endpoint mapping, schema validation, and model listings accept the rerank modality.
DeepInfra embedding routing
apps/gateway/src/embeddings/embeddings.ts, apps/gateway/src/chat-helpers.e2e.ts
DeepInfra embedding requests use /embeddings with provider-specific fields, while rerank models are excluded from general chat e2e tests.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant GatewayEmbeddings
  participant DeepInfra
  Client->>GatewayEmbeddings: Submit embedding request
  GatewayEmbeddings->>DeepInfra: POST /embeddings with input and model
  DeepInfra-->>GatewayEmbeddings: Return embedding response
  GatewayEmbeddings-->>Client: Return embedding response
Loading

Possibly related PRs

Suggested reviewers: steebchen

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is clear and relevant to the main change: adding DeepInfra-backed model definitions for Qwen3 Embedding 8B and BGE-M3.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch 4 times, most recently from db559e8 to 9df12d7 Compare July 22, 2026 18:02
@steebchen

Copy link
Copy Markdown
Member

Ran the e2e suite against this branch (TEST_MODELS="deepinfra/qwen3-embedding-8b,deepinfra/qwen3-reranker-8b" FULL_MODE=true pnpm test:e2e). Found three issues that need fixing — no changes were made on our side.

1. Blocker: the gateway no longer compiles

Adding "rerank" to ModelDefinition.output breaks pnpm build / pnpm test:e2e because the gateway mirrors that union in three places that weren't updated:

  • apps/gateway/src/lib/validate-model-output.ts:10 — the ModelOutput union is missing "rerank". Note OUTPUT_ENDPOINT just below it is a Record<ModelOutput, …>, so it also needs a rerank entry (e.g. { label: "a rerank", endpoint: "/v1/rerank" }).
  • apps/gateway/src/models/models.ts:26 — the zod enum for output_modalities is missing "rerank".
  • apps/gateway/src/models/models.ts:220 — the inline outputModalities union is missing "rerank".
gateway:build: src/lib/validate-model-output.ts(37,5): error TS2322: Type '"rerank"' is not assignable to type 'ModelOutput'.
gateway:build: src/models/models.ts(220,10): error TS2322: Type '"rerank"' is not assignable to type '…'

2. qwen3-embedding-8b 404s through the gateway

After locally patching the build errors above to let the suite run, the embeddings e2e test fails:

FAIL embeddings 'deepinfra/qwen3-embedding-8b' — expected 404 to be 200
embeddings response: { "detail": "Not Found" }

Root cause: the embeddings handler (apps/gateway/src/embeddings/embeddings.ts:899) builds ${baseUrl}/v1/embeddings for all non-Google providers, but DeepInfra's base URL is https://api.deepinfra.com/v1/openai, producing …/v1/openai/v1/embeddings, which DeepInfra 404s. Verified directly against DeepInfra with Qwen/Qwen3-Embedding-8B:

  • https://api.deepinfra.com/v1/openai/v1/embeddings → 404
  • https://api.deepinfra.com/v1/openai/embeddings → 200

DeepInfra is the first embeddings provider whose base URL already contains the /v1-style suffix, so the handler needs a provider-aware path here (the chat path already special-cases it: case "deepinfra": return ${url}/chat/completions`` in packages/actions/src/get-provider-endpoint.ts:648). Until this is fixed, the embedding model is dead on arrival through the gateway.

3. qwen3-reranker-8b leaks into the chat/responses e2e tests

The e2e model list builder (filteredModels in apps/gateway/src/chat-helpers.e2e.ts:180) excludes video-, audio-, OCR-, and embeddings-only models from the chat-completions tests, but has no exclusion for rerank-output models. So the reranker gets run through chat-basic, and responses e2e and fails with 400 (which is the gateway correctly rejecting a rerank model on a chat endpoint once validate-model-output knows about rerank):

FAIL chat-basic 'deepinfra/qwen3-reranker-8b' — expected 400 to be 200
FAIL responses single-turn 'deepinfra/qwen3-reranker-8b' — expected 400 to be 200
FAIL responses multi-turn 'deepinfra/qwen3-reranker-8b' — expected 400 to be 200

Please add a rerank exclusion to that filter (mirroring the embeddings one), since the reranker can't be exercised until the /v1/rerank endpoint lands.

What passed

  • packages/models/src/model-metadata.spec.ts — 4/4 unit tests pass with the new flag mapping.
  • All other scoped e2e files pass (the two models are correctly skipped where not applicable).
  • Pricing notation checks out (0.01e-6 = $0.01/M, 0.05e-6 = $0.05/M) and the DeepInfra key + external IDs are valid (verified with a direct API call).

@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch 2 times, most recently from b6a6cbb to 8c69131 Compare July 22, 2026 18:16
@steebchen

Copy link
Copy Markdown
Member

Re-ran the round on the updated branch (8c69131) — all three earlier issues are resolved and the scoped e2e suite is green.

TEST_MODELS="deepinfra/qwen3-embedding-8b,deepinfra/qwen3-reranker-8b" FULL_MODE=true pnpm test:e2e
Test Files  26 passed | 1 skipped (27)
     Tests  77 passed | 76 skipped (153)
  • Build compiles (gateway ModelOutput mirrors + OUTPUT_ENDPOINT updated).
  • embeddings 'deepinfra/qwen3-embedding-8b' passes — the DeepInfra-specific /embeddings path works end-to-end through the gateway.
  • The reranker no longer leaks into chat/responses e2e (new rerank exclusion in chat-helpers.e2e.ts works).
  • model-metadata.spec.ts unit tests: 4/4 pass.

One heads-up, not a blocker: during a first run the embeddings test timed out because DeepInfra was returning 429 {"code":"engine_overloaded","message":"Model busy, retry later"} for Qwen/Qwen3-Embedding-8B. It recovered within ~a minute and subsequent runs passed quickly, so it looks like transient capacity on DeepInfra's side — but if it recurs in CI it may be worth marking the mapping's stability accordingly.

@steebchen

Copy link
Copy Markdown
Member

One last thing before merge: the lint / run CI check is failing on a Prettier formatting issue — the new output union in packages/models/src/models.ts:650 exceeds the line width and needs to be wrapped multiline. Running pnpm format from the repo root will fix it. Everything else is green (build, tests, generate, and the scoped e2e suite I ran locally).

@steebchen
steebchen force-pushed the feat/deepinfra-qwen3-embedding-rerank branch from 8c69131 to ff5747c Compare July 22, 2026 20:02
@steebchen
steebchen marked this pull request as ready for review July 22, 2026 20:02
Copilot AI review requested due to automatic review settings July 22, 2026 20:02

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@vicovaro
vicovaro marked this pull request as draft July 22, 2026 20:04
@vicovaro
vicovaro marked this pull request as ready for review July 22, 2026 20:04
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@vicovaro

Copy link
Copy Markdown
Contributor Author

I have tried separately with DeepInfra's API key the Qwen3 8B embedded model, and it's also often buggy and not working. I have switched to BAAI/bge-m3 that always worked as in 99.99%. I think I will add that instead.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
apps/gateway/src/embeddings/embeddings.ts (1)

899-916: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Avoid duplicating the embedding request-body builder.

This branch duplicates the input, model, encoding_format, dimensions, and user construction that follows at Lines 918-931. Keep the provider-specific difference limited to URL selection and build the body once, preventing the two paths from drifting.

As per coding guidelines, apply DRY principles for code reuse.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/embeddings/embeddings.ts` around lines 899 - 916, Update the
isDeepInfra branch in the embeddings request flow to set only the
provider-specific upstreamUrl, then remove its duplicated requestBody
construction. Reuse the shared builder after the branch so input, model,
encoding_format, dimensions, and user are assembled once for both paths.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/models/src/models/alibaba.ts`:
- Around line 2937-2959: Update the Qwen3 Reranker 8B model entry in the Alibaba
model definitions to include the appropriate deactivatedAt value, keeping it
inactive until the /v1/rerank endpoint is available. Preserve the existing model
metadata and provider configuration.

---

Nitpick comments:
In `@apps/gateway/src/embeddings/embeddings.ts`:
- Around line 899-916: Update the isDeepInfra branch in the embeddings request
flow to set only the provider-specific upstreamUrl, then remove its duplicated
requestBody construction. Reuse the shared builder after the branch so input,
model, encoding_format, dimensions, and user are assembled once for both paths.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 8957e085-493e-4746-97bb-2f94edd49613

📥 Commits

Reviewing files that changed from the base of the PR and between 52b8993 and ff5747c.

📒 Files selected for processing (7)
  • apps/gateway/src/chat-helpers.e2e.ts
  • apps/gateway/src/embeddings/embeddings.ts
  • apps/gateway/src/lib/validate-model-output.ts
  • apps/gateway/src/models/models.ts
  • packages/models/src/model-metadata.spec.ts
  • packages/models/src/models.ts
  • packages/models/src/models/alibaba.ts

Comment thread packages/models/src/models/alibaba.ts Outdated
@vicovaro vicovaro changed the title feat(models): add Qwen3 Embedding 8B and Qwen3 Reranker 8B via DeepInfra feat(models): add Qwen3 Embedding, Qwen3 Reranker & BGE-M3 via DeepInfra Jul 22, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/models/src/models/baai.ts`:
- Line 17: Update the outputPrice value in the model definition to use the
per-token zero notation "0e-6" instead of "0", preserving the existing string
type and surrounding configuration.
- Line 8: Update the BAAI model description in the model metadata to remove the
claim that it supports the `dimensions` parameter, while preserving the
remaining capabilities and description.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 0122ad4f-2555-40c7-bd83-0d8820c28a62

📥 Commits

Reviewing files that changed from the base of the PR and between ff5747c and 0f21648.

📒 Files selected for processing (2)
  • packages/models/src/models.ts
  • packages/models/src/models/baai.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • packages/models/src/models.ts

Comment thread packages/models/src/models/baai.ts Outdated
Comment thread packages/models/src/models/baai.ts
@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch 3 times, most recently from 4fb3a3d to a1d7546 Compare July 22, 2026 20:30
@steebchen

Copy link
Copy Markdown
Member

One structural request before this merges: please either drop qwen3-reranker-8b from this PR, or implement /v1/rerank in it — shipping the model definition without the endpoint isn't a workable middle ground.

As merged, the reranker would be a catalog entry with zero working code path: calling it on chat completions 400s pointing at /v1/rerank, and /v1/rerank itself 404s because it doesn't exist. Meanwhile the model still surfaces everywhere:

  • It appears on the public /models directory, and isTextOutput() in packages/shared/src/components/models-directory/model-category-filters.ts classifies rerank-output models as text, so it would show up on /models/text and /models/cheapest as if it were a chat model.
  • The playground chat selector (apps/playground/src/app/playground-shell.tsx) and group-chat selectors only exclude embedding/media outputs, so the reranker is selectable for chat — where it can only error. The model detail page's "Try in Playground" CTA deep-links it into the chat playground too.
  • The sitemap generates an indexed /models/qwen3-reranker-8b page for a model nobody can call.
  • Catalog policy means models can never be removed once merged, only deactivated — so if the follow-up stalls, this is a permanently dead, mislabeled entry.

Both directions are fine by us:

  1. Remove for now (smaller): drop the qwen3-reranker-8b definition, the rerank flag/output type, and the related gateway/e2e plumbing from this PR, leaving just the DeepInfra qwen3-embedding-8b mapping — that part is verified working end-to-end and can merge immediately. Reintroduce the reranker in the follow-up PR alongside the endpoint.
  2. Implement here: add /v1/rerank (Cohere-compatible schema, DeepInfra translation — the same shape LiteLLM uses, so existing rerank clients work out of the box) in this PR, plus the frontend exclusions above (playground selectors, isTextOutput, CTA) so the model is correctly represented where it appears.

Which way do you want to go?

@vicovaro

Copy link
Copy Markdown
Contributor Author

I am choosing option 1.
I think I will remove anything related to rerank for now, as doing it properly would require to add the endpoint. That itself deserves another PR. Keeping only the Qwen and BAAI models for embedding.

thx for comment

@vicovaro vicovaro changed the title feat(models): add Qwen3 Embedding, Qwen3 Reranker & BGE-M3 via DeepInfra feat(models): add Qwen3 Embedding 8B and BGE-M3 via DeepInfra Jul 23, 2026
@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch 3 times, most recently from 09f7b77 to 5a5bdad Compare July 23, 2026 17:33
@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch 2 times, most recently from 40f04b0 to 3c28199 Compare July 23, 2026 17:35
@steebchen

Copy link
Copy Markdown
Member

Re-verified the updated branch (3c28199, rerank removed, BGE-M3 added) — everything is green.

TEST_MODELS="deepinfra/qwen3-embedding-8b,deepinfra/bge-m3" FULL_MODE=true pnpm test:e2e
Test Files  26 passed | 2 skipped (28)
     Tests  78 passed | 78 skipped (156)

Both embedding models return valid vectors end-to-end through the gateway's /v1/embeddings (the DeepInfra /embeddings path fix works for both). Lint and the model-metadata unit spec (4/4) pass locally as well. Thanks for the quick turnaround on dropping the rerank groundwork — this is now a clean embeddings-only PR.

@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch from 3c28199 to b5a2010 Compare July 23, 2026 18:28
Add two embedding model definitions served through the DeepInfra
provider and fix the embeddings upstream URL for DeepInfra.

- Add a DeepInfra provider mapping to the existing Qwen3 Embedding 8B
  model (32K context, 4096-dim output, MRL support, $0.010/1M tokens)
- Add BGE-M3 as a new BAAI family entry (8K context, 1024-dim output,
  dense/sparse/multi-vector retrieval, $0.010/1M tokens)
- Fix the DeepInfra embeddings upstream URL to use /embeddings instead
  of /v1/embeddings so it does not duplicate the /v1/openai segment
  already present in DeepInfra's base URL

Co-Authored-By: Claude <noreply@anthropic.com>
@vicovaro
vicovaro force-pushed the feat/deepinfra-qwen3-embedding-rerank branch 2 times, most recently from c5317a5 to 4bc15ee Compare July 23, 2026 18:33
@steebchen
steebchen merged commit f245685 into theopenco:main Jul 24, 2026
10 checks passed
@vicovaro
vicovaro deleted the feat/deepinfra-qwen3-embedding-rerank branch August 5, 2026 18:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants