Skip to content

feat: add Gemini embeddings support + fix missing OpenAI response fields (Letta compat) - #148

Closed
xuandung38 wants to merge 11 commits into
decolua:masterfrom
xuandung38:feat/gemini-embeddings-letta-fix
Closed

xuandung38 wants to merge 11 commits into
decolua:masterfrom
xuandung38:feat/gemini-embeddings-letta-fix

Conversation

@xuandung38

@xuandung38 xuandung38 commented Feb 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Two improvements to 9router:

  1. Gemini embeddings support — extend the existing /v1/embeddings endpoint with Google AI (Gemini) provider
  2. Fix: inject missing object/created fields — makes responses compatible with strict OpenAI clients

1. Gemini Embeddings Support

The existing /v1/embeddings endpoint (added in #146) only supports OpenAI passthrough. This PR adds Google AI (Gemini) as an additional embeddings provider.

Why

Many use cases require embeddings that are either free-tier or do not depend on a paid OpenAI subscription. Gemini offers high-quality embeddings (up to 3072 dimensions) with a generous free quota.

⚠️ Provider support at this time:

  • ✅ Google AI (Gemini) — fully supported in this PR
  • ✅ OpenAI — supported via passthrough (existing)
  • ❌ GitHub Copilot, Anthropic/Claude, and other providers do not have an embeddings API and are not supported

New models

Model Dimensions Notes
gemini/gemini-embedding-001 3072 Latest, highest quality
gemini/text-embedding-005 768 Efficient
gemini/text-embedding-004 768 Stable

Setup

Add a Google AI provider in your 9router provider config with a valid Gemini API key.
Get a free key at: https://aistudio.google.com/app/apikey

Usage

# Gemini embeddings
curl http://localhost:20128/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model": "gemini/gemini-embedding-001", "input": "Hello world"}'

# OpenAI embeddings (existing, unchanged)
curl http://localhost:20128/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model": "openai/text-embedding-3-small", "input": "Hello world"}'

Response follows standard OpenAI format:

{
  "object": "list",
  "data": [{"object": "embedding", "index": 0, "embedding": [0.123, ...]}],
  "model": "gemini-embedding-001",
  "usage": {"prompt_tokens": 2, "total_tokens": 2}
}

2. Fix: Inject Missing object / created Fields

Why

Some upstream providers (notably GitHub Copilot) return chat completion responses without the object and created fields. This causes strict OpenAI-compatible clients to reject the response with a validation error.

One concrete example: Letta (open-source memory layer) validates every LLM response against the OpenAI ChatCompletionResponse schema using Pydantic. A missing created or object field results in a 422 error, making it impossible to use 9router as Letta's LLM backend.

Fix

After receiving a response from the upstream provider, inject the missing fields if absent:

  • Non-streaming: "object": "chat.completion" and "created": <unix_timestamp>
  • Streaming chunks: "object": "chat.completion.chunk" and "created": <unix_timestamp>

This is a safe, additive change — providers that already include these fields are unaffected.


Files Changed

File Change
open-sse/services/embeddingsCore.js Gemini provider: different URL, query-param auth, normalized response
open-sse/config/providerModels.js Register 3 Gemini embedding models
open-sse/handlers/chatCore.js Inject object/created for non-streaming responses
open-sse/utils/stream.js Inject object/created for each SSE chunk (streaming passthrough)

Breaking Changes

None.

Testing

  • Gemini gemini-embedding-001 → 3072-dim vectors ✅
  • Letta integration: confirmed working end-to-end after the object/created fix ✅

- Add isGeminiProvider() helper for 'gemini' / 'google_ai_studio' providers
- buildEmbeddingsUrl(): Gemini uses generativelanguage.googleapis.com with
  API key as query param; single input → :embedContent, array → :batchEmbedContents
- buildEmbeddingsHeaders(): Gemini doesn't use Authorization header (key in URL)
- buildEmbeddingsBody(): Gemini format differs from OpenAI:
  single → { model, content: { parts: [{ text }] } }
  batch  → { requests: [{ model, content: { parts: [{ text }] } }] }
- normalizeEmbeddingsResponse(): convert Gemini response to OpenAI list format
  single: { embedding: { values: [] } } → data[0].embedding
  batch:  { embeddings: [{ values: [] }] } → data[].embedding
- Rebuild Gemini URL on retry (API key is in query param, not header)
- Simplify default case in buildEmbeddingsHeaders (redundant branch removed)
- Add Gemini embedding models to providerModels for /v1/models:
  gemini-embedding-001 (3072 dims), text-embedding-005, text-embedding-004 (legacy)
- Non-streaming: ensure object=chat.completion and created=<unix_ts>
  are always present before returning translatedResponse, even when
  the upstream (GitHub/OpenAI passthrough) omits them.
- Streaming passthrough: inject object=chat.completion.chunk and
  created=<unix_ts> into each parsed SSE chunk that lacks them,
  then re-serialise so the fields reach the client.
- Gemini/Claude paths already set these fields; no change needed there.
…re returning to client

GitHub Copilot provider adds a non-OpenAI 'padding' field to message objects.
Strict clients like Letta crash on unknown fields with 'str' object has no attribute model_dump'.

- chatCore.js: strip after translateNonStreamingResponse (non-streaming path)
- stream.js: strip from delta in passthrough mode (streaming path)

Allowed fields: role, content, tool_calls, tool_call_id, name, reasoning_content, refusal
@xuandung38
xuandung38 marked this pull request as draft February 18, 2026 13:16
…s/chat.js

Add sanitizeJsonResponse() wrapper applied to all successful non-streaming
responses from handleChatCore:
- Inject missing object: 'chat.completion' and created: <timestamp>
- Strip non-standard message fields (e.g. 'padding' from GitHub Copilot)
- Strip Azure content_filter_results from choices
- Strip Azure prompt_filter_results from root

This provides defense-in-depth alongside the existing fixes in
open-sse/handlers/chatCore.js, ensuring compatibility even if the
open-sse package is updated or the fixes are lost.
@xuandung38

Copy link
Copy Markdown
Contributor Author

WIP: please wait 🗡️

…dableStream lock

sanitizeJsonResponse() calls response.text() which consumes the body.
For text/event-stream responses this locks the ReadableStream, causing
"failed to pipe response" errors when Next.js tries to stream to the client.
@xuandung38
xuandung38 force-pushed the feat/gemini-embeddings-letta-fix branch from 8ad6aa2 to f258dd3 Compare February 19, 2026 13:59
@xuandung38
xuandung38 marked this pull request as ready for review February 19, 2026 14:18
@xuandung38

Copy link
Copy Markdown
Contributor Author

Ready for review and merge

@decolua

decolua commented Feb 20, 2026

Copy link
Copy Markdown
Owner

Cherry-picked commits 1, 2, 5, 6 into our fork. Thanks @xuandung38! 🙏

@decolua decolua closed this Feb 20, 2026
decolua pushed a commit that referenced this pull request Feb 20, 2026
Cherry-picked from #148 (author: xuandung38 / Hồ Xuân Dũng <me@hxd.vn>)

- Add Google AI (Gemini) embeddings support for /v1/embeddings endpoint
- Add Gemini embedding models: gemini-embedding-001, text-embedding-005, text-embedding-004
- Inject missing object/created fields for Letta and strict OpenAI clients
- Strip Azure-specific fields (prompt_filter_results, content_filter_results) from responses
- Fix Dockerfile: copy open-sse directory into Docker runner stage

Skipped: whitelist message field stripping (commit 3/7/8) — too aggressive for all providers
Skipped: default stream=false change (commit 9) — behavior change needs further review
Co-authored-by: Cursor <cursoragent@cursor.com>
kwanLeeFrmVi pushed a commit to kwanLeeFrmVi/9router that referenced this pull request Feb 26, 2026
Cherry-picked from decolua#148 (author: xuandung38 / Hồ Xuân Dũng <me@hxd.vn>)

- Add Google AI (Gemini) embeddings support for /v1/embeddings endpoint
- Add Gemini embedding models: gemini-embedding-001, text-embedding-005, text-embedding-004
- Inject missing object/created fields for Letta and strict OpenAI clients
- Strip Azure-specific fields (prompt_filter_results, content_filter_results) from responses
- Fix Dockerfile: copy open-sse directory into Docker runner stage

Skipped: whitelist message field stripping (commit 3/7/8) — too aggressive for all providers
Skipped: default stream=false change (commit 9) — behavior change needs further review
Co-authored-by: Cursor <cursoragent@cursor.com>
involvex added a commit to involvex/involvex-claude-router that referenced this pull request Feb 28, 2026
### Bug Fixes

* remove duplicated opencode prefix from model IDs and handle it in executor ([180f884](180f884))
* update opencode endpoint and model IDs ([6be8a03](6be8a03))

### Features

* add opencode provider ([03c846e](03c846e))

## [0.2.96](5645d0a...v0.2.96) (2026-02-24)

### Bug Fixes

*  GitHub Copilot model ([95fd950](95fd950))
* **auth:** allow HTTP for local network ([0a394d0](0a394d0))
* **auth:** prevent auto-login after logout ([49df3dc](49df3dc))
* **codex:** use user-agent detection for Droid CLI compatibility ([8c6e3b8](8c6e3b8))
* Correct indentation for clarity in chatCore and claude-to-openai response handlers ([fa06226](fa06226))
* correct token extraction for Claude non-streaming responses ([decolua#131](https://github.com/involvex/involvex-claude-router/issues/131)) ([9fbd6e6](9fbd6e6))
* **dashboard:** resolve 'Provider not found' for free providers ([45a4d3b](45a4d3b))
* **db:** improve error handling and null checks ([e6ef852](e6ef852))
* **gemini:** improve base64 image data parsing ([5645d0a](5645d0a))
* **github:** Implement dynamic fallback for Codex models requiring /responses endpoint ([decolua#127](https://github.com/involvex/involvex-claude-router/issues/127)) ([6913129](6913129))
* improve code formatting and reduce auto-refresh interval ([7f71916](7f71916))
* improve cursor auto-import reliability on macOS ([decolua#161](https://github.com/involvex/involvex-claude-router/issues/161)) ([d7e06c3](d7e06c3))
* **login:** avoid infinite loading on settings fetch failure ([01c9410](01c9410))
* **open-sse:** emit [DONE] in passthrough SSE mode ([decolua#142](https://github.com/involvex/involvex-claude-router/issues/142)) ([b9a6979](b9a6979))
* prevent race conditions in sticky round-robin ([3ad2f8d](3ad2f8d))
* resolve SonarQube findings and Next.js Image warnings ([7058b06](7058b06))
* update Codex executor for gpt-5.3-codex support ([d7d5dc9](d7d5dc9))

### Features

* add /v1/embeddings endpoint (OpenAI-compatible) ([decolua#146](https://github.com/involvex/involvex-claude-router/issues/146)) ([e1b8361](e1b8361)), closes [decolua#117](https://github.com/involvex/involvex-claude-router/issues/117)
* Add Anthropic Compatible provider support ([da5bdef](da5bdef))
* add API endpoint dimension to usage statistics dashboard ([decolua#152](https://github.com/involvex/involvex-claude-router/issues/152)) ([806bd4a](806bd4a))
* add Claude Opus 4.6 to GitHub Copilot provider ([decolua#97](https://github.com/involvex/involvex-claude-router/issues/97)) ([3d60597](3d60597))
* Add Claude Sonnet 4.6 to GitHub Copilot ([decolua#149](https://github.com/involvex/involvex-claude-router/issues/149)) ([4e2a3f8](4e2a3f8))
* add CLI entry points for claude-router commands ([62c632d](62c632d))
* add enable/disable toggle for provider connections ([ed796d2](ed796d2))
* add Gemini 3.1 Pro models to provider ([f2025cc](f2025cc))
* add Gemini embeddings support + Letta compatibility fixes ([a57a8ce](a57a8ce)), closes [decolua#148](decolua#148)
* add GLM 5 and MiniMax M2.5 models to providerModels.js; add Claude Sonnet 4.6 to CLI tools ([e1e5a81](e1e5a81))
* add GLM Coding (China) provider and Usage by API Keys statistics ([1ae4e31](1ae4e31))
* add GPT 4o to GitHub Copilot provider ([decolua#98](https://github.com/involvex/involvex-claude-router/issues/98)) ([c090bb0](c090bb0))
* add GPT 5.3 Codex Spark model to pricing and provider models ([decolua#133](https://github.com/involvex/involvex-claude-router/issues/133)) ([c7d4410](c7d4410))
* Add GPT 5.3 Codex to GitHub Copilot ([decolua#150](https://github.com/involvex/involvex-claude-router/issues/150)) ([c4aa424](c4aa424))
* add GPT-3.5 Turbo to GitHub Copilot provider ([e3dbd44](e3dbd44))
* add GPT-4 to GitHub Copilot provider ([6ade8ef](6ade8ef))
* add GPT-4o mini to GitHub Copilot provider ([053e490](053e490))
* add models management and router process control commands ([c65fd42](c65fd42))
* Add OpenAI-compatible provider nodes ([0a28f9f](0a28f9f))
* add password change functionality and dependencies ([23cfb19](23cfb19))
* add pause/resume functionality for API keys ([decolua#158](https://github.com/involvex/involvex-claude-router/issues/158)) ([73388a0](73388a0))
* add Qwen3.5 Coder Model configuration ([decolua#156](https://github.com/involvex/involvex-claude-router/issues/156)) ([f933dd9](f933dd9))
* add request logging functionality and usage metrics display ([e476907](e476907))
* add round-robin routing strategy ([9ebd7d3](9ebd7d3))
* add sticky round-robin routing strategy ([4f292aa](4f292aa))
* add support for GLM 5 (if) ([decolua#123](https://github.com/involvex/involvex-claude-router/issues/123)) ([03ab554](03ab554))
* add URL-based tab state persistence in usage page ([decolua#129](https://github.com/involvex/involvex-claude-router/issues/129)) ([6caef7f](6caef7f))
* allow custom user data directory via DATA_DIR environment variable ([d83bd86](d83bd86))
* **antigravity:** initial steps for Antigravity anti-ban alignment ([a229d79](a229d79)), closes [decolua#141](decolua#141)
* **antigravity:** integrate Antigravity tool with MITM support and update CLI tools ([2e854bd](2e854bd))
* **auth:** add model-level rate limit locking for multi-bucket providers ([decolua#120](https://github.com/involvex/involvex-claude-router/issues/120)) ([202fee7](202fee7)), closes [decolua#110](https://github.com/involvex/involvex-claude-router/issues/110)
* **auth:** Enhance authentication flow and settings management ([249fc28](249fc28))
* **cli-tools:** update CLI tools and add new models ([a2122e3](a2122e3))
* **cli-tools:** update default local endpoint port to 20128 ([6c41573](6c41573))
* **cli:** add new CLI package with basic scaffolding ([9cf4628](9cf4628))
* **cloud:** harden sync/auth flow, SSE fallback, and update changelog ([3d43983](3d43983))
* **codex:** add GPT 5.3, fix API translation, add thinking levels ([127475d](127475d))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([1c6dd6d](1c6dd6d))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([7b864a9](7b864a9))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([e9b0a73](e9b0a73))
* **config:** add Cloudflare MCP server and account ID configuration ([67297ea](67297ea))
* **cursor:** Add cursor Provider ([0a026c7](0a026c7))
* **cursor:** Integrate Cursor IDE support with OAuth import token flow ([137f315](137f315))
* **docker:** add Docker setup, environment examples, and architecture docs ([5e4a15b](5e4a15b))
* enhance disconnect handling and request tracking in chatCore.js ([decolua#126](https://github.com/involvex/involvex-claude-router/issues/126)) ([3d29b86](3d29b86))
* enhance request handling and error management in chatCore and streamToJsonConverter ([e2db638](e2db638))
* enhance usage stats with sortable columns and improved data handling ([bf6e09b](bf6e09b))
* Enhance usage tracking across response handlers ([a33924b](a33924b))
* **executors:**  Improved UI components for displaying provider limits and usage statistics in the dashboard. ([32aefe5](32aefe5))
* **iflow:** add IFlowExecutor with HMAC-SHA256 signature and enable models ([bd23ab4](bd23ab4))
* **iflow:** add kimi-k2.5 model support ([9e357a7](9e357a7))
* implement API key requirement toggle ([4cf25dc](4cf25dc))
* Implement buffer addition to usage tracking for improved context handling ([7881db8](7881db8))
* implement lazy loading for UsagePage with suspense fallback ([decolua#136](https://github.com/involvex/involvex-claude-router/issues/136)) ([05b09e6](05b09e6))
* implement provider connection reordering on create, update, and delete ([f2abcc6](f2abcc6))
* implement real project ID fetching for Antigravity ([decolua#170](https://github.com/involvex/involvex-claude-router/issues/170)) ([ea67742](ea67742))
* implement request tracking and enhance usage stats display ([e4f92cd](e4f92cd))
* implement usage tracking for AI requests ([9c3d6f4](9c3d6f4))
* Improve Antigravity quota monitoring and fix Droid CLI compatibility ([3c65e0c](3c65e0c))
* **open-sse:** add Claude Sonnet 4.6 ([b057c43](b057c43))
* OpenAI compatibility improvements & build fixes ([d9b8e48](d9b8e48)), closes [decolua#18](https://github.com/involvex/involvex-claude-router/issues/18)
* **provider:** add free providers and enhance error handling ([bdbe816](bdbe816))
* **providers:** add Minimax Coding (China) provider ([7c609d7](7c609d7))
* **providers:** add provider icons to dashboard ([60bd686](60bd686))
* **providers:** auto-validate API keys on save ([b275dfd](b275dfd))
* rename app to involvex-claude-router and add lenient JSON parsing ([bec206d](bec206d))
* **responses:** respect client streaming preference + string input support ([decolua#121](https://github.com/involvex/involvex-claude-router/issues/121)) ([ac7cedd](ac7cedd))
* **translator:** add thinking parameter support in OpenAI → Claude ([54e01d6](54e01d6))
* **ui:** add cost tracking to usage dashboard and pricing settings ([f302c88](f302c88))
* **ui:** add model support for custom providers and improve UX ([a7a52be](a7a52be))
* Update response handling and logging for improved usage tracking ([df0e1d6](df0e1d6))
* **usage:** implement cost tracking backend and pricing configuration ([a36afaa](a36afaa))
idbara pushed a commit to idbara/9router that referenced this pull request Apr 23, 2026
Cherry-picked from decolua#148 (author: xuandung38 / Hồ Xuân Dũng <me@hxd.vn>)

- Add Google AI (Gemini) embeddings support for /v1/embeddings endpoint
- Add Gemini embedding models: gemini-embedding-001, text-embedding-005, text-embedding-004
- Inject missing object/created fields for Letta and strict OpenAI clients
- Strip Azure-specific fields (prompt_filter_results, content_filter_results) from responses
- Fix Dockerfile: copy open-sse directory into Docker runner stage

Skipped: whitelist message field stripping (commit 3/7/8) — too aggressive for all providers
Skipped: default stream=false change (commit 9) — behavior change needs further review
Co-authored-by: Cursor <cursoragent@cursor.com>
luckystart79-lang pushed a commit to luckystart79-lang/minirouter that referenced this pull request May 2, 2026
Cherry-picked from decolua#148 (author: xuandung38 / Hồ Xuân Dũng <me@hxd.vn>)

- Add Google AI (Gemini) embeddings support for /v1/embeddings endpoint
- Add Gemini embedding models: gemini-embedding-001, text-embedding-005, text-embedding-004
- Inject missing object/created fields for Letta and strict OpenAI clients
- Strip Azure-specific fields (prompt_filter_results, content_filter_results) from responses
- Fix Dockerfile: copy open-sse directory into Docker runner stage

Skipped: whitelist message field stripping (commit 3/7/8) — too aggressive for all providers
Skipped: default stream=false change (commit 9) — behavior change needs further review
Co-authored-by: Cursor <cursoragent@cursor.com>
vibecoder11200 pushed a commit to vibecoder11200/9router that referenced this pull request Jun 13, 2026
Cherry-picked from decolua/9router#148 (author: xuandung38 / Hồ Xuân Dũng <me@hxd.vn>)

- Add Google AI (Gemini) embeddings support for /v1/embeddings endpoint
- Add Gemini embedding models: gemini-embedding-001, text-embedding-005, text-embedding-004
- Inject missing object/created fields for Letta and strict OpenAI clients
- Strip Azure-specific fields (prompt_filter_results, content_filter_results) from responses
- Fix Dockerfile: copy open-sse directory into Docker runner stage

Skipped: whitelist message field stripping (commit 3/7/8) — too aggressive for all providers
Skipped: default stream=false change (commit 9) — behavior change needs further review
Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants