Skip to content

fix(memory): generic embedding/rerank provider listing + runtime fallback for all configured providers - #11390

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.51from
rqzbeh:fix/cf-embedding-models
Sep 2, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.51from
rqzbeh:fix/cf-embedding-models

Conversation

@rqzbeh

@rqzbeh rqzbeh commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Problem

On dashboard/memory?tab=engine, the Embedding Model Quick select (and the Rerank selector) built their model lists only from a keyword heuristic over the CHAT catalog (AI_MODELS) plus OpenRouter live discovery. Providers whose embedding models are not in that catalog never appeared — even though many configured providers (mistral, gemini, nvidia nim, groq, vercel-ai-gateway, ...) serve embeddings on a standard OpenAI-compatible /embeddings endpoint — and typing such a model manually failed at runtime with "Unknown embedding provider".

Fix — one generic mechanism, no per-provider patches

  • Registry derivation (embeddingRegistry.ts): deriveEmbeddingProviderForChatProvider() turns ANY chat-registry entry with a /chat/completions base into an OpenAI-compatible /embeddings config. Curated EMBEDDING_PROVIDERS entries always win.
  • Runtime fallback: embeddings service resolves a derived config for unknown-but-configured providers instead of rejecting them; same for rerank via deriveRerankProviderForChatProvider() (Cohere-compatible /rerank).
  • Listings merge: Quick-select catalog merges curated registry models; listEmbeddingProviders() appends every CONFIGURED derivable provider (credential check) with an empty model catalog; new /api/memory/rerank-providers serves the rerank listing separately.
  • UI: both memory selectors fall back to a free-text provider/model input when no static models exist.

No provider is registered or special-cased by name — adding a provider to the chat registry automatically makes it embedding/rerank-capable here.

Validation

TDD: 5 new test files + 2 focused additions:

  • tests/unit/qdrant-quick-select-catalog.test.ts, tests/unit/memory-provider-listings.test.ts, tests/unit/rerank-provider-listings.test.ts, tests/unit/embedding-generic-provider-fallback.test.ts, tests/unit/rerank-generic-provider-fallback.test.ts, plus cases in tests/unit/embeddings-handler.test.ts

Results: focused suites 32/32 pass; wider embed/memory/rerank/qdrant suites 0 fail; npm run typecheck:core clean; ESLint clean on all changed files (4 pre-existing unused-var warnings on untouched base lines excluded).

Rebased on release/v3.8.50.


⚠️ base-red inherited: #9985 — two remaining check failures are inherited from release/v3.8.50 itself, not from this branch (verified byte-identical on base, absent from this diff):

  • No new ESLint warnings: unsuppressed @typescript-eslint/no-unused-vars (formatVideoTimestamp) in src/lib/guardrails/videoBridge.ts — file untouched here; error present on base tip with no suppression entry.
  • Fast Quality Gates (file-size): tests/unit/guardrails/videoBridgeResultCache.test.ts 1040 > cap 1000 — file untouched here, identical on base-ref and base tip.
    Per repo policy these belong to the freeze-gated basereds campaign, not this feature PR.

@rqzbeh
rqzbeh requested a review from diegosouzapw as a code owner August 24, 2026 12:53
@rqzbeh
rqzbeh marked this pull request as draft August 24, 2026 13:01
@rqzbeh
rqzbeh force-pushed the fix/cf-embedding-models branch from e092520 to 6f09962 Compare August 24, 2026 13:19
@rqzbeh
rqzbeh changed the base branch from release/v3.8.51 to release/v3.8.50 August 24, 2026 13:20
@rqzbeh rqzbeh changed the title fix(memory): list and serve embedding/rerank models from all configured providers fix(memory): generic embedding/rerank provider listing + runtime fallback for all configured providers Aug 24, 2026
@rqzbeh
rqzbeh force-pushed the fix/cf-embedding-models branch from 6f09962 to 542877e Compare August 24, 2026 13:23
@rqzbeh
rqzbeh marked this pull request as ready for review August 24, 2026 13:39
@rqzbeh
rqzbeh force-pushed the fix/cf-embedding-models branch from 54b83cb to 320cb6c Compare August 24, 2026 14:03
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.50 to release/v3.8.51 August 24, 2026 23:01
@diegosouzapw

Copy link
Copy Markdown
Owner

Re-homed to release/v3.8.51: v3.8.50 entered its release freeze, so the branch now belongs to the release captain and development continues on the next cycle. Nothing is wrong with this PR — it just needed a live base. No action needed from you; CI will re-run against the new base.

@diegosouzapw

Copy link
Copy Markdown
Owner

Nice mechanism — curated-entries-win derivation with credential-gated listings is the right shape, and all five of your new suites pass on the branch as pushed. Main blocker: the branch was synced against the release/v3.8.50 lineage, so the PR currently drags ~44 unrelated commits toward release/v3.8.51. Please re-derive from the current release/v3.8.51 tip and rebase your two commits there (expect small conflicts in the memory/embedding files and EngineTab). Also note our active base-red tracking moved to issue #11449 — re-check your CI exceptions against it after the rebase.

rqzbeh added 2 commits August 29, 2026 01:48
…ed providers

The memory Engine tab quick-select built its model list only from a
keyword heuristic over the chat catalog plus OpenRouter live discovery,
so configured providers whose embedding models are not in that catalog
(e.g. Cloudflare Workers AI cf/@cf/baai/bge-m3) never appeared, and
requests for them failed with 'Unknown embedding provider'.

- add cloudflare-ai to EMBEDDING_PROVIDERS with a requiresAccountId URL
  template resolved per-request via buildEmbeddingProviderUrl() (mirrors
  CloudflareAIExecutor on the chat side) plus the cf provider alias
- merge curated registry models into the quick-select catalog and add a
  generic fallback listing: any configured OpenAI-compatible chat
  provider without a curated entry appears with free-text model input
- same treatment for rerank: listRerankProviders + generic Cohere-
  compatible fallback, new /api/memory/rerank-providers endpoint,
  selectors fall back to free-text when no static models exist
- runtime: embeddings service/handler and rerank route/handler resolve
  derived providers instead of rejecting unlisted ones
Our PR fixed the last no-restricted-imports violation in
src/lib/memory/embedding/index.ts (localDb barrel import replaced by a
direct readCache import), so the frozen suppression entry for that file
became stale. The gate (run-eslint-json.mjs) hard-fails on stale entries,
so remove it to keep 'No new ESLint warnings' green.
@diegosouzapw
diegosouzapw force-pushed the fix/cf-embedding-models branch from bb89f24 to 0ad0764 Compare August 29, 2026 05:08
@diegosouzapw
diegosouzapw merged commit 451dd73 into diegosouzapw:release/v3.8.51 Sep 2, 2026
14 of 16 checks passed
diegosouzapw added a commit that referenced this pull request Sep 2, 2026
… + api-typecheck baseline ratchet (#12414)

* fix(memory): point the rerank-providers dynamic import at the real db module

#11390 landed with a dynamic import of the localDb barrel, which #12052 had
already removed from the base (and which Hard Rule #2 forbids) — the API
Route Typecheck gate reds on the tip with TS2307. getCachedProviderNodes
lives in src/lib/db/readCache.

* chore(quality): ratchet the api-typecheck baseline down (163 stale entries gone)

Regenerated with --update on a faithful npm ci environment (the .113 box)
against the current tip plus the rerank-providers import fix — the gate now
reads OK at 289 pre-existing errors, all baselined. No new entries added.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ured provider (diegosouzapw#11390)

On dashboard/memory?tab=engine the Embedding Model quick-select (and the rerank selector) built their lists from a keyword heuristic over the CHAT catalog (AI_MODELS) plus OpenRouter live discovery. Providers whose embedding models are not in that catalog never appeared — mistral, gemini, nvidia nim, groq, vercel-ai-gateway and others that serve embeddings on a standard OpenAI-compatible /embeddings endpoint — and typing such a model by hand failed at runtime with "Unknown embedding provider".

The fix is one generic mechanism rather than a list of per-provider patches: deriveEmbeddingProviderForChatProvider() turns any chat-registry entry with a /chat/completions base into an OpenAI-compatible /embeddings config, with curated EMBEDDING_PROVIDERS entries always winning; the embeddings service resolves a derived config for unknown-but-configured providers instead of rejecting them; deriveRerankProviderForChatProvider() does the same for Cohere-compatible /rerank; and both memory selectors fall back to a free-text provider/model input when no static catalog exists. No provider is special-cased by name, so adding one to the chat registry now makes it embedding- and rerank-capable here automatically.

Verified on the current release tip: merged clean, typecheck:core clean, check:cycles OK across 417 files, and 35/35 across the PR's five new suites (qdrant-quick-select-catalog, memory-provider-listings, rerank-provider-listings, embedding-generic-provider-fallback, rerank-generic-provider-fallback) plus the updated hard-session-lease-bypass-inventory and embeddings-handler.

Note: the base-red disclaimer in the description referenced diegosouzapw#9985 against release/v3.8.50 — that window is closed and the current tip carries no open base-red, so nothing was inherited here.

Thanks @rqzbeh — deriving the capability instead of enumerating providers is the version of this that stays correct as the registry grows.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… + api-typecheck baseline ratchet (diegosouzapw#12414)

* fix(memory): point the rerank-providers dynamic import at the real db module

diegosouzapw#11390 landed with a dynamic import of the localDb barrel, which diegosouzapw#12052 had
already removed from the base (and which Hard Rule diegosouzapw#2 forbids) — the API
Route Typecheck gate reds on the tip with TS2307. getCachedProviderNodes
lives in src/lib/db/readCache.

* chore(quality): ratchet the api-typecheck baseline down (163 stale entries gone)

Regenerated with --update on a faithful npm ci environment (the .113 box)
against the current tip plus the rerank-providers import fix — the gate now
reads OK at 289 pre-existing errors, all baselined. No new entries added.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants