fix(memory): generic embedding/rerank provider listing + runtime fallback for all configured providers - #11390
Conversation
e092520 to
6f09962
Compare
6f09962 to
542877e
Compare
54b83cb to
320cb6c
Compare
|
Re-homed to |
|
Nice mechanism — curated-entries-win derivation with credential-gated listings is the right shape, and all five of your new suites pass on the branch as pushed. Main blocker: the branch was synced against the release/v3.8.50 lineage, so the PR currently drags ~44 unrelated commits toward release/v3.8.51. Please re-derive from the current release/v3.8.51 tip and rebase your two commits there (expect small conflicts in the memory/embedding files and EngineTab). Also note our active base-red tracking moved to issue #11449 — re-check your CI exceptions against it after the rebase. |
…ed providers The memory Engine tab quick-select built its model list only from a keyword heuristic over the chat catalog plus OpenRouter live discovery, so configured providers whose embedding models are not in that catalog (e.g. Cloudflare Workers AI cf/@cf/baai/bge-m3) never appeared, and requests for them failed with 'Unknown embedding provider'. - add cloudflare-ai to EMBEDDING_PROVIDERS with a requiresAccountId URL template resolved per-request via buildEmbeddingProviderUrl() (mirrors CloudflareAIExecutor on the chat side) plus the cf provider alias - merge curated registry models into the quick-select catalog and add a generic fallback listing: any configured OpenAI-compatible chat provider without a curated entry appears with free-text model input - same treatment for rerank: listRerankProviders + generic Cohere- compatible fallback, new /api/memory/rerank-providers endpoint, selectors fall back to free-text when no static models exist - runtime: embeddings service/handler and rerank route/handler resolve derived providers instead of rejecting unlisted ones
Our PR fixed the last no-restricted-imports violation in src/lib/memory/embedding/index.ts (localDb barrel import replaced by a direct readCache import), so the frozen suppression entry for that file became stale. The gate (run-eslint-json.mjs) hard-fails on stale entries, so remove it to keep 'No new ESLint warnings' green.
bb89f24 to
0ad0764
Compare
451dd73
into
diegosouzapw:release/v3.8.51
… + api-typecheck baseline ratchet (#12414) * fix(memory): point the rerank-providers dynamic import at the real db module #11390 landed with a dynamic import of the localDb barrel, which #12052 had already removed from the base (and which Hard Rule #2 forbids) — the API Route Typecheck gate reds on the tip with TS2307. getCachedProviderNodes lives in src/lib/db/readCache. * chore(quality): ratchet the api-typecheck baseline down (163 stale entries gone) Regenerated with --update on a faithful npm ci environment (the .113 box) against the current tip plus the rerank-providers import fix — the gate now reads OK at 289 pre-existing errors, all baselined. No new entries added.
…ured provider (diegosouzapw#11390) On dashboard/memory?tab=engine the Embedding Model quick-select (and the rerank selector) built their lists from a keyword heuristic over the CHAT catalog (AI_MODELS) plus OpenRouter live discovery. Providers whose embedding models are not in that catalog never appeared — mistral, gemini, nvidia nim, groq, vercel-ai-gateway and others that serve embeddings on a standard OpenAI-compatible /embeddings endpoint — and typing such a model by hand failed at runtime with "Unknown embedding provider". The fix is one generic mechanism rather than a list of per-provider patches: deriveEmbeddingProviderForChatProvider() turns any chat-registry entry with a /chat/completions base into an OpenAI-compatible /embeddings config, with curated EMBEDDING_PROVIDERS entries always winning; the embeddings service resolves a derived config for unknown-but-configured providers instead of rejecting them; deriveRerankProviderForChatProvider() does the same for Cohere-compatible /rerank; and both memory selectors fall back to a free-text provider/model input when no static catalog exists. No provider is special-cased by name, so adding one to the chat registry now makes it embedding- and rerank-capable here automatically. Verified on the current release tip: merged clean, typecheck:core clean, check:cycles OK across 417 files, and 35/35 across the PR's five new suites (qdrant-quick-select-catalog, memory-provider-listings, rerank-provider-listings, embedding-generic-provider-fallback, rerank-generic-provider-fallback) plus the updated hard-session-lease-bypass-inventory and embeddings-handler. Note: the base-red disclaimer in the description referenced diegosouzapw#9985 against release/v3.8.50 — that window is closed and the current tip carries no open base-red, so nothing was inherited here. Thanks @rqzbeh — deriving the capability instead of enumerating providers is the version of this that stays correct as the registry grows.
… + api-typecheck baseline ratchet (diegosouzapw#12414) * fix(memory): point the rerank-providers dynamic import at the real db module diegosouzapw#11390 landed with a dynamic import of the localDb barrel, which diegosouzapw#12052 had already removed from the base (and which Hard Rule diegosouzapw#2 forbids) — the API Route Typecheck gate reds on the tip with TS2307. getCachedProviderNodes lives in src/lib/db/readCache. * chore(quality): ratchet the api-typecheck baseline down (163 stale entries gone) Regenerated with --update on a faithful npm ci environment (the .113 box) against the current tip plus the rerank-providers import fix — the gate now reads OK at 289 pre-existing errors, all baselined. No new entries added.
Problem
On
dashboard/memory?tab=engine, the Embedding Model Quick select (and the Rerank selector) built their model lists only from a keyword heuristic over the CHAT catalog (AI_MODELS) plus OpenRouter live discovery. Providers whose embedding models are not in that catalog never appeared — even though many configured providers (mistral, gemini, nvidia nim, groq, vercel-ai-gateway, ...) serve embeddings on a standard OpenAI-compatible/embeddingsendpoint — and typing such a model manually failed at runtime with "Unknown embedding provider".Fix — one generic mechanism, no per-provider patches
embeddingRegistry.ts):deriveEmbeddingProviderForChatProvider()turns ANY chat-registry entry with a/chat/completionsbase into an OpenAI-compatible/embeddingsconfig. CuratedEMBEDDING_PROVIDERSentries always win.deriveRerankProviderForChatProvider()(Cohere-compatible/rerank).listEmbeddingProviders()appends every CONFIGURED derivable provider (credential check) with an empty model catalog; new/api/memory/rerank-providersserves the rerank listing separately.provider/modelinput when no static models exist.No provider is registered or special-cased by name — adding a provider to the chat registry automatically makes it embedding/rerank-capable here.
Validation
TDD: 5 new test files + 2 focused additions:
tests/unit/qdrant-quick-select-catalog.test.ts,tests/unit/memory-provider-listings.test.ts,tests/unit/rerank-provider-listings.test.ts,tests/unit/embedding-generic-provider-fallback.test.ts,tests/unit/rerank-generic-provider-fallback.test.ts, plus cases intests/unit/embeddings-handler.test.tsResults: focused suites 32/32 pass; wider embed/memory/rerank/qdrant suites 0 fail;
npm run typecheck:coreclean; ESLint clean on all changed files (4 pre-existing unused-var warnings on untouched base lines excluded).Rebased on
release/v3.8.50.release/v3.8.50itself, not from this branch (verified byte-identical on base, absent from this diff):No new ESLint warnings: unsuppressed@typescript-eslint/no-unused-vars(formatVideoTimestamp) insrc/lib/guardrails/videoBridge.ts— file untouched here; error present on base tip with no suppression entry.Fast Quality Gates(file-size):tests/unit/guardrails/videoBridgeResultCache.test.ts1040 > cap 1000 — file untouched here, identical on base-ref and base tip.Per repo policy these belong to the freeze-gated basereds campaign, not this feature PR.