feat(cache): configurable dual-layer semantic caching with Redis/In-Memory vector stores (re-land of #12630) - #14159
Merged
Conversation
…#1) * feat(cache): implement dual-layer semantic caching with in-memory and redis vector stores Implements production-grade, configurable semantic caching for OmniRoute. - Dual-layer architecture: Layer 1 exact hash match (0 embedding latency) + Layer 2 vector cosine similarity search. - Backends: in-memory vector store with L2 normalization, LRU and TTL + Redis vector store adapter with fail-open fallback. - Embedding generation: conversation history normalization, system prompt exclusion, timeout protection. - Streaming support: serializable SSE stream synthesis ending in data: [DONE]\n\n. - Request overrides and telemetry headers: X-OmniRoute-Cache (HIT (exact) | HIT (semantic) | MISS), X-OmniRoute-Cache-Similarity, X-OmniRoute-Savings-Tokens, Cache-Control: no-cache, x-omniroute-no-cache, x-omniroute-cache-threshold, x-omniroute-cache-type, x-omniroute-cache-no-store, x-omniroute-cache-key. - Unit test coverage across dual-layer search, eviction, redis resilience, and streaming replay. * feat(cache): add default embedding client and live integration test for Lemonade server and Redis - Add embeddingBaseUrl and embeddingApiKey configuration options to SemanticCacheConfig. - Implement createDefaultEmbeddingGenerator for automatic OpenAI-compatible embedding integration. - Add live verification script scripts/ad-hoc/test-semantic-cache-lemonade.ts. - Add network-aware integration test tests/integration/semantic-cache-lemonade.test.ts for Lemonade harrier-oss-v1-0.6b and Redis vector store. * fix(cache): address CodeRabbit review recommendations on PR #1 - Multi-tenant partition isolation: support null sentinel in StoreFilter so anonymous requests cannot match authenticated entries. - Provider propagation: pass routed provider to non-streaming and streaming cache writes. - Embedding resilience: race generator with timeout promise in generateEmbeddingWithTimeout to guard against uncooperative generators. - Memory store consistency: replace older entries with identical hash on insert, and safe-guard hashToId deletion in removeEntry. - Redis store consistency: prune expired/missing entries from candidate sets during search/stats, and delete hash mapping conditionally. - Token telemetry: use managerResult.tokensSaved when cached response lacks usage data. - Clova batch timeout: add AbortSignal.timeout(FETCH_TIMEOUT_MS) to fetchClovaEmbeddingBatch. --------- Co-authored-by: John Smith <you@example.com>
… and sync Redis hit counters
PR #12630 (feat(cache): configurable dual-layer semantic caching) was opened 2026-09-03 with head at 1cd8897 and is now mergeable: CONFLICTING because upstream's release/v3.8.51 advanced 96 commits to ba597b6. This merge brings semantic-cache up to date so the PR becomes mergeable again. Three conflicts resolved: - open-sse/handlers/chatCore.ts: took upstream. Upstream's restructure preserves storeSemanticCacheResponse at line 5365 and saveIdempotency at line 5379; the PR's Phase 9.1/9.2 block at lines 5528-5543 is textually equivalent and superseded by the upstream version. - src/app/api/settings/cache-config/route.ts: combined. Took upstream's restructured file as base (empty-check guard, alwaysPreserveClientCache routed through flat settings — keeping upstream commit 8a95a2b's fix for the runtime no-op bug), then restored the PR's dual-layer cache content on top: resetSemanticCacheManager call, ensureSemanticCacheDbBridge() module-level side-effect, getEmbeddingOptions enrichment in GET, the 10 new cache schema fields (semanticCacheBackend, semanticCacheThreshold, semanticCacheEmbeddingProvider/Model/Dimension/BaseUrl/ApiKey, semanticCacheRedisUrl/Prefix, semanticCacheRequireZeroTemp), and matching CACHE_CONFIG_KEYS + DEFAULTS entries. alwaysPreserveClientCache stays in the flat-settings path per upstream's fix. - src/app/(dashboard)/dashboard/combos/page.tsx: took upstream. The useState+useEffect localStorage hydration pattern is replaced by useSyncExternalStore (no commit-after-mount, no hydration mismatch warning). Skipped from this PR: the bf8691a lint cleanup for tests/unit/volcengine-plan-binding-upsert.test.ts (upstream's file, not touched by the PR) will go in as a separate upstream PR. Forward-merge base-red inherited: #12732 (upstream's tip is red).
…-cache Merges the dual-layer semantic cache branch onto the current release tip so CI can run. Resolves two content conflicts by combining both sides rather than picking one: - open-sse/handlers/chatCore/semanticCache.ts: keeps the new dual-layer lookup (semanticCacheManager) with legacy SQLite fallback, and ports the release's #12309/#12734 fix (tool_choice/tools/ response_format folded into the cache signature) into the legacy fallback path and its matching hit-recording signature so both stay in sync. - src/lib/api/modelTestRunner.ts: keeps this branch's added embedding keyword detection (colbert/harrier-/nomic-embed) together with the release's isResponses/isNonChatGeneration routing added since. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…ped by this PR The dual-layer semantic cache branch's "catch up to upstream" merge commit accidentally dropped the react-hooks/immutability suppression for 5 UI test files this PR never touches (use-improve-prompt, use-presets, use-stream-metrics, use-structured-output, use-tools-builder). Verified with eslint that they still trip the rule and restored the entries; the other removed entries (tlsClientBase.ts, check-licenses.test.ts, combo-routing-engine.test.ts, src/lib/semanticCache.ts) were confirmed clean and left as-is. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Ports the single-layer cache's #12910 fix into the new dual-layer checkSemanticCache: finalize the exact pending request by id (finalizePendingScope) instead of an ambiguous (model, provider, connectionId) tuple, which could finalize the wrong in-flight request when connectionId is null or multiple requests share a connection. Co-authored-by: Jihyun Son <initguru@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…yer cache Ports the single-layer cache's #12885 fix into both write paths of the new dual-layer architecture: a response cut short by the output token ceiling (finish_reason "length"/"max_tokens") is a partial answer, not a reusable one, so it must never be written to the semantic cache. Caching it under a temperature:0 signature pinned the truncation for every later identical request. - src/lib/semanticCache.ts: new isTruncatedCompletion() predicate. - semanticCacheStore.ts / streamingSemanticCacheStore.ts: gate both the legacy setCachedResponse() write and the new dual-layer manager.store() write on it. Co-authored-by: Patryk Kopyciński <contact@patrykkopycinski.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…c-sync - open-sse/services/cache/redisVectorStore.ts: RedisLike was missing the optional disconnect() method that close() already calls as an ioredis fallback when quit() isn't available — 2 TS2339 errors under check:open-sse-typecheck. - scripts/check/check-env-doc-sync.mjs: allowlist LEMONADE_URL/KEY/MODEL, same class as the existing NVIDIA_BASE_URL/MODEL entry — they only configure the gated tests/integration/semantic-cache-lemonade.test.ts live test against an operator's local Lemonade server (self-skips when unreachable), never OmniRoute runtime config. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Two tests in tests/unit/model-embedding-discovery-and-cache.test.ts made a real HTTP call to the contributor's private Lemonade server (192.168.31.147) with no reachability guard — they failed with a 60s socket timeout on any machine without access to that LAN, including this one. tests/integration/semantic-cache-lemonade.test.ts already gates the same endpoint with an isEndpointReachable() skip; apply the identical pattern here so a unit test never depends on live network access. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…-cache re-land Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> # Conflicts: # changelog.d/fixes/12910-semantic-cache-exact-id-finalize.md # open-sse/handlers/chatCore/semanticCache.ts # open-sse/handlers/chatCore/streamingSemanticCacheStore.ts # src/app/(dashboard)/dashboard/providers/[id]/hooks/useModelImportHandlers.ts # src/lib/providerModels/modelDiscovery.ts # tests/unit/chatcore-semantic-cache.test.ts # tests/unit/semantic-cache-no-truncated-writes.test.ts
…ery contracts (#14159) The re-land of #12630 validated only its own tests and broke 15 existing ones. Three root causes, all fixed in production code so the tip's tests are untouched: 1. Dual-layer manager was ON by default. `DEFAULT_SEMANTIC_CACHE_CONFIG.enabled` was `true` and the DB bridge wired it to the pre-existing `semanticCacheEnabled` toggle (also default `true`), so every temperature=0 request called an embedding endpoint (lemonade @ localhost:13305 by default) on lookup AND store — two extra fetches per chat call (chat-combo-live-test, issue-agent-route-execution). Fix: new `semanticCacheVectorEnabled` setting (default false; types, DB keys, cache-config route, settings UI sub-toggle) and the bridge only enables the manager when BOTH toggles are on. Manager default is now `enabled: false`. With the opt-in off, chatCore behaves exactly like the legacy SQLite cache. 2. `X-OmniRoute-Cache` value changed from `HIT` to `HIT (exact)`/`HIT (semantic)`, and `cacheSource: "semantic_similarity"` fell through attemptLogging's `"semantic" | "upstream"` narrowing as "upstream". Fix: keep `HIT` verbatim (similarity hits are distinguished by X-OmniRoute-Cache-Similarity) and log both hit types as `cacheSource: "semantic"`. The PR's edits to chatcore-semantic-cache.test.ts are reverted to the tip version. 3. Model discovery stamped `modelType: "chat"` + `supportedInputTypes: ["text"]` on every model (import-mode diff noise, 6 catalog tests), and endpoint-based modality inference tagged a `["chat","embeddings"]` model as embedding-only (`apiFormat: "embeddings"`), dropping it from the chat catalog. Fix: an explicit chat endpoint vetoes the embedding/rerank/image heuristics; chat models keep the tip's row shape (metadata only stamped for non-chat modalities). Also: - Restore the tip's `isTruncatedStreamBody` dep in streamingSemanticCacheStore and widen the predicate to the assembled-object body chatCore actually passes (the PR's object-shaped streaming cases are appended to the tip's test file). - Pass `provider` from chatCore to both cache store paths — the manager scopes entries per provider on lookup, so writes without it could never hit. - test-embedding route: `validateBody()` has no `.response`; invalid payloads returned `undefined` (TS2339 in api-typecheck). Now a proper 400. - New guard: tests/unit/semantic-cache-vector-layer-opt-in.test.ts.
diegosouzapw
added a commit
that referenced
this pull request
Sep 21, 2026
…and #13874 The hard-lease bypass inventory froze every getProviderConnections / getProviderConnectionById site with a class; two landed on the tip without a golden update: - src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of #12630): read-only listing that feeds the semantic-cache embedding dropdown, same shape as the qdrant embedding-models route — class C. - src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an unrecoverable refresh error to detect credentials rotated by a concurrent Layer 2 refresh before deactivating — a state read, not dispatch; stays C. Refs #13866
This was referenced Sep 21, 2026
diegosouzapw
added a commit
that referenced
this pull request
Sep 21, 2026
Three squash merges on release/v3.8.51 lost the contributor attribution that the pipeline had preserved on the branches: - 7a92129 (#14159) re-landed #12630: the dual-layer semantic cache is @BillyOutlast's work; the squash author is the maintainer. - cd4c6f6 (#14162) re-landed #13180: the native Codex auto-resume fix is Felipe Reis's (@mdigitalbh81); the squash author is the maintainer. - 53a147c (#13295) shipped the same one-line resolvedExtensionEnd reorder that @ggiak landed first in #13036; the squash carried no trailer. The changelog fragment and these trailers record that credit in the branch history, the contributors graph and the release notes. Co-authored-by: BillyOutlast <172061051+BillyOutlast@users.noreply.github.com> Co-authored-by: Felipe Reis <mdigitalbh81@gmail.com> Co-authored-by: Giorgos Giakoumettis <giorgos@yiakoumettis.gr> Co-authored-by: ggiak <20743694+ggiak@users.noreply.github.com>
diegosouzapw
added a commit
that referenced
this pull request
Sep 21, 2026
…-09-21) Third reconciliation pass of the living [3.8.51] section against release/v3.8.50..release/v3.8.51 (0915890 → 06f1df9): 364 fragments folded, 104 bullets generated for commits without a fragment, 254 fragment bullets linked and credited by origin commit, 1,393 bullets total, 226 external contributors (0 missing on cross-check). Hand-corrected credits: #12885 → @patrykkopycinski, #14159 → @BillyOutlast (feature bullet), #12972 → @IAMBOBJIM.
Merged
diegosouzapw
added a commit
that referenced
this pull request
Sep 22, 2026
…-09-21) (#14371) Third reconciliation pass of the living [3.8.51] section against release/v3.8.50..release/v3.8.51 (0915890 → 06f1df9): 364 fragments folded, 104 bullets generated for commits without a fragment, 254 fragment bullets linked and credited by origin commit, 1,393 bullets total, 226 external contributors (0 missing on cross-check). Hand-corrected credits: #12885 → @patrykkopycinski, #14159 → @BillyOutlast (feature bullet), #12972 → @IAMBOBJIM.
diegosouzapw
added a commit
that referenced
this pull request
Sep 22, 2026
…ll, TS2677, ESLint (Refs #13866) (#14331) * fix(build): ship httpClientAbortGuard.mjs in the pack artifact; validate input_tokens with Zod Wave five of the release/v3.8.51 base-reds, part 1 — the two that matter. #14064 restored server-ws.mjs's import of ./httpClientAbortGuard.mjs and the assembleStandalone copy, but not the two pack-artifact policy entries that were lost with it. Without APP_STAGING_ALLOWED_EXACT_PATHS the prepublish prune deletes the file; without PACK_ARTIFACT_REQUIRED_PATHS nothing notices. Every boot of the published package would die with ERR_MODULE_NOT_FOUND — the 3.8.47 head-response-guard class. Both closure suites (9/9) now enforce it. #13910's /v1/responses/input_tokens read request.json() behind a hand-rolled typeof check. Hard Rule #7 wants the boundary on Zod; the t06 guard caught it. Same passthrough envelope the catch-all Responses route uses, since the counters below already walk the fields defensively. 9/9 on the route's suite. Five no-unused-vars left behind by the wave (cliRuntime execFileSync, arena test symbols and a type, compression rmSync, waitForServer req) are removed. The 'openwa routes removed without deprecation' entry from the #14101 run was an artifact of that PR trailing its base — the gate is clean on the tip. Refs #13866 * fix(compression): let anchored file-pack rules see the transformed text; align wave-5 guards Wave five of the release/v3.8.51 base-reds, part 2. One production defect. #12825 (Hungarian Caveman pack) stopped gating file-pack rules with the English keyword list and tested the rule's own regex instead — against `lowerResult`, a lower-cased copy of the ORIGINAL text that the loop never refreshed. An anchored pattern like leader_phrases' `^(?:i will|…)` therefore ran its prefilter on "sure, i will…", failed the anchor, and was skipped; the rule that strips "I will " from every English response was dead since the merge. The prefilter now sees the text as the rules so far have left it. New test fails on the tip and passes here; all Caveman suites, Hungarian included, are 95/95. A frozen no-unused-vars suppression on caveman.ts no longer had a target and is pruned. Two more TS2677 predicates of the kind #14101 fixed: #13910 (rerankProviderNodes.ts, `n is RerankProviderNodeRow` on a Record row) and #13957's mitm catalog (antigravity.ts, `c is DynamicCatalogModel` on a literal-or-null). Both narrow by NonNullable of the element's own type; the api-route typecheck was 285 against a baseline of 283 on the pristine tip. The rest are guards trailing legitimate changes: - #12663 made gemini-3.8-flash the catalog head; T28 pinned 3.7. - #13863 put mimo-v2.5 into the shared vision heuristic on purpose (the base model is multimodal, only the Pro variants are text-only). The safety test now asserts the real invariant: base and :free aliases yes, -pro no. - #12565 moved npm-prefix detection into cliRuntimeNpmPrefix.ts with a process-lifetime cache that importFresh() does not reset; the case resets it. #12565 also builds Windows candidates with path.win32 on purpose; the qodercli test compared against POSIX path.join. - #13990 (the 2 GB Docker image) copies better-sqlite3 with --chown; the guard matched the flag order literally. Now flag-order tolerant, still fails when --from=builder is removed. - #13378 reintroduced public/openference.svg under a name #11750 retired for missing provenance and swapped the Cerebras showcase cell for it. The cell is back and the asset is gone; whether the new drawing counts as provenance is the owner's call. Refs #13866 * fix(i18n): translate the 7 sidebar-pin and Claude low-priority keys into all 65 locales #7f1b4a5e (sidebar pinned items) and #1b2349de (Claude OAuth lower-priority / auto-reset) landed with their 7 new keys in en.json only, which the vi and pt-BR parity suites flag. Translated with the repo's own sync-ui-keys --translate-markers against the .113 i18n instance (codex/gpt-5.6-sol-low): +446 lines across 65 catalogs, zero __MISSING__ markers, placeholders intact. vi.json also has two keys reordered to mirror en.json; values unchanged. Refs #13866 * chore(quality): list the 8 covering tests the sixth wave added in stryker tap.testFiles 30 commits landed on release/v3.8.51 while wave five drained; eight new unit tests cover mutated modules and were not in tap.testFiles, so their mutant kills did not count and check:mutation-test-coverage --strict failed on the merged tree. Appended at the end of the list, nothing reordered. Refs #13866 * docs: document the five env vars of the 09-18 wave; regenerate the version-manager skill for the open-wa routes check:docs-all: BRIDGE_PORT, ROUTER_URL and CERT_DIR (bin/antigravity-bridge.mjs, #c74cea3d), OPENWA_SERVICE_PORT (src/lib/services/bootstrap.ts, #1e8c913c) and NEXT_PUBLIC_PORT (src/shared/hooks/useDisplayBaseUrl.ts, #d715190b) were read in code but absent from .env.example and docs/reference/ENVIRONMENT.md. Added next to their neighbours, with the defaults the code actually uses (open-wa is 8323, not the 201xx range the other services sit in). check:agent-skills-sync: the open-wa feature added eight /api/services/openwa/* routes to docs/openapi.yaml without regenerating skills/omni-version-manager/ SKILL.md. Regenerated with the repo generator; the diff is exactly those eight route sections. Refs #13866 * test: register the crash guard in the pack snapshot; inventory #13874's refresh-lane row read pack-artifact-policy pins the list of root runtime files check:pack-artifact must find in the tarball; dist/httpClientAbortGuard.mjs joined PACK_ARTIFACT_REQUIRED_PATHS in this PR and the snapshot follows. #13874 re-reads the connection row inside the Claude refresh lane so a queued health check does not POST a refresh token a Layer 2 refresh already rotated — a state read, inventoried like the family-cooldown lookup (tokenHealthCheck.ts 2 -> 3). Refs #13866 * chore(quality): list native-codex-auto-resume test in stryker tap.testFiles (#13180 landed without it) * fix(release): drain the seventh base-red wave of release/v3.8.51 (9 tests + pack-policy + dashboard-typecheck) Three production defects the tests caught: - rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's queue-budget gate as "0 ms left" and 503'd every protected request. - emergencyFallback: #14006 silently switched the budget-exhaustion target provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored. - claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only by casing; helpers renamed to claudeConnectionFieldValues.ts. Guards realigned to legitimate changes: #13874 rotation map (distinct token in the error test), #13350 origin-IP denylist, #13318 shared-catalog growth (counts by invariant), comboTargetKeyPolicy import in the telegram stub, the 22 README mirrors that #13940/#14106 stamped with the retired openference.svg (translated Cerebras cells recovered from history, hashes re-stamped), bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard typecheck regressions (typed pinned section, ComponentProps cast). Refs #13866. * test: type the #13848 Gemini pairing tests (no-explicit-any) and inventory the semantic-cache embedding picker's connection read Both arrived with the tip merge: #13848 added 13 explicit any casts to translator-openai-to-gemini.test.ts (no-explicit-any is an error under tests/), and 7a92129's embeddingOptions.ts reads provider connections once without a hard-session-lease inventory entry. Stale suppression count pruned for the test file only. Refs #13866. * test: split the #13848 turn-pairing cases out of translator-openai-to-gemini.test.ts The file sits exactly at its frozen size cap; typing the pairing tests (no-explicit-any) pushed it 14 lines over. The two cases are a coherent regression suite of their own, so they move to translator-openai-to-gemini-turn-pairing-13848.test.ts (registered in stryker tap.testFiles) instead of widening the baseline. * docs(env): document BRIDGE_PORT, ROUTER_URL, CERT_DIR, OPENWA_SERVICE_PORT and NEXT_PUBLIC_PORT (Refs #13866) check:env-doc-sync has been red on the release tip since these five vars reached code without their .env.example / ENVIRONMENT.md entries: bin/antigravity-bridge.mjs (BRIDGE_PORT, ROUTER_URL, CERT_DIR — #14006), src/lib/services/bootstrap.ts + api/services/openwa/_lib.ts (OPENWA_SERVICE_PORT) and src/shared/hooks/useDisplayBaseUrl.ts (NEXT_PUBLIC_PORT — #13533). Defaults and source files copied from the reads themselves. * chore(skills): regenerate omni-version-manager for the open-wa service routes (Refs #13866) check:agent-skills-sync (Merge integrity job) has been red on the tip since the open-wa embedded-service routes reached docs/openapi.yaml without the generated SKILL.md being refreshed. Output of scripts/skills/generate-agent-skills.mjs --apply, no hand edits: the eight /api/services/openwa/* operations. * fix(types): make the two TS2677 type predicates sound (Refs #13866) check:api-typecheck has been red on the tip with two "type predicate's type must be assignable to its parameter's type" errors: - src/app/api/v1/_shared/rerankProviderNodes.ts (#13733): the read cache hands back `Record<string, unknown> | null`, and an interface whose members are all optional is not assignable to an index-signature type. Narrow to the non-null record and assert the row shape afterwards. - src/mitm/handlers/antigravity.ts (#14006): the map callback returned `{ displayName: string }` while DynamicCatalogModel declares it optional, so the predicate could not be proven. Type the callback's return explicitly and filter on `!== null`. No runtime change; rerank-remote-provider-nodes / rerank-local-node-shapes / mitm-handler-antigravity stay green. * fix(lint): clear the 92 ESLint errors the lint gate reports on the tip (Refs #13866) - tests/unit/translator-openai-to-gemini.test.ts: #13848 / #13318 added 13 `any` casts/params on top of the 74 frozen for the file, so ESLint reported all 87. Typed them (GeminiRequestWithContents / GeminiToolPart, and the existing GeminiRequestWithConfig) and pruned the file's suppression to the new count of 71 — nothing else in eslint-suppressions.json changes. - no-unused-vars: execFileSync import (src/shared/services/cliRuntime.ts, #12565), getArenaEloSyncStatus + makeLeaderboardMap + ArenaLeaderboardMap (tests/unit/arena-elo-sync-redesign.test.ts, #13446), rmSync (compressionAnalyticsWriterFlatRate.test.ts, #13446), `req` → `_req` (waitForServer-slow-first-response.test.mjs). translator-openai-to-gemini 48/48; arena-elo-sync-redesign, compressionAnalyticsWriterFlatRate, waitForServer-slow-first-response green. * fix(compression): stop skipping anchored Caveman rules that only match after earlier rules #12825 (Hungarian pack) replaced the English keyword prefilter with a `rule.pattern.test(lowerText)` pre-check for every file-based rule, including the default `en` pack. `lowerText` is the ORIGINAL message, so anchored rules such as `leader_phrases` (`^i will …`) — which only match after `pleasantries` strips "Sure, " — were dropped before they could run. `caveman-v379` caught the regression ("I will ensure …" survived at full intensity). Tag file-based rules with their pack language in ruleLoader and let the keyword prefilter apply to `en`/built-in rules only; non-English packs (which reuse English rule names) simply run their localized regex, which is what the pre-test cost anyway. Drops the now-unused CAVEMAN_RULES import and prunes the already-stale `caveman.ts` no-unused-vars suppression (0 violations on the tip) that blocked the pre-commit hook for any change to this file. Refs #13866 * test(models): align catalog and vision-heuristic guards with the tip's intended contracts Three base-reds where the production change was deliberate and the pinned guard was simply not bumped by the PR that changed the contract: - agy-antigravity-shared-catalog-12724: #13318 added the three Gemini 3.8 Flash tiers (high/medium/low, no "-tiered" endpoint for 3.8) to the shared Antigravity/AGY base, 10 -> 13. Pin the new size in one constant and make the buildSurfaceCatalog delta assertions relative to it. - t28-model-catalog-updates: #12663 (issue #12638) registered gemini-3.8-flash at the head of the AI Studio fallback catalog as the current Flash default; assert 3.8 first and keep 3.7 present. - command-code-mimo-v2-5-safety: #13863 (issue #13847) added an explicit "mimo-v2.5" fragment to the shared vision heuristic so provider-qualified and `-free` aliases keep their vision flag. The guard's real concern (the "mimo-vl" fragment must not cover "mimo-v2.5") is asserted on the fragment itself; the bare id is now vision by heuristic on purpose, and the Pro text-only sibling stays excluded. Refs #13866 * test(cli): follow the #12565 cliRuntime module split in the npm-prefix and qodercli guards #12565 (issue #12563) moved the npm global-prefix cache out of cliRuntime.ts into cliRuntimeNpmPrefix.ts and built the Windows known-bin candidates with `path.win32` (cliRuntimeWindowsNode.ts) so they stay Windows-shaped when `process.platform` is mocked on a POSIX runner. Two pre-existing guards depended on the old layout: - cli-runtime-extended "resolves known binaries from npm global prefix": importFresh() only re-evaluates cliRuntime.ts; the prefix cache now lives in a module that stays shared across cases, so a real `npm config get prefix` from an earlier case was cached and the mocked execFileSync never ran. Reset the cache with the helper #12565 exported for exactly this in afterEach. - qodercli-windows-resolve-6263: compare against `path.win32.join` — identical to `path.join` on a real Windows host, which is the behaviour under test. Production behaviour is unchanged on both platforms. Refs #13866 * test(auto-update): write the source-mode log inside the test's own temp dir The launchAutoUpdate case pointed AUTO_UPDATE_LOG_PATH at a fixed, world-shared `/tmp/auto-update-source.log`. On the .113 runner the suite executes both as `root` and as `runner` (uid 1001): the file survives owned by whoever ran first (`-rw-r--r-- root root`), and the next `openSync(logPath, "a")` fails with EACCES for the other user. Reproduced locally by making the shared file read-only; production code is untouched (autoUpdate.ts last changed in #9354). Use a per-test mkdtemp path for the source-mode log and clean the whole temp root in the existing finally block. Refs #13866 * fix(dashboard): rename claudeConnectionFields.ts so it no longer case-collides with ClaudeConnectionFields.tsx #13074 added two modules to the provider-detail modals directory whose names differ only by casing: `ClaudeConnectionFields.tsx` (the component) and `claudeConnectionFields.ts` (the value/patch helpers). On a case-insensitive filesystem the pair breaks the webpack build (#6584 guard), and esbuild's resolver already picks the `.tsx` for the extension-less `./claudeConnectionFields` specifier, so the provider-detail client entry failed to bundle ("No matching export ... for import claudeConnectionFieldPatch"). Rename the helper module to `claudeConnectionFieldValues.ts` (the same naming the sibling `quotaScrapingFieldValues.ts` uses) and point the only importer, EditConnectionModal.tsx, at the new name. Greens tests/unit/case-collision-6584.test.ts and tests/unit/media-page-client-browser-bundle.test.ts. Refs #13866 * fix(build): allowlist dist/httpClientAbortGuard.mjs so the published tarball keeps the server-ws crash guard #14064 (re-land of #13636) made scripts/dev/standalone-server-ws.mjs import ./httpClientAbortGuard.mjs and taught assembleStandalone to copy the shared implementation next to dist/server-ws.mjs — but never registered the file in scripts/build/pack-artifact-policy.ts. The prepublish prune deletes anything outside APP_STAGING_ALLOWED_EXACT_PATHS, and check:pack-artifact only fails on PACK_ARTIFACT_REQUIRED_PATHS entries, so the next `omniroute` tarball would boot straight into ERR_MODULE_NOT_FOUND (the #7065 / tls-options class the closure tests exist to catch). Add the bare and dist/ entries to both lists and extend the required-paths snapshot in tests/unit/pack-artifact-policy.test.ts. Greens tests/unit/pack-artifact-entrypoint-closures.test.ts and tests/unit/pack-artifact-server-ws-closure.test.ts. Refs #13866 * test(docker): accept --chown=node:node on the better-sqlite3 runner COPY #14010 deliberately changed the runner-stage COPYs to `COPY --chown=node:node --from=builder ...` (ownership at copy time instead of a second ~2 GB `chown -R` overlay layer). The Dockerfile contract test still matched the old `COPY --from=builder /app/node_modules/better-sqlite3` prefix and went red on the tip even though the native-addon guard it protects is intact. Tolerate the optional --chown flag; every other assertion (node-gyp rebuild, both `test -f .../better_sqlite3.node` checks) is unchanged. Refs #13866 * fix(api): validate /v1/responses/input_tokens bodies with Zod (t06) #13167 added the local Responses token-count route with hand-rolled `typeof` checks on `request.json()`. Hard Rule #7 and the t06 gate (scripts/check/check-route-validation.mjs, mirrored by tests/unit/route-body-validation-t06.test.ts) require every route that reads request.json() to go through validateBody()/safeParse(), so the tip was red. Add `v1ResponsesInputTokensSchema` (pins the wire types the counter reads — model/instructions strings, input string-or-array, tools array — and lets unknown keys through since they are counted, never forwarded) and run the body through validateBody(); a type mismatch is now a 400 naming the field instead of a silently ignored key. Regression test added to tests/unit/responses-input-tokens-local-route.test.ts. Refs #13866 * fix(docs): drop the retired openference.svg asset reintroduced by #13378 `openference.svg` is one of the 78 provider assets retired for missing provenance (tests/unit/provider-assets-generic-fallback.test.mjs freezes that list and forbids any tracked surface from referencing a retired name). #13378 added a new hand-drawn `public/openference.svg` outside the manifest-audited public/providers/ tree and pointed the README free-tier table (plus the 22 i18n mirrors that carry the row) at it, which put the retired name back on a tracked surface and left an unaudited asset in the package. Use the generic fallback icon (`public/providers/cli-generic.svg`) the other provenance-less providers already use, delete the unaudited file, and adopt the mechanical README edit into .i18n-state.json (`i18n:run -- --adopt --files=README.md`, no API calls) so the i18n drift gate does not flag README.md as source-changed. Refs #13866 * test(lease): classify the two connection-query sites added by #14159 and #13874 The hard-lease bypass inventory froze every getProviderConnections / getProviderConnectionById site with a class; two landed on the tip without a golden update: - src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of #12630): read-only listing that feeds the semantic-cache embedding dropdown, same shape as the qdrant embedding-models route — class C. - src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an unrecoverable refresh error to detect credentials rotated by a concurrent Layer 2 refresh before deactivating — a state read, not dispatch; stays C. Refs #13866 * fix(sse): restore nvidia as the emergency budget-fallback provider #14006 (Antigravity MITM catalog injection) flipped EMERGENCY_FALLBACK_CONFIG.provider from "nvidia" to "groq" in one line, without touching ENVIRONMENT.md, .env.example, the chat.ts comment or the NVIDIA hosted-model snapshot, all of which still promise nvidia/openai/gpt-oss-120b. Operators without a Groq connection got the original 402 back instead of the free reroute, and chat-route-coverage ("uses the emergency fallback model on budget exhaustion" / "returns the primary budget error when emergency fallback also fails") went red on the tip. Put the documented default back; the #14006 bridge tests exercise bin/antigravity-bridge.mjs and do not read this config. Refs #13866 * fix(resilience): keep maxWaitMs=0 a "no queue deadline" sentinel #12902 released requestQueue.maxWaitMs=0 as the sentinel that disables the queue-wait deadline, but the #12715 queue-budget gate in withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget spent" and rejected every request on a protected connection with an immediate 503 queue-budget error — the exact opposite of what the setting promises. rate-limit-maxwaitms-disable-execution ("400ms job completes without 504") was red on the tip. When no caller budget is passed and the configured queue budget is 0, skip the gate, never arm the queue-wait timer and hand awaitProviderDefaultSlot no budget (it falls back to the window). Execution stays bounded by executionMaxWaitMs and the upstream fetch-start timeout, as before. Refs #13866 * test: align three fixtures with the #13874, #13861 and #13350 contracts Three base-reds that are deliberate contract changes, not defects: - executor-default-base "refreshCredentials swallows refresh errors": #13874 records rotations on the Layer 2 (no connectionId) refresh path too, so the "refresh-me" token the previous case already rotated was served from the rotation map without the network POST the test wanted to fail. Use a token nobody rotated. - telegram-keycache-bounded-13165: #13861 made comboTargetKeyPolicy import isModelBlockedByPatterns from db/apiKeys; the loader-stubbed module lacked it and the suite died at module load. Export an honest "not blocked" stub (the test has no blocked models). - upstream-headers-proxy-auth "ordinary headers are still allowed": #13350 forbids the whole origin-IP forwarding set upstream (covered by upstream-headers-sanitize). Swap x-forwarded-for for x-request-id. Refs #13866
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…emory vector stores (re-land of diegosouzapw#12630) (diegosouzapw#14159) Re-land of diegosouzapw#12630 by @BillyOutlast (their commits carried with authorship intact), merged via /merge-batch (2026-09-19) on top of the current `release/v3.8.51` tip. **Reconciled before landing — `0b34793e`.** The re-land had only run its own 78 tests; against the existing suites of the modules it touches it introduced **15 regressions** (all green on the pure tip, reproduced red with the PR). Root causes and fixes: - **Vector layer was on by default** (`semanticCacheConfig.ts` `enabled: true`, bridged from the legacy `semanticCacheEnabled` toggle) and its default embedding client called `http://localhost:13305/v1/embeddings` on every cacheable lookup *and* store — `chat-combo-live-test` ×2 and `issue-agent-route-execution` saw 3 fetches instead of 1. The layer is now **opt-in** via a new `semanticCacheVectorEnabled` setting (default `false`; sub-toggle in the Cache settings tab; `OMNIROUTE_SEMANTIC_CACHE_ENABLED=true` still works). With it off, `chatCore` behaves exactly like the legacy SQLite exact-match cache. - **`X-OmniRoute-Cache` changed from `HIT` to `HIT (exact)`/`HIT (semantic)`** — 5 chat-route/chatCore contract tests. Restored the legacy `HIT` value (similarity hits keep `X-OmniRoute-Cache-Similarity`); `cacheSource: "semantic_similarity"` also fell through `attemptLogging`'s narrowing as `"upstream"` and is now `"semantic"`. - **`normalizeDiscoveredModels` stamped `modelType: "chat"` + `supportedInputTypes: ["text"]` on every model** — kimi/vertex/reasoning-levels/provider-models/model-sync snapshots churned. Chat models keep the tip's exact shape; only non-chat modalities (or explicit `supportedInputTypes`) are stamped. - **`detectModelModality` classified `supportedEndpoints: ["chat","embeddings"]` as embedding** and dropped the model from the chat catalog. An explicit chat endpoint is now authoritative over the embedding/rerank/image heuristics. - Extras found on the way: `chatCore` now passes `provider` to both stores (the manager filters by provider on lookup, so writes without it could never hit); `test-embedding/route.ts` returned `undefined` on invalid payloads (`validateBody()` has no `.response`) → 400. - Tests: `chatcore-semantic-cache.test.ts` restored to the tip's contract; `semantic-cache-no-truncated-writes.test.ts` restored to the tip + the PR's two object-shaped streaming cases appended (with `isTruncatedStreamBody` now delegating to `isTruncatedCompletion` for object bodies — on the tip that guard was a no-op in production); new guard `semantic-cache-vector-layer-opt-in.test.ts`. **Evidence on the merged tree:** the 9 previously-red files + the PR's 7 test files: 270 pass / 0 fail / 2 skipped (Lemonade live, self-skip); `typecheck:core` exit 0; `check:open-sse-typecheck` 0 errors; `check-api-typecheck` only the two inherited errors (`rerankProviderNodes.ts`, `antigravity.ts`); file-size, changelog-integrity, docs-counts, complexity, cognitive-complexity, vitest-exclusions OK; the 5 removed eslint suppressions verified clean. **Owner decisions surfaced by the rework:** the PR wanted the hit type in the `X-OmniRoute-Cache` value — kept the legacy value; a separate header would be the non-breaking way. `modelDiscovery.ts` now considers `record.max_tokens` as an `inputTokenLimit` candidate (on several providers that is the *output* limit) — left as submitted, untested. The contributor's `/review/` `.gitignore` + eslint ignore entries were left as submitted. **Inherited, not from this PR:** the fast-path unit shard reds shared with every PR of this wave (vi locale parity, pack-artifact allowlists, `.env.example` sync, casing, budget fallback), `hard-session-lease-bypass-inventory`, the 5 `no-unused-vars` lint errors, the `omni-version-manager` generated-skill drift. Supersedes diegosouzapw#12630.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Re-land of #12630 by @BillyOutlast — the original fork (Heretek-AI) does not accept maintainer pushes, so this PR carries their commits with authorship intact.
What this is
Restructures OmniRoute's semantic cache into a configurable dual-layer architecture:
exact-hash matching (layer 1) plus configurable vector-similarity matching (layer 2),
with a pluggable backend (zero-dependency in-memory LRU, or Redis/Valkey with
RediSearch when available), a settings UI at
/dashboard/settings/cache, andmodel-discovery/modality guardrails so embedding models are excluded from the chat
endpoint. It extends the existing single-layer semantic cache
(
src/lib/semanticCache.ts) rather than replacing it from scratch, and explicitlyre-tests the historical cross-key isolation bug (#3740).
What the 2026-09-15 fix commits on top of the original PR applied
CONFLICTING DIRTY); 2 real content conflicts inopen-sse/handlers/chatCore/.had accidentally dropped were still needed (
use-improve-prompt,use-presets,use-stream-metrics,use-structured-output,use-tools-builder— none touched bythis PR) and restored them; the remaining removed entries (
tlsClientBase.ts,check-licenses.test.ts,combo-routing-engine.test.ts,src/lib/semanticCache.ts)were verified clean with
eslintand left removed.architecture so none of them regress in the transition:
checkSemanticCachenow finalizes the exactpending request by id instead of an ambiguous
(model, provider, connectionId)tuple.
isTruncatedCompletion()predicate gates both the legacy write path and the new dual-layer manager's
write path.
confirmed already covered by the new architecture's signature (test:
#12734: identical tool_choice/tools/response_format across requests still HITs).TS2339inredisVectorStore.ts(RedisLikewas missing the optionaldisconnect()methodclose()already calls as an ioredis fallback), and allowlisted theLEMONADE_URL/LEMONADE_KEY/LEMONADE_MODELtest-only env vars incheck-env-doc-sync.mjs(same class as the existingNVIDIA_BASE_URL/MODELentries — they only configure the self-skipping
tests/integration/semantic-cache-lemonade.test.tsagainst an operator's localLemonade server, never OmniRoute runtime config).
tests/unit/model-embedding-discovery-and-cache.test.tsmade an unguarded HTTP callto a contributor's private Lemonade server; applied the same
isEndpointReachable()self-skip pattern already used by the integration test.Validation run for this re-land
npm run typecheck:core— clean (0 errors).npx eslint --config eslint.config.mjs --suppressions-location config/quality/eslint-suppressions.json <touched files>— 0 errors (the frozen-suppression counts for pre-existing files match exactly: no new violations introduced).npm run check:complexity-ratchets(repo-wide ratchet, the actual CI gate) — no regression vs the frozen baseline.npm run check:file-size— clean, no new oversized files.npm run check:env-doc-sync— the only reported gap (BRIDGE_PORT,CERT_DIR,OPENWA_SERVICE_PORT,ROUTER_URL) is inherited from the base branch merge, unrelated to this PR's files.semantic-cache-dual-layer.test.ts(24/24),cache-config-route-8219.test.ts(6/6),lemonade-embedding-provider.test.ts,chatcore-semantic-cache.test.ts(16/16,including the fix(usage): finalize semantic cache hits by exact request id #12910 and fix(cache): fold the response output contract into the semantic cache signature #12309 regression cases),
semantic-cache-no-truncated-writes.test.ts,model-embedding-discovery-and-cache.test.ts.Supersedes #12630.