Skip to content

feat(cache): configurable dual-layer semantic caching with Redis/In-Memory vector stores (re-land of #12630) - #14159

Merged
diegosouzapw merged 13 commits into
release/v3.8.51from
feat/12630-semantic-cache-reland
Sep 19, 2026
Merged

diegosouzapw merged 13 commits into
release/v3.8.51from
feat/12630-semantic-cache-reland

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Re-land of #12630 by @BillyOutlast — the original fork (Heretek-AI) does not accept maintainer pushes, so this PR carries their commits with authorship intact.

What this is

Restructures OmniRoute's semantic cache into a configurable dual-layer architecture:
exact-hash matching (layer 1) plus configurable vector-similarity matching (layer 2),
with a pluggable backend (zero-dependency in-memory LRU, or Redis/Valkey with
RediSearch when available), a settings UI at /dashboard/settings/cache, and
model-discovery/modality guardrails so embedding models are excluded from the chat
endpoint. It extends the existing single-layer semantic cache
(src/lib/semanticCache.ts) rather than replacing it from scratch, and explicitly
re-tests the historical cross-key isolation bug (#3740).

What the 2026-09-15 fix commits on top of the original PR applied

  1. Rebase onto release/v3.8.51 — resolved as a merge (the branch showed
    CONFLICTING DIRTY); 2 real content conflicts in open-sse/handlers/chatCore/.
  2. Audited the eslint-suppressions.json diff — confirmed 5 entries this PR's merge
    had accidentally dropped were still needed (use-improve-prompt, use-presets,
    use-stream-metrics, use-structured-output, use-tools-builder — none touched by
    this PR) and restored them; the remaining removed entries (tlsClientBase.ts,
    check-licenses.test.ts, combo-routing-engine.test.ts, src/lib/semanticCache.ts)
    were verified clean with eslint and left removed.
  3. Ported the three targeted single-layer cache fixes into the new dual-layer
    architecture
    so none of them regress in the transition:
  4. Green the full gate: fixed a real pre-existing TS2339 in
    redisVectorStore.ts (RedisLike was missing the optional disconnect() method
    close() already calls as an ioredis fallback), and allowlisted the
    LEMONADE_URL/LEMONADE_KEY/LEMONADE_MODEL test-only env vars in
    check-env-doc-sync.mjs (same class as the existing NVIDIA_BASE_URL/MODEL
    entries — they only configure the self-skipping
    tests/integration/semantic-cache-lemonade.test.ts against an operator's local
    Lemonade server, never OmniRoute runtime config).
  5. Stopped 2 unit tests from calling a real LAN server —
    tests/unit/model-embedding-discovery-and-cache.test.ts made an unguarded HTTP call
    to a contributor's private Lemonade server; applied the same
    isEndpointReachable() self-skip pattern already used by the integration test.

Validation run for this re-land

  • npm run typecheck:core — clean (0 errors).
  • npx eslint --config eslint.config.mjs --suppressions-location config/quality/eslint-suppressions.json <touched files> — 0 errors (the frozen-suppression counts for pre-existing files match exactly: no new violations introduced).
  • npm run check:complexity-ratchets (repo-wide ratchet, the actual CI gate) — no regression vs the frozen baseline.
  • npm run check:file-size — clean, no new oversized files.
  • npm run check:env-doc-sync — the only reported gap (BRIDGE_PORT, CERT_DIR, OPENWA_SERVICE_PORT, ROUTER_URL) is inherited from the base branch merge, unrelated to this PR's files.
  • PR test suite — 78/78 passing (2 self-skipped, no local LAN access):
    semantic-cache-dual-layer.test.ts (24/24), cache-config-route-8219.test.ts (6/6),
    lemonade-embedding-provider.test.ts, chatcore-semantic-cache.test.ts (16/16,
    including the fix(usage): finalize semantic cache hits by exact request id #12910 and fix(cache): fold the response output contract into the semantic cache signature #12309 regression cases), semantic-cache-no-truncated-writes.test.ts,
    model-embedding-discovery-and-cache.test.ts.

Supersedes #12630.

BillyOutlast and others added 12 commits September 3, 2026 12:50
…#1)

* feat(cache): implement dual-layer semantic caching with in-memory and redis vector stores

Implements production-grade, configurable semantic caching for OmniRoute.

- Dual-layer architecture: Layer 1 exact hash match (0 embedding latency) + Layer 2 vector cosine similarity search.
- Backends: in-memory vector store with L2 normalization, LRU and TTL + Redis vector store adapter with fail-open fallback.
- Embedding generation: conversation history normalization, system prompt exclusion, timeout protection.
- Streaming support: serializable SSE stream synthesis ending in data: [DONE]\n\n.
- Request overrides and telemetry headers: X-OmniRoute-Cache (HIT (exact) | HIT (semantic) | MISS), X-OmniRoute-Cache-Similarity, X-OmniRoute-Savings-Tokens, Cache-Control: no-cache, x-omniroute-no-cache, x-omniroute-cache-threshold, x-omniroute-cache-type, x-omniroute-cache-no-store, x-omniroute-cache-key.
- Unit test coverage across dual-layer search, eviction, redis resilience, and streaming replay.

* feat(cache): add default embedding client and live integration test for Lemonade server and Redis

- Add embeddingBaseUrl and embeddingApiKey configuration options to SemanticCacheConfig.
- Implement createDefaultEmbeddingGenerator for automatic OpenAI-compatible embedding integration.
- Add live verification script scripts/ad-hoc/test-semantic-cache-lemonade.ts.
- Add network-aware integration test tests/integration/semantic-cache-lemonade.test.ts for Lemonade harrier-oss-v1-0.6b and Redis vector store.

* fix(cache): address CodeRabbit review recommendations on PR #1

- Multi-tenant partition isolation: support null sentinel in StoreFilter so anonymous requests cannot match authenticated entries.
- Provider propagation: pass routed provider to non-streaming and streaming cache writes.
- Embedding resilience: race generator with timeout promise in generateEmbeddingWithTimeout to guard against uncooperative generators.
- Memory store consistency: replace older entries with identical hash on insert, and safe-guard hashToId deletion in removeEntry.
- Redis store consistency: prune expired/missing entries from candidate sets during search/stats, and delete hash mapping conditionally.
- Token telemetry: use managerResult.tokensSaved when cached response lacks usage data.
- Clova batch timeout: add AbortSignal.timeout(FETCH_TIMEOUT_MS) to fetchClovaEmbeddingBatch.

---------

Co-authored-by: John Smith <you@example.com>
PR #12630 (feat(cache): configurable dual-layer semantic caching) was opened
2026-09-03 with head at 1cd8897 and is now mergeable: CONFLICTING because
upstream's release/v3.8.51 advanced 96 commits to ba597b6. This merge brings
semantic-cache up to date so the PR becomes mergeable again.

Three conflicts resolved:

- open-sse/handlers/chatCore.ts: took upstream. Upstream's restructure
  preserves storeSemanticCacheResponse at line 5365 and saveIdempotency at
  line 5379; the PR's Phase 9.1/9.2 block at lines 5528-5543 is textually
  equivalent and superseded by the upstream version.

- src/app/api/settings/cache-config/route.ts: combined. Took upstream's
  restructured file as base (empty-check guard, alwaysPreserveClientCache
  routed through flat settings — keeping upstream commit 8a95a2b's fix for
  the runtime no-op bug), then restored the PR's dual-layer cache content
  on top: resetSemanticCacheManager call, ensureSemanticCacheDbBridge()
  module-level side-effect, getEmbeddingOptions enrichment in GET, the 10
  new cache schema fields (semanticCacheBackend, semanticCacheThreshold,
  semanticCacheEmbeddingProvider/Model/Dimension/BaseUrl/ApiKey,
  semanticCacheRedisUrl/Prefix, semanticCacheRequireZeroTemp), and matching
  CACHE_CONFIG_KEYS + DEFAULTS entries. alwaysPreserveClientCache stays
  in the flat-settings path per upstream's fix.

- src/app/(dashboard)/dashboard/combos/page.tsx: took upstream. The
  useState+useEffect localStorage hydration pattern is replaced by
  useSyncExternalStore (no commit-after-mount, no hydration mismatch
  warning).

Skipped from this PR: the bf8691a lint cleanup for
tests/unit/volcengine-plan-binding-upsert.test.ts (upstream's file, not
touched by the PR) will go in as a separate upstream PR.

Forward-merge base-red inherited: #12732 (upstream's tip is red).
…-cache

Merges the dual-layer semantic cache branch onto the current release
tip so CI can run. Resolves two content conflicts by combining both
sides rather than picking one:

- open-sse/handlers/chatCore/semanticCache.ts: keeps the new
  dual-layer lookup (semanticCacheManager) with legacy SQLite fallback,
  and ports the release's #12309/#12734 fix (tool_choice/tools/
  response_format folded into the cache signature) into the legacy
  fallback path and its matching hit-recording signature so both stay
  in sync.
- src/lib/api/modelTestRunner.ts: keeps this branch's added embedding
  keyword detection (colbert/harrier-/nomic-embed) together with the
  release's isResponses/isNonChatGeneration routing added since.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…ped by this PR

The dual-layer semantic cache branch's "catch up to upstream" merge
commit accidentally dropped the react-hooks/immutability suppression
for 5 UI test files this PR never touches (use-improve-prompt,
use-presets, use-stream-metrics, use-structured-output,
use-tools-builder). Verified with eslint that they still trip the
rule and restored the entries; the other removed entries
(tlsClientBase.ts, check-licenses.test.ts, combo-routing-engine.test.ts,
src/lib/semanticCache.ts) were confirmed clean and left as-is.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Ports the single-layer cache's #12910 fix into the new dual-layer
checkSemanticCache: finalize the exact pending request by id
(finalizePendingScope) instead of an ambiguous (model, provider,
connectionId) tuple, which could finalize the wrong in-flight
request when connectionId is null or multiple requests share a
connection.

Co-authored-by: Jihyun Son <initguru@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…yer cache

Ports the single-layer cache's #12885 fix into both write paths of
the new dual-layer architecture: a response cut short by the output
token ceiling (finish_reason "length"/"max_tokens") is a partial
answer, not a reusable one, so it must never be written to the
semantic cache. Caching it under a temperature:0 signature pinned
the truncation for every later identical request.

- src/lib/semanticCache.ts: new isTruncatedCompletion() predicate.
- semanticCacheStore.ts / streamingSemanticCacheStore.ts: gate both
  the legacy setCachedResponse() write and the new dual-layer
  manager.store() write on it.

Co-authored-by: Patryk Kopyciński <contact@patrykkopycinski.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…c-sync

- open-sse/services/cache/redisVectorStore.ts: RedisLike was missing the
  optional disconnect() method that close() already calls as an ioredis
  fallback when quit() isn't available — 2 TS2339 errors under
  check:open-sse-typecheck.
- scripts/check/check-env-doc-sync.mjs: allowlist LEMONADE_URL/KEY/MODEL,
  same class as the existing NVIDIA_BASE_URL/MODEL entry — they only
  configure the gated tests/integration/semantic-cache-lemonade.test.ts
  live test against an operator's local Lemonade server (self-skips when
  unreachable), never OmniRoute runtime config.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Two tests in tests/unit/model-embedding-discovery-and-cache.test.ts
made a real HTTP call to the contributor's private Lemonade server
(192.168.31.147) with no reachability guard — they failed with a
60s socket timeout on any machine without access to that LAN,
including this one. tests/integration/semantic-cache-lemonade.test.ts
already gates the same endpoint with an isEndpointReachable() skip;
apply the identical pattern here so a unit test never depends on
live network access.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…-cache re-land

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

# Conflicts:
#	changelog.d/fixes/12910-semantic-cache-exact-id-finalize.md
#	open-sse/handlers/chatCore/semanticCache.ts
#	open-sse/handlers/chatCore/streamingSemanticCacheStore.ts
#	src/app/(dashboard)/dashboard/providers/[id]/hooks/useModelImportHandlers.ts
#	src/lib/providerModels/modelDiscovery.ts
#	tests/unit/chatcore-semantic-cache.test.ts
#	tests/unit/semantic-cache-no-truncated-writes.test.ts
…ery contracts (#14159)

The re-land of #12630 validated only its own tests and broke 15 existing ones.
Three root causes, all fixed in production code so the tip's tests are untouched:

1. Dual-layer manager was ON by default. `DEFAULT_SEMANTIC_CACHE_CONFIG.enabled`
   was `true` and the DB bridge wired it to the pre-existing `semanticCacheEnabled`
   toggle (also default `true`), so every temperature=0 request called an embedding
   endpoint (lemonade @ localhost:13305 by default) on lookup AND store — two extra
   fetches per chat call (chat-combo-live-test, issue-agent-route-execution).
   Fix: new `semanticCacheVectorEnabled` setting (default false; types, DB keys,
   cache-config route, settings UI sub-toggle) and the bridge only enables the
   manager when BOTH toggles are on. Manager default is now `enabled: false`.
   With the opt-in off, chatCore behaves exactly like the legacy SQLite cache.

2. `X-OmniRoute-Cache` value changed from `HIT` to `HIT (exact)`/`HIT (semantic)`,
   and `cacheSource: "semantic_similarity"` fell through attemptLogging's
   `"semantic" | "upstream"` narrowing as "upstream". Fix: keep `HIT` verbatim
   (similarity hits are distinguished by X-OmniRoute-Cache-Similarity) and log
   both hit types as `cacheSource: "semantic"`. The PR's edits to
   chatcore-semantic-cache.test.ts are reverted to the tip version.

3. Model discovery stamped `modelType: "chat"` + `supportedInputTypes: ["text"]`
   on every model (import-mode diff noise, 6 catalog tests), and endpoint-based
   modality inference tagged a `["chat","embeddings"]` model as embedding-only
   (`apiFormat: "embeddings"`), dropping it from the chat catalog. Fix: an explicit
   chat endpoint vetoes the embedding/rerank/image heuristics; chat models keep
   the tip's row shape (metadata only stamped for non-chat modalities).

Also:
- Restore the tip's `isTruncatedStreamBody` dep in streamingSemanticCacheStore and
  widen the predicate to the assembled-object body chatCore actually passes (the
  PR's object-shaped streaming cases are appended to the tip's test file).
- Pass `provider` from chatCore to both cache store paths — the manager scopes
  entries per provider on lookup, so writes without it could never hit.
- test-embedding route: `validateBody()` has no `.response`; invalid payloads
  returned `undefined` (TS2339 in api-typecheck). Now a proper 400.
- New guard: tests/unit/semantic-cache-vector-layer-opt-in.test.ts.
@diegosouzapw
diegosouzapw merged commit 7a92129 into release/v3.8.51 Sep 19, 2026
8 of 11 checks passed
diegosouzapw added a commit that referenced this pull request Sep 21, 2026
…and #13874

The hard-lease bypass inventory froze every getProviderConnections /
getProviderConnectionById site with a class; two landed on the tip without a
golden update:

- src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of
  #12630): read-only listing that feeds the semantic-cache embedding dropdown,
  same shape as the qdrant embedding-models route — class C.
- src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an
  unrecoverable refresh error to detect credentials rotated by a concurrent
  Layer 2 refresh before deactivating — a state read, not dispatch; stays C.

Refs #13866
diegosouzapw added a commit that referenced this pull request Sep 21, 2026
Three squash merges on release/v3.8.51 lost the contributor attribution that
the pipeline had preserved on the branches:

- 7a92129 (#14159) re-landed #12630: the dual-layer semantic cache is
  @BillyOutlast's work; the squash author is the maintainer.
- cd4c6f6 (#14162) re-landed #13180: the native Codex auto-resume fix is
  Felipe Reis's (@mdigitalbh81); the squash author is the maintainer.
- 53a147c (#13295) shipped the same one-line resolvedExtensionEnd reorder
  that @ggiak landed first in #13036; the squash carried no trailer.

The changelog fragment and these trailers record that credit in the branch
history, the contributors graph and the release notes.

Co-authored-by: BillyOutlast <172061051+BillyOutlast@users.noreply.github.com>
Co-authored-by: Felipe Reis <mdigitalbh81@gmail.com>
Co-authored-by: Giorgos Giakoumettis <giorgos@yiakoumettis.gr>
Co-authored-by: ggiak <20743694+ggiak@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Sep 21, 2026
…-09-21)

Third reconciliation pass of the living [3.8.51] section against release/v3.8.50..release/v3.8.51 (0915890 → 06f1df9): 364 fragments folded, 104 bullets generated for commits without a fragment, 254 fragment bullets linked and credited by origin commit, 1,393 bullets total, 226 external contributors (0 missing on cross-check). Hand-corrected credits: #12885 → @patrykkopycinski, #14159 → @BillyOutlast (feature bullet), #12972 → @IAMBOBJIM.
@diegosouzapw diegosouzapw mentioned this pull request Sep 21, 2026
diegosouzapw added a commit that referenced this pull request Sep 22, 2026
…-09-21) (#14371)

Third reconciliation pass of the living [3.8.51] section against release/v3.8.50..release/v3.8.51 (0915890 → 06f1df9): 364 fragments folded, 104 bullets generated for commits without a fragment, 254 fragment bullets linked and credited by origin commit, 1,393 bullets total, 226 external contributors (0 missing on cross-check). Hand-corrected credits: #12885 → @patrykkopycinski, #14159 → @BillyOutlast (feature bullet), #12972 → @IAMBOBJIM.
diegosouzapw added a commit that referenced this pull request Sep 22, 2026
…ll, TS2677, ESLint (Refs #13866) (#14331)

* fix(build): ship httpClientAbortGuard.mjs in the pack artifact; validate input_tokens with Zod

Wave five of the release/v3.8.51 base-reds, part 1 — the two that matter.

#14064 restored server-ws.mjs's import of ./httpClientAbortGuard.mjs and the
assembleStandalone copy, but not the two pack-artifact policy entries that
were lost with it. Without APP_STAGING_ALLOWED_EXACT_PATHS the prepublish
prune deletes the file; without PACK_ARTIFACT_REQUIRED_PATHS nothing notices.
Every boot of the published package would die with ERR_MODULE_NOT_FOUND — the
3.8.47 head-response-guard class. Both closure suites (9/9) now enforce it.

#13910's /v1/responses/input_tokens read request.json() behind a hand-rolled
typeof check. Hard Rule #7 wants the boundary on Zod; the t06 guard caught it.
Same passthrough envelope the catch-all Responses route uses, since the
counters below already walk the fields defensively. 9/9 on the route's suite.

Five no-unused-vars left behind by the wave (cliRuntime execFileSync, arena
test symbols and a type, compression rmSync, waitForServer req) are removed.
The 'openwa routes removed without deprecation' entry from the #14101 run was
an artifact of that PR trailing its base — the gate is clean on the tip.

Refs #13866

* fix(compression): let anchored file-pack rules see the transformed text; align wave-5 guards

Wave five of the release/v3.8.51 base-reds, part 2.

One production defect. #12825 (Hungarian Caveman pack) stopped gating file-pack
rules with the English keyword list and tested the rule's own regex instead —
against `lowerResult`, a lower-cased copy of the ORIGINAL text that the loop
never refreshed. An anchored pattern like leader_phrases' `^(?:i will|…)`
therefore ran its prefilter on "sure, i will…", failed the anchor, and was
skipped; the rule that strips "I will " from every English response was dead
since the merge. The prefilter now sees the text as the rules so far have left
it. New test fails on the tip and passes here; all Caveman suites, Hungarian
included, are 95/95. A frozen no-unused-vars suppression on caveman.ts no
longer had a target and is pruned.

Two more TS2677 predicates of the kind #14101 fixed: #13910
(rerankProviderNodes.ts, `n is RerankProviderNodeRow` on a Record row) and
#13957's mitm catalog (antigravity.ts, `c is DynamicCatalogModel` on a
literal-or-null). Both narrow by NonNullable of the element's own type; the
api-route typecheck was 285 against a baseline of 283 on the pristine tip.

The rest are guards trailing legitimate changes:

- #12663 made gemini-3.8-flash the catalog head; T28 pinned 3.7.
- #13863 put mimo-v2.5 into the shared vision heuristic on purpose (the base
  model is multimodal, only the Pro variants are text-only). The safety test
  now asserts the real invariant: base and :free aliases yes, -pro no.
- #12565 moved npm-prefix detection into cliRuntimeNpmPrefix.ts with a
  process-lifetime cache that importFresh() does not reset; the case resets it.
  #12565 also builds Windows candidates with path.win32 on purpose; the qodercli
  test compared against POSIX path.join.
- #13990 (the 2 GB Docker image) copies better-sqlite3 with --chown; the guard
  matched the flag order literally. Now flag-order tolerant, still fails when
  --from=builder is removed.
- #13378 reintroduced public/openference.svg under a name #11750 retired for
  missing provenance and swapped the Cerebras showcase cell for it. The cell is
  back and the asset is gone; whether the new drawing counts as provenance is
  the owner's call.

Refs #13866

* fix(i18n): translate the 7 sidebar-pin and Claude low-priority keys into all 65 locales

#7f1b4a5e (sidebar pinned items) and #1b2349de (Claude OAuth lower-priority /
auto-reset) landed with their 7 new keys in en.json only, which the vi and
pt-BR parity suites flag. Translated with the repo's own sync-ui-keys
--translate-markers against the .113 i18n instance (codex/gpt-5.6-sol-low):
+446 lines across 65 catalogs, zero __MISSING__ markers, placeholders intact.
vi.json also has two keys reordered to mirror en.json; values unchanged.

Refs #13866

* chore(quality): list the 8 covering tests the sixth wave added in stryker tap.testFiles

30 commits landed on release/v3.8.51 while wave five drained; eight new unit
tests cover mutated modules and were not in tap.testFiles, so their mutant
kills did not count and check:mutation-test-coverage --strict failed on the
merged tree. Appended at the end of the list, nothing reordered.

Refs #13866

* docs: document the five env vars of the 09-18 wave; regenerate the version-manager skill for the open-wa routes

check:docs-all: BRIDGE_PORT, ROUTER_URL and CERT_DIR (bin/antigravity-bridge.mjs,
#c74cea3d), OPENWA_SERVICE_PORT (src/lib/services/bootstrap.ts, #1e8c913c) and
NEXT_PUBLIC_PORT (src/shared/hooks/useDisplayBaseUrl.ts, #d715190b) were read
in code but absent from .env.example and docs/reference/ENVIRONMENT.md. Added
next to their neighbours, with the defaults the code actually uses (open-wa is
8323, not the 201xx range the other services sit in).

check:agent-skills-sync: the open-wa feature added eight /api/services/openwa/*
routes to docs/openapi.yaml without regenerating skills/omni-version-manager/
SKILL.md. Regenerated with the repo generator; the diff is exactly those eight
route sections.

Refs #13866

* test: register the crash guard in the pack snapshot; inventory #13874's refresh-lane row read

pack-artifact-policy pins the list of root runtime files check:pack-artifact
must find in the tarball; dist/httpClientAbortGuard.mjs joined
PACK_ARTIFACT_REQUIRED_PATHS in this PR and the snapshot follows.

#13874 re-reads the connection row inside the Claude refresh lane so a queued
health check does not POST a refresh token a Layer 2 refresh already rotated —
a state read, inventoried like the family-cooldown lookup (tokenHealthCheck.ts
2 -> 3).

Refs #13866

* chore(quality): list native-codex-auto-resume test in stryker tap.testFiles (#13180 landed without it)

* fix(release): drain the seventh base-red wave of release/v3.8.51 (9 tests + pack-policy + dashboard-typecheck)

Three production defects the tests caught:
- rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's
  queue-budget gate as "0 ms left" and 503'd every protected request.
- emergencyFallback: #14006 silently switched the budget-exhaustion target
  provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored.
- claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only
  by casing; helpers renamed to claudeConnectionFieldValues.ts.

Guards realigned to legitimate changes: #13874 rotation map (distinct token in
the error test), #13350 origin-IP denylist, #13318 shared-catalog growth
(counts by invariant), comboTargetKeyPolicy import in the telegram stub, the
22 README mirrors that #13940/#14106 stamped with the retired openference.svg
(translated Cerebras cells recovered from history, hashes re-stamped),
bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard
typecheck regressions (typed pinned section, ComponentProps cast).

Refs #13866.

* test: type the #13848 Gemini pairing tests (no-explicit-any) and inventory the semantic-cache embedding picker's connection read

Both arrived with the tip merge: #13848 added 13 explicit any casts to
translator-openai-to-gemini.test.ts (no-explicit-any is an error under
tests/), and 7a92129's embeddingOptions.ts reads provider connections
once without a hard-session-lease inventory entry. Stale suppression
count pruned for the test file only.

Refs #13866.

* test: split the #13848 turn-pairing cases out of translator-openai-to-gemini.test.ts

The file sits exactly at its frozen size cap; typing the pairing tests
(no-explicit-any) pushed it 14 lines over. The two cases are a coherent
regression suite of their own, so they move to
translator-openai-to-gemini-turn-pairing-13848.test.ts (registered in
stryker tap.testFiles) instead of widening the baseline.

* docs(env): document BRIDGE_PORT, ROUTER_URL, CERT_DIR, OPENWA_SERVICE_PORT and NEXT_PUBLIC_PORT (Refs #13866)

check:env-doc-sync has been red on the release tip since these five vars
reached code without their .env.example / ENVIRONMENT.md entries:
bin/antigravity-bridge.mjs (BRIDGE_PORT, ROUTER_URL, CERT_DIR — #14006),
src/lib/services/bootstrap.ts + api/services/openwa/_lib.ts (OPENWA_SERVICE_PORT)
and src/shared/hooks/useDisplayBaseUrl.ts (NEXT_PUBLIC_PORT — #13533).
Defaults and source files copied from the reads themselves.

* chore(skills): regenerate omni-version-manager for the open-wa service routes (Refs #13866)

check:agent-skills-sync (Merge integrity job) has been red on the tip since
the open-wa embedded-service routes reached docs/openapi.yaml without the
generated SKILL.md being refreshed. Output of
scripts/skills/generate-agent-skills.mjs --apply, no hand edits: the eight
/api/services/openwa/* operations.

* fix(types): make the two TS2677 type predicates sound (Refs #13866)

check:api-typecheck has been red on the tip with two "type predicate's
type must be assignable to its parameter's type" errors:

- src/app/api/v1/_shared/rerankProviderNodes.ts (#13733): the read cache
  hands back `Record<string, unknown> | null`, and an interface whose members
  are all optional is not assignable to an index-signature type. Narrow to the
  non-null record and assert the row shape afterwards.
- src/mitm/handlers/antigravity.ts (#14006): the map callback returned
  `{ displayName: string }` while DynamicCatalogModel declares it optional, so
  the predicate could not be proven. Type the callback's return explicitly and
  filter on `!== null`.

No runtime change; rerank-remote-provider-nodes / rerank-local-node-shapes /
mitm-handler-antigravity stay green.

* fix(lint): clear the 92 ESLint errors the lint gate reports on the tip (Refs #13866)

- tests/unit/translator-openai-to-gemini.test.ts: #13848 / #13318 added 13
  `any` casts/params on top of the 74 frozen for the file, so ESLint reported
  all 87. Typed them (GeminiRequestWithContents / GeminiToolPart, and the
  existing GeminiRequestWithConfig) and pruned the file's suppression to the
  new count of 71 — nothing else in eslint-suppressions.json changes.
- no-unused-vars: execFileSync import (src/shared/services/cliRuntime.ts,
  #12565), getArenaEloSyncStatus + makeLeaderboardMap + ArenaLeaderboardMap
  (tests/unit/arena-elo-sync-redesign.test.ts, #13446), rmSync
  (compressionAnalyticsWriterFlatRate.test.ts, #13446), `req` → `_req`
  (waitForServer-slow-first-response.test.mjs).

translator-openai-to-gemini 48/48; arena-elo-sync-redesign,
compressionAnalyticsWriterFlatRate, waitForServer-slow-first-response green.

* fix(compression): stop skipping anchored Caveman rules that only match after earlier rules

#12825 (Hungarian pack) replaced the English keyword prefilter with a
`rule.pattern.test(lowerText)` pre-check for every file-based rule, including
the default `en` pack. `lowerText` is the ORIGINAL message, so anchored rules
such as `leader_phrases` (`^i will …`) — which only match after `pleasantries`
strips "Sure, " — were dropped before they could run. `caveman-v379` caught the
regression ("I will ensure …" survived at full intensity).

Tag file-based rules with their pack language in ruleLoader and let the
keyword prefilter apply to `en`/built-in rules only; non-English packs (which
reuse English rule names) simply run their localized regex, which is what the
pre-test cost anyway. Drops the now-unused CAVEMAN_RULES import and prunes the
already-stale `caveman.ts` no-unused-vars suppression (0 violations on the tip)
that blocked the pre-commit hook for any change to this file.

Refs #13866

* test(models): align catalog and vision-heuristic guards with the tip's intended contracts

Three base-reds where the production change was deliberate and the pinned
guard was simply not bumped by the PR that changed the contract:

- agy-antigravity-shared-catalog-12724: #13318 added the three Gemini 3.8
  Flash tiers (high/medium/low, no "-tiered" endpoint for 3.8) to the shared
  Antigravity/AGY base, 10 -> 13. Pin the new size in one constant and make the
  buildSurfaceCatalog delta assertions relative to it.
- t28-model-catalog-updates: #12663 (issue #12638) registered gemini-3.8-flash
  at the head of the AI Studio fallback catalog as the current Flash default;
  assert 3.8 first and keep 3.7 present.
- command-code-mimo-v2-5-safety: #13863 (issue #13847) added an explicit
  "mimo-v2.5" fragment to the shared vision heuristic so provider-qualified and
  `-free` aliases keep their vision flag. The guard's real concern (the
  "mimo-vl" fragment must not cover "mimo-v2.5") is asserted on the fragment
  itself; the bare id is now vision by heuristic on purpose, and the Pro
  text-only sibling stays excluded.

Refs #13866

* test(cli): follow the #12565 cliRuntime module split in the npm-prefix and qodercli guards

#12565 (issue #12563) moved the npm global-prefix cache out of cliRuntime.ts
into cliRuntimeNpmPrefix.ts and built the Windows known-bin candidates with
`path.win32` (cliRuntimeWindowsNode.ts) so they stay Windows-shaped when
`process.platform` is mocked on a POSIX runner. Two pre-existing guards
depended on the old layout:

- cli-runtime-extended "resolves known binaries from npm global prefix":
  importFresh() only re-evaluates cliRuntime.ts; the prefix cache now lives in
  a module that stays shared across cases, so a real `npm config get prefix`
  from an earlier case was cached and the mocked execFileSync never ran. Reset
  the cache with the helper #12565 exported for exactly this in afterEach.
- qodercli-windows-resolve-6263: compare against `path.win32.join` — identical
  to `path.join` on a real Windows host, which is the behaviour under test.

Production behaviour is unchanged on both platforms.

Refs #13866

* test(auto-update): write the source-mode log inside the test's own temp dir

The launchAutoUpdate case pointed AUTO_UPDATE_LOG_PATH at a fixed, world-shared
`/tmp/auto-update-source.log`. On the .113 runner the suite executes both as
`root` and as `runner` (uid 1001): the file survives owned by whoever ran
first (`-rw-r--r-- root root`), and the next `openSync(logPath, "a")` fails
with EACCES for the other user. Reproduced locally by making the shared file
read-only; production code is untouched (autoUpdate.ts last changed in #9354).

Use a per-test mkdtemp path for the source-mode log and clean the whole temp
root in the existing finally block.

Refs #13866

* fix(dashboard): rename claudeConnectionFields.ts so it no longer case-collides with ClaudeConnectionFields.tsx

#13074 added two modules to the provider-detail modals directory whose names
differ only by casing: `ClaudeConnectionFields.tsx` (the component) and
`claudeConnectionFields.ts` (the value/patch helpers). On a case-insensitive
filesystem the pair breaks the webpack build (#6584 guard), and esbuild's
resolver already picks the `.tsx` for the extension-less `./claudeConnectionFields`
specifier, so the provider-detail client entry failed to bundle ("No matching
export ... for import claudeConnectionFieldPatch").

Rename the helper module to `claudeConnectionFieldValues.ts` (the same naming
the sibling `quotaScrapingFieldValues.ts` uses) and point the only importer,
EditConnectionModal.tsx, at the new name. Greens
tests/unit/case-collision-6584.test.ts and
tests/unit/media-page-client-browser-bundle.test.ts.

Refs #13866

* fix(build): allowlist dist/httpClientAbortGuard.mjs so the published tarball keeps the server-ws crash guard

#14064 (re-land of #13636) made scripts/dev/standalone-server-ws.mjs import
./httpClientAbortGuard.mjs and taught assembleStandalone to copy the shared
implementation next to dist/server-ws.mjs — but never registered the file in
scripts/build/pack-artifact-policy.ts. The prepublish prune deletes anything
outside APP_STAGING_ALLOWED_EXACT_PATHS, and check:pack-artifact only fails on
PACK_ARTIFACT_REQUIRED_PATHS entries, so the next `omniroute` tarball would
boot straight into ERR_MODULE_NOT_FOUND (the #7065 / tls-options class the
closure tests exist to catch).

Add the bare and dist/ entries to both lists and extend the required-paths
snapshot in tests/unit/pack-artifact-policy.test.ts. Greens
tests/unit/pack-artifact-entrypoint-closures.test.ts and
tests/unit/pack-artifact-server-ws-closure.test.ts.

Refs #13866

* test(docker): accept --chown=node:node on the better-sqlite3 runner COPY

#14010 deliberately changed the runner-stage COPYs to `COPY --chown=node:node
--from=builder ...` (ownership at copy time instead of a second ~2 GB
`chown -R` overlay layer). The Dockerfile contract test still matched the old
`COPY --from=builder /app/node_modules/better-sqlite3` prefix and went red on
the tip even though the native-addon guard it protects is intact. Tolerate the
optional --chown flag; every other assertion (node-gyp rebuild, both
`test -f .../better_sqlite3.node` checks) is unchanged.

Refs #13866

* fix(api): validate /v1/responses/input_tokens bodies with Zod (t06)

#13167 added the local Responses token-count route with hand-rolled
`typeof` checks on `request.json()`. Hard Rule #7 and the t06 gate
(scripts/check/check-route-validation.mjs, mirrored by
tests/unit/route-body-validation-t06.test.ts) require every route that reads
request.json() to go through validateBody()/safeParse(), so the tip was red.

Add `v1ResponsesInputTokensSchema` (pins the wire types the counter reads —
model/instructions strings, input string-or-array, tools array — and lets
unknown keys through since they are counted, never forwarded) and run the body
through validateBody(); a type mismatch is now a 400 naming the field instead
of a silently ignored key. Regression test added to
tests/unit/responses-input-tokens-local-route.test.ts.

Refs #13866

* fix(docs): drop the retired openference.svg asset reintroduced by #13378

`openference.svg` is one of the 78 provider assets retired for missing
provenance (tests/unit/provider-assets-generic-fallback.test.mjs freezes that
list and forbids any tracked surface from referencing a retired name). #13378
added a new hand-drawn `public/openference.svg` outside the manifest-audited
public/providers/ tree and pointed the README free-tier table (plus the 22
i18n mirrors that carry the row) at it, which put the retired name back on a
tracked surface and left an unaudited asset in the package.

Use the generic fallback icon (`public/providers/cli-generic.svg`) the other
provenance-less providers already use, delete the unaudited file, and adopt
the mechanical README edit into .i18n-state.json
(`i18n:run -- --adopt --files=README.md`, no API calls) so the i18n drift gate
does not flag README.md as source-changed.

Refs #13866

* test(lease): classify the two connection-query sites added by #14159 and #13874

The hard-lease bypass inventory froze every getProviderConnections /
getProviderConnectionById site with a class; two landed on the tip without a
golden update:

- src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of
  #12630): read-only listing that feeds the semantic-cache embedding dropdown,
  same shape as the qdrant embedding-models route — class C.
- src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an
  unrecoverable refresh error to detect credentials rotated by a concurrent
  Layer 2 refresh before deactivating — a state read, not dispatch; stays C.

Refs #13866

* fix(sse): restore nvidia as the emergency budget-fallback provider

#14006 (Antigravity MITM catalog injection) flipped
EMERGENCY_FALLBACK_CONFIG.provider from "nvidia" to "groq" in one line,
without touching ENVIRONMENT.md, .env.example, the chat.ts comment or the
NVIDIA hosted-model snapshot, all of which still promise
nvidia/openai/gpt-oss-120b. Operators without a Groq connection got the
original 402 back instead of the free reroute, and
chat-route-coverage ("uses the emergency fallback model on budget
exhaustion" / "returns the primary budget error when emergency fallback
also fails") went red on the tip.

Put the documented default back; the #14006 bridge tests exercise
bin/antigravity-bridge.mjs and do not read this config.

Refs #13866

* fix(resilience): keep maxWaitMs=0 a "no queue deadline" sentinel

#12902 released requestQueue.maxWaitMs=0 as the sentinel that disables
the queue-wait deadline, but the #12715 queue-budget gate in
withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget
spent" and rejected every request on a protected connection with an
immediate 503 queue-budget error — the exact opposite of what the
setting promises. rate-limit-maxwaitms-disable-execution ("400ms job
completes without 504") was red on the tip.

When no caller budget is passed and the configured queue budget is 0,
skip the gate, never arm the queue-wait timer and hand
awaitProviderDefaultSlot no budget (it falls back to the window).
Execution stays bounded by executionMaxWaitMs and the upstream
fetch-start timeout, as before.

Refs #13866

* test: align three fixtures with the #13874, #13861 and #13350 contracts

Three base-reds that are deliberate contract changes, not defects:

- executor-default-base "refreshCredentials swallows refresh errors":
  #13874 records rotations on the Layer 2 (no connectionId) refresh path
  too, so the "refresh-me" token the previous case already rotated was
  served from the rotation map without the network POST the test wanted
  to fail. Use a token nobody rotated.
- telegram-keycache-bounded-13165: #13861 made comboTargetKeyPolicy
  import isModelBlockedByPatterns from db/apiKeys; the loader-stubbed
  module lacked it and the suite died at module load. Export an honest
  "not blocked" stub (the test has no blocked models).
- upstream-headers-proxy-auth "ordinary headers are still allowed":
  #13350 forbids the whole origin-IP forwarding set upstream (covered
  by upstream-headers-sanitize). Swap x-forwarded-for for x-request-id.

Refs #13866
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…emory vector stores (re-land of diegosouzapw#12630) (diegosouzapw#14159)

Re-land of diegosouzapw#12630 by @BillyOutlast (their commits carried with authorship intact), merged via /merge-batch (2026-09-19) on top of the current `release/v3.8.51` tip.

**Reconciled before landing — `0b34793e`.** The re-land had only run its own 78 tests; against the existing suites of the modules it touches it introduced **15 regressions** (all green on the pure tip, reproduced red with the PR). Root causes and fixes:

- **Vector layer was on by default** (`semanticCacheConfig.ts` `enabled: true`, bridged from the legacy `semanticCacheEnabled` toggle) and its default embedding client called `http://localhost:13305/v1/embeddings` on every cacheable lookup *and* store — `chat-combo-live-test` ×2 and `issue-agent-route-execution` saw 3 fetches instead of 1. The layer is now **opt-in** via a new `semanticCacheVectorEnabled` setting (default `false`; sub-toggle in the Cache settings tab; `OMNIROUTE_SEMANTIC_CACHE_ENABLED=true` still works). With it off, `chatCore` behaves exactly like the legacy SQLite exact-match cache.
- **`X-OmniRoute-Cache` changed from `HIT` to `HIT (exact)`/`HIT (semantic)`** — 5 chat-route/chatCore contract tests. Restored the legacy `HIT` value (similarity hits keep `X-OmniRoute-Cache-Similarity`); `cacheSource: "semantic_similarity"` also fell through `attemptLogging`'s narrowing as `"upstream"` and is now `"semantic"`.
- **`normalizeDiscoveredModels` stamped `modelType: "chat"` + `supportedInputTypes: ["text"]` on every model** — kimi/vertex/reasoning-levels/provider-models/model-sync snapshots churned. Chat models keep the tip's exact shape; only non-chat modalities (or explicit `supportedInputTypes`) are stamped.
- **`detectModelModality` classified `supportedEndpoints: ["chat","embeddings"]` as embedding** and dropped the model from the chat catalog. An explicit chat endpoint is now authoritative over the embedding/rerank/image heuristics.
- Extras found on the way: `chatCore` now passes `provider` to both stores (the manager filters by provider on lookup, so writes without it could never hit); `test-embedding/route.ts` returned `undefined` on invalid payloads (`validateBody()` has no `.response`) → 400.
- Tests: `chatcore-semantic-cache.test.ts` restored to the tip's contract; `semantic-cache-no-truncated-writes.test.ts` restored to the tip + the PR's two object-shaped streaming cases appended (with `isTruncatedStreamBody` now delegating to `isTruncatedCompletion` for object bodies — on the tip that guard was a no-op in production); new guard `semantic-cache-vector-layer-opt-in.test.ts`.

**Evidence on the merged tree:** the 9 previously-red files + the PR's 7 test files: 270 pass / 0 fail / 2 skipped (Lemonade live, self-skip); `typecheck:core` exit 0; `check:open-sse-typecheck` 0 errors; `check-api-typecheck` only the two inherited errors (`rerankProviderNodes.ts`, `antigravity.ts`); file-size, changelog-integrity, docs-counts, complexity, cognitive-complexity, vitest-exclusions OK; the 5 removed eslint suppressions verified clean.

**Owner decisions surfaced by the rework:** the PR wanted the hit type in the `X-OmniRoute-Cache` value — kept the legacy value; a separate header would be the non-breaking way. `modelDiscovery.ts` now considers `record.max_tokens` as an `inputTokenLimit` candidate (on several providers that is the *output* limit) — left as submitted, untested. The contributor's `/review/` `.gitignore` + eslint ignore entries were left as submitted.

**Inherited, not from this PR:** the fast-path unit shard reds shared with every PR of this wave (vi locale parity, pack-artifact allowlists, `.env.example` sync, casing, budget fallback), `hard-session-lease-bypass-inventory`, the 5 `no-unused-vars` lint errors, the `omni-version-manager` generated-skill drift.

Supersedes diegosouzapw#12630.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants