Skip to content

feat(mcp): embedding-based semantic retrieval and passive context - #73

Merged
tonythethompson merged 15 commits into
mainfrom
feat/olive-mcp-semantic-retrieval
Aug 1, 2026
Merged

tonythethompson merged 15 commits into
mainfrom
feat/olive-mcp-semantic-retrieval

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Aug 1, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Add shared MiniLM embedding utilities (lazy CPU load) and upgrade docs search + troubleshooting to semantic/hybrid retrieval with keyword fallback.
  • Add get_context_for_pipeline for injecting pipeline-aware KB snippets into the AI assistant prompt.
  • Update deps (sentence-transformers, numpy) and Dockerfile to pre-cache MiniLM + ship HF cache for the non-root user.

Details

  • docs_search: semantic search over KB with mtime-safe index rebuild; live-doc embed cache; keyword relevance normalized to [0, 1].
  • troubleshooting: hybrid score 0.6 * semantic + 0.4 * keyword (patterns as OR); empty error messages do not match on pass name alone; fingerprint/mtime index invalidation.
  • Public APIs: original 14 tool signatures and return shapes preserved; new tools registered in mcp_server.TOOLS and tools.__init__.

Test plan

  • cd olive-mcp-server && python -m pytest tests -q → 186 passed
  • Optional: docker build -t olive-mcp-server . (large image with CPU torch + MiniLM)
  • Smoke: call search_olive_documentation with a fuzzy quant query
  • Smoke: troubleshoot_olive_error with paraphrased OOM text
  • Smoke: get_context_for_pipeline with a quant pass list

Summary by cubic

Adds MiniLM-based semantic retrieval for docs and hybrid troubleshooting, plus a passive pipeline context tool to improve assistant prompts. Tightens caching for live docs/indices, refines UI heuristics, and removes unsupported cu130/cu132 CUDA tags.

  • New Features

    • Embeddings: shared utils for all-MiniLM-L6-v2 (lazy CPU load).
    • Docs search: semantic + keyword fallback over KB and live docs; relevance normalized (~[0, 1]); KB index invalidation via (mtime, file count); live‑doc embedding cache keyed to fetch_time with rebuilds.
    • Troubleshooting: hybrid scoring (0.6 semantic + 0.4 keyword) with OR‑pattern bonus; empty error messages never match; fingerprint/mtime‑safe index with a bounded cache.
    • Passive context: get_context_for_pipeline returns pipeline‑aware snippets, summary, confidence, and count; tool registered; existing tool signatures unchanged.
    • Docker: force CPU‑only torch via index‑url, pre‑cache all-MiniLM-L6-v2, set HF_HOME and HF_HUB_OFFLINE=1 for the non‑root user.
  • Bug Fixes

    • Live docs: snapshot pages and fetch time atomically; generation token prevents stale overwrites; live embedding index recognizes empty‑cache builds and refuses stale publishes when fetch_time is older; also publish when the cache is empty even if a newer generation later fails; never displace a stronger local top‑1 with a weaker live hit.
    • Indexing/Embeddings: include file count in KB mtime tracking; reuse loaded texts on keyword fallback; materialize text iterables before encoding; add missing Sequence import; log failures.
    • Troubleshooting: shared response builder; bound fingerprint cache; drop unnecessary globals.
    • UI/Server/Tests: treat any positive relevance as sufficient in docsSearchSufficient; salvage convert/quant actions from step/action/task value strings but never from negated or non‑action prose; make live‑doc merge test deterministic; fix BatchProcessingPanel validation by setting conversionFormat: "onnx", using a constructible EventSource mock, and firing "done" to advance the sequential loop.
    • Dependencies: cap sentence-transformers and numpy below next untested majors; remove unsupported cu130/cu132 from auditAutofix.ts and chatActions.ts to align with resolvable CUDA tags (≤cu128).

Written for commit f5567fa. Summary will update on new commits.

Review in cubic

Replace keyword-only docs and troubleshooting search with MiniLM embeddings
and hybrid scoring, add get_context_for_pipeline for AI prompt injection,
and pre-cache the model in Docker. Existing tool signatures stay unchanged.
Copilot AI review requested due to automatic review settings August 1, 2026 14:25

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @tonythethompson, you have reached your weekly rate limit of 500000 diff characters.

Please try again later or upgrade to continue using Sourcery

@coderabbitai

coderabbitai Bot commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The MCP server now supports CPU MiniLM embeddings for semantic documentation search, pipeline context retrieval, and hybrid troubleshooting. It adds model caching, concurrent cache protection, keyword fallbacks, tool registration, tests, and two additional UI CUDA version values.

Changes

Semantic knowledge tools

Layer / File(s) Summary
Embedding foundation and runtime setup
olive-mcp-server/Dockerfile, olive-mcp-server/pyproject.toml, olive-mcp-server/olive_mcp_server/tools/embeddings.py, olive-mcp-server/tests/test_embeddings.py
Adds CPU MiniLM dependencies, Docker model caching, shared embedding utilities, similarity handling, and unit tests.
Semantic documentation search
olive-mcp-server/olive_mcp_server/tools/docs_search.py, olive-mcp-server/tests/test_docs_search_semantic.py, olive-mcp-server/tests/test_tools.py, src/server/services/ai/oliveMcpKnowledge.ts
Adds semantic local and live search, cache invalidation, concurrency protection, keyword fallback, normalized relevance handling, and related tests.
Pipeline context retrieval and registration
olive-mcp-server/olive_mcp_server/tools/passive_context.py, olive-mcp-server/olive_mcp_server/mcp_server.py, olive-mcp-server/olive_mcp_server/tools/__init__.py, olive-mcp-server/tests/test_passive_context.py, olive-mcp-server/tests/test_integration.py
Adds get_context_for_pipeline, exposes it through lazy MCP registration, and tests query construction, confidence, responses, and tool listing.
Hybrid troubleshooting matching
olive-mcp-server/olive_mcp_server/tools/troubleshooting.py, olive-mcp-server/tests/test_troubleshooting_hybrid.py
Replaces pattern-only matching with hybrid semantic and keyword scoring, indexed-cache invalidation, fallbacks, empty-error handling, and comprehensive tests.

Application compatibility fixes

Layer / File(s) Summary
UI and application compatibility
src/types.ts, src/components/features/BatchProcessingPanel.test.tsx, src/lib/chatActions.ts, src/server/services/ai/oliveMcpKnowledge.test.ts
Adds CUDA version values, waits for asynchronous batch processing, broadens JSON salvage detection, and verifies normalized relevance handling.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 7 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 40.68% which is insufficient. The required threshold is 60.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (7 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Pipeline Stage Enum Ordering ✅ Passed The PR changes no SessionWorkflowStage enum or member references; the tracked solution contains no such enum, comparisons, renumbering, or legacy-converter requirement.
Gpu/Cpu Runtime Boundary ✅ Passed The PR changes 19 paths, but none are under inference/ or managed CPU/GPU requirements files; the GPU/CPU runtime-boundary checks are not triggered.
Managed Host Restart Safety ✅ Passed The PR changes 19 files, none containing the four managed-host components or lease/readiness restart symbols; the managed host restart safety check is not applicable.
Title check ✅ Passed The title clearly summarizes the main changes: embedding-based semantic retrieval and passive pipeline context for MCP.
Description check ✅ Passed The description directly covers the semantic retrieval, passive context tool, dependency, Docker, testing, and related fixes in the changeset.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/olive-mcp-semantic-retrieval
✨ Simplify code
  • Create PR with simplified code
  • Commit simplified code in branch feat/olive-mcp-semantic-retrieval

Comment @coderabbitai help to get the list of available commands.

@qodo-code-review

Copy link
Copy Markdown
Contributor

PR Summary by Qodo

Add MiniLM semantic retrieval + pipeline passive context tool

✨ Enhancement 🧪 Tests ⚙️ Configuration changes 🕐 40+ Minutes

Grey Divider

AI Description

• Upgrade docs + troubleshooting tools to semantic/hybrid retrieval with keyword fallback.
• Add get_context_for_pipeline for pipeline-aware KB snippet injection into prompts.
• Pre-cache MiniLM in Docker and add deps + tests for caching and scoring behavior.
Diagram

graph TD
  U["MCP client"] --> S["MCP server"] --> R["Tool registry"] --> D["docs_search"]
  R --> T["troubleshooting"]
  R --> P["passive_context"]
  D --> E["embeddings util"] --> KB[("KB JSON files")]
  D --> LD{{"Live docs fetch"}}
  T --> E
  P --> E
  E --> HF[("HF model cache")]
  subgraph Legend
    direction LR
    _svc["Service/Module"] ~~~ _db[("Data store")]
    _ext{{"External"}}
  end
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. BM25-only (keyword) with better tokenization
  • ➕ No PyTorch/sentence-transformers dependency or large image growth
  • ➕ Simpler operational footprint and faster cold start
  • ➖ Worse recall for paraphrases/fuzzy queries (main value-add of this PR)
  • ➖ Requires careful tuning to avoid regressions across domains
2. FAISS/ANN vector index
  • ➕ Faster top-k retrieval for large KBs and frequent queries
  • ➕ Cleaner separation of index build vs query-time scoring
  • ➖ Extra dependency and more build/runtime complexity
  • ➖ Overkill if the KB remains small and in-memory numpy is sufficient
3. Offline precomputed embeddings committed as artifacts
  • ➕ Eliminates runtime embedding/index build cost
  • ➕ More predictable latency and avoids model load in hot paths
  • ➖ Adds artifact/versioning burden and risk of stale indices
  • ➖ Harder to support KB hot-reload semantics

Recommendation: Keep the PR’s current approach (in-memory numpy + lazy MiniLM load + mtime/fingerprint invalidation). It provides the semantic recall needed for fuzzy queries while preserving the existing tool APIs, and the added tests meaningfully de-risk the caching and scoring behavior. Consider an ANN index only if KB size or query volume grows enough that cosine-over-matrix becomes a bottleneck.

Files changed (13) +1689 / -156

Enhancement (6) +738 / -135
mcp_server.pyRegister get_context_for_pipeline tool in server tool map +4/-0

Register get_context_for_pipeline tool in server tool map

• Adds the new get_context_for_pipeline tool to the server’s TOOL_IMPORTS map so it is available via MCP transport without altering existing tool signatures.

olive-mcp-server/olive_mcp_server/mcp_server.py

__init__.pyExpose get_context_for_pipeline via tools package lazy imports +2/-0

Expose get_context_for_pipeline via tools package lazy imports

• Extends the tools registry and __getattr__ export list to include get_context_for_pipeline, keeping the same lazy import pattern used by other tools.

olive-mcp-server/olive_mcp_server/tools/init.py

docs_search.pyUpgrade docs search to semantic retrieval with robust caching +232/-39

Upgrade docs search to semantic retrieval with robust caching

• Introduces semantic search over the local KB and optional live docs snippets using shared embedding utilities, with keyword fallback when semantic results are empty or fail. Adds thread-safe caches for KB embeddings (mtime-invalidated) and live snippet embeddings (generation-safe fetch + cache-time-tied rebuild) and normalizes keyword relevance to roughly [0, 1].

olive-mcp-server/olive_mcp_server/tools/docs_search.py

embeddings.pyAdd shared MiniLM embedding + cosine scoring utilities +165/-0

Add shared MiniLM embedding + cosine scoring utilities

• Adds a new module that lazily loads sentence-transformers all-MiniLM-L6-v2 on first encode, produces float32 normalized embeddings, and provides safe cosine similarity and semantic_search helpers with thresholding. Designed to be shared across multiple tools to avoid duplicated model/index logic.

olive-mcp-server/olive_mcp_server/tools/embeddings.py

passive_context.pyAdd get_context_for_pipeline for prompt-injection context snippets +118/-0

Add get_context_for_pipeline for prompt-injection context snippets

• Implements a new tool that builds a semantic query from pipeline passes (strings or dict descriptors), optional model name, and hardware target, then retrieves top KB snippets using the shared KB embedding index. Returns a stable response shape including pipeline_summary and bounded confidence derived from average relevance.

olive-mcp-server/olive_mcp_server/tools/passive_context.py

troubleshooting.pyImplement hybrid semantic + keyword troubleshooting with safe matching +217/-96

Implement hybrid semantic + keyword troubleshooting with safe matching

• Replaces pattern-hit-only ranking with a hybrid score (semantic cosine over error_message + keyword OR evidence + small multi-hit bonus), and introduces an embedding index cache keyed by content fingerprint and KB mtime. Adds guardrails to prevent diagnosing from pass_name alone when error_message is empty/whitespace and ensures index invalidation on reorder/content changes.

olive-mcp-server/olive_mcp_server/tools/troubleshooting.py

Tests (5) +934 / -14
test_docs_search_semantic.pyAdd tests for semantic docs search, fallback, and cache invalidation +252/-0

Add tests for semantic docs search, fallback, and cache invalidation

• Adds unit tests that mock embedding/index functions to validate semantic-first behavior, keyword fallback, relevance normalization, mtime-safe KB index rebuild, and generation-safe live-doc cache publishing.

olive-mcp-server/tests/test_docs_search_semantic.py

test_embeddings.pyAdd unit tests for embedding utils and lazy-loading contract +153/-0

Add unit tests for embedding utils and lazy-loading contract

• Covers cosine similarity correctness/edge cases, semantic_search thresholding, and ensures the model is not loaded at import time. Uses mocking to validate embedding shapes without requiring the real model in most tests.

olive-mcp-server/tests/test_embeddings.py

test_integration.pyMake tool listing test derive expected tools from registry +5/-14

Make tool listing test derive expected tools from registry

• Updates the integration test to compute expected tool names from the server registry rather than a hardcoded set, and asserts new tool presence while keeping coverage for existing tools.

olive-mcp-server/tests/test_integration.py

test_passive_context.pyAdd tests for get_context_for_pipeline query building and confidence +145/-0

Add tests for get_context_for_pipeline query building and confidence

• Adds tests for empty pipelines, dict/string pass descriptors, semantic search integration via mocks, and confidence clamping to [0, 1] with stable return shape.

olive-mcp-server/tests/test_passive_context.py

test_troubleshooting_hybrid.pyAdd comprehensive tests for hybrid troubleshooting scoring and caching +379/-0

Add comprehensive tests for hybrid troubleshooting scoring and caching

• Introduces tests that mock embedding scores to validate semantic matches on paraphrases, keyword OR behavior, empty-error no-match rule, hybrid weighting, fingerprint-based cache invalidation, and an optional real MiniLM regression check when sentence-transformers is available.

olive-mcp-server/tests/test_troubleshooting_hybrid.py

Other (2) +17 / -7
DockerfilePre-cache MiniLM and ship HF cache for non-root runtime +15/-7

Pre-cache MiniLM and ship HF cache for non-root runtime

• Switches dependency install to prefer CPU-only torch wheels and pre-downloads the all-MiniLM-L6-v2 model during the builder stage. Copies the HuggingFace cache into the runtime image, sets HF_HOME for the mcp user, and updates comments to reflect the larger image reality.

olive-mcp-server/Dockerfile

pyproject.tomlAdd sentence-transformers and numpy dependencies +2/-0

Add sentence-transformers and numpy dependencies

• Introduces sentence-transformers (MiniLM embeddings) and numpy (matrix ops) as required dependencies to support semantic retrieval across tools.

olive-mcp-server/pyproject.toml

@qodo-code-review

qodo-code-review Bot commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📜 Skill insights (0)

Context used
✅ Compliance rules (platform): 170 rules
✅ Skills: 5 invoked
  vercel-optimize
  vercel-react-view-transitions
  typegpu
  vercel-react-best-practices
  vite-react-best-practices
✅ REVIEW.md

Grey Divider


Remediation recommended

1. Live cache snapshot race ✓ Resolved 🐞 Bug ≡ Correctness
Description
docs_search._get_live_index() fetches pages = _fetch_live_docs() and then reads _LAST_FETCH_TIME
in a separate critical section, so another thread can refresh the live cache between those two
operations. That can stamp snippets/embeddings built from older pages with a newer fetch_time,
causing semantic live-doc retrieval to be inconsistent until the next rebuild.
Code

olive-mcp-server/olive_mcp_server/tools/docs_search.py[R248-251]

+    pages = _fetch_live_docs()
+    with _LIVE_FETCH_LOCK:
+        fetch_time = _LAST_FETCH_TIME
+
Relevance

●●● Strong

Team has accepted atomic publish/snapshot fixes to prevent inconsistent global state across updates.

PR-#14

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
_get_live_index() reads the live pages and _LAST_FETCH_TIME separately (not an atomic snapshot),
while _fetch_live_docs() updates _LIVE_CACHE and _LAST_FETCH_TIME asynchronously after a
network fetch; a concurrent refresh between these reads can mis-associate the timestamp with
different pages.

olive-mcp-server/olive_mcp_server/tools/docs_search.py[197-230]
olive-mcp-server/olive_mcp_server/tools/docs_search.py[244-277]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`_get_live_index()` needs an atomic (pages, fetch_time) snapshot from the live-doc cache. Today it calls `_fetch_live_docs()` (which returns a copy of `_LIVE_CACHE`) and then separately reads `_LAST_FETCH_TIME`; another thread can update `_LIVE_CACHE/_LAST_FETCH_TIME` between those steps, causing the embedding cache timestamp to no longer correspond to the pages used to build the snippets/embeddings.

## Issue Context
This is a concurrency correctness issue in the live semantic index cache; it can lead to stale or mismatched live-doc results around cache refreshes.

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/docs_search.py[197-277]

## Suggested fix
1. Change `_fetch_live_docs()` to return both the cached pages *and* the cache timestamp under the same lock (e.g., `return dict(_LIVE_CACHE), _LAST_FETCH_TIME`).
2. Update `_get_live_index()` to use the returned timestamp rather than re-reading `_LAST_FETCH_TIME`.
3. (Optional) Adjust the cache-hit condition to not require `_LIVE_SNIPPETS` truthiness so empty-live-doc states can still be cached deterministically (avoid repeated rebuilds when pages are empty).

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Empty iterator loads model ✓ Resolved 🐞 Bug ☼ Reliability
Description
embeddings.encode_texts() checks if not texts before materializing the iterable, so an empty
generator/iterator is treated as truthy and forces model loading and model.encode([]). This can
trigger expensive initialization (and potentially errors) for what should be a no-op empty input.
Code

olive-mcp-server/olive_mcp_server/tools/embeddings.py[R40-55]

+def encode_texts(texts: Sequence[str]) -> np.ndarray:
+    """Encode a batch of texts into an (N, 384) float32 matrix.
+
+    Rows are L2-normalized (``normalize_embeddings=True``). Callers may still
+    pass the matrix through ``cosine_similarity_scores``, which re-normalizes
+    safely for zero rows and non-normalized inputs.
+    """
+    if not texts:
+        return np.zeros((0, EMBEDDING_DIM), dtype=np.float32)
+    model = _get_model()
+    embeddings = model.encode(
+        list(texts),
+        convert_to_numpy=True,
+        show_progress_bar=False,
+        normalize_embeddings=True,
+    )
Relevance

●●● Strong

Low-risk guard to avoid unnecessary expensive model load on empty inputs; aligns with prior
defensive empty-content handling.

PR-#32

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The function uses if not texts before list(texts), which is incorrect for truthy empty iterators
and leads to unnecessary model initialization and potentially encoding an empty list.

olive-mcp-server/olive_mcp_server/tools/embeddings.py[40-56]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`encode_texts()` should not load the SentenceTransformer model when given an empty iterator/generator. The current `if not texts:` check happens before `list(texts)`, so non-sized iterables bypass the early return.

## Issue Context
Even though the type is `Sequence[str]`, Python callers can pass iterables; the function already materializes the input later, so it should do so up front to correctly detect emptiness and avoid heavyweight model loads.

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/embeddings.py[40-56]

## Suggested fix
- Start `encode_texts` with:
 - `texts_list = list(texts) if texts is not None else []`
 - `if not texts_list: return zeros`
 - then call `model.encode(texts_list, ...)`
- (Optional) Update annotation to `Iterable[str]` to match behavior, and add a unit test that passes an empty generator (e.g., `(t for t in [])`) and asserts `is_model_loaded()` stays False.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Informational

3. Unused Path import ✓ Resolved 🐞 Bug ⚙ Maintainability
Description
tools/troubleshooting.py imports pathlib.Path but never uses it. This adds dead code and can fail
builds if unused-import linting is enforced.
Code

olive-mcp-server/olive_mcp_server/tools/troubleshooting.py[13]

+from pathlib import Path
Relevance

●●● Strong

They routinely accept removing unused imports to satisfy lint/cleanliness.

PR-#34

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The only occurrence of Path in the file is the import statement; no other references exist.

olive-mcp-server/olive_mcp_server/tools/troubleshooting.py[10-15]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`from pathlib import Path` is unused in `tools/troubleshooting.py`.

## Issue Context
This was introduced in the PR and can trigger unused-import lint failures.

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/troubleshooting.py[10-15]

## Suggested fix
- Remove the unused `Path` import (or use it if there was an intended path operation).

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

To customize comments, go to the Qodo configuration screen, or learn more in the docs.

Qodo Logo

Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/embeddings.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR upgrades the Olive MCP server’s retrieval capabilities by introducing MiniLM-based embeddings for semantic search (with keyword fallback), adding hybrid semantic+keyword troubleshooting scoring, and adding a passive pipeline-context tool for prompt injection. It also updates Python deps and the Docker image to support and pre-cache the embedding model.

Changes:

  • Add shared embedding utilities (tools/embeddings.py) and use them for semantic/hybrid retrieval in docs search and troubleshooting.
  • Add get_context_for_pipeline passive context tool (+ tests) and register it with the MCP server.
  • Update packaging (pyproject.toml) and Docker build to include and pre-cache the MiniLM model for non-root runtime.

Reviewed changes

Copilot reviewed 13 out of 13 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
olive-mcp-server/olive_mcp_server/tools/embeddings.py New shared lazy-loaded MiniLM embedding + cosine + semantic_search utilities
olive-mcp-server/olive_mcp_server/tools/docs_search.py Semantic search + keyword fallback; live-doc caching/indexing updates
olive-mcp-server/olive_mcp_server/tools/troubleshooting.py Hybrid semantic+keyword troubleshooting scoring + embedding index cache
olive-mcp-server/olive_mcp_server/tools/passive_context.py New pipeline-aware KB snippet retrieval tool
olive-mcp-server/olive_mcp_server/tools/init.py Expose/register new passive context tool
olive-mcp-server/olive_mcp_server/mcp_server.py Register new tool and expose tool import mapping
olive-mcp-server/pyproject.toml Add embedding dependencies (sentence-transformers, numpy)
olive-mcp-server/Dockerfile Pre-cache MiniLM in builder and ship HF cache for non-root user
olive-mcp-server/tests/test_embeddings.py New unit tests for embedding utilities and lazy load behavior
olive-mcp-server/tests/test_docs_search_semantic.py New tests for semantic docs search, cache invalidation, and live-doc race handling
olive-mcp-server/tests/test_troubleshooting_hybrid.py New tests for hybrid semantic+keyword troubleshooting behavior
olive-mcp-server/tests/test_passive_context.py New tests for get_context_for_pipeline return shape + confidence behavior
olive-mcp-server/tests/test_integration.py Update tool-list expectations for new tool registration
Suppressed comments (2)

olive-mcp-server/olive_mcp_server/tools/docs_search.py:251

  • Potential cache/index race: _get_live_index captures pages and _LAST_FETCH_TIME in separate critical sections, so a concurrent refresh can change _LAST_FETCH_TIME after pages is returned. That can stamp embeddings with a fetch_time that doesn't match the snippet content, causing stale embeddings to be treated as fresh.
    pages, fetch_time = _fetch_live_docs()

    with _LIVE_INDEX_LOCK:
        if (

olive-mcp-server/olive_mcp_server/tools/docs_search.py:349

  • The block that forces at least one live: result into results breaks strict top-k ranking: it can replace a higher-relevance local result with a lower-relevance live one. Since the sort key already prefers live docs on ties, this extra replacement should be removed (or gated by relevance).
        live_results
        and top_k > 0
        and results
        and not any(r["source"].startswith("live:") for r in results)
    ):
        best_live = max(live_results, key=lambda x: x["relevance"])
        if top_k == 1:
            results = [best_live]
        else:
            results = results[: top_k - 1] + [best_live]
            results.sort(
                key=lambda x: (-x["relevance"], 0 if x["source"].startswith("live:") else 1)
            )

    return {

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/tests/test_integration.py Outdated
@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo Fixer

✅ Committed (3) · ☑ Fixed (3)

Grey Divider

Commits pushed directly to this PR — no separate fix PR opened.

Process — 3 fixed
  • ☑ Fixed: Live cache snapshot race
  • ☑ Fixed: ad7154ee-5554-45fa-8bf9-b402ace555cd
  • ☑ Fixed: Unused Path import

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d91a90a04b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 14

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@olive-mcp-server/Dockerfile`:
- Around line 36-37: Update the Dockerfile’s Hugging Face cache copy to set
ownership directly with COPY --chown=mcp:mcp, remove the subsequent recursive
chown layer, and define HF_HUB_OFFLINE=1 in the runtime image so Hugging Face
lookups never require network access.
- Around line 13-16: Update the Dockerfile’s pip install command to constrain
torch to an explicit +cpu build from the PyTorch CPU index, while retaining the
regular PyPI index for other dependencies and the existing mcp<2 constraint.
Replace the current unprioritized torch resolution in the install flow without
changing unrelated package installation behavior.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py`:
- Around line 189-194: Update both exception handlers in the document search
flow, including the handler around the shown results logic and the matching
handler near the live-fetch fallback, to log the caught exception at debug or
warning level before invoking the existing keyword fallback. Preserve the
current fallback behavior and avoid leaving either exception silently swallowed.
- Around line 79-89: Update _kb_max_mtime to return a fingerprint containing
both the maximum searchable JSON mtime and the searchable file count, so
additions and deletions invalidate the cache even when the maximum mtime is
unchanged. Change _KB_INDEX_MTIME initialization and the related test fakes in
test_docs_search_semantic.py to use the new fingerprint type consistently,
preserving exclusion and OSError handling.
- Around line 336-349: Update the live-result insertion logic in the search
function so a live result is only selected when its relevance is competitive
with the local results it would replace. For top_k == 1, keep the existing local
result when it scores higher than best_live; for larger top_k values, do not
evict the lowest-ranked retained local result unless best_live meets or exceeds
its relevance, while preserving the existing relevance ordering and live-result
tie-breaking.
- Around line 176-194: Update _search_local to retain and reuse the
knowledge-base entries returned by get_or_build_kb_index when semantic_search
produces no results. Pass those entries to _keyword_search instead of calling
_load_kb_text; only invoke _load_kb_text when get_or_build_kb_index or semantic
search raises an exception and no indexed entries are available.

In `@olive-mcp-server/olive_mcp_server/tools/embeddings.py`:
- Line 10: Update the typing import in embeddings.py to import Sequence from
collections.abc instead of typing, while retaining Any from typing and
preserving the existing usage.
- Around line 22-32: Update _get_model to load SentenceTransformer strictly from
the local cache by passing local_files_only=True (and configure HF_HUB_OFFLINE=1
if required by the runtime), then invoke _get_model during application startup
so the model is warmed and missing-cache failures occur before tools serve
requests; retain the existing _model_lock singleton behavior.

In `@olive-mcp-server/olive_mcp_server/tools/passive_context.py`:
- Around line 96-106: Update the exception handler around get_or_build_kb_index
and semantic_search in the passive-context retrieval flow to log the caught
exception at warning level before returning the existing empty results. Preserve
the current return shape and fallback values, and use the module’s existing
logger if available.

In `@olive-mcp-server/olive_mcp_server/tools/troubleshooting.py`:
- Around line 416-455: Extract the duplicated diagnosis response construction
from the empty-message early return and the main return path into a shared
helper, passing best, matched_entry, matched_domain, applyable, pass_name, and
freq. Update both paths to call this helper, preserving the existing applyable
fallback behavior and all payload fields including frequency and relevant
quirks.
- Line 147: Remove the unnecessary global declaration for _ts_index_cache from
the surrounding troubleshooting function; retain the existing dictionary
mutation unchanged so lint passes without altering behavior.
- Around line 126-176: The troubleshooting loaders used by
_get_troubleshooting_index, specifically _cached_troubleshooting() and
_cached_studio_troubleshooting(), must detect JSON file mtime changes and reload
or invalidate their cached entries before indexing. Add an end-to-end test that
edits the troubleshooting file and verifies subsequent index loading reflects
the updated content rather than stale cached entries.

In `@olive-mcp-server/pyproject.toml`:
- Around line 14-15: Update the dependency bounds in pyproject.toml for
sentence-transformers and numpy to stop at the first untested release, rather
than allowing untested major versions; avoid a <6 upper bound unless CI covers
the 5.x API. Add CI checks that assert the resolved versions remain within the
tested ranges.

In `@olive-mcp-server/tests/test_docs_search_semantic.py`:
- Around line 211-252: Update
test_live_fetch_generation_ignores_stale_completion to call
docs_search._fetch_live_docs() with the patched fetcher, have fake_fetch
increment _LIVE_FETCH_GENERATION under _LIVE_FETCH_LOCK while the request is in
flight, and assert the stale result is not published. Extend the autouse fixture
to reset _LIVE_CACHE, _LAST_FETCH_TIME, and _LIVE_FETCH_GENERATION before and
after tests so live-cache state cannot leak.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 78bea72c-ca1c-4bd6-a90f-ebce1065c5a5

📥 Commits

Reviewing files that changed from the base of the PR and between 01d40df and d91a90a.

📒 Files selected for processing (13)
  • olive-mcp-server/Dockerfile
  • olive-mcp-server/olive_mcp_server/mcp_server.py
  • olive-mcp-server/olive_mcp_server/tools/__init__.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/passive_context.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
  • olive-mcp-server/pyproject.toml
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/tests/test_embeddings.py
  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/tests/test_passive_context.py
  • olive-mcp-server/tests/test_troubleshooting_hybrid.py
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

  • Trackdubllc/Trackdub (manual)
  • tonythethompson/QuickShell (manual)
  • tonythethompson/numan (manual)
  • tonythethompson/dependency-chain-substrate (manual)
📜 Review details
⏰ Context from checks skipped due to timeout. (4)
  • GitHub Check: cubic · AI code reviewer
  • GitHub Check: copilot-pull-request-reviewer
  • GitHub Check: Greptile Review
  • GitHub Check: python-tests
🧰 Additional context used
📓 Path-based instructions (4)
olive-mcp-server/pyproject.toml

📄 CodeRabbit inference engine (AGENTS.md)

Pin the mcp dependency to a version below 2 because mcp 2.x removes mcp.server.fastmcp and breaks imports and tests.

Files:

  • olive-mcp-server/pyproject.toml
**/*

📄 CodeRabbit inference engine (AGENTS.md)

Use the repository’s prescribed validation commands and preserve the CI order: lint, unit tests, server tests, integration tests, component tests, recipe validation, build, artifact assertion, production smoke testing, and CodeQL.

Files:

  • olive-mcp-server/pyproject.toml
  • olive-mcp-server/olive_mcp_server/mcp_server.py
  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/tests/test_passive_context.py
  • olive-mcp-server/olive_mcp_server/tools/passive_context.py
  • olive-mcp-server/Dockerfile
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/olive_mcp_server/tools/__init__.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/tests/test_embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/tests/test_troubleshooting_hybrid.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
olive-mcp-server/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Maintain compatibility with Python >=3.10 and the FastMCP server architecture when modifying the Olive MCP server.

Files:

  • olive-mcp-server/olive_mcp_server/mcp_server.py
  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/tests/test_passive_context.py
  • olive-mcp-server/olive_mcp_server/tools/passive_context.py
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/olive_mcp_server/tools/__init__.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/tests/test_embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/tests/test_troubleshooting_hybrid.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
olive-mcp-server/tests/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Run and maintain pytest coverage for the Olive MCP tools using python -m pytest tests -q.

Files:

  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/tests/test_passive_context.py
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/tests/test_embeddings.py
  • olive-mcp-server/tests/test_troubleshooting_hybrid.py
🪛 ast-grep (0.45.0)
olive-mcp-server/tests/test_embeddings.py

[error] 21-26: Command coming from incoming request
Context: subprocess.run(
[sys.executable, "-c", code],
capture_output=True,
text=True,
check=False,
)
Note: [CWE-78] Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection').

(subprocess-from-request)

🪛 GitHub Check: CodeFactor
olive-mcp-server/olive_mcp_server/tools/docs_search.py

[notice] 191-192: olive-mcp-server/olive_mcp_server/tools/docs_search.py#L191-L192
Try, Except, Pass detected. (B110)

🪛 Ruff (0.16.0)
olive-mcp-server/olive_mcp_server/tools/embeddings.py

[warning] 10-10: Import from collections.abc instead: Sequence

Import from collections.abc

(UP035)

olive-mcp-server/olive_mcp_server/tools/docs_search.py

[error] 191-192: try-except-pass detected, consider logging the exception

(S110)


[error] 226-227: try-except-pass detected, consider logging the exception

(S110)

olive-mcp-server/olive_mcp_server/tools/troubleshooting.py

[warning] 147-147: Using global for _ts_index_cache but no assignment is done

(PLW0602)

🔍 Remote MCP Context7, DeepWiki, GitHub Copilot

Review-relevant context

  • PR #73 is open, one commit, and changes 13 files (+1,689/−156). At retrieval time, validate had failed while Python tests and Docker build were still running; no review threads existed.
  • MCP history: PR #8 established the original 12-tool server and KB; PR #12 centralized KB loading and normalization; PR #65 introduced lazy tool imports and the diagnose_error dispatcher. These are the relevant compatibility baselines for the new registration.
  • sentence-transformers and numpy are added as unpinned lower-bounded runtime dependencies. The Docker build pre-caches all-MiniLM-L6-v2, copies installed packages/cache into the runtime image, and sets HF_HOME for the non-root user.
  • PyTorch’s documented CPU installation uses --index-url https://download.pytorch.org/whl/cpu; this PR uses --extra-index-url, so the resolver’s selected torch wheel should be verified.
  • Sentence Transformers supports local_files_only=True, but the implementation constructs SentenceTransformer(MODEL_NAME, device="cpu") without it; a missing/incomplete cache can therefore still trigger a download.
  • Cache review focus: local KB invalidation compares only the maximum JSON mtime, so changes to a non-newest file may not invalidate the index. Troubleshooting loaders are also lru_cache-backed, meaning mtime changes can rebuild embeddings from already-cached entries.

DeepWiki was attempted but could not index tonythethompson/Olive-Studio; the requested Babel-Player cross-language architecture is not involved in this Python-only PR.

🔇 Additional comments (24)
olive-mcp-server/olive_mcp_server/tools/troubleshooting.py (13)

500-516: 📐 Maintainability & Code Quality | ⚡ Quick win

See the duplication comment at lines 416-455; this block should call the same extracted helper.


4-25: LGTM!


35-54: LGTM!


102-124: LGTM!


178-214: LGTM!


240-252: LGTM!


273-302: LGTM!


312-312: LGTM!


334-338: LGTM!


347-387: LGTM!


498-498: LGTM!


525-525: LGTM!


541-541: LGTM!

olive-mcp-server/tests/test_troubleshooting_hybrid.py (1)

1-380: LGTM!

olive-mcp-server/olive_mcp_server/tools/embeddings.py (1)

40-99: LGTM!

Also applies to: 102-122, 125-165

olive-mcp-server/tests/test_embeddings.py (1)

11-30: LGTM!

Also applies to: 33-63, 66-104, 124-153

olive-mcp-server/olive_mcp_server/tools/docs_search.py (2)

32-47: LGTM!

Also applies to: 92-137, 149-173, 197-230, 244-277


351-354: 🗄️ Data Integrity & Integration

Resolve the count response contract before merge

Confirm whether count denotes total matches or returned results. The shown code sets it to len(combined), while results is truncated to top_k; align count with results or add total_matches.

olive-mcp-server/tests/test_docs_search_semantic.py (1)

24-44: LGTM!

Also applies to: 47-104, 108-149, 152-208

olive-mcp-server/olive_mcp_server/tools/passive_context.py (1)

15-29: LGTM!

Also applies to: 32-56, 59-94, 108-118

olive-mcp-server/olive_mcp_server/tools/__init__.py (1)

34-34: LGTM!

Also applies to: 178-178

olive-mcp-server/olive_mcp_server/mcp_server.py (1)

45-48: LGTM!

olive-mcp-server/tests/test_integration.py (1)

14-21: LGTM!

olive-mcp-server/tests/test_passive_context.py (1)

12-20: LGTM!

Also applies to: 23-34, 37-79, 82-103, 106-135, 138-145

Comment thread olive-mcp-server/Dockerfile Outdated
Comment thread olive-mcp-server/Dockerfile Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
Comment thread olive-mcp-server/pyproject.toml Outdated
Comment thread olive-mcp-server/tests/test_docs_search_semantic.py Outdated
tonythethompson and others added 2 commits August 1, 2026 07:34
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
olive-mcp-server/olive_mcp_server/tools/docs_search.py (1)

248-275: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Discard an old live-index build after a concurrent refresh.

_get_live_index captures fetch_time, builds embeddings outside _LIVE_INDEX_LOCK, and publishes the result without checking the current _LAST_FETCH_TIME. A newer fetch can complete during encoding, allowing the older build to overwrite the newer cache.

After build_kb_index, read _LAST_FETCH_TIME under _LIVE_FETCH_LOCK. Discard and retry the build when the timestamp changed. Add a concurrent-refresh test.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py` around lines 248 -
275, Update _get_live_index so that after build_kb_index completes, it reads
_LAST_FETCH_TIME while holding _LIVE_FETCH_LOCK and discards/retries the build
if it differs from the captured fetch_time, before publishing under
_LIVE_INDEX_LOCK. Add a test covering a concurrent refresh during embedding
construction and verify the stale build is not cached or returned.
♻️ Duplicate comments (2)
olive-mcp-server/olive_mcp_server/tools/embeddings.py (2)

10-10: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Import Iterable from collections.abc.

Ruff reports UP035 for from typing import Any, Iterable. Keep Any in typing and move Iterable to collections.abc. This lint failure blocks the first CI stage.

Proposed fix
-from typing import Any, Iterable
+from collections.abc import Iterable
+from typing import Any
#!/bin/bash
set -euo pipefail
rg -n '^from (typing|collections\.abc) import' olive-mcp-server/olive_mcp_server/tools/embeddings.py

As per coding guidelines: “preserve the CI order: lint, unit tests, server tests, integration tests, component tests, recipe validation, build, artifact assertion, production smoke testing, and CodeQL.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/embeddings.py` at line 10, Update the
imports in embeddings.py so Any remains imported from typing while Iterable is
imported from collections.abc, resolving Ruff’s UP035 lint violation.

Sources: Coding guidelines, Linters/SAST tools


22-32: 🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

Fail fast when the MiniLM cache is missing.

_get_model constructs SentenceTransformer(MODEL_NAME, device="cpu") without local_files_only=True. A missing or incomplete cache can trigger a Hugging Face download during the first semantic request. This makes offline and non-Docker deployments fail late.

Pass local_files_only=True or set HF_HUB_OFFLINE=1 in the runtime. Warm _get_model() during startup so cache failures occur before tools serve requests.

#!/bin/bash
set -euo pipefail
rg -n -A12 -B4 'def _get_model|SentenceTransformer\(' olive-mcp-server/olive_mcp_server/tools/embeddings.py
rg -n 'local_files_only|HF_HUB_OFFLINE|_get_model\(' olive-mcp-server --glob '*.py'
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/embeddings.py` around lines 22 - 32,
Update _get_model so SentenceTransformer loads the MiniLM model with
local_files_only=True, preventing network downloads and making missing or
incomplete caches fail immediately. Also invoke _get_model during application
startup, using the existing startup initialization path, so cache validation
completes before semantic tools accept requests.

Source: MCP tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py`:
- Around line 248-275: Update _get_live_index so that after build_kb_index
completes, it reads _LAST_FETCH_TIME while holding _LIVE_FETCH_LOCK and
discards/retries the build if it differs from the captured fetch_time, before
publishing under _LIVE_INDEX_LOCK. Add a test covering a concurrent refresh
during embedding construction and verify the stale build is not cached or
returned.

---

Duplicate comments:
In `@olive-mcp-server/olive_mcp_server/tools/embeddings.py`:
- Line 10: Update the imports in embeddings.py so Any remains imported from
typing while Iterable is imported from collections.abc, resolving Ruff’s UP035
lint violation.
- Around line 22-32: Update _get_model so SentenceTransformer loads the MiniLM
model with local_files_only=True, preventing network downloads and making
missing or incomplete caches fail immediately. Also invoke _get_model during
application startup, using the existing startup initialization path, so cache
validation completes before semantic tools accept requests.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: d3715a51-3cf5-4554-a7f1-8aed3491cd05

📥 Commits

Reviewing files that changed from the base of the PR and between d91a90a and 57919e8.

📒 Files selected for processing (3)
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

  • Trackdubllc/Trackdub (manual)
  • tonythethompson/QuickShell (manual)
  • tonythethompson/numan (manual)
  • tonythethompson/dependency-chain-substrate (manual)
💤 Files with no reviewable changes (1)
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
📜 Review details
⏰ Context from checks skipped due to timeout. (1)
  • GitHub Check: Greptile Review
⚠️ CI failures not shown inline (4)

GitHub Actions: CI / python-tests: feat(mcp): embedding-based semantic retrieval and passive context

Conclusion: failure

View job details

##[group]Run python -m pytest tests -q --tb=short
 �[36;1mpython -m pytest tests -q --tb=short�[0m
 shell: /usr/bin/bash -e {0}
 env:
   pythonLocation: /opt/hostedtoolcache/Python/3.12.13/x64
   PKG_CONFIG_PATH: /opt/hostedtoolcache/Python/3.12.13/x64/lib/pkgconfig
   Python_ROOT_DIR: /opt/hostedtoolcache/Python/3.12.13/x64
   Python2_ROOT_DIR: /opt/hostedtoolcache/Python/3.12.13/x64
   Python3_ROOT_DIR: /opt/hostedtoolcache/Python/3.12.13/x64
   LD_LIBRARY_PATH: /opt/hostedtoolcache/Python/3.12.13/x64/lib
 ##[endgroup]
 ........................................................................ [ 38%]
 ...............................................F........................ [ 77%]
 ..........................................                               [100%]
 =================================== FAILURES ===================================
 _______________ test_search_olive_documentation_with_live_source _______________
 tests/test_tools.py:173: in test_search_olive_documentation_with_live_source
     assert any("live:" in r["source"] for r in result["results"])
 E   assert False
 E    +  where False = any(<generator object test_search_olive_documentation_with_live_source.<locals>.<genexpr> at 0x7f06b46860c0>)
 =========================== short test summary info ============================
 FAILED tests/test_tools.py::test_search_olive_documentation_with_live_source - assert False
  +  where False = any(<generator object test_search_olive_documentation_with_live_source.<locals>.<genexpr> at 0x7f06b46860c0>)
 1 failed, 185 passed in 18.72s
 ##[error]Process completed with exit code 1.

GitHub Actions: CI / validate: feat(mcp): embedding-based semantic retrieval and passive context

Conclusion: failure

View job details

##[group]Run pnpm lint
 �[36;1mpnpm lint�[0m
 shell: /usr/bin/bash -e {0}
 env:
   PNPM_HOME: /home/runner/setup-pnpm/node_modules/.bin
 ##[endgroup]
 $ tsc --noEmit && eslint
 ##[error]src/lib/auditAutofix.ts(220,55): error TS2769: No overload matches this call.

GitHub Actions: CI / 1_python-tests.txt: feat(mcp): embedding-based semantic retrieval and passive context

Conclusion: failure

View job details

##[group]Run python -m pytest tests -q --tb=short
 �[36;1mpython -m pytest tests -q --tb=short�[0m
 shell: /usr/bin/bash -e {0}
 env:
   pythonLocation: /opt/hostedtoolcache/Python/3.12.13/x64
   PKG_CONFIG_PATH: /opt/hostedtoolcache/Python/3.12.13/x64/lib/pkgconfig
   Python_ROOT_DIR: /opt/hostedtoolcache/Python/3.12.13/x64
   Python2_ROOT_DIR: /opt/hostedtoolcache/Python/3.12.13/x64
   Python3_ROOT_DIR: /opt/hostedtoolcache/Python/3.12.13/x64
   LD_LIBRARY_PATH: /opt/hostedtoolcache/Python/3.12.13/x64/lib
 ##[endgroup]
 ........................................................................ [ 38%]
 ...............................................F........................ [ 77%]
 ..........................................                               [100%]
 =================================== FAILURES ===================================
 _______________ test_search_olive_documentation_with_live_source _______________
 tests/test_tools.py:173: in test_search_olive_documentation_with_live_source
     assert any("live:" in r["source"] for r in result["results"])
 E   assert False
 E    +  where False = any(<generator object test_search_olive_documentation_with_live_source.<locals>.<genexpr> at 0x7f06b46860c0>)
 =========================== short test summary info ============================
 FAILED tests/test_tools.py::test_search_olive_documentation_with_live_source - assert False
  +  where False = any(<generator object test_search_olive_documentation_with_live_source.<locals>.<genexpr> at 0x7f06b46860c0>)
 1 failed, 185 passed in 18.72s
 ##[error]Process completed with exit code 1.

GitHub Actions: CI / 3_validate.txt: feat(mcp): embedding-based semantic retrieval and passive context

Conclusion: failure

View job details

##[group]Run pnpm lint
 �[36;1mpnpm lint�[0m
 shell: /usr/bin/bash -e {0}
 env:
   PNPM_HOME: /home/runner/setup-pnpm/node_modules/.bin
 ##[endgroup]
 $ tsc --noEmit && eslint
 ##[error]src/lib/auditAutofix.ts(220,55): error TS2769: No overload matches this call.
🧰 Additional context used
📓 Path-based instructions (2)
olive-mcp-server/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Maintain compatibility with Python >=3.10 and the FastMCP server architecture when modifying the Olive MCP server.

Files:

  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
**/*

📄 CodeRabbit inference engine (AGENTS.md)

Use the repository’s prescribed validation commands and preserve the CI order: lint, unit tests, server tests, integration tests, component tests, recipe validation, build, artifact assertion, production smoke testing, and CodeQL.

Files:

  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
🪛 Ruff (0.16.0)
olive-mcp-server/olive_mcp_server/tools/embeddings.py

[warning] 10-10: Import from collections.abc instead: Iterable

Import from collections.abc

(UP035)

🔍 Remote MCP Context7, GitHub Copilot

Additional review context

  • PR #73’s latest checks show Docker build, CodeQL, security, and CodeFactor passing, but validate and python-tests failing.
  • Sentence Transformers supports local_files_only=True; this PR’s model construction does not use it, so a missing cache can still trigger network access at runtime.
  • The implementation’s encode_texts now materializes iterables before checking emptiness, avoiding model initialization for empty generators.
  • PyTorch’s documented CPU-only installation uses --index-url https://download.pytorch.org/whl/cpu; the Dockerfile uses --extra-index-url, so the selected wheel should be verified.
  • A prior automated review identified live-cache snapshot consistency and an unused import; the PR comments report both as fixed in follow-up commits.
  • The relevant architectural baseline is PR #65, which established the lazy MCP tool registry, documentation search, troubleshooting dispatcher, and Olive Studio’s MCP knowledge-gathering path.
🔇 Additional comments (3)
olive-mcp-server/olive_mcp_server/tools/docs_search.py (2)

79-89: 🗄️ Data Integrity & Integration

Verify the KB fingerprint contract before merge.

olive-mcp-server/tests/test_docs_search_semantic.py at Line 172-208 still returns a float from the fake _kb_max_mtime() and compares _KB_INDEX_MTIME to 200.0. If production now returns a composite fingerprint, the test contract is stale. If production still uses only maximum mtime, deleting a non-newest file does not invalidate the index.

Align _kb_max_mtime, _KB_INDEX_MTIME, and the test fake. Include file-count or file-identity data in the fingerprint.

#!/bin/bash
set -euo pipefail
rg -n -A25 -B5 'def _kb_max_mtime|_KB_INDEX_MTIME|get_or_build_kb_index' olive-mcp-server/olive_mcp_server/tools/docs_search.py
rg -n -A30 -B5 'test_kb_stale_build_does_not_poison_cache' olive-mcp-server/tests/test_docs_search_semantic.py

Source: Pipeline failures


197-225: LGTM!

Also applies to: 229-230, 297-300

olive-mcp-server/olive_mcp_server/tools/embeddings.py (1)

40-57: LGTM!

- Add cu130/cu132 to UIState.cudaVersion type to match auditAutofix's
  CUDA_VERSIONS set (TS2769 build failure)
- Fix test_search_olive_documentation_with_live_source: _fetch_live_docs
  now returns (pages, fetch_time) tuple, test mock still returned bare dict

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

5 issues found and verified against the latest diff

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="olive-mcp-server/olive_mcp_server/tools/troubleshooting.py">

<violation number="1" location="olive-mcp-server/olive_mcp_server/tools/troubleshooting.py:152">
P2: KB edits made while the server is running never affect troubleshooting results: an mtime change only rebuilds embeddings for the stale `@lru_cache` loader output. Invalidate/reload the parsed troubleshooting caches when these files change before building the new index.</violation>
</file>

<file name="olive-mcp-server/tests/test_docs_search_semantic.py">

<violation number="1" location="olive-mcp-server/tests/test_docs_search_semantic.py:232">
P3: This test doesn't exercise the code it claims to cover. It installs fake_fetch and patches fetch_official_docs, but never calls _fetch_live_docs; instead it manually writes generation-protected cache entries, reimplementing the guard by hand. As written it would pass even if the real generation guard in _fetch_live_docs regressed, giving false confidence in the out-of-order-completion fix. Consider driving the real path (e.g., call _fetch_live_docs twice under controlled fetch/TTL conditions using a real fetch queue) so the actual guard logic is what's verified, and remove the now-unused fake_fetch/patch setup.</violation>
</file>

<file name="olive-mcp-server/olive_mcp_server/tools/docs_search.py">

<violation number="1" location="olive-mcp-server/olive_mcp_server/tools/docs_search.py:89">
P2: KB updates to a non-newest file can remain invisible because the cache fingerprint is only the maximum mtime. Track every searchable filename and mtime (preferably `st_mtime_ns`) so any file-set change rebuilds the index.</violation>

<violation number="2" location="olive-mcp-server/olive_mcp_server/tools/docs_search.py:222">
P2: Concurrent cache refreshes can discard the only successful live-doc response when a later request fails, making default live search return no fresh docs. Coordinate a single in-flight refresh, or retain the latest successful completion rather than gating publication solely on the latest started request.</violation>

<violation number="3" location="olive-mcp-server/olive_mcp_server/tools/docs_search.py:339">
P2: Results no longer contain the top-ranked combined hits when every live result scores below the local cutoff: this branch replaces a higher-relevance local result with `best_live`. Keep the sorted `combined[:top_k]` list unless the API explicitly defines source diversity over ranking.</violation>
</file>

Shadow auto-approve: would not auto-approve because issues were found.

Re-trigger cubic

Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py
Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py Outdated
Comment thread olive-mcp-server/tests/test_docs_search_semantic.py
@greptile-apps

greptile-apps Bot commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR wires MiniLM-L6-v2 semantic embeddings into docs search and troubleshooting, adds a new get_context_for_pipeline passive-context tool, pre-caches the model in Docker, and fixes several existing issues (live-doc stale-overwrite race, BatchProcessingPanel test, chatAction salvage false-positives, docsSearchSufficient relevance threshold).

  • New Python tools: embeddings.py (lazy CPU singleton, cosine scoring), semantic/hybrid upgrade to docs_search.py and troubleshooting.py, and new passive_context.py — all with thread-safe generation-token and fingerprint caches.
  • Dockerfile: CPU-only torch is forced via --index-url, MiniLM is pre-cached in the builder stage and chown-copied to the mcp user, HF_HUB_OFFLINE=1 set for offline runtime.
  • Frontend fixes: salvageChatActionPatchFromLooseJson now matches quant/convert keywords in step/action/task value strings while blocking negated phrases; docsSearchSufficient correctly accepts any positive normalized relevance score; cu130/cu132 removed from validation sets to match the UIState[\"cudaVersion\"] type union.

Confidence Score: 5/5

Safe to merge; the only finding is a benign type mismatch in a test fixture that does not affect test outcomes or production behaviour.

The threading model in docs_search (generation tokens, double-checked locking outside/inside the KB index lock) and the troubleshooting fingerprint cache are correctly implemented and covered by dedicated race-condition tests. The Docker cache-copy and offline-mode env vars are straightforward. The one test-fixture issue is harmless in practice because _KB_EMBEDDINGS is also reset to None, short-circuiting the comparison before the mismatched value is ever read.

Files Needing Attention: olive-mcp-server/tests/test_passive_context.py — the autouse fixture resets _KB_INDEX_MTIME to a float instead of the expected tuple.

Important Files Changed

Filename Overview
olive-mcp-server/olive_mcp_server/tools/embeddings.py New shared embedding module: lazy-load MiniLM singleton, encode_texts/encode_query, cosine_similarity_scores, build_kb_index, and semantic_search — all well-tested with mocks.
olive-mcp-server/olive_mcp_server/tools/docs_search.py Upgraded to semantic search with mtime+count invalidation for KB index and generation-token guard for live-doc fetch races; keyword fallback preserved. Threading logic is carefully layered (build outside lock, double-check inside).
olive-mcp-server/olive_mcp_server/tools/troubleshooting.py Hybrid 0.6-semantic/0.4-keyword scoring with fingerprint-keyed index cache (max 8 entries); empty error_message short-circuits at multiple layers; _build_diagnosis_payload consolidates the response shape.
olive-mcp-server/olive_mcp_server/tools/passive_context.py New get_context_for_pipeline tool: normalizes pass descriptors, queries shared KB index, returns context_snippets + confidence clamped to [0,1]; straightforward and well-tested.
olive-mcp-server/Dockerfile CPU-only torch install via index-url, MiniLM pre-cached in builder and chown-copied to runtime user, HF_HOME and HF_HUB_OFFLINE=1 set for the mcp user. Image is now ~1.5–3 GB (documented).
src/lib/chatActions.ts salvageChatActionPatchFromLooseJson extended to also match conversion/quant keywords in step/action/task value strings; isNegated guard prevents skip-style false positives. cu130/cu132 removed to match UIState["cudaVersion"] type union.
src/server/services/ai/oliveMcpKnowledge.ts docsSearchSufficient threshold changed from relevance >= 1 to relevance > 0, correctly reflecting the new [0,1] normalized semantic score scale.
olive-mcp-server/tests/test_passive_context.py Good coverage of empty/string/dict pass descriptors and confidence bounds, but the autouse fixture resets _KB_INDEX_MTIME to a plain float -1.0 instead of the correct tuple (-1.0, -1).

Reviews (4): Last reviewed commit: "fix: don't widen cudaVersion type to mat..." | Re-trigger Greptile

Comment thread olive-mcp-server/olive_mcp_server/tools/docs_search.py
Comment thread olive-mcp-server/olive_mcp_server/tools/embeddings.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
Comment thread olive-mcp-server/olive_mcp_server/tools/troubleshooting.py

@coderabbitai coderabbitai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
olive-mcp-server/olive_mcp_server/tools/docs_search.py (3)

249-255: 🚀 Performance & Scalability | 🟡 Minor | ⚡ Quick win

Cache empty live results.

When _split_live_snippets returns an empty list, _LIVE_SNIPPETS is falsy. The cache-hit checks at Line 252 and Line 266 then miss even when _LIVE_EMBED_CACHE_TIME == fetch_time. Each live search rebuilds the empty index.

Use an explicit None sentinel or test _LIVE_SNIPPETS is not None.

Also applies to: 263-269

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py` around lines 249 -
255, Update the cache-hit checks in the live search flow around _LIVE_SNIPPETS
to treat an empty list as a valid cached result by testing whether
_LIVE_SNIPPETS is not None instead of relying on truthiness, including both
checks near the shown branches. Preserve the existing _LIVE_EMBEDDINGS and
_LIVE_EMBED_CACHE_TIME == fetch_time conditions.

263-273: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Prevent an older live snapshot from replacing a newer index.

This method builds outside _LIVE_INDEX_LOCK. A call can capture an older fetch_time, while another call publishes a newer embedding cache. The older call can then overwrite it at Line 271 through Line 273 because the publication path does not compare the current fetch generation.

Compare the current generation before publication. Discard or rebuild older snapshots.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py` around lines 263 -
273, Update the publication block guarded by _LIVE_INDEX_LOCK to compare the
current fetch generation with _LIVE_EMBED_CACHE_TIME before assigning
_LIVE_SNIPPETS, _LIVE_EMBEDDINGS, and _LIVE_EMBED_CACHE_TIME. If the existing
cache is newer than the captured fetch_time, discard the stale snapshot or
rebuild it; only publish snapshots that are not older than the current
generation.

180-191: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

Enforce offline model loading. The Dockerfile pre-caches MiniLM and grants mcp access, but SentenceTransformer(MODEL_NAME, device="cpu") leaves local_files_only=False. If the cache is missing or incomplete, the first local search can contact Hugging Face before keyword fallback. Set local_files_only=True or load the model during controlled startup, and test semantic search with network access disabled.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py` around lines 180 -
191, Update the model-loading path used by get_or_build_kb_index and
SentenceTransformer to enforce offline-only loading by setting
local_files_only=True. Ensure missing or incomplete cached models fail locally
and allow the existing _keyword_search fallback to run without contacting
Hugging Face, and add coverage for semantic search with network access disabled.

Source: MCP tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@olive-mcp-server/tests/test_tools.py`:
- Around line 169-172: Update the test around search_olive_documentation to
isolate the module-level _LIVE_SNIPPETS, _LIVE_EMBEDDINGS, and
_LIVE_EMBED_CACHE_TIME state before invoking it. Either reset those cache
variables or replace the mock’s fetch_time value with a timestamp unique to this
test, ensuring the live: assertion cannot reuse data from another test.

In `@src/types.ts`:
- Line 95: Remove cu130 and cu132 from the cudaVersion type and every related UI
state, chat action, audit autofix, and run-route input path. Ensure
RESOLVABLE_CUDA_TAGS and inferRequiredPackages remain aligned so only supported
tags through cu128 are accepted and forwarded as CUDA_VERSION; do not add CUDA
13 support.

---

Outside diff comments:
In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py`:
- Around line 249-255: Update the cache-hit checks in the live search flow
around _LIVE_SNIPPETS to treat an empty list as a valid cached result by testing
whether _LIVE_SNIPPETS is not None instead of relying on truthiness, including
both checks near the shown branches. Preserve the existing _LIVE_EMBEDDINGS and
_LIVE_EMBED_CACHE_TIME == fetch_time conditions.
- Around line 263-273: Update the publication block guarded by _LIVE_INDEX_LOCK
to compare the current fetch generation with _LIVE_EMBED_CACHE_TIME before
assigning _LIVE_SNIPPETS, _LIVE_EMBEDDINGS, and _LIVE_EMBED_CACHE_TIME. If the
existing cache is newer than the captured fetch_time, discard the stale snapshot
or rebuild it; only publish snapshots that are not older than the current
generation.
- Around line 180-191: Update the model-loading path used by
get_or_build_kb_index and SentenceTransformer to enforce offline-only loading by
setting local_files_only=True. Ensure missing or incomplete cached models fail
locally and allow the existing _keyword_search fallback to run without
contacting Hugging Face, and add coverage for semantic search with network
access disabled.
🪄 Autofix (Beta)

❌ Autofix failed (check again to retry)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: ec1cf954-4237-48fc-a74e-4da716e0fda1

📥 Commits

Reviewing files that changed from the base of the PR and between 57919e8 and c0786d9.

📒 Files selected for processing (4)
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/tests/test_tools.py
  • src/types.ts
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

  • Trackdubllc/Trackdub (manual)
  • tonythethompson/QuickShell (manual)
  • tonythethompson/numan (manual)
  • tonythethompson/dependency-chain-substrate (manual)
📜 Review details
⏰ Context from checks skipped due to timeout. (2)
  • GitHub Check: Greptile Review
  • GitHub Check: python-tests
🧰 Additional context used
📓 Path-based instructions (6)
olive-mcp-server/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Maintain compatibility with Python >=3.10 and the FastMCP server architecture when modifying the Olive MCP server.

Files:

  • olive-mcp-server/tests/test_tools.py
  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
olive-mcp-server/tests/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Run and maintain pytest coverage for the Olive MCP tools using python -m pytest tests -q.

Files:

  • olive-mcp-server/tests/test_tools.py
  • olive-mcp-server/tests/test_integration.py
**/*

📄 CodeRabbit inference engine (AGENTS.md)

Use the repository’s prescribed validation commands and preserve the CI order: lint, unit tests, server tests, integration tests, component tests, recipe validation, build, artifact assertion, production smoke testing, and CodeQL.

Files:

  • olive-mcp-server/tests/test_tools.py
  • olive-mcp-server/tests/test_integration.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • src/types.ts
src/**/*.{ts,tsx}

📄 CodeRabbit inference engine (CONTRIBUTING.md)

src/**/*.{ts,tsx}: Match existing naming, file layout, and TypeScript patterns in src/.
Put shared recipe logic in src/lib/, especially pipelineValidation.ts, oliveRecipeBuilder.ts, and recipePipeline.ts.

src/**/*.{ts,tsx}: Follow the React performance guidance in docs/REACT_BEST_PRACTICES.md, especially eliminating request waterfalls, avoiding barrel imports, and deferring non-critical third-party libraries.
Do not implement the listed backburner AI providers unless explicitly requested; prefer Custom or openai-compat for OpenAI-shaped hosts.

src/**/*.{ts,tsx}: Keep validation logic in shared libraries rather than duplicating it in UI cell helpers or inspectors.
Split the InputEnvironmentPanel, IHVIntegrationPanel, and ExecutionWorkspace mega-panels into feature folders with colocated hooks and tests.
Keep server and UI AI provider catalogs synchronized, preferably through a shared provider ID list or synchronization test; register new providers in both catalogs.
Add test coverage for recipe-graph/, passCatalog, oliveRecipeHub, jobHistoryStore, and vramEstimate, and strengthen component tests for the large panels.

Files:

  • src/types.ts
**/*.{ts,tsx,js,jsx}

📄 CodeRabbit inference engine (CONTRIBUTING.md)

**/*.{ts,tsx,js,jsx}: Place imports at the top of modules; use inline imports only for a documented circular dependency.
Run linting and ensure typecheck-related CI checks pass before submitting changes.
For UI or server changes, manually smoke-test development startup, recipe loading/building, validation banners, and live execution when execution behavior is touched.

Files:

  • src/types.ts
**/*.{ts,tsx}

📄 CodeRabbit inference engine (AGENTS.md)

**/*.{ts,tsx}: Use the project’s React 19, Vite, Express, and Tauri 2 conventions when modifying TypeScript or TSX application code.
Treat ESLint warnings as acceptable up to the configured limit; only lint errors or a non-zero lint exit indicate failure.

Files:

  • src/types.ts
🔍 Remote MCP Context7, DeepWiki, GitHub Copilot

Relevant review context

  • PR #73 modifies only the Olive MCP server, tests, Docker configuration, and src/types.ts; it does not cross the C#/Python inference boundary or touch workflow/session/host lifecycle code. DeepWiki could not index the referenced repository, so no architectural facts were obtained there.
  • The current diff uses SentenceTransformer(..., device="cpu") and encode(..., normalize_embeddings=True). Current Sentence Transformers documentation confirms both APIs, and also exposes local_files_only, which the implementation does not set; runtime model loading may therefore still attempt network access if the Docker cache is absent or incomplete.
  • The Dockerfile uses --extra-index-url https://download.pytorch.org/whl/cpu while installing all dependencies. The diff does not demonstrate that CUDA wheels cannot still be selected from the primary index; this remains worth validating against the resulting image.
  • The latest retrieved checks show CodeQL, CodeFactor, security, and auto-merge passing, while validate failed; python-tests and docker-build were still in progress at retrieval time.
  • Prior review findings reported in PR comments—live-cache snapshot race, empty-iterator model loading, and unused Path import—are marked fixed in follow-up commits.
  • CodeRabbit reported docstring coverage of 38.24% versus a 60% threshold, although this appeared as a warning in its pre-merge summary.
🔇 Additional comments (3)
olive-mcp-server/olive_mcp_server/tools/docs_search.py (2)

192-193: The previous swallowed-exception finding is still present.

The handler still suppresses embedding or index failures without logging before keyword fallback. This repeats the previous S110/B110 finding.


92-137: LGTM!

olive-mcp-server/tests/test_integration.py (1)

18-20: LGTM!

Comment thread olive-mcp-server/tests/test_tools.py Outdated
Comment thread src/types.ts Outdated
tonythethompson and others added 3 commits August 1, 2026 07:54
- src/types.ts: add missing cu130/cu132 CUDA version literals (unblocks
  tsc/lint, unrelated pre-existing drift vs auditAutofix.ts)
- tests/test_tools.py: fix live-doc test to match _fetch_live_docs' new
  (dict, float) return signature (was silently caught by broad except,
  failing the "live:" source assertion in CI)
- docs_search.py: kb mtime invalidation now includes file count so
  deletions/back-dated additions aren't missed; keyword fallback reuses
  already-loaded KB texts instead of re-reading from disk; log swallowed
  exceptions instead of silent pass; live-doc "freshness" injection no
  longer displaces a strictly better local/kept result (was unconditional
  at top_k==1)
- embeddings.py: fix missing Sequence import (was silently deferred by
  `from __future__ import annotations`, would NameError under runtime
  introspection)
- passive_context.py: log retrieval failures instead of returning an
  indistinguishable "no matches" response
- troubleshooting.py: drop unnecessary `global` on a dict-mutate-only
  cache, bound the fingerprint cache size, extract a shared response
  payload builder for the empty-message and matched-entry paths (they'd
  drifted apart)
- Dockerfile: force the CPU-only torch wheel (extra-index-url doesn't
  prioritize over PyPI's CUDA build), COPY --chown instead of a
  redundant chown -R layer, set HF_HUB_OFFLINE=1
- pyproject.toml: cap sentence-transformers/numpy below their next
  untested majors
- oliveMcpKnowledge.ts: fix docsSearchSufficient's relevance threshold,
  which still assumed the old raw-hit-count scale (>=1) after this PR
  normalized relevance to ~[0,1] — was silently making local KB results
  always look "insufficient"

Verified: 186/186 olive-mcp-server pytest, tsc --noEmit clean, pnpm lint
0 errors, 149/149 server vitest.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The previous version relied on the real embedding model ranking a
single synthetic live snippet above ~2500 real local KB entries for
"calibration data" — legitimately doesn't happen once the forced
live-injection bug was fixed, so the test failed in CI (which has the
real sentence-transformers model) even though it passed locally
(where the model isn't installed and the code falls back to keyword
search). Mock _search_local/_search_live directly instead, matching
the pattern used by the other merge-logic regression tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- docs_search.py:_get_live_index: cache-hit/publish checks used
  `_LIVE_SNIPPETS` truthiness, so an empty (but successfully built)
  live index was never recognized as cached and got rebuilt on every
  call. Also, publish had no ordering guard: a slow build for an
  older fetch_time could complete after a newer generation already
  published and clobber it, silently rolling the live cache back to
  stale content. Both fixed: cache checks now key off
  `_LIVE_EMBEDDINGS is not None`, and publish is skipped when the
  build's fetch_time is older than what's already cached.
- Added test_live_index_does_not_overwrite_newer_generation covering
  the race.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
olive-mcp-server/olive_mcp_server/tools/docs_search.py (2)

226-237: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

The generation guard can discard the only successful fetch.

The guard at Line 230 drops any completion that is not the newest generation. Consider this order: generation 1 starts, generation 2 starts, generation 2 fails, generation 1 returns valid pages. Generation 1 sees my_generation != _LIVE_FETCH_GENERATION and discards its result. _LIVE_CACHE stays empty, and both callers get no live documents even though one fetch succeeded.

Publish when the cache holds nothing, so a newer failure cannot suppress an older success.

♻️ Proposed fix
             if fetched:
                 with _LIVE_FETCH_LOCK:
-                    # Ignore stale completions if a newer fetch was started.
-                    if my_generation == _LIVE_FETCH_GENERATION:
+                    # Ignore stale completions if a newer fetch was started,
+                    # unless nothing has been published yet: a newer failure
+                    # must not suppress an older success.
+                    if my_generation == _LIVE_FETCH_GENERATION or not _LIVE_CACHE:
                         _LIVE_CACHE = fetched
                         _LAST_FETCH_TIME = time.monotonic()

Add a test that lets the newer generation fail while the older one succeeds.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py` around lines 226 -
237, Update the generation guard in the live-fetch completion logic around
_LIVE_CACHE so a successful fetched result is published when the cache is
currently empty, even if my_generation is older than _LIVE_FETCH_GENERATION;
retain the generation check for replacing an existing cache. Add a concurrency
test covering an older successful fetch and newer failed fetch, asserting the
successful pages populate the cache.

336-345: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Return the payload count

count: len(combined) reports merged candidates, not results returned after top_k truncation. Return count: len(results). The only application consumer checks count > 0 and does not depend on the exact total. Add total_matches only if required, and update the exact-key test if you add it.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py` around lines 336 -
345, Update the result payload construction following the `combined[:top_k]`
assignment in the search flow so its `count` field uses `len(results)` rather
than `len(combined)`, reflecting the truncated results returned to callers. Do
not add a separate `total_matches` field unless the implementation requires it.
olive-mcp-server/olive_mcp_server/tools/passive_context.py (1)

99-110: 🗄️ Data Integrity & Integration | 🔵 Trivial | ⚡ Quick win

The failure response is still indistinguishable from a genuine empty result.

The warning log addresses observability on the server. The payload does not. A retrieval failure and a real "nothing matched" both return confidence: 0.0 and snippet_count: 0. The frontend injects this into a system prompt, so a degraded knowledge base reads as an authoritative empty knowledge base. That is fake readiness at the API boundary.

get_context_for_pipeline is a new tool, so adding one field does not break the preserved 14 tool signatures.

♻️ Proposed fix
     try:
         kb_texts, embeddings = get_or_build_kb_index()
         results = semantic_search(
             query,
             kb_texts,
             embeddings,
             top_k,
             threshold=DEFAULT_THRESHOLD,
         )
+        retrieval_ok = True
     except Exception:
         logger.warning("KB retrieval failed for pipeline context", exc_info=True)
         results = []
+        retrieval_ok = False
     return {
         "context_snippets": results,
         "pipeline_summary": pipeline_summary,
         "confidence": confidence,
         "snippet_count": len(results),
+        "retrieval_ok": retrieval_ok,
     }

Set retrieval_ok to True on the early-return path at Lines 91-97 as well, and cover both cases in olive-mcp-server/tests/test_passive_context.py.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@olive-mcp-server/olive_mcp_server/tools/passive_context.py` around lines 99 -
110, Update get_context_for_pipeline so its response includes a retrieval_ok
field that is False when the get_or_build_kb_index or semantic_search try block
fails, while preserving the existing empty results payload; set retrieval_ok to
True on the early-return path and successful retrieval path. Add or update tests
in test_passive_context.py covering both the successful empty-result case and
the failure case.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@olive-mcp-server/Dockerfile`:
- Around line 18-20: Update the Dockerfile’s torch dependency in the pip install
command to an explicit CPU-wheel version available for CPython 3.12 Linux
x86_64, such as torch 2.9.1, and verify compatibility with
sentence-transformers>=3.0.0,<6.

In `@olive-mcp-server/tests/test_docs_search_semantic.py`:
- Around line 262-277: Update the threaded test around _fetch_live_docs so it
verifies t1 completed successfully after release_gen1 is set: assert that t1 is
no longer alive after t1.join(timeout=5), and assert that result_gen1 was
populated before checking the cache. Keep the existing generation-order and
cache assertions unchanged.

In `@src/server/services/ai/oliveMcpKnowledge.ts`:
- Around line 140-144: Add a regression case to the docsSearchSufficient tests
in oliveMcpKnowledge.test.ts using a positive relevance below 1, such as 0.2,
and assert that the result is sufficient. Keep the existing high-relevance
coverage intact so the test specifically distinguishes the normalized-score
behavior from the previous >=1 threshold.

---

Outside diff comments:
In `@olive-mcp-server/olive_mcp_server/tools/docs_search.py`:
- Around line 226-237: Update the generation guard in the live-fetch completion
logic around _LIVE_CACHE so a successful fetched result is published when the
cache is currently empty, even if my_generation is older than
_LIVE_FETCH_GENERATION; retain the generation check for replacing an existing
cache. Add a concurrency test covering an older successful fetch and newer
failed fetch, asserting the successful pages populate the cache.
- Around line 336-345: Update the result payload construction following the
`combined[:top_k]` assignment in the search flow so its `count` field uses
`len(results)` rather than `len(combined)`, reflecting the truncated results
returned to callers. Do not add a separate `total_matches` field unless the
implementation requires it.

In `@olive-mcp-server/olive_mcp_server/tools/passive_context.py`:
- Around line 99-110: Update get_context_for_pipeline so its response includes a
retrieval_ok field that is False when the get_or_build_kb_index or
semantic_search try block fails, while preserving the existing empty results
payload; set retrieval_ok to True on the early-return path and successful
retrieval path. Add or update tests in test_passive_context.py covering both the
successful empty-result case and the failure case.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c938bbaf-6842-4356-bb2f-e54f93667d4c

📥 Commits

Reviewing files that changed from the base of the PR and between c0786d9 and 0286504.

📒 Files selected for processing (9)
  • olive-mcp-server/Dockerfile
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/passive_context.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
  • olive-mcp-server/pyproject.toml
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/tests/test_tools.py
  • src/server/services/ai/oliveMcpKnowledge.ts
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

  • Trackdubllc/Trackdub (manual)
  • tonythethompson/QuickShell (manual)
  • tonythethompson/numan (manual)
  • tonythethompson/dependency-chain-substrate (manual)
📜 Review details
⏰ Context from checks skipped due to timeout. (5)
  • GitHub Check: Greptile Review
  • GitHub Check: python-tests
  • GitHub Check: validate
  • GitHub Check: docker-build
  • GitHub Check: security
🧰 Additional context used
📓 Path-based instructions (7)
src/**/*.{ts,tsx}

📄 CodeRabbit inference engine (CONTRIBUTING.md)

src/**/*.{ts,tsx}: Match existing naming, file layout, and TypeScript patterns in src/.
Put shared recipe logic in src/lib/, especially pipelineValidation.ts, oliveRecipeBuilder.ts, and recipePipeline.ts.

src/**/*.{ts,tsx}: Follow the React performance guidance in docs/REACT_BEST_PRACTICES.md, especially eliminating request waterfalls, avoiding barrel imports, and deferring non-critical third-party libraries.
Do not implement the listed backburner AI providers unless explicitly requested; prefer Custom or openai-compat for OpenAI-shaped hosts.

src/**/*.{ts,tsx}: Keep validation logic in shared libraries rather than duplicating it in UI cell helpers or inspectors.
Split the InputEnvironmentPanel, IHVIntegrationPanel, and ExecutionWorkspace mega-panels into feature folders with colocated hooks and tests.
Keep server and UI AI provider catalogs synchronized, preferably through a shared provider ID list or synchronization test; register new providers in both catalogs.
Add test coverage for recipe-graph/, passCatalog, oliveRecipeHub, jobHistoryStore, and vramEstimate, and strengthen component tests for the large panels.

Files:

  • src/server/services/ai/oliveMcpKnowledge.ts
**/*.{ts,tsx,js,jsx}

📄 CodeRabbit inference engine (CONTRIBUTING.md)

**/*.{ts,tsx,js,jsx}: Place imports at the top of modules; use inline imports only for a documented circular dependency.
Run linting and ensure typecheck-related CI checks pass before submitting changes.
For UI or server changes, manually smoke-test development startup, recipe loading/building, validation banners, and live execution when execution behavior is touched.

Files:

  • src/server/services/ai/oliveMcpKnowledge.ts
**/*.{ts,tsx}

📄 CodeRabbit inference engine (AGENTS.md)

**/*.{ts,tsx}: Use the project’s React 19, Vite, Express, and Tauri 2 conventions when modifying TypeScript or TSX application code.
Treat ESLint warnings as acceptable up to the configured limit; only lint errors or a non-zero lint exit indicate failure.

Files:

  • src/server/services/ai/oliveMcpKnowledge.ts
**/*

📄 CodeRabbit inference engine (AGENTS.md)

Use the repository’s prescribed validation commands and preserve the CI order: lint, unit tests, server tests, integration tests, component tests, recipe validation, build, artifact assertion, production smoke testing, and CodeQL.

Files:

  • src/server/services/ai/oliveMcpKnowledge.ts
  • olive-mcp-server/pyproject.toml
  • olive-mcp-server/tests/test_tools.py
  • olive-mcp-server/Dockerfile
  • olive-mcp-server/olive_mcp_server/tools/passive_context.py
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
olive-mcp-server/pyproject.toml

📄 CodeRabbit inference engine (AGENTS.md)

Pin the mcp dependency to a version below 2 because mcp 2.x removes mcp.server.fastmcp and breaks imports and tests.

Files:

  • olive-mcp-server/pyproject.toml
olive-mcp-server/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Maintain compatibility with Python >=3.10 and the FastMCP server architecture when modifying the Olive MCP server.

Files:

  • olive-mcp-server/tests/test_tools.py
  • olive-mcp-server/olive_mcp_server/tools/passive_context.py
  • olive-mcp-server/tests/test_docs_search_semantic.py
  • olive-mcp-server/olive_mcp_server/tools/embeddings.py
  • olive-mcp-server/olive_mcp_server/tools/docs_search.py
  • olive-mcp-server/olive_mcp_server/tools/troubleshooting.py
olive-mcp-server/tests/**/*.py

📄 CodeRabbit inference engine (AGENTS.md)

Run and maintain pytest coverage for the Olive MCP tools using python -m pytest tests -q.

Files:

  • olive-mcp-server/tests/test_tools.py
  • olive-mcp-server/tests/test_docs_search_semantic.py
🪛 Hadolint (2.14.0)
olive-mcp-server/Dockerfile

[warning] 18-18: Pin versions in pip. Instead of pip install <package> use pip install <package>==<version> or pip install --requirement <requirements file>

(DL3013)

🔍 Remote MCP DeepWiki, GitHub Copilot

Additional review context

  • Unresolved P1/P2 concern: each chat knowledge request launches a new Python MCP subprocess, so module-level model and embedding caches are discarded. Documentation search may re-encode ~2,500 KB leaves on every request, potentially approaching the client’s 45-second timeout.

  • Unresolved correctness concern: troubleshooting’s new mtime-aware embedding cache still calls pre-existing @lru_cache loaders. If troubleshooting JSON changes while the server runs, the index may rebuild from stale parsed entries.

  • Unresolved concurrency concern: if a newer live-doc refresh starts and fails after an older refresh succeeds, the generation guard can discard the only successful response, leaving live search empty.

  • Unresolved integration concern: cu130 and cu132 were added to UIState.cudaVersion, but review tooling found runtime/package resolution support only through cu128; the route may forward unsupported CUDA tags.

  • Current PR checks show python-tests, validate, docker-build, security, and Greptile Review in progress; CodeFactor and auto-merge passed.

  • DeepWiki could not index tonythethompson/Olive-Studio, so it provided no architectural validation.

🔇 Additional comments (12)
olive-mcp-server/olive_mcp_server/tools/troubleshooting.py (2)

138-172: Existing loader-cache invalidation issue.

file_mtime can trigger an index rebuild while load_troubleshooting() still returns entries from its pre-existing loader cache. A JSON edit can therefore re-embed stale entries. This is the same issue recorded in the previous review and should remain a separately scoped follow-up.


53-53: LGTM!

Also applies to: 173-179, 415-442, 467-469, 513-513

olive-mcp-server/Dockerfile (1)

43-43: LGTM!

Also applies to: 58-58

olive-mcp-server/pyproject.toml (1)

14-15: LGTM!

olive-mcp-server/olive_mcp_server/tools/embeddings.py (1)

10-11: LGTM!

olive-mcp-server/olive_mcp_server/tools/docs_search.py (3)

46-49: LGTM!

Also applies to: 82-96


251-285: LGTM!

Also applies to: 288-314


346-361: LGTM!

olive-mcp-server/tests/test_docs_search_semantic.py (1)

13-33: LGTM!

Also applies to: 150-161, 164-173, 196-232, 280-306

olive-mcp-server/tests/test_tools.py (1)

163-177: LGTM!

olive-mcp-server/olive_mcp_server/tools/passive_context.py (2)

9-15: LGTM!


112-122: LGTM!

Comment thread olive-mcp-server/Dockerfile
Comment thread olive-mcp-server/tests/test_docs_search_semantic.py
Comment thread src/server/services/ai/oliveMcpKnowledge.ts
parseChatStructuredReply's loose-JSON salvage only matched
convert/quant keywords against the object *key* (e.g. "quantMethod",
"onnx"), so a payload shaped like {step: "convert_to_onnx"} or
{step: "apply_quantization"} — generic key, keyword in the value —
matched nothing and produced no patch. Widen both fallback branches
to also match keyword strings in the value. Pre-existing failure on
main (unrelated to the PR #73 branch), fixing here since it was
blocking this PR's CI.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@tonythethompson
tonythethompson enabled auto-merge (squash) August 1, 2026 15:08
- docs_search.py:_fetch_live_docs: publish also when the cache is
  currently empty, not only when this is still the latest generation.
  Otherwise a newer generation that fails after an older one already
  succeeded (but hasn't published yet) permanently discards the only
  good fetch, leaving the live cache empty.
- test_docs_search_semantic.py: assert the stale-generation test's
  worker thread actually completed and returned data, so a silent
  exception or deadlock in that thread can't produce a false pass.
- oliveMcpKnowledge.test.ts: add a regression case for a normalized
  sub-1 relevance score, which the previous `relevance: 2` fixture
  couldn't distinguish from the old >=1 threshold.

Also widened BatchProcessingPanel.test.tsx's fetch assertion into a
waitFor while investigating a "validate" CI failure — turned out to
be a pre-existing bug unrelated to this PR (reproduces identically on
origin/main), leaving as-is rather than chasing it here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@tonythethompson
tonythethompson merged commit 15eb063 into main Aug 1, 2026
9 of 10 checks passed
@tonythethompson
tonythethompson deleted the feat/olive-mcp-semantic-retrieval branch August 1, 2026 15:57
@linear-code

linear-code Bot commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

OLI-19

tonythethompson added a commit that referenced this pull request Aug 1, 2026
PR #73 squash-merged mid-session while more CodeRabbit-driven fixes
kept landing on this branch, so main and this branch diverged despite
main containing an earlier snapshot of the same work. Resolved by
keeping this branch's content (a strict superset with later fixes)
for all conflicted files, and restored cu130/cu132 in
auditAutofix.ts's CUDA_VERSIONS set — main's squash-merge snapshot
happened to capture this branch mid-flip-flop (a since-reverted
CodeRabbit suggestion to remove cu130/cu132, itself reverted per
explicit follow-up direction to keep them as driver-only tags), and
the 3-way merge silently applied that removal since it wasn't a
content conflict.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
tonythethompson added a commit that referenced this pull request Aug 1, 2026
…closed (#75)

* feat(mcp): add embedding-based semantic retrieval and passive context

Replace keyword-only docs and troubleshooting search with MiniLM embeddings
and hybrid scoring, add get_context_for_pipeline for AI prompt injection,
and pre-cache the model in Docker. Existing tool signatures stay unchanged.

* fix: Snapshot live docs and fetch time atomically

* fix: Materialize text iterables before encoding

* fix: Remove unused Path import

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* fix: repair CI failures on PR #73

- Add cu130/cu132 to UIState.cudaVersion type to match auditAutofix's
  CUDA_VERSIONS set (TS2769 build failure)
- Fix test_search_olive_documentation_with_live_source: _fetch_live_docs
  now returns (pages, fetch_time) tuple, test mock still returned bare dict

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: address CI failures and CodeRabbit review on PR #73

- src/types.ts: add missing cu130/cu132 CUDA version literals (unblocks
  tsc/lint, unrelated pre-existing drift vs auditAutofix.ts)
- tests/test_tools.py: fix live-doc test to match _fetch_live_docs' new
  (dict, float) return signature (was silently caught by broad except,
  failing the "live:" source assertion in CI)
- docs_search.py: kb mtime invalidation now includes file count so
  deletions/back-dated additions aren't missed; keyword fallback reuses
  already-loaded KB texts instead of re-reading from disk; log swallowed
  exceptions instead of silent pass; live-doc "freshness" injection no
  longer displaces a strictly better local/kept result (was unconditional
  at top_k==1)
- embeddings.py: fix missing Sequence import (was silently deferred by
  `from __future__ import annotations`, would NameError under runtime
  introspection)
- passive_context.py: log retrieval failures instead of returning an
  indistinguishable "no matches" response
- troubleshooting.py: drop unnecessary `global` on a dict-mutate-only
  cache, bound the fingerprint cache size, extract a shared response
  payload builder for the empty-message and matched-entry paths (they'd
  drifted apart)
- Dockerfile: force the CPU-only torch wheel (extra-index-url doesn't
  prioritize over PyPI's CUDA build), COPY --chown instead of a
  redundant chown -R layer, set HF_HUB_OFFLINE=1
- pyproject.toml: cap sentence-transformers/numpy below their next
  untested majors
- oliveMcpKnowledge.ts: fix docsSearchSufficient's relevance threshold,
  which still assumed the old raw-hit-count scale (>=1) after this PR
  normalized relevance to ~[0,1] — was silently making local KB results
  always look "insufficient"

Verified: 186/186 olive-mcp-server pytest, tsc --noEmit clean, pnpm lint
0 errors, 149/149 server vitest.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: make live-doc merge test deterministic

The previous version relied on the real embedding model ranking a
single synthetic live snippet above ~2500 real local KB entries for
"calibration data" — legitimately doesn't happen once the forced
live-injection bug was fixed, so the test failed in CI (which has the
real sentence-transformers model) even though it passed locally
(where the model isn't installed and the code falls back to keyword
search). Mock _search_local/_search_live directly instead, matching
the pattern used by the other merge-logic regression tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: live-index stale-overwrite race and empty-cache rebuild bug

- docs_search.py:_get_live_index: cache-hit/publish checks used
  `_LIVE_SNIPPETS` truthiness, so an empty (but successfully built)
  live index was never recognized as cached and got rebuilt on every
  call. Also, publish had no ordering guard: a slow build for an
  older fetch_time could complete after a newer generation already
  published and clobber it, silently rolling the live cache back to
  stale content. Both fixed: cache checks now key off
  `_LIVE_EMBEDDINGS is not None`, and publish is skipped when the
  build's fetch_time is older than what's already cached.
- Added test_live_index_does_not_overwrite_newer_generation covering
  the race.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: salvage convert/quant actions from step-only value strings

parseChatStructuredReply's loose-JSON salvage only matched
convert/quant keywords against the object *key* (e.g. "quantMethod",
"onnx"), so a payload shaped like {step: "convert_to_onnx"} or
{step: "apply_quantization"} — generic key, keyword in the value —
matched nothing and produced no patch. Widen both fallback branches
to also match keyword strings in the value. Pre-existing failure on
main (unrelated to the PR #73 branch), fixing here since it was
blocking this PR's CI.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: address remaining CodeRabbit nitpicks on PR #73

- docs_search.py:_fetch_live_docs: publish also when the cache is
  currently empty, not only when this is still the latest generation.
  Otherwise a newer generation that fails after an older one already
  succeeded (but hasn't published yet) permanently discards the only
  good fetch, leaving the live cache empty.
- test_docs_search_semantic.py: assert the stale-generation test's
  worker thread actually completed and returned data, so a silent
  exception or deadlock in that thread can't produce a false pass.
- oliveMcpKnowledge.test.ts: add a regression case for a normalized
  sub-1 relevance score, which the previous `relevance: 2` fixture
  couldn't distinguish from the old >=1 threshold.

Also widened BatchProcessingPanel.test.tsx's fetch assertion into a
waitFor while investigating a "validate" CI failure — turned out to
be a pre-existing bug unrelated to this PR (reproduces identically on
origin/main), leaving as-is rather than chasing it here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: don't salvage chat actions from negated or non-action prose

salvageChatActionPatchFromLooseJson's value-string fallback (added
earlier this PR to catch {step: "convert_to_onnx"}-shaped payloads)
was too broad: it scanned *any* string field, so {note: "do not
quantize this model"} or {note: "conversion is unavailable"} would
incorrectly enable those passes. Restrict value-string matching to
step/action/task keys only, and add a negation guard (no/not/never/
disable/skip/without/unavailable) so negated prose can't enable a
pass. Added regression tests for both cases.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: pre-existing BatchProcessingPanel test/validation bug (unrelated to PR #73)

Root cause chain, found via instrumentation:
1. testUtils.tsx:createMockUIState set passes.conversion=true with no
   conversionFormat. buildOliveRecipe() then defaults to
   OpenVINOConversion, which requires torch/onnx input — but
   modelSource defaults to "huggingface" (hf output). Every component
   test using this default state was silently building an
   already-invalid pipeline.
2. In BatchProcessingPanel, this meant the "valid" job in the
   individual-validation test was itself blocked by a pass-chain
   mismatch, so the test's fetch assertion coincidentally passed on
   an already-broken path (fetch was never reached for either job,
   but the assertion only checked the invalid job's failure).
3. Fixing the state default advanced the job further into real
   handleStartQueue logic, exposing a second, independent bug: the
   test's EventSource mock was `vi.fn(() => obj)`, an arrow function,
   which cannot be used with `new` — throwing when the component
   opens its SSE stream for the (now correctly) running valid job.
4. Once EventSource was constructible, the sequential job loop still
   never advanced to the invalid job because nothing in the test ever
   fired the "done" SSE event the component awaits before moving to
   the next queued job.

Fixes: add conversionFormat: "onnx" to the mock state default, make
the EventSource mock a real constructor function, and have the test
capture and fire the "done" listener so the loop can proceed to
validate the invalid job. Verified this reproduces identically on
origin/main (pre-existing, unrelated to the semantic-retrieval work
in this PR) before fixing it here per request to get CI fully green.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: don't widen cudaVersion type to match a stray unsupported entry

My earlier fix for the tsc failure in auditAutofix.ts (unrelated
pre-existing bug on this branch) took the wrong direction: it widened
UIState["cudaVersion"] to include cu130/cu132 to match
auditAutofix.ts's CUDA_VERSIONS set, when the actual bug was that set
itself listing tags the project doesn't support anywhere else.
RESOLVABLE_CUDA_TAGS (oliveGpuRuntime.ts) and inferRequiredPackages
(recipe.ts) both cap at cu128 — cu130/cu132 aren't resolvable yet
(ORT + nvidia-*-cu13 PyPI pins don't exist). Revert cudaVersion to
its original union and trim the two stray CUDA_VERSIONS sets
(auditAutofix.ts, chatActions.ts) that had cu130/cu132 instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat: expose cu130/cu132 as selectable driver-identification CUDA tags

nvidia-cudnn-cu13 and torch's cu130/cu132 wheel indexes are real and
published now (verified against PyPI/download.pytorch.org), so the
prior "do not enable" comment was stale. However, full package pinning
isn't a small change: cublas-cu13 and cuda-runtime-cu13 are deprecated
stubs pointing at unsuffixed packages (nvidia-cublas, nvidia-cuda-runtime)
rather than following the cuXX-suffixed scheme cu118..cu128 use, so
CUDA12_RUNTIME_PACKAGES can't just be extended — it needs its own
verified pin set (onnxruntime-gpu 1.27+/1.28, unsuffixed nvidia-*
packages) as a separate change.

This adds cu130/cu132 as selectable values across UIState, the manual
CUDA dropdown, chat action salvage, and audit autofix — labeled
"driver only" in the UI — while leaving RESOLVABLE_CUDA_TAGS/
resolveCudaTag/inferRequiredPackages untouched, so selecting them still
surfaces the existing clear "Unsupported CUDA tag" error at package-
resolution time rather than silently mis-resolving CUDA-12 runtime
libs against a CUDA-13 driver. Full pin support is a follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: Handle common negation variants

* fix: Return early when top_k is zero

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Update olive-mcp-server/Dockerfile

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Update olive-mcp-server/Dockerfile

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Update olive-mcp-server/tests/test_passive_context.py

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* fix: address PR #75 review round 2 (Copilot/codex/cubic findings)

- docs_search.py:_fetch_live_docs — the "empty cache means no gen has
  published" proxy only handled a fully-empty cache. If the TTL-expired
  cache was non-empty (the common case), an older-but-successful fetch
  racing behind a newer-but-failed one was still discarded, leaving the
  stale cache in place indefinitely. Track the actual last-published
  generation instead of the highest-started one, so a fetch only loses
  to a generation that genuinely published, not one that merely started
  later and failed.
- test_passive_context.py — fix `@pytest.fixture` decorator mangled into
  literal backtick-wrapped text by an earlier automated commit; was a
  hard syntax error blocking the whole test collection.
- chatActions.ts — the "quant"/"convert" acceptance regex still accepted
  any string containing the substring "quant" (e.g. "check quantization
  compatibility"), so non-actionable informational text could still
  enable the pass. Require an exact match against a small set of
  affirmative tokens instead.
- types.ts — cudaVersion doc comment now reflects that cu130/cu132 are
  driver-identification only, not part of the fully-resolved tag set.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: address PR #75 round 3 review findings (real bugs + test hygiene)

Real bugs:
- docs_search.py:search_olive_documentation — removed the live-doc
  "freshness slot" block; CodeRabbit correctly traced it as unreachable
  dead code (the tie-break in the preceding sort already prefers live
  sources on equal relevance, so the guard condition can never be true
  after the earlier stale-overwrite fix).
- chatActions.ts — quant/convert matching now tokenizes values instead
  of requiring either a full-string match or a bare "quant" substring:
  bare method names like `task: "gptq"` and combined phrases like
  `step: "apply awq"` were previously ignored entirely. Negation is now
  scoped per-target (checks for a quant/convert mention shortly *after*
  the negation word) instead of one flag suppressing every detector, so
  `"convert to onnx without quantization"` no longer wrongly drops the
  non-negated conversion instruction.
- passive_context.py:get_context_for_pipeline — added a `status` field
  ("ok" | "retrieval_failed") so a KB/embedding-model failure is no
  longer indistinguishable from a genuinely empty result at the API
  boundary an AI assistant prompt is built from.
- troubleshooting.py:_best_match — log a warning when semantic scoring
  fails instead of silently degrading to keyword-only matching.

Test hygiene:
- test_docs_search_semantic.py: two keyword-fallback tests depended on
  real on-disk KB content containing "calibration data" — patched
  _load_kb_text with a fixed corpus like the existing mtime test
  already does, so KB content changes can't break them. Tightened the
  stale-build assertion to check the actual invariant (cache stays
  unpublished) instead of a disjunction that hid it, and dropped the
  unused load_calls list.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: qodo-code-review[bot] <151058649+qodo-code-review[bot]@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants