Skip to content

fix(workspace): thread document path through search results - #503

Merged
ilblackdragon merged 2 commits into
nearai:mainfrom
zmanian:fix/481-memory-search-document-path
Mar 3, 2026
Merged

ilblackdragon merged 2 commits into
nearai:mainfrom
zmanian:fix/481-memory-search-document-path

Conversation

@zmanian

@zmanian zmanian commented Mar 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

Fixes #481 - Memory search results link to chunk UUID instead of source file path.

  • Added document_path: String to RankedResult and SearchResult in search.rs
  • Updated FTS and vector search queries in repository.rs to SELECT d.path as document_path (JOIN already exists)
  • Updated the RRF fusion pipeline to propagate document_path through ChunkInfo
  • Updated libSQL backend search queries in db/libsql/workspace.rs similarly
  • Changed handlers/memory.rs to use r.document_path.clone() instead of r.document_id.to_string()
  • Added document_path to memory tool search output in builtin/memory.rs

Test plan

  • Search workspace memory via the web UI and verify results show file paths (e.g., "notes/todo.md") instead of UUIDs
  • Verify both FTS-only and hybrid (FTS + vector) search return correct paths
  • Run cargo test workspace::search::tests to confirm RRF unit tests pass
  • Test with libSQL backend (--features libsql) to verify parity

Co-Authored-By: Claude Opus 4.6 noreply@anthropic.com

Memory search results were showing chunk UUIDs instead of source file
paths. Thread document_path through RankedResult, SearchResult, and the
RRF fusion pipeline so handlers can display the actual file path.

Fixes nearai#481

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added scope: channel/web Web gateway channel scope: tool/builtin Built-in tools scope: workspace Persistent memory / workspace size: S 10-49 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Mar 3, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses an issue where memory search results were linking to chunk UUIDs instead of human-readable source file paths. The changes involve extending core data structures to include document paths, updating database queries across different backends to retrieve these paths, and ensuring they are correctly propagated and displayed in both the web UI and tool outputs. This significantly improves the usability and clarity of search results by providing direct file path references.

Highlights

  • Data Structure Enhancement: Added a document_path: String field to the RankedResult and SearchResult structs in search.rs to explicitly store the file path of source documents.
  • Database Query Updates: Modified Full-Text Search (FTS) and vector search queries in repository.rs and db/libsql/workspace.rs to select the d.path (document path) from the memory_documents table.
  • Result Propagation: Updated the Reciprocal Rank Fusion (RRF) pipeline to ensure document_path is correctly propagated through the ChunkInfo structure during search result processing.
  • Web UI Integration: Changed handlers/memory.rs to display the document_path directly in search results within the web UI, replacing the previously shown document_id (UUID).
  • Tool Output Consistency: Included the document_path in the output of the memory tool search in builtin/memory.rs, providing more informative results.
Changelog
  • src/channels/web/handlers/memory.rs
    • Updated the path field in SearchHit to use r.document_path instead of r.document_id.to_string() for displaying search results.
  • src/db/libsql/workspace.rs
    • Modified FTS SQL queries to include d.path in the SELECT statement.
    • Adjusted RankedResult mapping to correctly assign the retrieved d.path to document_path.
    • Modified vector search SQL queries to include d.path in the SELECT statement.
    • Adjusted RankedResult mapping for vector search to correctly assign the retrieved d.path to document_path.
  • src/tools/builtin/memory.rs
    • Added "path": r.document_path to the JSON output for memory search results.
  • src/workspace/repository.rs
    • Updated FTS SQL queries to select d.path as document_path.
    • Mapped the document_path from query results to the RankedResult struct.
    • Updated vector search SQL queries to select d.path as document_path.
    • Mapped the document_path from vector search query results to the RankedResult struct.
  • src/workspace/search.rs
    • Added a document_path: String field to the SearchResult struct.
    • Added a document_path: String field to the RankedResult struct.
    • Included document_path in the ChunkInfo struct used for reciprocal rank fusion.
    • Ensured document_path is populated when creating new ChunkInfo instances from FTS results.
    • Ensured document_path is populated when creating new ChunkInfo instances from vector search results.
    • Mapped info.document_path to SearchResult when converting ChunkInfo.
Activity
  • The pull request author, zmanian, has outlined a comprehensive test plan to verify the changes, including checking web UI results, FTS and hybrid search correctness, RRF unit tests, and libSQL backend parity.
  • The changes were co-authored by Claude Opus 4.6, indicating collaborative development.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces and propagates a document_path field throughout the memory search and retrieval system. The changes involve adding document_path to SearchResult and RankedResult structs, modifying SQL queries in libsql/workspace.rs and workspace/repository.rs to select the document path, and updating result parsing to populate this new field. The document_path is then utilized in the memory_search_handler, included in the JSON output of the MemorySearchTool, and integrated into the reciprocal_rank_fusion logic. The reviewer suggests performance optimizations by using .into_iter() instead of .iter() in src/channels/web/handlers/memory.rs and src/tools/builtin/memory.rs to avoid unnecessary cloning of document_path and content strings when the original results collection is consumed.

Comment thread src/channels/web/handlers/memory.rs Outdated
Comment on lines 162 to 167
.iter()
.map(|r| SearchHit {
path: r.document_id.to_string(),
path: r.document_path.clone(),
content: r.content.clone(),
score: r.score as f64,
})

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To improve performance, you can avoid cloning document_path and content. Since results is not used after this, you can use .into_iter() to consume it and move the string values into SearchHit.

Suggested change
.iter()
.map(|r| SearchHit {
path: r.document_id.to_string(),
path: r.document_path.clone(),
content: r.content.clone(),
score: r.score as f64,
})
.into_iter()
.map(|r| SearchHit {
path: r.document_path,
content: r.content,
score: r.score as f64,
})
References
  1. To improve performance, avoid unnecessary heap allocations by using iterators directly to consume collections and move values, rather than cloning.

Comment thread src/tools/builtin/memory.rs Outdated
@@ -100,6 +100,7 @@ impl Tool for MemorySearchTool {
"results": results.iter().map(|r| serde_json::json!({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since results is not used after this, you can use .into_iter() to consume the vector and move the string values (content and path) into the json! macro. This avoids unnecessary cloning and improves performance.

Suggested change
"results": results.iter().map(|r| serde_json::json!({
"results": results.into_iter().map(|r| serde_json::json!({
References
  1. To improve performance, avoid unnecessary heap allocations by using iterators directly to consume collections and move values, rather than cloning.

Address review feedback: consume results with into_iter() to move
String fields directly instead of cloning them.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@ilblackdragon
ilblackdragon merged commit d562dc8 into nearai:main Mar 3, 2026
14 checks passed
zmanian added a commit that referenced this pull request Mar 4, 2026
…rough RRF

Verifies that search results carry the source document's file path
through the RRF fusion pipeline, not the document UUID. Covers the
bug fixed in PR #503 / issue #481.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ilblackdragon pushed a commit that referenced this pull request Mar 6, 2026
…509)

* test(workspace): add regression test for document_path propagation through RRF

Verifies that search results carry the source document's file path
through the RRF fusion pipeline, not the document UUID. Covers the
bug fixed in PR #503 / issue #481.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Update src/workspace/search.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* chore: merge main and fix formatting

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
[skip-regression-check]

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
* fix(workspace): thread document path through search results

Memory search results were showing chunk UUIDs instead of source file
paths. Thread document_path through RankedResult, SearchResult, and the
RRF fusion pipeline so handlers can display the actual file path.

Fixes nearai#481

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: use into_iter to move values instead of cloning

Address review feedback: consume results with into_iter() to move
String fields directly instead of cloning them.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
…earai#509)

* test(workspace): add regression test for document_path propagation through RRF

Verifies that search results carry the source document's file path
through the RRF fusion pipeline, not the document UUID. Covers the
bug fixed in PR nearai#503 / issue nearai#481.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Update src/workspace/search.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* chore: merge main and fix formatting

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
[skip-regression-check]

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…earai#509)

* test(workspace): add regression test for document_path propagation through RRF

Verifies that search results carry the source document's file path
through the RRF fusion pipeline, not the document UUID. Covers the
bug fixed in PR nearai#503 / issue nearai#481.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Update src/workspace/search.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* chore: merge main and fix formatting

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
[skip-regression-check]

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: channel/web Web gateway channel scope: tool/builtin Built-in tools scope: workspace Persistent memory / workspace size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: memory search results link to chunk UUID instead of source file path

2 participants