Skip to content

fix(search): keep natural-language queries on the FTS5 path instead of crashing to the LIKE fallback - #354

Closed
Milgauss wants to merge 1 commit into
stephenschoettler:mainfrom
Milgauss:fix/fts5-natural-language-queries
Closed

Milgauss wants to merge 1 commit into
stephenschoettler:mainfrom
Milgauss:fix/fts5-natural-language-queries

Conversation

@Milgauss

@Milgauss Milgauss commented Jul 8, 2026

Copy link
Copy Markdown

Problem

Any conversational search query containing an apostrophe, comma, question mark, or most other punctuation (for example don't) is a syntax error to FTS5. store.search catches the error and silently degrades to the LIKE fallback, so in practice nearly every natural-language query runs on the slow, unranked fallback path instead of the FTS index.

Mechanism

sanitize_fts5_query (search_query.py:47) only replaces the fixed operator set _FTS5_SPECIAL_CHARS (search_query.py:36), leaving apostrophes, commas, and question marks in the unquoted query text. store.search (store.py:931) passes the result to messages_fts MATCH, FTS5 raises fts5: syntax error near "'", and the handler at store.py:997 logs "FTS message search failed, falling back to LIKE" and degrades. dag.py:427 has the same path for summary search.

Fix

Outside quoted phrases, keep only bareword-safe characters (alphanumerics, underscore, whitespace) and replace everything else with spaces. That is exactly the token alphabet the unicode61 tokenizer indexed, so the sanitized query matches the index instead of erroring; don't becomes don t, which matches rows containing don't. Quoted phrases pass through verbatim, as before.

The LIKE fallback must not use that strict sanitization: it matches raw stored content, so apostrophes, emoji, and CJK punctuation have to survive for the fallback to find them (an emoji-only query would otherwise produce zero terms and return nothing). The helper is therefore split in two: sanitize_fts5_query (strict, used by the FTS MATCH paths) and sanitize_like_terms_query (the previous operator-stripping behavior, used by _search_like in store.py and dag.py).

Verification

New tests cover: don't returning its row on the FTS path with no fallback warning, a full conversational sentence (what did I tell you, don't you remember?) searching without any OperationalError or fallback, quoted phrases with trailing conversational punctuation still matching as phrases, and sanitizer unit cases. On unpatched main the new tests fail with the exact production error (fts5: syntax error near ","). The full tests/test_lcm_core.py suite passes (288 tests), including the existing emoji and LIKE-fallback regression tests.

Incident evidence

One live store logged the FTS-to-LIKE fallback warning over 300 times, meaning nearly every conversational search had been silently running on the degraded fallback path.

…e LIKE fallback

sanitize_fts5_query only stripped a fixed set of FTS5 operator
characters, leaving apostrophes, commas, question marks and other
conversational punctuation in the unquoted query text. FTS5 parses
those as syntax errors ("fts5: syntax error near ..."), so store.search
silently degraded to the LIKE fallback for nearly any conversational
query, e.g. "don't". One production store logged the fallback warning
300+ times.

Fix: outside quoted phrases, keep only bareword-safe characters
(alphanumerics, underscore, whitespace). That is exactly the token
alphabet the unicode61 tokenizer indexed, so replacing everything else
with spaces matches the index instead of erroring. Quoted phrases pass
through verbatim, as before.

The LIKE fallback itself must NOT use that strict sanitization: it
matches raw stored content, so apostrophes, emoji and CJK punctuation
have to survive for the fallback to find them (emoji-only queries
would otherwise return nothing). Split the helper in two:
sanitize_fts5_query (strict, FTS MATCH paths) and
sanitize_like_terms_query (previous operator-stripping behavior, used
by the _search_like paths in store.py and dag.py).

Tests: new unit tests for the sanitizer plus store-level tests
asserting that "don't", a full conversational sentence, and a quoted
phrase with trailing punctuation all return results on the FTS path
with no "falling back to LIKE" warning. On unpatched code they fail
with the exact production symptom.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019uao4D1LKM35dKzivcsWVf
@Milgauss

Copy link
Copy Markdown
Author

Closing as superseded. Current main carries this fix in a more thorough shape than this PR: _fts5_safe_char turns every non-alphanumeric outside a quoted phrase into a separator (keeping combining marks, with NFC composition first), sanitize_like_query keeps the LIKE path on the weaker cleanup, and bare AND/OR/NOT/NEAR are neutralized. That landed with b2f228c ("recall: reduce a raw question to FTS5 terms before MATCH", #168) and its review-finding follow-ups, via #436.

Verified rather than assumed: this PR's three store.search regression tests (apostrophe query, full conversational question, quoted phrase plus conversational punctuation) pass unmodified against main at 10cbb78, and sanitize_fts5_query on main returns exactly the strings this PR's unit tests asserted (don't -> don t, the quoted-phrase case verbatim). Nothing left here to rebase. Thanks for taking the underlying problem on.

@Milgauss Milgauss closed this Aug 26, 2026
@Milgauss
Milgauss deleted the fix/fts5-natural-language-queries branch August 26, 2026 15:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant