Skip to content

fix(memory): sanitize FTS queries so natural-language recall works on libSQL (#7275) - #7285

Closed
serrrfirat wants to merge 4 commits into
mainfrom
implement-issue-7275-test
Closed

serrrfirat wants to merge 4 commits into
mainfrom
implement-issue-7275-test

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Closes Reborn: verify explicit persistent memory recall across conversations in production #7275: caller-level verification of explicit persistent-memory recall across conversations on the production composition path (standalone build over the embedded-libSQL root filesystem — the active shipping backend), not an in-memory harness.
  • Building that verification exposed a real production defect: the libSQL backend passes the memory search query verbatim to FTS5, which treats ? ! ( ) " + - * : ^ as operators. Any natural-language query containing them makes the MATCH expression invalid: memory_search fails with OperationFailed, and the proactive prompt-memory lane (whose query is the user's latest message) degrades to empty. User messages almost always carry punctuation — so explicit recall across conversations was unreliable in production. This is the likely root cause of the Memory not reliably recalled across conversations #7185 feedback.
  • Fix: sanitize queries in MemorySearchRequest::new (memory-native) to whitespace-joined alphanumeric tokens — the same treatment the unicode61 tokenizer applies to indexed content, so it is faithful, not lossy. Punctuation-only queries still fail as invalid input.

Change Type

  • Bug fix
  • Documentation (test evidence maps issue acceptance criteria)

Linked Issue

Closes #7275

Validation

  • cargo fmt --all -- --check
  • cargo clippy -p ironclaw_composition -p ironclaw_memory_native --tests (clean)
  • cargo test -p ironclaw_composition --lib — 505 passed
  • cargo test -p ironclaw_memory_native — all suites passed (52+39+13+25+4+4)
  • cargo test -p ironclaw_host_runtime --test memory_prompt_context + lib memory_* — passed
  • cargo test -p ironclaw_integration_tests --test reborn_group_memory — passed
  • cargo test -p ironclaw_integration_tests --test reborn_integration_wiring_parity — passed
  • Regression discipline: the new composition test fails without the sanitizer fix (verified by temporary revert)
  • Note: reborn_group_multiuser stack-overflows on the base commit too (pre-existing, reproduced with changes stashed)

Acceptance-criteria evidence (issue #7275)

Criterion Evidence
1. Write in conv A + provider-issued/durable write evidence memory_recall_across_conversations_on_production_path: tool response carries status:"written", path, append, content_length; tree read-back; full runtime teardown + rebuild over the same libSQL root
2. Conv B (different thread, same tenant/user/agent/project) finds marker via memory_search same test: result_count: 1, marker in content, search_scope marker
3. Conv B receives marker via proactive prompt-memory lane without a tool call same test: production memory_context_service (exact memory_lifecycle_consumers derivation build_reborn_runtime wires) returns the marker snippet for conv B's run scope; request mirrors the loop's own builder
4. Fails closed for different user + isolated scope axis memory_recall_fails_closed_for_other_users_and_isolated_scope_axes: cross-user and cross-project both return successful result_count: 0 and empty prompt lane; non-vacuous same-scope control finds the marker
5. Failure distinguishable from no-match memory_retrieval_failure_is_distinct_from_no_matching_memory + memory_prompt_lane_failure_emits_operator_diagnostic: no-match = successful invocation with result_count: 0; empty query = failed invocation (InputEncode); backend Unavailable = typed Err(Unavailable), prompt lane degrades to empty while the memory context lane retrieval failed operator diagnostic is emitted (captured via thread-local subscriber)
6. Verified on production composition path + active shipping backend all tests build build_runtime_substrate(local_filesystem_build_input(..)) (embedded libSQL, native memory provider)
7. Version documentation Complaint reported on a pre-#6345 build (#7185). Punctuation-triggered FTS failure is present on current main (post-#6345) and fixed by this PR. Verified on head 2e2439af4

Test Strategy

User behavior: recall of explicitly saved memory across conversations must not depend on whether the user's phrasing carries punctuation.

Risk areas:

  • Persistence
  • Cross-component behavior (memory provider ↔ prompt lane ↔ loop host ↔ tool surface)
  • Model behavior — Not applicable: no prompt/model behavior changed; the query normalization is provider-side
  • Browser — Not applicable: no frontend change
  • Security or permissions — Not applicable: scope isolation semantics unchanged (fail-closed assertions added)
  • External provider — Not applicable: native provider only

Tests added or updated:

  • Unit or contract: MemorySearchRequest sanitizer tests (memory-native search.rs)
  • Reborn integration: none added (existing reborn_group_memory + wiring_parity re-run green)
  • Recorded fixture: Not applicable
  • Browser E2E: Not applicable
  • Backend or runtime: composition factory caller-level tests over embedded libSQL
  • Live canary: Not applicable

What the tests prove: the full issue #7275 acceptance-criteria matrix above, on the production composition path; plus the FTS5 punctuation defect and its fix.

Commands run: see Validation.

… libSQL (#7275)

Issue #7275 asks for caller-level verification that explicit persistent
memory written in conversation A is searchable and proactively recalled
in conversation B on the production composition path. Building that
verification on the real embedded-libSQL standalone composition exposed
a genuine production defect: the libSQL backend passes the memory search
query verbatim to FTS5, which treats ?, !, (, ), quotes, and + - * : ^
as operators — any natural-language query containing them makes the
MATCH expression invalid, so memory_search fails with OperationFailed
and the proactive prompt-memory lane (query = the user's latest message)
degrades to empty. User messages almost always carry punctuation, so
explicit recall across conversations was unreliable in production.

- memory-native: sanitize queries in MemorySearchRequest::new to
  whitespace-joined alphanumeric tokens (the same treatment the
  unicode61 tokenizer applies to indexed content); punctuation-only
  queries still fail as invalid input.
- composition factory tests (issue #7275 acceptance criteria 1-6, on the
  production composition path with the shipping embedded-libSQL backend,
  not an in-memory harness):
  - conv A write with provider-issued write evidence (status/path/
    content_length), durable across a full runtime teardown/rebuild;
  - conv B (different thread, same tenant/user/agent/project) finds the
    marker via memory_search and receives it through the production
    prompt-memory lane (same derivation build_reborn_runtime wires)
    without any memory tool call, including natural-language queries
    with punctuation;
  - fail-closed for a different user and for an isolated project axis
    (successful empty results on both surfaces, with non-vacuous
    same-scope controls);
  - retrieval failure distinguishable from no-match: no-match is a
    successful invocation with result_count 0, invalid input fails with
    InputEncode, backend Unavailable errors degrade the prompt lane to
    empty while emitting the operator diagnostic (captured via a
    thread-local subscriber).
  - regression tests fail without the sanitizer fix.
@railway-app

railway-app Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7285 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw 🕒 Building (View Logs) Web Aug 6, 2026 at 1:06 pm

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7285 August 6, 2026 12:23 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 6, 2026
@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9d17aa4e-500b-4499-a01d-010370cf719f

📥 Commits

Reviewing files that changed from the base of the PR and between 79f2e5b and ccbd541.

📒 Files selected for processing (3)
  • crates/app/ironclaw_composition/src/factory/tests.rs
  • crates/extensions/packages/memory-native/src/search.rs
  • crates/substrates/ironclaw_filesystem/src/libsql.rs

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Improved memory recall across rebuilt conversations while maintaining caller and project isolation.
    • Added clearer diagnostics when memory services are unavailable or prompt processing is degraded.
  • Bug Fixes

    • Improved searches containing punctuation, Unicode, hyphenated terms, reserved operators, and quotation marks.
    • Prevented punctuation-only searches from producing invalid requests.
    • Improved distinction between no results, invalid input, and backend failures.

Walkthrough

The changes normalize punctuation-bearing memory searches and add production-path coverage for durable recall, scope isolation, typed failures, prompt-lane degradation, and operator diagnostics.

Changes

Memory search and production recall

Layer / File(s) Summary
FTS query normalization
crates/extensions/packages/memory-native/src/search.rs, crates/substrates/ironclaw_filesystem/src/libsql.rs
Memory requests sanitize non-alphanumeric input. LibSQL binds each FTS5 term as a quoted literal. Tests cover operators, Unicode, hyphens, empty input, and embedded quotes.
Cross-thread persistent recall
crates/app/ironclaw_composition/src/factory/tests.rs
Production-path tests verify durable writes, conversation rebuild persistence, cross-conversation search, proactive retrieval, and prompt snippets.
Isolation and failure behavior
crates/app/ironclaw_composition/src/factory/tests.rs
Tests cover user and project isolation, no-match results, invalid input, unavailable backends, graceful prompt-lane degradation, and structured diagnostics.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

  • nearai/ironclaw#7281: Modifies the same memory search sanitization, LibSQL FTS filtering, and production regression tests.

Suggested reviewers: benkurrek

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers the change, issue, validation, and test strategy but omits required security, database, blast-radius, rollback, review, and track sections. Add the missing template sections, including Security Impact, Database Impact, Blast Radius, Rollback Plan, Review Follow-Through, and Review track.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and clearly describes the FTS query sanitization fix.
Linked Issues check ✅ Passed The changes and evidence address all coding objectives and acceptance criteria in issue #7275, including durable recall, isolation, diagnostics, and production-path validation.
Out of Scope Changes check ✅ Passed The sanitizer and production-composition regression tests directly support issue #7275 and remain within its stated scope.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/extensions/packages/memory-native/src/search.rs (1)

66-86: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Punctuation-only input now produces a retrieval failure, which collides with the criterion-5 diagnostic contract.

Two consequences follow from sanitizing before the empty check:

  1. A query that contains only punctuation, emoji, or CJK punctuation now returns Err(InvalidId). Before this change it reached the backend. On the prompt lane the caller cannot distinguish this from a backend outage: crates/app/ironclaw_composition/src/factory/tests.rs at lines 1350-1355 asserts the lane emits memory context lane retrieval failed for any error. An operator reading that diagnostic cannot tell "the backend is down" from "the user typed only emoji". The PR states criterion 5 is exactly this distinction.
  2. value: query reports the sanitized value, which is always the empty string on this path. The original input never reaches the error. Keep that behavior if it is intentional redaction, but state it, because the field name implies the offending value.

Consider classifying "no searchable tokens" as a distinct outcome so the lane can treat it as a no-match rather than a retrieval failure.

This is a public constructor behavior change. Update the owning memory-search contract documentation.

As per coding guidelines: "Update the owning contract or documentation when behavior changes."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/extensions/packages/memory-native/src/search.rs` around lines 66 - 86,
Revise the memory-search flow around sanitize_fts_query so punctuation-only,
emoji-only, or otherwise tokenless input produces a distinct no-match outcome
rather than Err(HostApiError::InvalidId), allowing the prompt lane to
distinguish it from backend retrieval failures. Preserve the original query for
any error value that still represents invalid input, or explicitly document
intentional redaction if sanitized values remain. Update the owning public
memory-search contract documentation to describe this behavior.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/app/ironclaw_composition/src/factory/tests.rs`:
- Around line 1282-1304: Update the input-failure case in the caller-level
memory search test to submit a punctuation-only query, exercising the production
sanitizer in MemorySearchRequest::new rather than deserialization of a missing
field, and assert the resulting FailureKind (preserving a separate missing-field
case if its kind differs). Remove the tautological
MemoryServiceError::unavailable() kind assertion, and remove the
MemoryServiceErrorKind import if it is no longer used.
- Around line 1346-1355: Update the diagnostics assertion in the relevant test
to verify the exact error-kind rendering emitted by the ironclaw_host_runtime
diagnostic, in addition to the existing “memory context lane retrieval failed”
message check. Use the error-kind value established by the diagnostic
documentation or implementation so the test fails if that field is omitted.

In `@crates/extensions/packages/memory-native/src/search.rs`:
- Around line 11-31: Update the libSQL backend’s MATCH predicate to quote or
otherwise neutralize uppercase bareword operators produced by
sanitize_fts_query, while leaving sanitize_fts_query’s token output unchanged
for Postgres plainto_tsquery and in-memory substring searches; add caller-level
probes in crates/app/ironclaw_composition/src/factory/tests.rs:1079-1105 for
“AND staging”, “unlocks staging OR”, and “staging NOT”, and update the libSQL
search path in crates/extensions/packages/memory-native/src/search.rs:11-31
accordingly.

---

Outside diff comments:
In `@crates/extensions/packages/memory-native/src/search.rs`:
- Around line 66-86: Revise the memory-search flow around sanitize_fts_query so
punctuation-only, emoji-only, or otherwise tokenless input produces a distinct
no-match outcome rather than Err(HostApiError::InvalidId), allowing the prompt
lane to distinguish it from backend retrieval failures. Preserve the original
query for any error value that still represents invalid input, or explicitly
document intentional redaction if sanitized values remain. Update the owning
public memory-search contract documentation to describe this behavior.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5202f087-9406-4f95-906b-d88cbe2d8629

📥 Commits

Reviewing files that changed from the base of the PR and between 0c297cb and 79f2e5b.

📒 Files selected for processing (2)
  • crates/app/ironclaw_composition/src/factory/tests.rs
  • crates/extensions/packages/memory-native/src/search.rs

Comment thread crates/app/ironclaw_composition/src/factory/tests.rs Outdated
Comment thread crates/app/ironclaw_composition/src/factory/tests.rs
Comment thread crates/extensions/packages/memory-native/src/search.rs
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Reopening to trigger a fresh Railway preview deployment for the new head 79f2e5b (deploy-on-open integration).

@serrrfirat serrrfirat closed this Aug 6, 2026
@serrrfirat serrrfirat reopened this Aug 6, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7285 August 6, 2026 12:41 Destroyed
@ironloopai

ironloopai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #7285

🟢 Completed · Review submitted

1 actionable findings →

The sanitizer fixes punctuation-triggered FTS5 errors but still allows alphabetic FTS5 operators, leaving natural-language recall unreliable. The added tests do not cover these reserved tokens.

Automatic · PR opened + CI failed · attempt 1 of 3 · completed in 1m 40s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 0c297cb
  • Head: implement-issue-7275-test at 79f2e5b
  • Created: Aug 6, 2026, 12:45 PM UTC
  • Updated: Aug 6, 2026, 12:47 PM UTC
  • Run: 1ca9e0a1-88e8-4bec-95e1-4d236b32de8a
  • Latest attempt: 1 · Completed · 2dc83063-e636-4cd5-b97d-9e882b9a18d3

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #7285

⚠️ 1 finding · 1 blocking

The sanitizer fixes punctuation-triggered FTS5 errors but still allows alphabetic FTS5 operators, leaving natural-language recall unreliable. The added tests do not cover these reserved tokens.

Findings

  1. 🔴 High · FTS5 keyword operators still pass through sanitization — crates/extensions/packages/memory-native/src/search.rs:20-22
    Details are attached to the relevant diff.
Validation and technical details
  • Reviewed the complete trusted comparison refs/ironloop/base (0c297cb) to refs/ironloop/head (79f2e5b): 2 changed files, 617 insertions, 3 deletions.
  • Inspected both changed areas, crate-local guidance, MemorySearchRequest call sites, the filesystem repository search path, and libSQL FTS5 MATCH translation.
  • Reproduced with SQLite FTS5: OR, NOT, and AND each raise syntax errors; what should I NOT forget is parsed as a boolean expression and returns no match.
  • git diff --check refs/ironloop/base...refs/ironloop/head passed.
  • Focused Cargo tests could not be executed because cargo is unavailable in the review environment.
  • Base: main
  • Head: implement-issue-7275-test at 79f2e5b
  • Run: 1ca9e0a1-88e8-4bec-95e1-4d236b32de8a

Comment thread crates/extensions/packages/memory-native/src/search.rs
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

@ironloopai resolve

@ironloopai

ironloopai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

🧩 Resolve · PR #7285

🟢 Completed · Pull request updated

Commit ccbd541 →

Merged the trusted base without rewriting PR history and committed ccbd541. The update neutralizes FTS5 keyword operators, adds caller-level regression probes, exercises punctuation-only rejection through production composition, removes the tautological assertion, and verifies diagnostic error kinds.

Manual command by @serrrfirat · attempt 1 of 3 · completed in 3m 7s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 0c297cb
  • Head: implement-issue-7275-test at ccbd541
  • Created: Aug 6, 2026, 1:03 PM UTC
  • Updated: Aug 6, 2026, 1:06 PM UTC
  • Run: 86e9caf7-21cf-4810-9ecb-59a3085be2e2
  • Latest attempt: 1 · Completed · 40717ef5-8e89-44c4-95d8-59f348f43bdc

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7285 — ccbd541e Deployed Aug 6, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Reborn: verify explicit persistent memory recall across conversations in production

2 participants