Skip to content

fix(memory): sanitize FTS queries so natural-language recall works on libSQL (#7275) - #7289

Closed
serrrfirat wants to merge 4 commits into
mainfrom
implement-issue-7275-test
Closed

serrrfirat wants to merge 4 commits into
mainfrom
implement-issue-7275-test

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Closes Reborn: verify explicit persistent memory recall across conversations in production #7275: caller-level verification of explicit persistent-memory recall across conversations on the production composition path (standalone build over the embedded-libSQL root filesystem — the active shipping backend), not an in-memory harness.
  • Building that verification exposed a real production defect: the libSQL backend passes the memory search query verbatim to FTS5, which treats ? ! ( ) " + - * : ^ as operators. Any natural-language query containing them makes the MATCH expression invalid: memory_search fails with OperationFailed, and the proactive prompt-memory lane (whose query is the user's latest message) degrades to empty. User messages almost always carry punctuation — so explicit recall across conversations was unreliable in production. This is the likely root cause of the Memory not reliably recalled across conversations #7185 feedback.
  • Fix: sanitize queries in MemorySearchRequest::new (memory-native) to whitespace-joined alphanumeric tokens — the same treatment the unicode61 tokenizer applies to indexed content, so it is faithful, not lossy. Punctuation-only queries still fail as invalid input.

Change Type

  • Bug fix
  • Documentation (test evidence maps issue acceptance criteria)

Linked Issue

Closes #7275

Validation

  • cargo fmt --all -- --check
  • cargo clippy -p ironclaw_composition -p ironclaw_memory_native --tests (clean)
  • cargo test -p ironclaw_composition --lib — 505 passed
  • cargo test -p ironclaw_memory_native — all suites passed (52+39+13+25+4+4)
  • cargo test -p ironclaw_host_runtime --test memory_prompt_context + lib memory_* — passed
  • cargo test -p ironclaw_integration_tests --test reborn_group_memory — passed
  • cargo test -p ironclaw_integration_tests --test reborn_integration_wiring_parity — passed
  • Regression discipline: the new composition test fails without the sanitizer fix (verified by temporary revert)
  • Note: reborn_group_multiuser stack-overflows on the base commit too (pre-existing, reproduced with changes stashed)

Acceptance-criteria evidence (issue #7275)

Criterion Evidence
1. Write in conv A + provider-issued/durable write evidence memory_recall_across_conversations_on_production_path: tool response carries status:"written", path, append, content_length; tree read-back; full runtime teardown + rebuild over the same libSQL root
2. Conv B (different thread, same tenant/user/agent/project) finds marker via memory_search same test: result_count: 1, marker in content, search_scope marker
3. Conv B receives marker via proactive prompt-memory lane without a tool call same test: production memory_context_service (exact memory_lifecycle_consumers derivation build_reborn_runtime wires) returns the marker snippet for conv B's run scope; request mirrors the loop's own builder
4. Fails closed for different user + isolated scope axis memory_recall_fails_closed_for_other_users_and_isolated_scope_axes: cross-user and cross-project both return successful result_count: 0 and empty prompt lane; non-vacuous same-scope control finds the marker
5. Failure distinguishable from no-match memory_retrieval_failure_is_distinct_from_no_matching_memory + memory_prompt_lane_failure_emits_operator_diagnostic: no-match = successful invocation with result_count: 0; empty query = failed invocation (InputEncode); backend Unavailable = typed Err(Unavailable), prompt lane degrades to empty while the memory context lane retrieval failed operator diagnostic is emitted (captured via thread-local subscriber)
6. Verified on production composition path + active shipping backend all tests build build_runtime_substrate(local_filesystem_build_input(..)) (embedded libSQL, native memory provider)
7. Version documentation Complaint reported on a pre-#6345 build (#7185). Punctuation-triggered FTS failure is present on current main (post-#6345) and fixed by this PR. Verified on head d646c3bd8 (rebased onto current main, incl. #7286)

Test Strategy

User behavior: recall of explicitly saved memory across conversations must not depend on whether the user's phrasing carries punctuation.

Risk areas:

  • Persistence
  • Cross-component behavior (memory provider ↔ prompt lane ↔ loop host ↔ tool surface)
  • Model behavior — Not applicable: no prompt/model behavior changed; the query normalization is provider-side
  • Browser — Not applicable: no frontend change
  • Security or permissions — Not applicable: scope isolation semantics unchanged (fail-closed assertions added)
  • External provider — Not applicable: native provider only

Tests added or updated:

  • Unit or contract: MemorySearchRequest sanitizer tests (memory-native search.rs)
  • Reborn integration: none added (existing reborn_group_memory + wiring_parity re-run green)
  • Recorded fixture: Not applicable
  • Browser E2E: Not applicable
  • Backend or runtime: composition factory caller-level tests over embedded libSQL
  • Live canary: Not applicable

What the tests prove: the full issue #7275 acceptance-criteria matrix above, on the production composition path; plus the FTS5 punctuation defect and its fix.

Commands run: see Validation.

serrrfirat and others added 3 commits August 6, 2026 16:25
… libSQL (#7275)

Issue #7275 asks for caller-level verification that explicit persistent
memory written in conversation A is searchable and proactively recalled
in conversation B on the production composition path. Building that
verification on the real embedded-libSQL standalone composition exposed
a genuine production defect: the libSQL backend passes the memory search
query verbatim to FTS5, which treats ?, !, (, ), quotes, and + - * : ^
as operators — any natural-language query containing them makes the
MATCH expression invalid, so memory_search fails with OperationFailed
and the proactive prompt-memory lane (query = the user's latest message)
degrades to empty. User messages almost always carry punctuation, so
explicit recall across conversations was unreliable in production.

- memory-native: sanitize queries in MemorySearchRequest::new to
  whitespace-joined alphanumeric tokens (the same treatment the
  unicode61 tokenizer applies to indexed content); punctuation-only
  queries still fail as invalid input.
- composition factory tests (issue #7275 acceptance criteria 1-6, on the
  production composition path with the shipping embedded-libSQL backend,
  not an in-memory harness):
  - conv A write with provider-issued write evidence (status/path/
    content_length), durable across a full runtime teardown/rebuild;
  - conv B (different thread, same tenant/user/agent/project) finds the
    marker via memory_search and receives it through the production
    prompt-memory lane (same derivation build_reborn_runtime wires)
    without any memory tool call, including natural-language queries
    with punctuation;
  - fail-closed for a different user and for an isolated project axis
    (successful empty results on both surfaces, with non-vacuous
    same-scope controls);
  - retrieval failure distinguishable from no-match: no-match is a
    successful invocation with result_count 0, invalid input fails with
    InputEncode, backend Unavailable errors degrade the prompt lane to
    empty while emitting the operator diagnostic (captured via a
    thread-local subscriber).
  - regression tests fail without the sanitizer fix.
The FTS5-keyword probes (AND/OR/NOT) must assert the invocation SUCCEEDS
with a literal-token no-match (result_count 0), not result_count 1 — the
seeded document does not contain those literal tokens. The regression the
probes pin is the hard OperationFailed from FTS5 parsing barewords as
operators; that error is gone, and no-match is the correct outcome.
@railway-app

railway-app Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7289 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 6, 2026 at 2:09 pm

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7289 August 6, 2026 13:30 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 6, 2026
@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2663ee97-0501-4f14-af83-b3f71af85a0a

📥 Commits

Reviewing files that changed from the base of the PR and between d646c3b and f3099ff.

📒 Files selected for processing (1)
  • crates/app/ironclaw_composition/src/factory/tests.rs

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Improved memory search for natural-language, Unicode, and punctuation-containing queries.
    • Added safeguards so reserved search operators and special characters are treated as text.
    • Enhanced memory retrieval across conversations and runtime restarts.
  • Bug Fixes

    • Improved handling of empty, invalid, unmatched, and backend-failure search results.
    • Preserved isolation between users and projects during memory retrieval.
    • Added clearer diagnostic behavior when prompt-context retrieval is unavailable.

Walkthrough

The PR adds production-path memory recall coverage across runtime rebuilds and conversation scopes. It sanitizes native FTS queries and quotes libSQL FTS terms to neutralize operators and embedded quotes.

Changes

Memory recall and FTS handling

Layer / File(s) Summary
Native FTS query sanitization
crates/extensions/packages/memory-native/src/search.rs
Search requests remove non-alphanumeric separators, preserve Unicode tokens, and reject punctuation-only queries.
LibSQL FTS literalization
crates/substrates/ironclaw_filesystem/src/libsql.rs
FTS filters quote terms and escape embedded quotes. Regression tests cover reserved operators.
Production-path memory validation
crates/app/ironclaw_composition/src/factory/tests.rs
Integration tests cover persistent cross-conversation recall, prompt-lane retrieval, scope isolation, no-match results, invalid input, and backend failures.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

  • nearai/ironclaw#7281 — It modifies the same memory search sanitization and production recall test paths.
  • nearai/ironclaw#7285 — It implements overlapping FTS sanitization and production memory-recall coverage.
  • nearai/ironclaw#7286 — It modifies the same libSQL FTS handling path for index materialization and backfill reuse.

Sequence Diagram(s)

sequenceDiagram
  participant ConversationA
  participant MemoryService
  participant PersistentStore
  participant ConversationB
  participant PromptLane

  ConversationA->>MemoryService: Persist explicit memory
  MemoryService->>PersistentStore: Store and index memory
  ConversationB->>MemoryService: Search with sanitized query
  MemoryService->>PersistentStore: Execute quoted FTS query
  PersistentStore-->>MemoryService: Matching memory
  MemoryService-->>ConversationB: Return memory snippets
  PromptLane->>MemoryService: Retrieve prompt-context memories
  MemoryService->>PersistentStore: Execute scoped FTS query
  PersistentStore-->>PromptLane: Matching snippets or empty result
Loading

Suggested reviewers: benkurrek

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers the change, issue, validation, and test strategy but omits required security, database, blast-radius, rollback, and review sections. Complete the missing template sections, including Security Impact, Reborn Trust-Boundary Checklist, Database Impact, Blast Radius, Rollback Plan, Review Follow-Through, and Review track.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately describes the FTS query sanitization fix.
Linked Issues check ✅ Passed The PR provides evidence for all coding acceptance criteria in issue #7275, including durable recall, isolation, failure distinction, diagnostics, and production-path coverage.
Out of Scope Changes check ✅ Passed The backend fixes and production composition tests directly support issue #7275 and do not introduce unrelated code changes.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/app/ironclaw_composition/src/factory/tests.rs`:
- Around line 1157-1168: Add a same-user, same-project prompt-lane control
before the existing cross-user and cross-project negative assertions, using the
prompt-lane invocation path and a completed-status record that matches the
in-scope thread or scope. Assert that this control returns a hit, then retain
the existing negative assertions; update the nearby non-vacuity comment to state
that the control covers the prompt lane rather than only memory_search.

In `@crates/substrates/ironclaw_filesystem/src/libsql.rs`:
- Around line 2865-2876: Update fts5_literal_query to reject token-less input,
such as whitespace-only queries, before producing the MATCH expression.
Propagate this outcome through the Filter::Fts translation path as the
established typed Unsupported result (or equivalent empty-result behavior),
ensuring callers cannot bind an empty FTS5 query and that validation does not
rely on MemorySearchRequest::new.
- Line 2840: Make Filter::Fts semantics consistent across backends by defining
the intended query behavior in a shared contract case within the database root
filesystem tests. Exercise syntax such as abc*, col:term, and a OR b, and assert
the same expected row set for both libSQL and Postgres implementations, using
fts5_literal_query and the Postgres Filter::Fts path as the implementation
points to align.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 65e5d3b7-7505-48c6-9de9-0f7f58c96725

📥 Commits

Reviewing files that changed from the base of the PR and between 582a741 and d646c3b.

📒 Files selected for processing (3)
  • crates/app/ironclaw_composition/src/factory/tests.rs
  • crates/extensions/packages/memory-native/src/search.rs
  • crates/substrates/ironclaw_filesystem/src/libsql.rs

Comment on lines +1157 to +1168
// Non-vacuity control: same user + same project, different thread → found.
let control = invoke_json(
&services,
MEMORY_SEARCH_CAPABILITY_ID,
memory_context_for(MEMORY_SEARCH_CAPABILITY_ID, user, "conv-b", None),
serde_json::json!({"query": "isolation marker", "limit": 5}),
)
.await
.expect("control search succeeds");
assert_eq!(control["result_count"], serde_json::json!(1));

let lane = prompt_lane_service(&services).expect("native binding wires the prompt lane");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

The prompt-lane negatives are vacuous, and the doc comment overstates the control.

The control at Line 1158 exercises memory_search only. Lines 1200 and 1232 then assert the prompt lane returns empty for another user and another project. No assertion in this test proves the prompt lane returns anything at all for the in-scope user.

If prompt_lane_service, prompt_lane_turn_scope, or the resolver regressed to always-empty, both negatives still pass and the test reports isolation. The comment at Line 1127 claims the negatives are non-vacuous because of a same-test control; that holds for the search surface, not for the lane.

Add a same-scope lane hit before the negatives. crates/**/*.rs requires caller-level coverage at the nearest meaningful seam, and completed-status-only evidence is insufficient.

💚 Proposed fix: add a prompt-lane control
     let lane = prompt_lane_service(&services).expect("native binding wires the prompt lane");
+
+    // Non-vacuity control for the prompt lane itself: same user, same
+    // project, different thread must receive the marker. Without this, the
+    // negative lane assertions below pass even if the lane is always empty.
+    let control_scope = prompt_lane_turn_scope(user, "conv-b", None);
+    let control_snippets = lane
+        .load_memory_snippets(prompt_lane_request(&control_scope, user, "isolation marker?"))
+        .await
+        .expect("control prompt lane retrieval succeeds");
+    assert!(
+        control_snippets
+            .iter()
+            .any(|snippet| snippet.safe_summary.contains(MARKER)),
+        "the in-scope prompt lane must find the marker: {control_snippets:?}"
+    );
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
// Non-vacuity control: same user + same project, different thread → found.
let control = invoke_json(
&services,
MEMORY_SEARCH_CAPABILITY_ID,
memory_context_for(MEMORY_SEARCH_CAPABILITY_ID, user, "conv-b", None),
serde_json::json!({"query": "isolation marker", "limit": 5}),
)
.await
.expect("control search succeeds");
assert_eq!(control["result_count"], serde_json::json!(1));
let lane = prompt_lane_service(&services).expect("native binding wires the prompt lane");
// Non-vacuity control: same user + same project, different thread → found.
let control = invoke_json(
&services,
MEMORY_SEARCH_CAPABILITY_ID,
memory_context_for(MEMORY_SEARCH_CAPABILITY_ID, user, "conv-b", None),
serde_json::json!({"query": "isolation marker", "limit": 5}),
)
.await
.expect("control search succeeds");
assert_eq!(control["result_count"], serde_json::json!(1));
let lane = prompt_lane_service(&services).expect("native binding wires the prompt lane");
// Non-vacuity control for the prompt lane itself: same user, same
// project, different thread must receive the marker. Without this, the
// negative lane assertions below pass even if the lane is always empty.
let control_scope = prompt_lane_turn_scope(user, "conv-b", None);
let control_snippets = lane
.load_memory_snippets(prompt_lane_request(&control_scope, user, "isolation marker?"))
.await
.expect("control prompt lane retrieval succeeds");
assert!(
control_snippets
.iter()
.any(|snippet| snippet.safe_summary.contains(MARKER)),
"the in-scope prompt lane must find the marker: {control_snippets:?}"
);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/app/ironclaw_composition/src/factory/tests.rs` around lines 1157 -
1168, Add a same-user, same-project prompt-lane control before the existing
cross-user and cross-project negative assertions, using the prompt-lane
invocation path and a completed-status record that matches the in-scope thread
or scope. Assert that this control returns a hit, then retain the existing
negative assertions; update the nearby non-vacuity comment to state that the
control covers the prompt lane rather than only memory_search.

Source: Coding guidelines

});
};
params.push(libsql::Value::Text(query.clone()));
params.push(libsql::Value::Text(fts5_literal_query(query)));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# 1) Enumerate every Filter::Fts producer and the query values it passes.
set -euo pipefail

rg -nP --type=rust -C5 'Filter::Fts\s*\{' crates/

# 2) Show how the Postgres backend translates Filter::Fts for comparison.
fd -t f 'postgres.*\.rs' crates/substrates/ironclaw_filesystem/src \
  --exec rg -nP -C10 'Filter::Fts|plainto_tsquery|websearch_to_tsquery|to_tsquery' {}

# 3) Check whether the shared contract suite covers FTS at all.
fd -t f 'db_root_filesystem_contract.rs' crates/ --exec rg -nP -C4 'Fts|fts' {}

Repository: nearai/ironclaw

Length of output: 14642


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== libsql FTS helpers/context =="
sed -n '2800,2850p' crates/substrates/ironclaw_filesystem/src/libsql.rs
echo
sed -n '2870,2903p' crates/substrates/ironclaw_filesystem/src/libsql.rs

echo
echo "== memory-native full-text consumer =="
sed -n '920,945p' crates/extensions/packages/memory-native/src/repo/filesystem.rs

echo
echo "== postgres FTS handler =="
sed -n '2440,2466p' crates/substrates/ironclaw_filesystem/src/postgres.rs

echo
echo "== in-memory FTS helper/tests =="
sed -n '780,816p' crates/substrates/ironclaw_filesystem/src/in_memory.rs
sed -n '1528,1570p' crates/substrates/ironclaw_filesystem/src/in_memory.rs

echo
echo "== FTS terms in contract tests around first/last occurrences =="
sed -n '755,785p' crates/substrates/ironclaw_filesystem/tests/db_root_filesystem_contract.rs
sed -n '835,855p' crates/substrates/ironclaw_filesystem/tests/db_root_filesystem_contract.rs
sed -n '3760,3785p' crates/substrates/ironclaw_filesystem/tests/db_root_filesystem_contract.rs

echo
echo "== strict Filter::Fts producer/text matching (non-FullText) =="
rlib='rg -n --type=rust '\'"Filter::Fts\s*{|request\.full_text\(\)' crates/

Repository: nearai/ironclaw

Length of output: 11642


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== libsql FTS coverage tests =="
sed -n '2900,3075p' crates/substrates/ironclaw_filesystem/src/libsql.rs

echo
echo "== all Filter::Fts queries in tests =="
rg -n -C3 'Filter::Fts\s*\{|full_text\(\)\s*:' crates/substrates/ironclaw_filesystem crates/extensions/packages/memory-native/src/repo/filesystem.rs crates/extensions/packages/memory-native/** 2>/dev/null || true

echo
echo "== fts5_literal_query semantic comparison for representative inputs =="
python3 - <<'PY'
def to_postgres_phrase(s):
    return " ".join(f'{token.strip("*")}:*' for token in s.split())

def fts5_literal_query(s):
    return " ".join(f'"{token.replace(chr(34), chr(34)+chr(34))}"' for token in s.split())

for s in ["abc*", "col:term", "cat OR dog", "fox \"brown\""]:
    print(f"input={s!r}\npostgres={to_postgres_phrase(s)!r}\nlibsql={fts5_literal_query(s)!r}")
PY

Repository: nearai/ironclaw

Length of output: 14185


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== shared FTS contract loop setup =="
sed -n '720,782p' crates/substrates/ironclaw_filesystem/tests/db_root_filesystem_contract.rs

echo
echo "== all Filter::Fts occurrences =="
rg -n -C2 'Filter::Fts\s*{|full_text\(\)\s*:' crates/substrates/ironclaw_filesystem crates/extensions/packages/memory-native/src/repo/filesystem.rs crates/extensions/packages/memory-native 2>/dev/null || true

echo
echo "== backend-specific internal FTS tests =="
rg -n -C4 'fts5_literal_query|fts_.*match|full[_-]?text' crates/substrates/ironclaw_filesystem/src crates/substrates/ironclaw_filesystem/tests 2>/dev/null || true

Repository: nearai/ironclaw

Length of output: 9624


Make Filter::Fts serialization parity explicit across backends.

Filter::Fts is substrate-wide. crates/substrates/ironclaw_filesystem/src/libsql.rs::fts5_literal_query() literals every token before the FTS5 MATCH, while crates/substrates/ironclaw_filesystem/src/postgres.rs passes the raw query into plainto_tsquery. This makes abc*, col:term, and a OR b return different row sets on libSQL versus Postgres, and also diverge from the in-memory naive match. Add a shared contract case covering the intended Filter::Fts semantics in tests/db_root_filesystem_contract.rs so both SQL backends are pinned to one API.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/substrates/ironclaw_filesystem/src/libsql.rs` at line 2840, Make
Filter::Fts semantics consistent across backends by defining the intended query
behavior in a shared contract case within the database root filesystem tests.
Exercise syntax such as abc*, col:term, and a OR b, and assert the same expected
row set for both libSQL and Postgres implementations, using fts5_literal_query
and the Postgres Filter::Fts path as the implementation points to align.

Source: Coding guidelines

Comment on lines +2865 to +2876
/// Render a free-form FTS query as an implicit AND of FTS5 string literals.
///
/// Quoting each whitespace-delimited term keeps uppercase words such as
/// `AND`, `OR`, and `NOT` from being interpreted as FTS5 syntax. Doubling
/// embedded quotes follows FTS5's string-literal escaping rules.
fn fts5_literal_query(query: &str) -> String {
query
.split_whitespace()
.map(|token| format!("\"{}\"", token.replace('"', "\"\"")))
.collect::<Vec<_>>()
.join(" ")
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

Fail closed when the query has no tokens.

If query holds only whitespace, fts5_literal_query returns an empty string. That empty string is bound to MATCH ?, and FTS5 rejects an empty match expression as a syntax error. The caller then sees a Backend error where a typed Unsupported (or an empty result) is the correct outcome.

The memory lane guards this upstream in MemorySearchRequest::new. The filesystem substrate must not depend on one consumer's validator; Filter::Fts is reachable from any domain store.

🛡️ Proposed fix: reject a token-less FTS query in the translator
         Filter::Fts { key, query } => {
             let Some(fts_table) = fts_tables.get(key.as_str()) else {
                 return Err(FilesystemError::Unsupported {
                     path: path.clone(),
                     operation: FilesystemOperation::Query,
                 });
             };
-            params.push(libsql::Value::Text(fts5_literal_query(query)));
+            let match_expression = fts5_literal_query(query);
+            if match_expression.is_empty() {
+                return Err(FilesystemError::Unsupported {
+                    path: path.clone(),
+                    operation: FilesystemOperation::Query,
+                });
+            }
+            params.push(libsql::Value::Text(match_expression));
             out.push_str(&format!(
                 "(path IN (SELECT path FROM {fts_table} WHERE {fts_table} MATCH ?{}))",
                 params.len()
             ));
             Ok(())
         }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/substrates/ironclaw_filesystem/src/libsql.rs` around lines 2865 -
2876, Update fts5_literal_query to reject token-less input, such as
whitespace-only queries, before producing the MATCH expression. Propagate this
outcome through the Filter::Fts translation path as the established typed
Unsupported result (or equivalent empty-result behavior), ensuring callers
cannot bind an empty FTS5 query and that validation does not rely on
MemorySearchRequest::new.

Source: Coding guidelines

@ironloopai

ironloopai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #7289

🟢 Completed · Review submitted

Submitted review →

Reviewed the complete trusted base-to-head comparison. The memory query normalization and libSQL FTS literalization address punctuation and reserved-keyword failures without weakening scope isolation. The added composition tests cover durable cross-conversation recall, proactive retrieval, isolation, and distinguishable failure behavior. No actionable findings identified.

Automatic · PR opened + CI failed · attempt 1 of 3 · completed in 2m 27s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 582a741
  • Head: implement-issue-7275-test at d646c3b
  • Created: Aug 6, 2026, 1:35 PM UTC
  • Updated: Aug 6, 2026, 1:37 PM UTC
  • Run: 5e884d16-0167-4e0a-87cb-1adac4eaaf68
  • Latest attempt: 1 · Completed · 2d712d10-62f2-4b8c-9909-4bd685e7f907

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #7289

✅ No actionable findings

Reviewed the complete trusted base-to-head comparison. The memory query normalization and libSQL FTS literalization address punctuation and reserved-keyword failures without weakening scope isolation. The added composition tests cover durable cross-conversation recall, proactive retrieval, isolation, and distinguishable failure behavior. No actionable findings identified.

Validation and technical details
  • Inspected all three changed files and surrounding memory-native, filesystem backend, and production composition call paths.
  • Compared libSQL FTS behavior with the in-memory and PostgreSQL implementations and the shared Filter::Fts contract.
  • Verified the exact trusted comparison refs/ironloop/base...refs/ironloop/head and all three commits in that range.
  • git diff --check refs/ironloop/base...refs/ironloop/head completed cleanly.
  • Repository code graph was unavailable, so review used crate guidance and targeted live-code searches as prescribed.
  • Could not execute Rust tests because cargo is not installed in the review environment.
  • Base: main
  • Head: implement-issue-7275-test at d646c3b
  • Run: 5e884d16-0167-4e0a-87cb-1adac4eaaf68

@serrrfirat serrrfirat closed this Aug 6, 2026
@serrrfirat serrrfirat reopened this Aug 6, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7289 August 6, 2026 14:02 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Railway preview QA — PASS

Tested head: f3099ffd5272ef8a73af70a3d26534c27fb32415 (PR head, verified unchanged before and after the run)
Railway state: success at test time — Success - ironclaw-ironclaw-pr-7289.up.railway.app (deployment bound to the tested head SHA)
Preview URL: https://ironclaw-ironclaw-pr-7289.up.railway.app
Route: /chat (gateway token login, exact-head asset app-BiwQJlad.js)

Given/When/Then matrix

Case Acceptance Intended contract Actual contract exercised Observed result Status
Conversation A: user asks the assistant to save a marker to memory Required Explicit persistent-memory write succeeds; assistant confirms with a saved location Real deployed head, real model, real memory backend; 1 memory tool call "Saved. The staging rollback codename osprey-meridian-7 is now stored in memory at staging-rollback-codename. I'll have it ready when you need it." (worked 5s) PASS
Conversation B: NEW conversation (New button), user asks in natural language with punctuation: "what's the staging rollback codename?" Required Assistant recalls the marker from memory; no hard retrieval failure from the punctuation-laden query (the #7275 regression claim) Real deployed head, real model, fresh conversation (different thread id), same owner scope; 2 tool calls (memory/workspace search) "The staging rollback codename is osprey-meridian-7." (worked 4s) PASS
Cleanup Supplemental Remove the temporary test memory Model overwrote the staging-rollback-codename doc with a "deleted" marker "The memory note at staging-rollback-codename has been overwritten with a 'deleted' marker... Confirmed done." Done (note remains as deleted marker — throwaway preview)

Status derivation

  • Required cases: 2 executed against the intended contract (exact deployed head, real model, real memory backend, real punctuation-bearing natural-language query) — both PASSED.
  • PASS — every required case was executed against the intended contract and passed.
  • The live run does not isolate whether conv B's recall came from the proactive prompt-memory lane or an explicit memory_search tool call (the model exercised tools; the UI does not expose its prompt). The proactive-lane mechanism without any tool call is pinned by the deterministic caller-level test memory_recall_across_conversations_on_production_path in this PR (fails without the fix).

Notes

  • Earlier run on this PR's previous head (d646c3bd8, deployment since cancelled) produced the same result: conv A saved, conv B recalled "The staging rollback codename is osprey-meridian-7." — the marker was not carried over to this deployment (fresh preview volume per deployment), so the A→B pair was re-run end-to-end on the tested head.
  • Railway preview deployments for this PR were cancelled twice earlier — both times within seconds of the PR being closed on GitHub (7285 closed 13:08:40Z → deploy cancelled 13:08:47Z; 7289 closed 13:40:22Z → deploy cancelled 13:40:27Z). Deployment cancellation correlates with PR closure, not with the build.
  • CI note: workflow runs for this head were cancelled by the reopen churn (concurrency cancel-in-progress); a fresh run needs a push or a re-run. The only genuine test failure seen on earlier heads was the CLI smoke test serve_does_not_mount_cli_login_route_when_token_is_env_sourced (listener bind refused) — it passes locally and is unrelated to this diff.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

CI status: blocked by GitHub Actions outage

All currently-failing checks on head f3099ffd5 fail with Failed to resolve action download info: Error: Service Unavailable (jobs could not even download actions) — GitHub is in a Partial System Outage (githubstatus.com, ongoing since ~16:13Z, >6h). No code-related failures are present:

  • ✅ Code Style (fmt + clippy) — passed (fmt fix landed in f3099ffd5)
  • ✅ Tests (Reborn) — passed (the earlier serve_does_not_mount_cli_login_route_when_token_is_env_sourced listener-bind flake did not recur; it also passes locally)
  • ⛔ Reborn E2E / WebUI E2E / Platform & Compat / Detect code changes / affected-1 — all failed during the outage window with the Service Unavailable action-download error; reruns are queued but cannot start until the incident clears

Local verification on this head: cargo test -p ironclaw_composition --lib memory suite 20/20, ironclaw_memory_native all suites, ironclaw_filesystem 144 passed, fmt/clippy clean, smoke test passes.

Resume step (once githubstatus shows All Systems Operational): gh run rerun <run> --repo nearai/ironclaw --failed for runs 31108800186, 31108801051, 31108801914, 31108802987.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7289 — f3099ffd Deployed Aug 6, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Reborn: verify explicit persistent memory recall across conversations in production

2 participants