Skip to content

fix(memory): word/sentence-boundary aware fact truncation - #12383

Merged
diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.51from
kareem-jalal:fix/memory-extraction-boundary-truncation
Sep 18, 2026
Merged

diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.51from
kareem-jalal:fix/memory-extraction-boundary-truncation

Conversation

@kareem-jalal

Copy link
Copy Markdown
Contributor

Bug

src/lib/memory/extraction.ts truncates extracted memory facts with raw
character-offset slicing and no boundary awareness:

  • sanitizeMatch() did raw.trim().replace(/\s+/g, ' ').slice(0, MAX_FACT_LENGTH)
    — a hard cut that can land mid-word.
  • capExtractionText() did text.slice(-MAX_EXTRACTION_TEXT_LENGTH) on the
    front edge when input exceeds MAX_EXTRACTION_TEXT_LENGTH (64KB) — same
    problem, opposite edge.

These facts later get injected into LLM context as <system-reminder>
memory blocks, so a mid-word/mid-clause cut produces garbled fragments
(e.g. text starting with a stray closing parenthesis, or ending with no
punctuation and a half word).

Fix

  • sanitizeMatch() now backs the cut index off to the nearest clean
    boundary within a lookback window: prefers a sentence-ending mark
    (./!/?) so the fact reads as a complete clause, falls back to a
    plain whitespace/word boundary, and only falls back to the original hard
    cut when no boundary exists in the window (e.g. one long unbroken run of
    characters).
  • capExtractionText() applies the equivalent boundary-aware trim on the
    front edge of the kept tail (extends the cut forward to the next
    whitespace boundary rather than slicing mid-word).

This mirrors the boundary-aware truncation already used by
open-sse/services/compression/lite.ts (issue #8169) for tool-result
truncation, so the codebase now has one consistent pattern for this class
of problem.

Testing

  • New: tests/unit/memory-extraction-boundary-truncation.test.ts — covers
    a long match cut at a word boundary, a match cut preferentially at
    sentence-ending punctuation, short strings passing through untouched,
    the no-boundary-in-window fallback, and capExtractionText's
    boundary-aware tail behavior (both under and over the 64KB cap).
  • Ran existing tests/unit/memory-extraction.test.ts alongside the new
    file — all 28 tests pass, no regressions:
    ℹ tests 28
    ℹ pass 28
    ℹ fail 0
    
  • npx tsc --noEmit -p tsconfig.json — no new type errors introduced by
    this change.

No unrelated files were touched.

sanitizeMatch() and capExtractionText() previously did raw character-offset
slices (slice(0, MAX_FACT_LENGTH) / slice(-MAX_EXTRACTION_TEXT_LENGTH)) with
no boundary awareness, producing garbled mid-word/mid-clause fragments that
get injected into LLM context as memory facts.

- sanitizeMatch() now backs the cut off to the nearest sentence-ending
  punctuation (. ! ?) within a lookback window, falling back to a plain
  whitespace boundary, falling back to the original hard cut only when no
  boundary exists nearby.
- capExtractionText() applies the equivalent boundary-aware trim on the
  front edge of the kept tail.

Mirrors the boundary-aware truncation pattern already used by
open-sse/services/compression/lite.ts (diegosouzapw#8169) for tool-result truncation.

Adds tests/unit/memory-extraction-boundary-truncation.test.ts covering
word-boundary cuts, sentence-boundary preference, short-string passthrough,
the no-boundary-available fallback, and capExtractionText's tail behavior.
@diegosouzapw

Copy link
Copy Markdown
Owner

Nice fix — matching the same boundary-lookback pattern the repo already uses for #8169
(compressToolResults) is a good consistency call, and the sentence-punctuation-first /
word-boundary-fallback / original-cut-as-last-resort chain reads cleanly. Please add a
changelog fragment under changelog.d/fixes/ before merge.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @LeMonBLOCK — merging via the release merge-train. Validated in local merge-train (merge-train-20260918-120911-suite.log) on the devbox @ train tip 8c305709478b7052fb981a75852bbf88fe3a959d with the sibling PRs of this batch: typecheck:core, file-size, complexity, cognitive-complexity, changelog-integrity green; changed-area node:test 310/311 (the single red, hard-lease inventory, reproduces on the pure release tip) + vitest green (fast parity mode — full suite ran today on the tip via the base-red and 3b trains). Merged --admin per merge-gates §4/§7.

@diegosouzapw
diegosouzapw merged commit ad633c8 into diegosouzapw:release/v3.8.51 Sep 18, 2026
3 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…pw#12383)

* fix(memory): word/sentence-boundary aware truncation in extraction

sanitizeMatch() and capExtractionText() previously did raw character-offset
slices (slice(0, MAX_FACT_LENGTH) / slice(-MAX_EXTRACTION_TEXT_LENGTH)) with
no boundary awareness, producing garbled mid-word/mid-clause fragments that
get injected into LLM context as memory facts.

- sanitizeMatch() now backs the cut off to the nearest sentence-ending
  punctuation (. ! ?) within a lookback window, falling back to a plain
  whitespace boundary, falling back to the original hard cut only when no
  boundary exists nearby.
- capExtractionText() applies the equivalent boundary-aware trim on the
  front edge of the kept tail.

Mirrors the boundary-aware truncation pattern already used by
open-sse/services/compression/lite.ts (diegosouzapw#8169) for tool-result truncation.

Adds tests/unit/memory-extraction-boundary-truncation.test.ts covering
word-boundary cuts, sentence-boundary preference, short-string passthrough,
the no-boundary-available fallback, and capExtractionText's tail behavior.

* docs(changelog): add fragment for word/sentence-boundary fact truncation

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Sep 29, 2026
…d release date

- Every one of the 1972 cycle commits is covered by a bullet (reconcile-changelog
  vs origin/release/v3.8.50); fragments folded under [3.8.51] with the PR link
  and the author of the commit that added them.
- Re-land credit: #15019/#15020/#15119/#15120/#15121 credit @HouMinXi; #12383
  credits its author @kareem-jalal alongside the original credit.
- Contributors hall regenerated (301 external contributors).
- [3.8.51] dated 2026-09-29 in the root CHANGELOG and the 66 i18n mirrors.
@diegosouzapw diegosouzapw mentioned this pull request Sep 29, 2026
diegosouzapw added a commit that referenced this pull request Sep 30, 2026
The only bullet the sync-back removes from release/v3.8.52 is the #12383
line, re-credited to @kareem-jalal / @LeMonBLOCK in the v3.8.51
reconciliation; the other 630 changes are the bullets reconciled after
2026-09-15. Record the exact transition so check:changelog-integrity
accepts it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants