fix(memory): word/sentence-boundary aware fact truncation - #12383
diegosouzapw merged 2 commits into
Conversation
sanitizeMatch() and capExtractionText() previously did raw character-offset slices (slice(0, MAX_FACT_LENGTH) / slice(-MAX_EXTRACTION_TEXT_LENGTH)) with no boundary awareness, producing garbled mid-word/mid-clause fragments that get injected into LLM context as memory facts. - sanitizeMatch() now backs the cut off to the nearest sentence-ending punctuation (. ! ?) within a lookback window, falling back to a plain whitespace boundary, falling back to the original hard cut only when no boundary exists nearby. - capExtractionText() applies the equivalent boundary-aware trim on the front edge of the kept tail. Mirrors the boundary-aware truncation pattern already used by open-sse/services/compression/lite.ts (diegosouzapw#8169) for tool-result truncation. Adds tests/unit/memory-extraction-boundary-truncation.test.ts covering word-boundary cuts, sentence-boundary preference, short-string passthrough, the no-boundary-available fallback, and capExtractionText's tail behavior.
|
Nice fix — matching the same boundary-lookback pattern the repo already uses for #8169 |
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Thanks @LeMonBLOCK — merging via the release merge-train. Validated in local merge-train (merge-train-20260918-120911-suite.log) on the devbox @ train tip 8c305709478b7052fb981a75852bbf88fe3a959d with the sibling PRs of this batch: typecheck:core, file-size, complexity, cognitive-complexity, changelog-integrity green; changed-area node:test 310/311 (the single red, hard-lease inventory, reproduces on the pure release tip) + vitest green (fast parity mode — full suite ran today on the tip via the base-red and 3b trains). Merged --admin per merge-gates §4/§7. |
…pw#12383) * fix(memory): word/sentence-boundary aware truncation in extraction sanitizeMatch() and capExtractionText() previously did raw character-offset slices (slice(0, MAX_FACT_LENGTH) / slice(-MAX_EXTRACTION_TEXT_LENGTH)) with no boundary awareness, producing garbled mid-word/mid-clause fragments that get injected into LLM context as memory facts. - sanitizeMatch() now backs the cut off to the nearest sentence-ending punctuation (. ! ?) within a lookback window, falling back to a plain whitespace boundary, falling back to the original hard cut only when no boundary exists nearby. - capExtractionText() applies the equivalent boundary-aware trim on the front edge of the kept tail. Mirrors the boundary-aware truncation pattern already used by open-sse/services/compression/lite.ts (diegosouzapw#8169) for tool-result truncation. Adds tests/unit/memory-extraction-boundary-truncation.test.ts covering word-boundary cuts, sentence-boundary preference, short-string passthrough, the no-boundary-available fallback, and capExtractionText's tail behavior. * docs(changelog): add fragment for word/sentence-boundary fact truncation Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…d release date - Every one of the 1972 cycle commits is covered by a bullet (reconcile-changelog vs origin/release/v3.8.50); fragments folded under [3.8.51] with the PR link and the author of the commit that added them. - Re-land credit: #15019/#15020/#15119/#15120/#15121 credit @HouMinXi; #12383 credits its author @kareem-jalal alongside the original credit. - Contributors hall regenerated (301 external contributors). - [3.8.51] dated 2026-09-29 in the root CHANGELOG and the 66 i18n mirrors.
The only bullet the sync-back removes from release/v3.8.52 is the #12383 line, re-credited to @kareem-jalal / @LeMonBLOCK in the v3.8.51 reconciliation; the other 630 changes are the bullets reconciled after 2026-09-15. Record the exact transition so check:changelog-integrity accepts it.
Bug
src/lib/memory/extraction.tstruncates extracted memory facts with rawcharacter-offset slicing and no boundary awareness:
sanitizeMatch()didraw.trim().replace(/\s+/g, ' ').slice(0, MAX_FACT_LENGTH)— a hard cut that can land mid-word.
capExtractionText()didtext.slice(-MAX_EXTRACTION_TEXT_LENGTH)on thefront edge when input exceeds
MAX_EXTRACTION_TEXT_LENGTH(64KB) — sameproblem, opposite edge.
These facts later get injected into LLM context as
<system-reminder>memory blocks, so a mid-word/mid-clause cut produces garbled fragments
(e.g. text starting with a stray closing parenthesis, or ending with no
punctuation and a half word).
Fix
sanitizeMatch()now backs the cut index off to the nearest cleanboundary within a lookback window: prefers a sentence-ending mark
(
./!/?) so the fact reads as a complete clause, falls back to aplain whitespace/word boundary, and only falls back to the original hard
cut when no boundary exists in the window (e.g. one long unbroken run of
characters).
capExtractionText()applies the equivalent boundary-aware trim on thefront edge of the kept tail (extends the cut forward to the next
whitespace boundary rather than slicing mid-word).
This mirrors the boundary-aware truncation already used by
open-sse/services/compression/lite.ts(issue #8169) for tool-resulttruncation, so the codebase now has one consistent pattern for this class
of problem.
Testing
tests/unit/memory-extraction-boundary-truncation.test.ts— coversa long match cut at a word boundary, a match cut preferentially at
sentence-ending punctuation, short strings passing through untouched,
the no-boundary-in-window fallback, and
capExtractionText'sboundary-aware tail behavior (both under and over the 64KB cap).
tests/unit/memory-extraction.test.tsalongside the newfile — all 28 tests pass, no regressions:
npx tsc --noEmit -p tsconfig.json— no new type errors introduced bythis change.
No unrelated files were touched.