Repository navigation
fix(editor): account for NFC boundary composition in insert offset - #1954
Conversation
Cursor.modifyText builds newText from prefix + insert + suffix and hands it to
Cursor.fromText, which NFC-normalizes the whole string. But the new cursor
offset was computed as startOffset + insertString.normalize('NFC').length,
normalizing the insert in isolation. When the inserted text begins with a
combining mark that composes with the last character of the prefix (e.g. "e" +
U+0301 -> "é"), the normalized newText is one UTF-16 unit shorter than that
formula assumes, so the returned offset overshoots by one and the cursor lands
past the following text — the next keystroke then edits the wrong spot.
Measure the normalized prefix-plus-insert instead, so cross-boundary
composition is accounted for. Reduces to the previous behavior whenever no
boundary composition occurs (plain ASCII, astral emoji, insert at start).
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
📜 Recent review details⏰ Context from checks skipped due to timeout. (3)
🧰 Additional context used📓 Path-based instructions (4)**/*.{ts,tsx}📄 CodeRabbit inference engine (AGENTS.md)
Files:
**⚙️ CodeRabbit configuration file
Files:
**/*⚙️ CodeRabbit configuration file
Files:
{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}⚙️ CodeRabbit configuration file
Files:
🔇 Additional comments (2)
📝 WalkthroughWalkthroughChangesCursor NFC offset handling
Estimated code review effort: 2 (Simple) | ~10 minutes 🚥 Pre-merge checks | ✅ 7✅ Passed checks (7 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…wigpine#1954) Cursor.modifyText builds newText from prefix + insert + suffix and hands it to Cursor.fromText, which NFC-normalizes the whole string. But the new cursor offset was computed as startOffset + insertString.normalize('NFC').length, normalizing the insert in isolation. When the inserted text begins with a combining mark that composes with the last character of the prefix (e.g. "e" + U+0301 -> "é"), the normalized newText is one UTF-16 unit shorter than that formula assumes, so the returned offset overshoots by one and the cursor lands past the following text — the next keystroke then edits the wrong spot. Measure the normalized prefix-plus-insert instead, so cross-boundary composition is accounted for. Reduces to the previous behavior whenever no boundary composition occurs (plain ASCII, astral emoji, insert at start).
Problem
Cursor.modifyTextbuildsnewText = prefix + insert + suffixand hands it toCursor.fromText, which NFC-normalizes the whole string. But the new cursor offset was computed asstartOffset + insertString.normalize('NFC').length, normalizing the insert in isolation:When the inserted text begins with a combining mark that composes with the last character of the prefix (e.g.
"e"+U+0301→"é"), the normalizednewTextis one UTF-16 unit shorter than that formula assumes, so the offset overshoots by one and the cursor lands past the following text — the next keystroke edits the wrong spot.Reachability
Cursor.insert→modifyText;insertis the mainline typed/pasted-input path in the REPL editor (useTextInput.ts,useVimInput.ts,useSearchInput.ts). A pasted fragment starting with a combining mark (or an IME delivering a lone combining mark) with the cursor right after a base letter hits it.Fix
Measure the normalized prefix-plus-insert for the offset, so cross-boundary composition is accounted for. Reduces to the previous behavior whenever no boundary composition occurs.
Test
Cursor.insertwith a combining acute between"e"and"X"→ cursor at offset 1 over"éX"(was 2); plus non-composing ASCII and astral-emoji controls.Summary by CodeRabbit
Bug Fixes
Tests