Skip to content

fix(semantic tokens): fix delta underflow after multiline tokens - #513

Merged
16bit-ykiko merged 1 commit into
mainfrom
fix/semtok-multiline-delta
Jul 16, 2026
Merged

16bit-ykiko merged 1 commit into
mainfrom
fix/semtok-multiline-delta

Conversation

@16bit-ykiko

@16bit-ykiko 16bit-ykiko commented Jul 16, 2026 •

Copy link
Copy Markdown
Member

Problem

When a line contains tokens after a multiline token ends on it (code following the closing quote of a raw string literal, or after a */ of a multiline block comment), the semantic-tokens delta encoding is corrupted. SemanticTokenEncoder::append splits a multiline token into per-line entries whose last piece starts at column 0 of the end line, but last_start_character was unconditionally set to the token's start column on its FIRST line. The next token's deltaStart = char - last_start_character then either underflows u32 (runtime-measured char=4294967292, destroying highlighting for the whole line) or silently shifts columns when it stays positive. A secondary flaw: last_line advanced to the token's end line even when nothing was emitted there (token ending exactly at a newline).

Fix

Delta computation and prev-position bookkeeping fold into a single emit(line, character, ...) taking absolute positions — the only reader/writer of last_line/last_start_character, updated only when an entry is actually emitted. The split loop tracks each piece's absolute position; piece lengths and the splitting itself are unchanged.

Testing

  • Tests written FIRST and confirmed failing on the old code: CodeAfterRawString and CodeAfterMultilineComment decode the delta stream and assert exact absolute columns; the test decoders accumulate in 64-bit so an underflow can never wrap back into plausible values (which had masked this bug). Before the fix the raw-string case decoded a token start beyond the line end and the comment case landed at columns 6/10 instead of 10/14; both now decode exactly.
  • 3-agent pre-PR review (correctness / style / tests) passed with no blocking findings; the correctness pass traced same-end-line, later-line, consecutive-multiline, newline-terminated, CRLF and zero-length-piece paths, and notes the fix also repairs the last_line update when nothing is emitted on the end line.
  • Full local gates: format clean, unit 987, integration 283, smoke 3/3 — all green (RelWithDebInfo).

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Semantic token emission now uses a shared absolute-position encoder for single-line and multiline tokens. Tests widen position tracking to 64-bit, tolerate unmappable lines, and cover tokens following raw strings and multiline comments.

Changes

Semantic token positions

Layer / File(s) Summary
Absolute-position token emission
src/feature/semantic_tokens.cpp
SemanticTokenEncoder::append emits token pieces through a shared helper that computes LSP deltas and maintains the previous position.
Position decoding and regression coverage
tests/unit/feature/semantic_tokens_tests.cpp
Decoders use 64-bit coordinates, preserve invalid ranges for unmappable tokens, and validate positions after raw strings and multiline comments.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

  • clice-io/clice#442: Touches the same semantic-token emission and position-decoding pipeline.

Suggested reviewers: myriad-dreamin

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main fix: semantic-token delta underflow after multiline tokens.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/semtok-multiline-delta

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/feature/semantic_tokens.cpp (1)

540-562: 🩺 Stability & Availability | 🔵 Trivial | 💤 Low value

Defensively guard against out-of-order tokens.

Although the token stream is sorted and merged before reaching the encoder, adding a guard here prevents any theoretical upstream sorting bug from underflowing the LSP deltas. An underflow here would corrupt the entire file's subsequent syntax highlighting in the editor.

🛡️ Proposed defensive guard
               SymbolKind kind,
               std::uint32_t modifiers) {
         if(token_length == 0) {
             return;
         }
+
+        if(line < last_line || (line == last_line && character < last_start_character)) {
+            return;
+        }
 
         auto delta_line = line - last_line;
         auto delta_start = delta_line == 0 ? character - last_start_character : character;
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/feature/semantic_tokens.cpp` around lines 540 - 562, Update the
semantic-token encoder’s emit method to detect tokens that precede the
previously emitted position before subtracting line or character values, and
skip or otherwise safely handle those out-of-order tokens without appending
invalid deltas or updating last_line/last_start_character. Preserve the existing
zero-length filtering and normal ordered-token encoding behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@src/feature/semantic_tokens.cpp`:
- Around line 540-562: Update the semantic-token encoder’s emit method to detect
tokens that precede the previously emitted position before subtracting line or
character values, and skip or otherwise safely handle those out-of-order tokens
without appending invalid deltas or updating last_line/last_start_character.
Preserve the existing zero-length filtering and normal ordered-token encoding
behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 0c278d9a-d631-4bea-a6b1-7196df65a817

📥 Commits

Reviewing files that changed from the base of the PR and between e308cd6 and 842f9a2.

📒 Files selected for processing (2)
  • src/feature/semantic_tokens.cpp
  • tests/unit/feature/semantic_tokens_tests.cpp

@16bit-ykiko
16bit-ykiko merged commit e0f148f into main Jul 16, 2026
22 checks passed
@16bit-ykiko
16bit-ykiko deleted the fix/semtok-multiline-delta branch July 16, 2026 15:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant