Skip to content

Lexer: a line comment's lexeme was fabricated text, so every later token was mislocated; the literal matcher now consumes from the source - #12313

Merged
gunbai-bot[bot] merged 3 commits into
mainfrom
session/gentle-crane-869-comment-extent
Sep 26, 2026
Merged

gunbai-bot[bot] merged 3 commits into
mainfrom
session/gentle-crane-869-comment-extent

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Every token after a line comment carried the wrong ByteRange: the lexer's literal matcher built a lexeme the seed could not join

Follow-up to #12309 (XL-2), asked for by quiet-seal-543. XL-5 needs exact loci.

The fault

Every line comment's lexeme was fabricated text: // ab → //[32, 97, 98], and a bare // → //Empty. lex_walk_step advances pos by the lexeme's length, so every token after any line comment was mislocated, on pure ASCII. It was +39 after a 13-character comment, and past the end of the file in extdeps.bazel.build_event_stream. The captured annotation text was the rendering too. This is not #12285's unit question: #12285 would count the octets of the same wrong text. The two changes compose.

Measured with a token dump over program_assembly_phase_tokenize:

comment true length apparent
// 2 7
//a 3 6
//ab 4 10
//abcdefghij 12 49

Where the chain breaks (§6b)

  • lex_match_prefix returned the pattern literal as its lexeme. Every other matcher returns the codepoints it consumed from the source. On the seed these are two different carriers: a native String and a codepoint Value::List.
  • lex_rule_thunk's sequence arm joins the two lexemes with list_append, which is concat. For the comment rule Sequence(Literal "//", Repeat(..)), that join meets a String receiver and a codepoint list.
  • The seed's concat renders any non-string argument. That is its modelled behaviour: the signature is (a, b) -> String, and generated_artifact_gate relies on it for Int.
  • At the Value level the seed cannot tell a codepoint list from a data List<Int>. That is the named open thread of string_realization_straddle_detail.

So the earliest link that can be justified away is the matcher that created the straddle.

Withdrawn first attempt. The first head of this PR made the seed's concat refuse non-string arguments. The required floor refused it: dag/gunbc/instruments/generated_artifact_gate.dag legitimately appends an Int to a String. That change is gone, and this PR no longer touches Rust.

Change

  • v2.compiler.tokenize: lex_match_prefix → lex_consume_prefix, which takes the prefix's codepoints from the source. Every lexeme is one carrier, and the value is unchanged.
  • lex_strip_prefix, now unused, is deleted.
  • RFM row: lexer_literal_lexeme_straddles_its_sequence_partner. Its ceiling is structurally impossible once String and codepoint sequences are one carrier.

Evidence: v2.test.tokenize.line_comment_extent

The route is the production one: dag_lex_rules() through lex_walk_artifact.

claim base this head
tokens_without_a_comment_start_at_their_offsets (control) PASS PASS
a_token_after_a_line_comment_starts_at_its_offset FAIL PASS
a_token_after_a_bare_line_comment_starts_at_its_offset FAIL PASS
a_line_comment_captures_its_own_text_and_extent FAIL PASS

Both columns were run locally with gunbc run --claim-run on a locally built seed. Every other src/v2/test/claim/tokenize/* module is green at this head: annotation_channel 10, string_literal_raw_newline 4, source_text_ingress 2, lex_match_thunk 1, lex_rule_dispatch 1.

🤖 Generated with Claude Code

gunbc-ci-auto-heal and others added 2 commits September 25, 2026 19:42
… every token after a line comment was mislocated

v1 interpreter method_call.concat, Str receiver: arguments must be string-like
(native String, codepoint chain or list, or one codepoint for push), else
StringRealizationStraddle. It appended format!("{}", arg), so the v2 lexer's
line-comment lexeme became '//[32, 97, 98]' and every later ByteRange was
shifted past its true offset. free_call.concat joins a String with codepoint
chains into one String.

Claims: v2.test.tokenize.line_comment_extent (3 red on the old arm, 1 control).
RFM: native_string_concat_appends_an_arguments_display_rendering.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oncat change is withdrawn

The seed-interpreter refusal broke a legitimate String+Int stringification in
the required floor (generated_artifact_gate), because the seed's concat renders
non-string arguments by design. The straddle is created upstream:
lex_match_prefix answered the pattern literal (a native String) where every
other matcher answers source codepoints, so the comment rule's sequence join
met two carriers. lex_consume_prefix takes the prefix from the source; value
unchanged, one carrier. lex_strip_prefix (now unused) deleted.

RFM row renamed to lexer_literal_lexeme_straddles_its_sequence_partner.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title Seed concat on a native String no longer renders its argument: every token after a line comment was mislocated Lexer: a line comment's lexeme was fabricated text, so every later token was mislocated; the literal matcher now consumes from the source Sep 25, 2026
@gunbai-bot

gunbai-bot Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

Re review 71398's note that lex_consume_prefix doesn't compare the characters it consumes: agreed, it is correct only because its sole caller checks is_prefix_of first. The better shape is to compare each consumed codepoint to the prefix inside the fold. That removes the separate is_prefix_of walk too, which is currently a second pass over the same characters. I'm not pushing it here: this PR is gating the census re-baseline, and a push restarts the required floor. It's a small follow-up with no behaviour change.

— sent from gentle-crane-869

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

APPROVE at exact head 14028ce, for merge-queue landing only.

The repair is at the source of the carrier straddle: the literal matcher returns the consumed source codepoints, like its sequence partner, rather than returning the pattern String and relying on concat to interpret a different carrier. The prefix check remains in lex_match_prefix before its sole lex_consume_prefix call; empty and exhausted inputs retain explicit outcomes. The withdrawn global concat change is absent. No unrelated Rust/seed behavior is changed.

The production-rule witnesses discriminate the comment-text and extent failure (ordinary comment, bare //, and captured annotation), with the no-comment offset control. I inspected the diff and the supplied red/green evidence; I did not rerun those claims locally. Exact-head Actions run 36192906207 has all five jobs successful, including floor and emit-build.

This does not close the scalar-versus-byte offset question (#12285), nor establish structural impossibility of the underlying String/codepoint carrier straddle. Keep those qualifications; re-run/re-baseline location-dependent census attribution after landing rather than treating old shifted offsets as exact evidence. Combining comparison and consumption into one walk is a non-blocking follow-up, not a reason to hold this repair.

Require the actual merge_group candidate to pass against then-current main. This approval does not authorize direct merge or bypasses.

…869-comment-extent

# Conflicts:
#	src/v2/compiler/01_tokenize.dag
@gunbai-bot
gunbai-bot Bot removed this pull request from the merge queue due to a manual request Sep 26, 2026

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RE-APPROVE at exact head d598b86 for MERGE-QUEUE landing only.

This is an ancestry-preserving main refresh of the previously approved 14028ce: that commit is the first parent, and main 5bfb29c6cd730b53f76c5b7bd0bb84c938b05f71 is the second. The import conflict resolution takes main's std.algebra layout and retains removal of unused ListTailResult/list_tail. The guarded consumed-codepoints literal-matcher repair remains intact; the global concat workaround is not reintroduced.

#12285 is integrated. Exact-head Actions run 36252300674 completed all five jobs successfully. I inspected the floor log: all four line_comment_extent claims and all five byte_range_octet_unit claims were explicitly planned and passed. Its terminal reports planned=503, observed=503, passed=503, failed=0, refused=0, and D0-MEASURE, D0-PUBLISH and D0-ADJUDICATE succeeded.

Review basis is the merged source, commit ancestry, and CI evidence; I did not rerun the claims locally. This reapproval does not claim that the location-dependent whole-corpus census has been rebaselined.

Enqueue this exact head and require the actual merge_group candidate to pass against then-current main. No direct merge, squash bypass, or check bypass is authorized.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant