Repository navigation
Lexer: a line comment's lexeme was fabricated text, so every later token was mislocated; the literal matcher now consumes from the source - #12313
Conversation
… every token after a line comment was mislocated
v1 interpreter method_call.concat, Str receiver: arguments must be string-like
(native String, codepoint chain or list, or one codepoint for push), else
StringRealizationStraddle. It appended format!("{}", arg), so the v2 lexer's
line-comment lexeme became '//[32, 97, 98]' and every later ByteRange was
shifted past its true offset. free_call.concat joins a String with codepoint
chains into one String.
Claims: v2.test.tokenize.line_comment_extent (3 red on the old arm, 1 control).
RFM: native_string_concat_appends_an_arguments_display_rendering.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oncat change is withdrawn The seed-interpreter refusal broke a legitimate String+Int stringification in the required floor (generated_artifact_gate), because the seed's concat renders non-string arguments by design. The straddle is created upstream: lex_match_prefix answered the pattern literal (a native String) where every other matcher answers source codepoints, so the comment rule's sequence join met two carriers. lex_consume_prefix takes the prefix from the source; value unchanged, one carrier. lex_strip_prefix (now unused) deleted. RFM row renamed to lexer_literal_lexeme_straddles_its_sequence_partner. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Re review 71398's note that — sent from gentle-crane-869 |
briansrls
left a comment
There was a problem hiding this comment.
APPROVE at exact head 14028ce, for merge-queue landing only.
The repair is at the source of the carrier straddle: the literal matcher returns the consumed source codepoints, like its sequence partner, rather than returning the pattern String and relying on concat to interpret a different carrier. The prefix check remains in lex_match_prefix before its sole lex_consume_prefix call; empty and exhausted inputs retain explicit outcomes. The withdrawn global concat change is absent. No unrelated Rust/seed behavior is changed.
The production-rule witnesses discriminate the comment-text and extent failure (ordinary comment, bare //, and captured annotation), with the no-comment offset control. I inspected the diff and the supplied red/green evidence; I did not rerun those claims locally. Exact-head Actions run 36192906207 has all five jobs successful, including floor and emit-build.
This does not close the scalar-versus-byte offset question (#12285), nor establish structural impossibility of the underlying String/codepoint carrier straddle. Keep those qualifications; re-run/re-baseline location-dependent census attribution after landing rather than treating old shifted offsets as exact evidence. Combining comparison and consumption into one walk is a non-blocking follow-up, not a reason to hold this repair.
Require the actual merge_group candidate to pass against then-current main. This approval does not authorize direct merge or bypasses.
…869-comment-extent # Conflicts: # src/v2/compiler/01_tokenize.dag
briansrls
left a comment
There was a problem hiding this comment.
RE-APPROVE at exact head d598b86 for MERGE-QUEUE landing only.
This is an ancestry-preserving main refresh of the previously approved 14028ce: that commit is the first parent, and main 5bfb29c6cd730b53f76c5b7bd0bb84c938b05f71 is the second. The import conflict resolution takes main's std.algebra layout and retains removal of unused ListTailResult/list_tail. The guarded consumed-codepoints literal-matcher repair remains intact; the global concat workaround is not reintroduced.
#12285 is integrated. Exact-head Actions run 36252300674 completed all five jobs successfully. I inspected the floor log: all four line_comment_extent claims and all five byte_range_octet_unit claims were explicitly planned and passed. Its terminal reports planned=503, observed=503, passed=503, failed=0, refused=0, and D0-MEASURE, D0-PUBLISH and D0-ADJUDICATE succeeded.
Review basis is the merged source, commit ancestry, and CI evidence; I did not rerun the claims locally. This reapproval does not claim that the location-dependent whole-corpus census has been rebaselined.
Enqueue this exact head and require the actual merge_group candidate to pass against then-current main. No direct merge, squash bypass, or check bypass is authorized.
Every token after a line comment carried the wrong ByteRange: the lexer's literal matcher built a lexeme the seed could not join
Follow-up to #12309 (XL-2), asked for by quiet-seal-543. XL-5 needs exact loci.
The fault
Every line comment's lexeme was fabricated text:
// ab→//[32, 97, 98], and a bare//→//Empty.lex_walk_stepadvancesposby the lexeme's length, so every token after any line comment was mislocated, on pure ASCII. It was +39 after a 13-character comment, and past the end of the file inextdeps.bazel.build_event_stream. The captured annotation text was the rendering too. This is not #12285's unit question: #12285 would count the octets of the same wrong text. The two changes compose.Measured with a token dump over
program_assembly_phase_tokenize:////a//ab//abcdefghijWhere the chain breaks (§6b)
lex_match_prefixreturned the pattern literal as its lexeme. Every other matcher returns the codepoints it consumed from the source. On the seed these are two different carriers: a native String and a codepointValue::List.lex_rule_thunk's sequence arm joins the two lexemes withlist_append, which isconcat. For the comment ruleSequence(Literal "//", Repeat(..)), that join meets a String receiver and a codepoint list.concatrenders any non-string argument. That is its modelled behaviour: the signature is(a, b) -> String, andgenerated_artifact_gaterelies on it for Int.List<Int>. That is the named open thread ofstring_realization_straddle_detail.So the earliest link that can be justified away is the matcher that created the straddle.
Withdrawn first attempt. The first head of this PR made the seed's
concatrefuse non-string arguments. The required floor refused it:dag/gunbc/instruments/generated_artifact_gate.daglegitimately appends an Int to a String. That change is gone, and this PR no longer touches Rust.Change
v2.compiler.tokenize:lex_match_prefix→lex_consume_prefix, which takes the prefix's codepoints from the source. Every lexeme is one carrier, and the value is unchanged.lex_strip_prefix, now unused, is deleted.lexer_literal_lexeme_straddles_its_sequence_partner. Its ceiling is structurally impossible once String and codepoint sequences are one carrier.Evidence:
v2.test.tokenize.line_comment_extentThe route is the production one:
dag_lex_rules()throughlex_walk_artifact.tokens_without_a_comment_start_at_their_offsets(control)a_token_after_a_line_comment_starts_at_its_offseta_token_after_a_bare_line_comment_starts_at_its_offseta_line_comment_captures_its_own_text_and_extentBoth columns were run locally with
gunbc run --claim-runon a locally built seed. Every othersrc/v2/test/claim/tokenize/*module is green at this head: annotation_channel 10, string_literal_raw_newline 4, source_text_ingress 2, lex_match_thunk 1, lex_rule_dispatch 1.🤖 Generated with Claude Code