diff --git a/docs/plans/inner-cost-lanes-scoping.md b/docs/plans/inner-cost-lanes-scoping.md index 29349a41f9f..b0da30f5e38 100644 --- a/docs/plans/inner-cost-lanes-scoping.md +++ b/docs/plans/inner-cost-lanes-scoping.md @@ -13,42 +13,66 @@ resolve does per entry. ## Lane A — frontend construction +**LANDED.** `v1.compiler.tokenize` `process_escapes_loop` now consumes the pre-decoded code points. +The measured result and the two corrections this lane produced are recorded below; the scoping that +follows is kept because Lane B is still open and the two share this note. + **Contained slice: `process_escapes_loop`.** -`src/v1/stage0/src/v1_compiler_tokenize.rs:882`. It is the one function in its file still walking a -raw `String` by index while the rest of the file has migrated to the pre-decoded `SourceRef` / -`source_chars` idiom. That is the whole defect: the surrounding file already has the fix. +`v1.compiler.tokenize` `process_escapes_loop`. It was the one function in its file still walking a +raw `String` by index while the rest of the file had migrated to the pre-decoded `SourceRef` / +`source_chars` idiom. That is the whole defect: the surrounding file already had the fix. ### The cost shape, read off the primitives -Both primitives it calls are branch-on-ASCII (`src/v1/stage0/src/v1_rt.rs:251,266`): +Both primitives it called are branch-on-ASCII (`v1_rt` `char_at`, `string_length`): | call | ASCII | non-ASCII | |---|---|---| -| `char_at(s, pos)` | `bytes[pos]`, O(1) | `s.chars().nth(pos)`, **O(pos)** | -| `string_length(s)` | `s.len()`, O(1) | `s.chars().count()`, **O(n)** | - -`process_escapes_loop` calls `string_length` on the *same unchanging source* on every iteration — +| `char_at(s, pos)` | `is_ascii()` scan then `bytes[pos]`, **O(n)** | `is_ascii()` scan then `s.chars().nth(pos)`, **O(n)** | +| `string_length(s)` | `is_ascii()` scan then `s.len()`, **O(n)** | `s.chars().count()`, **O(n)** | + +**Correction (2026-07-31, measured).** An earlier version of this table read the ASCII column as +O(1) and concluded "ASCII input: O(n), non-ASCII input: O(n²)". That is wrong, and the error was in +reading the fast path as free: both primitives *begin* with `s.is_ascii()`, which scans the whole +string on **every call**, so the ASCII branch is O(n) per call too. A per-character loop over either +primitive is therefore quadratic **regardless of encoding**. Measured directly on `v1_rt::char_at` +over a pure-ASCII string, per-character loop: n=4,000 1.50ms, 8,000 5.66ms, 16,000 22.03ms, 32,000 +86.88ms — ~3.9× per doubling. Non-ASCII is not a different asymptotic class, only a worse constant +(`nth` walks as well). This matters beyond this slice: it means every surviving `char_at`-indexed +loop in the corpus is quadratic, not just the ones handling non-ASCII text. + +`process_escapes_loop` called `string_length` on the *same unchanging source* on every iteration — several times per iteration across its branches — and `char_at` at up to four offsets per iteration. -So the realized cost is: - -- **ASCII input: O(n)** in time, but with one heap `String` allocated *per character*, accumulated - into `Rc>` and `join`ed at the end. -- **Non-ASCII input: O(n²)**, twice over — once through `char_at`'s `nth`, once through the - re-counted `string_length`. - -The non-ASCII path is the one that bites, and it is not exotic: any string literal containing a -non-ASCII character anywhere flips the entire scan for that literal onto the quadratic branch, -including the ASCII portions of it. +Second, smaller correction: the defect was not confined to `process_escapes_loop`. `scan_string_body` +had already walked `source_chars` correctly, then `join`ed the characters into a `String` **purely so +that `process_escapes` could re-split it** — a decompress→recompress round trip across the +`StringScanResult` boundary. The fix had to cross that boundary to be real. ### The fix Migrate the loop onto `source_chars` (the pre-decoded array the rest of the file already builds), so -both `char_at` and `string_length` become O(1) regardless of encoding, and replace the -`Rc>` accumulator with a single `String` buffer. That removes the quadratic *and* the -per-character allocation in one motion — the accumulator-copy class DESIGN §6 already rules is always -fixed regardless of realized n, because "n is small here" is not a time-stable fact. +both `char_at` and `string_length` disappear from the escape path entirely, and drop the +`Rc>` accumulator. That removes the quadratic *and* the per-character allocation in one +motion — the accumulator-copy class DESIGN §6 already rules is always fixed regardless of realized n, +because "n is small here" is not a time-stable fact. + +Landed shape: `StringScanResult.content` carries `List`, so `scan_string_body` stops joining and +`process_escapes` stops re-splitting; the accumulator is code points converted once by +`chars_to_string`; the escape table compares code points, which is what the rest of the file already +did (`ch == 61 && next_ch == 62` for `=>`). + +**Measured, one string literal, non-ASCII with escapes (before → after):** + +| literal chars | before | after | speedup | +|---|---|---|---| +| 2,769 | 7.79 ms | 0.90 ms | 8.6× | +| 11,019 | 41.34 ms | 1.44 ms | 28.7× | +| 44,019 | 583.9 ms | 5.95 ms | 98.1× | + +The speedup grows with n because the quadratic term is gone: after the change, 2× the input costs +2.01× then 2.05× the time. ### Oracle @@ -58,6 +82,33 @@ string literal with non-ASCII content and escapes, which separates the two imple asymptotically, plus a RED that a wrong escape decode changes the output. Both halves execute, per DESIGN §5: a typecheck and a grep are not consumers. +Discharged by `src/v1/stage0/tests/tokenize_escape_receipt.rs`, executed on both sides of the change: + +- Equivalence, corpus grain: `regen_stage0 --verify` reports `regen_divergence_count=0`. Regen + re-tokenizes all of `src/v1` and `dag` and emits the seed; a byte-identical seed across a + 2-generation fixed point is the corpus token-stream equivalence, established by execution rather + than asserted. +- Equivalence, decode grain: `escape_decode_table`, `malformed_hex_escape_declines_rather_than_fabricating`, + `unknown_escape_passthrough_is_retained` — green against the pre-migration seed *and* after. +- Separation: `escape_cost_is_linear_in_literal_length` asserts a **ratio** (4× input must not cost + ≥8× time) so it does not encode one machine's speed. Against the pre-migration seed it reds at + **14.2×**; after, it passes. That RED was observed by running it, not predicted. + + It ships **`#[ignore]`d — a benchmark, not a gate** (review 45416). A wall-clock assertion can + fail correct code when the larger run is the one that catches contention, and gating correctness + on timing is against hermetic-first test discipline. The deterministic alternative does not + rescue it: the tree's only work counter (`v1_rt::take_text_lookup_chars_walked`) is behind the + non-default `text_lookup_work_counter` feature and does not instrument `char_at`/`string_length` + at all, so a counter-based test would be `#[cfg(feature = ...)]` and equally non-gating while + additionally changing a core primitive. + + The distinction that matters: the oracle was **discharged by execution** in the landing PR — red + observed before, green after — which is what DESIGN §5 asks for. What is deferred is the standing + regression *guard*, and the repo's native form for that is a structural lens over the `Node` tree, + as `v2.lens.complexity_accumulator_copy` is for the copied-accumulator class. Such a lens would + also cover the ~30 raw-index sites the audit below found, which a per-function test never could — + so it is the better instrument, not merely the substitute. + ### Beyond the slice The named generalization is `SourceCursor` / `TokenBuilder` / `FrozenTokenStream` and enforceable @@ -67,6 +118,34 @@ priced against a second measured instance rather than authored speculatively (th Task #3 tracks auditing the v2 parser for the same class; if it is clean, the generalization is worth less and should say so. +**Audit result (task #3, 2026-07-31): the v2 parser is clean, and the generalization is worth less +than it looked — but the class is not extinct, and it is bigger than the parser.** + +`v2.compiler.tokenize` and `v2.compiler.parse` contain **zero** `char_at` / `string_length` +occurrences. v2's `String` is `FreeMonoid`, a cons list; the tokenizer consumes it through +`fold_source` / `string_head` / `string_tail`, and `list_tail` is `Cons { tail: t } => TailFound`, +i.e. O(1) with structural sharing. There is no position to re-index and no length to re-count, so the +defect cannot be written in that shape. A `SourceCursor` abstraction would therefore be modelling a +problem v2's carrier already dissolved — it should not be authored for the parser's sake. + +What the audit *did* find is that the raw-index idiom survives at 28 sites across `src/v2`, none of +them in the parser. The concentration is in shell-emission validators — +`v2.compiler.emit_orchestration` (`char_at` loops over a path, a binder, a test operand) and +`extdeps.languages.bash_orch_if` — each pairing a per-character `char_at` loop with a hoisted +`string_length`. Under the corrected cost table above these are quadratic, not linear: the earlier +reading would have excused them as "ASCII, so O(1) per call", which is exactly the reasoning the +measurement refutes. `v1.compiler.tokenize` itself still has two more in `sentinel_prefix_matches` / +`sentinel_suffix_matches`. + +None of these were touched here: they are short-input call sites, so the realized cost is small +today, and DESIGN §6's bare-minimum-cost rule says a proven cost-shape defect is fixed regardless of +realized n — which makes them real work, not non-work. They are named rather than folded in because +the honest lever is now a different one: **the cheapest fix is `char_at` / `string_length` +themselves.** Removing the per-call `is_ascii()` scan (or caching the decode) makes all 30 sites +linear at once, without an abstraction and without touching 30 call sites — a `v1_rt` change whose +authority is `v1.runtime_rust`. That is the second measured instance the "price it against one" +condition asked for, and it points at the primitive rather than at `SourceCursor`. + --- ## Lane B — union resolve diff --git a/src/v1/01_tokenize.dag b/src/v1/01_tokenize.dag index 98f4e896066..09dc05f7463 100644 --- a/src/v1/01_tokenize.dag +++ b/src/v1/01_tokenize.dag @@ -243,9 +243,9 @@ fn scan_number(source: SourceRef, pos: TokPos) -> ScanResult { } type StringScanResult - = ClosedString { content: String, end_pos: Int } - | InterpolationStart { content: String, end_pos: Int } - | UnterminatedString { content: String, end_pos: Int } + = ClosedString { content: List, end_pos: Int } + | InterpolationStart { content: List, end_pos: Int } + | UnterminatedString { content: List, end_pos: Int } fn scan_string(source: SourceRef, pos: TokPos) -> ScanResult { let span_start = pos.pos @@ -253,7 +253,7 @@ fn scan_string(source: SourceRef, pos: TokPos) -> ScanResult { let result = scan_string_body(source: source, pos: body_start, acc: []) match result { ClosedString { content, end_pos } => { - let processed = process_escapes(raw: content) + let processed = process_escapes(chars: content) let token = make_token(text: processed, span: make_file_span(file: source.file, start: span_start, end: end_pos + 1), shape: ShLitStr) ScanResult { pos: end_pos + 1, token: token, @@ -261,7 +261,7 @@ fn scan_string(source: SourceRef, pos: TokPos) -> ScanResult { } } InterpolationStart { content, end_pos } => { - let processed = process_escapes(raw: content) + let processed = process_escapes(chars: content) let token = make_token(text: processed, span: make_file_span(file: source.file, start: span_start, end: end_pos + 1), shape: ShStrBegin) ScanResult { pos: end_pos + 1, @@ -270,7 +270,7 @@ fn scan_string(source: SourceRef, pos: TokPos) -> ScanResult { } } UnterminatedString { content, end_pos } => { - let processed = process_escapes(raw: content) + let processed = process_escapes(chars: content) let token = make_token(text: processed, span: make_file_span(file: source.file, start: span_start, end: end_pos), shape: ShUnknown) ScanResult { pos: end_pos, token: token, @@ -284,7 +284,7 @@ fn scan_str_cont(source: SourceRef, pos: TokPos, span_start: Int) -> ScanResult let result = scan_string_body(source: source, pos: pos.pos, acc: []) match result { ClosedString { content, end_pos } => { - let processed = process_escapes(raw: content) + let processed = process_escapes(chars: content) let token = make_token(text: processed, span: make_file_span(file: source.file, start: span_start, end: end_pos + 1), shape: ShStrEnd) ScanResult { pos: end_pos + 1, token: token, @@ -292,7 +292,7 @@ fn scan_str_cont(source: SourceRef, pos: TokPos, span_start: Int) -> ScanResult } } InterpolationStart { content, end_pos } => { - let processed = process_escapes(raw: content) + let processed = process_escapes(chars: content) let token = make_token(text: processed, span: make_file_span(file: source.file, start: span_start, end: end_pos + 1), shape: ShStrMid) ScanResult { pos: end_pos + 1, @@ -301,7 +301,7 @@ fn scan_str_cont(source: SourceRef, pos: TokPos, span_start: Int) -> ScanResult } } UnterminatedString { content, end_pos } => { - let processed = process_escapes(raw: content) + let processed = process_escapes(chars: content) let token = make_token(text: processed, span: make_file_span(file: source.file, start: span_start, end: end_pos), shape: ShUnknown) ScanResult { pos: end_pos, token: token, @@ -311,25 +311,25 @@ fn scan_str_cont(source: SourceRef, pos: TokPos, span_start: Int) -> ScanResult } } -fn scan_string_body(source: SourceRef, pos: Int, acc: List) -> StringScanResult { +fn scan_string_body(source: SourceRef, pos: Int, acc: List) -> StringScanResult { if pos >= source_len(source: source) { - UnterminatedString { content: join(acc, separator: ""), end_pos: pos } + UnterminatedString { content: acc, end_pos: pos } } else { - let ch = source_char(source: source, pos: pos) - if ch == "\"" { - ClosedString { content: join(acc, separator: ""), end_pos: pos } - } else if ch == "\\" { + let ch = source.source_chars[pos] + if ch == 34 { + ClosedString { content: acc, end_pos: pos } + } else if ch == 92 { if pos + 1 < source_len(source: source) { - let escaped = source_char(source: source, pos: pos + 1) - scan_string_body(source: source, pos: pos + 2, acc: list_push(list_push(acc, "\\"), escaped)) + let escaped = source.source_chars[pos + 1] + scan_string_body(source: source, pos: pos + 2, acc: list_push(list_push(acc, 92), escaped)) } else { - UnterminatedString { content: join(list_push(acc, "\\"), separator: ""), end_pos: pos + 1 } + UnterminatedString { content: list_push(acc, 92), end_pos: pos + 1 } } - } else if ch == "{" { + } else if ch == 123 { if should_start_interpolation(source: source, pos: pos) { - InterpolationStart { content: join(acc, separator: ""), end_pos: pos } + InterpolationStart { content: acc, end_pos: pos } } else { - scan_string_body(source: source, pos: pos + 1, acc: list_push(acc, "{")) + scan_string_body(source: source, pos: pos + 1, acc: list_push(acc, 123)) } } else { scan_string_body(source: source, pos: pos + 1, acc: list_push(acc, ch)) @@ -346,16 +346,25 @@ fn should_start_interpolation(source: SourceRef, pos: Int) -> Bool { } } -fn process_escapes(raw: String) -> String { - process_escapes_loop(source: raw, pos: 0, acc: []) +fn process_escapes(chars: List) -> String { + process_escapes_loop(source: chars, pos: 0, acc: []) } +data code_point_at_accessor_note: String = "Why this accessor exists rather than indexing inline. `04_access.check_index_access_node` types a list index as OPTIONAL (`with_optional_cardinality`) — an out-of-range index has no value — but the Rust realization of an index is a bare element that panics out of range. The two disagree, so wherever the emitter acts on the declared type it emits a coercion the realization cannot satisfy: passing `chars[i]` straight to a declared fn parameter emits `.expect(...)` on an `i64` (E0599). Untouched positions — a let binding, a comparison, arithmetic, a builtin call — happen to line up, which is why the corpus has exactly one prior index-as-argument site and it calls the BUILTIN `from_code_point`. This accessor is the same shape as `source_code_point` and `source_char` above: a declared `-> Int` return is what carries the value across a fn boundary. It works because return position is not checked against the body (recorded on the finalization carrier, #7481), so it is borrowed load-bearing behavior, not a guarantee. Dissolves together with those two when the index model and its realization are reconciled — either the index realizes as an Option or it types as the element — at which point return-position typechecking would red all three at once." + +fn code_point_at(chars: List, pos: Int) -> Int { + chars[pos] +} + +data escape_code_point_table_note: String = "The escape table below is written in code points because the whole tokenizer is: scan_token compares `ch == 61 && next_ch == 62` for `=>`, source_scan_to_eol compares `== 10`. The decimal literals here are the same alphabet: 34 double-quote, 92 backslash, 110 `n`, 116 `t`, 120 `x`, 123 open-brace, 125 close-brace; the values they resolve TO are 10 line-feed and 9 tab. This function used to compare one-character Strings instead, which is what forced it to re-index a raw String and made it the last quadratic scan in the file." + +data escape_receipt_seed_growth_mark: String = "🟡 Seed-growth mark (v1-test class): src/v1/stage0/tests/tokenize_escape_receipt.rs is hand-written Rust added to the v1 seed tree, and it is counted here rather than left silent. WHAT IT CARRIES, GATING: the escape decode table, the malformed-\\x decline, and the retained unknown-escape passthrough — all deterministic. WHAT IT CARRIES, NON-GATING: the cost separation that reds at 14.2x against the pre-migration implementation, the discriminating half of this lane's oracle, which a corpus equivalence check cannot supply because equivalence is satisfied by changing nothing. That half is an `#[ignore]`d benchmark rather than a required test, because gating correctness on wall clock can fail correct code when the larger run catches contention (review 45416), and the deterministic alternative does not rescue it: the only work counter in the tree sits behind the non-default `text_lookup_work_counter` feature and does not instrument `char_at`/`string_length` at all, so a counter-based test would be equally non-gating while also changing a core primitive. The oracle was still discharged BY EXECUTION in the landing PR — observed red before, green after — which is the DESIGN §5 bar; what is deferred is the standing regression GUARD, and the repo's native form for that is a structural lens over the Node tree, as v2.lens.complexity_accumulator_copy is for the copied-accumulator class. WHY HAND RUST AND NOT A .dag WITNESS: the .dag form is a claim witness, and enrolling one bumps gunbc.ci_spec's ci_floor_declared_resolve_count, which that gate's own note reserves for an operator-signed line; a session cannot raise it for its own receipt. So this is a DEFERRAL with a named owner, not a preference. It does not enter the hand-maintained stage0 ratchet: hand_maintained_stage0_filenames is built from seed-retained and emitter-produced `pub mod` basenames under src/v1/stage0/src/, and a tests/ target is not a crate module. DISSOLVES ON either trigger, whichever comes first: the resolve-count line is signed and these claims move to a dag/test/claim witness over the same tokenize entry point; or the v1 terminal deletion path reaches this tree, at which point the coverage must be carried per the v1-test-migration bar and the file deletes with src/v1. SEPARATE, SMALLER TRIGGER for the benchmark half alone: a structural lens that reds a per-character raw-String index walk (char_at / string_length in a loop) supersedes it, and would also cover the ~30 sites this lane's audit found still carrying that shape — at which point the `#[ignore]`d test deletes rather than being promoted. Precedent for this shape: namespace_occurrence_serde_seed_test_dissolution." + data hex_escape_note: String = "\\xNN decodes to the byte NN names. Before this arm every \\x fell through to the unknown-escape branch below, which preserves the backslash literally — so every ANSI code authored in .dag (extdeps.render.ansi.csi_esc, extdeps.render.terminal's reset, gunbc.ci_render's red/reset) emitted the six characters backslash-x-1-b instead of ESC, and no colour this repo authored had ever rendered. The witness that was supposed to catch it compared one undecoded literal against another, so it agreed with itself and could not fail (DESIGN §5 — a check satisfied by editing the declaration while the realization lies)." -data unknown_escape_passthrough_frontier: String = "🟡 dissolve-on: tokenizer_unknown_escape_strict_close (opened 2026-07-25) — the else-arm below resolves an UNRECOGNIZED escape to concat(backslash, next), so \\s silently means backslash-s and is indistinguishable from a typo'd escape. That is a closed-vocabulary violation (DESIGN §4: the vocabulary is closed; §5: a failure arm refuses, never widens) and it is knowingly RETAINED, not accepted: refusing today would red 142 occurrences across 38 files (textual census 2026-07-25 — grep over .dag, so an upper bound, and DESIGN's own caution about grep census applies), overwhelmingly regex fragments (\\' 28, \\. 19, \\| 13, \\/ 12, \\B 10) that intend the backslash to survive. NOT the four sites a \\x-only count suggests. DISSOLVES WHEN a raw-string form lands to carry regex fragments, the 142 occurrences migrate onto it, and this arm becomes a typed located parse refusal (the Rust rule: unknown escape is an error, a literal backslash is spelled \\\\). Because that is a PARSE-SEMANTICS change over a tokenizer every lane's probes compile through, the strict close lands with a corpus-wide compile receipt, never the site migration alone." +data unknown_escape_passthrough_frontier: String = "🟡 dissolve-on: tokenizer_unknown_escape_strict_close (opened 2026-07-25) — the else-arm below resolves an UNRECOGNIZED escape by pushing backslash and next unchanged, so \\s silently means backslash-s and is indistinguishable from a typo'd escape. That is a closed-vocabulary violation (DESIGN §4: the vocabulary is closed; §5: a failure arm refuses, never widens) and it is knowingly RETAINED, not accepted: refusing today would red 142 occurrences across 38 files (textual census 2026-07-25 — grep over .dag, so an upper bound, and DESIGN's own caution about grep census applies), overwhelmingly regex fragments (\\' 28, \\. 19, \\| 13, \\/ 12, \\B 10) that intend the backslash to survive. NOT the four sites a \\x-only count suggests. DISSOLVES WHEN a raw-string form lands to carry regex fragments, the 142 occurrences migrate onto it, and this arm becomes a typed located parse refusal (the Rust rule: unknown escape is an error, a literal backslash is spelled \\\\). Because that is a PARSE-SEMANTICS change over a tokenizer every lane's probes compile through, the strict close lands with a corpus-wide compile receipt, never the site migration alone." -fn hex_digit_value(ch: String) -> Int? { - let cp = code_point(ch) +fn hex_digit_value(cp: Int) -> Int? { if cp >= 48 && cp <= 57 { Present { value: cp - 48 } } else { @@ -371,39 +380,41 @@ fn hex_digit_value(ch: String) -> Int? { } } -fn hex_escape_char(hi: String, lo: String) -> String? { - match hex_digit_value(ch: hi) { - Present { value: h } => match hex_digit_value(ch: lo) { - Present { value: l } => Present { value: from_code_point(h * 16 + l) } +fn hex_escape_char(hi: Int, lo: Int) -> Int? { + match hex_digit_value(cp: hi) { + Present { value: h } => match hex_digit_value(cp: lo) { + Present { value: l } => Present { value: h * 16 + l } Absent => none } Absent => none } } -fn process_escapes_loop(source: String, pos: Int, acc: List) -> String { - if pos >= string_length(s: source) { - join(acc, separator: "") +fn process_escapes_loop(source: List, pos: Int, acc: List) -> String { + if pos >= count(source) { + chars_to_string(chars: acc, start: 0, end: count(acc)) } else { - let ch = char_at(s: source, pos: pos) - if ch == "\\" && pos + 1 < string_length(s: source) { - let next = char_at(s: source, pos: pos + 1) - if next == "x" && pos + 3 < string_length(s: source) { - match hex_escape_char(hi: char_at(s: source, pos: pos + 2), lo: char_at(s: source, pos: pos + 3)) { + let ch = code_point_at(chars: source, pos: pos) + if ch == 92 && pos + 1 < count(source) { + let next = code_point_at(chars: source, pos: pos + 1) + if next == 120 && pos + 3 < count(source) { + let hi = code_point_at(chars: source, pos: pos + 2) + let lo = code_point_at(chars: source, pos: pos + 3) + match hex_escape_char(hi: hi, lo: lo) { Present { value: decoded } => process_escapes_loop(source: source, pos: pos + 4, acc: list_push(acc, decoded)) Absent => - process_escapes_loop(source: source, pos: pos + 2, acc: list_push(acc, concat("\\", next))) + process_escapes_loop(source: source, pos: pos + 2, acc: list_push(list_push(acc, 92), next)) } } else { - let resolved = if next == "\"" { "\"" } - else if next == "\\" { "\\" } - else if next == "n" { "\n" } - else if next == "t" { "\t" } - else if next == "{" { "{" } - else if next == "}" { "}" } - else { concat("\\", next) } - process_escapes_loop(source: source, pos: pos + 2, acc: list_push(acc, resolved)) + let acc_with_escape = if next == 34 { list_push(acc, 34) } + else if next == 92 { list_push(acc, 92) } + else if next == 110 { list_push(acc, 10) } + else if next == 116 { list_push(acc, 9) } + else if next == 123 { list_push(acc, 123) } + else if next == 125 { list_push(acc, 125) } + else { list_push(list_push(acc, 92), next) } + process_escapes_loop(source: source, pos: pos + 2, acc: acc_with_escape) } } else { process_escapes_loop(source: source, pos: pos + 1, acc: list_push(acc, ch)) diff --git a/src/v1/stage0/src/v1_compiler_tokenize.rs b/src/v1/stage0/src/v1_compiler_tokenize.rs index b9cbaa7be5e..0349048e05e 100644 --- a/src/v1/stage0/src/v1_compiler_tokenize.rs +++ b/src/v1/stage0/src/v1_compiler_tokenize.rs @@ -601,12 +601,12 @@ pub fn scan_number(source: Rc, pos: Rc) -> Rc { #[derive(Debug, Clone, PartialEq, serde::Serialize, serde::Deserialize)] #[serde(tag = "_variant")] pub enum StringScanResult { - ClosedString { content: String, end_pos: i64 }, - InterpolationStart { content: String, end_pos: i64 }, - UnterminatedString { content: String, end_pos: i64 }, + ClosedString { content: Rc>, end_pos: i64 }, + InterpolationStart { content: Rc>, end_pos: i64 }, + UnterminatedString { content: Rc>, end_pos: i64 }, } impl StringScanResult { - pub fn content(&self) -> String { + pub fn content(&self) -> Rc> { match self { StringScanResult::ClosedString { content: __val, .. } => __val.clone(), StringScanResult::InterpolationStart { content: __val, .. } => __val.clone(), @@ -749,53 +749,51 @@ pub fn scan_str_cont(source: Rc, pos: Rc, span_start: i64) -> pub fn scan_string_body( mut source: Rc, mut pos: i64, - mut acc: Rc>, + mut acc: Rc>, ) -> Rc { loop { if (pos.clone() >= source_len(source.clone())) { break Rc::new(StringScanResult::UnterminatedString { - content: acc.clone().join(&"".to_string()), + content: acc.clone(), end_pos: pos.clone(), }); } else { - let ch = source_char(source.clone(), pos.clone()); - if (ch.clone() == "\"".to_string()) { + let ch = source.source_chars.clone()[(pos.clone()) as usize].clone(); + if (ch.clone() == 34) { break Rc::new(StringScanResult::ClosedString { - content: acc.clone().join(&"".to_string()), + content: acc.clone(), end_pos: pos.clone(), }); } else { - if (ch.clone() == "\\".to_string()) { + if (ch.clone() == 92) { if ((pos.clone() + 1) < source_len(source.clone())) { - let escaped = source_char(source.clone(), (pos.clone() + 1)); + let escaped = + source.source_chars.clone()[(pos.clone() + 1) as usize].clone(); { let __tco_0 = (pos + 2); - let __tco_1 = v1_rt::rc_list_push( - v1_rt::rc_list_push(acc, "\\".to_string()), - escaped.clone(), - ); + let __tco_1 = + v1_rt::rc_list_push(v1_rt::rc_list_push(acc, 92), escaped.clone()); pos = __tco_0; acc = __tco_1; continue; } } else { break Rc::new(StringScanResult::UnterminatedString { - content: v1_rt::rc_list_push(acc.clone(), "\\".to_string()) - .join(&"".to_string()), + content: v1_rt::rc_list_push(acc.clone(), 92), end_pos: (pos.clone() + 1), }); } } else { - if (ch.clone() == "{".to_string()) { + if (ch.clone() == 123) { if should_start_interpolation(source.clone(), pos.clone()) { break Rc::new(StringScanResult::InterpolationStart { - content: acc.clone().join(&"".to_string()), + content: acc.clone(), end_pos: pos.clone(), }); } else { { let __tco_0 = (pos + 1); - let __tco_1 = v1_rt::rc_list_push(acc, "{".to_string()); + let __tco_1 = v1_rt::rc_list_push(acc, 123); pos = __tco_0; acc = __tco_1; continue; @@ -828,8 +826,39 @@ pub fn should_start_interpolation(source: Rc, pos: i64) -> bool { } } -pub fn process_escapes(raw: String) -> String { - process_escapes_loop(raw.clone(), 0, Rc::new(vec![])) +pub fn process_escapes(chars: Rc>) -> String { + process_escapes_loop(chars.clone(), 0, Rc::new(vec![])) +} + +pub fn code_point_at_accessor_note() -> String { + thread_local! { + static CACHED: String = { + "Why this accessor exists rather than indexing inline. `04_access.check_index_access_node` types a list index as OPTIONAL (`with_optional_cardinality`) — an out-of-range index has no value — but the Rust realization of an index is a bare element that panics out of range. The two disagree, so wherever the emitter acts on the declared type it emits a coercion the realization cannot satisfy: passing `chars[i]` straight to a declared fn parameter emits `.expect(...)` on an `i64` (E0599). Untouched positions — a let binding, a comparison, arithmetic, a builtin call — happen to line up, which is why the corpus has exactly one prior index-as-argument site and it calls the BUILTIN `from_code_point`. This accessor is the same shape as `source_code_point` and `source_char` above: a declared `-> Int` return is what carries the value across a fn boundary. It works because return position is not checked against the body (recorded on the finalization carrier, #7481), so it is borrowed load-bearing behavior, not a guarantee. Dissolves together with those two when the index model and its realization are reconciled — either the index realizes as an Option or it types as the element — at which point return-position typechecking would red all three at once.".to_string() + }; + } + CACHED.with(|c: &String| c.clone()) +} + +pub fn code_point_at(chars: Rc>, pos: i64) -> i64 { + chars.clone()[(pos.clone()) as usize].clone() +} + +pub fn escape_code_point_table_note() -> String { + thread_local! { + static CACHED: String = { + "The escape table below is written in code points because the whole tokenizer is: scan_token compares `ch == 61 && next_ch == 62` for `=>`, source_scan_to_eol compares `== 10`. The decimal literals here are the same alphabet: 34 double-quote, 92 backslash, 110 `n`, 116 `t`, 120 `x`, 123 open-brace, 125 close-brace; the values they resolve TO are 10 line-feed and 9 tab. This function used to compare one-character Strings instead, which is what forced it to re-index a raw String and made it the last quadratic scan in the file.".to_string() + }; + } + CACHED.with(|c: &String| c.clone()) +} + +pub fn escape_receipt_seed_growth_mark() -> String { + thread_local! { + static CACHED: String = { + "🟡 Seed-growth mark (v1-test class): src/v1/stage0/tests/tokenize_escape_receipt.rs is hand-written Rust added to the v1 seed tree, and it is counted here rather than left silent. WHAT IT CARRIES, GATING: the escape decode table, the malformed-\\x decline, and the retained unknown-escape passthrough — all deterministic. WHAT IT CARRIES, NON-GATING: the cost separation that reds at 14.2x against the pre-migration implementation, the discriminating half of this lane's oracle, which a corpus equivalence check cannot supply because equivalence is satisfied by changing nothing. That half is an `#[ignore]`d benchmark rather than a required test, because gating correctness on wall clock can fail correct code when the larger run catches contention (review 45416), and the deterministic alternative does not rescue it: the only work counter in the tree sits behind the non-default `text_lookup_work_counter` feature and does not instrument `char_at`/`string_length` at all, so a counter-based test would be equally non-gating while also changing a core primitive. The oracle was still discharged BY EXECUTION in the landing PR — observed red before, green after — which is the DESIGN §5 bar; what is deferred is the standing regression GUARD, and the repo's native form for that is a structural lens over the Node tree, as v2.lens.complexity_accumulator_copy is for the copied-accumulator class. WHY HAND RUST AND NOT A .dag WITNESS: the .dag form is a claim witness, and enrolling one bumps gunbc.ci_spec's ci_floor_declared_resolve_count, which that gate's own note reserves for an operator-signed line; a session cannot raise it for its own receipt. So this is a DEFERRAL with a named owner, not a preference. It does not enter the hand-maintained stage0 ratchet: hand_maintained_stage0_filenames is built from seed-retained and emitter-produced `pub mod` basenames under src/v1/stage0/src/, and a tests/ target is not a crate module. DISSOLVES ON either trigger, whichever comes first: the resolve-count line is signed and these claims move to a dag/test/claim witness over the same tokenize entry point; or the v1 terminal deletion path reaches this tree, at which point the coverage must be carried per the v1-test-migration bar and the file deletes with src/v1. SEPARATE, SMALLER TRIGGER for the benchmark half alone: a structural lens that reds a per-character raw-String index walk (char_at / string_length in a loop) supersedes it, and would also cover the ~30 sites this lane's audit found still carrying that shape — at which point the `#[ignore]`d test deletes rather than being promoted. Precedent for this shape: namespace_occurrence_serde_seed_test_dissolution.".to_string() + }; + } + CACHED.with(|c: &String| c.clone()) } pub fn hex_escape_note() -> String { @@ -844,58 +873,54 @@ pub fn hex_escape_note() -> String { pub fn unknown_escape_passthrough_frontier() -> String { thread_local! { static CACHED: String = { - "🟡 dissolve-on: tokenizer_unknown_escape_strict_close (opened 2026-07-25) — the else-arm below resolves an UNRECOGNIZED escape to concat(backslash, next), so \\s silently means backslash-s and is indistinguishable from a typo'd escape. That is a closed-vocabulary violation (DESIGN §4: the vocabulary is closed; §5: a failure arm refuses, never widens) and it is knowingly RETAINED, not accepted: refusing today would red 142 occurrences across 38 files (textual census 2026-07-25 — grep over .dag, so an upper bound, and DESIGN's own caution about grep census applies), overwhelmingly regex fragments (\\' 28, \\. 19, \\| 13, \\/ 12, \\B 10) that intend the backslash to survive. NOT the four sites a \\x-only count suggests. DISSOLVES WHEN a raw-string form lands to carry regex fragments, the 142 occurrences migrate onto it, and this arm becomes a typed located parse refusal (the Rust rule: unknown escape is an error, a literal backslash is spelled \\\\). Because that is a PARSE-SEMANTICS change over a tokenizer every lane's probes compile through, the strict close lands with a corpus-wide compile receipt, never the site migration alone.".to_string() + "🟡 dissolve-on: tokenizer_unknown_escape_strict_close (opened 2026-07-25) — the else-arm below resolves an UNRECOGNIZED escape by pushing backslash and next unchanged, so \\s silently means backslash-s and is indistinguishable from a typo'd escape. That is a closed-vocabulary violation (DESIGN §4: the vocabulary is closed; §5: a failure arm refuses, never widens) and it is knowingly RETAINED, not accepted: refusing today would red 142 occurrences across 38 files (textual census 2026-07-25 — grep over .dag, so an upper bound, and DESIGN's own caution about grep census applies), overwhelmingly regex fragments (\\' 28, \\. 19, \\| 13, \\/ 12, \\B 10) that intend the backslash to survive. NOT the four sites a \\x-only count suggests. DISSOLVES WHEN a raw-string form lands to carry regex fragments, the 142 occurrences migrate onto it, and this arm becomes a typed located parse refusal (the Rust rule: unknown escape is an error, a literal backslash is spelled \\\\). Because that is a PARSE-SEMANTICS change over a tokenizer every lane's probes compile through, the strict close lands with a corpus-wide compile receipt, never the site migration alone.".to_string() }; } CACHED.with(|c: &String| c.clone()) } -pub fn hex_digit_value(ch: String) -> Option { - { - let cp = v1_rt::code_point(ch.clone()); - if ((cp.clone() >= 48) && (cp.clone() <= 57)) { - Some((cp.clone() - 48)) +pub fn hex_digit_value(cp: i64) -> Option { + if ((cp.clone() >= 48) && (cp.clone() <= 57)) { + Some((cp.clone() - 48)) + } else { + if ((cp.clone() >= 97) && (cp.clone() <= 102)) { + Some((cp.clone() - 87)) } else { - if ((cp.clone() >= 97) && (cp.clone() <= 102)) { - Some((cp.clone() - 87)) + if ((cp.clone() >= 65) && (cp.clone() <= 70)) { + Some((cp.clone() - 55)) } else { - if ((cp.clone() >= 65) && (cp.clone() <= 70)) { - Some((cp.clone() - 55)) - } else { - None - } + None } } } } -pub fn hex_escape_char(hi: String, lo: String) -> Option { +pub fn hex_escape_char(hi: i64, lo: i64) -> Option { match hex_digit_value(hi.clone()) { Some(h) => match hex_digit_value(lo.clone()) { - Some(l) => Some(v1_rt::from_code_point(((h.clone() * 16) + l.clone()))), + Some(l) => Some(((h.clone() * 16) + l.clone())), None => None, }, None => None, } } -pub fn process_escapes_loop(mut source: String, mut pos: i64, mut acc: Rc>) -> String { +pub fn process_escapes_loop( + mut source: Rc>, + mut pos: i64, + mut acc: Rc>, +) -> String { loop { - if (pos.clone() >= v1_rt::string_length(&source)) { - break acc.clone().join(&"".to_string()); + if (pos.clone() >= (source.clone().len() as i64)) { + break v1_rt::chars_to_string(&acc, 0, (acc.clone().len() as i64)); } else { - let ch = v1_rt::char_at(&source, pos.clone()); - if ((ch.clone() == "\\".to_string()) - && ((pos.clone() + 1) < v1_rt::string_length(&source))) - { - let next = v1_rt::char_at(&source, (pos.clone() + 1)); - if ((next.clone() == "x".to_string()) - && ((pos.clone() + 3) < v1_rt::string_length(&source))) - { - match hex_escape_char( - v1_rt::char_at(&source, (pos.clone() + 2)), - v1_rt::char_at(&source, (pos.clone() + 3)), - ) { + let ch = code_point_at(source.clone(), pos.clone()); + if ((ch.clone() == 92) && ((pos.clone() + 1) < (source.clone().len() as i64))) { + let next = code_point_at(source.clone(), (pos.clone() + 1)); + if ((next.clone() == 120) && ((pos.clone() + 3) < (source.clone().len() as i64))) { + let hi = code_point_at(source.clone(), (pos.clone() + 2)); + let lo = code_point_at(source.clone(), (pos.clone() + 3)); + match hex_escape_char(hi.clone(), lo.clone()) { Some(decoded) => { let __tco_0 = (pos + 4); let __tco_1 = v1_rt::rc_list_push(acc, decoded.clone()); @@ -905,35 +930,36 @@ pub fn process_escapes_loop(mut source: String, mut pos: i64, mut acc: Rc { let __tco_0 = (pos + 2); - let __tco_1 = v1_rt::rc_list_push( - acc, - v1_rt::concat("\\".to_string(), next.clone()), - ); + let __tco_1 = + v1_rt::rc_list_push(v1_rt::rc_list_push(acc, 92), next.clone()); pos = __tco_0; acc = __tco_1; continue; } } } else { - let resolved = if (next.clone() == "\"".to_string()) { - "\"".to_string() + let acc_with_escape = if (next.clone() == 34) { + v1_rt::rc_list_push(acc.clone(), 34) } else { - if (next.clone() == "\\".to_string()) { - "\\".to_string() + if (next.clone() == 92) { + v1_rt::rc_list_push(acc.clone(), 92) } else { - if (next.clone() == "n".to_string()) { - "\n".to_string() + if (next.clone() == 110) { + v1_rt::rc_list_push(acc.clone(), 10) } else { - if (next.clone() == "t".to_string()) { - "\t".to_string() + if (next.clone() == 116) { + v1_rt::rc_list_push(acc.clone(), 9) } else { - if (next.clone() == "{".to_string()) { - "{".to_string() + if (next.clone() == 123) { + v1_rt::rc_list_push(acc.clone(), 123) } else { - if (next.clone() == "}".to_string()) { - "}".to_string() + if (next.clone() == 125) { + v1_rt::rc_list_push(acc.clone(), 125) } else { - v1_rt::concat("\\".to_string(), next.clone()) + v1_rt::rc_list_push( + v1_rt::rc_list_push(acc.clone(), 92), + next.clone(), + ) } } } @@ -942,7 +968,7 @@ pub fn process_escapes_loop(mut source: String, mut pos: i64, mut acc: Rc 44,019 chars 584ms, ~3.9x per doubling; after, 5.95ms at 44,019. +//! +//! Why not gating: a wall-clock assertion can fail correct code when the larger run is the one +//! that catches contention, and gating correctness on timing is against the hermetic-first test +//! discipline (review 45416). The deterministic alternative does not rescue it either — the only +//! work counter in the tree (`v1_rt::take_text_lookup_chars_walked`) sits behind the non-default +//! `text_lookup_work_counter` feature, and `char_at`/`string_length` are not instrumented into it +//! at all, so a counter-based test would be `#[cfg(feature = ...)]` and equally non-gating while +//! also requiring a change to a core primitive. The durable regression guard for this class is a +//! structural lens over the `Node` tree, the way `v2.lens.complexity_accumulator_copy` guards the +//! copied-accumulator class — named as the dissolution trigger on +//! `escape_receipt_seed_growth_mark`, not authored here. +//! +//! Run it deliberately: +//! cargo test -p v1-compiler --release --test tokenize_escape_receipt -- --ignored --nocapture +//! +//! Both halves drive `tokenize`, the real consumer, rather than the escape helpers directly: +//! the helpers' signatures changed in the migration, and a receipt that could not run against +//! both sides could not have shown the separation. + +use std::time::{Duration, Instant}; +use v1_compiler::v1_compiler_tokenize::tokenize; +use v1_compiler::v1_std_core::TokenShape; + +/// Tokenize a single string literal and return the decoded token text. +fn decode_literal(body: &str) -> String { + let toks = tokenize( + format!("data x: String = \"{}\"", body), + "escape_receipt.dag".to_string(), + ); + let lit = toks + .iter() + .find(|t| t.shape == TokenShape::ShLitStr) + .unwrap_or_else(|| panic!("no ShLitStr token for body {:?}", body)); + lit.text.clone() +} + +#[test] +fn escape_decode_table() { + // The recognized escape vocabulary, exactly as `process_escapes_loop` declares it. + assert_eq!( + decode_literal(r"a\nb"), + "a\nb", + "\\n must decode to line feed" + ); + assert_eq!(decode_literal(r"a\tb"), "a\tb", "\\t must decode to tab"); + assert_eq!( + decode_literal(r"a\\b"), + r"a\b", + "\\\\ must decode to one backslash" + ); + assert_eq!( + decode_literal(r"a\{b"), + "a{b", + "escaped open brace must decode literally, not interpolate" + ); + assert_eq!( + decode_literal(r"a\}b"), + "a}b", + "escaped close brace must decode literally" + ); + + // \xNN. This is the arm whose absence silently emitted six characters where an ANSI + // escape belonged -- see `hex_escape_note` in src/v1/01_tokenize.dag. + assert_eq!( + decode_literal(r"\x1b[0m"), + "\u{1b}[0m", + "\\x1b must decode to ESC" + ); + assert_eq!(decode_literal(r"\x41"), "A", "\\x41 must decode to 'A'"); + assert_eq!(decode_literal(r"\x00"), "\u{0}", "\\x00 must decode to NUL"); + assert_eq!( + decode_literal(r"\x0d"), + "\r", + "\\x0d must decode to carriage return" + ); + + // Non-ASCII passes through untouched, and -- the point of the migration -- mixing it with + // escapes decodes identically to the pure-ASCII case. + assert_eq!( + decode_literal("héllo—🟡"), + "héllo—🟡", + "non-ASCII must pass through" + ); + assert_eq!( + decode_literal(r"é\n—\x41🟡"), + "é\n—A🟡", + "escapes must decode the same way when non-ASCII shares the literal" + ); +} + +/// The `\xNN` arm decodes only when BOTH digits are hex; otherwise it declines and the +/// backslash survives. This is the discriminating pair for the hex arm: an implementation +/// that fabricated a value for malformed input, or that dropped the backslash, fails here +/// while `escape_decode_table` alone would still pass. +#[test] +fn malformed_hex_escape_declines_rather_than_fabricating() { + assert_eq!( + decode_literal(r"\xzz"), + r"\xzz", + "non-hex digits must not decode" + ); + assert_eq!( + decode_literal(r"\x1z"), + r"\x1z", + "a single bad digit must not decode" + ); + // Truncated at end of literal: no digits to read at all. + assert_eq!(decode_literal(r"\x"), r"\x", "a bare \\x must not decode"); +} + +/// `\s` is not in the vocabulary and today resolves to backslash-s. That is a knowingly +/// RETAINED closed-vocabulary violation, not an accepted one -- `unknown_escape_passthrough_frontier` +/// in src/v1/01_tokenize.dag carries the census and the dissolution trigger. Pinned here so the +/// migration cannot change it silently; this assertion flips when that frontier closes. +#[test] +fn unknown_escape_passthrough_is_retained() { + assert_eq!(decode_literal(r"a\sb"), r"a\sb"); + assert_eq!(decode_literal(r"\."), r"\."); +} + +/// Best-of-`n` wall time for tokenizing one literal of `repeats` units. Minimum, not mean: +/// scheduler noise only ever adds time, so the minimum is the least noisy estimator here. +fn best_tokenize_time(repeats: usize, samples: usize) -> Duration { + // Non-ASCII (é, —, 🟡, ß) mixed with both escape shapes. The non-ASCII is what put the + // pre-migration implementation on its quadratic branch. + let unit = "é\\n—ü\\x41🟡ß"; + let mut body = String::with_capacity(unit.len() * repeats); + for _ in 0..repeats { + body.push_str(unit); + } + let src = format!("data x: String = \"{}\"", body); + + let mut best = Duration::MAX; + for _ in 0..samples { + let s = src.clone(); + let t0 = Instant::now(); + let toks = tokenize(s, "escape_cost.dag".to_string()); + let dt = t0.elapsed(); + assert!(!toks.is_empty()); + best = best.min(dt); + } + best +} + +/// NOT A GATE. `#[ignore]`d on purpose -- see the module header. This is the cost-shape +/// benchmark: run it by hand when touching the escape path, and read the printed ratio rather +/// than trusting the bound. It is kept executable because it is what established the result +/// (14.2x before, ~4x after), not because a wall-clock number belongs in a required suite. +#[test] +#[ignore = "wall-clock benchmark, not a correctness gate: run with --ignored"] +fn escape_cost_is_linear_in_literal_length() { + let small = best_tokenize_time(1_000, 3); + let large = best_tokenize_time(4_000, 3); + + // 4x the input. Linear => ~4x the time. Quadratic => ~16x. + let ratio = large.as_secs_f64() / small.as_secs_f64(); + println!( + "escape cost: 1k units {:?}, 4k units {:?} -> {:.2}x for 4x input", + small, large, ratio + ); + + // The seed before the migration measured 14.1-14.2x here; linear measures ~4x. + assert!( + ratio < 8.0, + "string-literal escape cost is superlinear: 1k units {:?}, 4k units {:?} (4x input, \ + {:.1}x time). A ratio near 16 means the scan is quadratic again -- most likely \ + something reintroduced raw-String indexing (char_at / string_length) into the \ + escape path instead of walking the pre-decoded code points.", + small, + large, + ratio + ); +}