From 7281901562a0ff5a738bc8b87a53595057db01cd Mon Sep 17 00:00:00 2001 From: gunbc-ci-auto-heal Date: Fri, 2 Oct 2026 22:20:39 +0000 Subject: [PATCH 1/9] body_lowering_fold: an if's arms are read at their slots of the if_expr row MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit body_lower_if_else_capture_optional searched the if's capture and, at any left element that was not `else`, recursed LEFT before right. For `if b { if c { X } else { Y } } else { Z }` it entered the then-arm and answered the nested if's `else { Y }` as the outer else: Z was dropped, Y lowered twice, and normalize ACCEPTED the module (silent wrongness, DESIGN §5). It was loud only when Y held a function value, whose binder occurrence then appeared twice: dag/std/materialization_ladder.dag value_materialization refused as normalize_reason_minted_occurrence_duplicated, one of the native seven's file refusals. body_lower_if_then_arm_optional searched the same way. Both arm readers are now projections of body_lower_if_row_arms_optional, which reads the two if_expr grammar rows by position (block: if cond fn_body else (if_expr | fn_body); then-form: if cond then expr else expr) and answers Absent for any other capture. The dead searching helper body_lower_if_fn_body_capture_optional is deleted. The arm contract is unchanged: an `else if` is answered as its if_expr shell, any other arm as its production's captured child. Claims: v2.test.claim.body_lowering.if_arm_position (an atom count over the normalized tree, plus the function-value forms). Row: the capture side is filed on gunbc.recurring_failure_mode else_arm_lowered_as_an_if_nested_inside_it. Co-Authored-By: Claude Opus 5.5 (1M context) --- ..._arm_lowered_as_an_if_nested_inside_it.dag | 7 +- src/v2/compiler/body_lowering_fold.dag | 207 ++++++++---------- .../body_lowering/if_arm_position_test.dag | 111 ++++++++++ 3 files changed, 203 insertions(+), 122 deletions(-) create mode 100644 src/v2/test/claim/body_lowering/if_arm_position_test.dag diff --git a/dag/gunbc/recurring_failure_mode/else_arm_lowered_as_an_if_nested_inside_it.dag b/dag/gunbc/recurring_failure_mode/else_arm_lowered_as_an_if_nested_inside_it.dag index e033a4d37a8..9876174b963 100644 --- a/dag/gunbc/recurring_failure_mode/else_arm_lowered_as_an_if_nested_inside_it.dag +++ b/dag/gunbc/recurring_failure_mode/else_arm_lowered_as_an_if_nested_inside_it.dag @@ -10,7 +10,8 @@ data else_arm_lowered_as_an_if_nested_inside_it: RecurringFailureMode = Recurrin receipts: [ "INVALID STATE: an `else` arm lowered to an `if` nested INSIDE it, as though the arm were `else if`. `v2.compiler.body_lowering_fold` `body_lower_if_else_capture_optional` answered what follows `else` with the first if_expr found by a DEEP search of the else part (`body_lower_find_production_shell_optional`), so `if u == 0 { .. } else { match k { A => x B { v: w } => if c { w } else { u } } }` lowered its else arm to the nested `if c {..} else {..}` alone: the match's scrutinee, its other arms and its patterns were dropped with no refusal.", "SPECIMEN, MEASURED ON MAIN 5bffaa1e82f (2026-09-30), PRE-EXISTING: `dag/gunbc/machine_intake/boot_image_export_ownership.dag` `export_ownership_standing` -- seven atoms absent (`observed_key`, `Absent`, `ExportOwnershipUnrecognizedAccount`, `observed_name`, `Present`, `value`, `k`). The module was refused on main for an unrelated guard; the structured-tail change (gunbc#12713) admitted it, and that PR's reference-conservation census surfaced the drop. A minimal fixture drops identically under main's fold, so the defect is independent of that change.", - "RUNG FOUND AT: below the floor (silent wrongness, and a correct program refused when a pattern binder the dropped arm bound is read in the kept if). RUNG NOW: structurally guaranteed at this link -- `body_lower_else_part_if_expr_optional` reads the else part's FIRST production shell without descending into any production (`body_lower_first_production_shell_optional`), which the grammar makes either an if_expr (an `else if`) or a fn_body, so an if nested inside the else body cannot be taken for the else part. CEILING: the same.", + "SECOND SITE, THE CAPTURE SIDE, MEASURED ON MAIN 1a013596270 (2026-10-02): the else-PART fix above left the reader that finds the else at all searching. `body_lower_if_else_capture_optional` walked the if's capture and, at any left element that was not the `else` token, recursed LEFT before right -- so for `if b { if c { X } else { Y } } else { Z }` it entered the then-arm and answered the NESTED if's `else { Y }` as the OUTER if's else. Z was dropped, Y lowered twice, and normalize ACCEPTED the module (an atom count over the normalized tree: Z 1, Y 3, against 2 and 2). It was loud only when Y held a function value, whose binder occurrence then appeared twice: `dag/std/materialization_ladder.dag` `value_materialization` refused as normalize_reason_minted_occurrence_duplicated, one of the native seven's file refusals. `body_lower_if_then_arm_optional` searched the same way (the first fn_body anywhere in the capture). Both arms are now projections of `body_lower_if_row_arms_optional`, which reads the two if_expr grammar rows by POSITION and answers Absent for any other capture, so a nested if cannot be taken for either arm.", + "RUNG FOUND AT: below the floor (silent wrongness, and a correct program refused when a pattern binder the dropped arm bound is read in the kept if). RUNG NOW: structurally guaranteed at this link -- `body_lower_else_part_if_expr_optional` reads the else part's FIRST production shell without descending into any production (`body_lower_first_production_shell_optional`), which the grammar makes either an if_expr (an `else if`) or a fn_body, so an if nested inside the else body cannot be taken for the else part. Both sides now hold at the same rung: the else part by its first production shell, and the arms by their slots in the if_expr row. CEILING: the same.", ], evidence: [ @@ -19,5 +20,9 @@ data else_arm_lowered_as_an_if_nested_inside_it: RecurringFailureMode = Recurrin DeclarationRef { module_path: "v2.test.claim.namespace_xl0.else_arm_nested_if", decl_name: "the_if_nested_in_the_match_arm_is_still_read", field: WholeDeclaration }, DeclarationRef { module_path: "v2.test.claim.namespace_xl0.else_arm_nested_if", decl_name: "a_real_else_if_chain_still_reads_its_nested_else", field: WholeDeclaration }, DeclarationRef { module_path: "v2.test.claim.namespace_xl0.else_arm_nested_if", decl_name: "the_declared_shape_resolves", field: WholeDeclaration }, + DeclarationRef { module_path: "v2.compiler.body_lowering_fold", decl_name: "body_lower_if_row_arms_optional", field: WholeDeclaration }, + DeclarationRef { module_path: "v2.test.claim.body_lowering.if_arm_position", decl_name: "an_if_nested_in_the_then_arm_keeps_the_outer_else_holds", field: WholeDeclaration }, + DeclarationRef { module_path: "v2.test.claim.body_lowering.if_arm_position", decl_name: "a_function_value_in_a_nested_else_lowers_once_holds", field: WholeDeclaration }, + DeclarationRef { module_path: "v2.test.claim.body_lowering.if_arm_position", decl_name: "an_else_if_chain_keeps_every_arm_control_holds", field: WholeDeclaration }, ], } diff --git a/src/v2/compiler/body_lowering_fold.dag b/src/v2/compiler/body_lowering_fold.dag index 1eb52cb08d6..acb37b7f224 100644 --- a/src/v2/compiler/body_lowering_fold.dag +++ b/src/v2/compiler/body_lowering_fold.dag @@ -7049,50 +7049,91 @@ fn body_lower_if_expr_capture_optional(spine: Node) -> Optional { } } -fn body_lower_if_fn_body_capture_optional(spine: Node) -> Optional { - match body_lower_find_captured(root: spine, emitted: ^dag_surface_fn_body) { - ParseSubtreeFound { captured: body } => Present { value: body } - ParseSubtreeAbsent => - if is_empty_conj_root(n: spine) { - Absent - } else { - match body_lower_match_spine_left(spine: spine) { - Present { value: left } => - match body_lower_match_spine_right(spine: spine) { - Present { value: right } => - match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { - Present { value: id } => - match id == ^dag_token_kw_else { - true => body_lower_if_fn_body_capture_optional(spine: right) - false => - match body_lower_if_fn_body_capture_optional(spine: left) { - Present { value: found } => Present { value: found } - Absent => body_lower_if_fn_body_capture_optional(spine: right) - } - } - Absent => - match body_lower_if_fn_body_capture_optional(spine: left) { - Present { value: found } => Present { value: found } - Absent => body_lower_if_fn_body_capture_optional(spine: right) - } - } - Absent => Absent - } - Absent => Absent - } +// THE if_expr ROW IS READ BY POSITION, NEVER SEARCHED. v2.extdeps.languages.dag declares exactly +// two rows: dag_grammar_if_block_expr seq(`if`, seq(cond, seq(fn_body, seq(`else`, if_expr | fn_body)))) +// and dag_grammar_if_then_expr seq(`if`, seq(cond, seq(`then`, seq(expr, seq(`else`, expr))))). Each +// arm is the element at its own slot. The readers this replaces SEARCHED the capture -- the else +// reader recursed into the LEFT subtree before the right whenever the token there was not `else`, so +// for `if b { if c { X } else { Y } } else { Z }` it entered the then-arm and answered the NESTED if's +// else `{ Y }` as the OUTER if's else: Z was dropped, Y lowered twice, and the module was ACCEPTED +// (gunbc.recurring_failure_mode else_arm_lowered_as_an_if_nested_inside_it, capture side). A capture +// that is neither row is Absent, and every caller refuses it as navigation, never guesses. +type IfRowArms { + then_arm: Node, + else_part: Node, +} + +fn body_lower_if_row_arms_optional(captured: Node) -> Optional { + let spine = body_lower_unwrap_surface_shell(node: captured) + match body_lower_if_row_slot_after_token(seq: spine, token: ^dag_token_kw_if) { + Absent => optional_absent() + Present { value: after_if } => + match body_lower_match_spine_right(spine: after_if) { + Absent => optional_absent() + Present { value: after_cond } => + match body_lower_if_row_slot_after_token(seq: after_cond, token: ^dag_token_kw_then) { + Present { value: after_then } => + match body_lower_match_spine_left(spine: after_then) { + Absent => optional_absent() + Present { value: then_expr } => + match body_lower_match_spine_right(spine: after_then) { + Absent => optional_absent() + Present { value: else_seq } => + body_lower_if_row_with_else(then_arm: body_lower_if_row_captured(node: then_expr), else_seq: else_seq) + } + } + Absent => + match body_lower_match_spine_left(spine: after_cond) { + Absent => optional_absent() + Present { value: then_body } => + match body_lower_match_spine_right(spine: after_cond) { + Absent => optional_absent() + Present { value: else_seq } => + body_lower_if_row_with_else(then_arm: body_lower_if_row_captured(node: then_body), else_seq: else_seq) + } + } + } } } } -fn body_lower_if_then_fn_body_optional(captured: Node) -> Optional { - let spine = body_lower_unwrap_surface_shell(node: captured) - match body_lower_if_fn_body_capture_optional(spine: spine) { - Present { value: first } => - match body_lower_find_captured(root: captured, emitted: ^dag_surface_fn_body) { - ParseSubtreeFound { captured: only } => Present { value: only } - ParseSubtreeAbsent => Present { value: first } +// seq(`else`, else_part): the else part, read as the arm readers have always answered it. +fn body_lower_if_row_with_else(then_arm: Node, else_seq: Node) -> Optional { + match body_lower_if_row_slot_after_token(seq: else_seq, token: ^dag_token_kw_else) { + Absent => optional_absent() + Present { value: else_part } => + optional_present(value: IfRowArms { then_arm: then_arm, else_part: body_lower_if_row_else_arm(else_part: else_part) }) + } +} + +// seq(, rest) -> rest, when the left element is exactly that keyword token. +fn body_lower_if_row_slot_after_token(seq: Node, token: Symbol) -> Optional { + match body_lower_match_spine_left(spine: seq) { + Absent => optional_absent() + Present { value: left } => + match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { + Present { value: id } => + if id == token { body_lower_match_spine_right(spine: seq) } else { optional_absent() } + Absent => optional_absent() } - Absent => Absent + } +} + +// An arm slot holds one production (fn_body or expr); the arm readers answer its captured child. +fn body_lower_if_row_captured(node: Node) -> Node { + let shell = body_lower_deep_unwrap_optional(node: node) + match parse_production_captured_child_optional(node: shell) { + Present { value: captured } => captured + Absent => shell + } +} + +// An `else if` arm is answered as its if_expr SHELL, so body_lower_if_arm_lowered sees an if and +// lowers it whole; any other else part is its production's captured child. +fn body_lower_if_row_else_arm(else_part: Node) -> Node { + match body_lower_else_part_if_expr_optional(else_part: else_part) { + Present { value: nested } => nested + Absent => body_lower_if_row_captured(node: else_part) } } @@ -7127,57 +7168,10 @@ fn body_lower_first_production_shell_optional(root: Node) -> Optional { } } -// AN `else if` ARM IS ANSWERED AS ITS if_expr SHELL, so body_lower_if_arm_lowered can see it is an -// if and lower it whole. Answering the shell's captured content instead hid the if_expr identity, and -// the arm's statement-sequence search found the nested then-arm: `if a { b } else if c { d } else { u }` -// lowered to Branch(a, b, d), and both the nested condition and its else reached no reference site. fn body_lower_if_else_capture_optional(captured: Node) -> Optional { - let spine = body_lower_unwrap_surface_shell(node: captured) - match is_empty_conj_root(n: spine) { - true => Absent - false => - match body_lower_match_spine_left(spine: spine) { - Present { value: left } => - match body_lower_match_spine_right(spine: spine) { - Present { value: right } => - match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { - Present { value: id } => - match id == ^dag_token_kw_else { - true => - match body_lower_else_part_if_expr_optional(else_part: right) { - Present { value: nested } => Present { value: nested } - Absent => - match body_lower_if_fn_body_capture_optional(spine: right) { - Present { value: body } => Present { value: body } - Absent => - match body_lower_if_expr_capture_optional(spine: right) { - Present { value: expr } => Present { value: expr } - Absent => Absent - } - } - } - false => - match body_lower_if_else_capture_optional(captured: left) { - Present { value: found } => Present { value: found } - Absent => body_lower_if_else_capture_optional(captured: right) - } - } - Absent => - match body_lower_if_else_capture_optional(captured: left) { - Present { value: found } => Present { value: found } - Absent => body_lower_if_else_capture_optional(captured: right) - } - } - Absent => Absent - } - Absent => - fold(spine.children, init: Absent, f: fn(acc, e) { - match acc { - Present { value: _ } => acc - Absent => body_lower_if_else_capture_optional(captured: e.target) - } - }) - } + match body_lower_if_row_arms_optional(captured: captured) { + Present { value: arms } => optional_present(value: arms.else_part) + Absent => optional_absent() } } @@ -7218,38 +7212,9 @@ fn body_lower_if_condition_optional(captured: Node) -> Optional { } fn body_lower_if_then_arm_optional(captured: Node) -> Optional { - let spine = body_lower_unwrap_surface_shell(node: captured) - match body_lower_find_captured(root: spine, emitted: ^dag_surface_fn_body) { - ParseSubtreeFound { captured: body } => Present { value: body } - ParseSubtreeAbsent => - match body_lower_if_then_fn_body_optional(captured: captured) { - Present { value: body } => Present { value: body } - Absent => - match body_lower_match_spine_left(spine: spine) { - Present { value: left } => - match body_lower_match_spine_right(spine: spine) { - Present { value: right } => - match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { - Present { value: id } => - match id == ^dag_token_kw_then { - true => body_lower_if_expr_capture_optional(spine: right) - false => - match body_lower_if_then_arm_optional(captured: left) { - Present { value: found } => Present { value: found } - Absent => body_lower_if_then_arm_optional(captured: right) - } - } - Absent => - match body_lower_if_then_arm_optional(captured: left) { - Present { value: found } => Present { value: found } - Absent => body_lower_if_then_arm_optional(captured: right) - } - } - Absent => Absent - } - Absent => Absent - } - } + match body_lower_if_row_arms_optional(captured: captured) { + Present { value: arms } => optional_present(value: arms.then_arm) + Absent => optional_absent() } } diff --git a/src/v2/test/claim/body_lowering/if_arm_position_test.dag b/src/v2/test/claim/body_lowering/if_arm_position_test.dag new file mode 100644 index 00000000000..f71c38cf274 --- /dev/null +++ b/src/v2/test/claim/body_lowering/if_arm_position_test.dag @@ -0,0 +1,111 @@ +module v2.test.claim.body_lowering.if_arm_position + +import v2.compiler.normalize { normalize } +import v2.compiler.parse { parse_module } +import v2.compiler.tokenize { tokenize } +import v2.extdeps.languages.dag { dag_language_model } +import v2.std.algebra { fold_list } +import v2.std.diagnostic { Accepted, Rejected, diagnostics_fatal_reason } +import v2.std.grammar { node_atom_identity_optional } +import v2.std.integer { Int } +import v2.std.language_model { LanguageModel } +import v2.std.live_tree { LiveTreeDisposition, SubstrateInputsOnly } +import v2.std.logic { Bool } +import v2.std.node { Node, node_subtree_nodes, symbol_lexeme } +import v2.std.optional { Absent, Optional, Present } +import v2.std.text { String } + +data live_tree_disposition: LiveTreeDisposition = SubstrateInputsOnly + +// AN if's ARMS ARE READ AT THEIR OWN SLOTS OF THE if_expr ROW. The subject is +// v2.compiler.body_lowering_fold body_lower_if_row_arms_optional, the one reader both arm readers +// project from. The else reader it replaced searched the if's capture and recursed into the left +// subtree first, so for an if nested in the THEN arm it answered the nested if's else as the outer +// if's else: the outer else was dropped, the nested else lowered twice, and normalize ACCEPTED the +// module (gunbc.recurring_failure_mode else_arm_lowered_as_an_if_nested_inside_it, capture side). +// +// THE MEASURE IS AN ATOM COUNT OVER THE NORMALIZED TREE, because the defect is silent: each marker is a +// parameter (one atom in the domain) plus one atom per use, so a marker used once counts 2 when its arm +// was kept exactly once. THE REDS under the searching reader: the outer else's marker counts 1 (dropped) +// and the nested else's counts 3 (lowered twice); and a function value in that nested else refused as +// normalize_reason_minted_occurrence_duplicated, the only form in which the defect was ever loud. +// CONTROLS, TRUE UNDER BOTH: a flat if and a real else-if chain keep every arm once. +// +// The stages are the native route's own, tokenize -> parse -> normalize, and nothing above them. +data iap_cached_lm: LanguageModel = dag_language_model() + +fn iap_normalized_root(src: String) -> Optional { + let lm = iap_cached_lm + match tokenize(text: src, file: ^if_arm_position_subject, rules: lm.lex) { + Rejected { diagnostics: _ } => Absent + Accepted { value: ts, diagnostics: _ } => + match parse_module(tokens: ts, grammar: lm.grammar) { + Rejected { diagnostics: _ } => Absent + Accepted { value: a, diagnostics: _ } => + match normalize(parse_tree: a.tree) { + Rejected { diagnostics: _ } => Absent + Accepted { value: tree, diagnostics: _ } => Present { value: tree.root } + } + } + } +} + +fn iap_normalize_outcome(src: String) -> String { + let lm = iap_cached_lm + match tokenize(text: src, file: ^if_arm_position_subject, rules: lm.lex) { + Rejected { diagnostics: d } => symbol_lexeme(sym: diagnostics_fatal_reason(d: d)) + Accepted { value: ts, diagnostics: _ } => + match parse_module(tokens: ts, grammar: lm.grammar) { + Rejected { diagnostics: d } => symbol_lexeme(sym: diagnostics_fatal_reason(d: d)) + Accepted { value: a, diagnostics: _ } => + match normalize(parse_tree: a.tree) { + Rejected { diagnostics: d } => symbol_lexeme(sym: diagnostics_fatal_reason(d: d)) + Accepted { value: _, diagnostics: _ } => "ACCEPTED" + } + } + } +} + +fn iap_count(root: Node, name: String) -> Int { + fold_list(xs: node_subtree_nodes(root: root), empty: 0, cons: fn(acc, n) { + match node_atom_identity_optional(node: n) { + Present { value: id } => if symbol_lexeme(sym: id) == name { acc + 1 } else { acc } + Absent => acc + } + }) +} + +// Every marker kept exactly once: its parameter atom plus its one use. +fn iap_each_arm_kept_once(src: String) -> Bool { + match iap_normalized_root(src: src) { + Absent => false + Present { value: root } => + (iap_count(root: root, name: "zzz_then") == 2) + && (iap_count(root: root, name: "zzz_inner_else") == 2) + && (iap_count(root: root, name: "zzz_outer_else") == 2) + } +} + +data src_nested_in_then: String = "module m.t\nfn t(b: Bool, c: Bool, zzz_then: Int, zzz_inner_else: Int, zzz_outer_else: Int) -> Int {\n if b {\n if c {\n zzz_then\n } else {\n zzz_inner_else\n }\n } else {\n zzz_outer_else\n }\n}\n" + +data src_flat_and_else_if: String = "module m.t\nfn t(b: Bool, c: Bool, zzz_then: Int, zzz_inner_else: Int, zzz_outer_else: Int) -> Int {\n if b {\n zzz_then\n } else if c {\n zzz_inner_else\n } else {\n zzz_outer_else\n }\n}\n" + +data src_lambda_in_nested_else: String = "module m.t\nfn g(p: Int) -> Bool {\n p > 3\n}\nfn h(w: fn(Int) -> Bool) -> Int {\n 0\n}\nfn t(b: Bool, c: Bool) -> Int {\n if b {\n if c {\n 2\n } else {\n h(w: p => g(p: p))\n }\n } else {\n 0\n }\n}\n" + +data src_fn_literal_in_nested_else: String = "module m.t\nfn g(p: Int) -> Bool {\n p > 3\n}\nfn h(w: fn(Int) -> Bool) -> Int {\n 0\n}\nfn t(b: Bool, c: Bool) -> Int {\n if b {\n if c {\n 2\n } else {\n h(w: fn(p) { g(p: p) })\n }\n } else {\n 0\n }\n}\n" + +test fn an_if_nested_in_the_then_arm_keeps_the_outer_else_holds() -> Bool { + iap_each_arm_kept_once(src: src_nested_in_then) +} + +test fn a_function_value_in_a_nested_else_lowers_once_holds() -> Bool { + iap_normalize_outcome(src: src_lambda_in_nested_else) == "ACCEPTED" +} + +test fn a_fn_literal_in_a_nested_else_lowers_once_holds() -> Bool { + iap_normalize_outcome(src: src_fn_literal_in_nested_else) == "ACCEPTED" +} + +test fn an_else_if_chain_keeps_every_arm_control_holds() -> Bool { + iap_each_arm_kept_once(src: src_flat_and_else_if) +} From 8351814bc0d11ce5ae710e0c6862dcc1cc396939 Mon Sep 17 00:00:00 2001 From: gunbc-ci-auto-heal Date: Fri, 2 Oct 2026 23:28:58 +0000 Subject: [PATCH 2/9] if_arm_position: one warm producer for the four claims The first floor run (37073686904) refused all four claims over the new-witness enrolment margin (353k-687k eval steps against 72300): each claim frame paid its own tokenize -> parse -> normalize. iap_verdicts computes the four verdicts once into four Bools (portable: no Node, no closure), each claim reads its field, and the producer is enrolled WARM in v2.workflow.floor_pure_producer_share beside eam_outcomes, the same ground. Co-Authored-By: Claude Opus 5.5 (1M context) --- .../body_lowering/if_arm_position_test.dag | 27 ++++++++++++++++--- src/v2/workflow/floor_pure_producer_share.dag | 4 +++ 2 files changed, 27 insertions(+), 4 deletions(-) diff --git a/src/v2/test/claim/body_lowering/if_arm_position_test.dag b/src/v2/test/claim/body_lowering/if_arm_position_test.dag index f71c38cf274..923f73cd13c 100644 --- a/src/v2/test/claim/body_lowering/if_arm_position_test.dag +++ b/src/v2/test/claim/body_lowering/if_arm_position_test.dag @@ -94,18 +94,37 @@ data src_lambda_in_nested_else: String = "module m.t\nfn g(p: Int) -> Bool {\n data src_fn_literal_in_nested_else: String = "module m.t\nfn g(p: Int) -> Bool {\n p > 3\n}\nfn h(w: fn(Int) -> Bool) -> Int {\n 0\n}\nfn t(b: Bool, c: Bool) -> Int {\n if b {\n if c {\n 2\n } else {\n h(w: fn(p) { g(p: p) })\n }\n } else {\n 0\n }\n}\n" +// ONE PRODUCER FOR EVERY CLAIM, enrolled WARM in v2.workflow.floor_pure_producer_share. Each verdict +// pays a full tokenize -> parse -> normalize of its own source, and the four claims would otherwise +// pay all four per claim frame; the stored value is four Bools (no Node, no closure), so it is portable. +type IapVerdicts { + nested_in_then_keeps_outer_else: Bool, + function_value_in_nested_else_lowers: Bool, + fn_literal_in_nested_else_lowers: Bool, + else_if_chain_keeps_every_arm: Bool, +} + +fn iap_verdicts() -> IapVerdicts { + IapVerdicts { + nested_in_then_keeps_outer_else: iap_each_arm_kept_once(src: src_nested_in_then), + function_value_in_nested_else_lowers: iap_normalize_outcome(src: src_lambda_in_nested_else) == "ACCEPTED", + fn_literal_in_nested_else_lowers: iap_normalize_outcome(src: src_fn_literal_in_nested_else) == "ACCEPTED", + else_if_chain_keeps_every_arm: iap_each_arm_kept_once(src: src_flat_and_else_if), + } +} + test fn an_if_nested_in_the_then_arm_keeps_the_outer_else_holds() -> Bool { - iap_each_arm_kept_once(src: src_nested_in_then) + iap_verdicts().nested_in_then_keeps_outer_else } test fn a_function_value_in_a_nested_else_lowers_once_holds() -> Bool { - iap_normalize_outcome(src: src_lambda_in_nested_else) == "ACCEPTED" + iap_verdicts().function_value_in_nested_else_lowers } test fn a_fn_literal_in_a_nested_else_lowers_once_holds() -> Bool { - iap_normalize_outcome(src: src_fn_literal_in_nested_else) == "ACCEPTED" + iap_verdicts().fn_literal_in_nested_else_lowers } test fn an_else_if_chain_keeps_every_arm_control_holds() -> Bool { - iap_each_arm_kept_once(src: src_flat_and_else_if) + iap_verdicts().else_if_chain_keeps_every_arm } diff --git a/src/v2/workflow/floor_pure_producer_share.dag b/src/v2/workflow/floor_pure_producer_share.dag index 6c9461b1448..72e16e87760 100644 --- a/src/v2/workflow/floor_pure_producer_share.dag +++ b/src/v2/workflow/floor_pure_producer_share.dag @@ -651,6 +651,9 @@ import v2.std.collection { List } // five claims over the enrolment margin at ~590ms each, every one of them paying the same ingest. // THE XL-2 ELSE-ARM ROW IS THE SAME GROUND AT FIVE SPECIMENS: v2.test.claim.namespace_xl0.else_arm_nested_if // eam_outcomes drives one front end over five inline modules and stores one verdict arm per module. +// THE IF-ARM-POSITION ROW IS THE SAME GROUND AT FOUR SPECIMENS: v2.test.claim.body_lowering.if_arm_position +// iap_verdicts tokenizes, parses and normalizes four inline modules and stores four Bools. Its first floor +// run (37073686904) refused all four claims over the enrolment margin, at 353k-687k eval steps each. // THE XL-2 REIFY-OPERAND ROW: v2.test.claim.body_lowering.reify_operand_refusal ror_verdicts tokenizes // and parses two inline modules and reifies two supplied bodies over each; the stored value is four // verdicts (no Node, no closure), portable for the same reason cav_outcomes is. @@ -827,6 +830,7 @@ data floor_cross_claim_pure_producers_warm: List = [ "v2.test.claim.body_lowering.match_position_structure.mps_normalized", "v2.test.claim.namespace_xl0.wildcard_arm_resolve.wc_outcomes", "v2.test.claim.namespace_xl0.else_arm_nested_if.eam_outcomes", + "v2.test.claim.body_lowering.if_arm_position.iap_verdicts", "v2.test.cli.v2_native_cli.cli_probe_trailing_outcome", "v2.test.emit.closure_emit_arrow_body_refusal.arrow_body_probe_verdict", "v2.test.emit.closure_emit_arrow_body_refusal.arrow_body_mixed_probe_verdict", From 4a1a9bd45f2c51bb38cfe27b68c61e9a59ce00cd Mon Sep 17 00:00:00 2001 From: gunbc-ci-auto-heal Date: Sat, 3 Oct 2026 00:12:03 +0000 Subject: [PATCH 3/9] nested_if_else_eval: the evaluated form of the nested-if else arm `if false { if true {1} else {2} } else {3}` is 3. The searching else reader took the nested `else { 2 }` as the outer else, so a body lowered through it answers 2. The mirror (`if true { if false {1} else {2} } else {3}` is 2) holds under either reader and is the control. The module is import-free: the floor evaluates it through the seed, and the native lane through the emitted compiler's own body lowering, where it is the discriminating red. It is not yet merge-blocking evidence for the v2 reader, since the native lane is not a required lane. Co-Authored-By: Claude Opus 5.5 (1M context) --- .../nested_if_else_eval_test.dag | 43 +++++++++++++++++++ 1 file changed, 43 insertions(+) create mode 100644 src/v2/test/claim/body_lowering/nested_if_else_eval_test.dag diff --git a/src/v2/test/claim/body_lowering/nested_if_else_eval_test.dag b/src/v2/test/claim/body_lowering/nested_if_else_eval_test.dag new file mode 100644 index 00000000000..5f23e05fa23 --- /dev/null +++ b/src/v2/test/claim/body_lowering/nested_if_else_eval_test.dag @@ -0,0 +1,43 @@ +module v2.test.claim.body_lowering.nested_if_else_eval + +// THE EVALUATED FORM OF THE NESTED-IF ELSE ARM (gunbc.recurring_failure_mode +// else_arm_lowered_as_an_if_nested_inside_it, capture side). The outer condition is false, so the +// program's value is the OUTER else, 3. The searching else reader took the nested if's `else { 2 }` as +// the outer if's else, so a body lowered through it answers 2: the dropped arm changes the evaluated +// result, not only the atom count v2.test.claim.body_lowering.if_arm_position measures. Import-free on +// purpose, so the native route prepares it cheaply: the floor evaluates it through the seed, and the +// native lane evaluates it through the emitted compiler's own body lowering, where it is the +// discriminating red for the arm readers. +fn nested_if_else_value() -> Int { + if false { + if true { + 1 + } else { + 2 + } + } else { + 3 + } +} + +// THE MIRROR: the outer condition is true, so the value is the nested if's own else, 2. It holds under +// either reader, so it is the control that the subject's red is the dropped OUTER else and nothing else. +fn nested_if_then_value() -> Int { + if true { + if false { + 1 + } else { + 2 + } + } else { + 3 + } +} + +test fn the_outer_else_of_an_if_nested_in_the_then_arm_is_the_value_holds() -> Bool { + nested_if_else_value() == 3 +} + +test fn the_nested_else_inside_the_then_arm_is_the_value_control_holds() -> Bool { + nested_if_then_value() == 2 +} From bf07eadafb7ba204d854671300988d937aa54365 Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Sat, 3 Oct 2026 00:22:54 +0000 Subject: [PATCH 4/9] v2 CLI: if-arm-reader-differential verb -- census of the pre-#13029 arm readers against the positional reader Co-Authored-By: Claude Opus 5.5 (1M context) --- src/v2/cli/compile_cli.dag | 86 ++- src/v2/cli/if_arm_reader_differential.dag | 526 ++++++++++++++++++ .../cli/if_arm_reader_differential_test.dag | 111 ++++ src/v2/test/claim/cli/v2_native_cli_test.dag | 8 + .../occurrence_file_attribution_test.dag | 2 + src/v2/workflow/floor_pure_producer_share.dag | 4 + 6 files changed, 736 insertions(+), 1 deletion(-) create mode 100644 src/v2/cli/if_arm_reader_differential.dag create mode 100644 src/v2/test/claim/cli/if_arm_reader_differential_test.dag diff --git a/src/v2/cli/compile_cli.dag b/src/v2/cli/compile_cli.dag index 8e48e1096b8..c7fe6f1967c 100644 --- a/src/v2/cli/compile_cli.dag +++ b/src/v2/cli/compile_cli.dag @@ -12,6 +12,13 @@ import v2.compiler.self_host.closure_emission { emit_closure_from_ingest_located } import v2.compiler.source_authority { SourceRootIngest, source_root_ingest_from_files } +import v2.cli.if_arm_reader_differential { + IardCensus, + IardCensusTaken, + IardGrammarUnprepared, + if_arm_reader_differential, + if_arm_reader_differential_text +} import gunbc.source_root_read { source_root_read, SourceRootFilesRead, SourceRootFilesRefused } import v2.extdeps.languages.c { c_target_model } import v2.extdeps.languages.rust { rust_target_model } @@ -135,6 +142,7 @@ fn cli_emit_target_roster_text() -> String { type CliVerb = CliEmitClosure { entry: String, source_roots: FreeMonoid, target: CliEmitTarget } + | CliIfArmReaderDifferential { source_roots: FreeMonoid } // A PLAN IS PARSED OR IT IS REFUSED, and the refusal carries its own reason symbol rather than a // rendered sentence, so the host renders it once and a witness can discriminate on the CAUSE. A @@ -156,6 +164,8 @@ type CliParseState | CliParseSeekRootValue { entry: String, roots: FreeMonoid, target: CliEmitTarget } | CliParseSeekEntryValue { entry: String, roots: FreeMonoid, target: CliEmitTarget } | CliParseSeekTargetValue { entry: String, roots: FreeMonoid } + | CliParseSeekDifferentialOption { roots: FreeMonoid } + | CliParseSeekDifferentialRootValue { roots: FreeMonoid } | CliParseFailed { reason: Symbol, detail: String } data cli_option_source_root: String = "--source-root" @@ -166,7 +176,12 @@ data cli_option_target: String = "--target" data cli_verb_emit: String = "emit" -data cli_usage: String = "usage: emit --entry --source-root [--source-root ]... [--target ]" +// THE NESTED-IF ARM-READER DIFFERENTIAL (v2.cli.if_arm_reader_differential). Named for exactly what +// it measures so it does not read as the general census, which lives on the native driver +// (v2.compiler.compile). It takes source roots only, and is deleted with its module. +data cli_verb_if_arm_reader_differential: String = "if-arm-reader-differential" + +data cli_usage: String = "usage: emit --entry --source-root [--source-root ]... [--target ] | if-arm-reader-differential --source-root [--source-root ]..." fn cli_parse_step(state: CliParseState, arg: String) -> CliParseState { match state { @@ -175,12 +190,32 @@ fn cli_parse_step(state: CliParseState, arg: String) -> CliParseState { CliParseSeekVerb => if arg == cli_verb_emit { CliParseSeekOption { entry: "", roots: Empty, target: cli_default_emit_target } + } else if arg == cli_verb_if_arm_reader_differential { + CliParseSeekDifferentialOption { roots: Empty } } else { CliParseFailed { reason: ^cli_unknown_verb, detail: concat(concat("unknown verb ", arg), concat(" -- ", cli_usage)) } } + CliParseSeekDifferentialOption { roots: roots } => + if arg == cli_option_source_root { + CliParseSeekDifferentialRootValue { roots: roots } + } else { + CliParseFailed { + reason: ^cli_unknown_option, + detail: concat(concat("unknown option ", arg), concat(" -- ", cli_usage)) + } + } + CliParseSeekDifferentialRootValue { roots: roots } => + if cli_arg_is_option(arg: arg) { + CliParseFailed { + reason: ^cli_source_root_missing_value, + detail: concat(cli_option_source_root, " takes a directory, and an option followed it") + } + } else { + CliParseSeekDifferentialOption { roots: Cons { head: arg, tail: roots } } + } CliParseSeekRootValue { entry: entry, roots: roots, target: target } => if cli_arg_is_option(arg: arg) { CliParseFailed { @@ -263,6 +298,20 @@ fn cli_parse_finish(state: CliParseState) -> CliPlan { reason: ^cli_target_missing_value, detail: concat(cli_option_target, " takes a target name, and none followed it") } + CliParseSeekDifferentialRootValue { roots: _ } => + CliPlanRefused { + reason: ^cli_source_root_missing_value, + detail: concat(cli_option_source_root, " takes a directory, and none followed it") + } + CliParseSeekDifferentialOption { roots: roots } => + if length(xs: roots) == 0 { + CliPlanRefused { + reason: ^cli_no_source_root, + detail: concat("if-arm-reader-differential needs at least one source root -- ", cli_usage) + } + } else { + CliPlanParsed { verb: CliIfArmReaderDifferential { source_roots: list_reverse(xs: roots) } } + } CliParseSeekOption { entry: entry, roots: roots, target: target } => if length(xs: roots) == 0 { CliPlanRefused { @@ -300,6 +349,7 @@ fn v2_cli_plan_source_roots(plan: CliPlan) -> FreeMonoid { CliPlanParsed { verb: verb } => match verb { CliEmitClosure { entry: _, source_roots: roots, target: _ } => roots + CliIfArmReaderDifferential { source_roots: roots } => roots } } } @@ -309,6 +359,7 @@ type CliRunOutcome | CliRunRefused { reason: Symbol, detail: String } | CliCensusRefused { reason: Symbol, detail: String, text: String } | CliUsageRefused { reason: Symbol, detail: String } + | CliDifferentialReported { text: String, observed: Bool, detail: String } // THE REFUSAL PRINTS ITS CAUSE, AND FOR ONE GENERATION IT DID NOT. The rejected arm already // CARRIED `d.head.reason` (now the FATAL reason -- the head was the first advisory that happened @@ -349,6 +400,15 @@ fn v2_cli_run(plan: CliPlan, ingest: SourceRootIngest) -> CliRunOutcome { CliUsageRefused { reason: reason, detail: detail } CliPlanParsed { verb: verb } => match verb { + CliIfArmReaderDifferential { source_roots: _ } => + if length(xs: ingest) == 0 { + CliRunRefused { + reason: ^cli_empty_ingest, + detail: "the source roots walked to zero .dag files; a census over nothing is not a census of zero" + } + } else { + cli_if_arm_reader_differential_outcome(census: if_arm_reader_differential(ingest: ingest)) + } CliEmitClosure { entry: entry, source_roots: _, target: target } => if length(xs: ingest) == 0 { CliRunRefused { @@ -393,6 +453,27 @@ fn v2_cli_run(plan: CliPlan, ingest: SourceRootIngest) -> CliRunOutcome { } } +// THE DIFFERENTIAL'S STATUS IS WHETHER THE POPULATION WAS WHOLLY OBSERVED, NEVER HOW MANY SITES +// DIFFER: the differing sites are the finding, not a failure. A file the census could not parse, or a +// grammar that did not prepare, leaves the population incomplete, and that fails the process so a +// partial count cannot be read as the whole one. +fn cli_if_arm_reader_differential_outcome(census: IardCensus) -> CliRunOutcome { + match census { + IardGrammarUnprepared { reason: _ } => + CliDifferentialReported { + text: if_arm_reader_differential_text(c: census), + observed: false, + detail: "the grammar did not prepare; no site was observed" + } + IardCensusTaken { tally: t } => + CliDifferentialReported { + text: if_arm_reader_differential_text(c: census), + observed: t.files_refused == 0, + detail: "some files did not parse; their sites are unobserved (see file_refused rows)" + } + } +} + // ONE RENDERING OF A LOCATED REFUSAL, for a refused closure and for a refused member alike: the whole // ordered chain, then the fatal's file and extent. A member's refusal fails the process as a closure // refusal does, so its status line owes the same located cause and not a summary sentence. @@ -437,6 +518,8 @@ fn v2_cli_exit(outcome: CliRunOutcome) -> ProcessExit { CliCensusRefused { reason: _, detail: detail, text: _ } => exit_failure(reason: detail) CliUsageRefused { reason: _, detail: detail } => ExitFailure { code: exit_code_misuse, reason: detail } + CliDifferentialReported { text: _, observed: observed, detail: detail } => + if observed { ExitSuccess } else { exit_failure(reason: detail) } } } @@ -446,6 +529,7 @@ fn v2_cli_outcome_text(outcome: CliRunOutcome) -> String { CliRunRefused { reason: _, detail: _ } => "" CliCensusRefused { reason: _, detail: _, text: text } => text CliUsageRefused { reason: _, detail: _ } => "" + CliDifferentialReported { text: text, observed: _, detail: _ } => text } } diff --git a/src/v2/cli/if_arm_reader_differential.dag b/src/v2/cli/if_arm_reader_differential.dag new file mode 100644 index 00000000000..313a7008c0e --- /dev/null +++ b/src/v2/cli/if_arm_reader_differential.dag @@ -0,0 +1,526 @@ +module v2.cli.if_arm_reader_differential + +import std.dissolution { DissolutionCondition, unbound_dissolution } +import std.algebra { Cons, Empty, FreeMonoid } +import std.occurrence_identity { occurrence_id_eq } +import v2.compiler.body_lowering_fold { + IfRowArms, + body_lower_deep_unwrap_optional, + body_lower_else_part_if_expr_optional, + body_lower_find_captured, + body_lower_if_expr_capture_optional, + body_lower_if_row_arms_optional, + body_lower_match_spine_left, + body_lower_match_spine_right, + body_lower_unwrap_surface_shell +} +import v2.compiler.parse { PreparedGrammar, parse_module_prepared } +import v2.compiler.program_assembly { dag_prepared_grammar } +import v2.compiler.source_authority { DagSourceReadWitness, SourceRootIngest } +import v2.compiler.tokenize { tokenize } +import v2.extdeps.languages.dag { + ParseSubtreeAbsent, + ParseSubtreeFound, + dag_lex, + dag_surface_kw_then_ident_from_captured, + parse_production_captured_child_optional, + parse_production_emitted_identity_optional +} +import v2.std.algebra { fold_list, length, list_reverse } +import v2.std.diagnostic { Accepted, Diagnostics, NonEmptyDiagnostics, Rejected, Textual, ByteRange, WholeFile, bind_outcome, diagnostics_fatal_reason } +import v2.std.grammar { node_atom_identity_optional } +import v2.std.integer { integer_int_to_decimal_string } +import v2.std.node { Node, Symbol, is_empty_conj_root, symbol_lexeme } +import v2.std.optional { Absent, Optional, Present, optional_absent, optional_present } +import v2.std.provenance { SpanIndex, node_occurrence_id_optional, span_index_textual_locus_optional } +import v2.std.text { String } + +// THE NESTED-IF ARM-READER DIFFERENTIAL: an OFFLINE ORACLE, NEVER A PRODUCTION ROUTE. +// +// Until gunbc#13029, v2.compiler.body_lowering_fold read an if's arms by SEARCHING its capture +// left-first. For an if nested in the THEN arm (or the condition) of another if, the else search +// entered the nested if and answered ITS else as the outer if's: the outer else was dropped, the +// nested else lowered twice, and normalize accepted (gunbc.recurring_failure_mode +// else_arm_lowered_as_an_if_nested_inside_it, capture side). #13029 replaced the searching readers +// with the positional body_lower_if_row_arms_optional. Two sites were loud; the silent population +// was unknown. This module measures it: for every if_expr capture in the ingest it asks both +// readers for each arm and reports every site where they answer a different occurrence. +// +// THE FROZEN READERS BELOW ARE main's body_lower_if_else_capture_optional, +// body_lower_if_fn_body_capture_optional, body_lower_if_then_arm_optional and +// body_lower_if_then_fn_body_optional as they stood before #13029, renamed and nothing else. Every +// helper they call is unchanged by #13029, so this is the old reader's logic exactly. They serve +// from history as a differential oracle (DESIGN section 3, replacement migrations): only this +// module's census calls them, and nothing on a lowering route reaches this module. They are +// deleted with it. +// +// WHERE IT RUNS (calm-boar-904's ruling, relayed by tidy-raven-393): as the +// `if-arm-reader-differential` verb of the native CLI door, v2.cli.compile_cli, because that door +// already acquires its plan's source roots natively. The interpreted parse of one large module took +// hours, which is why it is not a claim. +data if_arm_reader_differential_dissolve: DissolutionCondition = unbound_dissolution(description: "DISSOLVES WHEN every site this census reports is dispositioned (each named as harmless or repaired by #13029's positional reader) and the census has been run over the corpus; the module, its frozen readers and the CLI verb are then deleted together.") + +// --- the frozen pre-#13029 readers (offline oracle only) --- + +fn iard_old_fn_body_capture_optional(spine: Node) -> Optional { + match body_lower_find_captured(root: spine, emitted: ^dag_surface_fn_body) { + ParseSubtreeFound { captured: body } => Present { value: body } + ParseSubtreeAbsent => + if is_empty_conj_root(n: spine) { + Absent + } else { + match body_lower_match_spine_left(spine: spine) { + Present { value: left } => + match body_lower_match_spine_right(spine: spine) { + Present { value: right } => + match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { + Present { value: id } => + match id == ^dag_token_kw_else { + true => iard_old_fn_body_capture_optional(spine: right) + false => + match iard_old_fn_body_capture_optional(spine: left) { + Present { value: found } => Present { value: found } + Absent => iard_old_fn_body_capture_optional(spine: right) + } + } + Absent => + match iard_old_fn_body_capture_optional(spine: left) { + Present { value: found } => Present { value: found } + Absent => iard_old_fn_body_capture_optional(spine: right) + } + } + Absent => Absent + } + Absent => Absent + } + } + } +} + +fn iard_old_then_fn_body_optional(captured: Node) -> Optional { + let spine = body_lower_unwrap_surface_shell(node: captured) + match iard_old_fn_body_capture_optional(spine: spine) { + Present { value: first } => + match body_lower_find_captured(root: captured, emitted: ^dag_surface_fn_body) { + ParseSubtreeFound { captured: only } => Present { value: only } + ParseSubtreeAbsent => Present { value: first } + } + Absent => Absent + } +} + +fn iard_old_else_capture_optional(captured: Node) -> Optional { + let spine = body_lower_unwrap_surface_shell(node: captured) + match is_empty_conj_root(n: spine) { + true => Absent + false => + match body_lower_match_spine_left(spine: spine) { + Present { value: left } => + match body_lower_match_spine_right(spine: spine) { + Present { value: right } => + match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { + Present { value: id } => + match id == ^dag_token_kw_else { + true => + match body_lower_else_part_if_expr_optional(else_part: right) { + Present { value: nested } => Present { value: nested } + Absent => + match iard_old_fn_body_capture_optional(spine: right) { + Present { value: body } => Present { value: body } + Absent => + match body_lower_if_expr_capture_optional(spine: right) { + Present { value: expr } => Present { value: expr } + Absent => Absent + } + } + } + false => + match iard_old_else_capture_optional(captured: left) { + Present { value: found } => Present { value: found } + Absent => iard_old_else_capture_optional(captured: right) + } + } + Absent => + match iard_old_else_capture_optional(captured: left) { + Present { value: found } => Present { value: found } + Absent => iard_old_else_capture_optional(captured: right) + } + } + Absent => Absent + } + Absent => + fold(spine.children, init: Absent, f: fn(acc, e) { + match acc { + Present { value: _ } => acc + Absent => iard_old_else_capture_optional(captured: e.target) + } + }) + } + } +} + +fn iard_old_then_arm_optional(captured: Node) -> Optional { + let spine = body_lower_unwrap_surface_shell(node: captured) + match body_lower_find_captured(root: spine, emitted: ^dag_surface_fn_body) { + ParseSubtreeFound { captured: body } => Present { value: body } + ParseSubtreeAbsent => + match iard_old_then_fn_body_optional(captured: captured) { + Present { value: body } => Present { value: body } + Absent => + match body_lower_match_spine_left(spine: spine) { + Present { value: left } => + match body_lower_match_spine_right(spine: spine) { + Present { value: right } => + match node_atom_identity_optional(node: body_lower_deep_unwrap_optional(node: left)) { + Present { value: id } => + match id == ^dag_token_kw_then { + true => body_lower_if_expr_capture_optional(spine: right) + false => + match iard_old_then_arm_optional(captured: left) { + Present { value: found } => Present { value: found } + Absent => iard_old_then_arm_optional(captured: right) + } + } + Absent => + match iard_old_then_arm_optional(captured: left) { + Present { value: found } => Present { value: found } + Absent => iard_old_then_arm_optional(captured: right) + } + } + Absent => Absent + } + Absent => Absent + } + } + } +} + +// --- the comparison --- + +// HOW TWO READERS' ANSWERS FOR ONE ARM COMPARE. Agreement is OCCURRENCE IDENTITY, never content: +// `if a { if b { x } else { u } } else { u }` has two else arms with equal content and different +// occurrences, and the old reader answering the inner one is exactly the defect. An answer with no +// occurrence id cannot be compared, and is its own arm rather than being counted as agreement. +type IardArmComparison + = IardArmAgree + | IardArmDiffer { old: Node, new: Node } + | IardArmOldOnly { old: Node } + | IardArmNewOnly { new: Node } + | IardArmBothAbsent + | IardArmUnidentified + +fn iard_compare_arm(old: Optional, new: Optional) -> IardArmComparison { + match old { + Absent => + match new { + Absent => IardArmBothAbsent + Present { value: n } => IardArmNewOnly { new: n } + } + Present { value: o } => + match new { + Absent => IardArmOldOnly { old: o } + Present { value: n } => + match node_occurrence_id_optional(node: o) { + Absent => IardArmUnidentified + Present { value: oid } => + match node_occurrence_id_optional(node: n) { + Absent => IardArmUnidentified + Present { value: nid } => + if occurrence_id_eq(left: oid, right: nid) { IardArmAgree } else { IardArmDiffer { old: o, new: n } } + } + } + } + } +} + +fn iard_new_then(arms: Optional) -> Optional { + match arms { + Absent => optional_absent() + Present { value: a } => optional_present(value: a.then_arm) + } +} + +fn iard_new_else(arms: Optional) -> Optional { + match arms { + Absent => optional_absent() + Present { value: a } => optional_present(value: a.else_part) + } +} + +// ONE if_expr SITE: where it is, which declaration holds it, and how each arm compares. +type IardSite { + declaration: String + shell: Node + then_arm: IardArmComparison + else_arm: IardArmComparison +} + +fn iard_site(declaration: String, shell: Node, captured: Node) -> IardSite { + let arms = body_lower_if_row_arms_optional(captured: captured) + IardSite { + declaration: declaration, + shell: shell, + then_arm: iard_compare_arm(old: iard_old_then_arm_optional(captured: captured), new: iard_new_then(arms: arms)), + else_arm: iard_compare_arm(old: iard_old_else_capture_optional(captured: captured), new: iard_new_else(arms: arms)) + } +} + +// A module-scope declaration's name, read by the existing keyword-then-identifier reader; a +// declaration it cannot name (an import, a `test fn`) is reported by its production, never guessed. +fn iard_declaration_name(shell: Node, emitted: Symbol) -> String { + match parse_production_captured_child_optional(node: shell) { + Absent => concat("<", concat(symbol_lexeme(sym: emitted), ">")) + Present { value: captured } => + match dag_surface_kw_then_ident_from_captured(captured: captured) { + Present { value: name } => symbol_lexeme(sym: name) + Absent => concat("<", concat(symbol_lexeme(sym: emitted), ">")) + } + } +} + +fn iard_is_declaration(id: Symbol) -> Bool { + (id == ^dag_surface_fn_decl) + || (id == ^dag_surface_test_fn_decl) + || (id == ^dag_surface_data_decl) + || (id == ^dag_surface_type_decl) + || (id == ^dag_surface_alias_decl) + || (id == ^dag_surface_service_decl) + || (id == ^dag_surface_resource_decl) +} + +// EVERY if_expr IN THE TREE, IN PRE-ORDER, WITH ITS ENCLOSING DECLARATION. Nested ifs are visited: +// the walk does not stop at an if, because the defect's subject is exactly an if inside an if. +// Sites are prepended and the list reversed once (DESIGN section 6, bare minimum cost). +fn iard_collect(node: Node, declaration: String, acc: FreeMonoid) -> FreeMonoid { + let here = match parse_production_emitted_identity_optional(node: node) { + Absent => IardVisit { declaration: declaration, acc: acc } + Present { value: id } => + if id == ^dag_surface_if_expr { + match parse_production_captured_child_optional(node: node) { + Absent => IardVisit { declaration: declaration, acc: acc } + Present { value: captured } => + IardVisit { declaration: declaration, acc: Cons { head: iard_site(declaration: declaration, shell: node, captured: captured), tail: acc } } + } + } else if iard_is_declaration(id: id) { + IardVisit { declaration: iard_declaration_name(shell: node, emitted: id), acc: acc } + } else { + IardVisit { declaration: declaration, acc: acc } + } + } + fold(node.children, init: here.acc, f: fn(a, e) { + iard_collect(node: e.target, declaration: here.declaration, acc: a) + }) +} + +type IardVisit { + declaration: String + acc: FreeMonoid +} + +fn iard_sites(tree: Node) -> FreeMonoid { + list_reverse(xs: iard_collect(node: tree, declaration: "", acc: Empty)) +} + +// --- location --- + +// A NODE'S LOCATION, AS THE SPAN INDEX RESOLVES IT (v2.std.provenance is the one authority for +// following a derived occurrence to its source). The line is derived from the byte start over the +// file's own text; the byte range is carried beside it so a reader can check the derivation. +fn iard_location(node: Node, spans: SpanIndex, source: String) -> String { + match node_occurrence_id_optional(node: node) { + Absent => "" + Present { value: id } => + match span_index_textual_locus_optional(index: spans, id: id) { + Absent => "" + Present { value: locus } => + match locus { + Textual { file: _, extent: e } => + match e { + ByteRange { start: s, end: en } => + concat( + concat("line=", integer_int_to_decimal_string(value: iard_line_of(source: source, offset: s))), + concat(" bytes=", concat(integer_int_to_decimal_string(value: s), concat("..", integer_int_to_decimal_string(value: en)))) + ) + WholeFile => "" + } + _ => "" + } + } + } +} + +fn iard_line_of(source: String, offset: Int) -> Int { + length(xs: split(s: substring(s: source, start: 0, end: offset), delimiter: "\n")) +} + +// --- the per-file and whole-ingest census --- + +type IardFileOutcome + = IardFileRead { path: String, source: String, spans: SpanIndex, sites: FreeMonoid } + | IardFileRefused { path: String, reason: Symbol } + +fn iard_file(prepared: PreparedGrammar, residue: Diagnostics, read: DagSourceReadWitness) -> IardFileOutcome { + let parsed = bind_outcome( + o: tokenize(text: read.source.carried, file: read.compilation_unit, rules: dag_lex()), + f: fn(stream) { parse_module_prepared(tokens: stream, prepared: prepared, validation_residue: residue) } + ) + match parsed { + Rejected { diagnostics: d } => IardFileRefused { path: read.artifact.file_path, reason: diagnostics_fatal_reason(d: d) } + Accepted { value: artifact, diagnostics: _ } => + IardFileRead { + path: read.artifact.file_path, + source: read.source.carried, + spans: artifact.span_index, + sites: iard_sites(tree: artifact.tree) + } + } +} + +fn iard_arm_differs(c: IardArmComparison) -> Bool { + match c { + IardArmAgree => false + IardArmBothAbsent => false + _ => true + } +} + +fn iard_arm_label(c: IardArmComparison) -> String { + match c { + IardArmAgree => "agree" + IardArmDiffer { old: _, new: _ } => "differ" + IardArmOldOnly { old: _ } => "old_only" + IardArmNewOnly { new: _ } => "new_only" + IardArmBothAbsent => "both_absent" + IardArmUnidentified => "unidentified" + } +} + +fn iard_arm_detail(c: IardArmComparison, spans: SpanIndex, source: String) -> String { + match c { + IardArmDiffer { old: o, new: n } => + concat(concat(" old_at=[", concat(iard_location(node: o, spans: spans, source: source), "]")), + concat(" new_at=[", concat(iard_location(node: n, spans: spans, source: source), "]"))) + IardArmOldOnly { old: o } => concat(" old_at=[", concat(iard_location(node: o, spans: spans, source: source), "]")) + IardArmNewOnly { new: n } => concat(" new_at=[", concat(iard_location(node: n, spans: spans, source: source), "]")) + _ => "" + } +} + +// ONE ROW PER (SITE, ARM) WHERE THE READERS DISAGREE. An arm where both agree, or both are absent, +// writes nothing; every other comparison -- including one that could not be identified -- is a row, +// because a census that dropped what it could not compare would undercount silently. +fn iard_arm_row(path: String, site: IardSite, arm: String, c: IardArmComparison, spans: SpanIndex, source: String) -> String { + concat( + concat("if_arm_differs file=", path), + concat( + concat(" decl=", site.declaration), + concat( + concat(" at=[", concat(iard_location(node: site.shell, spans: spans, source: source), "]")), + concat(concat(" arm=", arm), concat(concat(" ", iard_arm_label(c: c)), concat(iard_arm_detail(c: c, spans: spans, source: source), "\n"))) + ) + ) + ) +} + +type IardTally { + text: String + files_read: Int + files_refused: Int + sites: Int + else_differs: Int + then_differs: Int + sites_differing: Int +} + +fn iard_tally_site(t: IardTally, path: String, spans: SpanIndex, source: String, site: IardSite) -> IardTally { + let e = iard_arm_differs(c: site.else_arm) + let h = iard_arm_differs(c: site.then_arm) + let with_then = if h { concat(t.text, iard_arm_row(path: path, site: site, arm: "then", c: site.then_arm, spans: spans, source: source)) } else { t.text } + let with_else = if e { concat(with_then, iard_arm_row(path: path, site: site, arm: "else", c: site.else_arm, spans: spans, source: source)) } else { with_then } + IardTally { + text: with_else, + files_read: t.files_read, + files_refused: t.files_refused, + sites: t.sites + 1, + else_differs: if e { t.else_differs + 1 } else { t.else_differs }, + then_differs: if h { t.then_differs + 1 } else { t.then_differs }, + sites_differing: if e || h { t.sites_differing + 1 } else { t.sites_differing } + } +} + +fn iard_tally_file(t: IardTally, outcome: IardFileOutcome) -> IardTally { + match outcome { + IardFileRefused { path: p, reason: r } => + IardTally { + text: concat(t.text, concat(concat("file_refused file=", p), concat(concat(" reason=", symbol_lexeme(sym: r)), "\n"))), + files_read: t.files_read, + files_refused: t.files_refused + 1, + sites: t.sites, + else_differs: t.else_differs, + then_differs: t.then_differs, + sites_differing: t.sites_differing + } + IardFileRead { path: p, source: src, spans: spans, sites: sites } => + let counted = IardTally { + text: t.text, + files_read: t.files_read + 1, + files_refused: t.files_refused, + sites: t.sites, + else_differs: t.else_differs, + then_differs: t.then_differs, + sites_differing: t.sites_differing + } + fold_list(xs: sites, empty: counted, cons: fn(acc, s) { + iard_tally_site(t: acc, path: p, spans: spans, source: src, site: s) + }) + } +} + +// THE CENSUS'S ANSWER. The grammar failing to prepare is no observation at all, distinct from a +// file refusing: a refused file is counted and named, and the census goes on, so one run names +// every file it could not read. +type IardCensus + = IardGrammarUnprepared { reason: Symbol } + | IardCensusTaken { tally: IardTally } + +fn if_arm_reader_differential(ingest: SourceRootIngest) -> IardCensus { + match dag_prepared_grammar() { + Rejected { diagnostics: d } => IardGrammarUnprepared { reason: diagnostics_fatal_reason(d: d) } + Accepted { value: prepared, diagnostics: residue } => + IardCensusTaken { + tally: fold_list( + xs: ingest, + empty: IardTally { text: "", files_read: 0, files_refused: 0, sites: 0, else_differs: 0, then_differs: 0, sites_differing: 0 }, + cons: fn(acc, read) { iard_tally_file(t: acc, outcome: iard_file(prepared: prepared, residue: residue, read: read)) } + ) + } + } +} + +// THE TERMINAL LINE: the population and its denominators together, so a total of zero differing +// sites can be read against how many sites and files were actually observed. +fn if_arm_reader_differential_summary(t: IardTally) -> String { + concat( + concat("if_arm_reader_differential files_read=", integer_int_to_decimal_string(value: t.files_read)), + concat( + concat(" files_refused=", integer_int_to_decimal_string(value: t.files_refused)), + concat( + concat(" if_sites=", integer_int_to_decimal_string(value: t.sites)), + concat( + concat(" sites_differing=", integer_int_to_decimal_string(value: t.sites_differing)), + concat( + concat(" else_differs=", integer_int_to_decimal_string(value: t.else_differs)), + concat(concat(" then_differs=", integer_int_to_decimal_string(value: t.then_differs)), "\n") + ) + ) + ) + ) + ) +} + +fn if_arm_reader_differential_text(c: IardCensus) -> String { + match c { + IardGrammarUnprepared { reason: r } => concat("if_arm_reader_differential grammar_unprepared reason=", concat(symbol_lexeme(sym: r), "\n")) + IardCensusTaken { tally: t } => concat(t.text, if_arm_reader_differential_summary(t: t)) + } +} diff --git a/src/v2/test/claim/cli/if_arm_reader_differential_test.dag b/src/v2/test/claim/cli/if_arm_reader_differential_test.dag new file mode 100644 index 00000000000..e60b4834105 --- /dev/null +++ b/src/v2/test/claim/cli/if_arm_reader_differential_test.dag @@ -0,0 +1,111 @@ +module v2.test.claim.cli.if_arm_reader_differential + +import v2.cli.if_arm_reader_differential { + IardCensus, + IardCensusTaken, + IardGrammarUnprepared, + IardTally, + if_arm_reader_differential, + if_arm_reader_differential_text +} +import v2.compiler.source_authority { DagSourceReadWitness, SourceRootIngest } +import v2.std.artifact { Artifact, SourceFile } +import v2.std.cross_tree.import_model { DagTree } +import v2.std.live_tree { LiveTreeDisposition, SubstrateInputsOnly } +import v2.std.logic { Bool } +import v2.std.node { Symbol } +import v2.std.text { String } +import std.algebra { Cons, Empty } +import extdeps.communication.medium { Lossless, Medium } + +data live_tree_disposition: LiveTreeDisposition = SubstrateInputsOnly + +// THE DIFFERENTIAL'S COMPARATOR, DISCRIMINATED BEFORE IT IS TRUSTED WITH A POPULATION. +// +// THE SUBJECT is v2.cli.if_arm_reader_differential if_arm_reader_differential: for each if_expr +// capture it asks the frozen pre-gunbc#13029 readers and the positional +// body_lower_if_row_arms_optional for each arm and counts the sites where they answer a different +// occurrence. A census reporting zero is only evidence if its comparator can report one, so: +// +// THE POSITIVE CONTROL: `if a { if b { 1 } else { 2 } } else { 3 }`. The old else reader entered the +// then-arm and answered the NESTED else `{ 2 }` for the OUTER if; the positional reader answers +// `{ 3 }`. Exactly one site (the outer if) differs, on its else arm, and its row names the +// declaration and the line. +// THE NEGATIVE CONTROLS, where the old reader was right: an `else if` chain, and an if nested in the +// ELSE arm. Every site there agrees, so a comparator that reported everything as differing is red. +// A FILE THAT DOES NOT PARSE is counted as refused and named, never folded into zero sites. +data iard_nested_then_source: String = "module v2.test.iard_nested_then\n\nimport v2.std.logic { Bool }\nimport v2.std.integer { Int }\n\nfn iard_nested_then_probe(a: Bool, b: Bool) -> Int {\n if a {\n if b { 1 } else { 2 }\n } else {\n 3\n }\n}\n" + +data iard_agreeing_source: String = "module v2.test.iard_agreeing\n\nimport v2.std.logic { Bool }\nimport v2.std.integer { Int }\n\nfn iard_else_if_probe(a: Bool, b: Bool) -> Int {\n if a { 1 } else if b { 2 } else { 3 }\n}\n\nfn iard_nested_else_probe(a: Bool, b: Bool) -> Int {\n if a {\n 1\n } else {\n if b { 2 } else { 3 }\n }\n}\n" + +data iard_unparsable_source: String = "module v2.test.iard_unparsable\n\nfn iard_broken( -> {\n" + +fn iard_read(source: String, id: Symbol, unit: Symbol, path: String) -> DagSourceReadWitness { + DagSourceReadWitness { + source: Medium { carried: source, fidelity: Lossless }, + artifact: Artifact { kind: SourceFile, id: id, file_path: path }, + compilation_unit: unit, + source_root: DagTree + } +} + +data iard_nested_then_ingest: SourceRootIngest = Cons { + head: iard_read(source: iard_nested_then_source, id: ^iard_nested_then_artifact, unit: ^iard_nested_then_cu, path: "src/v2/test/fixture/cli/iard_nested_then.dag"), + tail: Empty +} + +data iard_agreeing_ingest: SourceRootIngest = Cons { + head: iard_read(source: iard_agreeing_source, id: ^iard_agreeing_artifact, unit: ^iard_agreeing_cu, path: "src/v2/test/fixture/cli/iard_agreeing.dag"), + tail: Cons { + head: iard_read(source: iard_unparsable_source, id: ^iard_unparsable_artifact, unit: ^iard_unparsable_cu, path: "src/v2/test/fixture/cli/iard_unparsable.dag"), + tail: Empty + } +} + +// ONE FRONT END PER INGEST, stored as the census's own portable answer (counts and rendered rows, +// no Node), enrolled WARM in v2.workflow.floor_pure_producer_share so each claim reads it. +type IardFixtureCensuses { + nested_then: IardCensus + agreeing: IardCensus +} + +fn iard_fixture_censuses() -> IardFixtureCensuses { + IardFixtureCensuses { + nested_then: if_arm_reader_differential(ingest: iard_nested_then_ingest), + agreeing: if_arm_reader_differential(ingest: iard_agreeing_ingest) + } +} + +fn iard_tally_holds(c: IardCensus, check: fn(IardTally) -> Bool) -> Bool { + match c { + IardGrammarUnprepared { reason: _ } => false + IardCensusTaken { tally: t } => check(t) + } +} + +test fn an_if_nested_in_the_then_arm_differs_on_the_outer_else_only_holds() -> Bool { + iard_tally_holds(c: iard_fixture_censuses().nested_then, check: fn(t) { + (t.files_read == 1) && (t.files_refused == 0) && (t.sites == 2) + && (t.sites_differing == 1) && (t.else_differs == 1) && (t.then_differs == 0) + }) +} + +test fn the_differing_row_names_the_declaration_the_arm_and_the_outer_if_line_holds() -> Bool { + let text = if_arm_reader_differential_text(c: iard_fixture_censuses().nested_then) + string_contains(s: text, pattern: "if_arm_differs file=src/v2/test/fixture/cli/iard_nested_then.dag decl=iard_nested_then_probe at=[line=7 ") + && string_contains(s: text, pattern: " arm=else differ old_at=[line=8 ") + && string_contains(s: text, pattern: " new_at=[line=9 ") +} + +test fn an_else_if_chain_and_an_if_nested_in_the_else_arm_agree_holds() -> Bool { + iard_tally_holds(c: iard_fixture_censuses().agreeing, check: fn(t) { + (t.sites == 4) && (t.sites_differing == 0) && (t.else_differs == 0) && (t.then_differs == 0) + }) +} + +test fn a_file_that_does_not_parse_is_counted_refused_and_named_holds() -> Bool { + iard_tally_holds(c: iard_fixture_censuses().agreeing, check: fn(t) { + (t.files_read == 1) && (t.files_refused == 1) + }) + && string_contains(s: if_arm_reader_differential_text(c: iard_fixture_censuses().agreeing), pattern: "file_refused file=src/v2/test/fixture/cli/iard_unparsable.dag reason=") +} diff --git a/src/v2/test/claim/cli/v2_native_cli_test.dag b/src/v2/test/claim/cli/v2_native_cli_test.dag index 19a3e8f1ec2..1ee7e96ec1f 100644 --- a/src/v2/test/claim/cli/v2_native_cli_test.dag +++ b/src/v2/test/claim/cli/v2_native_cli_test.dag @@ -18,6 +18,8 @@ import v2.cli.compile_cli { CliTargetRust, CliTargetSwift, CliUsageRefused, + CliDifferentialReported, + CliIfArmReaderDifferential, cli_emit_target_spelling, v2_cli_exit, v2_cli_parse, @@ -68,6 +70,7 @@ fn plan_first_root(plan: CliPlan) -> String { HeadAbsent => "" HeadFound { value: first } => first } + CliIfArmReaderDifferential { source_roots: _ } => "" } } } @@ -79,6 +82,8 @@ fn outcome_refusal_reason(outcome: CliRunOutcome) -> Symbol { CliUsageRefused { reason: reason, detail: _ } => reason CliEmitted { text: _, closure_size: _, refused_count: _, refusal_detail: _ } => ^cli_outcome_was_not_refused + CliDifferentialReported { text: _, observed: _, detail: _ } => + ^cli_outcome_was_not_refused } } @@ -88,6 +93,7 @@ fn outcome_emitted(outcome: CliRunOutcome) -> Bool { CliCensusRefused { reason: _, detail: _, text: _ } => false CliUsageRefused { reason: _, detail: _ } => false CliEmitted { text: _, closure_size: _, refused_count: _, refusal_detail: _ } => true + CliDifferentialReported { text: _, observed: _, detail: _ } => false } } @@ -259,6 +265,7 @@ fn plan_is_target(plan: CliPlan, expected: CliEmitTarget) -> Bool { match verb { CliEmitClosure { entry: _, source_roots: _, target: target } => cli_emit_target_spelling(t: target) == cli_emit_target_spelling(t: expected) + CliIfArmReaderDifferential { source_roots: _ } => false } } } @@ -470,5 +477,6 @@ test fn a_trailing_token_refusal_locates_its_fatal_in_the_file_holds() -> Bool { && string_contains(s: detail, pattern: "parse_g0_tokens_remain @ cli_probe_trailing_cu bytes ") CliUsageRefused { reason: _, detail: _ } => false CliEmitted { text: _, closure_size: _, refused_count: _, refusal_detail: _ } => false + CliDifferentialReported { text: _, observed: _, detail: _ } => false } } diff --git a/src/v2/test/claim/provenance/occurrence_file_attribution_test.dag b/src/v2/test/claim/provenance/occurrence_file_attribution_test.dag index a8ec55bd083..c8b8b9f66be 100644 --- a/src/v2/test/claim/provenance/occurrence_file_attribution_test.dag +++ b/src/v2/test/claim/provenance/occurrence_file_attribution_test.dag @@ -21,6 +21,7 @@ import v2.cli.compile_cli { CliRunRefused, CliTargetRust, CliUsageRefused, + CliDifferentialReported, v2_cli_run } import v2.std.algebra { list_reverse, list_starts_with } @@ -282,6 +283,7 @@ fn door_probe_cli_detail() -> String { CliRunRefused { reason: _, detail: detail } => detail CliCensusRefused { reason: _, detail: detail, text: _ } => detail CliUsageRefused { reason: _, detail: detail } => detail + CliDifferentialReported { text: _, observed: _, detail: detail } => detail } } diff --git a/src/v2/workflow/floor_pure_producer_share.dag b/src/v2/workflow/floor_pure_producer_share.dag index 72e16e87760..929074f132b 100644 --- a/src/v2/workflow/floor_pure_producer_share.dag +++ b/src/v2/workflow/floor_pure_producer_share.dag @@ -654,6 +654,9 @@ import v2.std.collection { List } // THE IF-ARM-POSITION ROW IS THE SAME GROUND AT FOUR SPECIMENS: v2.test.claim.body_lowering.if_arm_position // iap_verdicts tokenizes, parses and normalizes four inline modules and stores four Bools. Its first floor // run (37073686904) refused all four claims over the enrolment margin, at 353k-687k eval steps each. +// THE IF-ARM READER DIFFERENTIAL ROW IS THE SAME GROUND AT THREE SPECIMENS: +// v2.test.claim.cli.if_arm_reader_differential iard_fixture_censuses tokenizes and parses three inline +// modules and stores two censuses (counts and rendered rows, no Node), portable for the same reason. // THE XL-2 REIFY-OPERAND ROW: v2.test.claim.body_lowering.reify_operand_refusal ror_verdicts tokenizes // and parses two inline modules and reifies two supplied bodies over each; the stored value is four // verdicts (no Node, no closure), portable for the same reason cav_outcomes is. @@ -831,6 +834,7 @@ data floor_cross_claim_pure_producers_warm: List = [ "v2.test.claim.namespace_xl0.wildcard_arm_resolve.wc_outcomes", "v2.test.claim.namespace_xl0.else_arm_nested_if.eam_outcomes", "v2.test.claim.body_lowering.if_arm_position.iap_verdicts", + "v2.test.claim.cli.if_arm_reader_differential.iard_fixture_censuses", "v2.test.cli.v2_native_cli.cli_probe_trailing_outcome", "v2.test.emit.closure_emit_arrow_body_refusal.arrow_body_probe_verdict", "v2.test.emit.closure_emit_arrow_body_refusal.arrow_body_mixed_probe_verdict", From 73e468aa24f11af28e6a895775d49e53a093d78d Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Sat, 3 Oct 2026 00:45:05 +0000 Subject: [PATCH 5/9] compile_cli: the differential verb carries its own usage line; the emit usage the door control pins is unchanged Co-Authored-By: Claude Opus 5.5 (1M context) --- src/v2/cli/compile_cli.dag | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/src/v2/cli/compile_cli.dag b/src/v2/cli/compile_cli.dag index c7fe6f1967c..6b45306a672 100644 --- a/src/v2/cli/compile_cli.dag +++ b/src/v2/cli/compile_cli.dag @@ -181,7 +181,9 @@ data cli_verb_emit: String = "emit" // (v2.compiler.compile). It takes source roots only, and is deleted with its module. data cli_verb_if_arm_reader_differential: String = "if-arm-reader-differential" -data cli_usage: String = "usage: emit --entry --source-root [--source-root ]... [--target ] | if-arm-reader-differential --source-root [--source-root ]..." +data cli_usage: String = "usage: emit --entry --source-root [--source-root ]... [--target ]" + +data cli_differential_usage: String = "usage: if-arm-reader-differential --source-root [--source-root ]..." fn cli_parse_step(state: CliParseState, arg: String) -> CliParseState { match state { @@ -204,7 +206,7 @@ fn cli_parse_step(state: CliParseState, arg: String) -> CliParseState { } else { CliParseFailed { reason: ^cli_unknown_option, - detail: concat(concat("unknown option ", arg), concat(" -- ", cli_usage)) + detail: concat(concat("unknown option ", arg), concat(" -- ", cli_differential_usage)) } } CliParseSeekDifferentialRootValue { roots: roots } => @@ -307,7 +309,7 @@ fn cli_parse_finish(state: CliParseState) -> CliPlan { if length(xs: roots) == 0 { CliPlanRefused { reason: ^cli_no_source_root, - detail: concat("if-arm-reader-differential needs at least one source root -- ", cli_usage) + detail: concat("if-arm-reader-differential needs at least one source root -- ", cli_differential_usage) } } else { CliPlanParsed { verb: CliIfArmReaderDifferential { source_roots: list_reverse(xs: roots) } } From 1add706e9a662e3cf7970248d1804483566b1eaf Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Sat, 3 Oct 2026 05:55:49 +0000 Subject: [PATCH 6/9] if_arm_reader_differential: name every arm instead of a wildcard over a closed coproduct The floor refused NonFoldResidueRosterDiverged at iard_arm_detail, iard_arm_differs and iard_location: each matched a closed coproduct (IardArmComparison, Locus) with a wildcard. The arms are enumerated, so a new variant refuses to compile here instead of being absorbed. Co-Authored-By: Claude Opus 5.5 (1M context) --- src/v2/cli/if_arm_reader_differential.dag | 16 ++++++++++++---- 1 file changed, 12 insertions(+), 4 deletions(-) diff --git a/src/v2/cli/if_arm_reader_differential.dag b/src/v2/cli/if_arm_reader_differential.dag index 313a7008c0e..e2a2ca007aa 100644 --- a/src/v2/cli/if_arm_reader_differential.dag +++ b/src/v2/cli/if_arm_reader_differential.dag @@ -27,7 +27,7 @@ import v2.extdeps.languages.dag { parse_production_emitted_identity_optional } import v2.std.algebra { fold_list, length, list_reverse } -import v2.std.diagnostic { Accepted, Diagnostics, NonEmptyDiagnostics, Rejected, Textual, ByteRange, WholeFile, bind_outcome, diagnostics_fatal_reason } +import v2.std.diagnostic { Accepted, Diagnostics, NonEmptyDiagnostics, Rejected, Textual, NodeLocus, SourcePortLocus, InvariantPortLocus, DeclarationLocus, ByteRange, WholeFile, bind_outcome, diagnostics_fatal_reason } import v2.std.grammar { node_atom_identity_optional } import v2.std.integer { integer_int_to_decimal_string } import v2.std.node { Node, Symbol, is_empty_conj_root, symbol_lexeme } @@ -343,7 +343,10 @@ fn iard_location(node: Node, spans: SpanIndex, source: String) -> String { ) WholeFile => "" } - _ => "" + NodeLocus { anchor: _ } => "" + SourcePortLocus { anchor: _, file: _, extent: _ } => "" + InvariantPortLocus { anchor: _, invariant: _ } => "" + DeclarationLocus { declaration: _ } => "" } } } @@ -380,7 +383,10 @@ fn iard_arm_differs(c: IardArmComparison) -> Bool { match c { IardArmAgree => false IardArmBothAbsent => false - _ => true + IardArmDiffer { old: _, new: _ } => true + IardArmOldOnly { old: _ } => true + IardArmNewOnly { new: _ } => true + IardArmUnidentified => true } } @@ -402,7 +408,9 @@ fn iard_arm_detail(c: IardArmComparison, spans: SpanIndex, source: String) -> St concat(" new_at=[", concat(iard_location(node: n, spans: spans, source: source), "]"))) IardArmOldOnly { old: o } => concat(" old_at=[", concat(iard_location(node: o, spans: spans, source: source), "]")) IardArmNewOnly { new: n } => concat(" new_at=[", concat(iard_location(node: n, spans: spans, source: source), "]")) - _ => "" + IardArmAgree => "" + IardArmBothAbsent => "" + IardArmUnidentified => "" } } From a0b5b36f9570d73df7e0e28b5b79bdd13dd87f3c Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Sat, 3 Oct 2026 07:19:53 +0000 Subject: [PATCH 7/9] if_arm_reader_differential: locate sites by octet range only; a multibyte control calm-boar-904 (via tidy-raven-393): iard_line_of counted newlines in substring(source, 0, offset), but the offset is a UTF-8 octet count (the tokenizer's ByteRange) and the emitted substring indexes scalars, so any multibyte character before a site raised the reported line. No line-from-octet reader exists in the corpus, so the census reports the byte range only, which is the span index's own authority. New control: a nested if after 2-, 3- and 4-octet scalars, located at octet 237 (scalar 223). Co-Authored-By: Claude Opus 5.5 (1M context) --- src/v2/cli/if_arm_reader_differential.dag | 50 +++++++++---------- .../cli/if_arm_reader_differential_test.dag | 33 +++++++++--- src/v2/workflow/floor_pure_producer_share.dag | 6 +-- 3 files changed, 53 insertions(+), 36 deletions(-) diff --git a/src/v2/cli/if_arm_reader_differential.dag b/src/v2/cli/if_arm_reader_differential.dag index e2a2ca007aa..2b651e0ecda 100644 --- a/src/v2/cli/if_arm_reader_differential.dag +++ b/src/v2/cli/if_arm_reader_differential.dag @@ -26,7 +26,7 @@ import v2.extdeps.languages.dag { parse_production_captured_child_optional, parse_production_emitted_identity_optional } -import v2.std.algebra { fold_list, length, list_reverse } +import v2.std.algebra { fold_list, list_reverse } import v2.std.diagnostic { Accepted, Diagnostics, NonEmptyDiagnostics, Rejected, Textual, NodeLocus, SourcePortLocus, InvariantPortLocus, DeclarationLocus, ByteRange, WholeFile, bind_outcome, diagnostics_fatal_reason } import v2.std.grammar { node_atom_identity_optional } import v2.std.integer { integer_int_to_decimal_string } @@ -324,9 +324,13 @@ fn iard_sites(tree: Node) -> FreeMonoid { // --- location --- // A NODE'S LOCATION, AS THE SPAN INDEX RESOLVES IT (v2.std.provenance is the one authority for -// following a derived occurrence to its source). The line is derived from the byte start over the -// file's own text; the byte range is carried beside it so a reader can check the derivation. -fn iard_location(node: Node, spans: SpanIndex, source: String) -> String { +// following a derived occurrence to its source): the file's byte range, in the UTF-8 octets the +// tokenizer counts (v2.compiler.tokenize lexeme_utf8_octet_count). NO LINE IS DERIVED HERE. No +// line-from-octet reader exists in the corpus, and deriving one with `substring` is wrong: the +// emitted substring indexes Unicode scalars, so every multibyte character before a site moved the +// line it reported (calm-boar-904, review of gunbc#13051). The byte range is the authority; a reader +// converts it to a line with a byte-aware tool. +fn iard_location(node: Node, spans: SpanIndex) -> String { match node_occurrence_id_optional(node: node) { Absent => "" Present { value: id } => @@ -337,10 +341,7 @@ fn iard_location(node: Node, spans: SpanIndex, source: String) -> String { Textual { file: _, extent: e } => match e { ByteRange { start: s, end: en } => - concat( - concat("line=", integer_int_to_decimal_string(value: iard_line_of(source: source, offset: s))), - concat(" bytes=", concat(integer_int_to_decimal_string(value: s), concat("..", integer_int_to_decimal_string(value: en)))) - ) + concat("bytes=", concat(integer_int_to_decimal_string(value: s), concat("..", integer_int_to_decimal_string(value: en)))) WholeFile => "" } NodeLocus { anchor: _ } => "" @@ -352,14 +353,10 @@ fn iard_location(node: Node, spans: SpanIndex, source: String) -> String { } } -fn iard_line_of(source: String, offset: Int) -> Int { - length(xs: split(s: substring(s: source, start: 0, end: offset), delimiter: "\n")) -} - // --- the per-file and whole-ingest census --- type IardFileOutcome - = IardFileRead { path: String, source: String, spans: SpanIndex, sites: FreeMonoid } + = IardFileRead { path: String, spans: SpanIndex, sites: FreeMonoid } | IardFileRefused { path: String, reason: Symbol } fn iard_file(prepared: PreparedGrammar, residue: Diagnostics, read: DagSourceReadWitness) -> IardFileOutcome { @@ -372,7 +369,6 @@ fn iard_file(prepared: PreparedGrammar, residue: Diagnostics, read: DagSourceRea Accepted { value: artifact, diagnostics: _ } => IardFileRead { path: read.artifact.file_path, - source: read.source.carried, spans: artifact.span_index, sites: iard_sites(tree: artifact.tree) } @@ -401,13 +397,13 @@ fn iard_arm_label(c: IardArmComparison) -> String { } } -fn iard_arm_detail(c: IardArmComparison, spans: SpanIndex, source: String) -> String { +fn iard_arm_detail(c: IardArmComparison, spans: SpanIndex) -> String { match c { IardArmDiffer { old: o, new: n } => - concat(concat(" old_at=[", concat(iard_location(node: o, spans: spans, source: source), "]")), - concat(" new_at=[", concat(iard_location(node: n, spans: spans, source: source), "]"))) - IardArmOldOnly { old: o } => concat(" old_at=[", concat(iard_location(node: o, spans: spans, source: source), "]")) - IardArmNewOnly { new: n } => concat(" new_at=[", concat(iard_location(node: n, spans: spans, source: source), "]")) + concat(concat(" old_at=[", concat(iard_location(node: o, spans: spans), "]")), + concat(" new_at=[", concat(iard_location(node: n, spans: spans), "]"))) + IardArmOldOnly { old: o } => concat(" old_at=[", concat(iard_location(node: o, spans: spans), "]")) + IardArmNewOnly { new: n } => concat(" new_at=[", concat(iard_location(node: n, spans: spans), "]")) IardArmAgree => "" IardArmBothAbsent => "" IardArmUnidentified => "" @@ -417,14 +413,14 @@ fn iard_arm_detail(c: IardArmComparison, spans: SpanIndex, source: String) -> St // ONE ROW PER (SITE, ARM) WHERE THE READERS DISAGREE. An arm where both agree, or both are absent, // writes nothing; every other comparison -- including one that could not be identified -- is a row, // because a census that dropped what it could not compare would undercount silently. -fn iard_arm_row(path: String, site: IardSite, arm: String, c: IardArmComparison, spans: SpanIndex, source: String) -> String { +fn iard_arm_row(path: String, site: IardSite, arm: String, c: IardArmComparison, spans: SpanIndex) -> String { concat( concat("if_arm_differs file=", path), concat( concat(" decl=", site.declaration), concat( - concat(" at=[", concat(iard_location(node: site.shell, spans: spans, source: source), "]")), - concat(concat(" arm=", arm), concat(concat(" ", iard_arm_label(c: c)), concat(iard_arm_detail(c: c, spans: spans, source: source), "\n"))) + concat(" at=[", concat(iard_location(node: site.shell, spans: spans), "]")), + concat(concat(" arm=", arm), concat(concat(" ", iard_arm_label(c: c)), concat(iard_arm_detail(c: c, spans: spans), "\n"))) ) ) ) @@ -440,11 +436,11 @@ type IardTally { sites_differing: Int } -fn iard_tally_site(t: IardTally, path: String, spans: SpanIndex, source: String, site: IardSite) -> IardTally { +fn iard_tally_site(t: IardTally, path: String, spans: SpanIndex, site: IardSite) -> IardTally { let e = iard_arm_differs(c: site.else_arm) let h = iard_arm_differs(c: site.then_arm) - let with_then = if h { concat(t.text, iard_arm_row(path: path, site: site, arm: "then", c: site.then_arm, spans: spans, source: source)) } else { t.text } - let with_else = if e { concat(with_then, iard_arm_row(path: path, site: site, arm: "else", c: site.else_arm, spans: spans, source: source)) } else { with_then } + let with_then = if h { concat(t.text, iard_arm_row(path: path, site: site, arm: "then", c: site.then_arm, spans: spans)) } else { t.text } + let with_else = if e { concat(with_then, iard_arm_row(path: path, site: site, arm: "else", c: site.else_arm, spans: spans)) } else { with_then } IardTally { text: with_else, files_read: t.files_read, @@ -468,7 +464,7 @@ fn iard_tally_file(t: IardTally, outcome: IardFileOutcome) -> IardTally { then_differs: t.then_differs, sites_differing: t.sites_differing } - IardFileRead { path: p, source: src, spans: spans, sites: sites } => + IardFileRead { path: p, spans: spans, sites: sites } => let counted = IardTally { text: t.text, files_read: t.files_read + 1, @@ -479,7 +475,7 @@ fn iard_tally_file(t: IardTally, outcome: IardFileOutcome) -> IardTally { sites_differing: t.sites_differing } fold_list(xs: sites, empty: counted, cons: fn(acc, s) { - iard_tally_site(t: acc, path: p, spans: spans, source: src, site: s) + iard_tally_site(t: acc, path: p, spans: spans, site: s) }) } } diff --git a/src/v2/test/claim/cli/if_arm_reader_differential_test.dag b/src/v2/test/claim/cli/if_arm_reader_differential_test.dag index e60b4834105..a349787dd33 100644 --- a/src/v2/test/claim/cli/if_arm_reader_differential_test.dag +++ b/src/v2/test/claim/cli/if_arm_reader_differential_test.dag @@ -30,7 +30,7 @@ data live_tree_disposition: LiveTreeDisposition = SubstrateInputsOnly // THE POSITIVE CONTROL: `if a { if b { 1 } else { 2 } } else { 3 }`. The old else reader entered the // then-arm and answered the NESTED else `{ 2 }` for the OUTER if; the positional reader answers // `{ 3 }`. Exactly one site (the outer if) differs, on its else arm, and its row names the -// declaration and the line. +// declaration and the byte range of the site and of each arm. // THE NEGATIVE CONTROLS, where the old reader was right: an `else if` chain, and an if nested in the // ELSE arm. Every site there agrees, so a comparator that reported everything as differing is red. // A FILE THAT DOES NOT PARSE is counted as refused and named, never folded into zero sites. @@ -38,6 +38,8 @@ data iard_nested_then_source: String = "module v2.test.iard_nested_then\n\nimpor data iard_agreeing_source: String = "module v2.test.iard_agreeing\n\nimport v2.std.logic { Bool }\nimport v2.std.integer { Int }\n\nfn iard_else_if_probe(a: Bool, b: Bool) -> Int {\n if a { 1 } else if b { 2 } else { 3 }\n}\n\nfn iard_nested_else_probe(a: Bool, b: Bool) -> Int {\n if a {\n 1\n } else {\n if b { 2 } else { 3 }\n }\n}\n" +data iard_multibyte_source: String = "module v2.test.iard_multibyte\n\nimport v2.std.logic { Bool }\nimport v2.std.integer { Int }\nimport v2.std.text { String }\n\ndata iard_multibyte_label: String = \"é—日本語 ✓ 𝄞\"\n\nfn iard_multibyte_probe(a: Bool, b: Bool) -> Int {\n if a {\n if b { 1 } else { 2 }\n } else {\n 3\n }\n}\n" + data iard_unparsable_source: String = "module v2.test.iard_unparsable\n\nfn iard_broken( -> {\n" fn iard_read(source: String, id: Symbol, unit: Symbol, path: String) -> DagSourceReadWitness { @@ -54,6 +56,11 @@ data iard_nested_then_ingest: SourceRootIngest = Cons { tail: Empty } +data iard_multibyte_ingest: SourceRootIngest = Cons { + head: iard_read(source: iard_multibyte_source, id: ^iard_multibyte_artifact, unit: ^iard_multibyte_cu, path: "src/v2/test/fixture/cli/iard_multibyte.dag"), + tail: Empty +} + data iard_agreeing_ingest: SourceRootIngest = Cons { head: iard_read(source: iard_agreeing_source, id: ^iard_agreeing_artifact, unit: ^iard_agreeing_cu, path: "src/v2/test/fixture/cli/iard_agreeing.dag"), tail: Cons { @@ -66,12 +73,14 @@ data iard_agreeing_ingest: SourceRootIngest = Cons { // no Node), enrolled WARM in v2.workflow.floor_pure_producer_share so each claim reads it. type IardFixtureCensuses { nested_then: IardCensus + multibyte: IardCensus agreeing: IardCensus } fn iard_fixture_censuses() -> IardFixtureCensuses { IardFixtureCensuses { nested_then: if_arm_reader_differential(ingest: iard_nested_then_ingest), + multibyte: if_arm_reader_differential(ingest: iard_multibyte_ingest), agreeing: if_arm_reader_differential(ingest: iard_agreeing_ingest) } } @@ -90,11 +99,23 @@ test fn an_if_nested_in_the_then_arm_differs_on_the_outer_else_only_holds() -> B }) } -test fn the_differing_row_names_the_declaration_the_arm_and_the_outer_if_line_holds() -> Bool { - let text = if_arm_reader_differential_text(c: iard_fixture_censuses().nested_then) - string_contains(s: text, pattern: "if_arm_differs file=src/v2/test/fixture/cli/iard_nested_then.dag decl=iard_nested_then_probe at=[line=7 ") - && string_contains(s: text, pattern: " arm=else differ old_at=[line=8 ") - && string_contains(s: text, pattern: " new_at=[line=9 ") +// THE EXPECTED EXTENTS ARE THE SOURCE'S UTF-8 OCTETS: the outer `if` token, the nested else's `{` +// (what the old reader answered) and the outer else's `{` (the real arm). +test fn the_differing_row_names_the_declaration_the_arm_and_each_byte_range_holds() -> Bool { + string_contains( + s: if_arm_reader_differential_text(c: iard_fixture_censuses().nested_then), + pattern: "if_arm_differs file=src/v2/test/fixture/cli/iard_nested_then.dag decl=iard_nested_then_probe at=[bytes=148..150] arm=else differ old_at=[bytes=175..176] new_at=[bytes=190..191]\n" + ) +} + +// MULTIBYTE UTF-8 BEFORE THE SITE (calm-boar-904's objection to gunbc#13051): the label holds +// 2-, 3- and 4-octet scalars, so the outer `if` is at octet 237 but at scalar 223. A location +// counted in scalars anywhere on the route reports 223 and is red here. +test fn multibyte_source_before_a_site_is_located_in_octets_holds() -> Bool { + string_contains( + s: if_arm_reader_differential_text(c: iard_fixture_censuses().multibyte), + pattern: "if_arm_differs file=src/v2/test/fixture/cli/iard_multibyte.dag decl=iard_multibyte_probe at=[bytes=237..239] arm=else differ old_at=[bytes=264..265] new_at=[bytes=279..280]\n" + ) } test fn an_else_if_chain_and_an_if_nested_in_the_else_arm_agree_holds() -> Bool { diff --git a/src/v2/workflow/floor_pure_producer_share.dag b/src/v2/workflow/floor_pure_producer_share.dag index 5c964b0d7b6..cb993383fbb 100644 --- a/src/v2/workflow/floor_pure_producer_share.dag +++ b/src/v2/workflow/floor_pure_producer_share.dag @@ -655,9 +655,9 @@ import v2.std.collection { List } // iap_verdicts tokenizes, parses and normalizes four inline modules and stores four Bools. Unshared, // each claim frame paid its own tokenize -> parse -> normalize, over the required floor's new-witness // enrolment margin (v2.workflow.required_floor claim_eval_step_budget_for_identity). -// THE IF-ARM READER DIFFERENTIAL ROW IS THE SAME GROUND AT THREE SPECIMENS: -// v2.test.claim.cli.if_arm_reader_differential iard_fixture_censuses tokenizes and parses three inline -// modules and stores two censuses (counts and rendered rows, no Node), portable for the same reason. +// THE IF-ARM READER DIFFERENTIAL ROW IS THE SAME GROUND AT FOUR SPECIMENS: +// v2.test.claim.cli.if_arm_reader_differential iard_fixture_censuses tokenizes and parses four inline +// modules and stores three censuses (counts and rendered rows, no Node), portable for the same reason. // THE XL-2 REIFY-OPERAND ROW: v2.test.claim.body_lowering.reify_operand_refusal ror_verdicts tokenizes // and parses two inline modules and reifies two supplied bodies over each; the stored value is four // verdicts (no Node, no closure), portable for the same reason cav_outcomes is. From ee8b94e7722bd0f6793142ffd7da0a888c300b01 Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Sat, 3 Oct 2026 07:29:22 +0000 Subject: [PATCH 8/9] recurring_failure_mode: octet_offset_consumed_by_a_scalar_indexed_text_operation The class #13051's review caught: an octet ByteRange offset handed to scalar-indexed substring. The specimen is repaired in this PR; a read of all ~486 substring calls in dag/ and src/v2 found no other site. Rung: no wall; trigger: a distinct octet-offset type for ByteRange. Co-Authored-By: Claude Opus 5.5 (1M context) --- ...med_by_a_scalar_indexed_text_operation.dag | 22 +++++++++++++++++++ 1 file changed, 22 insertions(+) create mode 100644 dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag diff --git a/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag b/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag new file mode 100644 index 00000000000..75352731e3b --- /dev/null +++ b/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag @@ -0,0 +1,22 @@ +module gunbc.recurring_failure_mode.octet_offset_consumed_by_a_scalar_indexed_text_operation + +import std.types { NonEmptyStr } +import std.decl_ref { DeclarationRef, WholeDeclaration } +import gunbc.recurring_failure_mode { RecurringFailureMode } + +data octet_offset_consumed_by_a_scalar_indexed_text_operation: RecurringFailureMode = RecurringFailureMode { + identity: "octet_offset_consumed_by_a_scalar_indexed_text_operation" as NonEmptyStr, + + receipts: [ + "INVALID STATE: an offset measured in UTF-8 OCTETS is handed to a text operation that indexes Unicode SCALARS. The v2 tokenizer counts every walk offset in octets (v2.compiler.tokenize lexeme_utf8_octet_count), so every v2.std.diagnostic ByteRange the span index resolves to is an octet range. `substring` and `string_length` index scalars: the emitted runtime walks chars(). Both are plain Int, so nothing stops an octet offset from reaching substring. On ASCII text the two agree, which is why the defect passes every ASCII fixture and is silent on the first multibyte character.", + "SPECIMEN, CAUGHT IN REVIEW BEFORE LANDING (gunbc#13051, calm-boar-904): v2.cli.if_arm_reader_differential derived each site's line as the count of line feeds in substring(source, 0, ByteRange.start). Every multibyte scalar before a site moved its reported line. The census COUNT was unaffected; only the rendered locations were wrong. Repaired by reporting the byte range alone, because no line-from-octet reader exists in the corpus. The control is a nested if after 2-, 3- and 4-octet scalars, located at octet 237, which is scalar 223.", + "POPULATION, read at the same revision over dag/ and src/v2/ (about 486 substring calls, each offset traced to its producer): no other site passes an octet offset to substring. Two neighbours look like the class but are not. src/v2/lens/text_string_importer_census.dag newline_offsets counts octets over utf8_encode_bytes and compares them with ByteRange.start: octets against octets. dag/std/import.dag import_strip_step slices with import-statement spans from the v1 tokenizer, which indexes its source_chars code-point list: scalars against scalars. One site agrees only by its input: dag/gunbc/runner/runner_microvm.dag jit_drive_content subtracts an octet count from a substring end over a string of line feeds only, where octets and scalars coincide.", + "RUNG FOUND AT: below the floor (a plausible, wrong location rendered with no refusal). RUNG NOW: no wall. The population is zero by a one-off read, not by any gate, so a new site can be written today. CEILING: structurally impossible. An octet offset and a scalar index are different quantities and can be different types. NEXT-RUNG TRIGGER: v2.std.diagnostic ByteRange carries a distinct octet-offset type that substring does not accept, with a sanctioned conversion that walks the UTF-8 table (std.unicode.types unicode_scalar_utf8_octet_count). A source-locus renderer that wants lines then derives them from octets through that one conversion.", + ], + + evidence: [ + DeclarationRef { module_path: "v2.test.claim.cli.if_arm_reader_differential", decl_name: "multibyte_source_before_a_site_is_located_in_octets_holds", field: WholeDeclaration }, + DeclarationRef { module_path: "v2.compiler.tokenize", decl_name: "lexeme_utf8_octet_count", field: WholeDeclaration }, + DeclarationRef { module_path: "v2.cli.if_arm_reader_differential", decl_name: "iard_location", field: WholeDeclaration }, + ], +} From 464b7dd6801509c7d640a0b23a7c3975a2c277fb Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Sat, 3 Oct 2026 08:05:22 +0000 Subject: [PATCH 9/9] if_arm_reader_differential, octet RFM row: the one line-from-octet reader was a lens's, not absent The earlier wording said no line-from-octet reader existed in the corpus. One did, private to v2.lens.text_string_importer_census (newline_offsets); swift-lynx-592 is moving it into v2.std.source_position. The census still reports octet ranges only until that authority lands. Co-Authored-By: Claude Opus 5.5 (1M context) --- ...t_consumed_by_a_scalar_indexed_text_operation.dag | 2 +- src/v2/cli/if_arm_reader_differential.dag | 12 +++++++----- 2 files changed, 8 insertions(+), 6 deletions(-) diff --git a/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag b/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag index 75352731e3b..05900421b0c 100644 --- a/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag +++ b/dag/gunbc/recurring_failure_mode/octet_offset_consumed_by_a_scalar_indexed_text_operation.dag @@ -9,7 +9,7 @@ data octet_offset_consumed_by_a_scalar_indexed_text_operation: RecurringFailureM receipts: [ "INVALID STATE: an offset measured in UTF-8 OCTETS is handed to a text operation that indexes Unicode SCALARS. The v2 tokenizer counts every walk offset in octets (v2.compiler.tokenize lexeme_utf8_octet_count), so every v2.std.diagnostic ByteRange the span index resolves to is an octet range. `substring` and `string_length` index scalars: the emitted runtime walks chars(). Both are plain Int, so nothing stops an octet offset from reaching substring. On ASCII text the two agree, which is why the defect passes every ASCII fixture and is silent on the first multibyte character.", - "SPECIMEN, CAUGHT IN REVIEW BEFORE LANDING (gunbc#13051, calm-boar-904): v2.cli.if_arm_reader_differential derived each site's line as the count of line feeds in substring(source, 0, ByteRange.start). Every multibyte scalar before a site moved its reported line. The census COUNT was unaffected; only the rendered locations were wrong. Repaired by reporting the byte range alone, because no line-from-octet reader exists in the corpus. The control is a nested if after 2-, 3- and 4-octet scalars, located at octet 237, which is scalar 223.", + "SPECIMEN, CAUGHT IN REVIEW BEFORE LANDING (gunbc#13051, calm-boar-904): v2.cli.if_arm_reader_differential derived each site's line as the count of line feeds in substring(source, 0, ByteRange.start). Every multibyte scalar before a site moved its reported line. The census COUNT was unaffected; only the rendered locations were wrong. Repaired by reporting the byte range alone. The corpus's one line-from-octet reader was private to a lens (v2.lens.text_string_importer_census newline_offsets), not a shared authority, and swift-lynx-592 is moving it into v2.std.source_position. When that lands it is the sanctioned conversion this row's trigger names. The control is a nested if after 2-, 3- and 4-octet scalars, located at octet 237, which is scalar 223.", "POPULATION, read at the same revision over dag/ and src/v2/ (about 486 substring calls, each offset traced to its producer): no other site passes an octet offset to substring. Two neighbours look like the class but are not. src/v2/lens/text_string_importer_census.dag newline_offsets counts octets over utf8_encode_bytes and compares them with ByteRange.start: octets against octets. dag/std/import.dag import_strip_step slices with import-statement spans from the v1 tokenizer, which indexes its source_chars code-point list: scalars against scalars. One site agrees only by its input: dag/gunbc/runner/runner_microvm.dag jit_drive_content subtracts an octet count from a substring end over a string of line feeds only, where octets and scalars coincide.", "RUNG FOUND AT: below the floor (a plausible, wrong location rendered with no refusal). RUNG NOW: no wall. The population is zero by a one-off read, not by any gate, so a new site can be written today. CEILING: structurally impossible. An octet offset and a scalar index are different quantities and can be different types. NEXT-RUNG TRIGGER: v2.std.diagnostic ByteRange carries a distinct octet-offset type that substring does not accept, with a sanctioned conversion that walks the UTF-8 table (std.unicode.types unicode_scalar_utf8_octet_count). A source-locus renderer that wants lines then derives them from octets through that one conversion.", ], diff --git a/src/v2/cli/if_arm_reader_differential.dag b/src/v2/cli/if_arm_reader_differential.dag index 2b651e0ecda..243dd232f63 100644 --- a/src/v2/cli/if_arm_reader_differential.dag +++ b/src/v2/cli/if_arm_reader_differential.dag @@ -325,11 +325,13 @@ fn iard_sites(tree: Node) -> FreeMonoid { // A NODE'S LOCATION, AS THE SPAN INDEX RESOLVES IT (v2.std.provenance is the one authority for // following a derived occurrence to its source): the file's byte range, in the UTF-8 octets the -// tokenizer counts (v2.compiler.tokenize lexeme_utf8_octet_count). NO LINE IS DERIVED HERE. No -// line-from-octet reader exists in the corpus, and deriving one with `substring` is wrong: the -// emitted substring indexes Unicode scalars, so every multibyte character before a site moved the -// line it reported (calm-boar-904, review of gunbc#13051). The byte range is the authority; a reader -// converts it to a line with a byte-aware tool. +// tokenizer counts (v2.compiler.tokenize lexeme_utf8_octet_count). NO LINE IS DERIVED HERE. The one +// line-from-octet reader in the corpus is private to a lens (v2.lens.text_string_importer_census +// newline_offsets), and swift-lynx-592 is moving it into v2.std.source_position as the shared +// authority. Deriving a line with `substring` instead is wrong: the emitted substring indexes Unicode +// scalars, so every multibyte character before a site moved the line it reported (calm-boar-904, +// review of gunbc#13051). Until that authority lands, the byte range is the answer. A line, if it +// returns, comes from that authority and from no second reader. fn iard_location(node: Node, spans: SpanIndex) -> String { match node_occurrence_id_optional(node: node) { Absent => ""