From 568e53ebf81afc9f8e5f727e312e51c140dc6fd0 Mon Sep 17 00:00:00 2001 From: gunbc-ci-auto-heal Date: Tue, 1 Sep 2026 23:39:02 +0000 Subject: [PATCH] Completeness is an identity join: the count comparison accepted a stale row standing in for an unrostered blob `test.claim.language_source_scaffold_index_test.compiler_tests_rust_blobs_are_all_rostered` was RED on main and ENROLLED, not unenrolled: it sits in `floor_expected_red_chunk_live_tree_admission` in `v2.workflow.floor_expected_red`, and its entry declares `ReadsLiveTree`, which in `entry_eligible_for_discovery_skip_before_resolve` means it never predict-skips. So it genuinely executed and genuinely failed every required run. THE GAP WAS NOT THE FOUR THE COUNTS SUGGESTED. The witness asserted `declared_fn_count(ct_) == rostered_count_for(...)`, 44 against 40. Joined by IDENTITY the residues are of two kinds: FIVE live blobs unrostered (ct_fixture_closure_rustc_discrimination_test, ct_import_lines_follow_resolved_binding_identity_test, ct_witness_carrier_declines_non_witness_expected_type_test, ct_generic_param_declines_fail_closed_unwrap_test, ct_shell_service_output_projection_known_hole_probe_test) and ONE row that outlived its blob -- ct_caret_parse_smoke_native_witness_tests, which #8532 deleted from the carrier while leaving the roster row standing. Rostering four of the five would have balanced the counts at 44 and GREENED the witness with one blob still unmarked and one row still naming a declaration that does not exist. The count was never the claim; it was a necessary condition of the claim being read as the claim. DESIGN section 5 already says this outright: completeness is an identity join, not a count equality. DISPOSITION. The five blobs are hand-authored Rust assertion blobs, the same class as their rostered neighbours, and carry `compiler_tests_rust_hand_assertion_scaffold_trigger`. The stale row is removed. THE CLAIM MADE HONEST. The witness now names the two residues separately -- declared-not-rostered (a blob landed unmarked) and rostered-not-declared (a row outlived its blob) -- and asserts each is empty, so the two directions red with distinct meanings and neither can pay for the other. Cardinality survives as a third conjunct answering the one question containment cannot, a DUPLICATE within one side. The `rt_` arm gets the same treatment. `head_before`'s unreachable Absent arm yields a spelling no roster row can carry rather than fabricating a plausible name: the failure arm refuses, it does not widen. THE EVIDENCE DOES NOT STOP AT THE REPAIRED TREE. A repaired population makes both live arms permanently green and the join indistinguishable from the count it replaced, so five discriminating controls run over authored fixtures, including the equal-counts-different-identities case that is exactly what the old form accepted. Delisted from the expected-red roster, since that roster self-empties on pass. Executed evidence, `gunbc run` against the live tree (BuildBuddy runners expose no cgroup memory limit, so `gunbc run` refuses there under `gunbc.host_budget_source`; this ran in the session container): all ten witnesses in the file return `true`, including compiler_tests_rust_blobs_are_all_rostered, and each of the four `the_join_refuses_*` controls returns `true`, i.e. actively refuses. DESIGN.md and docs/design-ledgers.md are regenerated through `dag/gunbc/instruments/generated_artifact_gate.dag main_wet`, carrying the new `compensating_errors_cancel_in_the_aggregate` recurring-failure-mode row. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01P7mphvNU1JoCbrowqDM5Zg --- DESIGN.md | 1 + dag/gunbc/language_source_scaffold_index.dag | 30 +++++- dag/gunbc/recurring_failure_mode.dag | 3 + .../language_source_scaffold_index_test.dag | 95 ++++++++++++++++--- docs/design-ledgers.md | 1 + src/v2/workflow/floor_expected_red.dag | 2 +- 6 files changed, 113 insertions(+), 19 deletions(-) diff --git a/DESIGN.md b/DESIGN.md index 5b4814c153f..c5c47f262e5 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -225,6 +225,7 @@ One row per class, each carrying its recognition rule and its receipts, in [docs - `review_summary_inverts_roles_and_affirms_the_join` - `accepted_source_emits_uncompilable_target` - `incidental_denominator_as_wall` +- `compensating_errors_cancel_in_the_aggregate` ## Building & checks diff --git a/dag/gunbc/language_source_scaffold_index.dag b/dag/gunbc/language_source_scaffold_index.dag index 2f9a1cdc83e..db0e37e71f5 100644 --- a/dag/gunbc/language_source_scaffold_index.dag +++ b/dag/gunbc/language_source_scaffold_index.dag @@ -394,8 +394,28 @@ data ct_row_ct_contracts_sidecar_witness_test: LanguageSourceScaffoldRow = Langu disposition: compiler_tests_rust_hand_assertion_scaffold_trigger } -data ct_row_ct_caret_parse_smoke_native_witness_tests: LanguageSourceScaffoldRow = LanguageSourceScaffoldRow { - carrier_module: "v1.compiler.compiler_tests_rust", blob_decl: "ct_caret_parse_smoke_native_witness_tests", +data ct_row_ct_fixture_closure_rustc_discrimination_test: LanguageSourceScaffoldRow = LanguageSourceScaffoldRow { + carrier_module: "v1.compiler.compiler_tests_rust", blob_decl: "ct_fixture_closure_rustc_discrimination_test", + disposition: compiler_tests_rust_hand_assertion_scaffold_trigger +} + +data ct_row_ct_import_lines_follow_resolved_binding_identity_test: LanguageSourceScaffoldRow = LanguageSourceScaffoldRow { + carrier_module: "v1.compiler.compiler_tests_rust", blob_decl: "ct_import_lines_follow_resolved_binding_identity_test", + disposition: compiler_tests_rust_hand_assertion_scaffold_trigger +} + +data ct_row_ct_witness_carrier_declines_non_witness_expected_type_test: LanguageSourceScaffoldRow = LanguageSourceScaffoldRow { + carrier_module: "v1.compiler.compiler_tests_rust", blob_decl: "ct_witness_carrier_declines_non_witness_expected_type_test", + disposition: compiler_tests_rust_hand_assertion_scaffold_trigger +} + +data ct_row_ct_generic_param_declines_fail_closed_unwrap_test: LanguageSourceScaffoldRow = LanguageSourceScaffoldRow { + carrier_module: "v1.compiler.compiler_tests_rust", blob_decl: "ct_generic_param_declines_fail_closed_unwrap_test", + disposition: compiler_tests_rust_hand_assertion_scaffold_trigger +} + +data ct_row_ct_shell_service_output_projection_known_hole_probe_test: LanguageSourceScaffoldRow = LanguageSourceScaffoldRow { + carrier_module: "v1.compiler.compiler_tests_rust", blob_decl: "ct_shell_service_output_projection_known_hole_probe_test", disposition: compiler_tests_rust_hand_assertion_scaffold_trigger } @@ -485,7 +505,11 @@ data language_source_scaffold_roster: List = [ ct_row_ct_constructor_call_admission_qualified_caller_test, ct_row_ct_sole_constructor_fieldless_witness_test, ct_row_ct_contracts_sidecar_witness_test, - ct_row_ct_caret_parse_smoke_native_witness_tests, + ct_row_ct_fixture_closure_rustc_discrimination_test, + ct_row_ct_import_lines_follow_resolved_binding_identity_test, + ct_row_ct_witness_carrier_declines_non_witness_expected_type_test, + ct_row_ct_generic_param_declines_fail_closed_unwrap_test, + ct_row_ct_shell_service_output_projection_known_hole_probe_test, ct_row_ct_call_shape_wall_witness_test, ct_row_ct_call_shape_duplicate_wall_witness_test, ct_row_ct_call_deficit_red_witness_test, diff --git a/dag/gunbc/recurring_failure_mode.dag b/dag/gunbc/recurring_failure_mode.dag index 527f1799650..12eb7ae7f58 100644 --- a/dag/gunbc/recurring_failure_mode.dag +++ b/dag/gunbc/recurring_failure_mode.dag @@ -169,6 +169,8 @@ data sealing_property_erases_structure: RecurringFailureMode = RecurringFailureM data censored_estimator_drops_its_own_tail: RecurringFailureMode = RecurringFailureMode { identity: "censored_estimator_drops_its_own_tail" as NonEmptyStr, authored: "**censored estimator drops its own tail** (an estimate of a variable is computed over the observations that SURVIVED a threshold on that same variable, so the sample structurally excludes its own extreme and the estimate is biased low with no bound. It is the neighbour of `instrument_output_read_as_subject_content` moved one step earlier: there a REPORT is complete for the reporter and read as complete for the consumer; here the truncation is in the ESTIMATOR'S OWN DOMAIN, and no amount of reading the instrument's output correctly recovers what the filter removed. SPECIMEN, both directions, 2026-09-01: the required floor's per-run cost inflation was estimated by pairing every claim that reported a cost in BOTH attempts of one tree -- median 1.053, p90 1.107. A claim that exceeds the cpu ceiling goes INTERRUPTED-BEFORE-VERDICT AND REPORTS NO COST, so the pairing dropped exactly the two most inflated observations; recovered from the ceiling itself, their inflation was at least 500/418 = 1.196, censored from above. THE SECOND HALF IS WHY THE ROW EXISTS: a peer lane, correcting a DIFFERENT defect in the same measurement, took the biased p90 by relay and derived an at-risk population of ONE -- an understatement arriving with the credibility of a retraction, which is the artifact nobody re-audits. Both counts were lower bounds and neither was labelled as one. RECOGNITION RULE: whenever an estimate is computed over rows that COMPLETED, RETURNED, PASSED, or were OBSERVED, ask what the incomplete rows would have contributed -- and if the incompleteness is caused by the very variable being estimated, the statistic is a floor and must be reported as one. The tell is a filter and an estimand naming the same quantity: cost estimated over rows that finished, latency over requests that did not time out, size over responses that were not truncated. REMEDY: report the censored bound rather than the sample statistic, and recover the excluded observations from the THRESHOLD they crossed, which is a real datum -- a row killed at 500ms is not missing, it is known to be above 500.)", evidence: [] } data restoration_promise_names_a_route_that_does_not_exist: RecurringFailureMode = RecurringFailureMode { identity: "restoration_promise_names_a_route_that_does_not_exist" as NonEmptyStr, authored: "**a restoration promise names a route that does not exist** (a mechanism withholds, defers or dormant-marks a population and tells the reader it becomes observable again under some named condition — and that condition names a RUN, LANE, CADENCE OR SWEEP the tree does not contain. Nothing refuses, because the promise is a string in a diagnostic rather than a citation anything resolves; the population is not silently dropped, which is what makes the class survive review — it is dropped WITH A RECEIPT, and the receipt is what stops anyone looking. **THE BOUNDARY AGAINST `unbacked_execution_claim` IS THE TENSE AND IT DECIDES THE REMEDY.** That class is prose asserting a relation that RUNS NOW; this is prose asserting a relation that WILL run — a future condition, so `git log --all -S` over the declaration form finds nothing to past-tense and the origin arms there do not apply. It is also not `absorbing_fallback`: nothing widens, the withhold is precise and correctly counted. It is 4b(3)'s trigger trap with the polarity inverted — there a trigger names LESS than the capability it restores and gets satisfied while the capability stays dead; here the trigger names a capability whose PRECONDITION IS ALREADY FALSE, so it can never be satisfied at all and the row waits forever in a state that reads as temporary. **SPECIMEN WITH A RECEIPT, measured 2026-09-01 on head 4c6c509e and unrepaired at authoring.** `v1_compiler.cli_run.required_floor_runner` `suppress_withheld` removes enrolled expected-red identities whose module sits outside the required gate, printing that their enrolment `becomes observable again when the gate roster admits the module or in the whole-corpus receipts run`. THAT RUN DOES NOT EXIST: the phrase occurs exactly once in the tree, inside the message that promises it, and the repository carries three workflows of which none is it. 39 identities across 23 modules sit under that promise. Each is enrolled on `v2.workflow.floor_expected_red`, whose own header states what an enrolment asserts — that the identity REACHES ITS SUBJECT AND ANSWERS, and that a row belongs there only while someone is fixing it — so every one of the 39 asserts `runs, fails, someone is fixing it` about a row no run reaches. The only surviving route by which one executes is the changed-witness override, i.e. somebody editing it. **THE POPULATION IS DERIVABLE AND IS DELIBERATELY NOT TRANSCRIBED HERE**: it is `v2.workflow.floor_expected_red` `floor_expected_red_roster` minus the identities whose module matches `v2.workflow.required_floor` `required_gate_prefixes` — two authorities and a set difference, so it re-derives instead of rotting. **THE SAME ROSTER'S HEADER ALREADY RECORDS THE ANCESTOR OF THIS MISTAKE**, which is why it is a class: 101 rows were held there as agreement while never reaching their subject, and were reclassified into `v2.workflow.floor_route_gap` on 2026-08-20 once `ExpectedRedArm` was taught to refuse `HostEffectRefused`, `HostToolUnresolved` and an interrupted budget. That repair closed the arm where a NON-VERDICT was read as agreement; this class is the same harm one step earlier, where a row never reaches an arm at all and a sentence promises it will. **RECOGNITION RULE: whenever a diagnostic says a withheld thing becomes observable again `in`/`under`/`by` some named run, grep the tree for that name and require an executing consumer — a workflow job, a scheduled entry point, an actuator argv. If the only occurrence is the promise itself, the population is dormant forever and the honest states are two: admit the row cannot be observed on any cadence, or delete the enrolment. **RUNG: 1 (mitigatable) — the withhold is counted and located, which is the whole of what holds. CEILING: 3, since `whether a named route exists` is decidable from the workflow and entry-point authorities the tree already carries. NEXT-RUNG TRIGGER, a CAPABILITY and not an artifact: restoration conditions expressed as a resolvable citation to an executing consumer rather than as prose, so a promise naming no route fails to compile — writing this particular sentence better retires nothing.)" as NonEmptyStr, evidence: [] } +data compensating_errors_cancel_in_the_aggregate: RecurringFailureMode = RecurringFailureMode { identity: "compensating_errors_cancel_in_the_aggregate" as NonEmptyStr, authored: "**compensating errors cancel in the aggregate** (a completeness check compares a COUNT on each side instead of joining the two populations by IDENTITY, so a member missing from one side and a phantom member on the other cancel exactly, and the check reports the coverage it was written to refuse. The green is not a near miss: it is the check confirming a roster that covers nothing of the sort, and it is louder the longer both defects stand, because each new member added to either side keeps the totals in step. Distinct from `executed_conjunct_discriminates_nothing`, where the conjunct can never go red: this one goes red readily on a SINGLE defect and is blind only to defects in opposing directions -- which is precisely the pair a long-lived roster accumulates, since the same neglect that leaves a new member unrostered leaves a deleted member rostered. Receipt (2026-09-01): `test.claim.language_source_scaffold_index_test.compiler_tests_rust_blobs_are_all_rostered` asserted `declared_fn_count(ct_) == rostered_count_for(v1.compiler.compiler_tests_rust)`. The tree carried 44 `ct_` blobs and 40 roster rows, so it was RED and enrolled in `v2.workflow.floor_expected_red` -- but the residues were not 4 of one kind. FIVE live blobs were unrostered and ONE row, `ct_caret_parse_smoke_native_witness_tests`, had outlived the blob #8532 deleted from the carrier. Rostering four of the five would have balanced the counts at 44 and greened the witness with one blob still unmarked and one row still naming a declaration that does not exist. The count was never the claim; it was a NECESSARY CONDITION of the claim being read as the claim. **Recognition rule: for any check whose sentence contains \"every\", \"all\" or \"covers\", ask what it would report if one member were missing from one side AND one phantom stood on the other. If the answer is \"green\", the check measures a cardinality and reports a coverage, and the two are joined by an assumption nothing enforces.** The remedy is the join: name the two residues separately -- present-not-covered and covered-not-present -- and assert each is empty, so the two directions red with distinct meanings and neither can pay for the other. Cardinality then survives as a third conjunct answering the one question containment cannot, a DUPLICATE within one side, and it is no longer the assertion. The evidence must not stop at the repaired tree, since a repaired population makes both live arms permanently green and the join indistinguishable from the count it replaced: the discriminating REDs belong over authored fixtures, with the equal-counts-different-identities case enrolled explicitly, because that is the one case the previous form accepted.", evidence: [] } + data incidental_denominator_as_wall: RecurringFailureMode = RecurringFailureMode { identity: "incidental_denominator_as_wall" as NonEmptyStr, authored: "**incidental denominator as a wall** (a destructive operation is safe only because an upstream filter written to answer a DIFFERENT question happens never to hand it the dangerous input. Nothing is wrong today and nothing declares why, so the safety is a coincidence with no authority, no diagnostic and no test — and it dissolves silently the moment the upstream question changes, which is a change nobody reads as touching safety. Distinct from `authority_substitution`, where PROSE invents an A-to-B arrow: here the CODE leans on an arrow that genuinely holds and that neither carrier claims, so grepping the guarded operation finds no guard to have gotten wrong. Receipt (2026-09-01): `v1_compiler.required_regen_host` `install_convergence_stage_with_backend` copies `candidate_src/` over `stage0_src/` for whatever roster it is handed, asking nothing about the artifact. The emitter really produces a bare `Cargo.toml` (`v1.compiler.emit_rust`); `write_emitted_tree` writes it into the candidate tree verbatim because it is not rustfmt-normalizable; `produce_candidate_manifest` rows it as a `GeneratedSurface` declared by that emitter. Every stage between the emitter and the copy admits it. The sole reason it never became an install target was `is_compared_generated_basename`, whose job is to denominate the DRIFT COMPARISON. **THE HARM CLAIM IS THE PART THIS ROW GOT WRONG FIRST, and the correction is the instructive half.** The finding was briefed, amplified through a manager and written into a work-item title as a destructive overwrite of the 172-line hand-maintained `src/v1/stage0/Cargo.toml`. It is not: the installer joins destinations under `stage0_src = workspace/src/v1/stage0/src`, so a bare `Cargo.toml` basename resolves one directory BELOW the manifest, to a path that does not exist. Widening the population alone writes debris rather than destruction ONLY IF the join is the whole story, and THAT CLAIM WAS ITSELF AN UNDER-ENUMERATION, caught in review of the repair: `fs::copy` follows the DESTINATION link, git tracks symlinks, so a committed stage0 entry that is a symlink reaches any path in the tree with the join untouched. The lesson generalizes past this row: a harm correction that names ONE mechanism and treats it as the requirement is the same coincidence-as-wall shape turned on the analysis -- the honest form enumerates the ROUTES to the destination and says which are closed, since `resolves outside X` is a property of the filesystem and not of the path expression. Two readers reached the destructive reading independently and agreement read as confirmation, which is why the deciding observation is a JOIN EXPRESSION and not a semantic comparison of the two files — the artifacts genuinely differ, and that true fact was fused with a destination claim it never established. **The surviving harm is exemption, not destruction, and it is the one worth naming:** a declared `GeneratedSurface` that the comparison filter hides is invisible to the fixed point (`first_generation_equal`, `changed_paths`) and unreachable by the one authority that would force its gap to be counted — `v2.compiler.self_host.stage0_crate_layout` `emitter_produced_divergent_registrations`, which already enforces in three directions and holds `main.rs` for exactly this relation. So the filter did not accidentally PROTECT the artifact; it accidentally EXEMPTED it. Beside it sat a second, CALLER-LESS installer, `install_candidate_paths`, equally unguarded and invisible to review because a crate-wide dead-code allow kept it compiling: an unguarded operation with no consumer still reads as sanctioned precedent for the next author. **Recognition rule: for every operation that mutates authoritative bytes, ask what it would do if handed the worst member of the population its input is DRAWN FROM — not the population it is called with today — then name the function that forbids that member and the question that function was written to answer. If those two questions differ, the wall is incidental and the operation has none.** The remedy is an admission at the mutation boundary itself, over the whole roster before the first byte moves and refusing rather than skipping, since a per-item skip re-exports the same silence one layer in; and it is honestly a mechanically-preventable rung, whose ceiling is structural impossibility once the plan projects a typed surface identity instead of a `List`. **THE REMEDY HAS ITS OWN TRAP, AND THIS ROW WALKED INTO IT BEFORE CATCHING IT: restating the incidental denominator as the guard's own rule.** The first repair refused every non-`.rs` install target, which reads as a wall and is the SAME defect dualised — the accidental denominator promoted from coincidence to policy, in a second location, now refusing the correct end state by construction, since the emitted manifest is stage0's own and is supposed to converge. **Recognition rule for the repair: state the arm's reason, and check whether it forbids the state the system is trying to REACH.** If it does, the guard has ratified the coincidence instead of replacing it, and the honest arms are the ones the boundary can answer from its own facts — here a path that cannot address its own destination, and a hand-authored mirror — while the artifact-kind question goes back to the authority that owns dispositions.)", evidence: [] } data restored_bytes_reviewed_as_authorship: RecurringFailureMode = RecurringFailureMode { identity: "restored_bytes_reviewed_as_authorship" as NonEmptyStr, authored: "**restored bytes reviewed as authorship** (a clobber repair returns content a prior PR already merged and a prior reviewer already accepted, but a diff has no way to say so: the restored lines are ADDED LINES, so a reviewer attributes them to the repairing PR and reviews them as new work. Specimen: gunbc#9937 restored two rows that gunbc#9868 had reverted off a stale base, and review 58203 filed REQUEST_CHANGES against a DESIGN section 6 transcribed-count defect in bytes that are character-for-character gunbc#9888's merged row -- a finding correct on its merits and misaddressed by one PR. **THE SECOND-ORDER EFFECT IS WHY THIS IS A CLASS AND NOT AN ANECDOTE: it makes clobber repairs EXPENSIVE TO LAND, which biases the fleet toward leaving clobbers standing.** The repairing lane inherits review debt for content it did not write, and the only way to discharge it inside the repair is to edit another lane's cells from a base lacking their context -- which is precisely the failure the repair exists to undo, committed a second time with better manners. RECOGNITION RULE: a diff whose added lines are byte-identical to a prior merged blob. It is DECIDABLE rather than a judgement call -- `cmp` against `git show :` answers it, and provenance for the whole region comes from `git log -S` naming the commit that added the text and the commit that removed it. THE REMEDY IS THE AUTHOR'S AND IT IS CHEAP: state the restoration ahead of the diff and carry the byte-identity receipt, so a reviewer meets provenance before content; without it the same finding is re-filed against the same bytes on every re-review, and the loop is paid for twice. A finding against restored bytes is well-formed and BINDS THE DESTINATION: provenance decides who answers it and in which diff, and never whether it is answered. **THE WRONG READING OF THIS ROW WAS AUTHORED AND REFUSED IN ITS OWN FILING** (review 58220): an earlier revision said such a finding 'belongs in its own diff', which reads as licence to merge a known defect behind an authorship argument -- DESIGN section 5's hard reject, and inverted, since a restoration is exactly the motion that puts the bytes back on main. What actually happened on the specimen is the rule: the reviewer's section 6 finding was correct, the repairing lane argued for routing it elsewhere, the owning manager OVERRULED that, and the defect was repaired inside the restoration under a ruling -- which is what made the edit a decision rather than one lane rewriting another's cells. Routing is a question about AUTHORITY and context, decided by whoever owns both lanes; the bytes are reviewed at their destination either way. **Rung: mitigatable** -- the author states the restoration and carries a byte-identity receipt, and a reviewer may still miss it. **Ceiling: mechanically preventable**, because the recognition rule is decidable: a tool that joins each added line against the merged blobs of the commits git log -S names can annotate a diff as restored before a human reads it. It is not structural -- nothing can make a restoration unwritable, nor should it. **Next-rung trigger, at capability grain: the review surface distinguishes restored bytes from authored bytes without the author asserting it**, since an author-supplied notice is exactly the evidence a tired or adversarial diff will omit. **THE RECEIPT MUST BE TAKEN AT THE DESTINATION AND NOT AT THE SOURCE, and that is the transferable half.** `cmp` decides the git case because it reads the bytes where they LANDED, not the bytes an author meant to send; the general shape is ANY step between assertion and delivery altering content while the delivery reports success -- and the step is not always the one a reader would blame first. Measured the same day in three places, and the THIRD one is a caution about naming the culprit too early: git's merge region dropped a row outside the conflict markers; a review read merged bytes as authored ones; and two sessions' messages arrived with a backticked identifier deleted and a dollar-variable replaced by a plausible shell path, which BOTH of them first attributed to the message transport. That attribution was wrong, and it was settled by measurement rather than by reasoning: a third session kept the local source of a sent message, read the delivered body back out of the store, and diffed them -- identical but for a trailing newline, with backticks, a backslash continuation and a dollar-brace token all intact. The corruption was in the SENDER'S OWN SHELL, expanding an inline double-quoted argument before the tool ever saw it, and the proposed remedy of banning backticks would have cost a real notation to avoid a hazard that lives in a construction nobody has to use. A deletion leaves a gap a reader might notice; a substitution leaves a syntactically fine claim that is simply false, which is why source-side confidence is not evidence at all.)", evidence: [] } @@ -228,4 +230,5 @@ data recurring_failure_mode_roster: List = [ review_summary_inverts_roles_and_affirms_the_join, accepted_source_emits_uncompilable_target, incidental_denominator_as_wall, + compensating_errors_cancel_in_the_aggregate, ] diff --git a/dag/test/claim/language_source_scaffold_index_test.dag b/dag/test/claim/language_source_scaffold_index_test.dag index 627da8b207d..5b50cd3c291 100644 --- a/dag/test/claim/language_source_scaffold_index_test.dag +++ b/dag/test/claim/language_source_scaffold_index_test.dag @@ -9,7 +9,6 @@ import v2.std.live_tree { LiveTreeDisposition, ReadsLiveTree } import gunbc.language_source_scaffold_index { LanguageSourceScaffoldRow, language_source_scaffold_roster, - language_source_scaffold_roster_size, language_source_scaffold_roster_all_dispositioned, language_source_scaffold_row_is_dispositioned, rust_pair_completion_spelling_scaffold_trigger @@ -29,9 +28,22 @@ data live_tree_disposition: LiveTreeDisposition = ReadsLiveTree // language_source_scaffold_row_is_dispositioned. (2) COVERAGE: the roster must actually cover its // carriers. Asserting hardcoded census counts alone would let a new rt_/ct_ blob land unmarked and // stay green until someone bumped the number by hand (review 43161), so coverage is read from the -// LIVE TREE: the rt_ and ct_ declaration counts in the carrier sources are compared against the -// roster, and a newly added blob reds this witness instead of sitting unrostered. The comparison is -// EXACT equality with no offset: an earlier version used declared == rostered + 1 to absorb +// LIVE TREE: both populations — the rt_/ct_ declarations in the carrier sources and the roster rows +// naming that carrier — are derived on every run, and neither side is a literal. +// +// COVERAGE IS AN IDENTITY JOIN, NOT A COUNT EQUALITY (DESIGN §5). The earlier form compared +// declared COUNT to rostered COUNT, and main proved that weaker: the roster carried +// ct_caret_parse_smoke_native_witness_tests, a blob #8532 had deleted from the carrier, while five +// live blobs sat unrostered. A count comparison cannot tell "44 declared, 44 rostered, same names" +// from "44 declared, 44 rostered, one stale row standing in for one unrostered blob" — the stale +// row would have MASKED an unrostered blob and this witness would have gone green on a roster that +// covered nothing of the sort. So the assertion now names the two residues: declared-not-rostered +// (a blob landed unmarked) and rostered-not-declared (a row outlived its blob). Both must be empty, +// and each reds with a distinct meaning. The residual count conjunct is not a change detector — +// with both containments holding it can only fail on a DUPLICATE name within one population, which +// is the one defect identity containment alone cannot see. +// +// The comparison carries no offset: an earlier version used declared == rostered + 1 to absorb // ct_coercion_tests, which was itself an absorbing fallback (review 43177) — that blob is a hybrid // (row-driven core via extract_coercion_tests, but it still emits a hand-authored header and // aggregates three hand blobs), so it is rostered like any other and the fudge is gone. @@ -45,13 +57,66 @@ data legitimate_terminal_control_row: LanguageSourceScaffoldRow = LanguageSource disposition: Terminal { reason: "named irreducible intrinsic kernel" } } -fn rostered_count_for(carrier: String) -> Int { - language_source_scaffold_roster |> filter(r => r.carrier_module == carrier) |> count +fn rostered_names_for(carrier: String) -> List { + language_source_scaffold_roster |> filter(r => r.carrier_module == carrier) |> map(r => r.blob_decl) } -fn declared_fn_count(path: String, prefix: String) -> Int { +// The head of a declaration segment, up to its parameter list. `split` on a present delimiter +// always yields at least one element, so Absent is unreachable here — but the arm may not fabricate +// a plausible name (DESIGN §5), so it yields a spelling no roster row can carry and the join reds. +// The failure arm refuses; it does not widen. +fn head_before(s: String, delimiter: String) -> String { + match s |> split(delimiter: delimiter) |> first { + Present { value: v } => v + Absent => "" + } +} + +fn declared_blob_names(path: String, decl_prefix: String) -> List { let src = filesystem_read(path: path) - (src.content |> split(delimiter: prefix) |> count) - 1 + src.content + |> split(delimiter: concat("\nfn ", decl_prefix)) + |> skip(1) + |> map(seg => concat(decl_prefix, head_before(s: seg, delimiter: "("))) +} + +fn names_absent_from(xs: List, ys: List) -> List { + xs |> filter(x => (ys |> filter(y => y == x) |> count) == 0) +} + +fn identity_join_holds(declared: List, rostered: List) -> Bool { + (names_absent_from(xs: declared, ys: rostered) |> count) == 0 + && (names_absent_from(xs: rostered, ys: declared) |> count) == 0 + && (declared |> count) == (rostered |> count) +} + +// THE DISCRIMINATING CONTROL FOR THE JOIN ITSELF, over authored fixtures rather than the live tree +// (DESIGN §4b(1): a rung is established by an executed RED plus an accepted positive control, and +// the live-tree arms cannot supply the RED once the tree is repaired). The middle case is the one +// main actually shipped: EQUAL COUNTS, one stale row standing in for one unrostered blob. A count +// comparison accepts it; this join refuses it, and refuses it from BOTH directions separately, so a +// repair that checked only one containment would go red here. + +data control_declared: List = ["ct_a", "ct_b", "ct_c"] + +test fn the_join_refuses_equal_counts_with_different_identities() -> Bool { + !identity_join_holds(declared: control_declared, rostered: ["ct_a", "ct_b", "ct_stale"]) +} + +test fn the_join_refuses_an_unrostered_blob() -> Bool { + !identity_join_holds(declared: control_declared, rostered: ["ct_a", "ct_b"]) +} + +test fn the_join_refuses_a_row_that_outlived_its_blob() -> Bool { + !identity_join_holds(declared: control_declared, rostered: ["ct_a", "ct_b", "ct_c", "ct_stale"]) +} + +test fn the_join_refuses_a_duplicate_row_masking_an_unrostered_blob() -> Bool { + !identity_join_holds(declared: control_declared, rostered: ["ct_a", "ct_a", "ct_b", "ct_c"]) +} + +test fn the_join_accepts_the_same_population_in_any_order() -> Bool { + identity_join_holds(declared: control_declared, rostered: ["ct_c", "ct_a", "ct_b"]) } test fn language_source_scaffold_roster_is_fully_dispositioned() -> Bool { @@ -59,15 +124,15 @@ test fn language_source_scaffold_roster_is_fully_dispositioned() -> Bool { } test fn runtime_rust_blobs_are_all_rostered() -> Bool { - let declared = declared_fn_count(path: runtime_rust_source_path, prefix: "\nfn rt_") - let rostered = rostered_count_for(carrier: "v1.compiler.runtime_rust") - declared == rostered + let declared = declared_blob_names(path: runtime_rust_source_path, decl_prefix: "rt_") + let rostered = rostered_names_for(carrier: "v1.compiler.runtime_rust") + identity_join_holds(declared: declared, rostered: rostered) } test fn compiler_tests_rust_blobs_are_all_rostered() -> Bool { - let declared = declared_fn_count(path: compiler_tests_rust_source_path, prefix: "\nfn ct_") - let rostered = rostered_count_for(carrier: "v1.compiler.compiler_tests_rust") - declared == rostered + let declared = declared_blob_names(path: compiler_tests_rust_source_path, decl_prefix: "ct_") + let rostered = rostered_names_for(carrier: "v1.compiler.compiler_tests_rust") + identity_join_holds(declared: declared, rostered: rostered) } test fn reasoned_terminal_is_accepted() -> Bool { @@ -86,4 +151,4 @@ test fn pair_completion_spelling_binds_the_derivation_authority() -> Bool { } } -data coverage_scan_dissolve_on: DissolutionCondition = unbound_dissolution(description: "dissolve-on: declared_fn_count below. It counts declarations by splitting the carrier SOURCE TEXT on a fn-name prefix, which is string-shape scanning and is anemic against the Node tree the substrate already has — the same class of debt as pair_completion_uses_rhs, and marked the same way rather than left implicit (review 43195). It is a deliberate interim, strictly better than the hardcoded censu") +data coverage_scan_dissolve_on: DissolutionCondition = unbound_dissolution(description: "dissolve-on: declared_blob_names below. It reads declaration IDENTITIES by splitting the carrier SOURCE TEXT on a fn-name prefix, which is string-shape scanning and is anemic against the Node tree the substrate already has — the same class of debt as pair_completion_uses_rhs, and marked the same way rather than left implicit (review 43195). It is a deliberate interim, strictly better than the hardcoded censu") diff --git a/docs/design-ledgers.md b/docs/design-ledgers.md index dd96a7e5f98..005e6c8c76c 100644 --- a/docs/design-ledgers.md +++ b/docs/design-ledgers.md @@ -53,6 +53,7 @@ The landing measurement partitions the 31 parser-visible identities into **2 cit - **a review summary that inverts real roles and then affirms the join** (a review names files that all exist, assigns each the wrong role, and closes by asserting the one property it was uniquely positioned to test. Specimen: review 58215 approved gunbc#9937 at `fafd2adc7e1` with *'restoration of two ledger rows to recurring_failure_mode.dag ... plus an accompanying prose update in the gap-analysis doc ... Diff is narrow and matches the PR title.'* The carrier held ONE added row, not two; that row was a NEW class row, not a restoration of anything; the actual restoration was the docs file it demoted to 'accompanying'; and the diff did NOT match the title, because the `.dag` file had been swept in by a staged-change leak the PR body never declared. **DISTINGUISH IT FROM A HALLUCINATED SUBJECT, which is the neighbouring failure and a much easier one:** review 58070 described a `match` arm that appeared nowhere in its diff, and an invention is refutable by grep. Here EVERY NOUN IS PRESENT IN THE TREE and only the relations between them are wrong, so the usual defence -- check that the cited things exist -- returns all-clear. **THE AFFIRMATION IS THE HARM AND NOT THE MISDESCRIPTION.** A lane cannot audit its own title-versus-diff agreement with fresh eyes; a reviewer can, and this one asserted that agreement in the exact PR where it failed. A review that stays silent on a property leaves the check undone; a review that affirms it wrongly marks the check DONE, which is worse and is why this is not merely a low-quality summary. RECOGNITION RULE, and it is decidable rather than a matter of reading care: join the changed-FILE set against what the title and body claim to change, at file grain, and refuse any file the body does not account for. That join is mechanical, needs no understanding of the diff, and would have fired here. **THE BACKSTOP QUESTION IS THE SECOND HALF OF THE ROW, and its answer is honest rather than comforting:** the leaked roster row carried no regenerated `DESIGN.md` or `docs/design-ledgers.md`, and by construction the required build lane's generated-artifact phase compares every `CommitRequired` projection against its authority, so it WOULD have refused. It did not get the chance -- that lane was CANCELLED on the leaked sha when the next push superseded it, and the leak was caught by the author re-reading the diff. So the backstop is real and was not what caught it, and a row claiming the gate held here would be asserting an execution that never ran. **Rung: mitigatable** -- the harm is contained only by a second reader noticing, and on the specimen the second reader was the author. **Ceiling: mechanically preventable**, and no higher: whether a summary DESCRIBES a diff correctly is not decidable, so no construction forbids the bad summary -- but the one clause that did the damage is decidable, because a changed-file set and a body are both machine-readable. **Next-rung trigger, at capability grain: no review verdict is admissible unless a changed-FILE-set join against the declared scope has executed and passed** -- not a reviewer instructed to check it, which is the same assertion that failed here, and not a lint on summary wording, which would ratchet the prose while leaving the join undone.) - **accepted source emits uncompilable target** (INVALID STATE: a .dag construction the front end accepts with zero blocking diagnostics, whose emission is a target program the target compiler refuses. The instance filed here is a coproduct's UNIT VARIANT standing in a NON-APPLIED TYPE POSITION -- a field, parameter or return type spelled with a constructor rather than a type. HARM: section 5 silent wrongness on the source-to-target path. gunbc says Accepted and hands over a program that cannot build, so the only wall that fires belongs to the target's compiler; where a target has no such wall, or where the emitted artifact is never compiled, nothing fires at all. Section 7 makes this the seed's floor and not the target's problem: the .dag graph is the authority and Rust is one realization, so 'rustc catches it' is exactly the outsourcing this project exists to end. THE SPECIMEN IS CARRIED IN THIS ROW RATHER THAN CITED, because the fixture that produced it was never committed to this repository (neat-otter-332, 2026-09-01). Its emitted artifacts are held by that lane's own receipts; no path is given here, since a path outside this repository is not a citation a later reader can resolve. Authored source, two modules: `scope.provider` declares `type Quantity = Time | Memory`; `scope.consumer` does `import scope.provider { Quantity, Time }`, `type NonApplied { value: Time }`, `fn field_as_quantity(subject: NonApplied) -> Quantity { subject.value }`. gunbc accepts it. DISTINGUISHING FACT, and the reason this class must not be read off the target's error code: WHAT rustc says is decided by what else the emitter minted, not by the source defect. Under the NARROW emitter classifier the consumer emits `pub use crate::scope_provider::{Quantity};` then `use crate::scope_provider::Quantity::{Time};` with no marker binding, and `pub struct NonApplied { pub value: Time }` refuses `error[E0573]: expected type, found variant Time` at `src/scope_consumer.rs:15:16` on `pub value: Time,`, label `not a type`, help `consider importing this struct instead`. Under the BROAD classifier the consumer instead emits `pub use crate::scope_provider::{Time};` beside `{Quantity}` -- the provider having emitted `pub enum Quantity { Time, Memory }` AND `pub struct Time;` -- so the field declaration COMPILES, and the refusal moves to the function: `error[E0308]: mismatched types` at `src/scope_consumer.rs:20:5` on `subject.value.clone()`, `expected Quantity, found Time`, against `expected Quantity because of return type`. INDEPENDENTLY REPRODUCED AT PARAMETER POSITION on this branch, 2026-09-01 on gunbc baeabbbf80, by a single-file source handed to the compiler: `type StampMode = StampClass | StampOther` with `fn take(stamp: StampClass) -> StampMode { stamp }`. `gunbc compile --target rust` exits 0 with 0 blocking errors, emits `pub fn take(stamp: StampClass) -> StampMode` beside `pub struct StampClass;`, and `cargo check` on the emitted crate refuses `E0308 expected StampMode, found StampClass`. One source defect; E0573, E0308-at-the-parent, and E0308-at-the-body depending on emission. THE MASK IS THE SHARP HALF, and it is fabricated plausible output rather than a lucky green: the broad classifier's marker import made the FIELD-ONLY source compile by substituting a distinct struct for a variant. It never preserved the modelled meaning, and extending the SAME accepted source across its declared parent boundary -- the function returning `Quantity` -- is what exposes it. So the narrow classifier's E0573 is this class becoming VISIBLE, not a regression the narrowing introduced. GENERAL FORM: a green obtained because the emitter manufactured a target-only entity for a source name is not evidence the source is well-typed, and the discriminator is to extend the source past the boundary the manufactured entity does not model. THE ONE .dag-SIDE SIGNAL IS ABOUT THE WRONG QUESTION: on the parameter reproduction the compile printed the ADVISORY `unlisted import use 'StampClass' (referenced but not in any import's name list)`. It fires because the name was not found among types -- the front end reached the exact fact that decides this case and reported it as import hygiene. It is not a partial wall: it is advisory, it names listing rather than type position, and the field specimen above IMPORTS `Time` explicitly, so it does not fire there at all (see `diagnostic_name_mechanism_silent`). RUNG FOUND AT: below the ladder, established by execution at the emission boundary -- the compile accepts and the emitted crate refuses. Section 4b's rung-1 mitigations are absent: nothing at the .dag boundary is typed, located or countable about this construction. CEILING: 4, structurally impossible, and the reason is that constructor identity and type identity are two distinct modelled facts whose membership is decidable from the coproduct declaration -- a variant name is reachable from that declaration as an ARM and never as a type, so the type-position slot has no constructor that admits it. No undecidable predicate is involved, so this is a wall now and not a ratchet. NEXT-RUNG TRIGGER, phrased as the capability that retires the row rather than an artifact that would contribute to one: type-position name resolution that consults the TYPE namespace alone and refuses a name resolving to a constructor with a located diagnostic, sufficient that NO Accepted program contains a variant name in any non-applied type position -- field, parameter or return. Repairing either specimen, narrowing the emitter classifier, or adding a fixture satisfies less than that and does not retire this row. RECOGNITION RULE: when the target compiler names a symbol at a type position, check whether that symbol is declared as an ARM of a coproduct in the source; if it is, the defect is in accepted .dag and the target compiler is the only wall that fired. SCOPE STATED RATHER THAN GENERALISED: executed for a unit arm at a field type and at a parameter type, Rust target only. Record-shaped arms, applied positions such as `List