RFM: classify the three residual reds (single_arm_match regressed by #13056) - #13327
Conversation
…match regressed by #13056; cwc and accumulator rostered) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head 01234d88492b4eff7c7e1c042b98fbcafe315b4f, against DESIGN.md §§3, 4b(1)/(3), 5 and 6b. These are corrections to the RFM classification and proposed restoration conditions, NOT a request to move the separately dispatched compiler/lens repairs into this PR.
The reported single_arm_match bisect is useful and the CWC contradictory pair is present in source: the two claims return P and !P for the identical input. Keeping the CWC resolve cause unbisected and the comma-versus-newline discriminator unestablished is appropriate. Two load-bearing statements in the new rows do not match their owning interfaces.
accumulator_copy_lens_binds_no_carrierconflates 'no Poly2Suspect' with 'clean/no finding', and its restoration trigger can already be satisfied by the broken path.
At this exact head, v2.lens.complexity_accumulator_copy.analyze::analyze_lowered_loop handles an absent carrier by emitting:
Unclassifiable { at: node, cause: ^fold_accumulator_unread }
The surface-fold absent-carrier branch in analyze does likewise. accumulator_copy_report retains the findings, and finding_is_suspect deliberately returns false for Unclassifiable while finding_is_refusal returns true. Consequently a failed has-suspect assertion does NOT establish that the lens answered clean. The row's sentence 'A lens that cannot bind answers clean' and silent-wrongness classification need either correction to a loss of supported classification with a refusal, or an identified downstream consumer that actually discards that refusal. Do not assign that latter defect to the lens without the consumer evidence.
The same row says every clean control passes, but its own five-pass roster excludes let_fresh_binding_is_provably_clean, an existing clean-control test in the named module. Preserve the actual mixed population rather than using the blanket claim to infer one mechanism for all 19 failures. report_counters_state_the_domain is also a conjunction over two different snippets, not a standalone observation that carriers_bound is zero.
Most importantly, 'binds the carrier ... OR refuses that Loop by name' is not a sufficient restoration trigger for the lost positive classification: the current absent-carrier branch already refuses that Loop by name. Require restoration of binding and discriminating classification for the supported inline-fold positives, while retaining the distinct clean and unsupported/refused controls. Narrow unresolved causal claims rather than requiring a bisect to land this record.
- The cost-row recommendation reintroduces a meaning the debt authority expressly rejects.
v2.workflow.floor_cost_debt distinguishes PROVEN passing/over-cost evidence from CENSORED interruption before verdict. Its header explicitly records and retracts the former universal claim that every enrolled witness passes. This PR nevertheless states 'A cost row asserts the claim would pass given time.' That is false for the censored chunk implicated here.
The useful proposed policy is narrower: an observed semantic FAIL must not remain represented solely as a cost-only/censored disposition; an interrupted run still establishes no semantic verdict. State that without promoting unknown to passing. A run at admission and restoration supplies evidence at those revisions, but cannot catch a regression introduced while the row remains withheld—the very temporal gap this incident illustrates. Mark that proposal's coverage as snapshot-limited and retain the existing broader execution/restoration obligation, rather than implying these two checkpoints prevent the intervening masked regression. No new gate implementation is requested in this RFM-only cut.
Reviewed the complete four-file diff, exact-head DESIGN, cost-debt authority, witness predicates and analyzer paths. The PR's execution/bisect numbers remain author-reported receipts; I did not execute local tests or reproduce the bisect. Exact-head witnesses run 37245335040 was still in progress at review.
…ficient trigger; cost-row policy narrowed (review) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
briansrls
left a comment
There was a problem hiding this comment.
APPROVE at exact head ff46b2a02eb094825321166d87d74974cc3e9c44, against DESIGN.md §§3, 4b(1)/(3)/(4), 5 and 6b. Both blockers from review 5408873037 are closed.
Compared the complete one-commit, two-file correction from 01234d88492b4eff7c7e1c042b98fbcafe315b4f and re-read the full four-file PR diff. DESIGN.md, the analyzer, witness predicates and cost-debt authority are unchanged from the prior review.
-
The accumulator row now distinguishes absence of Poly2Suspect from a clean result. It names the existing Unclassifiable/fold_accumulator_unread refusal, does not invent a downstream consumer that discards it, and records the 19 failures as a mixed population including positive, refusal, clean and conjunctive controls. The carrier-binding explanation and #13028 attribution remain hypotheses, not a measured common cause for every failure. The restoration trigger now requires carrier binding AND Poly2Suspect for the supported inline-fold positives, preserving clean and unsupported/refusal controls. The already-existing named refusal can no longer satisfy that trigger.
-
The cost-row proposal now preserves the authority's PROVEN-versus-CENSORED distinction: interruption supplies no semantic verdict, and an observed FAIL must not remain represented solely as cost/censored debt. Admission/restoration checks are explicitly snapshot-limited and supplement, rather than replace, the existing execution/restoration obligation. No implemented gate or continuous coverage is claimed.
The unchanged single_arm_match bisect report and CWC diagnosis retain their prior disposition; this approval does not independently reproduce those measurements or establish the unbisected causes. The PR records the residual defects and restoration requirements only. It does not repair the compiler/lens, discharge the outstanding controls, or approve a new gate implementation; those remain separate work.
Source review only; no local tests or bisect runs performed by this reviewer. Exact-head witnesses run 37246356964 is still in progress. No semantic blocker remains on this RFM-only change; land after the required exact-head checks pass.
Fifth PR of the floor-unimported-bare-provider triage: the three residual reds, classified. RFM rows only; no code changes and no repairs.
v2.test.claim.body_lowering.single_arm_matchtwo_arm_match_still_accepts_holdsandthree_arm_match_keeps_every_arm_holdsfail from fd30a50. The two-arm source parses and normalizes Accepted;conserved_normalizerejects on a dropped atom reference.match_arms_after_the_second_are_dropped_at_v2_body_lowering, whose 4b(4) evidence is now redv2.test.long.cwc_resolution_probecwc_cwc_module_resolve_acceptsis a regression candidate (lve_anonymous_record_expected_type_not_record), not bisectedanonymous_record_literal_refused_against_a_declared_record_typev2.test.long.accumulator_copy_fold_analysisUnclassifiable) rather than answering clean. The 19 failures are a mixed population; not bisectedaccumulator_copy_lens_binds_no_carrierMasked failure: the single_arm_match identities sit in a CENSORED
floor_cost_debtchunk, which asserts no verdict. Proposal: an observed semantic FAIL must not stay represented only as a cost/censored row. A check at admission and restoration is snapshot-limited, so it supplements the existing execution obligation rather than replacing it.Instrument:
claim_batchover every test fn, run on BuildBuddy withGUNBC_MEMORY_BUDGET_BYTESforwarded. The diagnostic reasons came from temporary probe claims, which are not in this diff.🤖 Generated with Claude Code