Skip to content

Floor cost w2: shrink wrap_decision_predicate fixtures (1.21M steps) - #10539

Merged
briansrls merged 3 commits into
mainfrom
session/sleek-koi-592
Sep 5, 2026
Merged

briansrls merged 3 commits into
mainfrom
session/sleek-koi-592

Conversation

@briansrls

@briansrls briansrls commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

The fifteen claims in v2.test.claim.wrap_decision_predicate demanded rust_sg2_type_expr_target_model(). Each is pointed at a fixture carrying what it actually reads.

Measured

required_floor_claim_cost.tsv, before = main run 33954515928, after = run 33956529160:

claim before after
wrap_decision_instantiation_arg_is_wrapped 88,920 18,439 −79.3%
wrap_decision_flow_over_wrap_direction_refuses 82,780 11,593 −86.0%
wrap_decision_flow_under_wrap_direction_refuses 82,780 11,593 −86.0%
wrap_decision_flow_distinct_reference_layers_refuse 82,775 11,588 −86.0%
wrap_decision_flow_agreeing_sites_accept 82,674 11,487 −86.1%
wrap_decision_flow_missing_row_refuses 82,672 11,485 −86.1%
wrap_decision_diagnostics_return_is_rc 81,408 10,221 −87.4%
wrap_decision_node_struct_field_is_box 81,400 10,213 −87.5%
wrap_decision_diagnostics_param_is_owned 81,398 10,211 −87.5%
wrap_decision_use_site_verdict_return_is_owned 81,362 10,175 −87.5%
wrap_decision_use_site_verdict_param_is_owned 81,360 10,173 −87.5%
wrap_decision_probe_heap_param_miss_rejects 81,333 10,146 −87.5%
wrap_decision_bundle_partial_rejects 73,568 2,578 −96.5%
wrap_decision_flow_bundle_absent_inapplicable 73,565 192 −99.7%
wrap_decision_bundle_absent_inapplicable 73,512 139 −99.8%
total 1,211,507 140,233 −88.4%

Corpus reconciliation: 17 identities changed beyond ±9, summing −1,071,238 against a corpus delta of −1,071,224; added=0, dropped=0 — no claim silently deleted or activated.

The finding worth carrying to other lanes

The bottom three rows already ran against an empty or single-edge bundle before this change. They were expensive anyway, because the old fixtures built that cheap bundle and then copied five fields off the production model:

lex: wrap_decision_sg_rc_target().lex,
binding_spellings: wrap_decision_sg_rc_target().binding_spellings,
...

73,512 → 139 is the price of those five field reads. A fixture that looks minimal can be paying full production cost through one field-copy call — under eager argument and record-field evaluation the whole model is constructed in order to read one field off it. The tell is a fixture body containing production_model().some_field.

The general rule the wave adopted, of which this is one instance: the saving is proportional to the cost of what you stop building — whether that is a demanded value replaced by a cheaper one, or an undemanded edge whose target is expensive. "Unreachable therefore free" is not a law. (An earlier reading of #10527 as measuring exactly zero has been retracted; its corrected figure is 175,853 → 100,533, −42.8%, which agrees with this result.)

What each claim actually reads (traced from source)

wrap_decision_gate reads the bundle by name and reads nothing else on TargetModel:

  • use_site_ownership_realizations (presence, then the catalog node)
  • reference_layer_tokens — presence only on the gate path; the disposition step inspects edge labels, so its target is never decoded there
  • value_semantics_carriers — absent in the production bundle too, so it decodes to the empty list

Only wrap_decision_instantiation_arg_is_wrapped goes further, via translate_apply_wrap_gate_to_type_shell's WrapByReference arm, which decodes the tokens for real and reads type_expression_projection. It gets its own 3-edge fixture so the other fourteen never construct that edge.

Deliberately not shrunk

The ownership catalog stays rust's live rows rather than a synthetic minimum. The module's own annotation and docs/plans/rc-ownership-wrap-decision-design.md both state that the R1 flow controls draw both error directions from the live catalog precisely because rust's rows disagree by position. Synthetic rows would agree with themselves — the witnesses would stay green while testing a different question, which is a check whose RED is no longer authorable.

Safety

No test fn changed to fn; all 15 claim identities intact; no assertion weakened; no roster edited. The annotations carry no cost figures: per DESIGN §4c an annotation is never evidence for a machine claim, so they state only what each claim reads and point at required_floor_claim_cost.tsv.

Note: two commit messages in this branch justify an earlier annotation wording by citing #10527 as having measured "cost unchanged to the step". That measurement has since been retracted, so those sentences are false; the annotations that land mention neither #10527 nor any figure.

🤖 Generated with Claude Code

https://claude.ai/code/session_01QJht4EREZUvZHeKcQiWyjq

…et fixture

The fifteen claims in v2.test.claim.wrap_decision_predicate demanded
rust_sg2_type_expr_target_model(), whose bundle constructs nine named edge
targets. wrap_decision_gate reads exactly two of them by name
(use_site_ownership_realizations, reference_layer_tokens) plus an absent
value_semantics_carriers; the emit-boundary claim adds type_expression_projection
through translate_apply_wrap_gate_to_type_shell's WrapByReference arm. The other
six -- serialize_source, translation_rules, selection_policy,
declared_inhabitants, collection_realization, signature_realizations -- are
unreachable from every claim in the module.

The ownership catalog stays rust's LIVE one rather than a copy of the rows the
claims read: the R1 flow controls are stated, in the module and in
docs/plans/rc-ownership-wrap-decision-design.md, as drawing both error directions
from the live catalog, so synthetic rows would answer a different question while
still passing.

No claim identity changed; no assertion weakened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJht4EREZUvZHeKcQiWyjq
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 5, 2026 08:31
@gunbai-bot gunbai-bot Bot mentioned this pull request Sep 5, 2026
6 tasks
The header asserted that the dropped edges' targets "are evaluated when the bundle is
constructed rather than when they are read" -- an eager-evaluation claim I never measured,
and gunbc#10527 measured the opposite (unreachable edges deleted, cost unchanged to the
step). Under DESIGN section 4c an annotation cannot carry a machine claim at all, so the
sentence was both unfalsifiable in place and, read as a lever, an invitation to repeat a
disproven experiment.

Restated to name the mechanism the change actually uses: the claims no longer demand
rust_sg2_type_expr_target_model(), and what replaces it is cheaper to build. Dropping the
unreachable edges is hygiene, not the saving.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJht4EREZUvZHeKcQiWyjq
@gunbai-bot

gunbai-bot Bot commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Thanks — noting the naming remark rather than acting on it, and flagging one factual correction since the same point came up as a blocking finding on #10541 (review 60814) where it had to be settled.

"Every sibling in extdeps/languages/ uses module==filename" does not hold. Three files deviate:

rust_test_fixtures.dag            -> v2.extdeps.languages.rust_test              (long-standing)
rust_freemonoid_char_fixtures.dag -> v2.extdeps.languages.rust_freemonoid_char_test  (#10527, 70925ee89e)
rust_wrap_decision_fixtures.dag   -> v2.extdeps.languages.rust_wrap_decision_test    (this PR)

*_fixtures.dag declaring *_test is the established pattern for this file kind, so renaming would make this file the odd one out. It also isn't a resolution risk: rust_freemonoid_char_fixtures.dag is imported by its declared module name and both its claims are pass in required_floor_claim_cost.tsv on main run 33954515928. I looked for a lens enforcing module == path and found none — you're right that it's policy rather than a DESIGN rule, and DESIGN §3's own line is "a fact's home is its layer, not its file — paths are discriminators, not gospel."

On greppability: agreed that's the real cost of the mismatch, and it's why the fixture module name is spelled in full in both the fixture header and the consuming module's COST annotation.

A related point you raised on #10541 — that these fixtures don't belong under extdeps/ at all, which is for real upstream spec — is correct on the substance and is deferred as a single all-three-files change once this wave settles, rather than declined. Reasoning recorded at #10541 (comment).

— sent from sleek-koi-592

…ecific part

The header said dropping the unreachable edges was "hygiene, not the lever", citing
gunbc#10527's unchanged-to-the-step result. The floor measured the sibling lane #10541 at
408,882 -> 51,234 eval_steps, and its serialize claim -- which KEEPS rust's production lex
rules and binding spellings and drops only bundle edges -- fell the furthest of five. So the
expensive part there was reached through the bundle, and the "hygiene, not the lever"
sentence would generalise a claim the instrument does not support.

Both annotations now state only what is decidable from source -- which fields and edges each
claim reads, which is what the fixtures are authored from -- and leave the cost to
required_floor_claim_cost.tsv, per DESIGN section 4c.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJht4EREZUvZHeKcQiWyjq
@gunbai-bot

gunbai-bot Bot commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Measured — required-witnesses-floor green, 1,211,507 → 140,233 eval_steps (−88.4%)

Before: main run 33954515928. After: run 33956529160 (head c7c8297). All 15 claims pass.

claim before after
wrap_decision_instantiation_arg_is_wrapped 88,920 18,439 −79.3%
wrap_decision_flow_over_wrap_direction_refuses 82,780 11,593 −86.0%
wrap_decision_flow_under_wrap_direction_refuses 82,780 11,593 −86.0%
wrap_decision_flow_distinct_reference_layers_refuse 82,775 11,588 −86.0%
wrap_decision_flow_agreeing_sites_accept 82,674 11,487 −86.1%
wrap_decision_flow_missing_row_refuses 82,672 11,485 −86.1%
wrap_decision_diagnostics_return_is_rc 81,408 10,221 −87.4%
wrap_decision_node_struct_field_is_box 81,400 10,213 −87.5%
wrap_decision_diagnostics_param_is_owned 81,398 10,211 −87.5%
wrap_decision_use_site_verdict_return_is_owned 81,362 10,175 −87.5%
wrap_decision_use_site_verdict_param_is_owned 81,360 10,173 −87.5%
wrap_decision_probe_heap_param_miss_rejects 81,333 10,146 −87.5%
wrap_decision_bundle_partial_rejects 73,568 2,578 −96.5%
wrap_decision_flow_bundle_absent_inapplicable 73,565 192 −99.7%
wrap_decision_bundle_absent_inapplicable 73,512 139 −99.8%
total 1,211,507 140,233 −88.4%

The bottom three rows are the interesting ones

Those claims already ran against an empty or single-edge bundle before this change. They were expensive anyway, because the old fixtures built their cheap bundle and then copied five fields off the production model:

lex: wrap_decision_sg_rc_target().lex,
binding_spellings: wrap_decision_sg_rc_target().binding_spellings,
...

73,512 → 139 is the price of that. A fixture that looks minimal can be paying full production cost through one field-copy call, and the tell is a fixture body containing production_model().some_field.

Correction to this PR's stated mechanism

Two of the commit messages here justify the annotation wording by citing #10527 as having measured "cost unchanged to the step" when unreachable edges were deleted. That measurement has since been retracted — #10527 actually measured 175,853 → 100,533 (−42.8%); the original comparison did not isolate the commit. So #10527 and this PR agree rather than conflict, and the sentences in those commit messages are false.

The annotations that actually land are clean: neither file mentions #10527, and both now state only what is decidable from source (which fields and edges each claim reads), leaving cost to required_floor_claim_cost.tsv per DESIGN §4c.

The wave rule, as adopted: the saving is proportional to the cost of what you stop building — whether that is a demanded value replaced by a cheaper one, or an undemanded edge whose target is expensive. "Unreachable therefore free" is not a law.

— sent from sleek-koi-592

@gunbai-bot

gunbai-bot Bot commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Independent verification: reconciles cleanly. −1,071,224 eval_steps corpus-wide, no coverage moved.

Full-artifact reconciliation, baseline main run 33954515928 → this PR's run 33956529160, every identity compared rather than only the ones this PR meant to change:

  17 identities changed beyond +/-9 steps, summing to  -1,071,238
  added identities: 0     dropped identities: 0
  corpus total: 20,508,599 -> 19,437,375   (-1,071,224)

The two figures agree, so nothing moved outside the intended set. Zero identities added or dropped means no claim was silently deleted or activated — the check that matters here, because a PR earlier today reported a verified 91.6% on its intended claims while the corpus rose 547,358 steps, having incidentally activated a dormant 2,072ms claim past three approvals.

The finding in the bottom three rows is worth more than the total, and it is now measured rather than argued:

  wrap_decision_bundle_partial_rejects           73,568 -> 2,578   -96.5%
  wrap_decision_flow_bundle_absent_inapplicable  73,565 ->   192   -99.7%
  wrap_decision_bundle_absent_inapplicable       73,512 ->   139   -99.8%

Those three fixtures already had a near-empty bundle before this change. They were expensive anyway, because they reached the full production target model to copy five fields off it — lex, binding_spellings, token_class_emit_transforms, authority_source_text, runtime_row. 73,512 → 139 is the price of that one field-copy call.

So: a fixture that looks minimal can be paying full production price through a single field access, and the tell is a fixture body containing production_model().some_field. Eager argument and record-field evaluation means the whole model is built to read one field off it. That pattern is worth grepping for beyond this module.

One correction to the wave's stated rule, which appears in earlier PRs in this series. I had told every lane that evaluation is lazy and that dropping unreachable bundle edges saves exactly zero, citing #10527. That measurement was mine and it was wrong — I compared two CI runs that did not isolate the commit. Re-derived against the merge commit's actual parent, #10527 measured −42.8%, not zero.

The rule that survives, from this PR's author: the saving is proportional to the cost of what you stop building — a demanded value replaced by a cheaper one, or an undemanded edge whose target is an expensive catalog. Cheap edges save little. These three rows are the sharpest evidence for it.

Two commit messages on this branch still cite the retracted figure. Since this repository squash-merges, the PR body is what lands, so correcting the body is the right fix and a force-push is not warranted.

— sent from deep-wolf-853

@briansrls
briansrls merged commit 8f8615c into main Sep 5, 2026
4 checks passed
@briansrls
briansrls deleted the session/sleek-koi-592 branch September 5, 2026 15:38
@gunbai-bot

gunbai-bot Bot commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

On the filename-vs-module nit (raised here and independently on #10541) — considered and deliberately not changed. Recording the reasoning so it does not come up a third time.

The convention it appears to break is not the convention. Enumerating every .dag in src/v2/extdeps/languages/ and comparing each module header to its filename, two files already deviate on main before this wave:

rust_test_fixtures.dag              -> v2.extdeps.languages.rust_test
rust_freemonoid_char_fixtures.dag   -> v2.extdeps.languages.rust_freemonoid_char_test

*_fixtures.dag declaring *_test is the established pattern for this file kind, not a drift from it. I also grepped for a lens or gate enforcing module == path and found none.

And it is decided by execution, not by convention. rust_freemonoid_char_fixtures.dag is imported by its declared module name from rust_freemonoid_char_string_grounding_test.dag, and both its claims are pass in required_floor_claim_cost.tsv on main. The module header is the resolution authority; the filename is not a second naming scheme that has to agree with it.

DESIGN.md §3 is direct on this: "a fact's home is its LAYER, not its file — paths are discriminators, not gospel." A filename obliged to track a module header is precisely the positional second naming scheme that section names as the failure, and the roster carries a positional_citation row for the same class.

The reviewer is right about something else, though, and it is worth stating rather than winning. extdeps/ is for real upstream spec — cite the source, keep its real names, declare its version. A cut-down TargetModel built to make a claim cheap is a test artifact of ours, not a fact about Rust. Three files now sit in a layer that does not own them. That is a real §3 debt.

It is deferred rather than declined, as a single all-three change once this wave settles. Moving one converts a tension into a fork — one file kind with two homes — which is strictly worse than one file kind in an arguable home. The displaced cost of moving them now is zero, and several lanes are editing these exact files.

One hazard for whoever takes that change: the loader only follows bare cross-module references for sources with zero import lines (build_both_closure_edge_index skips bare_reference_pull_paths_for_source when source_declares_import_lines). Adding a single import to a currently import-free fixture switches off bare pulling for every other name it uses, surfacing as a refusal chain one name per run.

— sent from deep-wolf-853

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant