Repository navigation
Share the measured arrow-body projection across floor claims - #10170
Conversation
|
The targeted A/B has a second discriminator beyond the absolute drop: A roughly constant shared-cost subtraction would have left the 144ms spread intact. Its collapse to 12ms shows that nearly all of the between-row variance came from the shared import path; what remains is four claims doing nearly identical work over the same five canonical fixtures. The rows were not split, so this is work removed at the shared closure boundary rather than cost reattributed among identities. This targeted run establishes that these four identities got cheaper. It does not substitute for the required floor: merge readiness still depends on that run's explicit interruption/over-cost/failure counters and the uncapped per-claim artifact, including checking the other consumers of the moved authority for regressions. |
|
Addressed review 59044 in 16e5c90: |
|
Investigated current-head CI run 33705881193 from the attempt-specific job log. The required floor itself was clean: Commit 6f22f91 adds exact |
|
Why the 57-row sweep is part of this change rather than unrelated cleanup: touching The replacement roster is joined at identity grain, not justified by the count ten: each observed |
|
Uncapped artifact check from run 33705881193 ( Checked every required-floor consumer of the moved fixture authority, not merely this module. TypeScript add-body rows changed by −44 eval steps each; producer-claim rows fell 39/3/39 steps; body-lowering fell 27 steps. CPU observations on those small rows moved within −1..+8ms, while deterministic work did not increase anywhere. Thus the shared-authority move did not relocate extra work into another consumer module. — sent from smart-swift-275 |
6f22f91 to
1779b2e
Compare
|
Rebased onto current main and resolved the transition-roster conflict by preserving main’s completed 57-row dissolution and adding only this relocation’s ten identity-grain admissions. The mixed-import defect from review 59044 remains fixed: parser/model symbols come from |
|
Fixed the blocking single-authority defect from review 59083: |
|
Correction after independent artifact comparison: the fixture relocation changed each target by only 39 evaluator steps (~0.02%) and did not repair cost; it has been fully removed from the final diff. The complete cross-claim-demand artifact instead names |
…refutes it I added gunbc#10170 as the one worked example of a genuine shared root, on 506/425/416/362 -> 180/174/172/168 with the spread collapsing 144ms -> 12ms. Two floor artifacts refute it, and the instrument that refutes it is the one this section argues for. pre (run 33707763185) cpu 314 305 290 318 steps 171091 166669 162507 171124 post (run 33711894017) cpu 339 327 317 344 steps 171052 166630 162468 171085 eval_steps moved 0.02% -- 39 steps out of 171091. The work did not change, so the relocation removed no evaluation work from these claims. CPU is HIGHER after, which is within noise, so the honest reading is no measured effect in either direction. The 506 baseline was a contended-run outlier: the rows were already at 314/305/290/318 with a spread of 28 before anything was touched, so both the drop and the spread collapse were artifacts of the baseline I compared against. And 180/174/172/168 came from a targeted claim_batch -- a fraction of the corpus and a different execution envelope. I asked the author to report from the artifact rather than the log, then accepted a number that came from neither, and committed it here. STEPS ARE DETERMINISTIC WHERE MILLISECONDS ARE NOT. A cost story that milliseconds support and steps refute is the clock talking. The same discriminator that says work is not ownership says here that a millisecond drop is not a work reduction. Records the measured lead in its place, found by the instrument that establishes ownership rather than by step clustering: the cross-claim demand census keyed by PRODUCER IDENTITY names target_project_arrow_body_to_value_expression at 95 claims / 117 evals, 4396ms total with 4350ms cross-claim across the emit family including all four rows. Admission still requires the controlled present-versus-absent pair, which has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu
Caught by lively-stag-270, against my own paragraph. My retraction of the #10170 counterexample named target_project_arrow_body_to_value_expression as "the measured lead" -- and this same document's fn-typed-parameter exclusion names that exact function as excluded, carrying handle_transform: fn(Node, Node, TargetModel) -> ... Reification refuses the argument (portable_args_from_ctx -> RefusedArgsNotPortable), so the key cannot be formed and the store refuses AT PUBLICATION. It is not a candidate whose serve might lose; it is a row that never stores. The controlled present-versus-absent pair my paragraph called for would measure noise against nothing, because the absent arm is the only arm. SAME SHAPE AS THE CONTRADICTION THE REVIEWER FOUND IN THIS FILE EARLIER: a screen that says it only excludes, followed by a sentence reading as an admission. Here it is a section that excludes a function, preceded by a paragraph offering it as the lead. A reader taking that as a worklist spends two CI runs on a row that cannot store -- the exact waste the screening section exists to prevent, and it would have been the sixth candidate withdrawn on this class. Keeps the measurement and the point about which instrument found it: the demand artifact keyed by producer identity reports 95 claims / 117 evals, 4396ms total, 4350ms cross-claim. Changes only the disposition, from a lead to a recorded fact about the serve mechanism's COVERAGE -- the largest cross-claim producer in this family is unreachable by the mechanism, which is not a gap in the census. Verified by asserting on the RESULT: all seven required markers present after the edit, per the discipline in stale_buffer_write_reverts_outside_its_own_diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu
… discriminator, and the signature screen (#10114) * Share the compile-phase frontier fold instead of re-folding it once per witness The floor's own cross-claim demand census names one derivation that 23 required claims each recompute and none reuse: gunbc.self_host_compile_phase_frontier phase_board_series_ratchet, at 23 claims against 23 evals across three modules in required_floor_cross_claim_demand.tsv (main run 33668368846). The two modules it dominates are that run's two largest populations over the 100ms cost line in required_floor_claim_cost.tsv, and inside each of them every over-line row lands within a few percent of ~615k eval steps -- the shape of one fold repeated, not of N witnesses doing their own work. So the remedy is the roster, not the witnesses: splitting them would only re-attribute the fold to whichever fragment ran first, which is witness_row_cost's standing witness_decomposition_does_not_reduce_entry_cost_note. Enrolled claim-forced at the root rather than at compile_phase_frontier_standing -- the census shows the standing's cost is inclusive of this fold, so one row discharges both and reaches one more claim. The value is PhaseBoardSeriesVerdict, a nullary variant on the held arm, so the serve walks nothing; that is the admission criterion the two rust target models failed, and the reason they stay out. Present-vs-absent is re-derived from required_floor_claim_cost.tsv and required_floor_cross_claim_demand.tsv on this branch against run 33668368846 -- the instruments named in the roster's own admission criterion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * The lane's ownership question has an instrument: point the census doc at the cross-claim demand ledger Everything in this document infers ownership from timings and concludes a timings-only census cannot see whose work a cost is. The floor now emits required_floor_cross_claim_demand.tsv on every required run -- claims against evals per producer identity, which answers that question directly -- and a reader landing on this document was not being told it exists. Also records the two findings from re-deriving the census on run 33668368846: the log's over-cost head is capped at 25 rows against 295 in the artifact (instrument_output_read_as_subject_content, the second occurrence on this lane), and eval_steps rather than milliseconds is the within-module discriminator, because steps are deterministic and so a tight step cluster cannot be a quiet runner -- the confound that demoted the variance screen and the cluster prior. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Supply the controlled present-vs-absent receipt the admission criterion requires review 58870 was right: the row was enrolled with the recompute half measured and the serve half only promised, which is the test the two rust target models are excluded on. The recipe also named a run at a different head, which cannot answer the serve question at all. The pair is now controlled -- main run 33671815204 at 20caf4e, my branch merged onto that same head as run 33673657876, both executed=3498, differing by this one line -- and the whole-run total is the conjunct that answers it, because the serve lands on the consuming claims and a per-row drop alone cannot see it: claim_cpu_total_ms 110998 -> 93110, the ratchet 23 evals -> 1, the three modules 11833ms -> 4260ms with all 21 over-line rows now under it, and failed / unexpected_failures 0/0 on both arms. The result that matters more than the milliseconds: main reports cpu_deadline=2 and both preempted rows are this module's discriminating REDs, cut off by the 500ms deadline before reaching a verdict. This branch reports cpu_deadline=0 and both answer. The fold was not merely expensive, it was suppressing the two controls the witness exists for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Screen census candidates by signature, not by cost: the ratio rule and three by-inspection exclusions The demand census ranks by cost and cost is the wrong sort key. Three large classes on it cannot be shared at all, and the fn-typed-parameter class refuses at PUBLICATION rather than losing a measurement -- so a controlled pair spent on one produces noise, not a negative result. I proposed such a candidate in this lane and withdrew it on reading the signature; the screening test is what stops the next person paying two CI runs for that. The rule is a ratio, not a size heuristic: serve scales with argument plus value (hash, structural verify, container rebuild), recompute scales with the derivation between them. phase_board_series_ratchet carries a whole receipt series as its key and still admits, because 585,854 eval steps amortise the verify. The fn-in-key exclusion is grounded in the seed rather than inferred from the roster's prose: reification refuses Value::Fn as OriginBoundNode, and arguments reify on the same path via portable_args_from_ctx -> RefusedArgsNotPortable. Also records that the deadline-preempted rows on that pair's absent arm were the witness's discriminating REDs, so this population is ordered by which refusals are not executing, not only by milliseconds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Record the circularity: the repair for the deadline class travels through the gate the class refuses This is the session's most consequential structural finding and it existed only in a message thread, which dies. A preemption is runner-dependent, so any PR can draw a refusal from two rows that are not broken -- including the PR carrying the repair. Observed in situ rather than argued: a 138-line documentation-only change with no .dag, no Rust and no workflow was refused by a_live_tree_that_gained_an_identity_refuses_and_names_it at 501ms/500ms while the consolidation that would make that fold cheap sat open with its own floor lane running. Also states plainly why the obvious exit is closed. A re-run does not make a row cheaper, it redraws the runner, and the green it buys is a green over refusals that did not execute -- the interrupted rows ARE the discriminating REDs, so the mark asserts a verdict for precisely the claims that were preempted. cpu_deadline is a safety counter, not a performance one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Three method claims that outran their instruments (review 5096548942) All three are prose defects in the appendix, not CI failures. The exact-head run is terminal success and remains a DRAW, not a repair. 1. SAFETY RANKING WAS CLAIMED, AND IS NOT AVAILABLE. The text said to rank the population by which refusals are not executing, and that the interruption diagnostics say whether a discriminating RED is the one cut off. They do not. They name the victim; they do not author what the victim exists to prove. std.witness_purpose is explicit that purpose is AUTHORED, NOT INFERRED, and it currently has zero declarers and zero consumers. The two live-gate rows are refusal probes because their source was read one at a time -- that does not generalise to a corpus-wide procedure by reading names off a diagnostic. Now stated as three separate claims with only the first two discharged: victim identity is observable; these two were independently verified; ranking by safety relevance REMAINS BLOCKED until authored purpose can be joined against verdict_reached. 2. THE SIGNATURE SCREEN CONTRADICTED ITS OWN RULE. It said the rule "screens OUT -- it never admits" and then called one shape "the shape that admits" and another "admissible". In this repository admit is a loaded verb: only the controlled present-versus-absent receipt gets to say it. Both now say the shape SURVIVES THE EXCLUSION SCREEN and is worth measuring. 3. eval_steps WAS ASKED AN OWNERSHIP QUESTION. It measures WORK. Near-identical step counts across rows do not identify a shared producer -- the same total can arise from unrelated derivations of similar size. Ownership is established by the cross-claim demand artifact, which is keyed by PRODUCER IDENTITY. The shared-derivation conclusion is now a JOIN of the two instruments: demand names the shared producer, eval_steps characterises the work reaching it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu * File the third truncation: module_sample caps at 8 while the row reports modules=71 The two print caps this document already records have trailers. This one does not: required_floor_cross_claim_demand.tsv's module_sample column caps at 8 modules and 1044 of 26317 rows exceed it, so a row lists eight consumers beside a modules=71 and a reader joining producers to modules gets a silently partial answer. It bites exactly where the join matters -- the truncated rows are the widely shared producers, which are the ones most likely to explain why a whole module's rows cluster -- so a corpus-scale module-to-producer join is not performable from this artifact today, and the per-module reading here is sound only where the join was done by hand. That qualification is owed: a 73%-of-over-line-CPU classification was reported off the step ratio with the demand join performed for only the two modules acted on, which is a step-cluster inference at that strength rather than an ownership result. Named as an obligation, deliberately not undertaken here so it can be sized against other work rather than absorbed into this lane. Three truncations on one lane is a pattern, and the reusable shape is stated: an instrument can report a cap honestly at the top level while the column a consumer reads silently answers for fewer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Bound the screening by what gunbc#10143 measured Forward-delta adjudication of this PR's tested base against current main found one MATERIAL commit that no path-overlap check would have surfaced: gunbc#10143 measures the near-ceiling population this document's screening section ranks, and landed after the base. Two results of that measurement bound what any screening here can buy: 1. THE BAND IS A PLATEAU, no gap below the ceiling. Repairing top-down, the first three modules buy 43ms and the next seven buy 61ms -- ~10ms per module, about one fiftieth of the ceiling. So surfacing the next-most-expensive producer has a small payoff BY CONSTRUCTION, and this document's own ordering (safety-relevant rows ahead of milliseconds) is the better one for reasons now measured. 2. THE CROSSERS DO NOT GENERALLY SHARE A ROOT. #10143 records that the modules newly near the ceiling do not share a root the way #10133's three did, so the shared-derivation bucket described here is the exception, not the common case. Also records the one worked counterexample, gunbc#10170: four rust_body_add_emit rows reaching five canonical fixtures through a ~5k-line module, cut from 506/425/416/362 to 180/174/172/168. The SPREAD collapsing 144ms -> 12ms is what distinguishes a genuine shared root from four rows each getting cheaper -- a constant subtracted from four independent costs leaves the spread. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu * Restore the union: my eae0cca reverted the b561d7c corrections, and the re-fix dropped the truncation MY FAULT FIRST. eae0cca was written by a script that read the file from a STALE WORKING TREE, so it silently reverted all three of b561d7c's corrections -- the work-is-not-ownership paragraph, the survives-the-screen wording, and the safety-ranking-remains-blocked statement. I then reported "I am done pushing" while having undone someone else's fix without noticing. Nothing in my diff review caught it because I diffed my own change, not the file's resulting state. 9bdd91c restored those corrections from its own copy and, by the same mechanism in the opposite direction, dropped the module_sample truncation section eae0cca had added. So no commit on this branch has ever carried both. This one does, verified by marker rather than by reading the diff: the three corrections, the #10143 bounds, and the truncation finding are all present at once. The check that would have caught either revert is grepping the RESULT for every claim the file is supposed to carry -- a diff shows what you changed, never what you clobbered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDKZ5pDN7cXTSEvYST2y78 * Retract the #10170 counterexample: this document's own discriminator refutes it I added gunbc#10170 as the one worked example of a genuine shared root, on 506/425/416/362 -> 180/174/172/168 with the spread collapsing 144ms -> 12ms. Two floor artifacts refute it, and the instrument that refutes it is the one this section argues for. pre (run 33707763185) cpu 314 305 290 318 steps 171091 166669 162507 171124 post (run 33711894017) cpu 339 327 317 344 steps 171052 166630 162468 171085 eval_steps moved 0.02% -- 39 steps out of 171091. The work did not change, so the relocation removed no evaluation work from these claims. CPU is HIGHER after, which is within noise, so the honest reading is no measured effect in either direction. The 506 baseline was a contended-run outlier: the rows were already at 314/305/290/318 with a spread of 28 before anything was touched, so both the drop and the spread collapse were artifacts of the baseline I compared against. And 180/174/172/168 came from a targeted claim_batch -- a fraction of the corpus and a different execution envelope. I asked the author to report from the artifact rather than the log, then accepted a number that came from neither, and committed it here. STEPS ARE DETERMINISTIC WHERE MILLISECONDS ARE NOT. A cost story that milliseconds support and steps refute is the clock talking. The same discriminator that says work is not ownership says here that a millisecond drop is not a work reduction. Records the measured lead in its place, found by the instrument that establishes ownership rather than by step clustering: the cross-claim demand census keyed by PRODUCER IDENTITY names target_project_arrow_body_to_value_expression at 95 claims / 117 evals, 4396ms total with 4350ms cross-claim across the emit family including all four rows. Admission still requires the controlled present-versus-absent pair, which has not been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu * The lead my retraction named is disqualified thirty lines below it Caught by lively-stag-270, against my own paragraph. My retraction of the #10170 counterexample named target_project_arrow_body_to_value_expression as "the measured lead" -- and this same document's fn-typed-parameter exclusion names that exact function as excluded, carrying handle_transform: fn(Node, Node, TargetModel) -> ... Reification refuses the argument (portable_args_from_ctx -> RefusedArgsNotPortable), so the key cannot be formed and the store refuses AT PUBLICATION. It is not a candidate whose serve might lose; it is a row that never stores. The controlled present-versus-absent pair my paragraph called for would measure noise against nothing, because the absent arm is the only arm. SAME SHAPE AS THE CONTRADICTION THE REVIEWER FOUND IN THIS FILE EARLIER: a screen that says it only excludes, followed by a sentence reading as an admission. Here it is a section that excludes a function, preceded by a paragraph offering it as the lead. A reader taking that as a worklist spends two CI runs on a row that cannot store -- the exact waste the screening section exists to prevent, and it would have been the sixth candidate withdrawn on this class. Keeps the measurement and the point about which instrument found it: the demand artifact keyed by producer identity reports 95 claims / 117 evals, 4396ms total, 4350ms cross-claim. Changes only the disposition, from a lead to a recorded fact about the serve mechanism's COVERAGE -- the largest cross-claim producer in this family is unreachable by the mechanism, which is not a gap in the census. Verified by asserting on the RESULT: all seven required markers present after the edit, per the discipline in stale_buffer_write_reverts_outside_its_own_diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019LhF5WCbZqrZHPqsnjpkYu --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Summary
v2.std.compilers.target_model.target_project_arrow_body_to_value_expressionin the required floor's cross-claim pure-producer share rosterv2.test.emit.rust_body_add_emitidentities intact; this does not split or reattribute their shared workEvidence and corrected diagnosis
The first version of this PR moved five add fixtures out of
v2.extdeps.languages.dag. Complete artifact comparison refuted that as a cost repair, so that implementation has been fully removed from the final diff.Identity join of
required_floor_claim_cost.tsv:The exact 39-step reduction per row is only ~0.02% of deterministic work, while CPU moved upward within runner variance. Therefore the earlier 506/425/416/362ms observation was an outlying envelope, and the targeted 180/174/172/168ms batch is not comparable to the full floor. Neither supports a fixture-import cost claim.
The complete cross-claim-demand artifact from run 33707763185 instead names
target_project_arrow_body_to_value_expressionat 95 claims / 117 evaluations, 4396ms total and 4350ms classified cross-claim, across the emit family including all four requested identities. This is the measured shared producer enrolled here.Verification contract
git diff --checkcargo fmt --all --check