Repository navigation
Report that the work-elimination alpha cannot be priced, proven three ways - #8641
Conversation
… ways The strategy corpus names work elimination as the strongest of seven alpha sources because it is the only one a competitor cannot buy, and the only one with no measurement behind it. It cannot be measured today, and this records why with the evidence rather than leaving it to be rediscovered. FIRMLY ESTABLISHED. Baseline entries is 9776 witnesses on every run -- an IDENTITY rather than an observation, since the fold has no selection concept to configure. Baseline wall over the last 20 successful main runs: p50 28.3, mean 28.9, p90 31.3 minutes, with three limits attached so it cannot be promoted -- one job on one repo (ours, arm64, warm), wall of the whole job including checkout and cargo build rather than billed unit-minutes, and successful runs only because a cancelled run measures when someone pushed. That number explicitly does NOT fill typical_job_minutes, which is a planning figure about CUSTOMER jobs; replacing it would retarget a load-bearing figure onto the heaviest job we own. An honest -1 beats a confident wrong subject. ADMITTED IS NOT COMPUTABLE, three routes, each verified: the .dag authority is per-entry over an OOM-class live-pool closure; the Rust twins entry_file_touched_via_import_closure and compile_clean_scope_plan_for_ci are private with zero references under src/v1/stage0/src/bin; and the only pub surfaces return Bool, answering "is it clean" and never "which entries would be selected". AND UNIT-MINUTES ARE UNAVAILABLE FOR BOTH ARMS, not just the admitted one -- a correction to how this was previously held. write_witness_row_cost_receipt is called once inside the batch walk, and run() returns early into run_required_floor before that walk is entered, so the path CI runs never reaches it. run_required_floor times only phases; RequiredFloorOutcome has no duration fields; the floor log times 1022 of 9776 rows. The log-parse route was deliberately not computed: those rows are selected FOR being slow, so summing them is survivorship bias pointing the expensive way -- inflating baseline, inflating the alpha, in the direction we would want the answer to go. ONE CUT, THREE CONSEQUENCES: unexercised, unproduced, unreachable. The alpha went dark and the light that would have shown it went dark in the same motion, which is why no gap was visible. ALSO CORRECTS DESIGN AT THE MODEL, not the emission. The floor-cut paragraph said the compile-clean scope authority and its import-closure selection are live and unmodified, and that what is gone is the CI job that invoked them rather than the authority. Both halves are true and together they license a false inference: surviving a cut and being callable are different properties, and only the first was checked. A reader planning against that sentence would budget an afternoon and find no caller -- premise contamination of the same class that paragraph already corrects itself for twice. DESIGN.md is regenerated from design_document.dag rather than hand-edited; the diff is one line. selected_software_execution_alpha stays WorkAvoidedUnmeasured, with a sharpened obligation that names what is missing rather than asking for a measurement in general. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012q31BK3okLA8vG4kdTWtBf
…iment Main's #8618 edited the same floor-cut paragraph this branch corrects, so design_document.dag is resolved by taking main's version and re-applying only the reachability clause on top. DESIGN.md is REGENERATED from the resolved model rather than hand-merged -- it is a generated projection and hand-editing it is what auto-heal reverts. Also folds in a correction to the FUTURE measurement design, which is worth carrying in the report because the obvious experiment is wrong in the direction that flatters us and somebody will rebuild it once instrumentation lands. Summing full-run durations of the selected rows is NOT the counterfactual. The full run carries shared preparation, first-toucher attribution, cross-claim memoization and order-dependent warm state, so if A pays a preparation that warms C, C's full-run duration is C's cost GIVEN A RAN -- and executing {C} alone would make C pay it itself. That understates selected cost and overstates the advantage, the same bias direction as the log-parse trap this report already refuses. What it requires instead is PAIRED EXECUTION with both arms observed and neither reconstructed: C_full executing the complete roster, C_sel executing the selected population in its OWN FRESH PROCESS including selection, preparation and finalization, and alpha = 1 - C_sel/C_full matched on subject, runtime closure, execution class and roster authority. And the selection entry point must publish a RECEIPT rather than a count -- subject, complete-roster digest, selected identities, selected-roster digest, per-identity basis, selector identity -- with the count a projection of it. A scalar cannot say which identities, whether cost rows join to them, or whether two selector versions picked different populations of the same size. That is the same collapse as the Bool surface this report criticises, one value up, and asking for a count would have reproduced it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012q31BK3okLA8vG4kdTWtBf
Two conflicts, resolved in the required order: the authority first, then its projection. dag/gunbc/design_document.dag — both sides had unique content, so neither side could be taken. This branch corrects a false-inference clause (surviving a cut and being callable are different properties); main independently added a passage explaining that `executed` counts a witness reaching the fold, not its assertion running. Taking main's side would have restored the very false inference this PR exists to remove — in a clause about premise contamination, which would have been self-refuting. The resolution is main's line with its 245-char stale clause replaced by this branch's 1215-char correction, verified by assertion rather than by reading: both sides' unique content present, stale clause absent, 8178 - 245 + 1215 = 9148. DESIGN.md is the generated projection of that authority and was REGENERATED, never hand-merged, per the generated-artifact driver's refusal. The regenerated projection carries both sides' content. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012q31BK3okLA8vG4kdTWtBf
|
Queue hold — authority-touching PRs (operator ruling, 2026-08-20) This PR modifies This is a queue hold, not a judgement on the change. Nothing here is wrong and nothing is being asked of you. The operator is merging manually, so the hold is enforced at the merge hand — you do not need to do anything to comply, and this comment is a courtesy so you are not surprised by a merge that does not come. Why the hold exists. A regeneration repair's entire content is "the derived files match the authorities as of now." Its correctness is indexed to a moment, so any authority merge landing while it is in flight invalidates part of it — silently, without touching a line its author wrote. Against a moving queue it cannot converge, because the target moves faster than build → regen → push → CI. The remedy has to be a queue policy rather than more effort from the repair author. Expected duration: short. The repair ( If your CI is currently red at "Regen fixed point: first generation matches committed candidate", that is very likely inherited rather than yours. Main has been red at that step since One trap worth knowing while reading that step: the step named "Regen fixed point" runs — sent from smart-ram-730 |
|
Hold LIFTED — the regen repair has landed and verified. The authority-touching hold posted on this PR earlier is over. Nothing is being asked of you; this is the follow-up to that notice so it does not sit here reading as still-active. What cleared it. gunbc#8677 merged as First green at step 6 since If your CI is still red at that step, it is a stale run from while main was broken. A re-run against current main should clear it. If it does not, the remaining failure is genuinely yours or a third cause — read the step output rather than the outcome, because that step has produced at least four distinct causes in the last day (inherited drift, own drift, an One correction to the earlier notice, since it circulated on this PR: step 7 is not a cheap receipt read. It performs a full second emit pass and took longer than step 6 on this run — twelve minutes and counting versus six. What it reads from the prior receipt rather than recomputing is the single value — sent from smart-ram-730 |
…ent-state # Conflicts: # DESIGN.md # dag/gunbc/design_document.dag
A report that a measurement cannot be taken, established three ways — plus the DESIGN correction that finding forces.
The strategy corpus names work elimination as the strongest of seven alpha sources "because it is the only one a competitor cannot buy", and the only one with no measurement behind it. It cannot be measured today.
What is firmly established
Baseline, entries: 9,776 witnesses every run — an identity, not an observation. The fold has no selection concept to configure.
Baseline, wall (n=20 successful main runs): min 26.5 · p50 28.3 · p90 31.3 · max 31.6 · mean 28.9 min — with three limits attached so it cannot be promoted: one job on one repo (ours, arm64, warm); wall of the whole job including checkout and
cargo build, not billed unit-minutes; successful runs only.That number explicitly does not fill
typical_job_minutes, which is a planning figure about customer jobs. Replacing it would retarget a load-bearing figure onto the heaviest job we own.Admitted: not computable. Three routes, each verified —
.dagauthorityentry_file_touched_via_import_closure,compile_clean_scope_plan_for_ci— private, zero references undersrc/v1/stage0/src/binpubentries returnBool— is it clean, never which entries would be selectedUnit-minutes are unavailable for both arms
A correction to how this was previously held.
write_witness_row_cost_receiptis called once inside the batch walk;run()returns early intorun_required_floorbefore that walk is entered, so the path CI runs never reaches it.run_required_floortimes only phases,RequiredFloorOutcomehas no duration fields, and the floor log times 1,022 of 9,776 rows.The log-parse route was deliberately not computed — those rows are selected for being slow, so summing them is survivorship bias pointing the expensive way: inflating baseline, inflating the alpha, in the direction we'd want the answer to go.
One cut, three consequences
unexercised (selection deleted at the root) · unproduced (no per-claim duration) · unreachable (live authority, no caller,
Bool-only surface).The DESIGN correction
The floor-cut paragraph said the scope authority and its import-closure selection are "live and unmodified by the cut — what is gone is the CI job that invoked them, not the authority." Both halves are true and together they license a false inference: surviving a cut and being callable are different properties, and only the first was checked. A reader planning against it would budget an afternoon and find no caller — premise contamination of the same class that paragraph already corrects itself for twice.
Corrected at the model (
gunbc.design_document), withDESIGN.mdregenerated rather than hand-edited. One-line diff.Carrier
selected_software_execution_alphastaysWorkAvoidedUnmeasured—work_avoided_permilledivides admitted by baseline, so feeding it entry counts would yield a confident per-mille from the wrong quantity. Its obligation is sharpened to name what's missing rather than asking for a measurement in general.🤖 Generated with Claude Code
https://claude.ai/code/session_012q31BK3okLA8vG4kdTWtBf