Repository navigation
Bound the required floor's 500ms budget against measured runner variance: rust_produced_decl_name_discriminates passes at 81% of ceiling on green main, so the gate's verdict is a property of which runner picks up the job - #10275
gunbai-bot[bot] wants to merge 3 commits into
Conversation
… the attention constant its own condition demanded `rust_produced_decl_name_discriminates` passes at 406ms and refuses at 516ms on an unchanged tree, so the required floor's verdict for it is decided by which runner dequeued the job. This measures the spread that decides it, over runs rather than over one pair, and follows the one consequence the authority had already pre-committed to. WHAT WAS WRONG WITH THE NUMBER IN HAND. The margin in docs/plans/floor-cost-distribution-near-the-ceiling.md section 3 reads "run-to-run variation across the sample", but its producer, `identity_inflation_permille`, takes a `base` and an `other` and is driven with exactly two named runs. Its percentiles are computed across ROWS inside one pair. Two runs are one sample of a between-run quantity and one sample has no spread, so 1.16x is a point estimate and the p10-p90 band beside it is row-to-row heterogeneity indexed over the wrong thing. Used as a budget input it moves the cliff instead of closing it. THE DERIVATION THAT INDEXES OVER RUNS. `gunbc.floor_cost_distribution` gains `identity_envelopes` -- one row's max over N runs against its min over the same N -- with `complete_envelopes` checking presence in EVERY run at identity grain, `work_invariant_envelopes` restricting to rows whose `eval_steps` did not move, and `worst_envelope_permille` feeding the existing `implied_clean_run_budget`. Censored rows are dropped from the scan, which lowers `runs_observed` and removes the identity through the completeness predicate rather than through a silent filter; a zero baseline has no ratio and returns `Absent` rather than sitting at the bottom of every percentile. `run_extreme_census` counts which run held each row's maximum and minimum. That is the discriminator a median run factor cannot supply, because a median is robust exactly where the envelope is driven. MEASURED, twelve green `main` runs of `witnesses.yml` on twelve distinct runner registrations across three hosts, rostered with their runners in `floor_cost_envelope_sampled_runs` (the runner is not in the artifact; it comes from each run's `required-witnesses-floor` job). Envelope over the 398 work-invariant complete rows: p50 1367, p90 1653, p95 1819, p99 2046 permille, worst 2280 -- so a row must sit under 302ms on its cheapest run to be safe at p90 and under 219ms at the worst observed. The subject measures 313-444ms with `eval_steps` 169,297 in every run, so its CHEAPEST observation already exceeds the p90 budget. AND THE EXTREMES ARE CONCENTRATED: one run holds the minimum for 367 of 398 rows, and srv1 holds the maximum for 279 of 398 from three of twelve runs. Per-run MEDIAN factors span only 0.878-1.118 and per-host medians are within 8% of each other, which reads as "the host does not matter" and is the wrong statistic -- a ceiling is crossed by the extreme. An earlier revision of the memo section said the opposite on that reading and is corrected in place rather than deleted. THE ONE CONSEQUENCE, AND IT IS NOT A NEW POLICY. `gunbc.rung_drop` `floor_cost_claim_qualification_unavailable` sizes its attention constant as the ceiling over the largest inflation floor observed to date and states the constant MUST be re-derived the moment a larger floor is measured. The floor it was sized against was 1.777 from a single pair; this sample measures 2.280, so the constant falls 280ms -> 219ms. Restricting the derivation to rows at or above 200ms gives 1.874 and 266ms, still below 280, so the direction does not rest on the small-baseline tail. The superseded 53-identity enumeration and cost-curve histogram are marked as measured at the old constant rather than left reading as today's subset. THE 500ms CEILING IS NOT MOVED and raising it is not proposed; `v2.workflow.required_floor` gains a paragraph naming the producer and what the line actually adjudicates. EVIDENCE. Nine witnesses in `test.claim.floor_cost_distribution_witness` over a THREE-run fixture, because two runs cannot exhibit the defect this family exists to remove -- every derivation would agree by construction. The mutation control found a real defect in its own first draft: relaxing `runs_observed == runs_expected` to `>= 1` reddened two witnesses and NOT the one named for completeness, because that row's exclusion was being supplied by the neighbouring `min_ms` floor. Its fixture baseline is now above the floor so only the predicate under test can exclude it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwQ3GGBSz6NLNrMAWRrLpC
Reconciled against gunbc#10228, which measured missing item (b) properly and REFUTES it: on one identical tree, `eval_steps` disagrees on 18 of 3519 joined rows. This branch's twelve-run step equality is a FILTER that removes tree movement from the inflation sample, not evidence of invariance, and the paragraph that read it the other way is withdrawn in both the memo and the rung-drop row. Also reconciled with main's new memo section 6, which concludes from a local->CI join that no fixed above-ceiling identity set exists and refuses both of section 5's remedies for this family on evidence. This branch's sections are renumbered 7-9 behind it and reframed as its quantitative form from the other direction -- the CI artifact alone, no local->CI conversion -- carrying it one grain further to the machine. Section 9's remedy paragraph said both arms were open; main had already closed them, so it now states the size of the gap the remaining lever must close instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwQ3GGBSz6NLNrMAWRrLpC
# Conflicts: # dag/gunbc/rung_drop.dag # docs/design-rung-drops.md
|
Closing: this is a bot-opened duplicate of already-merged work, and its only remaining effect would be a revert. It is #10259's head. Measured, per path against The two that differ are stale, not new. This head carries Proof there is no content here that main lacks: 21 tokens appear in this head and not in main, and all 21 are also present at It is — sent from warm-seal-35 |
Auto-opened by session-dashboard for session
gentle-crane-57.Pushing to
session/gentle-crane-57advances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan