Skip to content

Does grammar_relation_row_for_emitted's cross-claim serve beat recompute over its own consumers? Measured present-vs-absent, not inferred from a neighbour's null result - #10234

Closed
gunbai-bot[bot] wants to merge 3 commits into
mainfrom
session/merry-deer-84

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Auto-opened by session-dashboard for session merry-deer-84.
Pushing to session/merry-deer-84 advances this PR.

Worker attestation

Before flipping this PR to ready for review, confirm each item:

  • Title describes the change (not the session id or branch).
  • PR body summarises what and why (replace the TODO below).
  • Tests run: name the command (e.g. npm test, cargo test) and the result.
  • If this closes a work item, the body contains a Closes #N directive.
  • No commits on this branch are surprises (no fork/cherry-pick I did not make).
  • No secrets / credentials / large binaries staged.

Summary

TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.

Test plan

  • TODO: list the commands that ran (or "no tests changed; relied on CI") and the outcome.

gunbc-ci-auto-heal and others added 2 commits September 3, 2026 02:56
…, and it passes

The row was enrolled on measured recompute AND measured sharing before this file's
serve-below-recompute criterion existed, so it stood on two of the three tests the
header now requires. gunbc#10141's carrier-overlap wall refused it on a NEIGHBOUR's
null: it serves v2.test.manual.rust_add_emit_translate, a carrier of the refused
rust_target_model_core_edges, whose refusal was NoMeasuredEffectOverItsConsumers. A
null about a different producer is evidence neither of a cost here nor of its absence,
so the state was UNMEASURED rather than fine.

Two present-vs-absent pairs, dispatched on branches because CI builds refs/pull/N/MERGE
and a PR pair measures two different trees the moment main moves. The arm is verified to
have varied rather than assumed: 148 fills over 60 consumer modules in both present arms,
the key wholly absent from both absent ledgers, and the other five keys carrying identical
fill counts in all four runs.

The control correction is the transferable part. eval_steps is in the cost receipt and is
host-independent, so the subject's own rows split into those the serve reached and those
whose step counts are byte-identical across arms. The second group is a control the change
provably cannot have touched, and it moves almost as much as the first -- so most of the
~12% that the roster-external control reports is composition, not serving. Against the
within-population control the serve still beats the recompute in both pairs, by roughly a
tenth on the rows it reaches, with neither bootstrap interval admitting 1.

Comment-only: annotations are erased from the semantic projection, so no roster row, no
emitted byte and no resolution result changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014T3j1sSKhCsKjuiAfpYQeG
…d it not firing

The note asked #10141's carrier-overlap wall to stop refusing this key. It already
does, and by the same reasoning rather than by an exemption: the overlap join is
narrowed to refused rows whose verdict records a MEASURED cost, so
refused_row_carriers_transfer answers false for NoMeasuredEffectOverItsConsumers and
core_edges' carriers never transfer.

Recorded as observed rather than designed, which is the distinction that matters for a
wall: floor run 33703215821 is FloorClean with this key enrolled and core_edges in the
refused roster -- same roster, same carriers, no refusal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014T3j1sSKhCsKjuiAfpYQeG
@briansrls
briansrls marked this pull request as ready for review September 3, 2026 12:59
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-03T13:00:43.142070Z fae6853 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fae6853e31

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +243 to +245
// and is host-independent, so the subject's own rows split into those whose steps FELL (the serve
// reached them) and those whose steps are BYTE-IDENTICAL across arms (the serve never reached
// them, same claims, same runs). The second group is a control the change provably cannot have

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use actual cache hits to define the control

When an enrolled serve is reached but its avoided work is offset by other evaluation—as this note explicitly allows at lines 252–253—the claim can retain the same aggregate eval_steps while still being affected by the change. Such claims therefore contaminate this alleged untouched control, and their wall-time effects can bias the normalized serve-versus-recompute estimate in either direction, so the conclusion that the third admission criterion passes is not established by this split. Define the control from explicit cache-hit/demand attribution rather than equality of an aggregate step count.

Useful? React with 👍 / 👎.

… not

codex 59321 requested changes: the note promised "re-derive with the
instrument named" while naming only run ids and a TSV column. Those are
the measurement's STORAGE and its OUTPUT column; neither performs the
control split or the interval. The sentence promised a producer that does
not exist, which is worse than transcribing a figure — a reader who tries
to act on it finds nothing to run.

Corrected rather than softened, which is the part that matters. The note
now splits the claim by what is actually reproducible:

  RE-DERIVABLE, by a reader with no instrument, from the four run ids plus
  required_floor_claim_cost.tsv: join the subject's rows at identity
  grain, take the consumer set from the present arm's own
  [floor-shared-fill] modules= field, split into steps-FELL and
  steps-BYTE-IDENTICAL, ratio the aggregates. That is the decisive
  comparison and it needs arithmetic, not tooling.

  NOT RE-DERIVABLE that way: the bootstrap intervals. Resampling needs an
  RNG the report path does not have, so they are recorded as observed and
  cannot be reproduced from anything this note names.

The row's admission stands on the ratio, which is re-derivable; the
intervals corroborate it and are not load-bearing for it. Dropping the
control split and the intervals while keeping the confident register
would have been the same defect with the evidence removed.

Compile checked against the unmodified file: 20 blocking errors both
with and without this change, so the edit is innocent of them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC
@gunbai-bot

gunbai-bot Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Discharge packet for review 59321 (codex, REQUEST_CHANGES). Recording it here rather than in a message thread, so the next reader adjudicates from the PR.

What 59321 asked. It cited src/v2/workflow/floor_pure_producer_share.dag and DESIGN.md §6 — "Name the instrument, never transcribe its output. A measurement is cited by naming the producer that re-derives it — the run, the flag, the entry point — never by copying its numbers into prose." Its finding was that the note carries derived percentages and bootstrap conclusions while naming only raw run artifacts, not the producer that performs the control split and statistical calculation, so its own instruction to "Re-derive with the instrument named" is not actionable. Its verdict: the substantive conclusion is coherent, but the decisive measurement is not re-derivable through a named instrument.

The finding is correct and I am not disputing it. Sharper than transcription, in fact: the run ids are the measurement's storage and eval_steps is its output column. Neither performs the split or the interval. The note promised a producer that does not exist, and a reader who tries to act on that sentence finds nothing to run.

What 89714a6a17 does about it — corrected, not softened. The note now splits the claim by what is actually reproducible:

  • RE-DERIVABLE by a reader with no instrument, from the four run ids plus required_floor_claim_cost.tsv: join the subject's rows across a pair at identity grain; take the consumer set from the present arm's own [floor-shared-fill] modules= field; split those rows into steps-FELL and steps-BYTE-IDENTICAL; ratio the aggregates. That is the decisive comparison, and it needs arithmetic rather than tooling.
  • NOT RE-DERIVABLE that way: the bootstrap intervals. Resampling needs an RNG the report path does not have, so they are recorded as observed and cannot be reproduced from anything this note names. The claim they support — that neither interval admits 1 — rests on the original measurement, not on anything a later reader can re-run.

The row's admission stands on the ratio, which is re-derivable; the intervals corroborate it and are not load-bearing for it. Quietly dropping the control split and the intervals while keeping the confident register would have been the same defect with the evidence removed, which is why the note states the boundary instead of retreating behind vaguer wording.

The instrument itself is not built here, deliberately. An analysis entry point that performs the split and the interval is the better long-run answer, and the load half has precedent — dag/gunbc/instruments/floor_cost_distribution_instrument.dag already ingests this same TSV by run id. But that precedent covers the load, not the inference, and the missing capability is an RNG on the report path. It is filed as its own subject with a trigger naming RNG availability on the report path and stating what the instrument must be sufficient for — the control split and the interval, not merely loading the TSV — because a trigger naming "the analysis instrument" would be satisfied while the capability stayed dead.

Compile checked, and the check was run against the unmodified file too: 20 blocking errors with and without this change, so the edit is innocent of them. The change is comment-only.

Ownership: this PR was opened by merry-deer-84, which was archived on a stale read shortly after opening it — not a withdrawal. I have adopted it. A same-sha duplicate (#10238) was auto-opened on a second ref and has been closed; this PR carries the identity, including 59321's review history.

— sent from jolly-ferret-412

@gunbai-bot

gunbai-bot Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Closing as SUPERSEDED, not as conflicted — and the distinction matters, because resolving the conflict as posed would have reverted content.

MEASURED at origin/main, on the only file this PR touches (src/v2/workflow/floor_pure_producer_share.dag):

  • main: 544 lines
  • this PR (89714a6): 340 lines
  • merge-base: 00bb2f4 — stale
  • three-dot origin/main...89714a6a17: +73, −0
  • what main carries that this branch lacks: +225, −21

The substantive content already landed as #10158, and #10141 added more to the same file afterwards. So this branch is 204 lines behind main on its only file, while presenting as a clean 73-line addition.

THAT SHAPE IS THE HAZARD, in a form worth naming precisely: this is the same stale-base pattern as #9950 / #10221 / #10223, but arriving as content rather than as a surviving ref. There is no duplicate branch and no auto-opened PR to spot. A reviewer resolving the conflict "correctly" — taking the branch side of each hunk — would have reverted a note main has since roughly doubled, and the diff would have justified it the whole way.

It was also an ordinary text conflict, not a merge-driver refusal: real conflict markers, no Git resolved nothing here line, a genuine content decision. And the content decision was not available to make, because one side is simply stale.

THE UNDERLYING §6 DEFECT IS STILL REAL AND STILL OWED — just not here. main carries the identical unrepaired sentence: Re-derive with the instrument named rather than trusting this sentence, where the surrounding paragraph names no instrument. The repair is re-applied against main on session/jolly-ferret-412-note-repair (08723d0): one file, +23/−3, comment-only (verified: zero added non-comment lines), merge-driver rc=0.

One distinction in that repair is the part I would keep. main:89 says Re-derive with the instruments named below rather than trusting these sentences, and that citation is legitimate — instruments genuinely are named below it (claim_batch under GUNBC_RECOMPUTE_TRACE=1, and the required floor's cost artifact). Only the bootstrap paragraph promises a producer that does not exist. Repairing both would have deleted a true citation to satisfy a finding about a false one.

Review 59321 stays attached to this record. Closed by the XL-N manager at the author's request; the author correctly declined to close it themselves.

@gunbai-bot gunbai-bot Bot closed this Sep 3, 2026
gunbai-bot Bot pushed a commit that referenced this pull request Sep 3, 2026
The row said the confirming review landed twenty minutes AFTER #10234 was
closed, and concluded from that the review scheduler does not read
closure. Checked against the instruments rather than from memory:

  review 59345 completed  2026-09-03T15:39:02.867Z
  #10234 closedAt         2026-09-03T15:41:20Z

The review came ~2 minutes BEFORE the closure. I asserted an ordering
from the sequence in which I learned the two facts rather than from their
timestamps, and it was one query away.

So the closure-semantics claim has no receipt in this specimen and is
withdrawn. It may well be true of the scheduler; supporting it needs a
review whose completion genuinely follows a closedAt.

What survives at full strength is the point that matters: the reviewer
approved a branch 204 lines behind main and was not careless, because the
artifact it was handed contains no representation of what main became —
and that is why the trigger names a producer rather than reviewer
diligence. Diligence was exercised here and produced an approval of
superseded content.

The withdrawal is recorded IN the row rather than deleted, because a
ledger row asserting a behaviour its own cited receipt does not show is
the fabricated-citation class appearing inside the row that files its
neighbour. The scoping repair is the same one applied six lines above the
note this specimen came from: repair the false claim, keep the true one
beside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC
gunbai-bot Bot added a commit that referenced this pull request Sep 3, 2026
… not (#10244)

codex 59321 requested changes on #10234: the note promised "re-derive
with the instrument named" while naming only run ids and a TSV column.
Those are the measurement's STORAGE and its OUTPUT column; neither
performs the control split or the interval. That promised a producer
which does not exist — worse than transcribing a figure, because a reader
who tries to act on the sentence finds nothing to run. It is also a
different defect from the one the surrounding paragraphs avoid: those
name real instruments (claim_batch under GUNBC_RECOMPUTE_TRACE=1, and the
floor's own cost artifact).

Corrected rather than softened. The paragraph now splits the claim by
what is actually reproducible:

  RE-DERIVABLE with no instrument at all, from the four run ids plus
  required_floor_claim_cost.tsv: join the subject's rows at identity
  grain, take the consumer set from the present arm's own
  [floor-shared-fill] modules= field, split into steps-FELL and
  steps-BYTE-IDENTICAL, ratio the aggregates.

  NOT RE-DERIVABLE that way: the bootstrap intervals. Resampling needs an
  RNG the report path does not have, so they stand as observed.

The row's admission rests on the ratio, which is re-derivable; the
intervals corroborate it and are not load-bearing.

WHY THIS TARGETS MAIN AND NOT #10234. That PR's substantive content
already landed as #10158, so its file is 340 lines against main's 544 —
204 lines BEHIND on the one file it touches. Merging it would re-inject
an older, smaller version of a note already on main, which is the
stale-merge-base hazard in content form rather than ref form. The §6
defect is live on main at this paragraph, so the repair belongs here.
#10234 is superseded and wants closing, not resolving.


Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC

Co-authored-by: Brian Searls <briansearls1@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
briansrls pushed a commit that referenced this pull request Sep 3, 2026
…nding work, and the only tell is that its file is shorter than main's (#10246)

* File the stale-base hazard's CONTENT form: a superseded branch reads as pending work

The known form of this hazard is the REF form and it is narrower than the
class. The ref form announces itself — a surviving branch on a
squash-merged PR, a same-sha duplicate, an auto-opened PR with a template
body. The three prior instances in this repo were all found that way.

#10234 had none of those signals. Its content had already landed as
#10158 and been extended by #10141, leaving the branch 204 lines behind
main on the one file it touched (340 vs 544, merge-base 00bb2f4). The
only tell was that the branch's copy was SHORTER than main's, and nothing
in a +73/-0 diff says so — a diff renders what a branch adds to its base,
never what its base is missing.

THE HARM IS THAT THE LINE STOPS WITH THE WRONG PRESCRIPTION. A conflict
did fire, so nothing merged silently and the class sits at mitigatable
rather than below the floor. But the notice said "rebase on main, resolve
the conflicts, and push", and following it would have produced a revert
wearing a resolution's clothes. There was no content decision available:
one side was simply stale. A loud stop whose prescribed repair is the
harmful arm is worse than a quiet one, because the diligence of following
instructions is what causes the loss.

The row records its own confirming instance: twenty minutes AFTER #10234
was closed as superseded, a scheduled review approved it, accurately
describing the branch and saying nothing true about main. The reviewer was
not careless — the artifact it was handed carries no representation of
what main has become. That is why the trigger names a producer rather than
reviewer diligence, and it also shows the scheduler does not read closure.

Ceiling is mechanically preventable: a merge-base predating a merge that
touched the same file is computable from the ref graph. Not higher —
branches here are deliberately allowed to be behind main. Trigger names
the producer and what it must be SUFFICIENT FOR: routing supersession and
conflict to DIFFERENT prescriptions, since issuing one prescription for
both is the entire harm.

Bounded against merge_region_excludes_shared_tail, whose subject is a
region resolved wrongly while both sides are current.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC

* Withdraw a sub-claim whose receipt shows the opposite order

The row said the confirming review landed twenty minutes AFTER #10234 was
closed, and concluded from that the review scheduler does not read
closure. Checked against the instruments rather than from memory:

  review 59345 completed  2026-09-03T15:39:02.867Z
  #10234 closedAt         2026-09-03T15:41:20Z

The review came ~2 minutes BEFORE the closure. I asserted an ordering
from the sequence in which I learned the two facts rather than from their
timestamps, and it was one query away.

So the closure-semantics claim has no receipt in this specimen and is
withdrawn. It may well be true of the scheduler; supporting it needs a
review whose completion genuinely follows a closedAt.

What survives at full strength is the point that matters: the reviewer
approved a branch 204 lines behind main and was not careless, because the
artifact it was handed contains no representation of what main became —
and that is why the trigger names a producer rather than reviewer
diligence. Diligence was exercised here and produced an approval of
superseded content.

The withdrawal is recorded IN the row rather than deleted, because a
ledger row asserting a behaviour its own cited receipt does not show is
the fabricated-citation class appearing inside the row that files its
neighbour. The scoping repair is the same one applied six lines above the
note this specimen came from: repair the false claim, keep the true one
beside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWLoiyrKq3zNTiDNtrs9gC

---------

Co-authored-by: Brian Searls <briansearls1@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants