Skip to content

Restore the shape777 report schema the committed evidence and verifier read - #46

Open
pxtroniwnl wants to merge 1 commit into
TheoLeeCJ:masterfrom
pxtroniwnl:fix/shape777-report-schema
Open

pxtroniwnl wants to merge 1 commit into
TheoLeeCJ:masterfrom
pxtroniwnl:fix/shape777-report-schema

Conversation

@pxtroniwnl

Copy link
Copy Markdown

Summary

benchmarks/shape777.py and benchmarks/shape777_reranker.py no longer emit the
schema that the committed evidence uses. results/raw/shape777-direct.json and
results/raw/shape777-reranker.json were produced by an earlier, richer version
of these scripts, and in the meantime the generators lost fields, mode names and
top-level keys.

The important part is not that the JSON looks different. It is that
benchmarks/verify_published.py reads those raw files positionally and
re-derives all 69 headline claims from them, while the scripts that are supposed
to regenerate them had drifted. The published numbers could no longer be
reproduced from a run of the repository as it stands, and nothing detected it:
the verifier checks the committed bytes, never the generator.

This restores the contract the evidence was actually written against, and adds
the first test that ties the two together.

What changed

benchmarks/shape777.py

  • report version shape777-direct-v1
  • mode names restored to fresh_batch1, serial_prefix, parallel_suffix,
    matching the evidence
  • COMPARISON_TOLERANCE / COMPARISON_NOTE, and the per-mode and top-level field
    sets that the verifier's positional reads expect
  • judgments_per_second, p50/p95 latency, and the states field
  • cache_hits for the suffix mode and true suffix token counts, so the
    throughput denominator is the number of real decoded tokens rather than a
    padded one

benchmarks/shape777_reranker.py

  • report version shape777-reranker-v1
  • timing_scope, states, max_tokens, yes_no_pair_forwards

Two judgement calls worth reviewing

passed means exact reproduction, not "within tolerance". The schema has a
passed flag per mode and a tolerance of 1.0, and the tolerance is explicitly
documented as non-gating. I therefore computed passed as bit-exact agreement
between compared cells and kept the tolerance as a recorded, non-gating field.
If the intent was for passed to be tolerance-based, that is a one-line change,
but it would change the meaning of a headline field and I did not want to make
that call silently.

mean_probability_difference is the mean over cells, not over tokens. I
verified this empirically against the committed evidence rather than guessing:
the per-cell means reproduce the published value, per-token weighting does not.

Verification

tests/test_benchmark_schema.py adds 6 tests that assert the emitted document
against the committed evidence. All 6 fail against the current scripts and pass
with the restored ones, which is the property that matters here: they encode the
contract rather than restating the implementation.

Also checked by hand: both scripts compile, --help works, the create-only
guard still refuses to overwrite an existing output path, and the
37 states x 21 options fixture validates.

pytest -q on this branch: 71 passed, 3 skipped.
python benchmarks/verify_published.py: 69 claims verified.
cd results/raw && sha256sum -c SHA256SUMS: 23 of 23 OK, zero failures.

Limitations

No files under results/ are touched, no checksum changes, and no reported
number is altered by this PR.

…r read

benchmarks/shape777.py no longer produced a report verify_published.py can
consume. It emitted version shape777-published-v1, modes fresh and
parallel_shared, and decisions_per_second, while results/raw/shape777-direct.json
and the verifier at verify_published.py:89-96 use shape777-direct-v1, modes
fresh_batch1 and parallel_suffix, and judgments_per_second. The verifier zips
summary claims against raw results positionally, so a fresh run no longer lined
up with the claims it is meant to reproduce.

The rename also dropped evidence fields rather than renaming them: states,
judgments, the p95 latency, cache_hits, padded_suffix_tokens and
true_suffix_tokens were lost, as were passed, mean_probability_difference,
tolerance and the non-gating note in comparisons_to_fresh, plus the top-level
max_tokens and timing_scope. shape777_reranker.py had drifted the same way,
losing max_tokens, timing_scope, states and yes_no_pair_forwards and reporting
version shape777-reranker-published-v1.

Restore both scripts to the published v1 contract and declare it as module
constants so the schema is checkable. The report labels the unbatched mode
fresh_batch1 while its committed prediction rows say fresh, so the two
vocabularies stay distinct instead of being unified.

comparisons_to_fresh.passed cannot be reconstructed from the committed data,
where it is false under either reading, so it is defined as exact reproduction
of fresh, which is the only interpretation consistent with the note stating the
tolerance is non-gating. mean_probability_difference is the mean over option
cells rather than over decisions; both were confirmed against the committed
row-level predictions.

tests/test_benchmark_schema.py pins the committed reports to those constants
and fails if the scripts drift again, and checks that the keys
verify_published.py reads are still emitted. All six tests fail against the
previous scripts.

No committed evidence is touched: results/raw, results/raw/SHA256SUMS and
results/phase1-summary.json are unchanged, and verify_published.py still
reproduces all 69 claims.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant