Repository navigation
Restore the shape777 report schema the committed evidence and verifier read - #46
Open
pxtroniwnl wants to merge 1 commit into
Open
pxtroniwnl wants to merge 1 commit into
pxtroniwnl wants to merge 1 commit into
Conversation
…r read benchmarks/shape777.py no longer produced a report verify_published.py can consume. It emitted version shape777-published-v1, modes fresh and parallel_shared, and decisions_per_second, while results/raw/shape777-direct.json and the verifier at verify_published.py:89-96 use shape777-direct-v1, modes fresh_batch1 and parallel_suffix, and judgments_per_second. The verifier zips summary claims against raw results positionally, so a fresh run no longer lined up with the claims it is meant to reproduce. The rename also dropped evidence fields rather than renaming them: states, judgments, the p95 latency, cache_hits, padded_suffix_tokens and true_suffix_tokens were lost, as were passed, mean_probability_difference, tolerance and the non-gating note in comparisons_to_fresh, plus the top-level max_tokens and timing_scope. shape777_reranker.py had drifted the same way, losing max_tokens, timing_scope, states and yes_no_pair_forwards and reporting version shape777-reranker-published-v1. Restore both scripts to the published v1 contract and declare it as module constants so the schema is checkable. The report labels the unbatched mode fresh_batch1 while its committed prediction rows say fresh, so the two vocabularies stay distinct instead of being unified. comparisons_to_fresh.passed cannot be reconstructed from the committed data, where it is false under either reading, so it is defined as exact reproduction of fresh, which is the only interpretation consistent with the note stating the tolerance is non-gating. mean_probability_difference is the mean over option cells rather than over decisions; both were confirmed against the committed row-level predictions. tests/test_benchmark_schema.py pins the committed reports to those constants and fails if the scripts drift again, and checks that the keys verify_published.py reads are still emitted. All six tests fail against the previous scripts. No committed evidence is touched: results/raw, results/raw/SHA256SUMS and results/phase1-summary.json are unchanged, and verify_published.py still reproduces all 69 claims.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
benchmarks/shape777.pyandbenchmarks/shape777_reranker.pyno longer emit theschema that the committed evidence uses.
results/raw/shape777-direct.jsonandresults/raw/shape777-reranker.jsonwere produced by an earlier, richer versionof these scripts, and in the meantime the generators lost fields, mode names and
top-level keys.
The important part is not that the JSON looks different. It is that
benchmarks/verify_published.pyreads those raw files positionally andre-derives all 69 headline claims from them, while the scripts that are supposed
to regenerate them had drifted. The published numbers could no longer be
reproduced from a run of the repository as it stands, and nothing detected it:
the verifier checks the committed bytes, never the generator.
This restores the contract the evidence was actually written against, and adds
the first test that ties the two together.
What changed
benchmarks/shape777.pyshape777-direct-v1fresh_batch1,serial_prefix,parallel_suffix,matching the evidence
COMPARISON_TOLERANCE/COMPARISON_NOTE, and the per-mode and top-level fieldsets that the verifier's positional reads expect
judgments_per_second, p50/p95 latency, and thestatesfieldcache_hitsfor the suffix mode and true suffix token counts, so thethroughput denominator is the number of real decoded tokens rather than a
padded one
benchmarks/shape777_reranker.pyshape777-reranker-v1timing_scope,states,max_tokens,yes_no_pair_forwardsTwo judgement calls worth reviewing
passedmeans exact reproduction, not "within tolerance". The schema has apassedflag per mode and a tolerance of 1.0, and the tolerance is explicitlydocumented as non-gating. I therefore computed
passedas bit-exact agreementbetween compared cells and kept the tolerance as a recorded, non-gating field.
If the intent was for
passedto be tolerance-based, that is a one-line change,but it would change the meaning of a headline field and I did not want to make
that call silently.
mean_probability_differenceis the mean over cells, not over tokens. Iverified this empirically against the committed evidence rather than guessing:
the per-cell means reproduce the published value, per-token weighting does not.
Verification
tests/test_benchmark_schema.pyadds 6 tests that assert the emitted documentagainst the committed evidence. All 6 fail against the current scripts and pass
with the restored ones, which is the property that matters here: they encode the
contract rather than restating the implementation.
Also checked by hand: both scripts compile,
--helpworks, the create-onlyguard still refuses to overwrite an existing output path, and the
37 states x 21 options fixture validates.
pytest -qon this branch: 71 passed, 3 skipped.python benchmarks/verify_published.py: 69 claims verified.cd results/raw && sha256sum -c SHA256SUMS: 23 of 23 OK, zero failures.Limitations
with 6 GB of VRAM, which cannot hold Qwen3.5-4B in BF16, and I did not install
torchbecause the CUDA wheel alone is 5-6 GB against 7.2 GB free. Therestored code is therefore validated against the committed evidence and by
static checks, not by a fresh run. Someone with adequate hardware should
confirm the regenerated report matches the committed one before this merges.
shape777-direct.predictions.jsonlalso carriesbatch_size,padded_tokens,allowed_token_mass,answer_token_idsandfull_vocab_argmax_id, and the currentscore_sharedresult does not emitthem. I left that alone to keep this PR to the report contract, but the drift
is not fully closed by this change and is worth a follow-up.
Add Apple Metal (MPS) and CPU support with auto device selection #6, which also touch
benchmarks/shape777.py. Expect a rebase.No files under
results/are touched, no checksum changes, and no reportednumber is altered by this PR.