Skip to content

Add numerical and logical-state conformance support - #10

Draft
Terrydaktal wants to merge 2 commits into
magiccodingman:mainfrom
Terrydaktal:feat/numerical-state-conformance-support
Draft

Terrydaktal wants to merge 2 commits into
magiccodingman:mainfrom
Terrydaktal:feat/numerical-state-conformance-support

Conversation

@Terrydaktal

@Terrydaktal Terrydaktal commented Sep 17, 2026 •

Copy link
Copy Markdown

Problem and scope

Differences between serial target decode and speculative verification are difficult to localize when comparisons stop at generated text. This adds the numerical and logical-state conformance support used in the published M1/M8 investigation, following the invitation in #8.

The package lives under benchmarks/conformance and leaves serving defaults unchanged. It preserves the namespace and source identities used by the existing sealed workers. The companion repair submission #11 supplies the actual pinned native adapters and replay drivers, including common-input native stage substitution, explicit cache restoration and a checked bridge back to graph-enabled release outputs.

Included

  • Aligned top-1/10/20 comparison: membership, ordering, overlap, retained scores, inclusive boundary ties and full-logit digests are separate measurements.
  • Explicit logical-state and tentative-output checks, negative controls, source/binary binding, exclusive GPU leases and checkpointed worker supervision.
  • Eager/compiled execution-mode admission, isolated operator comparison, first-divergence evidence and source-bound numerical interventions.
  • Small scoped control-plane proofs, with unsupported or incomplete evidence kept distinct from a pass.
  • Synthetic CPU regressions, a declared GDN arithmetic contract, source provenance and reproducible uv dependencies. Private traces and aggregate publication receipts remain separate.

This is initially a draft to review the package boundary and retained dependency closure. It is not a claim that finite replay proves correctness for every input.

Validation

The support suite and companion repair adapters passed 412 CPU tests, with one skipped. The skip requires a separately retained compiler artifact; it does not count as native validation. Source provenance and syntax checks passed. No private Pi text or token fixture is included.

uv sync --project benchmarks/conformance --group dev
uv run --project benchmarks/conformance pytest benchmarks/conformance/tests -q

Before: original eager M1 versus original compiled M8. After both fixes: final eager M1 versus final compiled M8 (Fix 1 + Fix 2). Each revision has a fresh eager M1 reference; all four runs use the same 23 Pi continuations and full BF16 target head. Compiled M8 uses Inductor and piecewise GPU graphs.

Same set means the same tokens regardless of order. Same order also requires identical ranking. Mean shared is the average number of shared tokens per position.

Prediction Before: same set Before: same order Before: mean shared After both fixes: same set After both fixes: same order After both fixes: mean shared
Top 1 9,838 / 10,000 (98.38%) 9,838 / 10,000 (98.38%) 0.9838 / 1 10,000 / 10,000 (100%) 10,000 / 10,000 (100%) 1 / 1
Top 10 4,878 / 10,000 (48.78%) 761 / 10,000 (7.61%) 9.3516 / 10 10,000 / 10,000 (100%) 10,000 / 10,000 (100%) 10 / 10
Top 20 2,313 / 10,000 (23.13%) 4 / 10,000 (0.04%) 18.6883 / 20 10,000 / 10,000 (100%) 10,000 / 10,000 (100%) 20 / 20

After both fixes, full-vocabulary hashes match at 10,000/10,000 decode positions and 23/23 initial-prefill predictions. Top-1/10/20 retained scores and inclusive boundary-tie sets also match throughout.

This is an all-seven-accepted forced replay on a pinned configuration; it does not qualify arbitrary dependency updates or every rejection width. The separate 320-position eager-M8/compiled-M8 and isolated-stage measurements remain in the report.

Technical report, original/fixed comparisons, complete compiled stage timing table and aggregate evidence.

Import the source-bound support used to diagnose the D7 M1/M8 divergence. Separate top-1/10/20 membership, ordering, retained scores, tie completeness and full-row digests; reject empty, misaligned, incomplete or changed evidence rather than reporting a misleading pass.

Preserve explicit logical state, tentative-output checks, negative controls, artifact/source binding, GPU exclusion leases, checkpointed worker supervision and scoped control-plane proofs. Keep raw private traces distinct from aggregate reports and preserve module identities required by existing qualification bundles.

Package the dependency closure under benchmarks/conformance with an explicit GDN arithmetic contract, source hashes, synthetic CPU regressions, reproducible uv dependencies and an optional CPU-only tensor test extra. Ignore virtual environments, binary outputs and private trace files. Document the execution order, inputs, outputs and distinction between tested samples and mathematical proof.

Validation: 145 support CPU tests passed. Combined with the companion repair adapters, 328 CPU tests passed with one explicitly skipped retained-compiler-artifact check. Syntax, source provenance and publication file-scope checks passed. The separate published pinned-stack GPU result covers 10,000 aligned decode positions and 23 prefills; this import does not alter serving defaults.
Signed-off-by: Terrydaktal <9lewis9@gmail.com>
…support

Record the effective Inductor precision-cast setting as well as its environment
flag, so a requested rounding change cannot be mistaken for an observed one.
Admit execution-mode comparisons only when the fixture, capacity, repairs,
compiler/runtime sources and all other numerical controls remain bound.

Add explicit BF16-cast and native RoPE nearest-even interventions while retaining
both original and normalized receipt identities. Locate attention sub-boundaries
and compare common-input stage substitutions with checked remainder execution;
reject missing state, incomplete coverage and unobserved or unrelated changes.

Include the CPU regressions and complete support dependency closure needed by
the companion native replay drivers. Point documentation at the public report
and distinguish the compiled M1/M8 repair from the later eager/compiled repair.

Validation: combined support and companion repair suite: 412 passed, one skipped.
The skipped check needs a retained compiler artifact and is not native validation.

Signed-off-by: Terrydaktal <9lewis9@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant