Skip to content

Distinguish predictive claims and retain falsification-driven narrowing - #11113

Merged
briansrls merged 15 commits into
mainfrom
session/proud-otter-590
Sep 12, 2026
Merged

briansrls merged 15 commits into
mainfrom
session/proud-otter-590

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Local execution on source head 1979a0a011f passed all 21 behavior/render controls, with two mutation RED/green counterexamples: suppressing falsifier recognition and restoring the old hard-coded spare outcome each caused its relevant test to fail with exit 1; restoring production source made each pass. The constructor probe also passed three real compiler-forgery refusals and its readable-model control. Both machine-intake CLIs executed successfully before the final duration-join change, and the current 21-control probe exercises their rendering. Compiler: locally built ARM64 4acc40347e5, with no compiler-source changes between that revision and this head. Scoped execution details.

CI attempt 2 evaluated the unchanged head and passed. Run 34667218809, attempt 2 reports FloorClean, planned=3779, executed=3779, not_attempted=0, claims_failed=0, and unexpected_failures=0. The log records the predictive-claim witness identities in the evaluation fold. Build, regeneration, floor, and aggregate checks all passed. Review 64256 approves this head with zero current request changes.

Attempt 1 remains a separate pre-evaluation capacity refusal, not evaluated-and-red. required-witnesses-floor refused MemoryStallRefusedPageThrash; GitHub displayed that refusal and its dependent aggregate as two FAILUREs. No claim in this PR was evaluated by that attempt either way. Its preserved artifact is gunbc-floor-outcome-34667218809-1-required-witnesses-floor (artifact, receipt details). The successful unchanged-head retry does not erase that refusal or establish uniform runner capacity.

A vendor support table can be cited correctly while a derived hardware prediction overclaims what it establishes. Predictive claims now carry their factual basis, falsifier, test receipts, registration and confound standing, and a continuing testing obligation. Constructor admission requires recording before test start; a counterexample blocks further testing until a distinct, strictly narrower revision retains the killed version and its receipt.

The machine-intake consumers cover both reports from September 11:

  • The 32-DIMM restart is a weak, confounded conversation-only match, with no invented preregistration or deadline.
  • The spare-screen prediction's original identity a47ce298890ac8dca72cda501b8a1225fcb740ec was amended to published identity e72f8930cf647ef15abed0da676853e2c3ce51c2. Review 64186 exposed the dangling original citation. Its unchanged object is now retrievable through refs/tags/evidence/mtcollins1-spare-screen-a47ce298890; a fresh empty-repository fetch verified the commit and artifact blob. The specimen uses the separate PostOutcomeArchivedPrediction arm, carrying corrected identity history, the distinct author/committer times, mixed commit content, and ordering testimony. Post-outcome archival publication earns neither ordinary committed-snapshot nor preregistered-match standing. Its reported restart falsifies shape sufficiency. The replacement scope retains the reported original module set and remains untested; shape necessity and faulty-module versus combination-effect attribution remain unestablished.

The CLI reports preserve standing, confounds, superseded predictions, and the next obligation. The census and sibling failure-mode attachment is explicitly pending #10965, which is unmerged; no unresolved declaration reference or duplicate sibling row is introduced. DESIGN §4b's compiler guarantee ladder remains a separate axis.

The first scoped run exposed the existing callable-field generic-substitution stall: PredictionRestriction<S>.admits retained S on five concrete configuration field accesses. The consumers now use the corpus's declared named, typed predicate form, explicitly authorized by the parent; the existing stall row gains this measured specimen and inference trace. No compiler semantics change is included. Refusal receipt.

Canonical exhaustive folds own verdict, registration, discrimination, and observation-window elimination. The standard queries and CLI readers consume those folds. The boundary with std.claim_evidence is explicit: raw protocol reports lack the grounded inference and PROV inputs required to become admitted descriptive evidence; test consistency does not manufacture an EvidenceSupports link or claim readiness.

The spare-screen restart classification derives from the reported duration and the prediction’s window through its canonical fold. A 400-second restart, a shortened window, or a missing window cannot be authored as an in-window falsifier. The approximate historical window stays approximate and earns no witnessed preregistration standing.

@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 12, 2026 00:26
@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Addressed review 64186 in 7b52cfa. The original commit existed locally but was detached from published branch history; that was a real retrieval defect.

The unchanged original object is now published at refs/tags/evidence/mtcollins1-spare-screen-a47ce298890. A fresh empty repository fetched that ref from origin and resolved commit a47ce298890ac8dca72cda501b8a1225fcb740ec; the prediction file's blob is b8f78347192f92c570ac57b90a0d22bd0e949067, identical to the original local object. Read the original artifact at that commit.

The machine-intake CLI now carries this retrieval location, and the failure-mode receipt records the review finding and repair. Both explicitly identify this as post-outcome archival publication, which supplies retrievability but does not prove pre-outcome publication. No timestamp or original content was rewritten. The classification remains CommittedPredictionWithoutWitnessedOrder; it is not upgraded to witnessed preregistration. The path is read at the pinned commit, not assumed to exist in the current checkout.

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Further correction after the parent explained the amendment: 971afc1 introduces PostOutcomeArchivedPrediction as a distinct registration arm. The source now retains both original and published identities, author versus amended committer time, the unrelated refusal fix in the same commit, and the parent's ordering testimony. Archival matches have a separate count and cannot earn ordinary committed-snapshot or preregistered-match standing. The specimen uses this weaker arm.

The original object remains retrievable through the evidence tag verified above. Updated behavioral, construction, and CLI probes are running; this comment does not claim they have passed yet.

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Addressed review 64208 in 93876dc through the explicit §3b/§4c boundary the review permits. I checked the full std.claim_evidence carrier and admission implementation before choosing this resolution.

The source annotations now explain why these protocol receipts are not automatically EvidenceLink/EvidenceInference values: the historical inputs provide reported outcomes and partial ordering/identity evidence, not grounded PROV entity/activity/agent, inference-rule, fidelity, and whole-link admission evidence. Synthesizing those dimensions to obtain a link would invent provenance. A matching restart remains a test-protocol match with confounds; it is not automatically EvidenceSupports for the physical claim.

The annotations also locate each adjacent concept:

  • PredictionTestDiscrimination is domain vocabulary suitable for the canonical link's open Independence parameter when a grounded link exists.
  • Registration adds publication/start/outcome ordering and artifact-identity limitations; these are not expressed by PROV IDs or freshness's maximum conclusion.
  • The test verdict compares the reported outcome with the declared expected/falsifying outcomes. Descriptive inference direction and admission insufficiency remain owned by std.claim_evidence.
  • Predictive standing preserves test counts and narrowed revision history. It is not the four-valued support/challenge assessment, an evidence-admission result, or a ClaimReadinessReceipt.

The boundary expressly requires future consumers that admit descriptive evidence or establish readiness to supply the missing grounds and use the canonical EvidenceLink/EvidenceInference and admission policy. No automatic match-to-support adapter is provided.

This commit changes standalone source annotations only; executable behavior remains covered by the passing 20-control scoped run. Pre-push formatting passed. CI and renewed review are still required for the current head.

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Addressed review 64222 by routing the readers through canonical exhaustive folds:

  • prediction_test_verdict_fold owns verdict elimination.
  • prediction_registration_fold owns registration elimination.
  • prediction_test_discrimination_fold owns test-design elimination.

predictive_falsifiers, receipt/commit chronology admission, all match-strength counts, and the CLI renderer now derive their results through these folds. There is one direct match per coproduct across the standard carrier and renderer, with no wildcard branch in those authorities. A new variant therefore requires an explicit handler at its canonical fold.

The existing 20 behavioral/rendering controls, constructor-forgery probe, and both CLI consumers are rerunning remotely with the kernel-enforced 8 GiB limit. No passing result is claimed for this refactor yet.

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Addressed review 64241 in 1979a0a. The actual archived-receipt consumer now calls spare_screen_restart_outcome with the reported restart duration and the prediction’s own bound. The canonical window fold derives the outcome; a later restart or absent window yields SpareScreenOutcomeUnknown. The historical approximate window stays approximate and does not earn preregistration standing.

The expanded 21-control probe passed locally. Its new control sends a 400000 ms report through the actual model consumer and requires one inconclusive observation, zero matches and zero narrowed revisions. It also checks missing and shortened windows. The existing 290000 ms specimen still falsifies and preserves the narrowed lineage. A deliberate reintroduction of the hard-coded outcome is running as a RED control.

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Current source head 1979a0a011f passes all 21 behavior/render controls. Local ARM64 compiler: built from 4acc40347e5, confirmed by --version; no compiler-source changes between those heads. Existing kernel memory.max: 33578549248 bytes. No budget override or cgroup modification.

The review fixes now have executing counterexamples:

  • Review 64222: suppressing falsifier recognition made falsification_blocks_further_testing_until_narrowing report FAIL and exit 1; restoring the source made it report PASS and exit 0.
  • Review 64241: replacing the derived outcome with the former RestartInsideDeclaredWindow literal made a_restart_outside_the_reported_window_cannot_falsify_the_spare_screen report FAIL and exit 1; restoring the derived join made it report PASS and exit 0.

The constructor probe also passed its three actual compiler refusals and readable-model control; both machine-intake CLIs executed successfully before the duration-join change, and their rendering remains exercised by the current 21-control probe. The tree is restored and clean. CI run 34667218809 and a new-head review remain outstanding; these scoped results do not claim full-floor execution.

— sent from proud-otter-590

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

CI run 34667218809, attempt 1 passed build and regeneration but refused the witness floor before claim evaluation. The saved gunbc-floor-outcome-34667218809-1-required-witnesses-floor artifact reports class=infra, signature=MemoryStallRefusedPageThrash, exit=1.

The terminal-ledger authority could not resolve while typechecking v2.workflow.floor2_prepared_subject: runner srv4-14, admitted budget 16106127360 bytes from cgroup memory.high, 256786 major faults/minute with 4% CPU share over 37917 ms. No claim-cost or admission-roster receipt was produced. This is not a semantic verdict on the diff, and the local scoped passes do not replace the missing full-floor result.

Review 64256 approves the current head, with no current request changes. Merge remains blocked. The attempt logs and artifact are preserved before any rerun; a capacity/placement remedy has been requested from the parent. No refusal threshold or source behavior has been relaxed.

— sent from proud-otter-590

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

Clarification to the preserved attempt receipt: no CI claim verdict exists for this head; its claims are UNEVALUATED by the required floor, not evaluated-and-red. The failed GitHub check records a pre-evaluation infrastructure refusal. The separately reported scoped controls did execute and passed.

The parent reports an independent run still in strict preparation at 17.0 GiB RSS / 18.7 GiB cgroup charge, beside this attempt’s 15 GiB admitted budget. That supports a floor-memory deficit, not merely transient contention; waiting for a quieter runner is not an established remedy. The parent has escalated the remedy to the operator under msg_b0e46972-acec-4397-92d6-b6b1a0b0097b. This lane will remain idle: no CI retry, threshold change, or carrier edit to evade the refusal.

— sent from proud-otter-590

@gunbai-bot

gunbai-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor Author

The unchanged 1979a0a011f head obtained an evaluated green in CI attempt 2. The floor reports FloorClean, 3779 planned / 3779 executed / 0 not attempted, 0 failed claims, and 0 unexpected failures. Predictive-claim witness identities appear in the evaluation log. All four CI checks are successful; GitHub reports mergeStateStatus=CLEAN.

Review 64256 approves this head, with zero current request changes. Attempt 1’s pre-evaluation refusal is preserved and remains a distinct outcome. No source, closure, or threshold changed for this retry.

Ready for operator merge; this lane will not merge it itself.

— sent from proud-otter-590

@briansrls
briansrls merged commit 2abbb83 into main Sep 12, 2026
6 of 8 checks passed
@briansrls
briansrls deleted the session/proud-otter-590 branch September 12, 2026 04:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant