Repository navigation
Distinguish predictive claims and retain falsification-driven narrowing - #11113
Conversation
|
Addressed review 64186 in 7b52cfa. The original commit existed locally but was detached from published branch history; that was a real retrieval defect. The unchanged original object is now published at The machine-intake CLI now carries this retrieval location, and the failure-mode receipt records the review finding and repair. Both explicitly identify this as post-outcome archival publication, which supplies retrievability but does not prove pre-outcome publication. No timestamp or original content was rewritten. The classification remains |
|
Further correction after the parent explained the amendment: 971afc1 introduces The original object remains retrievable through the evidence tag verified above. Updated behavioral, construction, and CLI probes are running; this comment does not claim they have passed yet. |
|
Addressed review 64208 in 93876dc through the explicit §3b/§4c boundary the review permits. I checked the full The source annotations now explain why these protocol receipts are not automatically The annotations also locate each adjacent concept:
The boundary expressly requires future consumers that admit descriptive evidence or establish readiness to supply the missing grounds and use the canonical This commit changes standalone source annotations only; executable behavior remains covered by the passing 20-control scoped run. Pre-push formatting passed. CI and renewed review are still required for the current head. |
|
Addressed review 64222 by routing the readers through canonical exhaustive folds:
The existing 20 behavioral/rendering controls, constructor-forgery probe, and both CLI consumers are rerunning remotely with the kernel-enforced 8 GiB limit. No passing result is claimed for this refactor yet. |
|
Addressed review 64241 in 1979a0a. The actual archived-receipt consumer now calls The expanded 21-control probe passed locally. Its new control sends a 400000 ms report through the actual model consumer and requires one inconclusive observation, zero matches and zero narrowed revisions. It also checks missing and shortened windows. The existing 290000 ms specimen still falsifies and preserves the narrowed lineage. A deliberate reintroduction of the hard-coded outcome is running as a RED control. |
|
Current source head The review fixes now have executing counterexamples:
The constructor probe also passed its three actual compiler refusals and readable-model control; both machine-intake CLIs executed successfully before the duration-join change, and their rendering remains exercised by the current 21-control probe. The tree is restored and clean. CI run 34667218809 and a new-head review remain outstanding; these scoped results do not claim full-floor execution. — sent from proud-otter-590 |
|
CI run 34667218809, attempt 1 passed build and regeneration but refused the witness floor before claim evaluation. The saved The terminal-ledger authority could not resolve while typechecking Review 64256 approves the current head, with no current request changes. Merge remains blocked. The attempt logs and artifact are preserved before any rerun; a capacity/placement remedy has been requested from the parent. No refusal threshold or source behavior has been relaxed. — sent from proud-otter-590 |
|
Clarification to the preserved attempt receipt: no CI claim verdict exists for this head; its claims are UNEVALUATED by the required floor, not evaluated-and-red. The failed GitHub check records a pre-evaluation infrastructure refusal. The separately reported scoped controls did execute and passed. The parent reports an independent run still in strict preparation at 17.0 GiB RSS / 18.7 GiB cgroup charge, beside this attempt’s 15 GiB admitted budget. That supports a floor-memory deficit, not merely transient contention; waiting for a quieter runner is not an established remedy. The parent has escalated the remedy to the operator under msg_b0e46972-acec-4397-92d6-b6b1a0b0097b. This lane will remain idle: no CI retry, threshold change, or carrier edit to evade the refusal. — sent from proud-otter-590 |
|
The unchanged Review 64256 approves this head, with zero current request changes. Attempt 1’s pre-evaluation refusal is preserved and remains a distinct outcome. No source, closure, or threshold changed for this retry. Ready for operator merge; this lane will not merge it itself. — sent from proud-otter-590 |
Local execution on source head
1979a0a011fpassed all 21 behavior/render controls, with two mutation RED/green counterexamples: suppressing falsifier recognition and restoring the old hard-coded spare outcome each caused its relevant test to fail with exit 1; restoring production source made each pass. The constructor probe also passed three real compiler-forgery refusals and its readable-model control. Both machine-intake CLIs executed successfully before the final duration-join change, and the current 21-control probe exercises their rendering. Compiler: locally built ARM644acc40347e5, with no compiler-source changes between that revision and this head. Scoped execution details.CI attempt 2 evaluated the unchanged head and passed. Run 34667218809, attempt 2 reports
FloorClean,planned=3779,executed=3779,not_attempted=0,claims_failed=0, andunexpected_failures=0. The log records the predictive-claim witness identities in the evaluation fold. Build, regeneration, floor, and aggregate checks all passed. Review 64256 approves this head with zero current request changes.Attempt 1 remains a separate pre-evaluation capacity refusal, not evaluated-and-red.
required-witnesses-floorrefusedMemoryStallRefusedPageThrash; GitHub displayed that refusal and its dependent aggregate as two FAILUREs. No claim in this PR was evaluated by that attempt either way. Its preserved artifact isgunbc-floor-outcome-34667218809-1-required-witnesses-floor(artifact, receipt details). The successful unchanged-head retry does not erase that refusal or establish uniform runner capacity.A vendor support table can be cited correctly while a derived hardware prediction overclaims what it establishes. Predictive claims now carry their factual basis, falsifier, test receipts, registration and confound standing, and a continuing testing obligation. Constructor admission requires recording before test start; a counterexample blocks further testing until a distinct, strictly narrower revision retains the killed version and its receipt.
The machine-intake consumers cover both reports from September 11:
a47ce298890ac8dca72cda501b8a1225fcb740ecwas amended to published identitye72f8930cf647ef15abed0da676853e2c3ce51c2. Review 64186 exposed the dangling original citation. Its unchanged object is now retrievable throughrefs/tags/evidence/mtcollins1-spare-screen-a47ce298890; a fresh empty-repository fetch verified the commit and artifact blob. The specimen uses the separatePostOutcomeArchivedPredictionarm, carrying corrected identity history, the distinct author/committer times, mixed commit content, and ordering testimony. Post-outcome archival publication earns neither ordinary committed-snapshot nor preregistered-match standing. Its reported restart falsifies shape sufficiency. The replacement scope retains the reported original module set and remains untested; shape necessity and faulty-module versus combination-effect attribution remain unestablished.The CLI reports preserve standing, confounds, superseded predictions, and the next obligation. The census and sibling failure-mode attachment is explicitly pending #10965, which is unmerged; no unresolved declaration reference or duplicate sibling row is introduced. DESIGN §4b's compiler guarantee ladder remains a separate axis.
The first scoped run exposed the existing callable-field generic-substitution stall:
PredictionRestriction<S>.admitsretainedSon five concrete configuration field accesses. The consumers now use the corpus's declared named, typed predicate form, explicitly authorized by the parent; the existing stall row gains this measured specimen and inference trace. No compiler semantics change is included. Refusal receipt.Canonical exhaustive folds own verdict, registration, discrimination, and observation-window elimination. The standard queries and CLI readers consume those folds. The boundary with
std.claim_evidenceis explicit: raw protocol reports lack the grounded inference and PROV inputs required to become admitted descriptive evidence; test consistency does not manufacture anEvidenceSupportslink or claim readiness.The spare-screen restart classification derives from the reported duration and the prediction’s window through its canonical fold. A 400-second restart, a shortened window, or a missing window cannot be authored as an in-window falsifier. The approximate historical window stays approximate and earns no witnessed preregistration standing.