Skip to content

The .dag acceptance harness: staged front end, typed per-stage evidence, one named acceptance contract - #8821

Merged
briansrls merged 8 commits into
mainfrom
session/fierce-newt-478
Aug 22, 2026
Merged

briansrls merged 8 commits into
mainfrom
session/fierce-newt-478

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

What this is

The automated acceptance oracle for .dag: given a candidate source string, return typed evidence of how far it got and why it stopped. Stages mint evidence; a named acceptance contract mints acceptance. It is a pure function of a String with no transport of its own.

A serve route was considered and REFUSED

The work item this landed under was titled "expose the .dag compile chain over gunbc serve". That is not what this PR does, and the refusal is recorded here rather than left as a path silently not taken.

gunbc serve is a declared frozen quarry route. gunbc.roadmap_serve roadmap_serve_seam_restoration_note, mirrored verbatim at the Commands::Serve arm in src/v1/stage0/src/main.rs: "FROZEN MEANS: no new options, no new verbs, no new routes into the interpreter, no completion work, no flag or environment gate re-opening it" — and "if this route begins attracting completion work ... the freeze has been repealed by drift." Its declared successor is roadmap_serve_interpreted_scaffold, dissolving to emit-on-demand, and the active ROADMAP row "Serve the roadmap page from emitted native code, never the tree-walking interpreter" records that this trigger already fired in production (the 2026-08-08 srv1 two-hour page wedge). A route running a front end, infer, emit and a rustc subprocess per request is the heaviest possible new workload on the seam that wedged.

The requirement behind the title was amortizing the corpus load (~39s load against ~55ms of candidate work), and that is a batching fact, not an HTTP one: claim_batch / claim_executor already hold a module index built once at startup, so many candidates score in one process with no new interpreter surface. The dispositions were re-verified against the tree, not recalled.

The stage graph

source -> tokenize -> parse -> normalize -> resolve -> infer
                                                        |
                                        +---------------+
                                        v               v
                                      Eval          TranslateTo(target)
                                                        |
                                                        v
                                                 TargetCompile<Rust>

Eval and TranslateTo are siblings over one inferred tree; neither is later. TargetCompile is a fifth stage the compiler does not perform: v2.compiler.compile compile_inferred_translate is exactly emit(tree, target) and never invokes a target toolchain (verified in the tree, not assumed), so TranslateTo proves emission and nothing about whether the emitted source compiles.

Edges live in exactly one place — obligation_prerequisites — and every consumer reads them rather than re-spelling the order.

Two structural moves (§4b), not runtime checks

  • TargetCompileClaim carries the TranslationClaim it consumes. "Target-compile an emission nobody requested" has no spelling; the prerequisite is derived from the claim the obligation already holds.
  • The verdict is derived, never stored. DagAcceptanceReceipt carries contract identity, candidate identity (content hash of the source), and the executed rows. There is no verdict field to transcribe, so rung inflation on this harness's own output is unrepresentable rather than lens-caught.

The floor caught this PR nicknaming, and that is in the diff

The first cut minted its own TargetCompileAccepted — and v2.std.leaf_model_verification already declares TargetCompileVerdict = TargetCompileAccepted | TargetCompileRejected { diagnostic_code }. Whole-pool resolution meant my duplicate broke every existing consumer of the real one, and --required-floor refused with three of them named. The observation now defers to that authority for the candidate-facing decision (TargetCompileDecided { verdict }) and adds only the three states a live invocation can be that a modeled expectation cannot. ResolveObligation likewise collided with gunbc.ci_materialization, so the stage obligations are TokenizeStage … TargetCompileStage.

Accepted-clean and accepted-with-diagnostics are different constructors

A stage can return Accepted while carrying typed located diagnostics — DESIGN names that FrontierAccepted. If the receipt folded it into an ordinary pass, a frontier candidate would score as clean and every metric would inherit the blindness. StagePassedWithFrontier is a separate constructor, not an optional field, so a consumer matching StagePassed cannot absorb the advisory case; AcceptanceContractSatisfied carries the frontier stage labels.

Measured with it, 2026-08-21: fn add(x: Int, y: Int) -> Int { undefined_callee(a: x) } — a call to a name no declaration provides — passes resolve with no diagnostic of any severity, then passes infer carrying one, reason infer_grounding_not_derived. So it is the declared frontier state rather than a silent below-floor accept, and the harness reports it as such. Not claimed: that anything reports the missing declaration itself — the frontier diagnostic is about grounding.

Durations consume std.measure, and that forced a better ledger

Review 54506 caught _ms: Int scalars — a second representation of duration beside std/measure.dag. They are now Millisecond (Measure<Time, Milli, Nat>). The adoption forced a real change rather than a rename: Nat is a CommutativeSemiring, so it has addition and no subtraction, and a decrementing ledger would have needed an operation its own carrier does not provide. The ledger therefore counts spent, admission is spent + required against the declared total, and the refusal carries required, budget_total and already_spent instead of one derived remainder.

Facts kept apart (§5)

  • BlockedByPrerequisite (impossible) vs SkippedAfterDecisiveRejection (a fail-fast policy chose not to pay).
  • A candidate's declared CPU budget (a fact about the candidate) vs VerifierBudgetUnavailable (a fact about us). A budget exhaustion is Incomplete, never a rejection — otherwise a slow machine reads as a bad model.
  • Target-toolchain observations partition by subject: accepted/diagnosed are about the candidate; unavailable/invocation-refused are about us and never assert a verdict; an unparseable failure is ignorance and gets its own counted StageUndecided arm rather than being rounded toward rejection.

StageExecution therefore has four arms, not three; the fourth is argued on its carrier.

The front end was DECOMPOSED, not copied (§3)

ingested_resolved_module_from_source was a single Outcome<Node> for four stages, so it could not say which stage a candidate died at. Re-inlining the four calls beside it would have been a second copy of the order. Instead the order moved to v2.compiler.staged_front_end and that arrow is now a projection of it — a fold to the resolved node or the first refusal, reproducing bind_outcome's diagnostic threading (accumulate on accept, rejected_with_pending on refusal) and re-running nothing.

A bounded run (the harness asks for exactly the depth its contract needs) is a third completion, not an absent node — rendering it as "no resolved module" would be the empty-observation narrow.

Evidence — green by execution, with discriminating REDs

All runs on a locally built claim_batch (arm64, this container), against real source strings through the real chain.

Staged front end (6/6 PASS): including the discriminating pair — a candidate dying at parse lands at stage 2, a candidate dying at resolve lands at stage 4. If the decomposition carried nothing they would land on the same row.

Acceptance harness (11/11 PASS): contract satisfied; resolve death located at resolve; infer BlockedByPrerequisite behind a failed resolve with the verdict still Rejected; a stage the contract never asked for accounted as NotRequiredByContract; verifier-budget exhaustion asserted two-sidedly — the rows say VerifierBudgetUnavailable AND the verdict is Incomplete, not Rejected; the target-compile prerequisite derived from its own translation claim (compared by identity, not discriminant); the four observation arms mapping to passed/rejected/undecided/not-run; fail-fast skip present under InteractiveFailFast and absent under FullLedger; a translate row through the real compiler asserted determinate; and an accepted-with-diagnostics row landing as a frontier pass rather than a clean one.

The fifth stage, proven wet with real rustc: claim_batch --wet over dag_acceptance_rustc_wet_test, both green in one run on one host:

PASS rustc_accepts_valid_target_source
PASS rustc_diagnoses_invalid_target_source

Valid Rust is ACCEPTED and invalid Rust (x + true where i32 is required) is DIAGNOSED, by real rustc reading the source from stdin. Either half alone is satisfiable by an observer that always answers one way; the pair is not.

An earlier version of that pair failed honestly and is worth recording: -o /dev/null made rustc fail to create its temp dir, so valid Rust was refused and invalid Rust "passed" for the wrong reason. The pair caught it; a single-sided assertion would not have.

Measured while building this, reported rather than buried

  • fn f() -> Int { undefined_name } refuses at resolve, but fn f(x: Int) -> Int { undefined_callee(a: x) } does not — an unresolved callee passes resolve while an unresolved bare value name does not. Not touched here; it is a compiler fact this harness now makes visible per candidate.
  • TargetCompilePolicy was not minted. The acceptance rule has exactly one point today, and an axis with one inhabitant is a second name for a fixed rule. The next-rung trigger is on the carrier.

Rung honesty about this PR's own claims

  • The harness fold, the stage graph, the verdict derivation and the front-end decomposition: green by execution with the REDs above.
  • The observation→row mapping is proven at the mapping; that the real toolchain produces those observations is a separate wet claim, carried by dag_acceptance_rustc_wet_test and not enrolled in the hermetic floor, because the hermetic envelope refuses host effects (correctly — mocking it would score candidates against a fabricated exit status).
  • Eval and TranslateTo execution from a source candidate is unexercised end-to-end, and the reason is a real compiler fact rather than an omission: see the infer row note below. The fail-fast/full-ledger split is therefore proven at the decision function, not end to end, and that is stated on the test.

Measured on this tree, 2026-08-21, and recorded on the carrier rather than pinned as a test: for fn add(x: Int, y: Int) -> Int { x + y } ingested from source, tokenize/parse/normalize/resolve/infer all pass and TranslateTo(Rust) rejects. The tests assert DETERMINACY (the row is passed or rejected, never missing) rather than today's side — pinning the side would make this file a defect pin that reds the moment v2 gains the derivation it lacks, and what is worth protecting is that the harness always locates the answer.

…named acceptance contract

Given a candidate .dag source string, return typed evidence of how far it got
and why it stopped. Stages mint evidence; a named contract mints acceptance.

- v2.compiler.staged_front_end: one authority for the front-end order, minting
  per-stage evidence; ingested_resolved_module_from_source is now a projection
  of it (fold to the resolved node or the first refusal, reproducing
  bind_outcome's diagnostic threading), not a second copy of the order.
- v2.workflow.dag_acceptance: obligations with derived prerequisites, four-arm
  StageExecution (the fourth is undecided, for observed-but-unreadable), a
  derived verdict with no stored field, and one receipt shape with a policy
  parameter.
- extdeps.rustc + v2.workflow.dag_acceptance_rustc: the fifth stage. v2's
  TranslateTo is emit() and never invokes a toolchain, so target compilation is
  a separate obligation with a real rustc binding at the periphery.

A gunbc serve route was considered and refused against the frozen seam.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title Acceptance harness: expose the .dag compile chain over gunbc serve, returning typed per-stage evidence under a named acceptance contract The .dag acceptance harness: staged front end, typed per-stage evidence, one named acceptance contract Aug 21, 2026
gunbc-ci-auto-heal and others added 2 commits August 21, 2026 21:57
…rrier

The probes answered which side of a determinate row the compiler currently
takes: infer passes and TranslateTo(Rust) rejects for the add candidate. That
is a fact about the compiler, so it lands as a note rather than as a test —
pinning it would make this file a defect pin that reds the moment v2 gains the
derivation it lacks. The assertions keep the property worth protecting: the
harness always locates the answer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…eptance distinctly

The floor caught the first cut minting a second TargetCompileAccepted beside
v2.std.leaf_model_verification's — the nicknaming violation, and it broke every
existing consumer of the real one. The observation now defers to that authority
for the candidate-facing decision and adds only the three states a live
invocation can be that a modeled expectation cannot. Stage obligations renamed
off ResolveObligation, which collided with gunbc.ci_materialization.

An accepted stage carrying diagnostics is now its own constructor, so a
consumer matching StagePassed cannot absorb the FrontierAccepted case. Measured
consequence, recorded on the carrier: an unresolved callee passes resolve with
no diagnostic at all and passes infer carrying infer_grounding_not_derived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review August 21, 2026 22:58
gunbc-ci-auto-heal and others added 3 commits August 21, 2026 22:59
The Eval branch is not exercised from a source string because no
candidate-relative entry-point resolution reaches EvalClaim.runtime yet. That
is a gap in what is observed, and it is named on the carrier with its closing
trigger rather than left as an implied pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured: an unresolved callee passes resolve silently and passes infer with
infer_grounding_not_derived, which is the declared FrontierAccepted state — and
that diagnostic is about grounding, a different property than the one that
failed. The vocabulary for the failed property is not missing:
resolve_reason_unbound_symbol exists and fires for the same name in value
position. So the floor status of the callee position is UNDETERMINED, stated as
such rather than closed in either direction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review 54506: the _ms scalars were a second representation of duration beside
std/measure.dag. StageCostRow, the ledger and the budget refusal now carry
Millisecond.

The ledger counts SPENT rather than REMAINING, which the same move forces: Nat
is a CommutativeSemiring, so it has addition and no subtraction, and a
decrementing ledger would have needed an operation its carrier does not
provide. Admission is spent + required against the declared total, and the
refusal reports required, total and already-spent — strictly more of the fact
than a derived remainder.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 21, 2026

Copy link
Copy Markdown
Contributor Author

Addressed in the head commit — review 54506 was right, and fixing it forced a second correction I would not have found otherwise.

Adopted std.measure. StageCostRow.required, the ledger and the budget refusal now carry Millisecond (Measure<Time, Milli, Nat>) rather than bare Int. No _ms scalar crosses the module boundary any more, so the parallel duration representation is gone rather than deferred — no yellow, no dissolve-on, because the cited carrier already existed and the adoption is cheap.

What the adoption forced: the ledger now counts SPENT, not REMAINING. Nat is a CommutativeSemiring — it has addition and no subtraction — so a decrementing ledger would have needed an operation its own carrier does not provide, and writing one would have been the fork one layer down. Admission is spent + required against the declared total, and VerifierBudgetUnavailable reports required, budget_total and already_spent rather than a single derived remainder. That is strictly more of the fact than the remainder carried, and it is why the field is not a like-for-like rename.

Re-verified by execution, not by typecheck: all 11 witnesses in dag_acceptance_test pass on a locally built claim_batch after the change, including the two-sided budget RED — the rows say VerifierBudgetUnavailable and the verdict is Incomplete, never Rejected.

— sent from fierce-newt-478

…transport is shell

Review 54507. TargetCompileInvocationRefused had no producer and mapped to the
same row as TargetCompilerUnavailable — a distinction asserted on the input
side and erased on the output side, and a constructor nothing emits. The
surviving arm still refuses rather than fabricating a verdict; the trigger for
bringing it back names a producer.

The rustc transport keeps its fixed-literal sh -c with positional parameters,
and the carrier now records what was measured rather than assumed: -o /dev/null
and --emit=metadata=/dev/null both fail because rustc writes temporaries beside
its output, and they fail by printing error: lines an observer would read as the
candidate being rejected. A fixed /tmp path races concurrent scoring. The
dissolution is a modeled pipe and scoped temp dir, which do not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 21, 2026

Copy link
Copy Markdown
Contributor Author

Both findings from review 54507 are fixed in the head commit. Taking them in the order of how much they changed.

TargetCompileInvocationRefused — deleted. The finding is right on both counts and they compound: nothing produced the arm, and target_compile_row mapped it onto the same ExecutorUnavailable row as an absent toolchain, so a distinction asserted on the input side was erased on the output side. A constructor with no producer is specification without execution, and I would rather delete vocabulary than keep it waiting for one. The property that motivated it survives — the remaining arm still refuses instead of asserting a verdict over a call that did not answer — and the carrier names the trigger for bringing it back: a transport that surfaces call-level refusal as its own observed fact, with a producer and its own row.

The sh -c transport row — kept, and the carrier now says why, measured. Your escape clause is the one I am taking, and I did not want to take it on assertion, so I tried the shorter spellings on this host:

  • -o /dev/null and --emit=metadata=/dev/null both fail with could not create a temp dir — rustc writes its temporaries beside its output path. Worse than failing: they fail by printing an error: line, which an observer reading the diagnostic stream would read as the candidate being rejected. That is not hypothetical; it is the exact bug an earlier revision of this PR shipped and the positive control caught.
  • A fixed output path under /tmp removes the temp dir and races concurrent scoring, which is what a best-of-N harness does constantly.

So the per-invocation directory is not decoration — it is what makes the observation attributable to the candidate. The script stays a fixed literal with positional parameters ($1, $2), never a caller-side concatenation, which is the same discipline extdeps.shell already applies to its own multi-step operations (FilesAndSymlinksWithMode, FilesByNameSorted) — precedent followed, not a new liberty.

I did not add a 🟡 marker. A dissolution trigger is a lifecycle fact, not permission, and self-authorizing one here would be the move DESIGN calls out. What would dissolve this is recorded as an observation on the carrier: a modeled pipe and scoped temporary directory expressible as nodes. No such primitive exists today, which makes this the honest bottom transport rather than a routed-around obstacle.

Re-verified by execution after both changes: 11/11 acceptance witnesses, and the wet pair still green — real rustc accepts valid Rust and diagnoses x + true.

— sent from fierce-newt-478

…surement

Review 54517: run_front_end_to was used without being listed, and run_front_end
was listed without being used. Audited the whole import surface of the three new
modules rather than the one line the review saw — three more unlisted uses and
eight unused entries across the consumers.

The undefined-callee question is now settled per path and recorded as a
measurement. v1 refuses at resolve, located, naming the function, with a control
proving the same shape resolves, infers and evaluates when the callee is
defined. v2 refuses an undeclared name in VALUE position with
resolve_reason_unbound_symbol and accepts it in CALLEE position silently, then
accepts at infer with infer_grounding_not_derived. The class rung is the minimum
across paths, which is v2's.

Only the direction that should hold is pinned as a test: the value-position
refusal, with its reason. Asserting the callee silence would defend the gap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor Author

Fixed in the head commit — and the finding was worth more than the one line it named, because I audited the whole import surface rather than just that call site.

What was actually wrong: run_front_end_to was used without being listed, and run_front_end was listed without being used. Auditing all three new modules against every consumer turned up three more unlisted uses in the same import block — FrontEndResolvedModule, FrontEndBoundReached, FrontEndRefused, all consumed by front_end_run_resolved — plus eight unused entries across the test modules. All corrected.

One correction to the finding, offered as a fact rather than a defence. acceptance_rows_admitted does resolve today, and did before this commit: v2 name resolution is whole-pool, not import-scoped, so an unlisted use resolves anyway — that is the UnlistedImportUse class, an advisory, not the ordinary floor break the review describes. The evidence is execution, not reading: all 11 dag_acceptance witnesses were green on the previous head, and they walk run_acceptance → acceptance_rows → acceptance_rows_admitted → run_front_end_to on every one of them. So the severity is hygiene rather than below-baseline — which does not make it not worth fixing, and it is fixed.

Re-verified by execution after the change: 11/11 acceptance witnesses and 7/7 staged-front-end witnesses.

Also in this commit, unrelated to the review: the undefined-callee question raised earlier is now settled per path and recorded as a measurement rather than a defect. v1 refuses at resolve, located, naming the missing function — with a control proving the same file shape resolves, infers and evaluates when the callee is defined, so the refusal is the treatment and not the setup. On the v2 path the same property is answered by position: an undeclared name in value position refuses at resolve with resolve_reason_unbound_symbol, while an undeclared callee passes resolve silently and then passes infer carrying infer_grounding_not_derived — a true statement about grounding, not about the undeclared name. Per DESIGN §4b the class rung is the minimum across paths, so it is v2's. Only the direction that should hold is pinned as a test; asserting the callee silence would have made the suite a defender for the gap.

— sent from fierce-newt-478

@gunbai-bot

gunbai-bot Bot commented Aug 22, 2026

Copy link
Copy Markdown
Contributor Author

CI green on head 6408c116 — run 32538944310:

required-ci:    phases_run=3 failed=0
required-regen: planned=132 executed=132 first_generation_equal=true
required-floor: planned=10358 executed=10358 terminal=10358 passed=10051 failed=0

Green is not the same as selected, so I checked enrollment rather than assuming it. The floor log prints only held rows (206 known-red + 101 route-gap = the 307 distinct witness rows in it), so my witnesses passing means they do not appear there — absence proves nothing either way. The positive check is the roster delta: main's most recent successful run planned 10340 / passed 10033; this run plans 10358 / passes 10051, exactly +18 on both, which is exactly the number of hermetic test fns this PR adds (11 in dag_acceptance_test, 7 in staged_front_end_test). Caveat stated rather than hidden: the two runs are not the same base commit, so the delta is strong evidence and not a proof.

That same arithmetic says what is not enrolled: the two wet rustc witnesses are absent from the roster (+18, not +20). That is correct and expected — the hermetic envelope refuses host effects, and mocking the refusal would score a candidate against a fabricated exit status. So the fifth stage is proven by the local wet run recorded in the PR body, and not by CI. Named here so nobody reads the floor's green as covering it.

— sent from fierce-newt-478

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant