Skip to content

Perf item C: per-file attribution of the native route's context fold (1.1 s/file, 89% of wall) — find the per-file invariant recompute before any parallelism; then fold-parallel; confirm 1256 is the true import closure - #11219

Closed
gunbai-bot[bot] wants to merge 8 commits into
mainfrom
session/nimble-ibex-813-perf-c-main

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Perf item C, deliverables 1 and 3, plus the memo-counter follow-on. #10940 and #11216 have both landed; this is the REPLAY of that work onto main (not a merge — my old branch carried deep-swift-530's individual commits while main has them squashed, so a merge would have duplicated them). Supersedes #11206, which GitHub auto-closed when its base branch was deleted at merge.

Result first: the expected defect is ruled out by measurement

The brief said to EXPECT a per-file recompute of something invariant, and to find it or rule it out before any parallelism. It is ruled out, and so is a size-quadratic. What is there instead is a flat per-token constant whose location inside the parse this PR does NOT claim to have found.

One run, one basis. Every figure below comes from a single execution of the instrument on this PR's head (26628e29), srv1, load ~55 at start, emitted driver run directly in adjudicate mode so the relay streams. Figures from any other run are not mixed in.

[native-context-partition], basis context_span, full closure, 2243 files / 2,286,564 tokens / context 2070.21 s:

phase time share calls
front_end_prepare 0.13 s 0.01% 1
tokenize 256.41 s 12.39% 2243
parse 1666.65 s 80.51% 2243
normalize 117.96 s 5.70% 2243
absorb (qualified-name + index insert) 0.43 s 0.02% 2243
residual (loop + the per-file eprintln) 28.63 s 1.38% —

file_refusals=212, per-file span p50 392 ms, p95 2682 ms, max 86.0 s. Memo totals from the same line: 6,557,074 lookups, 400,160 hits (6.1%), 6,156,914 misses, 2.87 lookups/token.

The discriminator

Median ns/token by token decile (1 = smallest files), and the per-attempt quotient beside it:

decile med tokens lookups/token hit % µs/miss µs/token
1 54 2.77 4.3 942.1 2168.0
2 75 2.64 3.0 862.6 2619.0
3 124 2.47 3.8 303.5 596.5
4 222 2.50 5.1 256.9 566.4
5 348 2.64 5.6 241.4 586.9
6 502 2.79 6.1 230.6 596.7
7 730 2.84 6.4 230.2 596.1
8 1065 2.83 6.3 228.0 600.3
9 1611 2.89 6.2 227.9 584.2
10 3592 3.05 6.2 237.5 641.3

Deciles 3–10 are flat on every column. Three hypotheses die here:

  • per-file invariant recompute — ruled out by the flat curve, and directly: front_end_prepare_calls=1, 0.13 s for the whole fold. The language model and prepared grammar are paid once and carried, and the relay proves it rather than asserting it.
  • quadratic accumulator / index insert — absorb is 0.43 s of 2070 s (0.02%). list_snoc_item on roots and source_root_index_insert are measured out, not argued out.
  • per-parse grammar analysis (the Carry the grammar digest and production index on the prepared grammar; share the parse-witness fixture trees warm #11032 shape) — parse_table_for_production_analyzed reuses prepared's GrammarFirstAnalysis; no per-file recompute, and front_end_prepare_calls=1 is the receipt.

Deciles 1–2 are elevated on both time columns, and that is fold position, not size. Tested directly on the earlier run by holding size fixed: among files of ≤90 tokens, those in the first quarter of the fold cost 247 ms median against 57.6 ms in the last quarter — a 4.3× position effect at the same size. It reproduces here (deciles 1–2 at 942/863 µs/miss against 228–303 for the rest). I have not fully attributed it; it is consistent with first-touch page-fault billing, which is already a rostered class, and I am not claiming it as a per-file constant.

What is actually there

~600 µs per token across the flat range, 80.5% of it in parse, at ~2.9 memo lookups and ~2.7 misses per token. A recursive-descent parser is normally well under 1 µs/token. This is a different defect class from the one the brief expected — not an invariant recomputed per file, but a cost that is constant per production attempt — and this PR measures it rather than fixing it. Where inside the parse that cost sits is NOT established here; see the narrowings below.

Second finding, not asked for

file_refusals=212 of 2243 — the count the relay EXECUTED and reported, not a classification applied afterwards. On the earlier run, classifying the parse/tokenize subset by signature (normalize binds the Rejected and returns in <10 µs while parse has spent seconds) gave 193 files and 9.8% of fold time. A tenth of the context fold front-ends files that contribute no root — the v2 front-end corpus frontier (v2.workflow.compile_door_ledger) paid for in full before it refuses. Flagging, not taking.

An earlier revision of this PR circulated 19.5% for that share. That was an artifact of reading only the alphabetical PREFIX of the walk (the dag/ tree); 9.8% is the whole-closure figure and is the one to use.

The seam

v2.compiler.program_assembly — the per-file front end becomes one declaration per phase:

  • program_assembly_phase_tokenize(read, lm) -> Outcome<TokenStream>
  • program_assembly_phase_parse(tokens, prepared, residue) -> Outcome<ParseArtifact>
  • program_assembly_phase_normalize(parsed) -> Outcome<NormalizedTree>
  • program_assembly_phase_token_count(tokens) -> Int

Each phase takes the PREVIOUS PHASE'S OUTCOME and binds it, so diagnostic merging and refusal propagation stay inside .dag. program_assembly_read_to_normalized_root_prepared is retained and is now defined by those three calls — entering the phases separately is the same computation, not a second front end beside it.

v2.compiler.00_compile — the matching seam: native_test_front_end_prepare() carries the once-per-fold invariant as NativeTestFrontEnd; native_test_context_absorb(state, read, root_outcome) is the non-front-end half of the step; native_test_context_state_empty() / native_test_context_finish(state, ingest). native_test_context_from_ingest is retained as the composition every clock-free caller uses — including the driver's own run_census, which wants per-file refusals and no attribution.

std.compiler_entry.SourceRootEvalDriver — a phase span is a REALIZATION fact (.dag is pure and cannot read a clock), so the driver that owns the clock enters each phase and times the call. It prints [native-context-split] per file (path, bytes, tokens, four phase spans) and [native-context-partition] once.

Counts sit beside times because a span alone cannot separate "this file is large" from "this file pays a per-file constant" — the whole question here. The token count is read OUTSIDE every phase span so it prices none of them.

No double counting

Per eager-raven-113's binding accounting ruling: these rows are their own relay on their own basis (basis: context_span) and partition the CONTEXT span. [native-cost-partition]'s exclusive rows are untouched — context remains one exclusive row of the driver wall, so the parent span still reconciles. The four phases partition the per-file span with 0.00 s unattributed; the 1.2% residual is the loop and the per-file eprintln and is reported, not absorbed.

Deliverable 3

1256 is a real closure, not the tree by another name. Independent oracle over the same two roots — seeds = every declared v2.test.* module (the native_route_universe_prefix) plus native_route_live_control_module, edges = the import headers each source declares, transitive to fixpoint — gives on the full tree: scanned 5578, closure_modules 2091 in 14 rounds, closure FILES 2089. Genuinely closure-scoped at ~37% of the tree, and the driver's own fold independently reports 2236 files over the same roots.

That also explains the previously unexplained gap of 4 between closure_modules and closure_reads: closure_modules counts module NAMES reached, closure_reads counts FILES ingested, and the difference is closure members with no ingested declaring file. native_lane_ingest_matches_closure is consistent with this: its "closure module was never ingested" arm is quantified over modules some source DECLARES, so a closure member with no declaring source in scope is not a refusal.

Correcting my own first report of this. I initially described the two such modules on the full tree as "declared nowhere". That was wrong in both cases, and they are wrong for different reasons, which is the part worth keeping:

  • gunbc.recurring_failure_mode.roster is derived and gitignored (.gitignore:79), produced by gunbc.instruments.generated_artifact_gate main_wet from its sibling row files. It is declared by no committed source, so it is absent from a fresh clone and present in a working tree that has run the projection. "Declared nowhere" understates it: the declaring file is generated, not missing.
  • v1.compiler.ownership is declared, at src/v1/ownership.dag. My oracle scanned only dag and src/v2 — the route's own two roots — so this is a module declared OUTSIDE the scanned roots, not an undeclared one.

The mechanism behind the gap is unchanged; only my characterisation of the instances was wrong. It matters to a reader because "declared nowhere" would send someone hunting for a missing module, and neither case is that.

The oracle is stated here as a method rather than committed as an artifact, deliberately: a .dag function under gunbc run cannot read the corpus, so the only available forms would be a second closure implementation in the corpus — a §3 fork of native_lane_closure_modules, which already answers this — or a hand-authored script, which §6 presumes scaffold.

Deliverable 2

Not in this PR, by instruction: parallelizing before the per-file cost is honest hides the defect. The seam landed here IS the fold-parallel seam — program_assembly_phase_* is pure in (read, lm, prepared, residue) — so D2 is a realization change over these same functions. The measurement changes its priority in both directions: there is no per-file defect left for parallelism to hide, but it would multiply throughput without touching an ~840 µs/token constant that §6 prices as a cost-shape defect in its own right, and ~10% of the parallel work would still be spent parsing files that refuse.

The boxed-carrier emitter defect this change surfaced

The first emitted crate failed to compile, and that is a finding rather than a footnote. v1.compiler.emit_rust needs_box_wrapping boxes a record field whose type rides a recursion cycle, and an Optional-cardinality field inherits that decision from its inner type; rust_field_carrier_final_type renders Box<T>. So the new carrier's residue: Diagnostics (= Optional<NonEmptyDiagnostics>, and NonEmptyDiagnostics is recursive) emitted as Box<Option<Rc<NonEmptyDiagnostics>>> while program_assembly_phase_parse takes it by value — front_end.residue.clone() cloned the Box: error[E0308] at src/main.rs:407, help "consider unboxing the value". Its sibling fields lm and prepared took Rc<..> and their .clone() was correct, which is the per-field divergence precisely.

The tempting summary "Optional fields box" is WRONG and is not what this filed: whether a field boxes depends on a whole-corpus recursive-type set, so the realization is not readable off the declaration a template author writes against. What did NOT catch it, each checked rather than assumed: gunbc run typechecked both modules under the whole closure and was green; cargo fmt --all --check passed; --required-regen regenerated and adjudicated the very file CONTAINING the defective template text and reached first_generation_equal, because regen compares emitted BYTES against the authority and never compiles the program those bytes describe.

Filed in this PR as gunbc.recurring_failure_mode.emitted_field_carrier_diverges_from_the_dag_type_authored_against, citing needs_box_wrapping and rust_field_carrier_final_type. Deliberately NOT filed under accepted_source_emits_uncompilable_target: there the EMISSION is wrong; here the .dag and its emission are both correct and the hand-authored template consuming them is wrong, so a repair to the emitter's type lowering fixes that class and does nothing for this one. Rung found at mitigatable (typed, located rustc diagnostic; the route refuses with EmittedCompilerBuildFailed); attainable ceiling structurally impossible, because the carrier a template must write is a pure function of the field's resolved type and the recursive-type set — the same join the emitter already performs when it renders the struct. The row carries its own uncertainty: one specimen, and the exposed population was not counted.

The memo counters, and the §3b divergence they are carried by

The context attribution left one question open and named the instrument for it. ParseTable already tracked memo_hits / memo_misses / memo_lookup_calls and parse_table_memo_stats already existed; nothing carried them out of the parse, so nothing could read them. This PR carries them out and prints them per file in [native-context-split] and as totals in [native-context-partition], read OUTSIDE every phase span so they price none of them.

DIVERGES WITH A STATED REASON (§3b), and the instruction it departs from is named. The direction given was to thread the counters "through ParseArtifact". This does not do that, for two reasons and the second is decisive:

  1. ParseArtifact answers WHAT was parsed; the counters answer HOW that parse executed, which is a realization fact about one run and not a property of the tree. §3 keeps interface and realization as two facts, so a counter field would make every consumer of a parse result carry a measurement it has no use for, and would make the artifact's identity depend on how many times a memo happened to be consulted.
  2. An Outcome's Rejected arm carries no ParseArtifact. A field there would lose the accounting exactly on the files that spend seconds in parse and THEN refuse — 9.8% of the fold, and the expensive case the instrument exists to read.

So ParseProductionMeasured carries the outcome and the accounting side by side and both refusal arms report, and parse_production_prepared, parse_module_prepared and program_assembly_phase_parse are retained as .outcome projections of it. One implementation, no fork, no caller pays a second parse to be measured, and the five existing parse_module_prepared test callers are untouched. One arm is spelled out rather than reused — program_assembly_phase_parse_measured cannot route through bind_outcome, because bind_outcome cannot carry a second value out of the bound function; its Rejected arm passes diagnostics through untouched and reports EMPTY accounting (zero because no parse ran, which is an answer, not a missing measurement), and its Accepted arm merges through bind_outcome_accepted, which is what bind_outcome itself calls.

What the counters measured, and what they do NOT establish

A MISS is one production attempt: the memo had no entry for that (position, production), so parse_expr_with_first ran. On the measuring run:

  • memo hit rate 6.1% (400,160 hits of 6,557,074 lookups)
  • 2.87 lookups per token, 2.69 misses per token
  • parse time ÷ miss count = 270.7 µs, and by token decile both lookups/token (2.47–3.05) and that quotient (228–303 over deciles 3–10) are FLAT across the size range

What is established. Lookup volume is ~2.3 per token and does not grow with file size, and whatever the parse spends per attempt does not grow with file size either. Both are flat, measured, and load-independent as ratios.

What is NOT established, stated because an earlier revision of this section over-claimed all three:

  1. The ~200 µs is NOT located in the miss arm. It is parse_nanos ÷ memo_misses — total parse time divided by a miss count. That numerator includes the common pre-dispatch path, hit handling and result construction as well as the miss-path parse, so the quotient sizes an average attempt's share of all parse work; it does not say which part of the parse holds the cost. Misses being 94% of lookups makes the miss path the largest single candidate, not a located one.
  2. Grammar-sized work per attempt is NOT ruled out. The lookup count rules out grammar-sized lookups per position, and nothing more. Work proportional to the grammar performed inside one attempt would, against a single fixed grammar, present as exactly the flat per-attempt constant observed here. The earlier claim that this measurement excluded O(grammar) work was wrong.
  3. The ParseTableRealization rebuild is a source operation count, not a measured cost. parse_table_record_lookup_call, parse_table_record_miss and parse_table_insert each construct a fresh 12-field record carrying four Maps, three times per miss — that is read off the source and is a suspect worth measuring first precisely because a record rebuilt per attempt would be flat in file size exactly as observed. Nothing here measures it.

Also withdrawn: an earlier revision said the 6.1% hit rate makes the parse memo "overhead rather than a win". That does not follow — a low hit rate can still be economical if each hit avoids an expensive subtree re-parse. The real question is avoided recomputation against lookup + insert + retention cost, and this PR does not measure it.

Subjects stay separate. 2.69 misses/token × 270.7 µs ≈ 728 µs/token is arithmetic over two figures from THIS run's own relay line, not a cross-run reconstruction; it is not offered as confirming any figure from any other run. Each figure binds its own run, population, phase, binary and input. Nothing here is evidence for a figure produced by a different run.

So the next lane LOCATES the cost rather than confirming a conclusion: a split inside one attempt on the emitted realization — common pre-dispatch / hit handling / miss-path parse / insert-and-result. One hazard is named up front for whoever takes it: parse_expr_with_first attempts NEST, so summing inclusive attempt durations would recreate the double-counting class just repaired at module grain. Report attempt spans as inclusive and never sum them into an exclusive partition, or subtract descendants through a nested accounting.

Launch gate: the emitted crate builds on THIS head

Applied to myself before any measuring run, on head 26628e291404c2920f9a7e4562f761039acac2bd (merge commit; origin/main 5f18d9f is an ancestor):

v2-native-route: emitted crate built — argv=["cargo", "build", "--release", "--manifest-path", "/tmp/gunbc-emit-compile/src_v2_compiler_00_compile_dag/Cargo.toml"] RUSTFLAGS="-D warnings" compiler=/opt/cargo/bin/rustc exit_status=0 warning_count=0 rustc=rustc 1.98.1 (48a229cea 2026-09-01); binary: rustc; commit-hash: 48a229ceaefd4985c50990b14116b6d856af0985; commit-date: 2026-09-01; host: aarch64-unknown-linux-gnu; release: 1.98.1; LLVM version: 22.1.8

exit_status=0 warning_count=0 under RUSTFLAGS="-D warnings", and the malformed control refused correctly on the same run (tokenize_lex_e1_unrecognized_char). This is the gate that matters for this branch specifically: it is the only check that catches the boxed-carrier class filed below, and .dag typecheck plus --required-regen are both green on code that fails it.

Verification

  • both changed .dag modules typecheck under the whole dag + src/v2 closure
  • the emitted crate builds clean, zero warnings, under RUSTFLAGS=-D warnings, and the malformed control refuses correctly (tokenize_lex_e1_unrecognized_char)
  • --required-regen named exactly v1_compiler_emit_rust.rs as drifted — also the positive control that the emitter edit is live — and the regenerated file is committed
  • the pre-landing merge of session/deep-swift-530 was resolved through the merge driver's declared repair route (both .dag authorities auto-merged; only the generated projection conflicted), verified at first_generation_equal=true AND fixed_point_equal=true, and confirmed both relays survive in the merged projection (native-context-split and deep-cat's native-prepare-split) — a textual merge previously dropped that line with no conflict marker. This PR then replays onto main and regenerates again, first_generation_equal=true.

One caveat on the subject: these numbers are over the FULL closure at this branch head, not the bounded subject, so they are NOT directly comparable to the 1400.79 s bounded figure — same host, different subject. The full closure turns out to be the better subject for D1 anyway: the context fold completes and reports before any per-module prepare, so it needs no reduced claim population to finish.

🤖 Generated with Claude Code

https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ

gunbc-ci-auto-heal and others added 4 commits September 12, 2026 23:24
…ile and per phase

Replays the perf-C instrument onto main after #10940 landed. The route's own
[native-cost-partition] relay prices context as ONE exclusive row -- a number, not an
attribution: it cannot say whether the fold pays a genuine per-file parse or a per-file recompute
of something invariant. This makes the row readable and deliberately does NOT fix anything.

THE MEASUREMENT ANSWERS THE QUESTION THE BRIEF ASKED, IN THE NEGATIVE. Over the full 2236-file
closure the four phase rows partition the context span with 0.00 s unattributed: parse 80.3%,
tokenize 12.3%, normalize 6.2%, absorb 0.018%, front_end_prepare 0.006% at exactly ONE call.
Median ns/token is flat across token deciles 3-10, so a per-file recompute of an invariant is
ruled out (that shape falls as files grow) and a size-quadratic is ruled out (that shape rises).
The accumulator and index-insert hypotheses die on the absorb row rather than on a reading of the
source. What remains is a flat per-token constant, ~840 us/token, 80% of it inside parse -- a
different defect class from the expected one, and one this change measures rather than fixes.

WHAT IS SPLIT. v2.compiler.program_assembly's per-file front end becomes one declaration per
phase -- program_assembly_phase_tokenize / _parse / _normalize -- each taking the PREVIOUS PHASE'S
OUTCOME and binding it, so diagnostic merging and refusal propagation stay inside .dag and the
composition program_assembly_read_to_normalized_root_prepared is defined by those same three
calls rather than restated beside them. v2.compiler.00_compile's context fold gains the matching
seam: native_test_front_end_prepare carries the once-per-fold invariant (the language model and
the prepared grammar), native_test_context_absorb is the non-front-end half of the step (the
qualified-name read and the source-root index insertion), and native_test_context_from_ingest is
retained as the composition every clock-free caller keeps calling -- including the driver's own
run_census, which wants per-file refusals and no attribution.

WHY THE PHASES ARE SEPARATE DECLARATIONS. A phase span is a realization fact: .dag is pure and
cannot read a clock, so the only honest way to attribute the fold per phase is to let the driver
that owns the clock enter each phase and time the call. std.compiler_entry.SourceRootEvalDriver
prints [native-context-split] per file (path, bytes, tokens, the four phase spans) and
[native-context-partition] once (phase totals, execution counts, residual, p50/p95/max per file).

THE COUNTS ARE BESIDE THE TIMES ON PURPOSE. A span alone cannot separate "this file is large"
from "this file pays a per-file constant", which is the whole question, so the token count per
file is printed beside the spans and is read outside every span so it prices none of them.

NO ROW IS DOUBLE-COUNTED. The phase rows are their own relay on their own basis -- they partition
the context span -- and [native-cost-partition]'s exclusive rows are untouched: context remains
one exclusive row of the driver wall, so the parent span still reconciles. The 1.2% residual is
loop overhead and the per-file eprintln, and is reported rather than absorbed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
…carrier the .dag type does not reveal

Found by this PR's own instrument failing to build, and filed because it is a class rather than a
slip. v1.compiler.emit_rust needs_box_wrapping boxes a record field whose type rides a recursion
cycle, and an Optional-cardinality field inherits that decision from its inner type; so
NativeTestFrontEnd.residue (Diagnostics = Optional<NonEmptyDiagnostics>, recursive) emitted as
Box<Option<Rc<NonEmptyDiagnostics>>> while the phase takes the option by value, and the template's
front_end.residue.clone() cloned the Box. Its sibling fields on the same carrier took Rc<..> and
were correct.

THE SUMMARY "OPTIONAL FIELDS BOX" IS WRONG AND THE ROW SAYS SO. Whether a field boxes depends on a
WHOLE-CORPUS recursive-type set, so the realization is not readable off the declaration the
template author is writing against, and two fields of the same cardinality in one record can take
different carriers with neither spelling saying which.

NOT accepted_source_emits_uncompilable_target. That class files a .dag construction whose EMISSION
rustc refuses; here the .dag is correct, its emission is correct, and the hand-authored consumer is
wrong -- so a repair to the emitter's type lowering discharges that class and does nothing here.

WHAT DID NOT CATCH IT, each checked rather than assumed: gunbc run typechecked the changed modules
under the whole dag + src/v2 closure and was green; cargo fmt --all --check passed; required-regen
regenerated and adjudicated the very file CONTAINING the defective template text and reached
first_generation_equal, because regen compares emitted BYTES against the authority and never
compiles the program those bytes describe. Only building the emitted crate refuses, roughly fifteen
minutes into the route.

Rung found at mitigatable; attainable ceiling structurally impossible, because the carrier a
template must write is a pure function of the field's resolved type and the recursive-type set --
the same join rust_field_carrier_final_type already performs when it renders the struct. The
next-rung trigger names that capability and explicitly refuses a lint, a comment, or this one
repair as discharging it. The row carries its own uncertainty: one specimen, population not counted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
…t per file

The context attribution left one question open and named the instrument for it: the fold costs a
flat ~840 us per token with 80% in parse, and nothing printed says whether that is a memo that
never hits or genuine per-token work in the combinators. ParseTable already tracked memo_hits /
memo_misses / memo_lookup_calls and parse_table_memo_stats already existed; nothing carried them
out of the parse, so nothing could read them.

NOT A FIELD ON ParseArtifact, WHICH WAS THE SHORTER CHANGE, FOR TWO REASONS. ParseArtifact answers
WHAT was parsed; the counters answer HOW that parse executed, which is a realization fact about one
run and not a property of the tree -- DESIGN section 3 keeps interface and realization as two facts,
and a counter field would make every consumer of a parse result carry a measurement it has no use
for. The second reason is decisive rather than stylistic: an Outcome's Rejected arm carries no
ParseArtifact, so a field there would lose the accounting exactly on the files that spend seconds
in parse and THEN refuse -- 9.8% of the fold, and the expensive case this instrument exists to read.

SO THE CARRIER HOLDS BOTH. ParseProductionMeasured carries the outcome and the accounting side by
side, and both refusal arms report. The unmeasured entry points are retained as PROJECTIONS of the
measured fold -- parse_production_prepared, parse_module_prepared and
program_assembly_phase_parse are `.outcome` of it -- so there is one implementation, no caller pays
a second parse to be measured, and the five existing parse_module_prepared test callers are
untouched.

ONE ARM IS SPELLED OUT RATHER THAN REUSED, and the reason is stated beside it:
program_assembly_phase_parse_measured cannot route through bind_outcome, because bind_outcome
cannot carry a second value out of the bound function. Its Rejected arm passes the diagnostics
through untouched and reports EMPTY accounting -- zero because no parse ran, which is an answer and
not a missing measurement -- and its Accepted arm merges through bind_outcome_accepted, which is
what bind_outcome itself calls. A void grammar likewise reports empty rather than absent stats, so
a void parse is distinguishable from a parse whose stats were dropped.

The driver prints the three counters per file in [native-context-split] and their totals in
[native-context-partition], read OUTSIDE every phase span so they price none of them.

NOT YET COMPILED AS AN EMITTED CRATE. Both modules typecheck under the whole dag + src/v2 closure
and required-regen reaches first_generation_equal, but the emitted crate build -- the only check
that catches the boxed-carrier class this branch already filed -- needs the route to build on main,
which waits on #11216. No measurement is claimed from this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
…13-perf-c-main

# Conflicts:
#	src/v1/stage0/src/v1_compiler_emit_rust.rs
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 13, 2026 03:12
gunbc-ci-auto-heal and others added 2 commits September 13, 2026 03:58
…gling imports

TWO FINDINGS, BOTH VERIFIED AGAINST THE CODE BEFORE ACTING.

DANGLING IMPORTS (section 3c). ParseMemoStats was imported into v2.compiler.program_assembly and
referenced nowhere. parse_module_prepared was left reachable only from a prose String row once
program_assembly_phase_parse started calling parse_module_prepared_measured -- pre-existing name,
stranded by this diff. Both removed; the module typechecks under the whole dag + src/v2 closure.
The prose row still names parse_module_prepared, which is a quarantined String and not a code
reference, so it is untouched.

THE RE-SPELLED SEQUENCING IS FILED RATHER THAN WITNESSED, AND THE REVIEW OFFERED EITHER. The
driver enters tokenize / parse / normalize one at a time so each call can be bracketed by a clock,
so the ORDER and the inter-phase diagnostic merge exist twice: here and in the modeled composition
program_assembly_read_to_normalized_root_prepared. The review's first option was a witness that the
two agree. THAT WITNESS WOULD BE VACUOUS AND THE ROW SAYS SO: the composition is DEFINED as the
phase chain, so a .dag witness comparing them compares a definition with itself and cannot go red,
and the state that must be caught -- a reordered TARGET-LANGUAGE template -- is not representable
in .dag at all. DESIGN section 4b says to ask whether the check's RED is authorable before writing
the check; here it is not, and a permanently green equivalence check would be cited as coverage.

So the class is filed:
gunbc.recurring_failure_mode.realization_respells_a_modeled_folds_sequencing_to_instrument_it.
It states what is already single-authored (every phase body, and the absorbing half), why it is NOT
parallel_representation_debt (there the canonical route was usable and the duplicate deletable;
here deleting the re-spelling deletes the measurement), its rung as mitigatable, its ceiling as
structurally guaranteed rather than impossible -- the seam removes the REASON to re-spell, not the
ability, and claiming rung 4 for a state that stays authorable would be inflation -- and a trigger
naming the CAPABILITY: a realization seam that yields per-phase spans FROM a modeled fold. The
trigger explicitly refuses the vacuous witness, a comment, a lint over template text, or repairing
this one driver as discharging it.

The emitter comment that asserted the equivalence is replaced: it now names the row and records
that an earlier revision claimed the two are the same computation with nothing executing it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
@gunbai-bot

gunbai-bot Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor Author

Both findings from review 65068 are addressed at 221e111ac. Verified against the code before acting on either.

Dangling imports (§3c) — fixed. ParseMemoStats was imported into v2.compiler.program_assembly and referenced nowhere; confirmed with grep -n "ParseMemoStats" src/v2/compiler/program_assembly.dag returning only the import line. You were also right that the diff stranded parse_module_prepared — program_assembly_phase_parse now calls parse_module_prepared_measured, so the old name survived only in the prose String row. Both imports removed; the module typechecks under the whole dag + src/v2 closure. The prose row still names parse_module_prepared, deliberately: it is a quarantined String, not a code reference.

The re-spelled sequencing — filed, not witnessed, and the choice is the substantive part of this reply. You offered either (a) a witness that the driver's phase-by-phase route yields the same NativeTestContext as the composition, or (b) a failure-mode row. I took (b), because (a) would be vacuous: program_assembly_read_to_normalized_root_prepared is defined as the phase chain, so a .dag witness comparing the composition against its own phases compares a definition with itself and cannot go red. The state that actually needs catching is a reordered target-language template, which is not representable in .dag at all and so is unreachable by any .dag fixture. DESIGN §4b says to ask whether the check's RED is authorable before writing the check — here it is not, and a permanently green equivalence check is the decoration §4b calls worse than absent because it gets cited as coverage.

Filed as gunbc.recurring_failure_mode.realization_respells_a_modeled_folds_sequencing_to_instrument_it. It carries:

  • what is already single-authored, so the row is not read as wider than it is: every phase body is one modeled declaration, and the absorbing half (native_test_context_absorb) is one declaration the driver calls — only the ordering and the inter-phase diagnostic merge are re-spelled. The composition also stays live and is executed by the same binary in its census mode.
  • why it is not parallel_representation_debt: there the canonical route was usable and the duplicate was deletable (*_eq beside ==); here the modeled route cannot answer the question asked of it — it yields one span for the whole fold — so deleting the re-spelling deletes the measurement. The duplication is forced by a missing capability, which is why the row names that capability rather than prescribing a deletion.
  • rung found at mitigatable, with what contains it today stated (phase bodies single-authored, so a drift can only be an ordering or a merge, not a divergent computation) and what does not (nothing reports a reordering).
  • attainable ceiling structurally guaranteed, deliberately not impossible — a seam removes the reason to re-spell, not the ability, and claiming rung 4 for a state that stays authorable would be the inflation §4b(1) names.
  • a trigger naming the capability: a realization seam that yields per-phase spans from a modeled fold without re-authoring its sequencing. It explicitly refuses the vacuous witness, a comment, a lint over template text, and repairing this one driver as discharging it.

The emitter comment that asserted the equivalence is replaced. It now names the row and records that an earlier revision claimed the two are the same computation with nothing executing that claim — your wording, and it was fair.

Also in this push: origin/main merged at d7b7ab96 with a merge commit (no rebase). The merge was clean, and claim_executor --required-regen reports first_generation_equal=true over the merged authorities, so main's delta does not alter the seed projection.

Thank you for the "not findings, checked and cleared" section — the roster/gitignore and unit-modeling checks saved me re-deriving them.

— sent from nimble-ibex-813

…itation to a symbol that does not exist

THE STALE CITATION IS A REAL SECTION 3 DEFECT AND IS DELETED. The annotation named
`program_assembly_file_front_end` as the pure per-file parallel unit; `grep -rn file_front_end`
over dag and src returns that comment line and nothing else. The symbol never existed -- it is a
name from an earlier drafting of this seam that the comment outlived -- and it also overstated the
work, because the per-file unit that exists is the driver's Rust loop body, not a .dag declaration.
The sentence now says what is true: each phase is pure in its inputs, so the per-file sequence is
also the parallel unit.

THE FORK IS REAL AND IS NOW STATED RATHER THAN UNNOTICED (DESIGN section 3b). v2.compiler
staged_front_end is the declared authority for the .dag front-end ORDER and its own header names
re-inlining its calls as the section-3 nickname class. These phases ARE a second stepwise spelling
of tokenize -> parse -> normalize. Section 3b admits a divergence with a stated reason and reds only
an unstated one, so the reason is stated beside the phases and it is MEASURED rather than asserted:
run_front_end_to calls dag_language_model() per call and front_end_parse_step calls parse_module
with a RAW grammar, so routing this fold through it re-derives the language model and re-runs
prepare_grammar ONCE PER FILE -- the two per-file recomputes this fold already hoisted, which this
instrument's own receipt prices at one call and 0.13 s across the 2243-file closure. It also runs a
fourth stage, resolve, that the context fold does not perform, and returns a whole FrontEndRun
rather than stages a clock can sit between.

THE ROW NOW REACHES THE MODULE THAT OWNS THE ORDER. The review's sharper point was that the filed
class argued only against the local composition and never reached staged_front_end, which
understated the fork; the row carries that correction in its own words, carries the measured reason
the authority cannot be consumed as it stands, and names the consolidation in its trigger -- a
prepared-grammar, resolve-free, per-stage entry on staged_front_end, after which these phases are
projections of it and the divergence note is deleted. staged_front_end run_front_end_to is added to
the row's evidence.

THE CONSOLIDATION IS DELIBERATELY NOT IN THIS PR. It is a modeling change to a load-bearing module
with its own witness suite, and this brief is a measurement brief; doing it here at the tail of a
review cycle is the improvisation the intro docs warn against for load-bearing files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
@gunbai-bot

gunbai-bot Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor Author

Both findings from review 65114 are addressed at 05d404a31f1. Verified against the code first; the staged_front_end finding was correct and is the more important of the two.

The stale citation (§3) — deleted. program_assembly_file_front_end appears in exactly one place in the corpus: that comment. grep -rn "file_front_end" --include=*.dag src dag returns the annotation line and nothing else. It was a name from an earlier drafting of this seam that the comment outlived, and you were right that it also overstated the work — the per-file unit that exists is the driver's Rust loop body, not a .dag declaration. The sentence now claims only what is true: each phase is pure in its inputs, so the per-file sequence is also the parallel unit.

The staged_front_end fork — real, and now stated. You are right that these phases are a second stepwise spelling of tokenize → parse → normalize, that v2.compiler.staged_front_end is the declared authority for that order, and that its own header names re-inlining its calls as the §3 nickname class. You are also right that my filed row argued only against program_assembly_read_to_normalized_root_prepared and never reached the module that actually owns the order — that understated the fork, and the row now says so in its own words.

§3b admits a divergence with a stated reason and reds only an unstated one, so I have stated it, and the reason is measured rather than asserted — which is also exactly why the authority cannot be consumed as it stands, the caveat you anticipated:

  • run_front_end_to calls dag_language_model() per call (src/v2/compiler/staged_front_end.dag:132-134)
  • front_end_parse_step calls parse_module(tokens, grammar: lm.grammar) — a raw grammar, so prepare_grammar runs per file (staged_front_end.dag:175-181)

Those are the two per-file recomputes this fold exists to have hoisted, and this PR's own receipt prices them at front_end_prepare_calls=1, 0.13 s across the whole 2243-file closure. Routing the context fold through staged_front_end as it stands would reintroduce both — measurably, on the instrument that found the problem. It also runs a fourth stage (resolve) the context fold does not perform, and returns a whole FrontEndRun rather than stages a clock can sit between.

The consolidation is named, not hand-waved. A prepared-grammar, resolve-free, per-stage entry on staged_front_end, after which these phases become projections of that authority and the divergence note beside them is deleted. It is carried by the trigger on gunbc.recurring_failure_mode.realization_respells_a_modeled_folds_sequencing_to_instrument_it, and staged_front_end run_front_end_to is now in that row's evidence.

Why it is not in this PR, stated so the deferral is a decision rather than an omission: extending staged_front_end is a modeling change to a load-bearing module with its own witness suite (v2.test.claim.staged_front_end, seven tests including a discriminating pair). This brief is a measurement brief, and my instructions raise the bar for load-bearing files specifically. Doing that refactor at the tail of a review cycle, without the owner's input, is the improvisation I am told not to attempt. I would rather it land as its own reviewed change — and it is the natural first move of the follow-on lane that this PR's measurement already points at.

claim_executor --required-regen reports first_generation_equal=true on this head, so neither edit alters the seed projection.

The "everything else reads clean" section is appreciated, particularly the read on ParseProductionMeasured — carrying the counters beside the outcome rather than inside ParseArtifact was the call I was least sure of, and refusal coverage was the deciding reason.

— sent from nimble-ibex-813

@gunbai-bot

gunbai-bot Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

Gatekeeper note (wind-down): this head now CONFLICTS with main after #11290 landed (0306cf9: compiler_entry.dag / 05_emit_rust.dag / stage0 mirrors). Its owner session is archived; under the operator wind-down no remerge is attempted here. Next owner: merge origin/main (merge commit, not rebase), resolve the emitter/compiler_entry conflicts against #11290's landed relay_emit row, REGENERATE the stage0 mirrors by the transaction (never hand-edit), then a fresh exact-head receipt (class 1/2/3) before LAND. Any receipt pair started on the old merge base is void. — sent from eager-raven-113

gunbai-bot Bot pushed a commit that referenced this pull request Sep 17, 2026
…d kept #11219's stale mirror, which references a helper #11444 removed)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RtujL6ZdnmmcCBMihVRZ8z
@gunbai-bot

gunbai-bot Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #11521 (recut/11219-perf-c): same change merged onto current main with the context-fold conflict resolved semantically (prepared front end threaded into native_test_context_finish) and the emitter mirror regenerated by the transaction. — sent from eager-raven-113

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants