Repository navigation
Perf item C: per-file attribution of the native route's context fold (1.1 s/file, 89% of wall) — find the per-file invariant recompute before any parallelism; then fold-parallel; confirm 1256 is the true import closure - #11219
gunbai-bot[bot] wants to merge 8 commits into
Conversation
…ile and per phase Replays the perf-C instrument onto main after #10940 landed. The route's own [native-cost-partition] relay prices context as ONE exclusive row -- a number, not an attribution: it cannot say whether the fold pays a genuine per-file parse or a per-file recompute of something invariant. This makes the row readable and deliberately does NOT fix anything. THE MEASUREMENT ANSWERS THE QUESTION THE BRIEF ASKED, IN THE NEGATIVE. Over the full 2236-file closure the four phase rows partition the context span with 0.00 s unattributed: parse 80.3%, tokenize 12.3%, normalize 6.2%, absorb 0.018%, front_end_prepare 0.006% at exactly ONE call. Median ns/token is flat across token deciles 3-10, so a per-file recompute of an invariant is ruled out (that shape falls as files grow) and a size-quadratic is ruled out (that shape rises). The accumulator and index-insert hypotheses die on the absorb row rather than on a reading of the source. What remains is a flat per-token constant, ~840 us/token, 80% of it inside parse -- a different defect class from the expected one, and one this change measures rather than fixes. WHAT IS SPLIT. v2.compiler.program_assembly's per-file front end becomes one declaration per phase -- program_assembly_phase_tokenize / _parse / _normalize -- each taking the PREVIOUS PHASE'S OUTCOME and binding it, so diagnostic merging and refusal propagation stay inside .dag and the composition program_assembly_read_to_normalized_root_prepared is defined by those same three calls rather than restated beside them. v2.compiler.00_compile's context fold gains the matching seam: native_test_front_end_prepare carries the once-per-fold invariant (the language model and the prepared grammar), native_test_context_absorb is the non-front-end half of the step (the qualified-name read and the source-root index insertion), and native_test_context_from_ingest is retained as the composition every clock-free caller keeps calling -- including the driver's own run_census, which wants per-file refusals and no attribution. WHY THE PHASES ARE SEPARATE DECLARATIONS. A phase span is a realization fact: .dag is pure and cannot read a clock, so the only honest way to attribute the fold per phase is to let the driver that owns the clock enter each phase and time the call. std.compiler_entry.SourceRootEvalDriver prints [native-context-split] per file (path, bytes, tokens, the four phase spans) and [native-context-partition] once (phase totals, execution counts, residual, p50/p95/max per file). THE COUNTS ARE BESIDE THE TIMES ON PURPOSE. A span alone cannot separate "this file is large" from "this file pays a per-file constant", which is the whole question, so the token count per file is printed beside the spans and is read outside every span so it prices none of them. NO ROW IS DOUBLE-COUNTED. The phase rows are their own relay on their own basis -- they partition the context span -- and [native-cost-partition]'s exclusive rows are untouched: context remains one exclusive row of the driver wall, so the parent span still reconciles. The 1.2% residual is loop overhead and the per-file eprintln, and is reported rather than absorbed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
…carrier the .dag type does not reveal Found by this PR's own instrument failing to build, and filed because it is a class rather than a slip. v1.compiler.emit_rust needs_box_wrapping boxes a record field whose type rides a recursion cycle, and an Optional-cardinality field inherits that decision from its inner type; so NativeTestFrontEnd.residue (Diagnostics = Optional<NonEmptyDiagnostics>, recursive) emitted as Box<Option<Rc<NonEmptyDiagnostics>>> while the phase takes the option by value, and the template's front_end.residue.clone() cloned the Box. Its sibling fields on the same carrier took Rc<..> and were correct. THE SUMMARY "OPTIONAL FIELDS BOX" IS WRONG AND THE ROW SAYS SO. Whether a field boxes depends on a WHOLE-CORPUS recursive-type set, so the realization is not readable off the declaration the template author is writing against, and two fields of the same cardinality in one record can take different carriers with neither spelling saying which. NOT accepted_source_emits_uncompilable_target. That class files a .dag construction whose EMISSION rustc refuses; here the .dag is correct, its emission is correct, and the hand-authored consumer is wrong -- so a repair to the emitter's type lowering discharges that class and does nothing here. WHAT DID NOT CATCH IT, each checked rather than assumed: gunbc run typechecked the changed modules under the whole dag + src/v2 closure and was green; cargo fmt --all --check passed; required-regen regenerated and adjudicated the very file CONTAINING the defective template text and reached first_generation_equal, because regen compares emitted BYTES against the authority and never compiles the program those bytes describe. Only building the emitted crate refuses, roughly fifteen minutes into the route. Rung found at mitigatable; attainable ceiling structurally impossible, because the carrier a template must write is a pure function of the field's resolved type and the recursive-type set -- the same join rust_field_carrier_final_type already performs when it renders the struct. The next-rung trigger names that capability and explicitly refuses a lint, a comment, or this one repair as discharging it. The row carries its own uncertainty: one specimen, population not counted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
…t per file The context attribution left one question open and named the instrument for it: the fold costs a flat ~840 us per token with 80% in parse, and nothing printed says whether that is a memo that never hits or genuine per-token work in the combinators. ParseTable already tracked memo_hits / memo_misses / memo_lookup_calls and parse_table_memo_stats already existed; nothing carried them out of the parse, so nothing could read them. NOT A FIELD ON ParseArtifact, WHICH WAS THE SHORTER CHANGE, FOR TWO REASONS. ParseArtifact answers WHAT was parsed; the counters answer HOW that parse executed, which is a realization fact about one run and not a property of the tree -- DESIGN section 3 keeps interface and realization as two facts, and a counter field would make every consumer of a parse result carry a measurement it has no use for. The second reason is decisive rather than stylistic: an Outcome's Rejected arm carries no ParseArtifact, so a field there would lose the accounting exactly on the files that spend seconds in parse and THEN refuse -- 9.8% of the fold, and the expensive case this instrument exists to read. SO THE CARRIER HOLDS BOTH. ParseProductionMeasured carries the outcome and the accounting side by side, and both refusal arms report. The unmeasured entry points are retained as PROJECTIONS of the measured fold -- parse_production_prepared, parse_module_prepared and program_assembly_phase_parse are `.outcome` of it -- so there is one implementation, no caller pays a second parse to be measured, and the five existing parse_module_prepared test callers are untouched. ONE ARM IS SPELLED OUT RATHER THAN REUSED, and the reason is stated beside it: program_assembly_phase_parse_measured cannot route through bind_outcome, because bind_outcome cannot carry a second value out of the bound function. Its Rejected arm passes the diagnostics through untouched and reports EMPTY accounting -- zero because no parse ran, which is an answer and not a missing measurement -- and its Accepted arm merges through bind_outcome_accepted, which is what bind_outcome itself calls. A void grammar likewise reports empty rather than absent stats, so a void parse is distinguishable from a parse whose stats were dropped. The driver prints the three counters per file in [native-context-split] and their totals in [native-context-partition], read OUTSIDE every phase span so they price none of them. NOT YET COMPILED AS AN EMITTED CRATE. Both modules typecheck under the whole dag + src/v2 closure and required-regen reaches first_generation_equal, but the emitted crate build -- the only check that catches the boxed-carrier class this branch already filed -- needs the route to build on main, which waits on #11216. No measurement is claimed from this commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
…13-perf-c-main # Conflicts: # src/v1/stage0/src/v1_compiler_emit_rust.rs
…gling imports TWO FINDINGS, BOTH VERIFIED AGAINST THE CODE BEFORE ACTING. DANGLING IMPORTS (section 3c). ParseMemoStats was imported into v2.compiler.program_assembly and referenced nowhere. parse_module_prepared was left reachable only from a prose String row once program_assembly_phase_parse started calling parse_module_prepared_measured -- pre-existing name, stranded by this diff. Both removed; the module typechecks under the whole dag + src/v2 closure. The prose row still names parse_module_prepared, which is a quarantined String and not a code reference, so it is untouched. THE RE-SPELLED SEQUENCING IS FILED RATHER THAN WITNESSED, AND THE REVIEW OFFERED EITHER. The driver enters tokenize / parse / normalize one at a time so each call can be bracketed by a clock, so the ORDER and the inter-phase diagnostic merge exist twice: here and in the modeled composition program_assembly_read_to_normalized_root_prepared. The review's first option was a witness that the two agree. THAT WITNESS WOULD BE VACUOUS AND THE ROW SAYS SO: the composition is DEFINED as the phase chain, so a .dag witness comparing them compares a definition with itself and cannot go red, and the state that must be caught -- a reordered TARGET-LANGUAGE template -- is not representable in .dag at all. DESIGN section 4b says to ask whether the check's RED is authorable before writing the check; here it is not, and a permanently green equivalence check would be cited as coverage. So the class is filed: gunbc.recurring_failure_mode.realization_respells_a_modeled_folds_sequencing_to_instrument_it. It states what is already single-authored (every phase body, and the absorbing half), why it is NOT parallel_representation_debt (there the canonical route was usable and the duplicate deletable; here deleting the re-spelling deletes the measurement), its rung as mitigatable, its ceiling as structurally guaranteed rather than impossible -- the seam removes the REASON to re-spell, not the ability, and claiming rung 4 for a state that stays authorable would be inflation -- and a trigger naming the CAPABILITY: a realization seam that yields per-phase spans FROM a modeled fold. The trigger explicitly refuses the vacuous witness, a comment, a lint over template text, or repairing this one driver as discharging it. The emitter comment that asserted the equivalence is replaced: it now names the row and records that an earlier revision claimed the two are the same computation with nothing executing it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
|
Both findings from review 65068 are addressed at Dangling imports (§3c) — fixed. The re-spelled sequencing — filed, not witnessed, and the choice is the substantive part of this reply. You offered either (a) a witness that the driver's phase-by-phase route yields the same Filed as
The emitter comment that asserted the equivalence is replaced. It now names the row and records that an earlier revision claimed the two are the same computation with nothing executing that claim — your wording, and it was fair. Also in this push: Thank you for the "not findings, checked and cleared" section — the roster/gitignore and unit-modeling checks saved me re-deriving them. — sent from nimble-ibex-813 |
…itation to a symbol that does not exist THE STALE CITATION IS A REAL SECTION 3 DEFECT AND IS DELETED. The annotation named `program_assembly_file_front_end` as the pure per-file parallel unit; `grep -rn file_front_end` over dag and src returns that comment line and nothing else. The symbol never existed -- it is a name from an earlier drafting of this seam that the comment outlived -- and it also overstated the work, because the per-file unit that exists is the driver's Rust loop body, not a .dag declaration. The sentence now says what is true: each phase is pure in its inputs, so the per-file sequence is also the parallel unit. THE FORK IS REAL AND IS NOW STATED RATHER THAN UNNOTICED (DESIGN section 3b). v2.compiler staged_front_end is the declared authority for the .dag front-end ORDER and its own header names re-inlining its calls as the section-3 nickname class. These phases ARE a second stepwise spelling of tokenize -> parse -> normalize. Section 3b admits a divergence with a stated reason and reds only an unstated one, so the reason is stated beside the phases and it is MEASURED rather than asserted: run_front_end_to calls dag_language_model() per call and front_end_parse_step calls parse_module with a RAW grammar, so routing this fold through it re-derives the language model and re-runs prepare_grammar ONCE PER FILE -- the two per-file recomputes this fold already hoisted, which this instrument's own receipt prices at one call and 0.13 s across the 2243-file closure. It also runs a fourth stage, resolve, that the context fold does not perform, and returns a whole FrontEndRun rather than stages a clock can sit between. THE ROW NOW REACHES THE MODULE THAT OWNS THE ORDER. The review's sharper point was that the filed class argued only against the local composition and never reached staged_front_end, which understated the fork; the row carries that correction in its own words, carries the measured reason the authority cannot be consumed as it stands, and names the consolidation in its trigger -- a prepared-grammar, resolve-free, per-stage entry on staged_front_end, after which these phases are projections of it and the divergence note is deleted. staged_front_end run_front_end_to is added to the row's evidence. THE CONSOLIDATION IS DELIBERATELY NOT IN THIS PR. It is a modeling change to a load-bearing module with its own witness suite, and this brief is a measurement brief; doing it here at the tail of a review cycle is the improvisation the intro docs warn against for load-bearing files. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ
|
Both findings from review 65114 are addressed at The stale citation (§3) — deleted. The §3b admits a divergence with a stated reason and reds only an unstated one, so I have stated it, and the reason is measured rather than asserted — which is also exactly why the authority cannot be consumed as it stands, the caveat you anticipated:
Those are the two per-file recomputes this fold exists to have hoisted, and this PR's own receipt prices them at The consolidation is named, not hand-waved. A prepared-grammar, resolve-free, per-stage entry on Why it is not in this PR, stated so the deferral is a decision rather than an omission: extending
The "everything else reads clean" section is appreciated, particularly the read on — sent from nimble-ibex-813 |
|
Gatekeeper note (wind-down): this head now CONFLICTS with main after #11290 landed (0306cf9: compiler_entry.dag / 05_emit_rust.dag / stage0 mirrors). Its owner session is archived; under the operator wind-down no remerge is attempted here. Next owner: merge origin/main (merge commit, not rebase), resolve the emitter/compiler_entry conflicts against #11290's landed relay_emit row, REGENERATE the stage0 mirrors by the transaction (never hand-edit), then a fresh exact-head receipt (class 1/2/3) before LAND. Any receipt pair started on the old merge base is void. — sent from eager-raven-113 |
…d kept #11219's stale mirror, which references a helper #11444 removed) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RtujL6ZdnmmcCBMihVRZ8z
|
Superseded by #11521 (recut/11219-perf-c): same change merged onto current main with the context-fold conflict resolved semantically (prepared front end threaded into native_test_context_finish) and the emitter mirror regenerated by the transaction. — sent from eager-raven-113 |
Perf item C, deliverables 1 and 3, plus the memo-counter follow-on. #10940 and #11216 have both landed; this is the REPLAY of that work onto main (not a merge — my old branch carried deep-swift-530's individual commits while main has them squashed, so a merge would have duplicated them). Supersedes #11206, which GitHub auto-closed when its base branch was deleted at merge.
Result first: the expected defect is ruled out by measurement
The brief said to EXPECT a per-file recompute of something invariant, and to find it or rule it out before any parallelism. It is ruled out, and so is a size-quadratic. What is there instead is a flat per-token constant whose location inside the parse this PR does NOT claim to have found.
One run, one basis. Every figure below comes from a single execution of the instrument on this PR's head (
26628e29), srv1, load ~55 at start, emitted driver run directly inadjudicatemode so the relay streams. Figures from any other run are not mixed in.[native-context-partition], basiscontext_span, full closure, 2243 files / 2,286,564 tokens / context 2070.21 s:front_end_preparetokenizeparsenormalizeabsorb(qualified-name + index insert)eprintln)file_refusals=212, per-file span p50 392 ms, p95 2682 ms, max 86.0 s. Memo totals from the same line: 6,557,074 lookups, 400,160 hits (6.1%), 6,156,914 misses, 2.87 lookups/token.The discriminator
Median ns/token by token decile (1 = smallest files), and the per-attempt quotient beside it:
Deciles 3–10 are flat on every column. Three hypotheses die here:
front_end_prepare_calls=1, 0.13 s for the whole fold. The language model and prepared grammar are paid once and carried, and the relay proves it rather than asserting it.absorbis 0.43 s of 2070 s (0.02%).list_snoc_itemon roots andsource_root_index_insertare measured out, not argued out.parse_table_for_production_analyzedreusesprepared'sGrammarFirstAnalysis; no per-file recompute, andfront_end_prepare_calls=1is the receipt.Deciles 1–2 are elevated on both time columns, and that is fold position, not size. Tested directly on the earlier run by holding size fixed: among files of ≤90 tokens, those in the first quarter of the fold cost 247 ms median against 57.6 ms in the last quarter — a 4.3× position effect at the same size. It reproduces here (deciles 1–2 at 942/863 µs/miss against 228–303 for the rest). I have not fully attributed it; it is consistent with first-touch page-fault billing, which is already a rostered class, and I am not claiming it as a per-file constant.
What is actually there
~600 µs per token across the flat range, 80.5% of it in parse, at ~2.9 memo lookups and ~2.7 misses per token. A recursive-descent parser is normally well under 1 µs/token. This is a different defect class from the one the brief expected — not an invariant recomputed per file, but a cost that is constant per production attempt — and this PR measures it rather than fixing it. Where inside the parse that cost sits is NOT established here; see the narrowings below.
Second finding, not asked for
file_refusals=212of 2243 — the count the relay EXECUTED and reported, not a classification applied afterwards. On the earlier run, classifying the parse/tokenize subset by signature (normalize binds theRejectedand returns in <10 µs while parse has spent seconds) gave 193 files and 9.8% of fold time. A tenth of the context fold front-ends files that contribute no root — the v2 front-end corpus frontier (v2.workflow.compile_door_ledger) paid for in full before it refuses. Flagging, not taking.An earlier revision of this PR circulated 19.5% for that share. That was an artifact of reading only the alphabetical PREFIX of the walk (the
dag/tree); 9.8% is the whole-closure figure and is the one to use.The seam
v2.compiler.program_assembly— the per-file front end becomes one declaration per phase:program_assembly_phase_tokenize(read, lm) -> Outcome<TokenStream>program_assembly_phase_parse(tokens, prepared, residue) -> Outcome<ParseArtifact>program_assembly_phase_normalize(parsed) -> Outcome<NormalizedTree>program_assembly_phase_token_count(tokens) -> IntEach phase takes the PREVIOUS PHASE'S OUTCOME and binds it, so diagnostic merging and refusal propagation stay inside
.dag.program_assembly_read_to_normalized_root_preparedis retained and is now defined by those three calls — entering the phases separately is the same computation, not a second front end beside it.v2.compiler.00_compile— the matching seam:native_test_front_end_prepare()carries the once-per-fold invariant asNativeTestFrontEnd;native_test_context_absorb(state, read, root_outcome)is the non-front-end half of the step;native_test_context_state_empty()/native_test_context_finish(state, ingest).native_test_context_from_ingestis retained as the composition every clock-free caller uses — including the driver's ownrun_census, which wants per-file refusals and no attribution.std.compiler_entry.SourceRootEvalDriver— a phase span is a REALIZATION fact (.dagis pure and cannot read a clock), so the driver that owns the clock enters each phase and times the call. It prints[native-context-split]per file (path, bytes, tokens, four phase spans) and[native-context-partition]once.Counts sit beside times because a span alone cannot separate "this file is large" from "this file pays a per-file constant" — the whole question here. The token count is read OUTSIDE every phase span so it prices none of them.
No double counting
Per eager-raven-113's binding accounting ruling: these rows are their own relay on their own basis (
basis: context_span) and partition the CONTEXT span.[native-cost-partition]'s exclusive rows are untouched —contextremains one exclusive row of the driver wall, so the parent span still reconciles. The four phases partition the per-file span with 0.00 s unattributed; the 1.2% residual is the loop and the per-fileeprintlnand is reported, not absorbed.Deliverable 3
1256is a real closure, not the tree by another name. Independent oracle over the same two roots — seeds = every declaredv2.test.*module (thenative_route_universe_prefix) plusnative_route_live_control_module, edges = theimportheaders each source declares, transitive to fixpoint — gives on the full tree: scanned 5578, closure_modules 2091 in 14 rounds, closure FILES 2089. Genuinely closure-scoped at ~37% of the tree, and the driver's own fold independently reports 2236 files over the same roots.That also explains the previously unexplained gap of 4 between
closure_modulesandclosure_reads:closure_modulescounts module NAMES reached,closure_readscounts FILES ingested, and the difference is closure members with no ingested declaring file.native_lane_ingest_matches_closureis consistent with this: its "closure module was never ingested" arm is quantified over modules some source DECLARES, so a closure member with no declaring source in scope is not a refusal.Correcting my own first report of this. I initially described the two such modules on the full tree as "declared nowhere". That was wrong in both cases, and they are wrong for different reasons, which is the part worth keeping:
gunbc.recurring_failure_mode.rosteris derived and gitignored (.gitignore:79), produced bygunbc.instruments.generated_artifact_gatemain_wetfrom its sibling row files. It is declared by no committed source, so it is absent from a fresh clone and present in a working tree that has run the projection. "Declared nowhere" understates it: the declaring file is generated, not missing.v1.compiler.ownershipis declared, atsrc/v1/ownership.dag. My oracle scanned onlydagandsrc/v2— the route's own two roots — so this is a module declared OUTSIDE the scanned roots, not an undeclared one.The mechanism behind the gap is unchanged; only my characterisation of the instances was wrong. It matters to a reader because "declared nowhere" would send someone hunting for a missing module, and neither case is that.
The oracle is stated here as a method rather than committed as an artifact, deliberately: a
.dagfunction undergunbc runcannot read the corpus, so the only available forms would be a second closure implementation in the corpus — a §3 fork ofnative_lane_closure_modules, which already answers this — or a hand-authored script, which §6 presumes scaffold.Deliverable 2
Not in this PR, by instruction: parallelizing before the per-file cost is honest hides the defect. The seam landed here IS the fold-parallel seam —
program_assembly_phase_*is pure in(read, lm, prepared, residue)— so D2 is a realization change over these same functions. The measurement changes its priority in both directions: there is no per-file defect left for parallelism to hide, but it would multiply throughput without touching an ~840 µs/token constant that §6 prices as a cost-shape defect in its own right, and ~10% of the parallel work would still be spent parsing files that refuse.The boxed-carrier emitter defect this change surfaced
The first emitted crate failed to compile, and that is a finding rather than a footnote.
v1.compiler.emit_rustneeds_box_wrappingboxes a record field whose type rides a recursion cycle, and an Optional-cardinality field inherits that decision from its inner type;rust_field_carrier_final_typerendersBox<T>. So the new carrier'sresidue: Diagnostics(=Optional<NonEmptyDiagnostics>, andNonEmptyDiagnosticsis recursive) emitted asBox<Option<Rc<NonEmptyDiagnostics>>>whileprogram_assembly_phase_parsetakes it by value —front_end.residue.clone()cloned the Box:error[E0308]atsrc/main.rs:407, help "consider unboxing the value". Its sibling fieldslmandpreparedtookRc<..>and their.clone()was correct, which is the per-field divergence precisely.The tempting summary "Optional fields box" is WRONG and is not what this filed: whether a field boxes depends on a whole-corpus recursive-type set, so the realization is not readable off the declaration a template author writes against. What did NOT catch it, each checked rather than assumed:
gunbc runtypechecked both modules under the whole closure and was green;cargo fmt --all --checkpassed;--required-regenregenerated and adjudicated the very file CONTAINING the defective template text and reachedfirst_generation_equal, because regen compares emitted BYTES against the authority and never compiles the program those bytes describe.Filed in this PR as
gunbc.recurring_failure_mode.emitted_field_carrier_diverges_from_the_dag_type_authored_against, citingneeds_box_wrappingandrust_field_carrier_final_type. Deliberately NOT filed underaccepted_source_emits_uncompilable_target: there the EMISSION is wrong; here the.dagand its emission are both correct and the hand-authored template consuming them is wrong, so a repair to the emitter's type lowering fixes that class and does nothing for this one. Rung found at mitigatable (typed, located rustc diagnostic; the route refuses withEmittedCompilerBuildFailed); attainable ceiling structurally impossible, because the carrier a template must write is a pure function of the field's resolved type and the recursive-type set — the same join the emitter already performs when it renders the struct. The row carries its own uncertainty: one specimen, and the exposed population was not counted.The memo counters, and the §3b divergence they are carried by
The context attribution left one question open and named the instrument for it.
ParseTablealready trackedmemo_hits/memo_misses/memo_lookup_callsandparse_table_memo_statsalready existed; nothing carried them out of the parse, so nothing could read them. This PR carries them out and prints them per file in[native-context-split]and as totals in[native-context-partition], read OUTSIDE every phase span so they price none of them.DIVERGES WITH A STATED REASON (§3b), and the instruction it departs from is named. The direction given was to thread the counters "through
ParseArtifact". This does not do that, for two reasons and the second is decisive:ParseArtifactanswers WHAT was parsed; the counters answer HOW that parse executed, which is a realization fact about one run and not a property of the tree. §3 keeps interface and realization as two facts, so a counter field would make every consumer of a parse result carry a measurement it has no use for, and would make the artifact's identity depend on how many times a memo happened to be consulted.Outcome'sRejectedarm carries noParseArtifact. A field there would lose the accounting exactly on the files that spend seconds in parse and THEN refuse — 9.8% of the fold, and the expensive case the instrument exists to read.So
ParseProductionMeasuredcarries the outcome and the accounting side by side and both refusal arms report, andparse_production_prepared,parse_module_preparedandprogram_assembly_phase_parseare retained as.outcomeprojections of it. One implementation, no fork, no caller pays a second parse to be measured, and the five existingparse_module_preparedtest callers are untouched. One arm is spelled out rather than reused —program_assembly_phase_parse_measuredcannot route throughbind_outcome, becausebind_outcomecannot carry a second value out of the bound function; itsRejectedarm passes diagnostics through untouched and reports EMPTY accounting (zero because no parse ran, which is an answer, not a missing measurement), and itsAcceptedarm merges throughbind_outcome_accepted, which is whatbind_outcomeitself calls.What the counters measured, and what they do NOT establish
A MISS is one production attempt: the memo had no entry for that (position, production), so
parse_expr_with_firstran. On the measuring run:lookups/token(2.47–3.05) and that quotient (228–303 over deciles 3–10) are FLAT across the size rangeWhat is established. Lookup volume is ~2.3 per token and does not grow with file size, and whatever the parse spends per attempt does not grow with file size either. Both are flat, measured, and load-independent as ratios.
What is NOT established, stated because an earlier revision of this section over-claimed all three:
parse_nanos ÷ memo_misses— total parse time divided by a miss count. That numerator includes the common pre-dispatch path, hit handling and result construction as well as the miss-path parse, so the quotient sizes an average attempt's share of all parse work; it does not say which part of the parse holds the cost. Misses being 94% of lookups makes the miss path the largest single candidate, not a located one.ParseTableRealizationrebuild is a source operation count, not a measured cost.parse_table_record_lookup_call,parse_table_record_missandparse_table_inserteach construct a fresh 12-field record carrying fourMaps, three times per miss — that is read off the source and is a suspect worth measuring first precisely because a record rebuilt per attempt would be flat in file size exactly as observed. Nothing here measures it.Also withdrawn: an earlier revision said the 6.1% hit rate makes the parse memo "overhead rather than a win". That does not follow — a low hit rate can still be economical if each hit avoids an expensive subtree re-parse. The real question is avoided recomputation against lookup + insert + retention cost, and this PR does not measure it.
Subjects stay separate. 2.69 misses/token × 270.7 µs ≈ 728 µs/token is arithmetic over two figures from THIS run's own relay line, not a cross-run reconstruction; it is not offered as confirming any figure from any other run. Each figure binds its own run, population, phase, binary and input. Nothing here is evidence for a figure produced by a different run.
So the next lane LOCATES the cost rather than confirming a conclusion: a split inside one attempt on the emitted realization — common pre-dispatch / hit handling / miss-path parse / insert-and-result. One hazard is named up front for whoever takes it:
parse_expr_with_firstattempts NEST, so summing inclusive attempt durations would recreate the double-counting class just repaired at module grain. Report attempt spans as inclusive and never sum them into an exclusive partition, or subtract descendants through a nested accounting.Launch gate: the emitted crate builds on THIS head
Applied to myself before any measuring run, on head
26628e291404c2920f9a7e4562f761039acac2bd(merge commit;origin/main5f18d9f is an ancestor):exit_status=0 warning_count=0underRUSTFLAGS="-D warnings", and the malformed control refused correctly on the same run (tokenize_lex_e1_unrecognized_char). This is the gate that matters for this branch specifically: it is the only check that catches the boxed-carrier class filed below, and.dagtypecheck plus--required-regenare both green on code that fails it.Verification
.dagmodules typecheck under the wholedag+src/v2closureRUSTFLAGS=-D warnings, and the malformed control refuses correctly (tokenize_lex_e1_unrecognized_char)--required-regennamed exactlyv1_compiler_emit_rust.rsas drifted — also the positive control that the emitter edit is live — and the regenerated file is committedsession/deep-swift-530was resolved through the merge driver's declared repair route (both.dagauthorities auto-merged; only the generated projection conflicted), verified atfirst_generation_equal=trueANDfixed_point_equal=true, and confirmed both relays survive in the merged projection (native-context-splitand deep-cat'snative-prepare-split) — a textual merge previously dropped that line with no conflict marker. This PR then replays onto main and regenerates again,first_generation_equal=true.One caveat on the subject: these numbers are over the FULL closure at this branch head, not the bounded subject, so they are NOT directly comparable to the 1400.79 s bounded figure — same host, different subject. The full closure turns out to be the better subject for D1 anyway: the context fold completes and reports before any per-module prepare, so it needs no reduced claim population to finish.
🤖 Generated with Claude Code
https://claude.ai/code/session_019gufjy7FmMzjTRwAnUgZcZ