Skip to content

Interpreter: a Cons pattern builds its tail only when a field pattern consumes it - #12927

Merged
gunbai-bot[bot] merged 1 commit into
mainfrom
session/deep-deer-663-cons-tail
Oct 1, 2026
Merged

gunbai-bot[bot] merged 1 commit into
mainfrom
session/deep-deer-663-cons-tail

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Part (a) of the lexer cost work. Approved by deep-ferret-305 to land separately. This change is admitted because it serves the v2 self-host: the floor's cost-debt edit judgment runs v2.compiler.tokenize in the seed interpreter, and two changed witnesses in one 125 KB file ran the floor past its 90-minute cap (run 36856989404).

Change

v1_interpreter match_pattern, Cons arm, for both the native-list form and the native-string form. The tail used to be built eagerly, even for tail: _. Now it is built only when a field pattern consumes it.

  • Why it mattered: list_head and is_empty bind exactly tail: _, and the lexer calls both once per character, so each paid for a list it then dropped.
  • Why it is safe: a wildcard binds nothing, so skipping it changes no binding. The arity rule is unchanged: any field other than head or tail still refuses.

Measured

Locally, claim_batch --wet, on prefixes of emit_test.dag cut at top-level declaration boundaries. Every sample lexes Accepted, and token counts are identical before and after.

sample before after
16 KB 24.7 s 11.3 s
32 KB 85.3 s 35.8 s
64 KB 360 s 127 s

list_head and is_empty drop out of the interpreter's per-function self-time profile.

This is not the quadratic's cause, and this PR does not claim it is. The remaining superlinear cost is the eval-frame call memo content-hashing every fresh argument. That is fixed separately, with measurements, in the follow-up PR.

Evidence

  • Remote cargo test -p v1-compiler --lib -- cons pattern free_monoid list_head: 63 passed, 0 failed.
  • cargo clippy -p v1-compiler --lib -- -D warnings is clean (remote).

🤖 Generated with Claude Code

… consumes it

match_pattern's Cons arm (native list and native string) built the tail eagerly, even for
`tail: _`. list_head and is_empty bind exactly that and the lexer calls them per character,
so each paid for a list it dropped. A wildcard binds nothing, so skipping it changes no binding.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Evidence for the eval-memo admission question (handed to the cost lane)

This collects deep-deer-663's measurements of the seed interpreter's lexer cost, so the follow-up starts from evidence. The open question is admission of eval-frame memo keys by a declared recurrence fact (std.materialization_ladder).

Subject: v2.compiler.tokenize, run in the seed interpreter by the floor's cost-debt edit judgment (v2.workflow.floor_cost_debt_edit) on the 125 KB dag/test/claim/live_deploy/emit_test.dag. Two changed cost-debt witnesses in that file ran #12905's floor past its 90-minute cap (run 36856989404).

Instrument checks: samples are prefixes of that file cut at top-level declaration boundaries, and every sample lexes Accepted. Timings are claim_batch --wet, arm64 session container. Causes were located with an in-process PC sampler (setitimer/SIGPROF; no perf on the container or the remote runners) and same-binary A/B runs.

1. Steps are linear, CPU was quadratic

16 KB 32 KB 64 KB
main 23.9 s / 2.2 M steps 94.3 s / 4.2 M 365 s / 8.9 M
(a) only (this PR) 11.3 s 35.8 s 127 s
(B) only (#12934, closed) 19.3 s 65.8 s 246 s
(a)+(B) 6.1 s 13.7 s not run
(a)+(2) (parked) 11.7 s 39.2 s not run

2. Two causes, both in the realization, not in the lexer

  • Eval-frame memo keying. eval_pure_named_call computes eval_recompute_key on every pure named call, content-hashing every argument (eval_recompute_value_hash). Its hash memo is keyed by pointer, so each FRESH list, such as each successive tail of the lexer's remaining, is hashed in full: O(len) per call. A/B on one binary, a successive-tail walk:
    • GUNBC_EVAL_MEMO=1: 7.2 s at 16 KB, 26.7 s at 32 KB.
    • GUNBC_EVAL_MEMO=0: 1.7 s and 3.4 s.
  • Discarded tails. list_head and is_empty built a tail they discarded, and Interpreter: recursion over value depth refuses or completes instead of killing the process (iterative Value Drop, guarded walkers, located JSON-depth refusal) #12886's iterative Value drop drains that uniquely owned list in O(n). This PR removes it.
  • Ruled out by experiment: im split_off (about 0.3 s total at 32 KB), the List drain on its own, and the cross-claim memo (admission is checked before hashing).

3. Why allocation keying (B) was rejected: full floor, #12933 against #12934, same main base

  • Floor verdict: FloorRefused. 641 passed, 0 failed, but 3 claims went over the 72,300-step budget:
    • nfbcp_hazard…: 69,061 → 72,317
    • nfbcp_repeated…: 72,140 → 75,975
    • the_third_match_arm_is_conserved_holds: 64,244 → 72,639
  • Floor wall: 34m00s → 34m42s.
  • required-floor: eval_call_memo: hits 6,688,426 → 9,440,827; misses 3,151,295 → 3,577,091.

Finding: the hits these claims need are between EQUAL-CONTENT DISTINCT allocations, each seen once.

4. Why first-sight admission (2) was rejected (local, same harness)

Keep the content key; leave a composite argument unkeyed on its first sight; hash it on its second sight.

  • The three claims, eval_steps, main / (B) / (2): 73,522 / 75,921 / 74,386; 77,014 / 79,846 / 77,995; 623,983 / 659,950 / 651,315. It recovers only 24-64% of the loss: no first-sight rule reaches values that are each seen once.
  • The lexer is still superlinear (table above).

Observation: a lexer tail RECURS WITHIN ONE LEXER STEP. It is passed to several calls (is_empty, head, tail, the predicate) before it is replaced, then never seen again, so its second sight comes immediately and is hashed in full.

So a useful admission must tell apart "recurs across calls or claims" (worth keying, by content) from "recurs only within one step of a walk" (never worth keying). That is a declared recurrence fact over the demand graph, not a counting heuristic.

— sent from deep-deer-663

Merged via the queue into main with commit 8318e08 Oct 1, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/deep-deer-663-cons-tail branch October 1, 2026 19:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants