Skip to content

concat realizes the persistent append its contract declares (normalize O(n·depth) → O(n log n); ampere 147s → 15.5s) - #13461

Merged
gunbai-bot[bot] merged 2 commits into
mainfrom
concat-persistent-append
Oct 7, 2026
Merged

gunbai-bot[bot] merged 2 commits into
mainfrom
concat-persistent-append

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

The v2 native census took hours and looked hung (runs 37348196148 and 37197235156). The cause was concat's realization. It was supposed to be cheap and was not.

What was wrong

std.primitives concat_contract declares that concat costs log(n + m) on the persistent carrier. The comment next to the rc_* ops in v1.compiler.runtime_rust rt_rc_container_ops says the same. But the code did not do that:

  • rc_list_concat ran Rc::make_mut(&mut a).extend(b.iter().cloned()), pushing b onto a one element at a time.
  • The bare-Vec concat and list_concat used the same extend pattern.
  • The interpreter's free concat arms were worse. They copied both operands into a new vector.

So every concat cost the full length of its right operand.

v2.std.node node_subtree_nodes is a bottom-up fold whose step is concat(acc, sub). With that realization its cost was O(n · depth). v2.compiler.normalize normalize_with calls it on every parse tree. A comma list parses as a right-nested chain about N deep, so one large list literal was quadratic.

perf on the emitted compiler, running census over extdeps.cpu.ampere_altra_package_table4_raw alone, showed:

  • 96.6% inclusive time under node_subtree_nodes;
  • 82% under rc_list_concat;
  • self time in im chunk make_mut, drop_slow and Rrb push_back.

The interpreter did not show this. With its call memo off, eval-step counts stayed linear (×2.00 per doubling), because the extra work happened inside a single counted step. Only a profile of the emitted binary named it.

The fix

  • Runtime template. In src/v1/runtime_rust.dag, concat is now an im::Vector append onto a clone of b, an O(1) structural share that never iterates b. This covers rc_list_concat, the V2Concat impl for Vec<T>, and list_concat.
  • Audit of the sibling list ops. rc_list_push, list_push and list_snoc_item were already fine. Map merge, set union and reverse iterate their input, but they are not list concat, so they are out of scope here.
  • Generated files. v1_compiler_runtime_rust.rs and v1_rt.rs come from one claim_executor --required-regen run that reached a fixed point in 3 rounds. They were not hand-edited.
  • Interpreter. In v1_interpreter.rs, the free concat arm and the binary-operator concat path now use value_to_list_carrier plus an RRB append. That is the shape the method and list_concat arms already had. The copy counters are charged only when a free-monoid chain has to be materialized.
  • Annotation correction. gunbc.build_target aggregate_findings carried an annotation that checked rc_list_concat and said it was cheap. It looked at the clone and missed the extend after it. The annotation now records that error.
  • New RFM row: realization_violates_its_declared_complexity_contract. It is separate from primitive_operation_carries_two_cost_authorities: there the declarations disagree, here they agree and the realization contradicts both. The row's next-rung trigger is a scaling control run against the emitted realization of each contracted primitive.

Controls

Both sides ran on main 8e95149ecd with the same script, in separate remote dispatches. The "before" side checks out main's src/v1. Times are for the emitted compiler's census over a root containing just that file.

input before after
list literal, 1,000 records 5.9 s 3.3 s
list literal, 2,000 records 29.6 s (×5.0) 6.8 s (×2.04)
list literal, 4,000 records 123.1 s (×4.16) 15.2 s (×2.24)
ampere_altra_package_table4_raw.dag (4,926 rows, 54k tokens) 147.1 s 15.5 s

Semantic equality. I fingerprinted the normalize output on both sides. The fingerprint is the root's content_hash, the node count, and a sum over every node's occurrence identity. Both sides gave identical results:

  • 300-record list: 1d3d229b95d4e5ce/5602/37759245 before and after.
  • extdeps.ampere.nvparam: e7970b81c12a6cc3/9799/179843194 before and after.

Out of scope

node_subtree_nodes is called about 200 times across the corpus to build a whole-tree list that each caller then scans. That is linear but redundant demand, and it is tracked as a separate follow-up.

Merge

This PR touches src/v1, so it waits on the src/v1 hold until #13388 lands.

🤖 Generated with Claude Code

gunbc-ci-auto-heal and others added 2 commits October 6, 2026 08:41
rc_list_concat and the bare-Vec concat/list_concat extended the left
operand one element at a time, so concat cost the length of its right
operand against concat_contract's log(n + m). node_subtree_nodes, a
bottom-up concat fold called on every parse tree by normalize_with, was
therefore O(n * depth): 161 s on one 4,926-row list literal on the
emitted compiler. Realize concat as an im::Vector append in
runtime_rust.dag and in the interpreter's free concat arms; file the
class; correct the build_target annotation that had cleared it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…laim_executor --required-regen, fixed point in 3 rounds)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit a91f7a7 Oct 7, 2026
5 checks passed
@gunbai-bot
gunbai-bot Bot deleted the concat-persistent-append branch October 7, 2026 05:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants