Skip to content

Strings: a short + chain no longer pays for a deferred representation it never takes - #3571

Merged
lahma merged 1 commit into
sebastienros:mainfrom
lahma:fix/3527-short-concat-eager
Sep 1, 2026
Merged

lahma merged 1 commit into
sebastienros:mainfrom
lahma:fix/3527-short-concat-eager

Conversation

@lahma

@lahma lahma commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

date-format-xparb pays +3.91% since #3386 (#3527), and the profile in that issue put the cost under AdditionChainExpression: RhpNewVariableSizeObject under the chain ×2.0, RhpNewPtrArrayFast ×3.3, allocation events +23%, while String.Concat — the cost the deferred copy was meant to displace — did not move at all. The workload allocates more under the new lane, not less.

The cutoff was never the problem

The issue proposed adding a length cutoff to the chain lane. There already is one, and it was already doing its job: JsString.MinDeferredConcatenationLength is 512 characters, and both the pairwise + (inside JsString.Concat) and the flattened chain (ConcatThree / ConcatFour / ConcatMany) test the summed operand lengths against it. xparb's chains are "2007-01-01" and "Monday, January 01, 2007 1:11:11 AM" — ten and thirty-five characters. They were on the eager side of the cutoff the entire time, and no threshold value can change that. Moving the line would only have pushed more chains onto the deferred path, which is the expensive one here.

What they paid for was the shape of the path below the line, which #3386 changed without meaning to:

before #3386 after #3386
operand coercion TypeConverter.ToString → a string TypeConverter.ToJsString → a string plus a JsString wrapper, for every operand that is not already a string
ConcatMany storage one string[] a JsString[] and a string[] built from it for string.Concat

xparb's "l, F d, Y g:i:s A" compiles to one fifteen-operand chain. Per evaluation that is a second 144-byte array plus a wrapper per numeric operand — getFullYear(), the hour expression — every byte of it dead before the result was copied out. That is the doubled variable-size allocations and the tripled pointer-array allocations the profile measured, and none of it existed before #3386.

What changed

An operand is now held in whichever form its coercion already produced: the JsString it was — whose representation has to survive, because flattening it is precisely the copy the deferred lane exists to avoid — or the plain text a non-string primitive coerced to. Both answer their length without materializing anything, and the length is all the cutoff is decided on, so holding the cheaper form until the decision is made gives nothing up.

Below the cutoff the result is then joined through a ValueStringBuilder over a cutoff-sized stack buffer — free, since the assembly carries [module: SkipLocalsInit], and it cannot grow because the cutoff bounds it — so there is no second array and no wrapper. Above the cutoff the fold through JsString.Concat is untouched.

The cutoff itself stays at 512, and its remarks now carry the arithmetic that was missing: a RopeString is 56 bytes on 64-bit and does not remove the flat allocation, it postpones it, so a result read once costs 56 bytes and an indirection more than copying it would have. The node earns them back only across iterations. Which is why the line sits far above the point where a node is merely affordable — and why what the path below the line allocates matters more than where the line is.

Every lane that can reach the deferred representation

JsString.Concat is the only producer of RopeString, and it has exactly two callers.

lane verdict
ApplyAdditionToPrimitives — the pairwise +, and the pairwise fold a numeric chain falls back to covered. Same 512 cutoff as before. Two operands that are already strings now take a branch that coerces nothing at all; the mixed pair (n + "x") coerces to text and builds the wrapper only on the branch that goes on to defer.
AdditionChainExpression — the flattened a + b + c […], both the direct path and the resumed one, which share the same three joins covered, as described above.
+= — JintAssignmentExpression exempt. Builds JsString.ConcatenatedString, the StringBuilder-backed mutable accumulator. Never calls JsString.Concat, never produces a RopeString. #3386 did not touch it and neither does this.
String.prototype.concat exempt, for the same reason: EnsureCapacity / Append on ConcatenatedString.
JsString.CreateSliced → SlicedString out of scope. A different deferred representation (a zero-copy view) with its own cutoff and retention budget, not introduced by #3386 and not implicated by the profile.

Semantics

Unchanged, deliberately and checkably. Every path produces the same characters through the same JsString.Create, so the same shared instances are returned for the empty and single-character results, and the same representation decision is taken on the same numbers — typeof, .length, indexing, equality and hashing all see exactly what they saw. Coercion order is preserved (left operand before right; ToString of a symbol still throws from the same side), and the length guard still runs on the summed lengths before anything is built. No public API, no observable behaviour change, so no docs/v5-migration.md entry.

Failing first

AShortChainDoesNotPayForTheDeferredRepresentation measures bytes per evaluation of an eleven-operand short chain as a delta of GC.GetAllocatedBytesForCurrentThread, which says what a profile says without depending on what else the machine is doing. Measured against this branch and against main with only the three runtime files reverted:

framework main this branch
net10.0 576 B 272 B −52.8%
net8.0 576 B 272 B −52.8%
net472 712 B 296 B −58.4%

The ceiling is 400 B — 35% above the largest after-figure, 30% below the smallest before-figure — so it fails on every framework against unfixed code, and neither side is close enough to the line to be fragile.

The asymptotic guarantee #3386 exists for is pinned as before by AccumulatingWithPlusIsNoLongerQuadratic (allocation ratio at 4,000 vs 8,000 iterations under 3.0, plus a 16 MB absolute ceiling — the old shape allocated 640 MB), now extended with the two five-operand chain shapes that reach ConcatMany: s = s + chunk + chunk + chunk + chunk and s = chunk + chunk + chunk + chunk + s. Both stay linear.

Also added: the mixed coerced/string chain at every arity that has its own join (3, 4, 5, 6, 11) asserted character-for-character, and the two-operand mixed pair asserted on both sides of the cutoff — flat below it, RopeString above it in both leanings.

Verification

  • dotnet build -c Release — clean, 0 warnings, all five target frameworks.
  • Jint.Tests — 11,625 passed on net10.0 and net8.0, 8,245 on net472, 0 failed.
  • Jint.Tests.PublicInterface — 3,355 / 3,345 / 2,722 passed on net10.0 / net8.0 / net472, 0 failed.
  • Jint.Tests.CommonScripts — 28/28 on net10.0 and net472. These are the SunSpider scripts themselves, date-format-xparb included, run for correctness.
  • Jint.Tests.SourceGenerators — 71/71.
  • Jint.Tests.Test262 — 102,537 passed, 0 failed, 151 skipped of 102,688. Exactly the control.

Benchmarks — not run

Nothing here was measured with BenchmarkDotNet; this machine runs agents concurrently. The paired gates belong to whoever runs them, and these are the predictions the change should be judged against:

  1. date-format-xparb, string-tagcloud, crypto-md5 (SunSpider) — should recover. These are the three rows SunSpider: date-format-xparb pays +3.9% and bitops-3bit +1.8% for the deferred-copy string representation #3527 records as having paid, and all three are built from short concatenations; xparb is the one with a measured mechanism and should move most.
  2. ObjectRegExp (Dromaeo) — must keep Strings: a long + defers its copy, so s = s + x is linear like s += x #3386's −8.9% to −11.5%. Nothing on the deferring path changed: the cutoff, the fold and RopeString are all as they were, and a chain that crosses 512 characters takes the branch it took yesterday. A regression here would mean the operand-form change moved the cutoff decision, which is the one thing it is built not to do.
  3. StringConcatLargeBenchmark's Assign* lanes — flat. They accumulate past the cutoff within their first iterations and are deferred from then on.
  4. ChainSmallThree / ChainSmallSix — should improve slightly, being exactly the eager-chain shape this touches.
  5. The += rows (AppendSmallChunks, AppendLargeChunks, BuildLargeThenScan) — must be flat. Not one line on their path changed.

The other half of #3527

bitops-3bit-bits-in-byte — the +1.81% row — does not reproduce, and needs no code. The re-profile in the analysis comment found the candidate capture had fewer total samples than base (150,057 vs 152,123) with no function moved beyond ±0.3%; the row contains no string work in its loop, and the original figure was measured with a virus scanner active. The issue's own framing allows "not reproducible" as an outcome for that half, which is why this closes the issue rather than leaving it open on that row.

Fixes #3527.

🤖 Generated with Claude Code

https://claude.ai/code/session_014W5mbjGhyvgAS4pivXoc4S

…on it never takes

sebastienros#3386 gave a long `+` a deferred copy, and SunSpider's date-format-xparb
row paid 3.9% for it. The cutoff that was supposed to keep short
concatenations out of that trade was already there and already working —
`JsString.MinDeferredConcatenationLength`, 512 characters, honoured by
both the pairwise `+` and the flattened chain — and xparb's chains are
ten to thirty-five characters, so they were on the eager side of it the
whole time. What they paid for was not the deferral. It was the shape of
the eager path.

Two things had moved onto it. The chain lane coerced every operand with
`TypeConverter.ToJsString` before the cutoff was read, which allocates a
`JsString` wrapper for every operand that is not already a string — a
number, which is most of what a formatted date is made of — and the
eager join then unwrapped every one of them again. And `ConcatMany`
built a `string[]` to hand to `string.Concat` on top of the `JsString[]`
the operands already lived in, so a chain allocated two arrays where the
pre-sebastienros#3386 lane allocated one. For xparb's fifteen-operand long-format
chain that is a second 144-byte array plus a wrapper per numeric
operand, all of it dead before the result was copied out: the doubled
variable-size allocations and the +23% allocation events the profile in
sebastienros#3527 measured under the chain node.

An operand is now carried in whichever form its coercion already
produced — the `JsString` it was, whose representation has to survive
because flattening it is the copy the deferred lane exists to avoid, or
the plain text a non-string primitive coerced to. Both answer their
length without materializing anything, which is all the cutoff is
decided on, so nothing is given up by holding the cheaper form until the
decision is made. Below the cutoff the result is then joined through a
`ValueStringBuilder` over a cutoff-sized stack buffer, which the
assembly's `SkipLocalsInit` leaves unzeroed, so there is no second array
and no wrapper. Above it the fold is unchanged.

The pairwise `+` gets the same treatment from the other end: two operands
that are already strings — a `+` between two string expressions — now
take a branch that coerces nothing at all, and the mixed pair (`n + "x"`)
coerces to text and builds the wrapper only if it goes on to defer.

`+=` and `String.prototype.concat` are untouched and exempt: both build
`JsString.ConcatenatedString`, the mutable builder, and never enter the
deferred representation at all.

Bytes per evaluation of an eleven-operand short chain, as a delta of a
thread-local allocation counter: 576 to 272 on net10.0 and net8.0, 712
to 296 on net472. `AShortChainDoesNotPayForTheDeferredRepresentation`
pins it at 400, between the two with a third of the distance on either
side; the asymptotic guarantee sebastienros#3386 exists for is pinned as before by
`AccumulatingWithPlusIsNoLongerQuadratic`, now including the five-operand
chain shapes that reach `ConcatMany`.

Fixes sebastienros#3527.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014W5mbjGhyvgAS4pivXoc4S
@lahma

lahma commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator Author

Gate verdict: PASS on all three measurements (paired, quiet box, baseline = merge-base a50731039, 95% CI must exclude zero).

SunSpider, 26 rows, 8 rounds: zero regressions. string-fasta −1.18 [−2.70, −0.02] FASTER; the three rows this PR targets all lean the right way against a main that already carries the downstream workarounds: date-format-xparb −1.88, crypto-md5 −0.68, string-tagcloud +0.05. Full table in the gate log.

Dromaeo, 24 rows, 8 rounds: 23 no-change, and one borderline flag — ObjectRegExp[True,False] +1.35 [+0.05, +3.57] — whose final rounds ran while an unrelated workload held one core. Re-measured as its own 12-round pair on a fully quiet box:

row median % 95% CI sign verdict
ObjectRegExp[False,False] +1.45 [−0.48, +3.63] 8/12 no change
ObjectRegExp[False,True] −0.72 [−2.64, +1.06] 6/12 no change
ObjectRegExp[True,False] +0.53 [−1.58, +3.38] 6/12 no change
ObjectRegExp[True,True] +0.88 [−0.50, +1.30] 7/12 no change

The flag does not reproduce under clean conditions — #3386's deferred-copy wins hold, exactly as the deferring branch being byte-identical predicts. With the allocation evidence in the PR body (−52.8%/−58.4% bytes per short-chain evaluation) and every suite at its control, this merges.

@lahma
lahma merged commit ef706d7 into sebastienros:main Sep 1, 2026
7 checks passed
lahma added a commit to lahma/jint that referenced this pull request Oct 5, 2026
…er pays for a deferred representation it never takes

Backport of PR sebastienros#3571 (commit ef706d7) from main.

sebastienros#3386 cost SunSpider's date-format-xparb 3.9% on main (sebastienros#3527), and not
because of the deferral: xparb's chains are ten to thirty-five characters,
far below the 512-character cutoff. What moved onto the eager path was its
shape. The chain lane coerced every operand with TypeConverter.ToJsString
before the cutoff was read, allocating a JsString wrapper per non-string
operand, and ConcatMany built a string[] on top of the JsString[] the
operands already lived in.

An operand is now carried in whichever form its coercion produced - the
JsString it already was, or the plain text a non-string primitive coerced
to - both of which answer their length without materializing anything, and
below the cutoff the result is joined through a ValueStringBuilder over a
cutoff-sized stack buffer: no second array and no wrapper. Above the cutoff
the fold through JsString.Concat is unchanged. The pairwise `+` gets the same
treatment: two string operands take a branch that coerces nothing, and the
mixed pair (`n + "x"`) builds the wrapper only if it goes on to defer.
TypeConverter.ToStringNonString becomes internal for that. `+=` and
String.prototype.concat are untouched.

Adapted for 4.x:

- Engine hunks (JsString.cs, JintBinaryExpression.cs, TypeConverter.cs) apply
  verbatim; ValueStringBuilder and the assembly's SkipLocalsInit are already
  on this branch.
- StringConcatenationTests: transcribed from NUnit to xUnit v3 ([Test] to
  [Fact], [TestCase] to [Theory] with [InlineData]); the allocation guard
  asks TryGetAllocatedBytesForCurrentThread, as in the previous commit.

Evidence: bytes per evaluation of the eleven-operand short chain
(y + '-' + m + '-' + d + ' ' + y + ':' + m + ':' + d), measured as a delta of
the thread-local allocation counter over 20,000 evaluations - unfixed 4.x
272 on net10.0 and 408 on net472; with this package 272 and 296, the same
after-figures main measured. So a short chain now costs no more than it did
on 4.x before sebastienros#3386, and 27% less on net472. The ported tests against
unfixed 4.x: the eleven new cases fail 3 on net10.0 and 4 on net472 -
AccumulatingWithPlusIsNoLongerQuadratic's two five-operand rows (2,562,225,160
and 2,562,134,112 B for 8,000 iterations against a 16 MB bound),
APairWithOneCoercedOperandTakesBothSidesOfTheCutoff at its "is a RopeString"
premise, and on net472 AShortChainDoesNotPayForTheDeferredRepresentation at
408 B against its 400 B ceiling. The seven
AChainMixingCoercedAndStringOperandsProducesTheSameCharacters rows pass,
pinning characters that must not change. With the package all pass on both
frameworks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanPJbBD7pQC9fRiTpHxvs
lahma added a commit to lahma/jint that referenced this pull request Oct 5, 2026
…er pays for a deferred representation it never takes

Backport of PR sebastienros#3571 (commit ef706d7) from main.

sebastienros#3386 cost SunSpider's date-format-xparb 3.9% on main (sebastienros#3527), and not
because of the deferral: xparb's chains are ten to thirty-five characters,
far below the 512-character cutoff. What moved onto the eager path was its
shape. The chain lane coerced every operand with TypeConverter.ToJsString
before the cutoff was read, allocating a JsString wrapper per non-string
operand, and ConcatMany built a string[] on top of the JsString[] the
operands already lived in.

An operand is now carried in whichever form its coercion produced - the
JsString it already was, or the plain text a non-string primitive coerced
to - both of which answer their length without materializing anything, and
below the cutoff the result is joined through a ValueStringBuilder over a
cutoff-sized stack buffer: no second array and no wrapper. Above the cutoff
the fold through JsString.Concat is unchanged. The pairwise `+` gets the same
treatment: two string operands take a branch that coerces nothing, and the
mixed pair (`n + "x"`) builds the wrapper only if it goes on to defer.
TypeConverter.ToStringNonString becomes internal for that. `+=` and
String.prototype.concat are untouched.

Adapted for 4.x:

- Engine hunks (JsString.cs, JintBinaryExpression.cs, TypeConverter.cs) apply
  verbatim; ValueStringBuilder and the assembly's SkipLocalsInit are already
  on this branch.
- StringConcatenationTests: transcribed from NUnit to xUnit v3 ([Test] to
  [Fact], [TestCase] to [Theory] with [InlineData]); the allocation guard
  asks TryGetAllocatedBytesForCurrentThread, as in the previous commit.

Evidence: bytes per evaluation of the eleven-operand short chain
(y + '-' + m + '-' + d + ' ' + y + ':' + m + ':' + d), measured as a delta of
the thread-local allocation counter over 20,000 evaluations - unfixed 4.x
272 on net10.0 and 408 on net472; with this package 272 and 296, the same
after-figures main measured. So a short chain now costs no more than it did
on 4.x before sebastienros#3386, and 27% less on net472. The ported tests against
unfixed 4.x: the eleven new cases fail 3 on net10.0 and 4 on net472 -
AccumulatingWithPlusIsNoLongerQuadratic's two five-operand rows (2,562,225,160
and 2,562,134,112 B for 8,000 iterations against a 16 MB bound),
APairWithOneCoercedOperandTakesBothSidesOfTheCutoff at its "is a RopeString"
premise, and on net472 AShortChainDoesNotPayForTheDeferredRepresentation at
408 B against its 400 B ceiling. The seven
AChainMixingCoercedAndStringOperandsProducesTheSameCharacters rows pass,
pinning characters that must not change. With the package all pass on both
frameworks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanPJbBD7pQC9fRiTpHxvs
lahma added a commit that referenced this pull request Oct 5, 2026
…uilding, safe across engines and charged to LimitMemory (#4160)

* Backport #3386 to 4.x: Strings: a long `+` defers its copy, so `s = s + x` is linear like `s += x`

Backport of PR #3386 (commit 9d80990) from main.

`s += x` and `s = s + x` mean the same thing and did not cost the same
thing. The compound form builds into `JsString.ConcatenatedString`, which is
`StringBuilder`-backed and amortised linear; a plain `+` coerced both operands
to `string` and returned `JsString.Create(string.Concat(left, right))`, so
every iteration of an accumulator loop copied the whole left operand, and
prepending (`s = x + s`) had no fast path at all.

`JsString.RopeString` is an immutable deferred form: two operands and a total
length, flattened once on the first read that needs characters, memoized, and
with both operand references released at that point. `JsString.Concat` builds
one when the result reaches 512 characters and concatenates flat below that.
`Length` - and so truthiness and the length comparison string equality
performs first - is answered from the node; everything else flattens,
`this[int]` included, since descending per character would be a new
quadratic. The flatten walk is iterative over a heap array and descends
right-first, so depth is not capped and cannot overflow the stack. The
flattened `a + b + c` chain gets the same decision, so `s = s + a + b` is
linear too, and the MaxLength guard still runs on the summed lengths before
anything is built. `+=` is untouched. Also adds the four Assign* lanes to
StringConcatLargeBenchmark.

Adapted for 4.x:

- JsString.cs, class remarks (conflict): main rewrote the subclassing
  paragraphs around LazyJsString, which is v5's #3340 and absent here. Kept
  4.x's two paragraphs and their contract; only the list of Jint's own
  representations gains the deferred one.
- JsValueExtensions.cs (conflict; the file is Jint/JsValueExtensions.cs on
  this branch): the "InternalTypes.String is set only by" list gains
  RopeString, without main's LazyJsString; 4.x's `IsString` cref is kept.
- LazyJsString.cs (absent): main's hunk there is a comment. It adds RopeString
  to the list of Jint's own lazy strings that are built on JsString directly
  rather than on LazyJsString, which is what lets LazyJsString's declared-
  length verification run with no out-of-assembly gate. The invariant behind
  it is one a node depends on: it answers its own Length as the sum of its
  operands' Lengths, sizes the flatten buffer from that sum and writes each
  operand's text at an offset computed from it, so every operand's Length has
  to be the length of its text. On main a host string is a LazyJsString whose
  declared length host-contract verification checks against what it produces.
  On 4.x the same role is played by a host's direct JsString subclass - the
  documented subclassing contract, pinned by LazyHostStringTests and by
  RavenApiUsageTests' CustomString - whose Length is its own unverified claim.
  So main's `Immutable` (snapshot a ConcatenatedString, hold everything else)
  becomes `IsRetainable`: a node holds only Jint's own immutable
  representations - an exact JsString, a SlicedString or a RopeString - and a
  ConcatenatedString or a host subclass is snapshotted at the `+`, with the
  node's length summed from what it actually holds. That keeps a host's
  ToString() at the `+`, where 4.x has always called it, and keeps a host
  whose Length disagrees with its text producing that text. The two tests
  added to Jint.Tests.PublicInterface's LazyHostStringTests pin both; run
  against main's rule verbatim, all three cases fail - the host is not
  materialized at the `+` (count 0, expected 1), a host reporting 5 for "ab"
  flattens to "\0\0\0ab...", and one reporting 3 for "abcd" throws
  ArgumentOutOfRangeException out of RopeString.CopyInto. The accumulator a
  loop builds is always Jint's own after its first deferred `+`, so the
  snapshot does not reintroduce a copy per iteration.
- README.md (conflict): main's hunk edits the "Lazy strings" section, which
  this branch's README does not have, and this README never described
  `s = s + x` as quadratic or recommended `+=` over `+`; dropped.
- docs/v5-migration.md: absent on this branch; dropped.
- StringConcatenationTests: this branch's GCPolyfills has no
  AllocatedBytesForCurrentThreadIsSupported (main's #3036); the guard asks
  TryGetAllocatedBytesForCurrentThread instead. The file was already xUnit on
  main at #3386.

Evidence - the ported tests against unfixed 4.x (engine files at upstream/4.x
plus a never-committed compile stub declaring RopeString and the cutoff
constant), net10.0 and net472 alike: in StringConcatenationTests,
StringRepresentationKeyTests and StringLengthLimitTests 32 fail and 55 pass.
Three of the failures are the asymptotic claim itself -
AccumulatingWithPlusIsNoLongerQuadratic allocates 640,758,656 B
(`s = s + chunk`), 640,744,128 B (`s = chunk + s`) and 1,281,226,960 B
(`s = s + chunk + chunk`) for 8,000 iterations against a 16 MB bound; the
other 29 stop at the "is a RopeString" premise assertion. The 55 passes pin
what must not change (characters, order, keys, the existing rows of both
older classes). AVeryDeepTreeFlattensWithoutRecursing (4x10^10 characters
copied at 200,000 iterations without the fix) and
ConcatenationPastTheLimitIsACatchableRangeError (about 1.3 GB of flat
strings without the fix) were not run against unfixed code on this shared
machine. With the package all of them pass on both frameworks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanPJbBD7pQC9fRiTpHxvs

* Backport #3571 to 4.x: Strings: a short `+` chain no longer pays for a deferred representation it never takes

Backport of PR #3571 (commit ef706d7) from main.

#3386 cost SunSpider's date-format-xparb 3.9% on main (#3527), and not
because of the deferral: xparb's chains are ten to thirty-five characters,
far below the 512-character cutoff. What moved onto the eager path was its
shape. The chain lane coerced every operand with TypeConverter.ToJsString
before the cutoff was read, allocating a JsString wrapper per non-string
operand, and ConcatMany built a string[] on top of the JsString[] the
operands already lived in.

An operand is now carried in whichever form its coercion produced - the
JsString it already was, or the plain text a non-string primitive coerced
to - both of which answer their length without materializing anything, and
below the cutoff the result is joined through a ValueStringBuilder over a
cutoff-sized stack buffer: no second array and no wrapper. Above the cutoff
the fold through JsString.Concat is unchanged. The pairwise `+` gets the same
treatment: two string operands take a branch that coerces nothing, and the
mixed pair (`n + "x"`) builds the wrapper only if it goes on to defer.
TypeConverter.ToStringNonString becomes internal for that. `+=` and
String.prototype.concat are untouched.

Adapted for 4.x:

- Engine hunks (JsString.cs, JintBinaryExpression.cs, TypeConverter.cs) apply
  verbatim; ValueStringBuilder and the assembly's SkipLocalsInit are already
  on this branch.
- StringConcatenationTests: transcribed from NUnit to xUnit v3 ([Test] to
  [Fact], [TestCase] to [Theory] with [InlineData]); the allocation guard
  asks TryGetAllocatedBytesForCurrentThread, as in the previous commit.

Evidence: bytes per evaluation of the eleven-operand short chain
(y + '-' + m + '-' + d + ' ' + y + ':' + m + ':' + d), measured as a delta of
the thread-local allocation counter over 20,000 evaluations - unfixed 4.x
272 on net10.0 and 408 on net472; with this package 272 and 296, the same
after-figures main measured. So a short chain now costs no more than it did
on 4.x before #3386, and 27% less on net472. The ported tests against
unfixed 4.x: the eleven new cases fail 3 on net10.0 and 4 on net472 -
AccumulatingWithPlusIsNoLongerQuadratic's two five-operand rows (2,562,225,160
and 2,562,134,112 B for 8,000 iterations against a 16 MB bound),
APairWithOneCoercedOperandTakesBothSidesOfTheCutoff at its "is a RopeString"
premise, and on net472 AShortChainDoesNotPayForTheDeferredRepresentation at
408 B against its 400 B ceiling. The seven
AChainMixingCoercedAndStringOperandsProducesTheSameCharacters rows pass,
pinning characters that must not change. With the package all pass on both
frameworks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanPJbBD7pQC9fRiTpHxvs

* Backport #4130 to 4.x: Build a bound function's name as a rope so a chain of binds stays linear

Backport of PR #4130 (commit 5eac242) from main.

SetFunctionName composed step 4's "prefix name" with a CLR string
concatenation, which flattens whatever it is handed. Binding an already-bound
function therefore materialized "bound bound ... f" afresh at every level and
left that copy in the level's own name descriptor, so N binds retained about
3*N^2 characters (#4129). The prefixed name is now a JsString.Concat node over
the level below - O(1) to build, flattened once if something ever reads the
text - and every observable name is byte-identical.

Adapted for 4.x:

- On this branch BindFunction derives from ObjectInstance, not Function (main
  made it a Function in #3658, which is not here), so Function.prototype.bind
  names the function it creates through ObjectInstance.SetFunctionName, which
  main's hunk does not touch and which still did `prefix + " " + name`. With
  main's Function.cs hunk alone the bind chain stayed exactly as quadratic
  (measured: 9.5, 36.9, 145.3, 576.7 MB allocated for 1,250, 2,500, 5,000 and
  10,000 levels). Function.PrefixName becomes internal rather than private and
  ObjectInstance.SetFunctionName calls it, so both overloads defer. Main's
  Function.cs hunk applies verbatim otherwise; it still covers the get/set
  accessor names and ShadowRealm's wrapped functions, which are Functions.
- BoundFunctionNameTests: transcribed from NUnit to xUnit v3. The wedge
  ceiling is 256 MB instead of main's 768 MB: with the fix the whole script
  allocates 53.0 MB on net10.0 and 55.3 MB on net472, and against the defect
  the per-level copies are all retained, so main's ceiling would let a run
  against unfixed code hold about 770 MB of live text before it tripped.
  256 MB trips at level 6,649 (net10.0) / 6,648 (net472) with a peak working
  set of 291 / 284 MB. Added ABoundNameLongEnoughToDeferIsANodeOverTheLevelBelow,
  which asserts the premise directly - a 100-level bound name is a RopeString -
  because on this branch the route bind takes is the one main's hunk misses.

Evidence - the bind chain, measured as allocated and retained bytes around
`for (i < N) f = f.bind(null)` on net10.0:

  depth     unfixed 4.x allocated / retained    with this package
  1,250       9.5 MB /   9.4 MB                   0.6 MB / 0.5 MB
  2,500      36.9 MB /  36.7 MB                   1.1 MB / 1.0 MB
  5,000     145.3 MB / 145.0 MB                   2.2 MB / 1.9 MB
  100,000   (not run: about 60 GB)               45.5 MB / 36.7 MB, 98 ms

Unfixed quadruples per doubling; fixed doubles. net472 unfixed: 9.7, 37.1,
145.6 MB for the same three depths. The ported tests against unfixed 4.x fail
2 of 8 on both frameworks - ADeepChainOfBoundFunctionsDoesNotCopyTheNameAtEveryLevel
with MemoryLimitExceededException at 256 MB, and the new premise test at "is a
RopeString" - and the six APrefixedNameIsTheSameTextItAlwaysWas rows pass,
pinning the names that must not change. With the package all 8 pass on both.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanPJbBD7pQC9fRiTpHxvs

* Backport #4164 to 4.x: Fall back to a rope's published text when a concurrent flatten has released its operands

RopeString.Flatten memoized its text and then released both operands with
plain writes, while CopyInto tested the memo and then dereferenced the
operands. A second thread finishing its own flatten of the same node between
those two reads left the walk holding a null operand, and the next
ToString() threw NullReferenceException. On 4.x, too, preparation folds a
literal-plus-literal into one constant on the shared Prepared<Script>, so
every engine running a preparation reads the same RopeString.

The engine change applies as is. The tests are transcribed from NUnit to
xUnit v3: the two tests are async and wait on a local two-minute wedge
ceiling through Task.WhenAny (4.x has no TestBudgets, and xUnit1031 rejects
a blocking Task.WaitAll).

(cherry picked from commit 78e9bb0)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Backport #4173 to 4.x: Charge a deferred string concatenation to LimitMemory when it is built

A long `a + b` returns a RopeString, whose characters are allocated by whoever
flattens it, and that can be a host's ToString() after every constraint is
disarmed. JsString.Concat now charges the memory limit on its deferring
branches: the shorter operand, which keeps `s = s + x` linear, or the whole
length, refused before the node exists, when the flat form alone exceeds the
budget. A constant fold evaluates with an engine-less context and is not
charged.

Adapted for 4.x's MemoryLimitConstraint, which predates main's operation
state and async segments: it measures the thread's allocation counter
against a per-entry baseline. The charge is kept in a _deferredBytes total
that Reset() clears with the baseline and Check() adds to the measured
allocation. A result whose flat form alone exceeds the budget calls Check()
and then throws regardless, because Check() measures nothing on a thread
other than the one the entry started on. The engine finds its constraint
once at construction (Engine._memoryLimitConstraint). 4.x's ConcatSnapshotting
branch, which main does not have, is charged too, for the lengths the node
actually holds. Function.PrefixName is called from ObjectInstance.SetFunctionName
on 4.x as well, and both pass the engine's evaluation context.

Not ported: main's AllocatedBytes doc line and the
TheCharactersADeferredConcatenationAppendsAreReportedAsAllocated test
(MemoryLimitConstraint.AllocatedBytes is not public API on 4.x), the
docs/guide and migration-guide hunks (absent on 4.x; the README's
constraints section gets the paragraph instead) and the Jint/Constraints/AGENTS.md
hunk (absent on 4.x). Tests transcribed from NUnit to xUnit v3.

Adapted from 7d9cdcc

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SunSpider: date-format-xparb pays +3.9% and bitops-3bit +1.8% for the deferred-copy string representation

1 participant