Consolidate the ANSI escape grammar across stringWidth, stripANSI, sliceAnsi and wrapAnsi - #33488
Conversation
|
Warning Review limit reached
Next review available in: 19 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (1)
WalkthroughChangesThe PR regenerates Unicode width tables, expands ANSI parsing for C1/UTF-8/control-string forms, propagates Latin-1 ambiguous-width handling, refactors slicing and wrapping state management, and adds regression coverage for width, escape, SGR, hyperlink, trimming, and ellipsis behavior. ChangesString width and ANSI behavior
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Updated 8:05 PM PT - Jul 16th, 2026
@Jarred-Sumner, your commit 62b75ff is building: |
|
Found 1 issue this PR may fix:
🤖 Generated with Claude Code |
Checked, and this PR doesn't fix it — #8329 is stale. It was filed about a ( |
|
Benchmarked the Latin-1 kernel I'd replaced, and the first cut was 2.5-3.6x slower on escape-heavy input, which is exactly what the single-pass kernel exists to prevent. Pushed a second commit that keeps it. What changed: CSI and OSC still come straight off the chunk bitmasks. Only the forms the kernel didn't know about (two-byte, nF, ST-terminated control strings, an ESC that re-introduces a sequence) go through a scalar mirror of the recognizer. The escape block moved out of line — inlining it into the chunk loop costs the no-escape path ~40% of its throughput. Along the way the fuzz cross-check against the scalar reference caught a real bug:
|
| case | base | this PR | |
|---|---|---|---|
| no-ANSI 16KB latin1 | 0.336 us | 0.352 us | same |
| no-ANSI 16KB latin1 (é) | 0.336 us | 0.353 us | same |
| no-ANSI utf16 ascii 16K | 0.326 us | 0.347 us | same |
| no-ANSI utf16 CJK 8K | 1.288 us | 1.289 us | same |
| short line (typical TUI) | 0.055 us | 0.058 us | same |
| dense SGR 16KB | 8.839 us | 10.168 us | 15% slower |
| truecolor SGR 16KB | 5.902 us | 7.429 us | 26% slower |
| OSC-8 hyperlinks 16KB | 3.509 us | 4.474 us | 28% slower |
| utf16 CJK + SGR | 40.278 us | 32.905 us | 18% faster |
Every realistic call shape is at parity. The escape-dense 16KB rows pay one out-of-line call per chunk that contains an escape; the UTF-16 path gets faster because the old per-code-unit escape state machine is gone.
benchmark script
const latin1 = (s, n) => Buffer.alloc(s.length * n, s, "latin1").toString("latin1");
const utf16 = (s, n) => { let o = ""; for (let i = 0; i < n; i++) o += s; return ("\u{1F600}" + o).slice(2); };
const cases = [
["no-ANSI 16KB latin1", latin1("a", 16384), 2000],
["dense SGR (bash prompt)", latin1("\x1b[31mword\x1b[0m \x1b[32mword\x1b[0m \x1b[33mword\x1b[0m", 400), 2000],
["truecolor SGR", latin1("\x1b[38;2;255;128;64mhello\x1b[39m", 800), 2000],
["OSC-8 hyperlinks", latin1("\x1b]8;;https://example.com/xyz\x07click here\x1b]8;;\x07 ", 400), 2000],
["short line (typical TUI)", "\x1b[1;32mok\x1b[0m \x1b[2mpackages installed\x1b[0m", 200000],
["utf16 CJK + SGR", utf16("\x1b[31m安宁\x1b[0m hello ", 500), 2000],
];
for (const [name, s, iters] of cases) {
for (let i = 0; i < 500; i++) Bun.stringWidth(s);
let best = Infinity;
for (let r = 0; r < 12; r++) {
const t0 = Bun.nanoseconds();
for (let i = 0; i < iters; i++) Bun.stringWidth(s);
best = Math.min(best, (Bun.nanoseconds() - t0) / iters);
}
console.log(name.padEnd(26), (best / 1000).toFixed(3) + " us");
}2a5a5ce to
7ed5a6d
Compare
|
CI caught a real one: The escape block is pure scalar bit juggling, but I'd marked it Rather than allowlist five near-identical functions that have no business containing vector instructions, it moves above Also rebased onto Corrected throughputSuperseding my earlier table — that measured the shape CI rejected. Same protocol (both from-scratch release builds, min of 6 runs × 12 batches):
Strings with no escapes — the overwhelmingly common shape — are at parity, and the UTF-16 path gets faster because the old per-code-unit escape state machine is gone. The escape-containing rows pay one out-of-line call per chunk that holds an escape; in absolute terms that's +12ns on a typical TUI line. That call is the price of keeping the block off the SIMD targets. If you'd rather have it back inline and allowlist the five symbols, say the word and I'll do that instead — it's a small change either way. Caveat on the numbers: the box was at load average 60+, so treat sub-10% deltas as noise. |
|
CI status on The build is marked failed only because two The previous build's failures were the same class (a The |
2bf0add to
0bbc39a
Compare
There was a problem hiding this comment.
Actionable comments posted: 6
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/generate-stringwidth-tables.mjs`:
- Around line 78-79: Update the Unicode range generation in
scripts/generate-stringwidth-tables.mjs so isZeroWidth() includes all
General_Category Cf (format control) code points, including the listed ranges,
alongside the existing nonspacing and enclosing mark handling. Ensure the
generated zero-width table and lookup path consume this Cf range without
removing current Mn/Me coverage.
In `@src/jsc/bindings/sliceAnsi.cpp`:
- Around line 490-539: Update parseEscapeSequence to recognize CAN (0x18), SUB
(0x1A), and C1 ST (0x9C) immediately after a bare ESC as abort-to-ground
sequences. Consume these bytes together with the ESC, matching
ANSI::consumeANSI() behavior and preventing the literal ESC from reaching
sliceAnsi output.
In `@src/jsc/bindings/wrapAnsi.cpp`:
- Around line 230-250: Update wrapWord in src/jsc/bindings/wrapAnsi.cpp at lines
230-250 to iterate over complete grapheme clusters, keeping combining marks and
ZWJ emoji together, and only create a new row when the current row contains
visible content so an oversized first cluster does not produce an empty row.
Update the expectations in test/js/bun/util/wrapAnsi.test.ts at lines 103-107 to
require family emoji remain indivisible and add coverage for a base character
with a combining mark.
In `@test/js/bun/util/sliceAnsi-fuzz.test.ts`:
- Line 5: Update the workload-scaling condition in the sliceAnsi fuzz test to
use isDebug || isASAN, ensuring both debug and ASAN lanes select the reduced
iteration count while preserving the existing release workload.
In `@test/js/bun/util/wrapAnsi.test.ts`:
- Around line 57-59: Align the public contract for Bun.wrapAnsi with the
character-by-character behavior established by the wordWrap false test: update
the wordWrap false documentation in the wrapAnsi documentation to state that
text also wraps at the configured width, not only at explicit newlines. Preserve
the existing behavior and clarify the contract rather than changing the
implementation.
- Line 232: Add a hard-wrap test case in the wrapAnsi coverage alongside the
existing C1 CSI cases using a non-SGR final byte such as K (for example,
"\x9b2Kabcdef"). Assert that the complete zero-width CSI sequence remains intact
and K is not split during wrapping, covering the required 0x40–0x7E final-byte
behavior.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: a04d0e30-286c-4543-9b91-2fc8478b898d
📒 Files selected for processing (17)
docs/runtime/utils.mdxscripts/generate-stringwidth-tables.mjssrc/bun_core/string/immutable/visible.rssrc/bun_core/string/mod.rssrc/jsc/bindings/ANSIHelpers.hsrc/jsc/bindings/highway_strings.cppsrc/jsc/bindings/sliceAnsi.cppsrc/jsc/bindings/stringWidth.cppsrc/jsc/bindings/stringWidth.hsrc/jsc/bindings/stringWidthTables.hsrc/jsc/bindings/stripANSI.cppsrc/jsc/bindings/wrapAnsi.cpptest/js/bun/util/sliceAnsi-fuzz.test.tstest/js/bun/util/sliceAnsi.test.tstest/js/bun/util/stringWidth.test.tstest/js/bun/util/stripANSI.test.tstest/js/bun/util/wrapAnsi.test.ts
Bun.stringWidth, Bun.stripANSI, Bun.sliceAnsi and Bun.wrapAnsi each carried their own escape recognizer and disagreed on what a terminal displays. Drive all four through the shared grammar in ANSIHelpers.h: CSI (incl. 0x9B), OSC (incl. 0x9D), the DCS/SOS/PM/APC control strings (incl. 0x90/0x98/0x9E/0x9F), nF and Fe/Fs escapes, with VT500 abort semantics — an ESC, CAN, SUB or C1 ST inside a sequence ends it instead of swallowing the following text. Width classification is regenerated from Unicode 17.0 (nonspacing marks zero-width across all scripts, conjoining Hangul jamo, new emoji), the Latin-1 path honors ambiguousIsNarrow: false, and a variation selector no longer widens a zero-width or non-emoji base. sliceAnsi tracks SGR state per attribute slot (bold+dim, single+double underline stack instead of evicting), parses SGR 58/59 underline color as an extended color, and treats an empty parameter as 0. wrapAnsi tokenizes with the shared recognizer, keeps OSC-8 hyperlinks intact across row breaks (ST terminator and id= params included), does not break on a bare CR, and stops coercing falsy non-false options. The UTF-16 width accumulator scans each visible run only up to the next escape, removing an O(escapes x remainder) rescan on multi-line colored text.
Under an ellipsis, keep escape sequences ordered with the speculative-zone content by writing the zone in place and restoring an SGR/hyperlink snapshot when it is discarded, so a fitting styled string is not reordered and a discarded zone drops its trailing ANSI with it. Detect a cut when the final wide cluster overflows the end budget at EOF, stop treating trailing zero-width clusters as an end cut, and add the degenerate fallback the fast path already had: a range that admits no visible content returns just the ellipsis. The ellipsis inherits the active style, consistent with the known-cut and negative-index paths.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/jsc/bindings/sliceAnsi.cpp (1)
289-302: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winDo not track overflowed SGR parameters using only the first code.
Line 291 converts overflow into the colon path, which records the entire sequence as the style represented by
params[0]. For a valid long sequence such as31;0;...m, this incorrectly leaves red active even though a later reset cleared it. Return without updating state whenparams.overflowis set.Proposed fix
- if (params.overflow) hasColon = true; + if (params.overflow) + return;🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/jsc/bindings/sliceAnsi.cpp` around lines 289 - 302, In the SGR parsing flow, update the params.overflow handling before the hasColon branch to return immediately without calling state.applyStart or otherwise updating tracked style state. Remove the conversion of overflow to hasColon, while preserving existing handling for non-overflowed colon sequences.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@src/jsc/bindings/sliceAnsi.cpp`:
- Around line 289-302: In the SGR parsing flow, update the params.overflow
handling before the hasColon branch to return immediately without calling
state.applyStart or otherwise updating tracked style state. Remove the
conversion of overflow to hasColon, while preserving existing handling for
non-overflowed colon sequences.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 503f1266-493b-4c5c-9ccd-e4c0de0aa396
📒 Files selected for processing (2)
src/jsc/bindings/sliceAnsi.cpptest/js/bun/util/sliceAnsi.test.ts
- sliceAnsi-fuzz: scale the O(n) time bound on ASAN too, not only debug (the release ASAN lane took the 1000-iteration path and hit the 5s bound). - stripANSI: the heapStats string count can drop between the two reads when GC collects an unrelated string, so the new "standalone C1 ST is not stripped" check is one-sided (<=) with a full GC before the baseline. - docs/bun.d.ts: `wordWrap: false` breaks every line at the column width, not "only at explicit newlines" -- that is what Bun and npm wrap-ansi both do and what the new test asserts. - wrapAnsi.test: cover C1 CSI hard wrap with a non-SGR final byte (`K`). No src/ changes. Rebased onto main for #32498 / #34418 / #34423.
87158c2 to
2ca6ccb
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/jsc/bindings/wrapAnsi.cpp (1)
577-613: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winUse seam-aware width before deciding to wrap.
Lines 592–603 compare the additive
rowLength + wordLenbefore checking whether the separator and word fuse into one grapheme. For example,"aa \u20E3bb"has visible width 6, but at 6 columns this code calculates 7 and wraps it unnecessarily. Compute the projected combined width before the overflow branches, not only after appending.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/jsc/bindings/wrapAnsi.cpp` around lines 577 - 613, Update the wrapping logic around the overflow checks in the visible word-processing routine to compute the projected seam-aware width using the same separator/word fusion rules as the final lastRowWidth update. Use that projected width, rather than rowLength + wordLen, for both overflow branches so fused graphemes such as a space followed by a combining mark do not wrap unnecessarily; preserve the existing append and lastRowWidthDirty behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/js/bun/util/stringWidth.test.ts`:
- Around line 1103-1129: Update the alphabet used by the randomized consistency
invariant to include the C1 control bytes for CSI, OSC, DCS, SOS, PM, and APC,
and revise the adjacent comment to state that these introducers are covered.
Keep the existing ESC and printable/control character coverage unchanged so the
cross-check exercises both grammars through consume-to-EOF behavior.
In `@test/js/bun/util/stripANSI.test.ts`:
- Around line 546-553: Update this zero-copy test to run garbage collection
before capturing the baseline string count, then require the post-call count to
equal the baseline instead of allowing a decrease. Preserve the existing result
identity assertion and follow the exact heap-count assertion pattern used by the
nearby large-input test.
---
Outside diff comments:
In `@src/jsc/bindings/wrapAnsi.cpp`:
- Around line 577-613: Update the wrapping logic around the overflow checks in
the visible word-processing routine to compute the projected seam-aware width
using the same separator/word fusion rules as the final lastRowWidth update. Use
that projected width, rather than rowLength + wordLen, for both overflow
branches so fused graphemes such as a space followed by a combining mark do not
wrap unnecessarily; preserve the existing append and lastRowWidthDirty behavior.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 159ada9f-cd58-4abf-80bf-4411c726ec44
📒 Files selected for processing (18)
docs/runtime/utils.mdxpackages/bun-types/bun.d.tsscripts/generate-stringwidth-tables.mjssrc/bun_core/string/immutable/visible.rssrc/bun_core/string/mod.rssrc/jsc/bindings/ANSIHelpers.hsrc/jsc/bindings/highway_strings.cppsrc/jsc/bindings/sliceAnsi.cppsrc/jsc/bindings/stringWidth.cppsrc/jsc/bindings/stringWidth.hsrc/jsc/bindings/stringWidthTables.hsrc/jsc/bindings/stripANSI.cppsrc/jsc/bindings/wrapAnsi.cpptest/js/bun/util/sliceAnsi-fuzz.test.tstest/js/bun/util/sliceAnsi.test.tstest/js/bun/util/stringWidth.test.tstest/js/bun/util/stripANSI.test.tstest/js/bun/util/wrapAnsi.test.ts
👮 Files not reviewed due to content moderation or server errors (10)
- scripts/generate-stringwidth-tables.mjs
- src/jsc/bindings/stringWidthTables.h
- src/jsc/bindings/stringWidth.h
- src/jsc/bindings/stringWidth.cpp
- src/bun_core/string/immutable/visible.rs
- src/bun_core/string/mod.rs
- src/jsc/bindings/ANSIHelpers.h
- src/jsc/bindings/highway_strings.cpp
- src/jsc/bindings/sliceAnsi.cpp
- src/jsc/bindings/stripANSI.cpp
The comment was stale: the consolidated recognizer treats C1 CSI/OSC/DCS/ SOS/PM/APC/ST the same across stringWidth and stripANSI now, so the randomized consistency invariant can cover them.
|
Re the outside-diff comment on The space + U+20E3 fuse into one grapheme cluster, so "aa " + "⃣bb" is only 6 columns, but the additive check sees 3 + 4 = 7 and wraps. Leaving for the next |
Restore origin/main's normalization of a lone carriage return into a line break (\n, bare \r and a \r\n pair each produce one break). Claude Code's wrap-text maps wrapped output back onto the original string by relying on every \r becoming a break; leaving bare \r as ordinary content shifted that mapping by one per line.
…-width clusters (#34482) Follow-up to #33488 (which fixed #2515 for visible content). The start-side degenerate fallback still missed ranges whose plain-slice content is entirely zero-width. ## Repro ```js Bun.sliceAnsi("a\x1b[31m\t", 1, 3, {ellipsis: "…"}); // "" (expected "…") Bun.sliceAnsi("a\x1b[31m\t", 1, 3); // "\x1b[31m\t\x1b[39m" (plain slice keeps the tab) Bun.sliceAnsi("ЖЗИ", 1, 4, {ellipsis: ">>"}); // ">>" (already fixed by #33488) ``` ## Cause The bare-ellipsis fallback at `walkDone` tests `position > startBeforeBudget` to decide whether the requested range held content. `position` is the column *after* the last cluster, so a zero-width cluster (tab/LF/ZWSP) sitting exactly at `startBeforeBudget` leaves `position == startBeforeBudget` and the arm evaluates false. The start-ellipsis budget had pushed `start` past that cluster, so `include` never flipped, and the next line returns `emptyString()`. ## Fix Capture `position` (the last cluster's start column) before the EOF width finalize, and extend the start-cut arm with `!include && hasPrev && lastClusterCol >= startBeforeBudget`. This catches "only zero-width clusters reached the original start" without touching the hot walk loop. The `!include` guard scopes the change to the reported case: ranges where the start budget consumed the whole range. ## Verification `bun bd test test/js/bun/util/sliceAnsi.test.ts` → 166 pass; `sliceAnsi-fuzz.test.ts` → 48 pass. Fail-before verified by stashing `src/`. <!-- robobun:evidence:begin --> --- **[stamp-90s]** gate passed · iteration 1 · 2 files touched <details><summary>fails on main (without fix)</summary> ```console ASAN without fix: 1 FAILED $ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/bun/util/sliceAnsi.test.ts info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: checking for self-update (current version: 1.29.0) bun test v1.4.0 (1ed30a6) test/js/bun/util/sliceAnsi.test.ts: (pass) Bun.sliceAnsi > plain strings > slices ASCII string like String.prototype.slice [3.13ms] (pass) Bun.sliceAnsi > plain strings > returns empty string for empty input [2.02ms] (pass) Bun.sliceAnsi > plain strings > returns full string with no arguments beyond first [1.62ms] (pass) Bun.sliceAnsi > plain strings > start=0, end=0 returns empty [1.49ms] (pass) Bun.sliceAnsi > plain strings > start > end returns empty [1.33ms] (pass) Bun.sliceAnsi > plain strings > start beyond string length returns empty [1.50ms] (pass) Bun.sliceAnsi > plain strings > end beyond string length returns remainder [1.67ms] (pass) Bun.sliceAnsi > plain strings > negative start [2.36ms] (pass) Bun.sliceAns ... (truncated) release without fix: all passed bun test v1.4.0-canary.1 (75dce26) test/js/bun/util/sliceAnsi.test.ts: (pass) Bun.sliceAnsi > plain strings > slices ASCII string like String.prototype.slice [0.04ms] (pass) Bun.sliceAnsi > plain strings > returns empty string for empty input [0.01ms] (pass) Bun.sliceAnsi > plain strings > returns full string with no arguments beyond first [0.01ms] (pass) Bun.sliceAnsi > plain strings > start=0, end=0 returns empty (pass) Bun.sliceAnsi > plain strings > start > end returns empty [0.01ms] (pass) Bun.sliceAnsi > plain strings > start beyond string length returns empty [0.03ms] (pass) Bun.sliceAnsi > plain strings > end beyond string length returns remainder [0.01ms] (pass) Bun.sliceAnsi > plain strings > negative start [0.02ms] (pass) Bun.sliceAnsi > plain strings > negative end [0.03ms] (pass) Bun.sliceAnsi > plain strings > both negative (pass) Bun.sliceAnsi > plain strings > negative start keeps trailing zero-width and ANSI content [0.17ms] (pass) Bun.sliceAnsi > plain strings > negative end strictly before total width still cuts [0.02ms] (pass) Bun.sliceAnsi > plain strings > single character slice [0.02ms] (pass) Bun.sliceAnsi > ANSI color codes > slices color ... (truncated) ``` </details> <details><summary>passes on PR (with fix)</summary> ```console ASAN with fix: all passed $ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/bun/util/sliceAnsi.test.ts info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: checking for self-update (current version: 1.29.0) bun test v1.4.0 (1ed30a6) test/js/bun/util/sliceAnsi.test.ts: (pass) Bun.sliceAnsi > plain strings > slices ASCII string like String.prototype.slice [3.21ms] (pass) Bun.sliceAnsi > plain strings > returns empty string for empty input [2.00ms] (pass) Bun.sliceAnsi > plain strings > returns full string with no arguments beyond first [1.61ms] (pass) Bun.sliceAnsi > plain strings > start=0, end=0 returns empty [1.48ms] (pass) Bun.sliceAnsi > plain strings > start > end returns empty [1.29ms] (pass) Bun.sliceAnsi > plain strings > start beyond string length returns empty [1.57ms] (pass) Bun.sliceAnsi > plain strings > end beyond string length returns remainder [1.72ms] (pass) Bun.sliceAnsi > plain strings > negative start [2.35ms] (pass) Bun.sliceAns ... (truncated) release with fix: all passed $ bun scripts/build.ts --profile=release info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: checking for self-update (current version: 1.29.0) [configured] bun-profile → bun (stripped) in 681ms (unchanged) ninja: Entering directory `/workspace/bun/build/release' [1/7] cxx obj/unified/UnifiedSource-src_jsc_bindings-5.cpp.o [2/7] gen cpp.rs (cppbind) [2/7] cargo bun_bin → libbun_rust.a (--target x86_64-unknown-linux-gnu) info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: component rust-std is up to date nightly-2026-05-06-x86_64-unknown-linux-gnu unchanged - rustc 1.97.0-nightly (e95e73209 2026-05-05) info: checking for self-update (current version: 1.29.0) �[1m�[92m Blocking�[0m waiting for file lock on build directory �[1m�[92m Compiling�[0m bun_core v0.0.0 (/workspace/bun/src/bun_core) �[1m�[92m Compiling�[0m bun_errno v0.0.0 (/workspace/bun/sr ... (truncated) ``` </details> <details><summary>diff hotspot</summary> ``` src/jsc/bindings/sliceAnsi.cpp | 8 ++++++-- test/js/bun/util/sliceAnsi.test.ts | 24 ++++++++++++++++++++++++ 2 files changed, 30 insertions(+), 2 deletions(-) ``` </details> **gate history** · 3 passed · 0 rejected · iteration 1 <details><summary>evidence per changed file</summary> ``` file reads edits tests src/jsc/bindings/sliceAnsi.cpp 6 3 0 test/js/bun/util/sliceAnsi.test.ts 3 3 0 ``` </details> <!-- robobun:evidence:end --> --------- Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
Bun.stringWidth,Bun.stripANSI,Bun.sliceAnsiandBun.wrapAnsieach carried their own escape recognizer, and the four disagreed on what a terminal would display. This consolidates them onto the shared grammar inANSIHelpers.hand fixes the width / SGR / wrapping bugs that a terminal UI hits.Closes #34384, closes #34381, closes #34379, closes #34380, closes #34377, closes #34382, closes #34391, closes #34393, closes #34394.
Escape grammar (all four APIs)
ESC […<final 0x40–0x7E>0x9B…31mas textESC ]/ C10x9D… (BEL /ESC \/0x9C)0x9DunrecognizedESC P/X/^/_, C10x90/0x98/0x9E/0x9F… STESC <0x20–2F> <byte>(ESC ( B), Fe/FsESC <0x30–7E>(ESC 7,ESC c)(B0,7hi8visiblestripANSI("\x1b]0;title\x1b[31mtext")→"")"text"stringWidth(s) === stringWidth(stripANSI(s))and agreement withsliceAnsi/wrapAnsinow hold on all of the above, on Latin-1, UTF-16 and (viainspect.table/ the markdown renderer) UTF-8 strings — previously a JS string could measure differently before and after a wide character forced it to UTF-16.Repro (release, before → after)
Fix
ANSIHelpers.hconsumeANSI()Utf8mode where C1 =0xC2 0x9xand ST =0xC2 0x9C(a bare0x9Cis a continuation byte);sgrCloseCode/isSgrEndCodegain 20/21/51/52/58/73/74stringWidth.cppconsumeANSI; Latin-1ambiguousIsNarrow: falsepath; VS16 widens only Emoji-property baseshighway_strings.cppLatin-1 kernelorigin/main's single-pass loop kept; the fast-path gate also fails on0x7F–0x9F; a chunk containing any non-CSI/OSC introducer exits to the scalar recognizer (a call inside the loop spilled the caller-saved vector registers onto the fast path)stringWidthTables.hsliceAnsi.cppSGR statesliceAnsi.cppellipsiswrapAnsi.cppconsumeANSI(no more "inside escape until anm"), OSC-8 body never split (ST terminator +id=params, re-opened per row),trim/wordWraphonor!== falseonly, zero-width tail chars kept when trimming; a bare\rstill breaks a line as before (wrap-textmaps wrapped output back onto the source by relying on it)stripANSI.cpp0x9Ckeep zero-copy identityPerformance
Agent shaped input (the workload these exist for: 21 real TUI lines, 3–139 code units, per-call,
origin/mainvs this branch, local release, min of 3):stringWidth(s)sliceAnsi(s, 0, 40)wrapAnsi(s, 80)wrapAnsi(s, 40, {hard:true})stripANSI(s)stringWidth(s, {countAnsiEscapeCodes:true})Long single strings (4 KB–64 KB per call), where the SIMD kernels dominate:
The long-string escape shapes are 15–30 % behind
origin/main: the per-chunk cold path re-classifies the C1 range and the grammar-complete terminator sets. Growth is linear in every case (per-line cost on a 1600-line colored UTF-16 log: 0.024 µs/line, same as main). Short strings, which is nearly all real callers, are faster thanmain.Verification
[review] gate passed · iteration 5 · 18 files touched
fails on main (without fix)
passes on PR (with fix)
diff hotspot
gate history · 3 passed · 1 rejected · iteration 5
evidence per changed file