Conversation
Bun.stringWidth(s, {ambiguousIsNarrow: false}) returned different widths
for identical text depending on JSC's internal string encoding: "ééééé"
measured 5 when Latin-1-backed and 10 when UTF-16-backed. 43 bytes in
U+00A1..U+00FE are East Asian Ambiguous (§, °, ±, ×, ÷, é, ...) but the
Latin-1 width functions had no ambiguousAsWide parameter and every caller
(the stringWidth 8-bit path, wrapAnsi's stringWidth<Latin1>, sliceAnsi's
ellipsis sizing) silently dropped the option.
Thread the flag through visibleLatin1Width / visibleLatin1WidthExcludeANSI
and the Bun__visibleWidthExcludeANSI_latin1 C export. The per-byte wide
widths are derived at compile time from the same fused classification
table the UTF-16 path uses, with a static_assert that the narrow widths
agree with the existing scalar, so the two encodings cannot drift apart.
The default (ambiguousIsNarrow: true) keeps the SIMD fast path unchanged;
the opt-in wide mode adds a per-byte correction, falling back to a scalar
ANSI state machine only when escapes are present so OSC payload bytes are
not over-counted.
|
Updated 11:20 AM PT - Jul 16th, 2026
❌ @autofix-ci[bot], your commit 655db2c has 4 failures in
🧪 To try this PR locally: bunx bun-pr 34381That installs a local version of the PR into your bun-34381 --bun |
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 56 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (7)
Comment |
|
Found 1 issue this PR may fix:
🤖 Generated with Claude Code |
There was a problem hiding this comment.
I didn't find any bugs, but this touches a hot native path with an FFI signature change and a duplicated ANSI state machine, so it's worth a human look.
Checked: all callers of Bun__visibleWidthExcludeANSI_latin1 / visibleLatin1Width* updated for the new parameter (Rust FFI, wrapAnsi, sliceAnsi, stringWidth, UTF-8 path).
Checked: visibleLatin1WidthExcludeANSIAmbiguousWide matches VisibleLatin1WidthExcludeANSIScalar in highway_strings.cpp byte-for-byte, only the width lookup differs.
Checked: the narrow + (wide>>1) correction is sound — the static_assert proves narrow Latin-1 width ∈ {0,1}, so width-2 in the wide table implies ambiguous.
Checked: the no-ESC fast path is safe — the SIMD kernel only recognizes 7-bit ESC introducers, and all ambiguous Latin-1 bytes are ≥ 0xA1.
Extended reasoning...
Overview
Threads ambiguousAsWide through the Latin-1 width fast paths so Bun.stringWidth returns the same result for identical codepoints regardless of JSC's internal 8-bit/16-bit representation. Touches stringWidth.{h,cpp} (~120 lines added: a constexpr 256-byte table with static_assert guards, a countLatin1Ambiguous helper, and a scalar ANSI state machine for the wide+exclude-ANSI case), plus one-line parameter-threading in sliceAnsi.cpp, wrapAnsi.cpp, visible.rs, and mod.rs. Adds 7 tests (~100 lines).
Security risks
None. Pure computation over an in-memory byte span; no allocation, no I/O, bounded loops. The added state machine indexes input[i]/input[i+1] only after explicit i+1 < len checks.
Level of scrutiny
Moderate. The default behavior (ambiguousIsNarrow: true) is byte-for-byte unchanged — new code only runs on explicit opt-in. But this is native C++ in a documented performance-sensitive path (Bun.stringWidth), changes an extern "C" ABI consumed from both Rust and C++, and introduces a second copy of the CSI/OSC state machine (vs. the highway kernel's VisibleLatin1WidthExcludeANSIScalar) that could drift. A maintainer familiar with the highway_strings subsystem should confirm the duplicate-state-machine approach is preferred over parameterizing the SIMD kernel or exposing the existing scalar with a width callback.
Other factors
- All call sites of the changed signatures are updated (verified via grep across Rust and C++; the only remaining
false-hardcoded caller isvisibleUTF8WidthExcludeANSI, which is correct since its ASCII runs contain no bytes ≥ 0xA1). - The compile-time
static_assertproving narrow-table equivalence withvisibleLatin1WidthScalaris a nice guardrail against future drift between the fused table and the SIMD scalar. - Test coverage is thorough: all 256 Latin-1 bytes × both encodings × all option combinations, SIMD chunk boundaries, CSI/OSC interleaving, and npm
string-widthparity. - No prior human review comments; CI still building at time of review.
|
CI build #73973: 282/286 jobs passed. The stringWidth/wrapAnsi/sliceAnsi suites passed on every lane. The four red jobs are unrelated main breaks (worker-teardown ASAN leak in |
|
Superseded by #33488, which adds the Latin-1 |
|
Closing in favor of #33488. |
…iceAnsi and wrapAnsi (#33488) `Bun.stringWidth`, `Bun.stripANSI`, `Bun.sliceAnsi` and `Bun.wrapAnsi` each carried their own escape recognizer, and the four disagreed on what a terminal would display. This consolidates them onto the shared grammar in `ANSIHelpers.h` and fixes the width / SGR / wrapping bugs that a terminal UI hits. Closes #34384, closes #34381, closes #34379, closes #34380, closes #34377, closes #34382, closes #34391, closes #34393, closes #34394. ## Escape grammar (all four APIs) | Form | before | after | |---|---|---| | CSI `ESC [` … `<final 0x40–0x7E>` | ✓ (width only) | ✓ | | C1 CSI `0x9B` … | width counted `31m` as text | zero-width | | OSC `ESC ]` / C1 `0x9D` … (BEL / `ESC \` / `0x9C`) | `0x9D` unrecognized | ✓ | | DCS/SOS/PM/APC `ESC P/X/^/_`, C1 `0x90/0x98/0x9E/0x9F` … ST | payload counted as visible | zero-width | | nF `ESC <0x20–2F> <byte>` (`ESC ( B`), Fe/Fs `ESC <0x30–7E>` (`ESC 7`, `ESC c`) | `(B0`, `7hi8` visible | zero-width | | ESC / CAN / SUB / C1 ST *inside* a sequence | swallowed following text (`stripANSI("\x1b]0;title\x1b[31mtext")` → `""`) | aborts the sequence (VT500): `"text"` | `stringWidth(s) === stringWidth(stripANSI(s))` and agreement with `sliceAnsi`/`wrapAnsi` now hold on all of the above, on Latin-1, UTF-16 and (via `inspect.table` / the markdown renderer) UTF-8 strings — previously a JS string could measure differently before and after a wide character forced it to UTF-16. ## Repro (release, before → after) ```js Bun.stringWidth("\x9B31mhi\x9B39m") // 8 → 2 Bun.stringWidth("a\x1bP+q544e\x1b\\b") // 10 → 2 Bun.stringWidth("hi\x1b7\x1b8") // 4 → 2 Bun.stripANSI("text\x1b[3\x1b[0mmore") // "text0mmore" → "textmore" Bun.stripANSI("\x1b]0;title\x1b[31mtext") // "" → "text" Bun.sliceAnsi("\x1b[1m\x1b[2mLoading deps\x1b[22m", 0, 10) // bold lost → "\x1b[1m\x1b[2mLoading de\x1b[22m" Bun.sliceAnsi("\x1b[4m\x1b[58;5;196mERROR\x1b[59m\x1b[24m", 0, 5) // injected blink (SGR 5) → underline-color preserved Bun.sliceAnsi("ab漢", 0, 3, { ellipsis: "…" }) // "ab漢" (4 cols) → "ab…" Bun.sliceAnsi("ЖЗИ", 1, 4, { ellipsis: ">>" }) // "" → ">>" Bun.sliceAnsi("\x1b[31mabcd\x1b[39m", 0, 4, { ellipsis: "…" }) // reordered ANSI / spurious ellipsis → "\x1b[31mabcd\x1b[39m" Bun.wrapAnsi("x\x1b(0lqqqqqqqqqqk\x1b(B", 6, { hard: true }) // one width-13 row → wraps at 6 Bun.wrapAnsi("\x1b]8;;http://x/" + " word".repeat(1000), 40) // quadratic output growth → linear Bun.stringWidth("السَّلَامُ عَلَيْكُمْ") // 21 → 12 (Mn marks zero-width) Bun.stringWidth("한국어.txt".normalize("NFD")) // 15 → 10 Bun.stringWidth("café", { ambiguousIsNarrow: false }) // 4 → 5 (Latin-1 path honored the flag only after a UTF-16 force) Bun.stringWidth("") // 1 → 2 (Unicode 16/17 EAW) ``` ## Fix | Layer | change | |---|---| | `ANSIHelpers.h` `consumeANSI()` | full grammar + VT500 anywhere-transitions (ESC re-introduces, CAN/SUB/ST abort); a `Utf8` mode where C1 = `0xC2 0x9x` and ST = `0xC2 0x9C` (a bare `0x9C` is a continuation byte); `sgrCloseCode`/`isSgrEndCode` gain 20/21/51/52/58/73/74 | | `stringWidth.cpp` | UTF-8 and UTF-16 paths drive `consumeANSI`; Latin-1 `ambiguousIsNarrow: false` path; VS16 widens only Emoji-property bases | | `highway_strings.cpp` Latin-1 kernel | `origin/main`'s single-pass loop kept; the fast-path gate also fails on `0x7F–0x9F`; a chunk containing any non-CSI/OSC introducer exits to the scalar recognizer (a call inside the loop spilled the caller-saved vector registers onto the fast path) | | `stringWidthTables.h` | regenerated from Unicode 17.0 UCD (was EAW 15.1 + emoji 17.0 + a hand-picked Mn whitelist): all Mn/Me zero-width, jamo V/T zero-width, new emoji wide | | `sliceAnsi.cpp` SGR state | keyed by attribute slot (bold+dim, italic+fraktur, single+double underline share closes but stack), 58/59 as an extended color, empty param = 0, abort bytes match the recognizer | | `sliceAnsi.cpp` ellipsis | spec-zone content written in place with an SGR/hyperlink snapshot restored on discard (keeps ANSI ordered); wide-cluster EOF overflow is a cut; trailing zero-width clusters are not; a range admitting no visible content returns just the ellipsis; the ellipsis inherits the active style | | `wrapAnsi.cpp` | tokenizer walks with `consumeANSI` (no more "inside escape until an `m`"), OSC-8 body never split (ST terminator + `id=` params, re-opened per row), `trim`/`wordWrap` honor `!== false` only, zero-width tail chars kept when trimming; a bare `\r` still breaks a line as before ( `wrap-text` maps wrapped output back onto the source by relying on it) | | `stripANSI.cpp` | no-op inputs ending in a stray `0x9C` keep zero-copy identity | ## Performance Agent shaped input (the workload these exist for: 21 real TUI lines, 3–139 code units, per-call, `origin/main` vs this branch, local release, min of 3): | call | main | this PR | |---|---|---| | `stringWidth(s)` | 175 ns | **159 ns** | | `sliceAnsi(s, 0, 40)` | 398 ns | 412 ns | | `wrapAnsi(s, 80)` | 997 ns | **789 ns** | | `wrapAnsi(s, 40, {hard:true})` | 1076 ns | 964 ns | | `stripANSI(s)` | 37 ns | 37 ns | | `stringWidth(s, {countAnsiEscapeCodes:true})` | 194 ns | 193 ns | Long single strings (4 KB–64 KB per call), where the SIMD kernels dominate: | shape | main | this PR | |---|---|---| | ASCII, no escapes, ANSI kernel | 20.6 B/ns | 18.4 B/ns | | dense SGR (11 KB) | 3.28 µs | 4.12 µs | | truecolor lines (9.4 KB) | 1.30 µs | 1.72 µs | | OSC-8 links (12.6 KB) | 3.40 µs | 3.97 µs | | UTF-16 SGR every 15 units (7.2 KB) | 7.86 µs | 9.38 µs | | UTF-16 dense SGR + one wide char (11 KB) | 18.7 µs | 21.8 µs | The long-string escape shapes are 15–30 % behind `origin/main`: the per-chunk cold path re-classifies the C1 range and the grammar-complete terminator sets. Growth is linear in every case (per-line cost on a 1600-line colored UTF-16 log: 0.024 µs/line, same as main). Short strings, which is nearly all real callers, are faster than `main`. ## Verification ``` $ bun bd test test/js/bun/util/{stringWidth,stripANSI,sliceAnsi,sliceAnsi-fuzz,wrapAnsi}.test.ts \ test/js/bun/console/{bun-inspect-table,console-table}.test.ts 997 pass 0 fail # was 918 before this PR's test additions $ bun bd test test/js/bun/md 1065 pass 0 fail $ USE_SYSTEM_BUN=1 bun test test/js/bun/util/<file>.test.ts # canary, without the fix stringWidth 33 fail · wrapAnsi 39 fail · stripANSI 10 fail · sliceAnsi 14 fail (96 total) ``` <!-- robobun:evidence:begin --> --- **[review]** gate passed · iteration 5 · 18 files touched <details><summary>fails on main (without fix)</summary> ```console ASAN without fix: 94 FAILED $ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/bun/util/sliceAnsi-fuzz.test.ts test/js/bun/util/sliceAnsi.test.ts test/js/bun/util/stringWidth.test.ts test/js/bun/util/stripANSI.test.ts test/js/bun/util/wrapAnsi.test.ts info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: checking for self-update (current version: 1.29.0) bun test v1.4.0 (62b75ff) test/js/bun/util/sliceAnsi-fuzz.test.ts: (pass) sliceAnsi invariants > output width never exceeds requested range [983.76ms] (pass) sliceAnsi invariants > slice of stripped equals stripped slice (for 1-width chars) [676.01ms] (pass) sliceAnsi invariants > adjacent slices cover full visible string [342.56ms] (pass) sliceAnsi invariants > output is always well-formed UTF-16 [764.16ms] (pass) sliceAnsi invariants > full slice preserves visible content [272.68ms] (pass) sliceAnsi invariants > slicing a slice is idempotent on visible content [314.49ms] (pass) sliceAnsi invariants > ... (truncated) release without fix: 3 FAILED bun test v1.4.0-canary.1 (4f78471) test/js/bun/util/sliceAnsi-fuzz.test.ts: (pass) sliceAnsi invariants > output width never exceeds requested range [5.43ms] (pass) sliceAnsi invariants > slice of stripped equals stripped slice (for 1-width chars) [2.84ms] (pass) sliceAnsi invariants > adjacent slices cover full visible string [1.52ms] (pass) sliceAnsi invariants > output is always well-formed UTF-16 [3.87ms] (pass) sliceAnsi invariants > full slice preserves visible content [1.47ms] (pass) sliceAnsi invariants > slicing a slice is idempotent on visible content [1.55ms] (pass) sliceAnsi invariants > ellipsis output width respects budget [2.53ms] (pass) sliceAnsi adversarial > inputs near SIMD stride boundaries [0.21ms] (pass) sliceAnsi adversarial > C1 ST at SIMD boundary positions [0.07ms] (pass) sliceAnsi adversarial > unterminated CSI sequences don't hang or overread [0.09ms] (pass) sliceAnsi adversarial > many SGR codes don't overflow or quadratic-slow [66.02ms] (pass) sliceAnsi adversarial > huge SGR params don't overflow uint32 [0.08ms] (pass) sliceAnsi adversarial > SGR with many parameters [0.07ms] (pass) sliceAnsi adversarial > string of only zero-width ... (truncated) ``` </details> <details><summary>passes on PR (with fix)</summary> ```console ASAN with fix: all passed $ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/bun/util/sliceAnsi-fuzz.test.ts test/js/bun/util/sliceAnsi.test.ts test/js/bun/util/stringWidth.test.ts test/js/bun/util/stripANSI.test.ts test/js/bun/util/wrapAnsi.test.ts info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: checking for self-update (current version: 1.29.0) bun test v1.4.0 (62b75ff) test/js/bun/util/sliceAnsi-fuzz.test.ts: (pass) sliceAnsi invariants > output width never exceeds requested range [869.07ms] (pass) sliceAnsi invariants > slice of stripped equals stripped slice (for 1-width chars) [580.96ms] (pass) sliceAnsi invariants > adjacent slices cover full visible string [333.91ms] (pass) sliceAnsi invariants > output is always well-formed UTF-16 [739.83ms] (pass) sliceAnsi invariants > full slice preserves visible content [268.37ms] (pass) sliceAnsi invariants > slicing a slice is idempotent on visible content [307.68ms] (pass) sliceAnsi invariants > ... (truncated) release with fix: all passed $ bun scripts/build.ts --profile=release info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: checking for self-update (current version: 1.29.0) [configured] bun-profile → bun (stripped) in 753ms (unchanged) ninja: Entering directory `/workspace/bun/build/release' [1/9] gen generated_host_exports.rs generated_host_exports.rs: 91 exports (host=3, lazy=10, generic=78, rust=0); 243 extern-C blocks audited [2/9] gen cpp.rs (cppbind) [2/9] cargo bun_bin → libbun_rust.a (--target x86_64-unknown-linux-gnu) info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05) info: component rust-src is up to date info: component rust-std is up to date nightly-2026-05-06-x86_64-unknown-linux-gnu unchanged - rustc 1.97.0-nightly (e95e73209 2026-05-05) info: checking for self-update (current version: 1.29.0) �[1m�[92m Compiling�[0m bun_core v0.0.0 (/workspace/bun/src/bun_core) �[1m�[92m Compiling�[0m bun_errno v0.0.0 (/wor ... (truncated) ``` </details> <details><summary>diff hotspot</summary> ``` docs/runtime/utils.mdx | 10 +- packages/bun-types/bun.d.ts | 3 +- scripts/generate-stringwidth-tables.mjs | 185 ++-- src/bun_core/string/immutable/visible.rs | 16 +- src/bun_core/string/mod.rs | 2 +- src/jsc/bindings/ANSIHelpers.h | 224 +++-- src/jsc/bindings/highway_strings.cpp | 395 +++++---- src/jsc/bindings/sliceAnsi.cpp | 358 +++++--- src/jsc/bindings/stringWidth.cpp | 568 ++++++------ src/jsc/bindings/stringWidth.h | 4 + src/jsc/bindings/stringWidthTables.h | 1387 ++++++++---------------------- src/jsc/bindings/stripANSI.cpp | 5 + src/jsc/bindings/wrapAnsi.cpp | 631 +++++++------- test/js/bun/util/sliceAnsi-fuzz.test.ts | 43 +- test/js/bun/util/sliceAnsi.test.ts | 316 ++++++- test/js/bun/util/stringWidth.test.ts | 711 +++++++++++++-- test/js/bun/util/stripANSI.test.ts | 52 +- test/js/bun/util/wrapAnsi.test.ts | 254 ++++++ 18 files changed, 2963 insertions(+), 2201 deletions(-) ``` </details> **gate history** · 3 passed · 1 rejected · iteration 5 <details><summary>evidence per changed file</summary> ``` file reads edits tests docs/runtime/utils.mdx 1 2 0 packages/bun-types/bun.d.ts 1 1 0 scripts/generate-stringwidth-tables.mjs 1 2 0 src/bun_core/string/immutable/visible.rs 0 0 0 src/bun_core/string/mod.rs 0 0 0 src/jsc/bindings/ANSIHelpers.h 4 6 0 src/jsc/bindings/highway_strings.cpp 12 7 0 src/jsc/bindings/sliceAnsi.cpp 3 2 0 src/jsc/bindings/stringWidth.cpp 6 22 0 src/jsc/bindings/stringWidth.h 0 0 0 src/jsc/bindings/stringWidthTables.h 0 0 0 src/jsc/bindings/stripANSI.cpp 1 0 0 src/jsc/bindings/wrapAnsi.cpp 1 1 0 test/js/bun/util/sliceAnsi-fuzz.test.ts 1 2 0 test/js/bun/util/sliceAnsi.test.ts 1 2 0 test/js/bun/util/stringWidth.test.ts 2 16 0 (+ 2 more files) ``` </details> <!-- robobun:evidence:end --> --------- Co-authored-by: Jarred Sumner <jarred@jarredsumner.com>
Reproduction
Bun.stringWidthreturned different widths for identical text depending on whether JSC stored it as Latin-1 or UTF-16.Bun.sliceAnsihonored the option on Latin-1 input viaBun__codepointWidth, so the documented measure-then-cut pairing disagreed with itself.Cause
43 bytes in U+00A1..U+00FE are East Asian Ambiguous (§, °, ±, ×, ÷, é, è, ê, Þ, ß, ...), but
visibleLatin1Width/visibleLatin1WidthExcludeANSIhad noambiguousAsWideparameter. Every caller ofBun__visibleWidthExcludeANSI_latin1(theBun.stringWidth8-bit branch,wrapAnsi'sstringWidth<Latin1>,sliceAnsi's ellipsis sizing, the RustString::visible_width_exclude_ansi_colorswrapper) silently dropped the flag, under a comment claiming Latin-1 has no ambiguous-width characters.Fix
Thread
ambiguousAsWidethrough both Latin-1 width functions and the C export, and pass it from all four call sites. The per-byte wide-mode widths are a compile-timestd::array<uint8_t, 256>derived from the samefusedClassifytable the UTF-16 path uses, with astatic_assertthat the table's narrow Latin-1 widths match the existing scalar exactly, so the two encodings cannot drift apart.Performance: the default (
ambiguousIsNarrow: true) keeps the SIMD fast paths byte-for-byte unchanged. The opt-in wide mode oncountAnsiEscapeCodes: trueadds a scalar per-byte correction to the SIMD narrow result. The wide + exclude-ANSI mode uses SIMD + correction when noESCis present, and a scalar ANSI state machine (mirroringVisibleLatin1WidthExcludeANSIScalar) otherwise so ambiguous bytes inside OSC payloads are not over-counted.Verification
New tests in
test/js/bun/util/stringWidth.test.ts(ambiguousIsNarrow is encoding-independent on Latin-1 text) cover:stringWidth/sliceAnsiagreement on Latin-1 ambiguous textstring-widthfor Latin-1 ambiguous charactersAll 7 fail on
main, all pass with this change. ExistingstringWidth(154),wrapAnsi, andsliceAnsi(303) tests pass unchanged.[review] gate passed · iteration 0 · 7 files touched
fails on main (without fix)
passes on PR (with fix)
diff hotspot
gate history · 1 passed · 0 rejected · iteration 0
evidence per changed file