Skip to content

Bun.serve websocket: don't alias pub/sub topics that differ only in lone surrogates - #32891

Closed
robobun wants to merge 4 commits into
mainfrom
farm/36efef94/ws-topic-surrogate-aliasing
Closed

robobun wants to merge 4 commits into
mainfrom
farm/36efef94/ws-topic-surrogate-aliasing

Conversation

@robobun

@robobun robobun commented Jun 27, 2026

Copy link
Copy Markdown
Collaborator

Problem

Bun.serve WebSocket pub/sub keys topic names by converting the JS string to UTF-8 via a lossy path that replaces every unpaired surrogate with U+FFFD. Distinct JS strings "a\uD800b", "a\uDC00b", "a\uDFFFb", and the literal "a\uFFFDb" all produce the same byte key, so they alias onto the same topic:

const TA = "a\uD800b", TB = "a\uDC00b", TC = "a\uFFFDb";
// in websocket.open:
ws.subscribe(TA);               // subscribe ONLY to TA
ws.isSubscribed(TB);            // true  (never subscribed to TB)
ws.isSubscribed(TC);            // true  (never subscribed to TC)
server.subscriberCount(TB);     // 1
server.publish(TB, "x");        // delivered to the TA subscriber

Any application that builds topic names from user-controlled input (room names, user ids) can be made to receive, or have its publishes delivered to, a topic it never named by injecting a lone surrogate.

Root cause

JSValue::to_slice / ZigString::to_slice route 16-bit strings through strings::to_utf8_alloc, whose scalar fallback (append_wtf8_from_utf16) decodes with decode_utf16_with_fffd, mapping every unpaired surrogate to U+FFFD. The topic map in uWS then sees identical byte keys.

Fix

Add a WTF-8 conversion path (strings::to_wtf8_alloc, String::to_wtf8, ZigString::to_slice_wtf8, JSValue::to_slice_wtf8, JSString::to_slice_wtf8) that encodes unpaired surrogates as their 3-byte WTF-8 sequence (ED A0 80..ED BF BF) instead of replacing them. Distinct JS strings now produce distinct byte keys.

All topic-name call sites switch to the WTF-8 path:

  • ServerWebSocket: subscribe / unsubscribe / isSubscribed / publish / publishText / publishBinary (and the *_without_type_checks fast paths)
  • server: publish / subscriberCount

ws.subscriptions now decodes the stored WTF-8 bytes back to WTF-16 so the returned array round-trips the original JS string.

Well-formed strings (ASCII, Latin-1, valid surrogate pairs / emoji) encode to exactly the same bytes as before, so existing topic names are unchanged.

Verification

New test in test/js/bun/websocket/websocket-server.test.ts subscribes to "a\uD800b" and "a😶b" only, then asserts:

  • subscriberCount / isSubscribed report 0 / false for "a\uDC00b", "a\uFFFDb", "a\uDFFFb"
  • unsubscribe on those other topics returns false and leaves the real subscription intact
  • ws.subscriptions returns the exact original strings
  • server.publish / ws.publish / ws.publishText / ws.publishBinary to any of the other topics deliver nothing to the subscriber
  • publishing to the subscribed topics still delivers

Test fails on the released binary and passes with this change.

@coderabbitai

coderabbitai Bot commented Jun 27, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 58 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c06995fb-44df-4b62-901f-0eb621b4794d

📥 Commits

Reviewing files that changed from the base of the PR and between fb24aac and 317621b.

📒 Files selected for processing (12)
  • docs/guides/util/base64.mdx
  • docs/runtime/web-apis.mdx
  • src/bun_core/lib.rs
  • src/bun_core/string/immutable.rs
  • src/bun_core/string/mod.rs
  • src/bun_core/string/wtf.rs
  • src/jsc/JSString.rs
  • src/jsc/JSValue.rs
  • src/jsc/bindings/uws_bindings.cpp
  • src/runtime/server/ServerWebSocket.rs
  • src/runtime/server/server_body.rs
  • test/js/bun/websocket/websocket-server.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jun 27, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 9:36 AM PT - Jun 29th, 2026

❌ @robobun, your commit ab79e73 has 3 failures in Build #66795 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 32891

That installs a local version of the PR into your bun-32891 executable, so you can run:

bun-32891 --bun

Comment thread src/runtime/server/ServerWebSocket.rs

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the doc-comment fix — nothing else from me. Deferring to a human reviewer since this threads a new WTF-8 encoding path through core string infrastructure (unsafe simdutf FFI + a hand-rolled C++ WTF-8 decoder), which is more than I'm comfortable auto-approving.

Extended reasoning...

Overview

This PR fixes pub/sub topic aliasing in Bun.serve WebSockets: topic names that differ only in lone surrogates previously collapsed to the same byte key because the JS→bytes path replaced unpaired surrogates with U+FFFD. The fix adds a parallel WTF-8 encoding path — strings::to_wtf8_alloc, ZigString::to_slice_wtf8, WTFStringImplExt::to_wtf8, String::to_wtf8, JSString::to_slice_wtf8, JSValue::to_slice_wtf8 — and switches every topic-name call site (subscribe/unsubscribe/isSubscribed/publish/publishText/publishBinary on both ServerWebSocket and the server, plus the *_without_type_checks fast paths) to it. uws_bindings.cpp gains a hand-rolled topicBytesToString WTF-8→UTF-16 decoder so ws.subscriptions round-trips the original JS strings. 10 files changed plus a thorough new integration test.

Security risks

The bug being fixed is itself security-adjacent (user-controlled topic names with injected lone surrogates could alias onto another subscriber's topic). The fix introduces no new untrusted-input surfaces, but the new encoding primitives must be correct: to_wtf8_alloc wraps a simdutf FFI call inside unsafe { vec::fill_spare(...) } and falls back to a scalar encoder on SURROGATE; the C++ side adds a manual multi-byte decoder. Both only ever see bytes that Bun itself produced via the new encoder (uWS stores topic keys verbatim), so malformed-input handling is defensive rather than exposed — but encoding bugs here would be subtle and broadly reachable.

Level of scrutiny

Medium-high. This is not a config tweak or a localized bugfix — it adds a new encoding family to bun_core::strings / bun_core::string and the JSC bridge layer, with unsafe FFI and a bespoke C++ decoder. Well-formed strings are claimed (and appear) byte-identical to the existing UTF-8 path, so the blast radius for normal inputs is small, but the new primitives become part of the core string toolkit and deserve a human pass on the capacity math (len * 3 + 16), the simdutf-vs-scalar fallback, and the C++ decoder's branch coverage.

Other factors

  • My only prior comment (stale "UTF-8 slice" doc string on topic_dispatch) was addressed in 6eb39c5.
  • The bug-hunting pass found no issues this round.
  • The new test exercises every entry point (server.publish, ws.publish/publishText/publishBinary, subscribe/unsubscribe/isSubscribed, subscriberCount, ws.subscriptions round-trip) against four distinct surrogate-bearing topics plus a valid-pair control, which is solid coverage.
  • The single CI failure (sql-mysql.auth.test.ts on one Linux lane) is unrelated to anything this PR touches.

Net: looks correct to me, but the scope and the new unsafe/C++ encoding code put it outside what I'll auto-approve.

@robobun

robobun commented Jun 27, 2026 •

Copy link
Copy Markdown
Collaborator Author

Final CI status, after build 66795 completed on the current head (ab79e732).

Summary: 281 test lanes passed and 5 failed. Every failure is a tracked repo-wide CI issue, infrastructure, or an agent timeout that the same test passes on a sibling lane in the same build. None executes any code this PR changes, and test/js/bun/websocket/websocket-server.test.ts appears in no failure. I have stopped re-triggering; this is ready for review.

The failures, by lane:

  • alpine 3.23 x64 and x64-baseline: test/js/node/test/parallel/test-net-connect-memleak.js. This is CI: test-net-connect-memleak.js fails on half of PR builds on linux-x64-musl since June 28 ~23:00 UTC #33044, an open issue titled "CI: test-net-connect-memleak.js fails on half of PR builds on linux-x64-musl since June 28 ~23:00 UTC", i.e. a known repo-wide CI failure that began on main after this PR was opened. For reference, main build 66758 at fb24aac7, the exact commit this branch is rebased onto, has every test lane green.
  • darwin 14 aarch64: test/regression/issue/20965.test.ts (streaming file-response abort) timed out at 90s. The same test with the same binary passed on darwin 14 x64 in this same build, so this is the agent rather than the code; the darwin fleet produced three independent timeouts on this one build.
  • darwin 14 x64: test/js/bun/terminal/terminal.test.ts ("creates subprocess with terminal attached") timed out at 90s. Same darwin agent class.
  • darwin 26 aarch64: buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun', a 14 KB log in which no test ran. This is the fourth consecutive build in which this exact lane fails with this exact message (65676, 65723, 66778, 66795); it is that agent's artifact download, independent of any commit.

The Windows failures were all auto-annotated as flaky by CI and retried (bun-install-registry.test.ts, jsc-stress, napi.test.ts).

None of these can involve this diff: a net.connect GC test, an HTTP file-stream abort test, and a PTY test execute zero changed lines, since every code change here is in the WebSocket pub/sub topic path plus new string helpers that only those call sites use.

History, for completeness: across all five CI runs of this PR the only failure that was ever real was AsyncLocalStorage-tracking.test.ts segfaulting on the pre-rebase base, which main had already fixed in 0f9331d (#32184); rebasing resolved it. Every review comment raised has been addressed and its thread resolved.

@robobun
robobun force-pushed the farm/36efef94/ws-topic-surrogate-aliasing branch from 6eb39c5 to a1cf85e Compare June 28, 2026 01:27

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found beyond the resolved nit, but this adds new WTF-8 encoding infrastructure across bun_core/jsc (including an unsafe simdutf path) plus a hand-rolled WTF-8 decoder in C++, so it's worth a human pass on the encoding details.

Extended reasoning...

Overview

This PR fixes pub/sub topic aliasing in Bun.serve WebSockets: previously, topic names were keyed via lossy UTF-8 conversion that maps every unpaired surrogate to U+FFFD, so distinct JS strings like "a\uD800b" and "a\uDC00b" collided onto the same uWS topic. The fix introduces a parallel WTF-8 conversion path (to_wtf8_alloc, String::to_wtf8, ZigString::to_slice_wtf8, JSValue::to_slice_wtf8, JSString::to_slice_wtf8) that encodes lone surrogates as their 3-byte WTF-8 sequence, switches all topic-name call sites in ServerWebSocket.rs and server_body.rs to it, and adds a hand-rolled WTF-8 → UTF-16 decoder (topicBytesToString) in uws_bindings.cpp so ws.subscriptions round-trips. Ten files touched plus a thorough new integration test.

Security risks

The bug being fixed is itself security-adjacent (user-controlled topic names could alias onto unrelated topics). The fix doesn't introduce new attack surface that I can see — WTF-8 bytes stay internal to the uWS topic map and never go on the wire. The new unsafe block in to_wtf8_alloc mirrors the existing convert_utf16_to_utf8_append pattern exactly (same simdutf call, same fill_spare contract, same SURROGATE→0-commit fallback), and the len*3+16 capacity bound is correct for WTF-8 (each u16 → ≤3 bytes). The C++ decoder only ever sees bytes produced by our own encoder, so its lax continuation-byte handling isn't exploitable, but it is a fresh hand-rolled multi-byte decoder.

Level of scrutiny

Medium-high. This is not a mechanical change: it adds new public string-encoding API across three crate layers, a new unsafe simdutf block, and a custom C++ decoder, and it changes the keying semantics of a production pub/sub system. The implementation looks correct to me (decode_wtf16_raw + encode_wtf8_rune are pre-existing and the new code just composes them; the C++ side correctly handles the 4-byte → surrogate-pair case), but encoding edge cases and the design choice of WTF-8 keying are things a maintainer should sign off on.

Other factors

  • My earlier doc-comment nit was addressed (commit a1cf85e); no outstanding review comments remain.
  • CI failures are unrelated infra (MySQL docker harness) per the author's analysis; the new websocket test passes on all lanes that ran.
  • Test coverage for the new behavior is comprehensive (subscribe/unsubscribe/isSubscribed/subscriberCount/subscriptions round-trip plus all four publish entry points), and existing well-formed-string tests in the same file continue to exercise the unchanged-bytes guarantee.

@robobun

robobun commented Jun 28, 2026

Copy link
Copy Markdown
Collaborator Author

Addressing the two spots flagged for a closer look, here are the invariants they rely on, so a reviewer does not have to reconstruct them.

to_wtf8_alloc (the unsafe simdutf path):

  • utf16.len() * 3 is a hard upper bound on what simdutf can write, with or without an error: a single UTF-16 code unit encodes to at most 3 UTF-8 bytes, and a surrogate pair consumes 2 units to emit 4. So the with_capacity(len * 3 + 16) spare is sufficient even when the converter stops mid-input on SURROGATE. This is the same spare-sizing contract the existing convert_utf16_to_utf8_append already relies on for lone-surrogate input, just with a looser (larger) bound.
  • On SURROGATE we commit 0 bytes, discarding whatever simdutf scribbled, and the scalar pass rewrites from an empty Vec. The scalar pass composes two pre-existing helpers, decode_wtf16_raw (combines a valid lead+trail pair, otherwise returns the lone surrogate code unit verbatim) and encode_wtf8_rune. For a code point in 0xD800..=0xDFFF the latter emits ED A0 80..ED BF BF. Valid UTF-8 never produces those byte sequences, so the WTF-16 to WTF-8 mapping is injective: a lone-surrogate topic cannot collide with any well-formed topic, including a literal U+FFFD.
  • For well-formed input the fallback produces byte-identical output to simdutf, which is what keeps existing topic names keyed to the same bytes as before.

topicBytesToString in uws_bindings.cpp:

  • It only ever sees bytes that Bun's own encoder wrote into the uWS topic map (nothing else creates topic entries), so it is a left-inverse of the encoder rather than a general-purpose validator. It is still memory-safe on arbitrary input: every multi-byte arm bounds-checks (i+1 < n, i+2 < n, i+3 < n) before reading continuation bytes, and every iteration advances i by at least 1, so it cannot read out of bounds or loop forever. reserveInitialCapacity(n) is an upper bound on the UTF-16 length since each output code unit consumes at least one input byte.
  • The 3-byte arm is the load-bearing one: it reconstructs 0xD800..0xDFFF as a single char16_t. That is exactly why fromUTF8ReplacingInvalidSequences could not be kept, since it maps those sequences to U+FFFD and would break the ws.subscriptions round trip the test asserts.

Happy to restructure either piece if a different shape is preferred.

robobun added 2 commits June 29, 2026 13:44
…one surrogates

Topic names were converted from JS strings to bytes via a lossy
UTF-16 to UTF-8 path that replaces every unpaired surrogate with
U+FFFD. Distinct JS strings like "a\uD800b", "a\uDC00b",
"a\uDFFFb", and the literal "a\uFFFDb" all produced the same
byte key, so subscribing to one made the socket a subscriber of all
of them and publishes to any of them were delivered cross-topic.

Key topic names as WTF-8 instead: unpaired surrogates are encoded
as their 3-byte WTF-8 sequence so distinct JS strings map to
distinct byte keys. ws.subscriptions decodes WTF-8 back so the
round trip returns the original JS string. Well-formed strings
(ASCII, Latin-1, valid surrogate pairs) encode to the same bytes
as before.
@robobun
robobun force-pushed the farm/36efef94/ws-topic-surrogate-aliasing branch from a1cf85e to e96486c Compare June 29, 2026 13:51
@robobun

robobun commented Jun 29, 2026

Copy link
Copy Markdown
Collaborator Author

Rebased onto main (fb24aac7) to resolve the merge conflict; force-pushed as e96486ca0a. The diff is semantically unchanged, so for the record:

Main moved 45 commits, two of which landed on top of this change.

  • 0f9331d / bun_core: make bun_core::strings a direct alias of string::immutable #33035 restructured bun_core::strings (the old pub mod strings shim in lib.rs is gone; bun_core::strings is now a direct alias of string::immutable, with the implementations in strings_impl). This produced the only real conflict, in src/bun_core/string/immutable.rs: main added narrow_ascii_u16 to the same pub use crate::strings_impl::{...} list where I add to_wtf8_alloc. Resolved by taking the union. My two new functions (to_wtf8_alloc, append_wtf8_from_utf16_lossless) land inside strings_impl next to their existing siblings, matching the new layout.
  • 9f18300 / Bun.serve websocket: make publish() return 0/-1 on subscriber backpressure #32889 changed the return semantics of the same publish() functions this PR edits (their change is the return value, mine is the topic-key encoding). Git merged these cleanly since the intents are orthogonal, and I verified by running the whole suite rather than just my test: all 7 of Bun.serve websocket: make publish() return 0/-1 on subscriber backpressure #32889's new publish() return value reflects subscriber backpressure tests pass alongside the new lone-surrogate test, 106 pass / 0 fail across test/js/bun/websocket/websocket-server.test.ts.

cargo check -p bun_runtime is clean and cargo fmt produced no changes.

@mintlify

mintlify Bot commented Jun 29, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated (UTC)
bun 🟢 Ready View Preview Jun 29, 2026, 1:57 PM

💡 Tip: Enable Workflows to automatically generate PRs for you.

Comment thread src/runtime/server/ServerWebSocket.rs Outdated
…Text, not publishBinary

publish_text_without_type_checks copied its empty-topic guard from
publish_binary_without_type_checks, including the method name in the
error string. Every other topic error message in this file already
names its own function.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both earlier nits are addressed and I didn't find any new issues, but the new unsafe simdutf path in to_wtf8_alloc and the hand-rolled WTF-8→WTF-16 decoder in uws_bindings.cpp warrant a human pass.

Extended reasoning...

Overview

This PR fixes topic-name aliasing in Bun.serve WebSocket pub/sub: distinct JS strings that differ only in lone surrogates previously collapsed to the same uWS topic key because the JS→bytes conversion replaced every unpaired surrogate with U+FFFD. The fix introduces a WTF-8 encoding path (to_wtf8_alloc in bun_core, plus to_wtf8/to_slice_wtf8 wrappers up through JSValue/JSString), switches all ~10 topic-name call sites in ServerWebSocket.rs and server_body.rs to it, and adds a custom C++ WTF-8 decoder (topicBytesToString) so ws.subscriptions round-trips. A thorough new test exercises subscribe/unsubscribe/isSubscribed/subscriberCount/publish across all entry points. The two .mdx changes are autofix.ci whitespace reflow only.

Security risks

The bug being fixed is itself security-relevant (cross-topic message delivery via surrogate injection). The fix introduces new unsafe Rust (vec::fill_spare over a simdutf FFI call with a manually-justified len*3+16 capacity bound) and a hand-written multi-byte decoder in C++. The author's invariant write-up is convincing — the capacity bound holds, the C++ decoder bounds-checks every continuation byte and always advances, and reserveInitialCapacity(n) only covers the 4-byte arm by one (it emits 2 code units but consumes 4 input bytes, so n remains an upper bound). I don't see a memory-safety issue, but this is exactly the kind of code where a second pair of human eyes is the right call.

Level of scrutiny

Medium-high. This is not a mechanical change: it threads a new encoding semantics through five layers of the string stack, adds unsafe FFI interop, and replaces a WTF library call (fromUTF8ReplacingInvalidSequences) with a bespoke decoder. The design choice (WTF-8 keying rather than, say, rejecting ill-formed topic names) is reasonable and preserves backward compatibility for well-formed input, but it's a decision a maintainer should ratify.

Other factors

Both inline nits I raised on earlier revisions (stale "UTF-8 slice" doc comment; wrong method name in the publishText fast-path error) have been fixed in e96486c and ab79e73. The current bug-hunting pass found nothing. CI's only failure (test-net-connect-memleak.js) is an unrelated node-compat flake. The new test is comprehensive and the author has documented the safety invariants in-thread, which should make the human review quick.

@robobun

robobun commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

Closing as part of a cleanup of stale pull requests. This PR has had no new commits since 2026-06-29, it conflicts with main, and its last CI run failed. This is not a judgment on the fix itself. If the problem still reproduces on a current build, reopen this PR after a rebase or open a new one against main.

@robobun robobun closed this Sep 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant