Skip to content

Bump BoringSSL (oven-sh/boringssl#13 preview): CRYPTO_memcmp keeps an 8-bit accumulator under clang 23 - #43885

Draft
robobun wants to merge 1 commit into
mainfrom
robobun/42697d80/bump-boringssl-memcmp-barrier
Draft

robobun wants to merge 1 commit into
mainfrom
robobun/42697d80/bump-boringssl-memcmp-barrier

Conversation

@robobun

@robobun robobun commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Behaviour change: none

Problem

Fix

  • Pin BoringSSL to the head of Return CRYPTO_memcmp's accumulator through a value barrier boringssl#13 (51a84a81). It returns the accumulator through value_barrier_w, which keeps it 8 bits wide. It is the only commit on top of the current pin: 1 file, +6/-1.
  • This is a preview pin: CI builds and tests the fork change on every platform. It is a draft until Return CRYPTO_memcmp's accumulator through a value barrier boringssl#13 merges. Then the pin moves to the merged commit.
  • Linux x64 release builds, ns per call, 1.4.2 / main / this PR: timingSafeEqual at 1 KB 40.9 / 140.9 / 41.4, KeyObject.equals at 1 KB 34.1 / 132.7 / 37.8.
  • Verified: CRYPTO_memcmp in the release binary has 4 movdqu and no pmovzxbd (main: 2 pmovzxbd, 8-byte steps). crypto.test.ts, crypto.key-objects.test.ts, csrf.test.ts, password.test.ts, node-tls-connect.test.ts and node's test-crypto-timing-safe-equal.js pass.

Background

  • CRYPTO_memcmp is BoringSSL's constant-time compare. It XORs every byte pair into one accumulator and never exits early.
  • A value barrier is an empty inline-asm statement. The compiler cannot reason about a value that passes through it. Here it also stops the fold that widens the accumulator.
  • scripts/build/deps/boringssl.ts pins the fork by commit and fetches oven-sh/boringssl/archive/<sha>.tar.gz. A pull request head resolves there.
Notes

Numbers. ns per call, Xeon Platinum 8375C, linux x64 release builds of main (8d36bff) and of this branch, plus the 1.4.2 release. Each run takes the best of 25 batches per case. The table shows the minimum of 7 runs, interleaved across the three binaries. The host was loaded (load average 48 on 16 cores), so the medians are not stable. The minimums repeat: a second series of 5 interleaved runs differs by at most 0.6 ns up to 1 KB, 2.1 ns at 4 KB and 37 ns at 64 KB.

case 1.4.2 main this PR
timingSafeEqual 16 B 20.4 22.2 21.5
timingSafeEqual 64 B 23.3 28.9 24.7
timingSafeEqual 256 B 27.1 51.8 27.7
timingSafeEqual 1 KB 40.9 140.9 41.4
timingSafeEqual 4 KB 97.7 503.5 98.3
timingSafeEqual 64 KB 2241 8491 2241
KeyObject.equals 16 B 9.4 9.5 9.7
KeyObject.equals 64 B 15.8 23.5 16.8
KeyObject.equals 256 B 19.8 44.3 21.3
KeyObject.equals 1 KB 34.1 132.7 37.8
KeyObject.equals 4 KB 91.4 493.5 101.4
KeyObject.equals 64 KB 2239 8528 2236

bench/crypto/constant-time-compare.mjs (new, mitata) runs the same two calls at the same sizes. Only linux x64 is measured. The arm64 loop is also slower under clang 23 (16 tbl per 64 bytes), and the barrier restores the old arm64 loop in a standalone compile.

Why no new test. The change has no effect that a caller can observe except speed, and a wall-clock assertion is not stable on CI or under the debug and ASAN builds. The existing suites cover correctness of every caller. The Postgres SCRAM login test (test/js/sql/sql.test.ts) needs Docker and did not run locally. CI runs it.

Scope of the bump. gh api repos/oven-sh/boringssl/compare/41bf9b59...51a84a81 reports 1 commit ahead, 0 behind, and one file, crypto/mem.cc.

Relation to #43095. #43095 adds a runtime-dispatched Highway kernel for timingSafeEqual from 64 bytes. With this pin, 64 to 256 byte compares are already within 1 to 5 ns of that kernel (24.7 and 27.7 ns here, 21.5 and 22.4 ns there). The kernel's gain starts at about 1 KB: 41 ns against 26 ns, and 2241 ns against 849 ns at 64 KB, on a host with AVX-512.


no test proof · iteration 0 · no src or test change; test-proof not applicable

… barrier

Preview pin of oven-sh/boringssl#13, the one commit on top of the current pin.
Clang 23 otherwise folds the int conversion of the uint8_t accumulator into the
loop (llvm/llvm-project#222142) and compares 8 bytes per iteration instead of 32.

bench/crypto/constant-time-compare.mjs measures the two JS entry points that
reach CRYPTO_memcmp with inputs of any size: crypto.timingSafeEqual and
KeyObject.equals.
@robobun

robobun commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator Author

Status

Reproduced with release builds of main (8d36bff) and of this branch on linux x64, interleaved runs, best of 25 batches per case, minimum of 7 runs:

  • crypto.timingSafeEqual on 1 KB: 140.9 ns per call on main, 41.4 ns with this pin (40.9 ns on Bun 1.4.2)
  • KeyObject.equals on 1 KB keys: 132.7 ns on main, 37.8 ns with this pin (34.1 ns on Bun 1.4.2)
  • CRYPTO_memcmp in bun-profile: 2 pmovzxbd and 8-byte steps on main, 4 movdqu and 32-byte steps with this pin

bench/crypto/constant-time-compare.mjs in this PR runs both calls from 16 bytes to 64 KB.

This is a draft on purpose. It pins the head of oven-sh/boringssl#13, which is not merged yet. When it merges, the pin moves to the merged commit and the draft state ends.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants