Skip to content

Return CRYPTO_memcmp's accumulator through a value barrier - #13

Merged
Jarred-Sumner merged 1 commit into
oven-sh:masterfrom
robobun:robobun/crypto-memcmp-value-barrier
Sep 26, 2026
Merged

Jarred-Sumner merged 1 commit into
oven-sh:masterfrom
robobun:robobun/crypto-memcmp-value-barrier

Conversation

@robobun

@robobun robobun commented Sep 17, 2026 •

Copy link
Copy Markdown

Problem

Fix

  • Return x through value_barrier_w. A crypto_word_t is wider than int, so the fold does not apply and the accumulator stays 8 bits wide.
  • The loop does not change, so the function stays constant-time. The result is now also opaque to the compiler where CRYPTO_memcmp is inlined.
  • Verified standalone with clang 23.1.2 for x64 (-march=nehalem) and arm64: the old loop returns. Verified in a Bun linux x64 release build with ThinLTO: KeyObject.equals on 1 KB keys goes from 130 ns to 37 ns (31 ns on Bun 1.4.2).

Background

  • CRYPTO_memcmp is the constant-time compare behind AEAD tag checks and TLS MACs. In Bun it also serves crypto.timingSafeEqual, KeyObject.equals, CSRF tokens and the Postgres SCRAM check.
  • value_barrier_w (crypto/internal.h) returns its argument through an empty inline-asm statement. The compiler cannot reason about the value, but it still vectorizes the loop.
  • Bun pins this fork by commit in scripts/build/deps/boringssl.ts. A bun PR bumps the pin after this merges.
Notes

In the Bun release build, CRYPTO_memcmp disassembles to movdqu/pxor/por again, and KeyObject.equals on 64 KB keys goes from 9.0 µs to 2.3 µs (2.2 µs on 1.4.2). Without the barrier, x64 gets pmovzxbd + por on 4 x i32 and arm64 gets 16 tbl per 64 bytes.

Scalar IR from clang 23 without the barrier: %x = phi i32, zext i8 ... to i32, or i32. With the barrier: %x = phi i8, or i8, and one zext i8 ... to i64 after the loop. The fold in llvm/llvm-project#222142 is zext(trunc nuw X) -> X. It needs the extended type to match the type of X, which is int here. On a 32-bit target crypto_word_t is as wide as int, so this change does not help there. Bun builds 64-bit targets only.

Upstream BoringSSL has the same source and the same slowdown under clang 23.

Clang 23 folds the conversion of the uint8_t accumulator to int into the loop
(llvm/llvm-project#222142). The accumulator becomes 32 bits wide, and the loop
vectorizer then puts 4 bytes in a 128-bit vector instead of 16. A 1 KB compare
takes about 120 ns instead of about 26 ns.

value_barrier_w takes a crypto_word_t, which is wider than int, so the fold does
not apply and the accumulator stays 8 bits wide. The result is also opaque to the
compiler where CRYPTO_memcmp is inlined.
@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 6377be51-f0ad-4638-8f3c-999d7757ead8

📥 Commits

Reviewing files that changed from the base of the PR and between 41bf9b5 and 51a84a8.

📒 Files selected for processing (1)
  • crypto/mem.cc

Included review availability: Your plan provides up to 5 included reviews per hour; 0 remain after this review.


Walkthrough

CRYPTO_memcmp now converts its accumulated comparison result to int and returns it through value_barrier_w.

Changes

Constant-time comparison

Layer / File(s) Summary
Barrier-protected comparison return
crypto/mem.cc
CRYPTO_memcmp passes its accumulated result, converted to int, through value_barrier_w instead of returning it directly.

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to 51a84

The change preserves comparison semantics while improving compiler-generated performance, with no established merge-blocking risk.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: routing the CRYPTO_memcmp accumulator through a value barrier.
Description check ✅ Passed The description directly explains the performance problem, the value_barrier_w fix, constant-time behavior, target scope, and benchmark results.

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Sep 24, 2026

Copy link
Copy Markdown
Author

oven-sh/bun#43885 pins this commit (51a84a8) as a preview, so bun's CI builds and tests it on every platform. It is a draft until this PR merges.

Measured there on linux x64 release builds (ns per call, Bun 1.4.2 / main / with this commit): crypto.timingSafeEqual at 1 KB 40.9 / 140.9 / 41.4, KeyObject.equals at 1 KB 34.1 / 132.7 / 37.8, both at 64 KB 2240 / 8500 / 2240.

@robobun

robobun commented Sep 24, 2026

Copy link
Copy Markdown
Author

bun's CI passed with this commit pinned (oven-sh/bun#43885, Buildkite build 120273): 181 of 181 jobs. That covers the build and the full test suite on linux x64 and aarch64 (glibc, musl, android, ASAN), darwin x64 and aarch64, freebsd x64 and aarch64, windows x64 and aarch64, plus the baseline instruction scans.

@Jarred-Sumner
Jarred-Sumner merged commit 525782a into oven-sh:master Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants