Repository navigation
Conversation
… barrier Preview pin of oven-sh/boringssl#13, the one commit on top of the current pin. Clang 23 otherwise folds the int conversion of the uint8_t accumulator into the loop (llvm/llvm-project#222142) and compares 8 bytes per iteration instead of 32. bench/crypto/constant-time-compare.mjs measures the two JS entry points that reach CRYPTO_memcmp with inputs of any size: crypto.timingSafeEqual and KeyObject.equals.
Collaborator
Author
|
Status Reproduced with release builds of main (8d36bff) and of this branch on linux x64, interleaved runs, best of 25 batches per case, minimum of 7 runs:
This is a draft on purpose. It pins the head of oven-sh/boringssl#13, which is not merged yet. When it merges, the pin moves to the merged commit and the draft state ends. |
This was referenced Sep 24, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Behaviour change: none
Problem
CRYPTO_memcmpwith a 32-bit accumulator ([LoopVectorize] i8 OR-reduction vectorized with i32 lanes since InstSimplify zext(trunc nuw) fold (#204089) llvm/llvm-project#222142). The loop handles 8 bytes per iteration instead of 32. Constant-time compares from 64 bytes up are 1.3 to 3.8 times slower than on Bun 1.4.2.crypto.timingSafeEqual,KeyObject.equals, the CSRF,Bun.passwordand Postgres SCRAM checks, and BoringSSL's own AEAD and TLS MAC checks.Fix
value_barrier_w, which keeps it 8 bits wide. It is the only commit on top of the current pin: 1 file, +6/-1.timingSafeEqualat 1 KB 40.9 / 140.9 / 41.4,KeyObject.equalsat 1 KB 34.1 / 132.7 / 37.8.CRYPTO_memcmpin the release binary has 4movdquand nopmovzxbd(main: 2pmovzxbd, 8-byte steps).crypto.test.ts,crypto.key-objects.test.ts,csrf.test.ts,password.test.ts,node-tls-connect.test.tsand node'stest-crypto-timing-safe-equal.jspass.Background
CRYPTO_memcmpis BoringSSL's constant-time compare. It XORs every byte pair into one accumulator and never exits early.scripts/build/deps/boringssl.tspins the fork by commit and fetchesoven-sh/boringssl/archive/<sha>.tar.gz. A pull request head resolves there.Notes
Numbers. ns per call, Xeon Platinum 8375C, linux x64 release builds of main (8d36bff) and of this branch, plus the 1.4.2 release. Each run takes the best of 25 batches per case. The table shows the minimum of 7 runs, interleaved across the three binaries. The host was loaded (load average 48 on 16 cores), so the medians are not stable. The minimums repeat: a second series of 5 interleaved runs differs by at most 0.6 ns up to 1 KB, 2.1 ns at 4 KB and 37 ns at 64 KB.
timingSafeEqual16 BtimingSafeEqual64 BtimingSafeEqual256 BtimingSafeEqual1 KBtimingSafeEqual4 KBtimingSafeEqual64 KBKeyObject.equals16 BKeyObject.equals64 BKeyObject.equals256 BKeyObject.equals1 KBKeyObject.equals4 KBKeyObject.equals64 KBbench/crypto/constant-time-compare.mjs(new, mitata) runs the same two calls at the same sizes. Only linux x64 is measured. The arm64 loop is also slower under clang 23 (16tblper 64 bytes), and the barrier restores the old arm64 loop in a standalone compile.Why no new test. The change has no effect that a caller can observe except speed, and a wall-clock assertion is not stable on CI or under the debug and ASAN builds. The existing suites cover correctness of every caller. The Postgres SCRAM login test (
test/js/sql/sql.test.ts) needs Docker and did not run locally. CI runs it.Scope of the bump.
gh api repos/oven-sh/boringssl/compare/41bf9b59...51a84a81reports 1 commit ahead, 0 behind, and one file,crypto/mem.cc.Relation to #43095. #43095 adds a runtime-dispatched Highway kernel for
timingSafeEqualfrom 64 bytes. With this pin, 64 to 256 byte compares are already within 1 to 5 ns of that kernel (24.7 and 27.7 ns here, 21.5 and 22.4 ns there). The kernel's gain starts at about 1 KB: 41 ns against 26 ns, and 2241 ns against 849 ns at 64 KB, on a host with AVX-512.no test proof · iteration 0 · no src or test change; test-proof not applicable