From 51a84a8111ef25fab67079b7734c2fed3b9f91e2 Mon Sep 17 00:00:00 2001 From: robobun <117481402+robobun@users.noreply.github.com> Date: Thu, 17 Sep 2026 14:06:57 +0000 Subject: [PATCH] Return CRYPTO_memcmp's accumulator through a value barrier Clang 23 folds the conversion of the uint8_t accumulator to int into the loop (llvm/llvm-project#222142). The accumulator becomes 32 bits wide, and the loop vectorizer then puts 4 bytes in a 128-bit vector instead of 16. A 1 KB compare takes about 120 ns instead of about 26 ns. value_barrier_w takes a crypto_word_t, which is wider than int, so the fold does not apply and the accumulator stays 8 bits wide. The result is also opaque to the compiler where CRYPTO_memcmp is inlined. --- crypto/mem.cc | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/crypto/mem.cc b/crypto/mem.cc index 4320dc6998..e44d48aef5 100644 --- a/crypto/mem.cc +++ b/crypto/mem.cc @@ -358,7 +358,12 @@ int CRYPTO_memcmp(const void *in_a, const void *in_b, size_t len) { x |= a[i] ^ b[i]; } - return x; + // Bun: return `x` through a value barrier that is wider than `int`. Clang 23 + // otherwise folds the conversion of `x` to `int` into the loop, keeps a 32-bit + // accumulator, and vectorizes the loop with 4 bytes per 128-bit vector instead + // of 16 (https://github.com/llvm/llvm-project/issues/222142). The barrier also + // stops the compiler from reasoning about the result where this is inlined. + return static_cast(value_barrier_w(x)); } uint32_t OPENSSL_hash32(const void *ptr, size_t len) {